Files
spring-data-gemfire/src/main/asciidoc/reference/function-annotations.adoc
John Blum 6d34d24b2a SGF-768 - Replace all direct URL/links with Asciidoc variables.
Replace all product names with Asciidoc variables.

Replace all product versions with Asciidoc variables.
2018-07-25 22:34:25 -07:00

444 lines
23 KiB
Plaintext

[[function-annotations]]
= Annotation Support for Function Execution
Spring Data for {data-store-name} includes annotation support to simplify working with {data-store-name}
{x-data-store-docs}/developing/function_exec/chapter_overview.html[function execution].
Under the hood, the {data-store-name} API provides classes to implement and register {data-store-name}
{x-data-store-javadoc}/org/apache/geode/cache/execute/Function.html[functions]
that are deployed on {data-store-name} servers, which may then be invoked by other peer member applications
or remotely from cache clients.
Functions can execute in parallel, distributed among multiple {data-store-name} servers in the cluster, aggregating results
with the map-reduce pattern that are sent back to the caller. Functions can also be targeted to run on a single server
or region. The {data-store-name} API supports remote execution of functions targeted by using various predefined scopes:
on region, on members (in groups), on servers, and others. The implementation and execution of remote functions,
as with any RPC protocol, requires some boilerplate code.
Spring Data for {data-store-name}, true to Spring's core value proposition, aims to hide the mechanics of remote function execution
and let you focus on core POJO programming and business logic. To this end, Spring Data for {data-store-name} introduces
annotations to declaratively register the public methods of a POJO class as {data-store-name} functions along with the ability to
invoke registered functions (including remotely) by using annotated interfaces.
== Implementation Versus Execution
There are two separate concerns to address implementation and execution.
The first is function implementation (server-side), which must interact with the
{x-data-store-javadoc}/org/apache/geode/cache/execute/FunctionContext.html[`FunctionContext`]
to access the invocation arguments,
{x-data-store-javadoc}/org/apache/geode/cache/execute/ResultSender.html[`ResultsSender`],
and other execution context information. The function implementation typically accesses the cache and regions
and is registered with the
{x-data-store-javadoc}/org/apache/geode/cache/execute/FunctionService.html[`FunctionService`]
under a unique ID.
A cache client application invoking a function does not depend on the implementation. To invoke a function,
the application instantiates an
{x-data-store-javadoc}/org/apache/geode/cache/execute/Execution.html[`Execution`]
providing the function ID, invocation arguments, and the function target, which defines its scope:
region, server, servers, member, or members. If the function produces a result, the invoker uses a
{x-data-store-javadoc}/org/apache/geode/cache/execute/ResultCollector.html[`ResultCollector`]
to aggregate and acquire the execution results. In certain cases, a custom `ResultCollector` implementation
is required and may be registered with the `Execution`.
NOTE: 'Client' and 'Server' are used here in the context of function execution, which may have a different meaning
than client and server in {data-store-name}'s client-server topology. While it is common for an application using a `ClientCache`
to invoke a function on one or more {data-store-name} servers in a cluster, it is also possible to execute functions
in a peer-to-peer (P2P) configuration, where the application is a member of the cluster hosting a peer `Cache`.
Keep in mind that a peer member cache application is subject to all the constraints of being a peer member
of the cluster.
[[function-implementation]]
== Implementing a Function
Using {data-store-name} APIs, the `FunctionContext` provides a runtime invocation context that includes the client's
calling arguments and a `ResultSender` implementation to send results back to the client. Additionally,
if the function is executed on a region, the `FunctionContext` is actually an instance of `RegionFunctionContext`,
which provides additional information, such as the target region on which the function was invoked,
any filter (a set of specific keys) associated with the `Execution`, and so on. If the region is a `PARTITION` region,
the function should use the `PartitionRegionHelper` to extract only the local data.
By using Spring, you can write a simple POJO and use the Spring container to bind one or more of your POJO's
public methods to a function. The signature for a POJO method intended to be used as a function must generally
conform to the client's execution arguments. However, in the case of a region execution, the region data
may also be provided (presumably the data is held in the local partition if the region is a `PARTITION` region).
Additionally, the function may require the filter that was applied, if any. This suggests that the client and server
share a contract for the calling arguments but that the method signature may include additional parameters
to pass values provided by the `FunctionContext`. One possibility is for the client and server to share
a common interface, but this is not strictly required. The only constraint is that the method signature includes
the same sequence of calling arguments with which the function was invoked after the additional parameters
are resolved.
For example, suppose the client provides a `String` and an `int` as the calling arguments. These are provided
in the `FunctionContext` as an array, as the following example shows:
`Object[] args = new Object[] { "test", 123 };`
The Spring container should be able to bind to any method signature similar to the following (ignoring the return type for the moment):
[source,java]
----
public Object method1(String s1, int i2) {...}
public Object method2(Map<?, ?> data, String s1, int i2) {...}
public Object method3(String s1, Map<?, ?> data, int i2) {...}
public Object method4(String s1, Map<?, ?> data, Set<?> filter, int i2) {...}
public void method4(String s1, Set<?> filter, int i2, Region<?,?> data) {...}
public void method5(String s1, ResultSender rs, int i2);
public void method6(FunctionContest context);
----
The general rule is that once any additional arguments (that is, region data and filter) are resolved,
the remaining arguments must correspond exactly, in order and type, to the expected function method parameters.
The method's return type must be void or a type that may be serialized (as a `java.io.Serializable`,
`DataSerializable`, or `PdxSerializable`). The latter is also a requirement for the calling arguments.
The region data should normally be defined as a `Map`, to facilitate unit testing, but may also be of type region,
if necessary. As shown in the preceding example, it is also valid to pass the `FunctionContext` itself
or the `ResultSender` if you need to control how the results are returned to the client.
=== Annotations for Function Implementation
The following example shows how SDG's function annotations are used to expose POJO methods
as {data-store-name} functions:
[source,java]
----
@Component
public class ApplicationFunctions {
@GemfireFunction
public String function1(String value, @RegionData Map<?, ?> data, int i2) { ... }
@GemfireFunction("myFunction", batchSize=100, HA=true, optimizedForWrite=true)
public List<String> function2(String value, @RegionData Map<?, ?> data, int i2, @Filter Set<?> keys) { ... }
@GemfireFunction(hasResult=true)
public void functionWithContext(FunctionContext functionContext) { ... }
}
----
Note that the class itself must be registered as a Spring bean and each {data-store-name} Function is annotated
with `@GemfireFunction`. In the preceding example, Spring's `@Component` annotation was used, but you can register the bean
by using any method supported by Spring (such as XML configuration or with a Java configuration class when using Spring Boot).
This lets the Spring container create an instance of this class and wrap it in a
http://docs.spring.io/spring-data-gemfire/docs/current/api/org/springframework/data/gemfire/function/PojoFunctionWrapper.html[`PojoFunctionWrapper`].
Spring creates a wrapper instance for each method annotated with `@GemfireFunction`. Each wrapper instance shares
the same target object instance to invoke the corresponding method.
TIP: The fact that the POJO Function class is a Spring bean may offer other benefits, since it shares
the `ApplicationContext` with {data-store-name} components, such as the cache and regions. These may be injected into the class
if necessary.
Spring creates the wrapper class and registers the functions with {data-store-name}'s function service. The function ID used
to register each function must be unique. By using convention, it defaults to the simple (unqualified) method name.
The name can be explicitly defined by using the `id` attribute of the `@GemfireFunction` annotation.
The `@GemfireFunction` annotation also provides other configuration attributes, `HA` and `optimizedForWrite`,
which correspond to properties defined by {data-store-name}'s
{x-data-store-javadoc}/org/apache/geode/cache/execute/Function.html[`Function`] interface.
If the method's return type is `void`, then the `hasResult` property is automatically set to `false`.
Otherwise, if the method returns a value, the `hasResult` attributes is set to `true`.
Even for `void` return types, the annotation's `hasResult` attribute can be set to `true` to override this convention,
as shown in the `functionWithContext` method show previously. Presumably, the intention is to use the `ResultSender` directly
to send results to the caller.
The `PojoFunctionWrapper` implements {data-store-name}'s `Function` interface, binds method parameters, and invokes the target method
in its `execute()` method. It also sends the method's return value by using the `ResultSender`.
=== Batching Results
If the return type is an array or `Collection`, then some consideration must be given to how the results are returned.
By default, the `PojoFunctionWrapper` returns the entire array or `Collection` at once. If the number of elements
in the array or `Collection` is quite large, it may incur a performance penalty. To divide the payload into smaller,
more manageable chunks, you can set the `batchSize` attribute, as illustrated in `function2`, shown earlier.
TIP: If you need more control of the `ResultSender`, especially if the method itself would use too much memory
to create the `Collection`, you can pass the `ResultSender` or access it through the `FunctionContext` and use it directly
within the method to sends results back to the caller.
=== Enabling Annotation Processing
In accordance with Spring standards, you must explicitly activate annotation processing for `@GemfireFunction`
annotations. The following example activates annotation processing with XML:
[source,xml]
----
<gfe:annotation-driven/>
----
The following example activates annotations by annotating a Java configuration class:
[source,java]
----
@Configuration
@EnableGemfireFunctions
class ApplicationConfiguration { .. }
----
[[function-execution]]
== Executing a Function
A process that invokes a remote function needs to provide the function's ID, calling arguments, the execution target
(`onRegion`, `onServers`, `onServer`, `onMember`, or `onMembers`) and (optionally) a filter set. By using Spring Data for {data-store-name},
all you need do is define an interface supported by annotations. Spring creates a dynamic proxy
for the interface, which uses the `FunctionService` to create an `Execution`, invoke the `Execution`, and (if necessary) coerce
the results to the defined return type. This technique is similar to the way
Spring Data for {data-store-name}'s repository extension works. Thus, some of the configuration and concepts should be familiar.
Generally, a single interface definition maps to multiple function executions, one corresponding to each method
defined in the interface.
=== Annotations for Function Execution
To support client-side Function execution, the following SDG Function annotations are provided: `@OnRegion`,
`@OnServer`, `@OnServers`, `@OnMember`, and `@OnMembers`. These annotations correspond to the `Execution` implementations
provided by {data-store-name}'s
{x-data-store-javadoc}/org/apache/geode/cache/execute/FunctionService.html[`FunctionService`].
Each annotation exposes the appropriate attributes. These annotations also provide an optional
`resultCollector` attribute whose value is the name of a Spring bean implementing the
{x-data-store-javadoc}/org/apache/geode/cache/execute/ResultCollector.html[`ResultCollector`]
to use for the execution.
CAUTION: The proxy interface binds all declared methods to the same execution configuration. Although it is expected
that single method interfaces are common, all methods in the interface are backed by the same proxy instance
and therefore all share the same configuration.
The following listing shows a few examples:
[source,java]
----
@OnRegion(region="SomeRegion", resultCollector="myCollector")
public interface FunctionExecution {
@FunctionId("function1")
String doIt(String s1, int i2);
String getString(Object arg1, @Filter Set<Object> keys);
}
----
By default, the function ID is the simple (unqualified) method name. The `@FunctionId` annotation can be used
to bind this invocation to a different function ID.
=== Enabling Annotation Processing
The client-side uses Spring's classpath component scanning capability to discover annotated interfaces. To enable
function execution annotation processing in XML, insert the following element in your XML configuration:
[source,xml]
----
<gfe-data:function-executions base-package="org.example.myapp.gemfire.functions"/>
----
The `function-executions` element is provided in the `gfe-data` namespace. The `base-package` attribute is required,
to avoid scanning the entire classpath. Additional filters are provided as described in the Spring
http://docs.spring.io/spring/docs/current/spring-framework-reference/htmlsingle/#beans-scanning-filters[reference documentation].
Optionally, you can annotate your Java configuration class as follows:
[source,java]
----
@EnableGemfireFunctionExecutions(basePackages = "org.example.myapp.gemfire.functions")
----
[[function-execution-programmatic]]
== Programmatic Function Execution
Using the Function execution annotated interface defined in the previous section, simply auto-wire your interface
into an application bean that will invoke the Function:
[source,java]
----
@Component
public class MyApplication {
@Autowired
FunctionExecution functionExecution;
public void doSomething() {
functionExecution.doIt("hello", 123);
}
}
----
Alternately, you can use a function execution template directly. In the following example, the `GemfireOnRegionFunctionTemplate`
creates an `onRegion` function `Execution`:
.Using the `GemfireOnRegionFunctionTemplate`
====
[source,java]
----
Set<?, ?> myFilter = getFilter();
Region<?, ?> myRegion = getRegion();
GemfireOnRegionOperations template = new GemfireOnRegionFunctionTemplate(myRegion);
String result = template.executeAndExtract("someFunction", myFilter, "hello", "world", 1234);
----
====
Internally, function `Executions` always return a `List`. `executeAndExtract` assumes a singleton `List`
containing the result and attempts to coerce that value into the requested type. There is also
an `execute` method that returns the `List` as is. The first parameter is the function ID.
The filter argument is optional. The remaining arguments are a variable argument `List`.
[[function-execution-pdx]]
== Function Execution with PDX
When using Spring Data for {data-store-name}'s function annotation support combined with {data-store-name}'s
{x-data-store-docs}/developing/data_serialization/gemfire_pdx_serialization.html[PDX Serialization],
there are a few logistical things to keep in mind.
As explained earlier in this section, and by way of example, you should typically define {data-store-name} functions by using POJO classes
annotated with Spring Data for {data-store-name}
http://docs.spring.io/spring-data-gemfire/docs/current/api/org/springframework/data/gemfire/function/annotation/package-summary.html[function annotations],
as follows:
[source,java]
----
public class OrderFunctions {
@GemfireFunction(...)
Order process(@RegionData data, Order order, OrderSource orderSourceEnum, Integer count) { ... }
}
----
NOTE: The `Integer` type `count` parameter is arbitrary, as is the separation of the `Order` class and `OrderSource` enum,
which might be logical to combine. However, the arguments were setup this way to demonstrate the problem with
function executions in the context of PDX.
Your `Order` and `OrderSource` enum might be as follows:
[source,java]
----
public class Order ... {
private Long orderNumber;
private Calendar orderDateTime;
private Customer customer;
private List<Item> items
...
}
public enum OrderSource {
ONLINE,
PHONE,
POINT_OF_SALE
...
}
----
Of course, you can define a function `Execution` interface to call the 'process' {data-store-name} server function, as follows:
[source,java]
----
@OnServer
public interface OrderProcessingFunctions {
Order process(Order order, OrderSource orderSourceEnum, Integer count);
}
----
Clearly, this `process(..)` `Order` Function is being called from a client-side with an application based on `ClientCache`
(that is `<gfe:client-cache/>`). This implies that the function arguments must also be serializable.
The same is true when invoking peer-to-peer member functions (such as `@OnMember(s)) between peers in the cluster.
Any form of `distribution` requires the data transmitted between client and server (or peers) to be serialized.
Now, if you have configured {data-store-name} to use PDX for serialization (instead of Java serialization, for instance)
you can also set the `pdx-read-serialized` attribute to `true` in your configuration
of the {data-store-name} server(s), as follows:
[source,xml]
----
<gfe:cache ... pdx-read-serialized="true"/>
----
Alternatively, you can set the `pdx-read-serialized` attribute to `true` for a {data-store-name} cache client application, as follows:
[source,xml]
----
<gfe:client-cache ... pdx-read-serialized="true"/>
----
Doing so causes all values read from the cache (that is, regions) as well as information passed between client and servers
(or peers) to remain in serialized form, including, but not limited to, function arguments.
{data-store-name} serializes only application domain object types that you have specifically configured (registered)
either by using {data-store-name}'s
{x-data-store-javadoc}/org/apache/geode/pdx/ReflectionBasedAutoSerializer.html[`ReflectionBasedAutoSerializer`],
or specifically (and recommended) by using a "`custom`" {data-store-name}
{x-data-store-javadoc}/org/apache/geode/pdx/PdxSerializer.html[`PdxSerializer`]. If you use
Spring Data for {data-store-name}'s repository extension to Spring Data Common's repository abstraction and infrastructure,
you might even want to consider using Spring Data for {data-store-name}'s
http://docs.spring.io/spring-data-gemfire/docs/current/api/org/springframework/data/gemfire/mapping/MappingPdxSerializer.html[`MappingPdxSerializer`],
which uses an entity's mapping meta-data to determine data from the application domain object that are serialized
to the PDX instance.
What is less than apparent, though, is that {data-store-name} automatically handles Java `Enum` types regardless of whether they are
explicitly configured (that is, registered with a `ReflectionBasedAutoSerializer` using a regex pattern
and the `classes` parameter or are handled by a "`custom`" {data-store-name} `PdxSerializer`), despite the fact that Java enumerations
implement `java.io.Serializable`.
So, when you set `pdx-read-serialized` to `true` on {data-store-name} servers where the {data-store-name} functions
(including Spring Data for {data-store-name} function-annotated POJO classes) are registered, then you
may encounter surprising behavior when invoking the function `Execution`.
You might pass the following arguments when invoking the function:
[source,java]
----
orderProcessingFunctions.process(new Order(123, customer, Calendar.getInstance(), items), OrderSource.ONLINE, 400);
----
However, the {data-store-name} function on the server gets the following:
[source,java]
----
process(regionData, order:PdxInstance, :PdxInstanceEnum, 400);
----
The `Order` and `OrderSource` have been passed to the function as
{x-data-store-javadoc}/org/apache/geode/pdx/PdxInstance.html[PDX instances].
Again, this all happens because `pdx-read-serialized` is set to `true`, which may be necessary in cases where
the {data-store-name} servers interact with multiple different clients (for example, a combination of Java clients and native clients, such as C++, C#, and others).
This flies in the face of Spring Data for {data-store-name}'s strongly-typed function-annotated POJO class method signatures,
as you should reasonably expect application domain object types, not PDX serialized instances.
Consequently, Spring Data for {data-store-name} includes enhanced function support to automatically convert method arguments
type PDX to the desired application domain object types defined by the function method's
parameter types.
However, this also requires you to explicitly register a {data-store-name} `PdxSerializer` on the {data-store-name} Servers
where Spring Data for {data-store-name} function-annotated POJOs are registered and used, as the following example shows:
[source,java]
----
<bean id="customPdxSerializer" class="x.y.z.gemfire.serialization.pdx.MyCustomPdxSerializer"/>
<gfe:cache ... pdx-serializer-ref="customPdxSerializeer" pdx-read-serialized="true"/>
----
Alternatively, you can use {data-store-name}'s
{x-data-store-javadoc}/org/apache/geode/pdx/ReflectionBasedAutoSerializer.html[`ReflectionBasedAutoSerializer`]
for convenience. Of course, we recommend that, where possible, you use a custom `PdxSerializer` to maintain
finer-grained control over your serialization strategy.
Finally, Spring Data for {data-store-name} is careful not to convert your function arguments if you treat your function arguments
generically or as one of {data-store-name}'s PDX types, as follows:
[source,java]
----
@GemfireFunction
public Object genericFunction(String value, Object domainObject, PdxInstanceEnum enum) {
...
}
----
Spring Data for {data-store-name} converts PDX type data to the corresponding application domain types if and only if
the corresponding application domain types are on the classpath and the function-annotated POJO method expects it.
For a good example of custom, composed application-specific {data-store-name} `PdxSerializers` as well as appropriate
POJO function parameter type handling based on the method signatures, see Spring Data for {data-store-name}'s
https://github.com/spring-projects/spring-data-gemfire/blob/2.0.0.M2/src/test/java/org/springframework/data/gemfire/function/ClientCacheFunctionExecutionWithPdxIntegrationTest.java[`ClientCacheFunctionExecutionWithPdxIntegrationTest`] class.