Chat Models
Chat Models
Overview
Chat models enable agents to communicate with Large Language Models (LLMs) for natural language understanding, reasoning, and generation. In Flink Agents, chat models act as the “brain” of your agents, processing input messages and generating intelligent responses based on context, prompts, and available tools.
Getting Started
To use chat models in your agents, you need to define both a connection and setup using decorators/annotations, then interact with the model through events.
Resource Declaration
Flink Agents provides decorators in python and annotations in java to simplify chat model setup within agents:
@chat_model_connection/@ChatModelConnection
The @chat_model_connection decorator or @ChatModelConnection annotation marks a method that creates a chat model connection. This is typically defined once and shared across multiple chat model setups.
@chat_model_setup/@ChatModelSetup
The @chat_model_setup decorator or @ChatModelSetup annotation marks a method that creates a chat model setup. This references a connection and adds chat-specific configuration like prompts and tools.
Chat Events
Chat models communicate through built-in events:
- ChatRequestEvent: Sent by actions to request a chat completion from the LLM
- ChatResponseEvent: Received by actions containing the LLM’s response
Usage Example
Here’s how to define and use chat models in a workflow agent:
Python
class MyAgent(Agent):
@chat_model_connection
@staticmethod
def ollama_connection() -> ResourceDescriptor:
return ResourceDescriptor(
clazz=ResourceName.ChatModel.OLLAMA_CONNECTION,
base_url="http://localhost:11434",
request_timeout=30.0
)
@chat_model_setup
@staticmethod
def ollama_chat_model() -> ResourceDescriptor:
return ResourceDescriptor(
clazz=ResourceName.ChatModel.OLLAMA_SETUP,
connection="ollama_connection",
model="qwen3:8b",
temperature=0.7
)
@action(InputEvent.EVENT_TYPE)
@staticmethod
def process_input(event: Event, ctx: RunnerContext) -> None:
input_event = InputEvent.from_event(event)
# Create a chat request with user message
user_message = ChatMessage(
role=MessageRole.USER,
content=f"input: {input_event.input}"
)
ctx.send_event(
ChatRequestEvent(model="ollama_chat_model", messages=[user_message])
)
@action(ChatResponseEvent.EVENT_TYPE)
@staticmethod
def process_response(event: Event, ctx: RunnerContext) -> None:
chat_response = ChatResponseEvent.from_event(event)
response_content = chat_response.response.content
# Handle the LLM's response
# Process the response as needed for your use caseJava
public class MyAgent extends Agent {
@ChatModelConnection
public static ResourceDescriptor ollamaConnection() {
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.OLLAMA_CONNECTION)
.addInitialArgument("endpoint", "http://localhost:11434")
.build();
}
@ChatModelSetup
public static ResourceDescriptor ollamaChatModel() {
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.OLLAMA_SETUP)
.addInitialArgument("connection", "ollamaConnection")
.addInitialArgument("model", "qwen3:8b")
.build();
}
@Action(listenEventTypes = {InputEvent.EVENT_TYPE})
public static void processInput(Event event, RunnerContext ctx) throws Exception {
InputEvent inputEvent = InputEvent.fromEvent(event);
ChatMessage userMessage =
new ChatMessage(MessageRole.USER, String.format("input: {%s}", inputEvent.getInput()));
ctx.sendEvent(new ChatRequestEvent("ollamaChatModel", List.of(userMessage)));
}
@Action(listenEventTypes = {ChatResponseEvent.EVENT_TYPE})
public static void processResponse(Event event, RunnerContext ctx)
throws Exception {
ChatResponseEvent chatResponse = ChatResponseEvent.fromEvent(event);
String response = chatResponse.getResponse().getContent();
// Handle the LLM's response
// Process the response as needed for your use case
}
}Built-in Providers
Amazon Bedrock
Amazon Bedrock provides access to a wide range of foundation models from leading AI providers through a unified API. The Flink Agents Bedrock integration uses the Converse API, which provides a consistent interface across all supported models with native tool calling support. Authentication is handled via SigV4 using the AWS default credentials chain. No API keys are required.
Amazon Bedrock is only supported in Java currently. To use Amazon Bedrock from Python agents, see Using Cross-Language Providers.
Prerequisites
- An AWS account with Amazon Bedrock model access enabled for the models you plan to use
- IAM credentials configured via any method supported by the AWS Default Credentials Provider (environment variables,
~/.aws/credentials, IAM role, etc.)
BedrockChatModelConnection Parameters
Java
| Parameter | Type | Default | Description |
|---|---|---|---|
region | String | "us-east-1" | AWS region for the Bedrock service |
model | String | None | Default model ID (can be overridden per setup) |
max_retries | int | 5 | Maximum number of API retry attempts (retries on throttling, 429, 503) |
BedrockChatModelSetup Parameters
Java
| Parameter | Type | Default | Description | |
|---|---|---|---|---|
connection | String | Required | Reference to connection method name | |
model | String | Required | Bedrock model ID (e.g. "us.anthropic.claude-sonnet-4-20250514-v1:0") | |
prompt | Prompt \ | String | None | Prompt template or reference to prompt resource |
tools | List | None | List of tool names available to the model | |
temperature | double | 0.1 | Sampling temperature (0.0 to 1.0) | |
max_tokens | int | None | Maximum number of tokens to generate |
Usage Example
Java
public class MyAgent extends Agent {
@ChatModelConnection
public static ResourceDescriptor bedrockConnection() {
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.BEDROCK_CONNECTION)
.addInitialArgument("region", "us-east-1")
.build();
}
@ChatModelSetup
public static ResourceDescriptor bedrockChatModel() {
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.BEDROCK_SETUP)
.addInitialArgument("connection", "bedrockConnection")
.addInitialArgument("model", "us.anthropic.claude-sonnet-4-20250514-v1:0")
.addInitialArgument("temperature", 0.1d)
.addInitialArgument("max_tokens", 4096)
.build();
}
...
}Available Models
Amazon Bedrock supports models from multiple providers through a single API. Visit the Amazon Bedrock Model IDs documentation for the complete and up-to-date list of available models.
Some popular options include:
- Claude (Anthropic):
us.anthropic.claude-sonnet-4-6,us.anthropic.claude-opus-4-7,us.anthropic.claude-opus-4-6-v1 - Llama (Meta):
us.meta.llama4-scout-17b-16e-instruct-v1:0 - Mistral:
mistral.mistral-large-2402-v1:0 - Amazon Nova:
us.amazon.nova-pro-v1:0,us.amazon.nova-lite-v1:0
Model availability varies by AWS region and requires explicit model access enablement in the Bedrock console. Always check the Amazon Bedrock documentation for regional availability before implementing in production.
Current limitations: The integration uses text content blocks only. Extended thinking / reasoning content blocks (e.g. Claude extended thinking), citation blocks, and image / document content blocks are not yet supported.
Anthropic
Anthropic provides cloud-based chat models featuring the Claude family, known for their strong reasoning, coding, and safety capabilities.
Prerequisites
- Create an account at Anthropic Console
- Navigate to API Keys and create a new secret key
AnthropicChatModelConnection Parameters
Python
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key | str | Required | Anthropic API key for authentication |
max_retries | int | 3 | Maximum number of API retry attempts |
timeout | float | 60.0 | API request timeout in seconds |
Java
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key | String | Required | Anthropic API key for authentication |
timeout | int | None | Timeout in seconds for API requests |
max_retries | int | 2 | Maximum number of API retry attempts |
AnthropicChatModelSetup Parameters
Python
| Parameter | Type | Default | Description | |
|---|---|---|---|---|
connection | str | Required | Reference to connection method name | |
model | str | "claude-sonnet-4-20250514" | Name of the chat model to use | |
prompt | Prompt \ | str | None | Prompt template or reference to prompt resource |
tools | List[str] | None | List of tool names available to the model | |
max_tokens | int | 1024 | Maximum number of tokens to generate | |
temperature | float | 0.1 | Sampling temperature (0.0 to 1.0) | |
additional_kwargs | dict | {} | Additional Anthropic API parameters |
Java
| Parameter | Type | Default | Description | |
|---|---|---|---|---|
connection | String | Required | Reference to connection method name | |
model | String | "claude-sonnet-4-20250514" | Name of the chat model to use | |
prompt | Prompt \ | String | None | Prompt template or reference to prompt resource |
tools | List | None | List of tool names available to the model | |
max_tokens | long | 1024 | Maximum number of tokens to generate | |
temperature | double | 0.1 | Sampling temperature (0.0 to 1.0) | |
json_prefill | boolean | true | Prefill assistant response with “{” to enforce JSON output (disabled when tools are used) | |
strict_tools | boolean | false | Enable strict mode for tool calling schemas | |
additional_kwargs | Map<String, Object> | {} | Additional Anthropic API parameters (top_k, top_p, stop_sequences) |
Usage Example
Python
class MyAgent(Agent):
@chat_model_connection
@staticmethod
def anthropic_connection() -> ResourceDescriptor:
return ResourceDescriptor(
clazz=ResourceName.ChatModel.ANTHROPIC_CONNECTION,
api_key="<your-api-key>",
max_retries=3,
timeout=60.0
)
@chat_model_setup
@staticmethod
def anthropic_chat_model() -> ResourceDescriptor:
return ResourceDescriptor(
clazz=ResourceName.ChatModel.ANTHROPIC_SETUP,
connection="anthropic_connection",
model="claude-sonnet-4-20250514",
max_tokens=2048,
temperature=0.7
)
...Java
public class MyAgent extends Agent {
@ChatModelConnection
public static ResourceDescriptor anthropicConnection() {
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.ANTHROPIC_CONNECTION)
.addInitialArgument("api_key", "<your-api-key>")
.addInitialArgument("timeout", 120)
.addInitialArgument("max_retries", 3)
.build();
}
@ChatModelSetup
public static ResourceDescriptor anthropicChatModel() {
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.ANTHROPIC_SETUP)
.addInitialArgument("connection", "anthropicConnection")
.addInitialArgument("model", "claude-sonnet-4-20250514")
.addInitialArgument("temperature", 0.7d)
.addInitialArgument("max_tokens", 2048)
.build();
}
...
}Available Models
Visit the Anthropic Models documentation for the complete and up-to-date list of available chat models.
Some popular options include:
- Claude Sonnet 4.5 (claude-sonnet-4-5-20250929)
- Claude Sonnet 4 (claude-sonnet-4-20250514)
- Claude Sonnet 3.7 (claude-3-7-sonnet-20250219)
- Claude Opus 4.1 (claude-opus-4-1-20250805)
Model availability and specifications may change. Always check the official Anthropic documentation for the latest information before implementing in production.
Azure AI
Azure AI provides cloud-based chat models through Azure AI Inference API, supporting various models including Llama, Mistral, Phi, and other models deployed via Azure AI Studio.
Azure AI is only supported in Java currently. To use Azure AI from Python agents, see Using Cross-Language Providers.
Azure AI vs Azure OpenAI: Azure AI uses the Azure AI Inference API to access models deployed via Azure AI Studio (Llama, Mistral, Phi, etc.). If you want to use OpenAI models (GPT-4, etc.) hosted on Azure, see Azure OpenAI instead.
Prerequisites
- Create an Azure AI resource in the Azure Portal
- Obtain your endpoint URL and API key from the Azure AI resource
AzureAIChatModelConnection Parameters
Java
| Parameter | Type | Default | Description |
|---|---|---|---|
endpoint | String | Required | Azure AI service endpoint URL |
apiKey | String | Required | Azure AI API key for authentication |
AzureAIChatModelSetup Parameters
Java
| Parameter | Type | Default | Description | |
|---|---|---|---|---|
connection | String | Required | Reference to connection method name | |
model | String | Required | Name of the chat model to use (e.g., “gpt-4o”) | |
prompt | Prompt \ | String | None | Prompt template or reference to prompt resource |
tools | List[String] | None | List of tool names available to the model |
Usage Example
Java
public class MyAgent extends Agent {
@ChatModelConnection
public static ResourceDescriptor azureAIConnection() {
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.AZURE_CONNECTION)
.addInitialArgument("endpoint", "https://your-resource.inference.ai.azure.com")
.addInitialArgument("apiKey", "your-api-key-here")
.build();
}
@ChatModelSetup
public static ResourceDescriptor azureAIChatModel() {
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.AZURE_SETUP)
.addInitialArgument("connection", "azureAIConnection")
.addInitialArgument("model", "gpt-4o")
.build();
}
...
}Available Models
Azure AI supports various models through the Azure AI Inference API. Visit the Azure AI Model Catalog for the complete and up-to-date list of available models.
Some popular options include:
- GPT-4o (gpt-4o)
- GPT-4 (gpt-4)
- GPT-4 Turbo (gpt-4-turbo)
- GPT-3.5 Turbo (gpt-3.5-turbo)
Model availability and specifications may change. Always check the official Azure AI documentation for the latest information before implementing in production.
Azure OpenAI
Azure OpenAI provides access to OpenAI models (GPT-4, GPT-4o, etc.) through Azure’s cloud infrastructure, using the same OpenAI SDK with Azure-specific authentication and endpoints. This offers enterprise security, compliance, and regional availability while using familiar OpenAI APIs.
Azure OpenAI vs Azure AI: Azure OpenAI uses the OpenAI SDK to access OpenAI models (GPT-4, etc.) hosted on Azure. If you want to use other models like Llama, Mistral, or Phi deployed via Azure AI Studio, see Azure AI instead.
Prerequisites
- Create an Azure OpenAI resource in the Azure Portal
- Deploy a model in Azure OpenAI Studio
- Obtain your endpoint URL, API key, API version, and deployment name from the Azure portal
AzureOpenAIChatModelConnection Parameters
Python
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key | str | Required | Azure OpenAI API key for authentication |
api_version | str | Required | Azure OpenAI REST API version (e.g., “2024-02-15-preview”). See API versions |
azure_endpoint | str | Required | Azure OpenAI endpoint URL (e.g., https://{resource-name}.openai.azure.com) |
timeout | float | 60.0 | API request timeout in seconds |
max_retries | int | 3 | Maximum number of API retry attempts |
Java
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key | String | Required | Azure OpenAI API key for authentication |
api_version | String | Required | Azure OpenAI REST API version (e.g., “2024-02-01”). See API versions |
azure_endpoint | String | Required | Azure OpenAI endpoint URL (e.g., https://{resource-name}.openai.azure.com) — either a direct Azure resource or a proxy/gateway URL that fronts an Azure OpenAI service |
timeout | int | None | Timeout in seconds for API requests; must be greater than 0, otherwise ignored (SDK default applies) |
max_retries | int | None | Maximum number of API retry attempts; must be non-negative, otherwise ignored (SDK default applies) |
azure_url_path_mode | String | "AUTO" | Controls how the SDK constructs Azure OpenAI request URLs. One of "AUTO", "LEGACY", or "UNIFIED". Custom gateways that proxy Azure OpenAI typically need "LEGACY" to force the /openai/deployments/{model} path |
AzureOpenAIChatModelSetup Parameters
Python
| Parameter | Type | Default | Description | |
|---|---|---|---|---|
connection | str | Required | Reference to connection method name | |
model | str | Required | Name of OpenAI model deployment on Azure | |
model_of_azure_deployment | str | None | The underlying model name (e.g., ‘gpt-4’, ‘gpt-35-turbo’). Used for token metrics tracking | |
prompt | Prompt \ | str | None | Prompt template or reference to prompt resource |
tools | List[str] | None | List of tool names available to the model | |
temperature | float | None | Sampling temperature (0.0 to 2.0). Not supported by reasoning models | |
max_tokens | int | None | Maximum number of tokens to generate | |
logprobs | bool | False | Whether to return log probabilities of output tokens | |
additional_kwargs | dict | {} | Additional Azure OpenAI API parameters |
Java
| Parameter | Type | Default | Description | |
|---|---|---|---|---|
connection | String | Required | Reference to connection method name | |
model | String | Required | Azure deployment name (not the underlying OpenAI model name) | |
model_of_azure_deployment | String | None | The underlying model name (e.g., ‘gpt-4’, ‘gpt-4o’). Used solely for token metrics tracking | |
prompt | Prompt \ | String | None | Prompt template or reference to prompt resource |
tools | List | None | List of tool names available to the model | |
temperature | double | None | Sampling temperature (0.0 to 2.0). Not supported by reasoning models | |
max_tokens | int | None | Maximum number of tokens to generate (must be greater than 0) | |
logprobs | boolean | false | Whether to return log probabilities of output tokens | |
additional_kwargs | Map<String, Object> | {} | Additional Azure OpenAI API parameters (forwarded to the OpenAI request body) |
Usage Example
Python
class MyAgent(Agent):
@chat_model_connection
@staticmethod
def azure_openai_connection() -> ResourceDescriptor:
return ResourceDescriptor(
clazz=ResourceName.ChatModel.AZURE_OPENAI_CONNECTION,
api_key="<your-api-key>",
api_version="2024-02-15-preview",
azure_endpoint="https://your-resource.openai.azure.com"
)
@chat_model_setup
@staticmethod
def azure_openai_chat_model() -> ResourceDescriptor:
return ResourceDescriptor(
clazz=ResourceName.ChatModel.AZURE_OPENAI_SETUP,
connection="azure_openai_connection",
model="my-gpt4-deployment", # Your Azure deployment name
model_of_azure_deployment="gpt-4", # Underlying model for metrics
max_tokens=1000
)
...Java
public class MyAgent extends Agent {
@ChatModelConnection
public static ResourceDescriptor azureOpenAIConnection() {
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.AZURE_OPENAI_CONNECTION)
.addInitialArgument("api_key", "<your-api-key>")
.addInitialArgument("api_version", "2024-02-01")
.addInitialArgument("azure_endpoint", "https://your-resource.openai.azure.com")
.build();
}
@ChatModelSetup
public static ResourceDescriptor azureOpenAIChatModel() {
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.AZURE_OPENAI_SETUP)
.addInitialArgument("connection", "azureOpenAIConnection")
.addInitialArgument("model", "my-gpt4-deployment") // Your Azure deployment name
.addInitialArgument("model_of_azure_deployment", "gpt-4") // Underlying model for metrics
.addInitialArgument("temperature", 0.3d)
.addInitialArgument("max_tokens", 1000)
.build();
}
...
}Available Models
Azure OpenAI supports OpenAI models deployed through your Azure subscription. Visit the Azure OpenAI Models documentation for the complete and up-to-date list of available models.
Some popular options include:
- GPT-4o (gpt-4o)
- GPT-4 (gpt-4)
- GPT-4 Turbo (gpt-4-turbo)
- GPT-3.5 Turbo (gpt-35-turbo)
Model availability depends on your Azure region and subscription. Always check the official Azure OpenAI documentation for regional availability before implementing in production.
Ollama
Ollama provides local chat models that run on your machine, offering privacy, control, and no API costs.
Prerequisites
Ollama server 0.9.0 or higher is required.
- Install Ollama from https://ollama.com/
- Start the Ollama server:
ollama serve - Download a chat model:
ollama pull qwen3:8b
OllamaChatModelConnection Parameters
Python
| Parameter | Type | Default | Description |
|---|---|---|---|
base_url | str | "http://localhost:11434" | Ollama server URL |
request_timeout | float | 30.0 | HTTP request timeout in seconds |
Java
| Parameter | Type | Default | Description |
|---|---|---|---|
endpoint | String | "http://localhost:11434" | Ollama server URL |
requestTimeout | long | 10 | HTTP request timeout in seconds |
OllamaChatModelSetup Parameters
Python
| Parameter | Type | Default | Description | |
|---|---|---|---|---|
connection | str | Required | Reference to connection method name | |
model | str | Required | Name of the chat model to use | |
prompt | Prompt \ | str | None | Prompt template or reference to prompt resource |
tools | List[str] | None | List of tool names available to the model | |
temperature | float | 0.75 | Sampling temperature (0.0 to 1.0) | |
num_ctx | int | 2048 | Maximum number of context tokens | |
keep_alive | str \ | float | "5m" | How long to keep model loaded in memory |
think | bool \ | Literal[“low”, “medium”, “high”] | True | Whether enable model think |
extract_reasoning | bool | True | Extract reasoning content from response | |
additional_kwargs | dict | {} | Additional Ollama API parameters |
Java
| Parameter | Type | Default | Description | |
|---|---|---|---|---|
connection | String | Required | Reference to connection method name | |
model | String | Required | Name of the chat model to use | |
prompt | Prompt \ | String | None | Prompt template or reference to prompt resource |
tools | List[String] | None | List of tool names available to the model | |
think | Boolean \ | Literal[“low”, “medium”, “high”] | true | Whether enable model think |
extract_reasoning | Boolean | true | Extract reasoning content from response |
Usage Example
Python
class MyAgent(Agent):
@chat_model_connection
@staticmethod
def ollama_connection() -> ResourceDescriptor:
return ResourceDescriptor(
clazz=ResourceName.ChatModel.OLLAMA_CONNECTION,
base_url="http://localhost:11434",
request_timeout=120.0
)
@chat_model_setup
@staticmethod
def my_chat_model() -> ResourceDescriptor:
return ResourceDescriptor(
clazz=ResourceName.ChatModel.OLLAMA_SETUP,
connection="ollama_connection",
model="qwen3:8b",
temperature=0.7,
num_ctx=4096,
keep_alive="10m",
think=True,
extract_reasoning=True
)
...Java
public class MyAgent extends Agent {
@ChatModelConnection
public static ResourceDescriptor ollamaConnection() {
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.OLLAMA_CONNECTION)
.addInitialArgument("endpoint", "http://localhost:11434")
.addInitialArgument("requestTimeout", 120)
.build();
}
@ChatModelSetup
public static ResourceDescriptor ollamaChatModel() {
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.OLLAMA_SETUP)
.addInitialArgument("connection", "ollamaConnection")
.addInitialArgument("model", "qwen3:8b")
.addInitialArgument("extract_reasoning", true)
.build();
}
...
}Available Models
Visit the Ollama Models Library for the complete and up-to-date list of available chat models.
Some popular options include:
- qwen3 series (qwen3:8b, qwen3:14b, qwen3:32b)
- llama3 series (llama3:8b, llama3:70b)
- deepseek series (deepseek-r1, deepseek-v3.1)
- gpt-oss
Model availability and specifications may change. Always check the official Ollama documentation for the latest information before implementing in production.
OpenAI
OpenAI provides cloud-based chat models with state-of-the-art performance for a wide range of natural language tasks.
Prerequisites
- Create an account at OpenAI Platform
- Navigate to API Keys and create a new secret key
Completions API
OpenAICompletionsConnection Parameters
Python
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key | str | Required | OpenAI API key for authentication |
api_base_url | str | "https://api.openai.com/v1" | Base URL for OpenAI API |
max_retries | int | 3 | Maximum number of API retry attempts |
timeout | float | 60.0 | API request timeout in seconds |
default_headers | dict | None | Default headers for API requests |
reuse_client | bool | True | Whether to reuse the OpenAI client between requests |
Java
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key | String | Required | OpenAI API key for authentication |
api_base_url | String | "https://api.openai.com/v1" | Base URL for OpenAI API |
max_retries | int | 2 | Maximum number of API retry attempts |
timeout | int | None | Timeout in seconds for API requests |
default_headers | Map<String, String> | None | Default headers for API requests |
model | String | None | Default model to use if not specified in setup |
OpenAICompletionsSetup Parameters
Python
| Parameter | Type | Default | Description | |
|---|---|---|---|---|
connection | str | Required | Reference to connection method name | |
model | str | "gpt-3.5-turbo" | Name of the chat model to use | |
prompt | Prompt \ | str | None | Prompt template or reference to prompt resource |
tools | List[str] | None | List of tool names available to the model | |
temperature | float | 0.1 | Sampling temperature (0.0 to 2.0) | |
max_tokens | int | None | Maximum number of tokens to generate | |
logprobs | bool | None | Whether to return log probabilities per token | |
top_logprobs | int | 0 | Number of top token log probabilities to return (0-20) | |
strict | bool | False | Enable strict mode for tool calling and schemas | |
reasoning_effort | str | None | Reasoning effort level for reasoning models (“low”, “medium”, “high”) | |
additional_kwargs | dict | {} | Additional OpenAI API parameters |
Java
| Parameter | Type | Default | Description | |
|---|---|---|---|---|
connection | String | Required | Reference to connection method name | |
model | String | "gpt-3.5-turbo" | Name of the chat model to use | |
prompt | Prompt \ | String | None | Prompt template or reference to prompt resource |
tools | List | None | List of tool names available to the model | |
temperature | double | 0.1 | Sampling temperature (0.0 to 2.0) | |
max_tokens | int | None | Maximum number of tokens to generate | |
logprobs | boolean | None | Whether to return log probabilities per token | |
top_logprobs | int | 0 | Number of top token log probabilities to return (0-20) | |
strict | boolean | false | Enable strict mode for tool calling and schemas | |
reasoning_effort | String | None | Reasoning effort level for reasoning models (“low”, “medium”, “high”) | |
additional_kwargs | Map<String, Object> | {} | Additional OpenAI API parameters |
Usage Example
Python
class MyAgent(Agent):
@chat_model_connection
@staticmethod
def openai_connection() -> ResourceDescriptor:
return ResourceDescriptor(
clazz=ResourceName.ChatModel.OPENAI_COMPLETIONS_CONNECTION,
api_key="<your-api-key>",
api_base_url="https://api.openai.com/v1",
max_retries=3,
timeout=60.0
)
@chat_model_setup
@staticmethod
def openai_chat_model() -> ResourceDescriptor:
return ResourceDescriptor(
clazz=ResourceName.ChatModel.OPENAI_COMPLETIONS_SETUP,
connection="openai_connection",
model="gpt-4",
temperature=0.7,
max_tokens=1000
)
...Java
public class MyAgent extends Agent {
@ChatModelConnection
public static ResourceDescriptor openaiConnection() {
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.OPENAI_COMPLETIONS_CONNECTION)
.addInitialArgument("api_key", "<your-api-key>")
.addInitialArgument("api_base_url", "https://api.openai.com/v1")
.addInitialArgument("timeout", 60)
.addInitialArgument("max_retries", 3)
.build();
}
@ChatModelSetup
public static ResourceDescriptor openaiChatModel() {
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.OPENAI_COMPLETIONS_SETUP)
.addInitialArgument("connection", "openaiConnection")
.addInitialArgument("model", "gpt-4")
.addInitialArgument("temperature", 0.7d)
.addInitialArgument("max_tokens", 1000)
.build();
}
...
}Responses API
Responses API is only supported in Java currently. To use OpenAI Responses API from Python agents, see Using Cross-Language Providers.
OpenAIResponsesModelConnection Parameters
Java
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key | String | Required | OpenAI API key for authentication |
api_base_url | String | None | Base URL for OpenAI API (useful for proxies) |
max_retries | int | 2 | Maximum number of API retry attempts |
timeout | int | None | Timeout in seconds for API requests |
default_headers | Map<String, String> | None | Default headers for API requests |
model | String | None | Default model to use if not specified in setup |
OpenAIResponsesModelSetup Parameters
Java
| Parameter | Type | Default | Description | |
|---|---|---|---|---|
connection | String | Required | Reference to connection method name | |
model | String | "gpt-4o" | Name of the chat model to use | |
prompt | Prompt \ | String | None | Prompt template or reference to prompt resource |
tools | List | None | List of tool names available to the model | |
temperature | double | 0.1 | Sampling temperature (0.0 to 2.0) | |
max_tokens | int | None | Maximum number of tokens to generate | |
strict | boolean | false | Enable strict mode for tool calling schemas | |
reasoning_effort | String | None | Reasoning effort level for reasoning models (“low”, “medium”, “high”) | |
store | boolean | false | Whether to store the response for later retrieval | |
instructions | String | None | System-level instructions for the model | |
additional_kwargs | Map<String, Object> | {} | Additional Responses API parameters |
Usage Example
Java
public class MyAgent extends Agent {
@ChatModelConnection
public static ResourceDescriptor openaiResponsesConnection() {
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.OPENAI_RESPONSES_CONNECTION)
.addInitialArgument("api_key", "<your-api-key>")
.addInitialArgument("timeout", 120)
.addInitialArgument("max_retries", 3)
.build();
}
@ChatModelSetup
public static ResourceDescriptor openaiResponsesChatModel() {
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.OPENAI_RESPONSES_SETUP)
.addInitialArgument("connection", "openaiResponsesConnection")
.addInitialArgument("model", "gpt-4o")
.addInitialArgument("temperature", 0.3d)
.addInitialArgument("max_tokens", 2048)
.addInitialArgument("store", true)
.build();
}
...
}Available Models
Visit the OpenAI Models documentation for the complete and up-to-date list of available chat models.
Some popular options include:
- GPT-5 series (GPT-5, GPT-5 mini, GPT-5 nano)
- GPT-4.1
- gpt-oss series (gpt-oss-120b, gpt-oss-10b)
Model availability and specifications may change. Always check the official OpenAI documentation for the latest information before implementing in production.
Tongyi (DashScope)
Tongyi provides cloud-based chat models from Alibaba Cloud, offering powerful Chinese and English language capabilities.
Tongyi is only supported in Python currently. To use Tongyi from Java agents, see Using Cross-Language Providers.
Prerequisites
- Get an API key from Alibaba Cloud DashScope
TongyiChatModelConnection Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key | str | $DASHSCOPE_API_KEY | DashScope API key for authentication |
request_timeout | float | 60.0 | HTTP request timeout in seconds |
TongyiChatModelSetup Parameters
| Parameter | Type | Default | Description | |
|---|---|---|---|---|
connection | str | Required | Reference to connection method name | |
model | str | "qwen-plus" | Name of the chat model to use | |
prompt | Prompt \ | str | None | Prompt template or reference to prompt resource |
tools | List[str] | None | List of tool names available to the model | |
temperature | float | 0.7 | Sampling temperature (0.0 to 2.0) | |
extract_reasoning | bool | False | Extract reasoning content from response | |
additional_kwargs | dict | {} | Additional DashScope API parameters |
Usage Example
class MyAgent(Agent):
@chat_model_connection
@staticmethod
def tongyi_connection() -> ResourceDescriptor:
return ResourceDescriptor(
clazz=ResourceName.ChatModel.TONGYI_CONNECTION,
api_key="your-api-key-here", # Or set DASHSCOPE_API_KEY env var
request_timeout=60.0
)
@chat_model_setup
@staticmethod
def tongyi_chat_model() -> ResourceDescriptor:
return ResourceDescriptor(
clazz=ResourceName.ChatModel.TONGYI_SETUP,
connection="tongyi_connection",
model="qwen-plus",
temperature=0.7,
extract_reasoning=True
)
...Available Models
Visit the DashScope Models documentation for the complete and up-to-date list of available chat models.
Some popular options include:
- qwen-plus
- qwen-max
- qwen-turbo
- qwen-long
Model availability and specifications may change. Always check the official DashScope documentation for the latest information before implementing in production.
Using Cross-Language Providers
Flink Agents supports cross-language chat model integration, allowing you to use chat models implemented in one language (Java or Python) from agents written in the other language. This is particularly useful when a chat model provider is only available in one language (e.g., Tongyi is currently Python-only).
Limitations:
- Cross-language resources are currently supported only when running in Flink, not in local development mode
- Complex object serialization between languages may have limitations
How To Use
To leverage chat model supports provided in a different language, you need to declare the resource within a built-in cross-language wrapper, and specify the target provider as an argument:
- Using Java chat models in Python: Use
ResourceName.ChatModel.JAVA_WRAPPER_CONNECTIONandResourceName.ChatModel.JAVA_WRAPPER_SETUP, specifying the Java provider class via thejava_clazzparameter - Using Python chat models in Java: Use
ResourceName.ChatModel.PYTHON_WRAPPER_CONNECTIONandResourceName.ChatModel.PYTHON_WRAPPER_SETUP, specifying the Python provider via thepythonClazzparameter
Usage Example
Using Java Chat Model in Python
class MyAgent(Agent):
@chat_model_connection
@staticmethod
def java_chat_model_connection() -> ResourceDescriptor:
# In pure Java, the equivalent ResourceDescriptor would be:
# ResourceDescriptor.Builder
# .newBuilder(ResourceName.ChatModel.OLLAMA_CONNECTION)
# .addInitialArgument("endpoint", "http://localhost:11434")
# .addInitialArgument("requestTimeout", 120)
# .build();
return ResourceDescriptor(
clazz=ResourceName.ChatModel.JAVA_WRAPPER_CONNECTION,
java_clazz=ResourceName.ChatModel.Java.OLLAMA_CONNECTION,
endpoint="http://localhost:11434",
requestTimeout=120,
)
@chat_model_setup
@staticmethod
def java_chat_model() -> ResourceDescriptor:
# In pure Java, the equivalent ResourceDescriptor would be:
# ResourceDescriptor.Builder
# .newBuilder(ResourceName.ChatModel.OLLAMA_SETUP)
# .addInitialArgument("connection", "java_chat_model_connection")
# .addInitialArgument("model", "qwen3:8b")
# .addInitialArgument("prompt", "my_prompt")
# .addInitialArgument("tools", List.of("my_tool1", "my_tool2"))
# .addInitialArgument("extractReasoning", true)
# .build();
return ResourceDescriptor(
clazz=ResourceName.ChatModel.JAVA_WRAPPER_SETUP,
java_clazz=ResourceName.ChatModel.Java.OLLAMA_SETUP,
connection="java_chat_model_connection",
model="qwen3:8b",
prompt="my_prompt",
tools=["my_tool1", "my_tool2"],
extract_reasoning=True,
)
@action(InputEvent.EVENT_TYPE)
@staticmethod
def process_input(event: Event, ctx: RunnerContext) -> None:
input_event = InputEvent.from_event(event)
# Create a chat request with user message
user_message = ChatMessage(
role=MessageRole.USER,
content=f"input: {input_event.input}"
)
ctx.send_event(
ChatRequestEvent(model="java_chat_model", messages=[user_message])
)
@action(ChatResponseEvent.EVENT_TYPE)
@staticmethod
def process_response(event: Event, ctx: RunnerContext) -> None:
chat_response = ChatResponseEvent.from_event(event)
response_content = chat_response.response.content
# Handle the LLM's response
# Process the response as needed for your use caseUsing Python Chat Model in Java
public class MyAgent extends Agent {
@ChatModelConnection
public static ResourceDescriptor pythonChatModelConnection() {
// In pure Python, the equivalent ResourceDescriptor would be:
// ResourceDescriptor(
// clazz=ResourceName.ChatModel.OLLAMA_CONNECTION,
// request_timeout=120.0
// )
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.PYTHON_WRAPPER_CONNECTION)
.addInitialArgument("pythonClazz", ResourceName.ChatModel.Python.OLLAMA_CONNECTION)
.addInitialArgument("request_timeout", 120.0)
.build();
}
@ChatModelSetup
public static ResourceDescriptor pythonChatModel() {
// In pure Python, the equivalent ResourceDescriptor would be:
// ResourceDescriptor(
// clazz=ResourceName.ChatModel.OLLAMA_SETUP,
// connection="pythonChatModelConnection",
// model="qwen3:8b",
// tools=["tool1", "tool2"],
// extract_reasoning=True
// )
return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.PYTHON_WRAPPER_SETUP)
.addInitialArgument("pythonClazz", ResourceName.ChatModel.Python.OLLAMA_SETUP)
.addInitialArgument("connection", "pythonChatModelConnection")
.addInitialArgument("model", "qwen3:8b")
.addInitialArgument("tools", List.of("tool1", "tool2"))
.addInitialArgument("extract_reasoning", true)
.build();
}
@Action(listenEventTypes = {InputEvent.EVENT_TYPE})
public static void processInput(Event event, RunnerContext ctx) throws Exception {
InputEvent inputEvent = InputEvent.fromEvent(event);
ChatMessage userMessage =
new ChatMessage(MessageRole.USER, String.format("input: {%s}", inputEvent.getInput()));
ctx.sendEvent(new ChatRequestEvent("pythonChatModel", List.of(userMessage)));
}
@Action(listenEventTypes = {ChatResponseEvent.EVENT_TYPE})
public static void processResponse(Event event, RunnerContext ctx)
throws Exception {
ChatResponseEvent chatResponse = ChatResponseEvent.fromEvent(event);
String response = chatResponse.getResponse().getContent();
// Handle the LLM's response
// Process the response as needed for your use case
}
}Custom Providers
The custom provider APIs are experimental and unstable, subject to incompatible changes in future releases.
If you want to use chat models not offered by the built-in providers, you can extend the base chat classes and implement your own! The chat model system is built around two main abstract classes:
BaseChatModelConnection
Handles the connection to chat model services and provides the core chat functionality.
Python
class MyChatModelConnection(BaseChatModelConnection):
def chat(
self,
messages: Sequence[ChatMessage],
tools: List[Tool] | None = None,
**kwargs: Any,
) -> ChatMessage:
# Core method: send messages to LLM and return response
# - messages: Input message sequence
# - tools: Optional list of tools available to the model
# - kwargs: Additional parameters from model_kwargs
# - Returns: ChatMessage with the model's response
passJava
public class MyChatModelConnection extends BaseChatModelConnection {
/**
* Creates a new chat model connection.
*
* @param descriptor a resource descriptor contains the initial parameters
* @param resourceContext context for resolving resources (e.g., tools) by name and type
*/
public MyChatModelConnection(
ResourceDescriptor descriptor, ResourceContext resourceContext) {
super(descriptor, resourceContext);
// get custom arguments from descriptor
String endpoint = descriptor.getArgument("endpoint");
...
}
@Override
public ChatMessage chat(
List<ChatMessage> messages, List<Tool> tools, Map<String, Object> arguments) {
// Core method: send messages to LLM and return response
// - messages: Input message sequence
// - tools: Optional list of tools available to the model
// - arguments: Additional parameters from ChatModelSetup
// - Returns: ChatMessage with the model's response
}
}BaseChatModelSetup
The setup class acts as a high-level configuration interface that defines which connection to use and how to configure the chat model.
Python
class MyChatModelSetup(BaseChatModelSetup):
# Add your custom configuration fields here
@property
def model_kwargs(self) -> Dict[str, Any]:
# Return model-specific configuration passed to chat()
# This dictionary is passed as **kwargs to the chat() method
return {"model": self.model, "temperature": 0.7, ...}Java
public class MyChatModelSetup extends BaseChatModelSetup {
// Add your custom configuration fields here
@Override
public Map<String, Object> getParameters() {
Map<String, Object> params = new HashMap<>();
params.put("model", model);
...
// Return model-specific configuration passed to chat()
// This dictionary is passed as arguments to the chat() method
return params;
}
}Built-in Events and Actions
The built-in chat_model_action listens to ChatRequestEvent and ToolResponseEvent. To request a chat completion, send a ChatRequestEvent. If the model returns a final answer, the action sends a ChatResponseEvent.
If the model asks to call tools, chat_model_action sends a ToolRequestEvent instead of a final ChatResponseEvent. After the tools finish, it receives the matching ToolResponseEvent, appends the tool results to the chat history, and calls the model again. This loop continues until the model returns a final response. For details on how tools are executed, see Built-in Events and Actions in Tool Use.
评论
登录后参与评论
KnowForge