Chat Models

qianmoQqianmoQ· 更新于 2026-10-08· 阅读 104 分钟· 0 次阅读

登录后可跨设备保存划线和私人笔记登录

Chat Models

Overview

Chat models enable agents to communicate with Large Language Models (LLMs) for natural language understanding, reasoning, and generation. In Flink Agents, chat models act as the “brain” of your agents, processing input messages and generating intelligent responses based on context, prompts, and available tools.

Getting Started

To use chat models in your agents, you need to define both a connection and setup using decorators/annotations, then interact with the model through events.

Resource Declaration

Flink Agents provides decorators in python and annotations in java to simplify chat model setup within agents:

@chat_model_connection/@ChatModelConnection

The @chat_model_connection decorator or @ChatModelConnection annotation marks a method that creates a chat model connection. This is typically defined once and shared across multiple chat model setups.

@chat_model_setup/@ChatModelSetup

The @chat_model_setup decorator or @ChatModelSetup annotation marks a method that creates a chat model setup. This references a connection and adds chat-specific configuration like prompts and tools.

Chat Events

Chat models communicate through built-in events:

  • ChatRequestEvent: Sent by actions to request a chat completion from the LLM
  • ChatResponseEvent: Received by actions containing the LLM’s response

Usage Example

Here’s how to define and use chat models in a workflow agent:

Python

class MyAgent(Agent):

    @chat_model_connection
    @staticmethod
    def ollama_connection() -> ResourceDescriptor:
        return ResourceDescriptor(
            clazz=ResourceName.ChatModel.OLLAMA_CONNECTION,
            base_url="http://localhost:11434",
            request_timeout=30.0
        )

    @chat_model_setup
    @staticmethod
    def ollama_chat_model() -> ResourceDescriptor:
        return ResourceDescriptor(
            clazz=ResourceName.ChatModel.OLLAMA_SETUP,
            connection="ollama_connection",
            model="qwen3:8b",
            temperature=0.7
        )

    @action(InputEvent.EVENT_TYPE)
    @staticmethod
    def process_input(event: Event, ctx: RunnerContext) -> None:
        input_event = InputEvent.from_event(event)
        # Create a chat request with user message
        user_message = ChatMessage(
            role=MessageRole.USER,
            content=f"input: {input_event.input}"
        )
        ctx.send_event(
            ChatRequestEvent(model="ollama_chat_model", messages=[user_message])
        )

    @action(ChatResponseEvent.EVENT_TYPE)
    @staticmethod
    def process_response(event: Event, ctx: RunnerContext) -> None:
        chat_response = ChatResponseEvent.from_event(event)
        response_content = chat_response.response.content
        # Handle the LLM's response
        # Process the response as needed for your use case

Java

public class MyAgent extends Agent {
    @ChatModelConnection
    public static ResourceDescriptor ollamaConnection() {
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.OLLAMA_CONNECTION)
                .addInitialArgument("endpoint", "http://localhost:11434")
                .build();
    }

    @ChatModelSetup
    public static ResourceDescriptor ollamaChatModel() {
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.OLLAMA_SETUP)
                .addInitialArgument("connection", "ollamaConnection")
                .addInitialArgument("model", "qwen3:8b")
                .build();
    }

    @Action(listenEventTypes = {InputEvent.EVENT_TYPE})
    public static void processInput(Event event, RunnerContext ctx) throws Exception {
        InputEvent inputEvent = InputEvent.fromEvent(event);
        ChatMessage userMessage =
                new ChatMessage(MessageRole.USER, String.format("input: {%s}", inputEvent.getInput()));
        ctx.sendEvent(new ChatRequestEvent("ollamaChatModel", List.of(userMessage)));
    }

    @Action(listenEventTypes = {ChatResponseEvent.EVENT_TYPE})
    public static void processResponse(Event event, RunnerContext ctx)
            throws Exception {
        ChatResponseEvent chatResponse = ChatResponseEvent.fromEvent(event);
        String response = chatResponse.getResponse().getContent();
        // Handle the LLM's response
        // Process the response as needed for your use case
    }
}

Built-in Providers

Amazon Bedrock

Amazon Bedrock provides access to a wide range of foundation models from leading AI providers through a unified API. The Flink Agents Bedrock integration uses the Converse API, which provides a consistent interface across all supported models with native tool calling support. Authentication is handled via SigV4 using the AWS default credentials chain. No API keys are required.

Amazon Bedrock is only supported in Java currently. To use Amazon Bedrock from Python agents, see Using Cross-Language Providers.

Prerequisites

  1. An AWS account with Amazon Bedrock model access enabled for the models you plan to use
  2. IAM credentials configured via any method supported by the AWS Default Credentials Provider (environment variables, ~/.aws/credentials, IAM role, etc.)

BedrockChatModelConnection Parameters

Java

ParameterTypeDefaultDescription
regionString"us-east-1"AWS region for the Bedrock service
modelStringNoneDefault model ID (can be overridden per setup)
max_retriesint5Maximum number of API retry attempts (retries on throttling, 429, 503)

BedrockChatModelSetup Parameters

Java

ParameterTypeDefaultDescription
connectionStringRequiredReference to connection method name
modelStringRequiredBedrock model ID (e.g. "us.anthropic.claude-sonnet-4-20250514-v1:0")
promptPrompt \StringNonePrompt template or reference to prompt resource
toolsListNoneList of tool names available to the model
temperaturedouble0.1Sampling temperature (0.0 to 1.0)
max_tokensintNoneMaximum number of tokens to generate

Usage Example

Java

public class MyAgent extends Agent {
    @ChatModelConnection
    public static ResourceDescriptor bedrockConnection() {
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.BEDROCK_CONNECTION)
                .addInitialArgument("region", "us-east-1")
                .build();
    }

    @ChatModelSetup
    public static ResourceDescriptor bedrockChatModel() {
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.BEDROCK_SETUP)
                .addInitialArgument("connection", "bedrockConnection")
                .addInitialArgument("model", "us.anthropic.claude-sonnet-4-20250514-v1:0")
                .addInitialArgument("temperature", 0.1d)
                .addInitialArgument("max_tokens", 4096)
                .build();
    }

    ...
}

Available Models

Amazon Bedrock supports models from multiple providers through a single API. Visit the Amazon Bedrock Model IDs documentation for the complete and up-to-date list of available models.

Some popular options include:

  • Claude (Anthropic): us.anthropic.claude-sonnet-4-6, us.anthropic.claude-opus-4-7, us.anthropic.claude-opus-4-6-v1
  • Llama (Meta): us.meta.llama4-scout-17b-16e-instruct-v1:0
  • Mistral: mistral.mistral-large-2402-v1:0
  • Amazon Nova: us.amazon.nova-pro-v1:0, us.amazon.nova-lite-v1:0

Model availability varies by AWS region and requires explicit model access enablement in the Bedrock console. Always check the Amazon Bedrock documentation for regional availability before implementing in production.

Current limitations: The integration uses text content blocks only. Extended thinking / reasoning content blocks (e.g. Claude extended thinking), citation blocks, and image / document content blocks are not yet supported.

Anthropic

Anthropic provides cloud-based chat models featuring the Claude family, known for their strong reasoning, coding, and safety capabilities.

Prerequisites

  1. Create an account at Anthropic Console
  2. Navigate to API Keys and create a new secret key

AnthropicChatModelConnection Parameters

Python

ParameterTypeDefaultDescription
api_keystrRequiredAnthropic API key for authentication
max_retriesint3Maximum number of API retry attempts
timeoutfloat60.0API request timeout in seconds

Java

ParameterTypeDefaultDescription
api_keyStringRequiredAnthropic API key for authentication
timeoutintNoneTimeout in seconds for API requests
max_retriesint2Maximum number of API retry attempts

AnthropicChatModelSetup Parameters

Python

ParameterTypeDefaultDescription
connectionstrRequiredReference to connection method name
modelstr"claude-sonnet-4-20250514"Name of the chat model to use
promptPrompt \strNonePrompt template or reference to prompt resource
toolsList[str]NoneList of tool names available to the model
max_tokensint1024Maximum number of tokens to generate
temperaturefloat0.1Sampling temperature (0.0 to 1.0)
additional_kwargsdict{}Additional Anthropic API parameters

Java

ParameterTypeDefaultDescription
connectionStringRequiredReference to connection method name
modelString"claude-sonnet-4-20250514"Name of the chat model to use
promptPrompt \StringNonePrompt template or reference to prompt resource
toolsListNoneList of tool names available to the model
max_tokenslong1024Maximum number of tokens to generate
temperaturedouble0.1Sampling temperature (0.0 to 1.0)
json_prefillbooleantruePrefill assistant response with “{” to enforce JSON output (disabled when tools are used)
strict_toolsbooleanfalseEnable strict mode for tool calling schemas
additional_kwargsMap<String, Object>{}Additional Anthropic API parameters (top_k, top_p, stop_sequences)

Usage Example

Python

class MyAgent(Agent):

    @chat_model_connection
    @staticmethod
    def anthropic_connection() -> ResourceDescriptor:
        return ResourceDescriptor(
            clazz=ResourceName.ChatModel.ANTHROPIC_CONNECTION,
            api_key="<your-api-key>",
            max_retries=3,
            timeout=60.0
        )

    @chat_model_setup
    @staticmethod
    def anthropic_chat_model() -> ResourceDescriptor:
        return ResourceDescriptor(
            clazz=ResourceName.ChatModel.ANTHROPIC_SETUP,
            connection="anthropic_connection",
            model="claude-sonnet-4-20250514",
            max_tokens=2048,
            temperature=0.7
        )

    ...

Java

public class MyAgent extends Agent {
    @ChatModelConnection
    public static ResourceDescriptor anthropicConnection() {
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.ANTHROPIC_CONNECTION)
                .addInitialArgument("api_key", "<your-api-key>")
                .addInitialArgument("timeout", 120)
                .addInitialArgument("max_retries", 3)
                .build();
    }

    @ChatModelSetup
    public static ResourceDescriptor anthropicChatModel() {
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.ANTHROPIC_SETUP)
                .addInitialArgument("connection", "anthropicConnection")
                .addInitialArgument("model", "claude-sonnet-4-20250514")
                .addInitialArgument("temperature", 0.7d)
                .addInitialArgument("max_tokens", 2048)
                .build();
    }

    ...
}

Available Models

Visit the Anthropic Models documentation for the complete and up-to-date list of available chat models.

Some popular options include:

  • Claude Sonnet 4.5 (claude-sonnet-4-5-20250929)
  • Claude Sonnet 4 (claude-sonnet-4-20250514)
  • Claude Sonnet 3.7 (claude-3-7-sonnet-20250219)
  • Claude Opus 4.1 (claude-opus-4-1-20250805)

Model availability and specifications may change. Always check the official Anthropic documentation for the latest information before implementing in production.

Azure AI

Azure AI provides cloud-based chat models through Azure AI Inference API, supporting various models including Llama, Mistral, Phi, and other models deployed via Azure AI Studio.

Azure AI is only supported in Java currently. To use Azure AI from Python agents, see Using Cross-Language Providers.

Azure AI vs Azure OpenAI: Azure AI uses the Azure AI Inference API to access models deployed via Azure AI Studio (Llama, Mistral, Phi, etc.). If you want to use OpenAI models (GPT-4, etc.) hosted on Azure, see Azure OpenAI instead.

Prerequisites

  1. Create an Azure AI resource in the Azure Portal
  2. Obtain your endpoint URL and API key from the Azure AI resource

AzureAIChatModelConnection Parameters

Java

ParameterTypeDefaultDescription
endpointStringRequiredAzure AI service endpoint URL
apiKeyStringRequiredAzure AI API key for authentication

AzureAIChatModelSetup Parameters

Java

ParameterTypeDefaultDescription
connectionStringRequiredReference to connection method name
modelStringRequiredName of the chat model to use (e.g., “gpt-4o”)
promptPrompt \StringNonePrompt template or reference to prompt resource
toolsList[String]NoneList of tool names available to the model

Usage Example

Java

public class MyAgent extends Agent {
    @ChatModelConnection
    public static ResourceDescriptor azureAIConnection() {
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.AZURE_CONNECTION)
                .addInitialArgument("endpoint", "https://your-resource.inference.ai.azure.com")
                .addInitialArgument("apiKey", "your-api-key-here")
                .build();
    }

    @ChatModelSetup
    public static ResourceDescriptor azureAIChatModel() {
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.AZURE_SETUP)
                .addInitialArgument("connection", "azureAIConnection")
                .addInitialArgument("model", "gpt-4o")
                .build();
    }

    ...
}

Available Models

Azure AI supports various models through the Azure AI Inference API. Visit the Azure AI Model Catalog for the complete and up-to-date list of available models.

Some popular options include:

  • GPT-4o (gpt-4o)
  • GPT-4 (gpt-4)
  • GPT-4 Turbo (gpt-4-turbo)
  • GPT-3.5 Turbo (gpt-3.5-turbo)

Model availability and specifications may change. Always check the official Azure AI documentation for the latest information before implementing in production.

Azure OpenAI

Azure OpenAI provides access to OpenAI models (GPT-4, GPT-4o, etc.) through Azure’s cloud infrastructure, using the same OpenAI SDK with Azure-specific authentication and endpoints. This offers enterprise security, compliance, and regional availability while using familiar OpenAI APIs.

Azure OpenAI vs Azure AI: Azure OpenAI uses the OpenAI SDK to access OpenAI models (GPT-4, etc.) hosted on Azure. If you want to use other models like Llama, Mistral, or Phi deployed via Azure AI Studio, see Azure AI instead.

Prerequisites

  1. Create an Azure OpenAI resource in the Azure Portal
  2. Deploy a model in Azure OpenAI Studio
  3. Obtain your endpoint URL, API key, API version, and deployment name from the Azure portal

AzureOpenAIChatModelConnection Parameters

Python

ParameterTypeDefaultDescription
api_keystrRequiredAzure OpenAI API key for authentication
api_versionstrRequiredAzure OpenAI REST API version (e.g., “2024-02-15-preview”). See API versions
azure_endpointstrRequiredAzure OpenAI endpoint URL (e.g., https://{resource-name}.openai.azure.com)
timeoutfloat60.0API request timeout in seconds
max_retriesint3Maximum number of API retry attempts

Java

ParameterTypeDefaultDescription
api_keyStringRequiredAzure OpenAI API key for authentication
api_versionStringRequiredAzure OpenAI REST API version (e.g., “2024-02-01”). See API versions
azure_endpointStringRequiredAzure OpenAI endpoint URL (e.g., https://{resource-name}.openai.azure.com) — either a direct Azure resource or a proxy/gateway URL that fronts an Azure OpenAI service
timeoutintNoneTimeout in seconds for API requests; must be greater than 0, otherwise ignored (SDK default applies)
max_retriesintNoneMaximum number of API retry attempts; must be non-negative, otherwise ignored (SDK default applies)
azure_url_path_modeString"AUTO"Controls how the SDK constructs Azure OpenAI request URLs. One of "AUTO", "LEGACY", or "UNIFIED". Custom gateways that proxy Azure OpenAI typically need "LEGACY" to force the /openai/deployments/{model} path

AzureOpenAIChatModelSetup Parameters

Python

ParameterTypeDefaultDescription
connectionstrRequiredReference to connection method name
modelstrRequiredName of OpenAI model deployment on Azure
model_of_azure_deploymentstrNoneThe underlying model name (e.g., ‘gpt-4’, ‘gpt-35-turbo’). Used for token metrics tracking
promptPrompt \strNonePrompt template or reference to prompt resource
toolsList[str]NoneList of tool names available to the model
temperaturefloatNoneSampling temperature (0.0 to 2.0). Not supported by reasoning models
max_tokensintNoneMaximum number of tokens to generate
logprobsboolFalseWhether to return log probabilities of output tokens
additional_kwargsdict{}Additional Azure OpenAI API parameters

Java

ParameterTypeDefaultDescription
connectionStringRequiredReference to connection method name
modelStringRequiredAzure deployment name (not the underlying OpenAI model name)
model_of_azure_deploymentStringNoneThe underlying model name (e.g., ‘gpt-4’, ‘gpt-4o’). Used solely for token metrics tracking
promptPrompt \StringNonePrompt template or reference to prompt resource
toolsListNoneList of tool names available to the model
temperaturedoubleNoneSampling temperature (0.0 to 2.0). Not supported by reasoning models
max_tokensintNoneMaximum number of tokens to generate (must be greater than 0)
logprobsbooleanfalseWhether to return log probabilities of output tokens
additional_kwargsMap<String, Object>{}Additional Azure OpenAI API parameters (forwarded to the OpenAI request body)

Usage Example

Python

class MyAgent(Agent):

    @chat_model_connection
    @staticmethod
    def azure_openai_connection() -> ResourceDescriptor:
        return ResourceDescriptor(
            clazz=ResourceName.ChatModel.AZURE_OPENAI_CONNECTION,
            api_key="<your-api-key>",
            api_version="2024-02-15-preview",
            azure_endpoint="https://your-resource.openai.azure.com"
        )

    @chat_model_setup
    @staticmethod
    def azure_openai_chat_model() -> ResourceDescriptor:
        return ResourceDescriptor(
            clazz=ResourceName.ChatModel.AZURE_OPENAI_SETUP,
            connection="azure_openai_connection",
            model="my-gpt4-deployment",  # Your Azure deployment name
            model_of_azure_deployment="gpt-4",  # Underlying model for metrics
            max_tokens=1000
        )

    ...

Java

public class MyAgent extends Agent {
    @ChatModelConnection
    public static ResourceDescriptor azureOpenAIConnection() {
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.AZURE_OPENAI_CONNECTION)
                .addInitialArgument("api_key", "<your-api-key>")
                .addInitialArgument("api_version", "2024-02-01")
                .addInitialArgument("azure_endpoint", "https://your-resource.openai.azure.com")
                .build();
    }

    @ChatModelSetup
    public static ResourceDescriptor azureOpenAIChatModel() {
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.AZURE_OPENAI_SETUP)
                .addInitialArgument("connection", "azureOpenAIConnection")
                .addInitialArgument("model", "my-gpt4-deployment")          // Your Azure deployment name
                .addInitialArgument("model_of_azure_deployment", "gpt-4")   // Underlying model for metrics
                .addInitialArgument("temperature", 0.3d)
                .addInitialArgument("max_tokens", 1000)
                .build();
    }

    ...
}

Available Models

Azure OpenAI supports OpenAI models deployed through your Azure subscription. Visit the Azure OpenAI Models documentation for the complete and up-to-date list of available models.

Some popular options include:

  • GPT-4o (gpt-4o)
  • GPT-4 (gpt-4)
  • GPT-4 Turbo (gpt-4-turbo)
  • GPT-3.5 Turbo (gpt-35-turbo)

Model availability depends on your Azure region and subscription. Always check the official Azure OpenAI documentation for regional availability before implementing in production.

Ollama

Ollama provides local chat models that run on your machine, offering privacy, control, and no API costs.

Prerequisites

Ollama server 0.9.0 or higher is required.

  1. Install Ollama from https://ollama.com/
  2. Start the Ollama server: ollama serve
  3. Download a chat model: ollama pull qwen3:8b

OllamaChatModelConnection Parameters

Python

ParameterTypeDefaultDescription
base_urlstr"http://localhost:11434"Ollama server URL
request_timeoutfloat30.0HTTP request timeout in seconds

Java

ParameterTypeDefaultDescription
endpointString"http://localhost:11434"Ollama server URL
requestTimeoutlong10HTTP request timeout in seconds

OllamaChatModelSetup Parameters

Python

ParameterTypeDefaultDescription
connectionstrRequiredReference to connection method name
modelstrRequiredName of the chat model to use
promptPrompt \strNonePrompt template or reference to prompt resource
toolsList[str]NoneList of tool names available to the model
temperaturefloat0.75Sampling temperature (0.0 to 1.0)
num_ctxint2048Maximum number of context tokens
keep_alivestr \float"5m"How long to keep model loaded in memory
thinkbool \Literal[“low”, “medium”, “high”]TrueWhether enable model think
extract_reasoningboolTrueExtract reasoning content from response
additional_kwargsdict{}Additional Ollama API parameters

Java

ParameterTypeDefaultDescription
connectionStringRequiredReference to connection method name
modelStringRequiredName of the chat model to use
promptPrompt \StringNonePrompt template or reference to prompt resource
toolsList[String]NoneList of tool names available to the model
thinkBoolean \Literal[“low”, “medium”, “high”]trueWhether enable model think
extract_reasoningBooleantrueExtract reasoning content from response

Usage Example

Python

class MyAgent(Agent):

    @chat_model_connection
    @staticmethod
    def ollama_connection() -> ResourceDescriptor:
        return ResourceDescriptor(
            clazz=ResourceName.ChatModel.OLLAMA_CONNECTION,
            base_url="http://localhost:11434",
            request_timeout=120.0
        )

    @chat_model_setup
    @staticmethod
    def my_chat_model() -> ResourceDescriptor:
        return ResourceDescriptor(
            clazz=ResourceName.ChatModel.OLLAMA_SETUP,
            connection="ollama_connection",
            model="qwen3:8b",
            temperature=0.7,
            num_ctx=4096,
            keep_alive="10m",
            think=True,
            extract_reasoning=True
        )

    ...

Java

public class MyAgent extends Agent {
    @ChatModelConnection
    public static ResourceDescriptor ollamaConnection() {
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.OLLAMA_CONNECTION)
                .addInitialArgument("endpoint", "http://localhost:11434")
                .addInitialArgument("requestTimeout", 120)
                .build();
    }

    @ChatModelSetup
    public static ResourceDescriptor ollamaChatModel() {
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.OLLAMA_SETUP)
                .addInitialArgument("connection", "ollamaConnection")
                .addInitialArgument("model", "qwen3:8b")
                .addInitialArgument("extract_reasoning", true)
                .build();
    }

    ...
}

Available Models

Visit the Ollama Models Library for the complete and up-to-date list of available chat models.

Some popular options include:

  • qwen3 series (qwen3:8b, qwen3:14b, qwen3:32b)
  • llama3 series (llama3:8b, llama3:70b)
  • deepseek series (deepseek-r1, deepseek-v3.1)
  • gpt-oss

Model availability and specifications may change. Always check the official Ollama documentation for the latest information before implementing in production.

OpenAI

OpenAI provides cloud-based chat models with state-of-the-art performance for a wide range of natural language tasks.

Prerequisites

  1. Create an account at OpenAI Platform
  2. Navigate to API Keys and create a new secret key

Completions API

OpenAICompletionsConnection Parameters

Python

ParameterTypeDefaultDescription
api_keystrRequiredOpenAI API key for authentication
api_base_urlstr"https://api.openai.com/v1"Base URL for OpenAI API
max_retriesint3Maximum number of API retry attempts
timeoutfloat60.0API request timeout in seconds
default_headersdictNoneDefault headers for API requests
reuse_clientboolTrueWhether to reuse the OpenAI client between requests

Java

ParameterTypeDefaultDescription
api_keyStringRequiredOpenAI API key for authentication
api_base_urlString"https://api.openai.com/v1"Base URL for OpenAI API
max_retriesint2Maximum number of API retry attempts
timeoutintNoneTimeout in seconds for API requests
default_headersMap<String, String>NoneDefault headers for API requests
modelStringNoneDefault model to use if not specified in setup
OpenAICompletionsSetup Parameters

Python

ParameterTypeDefaultDescription
connectionstrRequiredReference to connection method name
modelstr"gpt-3.5-turbo"Name of the chat model to use
promptPrompt \strNonePrompt template or reference to prompt resource
toolsList[str]NoneList of tool names available to the model
temperaturefloat0.1Sampling temperature (0.0 to 2.0)
max_tokensintNoneMaximum number of tokens to generate
logprobsboolNoneWhether to return log probabilities per token
top_logprobsint0Number of top token log probabilities to return (0-20)
strictboolFalseEnable strict mode for tool calling and schemas
reasoning_effortstrNoneReasoning effort level for reasoning models (“low”, “medium”, “high”)
additional_kwargsdict{}Additional OpenAI API parameters

Java

ParameterTypeDefaultDescription
connectionStringRequiredReference to connection method name
modelString"gpt-3.5-turbo"Name of the chat model to use
promptPrompt \StringNonePrompt template or reference to prompt resource
toolsListNoneList of tool names available to the model
temperaturedouble0.1Sampling temperature (0.0 to 2.0)
max_tokensintNoneMaximum number of tokens to generate
logprobsbooleanNoneWhether to return log probabilities per token
top_logprobsint0Number of top token log probabilities to return (0-20)
strictbooleanfalseEnable strict mode for tool calling and schemas
reasoning_effortStringNoneReasoning effort level for reasoning models (“low”, “medium”, “high”)
additional_kwargsMap<String, Object>{}Additional OpenAI API parameters
Usage Example

Python

class MyAgent(Agent):

    @chat_model_connection
    @staticmethod
    def openai_connection() -> ResourceDescriptor:
        return ResourceDescriptor(
            clazz=ResourceName.ChatModel.OPENAI_COMPLETIONS_CONNECTION,
            api_key="<your-api-key>",
            api_base_url="https://api.openai.com/v1",
            max_retries=3,
            timeout=60.0
        )

    @chat_model_setup
    @staticmethod
    def openai_chat_model() -> ResourceDescriptor:
        return ResourceDescriptor(
            clazz=ResourceName.ChatModel.OPENAI_COMPLETIONS_SETUP,
            connection="openai_connection",
            model="gpt-4",
            temperature=0.7,
            max_tokens=1000
        )

    ...

Java

public class MyAgent extends Agent {
    @ChatModelConnection
    public static ResourceDescriptor openaiConnection() {
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.OPENAI_COMPLETIONS_CONNECTION)
                .addInitialArgument("api_key", "<your-api-key>")
                .addInitialArgument("api_base_url", "https://api.openai.com/v1")
                .addInitialArgument("timeout", 60)
                .addInitialArgument("max_retries", 3)
                .build();
    }

    @ChatModelSetup
    public static ResourceDescriptor openaiChatModel() {
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.OPENAI_COMPLETIONS_SETUP)
                .addInitialArgument("connection", "openaiConnection")
                .addInitialArgument("model", "gpt-4")
                .addInitialArgument("temperature", 0.7d)
                .addInitialArgument("max_tokens", 1000)
                .build();
    }

    ...
}

Responses API

Responses API is only supported in Java currently. To use OpenAI Responses API from Python agents, see Using Cross-Language Providers.

OpenAIResponsesModelConnection Parameters

Java

ParameterTypeDefaultDescription
api_keyStringRequiredOpenAI API key for authentication
api_base_urlStringNoneBase URL for OpenAI API (useful for proxies)
max_retriesint2Maximum number of API retry attempts
timeoutintNoneTimeout in seconds for API requests
default_headersMap<String, String>NoneDefault headers for API requests
modelStringNoneDefault model to use if not specified in setup
OpenAIResponsesModelSetup Parameters

Java

ParameterTypeDefaultDescription
connectionStringRequiredReference to connection method name
modelString"gpt-4o"Name of the chat model to use
promptPrompt \StringNonePrompt template or reference to prompt resource
toolsListNoneList of tool names available to the model
temperaturedouble0.1Sampling temperature (0.0 to 2.0)
max_tokensintNoneMaximum number of tokens to generate
strictbooleanfalseEnable strict mode for tool calling schemas
reasoning_effortStringNoneReasoning effort level for reasoning models (“low”, “medium”, “high”)
storebooleanfalseWhether to store the response for later retrieval
instructionsStringNoneSystem-level instructions for the model
additional_kwargsMap<String, Object>{}Additional Responses API parameters
Usage Example

Java

public class MyAgent extends Agent {
    @ChatModelConnection
    public static ResourceDescriptor openaiResponsesConnection() {
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.OPENAI_RESPONSES_CONNECTION)
                .addInitialArgument("api_key", "<your-api-key>")
                .addInitialArgument("timeout", 120)
                .addInitialArgument("max_retries", 3)
                .build();
    }

    @ChatModelSetup
    public static ResourceDescriptor openaiResponsesChatModel() {
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.OPENAI_RESPONSES_SETUP)
                .addInitialArgument("connection", "openaiResponsesConnection")
                .addInitialArgument("model", "gpt-4o")
                .addInitialArgument("temperature", 0.3d)
                .addInitialArgument("max_tokens", 2048)
                .addInitialArgument("store", true)
                .build();
    }

    ...
}

Available Models

Visit the OpenAI Models documentation for the complete and up-to-date list of available chat models.

Some popular options include:

  • GPT-5 series (GPT-5, GPT-5 mini, GPT-5 nano)
  • GPT-4.1
  • gpt-oss series (gpt-oss-120b, gpt-oss-10b)

Model availability and specifications may change. Always check the official OpenAI documentation for the latest information before implementing in production.

Tongyi (DashScope)

Tongyi provides cloud-based chat models from Alibaba Cloud, offering powerful Chinese and English language capabilities.

Tongyi is only supported in Python currently. To use Tongyi from Java agents, see Using Cross-Language Providers.

Prerequisites

  1. Get an API key from Alibaba Cloud DashScope

TongyiChatModelConnection Parameters

ParameterTypeDefaultDescription
api_keystr$DASHSCOPE_API_KEYDashScope API key for authentication
request_timeoutfloat60.0HTTP request timeout in seconds

TongyiChatModelSetup Parameters

ParameterTypeDefaultDescription
connectionstrRequiredReference to connection method name
modelstr"qwen-plus"Name of the chat model to use
promptPrompt \strNonePrompt template or reference to prompt resource
toolsList[str]NoneList of tool names available to the model
temperaturefloat0.7Sampling temperature (0.0 to 2.0)
extract_reasoningboolFalseExtract reasoning content from response
additional_kwargsdict{}Additional DashScope API parameters

Usage Example

class MyAgent(Agent):

    @chat_model_connection
    @staticmethod
    def tongyi_connection() -> ResourceDescriptor:
        return ResourceDescriptor(
            clazz=ResourceName.ChatModel.TONGYI_CONNECTION,
            api_key="your-api-key-here",  # Or set DASHSCOPE_API_KEY env var
            request_timeout=60.0
        )

    @chat_model_setup
    @staticmethod
    def tongyi_chat_model() -> ResourceDescriptor:
        return ResourceDescriptor(
            clazz=ResourceName.ChatModel.TONGYI_SETUP,
            connection="tongyi_connection",
            model="qwen-plus",
            temperature=0.7,
            extract_reasoning=True
        )

    ...

Available Models

Visit the DashScope Models documentation for the complete and up-to-date list of available chat models.

Some popular options include:

  • qwen-plus
  • qwen-max
  • qwen-turbo
  • qwen-long

Model availability and specifications may change. Always check the official DashScope documentation for the latest information before implementing in production.

Using Cross-Language Providers

Flink Agents supports cross-language chat model integration, allowing you to use chat models implemented in one language (Java or Python) from agents written in the other language. This is particularly useful when a chat model provider is only available in one language (e.g., Tongyi is currently Python-only).

Limitations:

  • Cross-language resources are currently supported only when running in Flink, not in local development mode
  • Complex object serialization between languages may have limitations

How To Use

To leverage chat model supports provided in a different language, you need to declare the resource within a built-in cross-language wrapper, and specify the target provider as an argument:

  • Using Java chat models in Python: Use ResourceName.ChatModel.JAVA_WRAPPER_CONNECTION and ResourceName.ChatModel.JAVA_WRAPPER_SETUP, specifying the Java provider class via the java_clazz parameter
  • Using Python chat models in Java: Use ResourceName.ChatModel.PYTHON_WRAPPER_CONNECTION and ResourceName.ChatModel.PYTHON_WRAPPER_SETUP, specifying the Python provider via the pythonClazz parameter

Usage Example

Using Java Chat Model in Python

class MyAgent(Agent):

    @chat_model_connection
    @staticmethod
    def java_chat_model_connection() -> ResourceDescriptor:
        # In pure Java, the equivalent ResourceDescriptor would be:
        # ResourceDescriptor.Builder
        #     .newBuilder(ResourceName.ChatModel.OLLAMA_CONNECTION)
        #     .addInitialArgument("endpoint", "http://localhost:11434")
        #     .addInitialArgument("requestTimeout", 120)
        #     .build();
        return ResourceDescriptor(
            clazz=ResourceName.ChatModel.JAVA_WRAPPER_CONNECTION,
            java_clazz=ResourceName.ChatModel.Java.OLLAMA_CONNECTION,
            endpoint="http://localhost:11434",
            requestTimeout=120,
        )

    @chat_model_setup
    @staticmethod
    def java_chat_model() -> ResourceDescriptor:
        # In pure Java, the equivalent ResourceDescriptor would be:
        # ResourceDescriptor.Builder
        #     .newBuilder(ResourceName.ChatModel.OLLAMA_SETUP)
        #     .addInitialArgument("connection", "java_chat_model_connection")
        #     .addInitialArgument("model", "qwen3:8b")
        #     .addInitialArgument("prompt", "my_prompt")
        #     .addInitialArgument("tools", List.of("my_tool1", "my_tool2"))
        #     .addInitialArgument("extractReasoning", true)
        #     .build();
        return ResourceDescriptor(
            clazz=ResourceName.ChatModel.JAVA_WRAPPER_SETUP,
            java_clazz=ResourceName.ChatModel.Java.OLLAMA_SETUP,
            connection="java_chat_model_connection",
            model="qwen3:8b",
            prompt="my_prompt",
            tools=["my_tool1", "my_tool2"],
            extract_reasoning=True,
        )

    @action(InputEvent.EVENT_TYPE)
    @staticmethod
    def process_input(event: Event, ctx: RunnerContext) -> None:
        input_event = InputEvent.from_event(event)
        # Create a chat request with user message
        user_message = ChatMessage(
            role=MessageRole.USER,
            content=f"input: {input_event.input}"
        )
        ctx.send_event(
            ChatRequestEvent(model="java_chat_model", messages=[user_message])
        )

    @action(ChatResponseEvent.EVENT_TYPE)
    @staticmethod
    def process_response(event: Event, ctx: RunnerContext) -> None:
        chat_response = ChatResponseEvent.from_event(event)
        response_content = chat_response.response.content
        # Handle the LLM's response
        # Process the response as needed for your use case

Using Python Chat Model in Java

public class MyAgent extends Agent {

    @ChatModelConnection
    public static ResourceDescriptor pythonChatModelConnection() {
        // In pure Python, the equivalent ResourceDescriptor would be:
        // ResourceDescriptor(
        //     clazz=ResourceName.ChatModel.OLLAMA_CONNECTION,
        //     request_timeout=120.0
        // )
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.PYTHON_WRAPPER_CONNECTION)
                .addInitialArgument("pythonClazz", ResourceName.ChatModel.Python.OLLAMA_CONNECTION)
                .addInitialArgument("request_timeout", 120.0)
                .build();
    }

      @ChatModelSetup
    public static ResourceDescriptor pythonChatModel() {
        // In pure Python, the equivalent ResourceDescriptor would be:
        // ResourceDescriptor(
        //     clazz=ResourceName.ChatModel.OLLAMA_SETUP,
        //     connection="pythonChatModelConnection",
        //     model="qwen3:8b",
        //     tools=["tool1", "tool2"],
        //     extract_reasoning=True
        // )
        return ResourceDescriptor.Builder.newBuilder(ResourceName.ChatModel.PYTHON_WRAPPER_SETUP)
                .addInitialArgument("pythonClazz", ResourceName.ChatModel.Python.OLLAMA_SETUP)
                .addInitialArgument("connection", "pythonChatModelConnection")
                .addInitialArgument("model", "qwen3:8b")
                .addInitialArgument("tools", List.of("tool1", "tool2"))
                .addInitialArgument("extract_reasoning", true)
                .build();
    }

    @Action(listenEventTypes = {InputEvent.EVENT_TYPE})
    public static void processInput(Event event, RunnerContext ctx) throws Exception {
        InputEvent inputEvent = InputEvent.fromEvent(event);
        ChatMessage userMessage =
                new ChatMessage(MessageRole.USER, String.format("input: {%s}", inputEvent.getInput()));
        ctx.sendEvent(new ChatRequestEvent("pythonChatModel", List.of(userMessage)));
    }

    @Action(listenEventTypes = {ChatResponseEvent.EVENT_TYPE})
    public static void processResponse(Event event, RunnerContext ctx)
            throws Exception {
        ChatResponseEvent chatResponse = ChatResponseEvent.fromEvent(event);
        String response = chatResponse.getResponse().getContent();
        // Handle the LLM's response
        // Process the response as needed for your use case
    }
}

Custom Providers

The custom provider APIs are experimental and unstable, subject to incompatible changes in future releases.

If you want to use chat models not offered by the built-in providers, you can extend the base chat classes and implement your own! The chat model system is built around two main abstract classes:

BaseChatModelConnection

Handles the connection to chat model services and provides the core chat functionality.

Python

class MyChatModelConnection(BaseChatModelConnection):

    def chat(
        self,
        messages: Sequence[ChatMessage],
        tools: List[Tool] | None = None,
        **kwargs: Any,
    ) -> ChatMessage:
        # Core method: send messages to LLM and return response
        # - messages: Input message sequence
        # - tools: Optional list of tools available to the model
        # - kwargs: Additional parameters from model_kwargs
        # - Returns: ChatMessage with the model's response
        pass

Java

public class MyChatModelConnection extends BaseChatModelConnection {

    /**
     * Creates a new chat model connection.
     *
     * @param descriptor a resource descriptor contains the initial parameters
     * @param resourceContext context for resolving resources (e.g., tools) by name and type
     */
    public MyChatModelConnection(
            ResourceDescriptor descriptor, ResourceContext resourceContext) {
        super(descriptor, resourceContext);
        // get custom arguments from descriptor
        String endpoint = descriptor.getArgument("endpoint");
        ...
    }

    @Override
    public ChatMessage chat(
            List<ChatMessage> messages, List<Tool> tools, Map<String, Object> arguments) {
        // Core method: send messages to LLM and return response
        // - messages: Input message sequence
        // - tools: Optional list of tools available to the model
        // - arguments: Additional parameters from ChatModelSetup
        // - Returns: ChatMessage with the model's response
    }
}

BaseChatModelSetup

The setup class acts as a high-level configuration interface that defines which connection to use and how to configure the chat model.

Python

class MyChatModelSetup(BaseChatModelSetup):
    # Add your custom configuration fields here

    @property
    def model_kwargs(self) -> Dict[str, Any]:
        # Return model-specific configuration passed to chat()
        # This dictionary is passed as **kwargs to the chat() method
        return {"model": self.model, "temperature": 0.7, ...}

Java

public class MyChatModelSetup extends BaseChatModelSetup {
    // Add your custom configuration fields here

    @Override
    public Map<String, Object> getParameters() {
        Map<String, Object> params = new HashMap<>();
        params.put("model", model);
        ...
        // Return model-specific configuration passed to chat()
        // This dictionary is passed as arguments to the chat() method
        return params;
    }
}

Built-in Events and Actions

The built-in chat_model_action listens to ChatRequestEvent and ToolResponseEvent. To request a chat completion, send a ChatRequestEvent. If the model returns a final answer, the action sends a ChatResponseEvent.

If the model asks to call tools, chat_model_action sends a ToolRequestEvent instead of a final ChatResponseEvent. After the tools finish, it receives the matching ToolResponseEvent, appends the tool results to the chat history, and calls the model again. This loop continues until the model returns a final response. For details on how tools are executed, see Built-in Events and Actions in Tool Use.

评论

登录后参与评论

正在加载评论…