This document covers the clients designed to interact with OpenAI's language models, both directly via the OpenAI API and through Microsoft's Azure OpenAI Service.
This class provides a client for OpenAI's chat completion models (e.g., GPT-3.5 Turbo, GPT-4).
- Purpose: It acts as an interface to OpenAI's chat models, handling request construction, API interaction, and response parsing. It standardizes these operations within the SEMOSS
genai_clientframework. - Relationship to Framework: It extends
AbstractOpenAiClient(which likely extendsAbstractTextGenerationClient), inheriting common OpenAI client setup, template management, and model limit considerations. It implements the coreask_call()via itsChatoperation delegate.
The OpenAITextGenerationClient (via its parent AbstractOpenAiClient) is initialized with:
api_key(str): The OpenAI API key.model_name(str): The specific OpenAI model ID (e.g., "gpt-4", "gpt-3.5-turbo").base_url(Optional[str]): For OpenAI-compatible APIs that are not hosted by OpenAI (e.g., local vLLM server, NVIDIA NIMs). If provided,api_keymight be set to "EMPTY".timeout(Optional[float]): Request timeout.max_retries(Optional[int]): Number of retries for API calls.model_type(Optional[str]): Can be "OPEN_AI" or "VLLM" to adjust for minor API differences, especially for structured output/JSON mode.use_max_tokens_param(Optional[bool]): If true, usesmax_tokensin requests; otherwise, usesmax_completion_tokens. Defaults toFalse.**kwargs: Passed toAbstractTextGenerationClientfor template and model limit setup.
The constructor initializes openai.OpenAI client and instantiates Instruct and Chat operation classes.
-
ask_call(self, **kwargs) -> AskModelEngineResponse:- This method is the primary interface for chat-like interactions.
- It delegates the actual API call to
self.chat_operation.ask(**kwargs). - The
Chat.ask()method (detailed underAbstractOpenAiClientdocumentation, but core logic resides ininference_call) handles:- Preparing the
messagespayload in the format OpenAI expects (list of dictionaries with "role" and "content"). - Managing various parameters like
temperature,max_tokens(ormax_completion_tokens),top_p,stream,tools,tool_choice. - Calling
self.client.chat.completions.create(). - Handling streaming responses by iterating through chunks and concatenating content.
- Parsing non-streaming responses, including handling
tool_callsif present. - Returning an
AskModelEngineResponsewith the text response, token counts, and message type ("CHAT" or "TOOL").
- Preparing the
-
inference_call(self, prefix: str, **kwargs) -> Tuple[str, int, str]: (Defined inAbstractOpenAiClientbut executed byOpenAiChatCompletioninstance)- This is a central method used by
Chat.ask(). - Structured Output/JSON Mode: If
schemais provided inkwargs, it calls_structured_output_call()to attempt to get JSON output adhering to the schema. - Tool Handling: If
toolsare provided, it setstool_choiceto "auto" if not specified and disables streaming. - Parameter Naming: Calls
resolve_token_param_naming()to use eithermax_tokensormax_completion_tokensbased onuse_max_tokens_param. - Model-Specific Kwargs: Calls
_update_model_specific_kwargs()to adjust parameters for compatibility with specific models like "o1-mini" (e.g., forcing temperature to 1.0, disabling streaming, converting system messages). - Makes the call to
self.client.chat.completions.create(). - Parses the response, handling both regular text and
tool_calls. - Returns the final text/tool result, response tokens, and message type.
- This is a central method used by
-
Structured Output Helpers:
_validate_structured_input(): Validates if a schema is a JSON string, dict, or Pydantic model._create_structured_response_format(): Creates theresponse_formatorguided_jsonparameter based onmodel_type(OpenAI vs. vLLM) and schema type._get_structured_output_response(): Makes the API call for structured output._structured_output_call(): Orchestrates the structured output process.
-
Token Limit Handling:
_truncate_by_tokens(): Truncates messages (oldest non-system first) if total tokens exceedsafe_window.check_token_limits(): Calculates prompt tokens, truncates if necessary, and adjustsmax_completion_tokensto fit the model's context window.
-
Image Handling (
_handle_image_params): (Defined inAbstractOpenAiClient)- Formats image inputs (URL or base64) into the structure expected by OpenAI's multimodal chat completion API.
- Uses the
openai.OpenAIclient from the official Python SDK. - Constructs requests for the
chat.completions.createendpoint. - Handles parameters like
model,messages,temperature,max_tokens,stream,tools,tool_choice,response_format.
This class provides a client for OpenAI models deployed via Microsoft Azure OpenAI Service.
-
Purpose: To enable interaction with OpenAI models using Azure-specific endpoints, API keys, and deployment names, while maintaining a consistent interface with other OpenAI clients.
-
Relationship to Framework: It extends
OpenAITextGenerationClient. This means it inherits most of the functionality for request preparation, response handling, streaming, tool use, and structured output. -
Key Differences:
- Initialization:
endpoint(str): The Azure OpenAI service endpoint URL (e.g.,https://your-resource-name.openai.azure.com/). Required.model_name(str, optional): While passed to the parent, for Azure, theazure_deploymentname (passed asdeployment_idinkwargsto the parent, or implicitly themodel_name) is often more critical for identifying the deployed model.api_key(str, default: "EMPTY"): The Azure OpenAI API key.api_version(str, default: "2023-07-01-preview"): The API version for Azure OpenAI.- The constructor calls the parent
OpenAiChatCompletionconstructor, passing along these Azure-specific parameters which are then used in_get_client.
- Client Instantiation (
_get_client):- This method is overridden to initialize and return an
openai.AzureOpenAIclient instance. - It uses the
api_key,azure_endpoint, andapi_versionfor configuration.
- This method is overridden to initialize and return an
- Tokenizer Initialization (
_get_tokenizer):- It attempts to get a tokenizer for
self.model_name. If this fails (e.g., ifmodel_nameis a deployment ID not directly recognized bytiktoken), it defaults to using a tokenizer for "gpt-3.5-turbo" (or anopenai_model_nameif provided ininit_args). This is important because Azure deployment names can be custom.
- It attempts to get a tokenizer for
- Initialization:
-
Leveraging Parent Class: Most other functionalities (API call logic, streaming, tool handling, structured output, token limit checks) are inherited directly from
OpenAITextGenerationClientandAbstractOpenAiClient. The key change is that these methods will operate using theAzureOpenAIclient instance.
These clients provide a standardized way to leverage OpenAI and Azure OpenAI models for advanced text generation tasks within SEMOSS, including chat, instruction following, tool use, and structured data extraction.