> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.mixlayer.com/choosing-an-inference-api/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.mixlayer.com/_mcp/server. # Choosing an Inference API > Decide between Chat Completions and Responses for your Mixlayer application Mixlayer supports both the OpenAI-compatible Chat Completions and Responses APIs. They run the same catalog of generation models, but use different request shapes and conversation patterns. > **Info** > > **Start with Chat Completions for broad client compatibility.** Choose > Responses when your application benefits from typed input and output items, > stored continuation, hosted tools, or Responses-style streaming. ## Chat Completions [Chat Completions](/api-reference/inference-ap-is/chat-completions/chat-completions) uses a `messages` array containing `system`, `user`, `assistant`, and `tool` turns. Choose it when: * You are connecting an existing OpenAI-compatible application or library. * Your application already stores and resends conversation history. * You want the most widely supported API shape across third-party tools. * You need standard function calling, structured output, vision, reasoning, or streaming without Responses-specific state. Your application owns conversation state. To continue a conversation, send the previous messages again with the next request. ## Responses [Responses](/api-reference/inference-ap-is/responses/responses) accepts a string or typed `input` items and returns typed output items and events. Choose it when: * You need `function_call` and `function_call_output` items. * You want to continue from a stored response with `previous_response_id`. * You use [server-side tools](/server-side-tools) such as web search. * You want typed streaming events or the [Responses WebSocket endpoint](/api-reference/inference-ap-is/responses/responses-websocket). * Your client already implements the OpenAI Responses API. Stored continuation is optional. Send `store: true` to make a response available for a later `previous_response_id`, subject to your organization's [data-retention policy](/zero-data-retention). ## Feature comparison | Capability | Chat Completions | Responses | | ------------------- | ---------------------- | ------------------------------------------ | | Primary input | `messages` | String or typed `input` items | | Conversation state | Resend message history | Resend input or use `previous_response_id` | | Function results | `tool` message | `function_call_output` item | | Hosted web search | `web_search_options` | `web_search` tool | | HTTP streaming | SSE chunks | Typed SSE events | | WebSocket transport | No | Yes | | Third-party support | Broadest | Growing; client must support Responses | Both APIs support function tools, structured text output, vision, and reasoning on compatible models. The item shapes differ, so follow the examples for the API you choose. ## Compatibility boundaries OpenAI compatibility describes the request and response shape; it does not mean every OpenAI field or tool type is supported. Use the generated API Reference as the source of truth for accepted fields and values. Unsupported input can be rejected even when an OpenAI client successfully connects. See [Client Libraries](/client-libraries) for configuration examples and [Tool Calling](/tool-calling) for the function-call loop in both APIs. Embeddings use a separate endpoint; see [Generating embeddings](/embeddings#generating-embeddings). > Decide between Chat Completions and Responses for your Mixlayer application