> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.mixlayer.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.mixlayer.com/_mcp/server.

# Regional Routing

> Understand Mixlayer's global architecture, choose a regional endpoint, and configure multi-region inference routing and fallback

Mixlayer runs a global inference platform across multiple regions. Send requests to `https://mixlayer.ai/v1` to reach our nearest edge, or use a regional hostname to choose where your traffic enters the platform.

By default, Mixlayer processes inference traffic in our nearest available region. You can customise this behaviour and how traffic falls back by configuring a custom routing policy for your organization.

## How global routing works

Each request has a **gateway region**, where your request enters Mixlayer's network, and an **inference region**, where the model runs. These can be the same region or different regions.

1. Your application connects to `mixlayer.ai`. The global endpoint routes the connection to our nearest edge.
2. Mixlayer selects an inference region using your organization's routing policy and the regions where the requested model is available.
3. If the selected model is not available in the first region, Mixlayer can try another eligible region, subject to your fallback policy.
4. Once inference starts, the request stays in the inference region where it started.

By default, inference routing is latency-based (`mode: "nearest"`). It prefers the gateway region when the model is available there, then regions in the same country by latency, then other regions by latency. These latency estimates are between the gateway and inference regions.

The default fallback policy allows other regions visible to your organization. Routing always respects your organization's region access and the model's regional availability.

## Choose where traffic enters

Use the global endpoint for automatic entry routing. To have traffic enter Mixlayer's network in a specific region, use its regional hostname with the same API paths and authentication:

| Endpoint                            | Entry routing                                                                                  |
| ----------------------------------- | ---------------------------------------------------------------------------------------------- |
| `https://mixlayer.ai/v1`            | Traffic enters Mixlayer's network at our nearest edge.                                         |
| `https://<region>.mixlayer.ai/v1`   | Traffic enters Mixlayer's network in the named region, where a regional endpoint is available. |
| `https://us-central.mixlayer.ai/v1` | Traffic enters Mixlayer's network in `us-central`.                                             |

For example, send a Responses API request with `us-central` as the gateway region:

```bash
curl https://us-central.mixlayer.ai/v1/responses \
  -H "Authorization: Bearer $MIXLAYER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/qwen3.5-4b-free",
    "input": "Tell me a fun fact about chihuahuas."
  }'
```

For an OpenAI-compatible client, set its `base_url` to the regional URL. See [Client Libraries](/client-libraries) for client setup examples.

> **Info**
>
> A regional hostname selects the gateway region. It does not pin inference to that region. Mixlayer still applies your organization's routing and fallback policy, and the model may run elsewhere.

## Find a model's regions

Regional model availability can differ. Before choosing inference regions, list the models available to your organization:

```bash
export ORG_ID="your-organization-id"

curl "https://api.mixlayer.com/v1/organizations/$ORG_ID/models" \
  -H "Authorization: Bearer $MIXLAYER_API_KEY"
```

Each model includes a `regions` array containing its deployment regions that your organization can access. To list models available in a particular region, add a `region` query parameter:

```bash
curl "https://api.mixlayer.com/v1/organizations/$ORG_ID/models?region=us-central" \
  -H "Authorization: Bearer $MIXLAYER_API_KEY"
```

This catalog describes model availability; your routing policy can narrow the regions actually used for inference further. An inference region must be accessible to your organization, host the requested model, and be reachable from the gateway region. If the policy lists `regions`, it must also be in that list.

## Configure your organization's routing policy

Routing policies apply across your organization's inference requests and are managed through the [Platform API](/api-reference/platform-api/platform-api).

Read the current policy:

```bash
curl "https://api.mixlayer.com/v1/organizations/$ORG_ID/routing_policy" \
  -H "Authorization: Bearer $MIXLAYER_API_KEY"
```

The response wraps the policy in `data`. A default policy looks like:

```json
{
  "data": {
    "mode": "nearest",
    "fallback": "all",
    "regions": []
  }
}
```

| Field      | Purpose                                                                                                                                                                        | Default   |
| ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------- |
| `mode`     | `nearest` uses latency-based routing. `ordered` uses the priority order in `regions`. `required` requires a `Mxl-Region` header on each request.                               | `nearest` |
| `fallback` | `all` allows fallback to other eligible regions. `same_country` limits fallback to the first selected inference region's country. `disabled` tries only the first region.      | `all`     |
| `regions`  | Limits inference to the listed region IDs. List order sets priority in `ordered` mode; `ordered` requires a non-empty list. An empty list or `null` uses every visible region. | `[]`      |

Read operations require an API key with `api-read`. Replacing the policy requires an organization owner; API keys require `api-admin`.

> **Warning**
>
> `PUT` replaces the whole routing policy. Omitted fields reset to their defaults. Send all three fields when you want to preserve explicit settings. A successful update returns `204 No Content`.

### Use latency-based routing

Use `nearest` to let Mixlayer choose the inference region automatically. This example explicitly restores the default routing and fallback behavior:

```bash
curl -X PUT "https://api.mixlayer.com/v1/organizations/$ORG_ID/routing_policy" \
  -H "Authorization: Bearer $MIXLAYER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "mode": "nearest",
    "fallback": "all",
    "regions": []
  }'
```

To limit latency-based routing to a subset of regions, supply that subset in `regions`. The list order does not affect `nearest` routing.

### Prefer regions in a fixed order

Use `ordered` when you want a primary inference region followed by a backup, regardless of latency. The following example prefers `us-central`, then `us-east`:

```bash
curl -X PUT "https://api.mixlayer.com/v1/organizations/$ORG_ID/routing_policy" \
  -H "Authorization: Bearer $MIXLAYER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "mode": "ordered",
    "fallback": "all",
    "regions": ["us-central", "us-east"]
  }'
```

The region IDs in this and the following examples are illustrative. Replace them with regions available to your organization and the model you intend to use.

If both regions host the model, Mixlayer tries `us-central` first and can fall back to `us-east` when the selected model is not available in `us-central`. If the model is deployed only in `us-east`, that is the first eligible region. Regions outside the list are never used for inference.

### Keep fallback within the same country

Use `same_country` to allow multi-region inference while keeping fallback in the country of the first selected inference region:

```bash
curl -X PUT "https://api.mixlayer.com/v1/organizations/$ORG_ID/routing_policy" \
  -H "Authorization: Bearer $MIXLAYER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "mode": "nearest",
    "fallback": "same_country",
    "regions": []
  }'
```

For example, if the first inference region is `us-central`, another eligible US region such as `us-east` can be a fallback. A region in another country is excluded even if it is otherwise available.

`same_country` follows the first selected **inference region's** country, which can differ from the gateway region or your application's location. To constrain inference to a particular country, also set `regions` to an explicit list of regions in that country. You can combine this with `ordered`:

```bash
curl -X PUT "https://api.mixlayer.com/v1/organizations/$ORG_ID/routing_policy" \
  -H "Authorization: Bearer $MIXLAYER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "mode": "ordered",
    "fallback": "same_country",
    "regions": ["us-central", "us-east"]
  }'
```

### Run inference in one region

To keep inference in one region, restrict `regions` to that region and disable fallback:

```bash
curl -X PUT "https://api.mixlayer.com/v1/organizations/$ORG_ID/routing_policy" \
  -H "Authorization: Bearer $MIXLAYER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "mode": "ordered",
    "fallback": "disabled",
    "regions": ["us-central"]
  }'
```

Requests fail if the model cannot be served in that region. They are not sent to another region. You can pair this policy with `us-central.mixlayer.ai` to select both the gateway region and the inference region.

## Prefer an inference region per request

Send the `Mxl-Region` request header to try a specific inference region first. The region must be allowed by your organization's policy and host the requested model. The header takes priority over both `nearest` and `ordered` selection; the fallback policy still applies.

```bash
curl -i https://mixlayer.ai/v1/responses \
  -H "Authorization: Bearer $MIXLAYER_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Mxl-Region: us-central" \
  -d '{
    "model": "qwen/qwen3.5-4b-free",
    "input": "Tell me a fun fact about chihuahuas."
  }'
```

The `Mxl-Inference-Region` response header identifies the region that served HTTP inference, including streaming responses and embeddings. The example uses `curl -i` to display response headers.

Use `mode: "required"` if every request must specify `Mxl-Region`. Combine it with `fallback: "disabled"` to require execution in the requested region. Responses WebSocket clients send `Mxl-Region` on the upgrade request; the serving region is not reported on that surface.

## When fallback happens

Fallback happens when the selected model is not available in a region at the start of inference. It does not restart an in-progress generation or retry a timeout or another inference error. Requests remain subject to platform attempt limits, so a long region list does not guarantee that every region will be tried.

If no eligible region can serve the model, the request fails instead of widening your routing policy. If the selected model is not available in any attempted region, the request returns `503` with the code `no_model_server_available`.