Skip to navigation

Regional Routing

Mixlayer runs a global inference platform across multiple regions. Send requests to https://mixlayer.ai/v1 to reach our nearest edge, or use a regional hostname to choose where your traffic enters the platform.

By default, Mixlayer processes inference traffic in our nearest available region. You can customise this behaviour and how traffic falls back by configuring a custom routing policy for your organization.

How global routing works

Each request has a gateway region, where your request enters Mixlayer’s network, and an inference region, where the model runs. These can be the same region or different regions.

  1. Your application connects to mixlayer.ai. The global endpoint routes the connection to our nearest edge.
  2. Mixlayer selects an inference region using your organization’s routing policy and the regions where the requested model is available.
  3. If the selected model is not available in the first region, Mixlayer can try another eligible region, subject to your fallback policy.
  4. Once inference starts, the request stays in the inference region where it started.

By default, inference routing is latency-based (mode: "nearest"). It prefers the gateway region when the model is available there, then regions in the same country by latency, then other regions by latency. These latency estimates are between the gateway and inference regions.

The default fallback policy allows other regions visible to your organization. Routing always respects your organization’s region access and the model’s regional availability.

Choose where traffic enters

Use the global endpoint for automatic entry routing. To have traffic enter Mixlayer’s network in a specific region, use its regional hostname with the same API paths and authentication:

EndpointEntry routing
https://mixlayer.ai/v1Traffic enters Mixlayer’s network at our nearest edge.
https://<region>.mixlayer.ai/v1Traffic enters Mixlayer’s network in the named region, where a regional endpoint is available.
https://us-central.mixlayer.ai/v1Traffic enters Mixlayer’s network in us-central.

For example, send a Responses API request with us-central as the gateway region:

curl https://us-central.mixlayer.ai/v1/responses \
-H "Authorization: Bearer $MIXLAYER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen/qwen3.5-4b-free",
"input": "Tell me a fun fact about chihuahuas."
}'

For an OpenAI-compatible client, set its base_url to the regional URL. See Client Libraries for client setup examples.

A regional hostname selects the gateway region. It does not pin inference to that region. Mixlayer still applies your organization’s routing and fallback policy, and the model may run elsewhere.

Find a model’s regions

Regional model availability can differ. Before choosing inference regions, list the models available to your organization:

export ORG_ID="your-organization-id"
curl "https://api.mixlayer.com/v1/organizations/$ORG_ID/models" \
-H "Authorization: Bearer $MIXLAYER_API_KEY"

Each model includes a regions array containing its deployment regions that your organization can access. To list models available in a particular region, add a region query parameter:

curl "https://api.mixlayer.com/v1/organizations/$ORG_ID/models?region=us-central" \
-H "Authorization: Bearer $MIXLAYER_API_KEY"

This catalog describes model availability; your routing policy can narrow the regions actually used for inference further. An inference region must be accessible to your organization, host the requested model, and be reachable from the gateway region. If the policy lists regions, it must also be in that list.

Configure your organization’s routing policy

Routing policies apply across your organization’s inference requests and are managed through the Platform API.

Read the current policy:

curl "https://api.mixlayer.com/v1/organizations/$ORG_ID/routing_policy" \
-H "Authorization: Bearer $MIXLAYER_API_KEY"

The response wraps the policy in data. A default policy looks like:

{
"data": {
"mode": "nearest",
"fallback": "all",
"regions": []
}
}
FieldPurposeDefault
modenearest uses latency-based routing. ordered uses the priority order in regions. required requires a Mxl-Region header on each request.nearest
fallbackall allows fallback to other eligible regions. same_country limits fallback to the first selected inference region’s country. disabled tries only the first region.all
regionsLimits inference to the listed region IDs. List order sets priority in ordered mode; ordered requires a non-empty list. An empty list or null uses every visible region.[]

Read operations require an API key with api-read. Replacing the policy requires an organization owner; API keys require api-admin.

PUT replaces the whole routing policy. Omitted fields reset to their defaults. Send all three fields when you want to preserve explicit settings. A successful update returns 204 No Content.

Use latency-based routing

Use nearest to let Mixlayer choose the inference region automatically. This example explicitly restores the default routing and fallback behavior:

curl -X PUT "https://api.mixlayer.com/v1/organizations/$ORG_ID/routing_policy" \
-H "Authorization: Bearer $MIXLAYER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "nearest",
"fallback": "all",
"regions": []
}'

To limit latency-based routing to a subset of regions, supply that subset in regions. The list order does not affect nearest routing.

Prefer regions in a fixed order

Use ordered when you want a primary inference region followed by a backup, regardless of latency. The following example prefers us-central, then us-east:

curl -X PUT "https://api.mixlayer.com/v1/organizations/$ORG_ID/routing_policy" \
-H "Authorization: Bearer $MIXLAYER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "ordered",
"fallback": "all",
"regions": ["us-central", "us-east"]
}'

The region IDs in this and the following examples are illustrative. Replace them with regions available to your organization and the model you intend to use.

If both regions host the model, Mixlayer tries us-central first and can fall back to us-east when the selected model is not available in us-central. If the model is deployed only in us-east, that is the first eligible region. Regions outside the list are never used for inference.

Keep fallback within the same country

Use same_country to allow multi-region inference while keeping fallback in the country of the first selected inference region:

curl -X PUT "https://api.mixlayer.com/v1/organizations/$ORG_ID/routing_policy" \
-H "Authorization: Bearer $MIXLAYER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "nearest",
"fallback": "same_country",
"regions": []
}'

For example, if the first inference region is us-central, another eligible US region such as us-east can be a fallback. A region in another country is excluded even if it is otherwise available.

same_country follows the first selected inference region’s country, which can differ from the gateway region or your application’s location. To constrain inference to a particular country, also set regions to an explicit list of regions in that country. You can combine this with ordered:

curl -X PUT "https://api.mixlayer.com/v1/organizations/$ORG_ID/routing_policy" \
-H "Authorization: Bearer $MIXLAYER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "ordered",
"fallback": "same_country",
"regions": ["us-central", "us-east"]
}'

Run inference in one region

To keep inference in one region, restrict regions to that region and disable fallback:

curl -X PUT "https://api.mixlayer.com/v1/organizations/$ORG_ID/routing_policy" \
-H "Authorization: Bearer $MIXLAYER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "ordered",
"fallback": "disabled",
"regions": ["us-central"]
}'

Requests fail if the model cannot be served in that region. They are not sent to another region. You can pair this policy with us-central.mixlayer.ai to select both the gateway region and the inference region.

Prefer an inference region per request

Send the Mxl-Region request header to try a specific inference region first. The region must be allowed by your organization’s policy and host the requested model. The header takes priority over both nearest and ordered selection; the fallback policy still applies.

curl -i https://mixlayer.ai/v1/responses \
-H "Authorization: Bearer $MIXLAYER_API_KEY" \
-H "Content-Type: application/json" \
-H "Mxl-Region: us-central" \
-d '{
"model": "qwen/qwen3.5-4b-free",
"input": "Tell me a fun fact about chihuahuas."
}'

The Mxl-Inference-Region response header identifies the region that served HTTP inference, including streaming responses and embeddings. The example uses curl -i to display response headers.

Use mode: "required" if every request must specify Mxl-Region. Combine it with fallback: "disabled" to require execution in the requested region. Responses WebSocket clients send Mxl-Region on the upgrade request; the serving region is not reported on that surface.

When fallback happens

Fallback happens when the selected model is not available in a region at the start of inference. It does not restart an in-progress generation or retry a timeout or another inference error. Requests remain subject to platform attempt limits, so a long region list does not guarantee that every region will be tried.

If no eligible region can serve the model, the request fails instead of widening your routing policy. If the selected model is not available in any attempted region, the request returns 503 with the code no_model_server_available.