Read more
AgentNuro · Control tier

AgentNuro Control: real-time AI cost and policy enforcement through the Gateway

AgentNuro Control adds a lightweight inline gateway between your applications and AI providers, so the recommendations from Insight execute automatically. Routing, caching, budget caps, rate limits, and a kill switch, enforced per request.

Request trace · live Enforced inline
Routegpt-4o-mini
Cachehit, skipped
Enforcewithin budget
Rate limitok
Telemetrylogged
2.8ms
added at p99
sample data, for illustration only
Cache hit rate
41%
Requests blocked by budget
3

What is an AI gateway?

An AI gateway is a proxy that sits between your applications and the AI providers they call, so every request passes through one place before it reaches a model. AgentNuro Control is AgentNuro’s gateway. It is where the routing, caching, budget, and rate limit decisions that Insight recommends get enforced, per request, without changes to your application logic.

Without a gateway
App AApp BApp C
OpenAIOpenAIClaudeClaudeGeminiGemini

Each app calls each provider on its own. No shared budget, no shared cache, no single kill switch.

With AgentNuro Control
App AApp BApp C
Gateway
OpenAIOpenAIClaudeClaudeGeminiGemini

One place enforces policy, budget, and caching for every app and every provider.

What does AgentNuro Control do?

Control sits inline between your apps and providers, so every capability below runs on every request, in real time, not on a schedule.

Model re-routing Route each request to the model that fits, automatically based on cost and quality, or by rules your team sets for specific workloads.
Semantic + exact caching Serve repeated or near-identical requests from cache instead of calling the provider again, whenever a cache hit is safe.
Tagging for governance Every request is tagged with the team, project, model, and agent behind it, so cost and policy decisions have context, not just a token count.
Policy enforcement Usage policy, which models a team can use, which data can leave the network, gets enforced on every request, not just reported after the fact.
Budget tracking and enforcement Track spend against budget in real time, per team, project, or key, and stop a request before it goes over instead of finding out at month end.
Auto and manual kill switch Cut off a provider, model, team, or key instantly, by hand when something looks wrong, or automatically when a rule you set is tripped.
Anomaly detection Unusual spend or usage, a spike, a new model showing up, a key behaving differently, gets flagged as it happens, not in next month’s bill.
Cost optimization Ongoing recommendations to right-size models, trim context, and route around waste, the same recommendations Insight surfaces, applied automatically.

Shown here are the core enforcement capabilities. Additional policies and controls ship regularly.

Outcomes
Cost Savings
Latency Held (p99)
Quality Maintained
High Cache Hit Rate

Does the Gateway add latency?

Yes, a small and fixed amount. AgentNuro Control adds processing time for routing, cache lookup, and policy checks before a request reaches the provider, typically a few milliseconds at p99. Cache hits skip the provider call entirely, which usually more than offsets the added overhead in total response time.

Added latency (p99)
<5ms
sample data, for illustration only
Cache hit rate
41%
sample data, for illustration only
Requests failed open, last 90 days
0
sample data, for illustration only

If the Gateway is unavailable

Control fails open. Requests fall back to calling the provider directly using the last known routing rule, so a Gateway outage pauses enforcement, not your application.

Full detail in the FAQ below.
App OpenAIProvider

How is Control different from Insight?

Insight is read-only. It shows you where AI spend is going. Control is inline. It acts on what Insight finds, automatically, on every request.

Control fits when…

  • The same waste (wrong model, no caching) shows up every review cycle
  • You need a budget cap or kill switch enforced per request, not just reported
  • Manually routing or rate limiting has become a full-time job
  • You’re ready to act on Insight’s data automatically, not just see it

Insight is enough if…

  • You’re still finding unattributed or unsanctioned AI spend
  • Teams are reviewing recommendations manually before acting
  • You need buy-in from finance or security before anything enforces automatically
  • You want zero latency and zero blast radius while you build trust in the data

Which providers and models does Control support?

AgentNuro Control routes across 83 model providers today, from the major labs to specialized inference platforms, covering chat, embeddings, image, and audio endpoints. Logos are on the way; names and capability tags below are accurate as of this list.

ClaudeAnthropicChat
AWS BedrockAWS - BedrockChat
Azure AI FoundryAzureChat
Azure AI FoundryAzure AIChat
Google VertexGoogle - Vertex AIChat
GeminiGoogle AI Studio - GeminiChat
OpenAIOpenAIChat
AIAI/ML APIChat
AIAI21Chat
AMAmazon NovaChat
ASAssemblyAIChat
AWAWS - PollyAudio
AWAWS - SagemakerChat
BABasetenChat
BLBlack Forest LabsImage
CECerebrasChat
CLCloudflare AI WorkersChat
COCognitionChat
COCohereChat
CRCrusoeChat
DADarkbloomChat
DADatabricksChat
DEDeepgramChat
DEDeepInfraChat
DEDeepseekChat
ELElevenLabsChat
FAFal AIChat
FEFeatherless AIChat
FIFireworks AIChat
FRFriendliAIChat
GIGigaChatChat
GMGMI CloudChat
GRGradientAIChat
GRGroq AIChat
HEHerokuChat
HUHuggingfaceChat
HYHyperbolicChat
IBIBM - Watsonx.aiChat
INInceptionChat
JIJina AIEmbeddings
LALambda AIChat
LELemonadeChat
LILibertAIChat
LLLlamaGateChat
MEMeta - Llama APIChat
MEMeta Model APIChat
MIMinimaxChat
MIMistral AI APIChat
MOMoonshotChat
MOMorphChat
NENebius AI StudioChat
NLNLP CloudChat
NONovita AIChat
NSNscaleChat
NVNvidia NIMChat
OCOCIChat
OLOllamaChat
OPOpenRouterChat
OVOVHCloud AI EndpointsChat
PEPerplexity AIChat
PIPinstripesChat
PUPublicAIChat
QWQwenCloudChat
RERecraftImage
REReplicateChat
RURunwayMLImage
SASambanovaChat
SASarvamChat
SCScalewayChat
SCSCX.aiChat
SNSnowflakeChat
SOSonioxAudio
STStability AIImage
TETencent TokenHubChat
TETensormeshChat
TOTogether AIChat
VV0Chat
VEVercel AI GatewayChat
VOVolcengineChat
VOVoyage AIEmbeddings
WAWandB InferenceChat
XAxAIChat
ZAZ.AI (Zhipu AI)Chat

Frequently asked questions

As a reverse proxy in front of your existing provider endpoints, or as a thin SDK wrapper around your existing client. Most teams route their first workload within a day, with the only application change being the base URL or client configuration.

Control fails open. If the Gateway can’t be reached, requests fall back to calling the provider directly using the last known routing rule, so an outage in the Gateway pauses enforcement, not your application.

Only a small one. Point your existing client at the Gateway’s endpoint, or wrap it with the SDK. Your prompts, models, and business logic stay the same.

Yes. AgentNuro is SOC 2 compliant, with independently audited controls covering access, data handling, and monitoring, the same controls that apply across Insight and Control.

See every AI dollar. Then control it automatically.

Route your first workload through Control Lite in minutes. No credit card required.