
Control AI access
Reduce shadow IT by giving your business one governed place to approve apps, providers, keys, and model access.
Msty Nexus - Inference gateway
Put one governed inference layer behind your AI tools for model routing, local runtimes, provider keys, usage visibility, and guardrails.
Product info
Msty Nexus sits between your applications and model endpoints. It standardizes the plumbing for local runtimes and online providers, protects credentials, applies routing and guardrails, and gives visibility into inference traffic.

Reduce shadow IT by giving your business one governed place to approve apps, providers, keys, and model access.

Keep sensitive data private and secure on local infrastructure, while still using online models when the job needs them.

Apply extra protections like scoped tokens, PII redaction, key blocking, and routing rules before requests reach models.
How it works
Nexus gives teams and power users a single place to connect AI tools without scattering provider keys and runtime settings across every application.
Bring local runtimes and hosted providers into one Nexus-managed catalog.
Standardize model settings so applications call stable names instead of fragile provider details.
Give each approved tool its own local token and revoke access without rotating every key.
Use compatible endpoints while Nexus tracks usage and keeps provider credentials protected.

Power features
Nexus is more than a gateway. It can choose the right model path, distribute traffic, and coordinate inference capacity so every request does not land on the most expensive model.

Send each request to the right model for the job. Keep expensive models for work that needs them, and route routine tasks to cheaper or local models to reduce token costs.

Distribute requests across available runtimes and machines so teams can keep throughput steady without pointing every app at one overloaded endpoint.

Coordinate multiple machines as shared inference capacity for larger local workloads, central servers, or teams that need more horsepower behind the gateway.
Features
Use Nexus as the operating layer for local runtimes, hosted providers, governed app access, and practical usage visibility.
Give applications one governed catalog across local runtimes and online providers.
Keep provider credentials in one controlled place instead of copying them across tools.
Connect approved applications through familiar OpenAI- and Anthropic-compatible endpoints.
Manage supported Ollama, llama.cpp, and MLX runtimes alongside hosted model access.
Issue scoped tokens per application instead of sharing one master key across every integration.
See requests, latency, and usage by application and provider from one place.
Review and approve new model or provider connections before applications can use them.
Keep the gateway off the LAN by default and open access only when you choose to.
Pricing
Nexus is available as a free download. Pro and Enterprise add managed scale, protection controls, logging, and team governance.
Download Nexus and run the local gateway on your machine.
Managed Nexus control for advanced access, protection, and observability.
Enterprise controls for larger teams, stricter environments, and managed deployments.
| Feature / Attribute | Free | Pro | Enterprise |
|---|---|---|---|
| Usage & Billing | |||
| Gateway Requests | Unlimited | Unlimited | Unlimited |
| Managed machines | 1 | 3 | Starting at 125 included |
| Administration | |||
| Desktop Console | |||
| Cloud console | - | ||
| Teams / RBAC | - | - | |
| Team / User rate limit | - | - | |
| Single Sign-On (SSO) | - | - | |
| Routing & Safety | |||
| Smart Routes | Limited | ||
| Load Balancer | - | ||
| Cluster | - | - | |
| Fallbacks | - | ||
| Logging & Support | |||
| Basic logging | |||
| Team/User logging | - | - | |
| Priority Support | - | - | |
Pricing FAQs
The number of machines or inference runtimes registered with Nexus to provide model inference, such as employee Macs, Mac Studios, or centralized servers.
Client devices do not automatically count as managed machines. A laptop simply consuming inference from a central server is a client; it only counts as a managed machine if it also provides inference through Nexus.
Humans and AI workloads are treated equally. A client can be a person, AI agent, application, workflow, or service.
Published plan capacity is based on the number of managed machines registered with Nexus.
No. Nexus does not charge based on the number of tokens processed and has no gateway request limits, so heavy AI usage does not create per-request fees or overages.
Nexus is designed for local and centralized AI. A deployment can range from a few shared inference servers to hundreds or thousands of employee devices contributing local inference.
Yes. Capacity can be enforced locally, allowing Nexus to support private and air-gapped environments without requiring usage telemetry to Msty.
Download
Download the desktop gateway for the machine that runs Nexus and routes approved apps, models, credentials, and local runtimes.