Semantic context routing & memory for Claude Agents.
Stop choking Claude’s 200k & 1M token context windows with noisy dumps. ContextRail AI is pure developer infrastructure that dynamically prunes irrelevant context, aligns prompt caches, and injects persistent cross-session episodic memory in under 50ms.
A high-velocity context highway for autonomous agents
Standard RAG dumps brute-force entire document chunks into the context window, inflating latency and Claude API costs. ContextRail AI introduces a semantic routing layer that filters, summarizes, and routes high-salience tokens.
Sub-50ms RAG Compression
Prunes syntactic fluff, redundant markdown, and low-salience paragraphs while preserving crucial technical needles. Compresses token footprints by up to 78% with zero hallucination penalty.
Cross-Session Agent Memory
Equip Claude with persistent episodic memory. Our hybrid vector-and-knowledge-graph index remembers previous user decisions, codebase invariants, and tool results across sessions.
Prompt Cache Boundary Alignment
Automatically formats prompts so static system guidelines and immutable tool schemas sit cleanly behind Anthropic's 5-minute Prompt Caching breakpoints, saving up to 90% on input token costs.
Zero Foundation Model Training
Your enterprise codebase and agent interactions are never used for training. Context is processed in volatile memory partitions with AES-256 tenant isolation and GDPR sovereignty.
Native Model Context Protocol (MCP)
First-class compliance with Anthropic's open MCP standard. Drop our official server card into Claude Desktop or Claude Code in seconds to expose memory and routing tools immediately.
Token Savings & Latency Telemetry
Inspect per-turn token economy, cache hit ratios, and millisecond routing latency directly from our developer console or via automated webhook alerts.
Connect to Claude Code & Claude Desktop in 60 seconds
ContextRail AI provides a verified Model Context Protocol server exposing real-time context routing and episodic memory tools directly to your local or cloud Claude runtimes.
Expose ContextRail directly in Claude
By adding ContextRail AI to your claude_desktop_config.json or Claude Code agent configuration, Claude gains automatic access to semantic memory queries and dynamic context pruning tools.
route_context — Salience reranking & pruning
compress_rag_context — Sub-50ms token reducer
retrieve_agent_memory — Cross-session vector lookup
store_agent_memory — Persistent episodic entity graph
$ claude mcp add --transport http contextrail https://api.contextrail.cloud/mcp
$ curl https://api.contextrail.cloud/.well-known/mcp/server-card.json
{
"mcpServers": {
"contextrail": {
"command": "npx",
"args": ["-y", "@contextrail/mcp-server"],
"env": {
"CONTEXTRAIL_API_KEY": "cr_live_your_api_key_here",
"CONTEXTRAIL_ENDPOINT": "https://api.contextrail.cloud/mcp",
"TARGET_MODEL": "claude-sonnet"
}
}
}
}
Measured performance on Claude Sonnet
Evaluated on real-world multi-repo refactoring and 100k+ token financial multi-document Q&A suites. ContextRail AI outperforms vanilla RAG and raw ingestion across all core dimensions.
| Routing Strategy | Context Reduction | Needle Recall | TTFT Latency | Cost / 10k Agent Turns |
|---|---|---|---|---|
|
ContextRail AI (MCP Engine) Semantic Reroute + Cache Breakpoint |
78.2% compressed | 99.4% accurate | 142 ms | $4.20 USD |
|
Standard Vector RAG (LangChain / LlamaIndex) Top-K chunk concatenation |
32.0% compressed | 76.5% accurate | 680 ms | $18.50 USD |
|
Raw Uncompressed Ingestion Dumping entire repo/docs into 200k window |
0% (Full dump) | 84.1% accurate | 1,850 ms | $38.90 USD |
Public HTTP API for programmatic context routing
Every MCP tool is also a documented REST endpoint, so you can call ContextRail AI from any runtime, CI job or backend service. TLS 1.3, bearer-key auth, RFC 9457 error bodies.
| Method | Path | Description | Auth | Rate limit |
|---|---|---|---|---|
| POST | /api/v1/context/route |
Rank and prune a context payload before injection | Bearer key | 120 req/min |
| POST | /api/v1/context/compress |
Sub-50ms semantic RAG chunk compression | Bearer key | 120 req/min |
| POST | /api/v1/memory/query |
Cross-session episodic memory lookup | Bearer key | 60 req/min |
| POST | /api/v1/memory/write |
Persist agent decisions into the vector-graph index | Bearer key | 60 req/min |
| GET | /api/v1/metrics/efficiency |
Token savings, cache hit ratio, routing latency | Bearer key | 120 req/min |
| POST | /api/v1/webhooks |
Register an HTTPS sink for budget and latency alerts | Bearer key | 10 req/min |
Idempotency-Key support on all POST routes. MCP transport: http streamable at https://api.contextrail.cloud/mcp. Scale plan raises limits to 1,000 req/min and per-tenant concurrency of 64.
Predictable SaaS plans for builders and enterprises
No convoluted credit tokens. Transparent monthly billing in USD with instant API key issuance and zero commitment.
Sandbox
For individual developers prototyping agents with Claude Code or local desktop scripts.
- 10,000 routed tokens/month
- 1 concurrent MCP agent connection
- Standard semantic compression
- Community Discord & GitHub support
Builder
For production startups building SaaS products on top of Claude Sonnet & Haiku.
- 2,500,000 routed tokens/month
- Up to 10 active MCP agents
- Semantic vector agent memory
- Anthropic Prompt Cache alignment
- Standard email support (sub-24h)
Scale
For heavy multi-agent fleets, high-frequency autonomous workflows, and enterprise scale.
- 15,000,000 routed tokens/month
- Unlimited concurrent MCP agent fleets
- Hybrid vector + knowledge graph memory
- Dedicated priority routing & webhooks
- 99.9% Uptime SLA & Slack engineering channel
Technical specifications & FAQs
Everything you need to know about ContextRail AI integration, security guarantees, and compatibility with Claude models.
@contextrail/mcp-server) and exposes a verified server card at /.well-known/mcp/server-card.json. You simply paste the snippet into your Claude configuration file and your assistant immediately acquires our routing and memory tools.
Building high-throughput infrastructure for the age of autonomous agents
ContextRail AI Software S.L. is a pure B2B software product engineered and operated by our systems team at Paseo de la Castellana 95, 28046 Madrid, Spain. We believe the biggest bottleneck facing autonomous agents is not model intelligence, but context bandwidth and retrieval bloat.
By treating context as a high-speed routing fabric rather than a static text bucket, we empower engineering teams to run complex multi-agent architectures on Claude with predictable costs, instant retrieval, and cross-session persistence.