FOUNDER & PARTNER PERK

Your agents pay for every token. Stop that.

Headroom compresses the context your coding agents send to the model: tool outputs, logs, files, search results. Same answers, a fraction of the tokens. Pick a perk and we'll set you up this week.

Claude CodeCodexGitHub CopilotVS CodeGemini+ any OpenAI / Anthropic API

With a 12-month Headroom Enterprise contract over $10,000. Companies of any size. One perk per company.

0T+
tokens saved per month on average, across Headroom users
0%
fewer prompt tokens in Google Cloud's agent benchmark
Gemini · 4 workloads
0%
task accuracy kept in that same benchmark
100% → 100%
0K
GitHub stars on the open-source core
Apache 2.0
Talk to us

Claim your perk

Tell us a little about your team and which agents you run. A founder will reply within a day to get you set up.

Perks apply to a 12-month Headroom Enterprise contract with a total value over $10,000. Open to companies of any size, one perk per company. We confirm eligibility and final details when you claim.

We only use this to set up your perk. Prefer email? hello@headroomlabs.ai

How it saves

Tool output rides along on every turn. We shrink it.

An agent's next request carries everything it has already read. Headroom compresses that context locally, before it reaches the model, and keeps every original one call away.

01

Compress. JSON, logs, code and prose each get the right compressor. Errors, anomalies and boundaries are kept.

02

Cache. The original stays on your machine with a retrieval reference.

03

Retrieve. If the model needs more, it calls headroom_retrieve and gets the exact source back.

claude · agent session
› why are payments failing?
● search_logs(service="payments", since="1h")
⎿ kept: FATAL connection pool exhausted · 218 affected
  full source → headroom_retrieve("ce1a52da…")
TOKENS
6,725−90%
Where Headroom fits

One local layer between your agents and the model

No new workflow. Point your agent's base URL at Headroom, or run headroom wrap claude. Compression runs on your machine.

Coding agents send requests through Headroom, which compresses context locally and forwards smaller requests to model providers. YOUR AGENTS RUNS LOCALLY MODEL PROVIDERS ▸ detect content type ▸ compress what's safe ▸ keep originals (CCR) 0 tokens saved
Claude Code · Codex · Copilot
↓ full context
Headroom
detect · compress · keep originals
↓ fewer tokens
Anthropic · OpenAI · Google · Bedrock
Proxy base_urlCLI headroom wrapSDK compress()MCP tools
Open source

Start free. The core is Apache 2.0.

Compression, the proxy, the CLI wrappers and MCP tools are all open source. Install it in a minute and watch your token bill drop.

74K stars
Python + TypeScript
Apache 2.0
headroomlabs-ai/headroom
pip install "headroom-ai[all]"
headroom wrap claude # Claude Code
headroom wrap codex # OpenAI Codex
headroom wrap copilot # GitHub Copilot
headroom proxy --port 8787 # anything else
Open source vs Enterprise

Your perk unlocks Enterprise

Open source gets you real savings. Enterprise is built for teams running agents at scale, where every percent of spend and every risky tool call matters.

CapabilityOpen sourceEnterprise
Cost savingsToken reduction on every requestCore compressionMore savings: extra compressors, tool search and harness tuning
Model routingRight model for each requestNot includedRoute requests to cheaper models when quality allows
Security & attack removalPrompt injection, risky tool callsNot includedDetect and strip attacks before they reach the model
IntegrationsWhere Headroom plugs inProxy, CLI wrappers, SDK, MCPDeeper: gateways, IDE fleets, SSO, admin policy, usage reporting
SupportWhen something needs a humanGitHub issues and DiscordComprehensive: dedicated engineers, onboarding, priority response