Relay Proxy :8787
Atlas active
claude · Max
62%
Nimbus
claude · Pro
91%
Orbit
codex · Team
28%
Compression saved 1.24M tokens today
Open Relay ⌘0
Connect Claude Code
Settings… ⌘,
Quit Relay ⌘Q
Relay

Overview

everything at a glance

Everything's on the wing

Four accounts enrolled, the proxy's up, and nothing live is near its limit. Vega needs a fresh sign-in; the other three are fine.

4
Accounts enrolled
2 healthy · 1 near cap · 1 needs sign-in
0
Tokens compressed today
lossless · ~34% of input
$0
Spend this month
≥ partial; 2 records unpriced
0
Account near cap
Nimbus · 91% of 5-hour

Accounts

Connections

Claude Code
ANTHROPIC_BASE_URL → :8787
Connected
Codex
Responses → :8787
The proxy port is sticky but it self-heals. If it moves, Relay rewrites each agent to the port it actually bound.

Live now

Flow

How each bound project's requests travel through compression to the account it's routing to right now. Faint lines are the map; a line lights up and carries packets only while a folder is actively sending.

Healthy Near reserve Capped
34%
Compression
lossless · 1.24M saved
Live. Accounts, weekly usage and compression totals are read straight from this Relay. A project moves to the next account in its chain once the one it's on crosses its reserve.

Accounts

Every enrolled login, and how much of its window is left. Relay moves new work down a project's chain before any of these hits its reserve, so a session never walks into a wall.

Enrolled · live usage

polling every 60s

Credits & gateways

OpenRouter
gateway · pay-as-you-go
$12.40
Vercel AI Gateway
gateway · key required
$8.10
Cursor
API credit + Auto
$40.00
Auto 72%

Bindings

A binding points a project folder at an ordered chain of accounts and pools. Work starts at the top and only drops to the next when an account crosses its reserve. Conversations already running stay on their warm account, so you keep the prompt cache.

Default

Any folder without an explicit binding
all accounts, by weekly headroom

Account pools

Personal
Atlas · Nimbus; reclaims earlier members when they recover

Project bindings

Spend

Estimated cost of routed and gateway traffic this month, split by model and project. Subscription usage is metered but not billed here; these figures are the paid slices only.

This month
Last 30 days
All time
This month
$0
≥ 2 records unpriced
OpenRouter
$0
gateway spend
Vercel AI Gateway
$0
gateway spend
Compression avoided an estimated $3.90 of gateway spend this month by never sending trimmed context.

By model · this month

By project · this month

Activity

A restart-safe, content-free record of what Relay did: account switches, limits reached, compression, and proxy recovery. No prompts or responses are ever stored here.

Recent events

Today
Atlas → Nimbus, switched ahead of the 5-hour reserve (predictive · re-warms ~4k)
09:38
Compressed 312K tokens of tool output, lossless
09:31
Nimbus reached its 5-hour limit; new work held for eligible accounts
08:57
Yesterday
Proxy recovered; re-bound :8787 and rewrote connected agents
18:12
Live sessions
~/Dev/perch on Atlas · opus-4.8 working
saved 18.6K tokens this session

Intelligence

Two optional layers, both off by default and fully auditable. Routing sends each task to the model that suits it; compression trims what gets sent without changing meaning. The essentials are here; open a section for the full rulebook.

Task routing
Relay classifies each prompt and sends it to the model in your rulebook. Off means every request runs on its own model, unchanged.
On
Router source
The small, cheap model that reads each prompt to classify it.
Active profile
A preset that sets the whole rulebook at once. Cache-first keeps conversations on their warm account.
Cache-first
Rules for
Apply globally, or override for a single project.
Ask before routing
Confirm each switch before it happens, rather than routing silently.

Task → model

Task typeRoutes toAlways
Design / UIclaude · sonnet-5 Ask
Refactorclaude · sonnet-5 Ask
Bulk editclaude · haiku-4.5 On
Debugclaude · opus-4.8 On
Docsclaude · haiku-4.5 Ask
Planningclaude · opus-4.8 Ask
Testsclaude · haiku-4.5 Ask
Data / JSONclaude · sonnet-5 Ask

Model roles

Route by role
Plan mode
While plan mode is active
Implement
When you leave plan mode
Subagent
A subagent's first turn
Review
Unavailable — no automatic signal
Cache-first Active
Strong evidence · keep conversations on their warm account
Applied
Economy
Provisional · cheapest capable model per task
Balanced
Provisional · quality on hard tasks, thrift elsewhere
Quality
Provisional · best available model everywhere
Planner → Builder Experiment
Staged: plan on opus, build on sonnet
High-value deliberation Experiment
Panel + judge on the hardest calls
Lower-cost diversity Experiment
Spread work across cheaper providers
Experiments have measured lift not yet established. Each shows its cache-rewarm, return-trip and multi-model-spend caveats before it applies, and any preset change can be undone.
Ask several models the same question, then a judge synthesises one answer. Read-only: ordinary model calls, no tools or web, 120-second timeout.
Conclusion
Research answer
Implementation plan
Code review
Decision compare
Question, decision, or plan to deliberate on…
Panel 1
Panel 2
Judge
Output limit
Warn if >
est. $0.18

Remembered "always" rules

Debug → always claude · opus-4.8
Bulk edit → always claude · haiku-4.5

Recent decisions

TimePromptModelHow
09:38"tidy the flow legend spacing"sonnet-5Auto
09:12"why does the proxy port move"opus-4.8Confirmed
08:44account switch · re-warms ~4kAtlas→NimbusHanded off

Mid-task hand-off Advanced

GLM-4.6 "drifted into design work" claude · sonnet-5 next turn

Confirmation style

In-CLI shows a numbered prompt in your terminal ([1] yes [2] keep [3] pick… [a] always); Seamless uses a native macOS notification you can act on without leaving the editor.
Best execution
Treat your chosen route as a quality grade, then let Relay find the cheapest way to keep it. Advisory for now: it explains what it would choose and why, and never acts on a warm conversation.
Advisory
Reference grade
The quality bar every cheaper candidate must clear. Price can never buy its way past this.
Quality floor
Of the tasks the reference solves, the candidate's conservative lower bound must retain at least this share.

Latest decision · explained from the record

Stayed on sol · high. The cheaper candidate met the benchmark floor but didn't clear the projected 92k-token cache re-warm plus uncertainty margin. Evidence: 148 matched local episodes + benchmark prior · policy best-exec-17 · updated 3d ago.
Counterfactual: would switch after ~3 more warm turns at current prices. Reconsidered at the next compaction or plan-to-build boundary.
Was this right?
Feedback joins the decision by an opaque id — a bounded metric, never your prompt. It tunes which candidates clear the floor for your workload, the same loop Ship's feedback API closes for theirs.
Compression
Trim historical tool output before it's sent. Lossless is byte-reversible; higher modes normalise or drop low-signal lines.
0tokens saved this session
Roughly 34% of what you'd have sent · 41 blocks · estimated (Anthropic's tokenizer isn't public).
Command logs
JSON payloads
Search output
Diffs

Protected, safe by default

Read / Edit / Write & file reads
Native, shell and MCP reads stay exact; edits need the original bytes.
Error output
Tracebacks kept verbatim so the model can recover.
System & tool definitions
Always exact.
Cached prefix
The already-cached prompt prefix is never rewritten.

Prompt-cache impact · 24h

86%
cache hits
≈4%
new-account miss
2.1M
Claude input
Cache hit is cache-read input over all Claude input. New-account miss attributes the prior-context portion re-written on the first completed turn after a switch; ≈ marks the unobservable same-account counterfactual.
1.8M
treatment cache reads
312
warm hits
6
profile changes held
Control vs treatment cohortlast 7 days
Control · 5% Off holdout
142 turns · avg 4.1s · repeated reads 18
Treatment · compression on
2.6K turns · avg 3.9s · repeated reads 21
Estimated $3.90 gateway spend avoided · API-equivalent priced-turn cost held flat.
Pinned Off holdout
Keep a slice of turns uncompressed to measure real lift.
Both layers fail open. If anything's uncertain, Relay sends the original, uncompressed, on the client's own model. Nothing here can make a request fail.

Settings

Connections, the proxy, polling, alerts, and appearance. The same controls back the ⌘, preferences window.

Agents & proxy

Claude Code
ANTHROPIC_BASE_URL → 127.0.0.1:8787
Connected
Codex
Responses → 127.0.0.1:8787 · ~/.codex/config.toml

Providers & gateways

OpenRouter
318 models · updated 6m ago
Connected
Vercel AI Gateway
API key
Cursor
API credit + Auto · login
Connected

Proxy

Port
8787
Auto-pick a free port
Falls forward if the port is taken.
Always on

Polling

Usage interval
60s
Balance interval
60s

Usage alerts

Notify when 5-hour usage crosses
85%
Notify when weekly usage crosses
90%
Alert on low gateway balance
Below this remaining credit.
$5

General

Appearance
Launch Relay at login
Menu-bar cycle interval
How fast the tray meter rotates between active accounts.
6s

About

Relay
3.0.0 (mock)
Proxy
127.0.0.1:8787running
Notifications
authorised