Fewer tokens upstream
Duplicate and oversized payloads are trimmed before they reach the model — often saving more than half of requested tokens.
A local proxy between Cursor / Claude Code and the model APIs. It removes duplicate or oversized context before the request goes upstream — then shows you the savings live.
Denser KPIs — lifetime tokens and dollars saved, spend velocity, strip rate, noise traffic, and model mix.
Full desktop screens — dark and light — for token throughput, disposition, traffic mix, and recent requests.
Token Saver runs on your machine. Editor traffic hits the proxy first; bulky or repeated context is stripped, the cleaned request is forwarded, and a dashboard tracks tokens saved, spend velocity, and model mix.
Duplicate and oversized payloads are trimmed before they reach the model — often saving more than half of requested tokens.
Watch requested vs saved vs forwarded tokens, save rate, traffic mix, and recent requests in dark or light UI.
A denser v2 board for lifetime totals, $/hr velocity, strip rate, noise traffic, and model mix at a glance.
Modes match how you already talk to models — MITM for default Cursor, reverse / OpenAI / Anthropic base URLs for BYOK and Claude Code.
mitm — intercept Cursor’s default HTTPS models (trust the local CA once).reverse / openai / anthropic — point Base URL at 127.0.0.1:8080, no CA.--dry-run to log savings without rewriting bodies.http://127.0.0.1:8081/ · ops board at /v2.AI Token Saver is a Universally Thinking project. Clone it, run it locally, or reach out for setup help.