Running Claude Code with open models used to mean hand-editing environment variables, standing up a LiteLLM proxy and hoping tool calls survived the translation. Together AI’s new Together Link CLI, released in beta on October 5, 2026, turns that into one install command and a tclaude alias. The pitch is that you keep Claude Code, Codex or OpenCode exactly as they are and cut your model bill by more than half.
I went through the launch post, the docs and Together’s pricing page, and then did the per-token arithmetic myself. The 50% claim holds for some sessions and falls apart for others, and the difference comes down to prompt caching.
What Is Together Link?
Together Link is a launcher that starts your existing coding agent with a temporary configuration pointing it at models hosted by Together AI. It does not replace Claude Code or Codex. It wraps them.
According to Together AI’s documentation, it supports:
- Claude Code (
togetherlink claude, shortcuttclaude) - Codex CLI (
togetherlink codex, shortcuttcodex) - OpenCode 2 (
togetherlink opencode) - Pi Code 0.80.8 or newer (
togetherlink pi) - Claude Desktop and Cowork (
togetherlink claude-desktop, beta) - ChatGPT Desktop (
togetherlink chatgpt, beta)
For the terminal agents, the docs say the configuration is created per launch and removed after the session. The docs state that “Together Link never modifies your normal configuration files”, so typing claude on its own still gives you Anthropic’s models with your normal login.
That design is the part I like most. Every previous way of doing this left a stray ANTHROPIC_BASE_URL in someone’s shell profile, and then three weeks later they wondered why Claude Code felt different.
Together Link is a wrapper, not a fork, so the agent you already know stays exactly the same and only the model behind it changes.
How to Set Up Claude Code With Open Models
The install is one line, per Together AI’s launch post:
curl -fsSL https://link.together.ai/install | bash
The docs say this installs the commands into ~/.local/bin and installs Bun if you do not already have it. You need macOS or Linux, Bash, curl, a Together API key, and the agent you plan to use already installed.
Before you run it, read the script. Piping a remote installer into Bash on a machine that holds your SSH keys deserves thirty seconds of attention, especially after the GitSpawn and Plugin4Shell agent vulnerabilities covered here on October 4.
Then set your key and launch:
togetherlink configure # or: export TOGETHER_API_KEY="..."
tclaude # Claude Code on the auto router
togetherlink --main zai-org/GLM-5.3 claude # pin one model
togetherlink usage --last 7d # spend across recent sessions
Two details from the docs that will save you some confusion. The --main flag must come before the tool name. Project .env files are not read for the API key, so a TOGETHER_API_KEY sitting in your repo’s .env will be ignored.
For headless runs in CI or scripts, the docs tell you to always append < /dev/null to togetherlink claude -p calls. A headless Claude Code process that inherits an open stdin blocks forever waiting for input, which is the kind of bug that eats an afternoon.
Setup takes a few minutes, and backing out means typing claude instead of tclaude.
How the Together Link Auto Router Works
By default, sessions use a model called auto. According to Together AI’s launch post, it reads the first task in a session and routes the whole session to one model:
- With an Anthropic API key: routes between Claude Opus 5.5 for hard problems and GLM 5.3 for the rest
- Without one: routes between GLM 5.3 and GLM 5.3 Flash
The docs add an important limit. The Opus route only applies to Claude sessions. Codex, OpenCode, Pi Code and ChatGPT Desktop sessions always stay on Together-hosted models.
Routing happens once per session, not per request. Together says this is deliberate, because switching models mid-session would throw away the prompt cache. I think that is the right call, and it is the same logic that makes long context window optimisation pay off: a warm cache is worth more than a slightly better model on any single turn.
The tradeoff is that the router judges a session by its opening message. If you start with “fix this typo” and the session turns into a three-hour refactor, you are on the cheap model for all of it. Start a new session when the task changes shape.
The auto router is a session-level bet based on your first prompt, so write that first prompt like it matters.
Together Link Pricing: Open Models vs Claude Opus 5.5
Here are Together AI’s serverless rates for the models in Together Link, next to Claude Opus 5.5’s API pricing.
| Model | Input /1M | Cached input /1M | Output /1M |
|---|---|---|---|
| Claude Opus 5.5 | $4.00 | $0.20 | $20.00 |
| Kimi K3 | $2.70 | $0.27 | $13.50 |
| GLM 5.3 | $1.40 | $0.26 | $4.40 |
| DeepSeek V4.1 Flash | $0.30 | $0.01 | $1.20 |
| GLM 5.3 Flash | $0.15 | $0.03 | $0.50 |
(Source: Together AI pricing page and Together Link docs, October 2026; Anthropic Claude Opus 5.5 pricing. Together’s pricing page lists DeepSeek V4.1 Flash cached input at $0.006.)
Look at the cached input column. Claude Opus 5.5’s cache reads at $0.20 per million are cheaper than both GLM 5.3 and Kimi K3 on Together. Anthropic cut the Opus cache-read multiplier with this generation, and that changes the maths for agent workloads, which are mostly cache reads.
To see what that does to the 50% claim, I priced one illustrative coding-agent session: 3 million input tokens, 90% of them cache hits, plus 100,000 output tokens. This is arithmetic on list prices, not a measured session, and it ignores cache-write charges.
| Model | Fresh input | Cached input | Output | Session total |
|---|---|---|---|---|
| Claude Opus 5.5 | $1.20 | $0.54 | $2.00 | $3.74 |
| Kimi K3 | $0.81 | $0.73 | $1.35 | $2.89 |
| GLM 5.3 | $0.42 | $0.70 | $0.44 | $1.56 |
| DeepSeek V4.1 Flash | $0.09 | $0.03 | $0.12 | $0.24 |
(Source: author’s calculation from the list prices above.)
GLM 5.3 comes out 58% cheaper than Opus in this example, so Together’s “over 50%” holds. Kimi K3 saves only 23%, because its output price is two thirds of Opus’s and its cache reads cost more. DeepSeek V4.1 Flash is about 94% cheaper.
Output-heavy sessions favour the open models more. Long, read-heavy sessions where the cache does most of the work favour them less.
The 50% saving is real for GLM 5.3 and the Flash models, but Kimi K3 on a cache-heavy session saves closer to a quarter.
Which Open Model Should You Pick in Claude Code?
Price is half the decision. The other half is whether the model can finish the task, because a cheap model that needs three attempts costs more than an expensive one that needs one.
The most recent public comparison of these exact models is in Reflection AI’s launch post for its new Beam model, published October 5, 2026. I covered that table in detail in the Reflection Beam benchmark breakdown, and the agentic coding rows are useful here:
| Model | DeepSWE v1.1 | Terminal-Bench v2.1 |
|---|---|---|
| DeepSeek V4.1 Flash | 74.2 | 90.6 |
| Kimi K3 | 68.0 | 88.3 |
| GLM 5.3 | 61.0 | 88.2 |
(Source: Reflection AI launch post, October 5, 2026. Scores are vendor-reported by Reflection, not by the model makers.)
For reference, Google’s Gemini 4 Argon launch table put Claude Opus 5.5 at 74.2% on DeepSWE v1.1. Different labs run these benchmarks in different harnesses, so two identical 74.2 scores from two different tables do not prove the models are equal. Still, the numbers suggest DeepSeek V4.1 Flash deserves more attention than its “Flash” label implies.
Here is how I would set it up:
- Default:
autowith an Anthropic key, so hard sessions still reach Opus 5.5 - Bulk, well-specified work like test generation, lint fixes and migrations with a clear pattern: pin
deepseek-ai/DeepSeek-V4.1-Flash - Long agentic sessions on an unfamiliar codebase: stay on Claude, or pin Kimi K3 if budget forces it
Without an Anthropic key, the auto router chooses only between GLM 5.3 and GLM 5.3 Flash, and never picks DeepSeek V4.1 Flash on its own. If the benchmark table above holds up in your own tests, pinning it by hand may beat auto.
Pin DeepSeek V4.1 Flash for bulk work, keep Opus in the loop for hard sessions, and measure your own success rate before trusting any of these tables.
Together Link Tradeoffs Nobody Mentions in the Launch Post
Four things matter before you roll this out to a team.
Your code goes to Together AI. That is how a hosted model works, but it means a new data processor for your repository. If your company approved Anthropic and nobody else, check before you run tclaude on the work laptop.
It is beta. The docs say commands, routing behaviour and the model list may change. Do not build CI pipelines that depend on the auto routing decisions staying the same.
Your Claude subscription does not apply. Together Link bills your Together API key, and the Opus route uses an Anthropic API key. If you are on a Claude Max plan, you are already paying a flat rate for Claude Code, and per-token open models may not save you anything. API users have cheaper levers to pull first, which the guide to reducing Claude API costs walks through.
Claude Code is tuned for Claude. Its system prompts, tool definitions and agent loop were built around Anthropic’s models. Open models can follow them, but expect more tool-call mistakes than on Opus or Sonnet. The same applies to Codex, which is tuned for OpenAI’s models, as covered in the Codex CLI vs Claude Code comparison.
Together Link makes sense for teams paying API rates at scale; it makes much less sense for an individual on a flat Claude subscription.
Is Together Link Worth Using?
For an engineering org spending thousands a month on Claude API tokens through Claude Code, yes, at least as an experiment. It takes a few minutes to set up, it leaves your existing configuration untouched, and togetherlink usage gives you the cost comparison per session without a spreadsheet.
The smartest rollout is narrow. Point it at the bulk, well-specified work first and keep Opus on the hard sessions. Measure task completion, not just the bill.
For a solo developer on a Claude Pro or Max plan, I would skip it for now. The open models are cheaper per token, but you are not paying per token, and the quality gap on long agentic sessions is still visible in the published numbers.
Together Link is the cleanest way to run Claude Code with open models in October 2026, and the savings are real as long as you check which model you are actually paying for.