Skip to main content
Posts tagged:

Software Development

Claude Code at Home

If you are lucky enough to have Claude code at work you get really spoiled. (I am.) You are the the bleeding edge of every update and can get an army of agents to tackle any problem. Unlike the real human world where we work a little, talk a little, get a cup of coffee and get back to work, the meter of token costs just keeps ticking.

And it’s expensive enough just to use power tools, but nothing works out of the box. You’re paying for the loop—the false starts, the “wait, let me open one more file,” the rewrite, the test, the cleanup, the second rewrite because the first rewrite wasn’t quite right. That loop is where real software gets built, and it’s also where credits vanish. It’s not that the models are overpriced in some abstract sense; it’s that the workflow is inherently iterative, and the billing model punishes iteration.

And the default model is to use the most expensive model for everything. While that’s convenient, it’s not what you should be doing at home when you have to choose between groceries or tokens.

Part of that is because these are days for being prodigal and testing a lot of things. You should be able to run things ten times in a row without feeling guilty. I want to chase a bug through a codebase, change my mind twice, and not have my brain doing mental math about whether I’m burning $8 or $80 today. I want to move fast, and I want the tool to be there the same way git is there—always available, always on, no drama.

But delegation is token-hungry by nature. The moment you let an agent do discovery—read files, skim configs, understand how tests run—you’ve entered the expensive zone. Not because the model is doing something complicated, but because the prompt becomes huge: policies, tool schemas, repo context, system instructions, plus your actual request. That’s before the model even starts thinking.

How do you actually pull this off? You let the strongest model be the boss and you assign everything else based on comparative advantage. For me that means Claude Code sits at the top as the supervisor and planner—it reads the repo, understands the architecture, and decides what needs to happen. OpenAI Codex lives one level down as the editor, taking tightly scoped instructions and producing clean diffs without wandering. And at the bottom is a locally running 7B Qwen 2.5 model through Ollama—CPU-bound, cheap, and always on—handling the boring, mechanical work. Judgment stays expensive and rare; execution becomes cheap and repeatable.

When I actually tried to do this, I found the bottleneck isn’t “can I run a model locally?” The bottleneck is “can I build a workflow that keeps my good tools while moving the expensive part out of the hot path?” Because you don’t want to give up Claude Code’s planning just to save money. You don’t want to give up Codex’s tight coding ergonomics either. You want to keep the good UI and the good behaviors, and just stop bleeding credits on every little edit.

That’s the problem statement in human terms: I want to use Claude and Codex like power tools, not like a casino meter. I want to delegate without anxiety. I want the loop back—fast iteration, lots of attempts, aggressive refactors—without that little voice saying “maybe don’t run it again.”

If you can solve that, you’re not just saving money. You’re restoring a way of working.

The good news is that the world is solving this problem. People have stopped thinking in terms of “pick one model” and started thinking in terms of “build a routing layer.” There’s been a genuine Cambrian explosion of models—open ones, semi-open ones, hosted ones—and developers are reacting the same way we always do when the ecosystem explodes: we put a gateway in front of it, normalize the interface, and swap engines behind the curtain. That’s why you see so much energy around OpenAI-compatible endpoints and proxies: you can point tools at one API shape and then decide later whether the request goes to Anthropic, OpenAI, Groq, a local Ollama box, or some hosted open model. LiteLLM is a clean example of this “gateway” pattern—one interface, lots of upstreams, plus routing/fallback ideas.

On the “hosted but cheaper” side, people are shopping for inference the way you shop for cloud compute: who can run decent models fast on specialized hardware at a lower price. Groq is the clearest “we are selling speed per dollar” play—very high token throughput and pricing that makes it attractive as a worker when you don’t need the absolute smartest model. OpenRouter is the marketplace version of the same instinct: one account, many models, and a routing/management layer so you can pick the right engine for the job (or let the router do it). They even publish guidance specifically about integrating Claude Code through their layer, which tells you how mainstream this “stick a router in front” approach has become.

Also popular are NVIDIA’s hosted endpoints, where you grab an NVIDIA API key and call models like Moonshot’s Kimi K2.5 through NVIDIA’s OpenAI-style chat/ completions interface. People mention it because it feels like cheating in the best way—GPU-backed inference, big context, modern agentic model, and the integration story is “just point your client at this base URL and use this model id.” NVIDIA’s own docs and blog posts are leaning into that exact pitch.

Meanwhile the other half of the world is going the opposite direction: “stop paying per token, I’ll run it myself.” Ollama on a workstation, LM Studio on a laptop, local quantized coding models, and a bunch of people stitching that into their editors so the ‘boring’ work has near-zero marginal cost. What’s interesting is that these two camps—cheap hosted inference and local inference—are converging on the same mental model: keep your premium model for judgment, and feed the grind to something cheaper, whether that’s a local box or a commodity inference provider. The tooling ecosystem is starting to assume you’ll do that, which is why everything is racing toward OpenAI-compatible APIs and “provider” abstractions.

What I ended up building is a split-brain workflow that feels obvious in hindsight: Claude stays in the driver’s seat as the planner and “reader,” and Codex becomes a local editor that only does the mechanical part—turning a very specific instruction into a clean diff. Under the hood, Codex isn’t talking to OpenAI at all; it’s pointed at an Ollama server running on the same machine, and Ollama is serving a small, cheap Qwen 2.5 7B model. The key design choice is that the local model never does discovery. It doesn’t roam your repo, it doesn’t decide what matters, it doesn’t try to be clever. Claude reads the codebase and decides exactly what should change; then it hands Codex a bounded edit task with the relevant file contents included, and Codex outputs a patch and stops. That’s the whole trick: you keep “judgment” expensive and rare, and you make “execution” cheap and repeatable.

To do this, install Ollama, then pull a model that’s small enough to be comfortable on CPU—Qwen 2.5 7B is a good starting point. The gotcha is that agentic tools like Codex have a surprisingly fat system prompt and tool schema, so a default 4k context model can choke in weird ways; the clean fix is to create a Codex-friendly variant with a larger context window. In Ollama that’s just a Modelfile and a new model name:

ollama pull qwen2.5:7b-instruct

cat > /tmp/Modelfile <<'EOF'
FROM qwen2.5:7b-instruct
PARAMETER num_ctx 12288
EOF

ollama create qwen2.5:7b-instruct-codex -f /tmp/Modelfile

Then you point Codex at Ollama using an OpenAI-compatible base URL, and you make sure Codex isn’t silently preferring cloud auth. Logging out of Codex cloud is the simplest way to prevent accidental spend, and then you set the config to use your local model by default:

codex logout 2>/dev/null || true

# ~/.codex/config.toml
model = "qwen2.5:7b-instruct-codex"
model_provider = "ollama"

[model_providers.ollama]
name = "Ollama"
base_url = "http://localhost:11434/v1"
wire_api = "responses"

[projects."/home/tim"]
trust_level = "untrusted"

[projects."/home/tim/code"]
trust_level = "trusted"

At that point, a single command is enough to prove you’re local:

codex exec "Reply with exactly: WORKER_READY"

And the “how to actually use it” piece is mostly discipline: Claude does discovery and planning, and when it’s time to change code it produces one tight Codex command that includes the file content inline and says “output a unified diff only, then stop.” That sounds restrictive, but it’s what makes the whole thing fast and reliable on a 7B CPU worker—and it’s what turns the stack from a token furnace into something you can iterate with all day.

Now we get to the part that actually matters: was this worth it?

Because if this whole dance saves pennies but costs minutes, then it’s clever but not practical. So I ran the numbers across comparable tasks—real edits, real delegation loops—not toy prompts.

On cost alone, delegation works. It is dramatically cheaper than running everything through Opus, and marginally cheaper than just using Sonnet directly.

Average Cost Per Task

ApproachAvg Costvs Opus
Opus 4.6$0.0779baseline
Sonnet 4.5$0.013982% cheaper
Delegated (Haiku + Codex)$0.013583% cheaper

That’s not subtle. Opus is expensive. It’s phenomenal, but it’s expensive. Sonnet is already 82% cheaper than Opus for these tasks. Delegation edges Sonnet out by a hair—83% cheaper than Opus—but the difference between Sonnet and Delegated is basically noise.

The more interesting story is speed.

Average Time Per Task

ApproachAvg Time
Sonnet 4.510.2s
Opus 4.612.2s
Delegated101.9s

Delegation is about 8.3× slower than just running Opus directly. And the bottleneck is exactly what you’d expect: Codex execution time on a local 7B CPU model. Once you leave the fast cloud inference path and move to a cheap worker, physics shows up. Thirty to one hundred twenty seconds per edit is normal.

So what does this mean in practice?

It means Sonnet 4.5 is a shockingly strong sweet spot. It’s almost as cheap as delegation—$0.0139 vs $0.0135 per task—and it’s an order of magnitude faster. If you’re optimizing purely for value per minute of your life, Sonnet direct execution is incredibly hard to beat.

Opus, meanwhile, is a premium product. At $0.0779 per task on average—5.6× more than Sonnet—it’s not what you use for mechanical edits. It’s what you use when you need the absolute best reasoning, when architecture matters, when ambiguity is high.

Now here’s the subtle part that changes the calculus: marginal cost.

When I tried to wire cloud Codex back into the stack, auth succeeded, the model resolved, and then I hit “quota exceeded.” That moment was clarifying. The architecture can’t depend on permission from a billing dashboard. If the middle layer disappears when credits run out, then it was never infrastructure — it was a subscription. So I reframed it: cloud Codex is an accelerator, not a dependency. The local worker is the default. Claude plans, the local model executes, and the system keeps running even if the cloud meter stops. If quota comes back, great — I get speed. If it doesn’t, nothing breaks. That shift — from rented capability to owned capability — is the difference between playing with AI and actually building with it.

So you end up with a very human tradeoff:

If you care about speed and you’re willing to pay $0.014 per task, Sonnet is phenomenal.

If you care about minimizing marginal cost and you can tolerate slower edits, delegation is almost free.

If you care about maximum reasoning quality, Opus is worth the premium.

The point of this architecture isn’t that delegation “wins” on every axis. It doesn’t. The point is that you now have control. You can route based on the task. You can pay for judgment and economize on repetition. You can choose whether you’re optimizing for time, money, or cognitive overhead.

And that’s the real unlock.

By 0 Comments