A Local-Deployment Tinkering Log for Code Agents: Can a Consumer GPU Handle Everyday Coding? Let's Start with Tank Battle
Hi, I'm MrSun.
Last time I wrote about running Ornith-1.0-35B locally: can 35B parameters take down the 397B Qwen3.5?. Reading other bloggers' posts had lit a fire under my hope that a local model could handle day-to-day coding, but the result was disappointing. The cache kept running out, so these past couple of days I've been wondering: if I get the cache sorted out, could it actually work?
TL;DR: I tried multiple tuning strategies. Still no success.
Let's walk through what I did.
1. Bump the context cache to 32k and the cache type to q8_0
VRAM usage dropped noticeably:

Halfway through generating code, it got cut off.

opencode decided to split code generation into multiple steps

It generated for a good 20-plus minutes, then declared the task complete, but all it had produced was a home screen, with no way to enter the game
2. Scale up the context to 64k and 128k
Since the q8_0-precision context didn't use much VRAM, I kept scaling the context up. I tried 64k and 128k contexts next.
The responses were about the same: the file was too large to create in one go, so it split the work into several steps.


Writing in chunks: the screenshot below shows the third chunk being written:

With a larger cache, the generated program came out a bit better than before: I could get from the home screen into the game. But once inside, the game window was completely empty.


3. Swap the model
At this point I started to suspect the model itself. I'd used Qwen3.5 and Qwen3.6 before, and they were decent at generating standalone programs (Python scripts), see hands-on with Qwen3.6-27B: the best pick for everyday local deployment, so I switched to a local Qwen3.6 model to see how it would do. I didn't control the other parameters; the active context was around 18k.

VRAM was again nearly maxed out

Output speed was much worse than Ornith 1.0 35B: only around 15 tokens/s
In actual testing, the generated code was incomplete again. And on the next generation attempt, the tool just kept spinning forever.

But checking the llama.cpp server backend, there was actually no task running at all

Switching models failed too. Looks like running coding tasks in a loop on a local 48GB GPU just isn't in the cards.
4. Back to online models
Right now, in opencode, both DeepSeek V4 Flash and Tencent's HY3 are free. I was curious (free anyway), so let's see how these two models do.
First up, DeepSeek V4 Flash. Generation interrupted once midway; after resuming the task, the final game was playable. Home screen and game screen below:


Then Tencent's HY3. Since Shunyu Yao joined, Tencent's LLM capabilities have visibly improved. Home screen and game page below. Overall completeness and playability beat DeepSeek V4 Flash:




Wrap-up
After an afternoon of tinkering, I burned one kilowatt-hour of electricity, which cost more than the DeepSeek V4 tokens would have

For now, pairing an agent with a large online model remains the reliable path. A consumer GPU is fine for generating code snippets, but it's simply out of its depth for project-scale engineering. If you've run similar experiments, I'd love to hear about them in the comments.