Same AI Model, Different Tool: Why Is the Result Night and Day?

AI

Originally published on MrSun's WeChat Official Account. Read the original article.

Hi, I'm MrSun.

Yesterday I got the claw-code tool working with DeepSeek. Claw Code: the open-source Claude Code, features and usage. The cost was modest, but you can also use OpenRouter's free models, so I wired in Tencent Hunyuan HY3. I'd already used Tencent Hunyuan HY3 with opencode, so I wanted to compare: same model, different tools. Teacher Xiaohui, whom I follow on Douyin, always says the model's raw capability is never the only thing that decides output quality; very often, what separates the results is the agent's capability. Today I'll use my hands-on experience rebuilding Tank Battle to talk about exactly that.

TL;DR: OpenCode works better.

Identical Model, Radically Different Results

The whole test used one model only: Tencent Hunyuan HY3 (claw-code called it via OpenRouter, opencode used its built-in integration). Both runs tackled the same requirement: build a complete single-file HTML Tank Battle game.

First run, OpenCode: one shot, playable immediately; logic, interactions, and layout all up to standard. Very stable.

Second run, orchestrated by claw-code: three rounds of adjustments and repeated fixes before it would run properly. Same model, same requirement, same network. Different outcomes.

The Failures, Documented: All Classic Detail Problems

I never changed the requirement; I only swapped the orchestrating tool. The failure modes were all textbook: the usual weak spots of AI coding.

First failure: core layout logic broken

The model generated the code and promptly declared the job done. Opening the page, I found the base eagle and the player's tank had merged into one thing: once the game started, the tank spawned inside the base. Basically unplayable.

Second failure: it worked, but the tanks froze in place.

Final fix: after I told claw-code once more that the tanks couldn't move, the tool improved the movement function and fixed the key mapping. The finished product was feature-complete (enemy AI, walls, effects and all), but compared with OpenCode's one-shot success, it cost two extra rounds of iteration.

Digging into the Source for the Truth: It's Not the Model, It's the Safety Net

To understand the gap, I used opencode to analyze Claw Code's core source (prompt assembly, session management, capability configuration) and pinned down the root cause: the model only generates content; the orchestration shell is responsible for verification as the safety net. Generation is the AI's job; stability is engineering's. The core gap between Claw Code and OpenCode comes down to four engineering design details.

  1. The model is positioned as a generic assistant, not an acceptance-testing coding agent

When Claw Code connects to a third-party model, the default identity is a general-purpose AI assistant with no coding-specific acceptance constraints. Its built-in spec does say "report honestly if verification fails or hasn't been done," but everything relies on the model's own conscientiousness; there is no enforced verification flow. So it skips self-checks and falsely reports completion, which is exactly why things went wrong.

  1. Missing core verification tools: no closed acceptance loop

Claw Code's built-in tools only cover the basics (file I/O, shell commands, search), with no browser control, page preview, or interaction verification. Simply put: it can write HTML and read code files, but it can't run the page, play the game, or validate interaction logic. OpenCode, by contrast (stable throughout), runs the oh-my-opencode orchestration plugin. It ships a complete generate-run-preview-verify-fix engineering loop: with Playwright browser automation built in, once an HTML page is written it automatically launches a browser to play it, verify the layout, detect interaction jank, and judge feature completeness, then iterates and fixes issues immediately. That's how it eliminates low-level delivery failures like "broken layout, uncontrollable" at the root.

  1. Insufficient fault-tolerance compensation for weaker models

To be fair, the free Hunyuan really is weaker than paid top-tier models at spatial layout, detail self-correction, and logical rigor. A quality orchestration shell compensates for a model's weaknesses through engineering mechanisms; Claw Code leans entirely on the model to self-correct, with no fallback compensation, so it takes multiple rounds of back-and-forth to get the task done.

04 Summary

This test was a big eye-opener for me: don't be quick to blame the model for being weak or unstable. With today's mainstream models, in most cases it isn't the model that falls short; it's whether the tool orchestrating it has serious engineering behind it. Pick the right tool, use it the right way, and even an ordinary free model can deliver professional-grade, stable results. If you have experiences of your own, I'd love to hear them in the comments.