AI Guides › Workbench
By Nigel Guy · 6 min read
Most people judge a new coding agent by its star count and the loudest comment thread, then switch because "free" sounds like a saving. Stars measure attention, not whether the tool survives your Tuesday afternoon refactor. DeepSeek Harness is a good test case, because it is real, it is open source, and it is also a young developer preview.
The rule: separate the harness from the model, price each one on its own, and switch only when a trial on your own repository beats your current setup.
DeepSeek Harness (command: dsh) is an open-source agent runtime from DeepSeek, released under the MIT licence. Its repository describes it as an agent harness built on the Cordis framework, with the idea that "everything is a plugin". Per InfoQ's report (20 August 2026), model adapters, tool registries, sandboxing, session state and user interfaces are loaded as separate plugins, and an append-only log records every message, tool call and token count so you can replay a session.
The repository's own README says it is in developer preview, that compatibility-breaking changes are expected, and that it is iterating fast. The GitHub page showed about 243,000 stars when we checked. We could not verify claims that it broke a particular star-growth record, so treat that part of the hype as unconfirmed.
A harness is the loop around a model: it reads files, runs commands and feeds results back. Claude Code and Codex are harnesses too. What differs is that this one is MIT-licensed and model-agnostic. Per DataCamp's tutorial, you add providers such as OpenAI or Anthropic under Settings, then Models, or point a custom provider at a compatible endpoint, and switch from the model picker without restarting.
| DeepSeek Harness | Your current paid agent | |
|---|---|---|
| Harness cost | Free (MIT licence) | Subscription or usage; check the £ price at checkout |
| Model cost | Pay per token to whichever provider you connect | Usually bundled |
| Best for | Tinkerers, people who want to inspect and swap every layer | People who want it to work on day one |
| Catch | Developer preview; plugin setup and docs are rough | Locked to one vendor's models and rules |
"Free" applies only to the harness. The model behind it still bills you. At time of writing, DeepSeek's pricing page lists two models, each with a 1M-token context:
| Model | Input, cache miss, per 1M tokens (off-peak) | Output per 1M tokens (off-peak) |
|---|---|---|
| deepseek-flash | $0.15 (about £0.11 at time of writing) | $0.60 (about £0.45 at time of writing) |
| deepseek-v4-pro | $0.66 (about £0.50 at time of writing) | $1.98 (about £1.49 at time of writing) |
Peak hours (01:00-04:00 and 06:00-10:00 UTC, weekdays) cost double. That window includes a UK working morning, so budget for the peak column. Prices move; check the current page and the £ figure at checkout.
Run this before you switch. It is a two-column ledger: what you give up, what you gain, each with evidence from your own work.
Step 1: Set up a trial that costs almost nothing. Install Node.js, then run npx @deepseek-ai/dsh web (the README's route). It opens a local web interface at http://127.0.0.1:3080. Set DEEPSEEK_API_KEY in your environment, or add the key under Settings, then Models. DataCamp suggests about $2 of credit (about £1.50 at time of writing) is plenty for a first test. Use a copy of a real repository, never your only copy.
Step 2: Pick three real tasks. One small bug fix, one multi-file change, one task that needs web research or an image. Run the same three in your current tool and in Harness.
Step 3: Fill in the ledger. For each task, record whether it finished without prompting, how many times you had to say "continue", what it cost, and whether the result passed your tests. DataCamp's tester reported the agent sometimes stopping mid-task and needing a nudge, and found plugin installation hard going. Your result may differ; that is the point of the ledger.
Step 4: Check what you would lose. Write down every feature of your current tool you use weekly: IDE integration, saved settings, team sharing, support. Plugins may or may not replace them.
Step 5: Decide by column, not by feeling. Switch if the gain column holds up on your tasks. Keep both if the harness wins on cost but loses on reliability. Check back in a month, because a preview changes quickly.
Say you are a freelance developer on a £20-a-month agent subscription, mostly fixing bugs in a mid-sized web app. Your ledger shows Harness with the cheaper flash model handled the bug fix and the multi-file change, but stopped twice on the longer one. The image task failed until you added a community vision plugin, because DataCamp reports the DeepSeek chat model in Harness is text-only. The verdict might be: use Harness for small, cheap tasks and keep the subscription for long ones. Your numbers will differ; the method is what transfers.
Use this with whichever assistant you trust. Fill in your current tool and what you do with it.
You are a sceptical engineering advisor helping me decide whether to adopt a new coding agent.
Context:
- My current coding agent: [CURRENT_TOOL]
- How I use it day to day: [YOUR_WORKFLOW, e.g. languages, repo size, solo or team]
- The tool I'm considering: DeepSeek Harness (open-source agent runtime, MIT licence, plugin-based, model-agnostic, currently a developer preview)
- Which model I'd run behind it: [MODEL_AND_PROVIDER]
- My monthly budget: [BUDGET_IN_POUNDS]
- Features I rely on weekly: [FEATURE_LIST]
Goal: tell me honestly what I would give up and what I would gain by switching, based on how I actually work. Do not assume that newer, free or open source means better.
Steps:
1. List any details about my workflow you still need. If anything important is missing, ask me before answering.
2. Build a two-column table: "What I would give up" and "What I would gain", each row tied to something I told you.
3. Flag every claim you cannot verify, including anything about current pricing or features. Tell me where to check.
4. Propose a three-task trial I can run this week, with a pass/fail test for each task.
5. Finish with a recommendation: switch, run both, or stay, and what evidence would change your mind.
Before answering, check that every row in your table traces back to my inputs rather than general opinion.