← Back to home@philipho01

dsh-llm-fidelity

DeepSeek Harness 插件:模型服务明确拒绝参数(400/422)时自动修正重发,对话不中断;用量统计防清零、防重复,如实可信。

Stars
0
Language
TypeScript
Created
Oct 1, 2026
Updated
Oct 1, 2026
GitHub repo

Introduction

dsh-llm-fidelity

English | 简体中文

Your AI agent breaks mid-task because a model rejected one optional parameter. Your usage dashboard shows numbers that don't add up. dsh-llm-fidelity fixes both — automatically, invisibly, and honestly.

A DeepSeek Harness plugin. Install it, change nothing else.

What changes after you install it

A real scenario, before and after:

You asked your Agent to fix a bug. Three steps in, the model service suddenly returns:

400 — 'stop' is not supported with this model.

Without the plugin — the conversation dies on the spot. You dig through provider docs, edit a config, start a new session, and re-explain the whole task. Switch providers next week? Same wall, different parameter.

With the plugin — that 400 is absorbed in milliseconds: the offending parameter is dropped and the identical request re-sent. What you see is the answer simply continuing. And the plugin remembers this provider's temperament — the next request never even hits that error. You never learn the error existed.

The same promise applies to 'max_tokens' is not supported…, 'temperature' is not supported…, reasoning effort is not supported…, and friends — plus the quieter pain of usage statistics being wiped to zero or counted twice by gateways. All of it stops at the plugin.

  • 🔧 Conversations that don't break — a rejected parameter is fixed and re-sent within the same request. You never see the error.
  • 💯 Numbers you can trust — usage statistics land in your accounting exactly as the provider reported them. No wiping, no double-counting, no invented figures.
  • 🚀 Zero config — smart defaults work out of the box; the plugin registers no model routes and never conflicts with official or community adapters.

Sound familiar?

If you have ever pointed an AI coding agent at a third-party model or gateway, you have almost certainly met one of these:

Unsupported parameter: 'max_tokens' is not supported with this model.
Use 'max_completion_tokens' instead.

Unsupported parameter: 'stop' is not supported with this model.

'temperature' is not supported with this model.

reasoning effort is not supported by this model

These are real, current, and everywhere. OpenAI's reasoning models (GPT-5 and GPT-6 series) reject max_tokens, stop, and temperature — and the rejections hit users through LiteLLM, Azure OpenAI, crewAI, and every OpenAI-compatible gateway in between. Passing stop: null does not help (crewAI #1908). The max_tokens → max_completion_tokens rename even carries a trap: the new field includes hidden reasoning tokens, so a naively converted budget returns empty responses that you still pay for.

What actually happens to you: your agent is three tool calls into a task, hits one of these errors, and the whole turn dies. You dig into provider docs, edit a config, restart, and pray the next provider doesn't reject something else.

And the quieter problem — your usage stats. Gateways re-send a zeroed usage snapshot at stream end, wiping the real numbers. The same snapshot arrives twice and gets counted twice. Cache-hit fields go missing. Your dashboard says 3.2M tokens; your invoice says 5.1M; nobody can explain the difference.

What dsh-llm-fidelity does about it

When a model explicitly rejects an optional parameter, the plugin fixes it and re-sends — within the same request, in milliseconds, invisibly.

You send:  { temperature: 0.7, stop: ["END"], ... }
Model:     400 — 'temperature' is not supported with this model.
Plugin:    (removes temperature, re-sends immediately)
Model:     200 — here is your answer.
You see:   the answer. Nothing else. Ever.

The fix list is deliberately short: temperature, reasoningEffort, stop, maxTokens (omitted, or clamped down when the error states an explicit output limit). Each rejected field is read from the provider's own error message — no guessing, no brute-force retries. And the plugin remembers which provider accepts which combination (1 hour, in-process), so the failure happens at most once per route. The second request already speaks that provider's dialect.

For the usage numbers, the plugin guards the statistics stream frame by frame: a later frame that omits fields can't erase earlier ones, an all-zero placeholder can't wipe real totals, a duplicate can't be counted twice. And one rule above all: if a number was never reported, the plugin never invents it. One missing reference number beats a fake one.

Why this one

There are a dozen ways to paper over a 400 error. This plugin chooses the boring, correct ones — on purpose:

dsh-llm-fidelity
Fixes only what the provider explicitly rejected✅ Reads the rejection from the error itself; never guesses. Context-overflow, quota, auth, and network errors are passed through honestly — a real problem should reach you.
Never touches your content✅ Messages, system prompt, tools, and constraints are untouchable. Only optional scalars (temperature / reasoningEffort / stop / maxTokens) may be dropped or clamped.
Bounded, no magic✅ At most 4 adaptations per request, each of which must actually change the request. No infinite retry loops, no silent protocol switching.
Learns✅ Per provider+model, in-process, 1-hour TTL. Fail once, never again on that route.
Plays well with others✅ Registers zero model routes and replaces zero services. Official adapters, community adapters, and the official dsh-llm-retry keep working untouched — 429 rate limits and server errors are its job, and this plugin steps aside for them.

Install

Compatibility: works with DSH 0.1.0-rc.6 and later, across the whole 0.2.0-rc line (tested against DeepSeek desktop 0.2.0-rc.2). If your DSH reports an incompatibility on install, check your version first.

dsh plugin add dsh-llm-fidelity

or any npm-based workflow:

pnpm add dsh-llm-fidelity

Then enable it in the plugin settings. Defaults are production-ready; every switch is optional:

OptionDefaultMeaning
usageFidelitytrueusage guarding (dedup, anti-wipe, field merging)
rectifytrueautomatic parameter-rejection adaptation
rectifyStatuses[400, 422]statuses treated as explicit parameter rejections
maxAdaptations4maximum automatic re-sends per request
strategyTtlMs3600000how long a learned provider preference is kept (ms)
strategyCacheSize256maximum learned preferences kept in memory
rectifyAuxiliaryfalsewhether background tasks (e.g. auto titling) are adapted too

Changing the configuration

The plugin installs as a profile bundle layer, and its configuration lives in the profile's patch file (on the desktop app: ~/.dsh/profiles/desktop/cordis.patch.yml). Override the row by its id llm-fidelity:

- id: llm-fidelity
  config:
    maxAdaptations: 2
    rectifyAuxiliary: true

Note: a patch's config replaces wholesale rather than merging — restate every key you want to keep (keys you omit fall back to the built-in defaults).

How it works (for the curious)

The plugin listens on the llm/stream event — the one road every model request travels. It registers no model routes and replaces no service; it adds two checks beside the road:

  • Parameter adaptation: watches for explicit parameter rejections, adjusts the offending optional field, and re-enters the queue through the public ctx.llm.stream() entry (the exact path a normal request takes, so internal state stays consistent). A re-send may only happen while nothing has been delivered downstream — once content reached you, errors pass through verbatim. No replay, no fabricated success.
  • Usage guarding: walks the statistics frames, merges missing fields, drops duplicates and placeholders, and hands the final numbers to the accounting layer right before the turn closes.

Honest boundaries

  • The plugin works at the unified data layer and sees standardized frames only. If a service never reported a number (say, cache hits), it cannot be recovered — this plugin keeps data clean, it cannot conjure data.
  • OAuth-based services, oversized contexts, and exhausted balances are out of adaptation scope by design.
  • The session log records the original request parameters; adaptation is transparent, and replay provenance follows the winning attempt.

License

MIT