junyi.is中文

Three models, one score

I fine-tuned an 8B twice and then swapped in a commercial model. All three scored 0.923. The eval was the problem.

I spent a few weeks distilling the conversation layer of a travel product into an 8B model and preparing it for release. Then I trained a second version on six times as much data, and it scored exactly the same. Then I replaced the whole thing with a commercial model from an unrelated vendor, and that scored exactly the same too.

Three systems with nothing in common, one number. It took me an embarrassingly long time to read that correctly.

What I built

The product routes a conversation through several model roles. One of them takes the user's turn and has to come back with a structured object the rest of the system can act on — slots filled, readiness computed, no leakage of internal reasoning into the reply. That role was on a hosted frontier model and I wanted it local, so I distilled it.

QLoRA 4-bit on Qwen3-8B, LLaMA-Factory, a single RTX 3090. The training data came from capturing the hosted model's behaviour on real conversation paths.

v1  = checkpoint-891
      3 epochs / 891 steps, cutoff 1536
      eval_loss 0.12487

merge to fp16 -> llama.cpp convert_hf_to_gguf -> f16 GGUF
-> ollama create --quantize q4_K_M   (5.0 GB)

The gate was 13 held-out conversations pulled off the production path. The metric I cared about was parseRate: the fraction of replies that parse into the required contract. v1 got 12 of 13, so 0.923, against a threshold of 0.9. Zero leakage, p50 637 ms. It passed the release gate.

The second version

v2 was the obvious next move: more data, more diversity. Seeds were diversified by a second model and soft-filtered by two independent judges, which took the training set to 28,810 samples — about six times v1.

v2  = checkpoint-3400
      2 epochs, cutoff 1600
      28,810 train samples  (~6x v1)
      eval_loss 0.2133

parseRate  0.923   <- identical to v1

Not close to v1. Identical. Still 12 of 13 — though not the same 12. v2 fixed one case and broke a different one, and the net was zero. My gate said “must beat 0.923”, v2 did not beat it, and I shelved it.

The conclusion I drew was that this task had saturated and more data wasn't going to move it. That conclusion was wrong, and I had no way of knowing it from the number I was looking at.

The third data point

A month later I swapped that role out for a hosted commercial model — different vendor, different architecture, not fine-tuned by me, nothing whatsoever in common with the other two. I ran the same 13 cases.

0.923.

At that point the pattern is not subtle. When three independent systems produce a number that agrees to the digit, you are not looking at three results. You are looking at one instrument.

The eval had a resolution of 7.7 points

Thirteen items, scored pass/fail. Every item is worth 1/13 of the score. Nothing smaller than 7.7 points can appear at all — a model that is genuinely five percent better returns a byte-identical number, and a model that trades one win for one loss returns that same number too. The gate wasn't ranking models. It was quantizing them into fourteen buckets and reporting which bucket.

Worse, I had a second instrument telling me something and I ignored it. eval_loss moved the wrong way, 0.12487 to 0.2133, which is what you would expect from a shorter schedule on a much larger and noisier set. Two measurements disagreed; I trusted the coarser one because it was the one written into the gate.

The rule for this is in my own operations handbook, and I wrote it about an unrelated incident: when several independent things deviate by exactly the same amount, the deviation is in the ruler, not in the things. I applied it correctly to a five-hour timezone offset and completely failed to apply it here, for weeks, because here the identical numbers looked like success rather than failure.

What I would do differently

Size the eval before training, not after. The resolution of an n-item pass/fail set is 1/n, so if the decision you intend to make is “is this three points better”, thirteen items cannot make it and no amount of care in running them will help.

Score with partial credit rather than pass/fail per item, so a reply that gets four slots right and one wrong is not the same event as a reply that produces nothing parseable. Compare paired on identical inputs and report per-case deltas, because a net of zero over one win and one loss is a completely different situation from thirteen identical outcomes, and an aggregate hides which one you have.

And write down what the gate is allowed to decide. Mine was built to answer “is this safe to ship”, which it did honestly. I then used it to answer “is this better”, which it was never able to answer at any point.

Notes from the training side

Things that cost me time and are not obvious from the docs:

  • ollama cannot convert Qwen3ForCausalLM from safetensors; the path has to go through llama.cpp's convert_hf_to_gguf.
  • The stock ollama template appends /no_think to the prompt, which is fatal for a contract task — the model stops emitting the structure. Needs a custom TEMPLATE.
  • Contract tasks want temperature=0. Anything above that trades parse rate for nothing you wanted.
  • preprocessing_num_workers must be 1, and the cutoff is worth setting from the measured token distribution rather than a round number — 1536 gave zero truncation on this set.
  • CUDA wheels have to come from download.pytorch.org/whl/cu124; plain PyPI silently installs a CPU build, and you find out at the first training step.

Coda

The training machine became unstable under sustained GPU load before any of this could be revisited. Power delivery was the leading suspect, but I never completed the component-level tests needed to diagnose the failed part. The machine is gone, production runs on hosted models, and my own weights never went in front of a user.

The hardware failing is the least interesting thing that happened. I had built a measuring device, used it to make a decision it could not support, and only noticed when a completely unrelated model walked in and returned the same number.