junyi.is中文

The training machine gave up first

The experiments still had questions left. The machine carrying them became unstable under full load, so I retired the whole system without pretending I had diagnosed the failed component.

I retired my training machine while the work it was built for was still unfinished.

The confirmed failure was not a dead graphics card, or a burned-out power supply, or any other neat component-level diagnosis. What I could actually observe was narrower: under sustained GPU load, the machine repeatedly became unstable and crashed. Power delivery was the leading suspect. I never completed the swap tests or repair work that would have pinned the fault on one part.

So I stopped calling a suspicion a diagnosis and retired the machine as a whole.

That distinction mattered more than I expected. A broken experiment rig creates pressure to explain itself quickly. The symptoms invite a satisfying sentence: the power supply failed, the GPU died, 350 watts killed it. Each version turns an incomplete record into a clean cause. None was supported by the evidence I had.

The less dramatic truth was enough: the machine could no longer be trusted to carry another long run.

What I had asked it to do

The machine was built around a single RTX 3090 and used for an already-retired training experiment, not for current production infrastructure. On that card I fine-tuned Qwen3-8B with 4-bit QLoRA. It was a consumer GPU doing a narrow, practical job: learning one structured conversation role inside a travel product.

The first run stopped at checkpoint 891. Its recorded evaluation loss was 0.12487. On a separate set of 13 pass-or-fail cases, it answered 12 correctly, which appeared as 0.923.

The second run used 28,810 training examples, roughly six times the first set, and stopped at checkpoint 3400. Its evaluation loss was recorded as 0.2133. On those same 13 binary cases, it also answered 12 correctly: 0.923 again.

Those numbers led to a different failure, the one I wrote about in “Three models, one score.” With only 13 binary items, one answer moves the score by about 7.7 percentage points. The evaluation could tell me whether a run landed in one of 14 coarse buckets. It could not reliably tell me whether the larger training set had produced a smaller improvement inside the same bucket.

That is background here, not the ending. The questions survived the evaluation. They also survived the machine.

My own weights cleared a pre-release gate, but they never served users. The product used hosted models instead. That fact made the retirement operationally easier, but intellectually stranger: I was shutting down a machine that still contained unresolved work, without pretending that the unresolved work had become worthless.

Full load is a condition, not a diagnosis

A 3090 operating near full load sits in roughly a 350-watt GPU power-draw class. That number describes the graphics card under load. It is not a measurement of the whole machine, and it is not a fault threshold.

The machine repeatedly lost stability during sustained full-load training. That pattern made power delivery the first thing I suspected. It did not prove the power supply was the failed component. I had no completed component swap, no repair report, and no cross-machine test establishing that the GPU itself was damaged or healthy.

I did not know which component had failed.

Keeping the machine in service would not have made that uncertainty more honest. It would only have attached new experiments to hardware I no longer trusted. A training run can consume hours before instability destroys the attempt, and a successful run can be harder to trust if the system underneath it has started failing unpredictably.

There is a point where “I can probably get one more run out of it” stops being persistence and becomes a way of borrowing confidence from the next result.

I was at that point.

Retiring a system with an incomplete story

On July 5, I retired the training machine and, according to the project archive, backed up the training artifacts, data, configuration, tools, and recovery materials. By July 16, the device was no longer in my possession.

The wording around the backup is deliberate. The project record says those materials were backed up. I did not independently restore them during the fact check for this article, so I cannot claim a verified recovery. A backup record and a successful restoration are different facts.

I kept the upstream material that mattered and did not insist on preserving every derived file. Checkpoints, adapters, merged weights, conversion outputs, the base material needed by the project, datasets, configuration, and recovery notes were recorded as retained. Some derived runtime artifacts could be rebuilt from earlier products and settings, so keeping every copy would have confused “exists somewhere” with “can be recovered deliberately.”

The physical environment disappeared. The record of how to reconstruct the work was supposed to remain.

That was the real retirement task. It was not to save a machine at any cost. It was to decide what would still matter if the machine vanished the next morning, preserve that material with its limitations attached, and stop depending on the rest.

The hardest item to preserve was uncertainty. The archive needed to say that power delivery was suspected, not proven. It needed to say that the 350-watt figure described GPU load, not a measured system draw. It needed to say that the artifacts were reported backed up, not independently restored. Otherwise the documentation would become more confident as the evidence became less available.

A retired system is especially vulnerable to false certainty. Once the hardware is gone, nobody can casually rerun the missing test. A guess written in the present tense hardens into project history.

What I refused to calculate

The obvious next paragraph was an economic comparison between owning the card and paying for hosted models. I deleted it.

I did not have a defensible ledger. I lacked a common accounting of purchase cost, residual value, electricity, training duration, failed reruns, repair time, hosted-model usage, and equivalent output quality. Without those, “the API was cheaper” would have been a preference dressed as arithmetic.

After retiring the machine, I chose not to maintain another dedicated training system, and the product continued with hosted models. That is a decision I made. It is not a universal cost conclusion.

Deleting the comparison improved the story because it left the decision where the evidence stopped. I did not need to prove that the replacement was economically optimal. I only needed to admit that the old machine had become an unreliable place to put new work.

The machine exited; the questions did not

The experiment left behind more than two checkpoints. It left a coarse evaluation that needed redesign, a set of training materials that should be recoverable according to the archive, and a lesson about ending systems without inventing a final diagnosis.

I used to think retirement followed understanding: first determine exactly what failed, then preserve everything important, then shut the system down. This machine forced a different order. The instability was real, the component cause remained uncertain, and waiting for a perfect explanation would have meant continuing to entrust work to an unreliable platform.

So I retired the platform first. I kept the evidence separate from the suspicion. I preserved what the project record said could carry the work forward. And I left the unresolved questions unresolved.

The training machine gave up first.

The work did not.