Eighty runs to change nothing, and a lever nobody could pull
The watering can went in this morning. Tonight was the part nobody films: finding out whether it broke the game. It did not, and proving that took eighty ten-day runs.
The question is sharper than it sounds. Crops and saplings now grow at a third of their pace until somebody pours a can on them, and the water comes from one well in the middle of the village — so every plant that wants watering is also a walk. That walk is time the player does not spend cutting wood, and wood is defence. The fear was that a farming idea had quietly become a tax on the guns.
The first answer looked bad and was wrong. A ten-seed run held eight villages where the old build held nine, which is exactly the kind of number that starts a panic and a redesign. A second analyst read the same logs without being told what to look for, and ran the test that mattered: the two runs disagree on three villages out of ten, two one way and one the other, which is a coin landing heads twice. The village that falls is a different village every time, and it always falls on the same day — the day is the structure, the name is noise.
So the question was asked properly: one build, one setting, forty seeds the work had never seen, three values of how much slower a dry plant grows. Ninety-two villages in a hundred survive when water is free; eighty-two when it is at its shipped strength; and the middle setting comes out below both, which is the giveaway that the differences are noise rather than a slope. The paired test agrees. Nothing moved.
What did show up, in the window where every seed is still alive, is the trade itself: with the dry penalty on, a village runs about three trees, eight logs and six planks poorer — and about eight grain richer. Wood down, food up. That is the bargain the mechanic was designed to strike, and it is good to see it in the numbers rather than in the hope.
Two smaller things, both of them the tools turning on their owner. The first: a run now prints the settings it actually resolved, because the engine’s way of setting one from the command line prints a reassuring line and then a second line that says it has created a dummy — and neither of those is proof. The value does land. But it should not take an experiment to know that, so now a single line in the log says what the run really believed.
The first time that line ran, it found a bug. One of the levers read back as “not there”. It had been declared inside the function that uses it, which in this language means the setting does not exist until that code has run once — so typing it into the console before the first wave answered “unrecognized”, and no run could ever prove which value it had used. It had been sitting there for days, in the one place where being able to change a number by hand matters most. Moved, and it reads back correctly now.
The second: a finished set of twenty runs was re-run against itself, deliberately with the machine buried under other work, to check that the results are repeatable. Nineteen came back identical line for line. One matched perfectly for nine days and then went its own way on the tenth. That is one in twenty — the same size as the effects these runs are asked to detect, which is precisely why tonight’s ten-point gap evaporated when it was tested properly. Knowing the ruler bends is worth more than another confident measurement taken with it.