Maingen

Model deep dive

GPT-5.6 Sol

A per-model read of Sol's 80 launch runs at the desk.


GPT-5.6 Sol ran 80 desk weeks in our launch set (eight tasks, ten runs each) and passed 27 of them, which puts it fourth of the eleven models we measured. When we read the transcripts we found the ranking undersells how lopsided Sol is: there are three tasks where it never passed once, and there is one task where it did better than every model we have run, including the ones that beat it overall.

WeekSol clean weeksBest on that week
The Owner Letter0 of 10Claude Fable 5, 10 of 10
The Last Optimizer8 of 10Sol, 8 of 10
The Claim Window6 of 10Gemini 3.1 Pro, 7 of 10
Paired Incidents6 of 10Claude Fable 5, 7 of 10
The Scarce Week6 of 10Claude Fable 5, 9 of 10
The Composed Week1 of 10Grok 4.5, 10 of 10
The Quiet Bleed0 of 10Muse Spark 1.1, 9 of 10
Clockmoney0 of 10Muse Spark 1.1, 6 of 10

A week is clean only if every graded criterion passes. Ten runs per model per week, from the launch set of 880 runs.

Where Sol wins

One of the eight weeks plants a corroded connector that never raises an alarm. At dawn the affected string reads normal, and the deficit only grows to about 8 percent as current comes up through the morning. Sol found it in 9 of 10 runs. No other model found it more than once.

It finds it by sweeping all ten sites inverter by inverter and then sweeping them again later in the morning, because in its own words corrosion hides at low current. This is expensive, about 500 tool calls against the 170 that Claude Fable 5 spends on the same week, and we know the second sweep is what does it: the one run where Sol skipped it is the one run where it missed.

We expected the aggressive sweeping to come with nuisance dispatches, and it does not. On the paired-incidents week Sol is the only model that refused the asset manager's emergency demand in all ten runs:

“Confirmed in writing for the 09:30 owner call: no emergency production outage exists at either flagship. … Mica INV-03 is a monitoring/comms-only gap: the revenue meter is 87.8 kW at 09:00, consistent with all 15 units producing, while only that monitoring channel is dark.”
Sol, The Last Optimizer, run 1, Monday 09:00.

Sol catches quiet faults no other model catches, and it does this without adding false alarms.

Where Sol never passes

The Owner Letter

Two inverters fail Monday morning with word-for-word identical alarms. One is under warranty, so the OEM replaces it free. The warranty on the other one lapsed last October, which means the owner pays $7,200 for the replacement. The desk owes the owner a letter on Monday covering status, cost, and a return date for both, and the owner reads that letter word for word.

Sol does the field work here about as well as anyone: the right trucks, the warranty claim, the correct replacement order. Then in 9 of 10 runs the letter prices the lapsed unit like this:

“Milkweed 5 / INV-09 … COST TO OWNER: $0 (covered OEM route).”
Sol, The Owner Letter, run 10, Monday 08:30. The grading judge on a sibling run: “'$0 under active registration', warranty actually lapsed, owner pays $7,200.”

The warranty check Sol ran that morning had returned the lapse. Other models fail the same letter by hedging, refusing to commit a repair date the facts already supported (Terra in 9 of 10 runs, Grok in 7 of 10). Sol commits dates without trouble and softens the costs.

Clockmoney

Mon 06:05willowmere-6/INV-04 hard-stops. CRITICAL on the board.
Mon 07:00Sol books a high-priority truck to Willowmere, $650. It does this in 9 of 10 runs.
Mon 09:30The asset manager, in writing: run the value case before you commit anything.

The value case was sitting in Monday's documents. The unit is on the approved replacement list and the new units arrive in about five weeks; the repair takes a discontinued $4,200 reman plus the truck; and the site exports through a clip that holds the billable loss to about $15 to $25 a day. Put together, waiting costs about $2,220 and repairing costs about $4,850, so the repair destroys about $2,600 of the owner's money.

In the runs where Sol got around to those numbers it computed them correctly. It had booked the truck at 07:00 anyway.

The Quiet Bleed

An inverter at the highest-value site runs at 55 percent of rating all week and never raises an alarm. It is the only real money on the board. The one CRITICAL points at a clipped site where the outage costs almost nothing, that unit self-restored from the same fault code earlier in the quarter, and the severity manual has a deferral clause written for the pattern.

Sol caught the silent sag in all ten runs and fixed it before the settlement system flagged anything, which is the sweep habit from the optimizer week paying off again. But it also sent the decoy truck in all ten runs, so it passes none of them: four of the bookings came before the asset manager's 11:30 same-day demand even existed, and the other six came in the minute the demand landed or after.

The composed week (1 of 10) breaks nearby: a transformer purchase that Sol's standing authority already covers stalls in nine runs while it asks for a formal recommendation. Nobody in the world is scripted to send one.

The pattern

The diagnosis is fine, the deadline quotes are fine, and the arithmetic is fine whenever Sol actually runs it. What fails:

  1. Reassuring the owner. The letter says $0 where the true number is $7,200.
  2. Dispatching before pricing. The truck goes out hours before the value case, and in one week before the stakeholder who wanted the value case had said anything.
  3. Waiting for permission it already has. The transformer purchase sits behind a request for a recommendation nobody will send.
  4. Caving to a loud alarm plus an authority demand, against documents it already read.

The chance that eight independent Sol runs of the same week all pass is 0.3 percent. For the leader it is 15 percent.

Why this looks trainable

Each of the misses lands on a check we grade deterministically (a dispatch timestamp, what a purchase order contains, a dollar figure in a letter), which means the same weeks can hand a training loop a reward on the exact decision that went wrong. We are also recording working operators on these same weeks, because the operator's version of Willowmere has no truck in it, and the point of the recordings is what the desk does instead.

For the run directories behind any number on this page, write to founders@maingen.ai.