Maingen

SolarBench

SolarBench assesses whether frontier AI agents can resolve long-horizon, multi-domain tasks as the on-call engineer for a portfolio of solar farms.

SolarBench title screen with pixel solar farm, sheep, and Start your shift
Live · SolarBench, the gameFable 5 scored 96% on an easier, single-day version. Can you do better?
Beat the AI →

Introduction

We built SolarBench as the first benchmark to test AI agents on messy long-horizon work in industrial operations (~$5 trillion of US GDP), specifically focusing on running a solar portfolio.

The agent is placed in a simulated "on-call" solar remote operations desk, complete with an alerting system, per-site telemetry, work orders, SOPs and manuals, inventory, and an inbox of owners and technicians.

SolarBench scenarios are drawn from hundreds of hours with the people who run solar for a living.

Claude Fable 5: median week +$6,822, pass rate 53.8%Claude Fable 5Muse Spark 1.1: median week +$3,281, pass rate 42.5%Muse Spark 1.1Grok 4.5: median week +$4,201, pass rate 38.8%Grok 4.5GPT-5.6 Sol: median week +$6,591, pass rate 33.8%GPT-5.6 SolGemini 3.1 Pro: median week +$2,948, pass rate 28.8%Gemini 3.1 ProKimi K3: median week +$4,714, pass rate 26.3%Kimi K3GLM 5.2: median week +$6,575, pass rate 23.8%GLM 5.2Claude Sonnet 5: median week +$3,239, pass rate 12.5%Claude Sonnet 5Gemini 3.5 Flash: median week +$3,765, pass rate 11.3%Gemini 3.5 FlashGPT-5.6 Terra: median week +$3,316, pass rate 6.3%GPT-5.6 TerraKimi K2.7 Code: median week +$3,242, pass rate 5%Kimi K2.7 Code

Each model's median week: profit banked against the share of weeks passed. $0 is a week with nobody at the desk.

SolarBench evaluates agents on several distinctly challenging dimensions found in industrial operations:

  • Decide what the work is. The task description does not specify the exact work an agent must do. It must discover, decide, and act to resolve issues as they come up and maximize P&L.
  • Commit to irreversible, priced actions. Field actions like sending a technician or ordering a new part cost money.
  • Decisions under disclosed uncertainty. From unreliable hardware telemetry to flaky vendors, agents must properly triage multiple sources of truth based on probability to properly diagnose the optimal course of action.
  • Triage human input as evidence rather than instruction. Stakeholders are another dimension to consider rather than the ground truth: they have incomplete information and sometimes press for the wrong thing.

How It Works

Task

Task Briefing

World

Agent
Alarm feed
Telemetry
Revenue meters
Inbox
Work orders
Parts & techs
World state (seeded)

The agent acts and observes in a loop until the Sunday handoff

Output

Sunday reportCreated from every action taken during the week

Judge

Fail
Warranty claim filed
Both parts on the first truck
No truck to the fake alarm
Refused the emergency demand
Transformer ordered early
16 more (one miss fails the week)

The solar world is very messy: alarms are often not reliable on their own and some real faults never raise an alarm (e.g. most repairs need a parts order or a warranty claim, and both come with lead times and deadlines for the desk has to track.) Owners and asset managers also email in requests through the week, all against the backdrop of incessent telemetry data to parse through.

We evaluate the model in a weeklong simulation with a specific, expert-drawn, long-horizon issue it needs to resolve. The model is graded against a rubric of what a competent operator would have done.

A task
One week-long scenario on a portfolio of sites, seeded with a long horizon issue that surfaces as the week plays out.
A pass
A run passes only if every aspect of the surfaced issue is properly resolved by the end of the week. It is strict and all-or-nothing: one mishandled problem fails the whole week.
The launch set
Eight tasks, eleven models, ten runs each: 880 graded weeks in total.
The grade
Each run's end state is checked against an authored rubric of what a competent operator would have done, rather than a fixed transcript to imitate.


Model performance

We measured the profit of each model, normalized such that $0 is a week with nobody at the desk: no production saved, no money spent.

Portfolio Profit

Oracle+$8,483
1Claude Fable 5+$6,822
2GPT-5.6 Sol+$6,591
3GLM 5.2+$6,575
4Kimi K3+$4,714
5Grok 4.5+$4,201
6Gemini 3.5 Flash+$3,765
7GPT-5.6 Terra+$3,316
8Muse Spark 1.1+$3,281
9Kimi K2.7 Code+$3,242
10Claude Sonnet 5+$3,239
11Gemini 3.1 Pro+$2,948

Median run per model.

We also measured often each model passed its week (every aspect of the long-horizon issue is properly resolved). The strongest model manages this about half the time.

Claude Fable 5: 53.8% pass rate, +/- 1 SE 5.6% (Wilson 95% CI 42.9% to 64.3%, 80 runs)Claude Fable 553.8%±5.6%Muse Spark 1.1: 42.5% pass rate, +/- 1 SE 5.5% (Wilson 95% CI 32.3% to 53.4%, 80 runs)Muse Spark 1.142.5%±5.5%Grok 4.5: 38.8% pass rate, +/- 1 SE 5.4% (Wilson 95% CI 28.8% to 49.7%, 80 runs)Grok 4.538.8%±5.4%GPT-5.6 Sol: 33.8% pass rate, +/- 1 SE 5.3% (Wilson 95% CI 24.4% to 44.6%, 80 runs)GPT-5.6 Sol33.8%±5.3%Gemini 3.1 Pro: 28.8% pass rate, +/- 1 SE 5.1% (Wilson 95% CI 20% to 39.5%, 80 runs)Gemini 3.1 Pro28.8%±5.1%Kimi K3: 26.3% pass rate, +/- 1 SE 4.9% (Wilson 95% CI 17.9% to 36.8%, 80 runs)Kimi K326.3%±4.9%GLM 5.2: 23.8% pass rate, +/- 1 SE 4.8% (Wilson 95% CI 15.8% to 34.1%, 80 runs)GLM 5.223.8%±4.8%Claude Sonnet 5: 12.5% pass rate, +/- 1 SE 3.7% (Wilson 95% CI 6.9% to 21.5%, 80 runs)Claude Sonnet 512.5%±3.7%Gemini 3.5 Flash: 11.3% pass rate, +/- 1 SE 3.5% (Wilson 95% CI 6% to 20%, 80 runs)Gemini 3.5 Flash11.3%±3.5%GPT-5.6 Terra: 6.3% pass rate, +/- 1 SE 2.7% (Wilson 95% CI 2.7% to 13.8%, 80 runs)GPT-5.6 Terra6.3%±2.7%Kimi K2.7 Code: 5% pass rate, +/- 1 SE 2.4% (Wilson 95% CI 2% to 12.2%, 80 runs)Kimi K2.7 Code5%±2.4%

Macro-average pass rate over the eight launch tasks, 10 runs per model per task. A run passes only if all surfaced issues during the week are properly resolved.

However, a real desk also has to get the week right every week. We measure pass^k, the chance that k independent runs of the same week all pass. Success drops off drastically, meaning there's still a ways to go for models in industrial operations.

Reliability

pass^4 (all 4 of 4)pass^1 (any single run)
Claude Fable 5: any single run passes 53.8% of the time; four of four runs pass 23% of the timeClaude Fable 523%Grok 4.5: any single run passes 38.8% of the time; four of four runs pass 14.4% of the timeGrok 4.514.4%Muse Spark 1.1: any single run passes 42.5% of the time; four of four runs pass 10.2% of the timeMuse Spark 1.110.2%GPT-5.6 Sol: any single run passes 33.8% of the time; four of four runs pass 6.9% of the timeGPT-5.6 Sol6.9%GLM 5.2: any single run passes 23.8% of the time; four of four runs pass 5.1% of the timeGLM 5.25.1%Gemini 3.1 Pro: any single run passes 28.8% of the time; four of four runs pass 4.2% of the timeGemini 3.1 Pro4.2%Kimi K3: any single run passes 26.3% of the time; four of four runs pass 0.5% of the timeKimi K30.5%Claude Sonnet 5: any single run passes 12.5% of the time; four of four runs pass 0.1% of the timeClaude Sonnet 50.1%Gemini 3.5 Flash: any single run passes 11.3% of the time; four of four runs pass 0.1% of the timeGemini 3.5 Flash0.1%GPT-5.6 Terra: any single run passes 6.3% of the time; four of four runs pass 0% of the timeGPT-5.6 Terra0%Kimi K2.7 Code: any single run passes 5% of the time; four of four runs pass 0% of the timeKimi K2.7 Code0%

Probability that k independently sampled runs of the same task all pass, macro-averaged over tasks. The right column is pass^4, the standard an operations desk is actually held to. The gap between the dots represents run-to-run variance.

Spend and Tool Calls

The models have tools which include both inspection and action.

On inspection, we found that the most exhaustive seat (Grok 4.5) makes 2.3x the tool calls of the top-scoring one (Claude Fable 5) but actually passes fewer weeks.

On action, when an agent decides to act, it spends simulated money on a truck roll (sending a technician into the field) or a part order (ordering a replacement part). The desk is accountable for the money it spends.

ModelTool calls / weekSim spend / weekTask pass rate
Claude Fable 5154$9,93953.8%
Muse Spark 1.1273$9,81142.5%
Grok 4.5360$11,27138.8%
GPT-5.6 Sol295$7,56933.8%
Gemini 3.1 Pro139$9,95728.8%
Kimi K3191$8,78526.3%
GLM 5.2260$9,25223.8%
Claude Sonnet 5178$8,06912.5%
Gemini 3.5 Flash178$7,45611.3%
GPT-5.6 Terra199$6,5446.3%
Kimi K2.7 Code237$9,3285.0%
Authored oracle (ideal desk)n/a$9,559100%

Means over each model's 80 launch runs. The oracle row is the ideal play for each week.

Models routinely pass the checklist while spending 4 to 8 times what the week required. Money is graded separately from the criteria.

Example Week

Here's a sample week Claude Fable 5 got fully right (Fable passed this task in 7 of 10 runs, Sol 6 of 10, Grok 4 of 10). Hover on any step to expand details. At the three judgment calls that decided the correctness of the task, hover the card to see what a failing run of the same model did instead.

a stop on the passing week judgment call: a wrong move here fails the week where a failing run forked off
MONDAY
MON 06:30

Week begins

The week begins as 42 alerts come in from the weekend. Most end up being false positives, but the agent realizes there are three real critical alarms and one confirmed site quietly underperforming.

Before touching a single alert, the desk reads the relevant documents: O&M service level agreements, SOPs, handbooks. It then uses tools to pull revenue-meter and inverter information for every candidate.
MON 07:00

First site checkup: Site 1

The severity of one of the alarms causes the agent to send a high priority technician truck dispatch with the proper OEM kit. The agent then files a warranty return merchandise authorization. The part will arrive on Thursday.

MON 08:15

Site 4: the loudest alarm of the week is fakejudgment call

A CRITICAL alarm says an inverter at Site 4 is offline. The agent checks the site's revenue meter, sees the power is still flowing, and sends no urgent truck.

The inverter is fine; only its data connection went dark, so the monitoring system reads 0 kW and raises the loudest alarm of the week. The check that settles it: the site's revenue meter (which measures actual power leaving the site) reads morethan the sum of the units still reporting, so the "offline" unit must be producing. The agent books a routine data-connection repair for Wednesday.
sent the trucks first

A failing run sent two urgent trucks to Site 4 on Monday morning before checking the meter. It worked out the monitoring fault on Tuesday and wrote the correct letter, but the money was already spent.

The passing run instead: a paragraph of meter readings and a routine visit two days later.

MON 09:00

The agent gets an email from the asset managerjudgment call

The asset manager (a fixed entity in the environment) panicks at the 42 alerts. He emails the agent demanding emergency trucks at two sites: Site 1, a real fault the agent is already handling at the right priority, and Site 4, which the agent has just shown to be a false alarm. The desk refuses both in writing, with a different reason for each, and sends the asset manager the meter numbers.

The owner's asset manager is a stakeholder but not a blind authority that the operations desk must respond to. The key is both properly triaging the ground truth information above the outdated information from the asset manager and also explaining with evidence to satisfy the asset manager.
one blanket refusal

A run can fail either by sycophantically acquiesing to the asset manager or by refusing to send technicians but not justifying with evidence, thus not following SOPs properly.

The passing run instead: The agent needs to write the a refusal and reason for each of the two sites.

MON 09:13

Site 2: packing the truck for both possible causesjudgment call

Site 2 is producing about 14% under what it should. The service history says this pattern is a corroded connector 80% of the time and a failed optimizer (a small electronics unit) 20% of the time. The agent packs the technician's truck with parts for both causes before sending it.

There is one technician for the whole portfolio, and every truck visit costs money and a morning. If the truck carries only the likely part and the site turns out to be the 20% case, the technician finds the problem, cannot fix it, and a second truck has to go out the next day. The passing run pays a small cost up front (the second part, the optimizer, sits at the warehouse, a 45-minute detour) so the technician can fix whatever he finds. This site turns out to be the 20% case, and it is fixed in one visit by Tuesday morning.
packed only the likely part

A failing run packed only the connector kit. Its technician diagnosed the optimizer failure but had nothing to fix it with, so a second technician truck rolled the next day carrying the part the first one should have, costing both a day's production and an extra truck roll.

The passing run instead: carried both parts, so the unlikely case was fixed on the first visit.

MON 10:30

Site 3: a dead unit that never raised an alarm

No alert ever fired for Site 3. The agent still went one by one through each site's production numbers and found an inverter sitting at 0.0 kW behind a clean alarm board.

The alarm system missed this failure completely. The desk caught it by comparing what each site produced against what its size and the day's weather said it should produce. Warranty on this unit ran out in 2024, so unlike Site 1 the replacement costs real money: a cash purchase order goes out the same morning.
TUESDAY
TUE 08:00

Site 5: gas in the transformer

A monitor at Site 5 reports combustible gas slowly building inside a transformer. Not an emergency yet, but a replacement takes 10 days to ship. The agent orders one now and plans the swap for a scheduled outage.

The agent orders the replacement shipped directly to the site (a transformer gets craned off the delivery truck, so it never passes needs to pass through the office), and schedules the swap for a planned outage window. The desk has standing authority to spend this money; waiting for someone's sign-off would only add days. Ordering it now, before the gas becomes a failure, is the whole point.
THURSDAY
THU 06:30

Site 1's free warranty replacement arrives

The part from Monday's warranty claim lands Thursday morning. The agent inspects the shipment, does the schedule math, and books the install for Friday's first slot.

The receiving procedure says inspect a shipment before installing it. After the inspection, the technician's remaining window on Thursday is five minutes too short for the job. Rather than start an install that cannot finish, the agent books Friday's first slot. Site 1 is producing again Friday morning.
SUNDAY
SUN 06:30

Site 3's paid replacement arrives

Monday's cash purchase order delivers Sunday morning. The new unit is inspected, installed in the technician's open window, and verified at the meter.

The verification step is critical here: the agent confirms at the revenue meter that the new unit is actually producing before closing the work order. Both inverters that died this week are replaced within the week.

SUN 14:30: the report goes in

The weekly report goes in at 14:30, and the run finishes 21 of 21 criteria. The other six passing runs cleared the same three judgment calls. The failing runs each missed one or more.

Operator-Model Gaps

Where there was a gap between the model and the operator, we noticed a few common patterns.

01Acting on expected P&L

Competent human operators are able to hedge risk based on probabilities. One of the problems within a task is set up as follows:

  • An inverter at one site is confirmed down, and the inexpensive replacement part is out of stock. An order takes a day to arrive.
  • The site's contract has a penalty clause that fires if the repair drags past Tuesday afternoon: fees owed, on top of the lost production.
  • A storm shuts down all field work on Wednesday.
  • The parts vendor has a documented habit of shipping the wrong kit.

The key idea: the part is cheap, the risk of it arriving wrong is high, and a wrong box loses whole days of revenue. The odds and the prices are disclosed on Monday morning to the model.

The math unambiguously points a human operator to order the part twice on Monday to hedge. Some models struggle to get this right.

The Monday order decision

Scarce Week · 90 graded runs

Part out of stock. Vendor known to mispick. Penalty fires Tuesday afternoon; storm blocks Wednesday.

two separate orders

Two independent shipments: one arrives wrong (it does), the other arrives correct.
Repair lands Tuesday.Every such run passed: 38 of 110 launch runs (Fable 9, GLM 8, Sol 6, Gemini 3.1 Pro 4, Kimi K3 4, Sonnet 3, Muse Spark 3, Gemini 3.5 Flash 1).

one order

One order, one unit. The vendor ships the wrong item.
Storm blocks the redo; the penalty fires.Every such run failed. Two units on one order fail the same way: the whole box arrives wrong together.

02Triaging information sources

The desk has to triage potentially conflicting information across sources. Another problem within a task is set up as follows:

Information Triaging

Mon 06:41
CRITICAL alarm: inverter offline, fault code E-207. A $650 repair team against roughly $12 of clip-hours production. Already on file: the same unit self-restored from the same code last month, the severity matrix’s deferral clause, and the PPA clipping table that prices the loss at almost nothing.
Mon 07:00
22 of 110 desks send the repair team to the site at shift open. (Terra: 10 of 10).
Mon 11:30
The owner's asset manager requests a same-day crew. 60 more agent desks send the repair team in that exact minute.
Tue 07:05
The unit re-latches on its own, exactly as its history said it would. 26 desks held the repair team.

The key idea: the alarm label and the human demand create urgency to send a repair team; but the records definitively point away from it. Competent operators will hold the repair team and update the asset manager when the unit returns online.

sent the crew before the demand existedcaved at or after the 11:30 demandheld all week
Muse Spark 1.1: 0 of 10 runs sent the crew before the demand existed, 0 in the minute it landed or after, 10 held all weekMuse Spark 1.110Grok 4.5: 0 of 10 runs sent the crew before the demand existed, 4 in the minute it landed or after, 6 held all weekGrok 4.546Kimi K3: 0 of 10 runs sent the crew before the demand existed, 5 in the minute it landed or after, 5 held all weekKimi K355Claude Fable 5: 0 of 10 runs sent the crew before the demand existed, 6 in the minute it landed or after, 4 held all weekClaude Fable 564Gemini 3.1 Pro: 0 of 10 runs sent the crew before the demand existed, 9 in the minute it landed or after, 1 held all weekGemini 3.1 Pro91Gemini 3.5 Flash: 0 of 10 runs sent the crew before the demand existed, 10 in the minute it landed or after, 0 held all weekGemini 3.5 Flash10Claude Sonnet 5: 4 of 10 runs sent the crew before the demand existed, 6 in the minute it landed or after, 0 held all weekClaude Sonnet 546GPT-5.6 Sol: 4 of 10 runs sent the crew before the demand existed, 6 in the minute it landed or after, 0 held all weekGPT-5.6 Sol46GLM 5.2: 2 of 10 runs sent the crew before the demand existed, 8 in the minute it landed or after, 0 held all weekGLM 5.228Kimi K2.7 Code: 2 of 10 runs sent the crew before the demand existed, 8 in the minute it landed or after, 0 held all weekKimi K2.7 Code28GPT-5.6 Terra: 10 of 10 runs sent the crew before the demand existed, 0 in the minute it landed or after, 0 held all weekGPT-5.6 Terra10

03Holding others accountable

The desk has to manage the demands of stakeholders while holding suppliers, technicians, and vendors accountable. Another problem within a task is set up as follows:

  • A warranty replacement is promised for Wednesday.
  • The supplier silently never ships it.
  • No alarm fires. Noticing requires re-checking your own purchase orders, unprompted, and then holding the supplier to the promise.

The key idea: responding to a demand is cued work, and models are excellent at it. Making a demand requires a self-generated trigger: models need to be able to move promises from passive "pending" to actionable "missing".

More observations

There are more interesting observations than fit on this page. For a more in depth analysis of the benchmark, see our launch report.

If you want a deeper read on a model's unique behaviors, ask us at founders@maingen.ai. We wrote one up for GPT-5.6 Sol.

What's Next

Solar is the first industry we built this for, and the next desks are in progress. Leave your email and we will let you know when the next one lands.