Currence's AI engine pulls in raw energy data from sources across the internet. But AI is famously unreliable, especially when it comes to high stakes investment or strategy decisions.
So we built an evals process. It's a rigorous, repeatable system that checks every value before it ships to the platform.
My last post covered how the engine builds that data in the first place, turning messy inputs like a spec sheet or news article into structured project records. This post covers what happens after: how we prove those values are right.
The problem with generic AI output
Generic large language models pull incorrect values and make unchecked assumptions. You can't know whether AI output is right until you've measured it against something you trust. In energy, the cost of a bad data point is particularlyhigh. Lenders and developers making billion-dollar infrastructure decisions need numbers they can rely on.
The other complication is that accuracy means something different for every field. For data centers, capacity within 5% is reliable. But for power plants, coordinates need to be accurate to the specific street or lot, because that's what determines which substation and transmission lines are nearby, and whether satellite imagery is useful.
Treating every field the same way produces a system that is mediocre across the board.
How evals work
Evals are human-led evaluation harnesses. You measure every field against a human-verified answer across a sample, iterate until accuracy clears a defined threshold, and only ship data you trust. For us that breaks into three steps:
Sample. We start by pulling a representative sample from our dataset. Six years of accumulated data gives us a strong base to draw from across different stages of data center development, geographies, and project sizes. A representative sample matters because a model tuned too heavily on one type of data won't hold up across the rest. We use tools like OxenAI to version and store our datasets, which makes it straightforward to track how samples update and evolve over time.
Human QA. That sample goes to our human QA team. Our expert researchers confirm that every field in the sample is accurate. This is the golden sample, the ground truth that all subsequent iterations are measured against. Getting this step right is everything.
Enrich and compare. We run the same sample through the enrichment agent, generate values for every field, and compare them against the human-verified answers. We use tools like Langfuse to track iterations — different prompts, different models, different tools — until 95% of rows match. That threshold is the bar before anything moves forward.
Calibration: from eval to production system
Hitting 95% accuracy in an eval is not enough on its own. The next step is calibration, which is what takes evals from a one-time test into a production system.
We tune the agent both to get the right value and to return its confidence in that value. Confidence scoring is a mix: an LLM-as-a-judge model assesses how well the answer is supported by the source data, and human-curated rules (such as source trustworthiness or expected value ranges) flag specific warning signs that should bump confidence down.
The goal is to make sure low confidence lines up with low accuracy. Once that's calibrated, we set a cutoff. Above the cutoff, we ship the data directly and keep it as current as possible. Below the cutoff, the data routes to a human researcher first for review. It's a slower path, but it means nothing goes on the platform that we're not confident in.
Three tiers of data
Not all data on the Currence platform carries the same level of scrutiny, and we're explicit about that.
Tier 1 is our gold standard. These are researcher-generated records with deep knowledge and original insights attached.
Tier 2 is human-in-the-loop data. It's agent-enriched or client-submitted, but every record either exceeds our confidence threshold or carries a human-stamped approval.
Tier 3 is AI-generated at scale. Broad coverage, no spot checks yet. We recently processed over 40,000 US mature power projects from the Energy Information Administration (EIA), and that dataset is sitting in Tier 3 while it works through the eval process. It's available for further analysis now, and it gets promoted to Tier 2 once it clears quality assurance.
Why this matters
Energy market intelligence is only as good as the data behind it. Building a research process from our analysts and proving it with evals is what makes our data comprehensive, consistent, and traceable. For teams financing and building the energy system, that's the foundation everything else runs on.
More than 90 teams already run on the Currence engine, including Microsoft, bp, Baker Hughes, Southern Company, HSBC, BBVA, Siemens Energy, Shell, BHP, B Capital, Galvanize, and Mitsui.
If you're interested in learning more about our coverage, request a 20-minute call with our team.



.png)