An independent evaluation by Caliper Lab found that Binocs examined roughly ten times more sources and three to eight times more of the market than leading frontier AI models on the same diligence research. This piece walks through what was tested, where Binocs led, where it tied, and what that means for anyone using AI in commercial due diligence.
The question deal teams should be asking
Most of the public conversation about AI in finance is a horse race between models. Is GPT ahead this quarter? Has Claude caught up? Which one writes the cleaner memo? For a deal team, that question matters less than it looks. The model leaderboard changes every few months. What does not change is when a partner points at one number on a slide and asks where it came from. That moment is the real test of any research tool. A well-written market summary built on a handful of sources reads beautifully until confirmatory diligence, a lender, or an LP asks it to hold up. What matters is how much of the market the system actually looked at, whether the figures can be traced, and whether they can go straight into a model.
Why an independent evaluation
Every AI vendor has a benchmark it wins. Buyers have no reliable way to tell which of those numbers mean anything, because in almost every case the vendor designed the test, ran it, and graded it. Enterprise software solved this problem years ago with independent analysts. AI applications have not had an equivalent.
Caliper Lab is built to fill that gap. It is an independent AI capability measurement firm that evaluates AI products on real professional workflows rather than academic benchmarks, publishes regardless of commercial relationships, and applies the same standard to every system it tests.
Caliper Lab independently evaluated Binocs against three frontier systems: Claude Opus 4.8, GPT-5.5, and GPT Deep Research. We did not run the test and we did not grade it.
Methodology: Each system produced six research reports. Caliper's pipeline then extracted every quantified claim from those outputs, 15,383 values in total, and tagged each one with the metric it measures, the entity it describes, its unit, its time period, its source, and whether it was stated firmly or hedged. Over 5,000 citations were probed to check whether the links actually resolve. The same automated pipeline ran across all four systems.
One design choice is worth calling out. Binocs produces structured slide output, while the other three produce prose, and tables naturally state firm numbers. Caliper ran the key commitment comparison on prose-extracted values only.
Finding 1: Coverage
Binocs cited 7,722 URLs across 833 distinct source domains, with 84% of those links still reachable when probed. The frontier models drew on between 64 and 96 domains. That is roughly a tenfold difference in how much of the available evidence each system consulted before reaching a conclusion.
Think about what a standalone model is doing when it answers a diligence prompt. It finds a handful of sources, often the most prominent ones, and builds the whole narrative around them. The result can read as authoritative while resting on what amounts to two or three analyst reports.
Volume alone would not settle this, though. A system can cite a thousand pages and still have its conclusion resting on two of them. So more imprtant here is concentration. Caliper measured how evenly each system's citations were spread using a Herfindahl index(HHI), the same measure used to check whether a market is dominated by one player. Zero means perfectly spread; one means everything traces back to a single source.
Binocs scored 0.012, and it held that level across more than 7,700 citations. The frontier models, working from a fraction of the sources, were two to three times more concentrated.
In diligence terms, this is about single points of failure. A conclusion built on a few sources falls apart the moment one of them is wrong, outdated, or challenged. A conclusion spread across hundreds of independent sources has no single report holding up the whole structure, and it leaves a trail an associate can follow when someone asks.
Finding 2: Depth, how much of the market was mapped
Binocs examined 1,062 distinct entities across its reports: competitors, customers, suppliers, products, and adjacent players. GPT Deep Research covered 344, GPT-5.5 covered 202, and Claude Opus 4.8 covered 131.
The distribution matters as much as the count. Caliper found that the frontier models tended to fixate on one market leader. Claude placed 16% of all its quantified analysis on a single company. Binocs most-covered entity accounted for 6% of its values, which produces a far more balanced picture of the competitive field.
Anyone who has worked on commercial diligence understands why this has consequences. Markets are rarely shaped by one company. They are shaped by challengers, customers, suppliers, substitutes, regulators, and technologies that are only starting to matter. When analysis concentrates on the obvious incumbent, the blind spots form exactly where the interesting findings tend to sit: a smaller competitor taking share, a niche segment growing faster than the headline market, or a shift in the supply chain.
The risk in diligence is that the team reaches a confident answer from an incomplete view of the market. Before judging any AI tool's conclusion, it is worth asking how much of the landscape it actually explored.
Finding 3: Figures a deal team can actually use
The difference between useful and unusable research output is often mundane. Can the number go straight into the model, or does someone have to work out what it means first?
Caliper looked at this in two ways. The first was commitment: whether a system states a firm figure or hedges with a range or an approximation. On the prose-only comparison, Binocs stated firm point estimates 82% of the time. GPT Deep Research followed at 72%, GPT-5.5 at 59%, and Claude Opus 4.8 at 34%, with the remaining two thirds of Claude's figures given as ranges or approximations.
A point estimate such as a market size of $4.2bn in 2024 can be dropped into a model. A range of $3bn to $5bn forces the analyst to pick an assumption. An approximation reduces accountability for the figure altogether. Because this comparison excluded Binocs tables, the gap reflects how the system commits to numbers, not the format it writes in.
The second test was units. Only 4% of Binocs values were missing a unit. Both GPT systems left close to 30% of their figures without unit, meaning nearly one in three needs surrounding context before it can be verified or compared. Claude was cleaner at 6%, but drew on just 30 unit types.
Binocs used 282 distinct unit types. That range is a sign of sector-specific analysis. In the satellite market Caliper used for the evaluation, that meant measuring orbital altitude, image resolution, revisit times, constellation sizes, and detection accuracy alongside the usual revenue, growth, and share figures. Standard financial vocabulary covers a summary. diligence needs the operating metrics of the industry being underwritten.
What we are still building
On internal consistency, Binocs and GPT Deep Research finished level. Caliper checked every case where the same metric for the same entity and period appeared more than once in a report, and measured how often the figures agreed. Binocs scored 66.7% and GPT Deep Research 66.0%, well ahead of GPT-5.5 at 34.1% and Claude Opus 4.8 at 20.5%. We read this as a real strength of agentic, multi source research, and one that the best frontier configuration shares. It is not a unique edge.
The report also showed that Binocs is weighted toward historical and current data. About 8% of its quantified values were forward-looking, against 28% for GPT Deep Research. That fits what Binocs is built to do, which is establish a verified factual baseline, but deal teams also need projections with stated assumptions. A forward-looking synthesis layer built on top of that baseline is in progress, and Caliper flagged inline citation depth as another area where we can tighten things up.
We would rather be measured by someone with no stake in the result and told plainly where the gaps are than win a test we designed for ourselves.
What this means for deal teams
Whichever tool your team uses, the Caliper framework gives you a practical set of questions to ask before trusting AI-generated research in a live process.
- How many distinct sources did it draw on, and is any single one carrying the conclusion?
- How much of the market did it map beyond the obvious leader?
- Can each material figure be traced to a source that still resolves?
- Are the numbers stated firmly, with units and periods, so they can go into a model as they stand?
- Does the report agree with itself when the same figure appears twice?
These are the properties that decide whether research survives Investment Committee, confirmatory diligence, and the questions that follow a close. On the first, second, and fourth, the independent evidence puts Binocs clearly ahead of frontier models answering the same prompt. On the fifth, it matches the strongest of them.
Binocs is built to produce the verified factual baseline an investment thesis stands on. The thesis itself remains the deal team's judgement, and it should. Our job is to make sure the foundation under that judgement holds when someone tries to break it.





