Benchmark · July 2026
The same AI, asked the same questions, on the same data.
We changed only one thing: whether the AI pulled its information through direct connectors or through BackEngine. 10 real business questions. Each answered three times by each method. Every answer graded against a verified list of the true facts, by a grader who didn't know which method wrote it. 135 AI runs in total.
BackEngine is the secure, permission-aware context layer for enterprise AI. It prepares governed, source-linked company context before an AI query, rather than asking a model to interpret raw records from direct connectors at query time. Learn how the BackEngine context layer works, see why direct MCP connectors are not enough, or review BackEngine security and permissions.
The number that matters
You only ask once. What are the odds the answer is right?
In real life, you ask a question once and act on the answer. So for every one of the 30 answers per method, we asked the simplest question there is: was the main conclusion right: the right account named, the right ranking, the full list?
Every single run
Spending more doesn't buy accuracy. Knowing where to look does.
Each dot is one of the 60 answers. Further right means more tokens spent. Higher up means more accurate. Direct-connector runs drift right and down; BackEngine runs bunch up in the cheap-and-right corner. Hover any dot for details.
Question by question
Where the gap comes from
Each question's scores, averaged across all its runs. BackEngine was more accurate on all 10 questions, found more of the facts on 9 of 10, and used fewer tokens on 9 of 10.
Of the facts each answer stated, the share that were true. BackEngine won all 10 questions.
All 60 answers, no cherry-picking
The full scoreboard, run by run
Each row is one question. Each cell is one answer. The grader never knew which method wrote which answer. And we're showing every result, including the ones we lost.
| Question | Direct connectors | BackEngine | ||||
|---|---|---|---|---|---|---|
| Q1 · Highest churn risk | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ |
| Q2 · Onboarding friction | ✗ | ✓ | ✗ | ✓ | ✓ | ~ |
| Q3 · Top competitor | ~ | ✗ | ~ | ✓ | ✓ | ✓ |
| Q4 · Capability comparisons | ✗ | ✓ | ✗ | ✓ | ~ | ~ |
| Q5 · Dormant accounts | ✗ | ✗ | ✗ | ✓ | ✓ | ✓ |
| Q6 · Top feature requests | ~ | ✗ | ✗ | ✓ | ~ | ✓ |
| Q7 · Expansion signals | ✗ | ✗ | ✗ | ✓ | ✓ | ✓ |
| Q8 · Champion loss | ✗ | ✗ | ✗ | ~ | ~ | ~ |
| Q9 · Renewal risk | ✗ | ✗ | ✗ | ✓ | ✓ | ✓ |
| Q10 · Pricing objections | ~ | ~ | ~ | ✓ | ✓ | ✓ |
| Total (30 runs each) | 3 ✓ · 6 ~ · 21 ✗ | 20 ✓ · 7 ~ · 3 ✗ | ||||
✓ the answer's main conclusion was right · ~ mostly right, with one real miss · ✗ wrong conclusion, or missed most of it.
Methodology
How we made sure this is real
A test is only as good as the rules that keep it fair. Here are ours, in order.
Both methods saw the exact same data
We froze the data first: everything measured as of July 7, 2026, looking back 90 days. Every run used that same slice of reality. Neither method could see anything the other couldn't.
10 questions, written down before we started
Real questions a sales or customer team asks every week: who might cancel, who's gone quiet, which renewals are at risk. The exact wording was locked in ahead of time and given to both methods word for word.
Three separate AIs that never talked to each other
Each part of the test ran as its own AI session, starting from a blank slate. No role ever saw another role's work.
The Direct-Connector Answerer
Answered each question using only the direct connectors. It was told to check every relevant account, work efficiently, and say when it couldn't see something instead of guessing.
The BackEngine Answerer
Answered each question using only BackEngine. It got the exact same instructions, word for word.
The Judge
Graded the answers without knowing which method wrote them. We stripped out every product name and gave the answers random labels first. The judge split each answer into its individual statements and checked each one against the verified fact list: true, false, or can't tell.
From that: accuracy = the share of statements that were true · facts found = the share of the important facts the answer included.
Every question answered three times
An AI rarely does the exact same thing twice. So each method answered every question three times. That is how we tell a real difference from a lucky run. Every run also logged its tool calls and token cost.
Every grade was checked twice
Each pair of answers was graded by two separate judges, with the labels swapped between them. If the two judges disagreed about which answer was better, the result did not count. We threw it out and re-ran it fresh. That happened 3 times in 30 pairs, and all three were near-ties that came out clean on the re-run.
The tally: 135 runs
60 answers and 72 grading passes, plus re-runs. For each method we report accuracy, facts found, tool calls, and tokens, averaged across all 30 answers, and counted question by question. With 10 questions, that shows a clear pattern, not mathematical proof, and we say so.
At your scale
What this adds up to over a year
More wrong facts with direct connectors
Based on each method's measured error rate (23.2% vs 7.6%). That's about 4.7 extra wrong facts every working day.
More AI spend with direct connectors
At Claude Sonnet pricing ($3 per million input tokens, $15 per million output), using each method's average tokens per question (137.3K vs 48.3K) and assuming 9 of every 10 tokens are input. A directional estimate, not a billing quote.
Run the test on your own data.
Same questions, your own tools and data. We'll show you the gap.
Demo →