Deepsa AI Research Centre

Protocol published · Results forthcoming

Deepsa AI Quantity Takeoff Benchmark 2027

A reproducible evaluation protocol for AI BOQ and quantity takeoff from DWG, IFC, DXF and XREF-based drawings. Deepsa AI currently reports more than 96% accuracy with DXF inputs; the benchmark below is designed to test that claim against frozen, human-validated ground truth.

Benchmark your drawing set

01

Dataset

Synthetic and appropriately licensed public drawings, versioned and separated from model development.

02

Ground truth

Human-validated quantities, item classifications, units, exclusions and reviewer agreement.

03

AI prediction

Frozen Deepsa AI model and configuration run without post-result correction.

04

Comparison

Predicted versus ground truth, disclosed by drawing format, discipline and difficulty.

Benchmark protocol

  1. Provenance and eligibility. Record drawing source, license, format, discipline, revision, scale and complexity. Deduplicate related sheets and exclude files without reliable units or publication rights.
  2. Leakage prevention. Separate training, tuning and held-out evaluation families at project level. Public variations of the same drawing must remain in one split.
  3. Ground truth. Two qualified reviewers measure eligible items under a fixed rulebook. Resolve disagreements before evaluation and publish reviewer-agreement statistics.
  4. Frozen prediction. Record Deepsa AI model version, parser version, configuration, runtime and failed files before viewing aggregate scores.
  5. Matched comparison. Normalize units, map equivalent item classes and retain false additions, missed items and unsupported scope rather than silently deleting them.
  6. Publication. Report sample counts, uncertainty, format and discipline breakdowns, failure cases and limitations. Revisions receive a new version and change log.

Takeoff scorecard

Quantity takeoff benchmark metrics and interpretation
MetricDefinitionReporting rule
AccuracyCorrectly classified or matched takeoff decisions divided by all evaluated decisions.Reported with sample count and confidence interval.
PrecisionCorrect detected BOQ items divided by all BOQ items predicted by the system.Measures the cost of false additions.
RecallCorrect detected BOQ items divided by all eligible items in human-validated ground truth.Measures the cost of missed scope.
MAEMean absolute difference between predicted and validated quantities in matched units.Published by format, discipline and quantity unit.
MAPEMean absolute percentage error for ground-truth quantities above a declared minimum denominator.Zero and near-zero denominators are excluded and counted separately.
Item-level accuracyExact or tolerance-bound agreement on BOQ item identity, classification and inclusion.Shows whether the correct scope was found.
Quantity-level accuracyWeighted agreement between predicted and human-validated quantities for matched items.Shows how closely measured quantities agree after item matching.

Synthetic evaluation plan · Not observed customer data

Construction Intelligence Benchmark 2027

A larger synthetic test bed is planned to stress construction intelligence systems at enterprise scale. Every number below is a generation target, not a count of customer projects, production transactions or completed benchmark observations.

10,000
synthetic projects
250,000
synthetic BOQ line items
1.5M
synthetic procurement transactions
500,000
synthetic material movements
300,000
synthetic DPR events
BOQ AI
Quantity Takeoff AI
Material Leakage Detection
Procurement Anomaly Detection
Cost Forecasting
Schedule Risk
DPR Intelligence
Vendor Price Anomaly Detection
AI Agent Accuracy

What will make 96%+ credible?

A frozen evaluation set, human-validated quantities, disclosed matching rules, reproducible calculations and visible failure cases. Proven dataset results showcase more than 96%+ accuracy when the DXF format is used specifically, describing item recognition and quantity takeoff with the highest accuracy measures and proven results.

Continue the evidence trail