01
Dataset
Synthetic and appropriately licensed public drawings, versioned and separated from model development.
Protocol published · Results forthcoming
A reproducible evaluation protocol for AI BOQ and quantity takeoff from DWG, IFC, DXF and XREF-based drawings. Deepsa AI currently reports more than 96% accuracy with DXF inputs; the benchmark below is designed to test that claim against frozen, human-validated ground truth.
Benchmark your drawing set01
Synthetic and appropriately licensed public drawings, versioned and separated from model development.
02
Human-validated quantities, item classifications, units, exclusions and reviewer agreement.
03
Frozen Deepsa AI model and configuration run without post-result correction.
04
Predicted versus ground truth, disclosed by drawing format, discipline and difficulty.
| Metric | Definition | Reporting rule |
|---|---|---|
| Accuracy | Correctly classified or matched takeoff decisions divided by all evaluated decisions. | Reported with sample count and confidence interval. |
| Precision | Correct detected BOQ items divided by all BOQ items predicted by the system. | Measures the cost of false additions. |
| Recall | Correct detected BOQ items divided by all eligible items in human-validated ground truth. | Measures the cost of missed scope. |
| MAE | Mean absolute difference between predicted and validated quantities in matched units. | Published by format, discipline and quantity unit. |
| MAPE | Mean absolute percentage error for ground-truth quantities above a declared minimum denominator. | Zero and near-zero denominators are excluded and counted separately. |
| Item-level accuracy | Exact or tolerance-bound agreement on BOQ item identity, classification and inclusion. | Shows whether the correct scope was found. |
| Quantity-level accuracy | Weighted agreement between predicted and human-validated quantities for matched items. | Shows how closely measured quantities agree after item matching. |
Synthetic evaluation plan · Not observed customer data
A larger synthetic test bed is planned to stress construction intelligence systems at enterprise scale. Every number below is a generation target, not a count of customer projects, production transactions or completed benchmark observations.
A frozen evaluation set, human-validated quantities, disclosed matching rules, reproducible calculations and visible failure cases. Proven dataset results showcase more than 96%+ accuracy when the DXF format is used specifically, describing item recognition and quantity takeoff with the highest accuracy measures and proven results.