toolcall.
ResearchSep 30, 2026, 08:25 UTC

EngiWorld shows frontier agents still struggle with engineering software

The benchmark covers CAD, simulation, manufacturing, BIM, electronics and 3D workflows, with the strongest tested model reaching only 44.3 points.

A new research benchmark called EngiWorld tests whether frontier AI agents can operate professional engineering software, not just solve isolated coding or browser tasks.

The benchmark includes 1,301 expert-curated tasks across CAD, CAE, CAM, BIM, electronic design automation and 3D visualization. It spans 26 professional software platforms and includes both GUI and command-line work, from choosing the right tool to completing open-ended design tasks.

The authors also built programmatic verifiers that check whether generated artifacts are geometrically valid, physically feasible and compliant with task rules. That matters because an engineering agent can appear to follow instructions while producing a design that cannot actually work.

The headline result is sobering: across seven frontier models, the strongest model reached an EngiScore of 44.3, and only 3.6% of multi-software attempts succeeded. The practical takeaway is that engineering remains a hard frontier for AI agents. Companies can use agents around professional design workflows, but the paper argues against assuming they are ready for unsupervised end-to-end engineering work.

Sources

agentsbenchmarksengineeringresearch