← Object gallery

Da Vinci Recursive Improvement CAD Harness

Drone sensor mount · 8 GPT-6 Astra iterations

PA12 (nominal) · 20 N load
Loading CAD…

Self-improvement methods

The agent improves its tools and working context between CAD attempts. Model weights stay fixed.

1. Reflect and remember

Retrieve prior designs and measurements; carry lessons into the next attempt.

2. Create and reuse tools

Write, test and save a wall-thickness utility for subsequent designs.

3. Evaluate and revise

Use independent CAD checks to guide revisions; retain the lightest passing design.

Research evidence

Reflection and memory. Closing the Consistency Gap · 8 Sep 2026 ↗ reported +16 percentage points in AppWorld tasks succeeding on all five runs, using stored diagnostic guidelines with ReAct/GPT-4.1; +13 points on similar tasks.

Reusable skills. SkillAlchemy · 24 Aug 2026 ↗ reported +19.9 percentage points pass rate over execution without skills across 87 SkillsBench tasks, by creating reusable skill packages from source material.

Evaluation and revision. AIDE² · 22 Sep 2026 ↗ found 7 successive agent improvements in 8 days by testing changes to its own code. Gains transferred to four held-out benchmarks.

Recent preprints; results are from other tasks, not validation of this CAD harness or measurements of each method’s contribution here.

MongoDB Atlas

Documents store designs, evaluations, tools and policies. Vector Search retrieves prior results. GridFS stores CAD files and source snapshots.

Agent harness setup

This sensor-mount study

  1. Set the design contract. A Python runner fixes the mounting footprint, PA12 material assumptions, 20 N load and acceptance limits. Astra writes CadQuery wrappers and chooses dimensions within the supplied cradle family.
  2. Build and measure. Candidate code runs in an isolated Docker container and exports STEP and GLB. A separate evaluator checks the STEP against the supported geometry family, measures volume and applies the same stress and deflection screen to every attempt.
  3. Reflect and reuse. Previous results, scoped Atlas Vector Search matches and accumulated lessons enter the next prompt. After two attempts, Astra created a wall-thickness utility; five validation cases passed before the saved tool was reused.
  4. Archive and select. Git versions source, tools and context. Atlas stores records and GridFS stores artifacts. The runner retains the lightest passing design; this page displays all eight attempts from an exported snapshot without making new model calls.

Full multi-agent workbench

The full harness coordinates structural and aerodynamic specialists against a shared vehicle specification. Two Python workers process durable jobs. Atlas candidate-insert triggers enqueue evaluation; evaluation-insert triggers enqueue reflection, with a recovery loop for missed delivery.

A meta-agent proposes reusable tools and versioned changes to orchestration logic, context policy and an isolated React policy panel. Independent tool tests and release checks run before activation. The evaluator and acceptance limits remain fixed outside those editable files.

The sensor study shown here uses a sequential Python loop with the same sandbox, archive and memory services; it does not use the full workbench’s two-specialist trigger queue.

Measurement scope

Mass uses STEP volume and nominal PA12 density. Stress and deflection are wall-strip estimates, not FEA. The VTOL is illustrative; measurements cover the mount only.

All 8 design iterations

Chronological order · best passing design highlighted
BEST PASSING MOUNT83.5 → 30.3 g63.7% mass reduction

Same checks for every design
Deflection ≤ 0.65 mm
Stress ≤ 28 MPa

01

Solid cradle

Baseline
Loading CAD…
Drag to rotate · Scroll to zoom
Iteration 01 · Solid cradle

6 mm solid base and 6 mm side walls.

Mass
83.5 g
vs. baseline
0.0%
Deflection est.
0.053 mm
Stress est.
0.66 MPa
Initial design explorationSTEP
02

Open-window cradle

Improved
Loading CAD…
Drag to rotate · Scroll to zoom
Iteration 02 · Open-window cradle

3.350 mm walls, oval windows and one base slot.

Mass
30.4 g
vs. baseline
−63.7%
Deflection est.
0.647 mm
Stress est.
4.51 MPa
Initial design explorationSTEP
03

Open-window refinement

Improved
Loading CAD…
Drag to rotate · Scroll to zoom
Iteration 03 · Open-window refinement

3.344 mm walls, oval windows and one base slot.

Mass
30.3 g
vs. baseline
−63.7%
Deflection est.
0.650 mm
Stress est.
4.52 MPa
Saved thickness tool usedSTEP
04

Open-window refinement

Passed
Loading CAD…
Drag to rotate · Scroll to zoom
Iteration 04 · Open-window refinement

3.344 mm walls, oval windows and one base slot.

Mass
30.3 g
vs. baseline
−63.7%
Deflection est.
0.650 mm
Stress est.
4.52 MPa
Saved thickness tool usedSTEP
05

Open-window refinement

Passed
Loading CAD…
Drag to rotate · Scroll to zoom
Iteration 05 · Open-window refinement

3.344 mm walls, oval windows and one base slot.

Mass
30.3 g
vs. baseline
−63.7%
Deflection est.
0.650 mm
Stress est.
4.52 MPa
Saved thickness tool usedSTEP
06

Open-window refinement

Best
Loading CAD…
Drag to rotate · Scroll to zoom
Iteration 06 · Open-window refinement

3.344 mm walls, oval windows and one base slot.

Mass
30.3 g
vs. baseline
−63.7%
Deflection est.
0.650 mm
Stress est.
4.52 MPa
Saved thickness tool usedSTEP
07

Ribbed cradle

Passed
Loading CAD…
Drag to rotate · Scroll to zoom
Iteration 07 · Ribbed cradle

Diagonal ribs, 3 base slots and 12 mm gussets.

Mass
31.7 g
vs. baseline
−62.1%
Deflection est.
0.650 mm
Stress est.
4.49 MPa
Saved thickness tool usedSTEP
08

Ribbed cradle

Passed
Loading CAD…
Drag to rotate · Scroll to zoom
Iteration 08 · Ribbed cradle

Diagonal ribs, 2 base slots and 6 mm gussets.

Mass
31.2 g
vs. baseline
−62.6%
Deflection est.
0.650 mm
Stress est.
4.49 MPa
Saved thickness tool usedSTEP
100 × 72 mm mounting footprint · 80 × 52 mm bolt patternFull harness ↗