How improvement works
The product separates the design agent from the evaluator:
sequenceDiagram
participant User
participant App as Local web / CLI
participant Worker
participant Model
participant CAD as Isolated CAD containers
participant Store as SQLite or Atlas
User->>App: YAML task or continuation
App->>Store: Frozen configuration, task, seed, image digest
Worker->>CAD: Build and independently evaluate baseline
CAD->>Store: STEP, GLB, metrics, violations
loop New iterations within budget
Worker->>Store: Retrieve compatible successes, failures, lessons and tools
Worker->>Model: Task, constraints, evidence and tested outputs
Model->>Store: Checkpoint proposal response
Worker->>CAD: Build, then independently evaluate exported STEP
CAD->>Store: Geometry, measurements and failures
Worker->>Model: Reflect on results
Model->>Store: Versioned lesson and next focus
Worker->>CAD: Test or invoke reusable numeric utility
Worker->>Store: Record tool evidence and progress event
Store-->>App: Live progress
end
Reflection changes the next proposal's context. Retrieval keeps relevant prior failures visible. Tested tools persist and can be reused across compatible runs. These mechanisms improve the search process; they do not retrain model weights or guarantee that each proposal improves performance.
The MVP's generated utility reports objective changes from measured results. It is independently tested and invoked, but never supplies the trusted physical score. Custom task evaluators can incorporate specialized solvers; the installed app and evaluator are not rewritten by the model.
Local mode uses SQLite documents, local content-addressed artifacts, and recent/lexical memory. Atlas mode uses the same records with GridFS and optional vector retrieval. Missing vector indexes trigger an explicit fallback event. Database Triggers are not needed for the product worker.
Run snapshots preserve parameters, source, task hash, runtime image digest, evaluation, reflection, tool versions, and API response usage. Git archives source snapshots. Checkpoints prevent completed model requests from being submitted again after a restart.
One managed worker executes one run at a time. The browser receives events and polls for state reconciliation. The server is localhost-only and rejects unexpected hosts and cross-origin mutation requests. Docker containers have no network, no credentials, limited resources, and no write access to the evaluator source.
Limits
Sensor evaluation uses a conservative wall-strip screen. The gripper uses a linear frame model and sampled travel checks. VTOL uses the established coupled aerodynamic/energy screening model; the live template does not automatically repeat the historical campaign's finalist convergence study. The UI labels estimates and never claims flight, fatigue, manufacturing, or certification validation.
The beta CST/spline experiment remains separate, unlisted, and unavailable as a product template. Archived campaigns remain unchanged.
Model generation follows the Responses API and validates the returned JSON before use. The current adapter uses JSON object output with local validation; see OpenAI structured output documentation for the distinction from strict JSON Schema outputs.