A living fleet, simulated — and every model benchmarked against it.

A multi-agent simulation of how a superyacht fleet really operates — 30+ specialist agents run complete charter seasons, and every decision is scored and recorded.

Agents
30+
30+ active · 100% ready
SOP library
318
100% approved
Runs
870+
scenario executions
Decisions
10,700+
agent turns recorded
Protocol msgs
4,400+
handoffs & escalations
Transactions
11,100+
map replay ledger
Vessels
50
fleet records
Benchmarks
160+
runs scored across models

One simulation, six reasons to run it.

Test the models
Benchmark every frontier and open-weight LLM on real yachting work.
Test the costing
Measure cost per season per model, so spend is known before production.
Improve agent prompting
Every run surfaces where prompts and procedures can be tightened.
Test with offline models
Run the same workload against local, offline models — LM Studio or Ollama.
Synthetic training data
Runs produce high-quality operational data to train our own models.
De-risk before production
Prove agents and processes here before anything reaches a client.

Everything the simulation produces,
in one place.

SOPs
Standard operating procedures that emerge from repeated agent behaviour.
Model benchmarks
A quality · latency · cost leaderboard across every model tested.
Live fleet map
Vessels moving through real ports as the charter season plays out.
Protocol messages
Every handoff and escalation agents exchange over the YachtingProtocol.
Agent stats
Per-agent activity, coverage and reliability across the roster.

Every model, scored on the
same charter season.

Quality against gold criteria picks the winner; cost and latency break ties.

Opus 4.7WINNER
94%SCORE
RANGE89–97%
P50 LATENCY3.1s
COST / SEASON$12.40
SUCCESS RATE98%
Kimi 2.7 Code
88%SCORE
RANGE82–92%
P50 LATENCY4.8s
COST / SEASON$2.10
SUCCESS RATE94%

Put the simulation to work
for your fleet.

Access is granted per operator. Tell us about your operation and we'll open the simulation and its benchmarks to you.