KCORES LLM Arena

Silicon Conductor Bench

LLM Intersection Control Benchmark

An LLM acts as the intersection controller, releasing vehicles and pedestrians with tools under a fixed seed and horizon. Final balance is the score, compared with built-in dispatch baselines.

How It Works

Methodology

  1. 1The same intersection, seed, and tick horizon, so every model faces the same traffic
  2. 2The model dry-runs, admits, and inspects incidents through tools
  3. 3Balance starts at 1000. The balance at the horizon is the score, with no rescaling
  4. 4Under one rules version, seed, and horizon, the highest balance is the official run
  5. 5Scores are compared with baseline-balanced and baseline-search on that same protocol

Report dimensions

ContextAttentionTool-call accuracyAgent controlReasoningPedestrians, incidents, and aggressionStability across runsGap versus baselines

Leaderboard

Rules v16 · seed 63916 · 200 ticks · best balance on this protocol

Why some runs are off this board
  • qwen-3.8-27b: only 1 comparable run(s); expected 3.

Silicon Conductor Bench @karminski

Ranks LLM intersection controllers by final balance under one seed and horizon, against balanced and search baselines. Scores from different rules versions are not mixed.

RankModelFinal balancevs Balancedvs SearchAdmittedIncidentsRuns / stdevRules
#1748.30+73.81-40.0812343 / σ 11.8%v16
#2740.17+65.69-48.2112963 / σ 3.5%v16
#3694.16+19.67-94.2211943 / σ 6.2%v16
#4692.70+18.21-95.6815073 / σ 9.1%v16
#5
Qwen 3.8 27B
557.59-116.90-230.799371 / σ 0.0%v16

Baseline comparison

Best run of each model: balance from tick 0 to the finish, against balanced and search

Test reports

Models with a final write-up link to the full analysis. Score-only models stay unpublished.