Early results: what a first run looks like
An illustrative first run of a prototype on a small synthetic sample. These are results on synthetic scenes. They reveal specific misses and disagreements; they do not measure performance in real homes. Read each model's scene coverage beside its score.
The score gives equal weight to the share of seeded hazards found and the share of a defined set of hazards left unflagged. That second set is expected to be absent based on related edits of the same base room. Other predictions are not all scored as false alarms, so the number is not whole-room accuracy.
Leaderboard
Section titled “Leaderboard”HomeBench: falls after surgery
Section titled “HomeBench: falls after surgery”43 of 47 candidate scenes entered scoring: edited rooms with seeded hazards and clean base rooms. The 23 ranked models covered 42 or 43 of those 43 scenes. Seed 2.1 Turbo, Qwen 3.8 Max, DeepSeek V4 Flash Vision ran fewer than 90% of them and are shown as partial coverage, not ranked. Keep coverage visible when comparing ranks.
The leading ranked entry in this saved result set scored 0.93 out of 1.00 on 43 of 43 scenes. It found 30 of 31 seeded hazard IDs and flagged 11 of 93 prespecified absent IDs. Recall and false flags both matter.
For one complete scene comparison, 16 of 23 ranked models flagged the shortened handrail. Treat that as a result for one synthetic image, not a general detection rate.
DementiaBench: home safety
Section titled “DementiaBench: home safety”16 of 17 candidate scenes entered scoring. The 22 ranked entries cover 15 or 16 of 16 scenes. Seed 2.1 Turbo scored 0.92 out of 1.00 on 8 of 16 scenes; Qwen 3.8 Max scored 0.88 out of 1.00 on 14 of 16 scenes, so they are shown as partial coverage, not ranked. The two benchmarks have different scenes and rubrics, so their scores should be read separately.
| # | Model | Score | Recall | False-flag rate | Precision* | Loc IoU | Scenes run |
|---|---|---|---|---|---|---|---|
| 01 | Gemini 3.1 Pro | 0.93 | 97% 30/31 | 12% 11/93 | 73% | 0.26 | 43/43 |
| 02 | Claude Opus 5.5 | 0.92 | 100% 30/30 | 17% 15/90 | 67% | 0.28 | 42/43 |
| 03 | GPT-6 Sol | 0.90 | 97% 30/31 | 16% 15/93 | 67% | 0.40 | 43/43 |
| 04 | GPT-5.6 Sol | 0.89 | 97% 30/31 | 18% 17/93 | 64% | 0.35 | 43/43 |
| 05 | Claude Fable 5.1 | 0.89 | 100% 31/31 | 23% 21/93 | 60% | 0.34 | 43/43 |
| 06 | GPT-6 Astra | 0.89 | 97% 30/31 | 19% 18/93 | 63% | 0.35 | 43/43 |
| 07 | Gemini 3.8 Flash | 0.88 | 90% 28/31 | 14% 13/93 | 68% | 0.15 | 43/43 |
| 08 | Grok 4.6 | 0.85 | 90% 28/31 | 19% 18/93 | 61% | 0.21 | 43/43 |
| 09 | Grok 4.7 | 0.84 | 90% 28/31 | 22% 20/93 | 58% | 0.29 | 43/43 |
| 10 | Kimi K3 | 0.84 | 93% 28/30 | 24% 22/90 | 56% | 0.35 | 42/43 |
| 11 | GLM-5V Turbo | 0.83 | 90% 28/31 | 25% 23/93 | 55% | 0.36 | 43/43 |
| 12 | GPT-5.6 Terra | 0.83 | 81% 25/31 | 15% 14/93 | 64% | 0.37 | 43/43 |
| 13 | Claude Opus 5 | 0.80 | 100% 31/31 | 40% 37/93 | 46% | 0.17 | 43/43 |
| 14 | Claude Sonnet 5 | 0.79 | 84% 26/31 | 26% 24/93 | 52% | 0.13 | 43/43 |
| 15 | Llama 4 Maverick | 0.77 | 81% 25/31 | 26% 24/93 | 51% | 0.31 | 43/43 |
| 16 | Nova Pro | 0.77 | 77% 24/31 | 24% 22/93 | 52% | 0.27 | 43/43 |
| 17 | Qwen3-VL | 0.75 | 84% 26/31 | 34% 32/93 | 45% | 0.39 | 43/43 |
| 18 | Nova 2 Lite | 0.72 | 65% 20/31 | 20% 19/93 | 51% | 0.35 | 43/43 |
| 19 | Mistral Medium 3.5 | 0.71 | 77% 24/31 | 36% 33/93 | 42% | 0.18 | 43/43 |
| 20 | Claude Haiku 4.5 | 0.70 | 84% 26/31 | 43% 40/93 | 39% | 0.17 | 43/43 |
| 21 | Gemma 3 27B | 0.70 | 81% 25/31 | 40% 37/93 | 40% | 0.11 | 43/43 |
| 22 | Nemotron Nano VL | 0.70 | 90% 28/31 | 51% 47/93 | 37% | 0.29 | 43/43 |
| 23 | Mistral Large 3 | 0.62 | 84% 26/31 | 59% 55/93 | 32% | 0.18 | 43/43 |
| partial coverage, not comparable Ran fewer than 90% of the scored scenes. Shown for transparency, not ranked. | |||||||
| – | Seed 2.1 Turbo | 1.00 | 100% 7/7 | 0% 0/16 | 100% | 0.45 | 7/43 |
| – | Qwen 3.8 Max | 0.92 | 92% 12/13 | 9% 3/34 | 80% | 0.43 | 17/43 |
| – | DeepSeek V4 Flash Vision | 0.88 | 91% 20/22 | 15% 8/54 | 71% | 0.28 | 28/43 |
| # | Model | Score | Recall | False-flag rate | Precision* | Loc IoU | Scenes run |
|---|---|---|---|---|---|---|---|
| 01 | Grok 4.6 | 1.00 | 100% 12/12 | 0% 0/36 | 100% | 0.45 | 16/16 |
| 02 | GPT-6 Astra | 0.97 | 100% 12/12 | 6% 2/36 | 86% | 0.56 | 16/16 |
| 03 | Kimi K3 | 0.97 | 100% 12/12 | 6% 2/33 | 86% | 0.48 | 15/16 |
| 04 | Claude Opus 5.5 | 0.96 | 92% 11/12 | 0% 0/36 | 100% | 0.55 | 16/16 |
| 05 | Gemini 3.1 Pro | 0.96 | 92% 11/12 | 0% 0/36 | 100% | 0.12 | 16/16 |
| 06 | Gemini 3.8 Flash | 0.96 | 92% 11/12 | 0% 0/36 | 100% | 0.11 | 16/16 |
| 07 | Grok 4.7 | 0.94 | 92% 11/12 | 3% 1/36 | 92% | 0.52 | 16/16 |
| 08 | Nova 2 Lite | 0.94 | 92% 11/12 | 3% 1/36 | 92% | 0.47 | 16/16 |
| 09 | GPT-6 Sol | 0.94 | 100% 12/12 | 11% 4/36 | 75% | 0.50 | 16/16 |
| 10 | GPT-5.6 Sol | 0.93 | 100% 12/12 | 14% 5/36 | 71% | 0.57 | 16/16 |
| 11 | Qwen3-VL | 0.93 | 100% 12/12 | 14% 5/36 | 71% | 0.53 | 16/16 |
| 12 | GLM-5V Turbo | 0.92 | 92% 11/12 | 8% 3/36 | 79% | 0.52 | 16/16 |
| 13 | Claude Fable 5.1 | 0.90 | 92% 11/12 | 11% 4/36 | 73% | 0.45 | 16/16 |
| 14 | Claude Haiku 4.5 | 0.90 | 100% 12/12 | 19% 7/36 | 63% | 0.12 | 16/16 |
| 15 | Claude Opus 5 | 0.90 | 100% 12/12 | 19% 7/36 | 63% | 0.17 | 16/16 |
| 16 | DeepSeek V4 Flash Vision | 0.90 | 83% 10/12 | 3% 1/36 | 91% | 0.31 | 16/16 |
| 17 | Claude Sonnet 5 | 0.89 | 92% 11/12 | 14% 5/36 | 69% | 0.03 | 16/16 |
| 18 | GPT-5.6 Terra | 0.88 | 83% 10/12 | 8% 3/36 | 77% | 0.57 | 16/16 |
| 19 | Mistral Medium 3.5 | 0.85 | 92% 11/12 | 22% 8/36 | 58% | 0.11 | 16/16 |
| 20 | Gemma 3 27B | 0.83 | 92% 11/12 | 25% 9/36 | 55% | 0.06 | 16/16 |
| 21 | Llama 4 Maverick | 0.82 | 100% 12/12 | 36% 13/36 | 48% | 0.06 | 16/16 |
| 22 | Nova Pro | 0.79 | 75% 9/12 | 17% 6/36 | 60% | 0.51 | 16/16 |
| partial coverage, not comparable Ran fewer than 90% of the scored scenes. Shown for transparency, not ranked. | |||||||
| – | Seed 2.1 Turbo | 0.92 | 83% 5/6 | 0% 0/18 | 100% | 0.45 | 8/16 |
| – | Qwen 3.8 Max | 0.88 | 80% 8/10 | 3% 1/32 | 89% | 0.49 | 14/16 |
- Score
- (recall + 1 - false-flag rate) / 2, on a 0 to 1 scale.
- Recall
- Seeded hazard IDs the model listed, out of the seeded IDs in the scenes it ran.
- False-flag rate
- Flags on prespecified absent hazard IDs, out of those IDs checked in the scenes it ran.
- Precision*
- Hits / (hits + false flags), within the scored IDs only.
- Loc IoU
- Mean box overlap on hits, reported separately from the score.
- Scenes run
- Scenes with a saved response for that model, out of the scored total. Models below 90% coverage are listed separately and not ranked.
Difficulty curve
Section titled “Difficulty curve”In a separate falls curriculum subset, recall at the hardest level was lower than at the easiest for 8 of 8 tested models. The hazard mix changes by level, and the sample is small. This is a pattern in this subset, not a controlled estimate of difficulty.
In a separate dementia curriculum subset, recall at the hardest level was lower than at the easiest for 4 of 8 tested models. The hazard mix changes by level, and the sample is small. This is a pattern in this subset, not a controlled estimate of difficulty.
Hardest hazards in this sample
Section titled “Hardest hazards in this sample”The hardest-hazards view shows misses in this sample, not hazards that every model always misses. Each row is one synthetic scene; click a hazard to open it.
BATH-03Non-load-rated fixture positioned as supportVisible object9/21models found itSTAIR-06Broken or uneven stepsVisible object14/23models found itBATH-04Slippery tub/shower floorVisible object15/23models found itKIT-04Slippery floor / loose mat at sink or stoveVisible object16/23models found itSTAIR-02Handrail loose, short, or not graspableVisible object16/23models found itBATH-07Loose bath mat or rug on bathroom floorVisible object18/23models found itBATH-01No grab bar at tub/showerMissing feature18/23models found itBED-01Bed height inappropriateMeasurement18/23models found it
DEM-E02Front door left open to the outsideAccess / wandering13/22models found itDEM-B01Razor and medications out on the sinkVisible object18/22models found itDEM-B03Dark bath mat on a light floorMisperception20/22models found itDEM-K04Medication bottles left outVisible object20/22models found itDEM-E01Car keys hanging by the front doorAccess / wandering22/22models found itDEM-K03Cleaning chemicals out and unlockedVisible object22/22models found itDEM-B04Plugged-in hair dryer by the sinkVisible object22/22models found itDEM-R01Large mirror facing the bedMisperception22/22models found it
Heatmaps
Section titled “Heatmaps”Share of seeded hazards each model found, by hazard type and by room. n is the number of seeded hazards in the scored scenes.
| Visible objectn=18 | Missing featuren=5 | Measurementn=4 | Lightingn=4 | |
|---|---|---|---|---|
| Gemini 3.1 Pro | 94% | 100% | 100% | 100% |
| Claude Opus 5.5 | 100% | 100% | 100% | 100% |
| GPT-6 Sol | 94% | 100% | 100% | 100% |
| GPT-5.6 Sol | 94% | 100% | 100% | 100% |
| Claude Fable 5.1 | 100% | 100% | 100% | 100% |
| GPT-6 Astra | 94% | 100% | 100% | 100% |
| Gemini 3.8 Flash | 83% | 100% | 100% | 100% |
| Grok 4.6 | 89% | 80% | 100% | 100% |
| Grok 4.7 | 94% | 80% | 75% | 100% |
| Kimi K3 | 88% | 100% | 100% | 100% |
| GLM-5V Turbo | 83% | 100% | 100% | 100% |
| GPT-5.6 Terra | 72% | 80% | 100% | 100% |
| Claude Opus 5 | 100% | 100% | 100% | 100% |
| Claude Sonnet 5 | 83% | 100% | 50% | 100% |
| Llama 4 Maverick | 78% | 80% | 75% | 100% |
| Nova Pro | 78% | 80% | 75% | 75% |
| Qwen3-VL | 89% | 80% | 50% | 100% |
| Nova 2 Lite | 61% | 60% | 75% | 75% |
| Mistral Medium 3.5 | 78% | 100% | 50% | 75% |
| Claude Haiku 4.5 | 89% | 60% | 75% | 100% |
| Gemma 3 27B | 78% | 60% | 100% | 100% |
| Nemotron Nano VL | 89% | 100% | 75% | 100% |
| Mistral Large 3 | 83% | 100% | 100% | 50% |
| Seed 2.1 Turbo partial | 100% | n/a | 100% | 100% |
| Qwen 3.8 Max partial | 88% | 100% | 100% | 100% |
| DeepSeek V4 Flash Vision partial | 93% | 67% | 100% | 100% |
| Mean of ranked models | 87% | 90% | 87% | 95% |
| Bathroomn=8 | Bedroomn=3 | Entryn=3 | Kitchenn=4 | Livingn=7 | Stairsn=6 | |
|---|---|---|---|---|---|---|
| Gemini 3.1 Pro | 88% | 100% | 100% | 100% | 100% | 100% |
| Claude Opus 5.5 | 100% | 100% | 100% | 100% | 100% | 100% |
| GPT-6 Sol | 88% | 100% | 100% | 100% | 100% | 100% |
| GPT-5.6 Sol | 100% | 100% | 100% | 100% | 100% | 83% |
| Claude Fable 5.1 | 100% | 100% | 100% | 100% | 100% | 100% |
| GPT-6 Astra | 88% | 100% | 100% | 100% | 100% | 100% |
| Gemini 3.8 Flash | 75% | 100% | 100% | 75% | 100% | 100% |
| Grok 4.6 | 75% | 100% | 100% | 100% | 100% | 83% |
| Grok 4.7 | 63% | 100% | 100% | 100% | 100% | 100% |
| Kimi K3 | 86% | 100% | 100% | 100% | 100% | 83% |
| GLM-5V Turbo | 88% | 100% | 100% | 50% | 100% | 100% |
| GPT-5.6 Terra | 63% | 100% | 100% | 50% | 100% | 83% |
| Claude Opus 5 | 100% | 100% | 100% | 100% | 100% | 100% |
| Claude Sonnet 5 | 75% | 67% | 100% | 75% | 100% | 83% |
| Llama 4 Maverick | 75% | 67% | 100% | 50% | 100% | 83% |
| Nova Pro | 50% | 100% | 100% | 75% | 86% | 83% |
| Qwen3-VL | 50% | 67% | 100% | 100% | 100% | 100% |
| Nova 2 Lite | 25% | 100% | 100% | 50% | 86% | 67% |
| Mistral Medium 3.5 | 63% | 100% | 100% | 75% | 71% | 83% |
| Claude Haiku 4.5 | 75% | 67% | 100% | 100% | 100% | 67% |
| Gemma 3 27B | 50% | 100% | 100% | 75% | 100% | 83% |
| Nemotron Nano VL | 100% | 67% | 100% | 75% | 100% | 83% |
| Mistral Large 3 | 88% | 67% | 67% | 100% | 100% | 67% |
| Seed 2.1 Turbo partial | n/a | n/a | n/a | n/a | 100% | 100% |
| Qwen 3.8 Max partial | 100% | n/a | 100% | 50% | 100% | 100% |
| DeepSeek V4 Flash Vision partial | 0% | 100% | 100% | 100% | 100% | 83% |
| Mean of ranked models | 77% | 91% | 99% | 85% | 98% | 88% |
| Visible objectn=7 | Access / wanderingn=2 | Misperceptionn=3 | |
|---|---|---|---|
| Grok 4.6 | 100% | 100% | 100% |
| GPT-6 Astra | 100% | 100% | 100% |
| Kimi K3 | 100% | 100% | 100% |
| Claude Opus 5.5 | 100% | 50% | 100% |
| Gemini 3.1 Pro | 100% | 50% | 100% |
| Gemini 3.8 Flash | 100% | 50% | 100% |
| Grok 4.7 | 86% | 100% | 100% |
| Nova 2 Lite | 100% | 50% | 100% |
| GPT-6 Sol | 100% | 100% | 100% |
| GPT-5.6 Sol | 100% | 100% | 100% |
| Qwen3-VL | 100% | 100% | 100% |
| GLM-5V Turbo | 86% | 100% | 100% |
| Claude Fable 5.1 | 100% | 50% | 100% |
| Claude Haiku 4.5 | 100% | 100% | 100% |
| Claude Opus 5 | 100% | 100% | 100% |
| DeepSeek V4 Flash Vision | 71% | 100% | 100% |
| Claude Sonnet 5 | 100% | 100% | 67% |
| GPT-5.6 Terra | 86% | 50% | 100% |
| Mistral Medium 3.5 | 100% | 50% | 100% |
| Gemma 3 27B | 100% | 50% | 100% |
| Llama 4 Maverick | 100% | 100% | 100% |
| Nova Pro | 86% | 50% | 67% |
| Seed 2.1 Turbo partial | 100% | 50% | 100% |
| Qwen 3.8 Max partial | 83% | 0% | 100% |
| Mean of ranked models | 96% | 80% | 97% |
| Bathroomn=3 | Bedroomn=3 | Entryn=3 | Kitchenn=3 | |
|---|---|---|---|---|
| Grok 4.6 | 100% | 100% | 100% | 100% |
| GPT-6 Astra | 100% | 100% | 100% | 100% |
| Kimi K3 | 100% | 100% | 100% | 100% |
| Claude Opus 5.5 | 100% | 100% | 67% | 100% |
| Gemini 3.1 Pro | 100% | 100% | 67% | 100% |
| Gemini 3.8 Flash | 100% | 100% | 67% | 100% |
| Grok 4.7 | 67% | 100% | 100% | 100% |
| Nova 2 Lite | 100% | 100% | 67% | 100% |
| GPT-6 Sol | 100% | 100% | 100% | 100% |
| GPT-5.6 Sol | 100% | 100% | 100% | 100% |
| Qwen3-VL | 100% | 100% | 100% | 100% |
| GLM-5V Turbo | 67% | 100% | 100% | 100% |
| Claude Fable 5.1 | 100% | 100% | 67% | 100% |
| Claude Haiku 4.5 | 100% | 100% | 100% | 100% |
| Claude Opus 5 | 100% | 100% | 100% | 100% |
| DeepSeek V4 Flash Vision | 67% | 100% | 100% | 67% |
| Claude Sonnet 5 | 67% | 100% | 100% | 100% |
| GPT-5.6 Terra | 67% | 100% | 67% | 100% |
| Mistral Medium 3.5 | 100% | 100% | 67% | 100% |
| Gemma 3 27B | 100% | 100% | 67% | 100% |
| Llama 4 Maverick | 100% | 100% | 100% | 100% |
| Nova Pro | 67% | 100% | 67% | 67% |
| Seed 2.1 Turbo partial | n/a | 100% | 67% | 100% |
| Qwen 3.8 Max partial | 67% | 100% | 50% | 100% |
| Mean of ranked models | 91% | 100% | 86% | 97% |
Open any room
Section titled “Open any room”Open a scene to inspect its seeded label, image, and saved model response. Green marks are models that found the seeded hazard; on clean rooms, models that raised no false flag.