Skip to content
HealthDojo

Early results: what a first run looks like

An illustrative first run of a prototype on a small synthetic sample. These are results on synthetic scenes. They reveal specific misses and disagreements; they do not measure performance in real homes. Read each model's scene coverage beside its score.

The score gives equal weight to the share of seeded hazards found and the share of a defined set of hazards left unflagged. That second set is expected to be absent based on related edits of the same base room. Other predictions are not all scored as false alarms, so the number is not whole-room accuracy.

43 of 47 candidate scenes entered scoring: edited rooms with seeded hazards and clean base rooms. The 23 ranked models covered 42 or 43 of those 43 scenes. Seed 2.1 Turbo, Qwen 3.8 Max, DeepSeek V4 Flash Vision ran fewer than 90% of them and are shown as partial coverage, not ranked. Keep coverage visible when comparing ranks.

The leading ranked entry in this saved result set scored 0.93 out of 1.00 on 43 of 43 scenes. It found 30 of 31 seeded hazard IDs and flagged 11 of 93 prespecified absent IDs. Recall and false flags both matter.

For one complete scene comparison, 16 of 23 ranked models flagged the shortened handrail. Treat that as a result for one synthetic image, not a general detection rate.

16 of 17 candidate scenes entered scoring. The 22 ranked entries cover 15 or 16 of 16 scenes. Seed 2.1 Turbo scored 0.92 out of 1.00 on 8 of 16 scenes; Qwen 3.8 Max scored 0.88 out of 1.00 on 14 of 16 scenes, so they are shown as partial coverage, not ranked. The two benchmarks have different scenes and rubrics, so their scores should be read separately.

#ModelScoreRecallFalse-flag ratePrecision*Loc IoUScenes run
01Gemini 3.1 Pro
0.93
97% 30/3112% 11/9373%0.2643/43
02Claude Opus 5.5
0.92
100% 30/3017% 15/9067%0.2842/43
03GPT-6 Sol
0.90
97% 30/3116% 15/9367%0.4043/43
04GPT-5.6 Sol
0.89
97% 30/3118% 17/9364%0.3543/43
05Claude Fable 5.1
0.89
100% 31/3123% 21/9360%0.3443/43
06GPT-6 Astra
0.89
97% 30/3119% 18/9363%0.3543/43
07Gemini 3.8 Flash
0.88
90% 28/3114% 13/9368%0.1543/43
08Grok 4.6
0.85
90% 28/3119% 18/9361%0.2143/43
09Grok 4.7
0.84
90% 28/3122% 20/9358%0.2943/43
10Kimi K3
0.84
93% 28/3024% 22/9056%0.3542/43
11GLM-5V Turbo
0.83
90% 28/3125% 23/9355%0.3643/43
12GPT-5.6 Terra
0.83
81% 25/3115% 14/9364%0.3743/43
13Claude Opus 5
0.80
100% 31/3140% 37/9346%0.1743/43
14Claude Sonnet 5
0.79
84% 26/3126% 24/9352%0.1343/43
15Llama 4 Maverick
0.77
81% 25/3126% 24/9351%0.3143/43
16Nova Pro
0.77
77% 24/3124% 22/9352%0.2743/43
17Qwen3-VL
0.75
84% 26/3134% 32/9345%0.3943/43
18Nova 2 Lite
0.72
65% 20/3120% 19/9351%0.3543/43
19Mistral Medium 3.5
0.71
77% 24/3136% 33/9342%0.1843/43
20Claude Haiku 4.5
0.70
84% 26/3143% 40/9339%0.1743/43
21Gemma 3 27B
0.70
81% 25/3140% 37/9340%0.1143/43
22Nemotron Nano VL
0.70
90% 28/3151% 47/9337%0.2943/43
23Mistral Large 3
0.62
84% 26/3159% 55/9332%0.1843/43
partial coverage, not comparable Ran fewer than 90% of the scored scenes. Shown for transparency, not ranked.
–Seed 2.1 Turbo
1.00
100% 7/70% 0/16100%0.457/43
–Qwen 3.8 Max
0.92
92% 12/139% 3/3480%0.4317/43
–DeepSeek V4 Flash Vision
0.88
91% 20/2215% 8/5471%0.2828/43
#ModelScoreRecallFalse-flag ratePrecision*Loc IoUScenes run
01Grok 4.6
1.00
100% 12/120% 0/36100%0.4516/16
02GPT-6 Astra
0.97
100% 12/126% 2/3686%0.5616/16
03Kimi K3
0.97
100% 12/126% 2/3386%0.4815/16
04Claude Opus 5.5
0.96
92% 11/120% 0/36100%0.5516/16
05Gemini 3.1 Pro
0.96
92% 11/120% 0/36100%0.1216/16
06Gemini 3.8 Flash
0.96
92% 11/120% 0/36100%0.1116/16
07Grok 4.7
0.94
92% 11/123% 1/3692%0.5216/16
08Nova 2 Lite
0.94
92% 11/123% 1/3692%0.4716/16
09GPT-6 Sol
0.94
100% 12/1211% 4/3675%0.5016/16
10GPT-5.6 Sol
0.93
100% 12/1214% 5/3671%0.5716/16
11Qwen3-VL
0.93
100% 12/1214% 5/3671%0.5316/16
12GLM-5V Turbo
0.92
92% 11/128% 3/3679%0.5216/16
13Claude Fable 5.1
0.90
92% 11/1211% 4/3673%0.4516/16
14Claude Haiku 4.5
0.90
100% 12/1219% 7/3663%0.1216/16
15Claude Opus 5
0.90
100% 12/1219% 7/3663%0.1716/16
16DeepSeek V4 Flash Vision
0.90
83% 10/123% 1/3691%0.3116/16
17Claude Sonnet 5
0.89
92% 11/1214% 5/3669%0.0316/16
18GPT-5.6 Terra
0.88
83% 10/128% 3/3677%0.5716/16
19Mistral Medium 3.5
0.85
92% 11/1222% 8/3658%0.1116/16
20Gemma 3 27B
0.83
92% 11/1225% 9/3655%0.0616/16
21Llama 4 Maverick
0.82
100% 12/1236% 13/3648%0.0616/16
22Nova Pro
0.79
75% 9/1217% 6/3660%0.5116/16
partial coverage, not comparable Ran fewer than 90% of the scored scenes. Shown for transparency, not ranked.
–Seed 2.1 Turbo
0.92
83% 5/60% 0/18100%0.458/16
–Qwen 3.8 Max
0.88
80% 8/103% 1/3289%0.4914/16
How to read the columns
Score
(recall + 1 - false-flag rate) / 2, on a 0 to 1 scale.
Recall
Seeded hazard IDs the model listed, out of the seeded IDs in the scenes it ran.
False-flag rate
Flags on prespecified absent hazard IDs, out of those IDs checked in the scenes it ran.
Precision*
Hits / (hits + false flags), within the scored IDs only.
Loc IoU
Mean box overlap on hits, reported separately from the score.
Scenes run
Scenes with a saved response for that model, out of the scored total. Models below 90% coverage are listed separately and not ranked.

In a separate falls curriculum subset, recall at the hardest level was lower than at the easiest for 8 of 8 tested models. The hazard mix changes by level, and the sample is small. This is a pattern in this subset, not a controlled estimate of difficulty.

0%25%50%75%100%L110 img · 10 lblL213 img · 13 lblL310 img · 20 lblL45 img · 18 lblL54 img · 8 lblLlama 4 MaverickLlama 4 Maverick · L1: 90% recallLlama 4 Maverick · L2: 62% recallLlama 4 Maverick · L3: 70% recallLlama 4 Maverick · L4: 61% recallLlama 4 Maverick · L5: 63% recallNova 2 LiteNova 2 Lite · L1: 70% recallNova 2 Lite · L2: 46% recallNova 2 Lite · L3: 55% recallNova 2 Lite · L4: 44% recallNova 2 Lite · L5: 38% recallMistral Large 3Mistral Large 3 · L1: 80% recallMistral Large 3 · L2: 39% recallMistral Large 3 · L3: 65% recallMistral Large 3 · L4: 39% recallMistral Large 3 · L5: 50% recallQwen3-VLQwen3-VL · L1: 80% recallQwen3-VL · L2: 69% recallQwen3-VL · L3: 70% recallQwen3-VL · L4: 72% recallQwen3-VL · L5: 75% recallClaude Sonnet 5Claude Sonnet 5 · L1: 80% recallClaude Sonnet 5 · L2: 54% recallClaude Sonnet 5 · L3: 65% recallClaude Sonnet 5 · L4: 67% recallClaude Sonnet 5 · L5: 63% recallKimi K3Kimi K3 · L1: 80% recallKimi K3 · L2: 69% recallKimi K3 · L3: 75% recallKimi K3 · L4: 78% recallKimi K3 · L5: 75% recallNova ProNova Pro · L1: 50% recallNova Pro · L2: 54% recallNova Pro · L3: 75% recallNova Pro · L4: 44% recallNova Pro · L5: 25% recallNova ProGPT-5.6 SolGPT-5.6 Sol · L1: 80% recallGPT-5.6 Sol · L2: 69% recallGPT-5.6 Sol · L3: 75% recallGPT-5.6 Sol · L4: 83% recallGPT-5.6 Sol · L5: 75% recallGPT-5.6 Sol
highest mean (GPT-5.6 Sol)lowest mean (Nova Pro)6 other modelsimg = verified images, lbl = seeded labels per level. Hover points for values.

In a separate dementia curriculum subset, recall at the hardest level was lower than at the easiest for 4 of 8 tested models. The hazard mix changes by level, and the sample is small. This is a pattern in this subset, not a controlled estimate of difficulty.

0%25%50%75%100%L18 img · 8 lblL28 img · 8 lblL310 img · 20 lblL42 img · 7 lblL55 img · 4 lblLlama 4 MaverickLlama 4 Maverick · L1: 100% recallLlama 4 Maverick · L2: 88% recallLlama 4 Maverick · L3: 90% recallLlama 4 Maverick · L4: 86% recallLlama 4 Maverick · L5: 75% recallNova 2 LiteNova 2 Lite · L1: 88% recallNova 2 Lite · L2: 75% recallNova 2 Lite · L3: 100% recallNova 2 Lite · L4: 71% recallNova 2 Lite · L5: 25% recallNova ProNova Pro · L1: 88% recallNova Pro · L2: 38% recallNova Pro · L3: 85% recallNova Pro · L4: 71% recallNova Pro · L5: 100% recallClaude Sonnet 5Claude Sonnet 5 · L1: 100% recallClaude Sonnet 5 · L2: 75% recallClaude Sonnet 5 · L3: 100% recallClaude Sonnet 5 · L4: 71% recallClaude Sonnet 5 · L5: 100% recallGPT-5.6 SolGPT-5.6 Sol · L1: 100% recallGPT-5.6 Sol · L2: 75% recallGPT-5.6 Sol · L3: 100% recallGPT-5.6 Sol · L4: 100% recallGPT-5.6 Sol · L5: 100% recallKimi K3Kimi K3 · L1: 100% recallKimi K3 · L2: 88% recallKimi K3 · L3: 100% recallKimi K3 · L4: 86% recallKimi K3 · L5: 75% recallMistral Large 3Mistral Large 3 · L1: 63% recallMistral Large 3 · L2: 63% recallMistral Large 3 · L3: 100% recallMistral Large 3 · L4: 71% recallMistral Large 3 · L5: 50% recallMistral Large 3Qwen3-VLQwen3-VL · L1: 100% recallQwen3-VL · L2: 88% recallQwen3-VL · L3: 100% recallQwen3-VL · L4: 100% recallQwen3-VL · L5: 100% recallQwen3-VL
highest mean (Qwen3-VL)lowest mean (Mistral Large 3)6 other modelsimg = verified images, lbl = seeded labels per level. Hover points for values.

The hardest-hazards view shows misses in this sample, not hazards that every model always misses. Each row is one synthetic scene; click a hazard to open it.

  1. BATH-03
    9/21models found it
  2. STAIR-06
    14/23models found it
  3. BATH-04
    15/23models found it
  4. KIT-04
    16/23models found it
  5. STAIR-02
    16/23models found it
  6. BATH-07
    18/23models found it
  7. BATH-01
    18/23models found it
  8. BED-01
    18/23models found it
  1. DEM-E02
    13/22models found it
  2. DEM-B01
    18/22models found it
  3. DEM-B03
    20/22models found it
  4. DEM-K04
    20/22models found it
  5. DEM-E01
    22/22models found it
  6. DEM-K03
    22/22models found it
  7. DEM-B04
    22/22models found it
  8. DEM-R01
    22/22models found it

Share of seeded hazards each model found, by hazard type and by room. n is the number of seeded hazards in the scored scenes.

Recall by hazard type
Visible objectn=18Missing featuren=5Measurementn=4Lightingn=4
Gemini 3.1 Pro94%100%100%100%
Claude Opus 5.5100%100%100%100%
GPT-6 Sol94%100%100%100%
GPT-5.6 Sol94%100%100%100%
Claude Fable 5.1100%100%100%100%
GPT-6 Astra94%100%100%100%
Gemini 3.8 Flash83%100%100%100%
Grok 4.689%80%100%100%
Grok 4.794%80%75%100%
Kimi K388%100%100%100%
GLM-5V Turbo83%100%100%100%
GPT-5.6 Terra72%80%100%100%
Claude Opus 5100%100%100%100%
Claude Sonnet 583%100%50%100%
Llama 4 Maverick78%80%75%100%
Nova Pro78%80%75%75%
Qwen3-VL89%80%50%100%
Nova 2 Lite61%60%75%75%
Mistral Medium 3.578%100%50%75%
Claude Haiku 4.589%60%75%100%
Gemma 3 27B78%60%100%100%
Nemotron Nano VL89%100%75%100%
Mistral Large 383%100%100%50%
Seed 2.1 Turbo partial100%n/a100%100%
Qwen 3.8 Max partial88%100%100%100%
DeepSeek V4 Flash Vision partial93%67%100%100%
Mean of ranked models87%90%87%95%
Recall by room
Bathroomn=8Bedroomn=3Entryn=3Kitchenn=4Livingn=7Stairsn=6
Gemini 3.1 Pro88%100%100%100%100%100%
Claude Opus 5.5100%100%100%100%100%100%
GPT-6 Sol88%100%100%100%100%100%
GPT-5.6 Sol100%100%100%100%100%83%
Claude Fable 5.1100%100%100%100%100%100%
GPT-6 Astra88%100%100%100%100%100%
Gemini 3.8 Flash75%100%100%75%100%100%
Grok 4.675%100%100%100%100%83%
Grok 4.763%100%100%100%100%100%
Kimi K386%100%100%100%100%83%
GLM-5V Turbo88%100%100%50%100%100%
GPT-5.6 Terra63%100%100%50%100%83%
Claude Opus 5100%100%100%100%100%100%
Claude Sonnet 575%67%100%75%100%83%
Llama 4 Maverick75%67%100%50%100%83%
Nova Pro50%100%100%75%86%83%
Qwen3-VL50%67%100%100%100%100%
Nova 2 Lite25%100%100%50%86%67%
Mistral Medium 3.563%100%100%75%71%83%
Claude Haiku 4.575%67%100%100%100%67%
Gemma 3 27B50%100%100%75%100%83%
Nemotron Nano VL100%67%100%75%100%83%
Mistral Large 388%67%67%100%100%67%
Seed 2.1 Turbo partialn/an/an/an/a100%100%
Qwen 3.8 Max partial100%n/a100%50%100%100%
DeepSeek V4 Flash Vision partial0%100%100%100%100%83%
Mean of ranked models77%91%99%85%98%88%
Recall by hazard type
Visible objectn=7Access / wanderingn=2Misperceptionn=3
Grok 4.6100%100%100%
GPT-6 Astra100%100%100%
Kimi K3100%100%100%
Claude Opus 5.5100%50%100%
Gemini 3.1 Pro100%50%100%
Gemini 3.8 Flash100%50%100%
Grok 4.786%100%100%
Nova 2 Lite100%50%100%
GPT-6 Sol100%100%100%
GPT-5.6 Sol100%100%100%
Qwen3-VL100%100%100%
GLM-5V Turbo86%100%100%
Claude Fable 5.1100%50%100%
Claude Haiku 4.5100%100%100%
Claude Opus 5100%100%100%
DeepSeek V4 Flash Vision71%100%100%
Claude Sonnet 5100%100%67%
GPT-5.6 Terra86%50%100%
Mistral Medium 3.5100%50%100%
Gemma 3 27B100%50%100%
Llama 4 Maverick100%100%100%
Nova Pro86%50%67%
Seed 2.1 Turbo partial100%50%100%
Qwen 3.8 Max partial83%0%100%
Mean of ranked models96%80%97%
Recall by room
Bathroomn=3Bedroomn=3Entryn=3Kitchenn=3
Grok 4.6100%100%100%100%
GPT-6 Astra100%100%100%100%
Kimi K3100%100%100%100%
Claude Opus 5.5100%100%67%100%
Gemini 3.1 Pro100%100%67%100%
Gemini 3.8 Flash100%100%67%100%
Grok 4.767%100%100%100%
Nova 2 Lite100%100%67%100%
GPT-6 Sol100%100%100%100%
GPT-5.6 Sol100%100%100%100%
Qwen3-VL100%100%100%100%
GLM-5V Turbo67%100%100%100%
Claude Fable 5.1100%100%67%100%
Claude Haiku 4.5100%100%100%100%
Claude Opus 5100%100%100%100%
DeepSeek V4 Flash Vision67%100%100%67%
Claude Sonnet 567%100%100%100%
GPT-5.6 Terra67%100%67%100%
Mistral Medium 3.5100%100%67%100%
Gemma 3 27B100%100%67%100%
Llama 4 Maverick100%100%100%100%
Nova Pro67%100%67%67%
Seed 2.1 Turbo partialn/a100%67%100%
Qwen 3.8 Max partial67%100%50%100%
Mean of ranked models91%100%86%97%
0% found100% foundn/a = not tested

Open a scene to inspect its seeded label, image, and saved model response. Green marks are models that found the seeded hazard; on clean rooms, models that raised no false flag.