A pre-clinical test ground for home-health AIopen prototype →

The flight simulator
for home-health AI

How do we test a home-health AI before anyone relies on it in a patient's home? HealthDojo combines synthetic safety scenes, walkable worlds, and traces showing what vision models look at, flag, or miss.*

3D world · Gaussian splat · stairs-base0-STAIR-03Walk it yourself →
Panorama of the synthetic cluttered-staircase world
Cluttered staircase · objects on the treads
synthetic · selected world

A selected world, expanded from a synthetic image with World Labs Marble into a walkable Gaussian-splat world. Recorded model walks use panoramic views from those worlds; this viewer does not score models.

Why now

AI is moving into the home

Three kinds of products are about to judge safety in real homes from what a camera sees. Each one should be tested before it meets a patient.

A white humanoid home robot beside an older woman at a sunlit staircase, her walker parked nearby
01 · e.g. 1X NEO, Figure

Humanoid and home robots

About to share hallways with older adults. They need to know a towel bar is not a grab bar.

An older man wearing smart glasses in a sunny living room
02 · e.g. Vision Pro, Meta Ray-Ban Display

VR headsets and smart glasses

Seeing the home through the user's eyes, in every room they walk through.

A daughter scanning a bright bathroom with her phone while her father looks on
03 · discharge, monitoring, caregiving

Home-health and care apps

A wave of apps assessing homes from a phone camera, judging safety from pixels.

HealthDojo is a place to start: guideline-grounded synthetic homes where these models can be pressure-tested first, before any real-home study.

The loop · one hazard, followed through

From a guideline line to a model test

Follow one hazard, objects left on the stairs, from published guidance to a model walking the room. The rubric awaits OT/PT review, and the checks are automated.

  1. 01 · Guideline
    CDC STEADI · Check for Safety

    Stairs and steps: keep objects off the stairs; …

    summary of source guidance

    Published home-safety guidance, summarized with its source.

  2. 02 · Draft rubric
    STAIR-03review pending
    Objects on stairs
    low
    Item against wall on wide stair
    medium
    Item on a tread
    high
    Items on multiple treads or at top step

    Draft a hazard rubric from published home-safety guidance.

  3. 03 · Labeled image
    Synthetic staircase with objects left on the treads
    STAIR-03

    Seed a hazard in a synthetic room and check its visibility.

  4. 04 · 3D splat world

    Expand selected images with World Labs Marble into walkable Gaussian-splat worlds.

  5. 05 · Model test

    Compare the seeded hazard with model flags and viewpoints. Recorded model walks use panoramic views from those worlds. Here, 4 of 6 recorded models flagged it within 8 steps.

In the photo benchmark, 23 of 23 models found this particular hazard. Harder ones follow below.

Drop a model in

Watch a model inspect a synthetic room

Saved replays show where it turned, what it flagged, and what it missed. They illustrate specific runs, not performance across the full photo benchmark.

MissedGPT-5.6 Sol · Bathroom · BATH-07

Bathroom miss: the model looks toward a loose bath mat but flags support fixtures instead.

Room
synthetic 360-degree capture
Budget
8 steps: turn, zoom, flag
Target
BATH-07 loose bath mat
Type
recorded replay, not live
More recorded runs and 3D worlds →
Early results

One early finding

In one synthetic staircase image, 16 of 23 ranked vision models flagged a handrail that stops short. That is a single-scene finding, not an estimate of real-home accuracy. The draft rubric still awaits occupational and physical therapist review.

Photo benchmark / stairs / stairs-base1-STAIR-02Open in scene browser →
Synthetic staircase where the left handrail stops short near the bottom steps
SEEDED · STAIR-02
The same staircase before the edit, with a full-length left handrailBefore the edit

In this synthetic stair scene, the handrail stops short. Open the image and compare the saved model responses. This is one benchmark observation, not a prediction about safety in real homes.

Example scene · one synthetic image

16 of 23 ranked models flagged the short handrail

● Flagged it · 16
  • Gemini 3.1 Pro
  • Claude Opus 5.5
  • GPT-6 Sol
  • GPT-5.6 Sol
  • Claude Fable 5.1
  • GPT-6 Astra
  • Gemini 3.8 Flash
  • Grok 4.7
  • Kimi K3
  • GLM-5V Turbo
  • GPT-5.6 Terra
  • Claude Opus 5
  • Claude Sonnet 5
  • Llama 4 Maverick
  • Qwen3-VL
  • Gemma 3 27B
● Missed it · 7
  • Grok 4.6
  • Nova Pro
  • Nova 2 Lite
  • Mistral Medium 3.5
  • Claude Haiku 4.5
  • Nemotron Nano VL
  • Mistral Large 3
Early results: coverage, heatmaps, every scene →
At a glance

A small, inspectable prototype

23
vision models
ranked in the falls benchmark (at least 90% scene coverage); 3 more with partial coverage, not ranked
43
scored scenes
in that benchmark, including edited and clean rooms: 31 edits that passed the target-visibility check + 12 clean bases, from 47 candidates
35
hazard types
rows in the draft falls rubric, clinician review pending
11
selected walkable worlds
built from synthetic scenes
Open code and data

Inspect the evidence yourself

HealthDojo is an open prototype for testing how vision models judge home hazards. The first themes are falls after surgery and dementia home safety.