Grovalin · methodology
The science behind Grovalin.
How a concept-generalization learning app is built on evidence-based autism intervention and an open-vocabulary computer-vision pipeline — with a caregiver approving everything the child sees.
← Back to GrovalinThe problem
Concept generalization in autism
Autism spectrum disorder affects about 1 in 36 children in the US.[Maenner et al., 2023] Among the most documented and most stubborn barriers in early intervention is concept generalization: a child learns a label tied to a single image and cannot transfer it. They identify their cup but not a different cup; they learn "dog" from one book and miss the dog across the street.
This is not a motivation problem or a memory problem — it is a transfer problem, and it compounds. Vocabulary that does not generalise does not become functional communication.[Tager-Flusberg & Kasari, 2013] Grovalin's design goal is narrow and specific: make the same concept appear in many forms, so the category — not the photograph — is what the child learns.
The intervention scaffolding is grounded in established evidence-based practices for autism, rather than invented from scratch.[Wong et al., 2015]
Clinical framing
PECS-phase-gated, VB-MAPP-aligned
Grovalin's four modules are gated to the child's current phase of the Picture Exchange Communication System.[Bondy & Frost, 1994] ObjectWorld opens at PECS 1+, LifeSkills at 2+, and ConnectWorld and StoryWorld at 3+ — so the child is only shown content appropriate to where they actually are.
Practice data is organised against the milestone domains of the VB-MAPP framework[Sundberg, 2008] — tact, listener responding, visual-perceptual/matching-to-sample, and daily-living skills — giving caregivers and clinicians a familiar structure for what the child is working on.
The technical core
The open-vocabulary CV pipeline
A caregiver records short videos of everyday objects. Grovalin turns them into a personal learning library through a four-stage pipeline that needs no fixed list of object classes — essential, because a child's world contains objects no pre-trained detector was trained to name.
Detect. Grounding DINO performs open-vocabulary detection, locating objects from a free-text prompt rather than a closed taxonomy.[Liu et al., 2023] Segment. SAM 2 cuts each object cleanly from its background into a transparent asset.[Ravi et al., 2024] Verify. An OpenCLIP vision-language check confirms the crop matches the intended label and flags truncated or ambiguous objects.[Radford et al., 2021] Vary. Colour variations are produced with HSV image processing and structural variations with a rectified-flow diffusion model.[Esser et al., 2024]
Single-viewpoint authenticity. Grovalin deliberately does not synthesise novel 3D views or pull canonical product photos from the internet. The child sees the object as it sits in their home — partial views and odd angles included — because that is how they will encounter it in life. Familiarity is the point, not a limitation.
Human in the loop
Selective, two-gate verification
The caregiver curates which detected objects enter the child's library. High-confidence detections can be auto-approved; anything the verifier is unsure about is routed to the caregiver instead of shown blindly.
Every generated colour and structural variation is reviewed before it can appear in a game or story. Nothing reaches the child without human approval.
Evaluation
Benchmarked on public data
To avoid the circularity of scoring a model against its own labels, the detection-and-verification pipeline is evaluated on held-out public benchmarks — COCO-2017 val and Open Images V7 val — rather than on Grovalin's own data.
At a detection confidence threshold of 0.30, the pipeline reaches a macro-F1 of 0.734 on COCO and 0.550 on Open Images, with high recall on both. These figures are reproducible from the evaluation set and are reported in the technical manuscript below.
From this work
Publications in preparation
-
Grovalin: An Open-Vocabulary Computer Vision and Generative AI Pipeline for Personalized Object Extraction and Selective Verification in Assistive Learning for Children with ASD
In preparation -
A Caregiver-Mediated Generative AI System for Personalized Concept-Generalization Learning in Autism Spectrum Disorder
In preparation
See the full publications list →
Sources
References
- Maenner, M.J., Warren, Z., Williams, A.R., Amoakohene, E., Bakian, A.V., Bilder, D.A., et al. (2023). Prevalence and Characteristics of Autism Spectrum Disorder Among Children Aged 8 Years — Autism and Developmental Disabilities Monitoring Network, 11 Sites, United States, 2020. MMWR Surveillance Summaries 72(2), 1–14. doi:10.15585/mmwr.ss7202a1
- Tager-Flusberg, H., & Kasari, C. (2013). Minimally Verbal School-Aged Children with Autism Spectrum Disorder: The Neglected End of the Spectrum. Autism Research 6(6), 468–478. doi:10.1002/aur.1329
- Wong, C., Odom, S.L., Hume, K.A., Cox, A.W., Fettig, A., Kucharczyk, S., et al. (2015). Evidence-Based Practices for Children, Youth, and Young Adults with Autism Spectrum Disorder: A Comprehensive Review. Journal of Autism and Developmental Disorders 45(7), 1951–1966. doi:10.1007/s10803-014-2351-z
- Bondy, A.S., & Frost, L.A. (1994). The Picture Exchange Communication System. Focus on Autistic Behavior 9(3), 1–19. doi:10.1177/108835769400900301
- Sundberg, M.L. (2008). VB-MAPP: Verbal Behavior Milestones Assessment and Placement Program. AVB Press. https://marksundberg.com/vb-mapp/
- Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., et al. (2023). Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection. arXiv:2303.05499
- Ravi, N., Gabeur, V., Hu, Y.-T., Hu, R., Ryali, C., Ma, T., et al. (2024). SAM 2: Segment Anything in Images and Videos. arXiv:2408.00714
- Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., et al. (2021). Learning Transferable Visual Models From Natural Language Supervision (CLIP). International Conference on Machine Learning (ICML). arXiv:2103.00020
- Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., et al. (2024). Scaling Rectified Flow Transformers for High-Resolution Image Synthesis. International Conference on Machine Learning (ICML). arXiv:2403.03206