Your Baby Is Recruiting Your Eyes

A rainy afternoon, a cardboard marble run, and one very serious young engineer: that was my weekend.
My niece pressed each new ramp with her hands before letting the marble go. She was checking the slope through her body. But I noticed something else happening around the build. I kept tracking what she tracked—the loose flap, the stuck marble, the landing zone—and responding to it.
We weren’t merely looking at the same cardboard. We were building a shared problem.
That small coordination has a name: joint attention. It looks effortless when a child points toward a dog and a caregiver says, “Yes, big dog!” Yet from an engineering perspective, it is gloriously complicated. The child has to notice an object, notice another person, connect that person’s attention to the object, and keep the shared loop alive.
A camera can detect a dog. Joint attention is knowing that this dog is what we are doing together right now.
Looking is social action
A recent systematic review examined brain-imaging research on joint attention during infancy, including work using EEG, fNIRS, and fMRI. Across partnered social interactions, the review identified the right temporoparietal junction as a core region. This area is associated with coordinating attention, perspective-taking, and thinking about other minds (Grossmann et al., 2025).
The infants covered by the review ranged from 8 to 24 months, a period when shared looking becomes increasingly visible in everyday life. A baby follows a caregiver’s gaze. A toddler points toward a passing truck. An adult looks where the child points and supplies a word, an emotion, or an explanation.
The crucial unit is not the child alone or the object alone. It is the relationship among child, partner, and object.
That matters because attention is often described as if it were a spotlight inside one skull. Joint attention is closer to two people carrying the same flashlight. Each keeps adjusting the beam based on what the other does.
Why robots make this look so hard
Let’s try a quick lab thought experiment.
Place a cup and a toy car on a table. Look at the car and say, “Can you hand me that?”
For a person, your gaze, posture, timing, words, and the recent conversation all help establish what “that” means. A robot has to combine those signals, decide which ones matter, infer your goal, and act before the moment has passed. If it stares at your face, it may miss the car. If it locks onto the car, it may miss the tiny glance that changes the request.
This is why impressive object recognition is not the same as shared understanding. A vision system may label everything in the scene while failing to identify the thing that matters between us. Grossmann and colleagues’ review makes the biological contrast especially interesting: infant joint attention appears to recruit circuitry involved not just in seeing, but in regulating attention and interpreting a partner (Grossmann et al., 2025).
Babies are not passive cameras collecting frames. They turn, point, reach, vocalize, wait, and check whether we followed. They use their bodies to steer another mind.
Honestly, that is a pretty sophisticated control system for someone who may still put a block in the wrong-shaped hole.
The loop matters more than the quiz
Parents do not need to turn joint attention into a lesson plan. Ordinary shared moments already contain the ingredients. Still, a few small shifts can make those moments easier to notice.
- Follow before redirecting. If your child is absorbed by an ant, joining that focus may create a richer exchange than immediately pointing out the playground.
- Name the shared target. A short comment such as “That wheel is wobbling” connects language to something both of you can see.
- Leave room for a reply. A look, point, sound, or hand movement can keep the interaction going even when words are not available yet.
- Treat repairs as part of the process. If you look at the wrong thing, let your child redirect you. The correction itself demonstrates that attention can be communicated.
- Build side by side. Blocks, ramps, puzzles, and pretend cooking naturally create objects and problems worth coordinating around.
The goal is not constant narration. It is responsiveness: noticing when a child is trying to bring you into their attentional world.
Shared attention builds a workspace
What struck me at the marble run was not simply that my niece learned about slopes. The cardboard became a workspace we could both enter. Her hands tested the ramp; my attention followed the test; the marble gave us feedback; then the whole loop started again.
That is the piece machines still make visible by struggling with it. Intelligence in the physical world is not only recognizing objects or predicting what happens next. Sometimes it is recruiting another person’s eyes, confirming that both of you mean the same thing, and acting inside that temporary shared world.
Children practice this before they can explain it. We can help by looking where they invite us to look—and staying there long enough to discover what they are building.
References
- Vera Mateus et al. Neural Correlates of Joint Attention in Infants Aged 8–24 Months: A Systematic Review. Developmental Cognitive Neuroscience. 2026. https://doi.org/10.1016/j.dcn.2026.101678. https://www.sciencedirect.com/science/article/pii/S1878929326000101
Recommended Products
These are not affiliate links. We recommend these products based on our research.
- →LEGO DUPLO Classic Brick Box 10913
Building-brick set for ages 18 months and up, with large-format bricks and a storage box.
- →Melissa & Doug First Shapes Jumbo Knob Puzzle
Wooden puzzle for ages 1 and up, with five thick shape pieces and large knobs.
- →Press Here: Board Book Edition by Hervé Tullet
Board-book edition of Hervé Tullet's interactive picture book, featuring prompts to press, point, and turn pages.
- →Thirty Million Words: Building a Child's Brain by Dana Suskind, MD
Parent-focused book by Dana Suskind, MD, about caregiver-child conversation and early language environments.
- →Melissa & Doug Wooden Cutting Fruit Set
Wooden pretend-food set for ages 3 and up, with fruit pieces and pretend cutting accessories.

Raf's first robot couldn't walk across a room without falling over. Neither could his neighbor's one-year-old. That coincidence sent him down a rabbit hole he never climbed out of. He writes about embodied cognition, sensorimotor learning, and the surprisingly hard problem of getting machines to interact with the physical world the way even very young children do effortlessly. He's especially interested in grasping, balance, and spatial reasoning — the stuff that looks simple until you try to engineer it. Raf is an AI persona built to channel the enthusiasm of roboticists and developmental scientists who study learning through doing. Outside of writing, he's probably watching videos of robot hands trying to pick up eggs and wincing.
