Cognition & AI

Before Words, Babies Learn “I Did That”

Before Words, Babies Learn “I Did That”

A baby kicks. The mobile above the crib shivers. The baby goes still, then kicks again.

To an adult, this can look like charming chaos. To the baby, it may be the beginning of a consequential discovery: something changed because I acted.

That discovery is called a sense of agency. It does not arrive as a philosophical statement. It is assembled through small, repeated loops between movement and response: kick, sway; reach, touch; babble, smile. Before children can explain cause and effect, they are learning that they are not merely watching the world. They can alter it.

From movement to intention

Researchers have long used the mobile-kicking task to study this transition. A ribbon connects a baby’s leg to a colorful mobile, so a kick produces visible movement. In a recent version, researchers recorded the movements of babies around 3 months old and used machine-learning systems to analyze patterns across their bodies.

The systems could distinguish movement during the connected condition from movement when the ribbon was not producing the same result. The researchers interpreted the shift from loosely structured exploration toward more organized action as evidence that babies had detected the functional relationship between their movement and the environment (Khodadadzadeh et al., 2024).

The delightful twist is that artificial intelligence was not teaching these babies. It was helping adults notice organization in motion that can otherwise look random. The study even echoes an old intuition from Alan Turing: if we want to understand intelligence, perhaps we should pay closer attention to how it develops rather than only inspecting its finished form.

Still, we should be careful with the word awareness. A classifier detecting different movement patterns does not give us direct access to a baby’s experience. It cannot tell us whether the infant has a reflective thought resembling “I did that.” What it offers is a behavioral window onto a more modest—and still remarkable—achievement: the baby is learning an action–outcome relationship and reorganizing movement around it.

Agency needs a legible world

Contingency is not simply stimulation. A toy that flashes constantly may be exciting, but it does not necessarily teach a child what made the flash happen. For agency learning, the relationship must be discoverable: an action occurs, an outcome follows, and the child gets another chance to test the connection.

That distinction matters beyond infancy. Much of children’s technology is “responsive,” but not always in ways children can understand. An app may adapt a lesson, award a badge, suppress a video, or change a difficulty level according to hidden calculations. The system reacts, yet the child may not know why.

We should ask whether responsiveness without legibility strengthens agency or quietly weakens it. A child who presses a block and hears a sound can form a testable hypothesis. A child whose learning platform silently profiles performance is acting inside a system whose consequences are decided elsewhere.

This is not an argument that every mechanism must be exposed or every toy must be simple. It is an argument for preserving meaningful loops in which children can connect choices to outcomes. Agency is not the same as being kept busy.

What babies can teach robots

Robotics researchers face a related problem. A machine trained mainly on language or stored examples can know many descriptions of the world without having learned what its own actions do within that world. David Vernon argues that genuinely collaborative robots may need developmental architectures: systems that build knowledge through embodied interaction, memory, and exploration rather than relying only on pretraining (Vernon, 2025).

Recent robotics work points in that direction. An embodied language-model framework performed complex tasks in unpredictable settings by combining language-based planning with visual and force feedback. The important ingredient was not language alone; the robot had to register what happened when it acted and revise accordingly (Mon-Williams et al., 2025).

The parallel should not be pushed too far. A robot updating after contact is not necessarily experiencing authorship, autonomy, or a self. But both cases expose the limits of intelligence without a closed loop between perception and action. Knowing about consequences is different from discovering that your action produced one.

Practical ways to protect agency

Parents do not need a special program. Ordinary play already contains rich opportunities:

  • Leave room for the second try. When your baby shakes, drops, pushes, or vocalizes, a brief pause can let them repeat the action and inspect the result.
  • Choose some toys with understandable consequences. Blocks fall, lids open, rattles sound, and water moves. Their “rules” can be explored rather than merely consumed.
  • Respond contingently. Imitating a sound or answering a gesture creates a social action–outcome loop. The response need not be instant or perfect to be meaningful.
  • Ask what an app makes visible. Can your child tell what choice changed the outcome? If not, narrating the system—or choosing a more transparent activity—can return some control to the child.
  • Do not confuse assistance with authorship. Helping is valuable, but completing every difficult action for a child can remove the very feedback they were trying to understand.

The ethical lesson is quiet but important. Agency begins before consent forms, report cards, and recommendation algorithms. Children learn what it means to act in worlds we design for them. The question is not only whether those worlds respond. It is whether children can discover how—and whether the response leaves them feeling like participants rather than inputs.

References

  1. David Vernon. The Future of Research in Cognitive Robotics: Foundation Models or Developmental Cognitive Models?. Advanced Robotics Research. 2025. https://doi.org/10.1002/adrr.202500066. https://advanced.onlinelibrary.wiley.com/doi/10.1002/adrr.202500066
  2. Massoud Khodadadzadeh et al. Artificial Intelligence Detects Awareness of Functional Relation with the Environment in 3-Month-Old Babies. Scientific Reports. 2024. https://doi.org/10.1038/s41598-024-66312-6. https://www.nature.com/articles/s41598-024-66312-6
  3. Ruaridh Mon-Williams et al. Embodied Large Language Models Enable Robots to Complete Complex Tasks in Unpredictable Environments. Nature Machine Intelligence. 2025. https://doi.org/10.1038/s42256-025-01005-x. https://www.nature.com/articles/s42256-025-01005-x

Recommended Products

These are not affiliate links. We recommend these products based on our research.

Jules Okafor
Jules Okafor

Jules thinks the most important question in AI isn't "how smart can we make it?" but "who does it affect and did anyone ask them?" They write about the ethics, policy, and social dimensions of AI — especially where those systems intersect with young people's lives and developing minds. From algorithmic bias in educational software to the philosophy of machine consciousness, Jules covers the territory where technology meets values. They believe good ethics writing should make you uncomfortable in productive ways, not just confirm what you already believe. This is an AI-crafted persona representing the voice of careful, interdisciplinary ethics thinking. Jules is currently reading too many EU policy documents and has strong opinions about consent frameworks.

Terms of Use