A new dual-process theory solves the mystery of dopamine ramps

Researchers have developed a new computational model to explain why dopamine levels steadily increase as individuals approach a predictable reward. By combining two distinct types of learning, the model demonstrates how the brain efficiently updates its expectations, resolving a long-standing puzzle in neuroscience. The findings were published in the journal eLife.

Dopamine is a chemical messenger in the brain associated with learning, motivation, and movement. For decades, the dominant framework in neuroscience has viewed dopamine as a signal for a reward prediction error. This error occurs when the outcome of a situation is better or worse than expected. If an animal receives an unexpected treat, dopamine neurons fire to signal a positive error. This signal helps update stored expectations in a brain region called the striatum. Over time, as a reward becomes entirely predictable, these prediction errors should drop to zero.

However, experiments measuring dopamine during spatial navigation tasks have revealed a pattern that contradicts this standard theory. As animals move closer to a known, predictable reward, their dopamine levels gradually climb in a continuous slope. Because the reward is already expected, traditional mathematical models struggle to explain why a prediction error would increase as the goal gets nearer.

Researchers Luke Priestley and Thomas Akam from the University of Oxford sought to resolve this contradiction. They built a computational model that features two separate learning processes working in tandem. The first process is the traditional, slow-learning system that relies on cached values stored in the basal ganglia. The second process is a fast, flexible system that actively infers values using an internal map or world model, likely housed in the brain’s frontal cortex.

Priestley and Akam proposed that these two systems interact in a very specific way to generate dopamine ramps. When the brain calculates a reward prediction error, it compares its current prediction against a new update target. In their model, the fast, inferred values only influence the update target. The current prediction relies entirely on the slow, cached values. Because the fast system already knows a reward is near while the slow system is still catching up, the gap between the update target and the prediction grows as the goal approaches. This growing gap produces the steady climb in dopamine.

The researchers first tested this asymmetrical dual-process model in a simulated linear track environment. They compared it against a standard model and a version where inferred values influenced both the prediction and the update target. The asymmetrical model learned the true value of the environment faster than the alternatives. It also successfully produced the ramping dopamine signals that the standard models failed to generate.

Next, Priestley and Akam simulated an environment where an artificial agent navigated between high and low rewards over thousands of trials. They modeled a previous experiment showing that dopamine ramps in mice diminish gradually after extensive training. The simulated agent replicated this long-term decline. As the slow-learning cached values eventually matched the fast-learning inferred values, the gap between them closed, causing the ramps to flatten over time.

The model also mirrored how dopamine behaves in completely novel environments. In biological experiments, animals do not show dopamine ramps the first time they explore a new maze, but the ramps appear quickly after a few successes. The simulated agents showed this exact rapid onset, demonstrating how the fast-learning internal map quickly shapes the prediction error.

The researchers then applied their model to a grid-like environment with multiple paths to a single destination. In real-world experiments, changing the amount of reward at a specific location instantly alters the dopamine ramp on the very next attempt, even if the animal takes a completely different route. The dual-process model successfully reproduced this global updating behavior. Because the fast-learning system uses a flexible mental map, it immediately applied the new reward information to all possible paths leading to that goal.

To test how unexpected events influence dopamine, the team simulated virtual reality experiments where animals were suddenly teleported closer to a goal or forced to move at different speeds. In the simulation, teleports caused sudden spikes in the simulated dopamine signal, with the size of the spike depending on how close the agent was teleported to the reward. Changing the speed of the agent altered the steepness of the ramp. These simulated responses matched actual biological recordings, supporting the idea that dopamine tracks momentary changes in expected value.

Finally, the researchers modeled spatial uncertainty by simulating a virtual reality task where the environment progressively darkened. In actual animal experiments, this darkening causes dopamine levels to rise in a hump shape rather than a steady ramp. The simulated agent produced these exact same shapes. As the visual environment darkened, the agent became less certain of its exact location, which distorted the fast system’s inferred value estimates and caused the prediction error to drop off before reaching the goal.

While the dual-process model unifies several puzzling observations, it relies on a few computational simplifications. The researchers assumed the model-based system focuses entirely on calculating the shortest path to a single, final goal. In reality, animals continue to behave and learn after a goal is reached, meaning the brain likely employs more generalized strategies.

The simulations also used a fixed parameter to arbitrate between the fast and slow learning systems. A biological brain likely adjusts this balance dynamically based on confidence, uncertainty, or past experience. The model also assumes that the distances between locations are known to the agent in advance, which may require separate navigation circuits to already be active.

Future research will need to verify the biological pathways that allow the frontal cortex to send these fast value inferences to dopamine-producing centers. By testing whether temporarily disabling specific brain circuits eliminates dopamine ramps, scientists could test if this dual-process architecture operates in living animals. Identifying these physical connections would reshape how scientists view the boundary between conscious planning and automatic habit formation in the brain.

The study, “Dopamine ramps as a normative consequence of dual-process control,” was authored by Luke Priestley and Thomas Akam.

Leave a comment
Stay up to date
Register now to get updates on promotions and coupons
Optimized by Optimole

Shopping cart

×