Summary

A bioRxiv preprint reports that reinforcement learning identified previously unreported ways to initiate reentry in models of excitable tissue. Learned stimulation patterns produced unidirectional propagation in cardiac monolayers, while whole-heart experiments offered preliminary evidence that they can influence early wave propagation.

A reinforcement-learning system has identified new ways of initiating reentry—the sustained circulation of excitation through excitable tissue—in a study posted as a bioRxiv preprint on September 14, 2026. The researchers found that stimulation patterns using individually subthreshold inputs could generate unidirectional propagation and reentry in computational models, then reproduced one class of the learned patterns in cardiac monolayers.

The work also reports preliminary whole-heart experiments in which the patterns influenced early propagation. It presents reinforcement learning as a tool for discovering mechanisms and generating experimental hypotheses, rather than as a clinical method for treating arrhythmias.

Contents

Why reentry matters in excitable tissue

Excitable media are systems in which activity can spread from one region to another. Cardiac tissue is one example: electrical excitation normally propagates across the heart and then subsides. Under some conditions, a wave can instead become unidirectional and circulate through tissue, creating reentry. Reentrant activity is a mechanism underlying many life-threatening cardiac arrhythmias.

The transition from a temporary wave to sustained reentry is difficult to explain because it depends on the timing and location of stimulation, the recovery state of nearby tissue and the geometry of the system. A stimulus that is too weak to activate tissue on its own may still alter local conditions in a way that affects a later wave.

The researchers treated this transition as a search problem. The objective was not simply to produce any excitation, but to find stimulation sequences that produced sustained activity while using as few stimuli as possible.

How the reinforcement-learning search worked

In reinforcement learning, an agent explores possible actions and receives feedback from the consequences. Here, the agent selected sequences of spatial stimulation patterns. It received a reward for sustained activity and a penalty related to the number of stimuli applied.

The team tested the approach with cellular automata in one-, two- and three-dimensional geometries. A cellular automaton represents a system as linked sites whose states change according to defined rules. This allowed the researchers to vary the structure of the excitable medium and examine how the learned strategies behaved in different settings.

The search recovered a previously described mechanism that combines superthreshold and subthreshold stimulation. A superthreshold input is strong enough to trigger activity directly, while a subthreshold input is weaker than that local activation threshold. The system also found two mechanisms based entirely on subthreshold stimulation.

One was sequential: subthreshold stimuli were delivered at different locations and times, with their combined effects producing the conditions for unidirectional propagation. The other was spatial: several sites that were individually subthreshold collectively initiated reentry.

In geometries containing boundaries and branches, the learned protocols also used structural source–sink asymmetries. In an excitable medium, an active region acts as a source of excitation while surrounding tissue forms a sink that receives it. Differences in geometry can change this balance, affecting whether a wave dies out or spreads.

New mechanisms and experimental tests

The computational findings were followed by optogenetic experiments in cardiac monolayers. Optogenetics uses light-sensitive proteins to control electrically excitable cells with light. In these experiments, the learned spatial subthreshold patterns reproducibly induced unidirectional propagation.

The researchers also carried out whole-heart experiments. These provided preliminary evidence that the learned patterns could shape early propagation in intact cardiac tissue, extending the test beyond simplified computational geometries and cell layers.

The significance of the study lies in the route to the mechanisms. Instead of specifying every possible initiation pathway in advance, the reinforcement-learning agent searched across timing, location and geometry and returned stimulation protocols that could be tested experimentally. This may provide a general way to study reentry in systems whose structure is more complicated than a regular laboratory model.

The report is a bioRxiv preprint and has not yet undergone peer review. Its evidence combines cellular-automaton simulations, cardiac monolayer experiments and preliminary whole-heart experiments; it is not a clinical study. The frequency of these mechanisms in biological hearts, and their relevance to particular human arrhythmias, remain subjects for further research.

Sources