A dog in a bathtub became a similarly colored goat. A cake turned into three sandwiches. These mistakes reveal the limits of an AI system that can otherwise reconstruct surprisingly detailed images from the brain activity of someone viewing them.
Called Brain-IT, the system comes from Michal Irani’s laboratory at the Weizmann Institute of Science in Rehovot, Israel. Its reconstruction study appeared in the proceedings of the 2026 International Conference on Learning Representations. Related work on a universal brain encoder was presented at the Conference on Cognitive Computational Neuroscience.
Together, the projects tackle two directions of the same problem: predicting brain responses to images and reconstructing images from those responses. The demonstrated task concerns pictures viewed during scanning, while applications involving dreams or patient communication remain research ambitions.

The work draws on the publicly available Natural Scenes Dataset, an unusually intensive collection of brain recordings. Eight carefully selected volunteers viewed thousands of photographs during repeated functional magnetic resonance imaging, or fMRI, sessions.
Each participant saw 9,000 to 10,000 distinct scenes across 30 to 40 sessions distributed over a year. Aggregated across participants, the dataset contains responses to 70,566 distinct images. Its depth provides extensive examples of how individual brains respond to visual information.
The recordings used a seven-tesla scanner and 1.8-millimeter spatial resolution. These measurements divide the brain into small volume units called voxels, rather than recording individual neurons.
Functional MRI measures blood-oxygenation changes associated with neural activity. It therefore provides an indirect signal, which researchers use to learn relationships between recorded responses and displayed pictures.
The volunteers were screened for data quality before joining the full experiment. That helped produce unusually useful recordings, but eight extensively sampled individuals still represent a small participant group. The dataset emphasizes depth within individuals rather than population-wide variation.
Earlier reconstruction systems could produce the right general subject without faithfully recovering its appearance. Recognizing bananas, for example, does not establish their position, shape or surrounding scene.

Brain-IT separates those challenges into complementary branches. One predicts semantic features, which describe the picture’s content. The other predicts structural features that help preserve its coarse layout, colors and outlines.
These predictions guide a diffusion model, which generates an image by progressively removing noise. The structural branch helps establish the starting arrangement, while the semantic branch steers the output toward appropriate objects.
The system also combines information from groups of voxels with similar functions. Its Brain-Interaction Transformer lets information pass between these groups before predicting localized image features.
The distinction matters because a convincing generated picture can still differ from what the participant actually saw. The objective is faithful reconstruction, including spatial relationships, rather than simply producing a plausible example of the right category.
Collecting enough paired photographs and brain recordings remains a major constraint. Every additional measured response requires a person, a scanner and experimental time.
The researchers addressed that shortage using an encoder that predicts fMRI responses from an image. A decoder performs the reverse task, estimating visual information from brain responses.

This creates a training loop that can use pictures never shown inside a scanner. An image enters the encoder, which predicts a brain response. The decoder then attempts to recover information about the original picture from that predicted response.
Comparing the reconstructed information with the starting image helps train the system. Irani reports that around 70% of the training data came from images without originally paired fMRI recordings.
Those extra examples expand the available training material, but their predicted responses are not new measurements of human brains. The process relies on relationships learned from actual recordings. Keeping that distinction clear prevents synthetic training data from being mistaken for additional experimental observations.
Brain anatomy and recorded responses differ between people, complicating attempts to reuse a model across participants. The team’s approach shares much of its learned machinery while adapting to a new person’s recordings.
In the reported tests, Brain-IT produced reconstructions using one hour of subject-specific fMRI data. The authors compared those results with earlier approaches trained on full 40-hour recordings and found comparable performance.
One-hour adaptation itself is not unprecedented. MindEye2, published in 2024, also demonstrated reconstruction with one hour of data from a new participant. Brain-IT’s experiments included comparisons with that system and MindTuner under the same limited-data condition.

For the full-data benchmark, the authors reported leading performance on seven of eight evaluated metrics. The main comparisons averaged results across four Natural Scenes Dataset participants, rather than representing a broad clinical evaluation.
Different measures assess different properties, including image structure and semantic resemblance. Strong performance across them supports the reconstruction claim, but it does not mean the system reproduces every picture exactly.
The associated universal encoder predicts responses across people and datasets. Its learned representations group voxels by functional similarity, offering another way to investigate visual processing.
Some groups respond preferentially to categories such as food, faces, text and outdoor scenes. These patterns may help researchers examine what visual information different brain regions represent. They do not establish that every person’s brain has identical anatomical boundaries for those functions.
Irani hopes to extend the work to imagined images, dreams, video and audio. Helping people who cannot move or speak communicate is another proposed application, alongside investigating experiences such as PTSD flashbacks.
Those possibilities have not been demonstrated by these image-viewing experiments. Reconstructing a photograph presented during scanning differs from recovering a private memory or translating an intended message.

“That’s something we don’t have yet,” Irani said of the hoped-for extensions. The present results offer a tool for studying perception, with further work needed to determine what other mental experiences it can capture.
The possibility of broader decoding has drawn attention from neuroethicists. Concerns center on whether future systems could extract mental information beyond what a person knowingly agreed to share.
Current fMRI research requires substantial equipment and participant involvement. These experiments do not demonstrate covert access to arbitrary thoughts or memories. Moving similar approaches to electroencephalography, or EEG, would require separate evidence that those signals support the intended decoding task.
EEG records electrical activity through electrodes and is being explored as another route to brain decoding. Experts warn that more accessible recordings could create consent and privacy risks if models infer additional information from them.
For Brain-IT, the immediate achievement remains narrower and measurable: better reconstruction of viewed images under controlled conditions. Its mistakes, calibration requirements and limited participant sample remain central to judging what that achievement means.
These resources examine visual reconstruction, language decoding, the underlying brain dataset and ethical safeguards for neurotechnology.
A massive 7T fMRI dataset to bridge cognitive neuroscience and artificial intelligence: Describes the Natural Scenes Dataset, including its intensive sampling, participant selection and imaging methods. (Nature Neuroscience, 2022)
MindEye2: Shared-Subject Models Enable fMRI-To-Image With 1 Hour of Data: Presents an earlier approach to transferring image reconstruction models across participants with limited calibration data. (Proceedings of Machine Learning Research, 2024)
Semantic reconstruction of continuous language from non-invasive brain recordings: Investigates language reconstruction from fMRI and examines the role of participant cooperation. (Nature Neuroscience, 2023)
Generative language reconstruction from brain recordings: Explores using decoded brain representations directly within a language model’s generation process. (Communications Biology, 2025)
Recommendation on the Ethics of Neurotechnology: Sets out ethical principles addressing consent, mental privacy and responsible uses of neurotechnology. (UNESCO, 2025)
Research findings are available online here.
The original story “New AI can reconstruct what you’re looking at by reading brain activity” is published in The Brighter Side of News.
Like these kind of feel good stories? Get The Brighter Side of News’ newsletter.
The post New AI can reconstruct what you’re looking at by reading brain activity appeared first on The Brighter Side of News.
Leave a comment
You must be logged in to post a comment.