A protein designed with artificial intelligence doubled targeted DNA insertion in cultured cells when incorporated into an experimental gene-editing system. The result emerged from a search through thousands of genomes, followed by laboratory tests of natural and synthetic proteins.
The study, involving Integra Therapeutics, Pompeu Fabra University and the Center for Genomic Regulation, appeared in Nature Biotechnology. It examines proteins called transposases, which move DNA between locations.
The findings expand the available tools for inserting genes, including into human immune cells. They also expose an important distinction: improving one step of gene editing does not necessarily improve every application.

PiggyBac transposases can cut DNA cargo out of a donor molecule and insert it into a genome. That capacity makes them useful for experiments requiring the addition of substantial DNA sequences.
However, conventional PiggyBac does not simply place genes wherever scientists choose. It favors sites containing the four-letter DNA sequence TTAA, which limits its targeting precision.
Much of the field has relied on a relatively small collection of known proteins. The team investigated whether organisms already carried a broader range of useful alternatives.
Researchers searched 31,565 eukaryotic genome assemblies in the National Center for Biotechnology Information database. They also examined PiggyBac sequences in Dfam, a database of mobile genetic elements.
The initial search identified 273,643 potential protein-coding sequences. Filtering for features associated with DNA movement narrowed the collection to 116,216 potentially functional elements, grouped into 13,693 subfamilies.
Those groups represent computational candidates, rather than thousands of experimentally verified enzymes. Their breadth nevertheless offered a much larger starting point for selecting and testing proteins.
The sequences came from organisms including insects, fungi, plants and mammals. About 60% were insect-derived, and the analysis separated the collection into five major groups.

The team selected 23 representative natural sequences for initial experiments in HEK293T cells, a widely used human cell line. Tests measured DNA excision and insertion using fluorescent reporters.
Nine candidates showed detectable activity in those initial tests. Two performed comparably to hyperactive PiggyBac, or HyPB, a version previously improved through laboratory evolution.
The results showed that useful activity could occur in proteins with substantially different sequences. They also demonstrated why a promising computational match still needs a laboratory test.
One candidate assembled from bat-derived sequences, for example, was inactive in these experiments. A different bat PiggyBac sequence had worked in earlier research, illustrating how sequence differences can affect performance.
The researchers further modified two leading proteins by removing sites associated with inhibitory phosphorylation. That change increased their activity, combining natural diversity with conventional protein engineering.
One newly identified protein, called Poetur, also performed well in primary human T cells. These immune cells are important targets for developing cell therapies, making their inclusion a useful test beyond an established laboratory cell line.

The next stage used the natural sequence collection to guide synthetic protein design. Researchers fine-tuned ProGen2, a protein language model, with more than 13,000 bioprospected sequences.
Such models work with amino acid sequences, the building blocks of proteins. Here, the team used the model to generate alternatives related to the already optimized HyPB protein.
Two model versions generated sequences in opposite directions along the protein chain. Each received a short stretch from one end of HyPB to provide a starting context.
The models produced more than 100,000 sequences. Computational filters assessed basic protein properties, PiggyBac-specific features and predicted structures before the team selected 22 candidates for experiments.
Those candidates differed from HyPB by 15 to 54 amino acid substitutions. All 22 showed excision activity, and seven were significantly more active at excising DNA than the laboratory-evolved comparator.
The experiments supplied evidence beyond favorable computer scores. They showed that selected synthetic sequences could function inside cells, although the selection process did not establish that every generated sequence would work.

The researchers also tested compatibility with FiCAT, a system that combines Cas9 with an engineered PiggyBac transposase. Cas9 cuts DNA at a chosen genomic site, while the transposase releases the DNA cargo from its donor molecule.
The cargo is then inserted at the break through a cellular DNA-repair process. Engineering the PiggyBac component changes its behavior from the conventional, freely integrating version.
Synthetic sequence 3277 improved targeted insertion twofold in the FiCAT experiments. The result demonstrates that an AI-generated protein could strengthen a component of a programmable insertion system.
The team also tested targeted insertion with leading synthetic candidates in mouse muscle precursor cells. Those experiments examined the TTR and PCSK9 genomic locations, extending testing beyond the original human cell line.
These are measurements of laboratory gene insertion, not evidence that a treatment improved health. They establish compatibility and activity under the tested conditions, while leaving therapeutic performance to future research.
Results in primary human T cells underscored how performance depended on the experiment. Researchers inserted DNA carrying a green fluorescent protein reporter to track successful integration.

Poetur and synthetic sequence 136 achieved higher nontargeted integration than HyPB. Sequence 3277, despite its stronger excision and targeted insertion results elsewhere, had similar nontargeted integration activity to HyPB in these cells.
That variation matters when choosing a protein for a particular editing job. An enzyme that releases cargo efficiently may not be the strongest performer when insertion is measured in another cell type.
The study therefore offers several candidates with different properties, rather than one universally superior replacement. Its natural sequence collection also provides material for investigating why those differences arise.
Cancer cell therapies and gene therapies for rare diseases are potential applications for improved gene-insertion tools. The experiments reported here did not establish a clinical benefit in either setting.
Higher activity also does not, by itself, demonstrate greater precision or safety. The authors identify the effect of AI-guided improvements on specificity as a crucial question for therapeutic protein development. That requires evaluating where DNA is inserted, alongside how often insertion occurs, before the new proteins can be assessed for a therapeutic role.
The findings support a discovery strategy that combines biodiversity searches, computational design and experimental validation. For now, its clearest achievement is a broader set of working DNA-moving proteins and improved insertion in selected laboratory tests.
These resources explore protein language models and experimental approaches to placing large DNA sequences into genomes.
ProGen2: Exploring the boundaries of protein language models: Introduces the protein language models underlying the sequence-design approach used in this research. (Cell Systems, 2023)
Design of highly functional genome editors by modelling CRISPR–Cas sequences: Examines AI-generated genome editors and experimentally evaluates their editing performance. (Nature, 2025)
Simulating 500 million years of evolution with a language model: Demonstrates a language-model approach to generating functional proteins beyond known natural sequences. (Science, 2025)
Find and cut-and-transfer (FiCAT) mammalian genome engineering: Describes the combination of Cas9 targeting and engineered PiggyBac cargo processing used for programmable DNA insertion. (Nature Communications, 2021)
Drag-and-drop genome insertion of large sequences without double-strand DNA cleavage using CRISPR-directed integrases: Presents a different experimental strategy for inserting large DNA payloads without creating double-strand breaks. (Nature Biotechnology, 2023)
Research findings are available online in the journal Nature Biotechnology.
The original story “Scientists use generative AI to build better proteins for editing DNA” is published in The Brighter Side of News.
Like these kind of feel good stories? Get The Brighter Side of News’ newsletter.
The post Scientists use generative AI to build better proteins for editing DNA appeared first on The Brighter Side of News.
Leave a comment
You must be logged in to post a comment.