University of Virginia School of Medicine scientists have identified a widespread source of error in CUT&Tag (Cleavage Under Targets & Tagmentation), a popular method for studying the epigenome, and built a free machine-learning tool to correct it. The tool is designed to improve the reliability of both conventional and single-cell data generated using this method, giving researchers a clearer view of how gene activity is controlled in health and disease and a stronger foundation for future diagnostic and drug-development work.

Although nearly every cell in the body carries the same DNA sequence, different cells rely on different sets of genes to maintain their identities and functions. Much of that control comes from the epigenome, chemical modifications and structural features of chromosomes that influence whether genes are turned on or off without altering the DNA sequence itself. CUT&Tag can map these epigenomic features efficiently from very small samples, even individual cells. But a team led by Chongzhi Zang, senior author of the study published in Nature Communications, found a “hidden bias” in the method that produces artifacts resembling genuine biological signals.

“The DNA sequence is like the sheet music. The epigenome determines which notes to play, when they are played and by what instruments,” Zang said. “When a technical artifact looks like a real signal, researchers can be led toward the wrong biological mechanism. That risk is especially serious in single-cell data, where the true signals are already sparse.”

The bias stems from the assay’s reliance on an enzyme called Tn5 transposase, which naturally favors open, accessible regions of the genome. After examining nearly 300 published datasets, Zang and colleagues found this preference can create misleading results that they described as “severe.” 

Search Antibodies
Search Now Use our Antibody Search Tool to find the right antibody for your research. Filter
by Type, Application, Reactivity, Host, Clonality, Conjugate/Tag, and Isotype.

In response, the team developed PATTY (Propensity Analyzer for Tn5 Transposase Yielded bias), a machine-learning tool that corrects the CUT&Tag bias in both conventional multi-cell data and sparser single-cell data. “Detecting true signals in noisy data is like finding a needle in a haystack, and it is even more difficult when many pieces of hay look like real needles,” Zang said. “PATTY does not detect signals simply by subtracting a background. It learns how a real needle differs from hay and uses the learned model to reduce the artifact while preserving real signals.”

PATTY is available as a free, open-source package on GitHub and Zenodo. Zang said, “We believe that bias correction should become a routine part of CUT&Tag analysis. Cleaner data can keep scientists from wasting time and resources pursuing technical artifacts and can make real biological differences easier to see. More broadly, PATTY provides a conceptual framework for correcting similar biases in other genomic technologies.”