Mixer Tornado Cash: What Analysts Can Still See On-Chain
Prepared by the editorial team. Updated August 31, 2026.
Research Notice: This guide is part of our fintech research series examining blockchain privacy tools and their regulatory context. It is informational and educational only, is not legal, financial or compliance advice, and does not endorse or instruct the use of any mixing service. Laws differ by jurisdiction and change over time; verify current rules for your location.
Mixer Tornado Cash runs on a public blockchain, so every deposit and every withdrawal is recorded permanently and can be read by anyone. What the protocol conceals is the correspondence between one deposit and one withdrawal, not the existence, size or timing of either. This article sets out what remains visible, how analysts reason from it, and why those inferences carry more uncertainty than the language around them implies.
What does the public ledger retain after a mixing transaction?
Everything except the link. The ledger keeps the depositing address, the amount, the block number and its timestamp, the fee paid, the contract that was called, and later the withdrawal with its recipient address and its own timing. The protocol removes one specific relationship from that record and removes nothing else at all.
Blockchains are append-only. Nothing written is deleted or amended, and every full node holds the same history, so analysis of this data is not privileged access and not surveillance in the conventional sense. It is reading a document published deliberately and designed to be replicable by anyone.
The contract’s own bookkeeping is public too. A deposit writes a commitment into the contract state and a withdrawal publishes a nullifier, and both values are visible to any observer. What neither value discloses is which commitment a given nullifier corresponds to, since the derivation runs one way and the proof asserts membership without naming the member.
Withdrawal transactions submitted by a relayer are equally public. A relayer is a third party that broadcasts the transaction and pays its network fee in exchange for a fee taken from the withdrawal, and such services tend to be persistent and recognisable. The submitting party is therefore visible even where the beneficiary relationship is not.
What can be inferred from amounts, timing and counterparties?
Four families of signal survive intact: the amounts and denomination patterns involved, the timing of each event, the counterparties sitting on either side of the pool, and behavioural regularities such as repeated intervals or characteristic transaction settings. None of these requires defeating any cryptography, because none of them was ever encrypted.
Amounts are the most immediately useful. Fixed denominations make the sizes themselves uninformative within a pool, but the combination of denominations in play, and the totals moving in and out over a period, can form a distinctive shape. Aggregate flows are simply arithmetic on public numbers.
Counterparties matter more than the pool itself. Funds usually arrive from and depart to addresses that already carry some identity, whether an exchange deposit address or a service that publishes its wallets. The endpoints are often the most informative part of the picture, and they sit outside anything the protocol governs.
How does heuristic clustering work at a conceptual level?
A heuristic is a rule that assigns probability rather than certainty. An analyst proposes a pattern that would usually indicate common control or a shared origin, applies it across the ledger, and produces a set of candidate links. The output is a ranked hypothesis, and its quality depends entirely on how well the rule matches real behaviour.
The classical example comes from older chains, where inputs spent together in one transaction were treated as probably belonging to one owner: usually right, occasionally badly wrong. On an account-based chain the analogous rules look at funding relationships, contract interactions and repeated activity between addresses.
Heuristics are then chained. One rule groups addresses, a second attaches an entity name from off-chain sources, a third propagates a risk characteristic across the group, and a fourth scores the result. Each stage inherits the uncertainty of the one before it, and the compounded error is rarely shown to whoever reads the output.
This is also how a probability becomes a label. Analytical products present the end result as a category attached to an address, and the category travels into compliance systems and into public commentary without its confidence interval. Treasury’s designation record, published alongside the August 2022 OFAC action notice, drew on this kind of attribution work, and researchers have contested parts of it since.
Why do clustering heuristics produce false positives?
Because a rule describing typical behaviour will also match atypical but entirely innocent behaviour. Two unrelated people can deposit the same denomination minutes apart. An address can receive funds it never asked for. When the underlying rate of illicit activity is low and transaction volume is high, even an accurate rule generates many wrong matches.
The arithmetic is worth spelling out because it is unintuitive. Suppose a rule is correct ninety-nine times in a hundred and one address in ten thousand is genuinely of concern. Applied to a million addresses, it flags the hundred genuine cases and also roughly ten thousand innocent ones. The overwhelming majority of alerts are wrong even though the rule performs well, and no improvement in the rule short of near-perfection changes that picture.
The consequences fall on people with no involvement in whatever prompted the rule. Unsolicited incoming transfers can attach a history to an address its owner never chose. Accounts get restricted and the person affected is often given a conclusion rather than a reason. False positives are not a marginal defect here; they are the normal operating condition.
How can you assess an analytics label attached to an address?
You assess it by identifying who produced it, finding the rule that triggered it, establishing how many transaction hops it spans, asking for an error rate rather than a confidence score, and escalating anything unresolved to a qualified professional. The procedure is for reviewing a claim someone else has made, not for interacting with any protocol.
Step 1: Identify who produced the label
Establish which firm or dataset produced the label and what method they publish about it, because a classification travels through downstream systems long after its origin has been forgotten. A label repeated by four systems that all bought it from one vendor is a single claim, not four.
Step 2: Find the rule behind the label
Ask which specific pattern triggered the classification, since a term such as mixer-linked can describe direct interaction, a single intermediate hop, or a distant statistical association. Those three situations differ enormously and are frequently displayed with identical wording.
Step 3: Ask how many hops it covers
Establish how many transaction hops separate the address from the activity being described, because risk described as inherited weakens sharply with distance and many systems do not display that distance. An address five hops away from a flagged event has almost certainly had no contact with it.
Step 4: Ask for the error rate, not the score
Request the false-positive rate the method produces on realistic data rather than the confidence score attached to a single result, because a score is an output of the model and not a measurement of it. A provider unwilling to describe its error behaviour has told you something useful anyway.
Step 5: Escalate an unresolved label
Bring a label you cannot resolve to a compliance or legal professional who can weigh it against your actual obligations, because a vendor classification is an input to a decision rather than the decision itself. Documenting the review and its date is part of a defensible process.
Visible signals and what each one can support
Different signals carry very different evidential weight, and that distinction is usually lost once they are aggregated into a single score. The table pairs each category of visible data with the strongest conclusion it can bear alone.
| Visible signal | What it can reasonably support |
|---|---|
| Deposit and withdrawal amounts | Whether a transaction fits a denomination pattern, and nothing about who sent it |
| Block timestamps | A bounded window of candidate origins, narrowing sharply when activity is sparse |
| Funding and receiving counterparties | Which identified venues sat on each side, but not the path between them |
| Repeated behavioural patterns | A hypothesis of common control, weaker as the user population grows |
| Relayer and fee details | That a third party submitted the transaction, which says little about the beneficiary |
Read as a set, the rows show why careful analysts call their output leads. Each signal supports a narrow statement, and stacking narrow statements yields a hypothesis that still needs corroboration from outside the ledger.
Frequently asked questions
Can an analyst prove which deposit funded a specific withdrawal?
Not from the chain alone in the general case, because the protocol publishes no link and the proof reveals nothing about which commitment was spent. What analysis can do is shrink the candidate list, sometimes to a single plausible option when a pool is nearly empty. A shrunken list is an inference rather than a demonstration.
Does receiving funds that touched a mixer make an address culpable?
Receiving funds is not by itself an act that carries intent, and a recipient generally has no control over what a sender did beforehand. Culpability in most legal systems turns on knowledge and purpose, not on transaction history alone. Anyone facing a concrete situation of this kind should raise it with qualified counsel rather than reasoning from the label.
Do analytics providers disagree with one another?
Yes, and the disagreements are informative. Different providers use different rules, different label taxonomies and different off-chain datasets, so the same address can carry conflicting classifications across products. A single provider agreeing with itself is not corroboration, which is why serious reviews compare independent sources.
Did the March 2025 delisting change how these addresses are labelled?
Not directly. Sanctions listing and commercial risk labelling are separate systems, and analytics products classify by observed behaviour rather than by list membership. Institutions run risk-based programs that continued to treat mixer-associated funds as elevated risk after the removal, so the label and the legal status can point in different directions.
