Classifying a surge: the order the questions get asked in
Surge classification fails far more often through bad ordering than through bad measurement. Ask the expensive behavioural question first and you will have built a narrative before the cheap structural fact arrives to contradict it. This note sets out the sequence this desk runs, why each step sits where it does, the two class boundaries that cause almost all of the trouble, and the shape a written classification has to take.
- Question
- In what order should the questions be asked when sorting a Solana volume surge into a class?
- Evidence used
- Recorded structural facts first, meaning pool creation, reserve changes and migration transactions, then behavioural distributions drawn from the same fixed window.
- Cannot show
- Which class is correct when two survive the sequence. The output is then two named candidates, not a choice between them.
- Falsified by
- A sample in which reversing the order, running behavioural tests before structural ones, produced the same class assignments at the same rate.
- Confidence
- Working. The ordering argument rests on how cheap and how decisive each check is, both of which are observable, but no measured comparison of orderings exists.
The short answer
Surge classification is a sequence, not a lookup table. The structural questions run first because they are cheap and, when they return something, decisive. The behavioural questions run afterwards because they are expensive to measure and every one of them carries an ordinary explanation that can defeat it on its own.
Getting that order wrong is the commonest failure in practice, and it is not a measurement problem. It is a narrative problem: an analyst who starts with the interesting behavioural question has usually built an explanation before the boring structural fact arrives, and explanations are extremely hard to dislodge once written down.
Why order matters more than method
Consider two analysts looking at the same window. The first checks whether pooled reserves moved, finds a large liquidity addition, and classifies the window as plumbing in about forty seconds. The second starts by measuring spacing, notices a suspiciously regular cadence, spends twenty minutes building a case for produced turnover, and only later discovers the liquidity event that explains everything.
Both analysts measured correctly. The second one produced a worse reading, and did so because of ordering alone. Worse, the second analyst now has a written argument to abandon, which is psychologically much harder than never having written it. Sequence is a defence against your own commitment, not just an efficiency measure.
There is a second reason. Structural facts are recorded on chain and can reach the Firm confidence band. Behavioural distributions are inferences and never reach it, however many of them agree. Running the checks that can produce a Firm result before the checks that cannot is simply putting the strongest available evidence first.
What would change my mind
If a sample of windows classified in reverse order produced the same assignments at the same rate as windows classified in this order, the ordering discipline would be ceremony. The falsifying observation is a paired comparison across a reasonable number of windows, which this desk has not run.
The sequence, step by step
Seven steps, run in this order every time. The first three are cheap enough to complete inside two minutes on a live window; the rest need the window to have closed or to have run long enough that the distributions mean something.
- Fix the window and the comparison periodWrite down a start, an end and the earlier period the window is being measured against, before looking at anything else. A window chosen after the data has been seen can support almost any conclusion without a single false sentence being written.
- Check the reservesDid pooled liquidity change during the window, and was a pool created or drained? This is a recorded fact rather than an inference. When it comes back positive it settles the class as plumbing and the sequence can stop.
- Check the venue spreadHow many programs carried the flow, and did one pool take nearly all of it? A surge that appears on one venue while a deeper venue stays quiet is a different event from one that appears everywhere at once.
- Measure end-of-window inventoryDid participants finish holding something different from what they started with? Inventory that returns to flat after heavy turnover rules out distribution and points at routing or produced flow. Inventory that ends strongly directional does the reverse.
- Measure address novelty and tail growthWhat share of trading addresses had never appeared in the pair, and was the address list still lengthening at the close? A crowd keeps producing addresses nobody has seen before. A fixed set of wallets exhausts itself.
- Measure spacing and size dispersionHow are the gaps between trades distributed, and how widely are sizes spread? This is where a configuration and a crowd diverge most, and also where the most ordinary explanations are available, so it runs late rather than early.
- Check lead-lag against the earlier moveDid this burst precede or follow the first one? An echo cannot precede what it echoes, so this is the only check that can eliminate the echo class outright, and it needs the comparison period fixed in step one.
The sequence is deliberately front-loaded with checks that can end it. Steps two and three between them resolve a meaningful share of windows with a recorded fact, and every window they resolve is a window nobody has to argue about later.
The two boundaries that cause trouble
Four of the six classes separate cleanly enough in practice. Two pairs do not, and knowing in advance which pairs are difficult is more useful than pretending the scheme is crisp.
| Boundary | What both produce | Best separator | Why it often fails |
|---|---|---|---|
| S4 Routing vs S5 Produced | Flat end inventory, a small recycled participant set, sizes inside a band | Whether a price gap existed and closed, and whether activity persisted once it had | Gaps across venues are hard to reconstruct after the fact without full coverage |
| S1 Attention vs S6 Echo | New addresses, directional inventory, wide-looking participation | Lead-lag against the earlier move in the comparison period | The comparison period often starts too late, cutting off the move being echoed |
| S3 Distribution vs S1 Attention | Directional inventory, real position change | Whether the flow concentrates against a few accounts or spreads across many | Concentration depends on grouping accounts, which has its own error rate |
The routing and produced boundary is the one worth understanding properly, because the two have almost the same footprint and completely different meanings. Intermediation exists to close a price difference and stops when the difference is gone. Produced activity has no such stopping condition; it stops when a budget does. Persistence in the absence of any gap is therefore the sharpest available separator.
The attention and echo boundary is easier to fix and easier to get wrong. It is fixed by choosing a comparison period that starts before the first move rather than at the exciting part. Analysts who begin watching at minute four routinely classify the echo as the event.
What sits behind the produced class
S5 is the only class with a purchasable supply chain, and understanding that chain sharpens the classification. Different tool classes produce visibly different footprints, and treating them as one category is why produced windows are so often misread.
The clearest distinction is between activity spread deliberately over time and activity packed into a single block. A tool that sequences swaps across many wallets on an interval leaves a spacing signature; a bundler that executes several transactions atomically leaves almost no spacing information at all, because everything lands in the same slot. A comparison of volume bot vs bundler behaviour is therefore a practical classification aid rather than a shopping exercise, since the two produce different shapes and require different checks.
The general point is that an automated Solana volume bot exposes a small number of parameters to an operator, and those parameters map directly onto the fields in step six. Wallet count constrains address novelty. Interval settings constrain spacing. A size band constrains dispersion. Knowing the parameter set is what makes a narrow band interpretable rather than merely odd.
Provisional rather than Working on the tool-class distinction, because it rests on how the software categories are described rather than on a measured comparison of their outputs. The atomic-versus-sequenced difference follows from how Solana transactions land, which is firmer, but the rest is reasoning about design rather than observation.
How a classification is written down
A classification that is not written down in a fixed shape drifts. The shape below is not bureaucracy; each line exists because leaving it out has a specific failure attached.
- The window, as a start and an end, with the time source named.
- The comparison period, stated separately, because the multiple depends on it entirely.
- Venue coverage, including the venues knowingly left out.
- The deduplication rule applied to multi-hop routes.
- Which steps of the sequence were run, including the ones that returned nothing.
- The class, or the two candidate classes, with no third option smuggled in as a hedge.
- The confidence band, with the reason it is not the band above.
- The single observation that would move the class, written before the prose.
- A plain sentence saying what the reading does not claim.
The falsifier line does the most work and is the most often skipped. Writing it before the explanatory prose forces a decision about what evidence would have been needed, which is a far harder question than describing what was found, and it has to be answered while the reading is still uncomfortable.
When undecidable is the answer
Undecidable is a finished result, not a failure to finish. It is the correct output when the fields that would separate the surviving candidates were not measurable in the window available, and it is common for windows read live.
The temptation is always to round undecidable up to the nearest interesting class, because undecidable is unsatisfying to publish and impossible to build a narrative on. That temptation is exactly why the band exists. A desk that never publishes undecidable is not being decisive; it is quietly converting absent evidence into confident readings.
Two situations produce it most reliably. Windows too short for address novelty and tail growth to have said anything, and windows where venue coverage is known to be incomplete in a way that could invert the reading. Both are stated in the record rather than glossed.
There is a third, less obvious case: windows where every field agrees but a single ordinary explanation accounts for all of them at once. Agreement among fields is only informative when the explanations that would defeat them are incompatible. If one scheduled rebalancing process would produce the regular spacing, the narrow size band and the flat inventory together, then three agreeing fields are really one observation counted three times, and the correct output is undecidable rather than a confident produced reading.
Naming two classes at once
When two classes survive the sequence, both are named. This looks like indecision and is the opposite: a reading that lists two live candidates has said something specific about four classes that were eliminated, and it tells a reader exactly which further observation would resolve it.
The failure mode it prevents is worse than it looks. Forced to pick one of two surviving candidates, an analyst will pick the more quotable, and the more quotable is almost always the one implying somebody did something deliberate. Over time that bias accumulates into a record that systematically overstates deliberate activity.
What would change my mind
If readings that named two classes were shown to be resolved correctly less often than readings that committed to one, the practice would be costing accuracy rather than protecting it. That comparison needs windows with independently known causes, which this desk does not have.
What the sequence cannot fix
Ordering the questions well does not manufacture evidence. If venue coverage is incomplete, running the steps in the right order produces a well-organised reading of the wrong data. The sequence protects against narrative bias and wasted effort, not against missing rows.
Nor does it help with a window that is still open. Steps five and six both need the distribution to have stabilised, and a distribution measured on the first ninety seconds of a burst is dominated by whichever participants happened to be fastest. That is a real limitation rather than an inconvenience, and it is why the triage procedure for live windows stops at the structural checks and says so plainly instead of pretending the behavioural fields are ready.
It also cannot reach intent. Every step above measures what accounts did. None of them touches why, who instructed it, or whether anybody involved considered it legitimate. A window that lands squarely in the produced class has been described, not accused, and the distinction is maintained deliberately throughout this site.
Finally, the sequence assumes the six classes are adequate. They are a working scheme, not a discovery, and a surge shape that none of them handles is a genuinely useful finding. The scheme is meant to be revisable, and a case that breaks it is worth more than a case that confirms it.
Questions the desk gets asked
Why not run all the checks at once and weigh them together?
Because the checks are not equally strong and weighing implies they are. A recorded pool creation inside the window is a fact; a narrow size band is a tendency with several ordinary explanations. Combining them into one score hides that difference and produces a number that looks more precise than the evidence underneath it. Running them in order keeps the strong evidence visibly strong.
What if the structural checks come back empty?
Then the class is behavioural by elimination, which is a weaker starting position and should be recorded as such. An empty structural result does not raise confidence in whatever comes next. It only means the cheap decisive route was unavailable and the reading now rests entirely on distributions that each carry an ordinary explanation.
How many fields should agree before a class is written down?
There is no threshold, and a threshold would be a false precision. What matters is whether the ordinary explanations for the fields that agree are compatible with each other. Three fields agreeing where one explanation accounts for all three is weak. Two fields agreeing where no single ordinary explanation covers both is stronger, despite being fewer.
Can a window change class as it develops?
Yes, and that is normal rather than a failure. A window classified on its first ten minutes may look different across forty, because address novelty and tail growth both need time to say anything. This is why each classification is recorded with its window, and why a later reassignment is published as a correction rather than treated as an embarrassment.
Does the sequence work for very small pairs?
Less well, and the reason is worth stating. On a thin pair, a handful of participants generate the entire distribution, so spacing and size dispersion are estimated from very few observations and are unstable. The structural checks still work normally. The behavioural ones should be reported with an explicit note that the sample was small.
Is there a shortcut for a live spike with no time to run the sequence?
The first two steps take under a minute between them and remove most of the risk of an embarrassing reading. Fix the window and the comparison, then check whether pooled reserves moved. Everything after that is refinement. A reading that stops after those two steps and says so is far more useful than a complete-sounding reading built in the wrong order.
Why does the desk publish undecidable results at all?
Because the alternative is publishing only the windows that happened to resolve, which quietly biases everything a reader sees towards cases where the fields lined up. Undecidable results are the honest majority outcome for short windows, and printing them keeps the visible record representative of what the method actually produces.
Filed in Surges by The Surge Watch Desk. Patterns described here come from protocol design and from public transaction data; every figure inside a worked example is invented, labelled as invented, and describes no real pair. How classes are defined and how confidence is worded is set out in the method note.