BlackMind
Research
data scienceon-chain analyticslearnabilityleakage-free modeling

The Learnability of pump.fun Meme Tokens

About 97% of new pump.fun coins go nowhere; the eventual graduates separate from the pack within the first five minutes, and a leakage-free model flags them live.

Jovonni L. PharrGeorgia Cyber Warfare Range / BlackMindAugust 2026

In brief

A rigorous, leakage-aware study of what predicts success among pump.fun meme tokens, spanning population base rates, learnability experiments, and a deployed real-time scorer. The through-line: learnability is a resource-allocation problem — the payoff comes from choosing the right sub-population, feature cost, and horizon, not from tuning a model. It culminates in a launch-time predictor validated live on coins it never trained on.

Key results

  • Measured on the ~1.25M tokens with resolved on-chain outcomes (within a ~1.4M-token population), the graduation base rate is 0.95% (about 1 in 105); a market-cap-sorted sample inflates apparent success roughly 13x by showing only winners.
  • The base rate is non-stationary — graduation collapsed from ~7% in early 2024 to ~0.5% by late 2024 — so random splits leak the future and walk-forward validation is mandatory.
  • Framing beats tuning: same features and model, posing the problem on 'creators with a prior graduate' lifts ROC-AUC 0.66 to 0.72 and top-1% lift 2x to 6.2x.
  • Cheap metadata caps at ROC ~0.66 while on-chain microstructure reaches ROC ~0.91 — early holder count separates future graduates at AUC 0.85, stable at 0.79-0.90 within every snapshot-age band (a passed leakage audit).
  • True first-5-minute features on 714,747 tokens give a deployable launch-time model: gradient boosting ROC 0.95 with 23x lift in the top 1%; watch the top 10% of new coins and catch 80% of future graduates.
  • In a live 16-minute run, 398 new coins were detected and 228 scored at 5-minute maturity; all 3 coins flagged at P > 0.5 graduated, and the top 4 by score were 4/4 real graduates.

Reframing the Question

pump.fun is a token factory — roughly 34 launches per minute, with 78% of creators launching exactly one token and a single wallet spawning 1,104. Almost everything dies, which makes 'which of these goes far?' the only interesting question and a treacherous one to pose.

The natural target — classifying 'rugs' — turns out to be ill-posed. Every meme token eventually goes to zero, and a prior model's rug label thresholded a feature that was also one of its inputs: textbook leakage. This project discards that framing and studies the questions that are well-defined instead: graduation, magnitude, timing and survival, and creator reputation. The discipline is stated up front as non-negotiable: features drawn strictly from the window [0, T], labels strictly from after T, and every headline signal subjected to a leakage audit.

The launch firehose and creator power law: most creators launch once, a tiny fraction drive the flood.
Fig.The launch firehose and creator power law: most creators launch once, a tiny fraction drive the flood.

The Population Base Rate

Measured on the population rather than a winners' sample, graduation runs about 0.95% — roughly one in 105. A market-cap-sorted sample inflates that apparent success by about 13x simply by omitting the losers, and this single sampling bias is what sinks most meme-coin models.

The rate also does not hold still. Graduation collapsed as the platform saturated: about 7% in early 2024, down to roughly 0.5% by late 2024, and 1.03% in the February-2025 bulk that dominates the dataset. A regime shift that large means a random train/test split leaks the future outright, so validation is walk-forward by construction — train on earlier launch cohorts, test on strictly later ones. The 1.3M-token set spans 13 monthly cohorts for exactly this purpose.

Sampling bias versus the honest population base rate.
Fig.Sampling bias versus the honest population base rate.

Framing Is the Lever

A series of learnability experiments probes not how to tune a model but which framing of the problem is tractable at all. Using only leakage-safe birth-time features (creator history, socials, launch timing) and out-of-time validation, cheap metadata already beats the base rate: gradient boosting reaches ROC-AUC about 0.66 with roughly 2.4x lift in the top 1%.

The key result changes only the population the problem is posed on, holding features and model fixed. Narrowing to 'tokens by a creator who has graduated before' jumps ROC from 0.66 to 0.72 and lift from 2x to 6.2x — choosing the right sub-problem beats tuning the model. Feature-family ablation adds texture: creator history and socials carry nearly all the signal, while launch timing is close to noise, and the learning curve plateaus around 100k tokens, so weak features hit a ceiling that more rows cannot raise.

How the framing of the population changes learnability.
Fig.How the framing of the population changes learnability.

The Learnability Frontier

The synthesis is a map, not a leaderboard. Cheap metadata — creator, social, and timing — is available for every token but caps at ROC around 0.66. On-chain microstructure — holder counts, bot activity, churn — reaches ROC around 0.91 but costs an RPC pass per token and exists only for a small sample. You trade along this frontier rather than escaping it.

The microstructure signal is sharp and survives scrutiny. Early holder count separates future graduates at AUC 0.85, and early transaction-frequency is a strong inverse signal: churn without a growing holder base marks the doomed. A leakage audit rules out a snapshot-timing artifact — the AUC holds at 0.79-0.90 within every snapshot-age band, and graduates were actually snapshotted younger, so the signal is real traction. The tractable design that falls out is a two-stage funnel: a cheap metadata screen over all ~1.3M tokens, conditioning on a proven subpopulation, then an expensive microstructure model on the survivors, where ROC around 0.9 is affordable because the set is now small.

The learnability map: cheap-but-capped metadata versus expensive-but-strong microstructure.
Fig.The learnability map: cheap-but-capped metadata versus expensive-but-strong microstructure.

Closing the Leakage Gap

The microstructure results above were descriptive, drawn from lifetime aggregates contemporaneous with the outcome — separability, not a launch-time edge. The deployable result closes that gap. A join-based Dune query extracted true first-five-minute trade features for 714,747 tokens launched in the last 45 days, with real graduation resolved on-chain for every one (2.96%). Every feature is computable live from the trade stream.

Trained on earlier launches and tested on strictly later ones, the model is a genuine launch-time predictor: gradient boosting ROC 0.95, deep MLP 0.94, logistic 0.91, with 23x lift in the top 1%. The live tells are legible — early buy count (AUC 0.85), trade count (0.83), and climbing reserves (0.77) flag future graduates, while SOL-per-trade is inverse (0.23): a few large trades beat broad organic buying. The practical read is a watchlist — monitor the top 10% of new coins and you catch 80% of all future graduates, since the winners separate early (278 trades versus 35 in the first five minutes).

The deployable first-5-minute model, validated out-of-time.
Fig.The deployable first-5-minute model, validated out-of-time.

Caught Live

The same features are served in production. Because PumpPortal's free tier streams only new-token events, an Alchemy path subscribes to the pump program's logs and decodes the on-chain CreateEvent and TradeEvent anchor logs directly — mint, buy/sell, trader, virtual-SOL reserves — maintaining each coin's leakage-free first-five-minute window and scoring it at maturity. What is trained is exactly what is served.

In a 16-minute live run, 398 new coins were detected and 228 scored at 5-minute maturity. The score distribution was calibrated — median P around 0.0018, correctly treating noise as near-zero. Resolving the actual on-chain outcome of every scored coin, 6 of 228 graduated (2.6%, right at the true base rate), and the ranking held: all 3 coins flagged at P > 0.5 graduated, and the top 4 by score were 4/4 real graduates, on coins the model had never trained on. A single 16-minute slice is a demonstration rather than a backtest, but it is the full pipeline working live: detect, track, score, catch.

Live scoring over Alchemy: graduates popping above the watchlist line amid the noise cloud.
Fig.Live scoring over Alchemy: graduates popping above the watchlist line amid the noise cloud.

Life After Graduation

A companion track studies what happens after a token migrates to the PumpSwap AMM, isolating post-graduation trading by a reserve threshold since the trade table does not tag phase directly. Graduation turns out to be the exit event, not the finish line: median post-graduation lifetime is about 1.2 minutes, with Kaplan-Meier estimates of 88% dead within an hour and 95% within a day. Upside is a thin power law — median run just 1.35x, only 4.1% reaching 5x or more — and the same pattern recurs: concentrated, low-breadth activity runs further than broad organic buying, with rugs concentrating in a pump-and-dump archetype at roughly 70x the rate of other cohorts.

The project is candid about scope. The money model's win rates are real and out-of-time (90.7% at the top-1% gate) but its payoff multiples are documented assumptions, not measured returns, so the dashboard plots conservative, base, and moonshot scenarios explicitly. The top-scored coins have already pumped by the five-minute mark, meaning late entry and smaller realized multiples, and the honest follow-up is to re-ground the P&L on realized entry-to-exit returns. The larger lesson stands regardless: learnability is a resource-allocation problem — align signal, sample size, feature cost, and the actual decision, and validate walk-forward because the regime shifts.

How far tokens run is a power law: peak magnitude concentrates in a thin tail, so upside must be modeled on the tail, not the mean.
Fig.How far tokens run is a power law: peak magnitude concentrates in a thin tail, so upside must be modeled on the tail, not the mean.

Abstract

This project studies which pump.fun meme tokens actually go far, using leakage-free datasets built from market and on-chain data across roughly 1.4M unique tokens (about 1.25M with resolved on-chain outcomes). Rather than classifying unreliable "rug" labels (whose prior thresholded a feature that was also a model input — textbook leakage), it reframes around well-posed questions: graduation, magnitude, timing/survival, and creator reputation. Headline statistics come from unbiased population cohorts, never outcome-selected samples: measured on the ~1.25M tokens with resolved outcomes, the honest graduation base rate is about 0.95%, and it is non-stationary, collapsing from roughly 7% in early 2024 to about 0.5% by late 2024, which makes walk-forward validation mandatory. A central finding is that the framing of the learning problem, not model tuning, is the dominant lever: cheap birth-time metadata caps at ROC-AUC around 0.66, but posing the problem on creators with a prior graduate raises it to 0.72 with 6.2x top-1% lift, while expensive on-chain microstructure reaches ROC around 0.91. Closing the leakage gap, true first-five-minute trade features on 714,747 tokens yield a genuine launch-time predictor (gradient boosting ROC 0.95, 23x top-1% lift), served live over an Alchemy WebSocket that decodes on-chain events directly. In a 16-minute live run on unseen coins, all three flagged coins graduated.

More figures

  • Graduation collapsed as pump.fun saturated — the non-stationarity that forces walk-forward validation.
    Fig. 1Graduation collapsed as pump.fun saturated — the non-stationarity that forces walk-forward validation.
  • The live five-minute tells: buy count, trade count, and reserve climb flag graduates; SOL-per-trade is inverse.
    Fig. 2The live five-minute tells: buy count, trade count, and reserve climb flag graduates; SOL-per-trade is inverse.
  • Label-free anomaly detection: IsolationForest ranks graduates at ~12x base rate in its top 1% (AUC 0.94).
    Fig. 3Label-free anomaly detection: IsolationForest ranks graduates at ~12x base rate in its top 1% (AUC 0.94).