Training a Text-to-Video Model from Scratch on Two GPUs
BlackMind's first video model: a 190M diffusion transformer trained from scratch on two consumer cards, measured on every checkpoint — a real caption effect, three bugs caught in our own code, and eight retractions kept.
In brief
BlackMind's first video model, trained from scratch on two consumer GPUs: it learns to turn noise into video-like scenes, develops a small but real two-seed caption effect, and is measured on every checkpoint by instruments strong enough to catch three bugs in our own code and retract eight claims.
Key results
- A 190.54M-parameter BAVDiT-B (20 layers, d=768) trained from random initialisation on two power-capped RTX PRO 4000 GPUs, one seed per card, over roughly three weeks of wall-clock (17 September to 6 October 2026) for about 378 productive GPU-hours each (~31 epochs over 58,692 filtered VGGSound clips); no downloaded generator weights and no LoRA at any point.
- The model learns a small but reproducible preference for its own clip's caption over another's: +0.33-0.36% of velocity-MSE, n=64 sign test 0.73-0.78 with p<=2e-4 on both seeds and both weight sets, yet only 0-1 of 16 fixed grid cells ever show a persistent, moving, on-prompt subject.
- Masked-mean pooling of the frozen text encoder made every caption 96.8% identical to every other; standardising by corpus mean and standard deviation pulled the 16 grid prompts apart from cosine 0.968 to 0.611, and the naive fix had to be guarded against a 640-norm null vector.
- The AdaLN modulation vector is 99.7% negative components, so conditioning rides on a handful of positive coordinates; one ordinary caption ('a video of metronome') cancelled that positive tail and drove a seed to predict a static clip (|v|max 0.95 vs 6.4 normal, loss 1.87 vs 0.09), and the shipped guards reduced but did not eliminate it.
- The strongest predictor of what the model can draw is whether the prompt's subject is a human (Spearman rho=+0.688, p=0.0032 over 16 classes), not how much data the class has (clip-count vs outcome rho=+0.41, p=0.113); classifier-free guidance was inadvertently off (cfg 1.0) for every evaluation in the run.
- 37% of GPU wall-clock produced zero gradient steps (178.9 h idle between legs, 54.8 h re-heat churn, 41.2 h on checkpoint saves), and eight claims were retracted under adversarial audit and kept in the paper as part of the result.
The Model and the Rig
This is the lab's first video model, trained end to end on our own hardware from random noise. BAVDiT-B is a 190.54M-parameter diffusion transformer, 20 layers at width 768, trained from random initialisation under a rectified-flow objective on velocity MSE with a timestep shift of 3.0. There are no downloaded generator weights and no LoRA at any point; a frozen Wan2.1 VAE supplies 13-frame latents of 1.63-second clips and a frozen Qwen3-1.7B supplies the text, but the generator itself is ours and starts from noise.
The hardware is two RTX PRO 4000 cards with 24 GB each, power-capped at 145 W, running one seed per GPU so that every claim carries two independent witnesses. Spatial attention is per-frame because global attention destroyed the image prior early; temporal AnimateDiff-style blocks sit every two layers and are a verified no-op at a single latent frame. After roughly 378 training-hours per seed the run sits at about 31 epochs over 58,692 VGGSound clips, step 36,048 on seed one and 37,516 on seed two.
The animation below is the honest record of what that bought: the 16-prompt grid for seed one as a timelapse from step 162, pure noise, to step 36,048, where scene-like structure has resolved. It is the thing to look at before any number, because in this project the pictures are the gate and the numbers are only the instrument.

Pictures Decide, Numbers Do Not
Measurement discipline is a genuine contribution of this work in its own right, and it rests on three instruments. The first is a live-weight paired eval run every 250 steps on 256 held-out clips with identical noise and timesteps across arms: the clip's own caption, another clip's caption, and a masked caption, read as a swap gap and a none gap. It runs on raw weights rather than EMA, because EMA lagged about 1,000 steps and blinded every earlier read. The second is a pre-registered gate of record, a per-clip sign test over 64 held-out clips with a 2,000-resample bootstrap, registered before the first result, passing only when the lower confidence bound clears 0.5.
The third instrument is strict visual grading: 16 fixed prompts rendered as a 4x4 grid every checkpoint and graded cell by cell, YES only for a persistent, moving, on-prompt subject, with a pass bar of at least 8 of 16. The owner's rule is that pictures decide and numbers do not, and the visual log runs to 2,235 lines of per-cell grades. The still strip below traces five checkpoints of that process from step 162 to 36,048.
The verdict the pictures return is blunt. The final 4x4 grid is video-like and temporally coherent, but it is mostly not on-prompt, and against the strict gate it passes 0 to 1 of 16 cells. That gap between a measurable text effect and a visibly empty grid is the whole story of the run, and we let the grid have the last word.

What Stands: A Small Caption Effect and a Human
Two results survive adversarial audit. The first is a small, real, reproducible caption effect: both seeds independently pass the pre-registered gate on both weight sets at the 144-hour checkpoints, with sign-test scores of 0.766, 0.781, 0.750 and 0.734 against roughly 0.5 earlier in the same run on the same clips and noise, p ranging down to 2e-5. The effect size is +0.33-0.36% of velocity-MSE, and on seed one raw the paired mean-delta confidence interval excludes zero at [+0.000208, +0.001045]. It is genuine, and it is small, and it has been flat since step 12,500.
The second is the most interesting finding we have and the one we can least explain. The strongest predictor of which prompts render is whether the subject is a human, at Spearman rho=+0.688, p=0.0032 over the 16 grid classes, with a Fisher test of 5 of 5 human classes against 4 of 7 non-human at p=0.034. People and people-scenes render; animals, fire, water and machines do not, and they fail regardless of how much data the class carries. For orientation, the first plotted point of any published text-conditioned DiT curve sits at 1.3e7 to 3.8e7 samples and we are at roughly 2e6; every T2V paper that shows caption following either pretrained on images-with-text first or initialised from a T2I model, and we did neither.
The grid below is where both results live. The caption effect is real in the numbers and the human finding is visible in which cells come to life, yet the final grid still passes 0 to 1 of 16 cells under the strict gate.

Three Bugs We Found in Our Own Code
Each of the three text-conditioning failures was found in our own code rather than in the literature. The first is pooling anisotropy: masked-mean pooling of the frozen encoder's states made every caption 96.8% identical to every other, with the distinguishing class carrying only 3.2% of the conditioning vector. Standardising by corpus mean and standard deviation pulled the 16 grid prompts from cosine 0.968 to 0.611, but the naive standardisation created a 640-norm null vector, because a dropped caption pools to exactly zero and (0-mean)/std is a 24-sigma outlier, which we then had to guard.
The second is an architectural fragility with a measured mechanism. The pooled vector fans out to three modulation sites, and 99.7% of its components are negative, so the whole AdaLN modulation rides on a handful of positive coordinates. A caption whose positive tail cancels drives every block to near-identity and the model predicts nothing: on one ordinary caption, 'a video of metronome', one seed collapsed from about step 10,000 with |v|max 0.95 against 6.4 normal and loss 1.87 against 0.09, and it poisoned that seed's entire eval mean. We shipped a null-relative delta guard OR'd with a modulation-norm-ratio test, the latter at AUC 1.000 against 0.973 for the amax test on a 53-row study.
The honest limit is that the modulation test is a high-noise detector: a degraded row's ratio falls with the timestep while a clean row's rises, so it catches full collapses at every noise level but partial degradation only above about t=0.6. The guards reduced but did not eliminate the problem, and 250 post-guard reads still show a none gap below -0.002 on seed two. The eval curves below show the clean result underneath all of this: held-out velocity-MSE falling from 0.93 to about 0.40 over roughly 31,000 steps.

Guidance, and Eight Retractions We Kept
Classifier-free guidance was off for the entire run: every grid ever judged was rendered at cfg 1.0. A proper sweep across both seeds and five guidance weights, with weights, seed and steps held fixed, failed its pre-registered bar of at least 3 of 16 cells gaining on both seeds; we got 1. But that one gain is the same cell on both seeds, the bowling prompt at weight 3 or above resolving into converging lanes, gutters, pinsetters, seating, a ball and a delivery posture, and the sub-threshold gains include seed two's fireworks finally rendering at night. Damage is a cliff at weight 5, crushed blacks and about 20% more flicker, not a slope, so we adopted weight 3.0. The honest reading is that guidance amplifies the scene the prompt already picked; it does not make the prompt pick a better subject.
Eight claims were retracted under adversarial audit, and we keep them in the paper because the retractions are part of the result. Among them: that the swap gap grew monotonically, which has been flat since step 12,500; that two seeds agreed to 0.00001 on a clean subset that exists in no artifact, when the on-disk seed-two receipts read +2.6% and +7.1% with intervals spanning zero; that the data filter removed exactly the classes we cannot generate, where the bar-detector mechanism is confirmed but the causal story is not supported at p=0.113; and that a run's first YES had emerged, downgraded at 9x zoom from a banjo to a keyed woodwind. We also retracted 'zero lost steps' once 150 steps turned out to have been re-executed across resumes, and a Contrastive Flow Matching trial that reverted by pre-registered rule.
The cfg sweep below is the receipt for the one decision guidance actually earned.

What Is Unresolved
Four things are genuinely open. We do not know why the caption effect sits at about 0.4% and stops, real on both seeds and flat. We do not know why humans render and animals, fire, water and machines do not, which is the most interesting finding in the work and the one with no mechanism. We do not know whether the proposed AdaLN additive-floor fix is right, a learnable, zero-initialised text branch added to the timestep modulation so that no caption can drive modulation below the unconditional model's, because it costs 1.8% more parameters and needs a fresh from-scratch run to test.
Three levers were never run by owner decision: REPA-style alignment to an external encoder, image-first pretraining on our own frames at roughly 13x the samples per GPU-day, and mid-layer or norm-averaged text features instead of the last hidden state. A downloaded 1.3B reference model is on-prompt for fireworks, owl, elephant, chicken and train at our exact bucket size, used only as a measurement control, which tells us the ceiling is ours and not the latent format's.
On the engineering side, 37% of card time produced zero steps, 178.9 hours idle between legs, 54.8 hours of re-heat churn after resumes and 41.2 hours on checkpoint saves that took 141 seconds each against 2 seconds for an honest copy, and a GPU-idle watchdog warned for 56 hours into a page nobody read. Two GPUs is also not 2x on one model: we ran one seed per card, and the scarce axis at 31 epochs is optimizer steps rather than samples, so cutting gradient accumulation offered 2.67x on both seeds for free. And a git push failed silently for two days because an animation crossed a hard 100 MB file limit. We report all of it because an honest account of a video model trained on two consumer GPUs is, mostly, an account of where the time and the certainty actually went.
Abstract
We trained a 190M-parameter text-to-video diffusion transformer from random initialisation on two consumer GPUs over roughly three weeks of wall-clock (17 September to 6 October 2026), about 378 productive GPU-hours per seed (~31 epochs over 58,692 VGGSound clips), measuring text conditioning with a pre-registered instrument on every checkpoint. The model learns a small but reproducible preference for its own caption over another clip's (+0.33-0.36% of velocity-MSE; n=64 sign test 0.73-0.78, p<=2e-4, both seeds and both weight sets) while remaining at 0-1 of 16 grid cells showing a persistent, on-prompt subject, and the strongest predictor of what renders is whether the subject is a human (rho=+0.688, p=0.0032), not how much data the class has. We report three conditioning bugs found in our own code, that guidance was inadvertently off for the entire run, that 37% of GPU time produced no gradient steps, and eight claims retracted under adversarial audit.
