BlackMind
Research
Image generationDiffusion transformersHonest measurement

Training a Text-to-Image Model from Scratch on Two GPUs

BlackMind's first image model: a latent diffusion transformer trained from scratch on CC12M on two consumer cards — a small but real prompt effect, a Muon optimiser win, and the bakeoffs that moved the loss but not the picture.

Jovonni L. PharrGeorgia Cyber Warfare Range / BlackMindSeptember 2026

In brief

BlackMind's first text-to-image model: a latent diffusion transformer trained from scratch on CC12M on two consumer GPUs. It learns to draw recognisable objects and scenes and provably prefers its own prompt, a Muon-hybrid optimiser beats all-AdamW, and a CLIP arm-similarity study shows that most interventions during the stage moved the training loss without moving the picture — with every retraction kept on purpose.

Key results

  • A DiT-S (43.67M) trained from random initialisation memorises 64 fixed CC12M images and, sampling from pure noise, routes every caption to its own image 64 of 64 times (100% against a 1.6% chance baseline) in 7.52 minutes — validating the ported flow-matching core under the Muon+WSD harness with no downloaded weights, no LoRA, and no pretrained image prior.
  • The full ~150-400M latent diffusion transformer, trained from scratch on our CC12M corpus, is built to clear a pre-registered Stage-1 gate: beat a frozen SD1.5-class open checkpoint on our own suites under our judge protocol, at matched inference compute — a pipeline-works gate, explicitly not a claim of parity with frontier image models.
  • An optimiser bakeoff at equal wall-clock gives Muon-hybrid the win over all-AdamW on held-out velocity-MSE (0.18333 vs 0.18609), AdamW losing despite completing ~8% more steps; the directional verdict is kept while the n=1 significance multiplier and a weight-decay confound on ~85% of parameters are withdrawn in an audit correction we keep in the body.
  • The central finding, a CLIP arm-similarity study, shows that three of six interventions — including the entire learning-rate sweep — changed the picture no more than a different training seed does; the most-cited bakeoff winner (Muon vs AdamW) is perceptually ambiguous (CI straddles the seed floor); and the one intervention that demonstrably moved the image distribution, quality-injection, is the one that lost (+0.00347 worse on the primary metric).
  • Classifier-free guidance was inadvertently off (w=1.0) for every image the program ever published; a sweep shows turning it on is correct (the w=1 and w=5 alignment CIs are disjoint) and w=3 is the alignment optimum, but the effect is mostly beautification (+70% saturation, flat sharpness, the subject unchanged — the car is already a car at w=1) rather than comprehension.
  • A held-out probe finds a small but real conditioning effect: 4 of 64 samples (6.2%) land nearest their own caption's ground-truth latent against a 1/64 = 1.6% chance baseline (exact binomial p=0.018) — real but weak on a coarse instrument, and every retraction from the stage is kept in the paper on purpose.

The Model and the Rig

This is M1, BlackMind's first image generator: a latent diffusion transformer trained from random initialisation on CC12M images, on two consumer graphics cards. There are no downloaded generator weights, no LoRA, and no pretrained image prior at any point in the run; the only frozen components are a VAE and a text encoder, used strictly as fixed instruments, and the transformer that does the drawing is ours, from scratch. The latents come from a frozen Wan2.1 VAE and the text pathway from a frozen Qwen3-1.7B; the objective is flow matching on a velocity-MSE loss in the rectified-flow formulation, with an SD3-style train-time timestep shift.

M1 is not a stand-alone project. It is the first real training rung of a staged video-model program, and the image model comes first by design, because the cheapest way to prove a diffusion-transformer training pipeline works is to make it draw a still frame before asking it to make a moving one. The program runs on a falsifiable ladder: every milestone is a pre-registered gate, passed only by an audited receipt, and failing a gate is treated as a result, not a delay. M1's gate is deliberately modest — our ~150-400M latent DiT, trained on our corpus, beats a frozen SD1.5-class open checkpoint on our suites under our judge protocol, at matched inference compute. It does not claim parity with modern frontier image models; it claims the pipeline works.

Two things make this worth writing up. The first is the model itself: from random noise it learns to draw recognisable objects and scenes — cars, garments, food, interiors — with distortion and garbled text, but unmistakably on the right track. The grid below is the honest state of the image model, 64 prompts the model never saw in training, sampled from pure noise. The second, and in our view just as valuable, is the measurement discipline we built around it, sharp enough to produce the central finding of the paper: that most of the interventions we ran during the stage moved the training loss without moving the picture.

What the model draws: the 8x8 held-out generation grid from the Stage-1 checkpoint, 64 prompts never seen in training, sampled from pure noise. Recognisable cars, garments, food and interiors with structural distortion and garbled text — the honest state of the image model, shown before any number.
Fig.What the model draws: the 8x8 held-out generation grid from the Stage-1 checkpoint, 64 prompts never seen in training, sampled from pure noise. Recognisable cars, garments, food and interiors with structural distortion and garbled text — the honest state of the image model, shown before any number.

The Trainer, Validated

Before any ranking claim, the trainer earns the right to be trusted by memorising a tiny fixed set and routing text to it correctly. Sixty-four CC12M images were centre-cropped to 256x256, encoded once through the frozen Wan2.1 VAE, and overfit with zero text-dropout; over 1,000 steps the smoothed loss fell from 1.2444 to 0.00353 — a 352x drop, monotone, no spikes, no NaN, in 7.52 minutes of wall-clock.

The go/no-go was the sampling test. Drawing 64 images from pure noise, each conditioned on one of the 64 cached caption embeddings, every single sample is nearest in latent space to its own caption's ground-truth image: 64 of 64, 100%, against a 1.6% chance baseline. We ran the first-skeptic check — this is not a decode-the-ground-truth bug, because the samples start from pure noise and the mean latent distance to the matching image is 3.11, not near zero. Conditioning, not tensor identity, selects the image, and the ported flow-matching core under the Muon+WSD harness is functionally correct. The smoke makes no claim about quality or generalisation; those are the first real GPU-hours.

The Bakeoffs: Muon Wins, and the Audit We Kept

With the pipeline validated, the stage ran two-arm bakeoffs, each with a pre-registered gate on held-out velocity-MSE. The optimiser verdict fired: Muon-hybrid beats all-AdamW at equal wall-clock, 0.18333 against 0.18609, and AdamW lost despite completing about 8% more steps, so its per-step disadvantage exceeds its throughput edge. This matches the sibling language-model program's Muon finding — the port's hot-learning-rate heritage transfers to flow-matching DiTs.

The audit correction is kept in the body because it bounds the claim honestly. The comparison is not a pure optimiser A/B: the Muon trunk matrices carried weight-decay 0.0 while all-AdamW carried 0.1, a regularisation confound on about 85% of the parameters, so this is a Muon-hybrid recipe against an all-AdamW recipe, not an isolated optimiser swap. The '9.2x seed noise' multiplier is withdrawn as a significance statement at n=1 per arm; what stands is directional — Muon led at every checkpoint and both AdamW arms landed above all four Muon runs. The learning-rate sweep was nearly flat, and here too we restated a verdict against ourselves: the stage first called the range 'insensitive', but the pre-registered rule required a gap no larger than the between-seed gap, and 0.00042 exceeded the 0.00030 seed-noise bound, so under the letter of the pre-registration muon-lr 0.01 was the screening winner and the convenient 'insensitive' framing is retracted.

The Arms Moved the Loss, Not the Picture

This is the result that, in the words of our own audit, should stop and reorient the program. Every gallery grid is the same 16 held-out prompts at the same seed, so a cell of two grids is two models' answer to one prompt. We encode each cell with CLIP ViT-L/14 and take the mean cosine between matched cells, judged against two anchors measured identically: a prompt floor (different prompts, same model: 0.7330) and a seed floor (same recipe, different training seed: mean 0.8238, but the two pairs disagree by 0.053, so a claim that an arm changed the picture must clear the conservative ~0.85 floor, not the average).

Three of six arm pairs — including the entire learning-rate sweep — are as similar to each other as two runs of the same recipe with a different seed. Those interventions moved held-out velocity-MSE without moving the picture. The sting is in the two headline results. Our most-cited bakeoff outcome, Muon ahead at every checkpoint, is perceptually the ambiguous one: AdamW versus Muon scores 0.8463 with a confidence interval straddling the conservative seed floor, so it is under-powered rather than a clear perceptual difference. And the one intervention that demonstrably changed the picture is the one that lost: quality-injection scored 0.7835, below every seed pair, the only arm pair to move the image distribution — and its pre-registered verdict was that injection hurt, 0.18226 against 0.17879 on the primary metric. The single lever that reached the pixels made the metric worse.

The limits are stated before anyone quotes this: 16 prompts per arm, a seed floor resting on two pairs that disagree by as much as the whole spread of effects, so this screen is under-powered by construction and no verdict here is settled. The lesson is not that the loss metric is useless; it is that a loss delta is not a picture delta, and the program now knows to demand both.

Guidance

An adversarial audit found that the trainer spent 10% of every batch training an unconditional branch the sampler never used: every image this program ever published was rendered at classifier-free guidance w=1.0 — that is, with guidance effectively off. Fixing it costs zero training GPU-hours; only the sampler changes. We ran a sweep on the same checkpoint and the same 16 held-out prompts, scoring the right criterion — whether the picture depends on its prompt, as a paired matched-minus-mismatched CLIPScore gap.

Three conclusions, one against us. Turning guidance on was correct and now has real evidence: the w=1 and w=5 confidence intervals do not overlap, and at w=1 the interval includes values indistinguishable from ignoring the prompt. On the point estimate w=3 is the optimum, not the w=5 that shipped — though the intervals from w=2 to w=7.5 overlap heavily, a plateau that 16 prompts cannot resolve, so the shipped value is arbitrary within a flat region rather than wrong. And the correction we keep: guidance is mostly a beautification knob, not a comprehension one — from w=1 to 7.5 saturation rises 70% while sharpness stays flat and the subject does not change; the car is already a car at w=1. The point stands confidently: guidance amplifies the scene the prompt already picked; it does not make the prompt pick a better subject.

The classifier-free guidance sweep on 16 held-out prompts. Prompt alignment rises from w=1 to w=3 and then plateaus; saturation rises 70% across the range while sharpness stays flat. w=3 is the alignment optimum, and the effect is real but mostly beautification, not comprehension.
Fig.The classifier-free guidance sweep on 16 held-out prompts. Prompt alignment rises from w=1 to w=3 and then plateaus; saturation rises 70% across the range while sharpness stays flat. w=3 is the alignment optimum, and the effect is real but mostly beautification, not comprehension.

A Small, Real Prompt Effect, and What We Kept

The clearest positive conditioning result comes from a held-out test designed under the audit's corrections. We draw 64 prompts at random from the held-out eval split, captions receipted and never seen in training, sample 64 images from noise, and ask how often each sample's nearest ground-truth-in-latent-space is its own caption's image. The result is 4 of 64, 6.2%, against a 1.6% chance baseline, with an exact binomial p of 0.018. The honest read is that the conditioning signal is real but weak: four-times chance on a coarse latent-distance instrument, with no quality claim attached. The matched ground-truth grid below makes the gap concrete — what the model gets right is coarse category and layout, what it misses is fine structure and text.

We keep every retraction on purpose: the recipe-locked optimiser claim downgraded to provisionally selected, the LR 'insensitive' framing restated, the w=5.0 ship justified after the fact on a non-significant screen, and the arm-similarity screen's own under-powering disclosed. What is unresolved is stated plainly. Does M1 clear its own gate? The pipeline is validated and the arms are run; the judge-and-metric scorecard against the SD1.5-class baseline is the remaining receipt, and until it exists M1 is not declared — the artifact is the truth, not the report. The optimiser verdict needs power — a trunk-weight-decay-matched Muon variant and at least three seeds would settle it. And why the prompt effect sits at 4 of 64 — corpus, text-encoder choice, or undertraining — is exactly what the later bakeoffs are designed to isolate. M1 is a confident result with a sharp lesson: the pipeline works, and the measurement discipline built around it is strong enough to tell us something uncomfortable and useful — that most of what we tuned moved a number without moving a picture. That is the knowledge the video rungs of the program are built on: demand both the loss and the picture, and never confuse one for the other.

The matched ground-truth images for the same 64 held-out prompts as the generation grid. Comparing the two shows what the model gets right (coarse category and layout) and what it does not (fine structure, text), and makes the 4-of-64 nearest-ground-truth result legible as a small, real effect rather than chance.
Fig.The matched ground-truth images for the same 64 held-out prompts as the generation grid. Comparing the two shows what the model gets right (coarse category and layout) and what it does not (fine structure, text), and makes the 4-of-64 nearest-ground-truth result legible as a small, real effect rather than chance.

Abstract

We trained BlackMind's first text-to-image model, M1, a latent diffusion transformer trained from random initialisation on CC12M, on two consumer GPUs, with no downloaded generator weights, no LoRA, and no pretrained image prior; the only frozen components are a VAE and a text encoder used strictly as instruments. M1 is the pipeline-validation rung of a staged video-model program, and we report the result and, with equal weight, the measurement discipline that framed it. A DiT-S memorises 64 fixed images and routes each caption to its own image 64 of 64 times from pure noise, validating the training core; the full ~150-400M model, trained on our corpus, is built to beat a frozen SD1.5-class open checkpoint on our own judge protocol at matched inference compute. A Muon-hybrid optimiser beats all-AdamW at equal wall-clock (held-out velocity-MSE 0.18333 vs 0.18609). A CLIP arm-similarity study shows that three of six interventions, including the entire learning-rate sweep, changed the picture no more than a change of random seed does, the most-cited bakeoff winner is perceptually ambiguous, and the one intervention that demonstrably moved the picture is the one that lost. A held-out probe finds a small but real conditioning effect (4 of 64 vs 1.6% chance, p=0.018). We keep every retraction on purpose.