BlackMind
Research
neural architecture searchevolutionary algorithmsAutoMLmechanism discovery

EDAG: Evolutionary Discovery of Neural Architectures and Mechanisms

A specification for an open-ended evolutionary engine meant not just to rewire known layers but to evolve the learning mechanisms themselves.

Jovonni L. PharrGeorgia Cyber Warfare Range / BlackMindJanuary 2026

In brief

EDAG is a design specification, not a benchmark paper: it lays out how an evolutionary search engine could discover deep learning architectures — and, more ambitiously, new learning mechanisms — with minimal human intervention. Its distinctive claim is that reordering existing layers is not enough; genuine novelty requires evolving state, update rules, and training dynamics as first-class objects. This entry reconstructs the design's motivating question, its search machinery, and the extension that turns it from an architecture-search tool into an open-ended mechanism-discovery framework.

Key results

  • Encodes each architecture as a direct genotype — a graph of layers with explicit parent links (illustrated by a four-layer Input to Conv2D to MaxPool2D to Dense CNN for 28x28 images) — so mutation and crossover operate on an explicit blueprint rather than an opaque vector.
  • Spans the full historical layer repertoire in one search space: dense/MLP, 1D/2D/3D convolution and pooling, residual and skip connections, inception-style parallel branches, RNN/LSTM/GRU cells, multi-head attention and Transformer blocks, GCN/GAT graph layers, and Titans-style long-term memory modules.
  • Ships four interchangeable search modes behind a single controller interface — steady-state/generational genetic algorithm (default), a policy-gradient RL controller, random search as a baseline, and hooks for Bayesian and differentiable (DARTS-style) NAS — all reusing one train-and-evaluate pipeline.
  • Cuts evaluation cost through weight inheritance (copying a parent's weights for unchanged layers so a child trains from a head-start), staged low-fidelity training with early stopping, and an optional learned performance-predictor surrogate to pre-screen mutations.
  • Preserves novelty structurally via age-based removal (killing the oldest individual rather than the worst, following regularized evolution), speciation, and a quality-diversity archive that keeps top performers within distinct behavioral niches.
  • Extends beyond topology by treating models as stateful programs with persistent internal state and distinct train-time vs inference-time behavior, and can promote recurring high-performing subgraphs into new first-class operators, expanding its own search space over time.

Designing the Designer

Most progress in deep learning still runs through human architects who hand-shape layers, connections, and training recipes, then tune. EDAG proposes to automate the design step itself — to treat the choice of architecture as an optimization problem in which the fitness of a candidate is simply how well the trained model performs on the target dataset. The stated ambition is to match or exceed human-designed models with minimal human intervention, going well beyond hyperparameter tuning to the structure of the network.

To keep the search domain-agnostic, EDAG abstracts every input as a numeric tensor. An image is a height-by-width-by-channels tensor, text is an encoded sequence of token IDs, a graph is adjacency plus feature matrices. Shape information is retained so that layer compatibility can be enforced, but no handcrafted features are supplied: the evolved networks are expected to build their own feature extraction. This unification lets a single evolutionary engine apply across images, text, audio, and graphs.

A Genome of Layers

Each candidate is represented as a direct encoding — an explicit blueprint listing every layer, its type, its hyperparameters, and its parent connections, in the spirit of NEAT's genome. The specification gives a concrete example: an Input of shape 28x28x1 feeding a 3x3 Conv2D with 16 filters and ReLU, then a 2x2 MaxPool2D, then a softmax Dense layer of 10 units — a minimal CNN the evolutionary operators can then grow.

The building-block library is deliberately exhaustive. It includes fully-connected layers of arbitrary depth; 1D/2D/3D convolution and pooling; residual and skip connections; inception-style parallel branches that merge; vanilla RNN, LSTM, and GRU cells; multi-head attention and full Transformer blocks; graph layers such as GCN and GAT; normalization and dropout; tunable activations; and specialized modules including external neural memory inspired by the Titans architecture for very long context. Layer-level hyperparameters — filter counts, kernel sizes, attention heads, unit counts — are part of the evolvable genotype alongside topology, so the same framework can in principle construct ResNet-like, LSTM-like, Transformer-like, and novel hybrid designs, and even weight-shared Siamese structures.

How the Population Evolves

The default engine is a genetic algorithm over a population of architectures. It is seeded deliberately from trivial models — logistic-regression-style input-to-output maps or single-hidden-layer MLPs — so that complexity is discovered incrementally rather than handed to the search. Individuals are trained on the dataset, scored by validation accuracy (the primary fitness), and then subjected to selection: the specification favors tournament selection with a steady-state loop in which two parents produce one child and one individual is removed to hold population size constant.

Mutation is the main source of variation, with operators that add a layer, remove a layer, modify a layer hyperparameter, alter connectivity and skip connections, adjust regularization, or introduce weight-sharing. Mutations are kept small so that complexity accumulates gradually. Crossover — swapping matching sub-networks between two parents, aligned by innovation-style IDs — is supported but optional and used sparingly, since random recombination can produce invalid graphs. Complexity is not directly penalized by default; the search is trusted to grow a network only when the added structure earns higher accuracy.

Making Evaluation Affordable

Training every candidate to convergence is the dominant cost, so the specification layers several economizations. Weight inheritance copies a parent's trained weights into the unchanged layers of a child, leaving only new layers randomly initialized, so an incremental mutant trains from a head-start rather than from scratch. Evaluation is staged: cheap short runs assess potential first, deeper training is spent only on promising or novel candidates, and full evaluation is reserved for survivors. Early stopping abandons clearly failing models, and an optional learned performance-predictor surrogate can pre-screen mutations before any full training.

Practical viability is enforced with guardrails rather than blanket complexity penalties: models that exceed a training-time budget are marked infeasible, and models that exceed available memory are discarded. Efficiency metrics — parameter count, estimated FLOPs, per-epoch time, inference latency — are always recorded and can be folded into a multi-objective fitness (for example via NSGA-II) when the user asks for Pareto-optimal accuracy-versus-cost trade-offs. Evaluation runs in parallel across available GPUs and cores, dispatched by a job scheduler.

Beyond Rewiring: Mechanism Discovery

The specification's second half argues that shuffling known blocks cannot, on its own, produce paradigm shifts like Nested Learning or test-time memory. EDAG therefore treats every candidate as a stateful computational program rather than a pure forward function: models may carry persistent internal state across timesteps and inference calls, define explicit update rules for that state, and behave differently at training time and inference time. Structure, internal state, state-update logic, parameter-update behavior, and auxiliary computation paths are all treated as separately evolvable dimensions, so novelty is not confined to topology.

Novelty is tracked by observed behavior — how representations change with depth, how gradients flow, how routing or attention patterns behave, how internal state evolves, how predictions respond to perturbation — so that two structurally similar models can still count as genuinely different. A quality-diversity archive preserves top performers within distinct behavioral niches instead of collapsing onto one global optimum, with age-based removal and speciation guarding against premature convergence. Recurring high-performing substructures can be abstracted into reusable composite mechanisms and promoted into the operator library, letting the system expand its own search space and build cumulatively on earlier discoveries. To keep novelty meaningful, candidates must clear minimum competence thresholds before entering the archive, filtering out untrainable or random behavior.

Two Runtimes, One Trail

EDAG is specified for parallel implementation in Python with PyTorch and in Rust with tch-rs, which binds the same libtorch backend. Python is intended for rapid prototyping of new layers and data handling; Rust for a safe, multi-threaded core engine that can train many networks in parallel and serve production AutoML use. The two are meant to cross-verify on small problems, with ONNX or TorchScript as an interchange format and consistent hyperparameters on both sides.

Because the design values interpretability, the system records the genealogy of every architecture. For the final best model it can reconstruct the full lineage — seed to descendant — annotating each step with the mutation applied and the fitness change it produced, alongside per-generation best and average scores. The evolved model can be exported as generated PyTorch module code, a saved state dict, ONNX, or TorchScript. This evolutionary trace is presented as a step-by-step account of how a design was reached, turning an otherwise opaque AutoML run into an auditable path.

Scope and Standing

This is a specification and design document, and it should be read as such: it defines representations, operators, search modes, efficiency strategies, and reporting, and it grounds its choices in prior results — large-scale and regularized evolution of image classifiers, RL-controller NAS, multi-objective NSGA-Net, performance-prediction and weight-inheritance NAS, categorical theories of architecture, and Titans-style test-time memory. It does not report EDAG's own trained models or measured accuracies; the numbers it cites are the design targets and the external precedents it builds on.

The stated goal is explicitly not to rediscover known architectures faster, but to operate as an open-ended discovery engine that can invent new uses of memory, state, and adaptation, develop families of models that share internal logic but differ in structure, and produce architectures humans did not explicitly design. Its listed future directions — meta-learning to guide the evolution, self-adaptation of the search's own parameters, and evolving the optimization procedure itself as part of a model's DNA — mark the boundary between what the specification commits to and what it leaves open.

Abstract

EDAG is a specification for an automated system that evolves deep neural network architectures for a given dataset, treating network design as an optimization problem in which fitness is model performance. Each candidate is encoded as a direct genotype — a graph of layers and connections spanning feed-forward, convolutional, recurrent, attention, graph, normalization, and memory-augmented components — and a population is refined through selection, mutation, crossover, and weight inheritance. The default search is a steady-state genetic algorithm seeded from trivial models and grown incrementally, but the search controller is abstracted behind an interface so that a reinforcement-learning controller, random search, Bayesian optimization, or differentiable NAS can be swapped in via configuration. Fitness is validation accuracy by default, with optional secondary penalties for latency, parameter count, or memory footprint, and with staged low-fidelity training, early stopping, and parent-to-child weight reuse to reduce evaluation cost. A second half of the specification extends the framework beyond permutations of known blocks: it treats every model as a stateful computational program, separates structure from internal state and update dynamics, tracks novelty by observed behavior rather than topology, maintains a quality-diversity archive across behavioral niches, and can abstract recurring high-performing substructures into new first-class operators. The system is planned for parallel implementation in Python (PyTorch) and Rust (tch-rs), and reports the full genealogy of the best model as an interpretable evolutionary trace.