arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2609.40258v1 [cs.NE] 30 Sep 2026

Large Language Model-Guided Evolutionary Discovery of Native Neural Architectures for Spiking Sequence Modeling

Ruoyu Zhao* Affiliation: Department of Computer Science, City University of Hong Kong    Jiaqi Wu* Affiliation: Department of Computer Science, City University of Hong Kong    Chenyu Zhu Affiliation: Department of Computer Science, City University of Hong Kong    Zhichao Lu Affiliation: Department of Computer Science, City University of Hong Kong Affiliation: {ruoyuzhao8-c, jwu395-c, chenyuzhu9-c}@my.cityu.edu.hk, zhichao.lu@cityu.edu.hk Corresponding author: Zhichao Lu (zhichao.lu@cityu.edu.hk). *Equal contribution.
Abstract

Spiking neural networks (SNNs) offer low-energy sequence modeling through sparse, event-driven computation. However, interactions among spike encoding, neuronal dynamics, and information propagation complicate architecture design. Existing SNN sequence models often adapt artificial neural network (ANN) architectures designed for real-valued activations, potentially underusing spike-based communication and temporal state updates, motivating automated discovery of native SNN architectures. Most evolutionary neural architecture search (ENAS) methods operate within predefined configuration spaces, limiting discovery to mechanisms expressible within those spaces. We introduce OpenArchEvo, which uses large language models (LLMs) to evolve executable architecture code in an open program space under spiking-projection constraints. In this space, code differences need not reflect architectural novelty, while direct performance evaluation requires costly training. We construct a three-view representation spanning code, design rationale, and a behavioral fingerprint to support novelty estimation and performance prediction. The search treats predicted performance and estimated novelty as two objectives, using surrogate predictions to select candidates for expensive training evaluations. With an estimated candidate-training cost of 132 V100 GPU-days, the search uncovers multiple native SNN architectures, exemplified by three designs featuring mechanisms such as spike-activity-dependent control of state updates and residual pathways. The discovered NeuroGate surpasses the ANN DeltaNet on WikiText-103, and the discovered architectures reduce estimated architecture-level arithmetic energy by up to 50.6×50.6\times (LoopMem) relative to a common dense Transformer (ANN) baseline. All code and all discovered architectures will be made publicly available soon.

1 Introduction

Refer to caption
Figure 1: Motivation and result preview for native SNN architecture discovery. (a) A qualitative map of LLM-based algorithm design by candidate-characterization difficulty and evaluation cost. (b) Partial rank agreement is shown by separately trained ANN and SNN implementations of the 97 ASI-Arch paired architectures on WikiText-2 (Kendall τb=0.2997\tau_{b}=0.2997; Appendix A.1). (c) NeuroGate achieves 26.4 WikiText-103 perplexity and a 31.7×\times estimated architecture-level arithmetic-energy reduction relative to a common dense Transformer (ANN) baseline (Table 1), under spiking-projection constraints (Appendix A.2). Red arrows mark the added pathway through which spike activity modulates the recurrent update and output gating.

Spiking neural networks (SNNs) communicate through sparse, event-driven spikes, offering a promising route to low-energy sequence modeling. Binary spike signals interact with continuous states that retain and integrate information over time. Architecture determines how spike activity updates these states and how stored information is propagated, shaping predictive performance and opportunities for sparse computation. Existing SNN sequence models include both adaptations of artificial neural network (ANN) architectures and mechanisms designed specifically for spiking computation (Zhu et al., 2024; Zhong et al., 2026; Shen et al., 2025). However, architectural choices effective in ANNs need not retain their advantages after adaptation to spiking computation. In our WikiText-2 comparison, ANN and SNN implementations of the 97 ASI-Arch paired architectures agree only partially in rank (Fig. 1(b)), motivating direct evaluation under spiking computation. The interactions among spike activity, state updates, and information pathways make useful combinations difficult to anticipate. This motivates an AutoML research question: how can search build on expert designs to discover native SNN architectures that match or exceed ANN performance while preserving substantial energy advantages?

Neural architecture search (NAS), particularly evolutionary NAS (ENAS), offers a route to this automated exploration (Liu et al., 2023; Li et al., 2024). Most ENAS methods vary operators, connections, and hyperparameters within predefined configuration spaces. These spaces permit new combinations of known components, but mechanisms outside their allowed operations and composition rules require redesigning the search representation. Large language models (LLMs) can generate and modify executable architecture code, enabling changes to internal computations and module interactions. Existing LLM-based NAS already supports architecture discovery through code evolution, including quality-diversity search and novelty checks (Chen et al., 2023; Nasir et al., 2024; Liu et al., 2025). Greater freedom to generate candidates, however, does not by itself yield effective architectural discovery. Figure 1(a) situates this problem within LLM-based algorithm design: code differences need not reflect meaningful architectural differences, while directly measuring task performance requires costly candidate training. The challenge is to characterize generated SNN architectures for novelty estimation and performance prediction, guiding training toward promising candidates while retaining distinct architectural alternatives.

To address this challenge, we introduce OpenArchEvo, an LLM-guided evolutionary method for native SNN architecture discovery. It searches over an open code implementation space for SNN blocks within a fixed outer model structure, subject to interface, causality, and spiking-projection constraints. We estimate each candidate’s architectural novelty relative to other architectures considered during search through three complementary views: code describes computational structure, the design rationale states the intended design, and a behavioral fingerprint combines architecture statistics with network responses measured before weight training. The fingerprint also supports surrogate performance prediction; predicted performance and estimated novelty form two objectives for selecting candidates for training. One discovered architecture, NeuroGate, features spike-activity-dependent modulation of recurrent updates and output gating. On WikiText-103, the discovered SNNs outperform the listed SNN baselines, and NeuroGate (Fig. 1(c)) attains 26.4 WikiText-103 perplexity versus 27.5 for ANN DeltaNet (lower is better). The discovered architectures reduce estimated architecture-level arithmetic energy by up to 50.6×50.6\times (LoopMem) relative to a common dense Transformer (ANN) baseline (Table 1).

Our main contributions are as follows:

  • •

    Native SNN architectures. We discover SNN architectures with evolved elements native to spiking computation, including spike-activity-dependent control in NeuroGate and HomeoResSSM and a state-norm feedback in LoopMem that benefits the spiking implementation, and examine these elements through SNN ablations and a non-spiking comparison.

  • •

    Three-view architecture representation. We construct a three-view representation using code, design rationale, and numerical behavioral fingerprints of architectural structure and behavior. We compare candidates using these views to estimate each candidate’s architectural novelty, while reusing the fingerprints for performance prediction.

  • •

    LLM-guided evolutionary discovery. We integrate this characterization into LLM-guided evolution over an open code implementation space, using surrogate-predicted performance and estimated novelty as two objectives to select architectures for training within a limited budget. Search diagnostics and ablations examine surrogate guidance and diversity maintenance.

2 Related Work

We review SNN architecture design, evolutionary neural architecture search, and representations for candidate comparison and performance prediction.

2.1 Spiking Sequence Models

Spiking neurons combine binary outputs with continuous membrane states that retain temporal information. Architectural choices determine how these signals are transformed and propagated. In vision models, Spikformer uses spike-based attention without softmax (Zhou et al., 2023), while Spike-driven Transformer places residual connections before spiking activations to preserve binary communication (Yao et al., 2023). These designs show how adapting information pathways to spiking computation can support performance with less computation.

For sequence modeling, SpikeGPT combines recurrent processing with binary spiking activations for language generation (Zhu et al., 2024). The SpikingSSMs architecture applies leaky integrate-and-fire (LIF) dynamics to state-space outputs, combining temporal memory with sparse synaptic computation (Shen et al., 2025). Dyn-SSM incorporates refractory LIF neurons with soft reset, preserving residual membrane potential after firing (Zhong et al., 2026).

Automated methods also demonstrate the value of SNN-specific architectural exploration. SNASNet searches forward and temporal feedback connections, while MSE-NAS explores neuron operations and multiscale connectivity (Kim et al., 2022; Pan et al., 2025). AutoSNN considers accuracy and spike count and reports improvements over its handcrafted baselines (Na et al., 2022). For spiking language models, EQ-SpikeLM combines evolutionary channel pruning with subsequent post-training quantization (Zhang et al., 2026). Further couplings among spike activity, sequence-state updates, and residual pathways offer opportunities to extend this accumulated design knowledge.

2.2 Evolutionary Neural Architecture Search

Evolutionary neural architecture search (ENAS) uses population-based variation and selection to optimize neural architectures (Liu et al., 2023; Li et al., 2024). Most methods vary operators, connections, widths, and depths within predefined representations. For example, EvoCNN evolves variable-length sequences of convolutional, pooling, and fully connected layers (Sun et al., 2020), while NSGA-Net evolves CNN blocks through predefined operation and connection choices (Lu et al., 2021). These encodings support new combinations of existing building blocks, but a mechanism outside their permitted operations and composition rules requires redesigning the search representation.

Genetic programming and grammar-based NAS support variable computational structures assembled from reusable primitives. CGP-CNN evolves convolutional architectures from layer-level components, while einspace uses typed primitives to define a broader compositional space (Suganuma et al., 2017; Ericsson et al., 2024). AutoML-Zero further explores model computation and learning rules constructed from basic mathematical operations (Real et al., 2020). High-level primitives supply architectural priors through predefined computational forms; finer primitives expose more computational choices but can make useful designs harder to find. Even when composition obeys type and interface constraints, finding useful architectures can require many expensive evaluations. This motivates using architectural knowledge to guide the generation and modification of candidate architecture code.

LLM-guided evolutionary search has become a general approach to automated algorithm design (Wu et al., 2025; Ma et al., 2026), evolving heuristics, metaheuristics, and scientific programs (Romera-Paredes et al., 2024; Liu et al., 2024; van Stein & Bäck, 2025; Novikov et al., 2025); ShinkaEvolve, for example, combines code-embedding similarity screening with an LLM novelty judge to reject redundant candidate programs (Lange et al., 2026). For neural architectures, LLMs use pretrained knowledge and natural-language design goals, constraints, and principles to guide configuration search or generate architecture code. LLMENAS adapts fitness functions within a predefined cell search space (Lai et al., 2026), while Design Principle Transfer uses natural-language principles to narrow subsequent searches (Zhou et al., 2025). EvoPrompting evolves complete classifiers and, for graph networks, selected computations within a fixed processor (Chen et al., 2023). LLMatic combines network-code evolution with quality-diversity search for image classification (Nasir et al., 2024). ASI-Arch evolves attention-layer implementations under interface and computation constraints; it checks novelty through rationale retrieval and LLM judgments before training, and includes LLM-assessed architectural quality in its fitness (Liu et al., 2025); Genesys discovers language-model architectures with LLM agents on a genetic-programming backbone (Cheng et al., 2025). These code-based approaches establish architecture discovery with mechanisms for diversity and novelty. Our work combines population-relative novelty and surrogate-predicted performance to select SNN architecture programs for training; Section 2.3 details the representations supporting these decisions.

2.3 Architecture Representation for Evolutionary Search

A search encoding specifies how architectures can be constructed; candidate descriptors summarize the resulting architectures for comparison and prediction. Implementation-level comparisons include the lexical, syntactic, and data-flow matches measured by CodeBLEU (Ren et al., 2020). Natural-language descriptions expose design intent, as in Evolution of Heuristics (EoH), which jointly evolves heuristic ideas and executable implementations (Liu et al., 2024). Execution-based descriptions capture observed behavior: phenotypic characterization records program responses to probe cases and supports surrogate modeling (Hildebrandt & Branke, 2015), while BehaveSim compares intermediate solution trajectories (Zhang & Lu, 2026). These views provide complementary, partial evidence: similar code need not implement the same computation, stated design intent need not match the implementation, and finite probes reveal only part of a program’s behavior.

For neural architectures, structural statistics and initialization-time forward and backward probes provide information before candidate training. Training-free proxies estimate trained performance from initialization-time activations or gradients without optimizing candidate weights (Abdelfattah et al., 2021). DCL-ENAS pretrains an architecture encoder without performance labels, then contrastively fine-tunes a predictor on evaluated architectures to predict their relative performance (Zhang et al., 2026b). In expressive grammar-based NAS, Transferrable Surrogates builds transferable predictors from training-free proxies and neural graph features, or from a fine-tuned language model, and uses them to filter candidates or serve directly as search objectives (Qin et al., 2025). Accurate performance prediction alone does not validate distances between architecture representations as measures of novelty.

Descriptors also determine which differences evolutionary search rewards. Novelty search rewards distance from previously observed behaviors, while MAP-Elites retains high-performing solutions within descriptor-defined regions (Lehman & Stanley, 2011; Mouret & Clune, 2015). In LLMatic, discussed in Section 2.2, the network archive uses width-to-depth ratio and floating-point operations (FLOPs) as descriptors (Nasir et al., 2024). BOP-Elites models both quality and descriptors with surrogates to select evaluations when these quantities are expensive to obtain (Kent et al., 2025). For SNNs, architectures with similar size and computational cost can nevertheless differ in how spike activity interacts with state updates and gating. We estimate novelty relative to a reference population by comparing code, design rationale, and a behavioral fingerprint combining architecture statistics with initialization-time probes. The same fingerprint supports performance prediction, and both estimates guide the selection of candidates for expensive training evaluations.

Refer to caption
Figure 2: Overview of OpenArchEvo. The inner loop generates feasible SNN architecture programs and uses surrogate predictions to guide evolution. The outer loop selects a small training batch using predicted performance and estimated novelty, then adds measured outcomes to the evaluated archive. Algorithms 1 and 2 in Appendix C give the procedures.

3 Method

OpenArchEvo discovers SNN architecture programs by coupling program-level variation with three-view architecture representation and surrogate-assisted evolution (Fig. 2). Code, design rationale, and a behavioral fingerprint provide complementary descriptions for comparing generated architectures. The fingerprint also supports near-duplicate screening and performance prediction, allowing an inner evolutionary loop to explore many candidates before an outer loop allocates expensive training evaluations. Performance–novelty selection encourages the retention of distinct architectural directions, and measured training outcomes guide subsequent search.

3.1 Preliminaries on SNNs

An SNN processes information through interconnected spiking neurons, whose internal states evolve over time and whose outputs are discrete binary events (Neftci et al., 2019). In a discrete-time description, a neuron’s binary output records whether it fires at each step, while its membrane potential remains continuous-valued. A common example is the leaky integrate-and-fire (LIF) neuron. In its hard-reset form, the pre-reset potential v~τ\tilde{v}_{\tau}, spike sτs_{\tau}, and membrane state vτv_{\tau} at step τ\tau satisfy

v~τ=λvτ−1+Iτ,sτ=𝕀[v~τ≥θ],vτ=(1−sτ)v~τ,\tilde{v}_{\tau}=\lambda v_{\tau-1}+I_{\tau},\quad s_{\tau}=\mathbb{I}[\tilde{v}_{\tau}\geq\theta],\quad v_{\tau}=(1-s_{\tau})\tilde{v}_{\tau}, (1)

where IτI_{\tau}, λ\lambda, and θ\theta denote input current, leakage factor, and firing threshold, and 𝕀⁡[⋅]\mathbb{I}[\cdot] denotes the indicator function. The neuron integrates the current with its decayed membrane state, emits a spike when the resulting potential reaches or exceeds the threshold, and resets its state to zero after firing.

Within the network, incoming spikes contribute to a neuron’s input current through weighted synaptic connections. The resulting spike response depends on both the incoming signals and the neuron’s preceding membrane state. Binary synaptic inputs also permit weighted sums to be evaluated by accumulating the weights associated with active spikes, providing opportunities for sparse computation (Davies et al., 2018; Horowitz, 2014).

3.2 Problem Formulation

Let 𝒜\mathcal{A} denote the search space of admissible SNN block architectures, each represented by executable code and feeding binary spikes to its parameter-dominant feature projections (the spiking-projection constraint; Appendix A.2). Architecture performance is task-dependent, so our main search combines WikiText-2 language modeling and ListOps hierarchical expression evaluation (Merity et al., 2017; Nangia & Bowman, 2018) to assess each design beyond a single task. Their weak rank agreement in our cross-task comparison supports their use as complementary evaluation signals (Appendix A.1). For a∈𝒜a\in\mathcal{A}, let ww collect the separate trainable parameter vectors of the full models instantiated with aa for the evaluation tasks, and let ℒtrain​(a,w)\mathcal{L}_{\mathrm{train}}(a,w) sum their training losses. We aggregate their evaluation scores into a single performance objective y⁡(a,w)y(a,w), with larger values indicating better performance. Let ℛ\mathcal{R} be a finite set of reference architectures, held fixed for each ranking, and let 𝒩⁡(a,ℛ)\mathcal{N}(a;\mathcal{R}) denote the estimated architectural novelty of aa relative to ℛ\mathcal{R}. To encourage distinct architectural alternatives alongside performance, we model discovery as a multiobjective optimization problem following NAS formulations (Lu et al., 2021; Lu et al., 2024) and novelty-based multiobjectivization (Mouret & Doncieux, 2012), mathematically as follows:

maximizea∈𝒜\displaystyle\underset{a\in\mathcal{A}}{\operatorname{maximize}} (y⁡(a,w∗​(a)),𝒩⁡(a,ℛ)),\displaystyle\bigl(y(a,w^{*}(a)),\mathcal{N}(a;\mathcal{R})\bigr), (2)
subject to\displaystyle\text{subject to} w∗​(a)∈arg​minw⁡ℒtrain​(a,w).\displaystyle w^{*}(a)\in\operatorname*{arg\,min}_{w}\mathcal{L}_{\mathrm{train}}(a,w).

In practice, fixed, finite training protocols approximate the lower-level optimization. Measuring a candidate’s fitness in the main search requires separate training on both datasets, increasing evaluation cost relative to training the same candidate on either task alone. Novelty can instead be estimated before weight training. To limit evaluation cost, each search run evaluates at most BB additional architectures under its prescribed task and training protocol, excluding those evaluated initially.

Figure 3: Three-view representation of an SNN architecture program. Code and design rationale describe its implementation and intended choices; a behavioral fingerprint combines architecture statistics with initialization-time response probes. Pairwise comparisons across the views estimate novelty. The fingerprint alone supplies architecture features for surrogate performance prediction and near-duplicate screening.

3.3 Three-View SNN Architecture Representation

To guide the discovery of useful native SNN architectures, we need to estimate architectural differences among generated candidates before committing to weight training. Source-code differences alone need not reflect changes in architectural mechanisms. We therefore propose a three-view representation for SNN architectures (Fig. 3): code describes implemented operations and dataflow; the design rationale states the intended architectural idea, following EoH’s idea–code pairing (Liu et al., 2024); and a numerical behavioral fingerprint records structural attributes and initialization-time responses. Together, these views provide partial evidence for comparing candidates generated in the code space.

Inspired by response-based program characterization (Hildebrandt & Branke, 2015; Zhang & Lu, 2026), we probe the instantiated SNN before weight training. Probe inputs, sequence shape, initialization procedure, and seed are fixed across candidate comparisons. We adapt SWAP’s sample-wise pattern count (Peng et al., 2024b) to binary spike outputs, yielding sample-wise spiking patterns (SWSP): each recorded neuron–sequence-position response forms a pattern across probe samples, and SWSP counts the distinct patterns. Alongside SWSP, we use mean spike activity across spiking layers (FireRate) and five established activation- and gradient-based proxies (Abdelfattah et al., 2021; Mellor et al., 2021), giving seven initialization-time measurements. We supplement them with four module statistics describing components and ten structural statistics describing dimension flow, computational organization, and resource allocation. We construct the behavioral fingerprint 𝐱⁡(a)\mathbf{x}(a) as a 21-dimensional feature vector by concatenating these measurements. Feature definitions are given in Appendix C.1.

For architectures a1a_{1} and a2a_{2}, we adopt CodeBLEU (Ren et al., 2020) and average its two comparison directions to obtain code similarity scode​(a1,a2)s_{\mathrm{code}}(a_{1},a_{2}). For rationale similarity srat​(a1,a2)s_{\mathrm{rat}}(a_{1},a_{2}), we encode each LLM-generated design rationale with the pretrained text-embedding model text-embedding-v4 (Alibaba Cloud, 2025) and compute cosine similarity between the resulting vectors. Fingerprint similarity sfp​(a1,a2)s_{\mathrm{fp}}(a_{1},a_{2}) uses cosine similarity after fixed feature-wise affine scaling of 𝐱⁡(a1)\mathbf{x}(a_{1}) and 𝐱⁡(a2)\mathbf{x}(a_{2}) (Appendix C.2). We define pairwise dissimilarity using fixed, untuned equal weights:

d(a1,a2)=1−13[\displaystyle d(a_{1},a_{2})=1-\tfrac{1}{3}\bigl[ scode​(a1,a2)+srat​(a1,a2)\displaystyle s_{\mathrm{code}}(a_{1},a_{2})+s_{\mathrm{rat}}(a_{1},a_{2}) (3)
+sfp(a1,a2)].\displaystyle+s_{\mathrm{fp}}(a_{1},a_{2})\bigr].

Inspired by novelty search (Lehman & Stanley, 2011), we estimate novelty as the mean dissimilarity to all other reference architectures:

𝒩⁡(a,ℛ)=∑a′∈ℛ∖{a}d⁡(a,a′)|ℛ∖{a}|,|ℛ∖{a}|>0.\mathcal{N}(a;\mathcal{R})=\frac{\sum_{a^{\prime}\in\mathcal{R}\setminus\{a\}}d(a,a^{\prime})}{|\mathcal{R}\setminus\{a\}|},\quad|\mathcal{R}\setminus\{a\}|>0. (4)

The surrogate in Section 3.4 uses evaluated architectures’ fingerprints and measured task outcomes to predict task performance from an untrained candidate’s 𝐱⁡(a)\mathbf{x}(a).

3.4 Surrogate-Assisted Evolutionary Search

To limit costly weight training, we use an inner loop for surrogate-guided architecture evolution and an outer loop for training allocation and archive updates, i.e., surrogate model management (Jin, 2011; Zhang et al., 2010). Let 𝒟0\mathcal{D}_{0} denote the initial archive of evaluated expert architectures, and 𝒟t\mathcal{D}_{t} the archive after outer iteration tt. Each archived architecture has an associated fingerprint and measured task outcomes. We adopt TabPFN-2.5 (Hollmann et al., 2025; Grinsztajn et al., 2025) as the surrogate, using these records from 𝒟t−1\mathcal{D}_{t-1} as labeled context. Given 𝐱⁡(a)\mathbf{x}(a), the surrogate predicts task metrics without training candidate aa. The same fitness mapping used for measured outcomes aggregates these predictions into y^t​(a)\hat{y}_{t}(a); hereafter, y⁡(a)y(a) denotes fitness measured after the prescribed training protocol.

We initialize the inner-loop populations from 𝒟t−1\mathcal{D}_{t-1} using NSGA-II’s nondominated sorting and crowding-distance truncation (Deb et al., 2002), based on measured fitness y⁡(a)y(a) and novelty 𝒩⁡(a,𝒟t−1)\mathcal{N}(a;\mathcal{D}_{t-1}). The selected architectures seed separate island populations, from which parents are sampled within or across islands. The LLM revises or recombines their code and design rationales. Feasible offspring pass fingerprint-based near-duplicate screening before surrogate evaluation. Predicted fitness guides subsequent parent selection and survival within islands, while accepted candidates accumulate in a pool for possible training.

Before selecting architectures for training, we shortlist candidates to avoid comparing every pair of generated architectures. We retain a subset prioritized by predicted performance and supplement it with candidates having high mean three-view dissimilarity to that subset. After screening against the evaluated archive for near-duplicates, we apply the same selection procedure using predicted fitness and novelty relative to the filtered pool, which remains fixed throughout selection. The resulting batch contains previously unevaluated architectures, limited by the maximum batch size KK and remaining budget. The selected architectures are trained under the prescribed task protocols and added to 𝒟t\mathcal{D}_{t} with their measured outcomes, supplying additional surrogate context and parent candidates for the next iteration. Across the prescribed iterations, at most BB additional architectures are evaluated under these protocols. Search returns the evaluated archive and its nondominated subset under measured fitness and novelty relative to the final archive. Appendix C gives the complete procedures and budget constraints.

4 Experiments

Table 1: Sequence-modeling performance and estimated energy reduction. WT103: WikiText-103 language modeling (Merity et al., 2017); LRA: Long Range Arena (Tay et al., 2021). Energy reduction is the estimated arithmetic energy of a dense Transformer of the same size divided by that of each SNN under a common architecture setting (Section 4.2 and Appendix F). Bold marks the best result in each column; “–” denotes unavailable results.
Model WT103 LRA accuracy (%)↑\uparrow Energy reduction↑\uparrow
PPL↓\downarrow ListOps Text Retrieval Image Pathfinder Path-X Avg.
ANN baselines
DeltaNet (Yang et al., 2024b) 27.5 62.2 85.1 91.3 89.8 94.2 94.9 86.2 –
S4D-Lin† (Gu et al., 2022) – 60.5 87.0 91.0 87.9 94.0 92.8∗ 85.5 –
SNN baselines
S6-based SNN† – 55.7 77.6 88.5 80.1 83.4 – – –
Dyn-SSM† (Zhong et al., 2026) 33.2 60.2 82.4 88.8 87.2 92.0 94.4 84.2 33.1×\times
SpikingDeltaNet 34.5 57.3 82.2 90.5 89.0 91.2 92.3 83.7 24.7×\times
SpikingMamba2 38.2 48.5 74.0 81.0 83.2 90.4 89.6 77.8 14.8×\times
Discovered SNNs (OpenArchEvo)
NeuroGate 26.4 61.8 82.4 90.1 91.2 93.6 95.3 85.7 31.7×\times
HomeoResSSM 28.1 46.2 84.8 92.3 87.6 86.3 85.2 80.4 17.3×\times
LoopMem 27.8 59.5 81.0 89.5 90.5 88.4 88.7 82.9 50.6×\times
†Task results quoted from Shen et al. (2025); Zhong et al. (2026). ∗S4D-Inv result.

We evaluate the discovered SNNs, examine their architectural mechanisms, and analyze how representation and search design contribute to the discovery.

4.1 Experimental Setup

Candidate blocks share the sequence-model wrapper described in Appendix D.1. During the main search, each selected candidate is trained separately for 30 epochs on WikiText-2 (WT2) language modeling and 25 epochs on one-third of the ListOps training set (Merity et al., 2017; Tay et al., 2021), producing distinct model weights for the two tasks. We instantiate the composite fitness as

y⁡(a)=12​ϕWT2​(PPLWT2​(a))+12​ϕListOps​(AccListOps​(a)),y(a)=\tfrac{1}{2}\phi_{\mathrm{WT2}}\!\left(\mathrm{PPL}_{\mathrm{WT2}}(a)\right)+\tfrac{1}{2}\phi_{\mathrm{ListOps}}\!\left(\mathrm{Acc}_{\mathrm{ListOps}}(a)\right), (5)

where the task transforms use SpikingDeltaNet as the 0.5000.500 fitness anchor (Appendix D.2).

Figure 4: Measured fitness of trained candidates during the main search, in evaluation order. E denotes the initial expert models (grey) and 1–10 the search iterations. The red trace shows the best-so-far fitness, including the experts, and stars mark its improvements. The dashed line marks the baseline fitness of 0.5000.500, attained by SpikingDeltaNet, the strongest expert.

The initial archive contains 20 expert SNN adaptations of linear recurrent and state-space architectures (Appendix C.3). Gemini-2.5-Flash (Comanici et al., 2025), selected after a comparison of candidate LLMs (Appendix C.5), is used as the LLM to generate candidates, with prompts in Appendix G, and the surrogate specified in Section 3.4 predicts their task metrics. The main search runs for 10 outer iterations with a batch limit of K=16K=16 and a budget of B=160B=160 architecture evaluations, each comprising both task-training runs; candidates that duplicate archived architectures are removed, so 151 new architectures are trained. Candidate weight training costs an estimated 132 V100 GPU-days (3,171 GPU-hours). Appendix C.4 gives the search configuration and resource accounting. The code, generated programs, and evaluated archive will be released.

To evaluate the selected architectures at full scale, we retrain them on WikiText-103 (WT103) (Merity et al., 2017) and the six Long Range Arena (LRA) tasks, whose sequences span 1K–16K tokens (Tay et al., 2021), following the task-specific configurations of SpikingSSMs (Shen et al., 2025); the WT103 configuration matches that of Dyn-SSM (Zhong et al., 2026). Unlike the search-time setting, ListOps is trained on the full training split and evaluated on the same test split, whereas WT103 and the other five LRA tasks are not used during search (Appendix D.1). The LRA average covers all six tasks. The ANN reference, DeltaNet, is among the strongest of the 20 experts and 97 ASI-Arch models trained under our common protocol. S4D-Lin and S6-based SNN results are quoted from Shen et al. (2025) and Dyn-SSM results from Zhong et al. (2026); all other models, including our spiking adaptations SpikingDeltaNet and SpikingMamba2, are trained under our protocol. The main search comprises one run; Sections 4.4 and 4.5 report repeated smaller-scale studies.

4.2 Main Results

Figure 4 traces the main search. Candidates that surpass the strongest expert appear in every iteration from the second onward, and the best fitness improves in five separate iterations, reaching 0.8720.872 at iteration 9. Progress is therefore sustained rather than confined to an early proposal, consistent with the archive supplying better parents and surrogate context as it grows. The nondominated front also advances in the two task metrics, eventually containing candidates that surpass SpikingDeltaNet on both WT2 and ListOps (Appendix C.6).

Refer to caption
Figure 5: HomeoResSSM and LoopMem. HomeoResSSM uses spike activity to modulate two residual branches. LoopMem feeds a normalized summary of its recurrent state back to the forget gate and adds activity penalties to the training loss. Red paths highlight activity and feedback connections; paths to the loss act during training. NeuroGate is shown in Fig. 1(c).

Using search-time measurements only, we select three architectures that represent distinct regions of the final Pareto front in composite fitness and novelty (Fig. 10): NeuroGate, the extreme solution with the highest fitness; HomeoResSSM, the knee point, which lies farthest from the line joining the two extreme solutions after normalization (Das, 1999; Zhang et al., 2015); and LoopMem, the most novel member that retains nontrivial performance on both tasks.

The selected architectures remain strong after full-scale retraining (Table 1). On WT103, all three outperform the listed SNN baselines, and NeuroGate also surpasses the ANN DeltaNet trained under our protocol (26.4 versus 27.5 perplexity), although its parameter-dominant projections receive binary spikes. On LRA, NeuroGate attains the highest average accuracy among the SNNs, between the two ANN baselines, and remains the highest among the SNNs without ListOps (90.5 versus 89.0), while HomeoResSSM gives the best Retrieval accuracy among all listed models. The discovered architectures thus combine competitive accuracy with distinct task profiles.

These results come with substantially lower estimated energy. Counting spike-driven projections as accumulate operations and continuous computation as multiply–accumulate operations, with each model’s measured firing rate and a shared model wrapper, the discovered architectures reduce the arithmetic energy of a dense Transformer of the same size by 1717–51×51\times (Appendix F, which also details the Dyn-SSM estimate). The spiking feed-forward network (FFN) and output head of the shared wrapper account for most spike-driven operations, so differences among the SNNs mainly reflect their firing rates and token-mixer arithmetic. NeuroGate fires about half as often as SpikingDeltaNet, so its additional control computation still yields a larger reduction, and LoopMem attains the largest reduction among all listed models. We next examine which evolved elements distinguish the discovered architectures and whether their effects depend on spiking computation.

4.3 Spike-Native Mechanisms in the Discovered Architectures

The discovered architectures mainly contain two kinds of elements that we regard as native to SNNs. ① Spike-activity-dependent computation takes spike activity itself as an input. NeuroGate maps the mean input spike activity of its query, key, and value projections at each position to gains on the delta-rule update coefficient and the output gate (Fig. 1(c)), and HomeoResSSM uses the spike activity of its state-space and channel-mixing branches to scale their residual contributions (Fig. 5(a)); we test this class on HomeoResSSM. Because each neuron integrates its input across tokens and resets on firing, this activity reflects membrane dynamics that have no counterpart in ANN units, so the control reads a signal that exists only in spiking units. ② Spiking-specific design choices do not read spikes explicitly, but depart from ANN practice in either direction: components that are uncommon in ANN design are added when they benefit the SNN, and components that are standard in ANNs are removed when they degrade it. Direct ANN-to-SNN adaptation introduces neither kind of change. LoopMem illustrates the first direction, feeding a normalized summary of its recurrent state back to the forget gate (state-norm feedback) (Fig. 5(b)); NeuroGate illustrates the second, omitting the sigmoid linear unit (SiLU) activation that DeltaNet applies after its query, key, and value convolutions.

Table 2: Ablations of evolved elements on WT2. Block ① tests whether a control relies on spike activity; block ② tests design choices that depart from ANN practice. Non-spiking variants remove the spiking neurons and all spike-dependent components. Δ\Delta is relative to the closest row above marked “–”.
Variant WT2 PPL↓\downarrow Δ\Delta
① Spike-activity-dependent computation
HomeoResSSM
Discovered 57.5 –
Control driven by continuous inputs 59.2 +1.7+1.7
Control removed 64.4 +6.9+6.9
② Spiking-specific design choices
NeuroGate
Discovered (SiLU removed) 57.2 –
SiLU restored 63.1 +5.9+5.9
LoopMem
Discovered 60.5 –
State-norm feedback removed 69.3 +8.8+8.8
Non-spiking, with feedback 61.3 –
Non-spiking, feedback removed 59.7 −1.6-1.6

Table 2 examines these elements on WT2. All variants, including the discovered architectures, are trained with the fixed protocol applied to every search candidate, and we compare the effect of each element within an implementation rather than perplexities across implementations. For ①, removing HomeoResSSM’s control raises perplexity by 6.9. Driving the same controller with the channel mean of the continuous inputs, from which the spikes are generated, recovers most of this benefit, yet the spike-driven control still attains 1.7 lower perplexity. For ②, restoring SiLU raises NeuroGate’s perplexity by 5.9. Removing LoopMem’s state-norm feedback raises perplexity by 8.8 in the SNN, whereas the same removal lowers the perplexity of the non-spiking variant by 1.6. LoopMem’s feedback thus has opposite effects in the two implementations, and the ANN-standard SiLU degrades the SNN.

From the search-space perspective, each control element combines several decisions: which signal to summarize, how to transform it, and which update, branch, or gate it modulates. In OpenArchEvo, the LLM can propose these decisions within one code revision of a parent program. Among seven published encodings, covering SNN cell and block spaces, recurrent-cell spaces, and a recursive architecture grammar, none can select a spike-statistic or state-summary input, a learned controller, and its target together; directly representing these elements requires extending the encoding (Appendix B, Table 5). Open-code search can express such relations, but the best architecture of our adapted ASI-Arch run, SpikingCondFuse, conditions its gate on hidden-state statistics rather than spike activity (Appendix D.4). Appendix E gives the equations of the evolved elements.

4.4 Comparison with Existing Search Methods

Figure 6: Inner-loop search comparison in six-iteration WT2 searches: (a) best fitness and (b) hypervolume of the nondominated set in fitness and novelty after each stage. The EoH and FunSearch adaptations replace our inner-loop search and share the outer evaluation and surrogate-selection loop. Curves show means over three runs; shaded regions indicate standard deviation.
Table 3: Outcomes and organization of the completed program-search runs. Downstream results compare SpikingCondFuse and NeuroGate, respectively.
Adapted ASI-Arch OpenArchEvo
Best fitness↑\uparrow 0.653 0.872
Top-10 mean fitness↑\uparrow 0.570 0.714
WT103 PPL↓\downarrow 29.6 26.4
ListOps accuracy↑\uparrow 55.3 61.8
Text accuracy↑\uparrow 82.3 82.4
Retrieval accuracy↑\uparrow 85.3 90.1
Energy reduction‡ 18.9×\times 31.7×\times
LLM agents 9 1
Prompt tokens† ∼\sim25K ∼\sim1.2K
API expenditure (USD) ∼\sim440 ∼\sim170

One search run per method; training budgets differ as stated in the text. †Prompt tokens per executable program. ‡Computed as in Table 1.

Inner-loop comparisons. We compare our inner-loop search with adaptations of EoH (Liu et al., 2024) and FunSearch (Romera-Paredes et al., 2024), which replace the inner loop while sharing our outer evaluation and surrogate-selection loop. Their original designs evaluate every generated program, which is impractical when each evaluation requires training an architecture. The EoH adaptation uses thought–code evolution in a single population, whereas the FunSearch adaptation uses code-based best-shot prompting with islands. Our inner loop features population management and redundancy control guided by the three-view representation and novelty. In the main search, the trained expert archive supplies the initial surrogate context and parents. Each WT2-only search instead begins with a warm-start stage that trains a first batch of diversified candidates. Each configuration then runs six outer iterations of WT2-only search, scored by a WT2-only fitness of the same form (Appendices D.2 and D.4). OpenArchEvo attains the highest final best fitness and hypervolume (Fig. 6). All configurations use the same total training budget.

End-to-end program-search comparison. We retain the agent workflow of ASI-Arch (Liu et al., 2025), adapt its evaluation tasks, and impose the spiking constraints. Table 3 reports the resulting search outcomes and representative architectures. The ASI-Arch and OpenArchEvo runs train 39 and 151 architectures over six and four weeks, respectively. OpenArchEvo uses one program generator; the API expenditure excludes candidate-training compute. The best and Top-10 mean fitness are higher in our completed run. Appendix D.4 provides the adaptation details.

4.5 Search Analysis and Ablation Studies

Figure 7: Screening raw proposals in the main search. Compilation failures and causal-check failures account for 18% and 3%; fingerprint-based near-duplicate screening removes 30%, leaving 49% accepted. Rounded percentages share the raw-proposal denominator.

Near-duplicate screening. Fingerprint-based near-duplicate screening removes 30% of raw proposals (Fig. 7). The shared initialization-time probes and architecture statistics identify these candidates before surrogate-guided selection.

Surrogate prediction. The surrogate studies were conducted before the main search and guided the choice of its surrogate and fingerprint (Appendix D.3). Initialization-time probes capture network responses and form the starting point for our surrogate-input study (Fig. 8(b)). Adding SWSP and firing rate to the classical proxies improves ranking on WT2, with a smaller change on ListOps. Supplementing these probes with module and structural statistics gives the strongest ranking along the tested sequence. With the complete fingerprint, TabPFN-2.5 ranks best among the tested regressors on both tasks (Fig. 8(a)). In these studies, prediction generally improves as the labeled context grows (Fig. 8(c)). Early in search, candidates are selected for training by a surrogate with little labeled context. In WT2-only searches without the warm-start stage, which start from the expert archive alone, the best fitness after six iterations is 0.56 instead of 0.68, and the gap persists across iterations (Appendix D.4).

Table 4: Selection variants in six-iteration WT2 searches (three runs). Statistics are computed over the final-iteration trained candidates pooled across runs: Top-1 and Top-5 give the lowest and the mean of the five lowest PPL, Novelty the mean novelty, and Variance the variance of WT2 fitness.
Selection method Top-1 PPL↓\downarrow Top-5 PPL↓\downarrow Novelty↑\uparrow Variance↓\downarrow
Evolution + surrogate 56.2 57.4 0.534 0.0281
Single island + surrogate 57.3 57.9 0.513 0.0272
Sampling + surrogate 57.3 57.8 0.520 0.0314
Sampling only 57.5 58.5 0.511 0.0414

Search ablations. We compare candidate-generation and selection variants in six-iteration WT2-only searches, each repeated three times. Besides the full configuration, which evolves candidates on multiple islands and selects them with the surrogate, the variants evolve a single island, sample candidates from the LLM without evolution before surrogate selection, or sample and select without the surrogate. The full configuration has the lowest Top-1 and Top-5 PPL and the highest novelty, while the single-island variant has the lowest fitness variance (Table 4).

Figure 8: Offline surrogate studies on WT2 (top, 223 architectures) and ListOps (bottom, 54). Kendall τ\tau compares regressors (a), progressive surrogate-input additions (b), and labeled-context sizes (c). The first two studies share repeated cross-validation splits. In (b), gray shades mark the measurement group added at each cumulative step, and the task color marks the complete fingerprint. The +SWSP step adds both SWSP and FireRate; +Par. adds the parameter count and +Comp the other module statistics, while +Dim, +Res., and +Graph add the dimension-flow, resource-allocation, and subgraph statistics (Table 6). Learning curves use fixed holdouts and 20 training-pool subsamples per size, with ±1​σ\pm 1\sigma bands. Appendix D.3 gives the protocols.

5 Discussion

Program evolution discovers changes in control and pathway design around established recurrent operators, including how spike activity regulates computation. It shifts part of architecture design from enumerating admissible mechanisms to specifying executable constraints and evaluating proposed programs. The resulting archive contains both usable architectures and design hypotheses for further study.

For expensive evolutionary search, the framework separates abundant program variation from limited training evaluations. A shared characterization supports redundancy filtering, population organization, and surrogate prediction, allowing these components to improve together. The evidence comprises one full search with smaller repeated studies; the ASI-Arch comparison uses unequal budgets. Broader validation across search runs and domains remains necessary, as fingerprints approximate redundancy and novelty depends on the reference population.

The discovered SNNs demonstrate that program evolution can identify competitive spiking sequence architectures with spike-activity-dependent computation and spiking-specific design choices. Their energy reductions are theoretical arithmetic estimates; realizing and measuring these reductions requires implementations on specific hardware that account for memory access and execution costs. The present search fixes the neuron model. Extending it to jointly evolve neuron dynamics and architecture, and evaluating broader and more demanding downstream tasks, are natural next steps. These results show that program search can discover competitive SNN architectures, which motivates extending it to those harder settings.

6 Conclusion

We introduced OpenArchEvo for automated discovery of native neural architectures, with spiking sequence modeling as a demanding test case. The discovered architectures combine competitive predictive performance with substantial estimated arithmetic-energy savings, demonstrating the value of exploring computations tailored to spike activity. Beyond the resulting models, the study shows how executable architectural proposals can become a cumulative record of empirically evaluated designs. The broader opportunity is to connect expressive program generation with meaningful architectural comparison and selective evaluation, enabling discovery in neural computing domains where useful mechanisms are difficult to specify in advance.

Acknowledgments

Large language models were used to polish the language of this manuscript; the authors take full responsibility for its content.

References

  • Abdelfattah et al. (2021) Mohamed S. Abdelfattah, Abhinav Mehrotra, Łukasz Dudziak, and Nicholas D. Lane. Zero-cost proxies for lightweight NAS. In International Conference on Learning Representations (ICLR), 2021.
  • Alibaba Cloud (2025) Alibaba Cloud. Embedding: text-embedding-v4. Alibaba Cloud Model Studio documentation, 2025.
  • Chen et al. (2023) Angelica Chen, David Dohan, and David So. EvoPrompting: Language models for code-level neural architecture search. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
  • Cheng et al. (2025) Junyan Cheng, Peter Clark, and Kyle Richardson. Language modeling by language models. In Advances in Neural Information Processing Systems (NeurIPS), 2025.
  • Comanici et al. (2025) Gheorghe Comanici, Eric Bieber, Mike Schaekermann, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025.
  • Das (1999) Indraneel Das. On characterizing the “knee” of the Pareto curve based on normal-boundary intersection. Structural Optimization, 18(2–3):107–115, 1999.
  • Davies et al. (2018) Mike Davies, Narayan Srinivasa, Tsung-Han Lin, Gautham Chinya, Yongqiang Cao, Sri Harsha Choday, Georgios Dimou, Prasad Joshi, Nabil Imam, Shweta Jain, et al. Loihi: A neuromorphic manycore processor with on-chip learning. IEEE Micro, 38(1):82–99, 2018.
  • Deb et al. (2002) Kalyanmoy Deb, Amrit Pratap, Sameer Agarwal, and T. Meyarivan. A fast and elitist multiobjective genetic algorithm: NSGA-II. IEEE Transactions on Evolutionary Computation, 6(2):182–197, 2002.
  • Ericsson et al. (2024) Linus Ericsson, Miguel Espinosa, Chenhongyi Yang, Antreas Antoniou, Amos Storkey, Shay B. Cohen, Steven McDonagh, and Elliot J. Crowley. einspace: Searching for neural architectures from fundamental operations. In Advances in Neural Information Processing Systems (NeurIPS), 2024.
  • Grinsztajn et al. (2025) Léo Grinsztajn, Klemens Flöge, Oscar Key, Felix Birkel, Philipp Jund, Brendan Roof, Benjamin Jäger, Dominik Safaric, Simone Alessi, Adrian Hayler, et al. TabPFN-2.5: Advancing the state of the art in tabular foundation models. arXiv preprint arXiv:2511.08667, 2025.
  • Gu et al. (2022) Albert Gu, Karan Goel, Ankit Gupta, and Christopher Ré. On the parameterization and initialization of diagonal state space models. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  • Hildebrandt & Branke (2015) Torsten Hildebrandt and Jürgen Branke. On using surrogates with genetic programming. Evolutionary Computation, 23(3):343–367, 2015.
  • Hollmann et al. (2025) Noah Hollmann, Samuel Müller, Lennart Purucker, Arjun Krishnakumar, Max Körfer, Shi Bin Hoo, Robin Tibor Schirrmeister, and Frank Hutter. Accurate predictions on small data with a tabular foundation model. Nature, 637(8045):319–326, 2025.
  • Horowitz (2014) Mark Horowitz. 1.1 Computing’s energy problem (and what we can do about it). In IEEE International Solid-State Circuits Conference (ISSCC), 2014.
  • Jin (2011) Yaochu Jin. Surrogate-assisted evolutionary computation: Recent advances and future challenges. Swarm and Evolutionary Computation, 1(2):61–70, 2011.
  • Kent et al. (2025) Paul Kent, Adam Gaier, Jean-Baptiste Mouret, and Juergen Branke. Bayesian optimization for quality diversity search with coupled descriptor functions. IEEE Transactions on Evolutionary Computation, 29(2):302–316, 2025.
  • Kim et al. (2022) Youngeun Kim, Yuhang Li, Hyoungseob Park, Yeshwanth Venkatesha, and Priyadarshini Panda. Neural architecture search for spiking neural networks. In European Conference on Computer Vision (ECCV), 2022.
  • Lai et al. (2026) Yutao Lai, Zicheng Cai, Lei Chen, Tongtao Ling, and Hai-Lin Liu. LLMENAS: Evolutionary neural architecture search via large language model guidance. IEEE Transactions on Evolutionary Computation, 30(4):1362–1376, 2026.
  • Lange et al. (2026) Robert Tjarko Lange, Yuki Imajuku, and Edoardo Cetin. ShinkaEvolve: Towards open-ended and sample-efficient program evolution. In International Conference on Learning Representations (ICLR), 2026.
  • Lehman & Stanley (2011) Joel Lehman and Kenneth O. Stanley. Abandoning objectives: Evolution through the search for novelty alone. Evolutionary Computation, 19(2):189–223, 2011.
  • Li et al. (2024) Nan Li, Lianbo Ma, Guo Yu, Bing Xue, Mengjie Zhang, and Yaochu Jin. Survey on evolutionary deep learning: Principles, algorithms, applications, and open issues. ACM Computing Surveys, 56(2):41:1–41:34, 2024.
  • Liu et al. (2024) Fei Liu, Xialiang Tong, Mingxuan Yuan, Xi Lin, Fu Luo, Zhenkun Wang, Zhichao Lu, and Qingfu Zhang. Evolution of heuristics: Towards efficient automatic algorithm design using large language model. In International Conference on Machine Learning (ICML), 2024.
  • Liu et al. (2025) Yixiu Liu, Yang Nan, Weixian Xu, Xiangkun Hu, Lyumanshan Ye, Zhen Qin, and Pengfei Liu. AlphaGo moment for model architecture discovery. arXiv preprint arXiv:2507.18074, 2025.
  • Liu et al. (2023) Yuqiao Liu, Yanan Sun, Bing Xue, Mengjie Zhang, Gary G. Yen, and Kay Chen Tan. A survey on evolutionary neural architecture search. IEEE Transactions on Neural Networks and Learning Systems, 34(2):550–570, 2023.
  • Lu et al. (2021) Zhichao Lu, Ian Whalen, Yashesh Dhebar, Kalyanmoy Deb, Erik D. Goodman, Wolfgang Banzhaf, and Vishnu Naresh Boddeti. Multiobjective evolutionary design of deep convolutional neural networks for image classification. IEEE Transactions on Evolutionary Computation, 25(2):277–291, 2021.
  • Lu et al. (2024) Zhichao Lu, Ran Cheng, Yaochu Jin, Kay Chen Tan, and Kalyanmoy Deb. Neural architecture search as multiobjective optimization benchmarks: Problem formulation and performance assessment. IEEE Transactions on Evolutionary Computation, 28(2):323–337, 2024.
  • Ma et al. (2026) Zeyuan Ma, Hongshu Guo, Yue-Jiao Gong, Jun Zhang, and Kay Chen Tan. Toward automated algorithm design: A survey and practical guide to meta-black-box-optimization. IEEE Transactions on Evolutionary Computation, 30(2):667–687, 2026.
  • Mellor et al. (2021) Joseph Mellor, Jack Turner, Amos Storkey, and Elliot J. Crowley. Neural architecture search without training. In International Conference on Machine Learning (ICML), 2021.
  • Merity et al. (2017) Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. In International Conference on Learning Representations (ICLR), 2017.
  • Mouret & Clune (2015) Jean-Baptiste Mouret and Jeff Clune. Illuminating search spaces by mapping elites. arXiv preprint arXiv:1504.04909, 2015.
  • Mouret & Doncieux (2012) Jean-Baptiste Mouret and Stéphane Doncieux. Encouraging behavioral diversity in evolutionary robotics: An empirical study. Evolutionary Computation, 20(1):91–133, 2012.
  • Na et al. (2022) Byunggook Na, Jisoo Mok, Seongsik Park, Dongjin Lee, Hyeokjun Choe, and Sungroh Yoon. AutoSNN: Towards energy-efficient spiking neural networks. In International Conference on Machine Learning (ICML), 2022.
  • Nangia & Bowman (2018) Nikita Nangia and Samuel R. Bowman. ListOps: A diagnostic dataset for latent tree learning. In Conference of the North American Chapter of the Association for Computational Linguistics: Student Research Workshop (NAACL SRW), 2018.
  • Nasir et al. (2024) Muhammad Umair Nasir, Sam Earle, Julian Togelius, Steven James, and Christopher W. Cleghorn. LLMatic: Neural architecture search via large language models and quality diversity optimization. In Genetic and Evolutionary Computation Conference (GECCO), 2024.
  • Neftci et al. (2019) Emre O. Neftci, Hesham Mostafa, and Friedemann Zenke. Surrogate gradient learning in spiking neural networks: Bringing the power of gradient-based optimization to spiking neural networks. IEEE Signal Processing Magazine, 36(6):51–63, 2019.
  • Novikov et al. (2025) Alexander Novikov, Ngân Vũ, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco J. R. Ruiz, Abbas Mehrabian, et al. AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv preprint arXiv:2506.13131, 2025.
  • Pan et al. (2025) Wenxuan Pan, Feifei Zhao, Guobin Shen, Bing Han, and Yi Zeng. Brain-inspired multiscale evolutionary neural architecture search for deep spiking neural networks. IEEE Transactions on Evolutionary Computation, 29(5):2258–2270, 2025.
  • Peng et al. (2024) Yameng Peng, Andy Song, Haytham M. Fayek, Vic Ciesielski, and Xiaojun Chang. SWAP-NAS: Sample-wise activation patterns for ultra-fast NAS. In International Conference on Learning Representations (ICLR), 2024.
  • Qin et al. (2025) Shiwen Qin, Gabriela Kadlecová, Martin Pilát, Shay B. Cohen, Roman Neruda, Elliot J. Crowley, Jovita Lukasik, and Linus Ericsson. Transferrable surrogates in expressive neural architecture search spaces. In International Conference on Automated Machine Learning (AutoML), 2025.
  • Real et al. (2020) Esteban Real, Chen Liang, David R. So, and Quoc V. Le. AutoML-Zero: Evolving machine learning algorithms from scratch. In International Conference on Machine Learning (ICML), 2020.
  • Ren et al. (2020) Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. CodeBLEU: A method for automatic evaluation of code synthesis. arXiv preprint arXiv:2009.10297, 2020.
  • Romera-Paredes et al. (2024) Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M. Pawan Kumar, Emilien Dupont, Francisco J. R. Ruiz, Jordan S. Ellenberg, Pengming Wang, Omar Fawzi, et al. Mathematical discoveries from program search with large language models. Nature, 625(7995):468–475, 2024.
  • Shen et al. (2025) Shuaijie Shen, Chao Wang, Renzhuo Huang, Yan Zhong, Qinghai Guo, Zhichao Lu, Jianguo Zhang, and Luziwei Leng. SpikingSSMs: Learning long sequences with sparse and parallel spiking state space models. In AAAI Conference on Artificial Intelligence (AAAI), 2025.
  • Suganuma et al. (2017) Masanori Suganuma, Shinichi Shirakawa, and Tomoharu Nagao. A genetic programming approach to designing convolutional neural network architectures. In Genetic and Evolutionary Computation Conference (GECCO), 2017.
  • Sun et al. (2020) Yanan Sun, Bing Xue, Mengjie Zhang, and Gary G. Yen. Evolving deep convolutional neural networks for image classification. IEEE Transactions on Evolutionary Computation, 24(2):394–407, 2020.
  • Tay et al. (2021) Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. Long range arena: A benchmark for efficient transformers. In International Conference on Learning Representations (ICLR), 2021.
  • van Stein & Bäck (2025) Niki van Stein and Thomas Bäck. LLaMEA: A large language model evolutionary algorithm for automatically generating metaheuristics. IEEE Transactions on Evolutionary Computation, 29(2):331–345, 2025.
  • Wu et al. (2025) Xingyu Wu, Sheng-Hao Wu, Jibin Wu, Liang Feng, and Kay Chen Tan. Evolutionary computation in the era of large language model: Survey and roadmap. IEEE Transactions on Evolutionary Computation, 29(2):534–554, 2025.
  • Yang et al. (2024) Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim. Parallelizing linear transformers with the delta rule over sequence length. In Advances in Neural Information Processing Systems (NeurIPS), 2024.
  • Yao et al. (2023) Man Yao, Jiakui Hu, Zhaokun Zhou, Li Yuan, Yonghong Tian, Bo Xu, and Guoqi Li. Spike-driven transformer. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
  • Zhang et al. (2026a) Malu Zhang, Wenjie Wei, Zijian Zhou, Wanlong Liu, Jie Zhang, Ammar Belatreche, and Yang Yang. Spike-driven lightweight large language model with evolutionary computation. IEEE Transactions on Evolutionary Computation, 30(4):1333–1346, 2026a.
  • Zhang et al. (2010) Qingfu Zhang, Wudong Liu, Edward Tsang, and Botond Virginas. Expensive multiobjective optimization by MOEA/D with Gaussian process model. IEEE Transactions on Evolutionary Computation, 14(3):456–474, 2010.
  • Zhang & Lu (2026) Rui Zhang and Zhichao Lu. Rethinking code similarity for automated algorithm design with LLMs. In International Conference on Learning Representations (ICLR), 2026.
  • Zhang et al. (2026b) Xian-Rong Zhang, Yue-Jiao Gong, Wei-Neng Chen, and Jun Zhang. Evolutionary neural architecture search with dual contrastive learning. Applied Soft Computing, 189, 2026b. Art. no. 114507.
  • Zhang et al. (2015) Xingyi Zhang, Ye Tian, and Yaochu Jin. A knee point-driven evolutionary algorithm for many-objective optimization. IEEE Transactions on Evolutionary Computation, 19(6):761–776, 2015.
  • Zhong et al. (2026) Yan Zhong, Ruoyu Zhao, Chao Wang, Jiaqi He, Qinghai Guo, Jianguo Zhang, Zhichao Lu, and Luziwei Leng. Dyn-SSM: Towards the efficient long sequence learning via bio-interpretable dynamics in spiking state space models. IEEE Transactions on Cognitive and Developmental Systems, pp. 1–15, 2026. Early access.
  • Zhou et al. (2025) Xun Zhou, Xingyu Wu, Liang Feng, Zhichao Lu, and Kay Chen Tan. Design principle transfer in neural architecture search via large language models. In AAAI Conference on Artificial Intelligence (AAAI), 2025.
  • Zhou et al. (2023) Zhaokun Zhou, Yuesheng Zhu, Chao He, Yaowei Wang, Shuicheng Yan, Yonghong Tian, and Li Yuan. Spikformer: When spiking neural network meets transformer. In International Conference on Learning Representations (ICLR), 2023.
  • Zhu et al. (2024) Rui-Jie Zhu, Qihang Zhao, Guoqi Li, and Jason K. Eshraghian. SpikeGPT: Generative pre-trained language model with spiking neural networks. Transactions on Machine Learning Research, 2024.

Appendix A Supporting Architecture Comparisons

A.1 Paired ANN and SNN Evaluation

We adapted the 106 architecture programs released by ASI-Arch (Liu et al., 2025) to the spiking-projection constraint below and trained the ANN and SNN implementations separately from scratch on WT2 and ListOps. Nine SNN runs produced non-finite loss; the remaining 97 architectures have complete paired results on both tasks. This pool is distinct from the 20 experts used to initialize our search. The WT2 ANN–SNN comparison uses test perplexity and tie-adjusted Kendall τb\tau_{b}, giving 0.29970.2997 on the 97 successful pairs.

On these same architecture identifiers, the cross-task rank correlation between WT2 perplexity and ListOps accuracy is 0.05160.0516 for SNNs and 0.30310.3031 for ANNs, with lower perplexity and higher accuracy oriented as better.

A.2 Scope of the Spiking Constraint

The spiking-projection constraint places binary spike inputs before the parameter-dominant feature projections. Membrane potentials, residual streams, recurrent states, and token-mixer arithmetic remain continuous. Lightweight control computations may also be continuous, including Comba’s closed-loop state feedback and LoopMem’s evolved state-norm feedback network. Thus the constraint does not imply that all arithmetic is spike-driven. Candidate blocks use the shared hard-reset LIF implementation in Eq. 1; the surrogate dynamic network (SDN) of SpikingSSMs (Shen et al., 2025) accelerates the neuron computation during search and evaluation. Candidates must also preserve the block interface, causality, and subquadratic sequence complexity.

Appendix B Architectural Changes and the Scope of Published Search Encodings

We compare the dependencies in the discovered blocks with the choices exposed by published search encodings. Table 5 records the available primitives and composition rules, together with the extensions needed to directly represent the relevant computation. This is an encoding-level comparison: the methods address different tasks and use different evaluation protocols. The criterion is whether a specified computational relation is selectable under the documented encoding, rather than whether another network could approximate its input–output function.

Table 5: Scope of documented search encodings. The final column identifies missing choices for directly representing the examined SNN computations. Source locations specify the versions and sections inspected; extending an encoding can change its scope.
Method and source location Encoded architectural choices Relation to the discovered computations
AutoSNN (Na et al., 2022), Sec. 4.1 Five block choices in a fixed backbone: skip, spiking convolution, and spiking residual blocks, with specified kernel sizes. The block menu needs a spike-statistic controller and its target connection to select the activity-dependent update or residual modulation.
SNASNet (Kim et al., 2022), Cell Search Strategy Four-node cells with zero, skip, convolution, and pooling operations on forward and cross-time backward edges. Temporal feedback is already permitted. A spike-statistic reduction and learned multiplicative control are additional operations beyond this edge menu.
MSE-NAS (Pan et al., 2025), Sec. III, Fig. 1 A multiscale genotype selects layer operations, excitatory/inhibitory types, motifs, and global connections under a specified decoder. These choices alter operations and connectivity; the decoder would need to expose activity-conditioned control equations and their attachment points.
EQ-SpikeLM (Zhang et al., 2026), Sec. IV-B.1, Eqs. (14)–(16) Per-layer preserved channel ratios for query–key, value, and feed-forward projections in a pretrained spiking language model. Channel pruning changes widths within the existing computation; it does not introduce a new spike-derived control dependency.
ENAS (efficient NAS), recurrent space (Pham et al., 2018), Secs. 2.1, 3.1 Predecessor and activation choices in a recurrent cell, with a prescribed highway-gating construction. Recurrence and multiplicative gating are present. Spike-statistic inputs and revisions to the controller’s equation or target require extending the template.
DARTS, recurrent space (Liu et al., 2019), Sec. 3.1.2 Operation choices over linear transforms and activations, identity, and zero, within a recurrent-cell template with highway bypasses. Selecting operations does not itself expose the spike-reduction and control-target relation; these must be added to the operations or template.
einspace (Ericsson et al., 2024), Secs. 3.1–3.3, 5 Recursive grammar for sequential, branching, routing, and computation modules, including matrix multiplication, summation, and concatenation. Composition is substantially broader than a fixed cell. The published grammar excludes recurrent computation, so the complete recurrent SNN blocks require a grammar extension.

The computational relations being compared. NeuroGate forms activity statistics from spikes and uses learned transformations to modulate the recurrent update coefficient and output-gate input. HomeoResSSM constructs separate activity-conditioned gains for the state-space and channel-mixer residual branches. Their distinguishing dependency is therefore spike statistic →\rightarrow learned control →\rightarrow a specified update or branch, within an otherwise inherited recurrent core. LoopMem provides a related case: a normalized summary of the recurrent state controls the forget gate, giving the dependency state summary →\rightarrow learned control →\rightarrow gate, which likewise requires a summary input and a controller attached to the gate.

Appendix C Search Implementation

Algorithms 1 and 2 specify the outer and inner loops. Each nondominated ranking uses a fixed reference set for novelty: the evaluated archive for parent selection and the filtered candidate pool for training-batch selection. NDℛ⁡(𝒮)\operatorname{ND}_{\mathcal{R}}(\mathcal{S}) denotes the nondominated subset of 𝒮\mathcal{S} under [y⁡(a),𝒩⁡(a,ℛ)][y(a),\,\mathcal{N}(a;\mathcal{R})]. The new-evaluation budget excludes the initial expert archive. Island survival changes the active parent population; accepted programs remain in the cumulative candidate pool.

Algorithm 1 Outer loop: selective real-training allocation
Input : Initial evaluated archive 𝒟0\mathcal{D}_{0}; rounds TT; pool target npooln_{\mathrm{pool}}; batch limit KK; new-evaluation budget BB
Output : Evaluated archive 𝒟T\mathcal{D}_{T} and ND𝒟T⁡(𝒟T)\operatorname{ND}_{\mathcal{D}_{T}}(\mathcal{D}_{T})
1 for t=1,…,Tt=1,\ldots,T do
     2 Refresh surrogate 𝒮t\mathcal{S}_{t} from 𝒟t−1\mathcal{D}_{t-1}
     3 Form 𝒫t​0⊆𝒟t−1\mathcal{P}_{t0}\subseteq\mathcal{D}_{t-1} of size |𝒟0||\mathcal{D}_{0}| using NSGA-II’s nondominated sorting and crowding-distance truncation on [y⁡(a),𝒩⁡(a,𝒟t−1)][y(a),\,\mathcal{N}(a;\mathcal{D}_{t-1})]
     4 𝒫t←\mathcal{P}_{t}\leftarrow Algorithm 2(𝒮t,𝒫t​0)(\mathcal{S}_{t},\mathcal{P}_{t0}) // surrogate-guided variation
     5 𝒫top←\mathcal{P}_{\mathrm{top}}\leftarrow globally best candidates, the best candidates of each island, and stepping stones that raised the best-so-far y^t\hat{y}_{t}, filled to the shortlist size by y^t\hat{y}_{t}
     6 𝒫dissim←\mathcal{P}_{\mathrm{dissim}}\leftarrow lowest mean similarity to 𝒫top\mathcal{P}_{\mathrm{top}} among remaining candidates
     7 𝒫f←𝒫top∪𝒫dissim\mathcal{P}_{f}\leftarrow\mathcal{P}_{\mathrm{top}}\cup\mathcal{P}_{\mathrm{dissim}}
     8 𝒫f←\mathcal{P}_{f}\leftarrow screen 𝒫f\mathcal{P}_{f} against the evaluated archive for near-duplicates
     9 Set remaining budget bt←B−|𝒟t−1∖𝒟0|b_{t}\leftarrow B-|\mathcal{D}_{t-1}\setminus\mathcal{D}_{0}|
     10 Select batch 𝒬t\mathcal{Q}_{t} of at most min⁡(K,bt)\min(K,b_{t}) candidates by the same rule on [y^t​(a),𝒩⁡(a,𝒫f)][\hat{y}_{t}(a),\,\mathcal{N}(a;\mathcal{P}_{f})]
     11 Train each a∈𝒬ta\in\mathcal{Q}_{t} on the proxy tasks; compute task metrics and y⁡(a)y(a)
     12 𝒟t←𝒟t−1∪𝒬t\mathcal{D}_{t}\leftarrow\mathcal{D}_{t-1}\cup\mathcal{Q}_{t}
     13 Store code, rationale, fingerprint, task metrics, and fitness for each new archive member
14 return 𝒟T\mathcal{D}_{T} and ND𝒟T⁡(𝒟T)\operatorname{ND}_{\mathcal{D}_{T}}(\mathcal{D}_{T})
Algorithm 2 Inner loop: surrogate-guided program evolution
Input : Surrogate 𝒮t\mathcal{S}_{t}; initial population 𝒫t​0\mathcal{P}_{t0}; target size npooln_{\mathrm{pool}}; validity predicate 𝒞\mathcal{C}; MM islands of capacity cislc_{\mathrm{isl}}
Output : Candidate pool 𝒫t\mathcal{P}_{t}
1 Spectrally cluster 𝒫t​0\mathcal{P}_{t0} into MM islands using Sim=1−d\operatorname{Sim}=1-d (one island in the warm-start stage)
2 𝒫t←∅\mathcal{P}_{t}\leftarrow\emptyset
3 while |𝒫t|<npool|\mathcal{P}_{t}|<n_{\mathrm{pool}} do
     4 Sample parents within or across islands; propose program aa and rationale via LLM mutation
     5 if program execution fails or 𝒞⁡(a)=0\mathcal{C}(a)=0 then continue
     6 Extract fingerprint 𝐱⁡(a)\mathbf{x}(a)
     7 if fingerprint nearly matches a parent then continue
     8 Predict task metrics using 𝒮t\mathcal{S}_{t}; compute y^t​(a)\hat{y}_{t}(a)
     9 Assign aa to island ImI_{m} of greatest mean similarity s¯m​(a)\bar{s}_{m}(a)
     10 if registration rejects a near-redundant proposal then continue
     11 Add aa to ImI_{m} and 𝒫t\mathcal{P}_{t}
     12 if |Im|≥2​cisl|I_{m}|\geq 2c_{\mathrm{isl}} then
         13 retain cislc_{\mathrm{isl}} active members by island survival
     14 if reset interval has elapsed then
         15 Rank islands by their best predicted fitness; clear the weaker half
         16 Seed each cleared island with the best member of a sampled retained island
17 return 𝒫t\mathcal{P}_{t}

C.1 Fingerprint and Three-View Comparison

The 21 measurements in Table 6 are extracted from the instantiated model with its fixed outer wrapper. Module statistics use the module inventory; structural measurements use the execution trace and tensor shapes. Initialization-time probes use a fixed ListOps minibatch, sequence shape, initialization procedure, and seed across candidates. The five classical proxies follow the corresponding definitions (Tanaka et al., 2020; Abdelfattah et al., 2021; Lee et al., 2019; Mellor et al., 2021).

For SWSP, collect binary responses in S∈{0,1}n×VS\in\{0,1\}^{n\times V}, with nn probe samples and VV recorded neuron–sequence-position responses. Following the sample-wise orientation of SWAP (Peng et al., 2024b),

SWSP=|{S:,v:v=1,…,V}|.\mathrm{SWSP}=\left|\{S_{:,v}:v=1,\ldots,V\}\right|. (6)

FireRate is the mean spike activity across the recorded spiking layers at initialization. Both are surrogate inputs measured before training.

Table 6: Complete 21-dimensional behavioral fingerprint used in Section 3.3. The first two groups form architecture statistics; the third comprises initialization-time probes. Type indicates extraction cost: C = code analysis (static), F = forward-based measurement, F+B = forward + backward computations.
Group Measurement Type Description
Module statistics Params C Total trainable parameters
SpkLinear C Spiking linear module count
Conv1d C 1-D convolution count
GateRatio C Gating layer fraction
Structural statistics Depth F Shape-changing layer count
Expand F Dim-increase steps
Contract F Dim-decrease steps
BneckRatio F Min-dim / model-dim
MaxExpand F Largest expansion factor
FFNexp F Avg FFN expansion ratio
FFNparam F FFN parameter fraction
FFNmac F FFN MAC fraction
MAC/Param F Compute density
SubgraphR F Largest FX subgraph ratio
Initialization-time probes SynFlow F+B Synaptic flow
GradNorm F+B Gradient ℓ2\ell_{2} norm
SNIP F+B Connection sensitivity
Jacobcov F+B Jacobian covariance
NASWOT F Activation overlap
SWSP F Sample-wise spiking patterns
FireRate F Spiking firing rate

C.2 Three-View Similarity

Code similarity averages full Python CodeBLEU in both reference–candidate directions, with equal weights on its four components (Ren et al., 2020). Rationale similarity uses cosine similarity between 2,048-dimensional text-embedding-v4 embeddings (Alibaba Cloud, 2025). Before fingerprint cosine similarity, each feature is transformed by

x~j=c⁡(2​xj−ℓjuj−ℓj−1),\widetilde{x}_{j}=c\left(2\,\frac{x_{j}-\ell_{j}}{u_{j}-\ell_{j}}-1\right), (7)

which maps [ℓj,uj][\ell_{j},u_{j}] to [−c,c][-c,c]; we set c=0.9c=0.9. The 21 reference pairs were calibrated on the 97 ASI-Arch paired architectures, satisfy uj>ℓju_{j}>\ell_{j}, and remain fixed throughout search. The transform is applied without clipping; cosine similarity is set to zero if either vector has zero norm. The three views have equal weights as in Eq. 3.

C.3 Initial Expert Archive

The initial archive 𝒟0\mathcal{D}_{0} contains 20 linear recurrent and state-space architectures from five families (Table 7), each adapted to the spiking-projection constraint in Appendix A.2. Eighteen are taken from the Flash Linear Attention library (Yang & Zhang, 2024); S4 (Gu et al., 2022) and MetaLA (Chou et al., 2024) extend coverage of state-space models and modern RNNs. The families span data-independent and data-dependent gating, delta-rule updates, and closed-loop state feedback. Designs that augment the recurrent state with an external memory hierarchy, such as the hierarchical memory for Mamba (Wang et al., 2026), are not included. All 20 are trained on both proxy tasks to supply initial parents and labeled surrogate context; the spiking implementations will be released with the code.

Table 7: The 20 architectures in the initial expert archive 𝒟0\mathcal{D}_{0}, each adapted to the spiking-projection constraint.
Model Family Venue Ref. Model Family Venue Ref.
RetNet Foundational arXiv 2023 (Sun et al., 2023) S4 SSM ICLR 2022 (Gu et al., 2022)
LightNet Foundational TMLR 2026 (Qin et al., 2026) Mamba (S6) SSM COLM 2024 (Gu & Dao, 2024)
GLA Modern RNN ICML 2024 (Yang et al., 2024a) Mamba2 (SSD) SSM ICML 2024 (Dao & Gu, 2024)
HGRN Modern RNN NeurIPS 2023 (Qin et al., 2023) DeltaNet Delta Rule NeurIPS 2024 (Yang et al., 2024b)
HGRN2 Modern RNN COLM 2024 (Qin et al., 2024) DeltaFormer Delta Rule arXiv 2025 (Zhong et al., 2025)
RWKV-6 Modern RNN COLM 2024 (Peng et al., 2024a) Gated DeltaNet Delta Rule ICLR 2025 (Yang et al., 2025)
RWKV-7 Modern RNN COLM 2025 (Peng et al., 2025) KDA Delta Rule arXiv 2025 (Kimi Team, 2025)
ABC Modern RNN ACL 2022 (Peng et al., 2022) DeltaProduct Delta Rule NeurIPS 2025 (Siems et al., 2025)
GSA Modern RNN NeurIPS 2024 (Zhang et al., 2024) MesaNet Advanced ICLR 2026 (von Oswald et al., 2026)
MetaLA Modern RNN NeurIPS 2024 (Chou et al., 2024) Comba Advanced NeurIPS 2025 (Hu et al., 2025)

C.4 Search Configuration and Training Cost

Table 8: Main-search configuration.
Setting Value
LLM / maximum generation tokens Gemini-2.5-Flash / 32,768
Outer iterations / accepted candidates per iteration 10 / 1,080
Batch limit KK / evaluation budget BB 16 / 160
Islands / retained capacity 10 / approximately 20 per island
Population trimming threshold Twice the retained capacity
Inter-island parent sampling probability 0.5
Reset interval / fraction of islands reset 3,600 seconds / one-half
Performance / dissimilarity shortlist sizes 30 / 30
Similarity early-exit tolerance 10−610^{-6}

Each iteration selects at most K=16K=16 candidates for training; candidates that duplicate archived architectures are removed, so some batches are smaller and 151 new architectures are trained in total. Each candidate-training job uses one NVIDIA V100 (32 GB). Approximate per-architecture training costs are one GPU-hour on WT2 and 20 GPU-hours on ListOps, giving 151​(1+20)=3,171151(1+20)=3{,}171 GPU-hours, or about 132 GPU-days, excluding initial-expert training, downstream retraining, and LLM inference.

C.5 LLM Selection

Figure 9: Comparison of candidate LLMs as program generators. Each axis is normalized across models, with 1.0 for the best and 0.2 for the worst.

Before the main search, candidate LLMs, including Gemini-2.5-Flash (Comanici et al., 2025), OpenAI o3 (OpenAI, 2025), and Qwen3-Max (Qwen Team, 2025), were compared as program generators on sampling success, performance and novelty of the generated architectures, generation time, and API expense (Fig. 9). No model is best on every axis. Because the search issues more than two thousand generation calls per iteration (1,080 accepted candidates at a 49% acceptance rate; Table 8 and Fig. 7), its time and expense scale with the number of calls. Gemini-2.5-Flash is best in time, expense, and sampling success and is therefore used in all reported searches.

C.6 Pareto-Front Evolution

The outer loop selects candidates by two objectives, composite fitness and novelty (Section 3.4). Fig. 10 shows the nondominated front of all trained architectures in these objectives after each iteration. The front advances along both axes. The selected architectures are marked where they first appear and on the final front: NeuroGate (iteration 9) has the highest fitness; HomeoResSSM (iteration 3) lies at the knee point, the member farthest from the line joining the two extreme solutions after min–max normalization of both objectives; and LoopMem (iteration 3) is the most novel member with nontrivial performance on both tasks, as the two more novel members reach only 18.7% and 17.8% ListOps accuracy.

The composite fitness itself combines two task objectives. Fig. 11 shows the corresponding front in ListOps accuracy and WT2 perplexity. At initialization, SpikingDeltaNet alone forms this front. As the search proceeds, the front extends in both directions, and four search candidates, including the one with the highest fitness, improve on SpikingDeltaNet in both metrics.

Figure 10: Evolution of the nondominated front in composite fitness and novelty during the main search. Each panel adds the architectures trained in one iteration (colored) to those of earlier iterations (faded); stars mark the nondominated architectures. Step 0 contains the initial expert models. Labels A–C mark the three selected architectures in the iteration where each first appears and on the final front.
Figure 11: Evolution of the nondominated front in ListOps accuracy and WT2 perplexity during the main search, displayed as in Fig. 10; the horizontal axis shows log10\log_{10} WT2 perplexity, decreasing to the right.

Appendix D Evaluation Protocols

D.1 Training Tasks and Fixed Model Components

Table 9: Proxy-training configurations used to compute search fitness.
Setting WT2 ListOps
Block layers / width 6 / 256 2 / 128
Attention heads 8 4
FFN expansion 4 2
Dropout 0.2 0.1
Batch size 16 32
Epochs 30 25
Learning rate 6×10−36\times 10^{-3} 10−310^{-3}
Weight decay 0.15 10−410^{-4}
Warmup steps 600 3000
Gradient clipping norm 2.0 2.0
Optimizer AdamW AdamW
Schedule Cosine Cosine

Search-time WT2 fitness uses validation perplexity. WT2 uses GPT-2 byte-level BPE, vocabulary size 50,257, and concatenated token streams chunked into length-512 sequences. ListOps uses a training-derived vocabulary, whitespace tokenization, and an appended end-of-sequence token. Search-time ListOps training uses one-third of the LRA training split, and accuracy is measured on the LRA test split. Downstream ListOps evaluation trains on the full training split and reports accuracy on the same test split. The other five LRA tasks are not used during search, and WT103 perplexity is reported on the test split.

The fixed wrapper embeds tokens, stacks candidate blocks, and applies final RMS normalization. Language modeling uses a vocabulary projection; ListOps uses mean pooling and a ten-class head. Table 9 specifies the proxy-training settings. AdamW uses β=(0.9,0.95)\beta=(0.9,0.95), cosine decay, and linear warmup. For downstream evaluation, we follow the task-specific training configurations released with SpikingSSMs (Shen et al., 2025) in its official repository, which builds on the S4 codebase.

D.2 Fitness Definition

The task transforms are anchored to SpikingDeltaNet, the strongest model on both proxy tasks among the 117 architectures evaluated before the search (the 20 experts and the 97 ASI-Arch paired architectures in Appendix A.1); its ANN counterpart, DeltaNet, is likewise the strongest ANN among them and serves as the ANN reference in Table 1. Let xWT2∘x^{\circ}_{\mathrm{WT2}} and xListOps∘x^{\circ}_{\mathrm{ListOps}} denote its WT2 perplexity and ListOps accuracy in percentage points. The relative improvements are

rWT2​(x)\displaystyle r_{\mathrm{WT2}}(x) =xWT2∘−xxWT2∘​TWT2​(x),\displaystyle=\frac{x^{\circ}_{\mathrm{WT2}}-x}{x^{\circ}_{\mathrm{WT2}}\,T_{\mathrm{WT2}}(x)}, TWT2​(x)\displaystyle T_{\mathrm{WT2}}(x) ={TWT2+,x≤xWT2∘,TWT2−,x>xWT2∘,\displaystyle=\begin{cases}T^{+}_{\mathrm{WT2}},&x\leq x^{\circ}_{\mathrm{WT2}},\\ T^{-}_{\mathrm{WT2}},&x>x^{\circ}_{\mathrm{WT2}},\end{cases} (8)
rListOps​(x)\displaystyle r_{\mathrm{ListOps}}(x) =x−xListOps∘xListOps∘​TListOps​(x),\displaystyle=\frac{x-x^{\circ}_{\mathrm{ListOps}}}{x^{\circ}_{\mathrm{ListOps}}\,T_{\mathrm{ListOps}}(x)}, TListOps​(x)\displaystyle T_{\mathrm{ListOps}}(x) ={TListOps+,x≥xListOps∘,TListOps−,x<xListOps∘,\displaystyle=\begin{cases}T^{+}_{\mathrm{ListOps}},&x\geq x^{\circ}_{\mathrm{ListOps}},\\ T^{-}_{\mathrm{ListOps}},&x<x^{\circ}_{\mathrm{ListOps}},\end{cases} (9)

and, for each task k∈{WT2,ListOps}k\in\{\mathrm{WT2},\mathrm{ListOps}\},

ϕk​(x)=σ⁡(κk​(x)​clip⁡(rk​(x),−1,1)),κk​(x)={κ+,rk​(x)≥0,κ−,rk​(x)<0,\phi_{k}(x)=\sigma\bigl(\kappa_{k}(x)\,\operatorname{clip}(r_{k}(x),-1,1)\bigr),\qquad\kappa_{k}(x)=\begin{cases}\kappa^{+},&r_{k}(x)\geq 0,\\ \kappa^{-},&r_{k}(x)<0,\end{cases} (10)

where σ⁡(z)=(1+e−z)−1\sigma(z)=(1+e^{-z})^{-1}; Eq. 5 averages the two task scores. We set xWT2∘=58.28x^{\circ}_{\mathrm{WT2}}=58.28, xListOps∘=53.0x^{\circ}_{\mathrm{ListOps}}=53.0, TWT2+=TListOps+=0.02T^{+}_{\mathrm{WT2}}=T^{+}_{\mathrm{ListOps}}=0.02, TWT2−=0.25T^{-}_{\mathrm{WT2}}=0.25, TListOps−=0.70T^{-}_{\mathrm{ListOps}}=0.70, κ+=3\kappa^{+}=3, and κ−=ln⁡3\kappa^{-}=\ln 3. Each anchor score maps to 0.50.5, and each task score lies in [σ⁡(−κ−),σ⁡(κ+)]≈[0.25,0.953][\sigma(-\kappa^{-}),\sigma(\kappa^{+})]\approx[0.25,0.953]. Because none of the other 19 experts surpasses the anchor on either task, the narrow improvement tolerance gives large rewards to candidates that do, while the wide, bounded loss side keeps failed candidates from dominating the scale.

WT2-only studies use ϕWT2\phi_{\mathrm{WT2}} alone with the same anchor and TWT2+=0.15T^{+}_{\mathrm{WT2}}=0.15, TWT2−=0.25T^{-}_{\mathrm{WT2}}=0.25, and κ+=κ−=3\kappa^{+}=\kappa^{-}=3.

D.3 Offline Surrogate Studies

These studies were conducted before the main search and guided the choice of its surrogate and fingerprint. Their architectures were drawn at random from the 97 ASI-Arch paired architectures, the 20 experts, and candidates generated with a preliminary surrogate; fewer were trained on ListOps, whose training is slower. The regressor and progressive-input studies use 223 WT2 architectures and 54 ListOps architectures. They share five-fold cross-validation splits repeated three times and report Kendall rank correlation with measured task outcomes. The progressive study starts with five classical probes, adds SWSP and FireRate together, and then adds module and structural statistics.

The labeled-context curves use fixed holdouts of 50 WT2 and ten ListOps architectures. At each context size, 20 subsets are drawn from the remaining pools of 173 and 44 architectures, respectively; the curves show mean correlation and one standard deviation.

D.4 Search-Method Comparisons

Inner-loop comparisons. The EoH and FunSearch adaptations and OpenArchEvo share the outer loop in Algorithm 1: the same expert archive, surrogate refresh, candidate filtering, and NSGA-II batch selection for training. Only the inner loop in Algorithm 2 differs. The EoH adaptation keeps design guidance in LLM variation but evolves one population without similarity-based islands or island resets. The FunSearch adaptation keeps the islands but removes design guidance from variation. The main search needs no warm-start stage, since its expert archive is trained on both tasks and supplies the initial surrogate context and parents. In the WT2-only studies, a warm-start stage precedes the first outer iteration: a single-island inner loop diversifies the expert archive, and a batch of its candidates is trained on WT2 to extend the surrogate context.

Each configuration runs a warm-start stage and six outer iterations of WT2-only search, three times, and uses the WT2-only fitness in Appendix D.2. After each stage, we record the best fitness among trained candidates and the hypervolume of the nondominated set in the novelty–fitness plane. For these comparisons, novelty is recomputed against one fixed reference set shared across methods, runs, and stages. Hypervolume also uses a common reference point, placed just below the minimum novelty and fitness observed across the runs. Curves show the mean over runs, and shaded bands show one sample standard deviation.

Warm-start stage. The OpenArchEvo w/o Warm-Start variant skips the warm-start stage and reallocates its training budget to iteration 1, keeping the total training budget equal; its Warm-Start value in Fig. 12 therefore equals the expert value. Its first selections for training thus rely on a surrogate whose labeled context contains only the expert archive. After six iterations, the variant reaches a best fitness of 0.557 versus 0.679 for OpenArchEvo and a hypervolume of 0.142 versus 0.215, and it remains below OpenArchEvo from the warm-start stage onward (Fig. 12).

Figure 12: Effect of the warm-start stage in six-iteration WT2 searches: (a) best fitness and (b) hypervolume of the nondominated set in fitness and novelty after each stage, computed as in Fig. 6. Curves show means over three runs; shaded regions indicate standard deviation.

ASI-Arch. For ASI-Arch (Liu et al., 2025), we retain the agent workflow and original search hyperparameters where applicable, replace the evaluation tasks, and add the spiking constraint to the prompts. Table 3 reports one run per method.

ASI-Arch draws on a literature-derived prior of about 100 papers and nine cooperating agents, whereas OpenArchEvo starts from 20 expert architectures with one program generator. The resulting SpikingCondFuse combines a delta-rule recurrent path with a local value path containing a five-tap finite-impulse-response filter; a conditioned gate fuses the paths using hidden-state statistics. Its downstream results are in Table 3.

Appendix E Discovered Architecture Details

The discovered implementations retain the recurrent operator families of the expert architectures from which they descend. NeuroGate adds two spiking activity-to-control projections and removes the SiLU activation after its query, key, and value short convolutions; it also adds a mean-spike penalty with weight 10−310^{-3} to the training loss. HomeoResSSM adds a separate spiking gate network to each residual branch, with hidden width dmodel/2d_{\mathrm{model}}/2. LoopMem adds a state-norm feedback to the forget gate of its closed-loop recurrence and penalties on the mean spike activity of its token-mixing and channel-mixing branches.

Spike-activity-dependent controls. NeuroGate forms 𝒔¯t=[s¯q,t,s¯k,t,s¯v,t]\bar{\boldsymbol{s}}_{t}=[\bar{s}_{q,t},\bar{s}_{k,t},\bar{s}_{v,t}] from mean query, key, and value spikes, and modifies the update coefficient and output-gate input as

βt←βt​[1+σ⁡(Pβ​(𝒔¯t))],gt←gt​[1+σ⁡(Pg​(𝒔¯t))],\beta_{t}\leftarrow\beta_{t}\bigl[1+\sigma(P_{\beta}(\bar{\boldsymbol{s}}_{t}))\bigr],\qquad g_{t}\leftarrow g_{t}\bigl[1+\sigma(P_{g}(\bar{\boldsymbol{s}}_{t}))\bigr], (11)

where each projection PP includes its spiking encoder. HomeoResSSM scales each branch output b⁡(x)b(x) before residual addition:

xnext=x+Dropout⁡(f​b​(x)),f=1+2​σ​(P2​(GELU⁡(P1​(s¯)))+b0),x_{\mathrm{next}}=x+\mathrm{Dropout}\bigl(f\,b(x)\bigr),\qquad f=1+2\sigma\!\left(P_{2}\!\left(\mathrm{GELU}(P_{1}(\bar{s}))\right)+b_{0}\right), (12)

where s¯\bar{s} is the branch’s mean spike activity and both projections are spiking. For each chunk, LoopMem offsets the forget-gate logit gg by a function of the normalized recurrent state SS and bounds the state norm after the update:

g←g+2​tanh⁡(F⁡(mean⁡(S/‖S‖F))),S←S/max⁡(‖S‖F/10, 1),g\leftarrow g+2\tanh\!\bigl(F(\operatorname{mean}(S/\|S\|_{F}))\bigr),\qquad S\leftarrow S/\max(\|S\|_{F}/10,\,1), (13)

where FF is a two-layer network of width 8 and norms are taken per head. It also adds the activity penalty γ⁡(s¯attn+s¯FFN)\gamma(\bar{s}_{\mathrm{attn}}+\bar{s}_{\mathrm{FFN}}) to the training loss, with γ=10−2\gamma=10^{-2}.

Appendix F Arithmetic Energy Accounting

Energy is compared for all models under one common architecture setting, independent of the task-specific training configurations: Nb=6N_{b}=6 blocks, model width dmodel=256d_{\mathrm{model}}=256, eight heads, FFN width dff=1024d_{\mathrm{ff}}=1024, sequence length L=512L=512, and vocabulary size V=50,257V=50{,}257. All SNNs share the same wrapper, whose FFN and output head receive spikes; the token embedding lookup is excluded. For spike-driven projection jj, let OjO_{j} denote its dense-equivalent operation count and ρ\rho its measured, operation-weighted mean input firing rate. We estimate

E=EAC​ρ​∑jOj+EMAC​NMAC,E=E_{\mathrm{AC}}\,\rho\sum_{j}O_{j}+E_{\mathrm{MAC}}\,N_{\mathrm{MAC}}, (14)

with EAC=0.9E_{\mathrm{AC}}=0.9 pJ and EMAC=4.6E_{\mathrm{MAC}}=4.6 pJ (Horowitz, 2014); ρ​∑jOj\rho\sum_{j}O_{j} is the number of synaptic operations (SOPs). All designated SpikingLinear projections, including the output head, are counted as accumulate operations. Continuous state updates, readouts, short convolutions, and feedback projections are counted as MACs; a spike-driven input projection does not make the subsequent recurrent computation spike-driven. Continuous scalar multiplications are included as MAC equivalents. Normalization, nonlinear functions, LIF/SDN execution, and memory access are excluded; the FFT arithmetic of PMBC is included for Dyn-SSM as specified below. NeuroGate uses the output-gate-enabled configuration.

Firing rates differ across architectures under the same data, wrapper, and training configuration, so we use each model’s measured, operation-weighted mean input firing rate of its spike-driven projections (Table 10). For Dyn-SSM (Zhong et al., 2026), the block combines its token mixer, an SSM convolution followed by a spike-driven 1×11\times 1 convolution from dmodeld_{\mathrm{model}} to 2​dmodel2d_{\mathrm{model}} channels and a GLU, with the shared spiking FFN. The token mixer’s SSM convolution is counted using FFT arithmetic: with the kernel spectrum cached and nfft=2​Ln_{\mathrm{fft}}=2L, a convolution costs 2​nfft​(log2⁡nfft+1)2n_{\mathrm{fft}}(\log_{2}n_{\mathrm{fft}}+1) MAC equivalents per channel for the transforms and pointwise product. For PMBC, we assume one boundary-compression iteration (M=1M=1). The algorithm computes one membrane-integration convolution and two convolutions per iteration for the upper and lower bounds, giving 1+2​M=31+2M=3 additional convolutions per channel (Zhong et al., 2026). These FFT operations add arithmetic beyond direct LIF updates and are included explicitly; the estimate does not cover the full implementation cost of LIF or SDN. This single-iteration energy scenario is separate from the published settings underlying the quoted task scores. The dense Transformer reference with the same depth and width counts

Nref=Nb​[L⁡(4​dmodel2+2​dmodel​dff)+2​L2​dmodel]+L​dmodel​VN_{\mathrm{ref}}=N_{b}\bigl[L(4d_{\mathrm{model}}^{2}+2d_{\mathrm{model}}d_{\mathrm{ff}})+2L^{2}d_{\mathrm{model}}\bigr]+Ld_{\mathrm{model}}V (15)

MACs, giving 45.1245.12 mJ; the energy reduction is Eref/EE_{\mathrm{ref}}/E.

Table 10: Arithmetic energy under the common architecture setting. ρ\rho is the measured, operation-weighted input firing rate; SOP and MAC counts are in billions per sequence; reductions are computed from unrounded counts.
Architecture ρ\rho (%) SOPs (G) MACs (G) Energy (mJ) Reduction
Dense Transformer – – 9.809 45.12 1.0×\times
Dyn-SSM 9.3 0.800 0.140 1.36 33.1×\times
SpikingDeltaNet 9.8 0.883 0.225 1.83 24.7×\times
SpikingMamba2 12.2 1.247 0.420 3.05 14.8×\times
SpikingCondFuse 8.3 0.782 0.367 2.39 18.9×\times
NeuroGate 4.6 0.424 0.226 1.42 31.7×\times
HomeoResSSM 7.2 0.736 0.422 2.60 17.3×\times
LoopMem 4.9 0.432 0.110 0.89 50.6×\times

For DeltaNet-derived programs, the MAC count includes the chunkwise delta-rule correction, local interactions, and recurrent-state read/write operations. Mamba2-derived programs include local SSD interactions, state propagation, decay scaling, and output gating. LoopMem includes its continuous state-norm feedback; SpikingCondFuse additionally includes the local filter and fusion arithmetic.

Appendix G Generation Prompts

The system and user prompt templates below retain the historical wording used in the experiments. They adapt the open-source ASI-Arch prompts (Liu et al., 2025) with SNN-specific constraints and interface requirements; placeholders in the user template are filled with the selected parent programs and rationales. The user template also asks the generator to consider structures that may reduce the firing rate. The operational search constraints are specified in the main text.

G.1 System Prompt

Instructions
You are an advanced AI architecture designer specializing in evolving neural network architectures through systematic experimentation and analysis, especially in sequence modeling architectures and neuronmorphic computing. Your PRIMARY responsibility is to IMPLEMENT working code modifications that improve model performance and possess efficient spiking neural network nature.
CRITICAL: Code Implementation First
YOU MUST IMPLEMENT YOUR DESIGN. A motivation without code implementation is useless. Your job is to:
1. First understand the current architecture, especially the usage and insertion of spiking neuron 2. Design and implement concrete code changes 3. Only then provide the motivation explaining your implementation Core Objectives 1. READ existing code 2. IMPLEMENT architectural modifications 3. Generate excellent spiking neural network for sequence modeling tasks 4. Ensure all changes maintain sub-quadratic complexity (avoiding O⁡(N2)O(N^{2}) softmax attention) 5. Ensure the inputs of ALL linear layers are binary signals producted by spike neurons (replacing all nn.Linear() by SpikingLinear()) 6. Write working, runnable code that integrates seamlessly with existing infrastructure 7. Provide clear motivation that explains the implemented changes Implementation Requirements • Dependency Consistency: Do not remove any of the existing import statements • Preserve SpikingLinear Class: The class SpikingLinear(nn.Module) must be kept exactly as it is, without modification or deletion • Spiking Requirements: Use spike neuron from src.models.spike.neuron before linear layers (uniformly use SpikingLinear()) • Correct Spike Output Format: A final_spikes tensor MUST be returned, not the firing rate • Complete Layer: Implement the full layer class including __init__ and forward methods • Preserve Signatures: Do NOT change forward() input/output signatures • Default Parameters: New features must have sensible defaults and be enabled by default • Compatible Input/Output Interface: Keep all the original input and output (especially final_spikes in output) • No Config Changes: Since config doesn’t evolve, use ALL the default parameters in __init__ • Keep Class Name: Always keep class name as SpikingFLABlock • Disable Decorators: Do not use @torch.compile decorators and bf.float16 for robustly training Technical Constraints 1. Complexity: Must be sub-quadratic (linear or O⁡(n​log⁡n)O(n\log n) acceptable) 2. Chunkwise Processing: Use chunk-based computation for efficiency 3. Mask Correctness: Ensure causal masking prevents future information leakage 4. Batch Size Independence: CRITICAL - Your code must work with ANY batch size • Never hardcode batch dimensions • Use dynamic shapes from input tensors • Avoid operations that assume specific batch/sequence dimensions • Ensure all tensor operations are batch-agnostic 5. Parameter Preservation: Keep core parameters like d_model, num_heads unchanged 6. Kwargs Support: Always include **kwargs in __init__ for compatibility Design Philosophy • Working Code Over Ideas: An implemented solution beats a theoretical one • Bold Changes: Make significant architectural modifications, not just tweaks • Evidence-Based: Ground modifications in experimental results and research • Simplification: When adding features, consider removing outdated ones • Theoretical Grounding: Every change needs solid theoretical justification Implementation Process 1. Read Current Code: Review and understand the existing implementation 2. Analyze Results: Identify specific weaknesses from training/test metrics 3. Architecture migration: Inherit the current insights, associate and migrate to Spiking Neural Network 4. Design Solution: Create a theoretically-grounded architectural change 5. Implement Code: Write the complete layer implementation 6. Document Motivation: Explain what you implemented and why Code Quality Standards • Clean, readable code with appropriate comments • Efficient tensor operations using PyTorch best practices • Proper initialization of new parameters • Correct gradient flow through all operations • Memory-efficient implementations • Batch-size agnostic operations Output Requirements • name: Model identifier starting with “spiking_fla_” • motivation: Clear explanation of WHAT you implemented and WHY

G.2 User Prompt

Neural Architecture Evolution Mission EXPERIMENTAL CONTEXT & HISTORICAL EVIDENCE
// Versioned historical context logic:
[Version 1]
{version_1_content}

We find that the below version outperforms [Version i].
[Version i+1]
{version_i+1_content}

// Per version content (depends on motivation flag):
### Motivation
{idea_i}
### Program
{markdown_code_block_i}
ARCHITECTURE EVOLUTION OBJECTIVE
Your mission is to create a breakthrough neural architecture that addresses critical performance limitations identified through experimental evidence while integrating cutting-edge research insights from sequence modeling architectures and neuronmorphic computing. Design and implement an innovative architecture that maintains computational efficiency and spiking neural network natures while achieving superior cognitive capabilities.
SYSTEMATIC EVOLUTION METHODOLOGY PHASE 1: Evidence-Based Analysis Framework 1.1 Architecture Forensics
Current State Assessment:
• Examine existing architectural implementations • Map computational mechanisms, design patterns, and information flow • Identify core algorithmic approaches and their theoretical foundations • Document interface constraints and compatibility requirements 1.2 Performance Pattern Recognition
Historical Evidence Analysis:
• Modeling Capability: Extract optimization challenges from results of loss (fitting) and accuracy/ppl (generalization) • Cognitive Potential: Identify capability gaps across two different tasks (ListOps for long sequence modeling, WikiText2 for natural language processing) • Bottleneck Identification: Pinpoint architectural elements limiting performance vs. those enabling strengths • Cross-Architecture Comparison: Analyze performance patterns across different experimental variants 1.3 Research Integration Strategy
Theoretical Foundation Building:
• Map research insights to observed performance limitations • Identify specific theoretical principles addressing architectural weaknesses • Synthesize multiple research findings for comprehensive enhancement opportunities • Validate theoretical applicability through experimental evidence correlation PHASE 2: Innovation Design Framework 2.1 Targeted Performance Engineering
Gap-Specific Solutions:
• Design architectural modifications targeting the most critical performance bottlenecks • Create mechanisms leveraging research insights for problematic capability domains • Balance multiple improvement objectives while maintaining architectural coherence • Ensure modifications address root causes rather than symptoms 2.2 Theoretical Grounding Protocol
Research-Driven Design:
• Ground all modifications in validated theoretical principles • Ensure mathematical and computational justification for proposed changes • Verify alignment with established research findings and best practices • Create novel combinations of insights for breakthrough potential 2.3 Efficiency Optimization Standards
Computational Constraints:
• Design using chunked computation patterns for scalability • Maintain sub-quadratic O⁡(N​log⁡N)O(N\log N) even O⁡(N)O(N) complexity throughout • Optimize memory usage through efficient processing strategies • Preserve performance gains within strict complexity bounds 2.4 Neuromorphic Computing Natures
Brain-like characteristics:
• Motivate the suitability of spiking neurons and ANN components on sequence modeling tasks • Consider the possible structure and connnection that might be beneficial for reducing the firing rate • Encourage SNN operations like SpikingLinear, discourage too much or complicated ANN inefficient operations like multiplication of matrices • Focus on the design of locality and statefulness of computation (state-based mechanisms like RNN/SSMs/Mamba) • Encourage a clean structure featuring high cohesion and low coupling, combined with the exclusive use of plausible operators PHASE 3: Implementation Excellence Protocol 3.1 Architecture Implementation Standards
Code Development Requirements:
• Implement the complete evolved architecture • Preserve interface compatibility (forward function signatures, __init__ **kwargs) • Add new parameters with sensible defaults (enabled by default for new features) • Remove or refactor existing features to prevent architectural bloat • Implement proper causal masking and information flow constraints 3.2 Quality Assurance Framework
Technical Excellence Standards:
• Disable @torch.compile decorators and bf.float16 for robust training • Preserve chunked processing patterns throughout the architecture • Ensure causal constraints prevent any information leakage • Verify sub-quadratic complexity in all implemented operations 3.3 Documentation and Justification
Innovation Communication:
• Create comprehensive motivation explaining evolution rationale • Connect experimental evidence to theoretical insights and implementation decisions • Justify expected improvements based on research findings • Provide clear reasoning for all architectural design choices TECHNICAL IMPLEMENTATION SPECIFICATIONS Spiking Neural Network Requirements • Dependency Consistency: Do not remove any of the existing import statements • Preserve SpikingLinear Class: The class SpikingLinear(nn.Module) must be kept exactly as it is, without modification or deletion. But the class SpikingMLP(nn.Module) can be removed if unnecessary • Spiking Linear Layer Requirements: Use spike neuron from src.models.spike.neuron before linear layers (uniformly use SpikingLinear()). Do not use other implementations of SNN like SpikingJelly • Correct Spike Output Format: A final_spikes tensor of all spiking neurons MUST be returned, not the firing rate. This tensor MUST be a flat, concatenated torch tensor of binary spikes, created for example by torch.cat([s.flatten() for s in all_spikes_list]). It must NOT be a firing rate, a list of tensors, or any other format. Critical Preservation Requirements • Class Structure: Maintain SpikingFLABlock class name and inheritance hierarchy • Interface Stability: Preserve exact forward function signature compatibility • Parameter Compatibility: Support **kwargs in __init__ for extensibility • Dimensional Consistency: Maintain d_model and core parameter structure Implementation Quality Standards • Chunked Processing: All sequence operations must utilize fixed-size chunking • Causal Integrity: Implement strict causal constraints in attention-like mechanisms • Complexity Bounds: Ensure O⁡(N​log⁡N)O(N\log N) or better for all operations • Memory Efficiency: Design for optimal memory usage with chunked patterns • Compilation Safety: Avoid @torch.compile on utility functions to prevent conflicts MANDATORY: Tensor Operations Robustness • einops.rearrange() Requirement: Replace ALL .view()/.reshape() with einops.rearrange() • Dynamic Dimension Handling: Never manually calculate dimensions - use einops inference • Batch Size Agnostic: All operations must work with ANY batch size • Runtime Shape Extraction: Get dimensions from tensor.shape at runtime, not config • Adaptive Processing: Design for actual tensor dimensions, not predetermined values Cross-Environment Robustness Standards • Universal Compatibility: Identical performance across training/evaluation/inference • Memory Adaptation: Graceful handling of varying memory constraints • Shape Tolerance: Robust operation with varying input dimensions • Resource Awareness: Automatic adaptation to available computational resources INNOVATION TARGET DOMAINS Primary Capability Enhancement Areas • Extended Context Memory: Revolutionary long-range dependency handling • Multi-Scale Information Integration: Enhanced temporal and semantic scale processing • Adaptive Computational Mechanisms: Dynamic adjustment based on input characteristics • Efficiency-Performance Optimization: Superior capabilities within complexity constraints • Cognitive Task Performance: Breakthrough improvements in reasoning and comprehension • Environmental Robustness: Consistent performance across execution contexts • Resource Efficiency: Optimal adaptation to computational constraints DELIVERABLE SPECIFICATIONS PRIMARY DELIVERABLE: Complete Implementation
Architecture Code (MANDATORY):
• Implementation: Create complete working architecture • Innovation Quality: Embed revolutionary architectural advances in functional code • Constraint Compliance: Preserve class structure, parameters, and interface compatibility • Technical Standards: Maintain sub-quadratic complexity, chunked processing, causal constraints • Robustness Implementation: Use einops.rearrange() universally, ensure batch size independence SECONDARY DELIVERABLE: Design Documentation
Architecture Description:
• Naming Convention: SpikingFLABlock_[innovation_identifier] reflecting core innovations • Motivation Document: Comprehensive explanation including: – Key architectural innovations and their implementation – Research insights applied and expected performance improvements – Design choice justification based on experimental evidence – Connection between theory, evidence, and implementation SUCCESS CRITERIA FRAMEWORK Critical Success Factors (Ranked by Priority) 1. Implementation Excellence: Successfully create breakthrough architecture 2. Constraint Adherence: Maintain class name, parameters, and interface compatibility 3. Technical Robustness: Ensure complexity bounds, chunked processing, causal constraints 4. Universal Compatibility: Use einops.rearrange() universally, support any batch size 5. Evidence-Based Innovation: Embed research insights addressing identified limitations 6. Performance Targeting: Implement solutions for specific weakness areas identified MISSION EMPHASIS
Your PRIMARY OBJECTIVE is implementing breakthrough architectural code that demonstrates robust performance across all execution environments and batch configurations. Create working innovations that directly address identified performance gaps through research-guided architectural evolution. Documentation serves as secondary validation of implemented innovations.
Begin your evolution process by examining the experimental evidence and identifying the most critical architectural improvement opportunities.

References

  • Abdelfattah et al. (2021) Mohamed S. Abdelfattah, Abhinav Mehrotra, Łukasz Dudziak, and Nicholas D. Lane. Zero-cost proxies for lightweight NAS. In International Conference on Learning Representations (ICLR), 2021.
  • Alibaba Cloud (2025) Alibaba Cloud. Embedding: text-embedding-v4. Alibaba Cloud Model Studio documentation, 2025.
  • Chou et al. (2024) Yuhong Chou, Man Yao, Kexin Wang, Yuqi Pan, Ruijie Zhu, Yiran Zhong, Yu Qiao, Jibin Wu, Bo Xu, and Guoqi Li. MetaLA: Unified optimal linear approximation to softmax attention map. In Advances in Neural Information Processing Systems (NeurIPS), 2024.
  • Comanici et al. (2025) Gheorghe Comanici, Eric Bieber, Mike Schaekermann, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025.
  • Dao & Gu (2024) Tri Dao and Albert Gu. Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality. In International Conference on Machine Learning (ICML), 2024.
  • Ericsson et al. (2024) Linus Ericsson, Miguel Espinosa, Chenhongyi Yang, Antreas Antoniou, Amos Storkey, Shay B. Cohen, Steven McDonagh, and Elliot J. Crowley. einspace: Searching for neural architectures from fundamental operations. In Advances in Neural Information Processing Systems (NeurIPS), 2024.
  • Gu & Dao (2024) Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. In Conference on Language Modeling (COLM), 2024.
  • Gu et al. (2022) Albert Gu, Karan Goel, and Christopher Ré. Efficiently modeling long sequences with structured state spaces. In International Conference on Learning Representations (ICLR), 2022.
  • Horowitz (2014) Mark Horowitz. 1.1 Computing’s energy problem (and what we can do about it). In IEEE International Solid-State Circuits Conference (ISSCC), 2014.
  • Hu et al. (2025) Jiaxi Hu, Yongqi Pan, Jusen Du, Disen Lan, Xiaqiang Tang, Qingsong Wen, Yuxuan Liang, and Weigao Sun. Improving bilinear RNNs with closed-loop control. In Advances in Neural Information Processing Systems (NeurIPS), 2025.
  • Kim et al. (2022) Youngeun Kim, Yuhang Li, Hyoungseob Park, Yeshwanth Venkatesha, and Priyadarshini Panda. Neural architecture search for spiking neural networks. In European Conference on Computer Vision (ECCV), 2022.
  • Kimi Team (2025) Kimi Team. Kimi Linear: An expressive, efficient attention architecture. arXiv preprint arXiv:2510.26692, 2025.
  • Lee et al. (2019) Namhoon Lee, Thalaiyasingam Ajanthan, and Philip H. S. Torr. SNIP: Single-shot network pruning based on connection sensitivity. In International Conference on Learning Representations (ICLR), 2019.
  • Liu et al. (2019) Hanxiao Liu, Karen Simonyan, and Yiming Yang. DARTS: Differentiable architecture search. In International Conference on Learning Representations (ICLR), 2019.
  • Liu et al. (2025) Yixiu Liu, Yang Nan, Weixian Xu, Xiangkun Hu, Lyumanshan Ye, Zhen Qin, and Pengfei Liu. AlphaGo moment for model architecture discovery. arXiv preprint arXiv:2507.18074, 2025.
  • Mellor et al. (2021) Joseph Mellor, Jack Turner, Amos Storkey, and Elliot J. Crowley. Neural architecture search without training. In International Conference on Machine Learning (ICML), 2021.
  • Na et al. (2022) Byunggook Na, Jisoo Mok, Seongsik Park, Dongjin Lee, Hyeokjun Choe, and Sungroh Yoon. AutoSNN: Towards energy-efficient spiking neural networks. In International Conference on Machine Learning (ICML), 2022.
  • OpenAI (2025) OpenAI. OpenAI o3 and o4-mini system card. System card, 2025.
  • Pan et al. (2025) Wenxuan Pan, Feifei Zhao, Guobin Shen, Bing Han, and Yi Zeng. Brain-inspired multiscale evolutionary neural architecture search for deep spiking neural networks. IEEE Transactions on Evolutionary Computation, 29(5):2258–2270, 2025.
  • Peng et al. (2024a) Bo Peng, Daniel Goldstein, Quentin Anthony, et al. Eagle and Finch: RWKV with matrix-valued states and dynamic recurrence. In Conference on Language Modeling (COLM), 2024a.
  • Peng et al. (2025) Bo Peng, Ruichong Zhang, Daniel Goldstein, Eric Alcaide, Xingjian Du, Haowen Hou, Jiaju Lin, Jiaxing Liu, Janna Lu, William Merrill, et al. RWKV-7 “Goose” with expressive dynamic state evolution. In Conference on Language Modeling (COLM), 2025.
  • Peng et al. (2022) Hao Peng, Jungo Kasai, Nikolaos Pappas, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz, and Noah A. Smith. ABC: Attention with bounded-memory control. In Annual Meeting of the Association for Computational Linguistics (ACL), 2022.
  • Peng et al. (2024b) Yameng Peng, Andy Song, Haytham M. Fayek, Vic Ciesielski, and Xiaojun Chang. SWAP-NAS: Sample-wise activation patterns for ultra-fast NAS. In International Conference on Learning Representations (ICLR), 2024b.
  • Pham et al. (2018) Hieu Pham, Melody Y. Guan, Barret Zoph, Quoc V. Le, and Jeff Dean. Efficient neural architecture search via parameters sharing. In International Conference on Machine Learning (ICML), 2018.
  • Qin et al. (2023) Zhen Qin, Songlin Yang, and Yiran Zhong. Hierarchically gated recurrent neural network for sequence modeling. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
  • Qin et al. (2024) Zhen Qin, Songlin Yang, Weixuan Sun, Xuyang Shen, Dong Li, Weigao Sun, and Yiran Zhong. HGRN2: Gated linear RNNs with state expansion. In Conference on Language Modeling (COLM), 2024.
  • Qin et al. (2026) Zhen Qin, Yuxin Mao, Xuyang Shen, Dong Li, Jing Zhang, Yuchao Dai, and Yiran Zhong. You only scan once: Efficient multi-dimension sequential modeling with LightNet. Transactions on Machine Learning Research, 2026. ISSN 2835-8856.
  • Qwen Team (2025) Qwen Team. Qwen3-Max: Just scale it. Model release blog, September 2025.
  • Ren et al. (2020) Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. CodeBLEU: A method for automatic evaluation of code synthesis. arXiv preprint arXiv:2009.10297, 2020.
  • Shen et al. (2025) Shuaijie Shen, Chao Wang, Renzhuo Huang, Yan Zhong, Qinghai Guo, Zhichao Lu, Jianguo Zhang, and Luziwei Leng. SpikingSSMs: Learning long sequences with sparse and parallel spiking state space models. In AAAI Conference on Artificial Intelligence (AAAI), 2025.
  • Siems et al. (2025) Julien Siems, Timur Carstensen, Arber Zela, Frank Hutter, Massimiliano Pontil, and Riccardo Grazzi. DeltaProduct: Improving state-tracking in linear RNNs via Householder products. In Advances in Neural Information Processing Systems (NeurIPS), 2025.
  • Sun et al. (2023) Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. Retentive network: A successor to transformer for large language models. arXiv preprint arXiv:2307.08621, 2023.
  • Tanaka et al. (2020) Hidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, and Surya Ganguli. Pruning neural networks without any data by iteratively conserving synaptic flow. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  • von Oswald et al. (2026) Johannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari, Songlin Yang, Maximilian Schlegel, Kaitlin Maile, Yanick Schimpf, Oliver Sieberling, Alexander Meulemans, Guillaume Lajoie, Rif A. Saurous, Charlotte Frenkel, Razvan Pascanu, Blaise Aguera y Arcas, and Joao Sacramento. MesaNet: Sequence modeling by locally optimal test-time training. In International Conference on Learning Representations (ICLR), 2026.
  • Wang et al. (2026) Qinwen Wang, Jieping Luo, Aoxiang Qin, Ruoyu Zhao, Jianxiong Tang, Wei Zhang, Zhichao Lu, and Luziwei Leng. Mamba with hierarchical memory: Solving representation bottleneck in long sequence modeling. arXiv preprint arXiv:2608.02347, 2026.
  • Yang & Zhang (2024) Songlin Yang and Yu Zhang. FLA: A Triton-based library for hardware-efficient implementations of linear attention mechanism. Software library, 2024.
  • Yang et al. (2024a) Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim. Gated linear attention transformers with hardware-efficient training. In International Conference on Machine Learning (ICML), 2024a.
  • Yang et al. (2024b) Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim. Parallelizing linear transformers with the delta rule over sequence length. In Advances in Neural Information Processing Systems (NeurIPS), 2024b.
  • Yang et al. (2025) Songlin Yang, Jan Kautz, and Ali Hatamizadeh. Gated delta networks: Improving Mamba2 with delta rule. In International Conference on Learning Representations (ICLR), 2025.
  • Zhang et al. (2026) Malu Zhang, Wenjie Wei, Zijian Zhou, Wanlong Liu, Jie Zhang, Ammar Belatreche, and Yang Yang. Spike-driven lightweight large language model with evolutionary computation. IEEE Transactions on Evolutionary Computation, 30(4):1333–1346, 2026.
  • Zhang et al. (2024) Yu Zhang, Songlin Yang, Ruijie Zhu, Yue Zhang, Leyang Cui, Yiqiao Wang, Bolun Wang, Freda Shi, Bailin Wang, Wei Bi, Peng Zhou, and Guohong Fu. Gated slot attention for efficient linear-time sequence modeling. In Advances in Neural Information Processing Systems (NeurIPS), 2024.
  • Zhong et al. (2025) Shu Zhong, Mingyu Xu, Tenglong Ao, and Guang Shi. Understanding transformer from the perspective of associative memory. arXiv preprint arXiv:2505.19488, 2025.
  • Zhong et al. (2026) Yan Zhong, Ruoyu Zhao, Chao Wang, Jiaqi He, Qinghai Guo, Jianguo Zhang, Zhichao Lu, and Luziwei Leng. Dyn-SSM: Towards the efficient long sequence learning via bio-interpretable dynamics in spiking state space models. IEEE Transactions on Cognitive and Developmental Systems, pp. 1–15, 2026. Early access.