Large Language Model-Guided Evolutionary Discovery of Native Neural Architectures for Spiking Sequence Modeling
Abstract
Spiking neural networks (SNNs) offer low-energy sequence modeling through sparse, event-driven computation. However, interactions among spike encoding, neuronal dynamics, and information propagation complicate architecture design. Existing SNN sequence models often adapt artificial neural network (ANN) architectures designed for real-valued activations, potentially underusing spike-based communication and temporal state updates, motivating automated discovery of native SNN architectures. Most evolutionary neural architecture search (ENAS) methods operate within predefined configuration spaces, limiting discovery to mechanisms expressible within those spaces. We introduce OpenArchEvo, which uses large language models (LLMs) to evolve executable architecture code in an open program space under spiking-projection constraints. In this space, code differences need not reflect architectural novelty, while direct performance evaluation requires costly training. We construct a three-view representation spanning code, design rationale, and a behavioral fingerprint to support novelty estimation and performance prediction. The search treats predicted performance and estimated novelty as two objectives, using surrogate predictions to select candidates for expensive training evaluations. With an estimated candidate-training cost of 132 V100 GPU-days, the search uncovers multiple native SNN architectures, exemplified by three designs featuring mechanisms such as spike-activity-dependent control of state updates and residual pathways. The discovered NeuroGate surpasses the ANN DeltaNet on WikiText-103, and the discovered architectures reduce estimated architecture-level arithmetic energy by up to (LoopMem) relative to a common dense Transformer (ANN) baseline. All code and all discovered architectures will be made publicly available soon.
1 Introduction
Spiking neural networks (SNNs) communicate through sparse, event-driven spikes, offering a promising route to low-energy sequence modeling. Binary spike signals interact with continuous states that retain and integrate information over time. Architecture determines how spike activity updates these states and how stored information is propagated, shaping predictive performance and opportunities for sparse computation. Existing SNN sequence models include both adaptations of artificial neural network (ANN) architectures and mechanisms designed specifically for spiking computation (Zhu et al., 2024; Zhong et al., 2026; Shen et al., 2025). However, architectural choices effective in ANNs need not retain their advantages after adaptation to spiking computation. In our WikiText-2 comparison, ANN and SNN implementations of the 97 ASI-Arch paired architectures agree only partially in rank (Fig. 1(b)), motivating direct evaluation under spiking computation. The interactions among spike activity, state updates, and information pathways make useful combinations difficult to anticipate. This motivates an AutoML research question: how can search build on expert designs to discover native SNN architectures that match or exceed ANN performance while preserving substantial energy advantages?
Neural architecture search (NAS), particularly evolutionary NAS (ENAS), offers a route to this automated exploration (Liu et al., 2023; Li et al., 2024). Most ENAS methods vary operators, connections, and hyperparameters within predefined configuration spaces. These spaces permit new combinations of known components, but mechanisms outside their allowed operations and composition rules require redesigning the search representation. Large language models (LLMs) can generate and modify executable architecture code, enabling changes to internal computations and module interactions. Existing LLM-based NAS already supports architecture discovery through code evolution, including quality-diversity search and novelty checks (Chen et al., 2023; Nasir et al., 2024; Liu et al., 2025). Greater freedom to generate candidates, however, does not by itself yield effective architectural discovery. Figure 1(a) situates this problem within LLM-based algorithm design: code differences need not reflect meaningful architectural differences, while directly measuring task performance requires costly candidate training. The challenge is to characterize generated SNN architectures for novelty estimation and performance prediction, guiding training toward promising candidates while retaining distinct architectural alternatives.
To address this challenge, we introduce OpenArchEvo, an LLM-guided evolutionary method for native SNN architecture discovery. It searches over an open code implementation space for SNN blocks within a fixed outer model structure, subject to interface, causality, and spiking-projection constraints. We estimate each candidate’s architectural novelty relative to other architectures considered during search through three complementary views: code describes computational structure, the design rationale states the intended design, and a behavioral fingerprint combines architecture statistics with network responses measured before weight training. The fingerprint also supports surrogate performance prediction; predicted performance and estimated novelty form two objectives for selecting candidates for training. One discovered architecture, NeuroGate, features spike-activity-dependent modulation of recurrent updates and output gating. On WikiText-103, the discovered SNNs outperform the listed SNN baselines, and NeuroGate (Fig. 1(c)) attains 26.4 WikiText-103 perplexity versus 27.5 for ANN DeltaNet (lower is better). The discovered architectures reduce estimated architecture-level arithmetic energy by up to (LoopMem) relative to a common dense Transformer (ANN) baseline (Table 1).
Our main contributions are as follows:
- •
Native SNN architectures. We discover SNN architectures with evolved elements native to spiking computation, including spike-activity-dependent control in NeuroGate and HomeoResSSM and a state-norm feedback in LoopMem that benefits the spiking implementation, and examine these elements through SNN ablations and a non-spiking comparison.
- •
Three-view architecture representation. We construct a three-view representation using code, design rationale, and numerical behavioral fingerprints of architectural structure and behavior. We compare candidates using these views to estimate each candidate’s architectural novelty, while reusing the fingerprints for performance prediction.
- •
LLM-guided evolutionary discovery. We integrate this characterization into LLM-guided evolution over an open code implementation space, using surrogate-predicted performance and estimated novelty as two objectives to select architectures for training within a limited budget. Search diagnostics and ablations examine surrogate guidance and diversity maintenance.
2 Related Work
We review SNN architecture design, evolutionary neural architecture search, and representations for candidate comparison and performance prediction.
2.1 Spiking Sequence Models
Spiking neurons combine binary outputs with continuous membrane states that retain temporal information. Architectural choices determine how these signals are transformed and propagated. In vision models, Spikformer uses spike-based attention without softmax (Zhou et al., 2023), while Spike-driven Transformer places residual connections before spiking activations to preserve binary communication (Yao et al., 2023). These designs show how adapting information pathways to spiking computation can support performance with less computation.
For sequence modeling, SpikeGPT combines recurrent processing with binary spiking activations for language generation (Zhu et al., 2024). The SpikingSSMs architecture applies leaky integrate-and-fire (LIF) dynamics to state-space outputs, combining temporal memory with sparse synaptic computation (Shen et al., 2025). Dyn-SSM incorporates refractory LIF neurons with soft reset, preserving residual membrane potential after firing (Zhong et al., 2026).
Automated methods also demonstrate the value of SNN-specific architectural exploration. SNASNet searches forward and temporal feedback connections, while MSE-NAS explores neuron operations and multiscale connectivity (Kim et al., 2022; Pan et al., 2025). AutoSNN considers accuracy and spike count and reports improvements over its handcrafted baselines (Na et al., 2022). For spiking language models, EQ-SpikeLM combines evolutionary channel pruning with subsequent post-training quantization (Zhang et al., 2026). Further couplings among spike activity, sequence-state updates, and residual pathways offer opportunities to extend this accumulated design knowledge.
2.2 Evolutionary Neural Architecture Search
Evolutionary neural architecture search (ENAS) uses population-based variation and selection to optimize neural architectures (Liu et al., 2023; Li et al., 2024). Most methods vary operators, connections, widths, and depths within predefined representations. For example, EvoCNN evolves variable-length sequences of convolutional, pooling, and fully connected layers (Sun et al., 2020), while NSGA-Net evolves CNN blocks through predefined operation and connection choices (Lu et al., 2021). These encodings support new combinations of existing building blocks, but a mechanism outside their permitted operations and composition rules requires redesigning the search representation.
Genetic programming and grammar-based NAS support variable computational structures assembled from reusable primitives. CGP-CNN evolves convolutional architectures from layer-level components, while einspace uses typed primitives to define a broader compositional space (Suganuma et al., 2017; Ericsson et al., 2024). AutoML-Zero further explores model computation and learning rules constructed from basic mathematical operations (Real et al., 2020). High-level primitives supply architectural priors through predefined computational forms; finer primitives expose more computational choices but can make useful designs harder to find. Even when composition obeys type and interface constraints, finding useful architectures can require many expensive evaluations. This motivates using architectural knowledge to guide the generation and modification of candidate architecture code.
LLM-guided evolutionary search has become a general approach to automated algorithm design (Wu et al., 2025; Ma et al., 2026), evolving heuristics, metaheuristics, and scientific programs (Romera-Paredes et al., 2024; Liu et al., 2024; van Stein & Bäck, 2025; Novikov et al., 2025); ShinkaEvolve, for example, combines code-embedding similarity screening with an LLM novelty judge to reject redundant candidate programs (Lange et al., 2026). For neural architectures, LLMs use pretrained knowledge and natural-language design goals, constraints, and principles to guide configuration search or generate architecture code. LLMENAS adapts fitness functions within a predefined cell search space (Lai et al., 2026), while Design Principle Transfer uses natural-language principles to narrow subsequent searches (Zhou et al., 2025). EvoPrompting evolves complete classifiers and, for graph networks, selected computations within a fixed processor (Chen et al., 2023). LLMatic combines network-code evolution with quality-diversity search for image classification (Nasir et al., 2024). ASI-Arch evolves attention-layer implementations under interface and computation constraints; it checks novelty through rationale retrieval and LLM judgments before training, and includes LLM-assessed architectural quality in its fitness (Liu et al., 2025); Genesys discovers language-model architectures with LLM agents on a genetic-programming backbone (Cheng et al., 2025). These code-based approaches establish architecture discovery with mechanisms for diversity and novelty. Our work combines population-relative novelty and surrogate-predicted performance to select SNN architecture programs for training; Section 2.3 details the representations supporting these decisions.
2.3 Architecture Representation for Evolutionary Search
A search encoding specifies how architectures can be constructed; candidate descriptors summarize the resulting architectures for comparison and prediction. Implementation-level comparisons include the lexical, syntactic, and data-flow matches measured by CodeBLEU (Ren et al., 2020). Natural-language descriptions expose design intent, as in Evolution of Heuristics (EoH), which jointly evolves heuristic ideas and executable implementations (Liu et al., 2024). Execution-based descriptions capture observed behavior: phenotypic characterization records program responses to probe cases and supports surrogate modeling (Hildebrandt & Branke, 2015), while BehaveSim compares intermediate solution trajectories (Zhang & Lu, 2026). These views provide complementary, partial evidence: similar code need not implement the same computation, stated design intent need not match the implementation, and finite probes reveal only part of a program’s behavior.
For neural architectures, structural statistics and initialization-time forward and backward probes provide information before candidate training. Training-free proxies estimate trained performance from initialization-time activations or gradients without optimizing candidate weights (Abdelfattah et al., 2021). DCL-ENAS pretrains an architecture encoder without performance labels, then contrastively fine-tunes a predictor on evaluated architectures to predict their relative performance (Zhang et al., 2026b). In expressive grammar-based NAS, Transferrable Surrogates builds transferable predictors from training-free proxies and neural graph features, or from a fine-tuned language model, and uses them to filter candidates or serve directly as search objectives (Qin et al., 2025). Accurate performance prediction alone does not validate distances between architecture representations as measures of novelty.
Descriptors also determine which differences evolutionary search rewards. Novelty search rewards distance from previously observed behaviors, while MAP-Elites retains high-performing solutions within descriptor-defined regions (Lehman & Stanley, 2011; Mouret & Clune, 2015). In LLMatic, discussed in Section 2.2, the network archive uses width-to-depth ratio and floating-point operations (FLOPs) as descriptors (Nasir et al., 2024). BOP-Elites models both quality and descriptors with surrogates to select evaluations when these quantities are expensive to obtain (Kent et al., 2025). For SNNs, architectures with similar size and computational cost can nevertheless differ in how spike activity interacts with state updates and gating. We estimate novelty relative to a reference population by comparing code, design rationale, and a behavioral fingerprint combining architecture statistics with initialization-time probes. The same fingerprint supports performance prediction, and both estimates guide the selection of candidates for expensive training evaluations.
3 Method
OpenArchEvo discovers SNN architecture programs by coupling program-level variation with three-view architecture representation and surrogate-assisted evolution (Fig. 2). Code, design rationale, and a behavioral fingerprint provide complementary descriptions for comparing generated architectures. The fingerprint also supports near-duplicate screening and performance prediction, allowing an inner evolutionary loop to explore many candidates before an outer loop allocates expensive training evaluations. Performance–novelty selection encourages the retention of distinct architectural directions, and measured training outcomes guide subsequent search.
3.1 Preliminaries on SNNs
An SNN processes information through interconnected spiking neurons, whose internal states evolve over time and whose outputs are discrete binary events (Neftci et al., 2019). In a discrete-time description, a neuron’s binary output records whether it fires at each step, while its membrane potential remains continuous-valued. A common example is the leaky integrate-and-fire (LIF) neuron. In its hard-reset form, the pre-reset potential , spike , and membrane state at step satisfy
| (1) |
where , , and denote input current, leakage factor, and firing threshold, and denotes the indicator function. The neuron integrates the current with its decayed membrane state, emits a spike when the resulting potential reaches or exceeds the threshold, and resets its state to zero after firing.
Within the network, incoming spikes contribute to a neuron’s input current through weighted synaptic connections. The resulting spike response depends on both the incoming signals and the neuron’s preceding membrane state. Binary synaptic inputs also permit weighted sums to be evaluated by accumulating the weights associated with active spikes, providing opportunities for sparse computation (Davies et al., 2018; Horowitz, 2014).
3.2 Problem Formulation
Let denote the search space of admissible SNN block architectures, each represented by executable code and feeding binary spikes to its parameter-dominant feature projections (the spiking-projection constraint; Appendix A.2). Architecture performance is task-dependent, so our main search combines WikiText-2 language modeling and ListOps hierarchical expression evaluation (Merity et al., 2017; Nangia & Bowman, 2018) to assess each design beyond a single task. Their weak rank agreement in our cross-task comparison supports their use as complementary evaluation signals (Appendix A.1). For , let collect the separate trainable parameter vectors of the full models instantiated with for the evaluation tasks, and let sum their training losses. We aggregate their evaluation scores into a single performance objective , with larger values indicating better performance. Let be a finite set of reference architectures, held fixed for each ranking, and let denote the estimated architectural novelty of relative to . To encourage distinct architectural alternatives alongside performance, we model discovery as a multiobjective optimization problem following NAS formulations (Lu et al., 2021; Lu et al., 2024) and novelty-based multiobjectivization (Mouret & Doncieux, 2012), mathematically as follows:
| (2) | ||||
In practice, fixed, finite training protocols approximate the lower-level optimization. Measuring a candidate’s fitness in the main search requires separate training on both datasets, increasing evaluation cost relative to training the same candidate on either task alone. Novelty can instead be estimated before weight training. To limit evaluation cost, each search run evaluates at most additional architectures under its prescribed task and training protocol, excluding those evaluated initially.
3.3 Three-View SNN Architecture Representation
To guide the discovery of useful native SNN architectures, we need to estimate architectural differences among generated candidates before committing to weight training. Source-code differences alone need not reflect changes in architectural mechanisms. We therefore propose a three-view representation for SNN architectures (Fig. 3): code describes implemented operations and dataflow; the design rationale states the intended architectural idea, following EoH’s idea–code pairing (Liu et al., 2024); and a numerical behavioral fingerprint records structural attributes and initialization-time responses. Together, these views provide partial evidence for comparing candidates generated in the code space.
Inspired by response-based program characterization (Hildebrandt & Branke, 2015; Zhang & Lu, 2026), we probe the instantiated SNN before weight training. Probe inputs, sequence shape, initialization procedure, and seed are fixed across candidate comparisons. We adapt SWAP’s sample-wise pattern count (Peng et al., 2024b) to binary spike outputs, yielding sample-wise spiking patterns (SWSP): each recorded neuron–sequence-position response forms a pattern across probe samples, and SWSP counts the distinct patterns. Alongside SWSP, we use mean spike activity across spiking layers (FireRate) and five established activation- and gradient-based proxies (Abdelfattah et al., 2021; Mellor et al., 2021), giving seven initialization-time measurements. We supplement them with four module statistics describing components and ten structural statistics describing dimension flow, computational organization, and resource allocation. We construct the behavioral fingerprint as a 21-dimensional feature vector by concatenating these measurements. Feature definitions are given in Appendix C.1.
For architectures and , we adopt CodeBLEU (Ren et al., 2020) and average its two comparison directions to obtain code similarity . For rationale similarity , we encode each LLM-generated design rationale with the pretrained text-embedding model text-embedding-v4 (Alibaba Cloud, 2025) and compute cosine similarity between the resulting vectors. Fingerprint similarity uses cosine similarity after fixed feature-wise affine scaling of and (Appendix C.2). We define pairwise dissimilarity using fixed, untuned equal weights:
| (3) | ||||
Inspired by novelty search (Lehman & Stanley, 2011), we estimate novelty as the mean dissimilarity to all other reference architectures:
| (4) |
The surrogate in Section 3.4 uses evaluated architectures’ fingerprints and measured task outcomes to predict task performance from an untrained candidate’s .
3.4 Surrogate-Assisted Evolutionary Search
To limit costly weight training, we use an inner loop for surrogate-guided architecture evolution and an outer loop for training allocation and archive updates, i.e., surrogate model management (Jin, 2011; Zhang et al., 2010). Let denote the initial archive of evaluated expert architectures, and the archive after outer iteration . Each archived architecture has an associated fingerprint and measured task outcomes. We adopt TabPFN-2.5 (Hollmann et al., 2025; Grinsztajn et al., 2025) as the surrogate, using these records from as labeled context. Given , the surrogate predicts task metrics without training candidate . The same fitness mapping used for measured outcomes aggregates these predictions into ; hereafter, denotes fitness measured after the prescribed training protocol.
We initialize the inner-loop populations from using NSGA-II’s nondominated sorting and crowding-distance truncation (Deb et al., 2002), based on measured fitness and novelty . The selected architectures seed separate island populations, from which parents are sampled within or across islands. The LLM revises or recombines their code and design rationales. Feasible offspring pass fingerprint-based near-duplicate screening before surrogate evaluation. Predicted fitness guides subsequent parent selection and survival within islands, while accepted candidates accumulate in a pool for possible training.
Before selecting architectures for training, we shortlist candidates to avoid comparing every pair of generated architectures. We retain a subset prioritized by predicted performance and supplement it with candidates having high mean three-view dissimilarity to that subset. After screening against the evaluated archive for near-duplicates, we apply the same selection procedure using predicted fitness and novelty relative to the filtered pool, which remains fixed throughout selection. The resulting batch contains previously unevaluated architectures, limited by the maximum batch size and remaining budget. The selected architectures are trained under the prescribed task protocols and added to with their measured outcomes, supplying additional surrogate context and parent candidates for the next iteration. Across the prescribed iterations, at most additional architectures are evaluated under these protocols. Search returns the evaluated archive and its nondominated subset under measured fitness and novelty relative to the final archive. Appendix C gives the complete procedures and budget constraints.
4 Experiments
| Model | WT103 | LRA accuracy (%) | Energy reduction | ||||||
|---|---|---|---|---|---|---|---|---|---|
| PPL | ListOps | Text | Retrieval | Image | Pathfinder | Path-X | Avg. | ||
| ANN baselines | |||||||||
| DeltaNet (Yang et al., 2024b) | 27.5 | 62.2 | 85.1 | 91.3 | 89.8 | 94.2 | 94.9 | 86.2 | – |
| S4D-Lin† (Gu et al., 2022) | – | 60.5 | 87.0 | 91.0 | 87.9 | 94.0 | 92.8∗ | 85.5 | – |
| SNN baselines | |||||||||
| S6-based SNN† | – | 55.7 | 77.6 | 88.5 | 80.1 | 83.4 | – | – | – |
| Dyn-SSM† (Zhong et al., 2026) | 33.2 | 60.2 | 82.4 | 88.8 | 87.2 | 92.0 | 94.4 | 84.2 | 33.1 |
| SpikingDeltaNet | 34.5 | 57.3 | 82.2 | 90.5 | 89.0 | 91.2 | 92.3 | 83.7 | 24.7 |
| SpikingMamba2 | 38.2 | 48.5 | 74.0 | 81.0 | 83.2 | 90.4 | 89.6 | 77.8 | 14.8 |
| Discovered SNNs (OpenArchEvo) | |||||||||
| NeuroGate | 26.4 | 61.8 | 82.4 | 90.1 | 91.2 | 93.6 | 95.3 | 85.7 | 31.7 |
| HomeoResSSM | 28.1 | 46.2 | 84.8 | 92.3 | 87.6 | 86.3 | 85.2 | 80.4 | 17.3 |
| LoopMem | 27.8 | 59.5 | 81.0 | 89.5 | 90.5 | 88.4 | 88.7 | 82.9 | 50.6 |
| †Task results quoted from Shen et al. (2025); Zhong et al. (2026). ∗S4D-Inv result. | |||||||||
We evaluate the discovered SNNs, examine their architectural mechanisms, and analyze how representation and search design contribute to the discovery.
4.1 Experimental Setup
Candidate blocks share the sequence-model wrapper described in Appendix D.1. During the main search, each selected candidate is trained separately for 30 epochs on WikiText-2 (WT2) language modeling and 25 epochs on one-third of the ListOps training set (Merity et al., 2017; Tay et al., 2021), producing distinct model weights for the two tasks. We instantiate the composite fitness as
| (5) |
where the task transforms use SpikingDeltaNet as the fitness anchor (Appendix D.2).
The initial archive contains 20 expert SNN adaptations of linear recurrent and state-space architectures (Appendix C.3). Gemini-2.5-Flash (Comanici et al., 2025), selected after a comparison of candidate LLMs (Appendix C.5), is used as the LLM to generate candidates, with prompts in Appendix G, and the surrogate specified in Section 3.4 predicts their task metrics. The main search runs for 10 outer iterations with a batch limit of and a budget of architecture evaluations, each comprising both task-training runs; candidates that duplicate archived architectures are removed, so 151 new architectures are trained. Candidate weight training costs an estimated 132 V100 GPU-days (3,171 GPU-hours). Appendix C.4 gives the search configuration and resource accounting. The code, generated programs, and evaluated archive will be released.
To evaluate the selected architectures at full scale, we retrain them on WikiText-103 (WT103) (Merity et al., 2017) and the six Long Range Arena (LRA) tasks, whose sequences span 1K–16K tokens (Tay et al., 2021), following the task-specific configurations of SpikingSSMs (Shen et al., 2025); the WT103 configuration matches that of Dyn-SSM (Zhong et al., 2026). Unlike the search-time setting, ListOps is trained on the full training split and evaluated on the same test split, whereas WT103 and the other five LRA tasks are not used during search (Appendix D.1). The LRA average covers all six tasks. The ANN reference, DeltaNet, is among the strongest of the 20 experts and 97 ASI-Arch models trained under our common protocol. S4D-Lin and S6-based SNN results are quoted from Shen et al. (2025) and Dyn-SSM results from Zhong et al. (2026); all other models, including our spiking adaptations SpikingDeltaNet and SpikingMamba2, are trained under our protocol. The main search comprises one run; Sections 4.4 and 4.5 report repeated smaller-scale studies.
4.2 Main Results
Figure 4 traces the main search. Candidates that surpass the strongest expert appear in every iteration from the second onward, and the best fitness improves in five separate iterations, reaching at iteration 9. Progress is therefore sustained rather than confined to an early proposal, consistent with the archive supplying better parents and surrogate context as it grows. The nondominated front also advances in the two task metrics, eventually containing candidates that surpass SpikingDeltaNet on both WT2 and ListOps (Appendix C.6).
Using search-time measurements only, we select three architectures that represent distinct regions of the final Pareto front in composite fitness and novelty (Fig. 10): NeuroGate, the extreme solution with the highest fitness; HomeoResSSM, the knee point, which lies farthest from the line joining the two extreme solutions after normalization (Das, 1999; Zhang et al., 2015); and LoopMem, the most novel member that retains nontrivial performance on both tasks.
The selected architectures remain strong after full-scale retraining (Table 1). On WT103, all three outperform the listed SNN baselines, and NeuroGate also surpasses the ANN DeltaNet trained under our protocol (26.4 versus 27.5 perplexity), although its parameter-dominant projections receive binary spikes. On LRA, NeuroGate attains the highest average accuracy among the SNNs, between the two ANN baselines, and remains the highest among the SNNs without ListOps (90.5 versus 89.0), while HomeoResSSM gives the best Retrieval accuracy among all listed models. The discovered architectures thus combine competitive accuracy with distinct task profiles.
These results come with substantially lower estimated energy. Counting spike-driven projections as accumulate operations and continuous computation as multiply–accumulate operations, with each model’s measured firing rate and a shared model wrapper, the discovered architectures reduce the arithmetic energy of a dense Transformer of the same size by – (Appendix F, which also details the Dyn-SSM estimate). The spiking feed-forward network (FFN) and output head of the shared wrapper account for most spike-driven operations, so differences among the SNNs mainly reflect their firing rates and token-mixer arithmetic. NeuroGate fires about half as often as SpikingDeltaNet, so its additional control computation still yields a larger reduction, and LoopMem attains the largest reduction among all listed models. We next examine which evolved elements distinguish the discovered architectures and whether their effects depend on spiking computation.
4.3 Spike-Native Mechanisms in the Discovered Architectures
The discovered architectures mainly contain two kinds of elements that we regard as native to SNNs. ① Spike-activity-dependent computation takes spike activity itself as an input. NeuroGate maps the mean input spike activity of its query, key, and value projections at each position to gains on the delta-rule update coefficient and the output gate (Fig. 1(c)), and HomeoResSSM uses the spike activity of its state-space and channel-mixing branches to scale their residual contributions (Fig. 5(a)); we test this class on HomeoResSSM. Because each neuron integrates its input across tokens and resets on firing, this activity reflects membrane dynamics that have no counterpart in ANN units, so the control reads a signal that exists only in spiking units. ② Spiking-specific design choices do not read spikes explicitly, but depart from ANN practice in either direction: components that are uncommon in ANN design are added when they benefit the SNN, and components that are standard in ANNs are removed when they degrade it. Direct ANN-to-SNN adaptation introduces neither kind of change. LoopMem illustrates the first direction, feeding a normalized summary of its recurrent state back to the forget gate (state-norm feedback) (Fig. 5(b)); NeuroGate illustrates the second, omitting the sigmoid linear unit (SiLU) activation that DeltaNet applies after its query, key, and value convolutions.
| Variant | WT2 PPL | |
| ① Spike-activity-dependent computation | ||
| HomeoResSSM | ||
| Discovered | 57.5 | – |
| Control driven by continuous inputs | 59.2 | |
| Control removed | 64.4 | |
| ② Spiking-specific design choices | ||
| NeuroGate | ||
| Discovered (SiLU removed) | 57.2 | – |
| SiLU restored | 63.1 | |
| LoopMem | ||
| Discovered | 60.5 | – |
| State-norm feedback removed | 69.3 | |
| Non-spiking, with feedback | 61.3 | – |
| Non-spiking, feedback removed | 59.7 | |
Table 2 examines these elements on WT2. All variants, including the discovered architectures, are trained with the fixed protocol applied to every search candidate, and we compare the effect of each element within an implementation rather than perplexities across implementations. For ①, removing HomeoResSSM’s control raises perplexity by 6.9. Driving the same controller with the channel mean of the continuous inputs, from which the spikes are generated, recovers most of this benefit, yet the spike-driven control still attains 1.7 lower perplexity. For ②, restoring SiLU raises NeuroGate’s perplexity by 5.9. Removing LoopMem’s state-norm feedback raises perplexity by 8.8 in the SNN, whereas the same removal lowers the perplexity of the non-spiking variant by 1.6. LoopMem’s feedback thus has opposite effects in the two implementations, and the ANN-standard SiLU degrades the SNN.
From the search-space perspective, each control element combines several decisions: which signal to summarize, how to transform it, and which update, branch, or gate it modulates. In OpenArchEvo, the LLM can propose these decisions within one code revision of a parent program. Among seven published encodings, covering SNN cell and block spaces, recurrent-cell spaces, and a recursive architecture grammar, none can select a spike-statistic or state-summary input, a learned controller, and its target together; directly representing these elements requires extending the encoding (Appendix B, Table 5). Open-code search can express such relations, but the best architecture of our adapted ASI-Arch run, SpikingCondFuse, conditions its gate on hidden-state statistics rather than spike activity (Appendix D.4). Appendix E gives the equations of the evolved elements.
4.4 Comparison with Existing Search Methods
| Adapted ASI-Arch | OpenArchEvo | |
| Best fitness | 0.653 | 0.872 |
| Top-10 mean fitness | 0.570 | 0.714 |
| WT103 PPL | 29.6 | 26.4 |
| ListOps accuracy | 55.3 | 61.8 |
| Text accuracy | 82.3 | 82.4 |
| Retrieval accuracy | 85.3 | 90.1 |
| Energy reduction‡ | 18.9 | 31.7 |
| LLM agents | 9 | 1 |
| Prompt tokens† | 25K | 1.2K |
| API expenditure (USD) | 440 | 170 |
One search run per method; training budgets differ as stated in the text. †Prompt tokens per executable program. ‡Computed as in Table 1.
Inner-loop comparisons. We compare our inner-loop search with adaptations of EoH (Liu et al., 2024) and FunSearch (Romera-Paredes et al., 2024), which replace the inner loop while sharing our outer evaluation and surrogate-selection loop. Their original designs evaluate every generated program, which is impractical when each evaluation requires training an architecture. The EoH adaptation uses thought–code evolution in a single population, whereas the FunSearch adaptation uses code-based best-shot prompting with islands. Our inner loop features population management and redundancy control guided by the three-view representation and novelty. In the main search, the trained expert archive supplies the initial surrogate context and parents. Each WT2-only search instead begins with a warm-start stage that trains a first batch of diversified candidates. Each configuration then runs six outer iterations of WT2-only search, scored by a WT2-only fitness of the same form (Appendices D.2 and D.4). OpenArchEvo attains the highest final best fitness and hypervolume (Fig. 6). All configurations use the same total training budget.
End-to-end program-search comparison. We retain the agent workflow of ASI-Arch (Liu et al., 2025), adapt its evaluation tasks, and impose the spiking constraints. Table 3 reports the resulting search outcomes and representative architectures. The ASI-Arch and OpenArchEvo runs train 39 and 151 architectures over six and four weeks, respectively. OpenArchEvo uses one program generator; the API expenditure excludes candidate-training compute. The best and Top-10 mean fitness are higher in our completed run. Appendix D.4 provides the adaptation details.
4.5 Search Analysis and Ablation Studies
Near-duplicate screening. Fingerprint-based near-duplicate screening removes 30% of raw proposals (Fig. 7). The shared initialization-time probes and architecture statistics identify these candidates before surrogate-guided selection.
Surrogate prediction. The surrogate studies were conducted before the main search and guided the choice of its surrogate and fingerprint (Appendix D.3). Initialization-time probes capture network responses and form the starting point for our surrogate-input study (Fig. 8(b)). Adding SWSP and firing rate to the classical proxies improves ranking on WT2, with a smaller change on ListOps. Supplementing these probes with module and structural statistics gives the strongest ranking along the tested sequence. With the complete fingerprint, TabPFN-2.5 ranks best among the tested regressors on both tasks (Fig. 8(a)). In these studies, prediction generally improves as the labeled context grows (Fig. 8(c)). Early in search, candidates are selected for training by a surrogate with little labeled context. In WT2-only searches without the warm-start stage, which start from the expert archive alone, the best fitness after six iterations is 0.56 instead of 0.68, and the gap persists across iterations (Appendix D.4).
| Selection method | Top-1 PPL | Top-5 PPL | Novelty | Variance |
|---|---|---|---|---|
| Evolution + surrogate | 56.2 | 57.4 | 0.534 | 0.0281 |
| Single island + surrogate | 57.3 | 57.9 | 0.513 | 0.0272 |
| Sampling + surrogate | 57.3 | 57.8 | 0.520 | 0.0314 |
| Sampling only | 57.5 | 58.5 | 0.511 | 0.0414 |
Search ablations. We compare candidate-generation and selection variants in six-iteration WT2-only searches, each repeated three times. Besides the full configuration, which evolves candidates on multiple islands and selects them with the surrogate, the variants evolve a single island, sample candidates from the LLM without evolution before surrogate selection, or sample and select without the surrogate. The full configuration has the lowest Top-1 and Top-5 PPL and the highest novelty, while the single-island variant has the lowest fitness variance (Table 4).
5 Discussion
Program evolution discovers changes in control and pathway design around established recurrent operators, including how spike activity regulates computation. It shifts part of architecture design from enumerating admissible mechanisms to specifying executable constraints and evaluating proposed programs. The resulting archive contains both usable architectures and design hypotheses for further study.
For expensive evolutionary search, the framework separates abundant program variation from limited training evaluations. A shared characterization supports redundancy filtering, population organization, and surrogate prediction, allowing these components to improve together. The evidence comprises one full search with smaller repeated studies; the ASI-Arch comparison uses unequal budgets. Broader validation across search runs and domains remains necessary, as fingerprints approximate redundancy and novelty depends on the reference population.
The discovered SNNs demonstrate that program evolution can identify competitive spiking sequence architectures with spike-activity-dependent computation and spiking-specific design choices. Their energy reductions are theoretical arithmetic estimates; realizing and measuring these reductions requires implementations on specific hardware that account for memory access and execution costs. The present search fixes the neuron model. Extending it to jointly evolve neuron dynamics and architecture, and evaluating broader and more demanding downstream tasks, are natural next steps. These results show that program search can discover competitive SNN architectures, which motivates extending it to those harder settings.
6 Conclusion
We introduced OpenArchEvo for automated discovery of native neural architectures, with spiking sequence modeling as a demanding test case. The discovered architectures combine competitive predictive performance with substantial estimated arithmetic-energy savings, demonstrating the value of exploring computations tailored to spike activity. Beyond the resulting models, the study shows how executable architectural proposals can become a cumulative record of empirically evaluated designs. The broader opportunity is to connect expressive program generation with meaningful architectural comparison and selective evaluation, enabling discovery in neural computing domains where useful mechanisms are difficult to specify in advance.
Acknowledgments
Large language models were used to polish the language of this manuscript; the authors take full responsibility for its content.
References
- Abdelfattah et al. (2021) Mohamed S. Abdelfattah, Abhinav Mehrotra, Łukasz Dudziak, and Nicholas D. Lane. Zero-cost proxies for lightweight NAS. In International Conference on Learning Representations (ICLR), 2021.
- Alibaba Cloud (2025) Alibaba Cloud. Embedding: text-embedding-v4. Alibaba Cloud Model Studio documentation, 2025.
- Chen et al. (2023) Angelica Chen, David Dohan, and David So. EvoPrompting: Language models for code-level neural architecture search. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
- Cheng et al. (2025) Junyan Cheng, Peter Clark, and Kyle Richardson. Language modeling by language models. In Advances in Neural Information Processing Systems (NeurIPS), 2025.
- Comanici et al. (2025) Gheorghe Comanici, Eric Bieber, Mike Schaekermann, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025.
- Das (1999) Indraneel Das. On characterizing the “knee” of the Pareto curve based on normal-boundary intersection. Structural Optimization, 18(2–3):107–115, 1999.
- Davies et al. (2018) Mike Davies, Narayan Srinivasa, Tsung-Han Lin, Gautham Chinya, Yongqiang Cao, Sri Harsha Choday, Georgios Dimou, Prasad Joshi, Nabil Imam, Shweta Jain, et al. Loihi: A neuromorphic manycore processor with on-chip learning. IEEE Micro, 38(1):82–99, 2018.
- Deb et al. (2002) Kalyanmoy Deb, Amrit Pratap, Sameer Agarwal, and T. Meyarivan. A fast and elitist multiobjective genetic algorithm: NSGA-II. IEEE Transactions on Evolutionary Computation, 6(2):182–197, 2002.
- Ericsson et al. (2024) Linus Ericsson, Miguel Espinosa, Chenhongyi Yang, Antreas Antoniou, Amos Storkey, Shay B. Cohen, Steven McDonagh, and Elliot J. Crowley. einspace: Searching for neural architectures from fundamental operations. In Advances in Neural Information Processing Systems (NeurIPS), 2024.
- Grinsztajn et al. (2025) Léo Grinsztajn, Klemens Flöge, Oscar Key, Felix Birkel, Philipp Jund, Brendan Roof, Benjamin Jäger, Dominik Safaric, Simone Alessi, Adrian Hayler, et al. TabPFN-2.5: Advancing the state of the art in tabular foundation models. arXiv preprint arXiv:2511.08667, 2025.
- Gu et al. (2022) Albert Gu, Karan Goel, Ankit Gupta, and Christopher Ré. On the parameterization and initialization of diagonal state space models. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
- Hildebrandt & Branke (2015) Torsten Hildebrandt and Jürgen Branke. On using surrogates with genetic programming. Evolutionary Computation, 23(3):343–367, 2015.
- Hollmann et al. (2025) Noah Hollmann, Samuel Müller, Lennart Purucker, Arjun Krishnakumar, Max Körfer, Shi Bin Hoo, Robin Tibor Schirrmeister, and Frank Hutter. Accurate predictions on small data with a tabular foundation model. Nature, 637(8045):319–326, 2025.
- Horowitz (2014) Mark Horowitz. 1.1 Computing’s energy problem (and what we can do about it). In IEEE International Solid-State Circuits Conference (ISSCC), 2014.
- Jin (2011) Yaochu Jin. Surrogate-assisted evolutionary computation: Recent advances and future challenges. Swarm and Evolutionary Computation, 1(2):61–70, 2011.
- Kent et al. (2025) Paul Kent, Adam Gaier, Jean-Baptiste Mouret, and Juergen Branke. Bayesian optimization for quality diversity search with coupled descriptor functions. IEEE Transactions on Evolutionary Computation, 29(2):302–316, 2025.
- Kim et al. (2022) Youngeun Kim, Yuhang Li, Hyoungseob Park, Yeshwanth Venkatesha, and Priyadarshini Panda. Neural architecture search for spiking neural networks. In European Conference on Computer Vision (ECCV), 2022.
- Lai et al. (2026) Yutao Lai, Zicheng Cai, Lei Chen, Tongtao Ling, and Hai-Lin Liu. LLMENAS: Evolutionary neural architecture search via large language model guidance. IEEE Transactions on Evolutionary Computation, 30(4):1362–1376, 2026.
- Lange et al. (2026) Robert Tjarko Lange, Yuki Imajuku, and Edoardo Cetin. ShinkaEvolve: Towards open-ended and sample-efficient program evolution. In International Conference on Learning Representations (ICLR), 2026.
- Lehman & Stanley (2011) Joel Lehman and Kenneth O. Stanley. Abandoning objectives: Evolution through the search for novelty alone. Evolutionary Computation, 19(2):189–223, 2011.
- Li et al. (2024) Nan Li, Lianbo Ma, Guo Yu, Bing Xue, Mengjie Zhang, and Yaochu Jin. Survey on evolutionary deep learning: Principles, algorithms, applications, and open issues. ACM Computing Surveys, 56(2):41:1–41:34, 2024.
- Liu et al. (2024) Fei Liu, Xialiang Tong, Mingxuan Yuan, Xi Lin, Fu Luo, Zhenkun Wang, Zhichao Lu, and Qingfu Zhang. Evolution of heuristics: Towards efficient automatic algorithm design using large language model. In International Conference on Machine Learning (ICML), 2024.
- Liu et al. (2025) Yixiu Liu, Yang Nan, Weixian Xu, Xiangkun Hu, Lyumanshan Ye, Zhen Qin, and Pengfei Liu. AlphaGo moment for model architecture discovery. arXiv preprint arXiv:2507.18074, 2025.
- Liu et al. (2023) Yuqiao Liu, Yanan Sun, Bing Xue, Mengjie Zhang, Gary G. Yen, and Kay Chen Tan. A survey on evolutionary neural architecture search. IEEE Transactions on Neural Networks and Learning Systems, 34(2):550–570, 2023.
- Lu et al. (2021) Zhichao Lu, Ian Whalen, Yashesh Dhebar, Kalyanmoy Deb, Erik D. Goodman, Wolfgang Banzhaf, and Vishnu Naresh Boddeti. Multiobjective evolutionary design of deep convolutional neural networks for image classification. IEEE Transactions on Evolutionary Computation, 25(2):277–291, 2021.
- Lu et al. (2024) Zhichao Lu, Ran Cheng, Yaochu Jin, Kay Chen Tan, and Kalyanmoy Deb. Neural architecture search as multiobjective optimization benchmarks: Problem formulation and performance assessment. IEEE Transactions on Evolutionary Computation, 28(2):323–337, 2024.
- Ma et al. (2026) Zeyuan Ma, Hongshu Guo, Yue-Jiao Gong, Jun Zhang, and Kay Chen Tan. Toward automated algorithm design: A survey and practical guide to meta-black-box-optimization. IEEE Transactions on Evolutionary Computation, 30(2):667–687, 2026.
- Mellor et al. (2021) Joseph Mellor, Jack Turner, Amos Storkey, and Elliot J. Crowley. Neural architecture search without training. In International Conference on Machine Learning (ICML), 2021.
- Merity et al. (2017) Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. In International Conference on Learning Representations (ICLR), 2017.
- Mouret & Clune (2015) Jean-Baptiste Mouret and Jeff Clune. Illuminating search spaces by mapping elites. arXiv preprint arXiv:1504.04909, 2015.
- Mouret & Doncieux (2012) Jean-Baptiste Mouret and Stéphane Doncieux. Encouraging behavioral diversity in evolutionary robotics: An empirical study. Evolutionary Computation, 20(1):91–133, 2012.
- Na et al. (2022) Byunggook Na, Jisoo Mok, Seongsik Park, Dongjin Lee, Hyeokjun Choe, and Sungroh Yoon. AutoSNN: Towards energy-efficient spiking neural networks. In International Conference on Machine Learning (ICML), 2022.
- Nangia & Bowman (2018) Nikita Nangia and Samuel R. Bowman. ListOps: A diagnostic dataset for latent tree learning. In Conference of the North American Chapter of the Association for Computational Linguistics: Student Research Workshop (NAACL SRW), 2018.
- Nasir et al. (2024) Muhammad Umair Nasir, Sam Earle, Julian Togelius, Steven James, and Christopher W. Cleghorn. LLMatic: Neural architecture search via large language models and quality diversity optimization. In Genetic and Evolutionary Computation Conference (GECCO), 2024.
- Neftci et al. (2019) Emre O. Neftci, Hesham Mostafa, and Friedemann Zenke. Surrogate gradient learning in spiking neural networks: Bringing the power of gradient-based optimization to spiking neural networks. IEEE Signal Processing Magazine, 36(6):51–63, 2019.
- Novikov et al. (2025) Alexander Novikov, Ngân Vũ, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco J. R. Ruiz, Abbas Mehrabian, et al. AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv preprint arXiv:2506.13131, 2025.
- Pan et al. (2025) Wenxuan Pan, Feifei Zhao, Guobin Shen, Bing Han, and Yi Zeng. Brain-inspired multiscale evolutionary neural architecture search for deep spiking neural networks. IEEE Transactions on Evolutionary Computation, 29(5):2258–2270, 2025.
- Peng et al. (2024) Yameng Peng, Andy Song, Haytham M. Fayek, Vic Ciesielski, and Xiaojun Chang. SWAP-NAS: Sample-wise activation patterns for ultra-fast NAS. In International Conference on Learning Representations (ICLR), 2024.
- Qin et al. (2025) Shiwen Qin, Gabriela Kadlecová, Martin Pilát, Shay B. Cohen, Roman Neruda, Elliot J. Crowley, Jovita Lukasik, and Linus Ericsson. Transferrable surrogates in expressive neural architecture search spaces. In International Conference on Automated Machine Learning (AutoML), 2025.
- Real et al. (2020) Esteban Real, Chen Liang, David R. So, and Quoc V. Le. AutoML-Zero: Evolving machine learning algorithms from scratch. In International Conference on Machine Learning (ICML), 2020.
- Ren et al. (2020) Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. CodeBLEU: A method for automatic evaluation of code synthesis. arXiv preprint arXiv:2009.10297, 2020.
- Romera-Paredes et al. (2024) Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M. Pawan Kumar, Emilien Dupont, Francisco J. R. Ruiz, Jordan S. Ellenberg, Pengming Wang, Omar Fawzi, et al. Mathematical discoveries from program search with large language models. Nature, 625(7995):468–475, 2024.
- Shen et al. (2025) Shuaijie Shen, Chao Wang, Renzhuo Huang, Yan Zhong, Qinghai Guo, Zhichao Lu, Jianguo Zhang, and Luziwei Leng. SpikingSSMs: Learning long sequences with sparse and parallel spiking state space models. In AAAI Conference on Artificial Intelligence (AAAI), 2025.
- Suganuma et al. (2017) Masanori Suganuma, Shinichi Shirakawa, and Tomoharu Nagao. A genetic programming approach to designing convolutional neural network architectures. In Genetic and Evolutionary Computation Conference (GECCO), 2017.
- Sun et al. (2020) Yanan Sun, Bing Xue, Mengjie Zhang, and Gary G. Yen. Evolving deep convolutional neural networks for image classification. IEEE Transactions on Evolutionary Computation, 24(2):394–407, 2020.
- Tay et al. (2021) Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. Long range arena: A benchmark for efficient transformers. In International Conference on Learning Representations (ICLR), 2021.
- van Stein & Bäck (2025) Niki van Stein and Thomas Bäck. LLaMEA: A large language model evolutionary algorithm for automatically generating metaheuristics. IEEE Transactions on Evolutionary Computation, 29(2):331–345, 2025.
- Wu et al. (2025) Xingyu Wu, Sheng-Hao Wu, Jibin Wu, Liang Feng, and Kay Chen Tan. Evolutionary computation in the era of large language model: Survey and roadmap. IEEE Transactions on Evolutionary Computation, 29(2):534–554, 2025.
- Yang et al. (2024) Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim. Parallelizing linear transformers with the delta rule over sequence length. In Advances in Neural Information Processing Systems (NeurIPS), 2024.
- Yao et al. (2023) Man Yao, Jiakui Hu, Zhaokun Zhou, Li Yuan, Yonghong Tian, Bo Xu, and Guoqi Li. Spike-driven transformer. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
- Zhang et al. (2026a) Malu Zhang, Wenjie Wei, Zijian Zhou, Wanlong Liu, Jie Zhang, Ammar Belatreche, and Yang Yang. Spike-driven lightweight large language model with evolutionary computation. IEEE Transactions on Evolutionary Computation, 30(4):1333–1346, 2026a.
- Zhang et al. (2010) Qingfu Zhang, Wudong Liu, Edward Tsang, and Botond Virginas. Expensive multiobjective optimization by MOEA/D with Gaussian process model. IEEE Transactions on Evolutionary Computation, 14(3):456–474, 2010.
- Zhang & Lu (2026) Rui Zhang and Zhichao Lu. Rethinking code similarity for automated algorithm design with LLMs. In International Conference on Learning Representations (ICLR), 2026.
- Zhang et al. (2026b) Xian-Rong Zhang, Yue-Jiao Gong, Wei-Neng Chen, and Jun Zhang. Evolutionary neural architecture search with dual contrastive learning. Applied Soft Computing, 189, 2026b. Art. no. 114507.
- Zhang et al. (2015) Xingyi Zhang, Ye Tian, and Yaochu Jin. A knee point-driven evolutionary algorithm for many-objective optimization. IEEE Transactions on Evolutionary Computation, 19(6):761–776, 2015.
- Zhong et al. (2026) Yan Zhong, Ruoyu Zhao, Chao Wang, Jiaqi He, Qinghai Guo, Jianguo Zhang, Zhichao Lu, and Luziwei Leng. Dyn-SSM: Towards the efficient long sequence learning via bio-interpretable dynamics in spiking state space models. IEEE Transactions on Cognitive and Developmental Systems, pp. 1–15, 2026. Early access.
- Zhou et al. (2025) Xun Zhou, Xingyu Wu, Liang Feng, Zhichao Lu, and Kay Chen Tan. Design principle transfer in neural architecture search via large language models. In AAAI Conference on Artificial Intelligence (AAAI), 2025.
- Zhou et al. (2023) Zhaokun Zhou, Yuesheng Zhu, Chao He, Yaowei Wang, Shuicheng Yan, Yonghong Tian, and Li Yuan. Spikformer: When spiking neural network meets transformer. In International Conference on Learning Representations (ICLR), 2023.
- Zhu et al. (2024) Rui-Jie Zhu, Qihang Zhao, Guoqi Li, and Jason K. Eshraghian. SpikeGPT: Generative pre-trained language model with spiking neural networks. Transactions on Machine Learning Research, 2024.
Appendix A Supporting Architecture Comparisons
A.1 Paired ANN and SNN Evaluation
We adapted the 106 architecture programs released by ASI-Arch (Liu et al., 2025) to the spiking-projection constraint below and trained the ANN and SNN implementations separately from scratch on WT2 and ListOps. Nine SNN runs produced non-finite loss; the remaining 97 architectures have complete paired results on both tasks. This pool is distinct from the 20 experts used to initialize our search. The WT2 ANN–SNN comparison uses test perplexity and tie-adjusted Kendall , giving on the 97 successful pairs.
On these same architecture identifiers, the cross-task rank correlation between WT2 perplexity and ListOps accuracy is for SNNs and for ANNs, with lower perplexity and higher accuracy oriented as better.
A.2 Scope of the Spiking Constraint
The spiking-projection constraint places binary spike inputs before the parameter-dominant feature projections. Membrane potentials, residual streams, recurrent states, and token-mixer arithmetic remain continuous. Lightweight control computations may also be continuous, including Comba’s closed-loop state feedback and LoopMem’s evolved state-norm feedback network. Thus the constraint does not imply that all arithmetic is spike-driven. Candidate blocks use the shared hard-reset LIF implementation in Eq. 1; the surrogate dynamic network (SDN) of SpikingSSMs (Shen et al., 2025) accelerates the neuron computation during search and evaluation. Candidates must also preserve the block interface, causality, and subquadratic sequence complexity.
Appendix B Architectural Changes and the Scope of Published Search Encodings
We compare the dependencies in the discovered blocks with the choices exposed by published search encodings. Table 5 records the available primitives and composition rules, together with the extensions needed to directly represent the relevant computation. This is an encoding-level comparison: the methods address different tasks and use different evaluation protocols. The criterion is whether a specified computational relation is selectable under the documented encoding, rather than whether another network could approximate its input–output function.
| Method and source location | Encoded architectural choices | Relation to the discovered computations |
|---|---|---|
| AutoSNN (Na et al., 2022), Sec. 4.1 | Five block choices in a fixed backbone: skip, spiking convolution, and spiking residual blocks, with specified kernel sizes. | The block menu needs a spike-statistic controller and its target connection to select the activity-dependent update or residual modulation. |
| SNASNet (Kim et al., 2022), Cell Search Strategy | Four-node cells with zero, skip, convolution, and pooling operations on forward and cross-time backward edges. | Temporal feedback is already permitted. A spike-statistic reduction and learned multiplicative control are additional operations beyond this edge menu. |
| MSE-NAS (Pan et al., 2025), Sec. III, Fig. 1 | A multiscale genotype selects layer operations, excitatory/inhibitory types, motifs, and global connections under a specified decoder. | These choices alter operations and connectivity; the decoder would need to expose activity-conditioned control equations and their attachment points. |
| EQ-SpikeLM (Zhang et al., 2026), Sec. IV-B.1, Eqs. (14)–(16) | Per-layer preserved channel ratios for query–key, value, and feed-forward projections in a pretrained spiking language model. | Channel pruning changes widths within the existing computation; it does not introduce a new spike-derived control dependency. |
| ENAS (efficient NAS), recurrent space (Pham et al., 2018), Secs. 2.1, 3.1 | Predecessor and activation choices in a recurrent cell, with a prescribed highway-gating construction. | Recurrence and multiplicative gating are present. Spike-statistic inputs and revisions to the controller’s equation or target require extending the template. |
| DARTS, recurrent space (Liu et al., 2019), Sec. 3.1.2 | Operation choices over linear transforms and activations, identity, and zero, within a recurrent-cell template with highway bypasses. | Selecting operations does not itself expose the spike-reduction and control-target relation; these must be added to the operations or template. |
| einspace (Ericsson et al., 2024), Secs. 3.1–3.3, 5 | Recursive grammar for sequential, branching, routing, and computation modules, including matrix multiplication, summation, and concatenation. | Composition is substantially broader than a fixed cell. The published grammar excludes recurrent computation, so the complete recurrent SNN blocks require a grammar extension. |
The computational relations being compared. NeuroGate forms activity statistics from spikes and uses learned transformations to modulate the recurrent update coefficient and output-gate input. HomeoResSSM constructs separate activity-conditioned gains for the state-space and channel-mixer residual branches. Their distinguishing dependency is therefore spike statistic learned control a specified update or branch, within an otherwise inherited recurrent core. LoopMem provides a related case: a normalized summary of the recurrent state controls the forget gate, giving the dependency state summary learned control gate, which likewise requires a summary input and a controller attached to the gate.
Appendix C Search Implementation
Algorithms 1 and 2 specify the outer and inner loops. Each nondominated ranking uses a fixed reference set for novelty: the evaluated archive for parent selection and the filtered candidate pool for training-batch selection. denotes the nondominated subset of under . The new-evaluation budget excludes the initial expert archive. Island survival changes the active parent population; accepted programs remain in the cumulative candidate pool.
C.1 Fingerprint and Three-View Comparison
The 21 measurements in Table 6 are extracted from the instantiated model with its fixed outer wrapper. Module statistics use the module inventory; structural measurements use the execution trace and tensor shapes. Initialization-time probes use a fixed ListOps minibatch, sequence shape, initialization procedure, and seed across candidates. The five classical proxies follow the corresponding definitions (Tanaka et al., 2020; Abdelfattah et al., 2021; Lee et al., 2019; Mellor et al., 2021).
For SWSP, collect binary responses in , with probe samples and recorded neuron–sequence-position responses. Following the sample-wise orientation of SWAP (Peng et al., 2024b),
| (6) |
FireRate is the mean spike activity across the recorded spiking layers at initialization. Both are surrogate inputs measured before training.
| Group | Measurement | Type | Description |
|---|---|---|---|
| Module statistics | Params | C | Total trainable parameters |
| SpkLinear | C | Spiking linear module count | |
| Conv1d | C | 1-D convolution count | |
| GateRatio | C | Gating layer fraction | |
| Structural statistics | Depth | F | Shape-changing layer count |
| Expand | F | Dim-increase steps | |
| Contract | F | Dim-decrease steps | |
| BneckRatio | F | Min-dim / model-dim | |
| MaxExpand | F | Largest expansion factor | |
| FFNexp | F | Avg FFN expansion ratio | |
| FFNparam | F | FFN parameter fraction | |
| FFNmac | F | FFN MAC fraction | |
| MAC/Param | F | Compute density | |
| SubgraphR | F | Largest FX subgraph ratio | |
| Initialization-time probes | SynFlow | F+B | Synaptic flow |
| GradNorm | F+B | Gradient norm | |
| SNIP | F+B | Connection sensitivity | |
| Jacobcov | F+B | Jacobian covariance | |
| NASWOT | F | Activation overlap | |
| SWSP | F | Sample-wise spiking patterns | |
| FireRate | F | Spiking firing rate |
C.2 Three-View Similarity
Code similarity averages full Python CodeBLEU in both reference–candidate directions, with equal weights on its four components (Ren et al., 2020). Rationale similarity uses cosine similarity between 2,048-dimensional text-embedding-v4 embeddings (Alibaba Cloud, 2025). Before fingerprint cosine similarity, each feature is transformed by
| (7) |
which maps to ; we set . The 21 reference pairs were calibrated on the 97 ASI-Arch paired architectures, satisfy , and remain fixed throughout search. The transform is applied without clipping; cosine similarity is set to zero if either vector has zero norm. The three views have equal weights as in Eq. 3.
C.3 Initial Expert Archive
The initial archive contains 20 linear recurrent and state-space architectures from five families (Table 7), each adapted to the spiking-projection constraint in Appendix A.2. Eighteen are taken from the Flash Linear Attention library (Yang & Zhang, 2024); S4 (Gu et al., 2022) and MetaLA (Chou et al., 2024) extend coverage of state-space models and modern RNNs. The families span data-independent and data-dependent gating, delta-rule updates, and closed-loop state feedback. Designs that augment the recurrent state with an external memory hierarchy, such as the hierarchical memory for Mamba (Wang et al., 2026), are not included. All 20 are trained on both proxy tasks to supply initial parents and labeled surrogate context; the spiking implementations will be released with the code.
| Model | Family | Venue | Ref. | Model | Family | Venue | Ref. |
|---|---|---|---|---|---|---|---|
| RetNet | Foundational | arXiv 2023 | (Sun et al., 2023) | S4 | SSM | ICLR 2022 | (Gu et al., 2022) |
| LightNet | Foundational | TMLR 2026 | (Qin et al., 2026) | Mamba (S6) | SSM | COLM 2024 | (Gu & Dao, 2024) |
| GLA | Modern RNN | ICML 2024 | (Yang et al., 2024a) | Mamba2 (SSD) | SSM | ICML 2024 | (Dao & Gu, 2024) |
| HGRN | Modern RNN | NeurIPS 2023 | (Qin et al., 2023) | DeltaNet | Delta Rule | NeurIPS 2024 | (Yang et al., 2024b) |
| HGRN2 | Modern RNN | COLM 2024 | (Qin et al., 2024) | DeltaFormer | Delta Rule | arXiv 2025 | (Zhong et al., 2025) |
| RWKV-6 | Modern RNN | COLM 2024 | (Peng et al., 2024a) | Gated DeltaNet | Delta Rule | ICLR 2025 | (Yang et al., 2025) |
| RWKV-7 | Modern RNN | COLM 2025 | (Peng et al., 2025) | KDA | Delta Rule | arXiv 2025 | (Kimi Team, 2025) |
| ABC | Modern RNN | ACL 2022 | (Peng et al., 2022) | DeltaProduct | Delta Rule | NeurIPS 2025 | (Siems et al., 2025) |
| GSA | Modern RNN | NeurIPS 2024 | (Zhang et al., 2024) | MesaNet | Advanced | ICLR 2026 | (von Oswald et al., 2026) |
| MetaLA | Modern RNN | NeurIPS 2024 | (Chou et al., 2024) | Comba | Advanced | NeurIPS 2025 | (Hu et al., 2025) |
C.4 Search Configuration and Training Cost
| Setting | Value |
|---|---|
| LLM / maximum generation tokens | Gemini-2.5-Flash / 32,768 |
| Outer iterations / accepted candidates per iteration | 10 / 1,080 |
| Batch limit / evaluation budget | 16 / 160 |
| Islands / retained capacity | 10 / approximately 20 per island |
| Population trimming threshold | Twice the retained capacity |
| Inter-island parent sampling probability | 0.5 |
| Reset interval / fraction of islands reset | 3,600 seconds / one-half |
| Performance / dissimilarity shortlist sizes | 30 / 30 |
| Similarity early-exit tolerance |
Each iteration selects at most candidates for training; candidates that duplicate archived architectures are removed, so some batches are smaller and 151 new architectures are trained in total. Each candidate-training job uses one NVIDIA V100 (32 GB). Approximate per-architecture training costs are one GPU-hour on WT2 and 20 GPU-hours on ListOps, giving GPU-hours, or about 132 GPU-days, excluding initial-expert training, downstream retraining, and LLM inference.
C.5 LLM Selection
Before the main search, candidate LLMs, including Gemini-2.5-Flash (Comanici et al., 2025), OpenAI o3 (OpenAI, 2025), and Qwen3-Max (Qwen Team, 2025), were compared as program generators on sampling success, performance and novelty of the generated architectures, generation time, and API expense (Fig. 9). No model is best on every axis. Because the search issues more than two thousand generation calls per iteration (1,080 accepted candidates at a 49% acceptance rate; Table 8 and Fig. 7), its time and expense scale with the number of calls. Gemini-2.5-Flash is best in time, expense, and sampling success and is therefore used in all reported searches.
C.6 Pareto-Front Evolution
The outer loop selects candidates by two objectives, composite fitness and novelty (Section 3.4). Fig. 10 shows the nondominated front of all trained architectures in these objectives after each iteration. The front advances along both axes. The selected architectures are marked where they first appear and on the final front: NeuroGate (iteration 9) has the highest fitness; HomeoResSSM (iteration 3) lies at the knee point, the member farthest from the line joining the two extreme solutions after min–max normalization of both objectives; and LoopMem (iteration 3) is the most novel member with nontrivial performance on both tasks, as the two more novel members reach only 18.7% and 17.8% ListOps accuracy.
The composite fitness itself combines two task objectives. Fig. 11 shows the corresponding front in ListOps accuracy and WT2 perplexity. At initialization, SpikingDeltaNet alone forms this front. As the search proceeds, the front extends in both directions, and four search candidates, including the one with the highest fitness, improve on SpikingDeltaNet in both metrics.
Appendix D Evaluation Protocols
D.1 Training Tasks and Fixed Model Components
| Setting | WT2 | ListOps |
|---|---|---|
| Block layers / width | 6 / 256 | 2 / 128 |
| Attention heads | 8 | 4 |
| FFN expansion | 4 | 2 |
| Dropout | 0.2 | 0.1 |
| Batch size | 16 | 32 |
| Epochs | 30 | 25 |
| Learning rate | ||
| Weight decay | 0.15 | |
| Warmup steps | 600 | 3000 |
| Gradient clipping norm | 2.0 | 2.0 |
| Optimizer | AdamW | AdamW |
| Schedule | Cosine | Cosine |
Search-time WT2 fitness uses validation perplexity. WT2 uses GPT-2 byte-level BPE, vocabulary size 50,257, and concatenated token streams chunked into length-512 sequences. ListOps uses a training-derived vocabulary, whitespace tokenization, and an appended end-of-sequence token. Search-time ListOps training uses one-third of the LRA training split, and accuracy is measured on the LRA test split. Downstream ListOps evaluation trains on the full training split and reports accuracy on the same test split. The other five LRA tasks are not used during search, and WT103 perplexity is reported on the test split.
The fixed wrapper embeds tokens, stacks candidate blocks, and applies final RMS normalization. Language modeling uses a vocabulary projection; ListOps uses mean pooling and a ten-class head. Table 9 specifies the proxy-training settings. AdamW uses , cosine decay, and linear warmup. For downstream evaluation, we follow the task-specific training configurations released with SpikingSSMs (Shen et al., 2025) in its official repository, which builds on the S4 codebase.
D.2 Fitness Definition
The task transforms are anchored to SpikingDeltaNet, the strongest model on both proxy tasks among the 117 architectures evaluated before the search (the 20 experts and the 97 ASI-Arch paired architectures in Appendix A.1); its ANN counterpart, DeltaNet, is likewise the strongest ANN among them and serves as the ANN reference in Table 1. Let and denote its WT2 perplexity and ListOps accuracy in percentage points. The relative improvements are
| (8) | ||||||
| (9) |
and, for each task ,
| (10) |
where ; Eq. 5 averages the two task scores. We set , , , , , , and . Each anchor score maps to , and each task score lies in . Because none of the other 19 experts surpasses the anchor on either task, the narrow improvement tolerance gives large rewards to candidates that do, while the wide, bounded loss side keeps failed candidates from dominating the scale.
WT2-only studies use alone with the same anchor and , , and .
D.3 Offline Surrogate Studies
These studies were conducted before the main search and guided the choice of its surrogate and fingerprint. Their architectures were drawn at random from the 97 ASI-Arch paired architectures, the 20 experts, and candidates generated with a preliminary surrogate; fewer were trained on ListOps, whose training is slower. The regressor and progressive-input studies use 223 WT2 architectures and 54 ListOps architectures. They share five-fold cross-validation splits repeated three times and report Kendall rank correlation with measured task outcomes. The progressive study starts with five classical probes, adds SWSP and FireRate together, and then adds module and structural statistics.
The labeled-context curves use fixed holdouts of 50 WT2 and ten ListOps architectures. At each context size, 20 subsets are drawn from the remaining pools of 173 and 44 architectures, respectively; the curves show mean correlation and one standard deviation.
D.4 Search-Method Comparisons
Inner-loop comparisons. The EoH and FunSearch adaptations and OpenArchEvo share the outer loop in Algorithm 1: the same expert archive, surrogate refresh, candidate filtering, and NSGA-II batch selection for training. Only the inner loop in Algorithm 2 differs. The EoH adaptation keeps design guidance in LLM variation but evolves one population without similarity-based islands or island resets. The FunSearch adaptation keeps the islands but removes design guidance from variation. The main search needs no warm-start stage, since its expert archive is trained on both tasks and supplies the initial surrogate context and parents. In the WT2-only studies, a warm-start stage precedes the first outer iteration: a single-island inner loop diversifies the expert archive, and a batch of its candidates is trained on WT2 to extend the surrogate context.
Each configuration runs a warm-start stage and six outer iterations of WT2-only search, three times, and uses the WT2-only fitness in Appendix D.2. After each stage, we record the best fitness among trained candidates and the hypervolume of the nondominated set in the novelty–fitness plane. For these comparisons, novelty is recomputed against one fixed reference set shared across methods, runs, and stages. Hypervolume also uses a common reference point, placed just below the minimum novelty and fitness observed across the runs. Curves show the mean over runs, and shaded bands show one sample standard deviation.
Warm-start stage. The OpenArchEvo w/o Warm-Start variant skips the warm-start stage and reallocates its training budget to iteration 1, keeping the total training budget equal; its Warm-Start value in Fig. 12 therefore equals the expert value. Its first selections for training thus rely on a surrogate whose labeled context contains only the expert archive. After six iterations, the variant reaches a best fitness of 0.557 versus 0.679 for OpenArchEvo and a hypervolume of 0.142 versus 0.215, and it remains below OpenArchEvo from the warm-start stage onward (Fig. 12).
ASI-Arch. For ASI-Arch (Liu et al., 2025), we retain the agent workflow and original search hyperparameters where applicable, replace the evaluation tasks, and add the spiking constraint to the prompts. Table 3 reports one run per method.
ASI-Arch draws on a literature-derived prior of about 100 papers and nine cooperating agents, whereas OpenArchEvo starts from 20 expert architectures with one program generator. The resulting SpikingCondFuse combines a delta-rule recurrent path with a local value path containing a five-tap finite-impulse-response filter; a conditioned gate fuses the paths using hidden-state statistics. Its downstream results are in Table 3.
Appendix E Discovered Architecture Details
The discovered implementations retain the recurrent operator families of the expert architectures from which they descend. NeuroGate adds two spiking activity-to-control projections and removes the SiLU activation after its query, key, and value short convolutions; it also adds a mean-spike penalty with weight to the training loss. HomeoResSSM adds a separate spiking gate network to each residual branch, with hidden width . LoopMem adds a state-norm feedback to the forget gate of its closed-loop recurrence and penalties on the mean spike activity of its token-mixing and channel-mixing branches.
Spike-activity-dependent controls. NeuroGate forms from mean query, key, and value spikes, and modifies the update coefficient and output-gate input as
| (11) |
where each projection includes its spiking encoder. HomeoResSSM scales each branch output before residual addition:
| (12) |
where is the branch’s mean spike activity and both projections are spiking. For each chunk, LoopMem offsets the forget-gate logit by a function of the normalized recurrent state and bounds the state norm after the update:
| (13) |
where is a two-layer network of width 8 and norms are taken per head. It also adds the activity penalty to the training loss, with .
Appendix F Arithmetic Energy Accounting
Energy is compared for all models under one common architecture setting, independent of the task-specific training configurations: blocks, model width , eight heads, FFN width , sequence length , and vocabulary size . All SNNs share the same wrapper, whose FFN and output head receive spikes; the token embedding lookup is excluded. For spike-driven projection , let denote its dense-equivalent operation count and its measured, operation-weighted mean input firing rate. We estimate
| (14) |
with pJ and pJ (Horowitz, 2014); is the number of synaptic operations (SOPs). All designated SpikingLinear projections, including the output head, are counted as accumulate operations. Continuous state updates, readouts, short convolutions, and feedback projections are counted as MACs; a spike-driven input projection does not make the subsequent recurrent computation spike-driven. Continuous scalar multiplications are included as MAC equivalents. Normalization, nonlinear functions, LIF/SDN execution, and memory access are excluded; the FFT arithmetic of PMBC is included for Dyn-SSM as specified below. NeuroGate uses the output-gate-enabled configuration.
Firing rates differ across architectures under the same data, wrapper, and training configuration, so we use each model’s measured, operation-weighted mean input firing rate of its spike-driven projections (Table 10). For Dyn-SSM (Zhong et al., 2026), the block combines its token mixer, an SSM convolution followed by a spike-driven convolution from to channels and a GLU, with the shared spiking FFN. The token mixer’s SSM convolution is counted using FFT arithmetic: with the kernel spectrum cached and , a convolution costs MAC equivalents per channel for the transforms and pointwise product. For PMBC, we assume one boundary-compression iteration (). The algorithm computes one membrane-integration convolution and two convolutions per iteration for the upper and lower bounds, giving additional convolutions per channel (Zhong et al., 2026). These FFT operations add arithmetic beyond direct LIF updates and are included explicitly; the estimate does not cover the full implementation cost of LIF or SDN. This single-iteration energy scenario is separate from the published settings underlying the quoted task scores. The dense Transformer reference with the same depth and width counts
| (15) |
MACs, giving mJ; the energy reduction is .
| Architecture | (%) | SOPs (G) | MACs (G) | Energy (mJ) | Reduction |
|---|---|---|---|---|---|
| Dense Transformer | – | – | 9.809 | 45.12 | 1.0 |
| Dyn-SSM | 9.3 | 0.800 | 0.140 | 1.36 | 33.1 |
| SpikingDeltaNet | 9.8 | 0.883 | 0.225 | 1.83 | 24.7 |
| SpikingMamba2 | 12.2 | 1.247 | 0.420 | 3.05 | 14.8 |
| SpikingCondFuse | 8.3 | 0.782 | 0.367 | 2.39 | 18.9 |
| NeuroGate | 4.6 | 0.424 | 0.226 | 1.42 | 31.7 |
| HomeoResSSM | 7.2 | 0.736 | 0.422 | 2.60 | 17.3 |
| LoopMem | 4.9 | 0.432 | 0.110 | 0.89 | 50.6 |
For DeltaNet-derived programs, the MAC count includes the chunkwise delta-rule correction, local interactions, and recurrent-state read/write operations. Mamba2-derived programs include local SSD interactions, state propagation, decay scaling, and output gating. LoopMem includes its continuous state-norm feedback; SpikingCondFuse additionally includes the local filter and fusion arithmetic.
Appendix G Generation Prompts
The system and user prompt templates below retain the historical wording used in the experiments. They adapt the open-source ASI-Arch prompts (Liu et al., 2025) with SNN-specific constraints and interface requirements; placeholders in the user template are filled with the selected parent programs and rationales. The user template also asks the generator to consider structures that may reduce the firing rate. The operational search constraints are specified in the main text.
G.1 System Prompt
G.2 User Prompt
References
- Abdelfattah et al. (2021) Mohamed S. Abdelfattah, Abhinav Mehrotra, Łukasz Dudziak, and Nicholas D. Lane. Zero-cost proxies for lightweight NAS. In International Conference on Learning Representations (ICLR), 2021.
- Alibaba Cloud (2025) Alibaba Cloud. Embedding: text-embedding-v4. Alibaba Cloud Model Studio documentation, 2025.
- Chou et al. (2024) Yuhong Chou, Man Yao, Kexin Wang, Yuqi Pan, Ruijie Zhu, Yiran Zhong, Yu Qiao, Jibin Wu, Bo Xu, and Guoqi Li. MetaLA: Unified optimal linear approximation to softmax attention map. In Advances in Neural Information Processing Systems (NeurIPS), 2024.
- Comanici et al. (2025) Gheorghe Comanici, Eric Bieber, Mike Schaekermann, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025.
- Dao & Gu (2024) Tri Dao and Albert Gu. Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality. In International Conference on Machine Learning (ICML), 2024.
- Ericsson et al. (2024) Linus Ericsson, Miguel Espinosa, Chenhongyi Yang, Antreas Antoniou, Amos Storkey, Shay B. Cohen, Steven McDonagh, and Elliot J. Crowley. einspace: Searching for neural architectures from fundamental operations. In Advances in Neural Information Processing Systems (NeurIPS), 2024.
- Gu & Dao (2024) Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. In Conference on Language Modeling (COLM), 2024.
- Gu et al. (2022) Albert Gu, Karan Goel, and Christopher Ré. Efficiently modeling long sequences with structured state spaces. In International Conference on Learning Representations (ICLR), 2022.
- Horowitz (2014) Mark Horowitz. 1.1 Computing’s energy problem (and what we can do about it). In IEEE International Solid-State Circuits Conference (ISSCC), 2014.
- Hu et al. (2025) Jiaxi Hu, Yongqi Pan, Jusen Du, Disen Lan, Xiaqiang Tang, Qingsong Wen, Yuxuan Liang, and Weigao Sun. Improving bilinear RNNs with closed-loop control. In Advances in Neural Information Processing Systems (NeurIPS), 2025.
- Kim et al. (2022) Youngeun Kim, Yuhang Li, Hyoungseob Park, Yeshwanth Venkatesha, and Priyadarshini Panda. Neural architecture search for spiking neural networks. In European Conference on Computer Vision (ECCV), 2022.
- Kimi Team (2025) Kimi Team. Kimi Linear: An expressive, efficient attention architecture. arXiv preprint arXiv:2510.26692, 2025.
- Lee et al. (2019) Namhoon Lee, Thalaiyasingam Ajanthan, and Philip H. S. Torr. SNIP: Single-shot network pruning based on connection sensitivity. In International Conference on Learning Representations (ICLR), 2019.
- Liu et al. (2019) Hanxiao Liu, Karen Simonyan, and Yiming Yang. DARTS: Differentiable architecture search. In International Conference on Learning Representations (ICLR), 2019.
- Liu et al. (2025) Yixiu Liu, Yang Nan, Weixian Xu, Xiangkun Hu, Lyumanshan Ye, Zhen Qin, and Pengfei Liu. AlphaGo moment for model architecture discovery. arXiv preprint arXiv:2507.18074, 2025.
- Mellor et al. (2021) Joseph Mellor, Jack Turner, Amos Storkey, and Elliot J. Crowley. Neural architecture search without training. In International Conference on Machine Learning (ICML), 2021.
- Na et al. (2022) Byunggook Na, Jisoo Mok, Seongsik Park, Dongjin Lee, Hyeokjun Choe, and Sungroh Yoon. AutoSNN: Towards energy-efficient spiking neural networks. In International Conference on Machine Learning (ICML), 2022.
- OpenAI (2025) OpenAI. OpenAI o3 and o4-mini system card. System card, 2025.
- Pan et al. (2025) Wenxuan Pan, Feifei Zhao, Guobin Shen, Bing Han, and Yi Zeng. Brain-inspired multiscale evolutionary neural architecture search for deep spiking neural networks. IEEE Transactions on Evolutionary Computation, 29(5):2258–2270, 2025.
- Peng et al. (2024a) Bo Peng, Daniel Goldstein, Quentin Anthony, et al. Eagle and Finch: RWKV with matrix-valued states and dynamic recurrence. In Conference on Language Modeling (COLM), 2024a.
- Peng et al. (2025) Bo Peng, Ruichong Zhang, Daniel Goldstein, Eric Alcaide, Xingjian Du, Haowen Hou, Jiaju Lin, Jiaxing Liu, Janna Lu, William Merrill, et al. RWKV-7 “Goose” with expressive dynamic state evolution. In Conference on Language Modeling (COLM), 2025.
- Peng et al. (2022) Hao Peng, Jungo Kasai, Nikolaos Pappas, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz, and Noah A. Smith. ABC: Attention with bounded-memory control. In Annual Meeting of the Association for Computational Linguistics (ACL), 2022.
- Peng et al. (2024b) Yameng Peng, Andy Song, Haytham M. Fayek, Vic Ciesielski, and Xiaojun Chang. SWAP-NAS: Sample-wise activation patterns for ultra-fast NAS. In International Conference on Learning Representations (ICLR), 2024b.
- Pham et al. (2018) Hieu Pham, Melody Y. Guan, Barret Zoph, Quoc V. Le, and Jeff Dean. Efficient neural architecture search via parameters sharing. In International Conference on Machine Learning (ICML), 2018.
- Qin et al. (2023) Zhen Qin, Songlin Yang, and Yiran Zhong. Hierarchically gated recurrent neural network for sequence modeling. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
- Qin et al. (2024) Zhen Qin, Songlin Yang, Weixuan Sun, Xuyang Shen, Dong Li, Weigao Sun, and Yiran Zhong. HGRN2: Gated linear RNNs with state expansion. In Conference on Language Modeling (COLM), 2024.
- Qin et al. (2026) Zhen Qin, Yuxin Mao, Xuyang Shen, Dong Li, Jing Zhang, Yuchao Dai, and Yiran Zhong. You only scan once: Efficient multi-dimension sequential modeling with LightNet. Transactions on Machine Learning Research, 2026. ISSN 2835-8856.
- Qwen Team (2025) Qwen Team. Qwen3-Max: Just scale it. Model release blog, September 2025.
- Ren et al. (2020) Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. CodeBLEU: A method for automatic evaluation of code synthesis. arXiv preprint arXiv:2009.10297, 2020.
- Shen et al. (2025) Shuaijie Shen, Chao Wang, Renzhuo Huang, Yan Zhong, Qinghai Guo, Zhichao Lu, Jianguo Zhang, and Luziwei Leng. SpikingSSMs: Learning long sequences with sparse and parallel spiking state space models. In AAAI Conference on Artificial Intelligence (AAAI), 2025.
- Siems et al. (2025) Julien Siems, Timur Carstensen, Arber Zela, Frank Hutter, Massimiliano Pontil, and Riccardo Grazzi. DeltaProduct: Improving state-tracking in linear RNNs via Householder products. In Advances in Neural Information Processing Systems (NeurIPS), 2025.
- Sun et al. (2023) Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. Retentive network: A successor to transformer for large language models. arXiv preprint arXiv:2307.08621, 2023.
- Tanaka et al. (2020) Hidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, and Surya Ganguli. Pruning neural networks without any data by iteratively conserving synaptic flow. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
- von Oswald et al. (2026) Johannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari, Songlin Yang, Maximilian Schlegel, Kaitlin Maile, Yanick Schimpf, Oliver Sieberling, Alexander Meulemans, Guillaume Lajoie, Rif A. Saurous, Charlotte Frenkel, Razvan Pascanu, Blaise Aguera y Arcas, and Joao Sacramento. MesaNet: Sequence modeling by locally optimal test-time training. In International Conference on Learning Representations (ICLR), 2026.
- Wang et al. (2026) Qinwen Wang, Jieping Luo, Aoxiang Qin, Ruoyu Zhao, Jianxiong Tang, Wei Zhang, Zhichao Lu, and Luziwei Leng. Mamba with hierarchical memory: Solving representation bottleneck in long sequence modeling. arXiv preprint arXiv:2608.02347, 2026.
- Yang & Zhang (2024) Songlin Yang and Yu Zhang. FLA: A Triton-based library for hardware-efficient implementations of linear attention mechanism. Software library, 2024.
- Yang et al. (2024a) Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim. Gated linear attention transformers with hardware-efficient training. In International Conference on Machine Learning (ICML), 2024a.
- Yang et al. (2024b) Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim. Parallelizing linear transformers with the delta rule over sequence length. In Advances in Neural Information Processing Systems (NeurIPS), 2024b.
- Yang et al. (2025) Songlin Yang, Jan Kautz, and Ali Hatamizadeh. Gated delta networks: Improving Mamba2 with delta rule. In International Conference on Learning Representations (ICLR), 2025.
- Zhang et al. (2026) Malu Zhang, Wenjie Wei, Zijian Zhou, Wanlong Liu, Jie Zhang, Ammar Belatreche, and Yang Yang. Spike-driven lightweight large language model with evolutionary computation. IEEE Transactions on Evolutionary Computation, 30(4):1333–1346, 2026.
- Zhang et al. (2024) Yu Zhang, Songlin Yang, Ruijie Zhu, Yue Zhang, Leyang Cui, Yiqiao Wang, Bolun Wang, Freda Shi, Bailin Wang, Wei Bi, Peng Zhou, and Guohong Fu. Gated slot attention for efficient linear-time sequence modeling. In Advances in Neural Information Processing Systems (NeurIPS), 2024.
- Zhong et al. (2025) Shu Zhong, Mingyu Xu, Tenglong Ao, and Guang Shi. Understanding transformer from the perspective of associative memory. arXiv preprint arXiv:2505.19488, 2025.
- Zhong et al. (2026) Yan Zhong, Ruoyu Zhao, Chao Wang, Jiaqi He, Qinghai Guo, Jianguo Zhang, Zhichao Lu, and Luziwei Leng. Dyn-SSM: Towards the efficient long sequence learning via bio-interpretable dynamics in spiking state space models. IEEE Transactions on Cognitive and Developmental Systems, pp. 1–15, 2026. Early access.