PyTorch
Safetensors
RWKV
English
spikegpt
spiking-neural-network
full-parameter-finetuning

SpikeGPT-216M-Mixture-Finetune

This repository contains a full-parameter fine-tuned version of ridger/SpikeGPT-OpenWebText-216M. Training consists of mixture SFT v0.7 followed by the lower-learning-rate v0.8 continuation. The released checkpoint is v0.8.

Model identity

Item Value
Base repository ridger/SpikeGPT-OpenWebText-216M
Base immutable revision 4039295cca3da1df0e5871f4bc7727b227496132
Base checkpoint SHA-256 024d2dab5f71b6bd8d4f3ef38bf9bd61c54b2ea8f9efe4b57517bf9b6b0328de
v0.7 checkpoint SHA-256 18e5565175e6658179d3ac3492ab14fce34dbeae5f9a612b2d9d566273cfef1c
Released v0.8 checkpoint SHA-256 (model.pth) 1256a913e566bc6247c73127e7c4ae2348f8fb034b1cf5de13315713ef083471
Released model.safetensors SHA-256 26ce68619b45b9fbd83257630da23039188ab189161062ffbfd997888e405bff
Tokenizer SHA-256 56ac4821e129d2c520fdaba60abd920fa852ada51b45c0dd52bbb6bd8c985ade
State tensors / parameters 330 / 215,399,424
Parameter dtype float32
Layers / embedding width / vocabulary 18 / 768 / 50,277
Training context 512 tokens

Recorded same-split evaluation

The validation split is fixed across all three checkpoints: 965 examples, 54,429 supervised assistant tokens, JSONL SHA-256 23f963aa7ce2038d6c0f104e01ab37a0613de14f6f7b6f5aebb9a2f2c1c670b5. Loss is mean assistant-token cross-entropy and perplexity is exp(loss).

Transition Loss before Loss after Loss Ξ” Loss Ξ”% PPL before PPL after PPL Ξ” PPL Ξ”%
Base β†’ v0.7 3.7950713298 3.1936229874 -0.6014483424 -15.848% 44.48140914 24.37658371 -20.10482543 -45.198%
v0.7 β†’ v0.8 3.1936229874 3.1826934650 -0.0109295224 -0.342% 24.37658371 24.11160995 -0.26497376 -1.087%
Base β†’ v0.8 3.7950713298 3.1826934650 -0.6123778648 -16.136% 44.48140914 24.11160995 -20.36979919 -45.794%

For base β†’ v0.8, exp(-loss) changes from 2.248% to 4.147% (1.8448Γ—), and the cross-entropy difference is 0.88347 bits per supervised token.

Training data

The deterministic mixture manifest SHA-256 is 7378939249535ff604e9eed66acea7414a8a35e63deb46ac649ce90a7bc32b52. It has 12,736 scheduled training occurrences, 9,934 unique source rows, and 1,166,515 supervised tokens per epoch. v0.7 processed two epochs (2,333,030 supervised-token exposures) in 3,184 optimizer steps. v0.8 processed one additional epoch (1,166,515) in 1,592 additional steps. Total exposure is 3,499,545 supervised tokens.

Component Immutable source / generator Unique selected rows Scheduled occurrences Supervised tokens / epoch Terms
Smol constraints HuggingFaceTB/smol-smoltalk@f73fe857d519ff6ac5af2ea67c4d3834da7b8bcc; source smol-contraints 8,000 8,000 788,086 Pinned dataset-card metadata: apache-2.0; no standalone dataset license file at the pinned revision
OASST1 complete-fit OpenAssistant/oasst1@fdf72ae0827c1cda404aff25b6603abec9e3399b 934 3,736 373,692 Apache-2.0 standalone license evidence
Repository-authored CC0 calibration generator SHA-256 9f02cf8a88f05c32c587443921f31fb3f0801fbe1a374d9e9aff54570ae47344 1,000 1,000 4,737 CC0-1.0

Validation contains 965 examples and 54,429 supervised tokens: 500 Smol test rows, 50 OASST1 validation rows, and 415 repository-authored CC0 validation rows. Complete sequences fit the 512-token context; response truncation is disabled. The combined mixture has component-scoped terms and the aggregate SPDX field NOASSERTION.

The CC0 component source has 1,778 generated training rows; the deterministic mixture selects 1,000 of them and includes all 415 validation rows.

The pinned Smol generation pipeline identifies Qwen/Qwen2.5-72B-Instruct and requires the notice Improved using Qwen. The copied pipeline source is huggingface/smollm@951394e9b214ce91e3223b2257a8eecb0a0d3d4d, path text/data/smoltalk/constraints/pipeline/pipeline.py.

Training configuration and runtime

Phase Steps LR Warmup Wall time Mean logged cumulative supervised tok/s Final cumulative supervised tokens Non-finite numeric metrics OOM markers in metrics
v0.7 3,184 3.0e-6 β†’ 3.0e-7 96 26m 25.7s 1512.9 2,333,030 0 0
v0.8 continuation 1,592 additional 1.0e-6 β†’ 1.0e-7 96 13m 23.5s 1522.3 1,166,515 0 0

The optimizer is AdamW (beta1=0.9, beta2=0.95, eps=1e-8). Microbatch is 1, gradient accumulation is 8, weight decay is 0.01, gradient clipping is 1.0, seed is 20260824, hardware is NVIDIA GB10, and the recorded runtime is PyTorch 2.13.0+cu130 with CUDA 13.0 and CuPy SNN backend.

Weight deltas

All 330 float32 state tensors were compared. Accumulations use float64. changed is exact float32 inequality; relative L2 is ||target - source||β‚‚ / ||source||β‚‚ and is displayed as a percentage.

Transition Changed values Mean absolute Ξ” RMS Ξ” L2 Ξ” Relative L2 Max absolute Ξ” Cosine similarity
base β†’ v0.7 215,392,078 (99.996590%) 0.000239008016 0.000514355541 7.54893509 0.142417% 0.004543975 0.999998989254
v0.7 β†’ v0.8 192,781,844 (89.499703%) 1.27409533e-05 2.70492657e-05 0.396988337 0.007490% 0.000258207321 0.999999997199
base β†’ v0.8 215,392,053 (99.996578%) 0.000248678666 0.000539643772 7.92007762 0.149419% 0.0047981441 0.999998887322

Base β†’ v0.8 parameter RMS changes from 0.3611604 to 0.3611292. Complete values and checkpoint hashes are in weight-delta.json.

Runtime format

The checkpoint uses the custom recurrent SpikeGPT runtime in runtime/. config.json is descriptive metadata and does not define a Transformers AutoModel implementation. Token 0 is EOS/conversation separator; token 1 is padding. Training serialization is <|endoftext|>User: ... Assistant: ...<|endoftext|>; loss is applied to assistant response and EOS tokens.

Files and integrity

  • model.safetensors: 26ce68619b45b9fbd83257630da23039188ab189161062ffbfd997888e405bff
  • model.pth: 1256a913e566bc6247c73127e7c4ae2348f8fb034b1cf5de13315713ef083471
  • 20B_tokenizer.json: 56ac4821e129d2c520fdaba60abd920fa852ada51b45c0dd52bbb6bd8c985ade
  • training-lineage.json: base β†’ v0.7 β†’ v0.8 artifact and checkpoint links
  • weight-delta.json: full-state numerical delta report
  • eval_before.json, eval_after.json, metrics.jsonl: v0.8 continuation records
  • training/phase1/: sanitized v0.7 run, config, evaluation, and metrics records
  • dataset_manifest.json, components/: split hashes, component manifests, license/card/pipeline evidence, Qwen license, and notices
  • source-snapshot.tar.gz, source-snapshot-manifest.json, runtime/: exact v0.8 source/runtime snapshot
  • release_manifest.json: SHA-256 and byte size for all preceding release files

Licenses and attribution

  • Fine-tuned weights: Apache-2.0; see LICENSE.
  • Base weights: ridger/SpikeGPT-OpenWebText-216M@4039295cca3da1df0e5871f4bc7727b227496132. Pinned model-card metadata declares Apache-2.0; the pinned repository has no standalone LICENSE file.
  • OASST1: Apache-2.0; see components/oasst1/OASST1-LICENSE.
  • Repository-authored calibration corpus: CC0-1.0; see components/cc0_calibration/SOURCE.md.
  • Smol constraints: pinned card metadata says apache-2.0; the pinned dataset revision has no standalone dataset license file. Evidence, generation-pipeline source, QWEN-LICENSE, and SOURCE-NOTICE.md are copied under components/smol_constraints/.
  • Qwen notice: Improved using Qwen. Qwen is licensed under the Qwen LICENSE AGREEMENT, Copyright (c) Alibaba Cloud. All Rights Reserved.
  • SpikeGPT source: BSD-2-Clause; see CODE-LICENSE-BSD-2-Clause.
  • RWKV-derived code, CUDA sources, and tokenizer provenance: Apache-2.0; see THIRD-PARTY-LICENSE-RWKV-Apache-2.0. Upstream commit: 9917078f936e1f45bff341cc58550fd03ec4e5eb.
  • SmolLM generation-pipeline code: Apache-2.0; see components/smol_constraints/SMOLLM-PIPELINE-LICENSE. This license copy applies to the pipeline code only.
  • Vendored and modified SpikingJelly 0.0.0.0.12: Open-Intelligence Open Source License V1.0; see THIRD-PARTY-LICENSE-SpikingJelly and THIRD-PARTY-LICENSE-SpikingJelly-CN. Upstream commit: 1171f5249a4ebeace6ab8d6a74d85579fafb93ed; PyPI wheel SHA-256: ffa5298947348d5928840a077dc2840c3f411b192401e1bdc4f2224883a8a51a. Section V disclosure URL: https://www.aitisa.org.cn/OI_Statement/.
  • Paper: Zhu et al., SpikeGPT: Generative Pre-trained Language Model with Spiking Neural Networks.
Downloads last month
8
Safetensors
Model size
0.2B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for TaewoooPark/SpikeGPT-216M-Mixture-Finetune

Finetuned
(1)
this model

Datasets used to train TaewoooPark/SpikeGPT-216M-Mixture-Finetune

Paper for TaewoooPark/SpikeGPT-216M-Mixture-Finetune