Ines-1 / ENVIRONMENT.md
Endikavi's picture
Ines-1 RC1 (private staging; release commit b7f5644)
61b6fb9 verified
|
Raw History Blame
3.3 kB

Canonical evaluation environment (Ines-1)

Every canonical number in this repository was measured in this environment. Other environments may differ in the last bits of bf16 arithmetic; see What changes the numbers.

GPU 1× NVIDIA H200 (143,771 MiB), driver 595.58.03 (CUDA driver 13.2)
Python 3.12.3
PyTorch 2.14.0+cu130 (CUDA runtime 13.0, cuDNN 9.24.0), triton 3.8.0; wheels from https://download.pytorch.org/whl/cu130
other packages requirements.lock (exact versions of the evaluation environment, plus safetensors to load the bundle)
weights / compute bf16 weights, forward under torch.autocast("cuda", dtype=torch.bfloat16)
readout one prompt per question, MiniV41.prefill(ids, num_logits=1), softmax over the option-letter logits (mini_v41_jev.Decider.decide, default path)
TF32 torch.backends.cuda.matmul.allow_tf32 = False, torch.backends.cudnn.allow_tf32 = True (PyTorch defaults), float32_matmul_precision = "highest"
determinism torch.use_deterministic_algorithms off, cudnn.deterministic = False, cudnn.benchmark = False (defaults)
attention (SDPA) default backend selection (flash, memory-efficient, math and cuDNN all enabled)
CUDA graphs / compile not used by the canonical readout; the HTTP server uses CUDA graphs for small prefill buckets (scripts/serve_jev.py), under the same autocast; no torch.compile

What changes the numbers (measured, 400 typed-decisions questions, released weights)

| change | max |Δp| | answers changed | |---|---:|---:| | PyTorch 2.13.0 vs 2.14.0, same code, same settings (default, math-only and efficient-only SDPA) | 0 | 0 / 400 | | bf16 autocast on vs off | 0.0705 | 8 / 400 | | default SDPA vs math-only SDPA (either PyTorch version) | 0.0506 | 8 / 400 | | this bundle vs the original training checkpoint and code, same environment and settings | 0 | 0 / 400 |

So the version of PyTorch is not what moves results; the precision policy (autocast) and the attention backend are. The batched path (Decider.decide(..., batched=True) and the HTTP server: right-padded rows, per-row decoder span) is not bit-identical to the canonical one; its accuracy through HTTP matched the canonical evaluation within bf16 noise.

Dockerfile

Dockerfile reproduces this environment as an image: python:3.12.3-slim-bookworm + pip install -r requirements.lock (PyTorch 2.14.0 from PyPI is the CUDA 13.0 build: it reports 2.14.0+cu130, with the CUDA and cuDNN libraries as pip wheels; the NVIDIA driver comes from the host). The model is not copied into the image; mount the repository at /model:

docker build -t ines-1-env .
docker run --rm --gpus all -v "$PWD":/model ines-1-env        # GPU: examples/quickstart.py, bf16 + autocast
docker run --rm -v "$PWD":/model ines-1-env                   # CPU (fp32)

Built once from scratch (--no-cache --pull, 2026-10-02) and tested on a copy of this repository: Python 3.12.3, torch 2.14.0+cu130, CUDA 13.0, cuDNN 9.24.0, numpy 2.5.3, tokenizers 0.23.2, safetensors 0.8.0; quickstart on CPU and GPU; on GPU (bf16 + autocast) the 400-question probe matched the canonical outputs exactly (0 answers changed, max |Δp| 6e-8), and the batched path differed in 8 of 400, as above.