whittle-45b-chat / README.md
logic65's picture
Whittle 45B ZeroGPU chat demo: stock transformers + whittle_load.py, model mounted at /models/w45
35b139d verified
|
Raw History Blame Contribute Delete
690 Bytes

A newer version of the Gradio SDK is available: 6.30.0

Upgrade
metadata
title: Whittle 45B
emoji: 🪵
colorFrom: green
colorTo: gray
sdk: gradio
sdk_version: 6.29.1
python_version: '3.12'
app_file: app.py
pinned: false
license: apache-2.0
models:
  - logic65/Whittle-Qwen-3.8-45B-A3B
short_description: Chat with Whittle-Qwen-3.8-45B-A3B on ZeroGPU

Chat demo for Whittle-Qwen-3.8-45B-A3B: stock transformers plus whittle_load.py, which maps Whittle's n-gram table (8 hash heads of 4,880,000 rows, stored in 5 row pieces) onto the stock Qwen4Exp layout. The model repo is mounted read-only at /models/w45; the 35 B body runs on a full RTX PRO 6000 and the 10 B table stays in CPU memory.