Instructions to use divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Huihui-gemma-4-31B-it-abliterated — 4-bit MLX for Apple Silicon
Uncensored Gemma 4 31B, local on Apple Silicon. A 4-bit MLX conversion of huihui-ai/Huihui-gemma-4-31B-it-abliterated, the abliterated Gemma 4 31B Instruct. 17.3 GB on disk — wants a 32 GB Mac. No cloud, no API key, no refusals.
| Size on disk | 17.3 GB (4 shards) |
| Mac RAM | 32 GB recommended, 24 GB for short contexts |
| Architecture | Gemma 4 |
What "abliterated" means
The refusal direction has been orthogonalized out of the weights, so the model answers instructions a stock instruct-tuned model would decline, while staying coherent on ordinary tasks. The abliteration here is huihui-ai's, not mine — this repo is the MLX conversion of their work.
This is an uncensored model. You are responsible for how you use it and for complying with the base model's license and applicable law.
Quick start
pip install mlx-lm
mlx_lm.generate --model divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx \
--prompt "Explain quantum entanglement to a 12 year old." --max-tokens 400
from mlx_lm import load, generate
model, tok = load("divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx")
text = tok.apply_chat_template(
[{"role": "user", "content": "Explain quantum entanglement to a 12 year old."}],
add_generation_prompt=True, tokenize=False)
print(generate(model, tok, prompt=text, max_tokens=400))
Also loads in LM Studio and anything else that reads MLX models.
Conversion details
- Quantization: 4-bit affine, group size 64, via
mlx_lm.convert - Format: MLX safetensors — no GGUF, no llama.cpp, no GPU needed
- Runs on: M1 / M2 / M3 / M4 / M5 Macs, entirely on-device
More abliterated MLX models
Part of the Abliterated MLX for Apple Silicon collection — Llama 3.3 70B, Gemma 4, Qwen3, Qwen3-VL, Hermes 4 and Muse Glimmer, all converted for Apple Silicon.
Credit
Base model and abliteration by huihui-ai. MLX 4-bit conversion by divinetribe.
Part of Claude Code Local
This model is one of the fighters in Claude Code Local (3.2k★), which runs Claude Code 100% on-device on Apple Silicon through an MLX-native Anthropic-API server. Not sure which local model to run as an agent? Check the Agent-12 local agent leaderboard: real agent tasks, judged by the filesystem, same hardware for every row.
Built by Matt Macosko in Arcata, CA. Open to work on local-AI and Apple Silicon inference: [email protected].
- Downloads last month
- 292
4-bit
Model tree for divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx
Base model
google/gemma-4-31B