Instructions to use sensix-zo/sensix-paite-4b-cpt-16bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Local Apps Settings
- Unsloth Desktop
Sensix Paite 4B CPT Base (16-bit)
This model represents Phase 1 of the Sensix Paite AI development pipeline. It is a Continued Pre-Training (CPT) model built on top of the Gemma 3 4B-IT architecture.
The primary objective of this model is vocabulary acquisition and syntax alignment for the Paite language, establishing a native-level linguistic foundation before advanced instruction or conversational fine-tuning.
Dataset Strategy: Avoiding "Preacher Bias"
A critical aspect of training regional languages is data quality. Many low-resource language datasets rely heavily on translated religious texts (like the Bible) or highly fragmented short sentences.
To prevent this model from developing a repetitive "robotic" tone or archaic "Preacher Bias," it was trained exclusively on the PERFECT_PAITE_DATA.jsonl dataset.
- Included Data: Modern news articles, contemporary essays, parallel dictionaries, folksongs, and long-form conversational paragraphs.
- Excluded Data: The Paite Bible and isolated short-sentence fragments were strictly removed to ensure the model learns modern, fluid, and logical conversational reasoning.
Training Parameters
The model was trained using Unsloth on a high-performance GPU environment with aggressive learning rates to force the base model to map new vocabulary effectively without lobotomizing its native English reasoning.
- Base Model: unsloth/gemma-3-4b-it
- Learning Rate: 2e-4 (Aggressive for vocab acquisition)
- LoRA Config: r=64, alpha=64
- Target Modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
- Epochs: 1
- Context Length: 2048 tokens
Technical Implementation: The Hard Merge
Standard Unsloth merging functions (save_pretrained_merged) have been known to cause weight-scrambling bugs in Gemma 3 architectures, often resulting in gibberish outputs (the "Calcium/Blades" bug).
To guarantee complete stability, this model was fused using an Official Hard Merge (model.merge_and_unload()).
- The LoRA adapter weights are physically baked into the base model.
- The model is saved in full 16-bit precision (bfloat16).
- It operates flawlessly in standard inference environments without logic collapse.
Usage
While this model is a "CPT Base," it inherits the underlying instruction chat template of the Gemma 3 IT model. It can be used for text completion, translation tasks, or as the starting point for further Supervised Fine-Tuning (SFT).
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "sensix-zo/sensix-paite-4b-cpt-16bit"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
# Example usage for text generation
messages = [
{"role": "user", "content": "Paite pau in thulim khat hon gelh in."}
]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=512, do_sample=True, temperature=0.7)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Limitations
Because this model has only completed Phase 1 (CPT), its conversational personality is not yet fully aligned. For natural, multi-turn dialogue, users are encouraged to use the downstream SFT or "Message-Format" versions of the Sensix Paite AI models.
- Downloads last month
- -