You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Sensix Paite 4B CPT Base (16-bit)

This model represents Phase 1 of the Sensix Paite AI development pipeline. It is a Continued Pre-Training (CPT) model built on top of the Gemma 3 4B-IT architecture.

The primary objective of this model is vocabulary acquisition and syntax alignment for the Paite language, establishing a native-level linguistic foundation before advanced instruction or conversational fine-tuning.

Dataset Strategy: Avoiding "Preacher Bias"

A critical aspect of training regional languages is data quality. Many low-resource language datasets rely heavily on translated religious texts (like the Bible) or highly fragmented short sentences.

To prevent this model from developing a repetitive "robotic" tone or archaic "Preacher Bias," it was trained exclusively on the PERFECT_PAITE_DATA.jsonl dataset.

  • Included Data: Modern news articles, contemporary essays, parallel dictionaries, folksongs, and long-form conversational paragraphs.
  • Excluded Data: The Paite Bible and isolated short-sentence fragments were strictly removed to ensure the model learns modern, fluid, and logical conversational reasoning.

Training Parameters

The model was trained using Unsloth on a high-performance GPU environment with aggressive learning rates to force the base model to map new vocabulary effectively without lobotomizing its native English reasoning.

  • Base Model: unsloth/gemma-3-4b-it
  • Learning Rate: 2e-4 (Aggressive for vocab acquisition)
  • LoRA Config: r=64, alpha=64
  • Target Modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
  • Epochs: 1
  • Context Length: 2048 tokens

Technical Implementation: The Hard Merge

Standard Unsloth merging functions (save_pretrained_merged) have been known to cause weight-scrambling bugs in Gemma 3 architectures, often resulting in gibberish outputs (the "Calcium/Blades" bug).

To guarantee complete stability, this model was fused using an Official Hard Merge (model.merge_and_unload()).

  • The LoRA adapter weights are physically baked into the base model.
  • The model is saved in full 16-bit precision (bfloat16).
  • It operates flawlessly in standard inference environments without logic collapse.

Usage

While this model is a "CPT Base," it inherits the underlying instruction chat template of the Gemma 3 IT model. It can be used for text completion, translation tasks, or as the starting point for further Supervised Fine-Tuning (SFT).

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "sensix-zo/sensix-paite-4b-cpt-16bit"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

# Example usage for text generation
messages = [
    {"role": "user", "content": "Paite pau in thulim khat hon gelh in."}
]

inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=512, do_sample=True, temperature=0.7)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Limitations

Because this model has only completed Phase 1 (CPT), its conversational personality is not yet fully aligned. For natural, multi-turn dialogue, users are encouraged to use the downstream SFT or "Message-Format" versions of the Sensix Paite AI models.

Downloads last month
-
Safetensors
Model size
4B params
Tensor type
F32
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sensix-zo/sensix-paite-4b-cpt-16bit

Finetuned
(1117)
this model