WorkRB: A Community-Driven Evaluation Framework for AI in the Work Domain
Paper • 2604.13055 • Published
How to use Aleksandruz/skillmatch-mpnet-curriculum-retriever with Transformers:
# Load model directly
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("Aleksandruz/skillmatch-mpnet-curriculum-retriever")
model = AutoModel.from_pretrained("Aleksandruz/skillmatch-mpnet-curriculum-retriever", device_map="auto")How to use Aleksandruz/skillmatch-mpnet-curriculum-retriever with sentence-transformers:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("Aleksandruz/skillmatch-mpnet-curriculum-retriever")
sentences = [
"The weather is lovely today.",
"It's so sunny outside!",
"He drove to the stadium."
]
embeddings = model.encode(sentences)
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]This model is a fine-tuned version of sentence-transformers/all-mpnet-base-v2 for matching job description sentences to skill definitions.
It was originally introduced and detailed in the paper: [From Retrieval to Ranking: A Two-Stage Neural Framework for Automated Skill Extraction] (https://ceur-ws.org/Vol-4046/RecSysHR2025-paper_5.pdf)
Additionally, this model was evaluated as part of the work-domain AI benchmark in the paper: [WorkRB: A Community-Driven Evaluation Framework for AI in the Work Domain] (https://huggingface.co/papers/2604.13055)
from transformers import AutoTokenizer, AutoModel
import torch
import torch.nn.functional as F
model_id = "Aleksandruz/skillmatch-mpnet-curriculum-retriever"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id)
def encode(texts):
enc = tokenizer(
texts,
padding=True,
truncation=True,
return_tensors="pt"
)
with torch.no_grad():
out = model(**enc)
# mean pooling
attn = enc["attention_mask"].unsqueeze(-1)
emb = (out.last_hidden_state * attn).sum(1) / attn.sum(1)
return F.normalize(emb, p=2, dim=1)
jobs = ["Looking for a data scientist with NLP experience"]
skills = ["Machine learning", "Natural language processing", "Java programming"]
job_emb = encode(jobs)
skill_emb = encode(skills)
scores = job_emb @ skill_emb.T
print(scores)
Base model
sentence-transformers/all-mpnet-base-v2