FinBERT is a pre-trained NLP model to analyze sentiment of financial text. It is built by further training the BERT language model in the finance domain, using a large financial corpus and thereby fine-tuning it for financial sentiment classification. Financial PhraseBank by Malo et al. (2014) is used for fine-tuning. For more details, please see the paper FinBERT: Financial Sentiment Analysis with Pre-trained Language Models and our related blog post on Medium. The model will give softmax outputs for three labels: positive, negative or neutral. About Prosus Prosus is a global consumer internet group and one of the largest technology investors in the world. Operating and investing globally…
Open-weight model · Text classification
SourceCodeAuthorCheck-SLM-10M
by Anthony Assi assix-research/SourceCodeAuthorCheck-SLM-10M
SourceCodeAuthorCheck-SLM-10M is an open-weight model for text classification from Anthony Assi, released under MIT License. Its published files total 29.5 MB.
A ~10 million parameter Small Language Model (SLM) Transformer designed for binary classification to detect whether Python source code was written by a human or generated by AI.
Model Card
By Anthony Assi, published under mit, revision ba10a5902d9a.
A ~10 million parameter Small Language Model (SLM) Transformer designed for binary classification to detect whether Python source code was written by a human or generated by AI. The model was trained on a dataset of human-written code extracted from GitHub (Q3 2017 and prior) and synthetic AI-generated code mimicking Q3 2026 generative AI paradigms. You can test the model interactively without writing any code by visiting our Gradio Web UI Space. To use this model locally, you must include the model architecture class in your script before loading the weights. Once the class is defined, you can automatically pull the weights from Hugging Face and run inference cleanly using this helper…
Read Anthony Assi's full model card
A ~10 million parameter Small Language Model (SLM) Transformer designed for binary classification to detect whether Python source code was written by a human or generated by AI.
The model was trained on a dataset of human-written code extracted from GitHub (Q3 2017 and prior) and synthetic AI-generated code mimicking Q3 2026 generative AI paradigms.
Live Demo
You can test the model interactively without writing any code by visiting our Gradio Web UI Space.
Usage
To use this model locally, you must include the model architecture class in your script before loading the weights.
1. Define the Architecture
import torch
import torch.nn as nn
from transformers import AutoTokenizer
from huggingface_hub import hf_hub_download
class SourceCodeAuthorCheck(nn.Module):
def __init__(self, vocab_size=50257, d_model=128, nhead=8, num_layers=4, dim_feedforward=512):
super().__init__()
self.embedding = nn.Embedding(vocab_size, d_model)
self.pos_encoder = nn.Parameter(torch.zeros(1, 1024, d_model))
encoder_layers = nn.TransformerEncoderLayer(
d_model=d_model, nhead=nhead, dim_feedforward=dim_feedforward, batch_first=True
)
self.transformer = nn.TransformerEncoder(encoder_layers, num_layers=num_layers)
self.fc = nn.Linear(d_model, 1)
def forward(self, input_ids, attention_mask):
seq_len = input_ids.size(1)
x = self.embedding(input_ids) + self.pos_encoder[:, :seq_len, :]
src_key_padding_mask = ~attention_mask.bool()
x = self.transformer(x, src_key_padding_mask=src_key_padding_mask)
mask_expanded = attention_mask.unsqueeze(-1).float()
sum_embeddings = torch.sum(x * mask_expanded, 1)
sum_mask = torch.clamp(mask_expanded.sum(1), min=1e-9)
pooled = sum_embeddings / sum_mask
return self.fc(pooled)
2. Load Weights and Run Inference
Once the class is defined, you can automatically pull the weights from Hugging Face and run inference cleanly using this helper function:
# Initialize device and tokenizer
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
tokenizer = AutoTokenizer.from_pretrained("gpt2")
tokenizer.pad_token = tokenizer.eos_token
# Download and load model weights from the Hugging Face Hub
model = SourceCodeAuthorCheck().to(device)
model_path = hf_hub_download(repo_id="assix-research/SourceCodeAuthorCheck-SLM-10M", filename="source_code_classifier.pth")
model.load_state_dict(torch.load(model_path, map_location=device, weights_only=True))
model.eval()
def predict_source(code_snippet):
"""Predicts whether a given Python snippet is human-written or AI-generated."""
inputs = tokenizer(
code_snippet,
return_tensors="pt",
truncation=True,
padding="max_length",
max_length=1024
).to(device)
with torch.no_grad():
if torch.cuda.is_available():
with torch.autocast(device_type='cuda', dtype=torch.bfloat16):
logits = model(inputs['input_ids'], inputs['attention_mask'])
else:
logits = model(inputs['input_ids'], inputs['attention_mask'])
prob = torch.sigmoid(logits).item()
verdict = "AI Generated" if prob > 0.5 else "Human Written"
print(f"Verdict: {verdict} (AI Probability: {prob:.1%})")
# --- Example Usage ---
sample_code = """
def calculate_factorial(n):
if n == 0:
return 1
return n * calculate_factorial(n-1)
"""
predict_source(sample_code)
Identity and Version
- Repository
- assix-research/SourceCodeAuthorCheck-SLM-10M
- Publisher
- Anthony Assi
- Task
- Text classification
- Modality
- Text
- Library
- Not stated by the source
- Parameters
- Not stated by the source
- Languages
- en
- Revision
- ba10a5902d9a06b1437a845ca110f3b5bb4865db
- First published
- 2026-10-02
- Last updated
- 2026-10-02
Files and Weights
4 files, 29.5 MB in total. The weights are 1 file totalling 29.4 MB in pth.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| source_code_classifier.pth | Weights | 29.4 MB | 1b4e8941a866 |
| inference.py | Configuration | 2.9 KB | — |
| README.md | Documentation | 4.0 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
License and Download
- License
- mit
- Access
- Open weights, no gate
- Download size
- 29.4 MB
Released by Anthony Assi through its official repository on Hugging Face. Read the license.
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 29.4 MB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About SourceCodeAuthorCheck-SLM-10M
Can I use SourceCodeAuthorCheck-SLM-10M commercially?
Yes. SourceCodeAuthorCheck-SLM-10M is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.
Similar Models
This is a RoBERTa-base model trained on ~124M tweets from January 2018 to December 2021, and finetuned for sentiment analysis with the TweetEval benchmark. The original Twitter-based RoBERTa model can be found here and the original reference paper is TweetEval. This model is suitable for English. 0 -> Negative; 1 -> Neutral; 2 -> Positive This sentiment analysis model has been integrated into TweetNLP. You can access the demo here.
FinBERT is a BERT model pre-trained on financial communication text. The purpose is to enhance financial NLP research and practice. It is trained on the following three financial communication corpus. The total corpora size is 4.9B tokens. More technical details on FinBERT: Click Link This released finbert-tone model is the FinBERT model fine-tuned on 10,000 manually annotated (positive, negative, neutral) sentences from analyst reports. This model achieves superior performance on financial tone analysis task. If you are simply interested in using FinBERT for financial tone analysis, give it a try. If you use the model in your academic work, please cite the following paper: Huang, Allen H.…
https://huggingface.co/cross-encoder/ms-marco-MiniLM-L-6-v2 with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).
This is a multilingual XLM-roBERTa-base model trained on ~198M tweets and finetuned for sentiment analysis. The sentiment fine-tuning was done on 8 languages (Ar, En, Fr, De, Hi, It, Sp, Pt) but it can be used for more languages (see paper for details). This model has been integrated into the TweetNLP library.
With this model, you can classify emotions in English text data. The model was trained on 6 diverse datasets (see Appendix below) and predicts Ekman's 6 basic emotions, plus a neutral class: 1) anger 2) disgust 3) fear 4) joy 5) neutral 6) sadness 7) surprise The model is a fine-tuned checkpoint of DistilRoBERTa-base. For a 'non-distilled' emotion model, please refer to the model card of the RoBERTa-large version. a) Run emotion model with 3 lines of code on single text example using Hugging Face's pipeline command on Google Colab: b) Run emotion model on multiple examples and full datasets (e.g.,.csv files) on Google Colab: Please reach out to [email protected] if you have any…
