What it is. Nitzotz reads a Hebrew message and answers questions you type about it: pick one of several options,
give a score on a scale, or say yes or no to a claim. For every answer it gives a probability you can trust, so you
know when it is sure and when it is guessing. It does not write text, so it cannot make things up. It runs on a normal
laptop, with no internet connection and no cost per question.
What it is for. Deciding what to do with incoming messages: is this a scam, what kind of message is it, which
department should get it, how urgent is it. It is not a chatbot and it is not built for long documents
(see Limitations).
In numbers. On 298 Hebrew messages it says correctly whether a message is a scam 91.3% of the
time. For context: 70% of those messages are not scams, so a model that always says "not a scam" would
score 70.5%. The number that matters more is the ranking: it gives real scams a higher probability than
legitimate messages 98% of the time (AUC 0.98).
What changed in version 1.1
The same model, with the same size, files and way of use. It was trained again with more data: reading questions,
the same questions in many different wordings, and harder scam examples (see Training data).
On the same frozen tests as version 1.0:
| Test |
Version 1.0 |
Version 1.1 |
Real difference? |
| Reading comprehension (Belebele, 900 questions) |
52.7% |
62.7% |
yes (p<0.001) |
| only the "which of these is NOT ..." questions (158) |
25.3% |
51.9% |
- |
| Scam or not (298 messages) |
82.5% |
91.3% |
yes (p<0.001) |
| hard cases only (65) |
63.1% |
80.0% |
yes (p=0.03) |
| Paying customer? (30 tickets) |
63.3% |
76.7% |
not clear (p=0.12) |
| Wording changes the answer: share of questions whose answer is not the same in all 6 wordings (1,786 questions, lower is better) |
9.8% |
4.4% |
- |
- Reading. The biggest change. Questions of the form "which of these is NOT said in the passage" were at the level
of a blind guess (25.3%); now 51.9%.
- Scams. Better overall and on the hard cases (scams written to look legitimate, and real messages that look like
scams). The scam threshold was chosen again on held-out training messages: 0.515 (version 1.0: 0.517).
- Wording. Asking the same question in other words now changes the answer much less often. In version 1.0
one everyday wording of the paying-customer question made it answer "yes" for all 30 tickets; now no wording
gives one answer to every ticket. The urgency question is still unstable (see Limitations).
- No test got clearly worse (paired exact McNemar against version 1.0 on the same questions: no drop with p < 0.05).
Try it
Python (the laya library, pip install laya):
import laya
agent = laya.load("BrainboxAI/nitzotz")
q = {"scam": {"type": "noul",
"instructions": "ההודעה מנסה לגרום לנמען ללחוץ על קישור, לשלם או למסור פרטים בלי סיבה לגיטימית.",
"criteria": {"true": "כן, זה ניסיון מרמה",
"false": "לא, זו הודעה לגיטימית, גם אם יש בה קישור או בקשת תשלום"}}}
print(agent.predict("החבילה שלך מעוכבת. לשחרור שלם 12.90 בקישור", q)["answers"]["scam"]["noul"])
Output: 0.8638, the probability that the claim ("the message tries to make the reader click a link, pay or hand
over details for no legitimate reason") is true.
The wording of the question matters. This is the exact question the scam test used, and the numbers on this card
are for it. In our tries a shorter wording ("the message is a scam attempt") gave clearly worse answers, for example
a high scam probability for a plain verification code. If you change the wording, check it on your own messages first.
laya.exe (the standalone binary from ggmlc releases, no Python
needed). Download the Q8 file, start the daemon, then send one JSON request per line; each answer comes back on one line:
hf download BrainboxAI/nitzotz nitzotz-q8_0.gguf --local-dir .
laya.exe daemon nitzotz-q8_0.gguf --device vulkan
{"id":"1","state":"היי, הפגישה מחר ב-10 עדיין בתוקף?","questions":{"scam":{"type":"noul","instructions":"ההודעה מנסה לגרום לנמען ללחוץ על קישור, לשלם או למסור פרטים בלי סיבה לגיטימית.","criteria":{"true":"כן, זה ניסיון מרמה","false":"לא, זו הודעה לגיטימית, גם אם יש בה קישור או בקשת תשלום"}}}}
{"status":"ready","model":"laya"}
{"model": "nitzotz", "family": "nitzotz", "route": "forced nitzotz", "answers": {"scam": {"type": "noul", "action": {"act_probability": 0.6666}, "confidence": 0.9708, "noul": 0.0292}}, "usage": {"input_tokens": 66, "output_tokens": 0, "latency_ms": 239.876}, "id": "1"}
noul is the probability that the claim is true. Several questions in one request are answered together in one pass.
Use --device cpu on a machine without a GPU.
Benchmarks
8 frozen test sets, 3,990 questions in total, locked by checksum before any training data existed. All three
models answered exactly the same questions. The two reference points are the other open laya models that read Hebrew:
RoeiG/laya-hebrew (licence CC-BY-NC-SA-4.0, non-commercial use only)
and laya-multilingual (Apache-2.0).
Full table. Accuracy, the best in each row in bold. The last two columns say whether Nitzotz's difference from that
model is real or could be luck (a paired exact McNemar test on the same questions): "better" or "worse" means
p < 0.05, "tie" means the difference could be chance.
| Test (questions) |
Nitzotz |
RoeiG (non-commercial) |
laya-multilingual |
Chance |
Nitzotz vs RoeiG |
Nitzotz vs laya-multilingual |
| Scam or not? (298 messages) |
91.3% |
32.2% |
50.3% |
50.0% |
better p<0.001 |
better p<0.001 |
| same, hard cases only (65) |
80.0% |
32.3% |
35.4% |
50.0% |
better p<0.001 |
better p<0.001 |
| Message type, 6 options (298) |
76.2% |
52.7% |
20.8% |
16.7% |
better p<0.001 |
better p<0.001 |
| same, hard cases only (65) |
63.1% |
44.6% |
24.6% |
16.7% |
tie p=0.05 |
better p<0.001 |
| Support ticket type, 5 options (30) |
86.7% |
80.0% |
56.7% |
20.0% |
tie p=0.69 |
better p=0.02 |
| Ticket urgency, 5 levels (30) |
80.0% |
36.7% |
33.3% |
20.0% |
better p=0.002 |
better p=0.003 |
| Paying customer? yes/no (30) |
76.7% |
56.7% |
50.0% |
50.0% |
better p=0.03 |
tie p=0.12 |
| Voice command intent, 20 options (500) |
90.0% |
72.6% |
47.4% |
5.0% |
better p<0.001 |
better p<0.001 |
| Voice command intent, 4 options (500) |
97.0% |
90.8% |
69.4% |
25.0% |
better p<0.001 |
better p<0.001 |
| News topic, 7 options (204) |
80.4% |
82.3% |
66.2% |
14.3% |
tie p=0.61 |
better p<0.001 |
| Does the passage support this answer? (600) |
94.8% |
94.2% |
48.8% |
50.0% |
tie p=0.70 |
better p<0.001 |
| Plausible answer the passage does not give (600) |
90.7% |
53.0% |
53.7% |
50.0% |
better p<0.001 |
better p<0.001 |
| Reading comprehension, 4 options (900) |
62.7% |
75.4% |
31.4% |
25.0% |
worse p<0.001 |
better p<0.001 |
On Belebele reading comprehension Nitzotz scores 62.7%, below RoeiG/laya-hebrew (75.4%); Nitzotz is built for
message decisions, not long-passage comprehension.
How to read it:
- MASSIVE (voice commands): Nitzotz was trained on MASSIVE's training commands. The test commands are different,
but written by the same people in the same style, so this test is easier for Nitzotz than for the others.
- HeQ: Nitzotz was trained on HeQ's training passages. The test passages are different ones, and training
passages that overlapped a test passage were removed.
- The support-ticket rows have only 30 questions. No difference there is reliable.
- The spam and scam test has a lean: 70% of its messages are not scams. The "chance" line (50%) is a
coin flip, not the best blind strategy.
Why you can trust it
The probabilities mean something. Take all the answers where Nitzotz said it was about 70 to 80% sure, and count
how many were right: the chart above does that for every confidence level, over 3,990 test questions.
Where the dots sit above the line, Nitzotz is more often right than it claims (it is modest); where they sit below, it
is over-confident. This is what lets you set thresholds (see the next section). The per-question-type temperatures
were fitted on 4,550 held-out training items, never on the test sets (calibration.json).
It reads the text. With the message removed and only the question left, its accuracy falls to
70.5% on the scam question and 7.7% on message type. So the answers come from the message, not
from the wording of the question.
It is fast on ordinary hardware. On one laptop (Intel Core Ultra 9 285H laptop, built-in Arc 140T GPU, Windows 11), one question at a time:
| Runtime |
Median over 50 check questions (about 135 tokens) |
Short message (38 tokens) |
Long input (430 tokens) |
| laya.exe, Q8 file, GPU (Vulkan) |
50 ms |
38 ms |
138 ms |
| laya.exe, F16 file, GPU (Vulkan) |
52 ms |
44 ms |
137 ms |
| Python (laya), GPU (PyTorch XPU) |
70 ms |
44 ms |
192 ms |
| laya.exe, Q8 file, CPU only (16 threads) |
340 ms |
186 ms |
1270 ms |
| Python (laya), CPU only |
342 ms |
134 ms |
1236 ms |
The short and long columns repeat one fixed input 30 times after 5 warm-up calls. Timings on this laptop change a
lot from one session to another (an earlier measurement of the same setup was several times slower), so treat these
as rough.
The GGUF files give almost the same answers as the Python model. Compared with the full-precision Python model on the
CPU:
| File, device |
Same top answer, 50 questions |
Largest probability gap |
Same top answer, 596 spam questions |
Largest gap |
Average gap |
| Q8, GPU |
50/50 |
0.0456 |
595/596 |
0.0189 |
0.00116 |
| F16, GPU |
50/50 |
0.0031 |
596/596 |
0.0023 |
0.00023 |
| Q8, CPU |
50/50 |
0.0563 |
595/596 |
0.0278 |
0.00185 |
| F16, CPU |
50/50 |
0.0015 |
596/596 |
0.0017 |
0.00020 |
A gap of 0.01 means, for example, 0.83 against 0.84. The few questions where the top answer changes are ones where
the two best answers were almost tied. Over both spam questions the Q8 file on the GPU is right 83.6% of the time,
against 83.7% for the Python model; the F16 file is closer (83.7%).
Use it in your business
The idea ("Ramzor", traffic light). Every incoming message gets one or more quick questions. The probability
decides what happens next:
- Green (confident it is fine): handle automatically, file it, tag it, route it.
- Yellow (not sure): send it to a person, or to a large language model if you use one.
- Red (confident it is a scam): block or quarantine it.
The drawing is a worked example on the 298 test messages. The upper threshold, 0.52, is the one chosen for the
scam question on 1000 held-out training messages (never on the test); the lower one, 0.35, was picked by hand. With them,
207 messages go to green (11 of them are in fact scams), 5 to yellow (3 scams), and 86 to red (12 of them are
in fact legitimate). So red should mean "quarantine and check", not "delete". Pick your own thresholds on a sample of
your own messages, and decide how many mistakes in green and red you can live with. At the 0.52 threshold alone,
the scam answer is right 91.3% of the time overall and 80.0% on the 65 hard cases (at 0.50: 91.3% and 80.0%).
The model is cheap enough to run on every message. The person (or the LLM) only sees the yellow part. Do not use it
as the only line of defence for decisions that can hurt someone.
Training data and transparency
Two training stages, on a rented GPU:
- Learning to read. HalleluBERT-large was first trained to find the answer to a question inside a passage, on
27,085 HeQ training questions (CC BY 4.0). Passages that overlapped a test passage were removed.
- Learning to decide. A laya decision head was put on top and the whole model was trained on 106,723 items
(102,173 for training, 4,550 held out to pick the best of 2 passes and to fit the temperatures).
| Source |
Items |
Licence |
How it was made |
| Synthetic: Message type, 6 classes |
14,000 |
teacher output, project-owned (DeepSeek MIT; Gemma Apache-2.0) |
written by DeepSeek V4.1 Flash, labelled independently by both models |
| Synthetic: Routing to a department (3 to 6 options) |
8,750 |
teacher output, project-owned (DeepSeek MIT; Gemma Apache-2.0) |
written by DeepSeek V4.1 Flash, labelled independently by both models |
| Synthetic: Yes/no claims about a message |
8,750 |
teacher output, project-owned (DeepSeek MIT; Gemma Apache-2.0) |
written by DeepSeek V4.1 Flash, labelled independently by both models |
| Synthetic: Urgency, 5 levels |
3,500 |
teacher output, project-owned (DeepSeek MIT; Gemma Apache-2.0) |
written by DeepSeek V4.1 Flash, labelled independently by both models |
| HeQ: is the proposed answer supported (yes/no) |
6,000 |
CC BY 4.0 |
the dataset's gold labels, train split only |
| HeQ: plausible answer to an unanswerable question |
9,750 (4,750 of them repeated copies, new in 1.1) |
CC BY 4.0 |
the dataset's gold labels, train split only |
| HeQ: 4-option reading |
5,000 |
CC BY 4.0 |
the dataset's gold labels, train split only |
| MASSIVE he-IL: voice command intent |
7,000 |
CC BY 4.0 |
the dataset's gold labels, train split only |
| New in 1.1: scam or not, disguised scams and real messages that look suspicious |
6,000 |
model output, project-owned (DeepSeek MIT; Gemma Apache-2.0) |
written by DeepSeek V4.1 Flash from scenario outlines written with GPT; kept only if both labellers agreed with the writer |
| New in 1.1: message type of the same messages |
5,374 |
model output, project-owned |
the type both labellers agreed on |
| New in 1.1: copies of existing items with the question reworded |
21,199 |
as the original item; the wordings were written with GPT |
question or options in one of 376 alternative wordings written with GPT; the label comes from the original item |
| New in 1.1: reading, 4 options ("which is NOT", reworded answers, several sentences) |
6,825 |
passages: FineWeb-2, ODC-By 1.0; questions: model output, project-owned |
written by DeepSeek V4.1 Flash, checked by Gemma 4 31B with the passage and again without it; dropped if it could be answered without the passage |
| New in 1.1: topic of a passage, 7 options |
4,000 |
passages: FineWeb-2, ODC-By 1.0; labels: model output |
labelled by DeepSeek V4.1 Flash and Gemma 4 31B |
| New in 1.1: is the sender an existing paying customer (yes/no) |
575 |
model output, project-owned; question wordings written with GPT |
support messages written by DeepSeek V4.1 Flash, checked by Gemma 4 31B |
- Synthetic messages (35,000 items). DeepSeek V4.1 Flash (MIT) wrote Israeli-style SMS, WhatsApp and email messages
from a plan (intended label, topic, tone, varied fake phone numbers and links). DeepSeek and Gemma 4 31B
(Apache-2.0) then each labelled every message on their own. A message was kept only if both agreed; the two models
agreed on 94% of them, and 93.7% of the 37,429 written messages were kept. Both ran
through DeepInfra (via OpenRouter), with data retention off and "thinking" mode off. The synthetic data is not
published.
- New in version 1.1. 6,825 reading questions on Hebrew web passages from FineWeb-2 (ODC-By 1.0): "which of
these is NOT true or NOT mentioned" (each paired with a positive question on the same passage), answers said in
other words, and answers that need several sentences. DeepSeek V4.1 Flash wrote them and Gemma 4 31B checked each one
twice, with the passage and without it; a question that could be answered without the passage was dropped.
6,000 messages that are hard to tell apart (disguised scams, and real messages that look suspicious), written by
DeepSeek and kept only if both labellers agreed with each other and with the writer. 4,000 passages labelled
with a topic, 575 support messages for the paying-customer question, and 4,750 repeated HeQ items so that the
older skills are not lost.
- Written with GPT. Part of the training data was written with OpenAI's GPT, through our ChatGPT subscription: 376 alternative wordings of the questions and 34 sets of alternative answer options, used in 30,488 training items (including all 575 paying-customer items), and 120 short scenario outlines from which DeepSeek V4.1 Flash wrote 6,000 messages. GPT wrote no message and no label. In total 33,148 of the 106,723 training items (31%) use text written with GPT. GPT output is covered by OpenAI's terms of use, not by an open licence.
- Open data. HeQ v1.1 (CC BY 4.0; about half of its questions are on Geektime articles, shared by the HeQ authors under
the same licence) and MASSIVE he-IL (CC BY 4.0), training splits only. FineWeb-2 Hebrew (ODC-By 1.0) passages for
the new reading and topic items.
- No leaks from the tests. Every training item was compared with every test question; anything sharing an
8-word run with a test text was dropped, and so were MASSIVE commands equal to a test command. The new items were
checked with stricter rules: every new passage against every test text (any shared 6-word run), and every new
question, option and wording against every test question and every reworded test question (exact match and
character similarity).
- Not used: no output of any other closed commercial chatbot, no DICTA model, no non-commercial or share-alike
data. GPT, a closed commercial model, was used only as described above.
- Size: encoder 357.1M parameters (HalleluBERT-large, MIT, fine-tuned); decision head 26.5M parameters,
trained from scratch.
About the tests.
| Test |
Questions |
Source and licence |
| Spam and business messages |
596 (298 messages, 2 questions each) |
in-house, Israeli SMS, WhatsApp and email style, 65 hard cases |
| Support tickets (triage30) |
90 (30 tickets, 3 questions each) |
in-house |
| MASSIVE he-IL, 20 and 4 options |
500 + 500 |
MASSIVE test split, CC BY 4.0 |
| SIB-200 news topic |
204 |
CC BY-SA 4.0, used for testing only |
| Belebele reading |
900 |
CC BY-SA 4.0, used for testing only |
| HeQ verify and unanswerable |
600 + 600 |
HeQ v1.1 test split, CC BY 4.0 |
Caveats that change how much to trust the numbers:
- The spam and business test was written by an AI model and checked by an AI model, not by a person. Real inboxes
will look different. It was frozen after that review (4 labels changed, 2 messages removed).
- Three training runs, three random seeds; this is one of them. It was picked on held-out training items (95.8% against 94.9% and 95.3%), not on the tests. The three runs are close on the large tests: scam or not 91.3% here against 91.3% and 90.9%, news topic 80.4% against 80.9% and 82.3%, reading 62.7% against 63.1% and 63.4%. This run and each of the other two give the same answer on 90.6% to 91.0% of all test questions. On the 30-ticket tests they differ more: ticket urgency 80.0% here against 66.7% and 66.7%.
- The training labels come from two AI models. Where both are wrong in the same way, Nitzotz learned their mistake.
- HeQ's wrong answers in the test were picked by code, not checked by a person.
Limitations
- Reading comprehension of longer passages is better, but still limited. Belebele: 62.7%, where a blind guess
gets 25%, and still below the best other laya model on this test (see Benchmarks). When the right
answer is written word for word in the passage it gets 79% (252 questions); when the answer is said in other
words it gets 56% (648 questions); on "which of these is NOT" questions 52% (158 questions). Do not ask it
whether a long document supports a claim.
- Checking an answer against a passage (HeQ) works. 94.8% with the passage; with the passage removed it falls to 50.3%, a coin flip. So on this kind of question it really reads the passage.
- Numbers, dates, amounts and rules: not trained and not measured. Compute them in code and pass the result in.
- Hard cases are still the weak spot: scams written to look legitimate (a "supplier" changing bank details, the
"CEO" asking for a transfer) and real messages that look like scams (a real bank alert with a link, a real
verification code). On the 65 hard cases the scam question is right 80.0% of the time (always answering "not a
scam" there would give 70.8%), with AUC 0.88. At the chosen threshold it still misses 5 of the 19 scams written to look
legitimate, and flags 8 of the 46 real messages that look like scams.
- Urgency is subjective, and its answer depends on the wording. Even the two teacher models matched the intended
urgency only about 62 to 64% of the time. When the urgency question is asked in other words, the answer changes on
43% of the 30 test tickets.
- The in-house tests and much of the training data were written by AI models, not by people. This includes the
spam and business test, the support-ticket test and the new reading questions. Real messages will look different.
- The support-ticket tests are small: 30 tickets per question, so one ticket moves a score by 3.3 points.
- Sarcasm and irony are probably read literally. Not measured.
- 512 tokens (roughly 300 to 400 Hebrew words) per question. A longer message is cut from the end without a warning.
- Hebrew only. Not trained or tested on English or Arabic.
- Not a safety system on its own. It makes mistakes in both directions. Keep a person in the loop for anything
that can hurt someone.
Licence and attribution
Apache-2.0 for the weights, the GGUF files and the code. Commercial use is allowed. Built on:
- HalleluBERT-large: the encoder, MIT.
- laya (NandhaKishorM, Convai Innovations): the decision-head architecture and the
runtime, Apache-2.0. No laya weights are used.
- HeQ: CC BY 4.0, by Webiks for MAFAT and the Israeli
National NLP Program (NNLP-IL); includes Geektime passages.
- MASSIVE: CC BY 4.0, Amazon (FitzGerald et al., 2022).
- DeepSeek V4.1 Flash (MIT) and Gemma 4 31B (Apache-2.0), used through DeepInfra as data writer and labellers.
- FineWeb-2 (Hebrew): ODC-By 1.0, Hugging Face; the passages of the new
reading and topic items.
- OpenAI GPT, through a ChatGPT subscription (OpenAI terms of use): question wordings and scenario outlines, as
described in the training data section.
- ggmlc for the GGUF files.
- Belebele and SIB-200 (CC BY-SA 4.0) were used only to test, never to train.
The full notice is in NOTICE.
## ניצוץ, בעברית


**מה זה.** ניצוץ קורא הודעה בעברית ועונה על שאלות שאתם מקלידים עליה: לבחור אחת מכמה אפשרויות, לתת ציון בסולם, או
לענות כן או לא על טענה. על כל תשובה הוא נותן הסתברות שאפשר לסמוך עליה, כך שיודעים מתי הוא בטוח ומתי הוא מנחש. הוא לא
כותב טקסט, ולכן הוא לא יכול להמציא דברים. הוא רץ על מחשב נייד רגיל, בלי אינטרנט ובלי תשלום על כל שאלה.
**בשביל מה.** להחליט מה עושים עם הודעות נכנסות: האם זו הונאה, איזה סוג הודעה זו, לאיזו מחלקה להעביר, כמה זה דחוף. זה
לא צ'אטבוט, והוא לא בנוי למסמכים ארוכים (ראו [מגבלות](#מגבלות)).
**במספרים.** על 298 הודעות בעברית הוא קובע נכון אם ההודעה היא הונאה ב-91.3% מהמקרים. בשביל
פרופורציה: 70% מההודעות האלה הן לא הונאה, כך שמודל שתמיד עונה "לא הונאה" היה מקבל 70.5%. המספר
שחשוב יותר הוא הדירוג: הוא נותן להונאה אמיתית הסתברות גבוהה יותר מאשר להודעה תקינה ב-98% מהמקרים (AUC 0.98).
### מה השתנה בגרסה 1.1
אותו מודל, באותו גודל, אותם קבצים ואותה דרך שימוש. הוא אומן מחדש עם יותר נתונים: שאלות קריאה, אותן שאלות בהרבה
ניסוחים שונים, ודוגמאות הונאה קשות יותר (ראו [נתוני האימון](#נתוני-האימון-ושקיפות)). על אותם מבחנים קפואים כמו בגרסה 1.0:
| מבחן | גרסה 1.0 | גרסה 1.1 | ההבדל אמיתי? |
|---|---|---|---|
| הבנת הנקרא (Belebele, 900 שאלות) | 52.7% | **62.7%** | כן (p<0.001) |
| רק שאלות "איזו מהבאות לא ..." (158) | 25.3% | **51.9%** | - |
| הונאה או לא (298 הודעות) | 82.5% | **91.3%** | כן (p<0.001) |
| רק המקרים הקשים (65) | 63.1% | **80.0%** | כן (p=0.03) |
| לקוח משלם? (30 פניות) | 63.3% | **76.7%** | לא ברור (p=0.12) |
| הניסוח משנה את התשובה: אחוז השאלות שהתשובה עליהן לא זהה בכל 6 הניסוחים (1,786 שאלות, נמוך = טוב) | 9.8% | **4.4%** | - |
- **קריאה.** השינוי הגדול ביותר. שאלות מהסוג "איזו מהבאות לא נאמרה בקטע" היו בגובה ניחוש עיוור (25.3%). עכשיו
51.9%.
- **הונאות.** טוב יותר בסך הכול וגם במקרים הקשים (הונאות שנכתבו כדי להיראות לגיטימיות, והודעות אמיתיות שנראות כמו הונאה).
סף ההונאה נבחר מחדש על הודעות אימון שהופרדו מראש: 0.515 (בגרסה 1.0: 0.517).
- **ניסוח.** כשמנסחים את אותה שאלה במילים אחרות, התשובה משתנה הרבה פחות. בגרסה 1.0 ניסוח יומיומי אחד של שאלת
הלקוח המשלם גרם לו לענות "כן" על כל 30 הפניות. עכשיו אף ניסוח לא נותן תשובה אחת לכל הפניות. שאלת
הדחיפות עדיין לא יציבה (ראו [מגבלות](#מגבלות)).
- אף מבחן לא נהיה גרוע יותר באופן ברור (McNemar מדויק מול גרסה 1.0 על אותן שאלות: אין ירידה עם p קטן מ-0.05).
### לנסות
**בפייתון** (הספרייה [laya](https://github.com/NandhaKishorM/laya), `pip install laya`):
import laya
agent = laya.load("BrainboxAI/nitzotz")
q = {"scam": {"type": "noul",
"instructions": "ההודעה מנסה לגרום לנמען ללחוץ על קישור, לשלם או למסור פרטים בלי סיבה לגיטימית.",
"criteria": {"true": "כן, זה ניסיון מרמה",
"false": "לא, זו הודעה לגיטימית, גם אם יש בה קישור או בקשת תשלום"}}}
print(agent.predict("החבילה שלך מעוכבת. לשחרור שלם 12.90 בקישור", q)["answers"]["scam"]["noul"])
הפלט: `0.8638`, ההסתברות שהטענה ("ההודעה מנסה לגרום לנמען ללחוץ על קישור, לשלם או למסור פרטים בלי סיבה לגיטימית")
נכונה.
**הניסוח של השאלה משנה.** זו בדיוק השאלה שבה השתמש מבחן ההונאות, והמספרים בכרטיס הזה הם עליה. בניסיונות שלנו ניסוח
קצר יותר ("ההודעה היא ניסיון הונאה") נתן תשובות גרועות בהרבה, למשל הסתברות גבוהה להונאה לקוד אימות רגיל. אם משנים את
הניסוח, בודקים אותו קודם על ההודעות שלכם.
**בלי פייתון, עם laya.exe** (תוכנה עצמאית מ-[ggmlc](https://github.com/monatis/ggmlc/releases/latest)). מורידים את קובץ
Q8, מפעילים, ושולחים בקשת JSON אחת בכל שורה. כל תשובה חוזרת בשורה אחת:
hf download BrainboxAI/nitzotz nitzotz-q8_0.gguf --local-dir .
laya.exe daemon nitzotz-q8_0.gguf --device vulkan
{"id":"1","state":"היי, הפגישה מחר ב-10 עדיין בתוקף?","questions":{"scam":{"type":"noul","instructions":"ההודעה מנסה לגרום לנמען ללחוץ על קישור, לשלם או למסור פרטים בלי סיבה לגיטימית.","criteria":{"true":"כן, זה ניסיון מרמה","false":"לא, זו הודעה לגיטימית, גם אם יש בה קישור או בקשת תשלום"}}}}
{"status":"ready","model":"laya"}
{"model": "nitzotz", "family": "nitzotz", "route": "forced nitzotz", "answers": {"scam": {"type": "noul", "action": {"act_probability": 0.6666}, "confidence": 0.9708, "noul": 0.0292}}, "usage": {"input_tokens": 66, "output_tokens": 0, "latency_ms": 239.876}, "id": "1"}
`noul` היא ההסתברות שהטענה נכונה. כמה שאלות בבקשה אחת נענות יחד, במעבר אחד. על מחשב בלי כרטיס מסך משתמשים ב-`--device cpu`.
### מבחנים
8 סטים של מבחן, 3,990 שאלות בסך הכול, שננעלו בטביעת אצבע לפני שנוצר פריט אימון אחד. שלושת המודלים ענו על
אותן שאלות בדיוק. שתי נקודות ההשוואה הן מודלי laya הפתוחים האחרים שקוראים עברית:
[RoeiG/laya-hebrew](https://huggingface.co/RoeiG/laya-hebrew) (רישיון CC-BY-NC-SA-4.0, **לשימוש לא מסחרי בלבד**)
ו-[laya-multilingual](https://huggingface.co/convaiinnovations/laya) (Apache-2.0).


**הטבלה המלאה.** אחוז התשובות הנכונות, הטוב ביותר בכל שורה מודגש. שתי העמודות האחרונות אומרות אם ההבדל של ניצוץ
מהמודל הזה אמיתי או שאולי זה מזל (מבחן McNemar מדויק על אותן שאלות בדיוק): "טוב יותר" או "חלש יותר" פירושו p
קטן מ-0.05, "תיקו" פירושו שההבדל יכול להיות מקרי.
| מבחן (מספר שאלות) | ניצוץ | RoeiG (לא מסחרי) | laya-multilingual | ניחוש | ניצוץ מול RoeiG | ניצוץ מול laya-multilingual |
|---|---|---|---|---|---|---|
| הונאה או לא? (298 הודעות) | **91.3%** | 32.2% | 50.3% | 50.0% | **טוב יותר**
p<0.001 | **טוב יותר**
p<0.001 |
| אותו דבר, רק המקרים הקשים (65) | **80.0%** | 32.3% | 35.4% | 50.0% | **טוב יותר**
p<0.001 | **טוב יותר**
p<0.001 |
| סוג ההודעה, 6 אפשרויות (298) | **76.2%** | 52.7% | 20.8% | 16.7% | **טוב יותר**
p<0.001 | **טוב יותר**
p<0.001 |
| אותו דבר, רק המקרים הקשים (65) | **63.1%** | 44.6% | 24.6% | 16.7% | תיקו
p=0.05 | **טוב יותר**
p<0.001 |
| סוג פניית תמיכה, 5 אפשרויות (30) | **86.7%** | 80.0% | 56.7% | 20.0% | תיקו
p=0.69 | **טוב יותר**
p=0.02 |
| דחיפות הפנייה, 5 רמות (30) | **80.0%** | 36.7% | 33.3% | 20.0% | **טוב יותר**
p=0.002 | **טוב יותר**
p=0.003 |
| לקוח משלם? כן/לא (30) | **76.7%** | 56.7% | 50.0% | 50.0% | **טוב יותר**
p=0.03 | תיקו
p=0.12 |
| כוונת פקודה קולית, 20 אפשרויות (500) | **90.0%** | 72.6% | 47.4% | 5.0% | **טוב יותר**
p<0.001 | **טוב יותר**
p<0.001 |
| כוונת פקודה קולית, 4 אפשרויות (500) | **97.0%** | 90.8% | 69.4% | 25.0% | **טוב יותר**
p<0.001 | **טוב יותר**
p<0.001 |
| נושא של ידיעה, 7 אפשרויות (204) | 80.4% | **82.3%** | 66.2% | 14.3% | תיקו
p=0.61 | **טוב יותר**
p<0.001 |
| האם הקטע תומך בתשובה? (600) | **94.8%** | 94.2% | 48.8% | 50.0% | תיקו
p=0.70 | **טוב יותר**
p<0.001 |
| תשובה סבירה שהקטע לא נותן (600) | **90.7%** | 53.0% | 53.7% | 50.0% | **טוב יותר**
p<0.001 | **טוב יותר**
p<0.001 |
| הבנת הנקרא, 4 אפשרויות (900) | 62.7% | **75.4%** | 31.4% | 25.0% | **חלש יותר**
p<0.001 | **טוב יותר**
p<0.001 |
בהבנת הנקרא של Belebele ניצוץ מקבל 62.7%, פחות מ-RoeiG/laya-hebrew (75.4%). ניצוץ בנוי להחלטות על הודעות, לא
להבנה של קטעים ארוכים.
איך לקרוא את זה:
- **MASSIVE** (פקודות קוליות): ניצוץ אומן על פקודות האימון של MASSIVE. פקודות המבחן אחרות, אבל נכתבו בידי אותם
אנשים ובאותו סגנון, ולכן המבחן הזה קל יותר לניצוץ מאשר לאחרים.
- **HeQ**: ניצוץ אומן על קטעי האימון של HeQ. קטעי המבחן אחרים, וקטעי אימון שחפפו לקטע מבחן הוסרו.
- **בשורות של פניות התמיכה יש רק 30 שאלות.** שום הבדל שם לא אמין.
- **מבחן הספאם וההונאות לא מאוזן**: 70% מההודעות בו הן לא הונאה. קו ה"ניחוש" (50%) הוא הטלת מטבע, לא
האסטרטגיה העיוורת הטובה ביותר.
### למה אפשר לסמוך עליו

**להסתברויות יש משמעות.** קחו את כל התשובות שבהן ניצוץ אמר שהוא בטוח בערך ב-70 עד 80%, וספרו כמה מהן היו נכונות.
הגרף עושה את זה לכל רמת ביטחון, על 3,990 שאלות מבחן. כשהנקודות מעל הקו, ניצוץ צודק יותר ממה שהוא אומר (הוא
צנוע). כשהן מתחת, הוא בטוח בעצמו יותר מדי. זה מה שמאפשר לקבוע ספים (בפרק הבא). הכיול נעשה על 4,550 פריטי אימון
שהופרדו מראש, אף פעם לא על המבחן (`calibration.json`).
**הוא באמת קורא את הטקסט.** כשמוחקים את ההודעה ומשאירים רק את השאלה, הדיוק יורד ל-70.5% בשאלת ההונאה
ול-7.7% בסוג ההודעה. כלומר התשובות באות מההודעה, לא מהניסוח של השאלה.

**הוא מהיר על חומרה רגילה.** על מחשב נייד עם מעבד Intel Core Ultra 9 285H וכרטיס המסך המובנה Arc 140T, ווינדוס 11, שאלה אחת בכל פעם:
| איך מריצים | חציון על 50 שאלות בדיקה (בממוצע 135 טוקנים) | הודעה קצרה (38 טוקנים) | קלט ארוך (430 טוקנים) |
|---|---|---|---|
| laya.exe, קובץ Q8, כרטיס מסך (Vulkan) | 50 ms | 38 ms | 138 ms |
| laya.exe, קובץ F16, כרטיס מסך (Vulkan) | 52 ms | 44 ms | 137 ms |
| פייתון (laya), כרטיס מסך (PyTorch XPU) | 70 ms | 44 ms | 192 ms |
| laya.exe, קובץ Q8, מעבד בלבד (16 תהליכונים) | 340 ms | 186 ms | 1270 ms |
| פייתון (laya), מעבד בלבד | 342 ms | 134 ms | 1236 ms |
בעמודות של ההודעה הקצרה והקלט הארוך אותה שאלה רצה 30 פעמים, אחרי 5 הרצות חימום. הזמנים על המחשב הזה משתנים הרבה בין
הפעלה להפעלה (מדידה קודמת של אותה הגדרה יצאה איטית פי כמה), אז אלה מספרים בקירוב.
**קובצי ה-GGUF נותנים כמעט את אותן תשובות כמו מודל הפייתון.** בהשוואה למודל הפייתון המלא על המעבד:
| קובץ, מכשיר | אותה תשובה מובילה, 50 שאלות | הפרש הסתברות מרבי | אותה תשובה מובילה, 596 שאלות ספאם | הפרש מרבי | הפרש ממוצע |
|---|---|---|---|---|---|
| Q8, כרטיס מסך | 50/50 | 0.0456 | 595/596 | 0.0189 | 0.00116 |
| F16, כרטיס מסך | 50/50 | 0.0031 | 596/596 | 0.0023 | 0.00023 |
| Q8, מעבד | 50/50 | 0.0563 | 595/596 | 0.0278 | 0.00185 |
| F16, מעבד | 50/50 | 0.0015 | 596/596 | 0.0017 | 0.00020 |
פער של 0.01 פירושו, למשל, 0.83 מול 0.84. השאלות המעטות שבהן התשובה המובילה משתנה הן כאלה שבהן שתי התשובות הטובות
היו כמעט שוות. בשתי שאלות הספאם יחד קובץ Q8 על כרטיס המסך צודק ב-83.6%, מול 83.7% למודל הפייתון. קובץ F16
קרוב יותר (83.7%).
### שימוש בעסק

**הרעיון ("רמזור").** כל הודעה נכנסת מקבלת שאלה מהירה אחת או כמה. ההסתברות מחליטה מה קורה הלאה:
- **ירוק** (בטוח שזה בסדר): טיפול אוטומטי, תיוק, תיוג, ניתוב.
- **צהוב** (לא בטוח): לבדיקה של אדם, או של מודל שפה גדול אם אתם משתמשים בו.
- **אדום** (בטוח שזו הונאה): חסימה או הסגר.
השרטוט הוא דוגמה על 298 הודעות המבחן. הסף העליון, 0.52, הוא הסף שנבחר לשאלת ההונאה על 1000 הודעות אימון
שהופרדו מראש (אף פעם לא על המבחן). הסף התחתון, 0.35, נבחר ביד. איתם 207 הודעות הולכות לירוק (11 מהן הן בעצם הונאה),
5 לצהוב (3 הונאות), ו-86 לאדום (12 מהן בעצם תקינות). כלומר אדום צריך להיות "הסגר ובדיקה", לא "מחיקה".
בחרו ספים משלכם על מדגם של ההודעות שלכם, והחליטו כמה טעויות בירוק ובאדום אתם מוכנים לקבל. בסף 0.52 לבדו, תשובת
ההונאה נכונה ב-91.3% מהמקרים בסך הכול וב-80.0% על 65 המקרים הקשים (בסף 0.50: 91.3% ו-80.0%).
המודל זול מספיק כדי להריץ אותו על כל הודעה. האדם (או מודל השפה) רואה רק את החלק הצהוב. **אל תשתמשו בו כקו הגנה
יחיד בהחלטות שיכולות לפגוע במישהו.**
### נתוני האימון ושקיפות
שני שלבי אימון, על כרטיס מסך שכור:
1. **ללמוד לקרוא.** HalleluBERT-large אומן קודם למצוא את התשובה לשאלה בתוך קטע, על 27,085 שאלות אימון של HeQ
(CC BY 4.0). קטעים שחפפו לקטע מבחן הוסרו.
2. **ללמוד להחליט.** מעליו הונח ראש החלטות של laya, וכל המודל אומן על 106,723 פריטים (102,173 לאימון,
ו-4,550 הופרדו מראש כדי לבחור את הטוב מבין 2 מעברים ולכייל את הטמפרטורות).
| מקור | פריטים | רישיון | איך נוצר |
|---|---|---|---|
| סינתטי: סוג הודעה, 6 סוגים | 14,000 | הפלט של המורה שייך לנו (DeepSeek, MIT; Gemma, Apache-2.0) | נכתב בידי DeepSeek V4.1 Flash, סומן בנפרד בידי שני המודלים |
| סינתטי: ניתוב למחלקה (3 עד 6 אפשרויות) | 8,750 | הפלט של המורה שייך לנו (DeepSeek, MIT; Gemma, Apache-2.0) | נכתב בידי DeepSeek V4.1 Flash, סומן בנפרד בידי שני המודלים |
| סינתטי: טענות כן/לא על הודעה | 8,750 | הפלט של המורה שייך לנו (DeepSeek, MIT; Gemma, Apache-2.0) | נכתב בידי DeepSeek V4.1 Flash, סומן בנפרד בידי שני המודלים |
| סינתטי: דחיפות, 5 רמות | 3,500 | הפלט של המורה שייך לנו (DeepSeek, MIT; Gemma, Apache-2.0) | נכתב בידי DeepSeek V4.1 Flash, סומן בנפרד בידי שני המודלים |
| HeQ: האם התשובה המוצעת נתמכת (כן/לא) | 6,000 | CC BY 4.0 | תוויות הזהב של המאגר, מפיצול האימון בלבד |
| HeQ: תשובה סבירה לשאלה שאין לה תשובה | 9,750 (מתוכם 4,750 עותקים חוזרים, חדש ב-1.1) | CC BY 4.0 | תוויות הזהב של המאגר, מפיצול האימון בלבד |
| HeQ: קריאה, 4 אפשרויות | 5,000 | CC BY 4.0 | תוויות הזהב של המאגר, מפיצול האימון בלבד |
| MASSIVE he-IL: כוונת פקודה קולית | 7,000 | CC BY 4.0 | תוויות הזהב של המאגר, מפיצול האימון בלבד |
| חדש ב-1.1: הונאה או לא, הונאות מוסוות והודעות אמיתיות שנראות חשודות | 6,000 | פלט מודלים, שייך לנו (DeepSeek, MIT; Gemma, Apache-2.0) | נכתב בידי DeepSeek V4.1 Flash לפי תרחישים שנכתבו עם GPT. נשמר רק אם שני המסמנים הסכימו עם הכותב |
| חדש ב-1.1: סוג ההודעה של אותן הודעות | 5,374 | פלט מודלים, שייך לנו | הסוג ששני המסמנים הסכימו עליו |
| חדש ב-1.1: עותקים של פריטים קיימים עם השאלה בניסוח אחר | 21,199 | כמו הפריט המקורי. הניסוחים נכתבו עם GPT | השאלה או האפשרויות באחד מ-376 ניסוחים חלופיים שנכתבו עם GPT. התווית לקוחה מהפריט המקורי |
| חדש ב-1.1: קריאה, 4 אפשרויות ("איזו מהבאות לא", תשובה במילים אחרות, כמה משפטים) | 6,825 | קטעים: FineWeb-2, ODC-By 1.0. שאלות: פלט מודלים, שייך לנו | נכתב בידי DeepSeek V4.1 Flash ונבדק בידי Gemma 4 31B עם הקטע ושוב בלי הקטע. נזרק אם אפשר היה לענות בלי הקטע |
| חדש ב-1.1: נושא של קטע, 7 אפשרויות | 4,000 | קטעים: FineWeb-2, ODC-By 1.0. תוויות: פלט מודלים | סומן בידי DeepSeek V4.1 Flash ו-Gemma 4 31B |
| חדש ב-1.1: האם השולח לקוח משלם קיים (כן/לא) | 575 | פלט מודלים, שייך לנו. ניסוחי השאלה נכתבו עם GPT | פניות תמיכה שנכתבו בידי DeepSeek V4.1 Flash ונבדקו בידי Gemma 4 31B |
- **הודעות סינתטיות (35,000 פריטים).** DeepSeek V4.1 Flash (MIT) כתב הודעות SMS, וואטסאפ ומייל בסגנון ישראלי
לפי תוכנית (התווית המתוכננת, נושא, משלב, מספרי טלפון וקישורים מזויפים ומגוונים). אחר כך DeepSeek ו-Gemma 4 31B
(Apache-2.0) סימנו כל הודעה, כל אחד לבד. הודעה נשמרה רק אם שניהם הסכימו. הם הסכימו על 94% מההודעות,
ונשמרו 93.7% מתוך 37,429 ההודעות שנכתבו. שניהם רצו דרך DeepInfra (דרך OpenRouter), בלי שמירת נתונים
ובלי מצב "חשיבה". הנתונים הסינתטיים לא מפורסמים.
- **חדש בגרסה 1.1.** 6,825 שאלות קריאה על קטעים מהרשת בעברית מתוך FineWeb-2 (ODC-By 1.0): "איזו מהבאות לא
נכונה או לא מוזכרת" (כל אחת עם שאלה חיובית צמודה על אותו קטע), תשובות שנאמרות במילים אחרות, ותשובות שדורשות כמה
משפטים. DeepSeek V4.1 Flash כתב אותן, ו-Gemma 4 31B בדק כל אחת פעמיים, עם הקטע ובלי הקטע. שאלה שאפשר היה לענות עליה בלי
הקטע נזרקה. 6,000 הודעות שקשה להבחין ביניהן (הונאות מוסוות, והודעות אמיתיות שנראות חשודות), שנכתבו בידי
DeepSeek ונשמרו רק אם שני המסמנים הסכימו זה עם זה וגם עם הכותב. 4,000 קטעים שסומנו לפי נושא, 575 פניות
תמיכה לשאלת הלקוח המשלם, ו-4,750 פריטי HeQ חוזרים, כדי שהיכולות הקודמות לא יישכחו.
- **נכתב עם GPT.** חלק מנתוני האימון נכתב עם GPT של OpenAI, דרך מנוי ChatGPT שלנו: 376 ניסוחים חלופיים של השאלות ו-34 סטים חלופיים של אפשרויות תשובה, שמופיעים ב-30,488 פריטי אימון (כולל כל 575 הפריטים של שאלת הלקוח המשלם), ו-120 תרחישים קצרים שלפיהם DeepSeek V4.1 Flash כתב 6,000 הודעות. GPT לא כתב אף הודעה ואף תווית. בסך הכול 33,148 מתוך 106,723 פריטי האימון (31%) משתמשים בטקסט שנכתב עם GPT. הפלט של GPT כפוף לתנאי השימוש של OpenAI, לא לרישיון פתוח.
- **נתונים פתוחים.** HeQ v1.1 (CC BY 4.0. בערך חצי מהשאלות שלו על כתבות של Geektime, ומחברי HeQ משתפים אותם באותו רישיון) ו-MASSIVE
he-IL (CC BY 4.0), רק מפיצולי האימון. קטעים בעברית מ-FineWeb-2 (ODC-By 1.0) לשאלות הקריאה והנושא החדשות.
- **בלי דליפה מהמבחנים.** כל פריט אימון הושווה לכל שאלת מבחן. כל מה שחולק רצף של 8 מילים עם טקסט מבחן נזרק, וגם
פקודות MASSIVE שזהות לפקודת מבחן. הפריטים החדשים נבדקו בכללים מחמירים יותר: כל קטע חדש מול כל טקסט מבחן (כל רצף משותף
של 6 מילים), וכל שאלה, אפשרות וניסוח חדשים מול כל שאלת מבחן וכל ניסוח חלופי של שאלת מבחן (התאמה מדויקת ודמיון תווים).
- **מה לא שימש:** שום פלט של צ'אטבוט מסחרי סגור אחר, שום מודל של DICTA, שום נתונים לא מסחריים או ברישיון "שיתוף זהה".
GPT, מודל מסחרי סגור, שימש רק כמתואר למעלה.
- **גודל:** המקודד 357.1 מיליון פרמטרים (HalleluBERT-large, MIT, אומן מחדש). ראש ההחלטות 26.5 מיליון פרמטרים,
אומן מאפס.
**על המבחנים.**
| מבחן | שאלות | מקור ורישיון |
|---|---|---|
| הודעות ספאם ועסקים | 596 (298 הודעות, 2 שאלות לכל אחת) | פנימי, בסגנון SMS, וואטסאפ ומייל ישראלי, 65 מקרים קשים |
| פניות תמיכה (triage30) | 90 (30 פניות, 3 שאלות לכל אחת) | פנימי |
| MASSIVE he-IL, 20 ו-4 אפשרויות | 500 + 500 | פיצול המבחן של MASSIVE, CC BY 4.0 |
| SIB-200, נושא ידיעה | 204 | CC BY-SA 4.0, רק לבדיקה |
| Belebele, הבנת הנקרא | 900 | CC BY-SA 4.0, רק לבדיקה |
| HeQ, אימות ו"אין תשובה" | 600 + 600 | פיצול המבחן של HeQ v1.1, CC BY 4.0 |
הסתייגויות שמשנות כמה לסמוך על המספרים:
- **מבחן הספאם והעסקים נכתב בידי מודל AI ונבדק בידי מודל AI, לא בידי אדם.** תיבות דואר אמיתיות ייראו אחרת. הוא ננעל
אחרי הבדיקה הזו (4 תוויות שונו, 2 הודעות הוסרו).
- **שלוש ריצות אימון, עם שלושה זרעים אקראיים, וזו אחת מהן.** היא נבחרה לפי פריטי אימון שהופרדו מראש (95.8% מול 94.9% ו-95.3%), לא לפי המבחנים. במבחנים הגדולים שלוש הריצות קרובות: הונאה או לא 91.3% כאן מול 91.3% ו-90.9%, נושא ידיעה 80.4% מול 80.9% ו-82.3%, הבנת הנקרא 62.7% מול 63.1% ו-63.4%. הריצה הזו וכל אחת מהשתיים האחרות נותנות אותה תשובה על 90.6% עד 91.0% מכל שאלות המבחן. במבחנים של 30 פניות ההבדלים גדולים יותר: דחיפות הפנייה 80.0% כאן מול 66.7% ו-66.7%.
- **תוויות האימון באות משני מודלי AI.** איפה ששניהם טועים באותו אופן, ניצוץ למד את הטעות שלהם.
- התשובות השגויות של HeQ במבחן נבחרו בקוד, ולא נבדקו בידי אדם.
### מגבלות
- **הבנת הנקרא של קטעים ארוכים השתפרה, אבל עדיין מוגבלת.** Belebele: 62.7%, כשניחוש עיוור מקבל 25%, ועדיין פחות ממודל
laya האחר הטוב ביותר במבחן הזה (ראו [מבחנים](#מבחנים)). כשהתשובה הנכונה כתובה בקטע מילה במילה הוא מקבל 79% (252 שאלות).
כשהתשובה נאמרת במילים אחרות הוא מקבל 56% (648 שאלות). בשאלות "איזו מהבאות לא" 52% (158 שאלות).
אל תשאלו אותו אם מסמך ארוך תומך בטענה.
- **בדיקה אם קטע תומך בתשובה (HeQ) עובדת.** 94.8% עם הקטע. כשמוחקים את הקטע זה יורד ל-50.3%, הטלת מטבע. כלומר בשאלות מהסוג הזה הוא באמת קורא את הקטע.
- **מספרים, תאריכים, סכומים וכללים**: לא אומן עליהם ולא נמדד. חשבו אותם בקוד והעבירו את התוצאה.
- **המקרים הקשים עדיין נקודת התורפה**: הונאות שנכתבו כדי להיראות לגיטימיות ("ספק" שמחליף פרטי חשבון בנק, "המנכ"ל"
שמבקש העברה), והודעות אמיתיות שנראות כמו הונאה (התראה אמיתית מהבנק עם קישור, קוד אימות אמיתי). על 65 המקרים הקשים
שאלת ההונאה צודקת ב-80.0% מהמקרים (תשובה קבועה "לא הונאה" הייתה מקבלת שם 70.8%), עם AUC 0.88. בסף שנבחר הוא
עדיין מפספס 5 מתוך 19 הונאות שנכתבו כדי להיראות לגיטימיות, ומתריע על 8 מתוך 46 הודעות אמיתיות שנראות כמו הונאה.
- **דחיפות היא עניין סובייקטיבי, והתשובה עליה תלויה בניסוח.** אפילו שני מודלי המורה הסכימו עם הדחיפות המתוכננת רק בכ-62
עד 64% מהמקרים. כששואלים את שאלת הדחיפות במילים אחרות, התשובה משתנה ב-43% מתוך 30 פניות המבחן.
- **המבחנים הפנימיים וחלק גדול מנתוני האימון נכתבו בידי מודלי AI, לא בידי אנשים.** זה כולל את מבחן הספאם והעסקים, את
מבחן פניות התמיכה ואת שאלות הקריאה החדשות. הודעות אמיתיות ייראו אחרת.
- **מבחני פניות התמיכה קטנים**: 30 פניות לכל שאלה, כך שפנייה אחת מזיזה ציון ב-3.3 נקודות.
- **ציניות ואירוניה** כנראה נקראות מילולית. לא נמדד.
- **512 טוקנים** (בערך 300 עד 400 מילים בעברית) לשאלה. הודעה ארוכה יותר נחתכת מהסוף בלי אזהרה.
- **עברית בלבד.** לא אומן ולא נבדק על אנגלית או ערבית.
- **זו לא מערכת הגנה לבד.** הוא טועה לשני הכיוונים. השאירו אדם בתהליך בכל דבר שיכול לפגוע במישהו.
### רישיון וקרדיטים
**Apache-2.0** למשקולות, לקובצי ה-GGUF ולקוד. מותר לשימוש מסחרי. בנוי על:
- [HalleluBERT-large](https://huggingface.co/HalleluBERT/HalleluBERT_large): המקודד, MIT.
- [laya](https://github.com/NandhaKishorM/laya) (NandhaKishorM, Convai Innovations): הארכיטקטורה של ראש ההחלטות וסביבת
ההרצה, Apache-2.0. לא נעשה שימוש במשקולות של laya.
- [HeQ](https://github.com/NNLP-IL/Hebrew-Question-Answering-Dataset): CC BY 4.0, של Webiks עבור מפא"ת ותוכנית ה-NLP
הלאומית (NNLP-IL). כולל קטעים מ-Geektime.
- [MASSIVE](https://huggingface.co/datasets/AmazonScience/massive): CC BY 4.0, אמזון (FitzGerald ואחרים, 2022).
- **DeepSeek V4.1 Flash** (MIT) ו-**Gemma 4 31B** (Apache-2.0), דרך DeepInfra, ככותב הנתונים וכמסמנים.
- [FineWeb-2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-2) (עברית): ODC-By 1.0, של Hugging Face. הקטעים של שאלות
הקריאה והנושא החדשות.
- **GPT של OpenAI**, דרך מנוי ChatGPT (תנאי השימוש של OpenAI): ניסוחי שאלות ותרחישים, כמתואר בפרק נתוני האימון.
- [ggmlc](https://github.com/monatis/ggmlc) לקובצי ה-GGUF.
- Belebele ו-SIB-200 (CC BY-SA 4.0) שימשו רק לבדיקה, אף פעם לא לאימון.
ההודעה המלאה בקובץ `NOTICE`.