LAMB (LAtin ModernBERT) is a Latin encoder-only model based on the ModernBERT architecture, pre-trained on nearly 24B Latin tokens, and ready for use with any Latin orthography. If you use this in your work, please cite: Paper pending shortly.
Independent publisher
Aidan
aimgo
Models
Emendator is a byt5-xl model finetuned to correct OCR artifacts in Latin text. This model cannot provide completely faithful reconstruction for all orthographies - on a large scale, it will shift the distribution of tokens towards that which it has been trained on. This is to say: Emendator will take editorial liberties with your data. As such, use it only in circumstances when the primary concern is only to recover intelligible Latin, not to recover intelligible Latin of a particular style. The model is intended to be used on segments of 250 characters. Anything else will compromise performance. Original: "atque optimo viro, peterem; superavi tamen dignitate Catilinam, gratia Galbam. Quod…
CaputEmendatoris is a projection head for Emendator trained to identify OCR artifacts in Latin text at a character level. You can use it to quantify the amount of damage to a sample or guide Emendator. The model is intended to be used on segments of 250 characters. Anything else will compromise performance. In initial testing, using 0.25 as a character probability threshold typically produced the best F1 score across all degrees of corruption. Orig: Cognoscenda virtute circumscripta est scientia, quae ad experientiam pertinet et ad rationem. OCR: C0gn0fccndauirtutccircurnfcriptacftfcientia:quacadcxpcricntiarnpcrtinct&adrationcrn« To use CaputEmendatoris, you can load it via the Transformers…