A research-oriented Albef prototype targeting Multitask. The included small setup documents defaults and file formats without presenting unverified performance numbers. - The Python file contains the model and runnable example or training entry point. - config.json records the generated architecture settings. - trainingargs.json records the default experiment recipe. - model.safetensors is a valid initialization checkpoint for smoke tests; it is not presented as a trained benchmark checkpoint. - No benchmark score is claimed in this repository. The included configuration uses adamw with a cosine schedule. These are starting values in the script, not evidence of a completed run. For a…
Open weights
bsd-3-clause
49,600 parameters
128 tokens
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
4B parameters
262,144 tokens
Nebium-Small is a 117-million-parameter causal Transformer trained for autoregressive next-chess-move prediction over Lichess UCI move sequences. - Rotary Position Embeddings (RoPE) on attention query and key projections ($\theta = 10000$) - SwiGLU feed-forward transformation - RMSNorm pre-normalization - Causal mask with padding token masking - Byte-Pair Encoding (BPE) tokenizer trained on UCI move plies $$L(N, D) = 1.69 + \frac{406.4}{N^{0.34}} + \frac{410.7}{D^{0.28}}$$ MIT License.
Open weights
mit
pytorch
Quantized and FP16 GGUF format binaries for Nebium-Small (117M parameters). Designed for low-latency CPU and GPU execution with llama.cpp and Ollama. MIT License.
Open weights
mit
Open weights
31.6B parameters
262,144 tokens
mlx
Open weights
apache-2.0
This repository contains a working research note about Multimodal Generation. It organizes motivation, related work, a falsifiable hypothesis, and an evaluation plan. It is not presented as a completed paper or a release of trained models. - the scope of the research question and likely confounders - a proposed comparison with matched baselines - concrete evaluation context such as task-appropriate public benchmarks named in the main note - reproducibility checks, failure modes, and open questions - topic-relevant references Start with summary.md for the full note. Sections labeled as plans or hypotheses should not be interpreted as experimental results. If results are added later, they…
Open weights
cc-by-4.0
16,576 parameters
128 tokens
Open weights
R
Model · Image and text to text
Ray
Open weights
apache-2.0
35.1B parameters
transformers
Open weights
Open weights
Access requested at publisher
mit