SAVRN
Search Contact SAVRN

Open-weight model · Text generation

Llama-3.1-8B-Instruct

by Meta Llama meta-llama/Llama-3.1-8B-Instruct

The Meta Llama 3.1 collection of multilingual large language models (LLMs) is a collection of pretrained and instruction tuned generative models in 8B, 70B and 405B sizes (text in/text out).

Parameters8B
Context
Weights32.1 GB
Licensellama3.1
AccessAccess requested at publisher
Monthly Downloads5.9M

Runs On

What it takes to serve Llama-3.1-8B-Instruct (8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 16.1 GB 19.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 8.0 GB 9.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 4.0 GB 4.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on Llama-3.1-8B-Instruct

At 16-bit this checkpoint needs 19.3 GB of memory, and the cheapest SAVRN Index listing that covers it is a single 192 GB MI300X at $1.85 per hour on-demand. That card is the cheapest answer at 8-bit (9.6 GB) and 4-bit (4.8 GB) too, so quantizing does not buy a cheaper hour; it buys room for more concurrent multilingual dialogue sessions on one card.

Commercial use is allowed under the Meta Llama 3.1 Community License, with attribution and Meta's Acceptable Use Policy observed, unless your products had more than 700 million monthly active users on the release date; then you request a license from Meta. Gated access means approval precedes download. Two checks: the page lists no context length, and the $1.85 hour must beat the Index's hosted rates, $0.02 in and $0.05 out per million tokens at DeepInfra and Novita, $0.06 both ways at Nscale.

Model Card

The Meta Llama 3.1 collection of multilingual large language models (LLMs) is a collection of pretrained and instruction tuned generative models in 8B, 70B and 405B sizes (text in/text out). The Llama 3.1 instruction tuned text only models (8B, 70B, 405B) are optimized for multilingual dialogue use cases and outperform many of the available open source and closed chat models on common industry benchmarks. Model Architecture: Llama 3.1 is an auto-regressive language model that uses an optimized transformer architecture. The tuned versions use supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to align with human preferences for helpfulness and safety.…

Excerpt from the card by Meta Llama, licensed llama3.1.

Identity and Version

Repository
meta-llama/Llama-3.1-8B-Instruct
Publisher
Meta Llama
Task
Text generation
Modality
Text
Library
transformers
Parameters
8B parameters
Languages
en, de, fr, it, pt, hi, es, th
Revision
0e9e39f249a16976918f6564b8830bc894c89659
First published
2024-07-18
Last updated
2024-09-25

Files and Weights

17 files, 32.1 GB in total. The weights are 5 files totalling 32.1 GB in pth, safetensors.

Weights5 files · 32.1 GB
Configuration5 files · 25.5 KB
Tokenizer3 files · 11.3 MB
Documentation3 files · 56.4 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00004.safetensorsWeights5.0 GB
model-00002-of-00004.safetensorsWeights5.0 GB
model-00003-of-00004.safetensorsWeights4.9 GB
model-00004-of-00004.safetensorsWeights1.2 GB
original/consolidated.00.pthWeights16.1 GB
config.jsonConfiguration855 B
generation_config.jsonConfiguration184 B
model.safetensors.index.jsonConfiguration23.9 KB
original/params.jsonConfiguration199 B
special_tokens_map.jsonConfiguration296 B
LICENSEDocumentation7.6 KB
README.mdDocumentation44.0 KB
USE_POLICY.mdDocumentation4.7 KB
.gitattributesRepository1.5 KB
original/tokenizer.modelTokenizer2.2 MB
tokenizer.jsonTokenizer9.1 MB
tokenizer_config.jsonTokenizer55.4 KB

License and Download

License
llama3.1
Access
Access requested at publisher
Download size
32.1 GB
Download from Meta Llama

Released by Meta Llama through Meta's Llama downloads.

Built From

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
Idavidrein/gpqa Task diamondMetric diamondComparison conditions not established 30.4 Model Card
Reported by a third party
Evaluated revision not stated 2026-01-27
LEXam-Benchmark/LEXam Task mcq_4_choicesMetric mcq_4_choicesComparison conditions not established 24.04 LEXam Leaderboard
Reported by a third party
Evaluated revision not stated 2026-06-02
LEXam-Benchmark/LEXam Task open_questionMetric open_questionComparison conditions not established 10 LEXam Leaderboard
Reported by a third party
Evaluated revision not stated 2026-06-02
openai/gsm8k Task gsm8kMetric gsm8kComparison conditions not established 84.5 Model Card
Reported by a third party
Evaluated revision not stated 2026-03-23
thamilvendhan/signalbench Task access_denyMetric access_denySetup family=access_deny; n=12Comparison conditions not established 0.5 thamilvendhan
Reported by a third party
Evaluated revision not stated 2026-07-08
thamilvendhan/signalbench Task bot_policyMetric bot_policySetup family=bot_policy; n=12Comparison conditions not established 0.4167 thamilvendhan
Reported by a third party
Evaluated revision not stated 2026-07-08
thamilvendhan/signalbench Task injectionMetric injectionSetup family=injection; n=12Comparison conditions not established 0.75 thamilvendhan
Reported by a third party
Evaluated revision not stated 2026-07-08
thamilvendhan/signalbench Task memory_labelMetric memory_labelSetup family=memory_label; n=12Comparison conditions not established 0.6667 thamilvendhan
Reported by a third party
Evaluated revision not stated 2026-07-08
thamilvendhan/signalbench Task srcMetric srcSetup SRC overall; deterministic action-based grader, no LLM judge; seed 0, n=75Comparison conditions not established 0.6333 signalbench raw per-item responses
Reported by a third party
Evaluated revision not stated 2026-07-08
thamilvendhan/signalbench Task timeMetric timeSetup family=time; n=12Comparison conditions not established 0.8333 thamilvendhan
Reported by a third party
Evaluated revision not stated 2026-07-08

Memory Requirements

PrecisionWeights in memory
As published32.1 GB
16-bit16.1 GB
8-bit8.0 GB
4-bit4.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Hosted Prices

HostInput / outputUnitObserved
DeepInfra$0.02 / $0.05input / output, per million tokensSep 18, 2026
Novita$0.02 / $0.05input / output, per million tokensSep 18, 2026
Nscale$0.06 / $0.06input / output, per million tokensSep 18, 2026

From the SAVRN Index.

Built on This Model

Compare Llama-3.1-8B-Instruct

Questions About Llama-3.1-8B-Instruct

How much GPU memory does Llama-3.1-8B-Instruct need?

About 19.3 GB at 16-bit and 4.8 GB at 4-bit: the weights (8B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Llama-3.1-8B-Instruct on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Llama-3.1-8B-Instruct commercially?

Yes, with conditions. Llama-3.1-8B-Instruct is released under Meta Llama 3.1 Community License. The Llama 3.1 Community License permits commercial use, except that a licensee whose products had more than 700 million monthly active users on the release date must request a license from Meta. It requires attribution as the license specifies and compliance with Meta's Acceptable Use Policy.

Similar Models

Model · Text generation

Meta-Llama-3-8B-Instruct

Meta Llama

Meta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8 and 70B sizes. The Llama 3 instruction tuned models are optimized for dialogue use cases and outperform many of the available open source chat models on common industry benchmarks. Further, in developing these models, we took great care to optimize helpfulness and safety. Model developers Meta Variations Llama 3 comes in two sizes — 8B and 70B parameters — in pre-trained and instruction tuned variants. Input Models input text only. Output Models generate text and code only. Model Architecture Llama 3 is an auto-regressive language…

Access requested at publisher llama3 8B parameters transformers

Model · Text generation

Llama-3.1-8B-Instruct-4bit

MLX Community

The Model mlx-community/Llama-3.1-8B-Instruct-4bit was converted to MLX format from meta-llama/Llama-3.1-8B-Instruct using mlx-lm version 0.21.4.

Open weights llama3.1 8B parameters 131,072 tokens mlx

This is the 8B (high-capacity flagship) member of the Med-LLaMA3 family introduced in the paper “Med-LLaMA3: Advancing Medical Question-Answering Through Parameter-Efficient Fine-Tuning of Large Language Models” (Applied Sciences, 2026). The family adapts the LLaMA-3 architecture to the medical domain by training only a small fraction of the base model’s parameters (4.01% for this 8B variant), achieving strong medical question-answering performance while keeping the memory footprint low — enabling development and inference on low-cost, consumer-grade hardware. The 8B variant is the high-capacity model for complex clinical reasoning. It attains a mean accuracy of 75.71% across the eight MMLU…

Open weights llama3.1 8B parameters 131,072 tokens transformers

A Llama-3-8B-Instruct checkpoint compressed with SVD-LLM to 50.0% of dense parameters, then edited by 4 of 10 rounds of iterative parameter-neutral swap selected by the gapiter rule (up to 0.1% of dense parameters per round; the full run's budget is 1.0%). This is a research artifact from a study of how SVD compression damages safety behaviour and which component-selection rule best repairs it. It is one cell of a grid over selection rules and budgets; it is not a general-purpose chat model. This checkpoint exists to measure safety/utility trade-offs under compression. Several arms in the grid are deliberately safety-degraded relative to Llama-3-8B-Instruct: compression alone raises…

Open weights llama3 8B parameters 8,192 tokens transformers

A Llama-3-8B-Instruct checkpoint compressed with SVD-LLM to 50.0% of dense parameters, then edited by 3 of 10 rounds of iterative parameter-neutral swap selected by the gapiter rule (up to 0.1% of dense parameters per round; the full run's budget is 1.0%). This is a research artifact from a study of how SVD compression damages safety behaviour and which component-selection rule best repairs it. It is one cell of a grid over selection rules and budgets; it is not a general-purpose chat model. This checkpoint exists to measure safety/utility trade-offs under compression. Several arms in the grid are deliberately safety-degraded relative to Llama-3-8B-Instruct: compression alone raises…

Open weights llama3 8B parameters 8,192 tokens transformers

A Llama-3-8B-Instruct checkpoint compressed with SVD-LLM to 50.0% of dense parameters, then edited by 2 of 10 rounds of iterative parameter-neutral swap selected by the gapiter rule (up to 0.1% of dense parameters per round; the full run's budget is 1.0%). This is a research artifact from a study of how SVD compression damages safety behaviour and which component-selection rule best repairs it. It is one cell of a grid over selection rules and budgets; it is not a general-purpose chat model. This checkpoint exists to measure safety/utility trade-offs under compression. Several arms in the grid are deliberately safety-degraded relative to Llama-3-8B-Instruct: compression alone raises…

Open weights llama3 8B parameters 8,192 tokens transformers