SWE-chat · Dataset Card
SWE-chat: Dataset Card
Written by Social And Language Technology Lab, published under odc-by, revision cb99dee5e013, read 2026-10-03. Shown as written; SAVRN's own facts about this dataset are on its page.
SWE-chat: Coding Agent Interactions From Real Users in the Wild
- Paper: arxiv.org/abs/2604.20779
- Website: swe-chat.com
Dataset Summary
SWE-chat captures real-world AI coding sessions from developers using AI coding assistants (Claude Code, Codex, Gemini CLI, and others via the Entire.io CLI). Each session includes the full conversation transcript, tool calls, thinking traces, code changes, and attribution of human vs. agent-authored code.
Dataset Size
| Table | Rows |
|---|---|
repositories |
707 |
checkpoints |
54,918 |
sessions |
17,968 |
session_logs |
17,968 |
commits |
62,333 |
conversations |
11,607,572 (229,909 user prompts) |
subagent_tasks |
546 |
conversations_subagents |
8,673 |
skill_definitions |
3,523 |
skill_invocations |
12,243 |
context_events |
80,958 |
transcripts/* (raw log files) |
17,819 |
transcripts_compact/* (compacted log files) |
3,760 |
transcripts_subagent/* (raw subagent log files) |
83 |
Tables
| Table | Description | Key |
|---|---|---|
| conversations | All transcript entries: user prompts, assistant responses, thinking, tool calls, tool results, metadata events | turn_id |
| sessions | One row per coding session with metadata, token usage, and attribution | session_id |
| session_logs | Per-session transcript pointer (transcript_path) plus auxiliary text fields |
session_id |
| checkpoints | Checkpoint snapshots linking sessions to commits | checkpoint_pk |
| commits | Checkpoint-to-commit associations with diffs and attribution | (checkpoint_pk, commit_sha) |
| repositories | Repository metadata, settings, and GitHub API enrichment | repo_id |
| subagent_tasks | Captured subagent task records | task_pk |
| conversations_subagents | Turns from subagent transcripts, kept separate from conversations |
turn_pk |
| skill_definitions | Distinct observed skill bodies and unresolved definition stubs | definition_id |
| skill_invocations | Explicit and inferred skill-use signals with execution scope and invoker | invocation_id |
| context_events | Instruction files, slash commands, plugins, reviews, and harness context | context_event_id |
Loading the Data
from datasets import load_dataset
# Load individual tables (each table is exposed as a named config)
conversations = load_dataset("SALT-NLP/SWE-chat", "conversations", split="train")
sessions = load_dataset("SALT-NLP/SWE-chat", "sessions", split="train")
commits = load_dataset("SALT-NLP/SWE-chat", "commits", split="train")
# Or load with pandas directly
import pandas as pd
df_conv = pd.read_parquet("conversations.parquet")
df_sessions = pd.read_parquet("sessions.parquet")
# Raw transcripts live under transcripts/{session_id}.jsonl
# session_logs.parquet holds the relative path + auxiliary text fields
# Compacted transcripts: transcripts_compact/{session_id}.jsonl
# Subagent transcripts: transcripts_subagent/{agent_id}__{tool_use_id}.jsonl
df_session_logs = pd.read_parquet("session_logs.parquet")
with open(df_session_logs["transcript_path"].iloc[0]) as f:
raw_transcript = f.read()
Quick Examples
import pandas as pd
conversations = pd.read_parquet("conversations.parquet")
sessions = pd.read_parquet("sessions.parquet")
# Compact user-assistant conversation (no tool calls/thinking)
conv = conversations[conversations.is_conversational]
print(conv[["conversation_turn_number", "role", "content"]].head(10))
# All user prompts
prompts = conversations[conversations.turn_type == 'user_prompt']
print(prompts[["content", "agent", "language", "prompt_intent", "prompt_pushback"]])
# All tool calls with extracted file paths
tools = conversations[conversations.turn_type == "tool_use"]
print(tools[["tool_name", "file_path", "command"]].head())
# Recorded parent-side subagent calls (inspect JSON arguments for task text).
delegations = conversations.loc[(conversations.turn_type == "tool_use") & (conversations.orchestration_kind == "subagent_spawn")]
print(delegations[["session_id", "turn_id", "tool_name", "tool_input_json"]].head().to_string(index=False))
# Agent thinking traces
thinking = conversations[conversations.turn_type == "assistant_thinking"]
print(thinking[thinking['content']!=''][["content", "agent", "model"]])
# Sessions with high agent-authored code percentage
high_agent = sessions[sessions.agent_percentage > 90]
# Join conversations to sessions for richer context
merged = conversations.merge(
sessions[["session_id", "repo_id", "agent_percentage"]],
on="session_id"
)
Dataset Schema
Relationship Map
repositories (1) ───-> (N) checkpoints [repo_id]
checkpoints (N) ←──-> (M) sessions [JSON arrays on both sides]
checkpoints (1) ───-> (N) commits [checkpoint_pk]
sessions (1) ───-> (1) session_logs [session_id]
sessions (1) ───-> (N) conversations [session_id]
checkpoints (1) ───-> (N) subagent_tasks [checkpoint_pk]
subagent_tasks (1) ───-> (N) conversations_subagents [task_pk]
skill_definitions (1) ───-> (N) skill_invocations [definition_id]
sessions (1) ───-> (N) skill_invocations/context_events [session_id, when available]
subagent_tasks.session_id and conversations_subagents.session_id hold the
parent session's id, since a subagent runs inside its parent's session. That is why subagent turns live in their own table: keyed on session_id they would collide with the parent's rows in conversations.
ID Strategy
| Entity | Primary Key | Format | Globally Unique? |
|---|---|---|---|
| Repository | repo_id |
owner/repo |
Yes |
| Checkpoint | checkpoint_pk |
{repo_id}#{checkpoint_id} |
Yes |
| Session | session_id |
UUID from agent | Yes |
| Commit association | (checkpoint_pk, commit_sha) |
checkpoint key + 40-hex SHA | Yes; a SHA may appear under multiple checkpoints |
| Conversation turn | turn_id |
{session_id}#{turn_number} |
Yes |
| Subagent task | task_pk |
{checkpoint_pk}#{tool_use_id} |
Yes |
| Subagent turn | turn_pk |
{task_pk}#{turn_number} |
Yes |
| Skill definition | definition_id |
stable hash-derived id | Yes |
| Skill invocation | invocation_id |
stable hash-derived id | Yes |
| Context event | context_event_id |
stable hash-derived id | Yes |
turn_idis unique within a release, but not stable across releases.turn_numberis a positional index over all rows, which can change if sessions update across release. Join onsession_id(stable) rather thanturn_idwhen comparing across releases.
1. conversations.parquet
One row per conversation turn (user prompt, assistant response, assistant thinking step, tool call, or metadata event). Filter with is_conversational=True for a clean user-assistant dialogue view.
Two turn columns:
- turn_number: sequential across ALL rows in the session (including tool calls)
- conversation_turn_number: sequential ONLY for is_conversational=True rows. NULL for tool/thinking rows
| Column | Type | Description |
|---|---|---|
turn_id |
string | PK. {session_id}#{turn_number} |
session_id |
string | FK -> sessions |
checkpoint_pk |
string | FK -> checkpoints (canonical) |
repo_id |
string | FK -> repositories |
user_id |
string | Commit author (sessions.user_id) |
turn_number |
int | 0-based sequential index across ALL rows |
conversation_turn_number |
int | Sequential index for conversational rows only. NULL for tool/thinking rows |
source_entry_uuid |
string | Native transcript entry/event identifier, when the format provides one. |
response_id |
string | Native logical API-response/message identifier shared by content-block rows from the same response. |
content_block_index |
int | 0-based block index within the source entry (0 for formats without block structure). |
role |
string | user, assistant, tool_use, tool_result, or metadata |
turn_type |
string | user_prompt, system_injected, assistant_response, assistant_thinking, tool_use, tool_result, summary, system_event, file_snapshot, progress, queue_operation |
is_conversational |
bool | True for user_prompt and assistant_response |
content |
string | Prompt text, response text, thinking text, tool input JSON, or result text (truncated to 10KB for tool_result) |
model |
string | Model name (e.g. claude-opus-4-6). NULL for user/tool_result/metadata rows |
timestamp |
timestamp | Entry timestamp |
input_tokens |
int | Per-message input tokens from API usage. Carried on the assistant_response row; for assistant turns that are only tool calls (no text), carried on the first tool_use row of that turn so per-turn token/cache columns are complete. NULL/0 elsewhere. |
output_tokens |
int | Per-message output tokens (same row placement as input_tokens). |
cache_creation_input_tokens |
int | Cache-creation tokens from API usage (same row placement as input_tokens). |
cache_read_input_tokens |
int | Cache-read tokens from API usage (same row placement as input_tokens). |
is_continuation |
bool | User turn starts with "This session is being continued" |
is_first_turn |
bool | First conversational turn in session |
word_count |
int | Word count of content |
char_count |
int | Character count of content |
tool_name |
string | Tool name for tool_use/tool_result rows, e.g.: Read, Write, Edit, Bash, Grep, Glob, ... |
tool_call_id |
string | The tool-call ID linking tool_use to tool_result |
file_path |
string | Extracted from file-modifying/reading tool input. |
command |
string | Extracted from Bash / run_command / shell / Execute / exec_command tool input |
pattern |
string | Extracted from Grep/Glob tool input |
tool_input_json |
string (JSON) | Full tool input parameters for tool_use rows |
orchestration_kind |
string | For rows that orchestrate other agents (e.g. a Task/Agent launch), which kind of orchestration it was. NULL for ordinary rows. |
category |
string | Research / Action / Orchestration / Other for tool rows |
bash_category |
string | git / package manager / test-build / file ops / other for Bash/shell tools |
queue_op_subtype |
string | For queue_operation rows only. One of: user_prompt_enqueued (user typed while agent busy), user_prompt_delivered (queued message was sent to agent), user_prompt_discarded (queued message was removed without delivery), task_notification (completed subagent result), other. NULL for all other turn types. |
agent |
string | Agent name (denormalized from session) |
strategy |
string | Strategy (denormalized from session) |
language |
string | Detected natural language of the user prompt (populated for user_prompt rows only). |
prompt_intent |
string | LLM-annotated intent of a user prompt (turn_type=user_prompt). One of: create new code, refactor, debug, understand, connect, git, test, other. NULL for non-prompt rows or unannotated prompts. See paper for more details. |
prompt_pushback |
string | LLM-annotated pushback class for a user prompt (turn_type=user_prompt). One of: correction, rejection, failure_report, non_pushback. NULL for non-prompt rows or unannotated prompts. See paper for more details. |
Row types produced from transcript entries
| Source | turn_type | role | is_conversational |
|---|---|---|---|
| User text message | user_prompt |
user |
True |
System-injected user message — any role: user event that is harness/system output rather than human-typed prose, except if it is exectud by the user, such as slash-commands or interruptions. |
system_injected |
user |
False |
| Assistant text blocks | assistant_response |
assistant |
True |
| Assistant thinking blocks | assistant_thinking |
assistant |
False |
| Assistant tool_use blocks | tool_use |
tool_use |
False |
| User tool_result blocks | tool_result |
tool_result |
False |
| Session summary entries | summary |
metadata |
False |
| System/hook event entries | system_event |
metadata |
False |
| File history snapshot entries | file_snapshot |
metadata |
False |
| Progress entries (hooks, etc.) | progress |
metadata |
False |
| Queue operation entries | queue_operation |
metadata |
False |
Note on queue_operation rows: Claude Code queues messages when the agent is busy. The queue_op_subtype column distinguishes: (a) user prompts typed while agent was running (user_prompt_enqueued), which may have been delivered (user_prompt_delivered) or silently discarded before reaching the agent (user_prompt_discarded); and (b) completed subagent results (task_notification). The content column contains the plain-text prompt or the <task-notification> XML payload; it is empty for dequeue and remove operations.
2. sessions.parquet
| Column | Type | Description |
|---|---|---|
session_id |
string | PK. UUID from agent |
repo_id |
string | FK -> repositories |
owner_id |
string | Repository owner (from repo_id split) |
user_id |
string | Commit author resolved via GitHub API. |
checkpoint_ids |
string (JSON) | All checkpoint_pks referencing this session |
canonical_checkpoint_pk |
string | FK -> checkpoints (source for dedup) |
agent |
string | e.g. "Claude Code", "Gemini CLI" |
strategy |
string | Entire CLI strategy (e.g. "auto", "manual") |
branch |
string | Git branch the session was recorded on |
created_at |
timestamp | Session creation time |
cli_version |
string | Version of the Entire CLI that recorded the session |
files_touched |
string (JSON) | List of file paths touched during the session |
files_touched_count |
int | Number of files touched |
checkpoints_count |
int | Number of checkpoints created during the session |
input_tokens |
int | Main-agent input tokens reported by Entire from the agent's native usage metadata. |
output_tokens |
int | Main-agent output tokens reported by Entire from native usage metadata. |
cache_creation_tokens |
int | Main-agent prompt-cache creation tokens reported by Entire. |
cache_read_tokens |
int | Main-agent prompt-cache read tokens reported by Entire. |
api_call_count |
int | Main-agent API calls reported by Entire. |
agent_lines |
float | Committed lines attributed to the agent by Entire's recorded attribution. |
human_added |
float | Human-added lines reported by Entire's attribution. |
human_modified |
float | Human-modified lines reported by Entire's attribution. |
human_removed |
float | Human-removed lines reported by Entire's attribution. |
total_committed |
float | Total committed lines in Entire's attribution record. |
agent_percentage |
float | Agent-authored percentage reported by Entire. Interpret with attribution_metric_version. |
transcript_identifier_at_start |
string | Transcript identifier at the start of the session |
transcript_path |
string | Path to the transcript file |
tool_call_count |
int | Total tool calls in this session |
unique_tools_count |
int | Number of distinct tools used |
research_count |
int | Count of research tool calls (Read/Grep/Glob/WebFetch/WebSearch/read_file/grep/glob/list_directory) |
action_count |
int | Count of action tool calls (Write/Edit/NotebookEdit/Bash/write_file/edit_file/run_command/etc.) |
first_write_position |
int | Position of the first file-modifying tool call. |
duration_seconds |
float | Estimated session duration in seconds (from transcript timestamps) |
turn_count |
int | Number of conversational turns |
prompt_count |
int | Non-continuation user turns |
content_hash |
string | SHA256 from content_hash.txt |
user_persona |
string | LLM-annotated persona of the session's primary user. One of: Expert Nitpicker, Vague Requester, Mind Changer, Other. NULL for unannotated sessions. See paper for more details. |
session_success |
string | LLM-annotated session success score (0-100, as a string). NULL for unannotated sessions. See paper for more details. |
viewer_url_hf |
string | Link to this session's transcript file in this Hugging Face repo |
viewer_url_entireio |
string | Link to view this session in the Entire web viewer |
agent_turn_id |
string | The CLI's own turn-correlation id: checkpoints produced by one agent turn share it. Distinct from conversations.turn_id. |
has_subagent_tokens |
bool | True when Entire reported a separate subagent-usage block. |
subagent_input_tokens |
int | Subagent input tokens reported separately by Entire; add to input_tokens for combined usage. |
subagent_output_tokens |
int | Separately reported subagent output tokens. |
subagent_cache_creation_tokens |
int | Separately reported subagent cache-creation tokens. |
subagent_cache_read_tokens |
int | Separately reported subagent cache-read tokens. |
subagent_api_call_count |
int | Separately reported subagent API calls. |
model |
string | LLM the session ran on, e.g. claude-opus-4-6. Empty string when the agent did not report one (the field is always written upstream). |
kind |
string | Session purpose. Empty for a normal session; otherwise e.g. agent_review, agent_investigate, imported. |
compact_transcript_delta_only |
bool | True for a compact transcript written by an old CLI that stored only one checkpoint's delta rather than the whole session. Such a file ships but must not be read as a complete session. |
checkpoint_transcript_start |
float | Line offset in the raw transcript where this checkpoint's slice begins. Lets you segment a transcript shared by several checkpoints. |
context_tokens |
float | Context tokens in use at session end, when the agent reports it. |
context_window_size |
float | Context window size, when the agent reports it. |
review_skills |
string (JSON) | Review skills that were run. Set only for review-kind sessions. |
review_prompt |
string | Prompt used for a review-kind session. |
summary_intent |
string | AI-generated summary: what the user wanted to accomplish. |
summary_outcome |
string | AI-generated summary: what was achieved. |
summary_learnings |
string (JSON) | AI-generated learnings, grouped into repo, code ([{path, line, end_line, finding}]) and workflow. |
summary_friction |
string (JSON) | AI-generated list of problems encountered. [] means none were found; empty string means no summary was generated. |
summary_open_items |
string (JSON) | AI-generated list of unfinished work / tech debt. Same empty-vs-[] distinction as above. |
agent_removed |
float | Committed deletions attributed to the agent. |
total_lines_changed |
float | Total committed line changes (adds + modifies + removes). The newer denominator; total_committed is the legacy additions-only one. |
attribution_metric_version |
float | 2 when agent_percentage uses total_lines_changed; 0/NULL for the legacy additions-only definition. |
attribution_binary_files_changed |
float | Binary files changed, excluded from line counts. |
attribution_binary_files_removed |
float | Binary files removed, excluded from line counts. |
attribution_type |
string | Attribution variant recorded upstream, when present. |
attribution_branch |
string | Branch the attribution was computed against, when present. |
prompt_attributions |
string (JSON) | Per-prompt line attribution: [{checkpoint_number, user_lines_added, user_lines_removed, agent_lines_added, agent_lines_removed, user_added_per_file}]. user_added_per_file keys are file paths. |
agents |
string | Additional agent identification recorded by some CLI versions (distinct from agent). |
commit_tree_hash |
string | Git tree hash the session was recorded against, when present. |
steps_count |
float | Step count recorded by some CLI versions. |
session_started_at |
string | Session start time, when the agent reports it separately from created_at. |
session_ended_at |
string | Session end time, when the agent reports it. |
Note: Session token and attribution fields come from metadata recorded by Entire from the agent integration. They are not reconstructed by summing parsed tool calls. By contrast, conversation-row usage is parsed from native transcript usage records, and git diff/history totals are computed locally.
2b. session_logs.parquet
One row per session. The raw JSONL/JSON transcript is stored as a file under transcripts/{session_id}.jsonl (or .json for Gemini CLI) and referenced by the transcript_path column. Join to sessions on session_id.
Some sessions additionally have a compact transcript under
transcripts_compact/{session_id}.jsonl (compact_transcript_path). Prefer
transcripts/ for analysis — it is the agent's own complete record; the compact
form is derived and lossy (no system turns or reasoning, and parallel tool
results may be missing).
| Column | Type | Description |
|---|---|---|
session_id |
string | PK / FK -> sessions |
transcript_path |
string | Path to the session's transcript file under transcripts/ |
compact_transcript_path |
string | Path to the CLI's compact transcript under transcripts_compact/, when one exists. A normalised, agent-agnostic rendering of the same session: user/assistant turns with tool calls inlined, but no system turns or reasoning. Empty when the session has none. |
context_md |
string | Full content of context.md |
session_metadata_raw |
string (JSON) | Raw metadata.json content |
3. checkpoints.parquet
| Column | Type | Description |
|---|---|---|
checkpoint_pk |
string | PK. {repo_id}#{checkpoint_id} |
checkpoint_id |
string | Checkpoint id: 12-char hex, or a 26-char ULID for newer CLI versions (see checkpoint_id_format) |
repo_id |
string | FK -> repositories |
session_pks |
string (JSON) | FK list -> sessions |
session_count |
int | Number of sessions in this checkpoint |
commit_shas |
string (JSON) | FK list -> commits |
commit_count |
int | Number of commits in this checkpoint |
commit_link_status |
string | Commit-trailer resolution result: linked, not_found, search_failed, or not_searched. not_found means no linking commit was reachable in the successfully fetched code/PR history; it does not mean the checkpoint made no code changes. |
author_user_ids |
string (JSON) | Unique user_ids of commit authors |
unique_author_count |
float | Number of distinct commit authors |
user_id |
string | Set when checkpoint has a single author. |
cli_version |
string | Version of the Entire CLI |
strategy |
string | Entire CLI strategy |
branch |
string | Git branch |
checkpoints_count |
int | From metadata: total checkpoints in the session |
files_touched |
string (JSON) | List of file paths touched |
files_touched_count |
int | Number of files touched |
cp_input_tokens |
int | Main-agent input tokens aggregated from Entire's native checkpoint metadata. |
cp_output_tokens |
int | Main-agent output tokens aggregated from Entire's native checkpoint metadata. |
cp_cache_creation_tokens |
int | Main-agent cache-creation tokens aggregated from Entire metadata. |
cp_cache_read_tokens |
int | Main-agent cache-read tokens aggregated from Entire metadata. |
cp_api_call_count |
int | Main-agent API calls aggregated from Entire metadata. |
cp_has_subagent_tokens |
bool | True when Entire reported a separate subagent-usage block. |
cp_subagent_input_tokens |
int | Separately reported subagent input tokens; add to cp_input_tokens for combined usage. |
cp_subagent_output_tokens |
int | Separately reported subagent output tokens. |
cp_subagent_cache_creation_tokens |
int | Separately reported subagent cache-creation tokens. |
cp_subagent_cache_read_tokens |
int | Separately reported subagent cache-read tokens. |
cp_subagent_api_call_count |
int | Separately reported subagent API calls. |
total_additions |
int | Sum of added lines across commits in this checkpoint |
total_deletions |
int | Sum of deleted lines across commits in this checkpoint |
checkpoint_backend |
string | Storage backend the checkpoint came from: git-branch (one shared branch) or git-refs (one git ref per checkpoint). Empty if not recorded. |
checkpoint_id_format |
string | hex (12-char) or ulid (26-char) checkpoint id. |
source_repo_id |
string | Repository the checkpoint data was fetched from. Differs from repo_id only when the repo keeps its checkpoints in a separate repository. |
cp_combined_agent_lines |
float | Checkpoint-level combined agent lines recorded by Entire. |
cp_combined_agent_removed |
float | Checkpoint-level combined agent removals recorded by Entire. |
cp_combined_human_added |
float | Checkpoint-level combined human additions recorded by Entire. |
cp_combined_human_modified |
float | Checkpoint-level combined human modifications recorded by Entire. |
cp_combined_human_removed |
float | Checkpoint-level combined human removals recorded by Entire. |
cp_combined_total_committed |
float | Entire's legacy additions-only combined denominator. |
cp_combined_total_lines_changed |
float | Entire's combined changed-line denominator. |
cp_combined_agent_percentage |
float | Combined agent percentage recorded by Entire (0-100). |
cp_combined_metric_version |
float | 2 when the percentage uses total_lines_changed; 0/NULL for the legacy definition. |
imported |
bool | True when the checkpoint was imported from pre-existing agent history rather than recorded live (read-only, commit-less). |
has_review |
bool | True when at least one session in this checkpoint was a review session. |
has_investigation |
bool | True when at least one session was an investigation session. |
checkpoint_version |
string | Upstream storage-format marker, e.g. branch-v1. |
migration_source_commit |
string | Set when the checkpoint was moved between storage backends by the CLI's migration tool; the commit it came from. |
cp_commit_tree_hash |
string | Git tree hash recorded with the checkpoint, when present. |
imported_commit_sha |
string | For imported checkpoints only: the commit they were anchored to. Not a foreign key into commits. |
session_file_paths |
string (JSON) | The checkpoint's sessions[] pointer array as written upstream: per-session paths to metadata, prompt, transcript, content_hash, context, and compact_transcript. |
checkpoint_metadata_raw |
string (JSON) | Raw metadata.json |
4. commits.parquet
| Column | Type | Description |
|---|---|---|
commit_sha |
string | Git commit SHA. Rows are checkpoint-commit associations; the key is (checkpoint_pk, commit_sha). NULL is retained on legacy-compatible commit_not_found sentinel rows. |
viewer_url_entireio |
string | Link to view this commit in the Entire web viewer |
checkpoint_pk |
string | FK -> checkpoints |
repo_id |
string | FK -> repositories |
commit_index |
int | 0-based within checkpoint |
num_commits |
int | Total commits for this checkpoint |
user_id |
string | Canonical user identity: GitHub username if resolved, else email, else author name. |
github_username |
string | GitHub username resolved via commit API. |
author_name |
string | Git author name |
author_email |
string | Git author email |
author_date |
timestamp | Git author timestamp |
commit_date |
timestamp | Git commit timestamp |
commit_message |
string | Commit message |
branch |
string | Git branch the commit was observed on |
is_agent_author |
bool | Author matches agent patterns |
files_changed_count |
int | Number of files touched in this commit |
total_additions |
int | Lines added in this commit |
total_deletions |
int | Lines removed in this commit |
files_changed |
string | Raw git name-status |
numstat |
string | Raw git numstat |
patch |
string | Full unified diff |
agent_changes |
string (JSON) | Agent file-modification tool calls. Each change includes session_id, change_position (the 0-based file-modification index within that session, not turn_number), and timestamp when available. |
file_attribution |
string (JSON) | Pipeline-computed per-file attribution (agent_only, human_only, or mixed) from matching transcript file-edit tool calls against the linked git diff. |
status |
string | ok, commit_not_found, or commit_search_failed. Rows with either non-ok status are retained compatibility sentinels and have a NULL commit_sha; use checkpoints.commit_link_status for checkpoint-level resolution state. |
5. repositories.parquet
| Column | Type | Description |
|---|---|---|
repo_id |
string | PK. owner/repo |
owner_id |
string | Owner part of repo_id |
name |
string | Short repo name |
url |
string | GitHub URL |
is_fork |
bool | Whether the repo is a fork on GitHub |
settings |
string (JSON) | Raw settings.json from the Entire CLI |
num_checkpoints |
int | Checkpoints for this repo in the dataset |
num_sessions |
int | Sessions for this repo in the dataset |
num_commits |
int | Commits for this repo in the dataset |
num_contributors_in_dataset |
int | Distinct commit authors for this repo in the dataset |
total_additions_in_dataset |
int | Lines added across dataset commits for this repo |
total_deletions_in_dataset |
int | Lines removed across dataset commits for this repo |
total_repo_commits_ever |
int | Commits reachable in the locally fetched code refs. |
total_repo_additions_ever |
int | Additions from local git log --shortstat over those refs. NULL when the history scan timed out. |
total_repo_deletions_ever |
int | Deletions from the same local history scan. NULL when it timed out. |
total_agent_commits_ever |
int | Reachable commits whose author matches the pipeline's agent-author patterns. |
total_agent_additions_ever |
int | Additions in those agent-author commits. NULL when the history scan timed out. |
total_agent_deletions_ever |
int | Deletions in those agent-author commits. NULL when the history scan timed out. |
last_scraped_at |
timestamp | When repo metadata was last scraped |
license_type |
string | Usable-license classification from the curated license registry |
checkpoints_repo |
string | The separate repository this repo pushes its checkpoints to, when it uses one; empty otherwise. Only code repos get a row here, so a checkpoints-only repository never appears in this table and never contributes a license_type. |
repo_github_metadata |
string (JSON) | Full GitHub /repos/{owner}/{repo} response |
repo_type_domain |
string | LLM-annotated repo domain. One of: application, devtools, other. See paper for more details. |
repo_type_audience |
string | LLM-annotated repo target audience. One of: enduser, developer, researchers, education. See paper for more details. |
6. subagent_tasks.parquet
Captured subagent task records (non-exhaustive list of delegations). Join to sessions on session_id — which is the parent session's id — or to
checkpoints on checkpoint_pk.
Older CLI versions recorded only a subagent's identity and its link to the parent session; newer ones also store the subagent's own transcript. Columns that only the newer format supplies are empty for the older rows, and such a task has no rows in conversations_subagents.
For broader historical coverage, select turn_type == "tool_use" and orchestration_kind == "subagent_spawn" from conversations.
| Column | Type | Description |
|---|---|---|
task_pk |
string | PK. {checkpoint_pk}#{tool_use_id} |
checkpoint_pk |
string | FK -> checkpoints |
repo_id |
string | FK -> repositories |
source_repo_id |
string | Repository the data was fetched from (see checkpoints.source_repo_id) |
session_id |
string | FK -> sessions. The parent session the subagent ran inside. Empty when it could not be resolved. |
session_link |
string | How session_id was determined: checkpoint_json (stated outright upstream), single (the checkpoint held exactly one session), transcript (taken from the subagent transcript), or ambiguous (several candidates, none provable — session_id is then empty). |
tool_use_id |
string | Id of the tool call that launched the subagent |
agent_id |
string | Subagent identifier |
subagent_type |
string | Subagent type, e.g. general-purpose. Empty for older records. |
task_description |
string | Free-text task the parent gave the subagent. PII-redacted. Empty for older records. |
files |
string (JSON) | Files the subagent touched. Empty for older records. |
files_count |
int | Number of files touched |
started_at |
string | When the subagent launch was observed |
completed_at |
string | When the subagent finished. Empty if it was still running when the checkpoint was written. |
duration_seconds |
float | completed_at − started_at. NULL if either is missing. |
in_flight |
bool | True when the subagent had started but not finished — its transcript is a snapshot, not the final one. |
transcript_unavailable_reason |
string | Why no transcript was stored, when the CLI could not retrieve one. |
transcript_path |
string | Path to the subagent transcript under transcripts_subagent/. Empty when none was stored. |
transcript_fingerprint |
string | "{size}:{sha256[:32]}" of the subagent transcript; changes when the transcript changes. |
turn_count |
int | Turns parsed from the subagent transcript (0 when there is none) |
tool_call_count |
int | Tool calls parsed from the subagent transcript |
7. conversations_subagents.parquet
Turns from subagent transcripts. Same idea as conversations but for work done
by a subagent, and deliberately a separate table: these rows carry the parent
session's session_id, so merging them into conversations would collide with
the parent's own turns.
Only available subagent transcripts contribute rows here. A child transcript's user role can contain instructions supplied by the parent agent or harness.
| Column | Type | Description |
|---|---|---|
turn_pk |
string | PK. {task_pk}#{turn_number} |
task_pk |
string | FK -> subagent_tasks |
checkpoint_pk |
string | FK -> checkpoints |
repo_id |
string | FK -> repositories |
session_id |
string | FK -> sessions (the parent session) |
agent_id |
string | Subagent identifier |
turn_number |
int | 0-based index within this subagent's transcript |
role |
string | user, assistant, tool_use, tool_result, or metadata |
turn_type |
string | Same vocabulary as conversations.turn_type |
content |
string | Turn text, tool input, or tool result. PII-redacted for user/assistant turns. |
ts |
string | Turn timestamp, when the transcript provides one |
input_tokens |
int | Input tokens for this turn, when reported |
output_tokens |
int | Output tokens for this turn, when reported |
8. skill_definitions.parquet
| Column | Description |
|---|---|
definition_id |
PK. Stable identifier for this definition/body version. |
skill_name |
Observed skill name, including a plugin prefix when present. |
unqualified_name |
Skill name without its plugin prefix. |
agent |
Coding agent associated with the observation. |
plugin_name |
Plugin namespace inferred from the qualified name or observed injection. |
scope |
Observed definition scope, such as plugin or unknown. |
skill_body |
Captured skill instructions; empty for an unobserved stub. |
body_available |
Whether the body was captured. |
definition_kind |
observed_version or unobserved_stub. |
definition_source |
Evidence source from which this catalog row was created. |
9. skill_invocations.parquet
| Column | Description |
|---|---|
invocation_id |
PK. Stable invocation/event identifier. |
definition_id |
FK to skill_definitions; may reference an unobserved stub. |
skill_name |
Observed skill name, including a plugin prefix when present. |
unqualified_name |
Skill name without its plugin prefix. |
plugin_name |
Plugin namespace, when present. |
repo_id |
FK to repositories. |
session_id |
Parent/main session identifier. |
task_pk |
FK to subagent_tasks for a subagent-scoped invocation; empty otherwise. |
subagent_id |
Native subagent identifier when available. |
execution_scope |
main or subagent. |
agent |
Coding agent associated with the invocation. |
invoker |
Actor that invoked the skill: user, agent, subagent, harness, or unknown. |
event_type |
Invocation signal type; use tool_invocation, prompt_invocation, and skill_injection for usage counts. |
source |
Native or parsed source of the signal. |
source_confidence |
explicit or inferred. |
turn_id |
Source turn identifier when available; not guaranteed to join to released conversations. |
tool_call_id |
Source tool-call identifier when available; not guaranteed to join to released conversations. |
timestamp |
Invocation timestamp when reported. |
transcript_start |
Native transcript start anchor when reported. |
transcript_end |
Native transcript end anchor when reported. |
native_event_id |
Native skill-event identifier when reported. |
source_signal |
Native or inferred signal label. |
source_agent |
Agent named by the native source, when present. |
native_tool_name |
Native tool name associated with the event. |
10. context_events.parquet
| Column | Description |
|---|---|
context_event_id |
PK. Stable event identifier. |
repo_id |
FK to repositories. |
session_id |
Parent/main session identifier when available. |
task_pk |
FK to subagent_tasks for subagent-scoped context. |
subagent_id |
Native subagent identifier when available. |
execution_scope |
main or subagent. |
agent |
Coding agent associated with the event. |
invoker |
Actor responsible for the event. |
event_type |
Context category, such as harness_injection, slash_command, instruction_file, skill_definition_load, plugin_skill_invocation, or review_prompt. |
name |
Normalized command, instruction-file, skill, or event name. |
plugin_name |
Plugin namespace when applicable. |
turn_id |
Source turn identifier when available. |
timestamp |
Event timestamp when available. |
source |
Extraction source. |
source_confidence |
explicit or inferred. |
source_table |
Provenance table; it may be an internal table not shipped in this release. |
source_row_id |
Identifier of the source row in source_table. |
Data Collection
Source
Data is collected from public GitHub repositories that use the Entire.io CLI to checkpoint their AI coding sessions. Checkpoints are stored either on a special branch ()entire/checkpoints/v1) or in one ref per checkpoint under refs/entire/checkpoints/, containing:
- Session metadata (agent, strategy, token usage, code attribution)
- Full conversation transcripts (JSONL for Claude Code, JSON for Gemini CLI, etc.)
- User prompts and context
Supported Agents
- Claude Code
- OpenAI Codex
- Gemini CLI
- Cursor
- OpenCode
- GitHub Copilot CLI
- Factory AI Droid
- Pi
PII Redaction
We redacted personally identifiable information in all user prompts, assistant responses, subagent prose/task descriptions, and released skills using Microsoft Presidio (named-entity detection) and TruffleHog (secret detection).
Deduplication Strategy
Sessions may appear in multiple checkpoints (a session spans checkpoint boundaries). We deduplicate on session_id, keeping the record with the highest output_tokens (most complete). The checkpoint_ids column preserves the full list of checkpoints each session appeared in.
Data Removal Requests
If you would like your data removed from SWE-chat, or if you encounter content that is illegal, you may request deletion.
To do so, please contact us via [email protected]. Please include repo_id, session_id, or turn_id values corresponding to the entries you wish to remove, and a brief explanation of the reason for removal.
SWE-chat dataset versions
| Version | Release date | Sessions | User prompts | Release commit |
|---|---|---|---|---|
| v1 | 2026-04-29 | 5,851 | 62,544 | f66cca9 |
| v2 | 2026-09-TODO | 17,968 | 229,909 | <sha> |
Citation
Please consider citing the following papers if you find this dataset useful:
@inproceedings{baumann2026swechat,
title={SWE-chat: Real-World AI Coding Sessions in the Wild},
author={Baumann, Joachim and Padmakumar, Vishakh and Li, Xiang and Yang, John and Yang, Diyi and Koyejo, Sanmi},
booktitle={Third Conference on Language Modeling},
year={2026},
url={https://arxiv.org/pdf/2604.20779}
}