Skip to content

For the complete documentation index, see llms.txt.

Trace object and sources

PIG reads native session files on an enrolled host and sends their contents to your analyzer. The analyzer stores the raw records and builds a canonical trace, a common representation of one agent session. This lets analysis work across different native formats while retaining the original evidence.

The collector supports three source families. Startup and terminal hooks can supply a transcript path directly. Background catch-up also discovers files in the following locations:

SourceHost valueDefault discovery locations
Claude Codeclaude~/.claude/projects/**/*.jsonl
Codexcodexsessions/**/*.jsonl and archived_sessions/**/*.jsonl beneath CODEX_HOME, or ~/.codex when unset
Claude Desktopclaude-desktop**/audit.jsonl beneath the two supported session directories in Claude’s application data

For Claude Desktop, the supported session directories are local-agent-mode-sessions and claude-code-sessions. Their parent application directory depends on the operating system:

PlatformClaude application directory
macOS~/Library/Application Support/Claude/
Windows%APPDATA%/Claude/, or ~/AppData/Roaming/Claude/ when APPDATA is unset
Linux$XDG_CONFIG_HOME/Claude/, or ~/.config/Claude/ when XDG_CONFIG_HOME is unset

These are discovery paths implemented by the collector. They do not mean every agent application is available on every platform or writes every listed format. Claude Desktop collection depends on the audit files being present and readable.

Claude Code’s generated hooks also attempt local Claude Desktop collection. Codex uses its own generated hooks. Cursor and Gemini currently distribute instructions without native trace collection. See Supported agents for the distinction.

The local ledger at ~/.promptless/instruction-hub/host-runtime-ledger.json records progress for each source. The collector reads new ranges, compresses their contents with gzip, and base64-encodes them for transport. It sends authenticated batches to the analyzer’s POST /v0/traces/batches?target=<host> endpoint, where <host> is a value from the table above.

The worker validates the source and target, stores the native data, and returns acknowledged ranges. Only a matching acknowledgment advances the host’s recorded progress. A retry can therefore resend a range whose response was lost. If the worker already accepted part of that range, the runtime checks the worker’s content proof before reconciling its offset.

Do not edit the ledger to force an upload or treat an HTTP success alone as evidence of completed analysis. Preserve native files while an outage is being resolved. File truncation, replaced content, malformed records, and oversized records require accounting beyond simply advancing a byte counter.

When a file is first discovered, collection can begin at its start. Existing history can therefore be uploaded after first enrollment. Catch-up skips very recent files while they are inside the collector’s idle grace period; a hook’s explicit transcript path receives priority.

The analyzer stores raw native artifacts in your trace bucket and maintains canonical session objects alongside them. PostgreSQL holds the associated metadata, ingestion progress, and analysis state.

A canonical object uses schema_version: 1 and contains these top-level fields:

FieldMeaning
schema_versionThe canonical object format version
sessionSource and session identity, agent type, observed models, and optional repository or parent-agent context
lifecycleSession state and timestamps, with native event time separate from ingestion time
sourcesSource file identities, raw artifact references, observed offsets, ingestion counts, and skipped or truncated content counts
rollupsObserved token usage, tool counts, compactions, and record accounting
context_artifactsTrace observations of instructions used, such as skill invocations or instruction-file references
preambleEvents before the first turn
pendingEvents not yet attached to a turn
turnsOrdered turns containing the session’s events

context_artifacts records what the trace reveals. It is separate from the instruction and repository snapshot prepared for analysis, described in Trust and data model.

Events have a kind, a global sequence number seq, and source evidence where available. The kinds are:

KindContent
user_messageUser input that opens a turn
assistant_messageAssistant output, with its observed audience where available
reasoningReasoning content exposed by the native trace
tool_callA tool invocation and its input
tool_resultA tool’s output or error
compactionA compaction summary
session_eventSession lifecycle or other agent machinery
otherA record not represented by a more specific event kind

Sequence numbers are contiguous in document order. Native events carry a src reference identifying their source entry, record, and byte offset. When the canonical view truncates a payload inline, the original uploaded bytes remain in the referenced raw artifact. Unknown event kinds and malformed records are different cases; parse coverage records what could not be decoded or modeled.

For an Acme pilot session, first check that the expected source and session ID appear in the analyzer. Then confirm the raw artifacts are stored, review ingestion or parse gaps, and check the analysis run’s status. A stored trace can still be waiting for the quiet period, repository resolution, or model access.

Use observability to investigate those stages. If the session never reaches the analyzer, start with the host’s enrollment and collection checks.