Memory System

LibreFang's memory system provides persistent storage, semantic search, and knowledge graph functionality.


Table of Contents


Overview

The LibreFang memory system includes:

  • SQLite Persistence - Structured KV storage
  • Vector Embeddings - Semantic search capability
  • Knowledge Graph - Entities and relationships
  • Session Management - Cross-channel memory
  • Session Compaction - Shrinks an active session when it grows past a threshold
  • Auto-Dream - Optional, opt-in background consolidation that asks agents to periodically reflect on and curate their own long-term memory. See Background Consolidation (Auto-Dream).
  • Usage Tracking - Cost and usage statistics

Architecture

┌─────────────────────────────────────────┐
│              Agent Loop                  │
└─────────────────┬───────────────────────┘
                  │
                  ▼
┌─────────────────────────────────────────┐
│           Memory Subsystem               │
├─────────────────────────────────────────┤
│  ┌─────────┐  ┌─────────┐  ┌─────────┐  │
│  │ Session │  │ Vector  │  │Knowledge│  │
│  │ Store   │  │ Search  │  │ Graph   │  │
│  └─────────┘  └─────────┘  └─────────┘  │
├─────────────────────────────────────────┤
│           SQLite Database               │
└─────────────────────────────────────────┘

Configuration

Basic Configuration

[memory]
decay_rate = 0.05
sqlite_path = "~/.librefang/data/librefang.db"

Advanced Configuration

[memory]
decay_rate = 0.05
sqlite_path = "~/.librefang/data/librefang.db"
vector_dimension = 1536
max_memory_items = 10000
auto_compact = true
FieldTypeDefaultDescription
decay_rateFloat0.05Memory confidence decay rate
sqlite_pathString~/.librefang/data/librefang.dbDatabase path
vector_dimensionInteger1536Vector dimension
max_memory_itemsInteger10000Maximum memory entries
auto_compactBooleantrueAuto compaction

Session Management

Create Session

# Create new session
librefang session create --name "research-project"

List Sessions

# List all sessions
librefang session list

Session Operations

# View session details
librefang session info <session-id>

# Delete session
librefang session delete <session-id>

# Compact session
librefang session compact <session-id>

# Export session
librefang session export <session-id> --format json

Memory Operations

Store Memory

# Store simple memory
librefang memory store <agent> "user:preference:theme" "dark"

# Store a key-value pair (positional: agent key value)
librefang memory store coder "project:version" "1.0"

# Equivalent long form using the set subcommand
librefang memory set coder "note:1" "Important note"

Search Memory

# Keyword search
librefang memory search "project"

# Vector search (semantic search)
librefang memory search --vector "find information about AI agents"

# Search with filters
librefang memory search "meeting" --tags "work" --limit 10

Memory Operations

# Read memory
librefang memory get <key>

# Update memory
librefang memory update <key> --value "new value"

# Delete memory
librefang memory delete <key>

# List all memories
librefang memory list --prefix "project:"

Agent-callable semantic memory

Semantic memory is not only something an agent is given — it is something an agent can ask. The memory_semantic_* tools reach the embedding-backed memories table, which is the store the agent loop already recalls from automatically before each turn.

ToolEffect
memory_semantic_search(query, limit?, min_confidence?, min_similarity?)Search remembered facts by meaning. Each result carries an id, and a similarity when embeddings ranked it.
memory_semantic_add(content)Record a durable fact. The extractor may distil, merge, or decline it; the result says which.
memory_semantic_forget(memory_id)Retract one memory so it stops being recalled.
memory_semantic_stats()Counts per level and category, plus whether the subsystem is switched on.
memory_semantic_duplicates()Report near-duplicate groups. Changes nothing.
memory_semantic_consolidate()Merge every near-duplicate group, keeping the newest of each and retracting the rest. Opt-in, see below.

Only memory_semantic_search ships in every request; the rest are reachable through tool_search / tool_load. All six are withheld when [proactive_memory] enabled = false, and gated a second time on the agent's declared memory_read / memory_write capability scopes.

Consolidation is opt-in, per agent

memory_semantic_consolidate is the only tool here whose reach is the agent's whole store rather than a row it named: it soft-deletes every member but the newest of each near-duplicate group, in one unattended call, with no undo the agent can reach. It is therefore withheld unless the agent's own manifest asks for it.

# {workspace}/agent.toml — NOT config.toml (#5476)
[proactive_memory]
allow_self_consolidation = true

An explicit capabilities.tools entry does not stand in for this, and neither does tools = ["*"]: naming a tool grants reach, and this switch grants permission to delete rows the caller never named. There is deliberately no kernel-global counterpart — a deployment-wide switch would arm destructive maintenance on agents nobody re-examined.

Leaving it off costs nothing an operator cannot recover. memory_semantic_duplicates still reports every group the merge would have touched, and POST /api/memory/agents/{id}/consolidate still performs it with a human deciding when.

Asking for nothing rather than noise

Recall over-fetches candidates, ranks them by cosine similarity, and truncates to the top few. With no floor, a sparse store fills that top-k with whatever exists — irrelevant memories are promoted on merit and land in the prompt as if they were answers.

min_similarity sets the floor below which a memory is not recalled at all, which makes an empty recall possible:

# ~/.librefang/config.toml — deployment default, applies to automatic recall too
[proactive_memory]
min_similarity = 0.3
# {workspace}/agent.toml — per-agent override
[proactive_memory]
min_similarity = 0.45

Resolution is per-call argument → agent override → deployment default → no floor. It is inert when no embedding provider is configured, because then nothing has been scored: distinct from min_confidence, which is decay-derived trust in a memory's content and says nothing about whether that memory answers this query.


LibreFang supports semantic search with vector embeddings:

# Semantic search example
librefang memory search --vector "machine learning techniques for text classification"

Similarity Threshold

[memory]
similarity_threshold = 0.75

Full-text fallback

When no embedding provider is reachable — including the deliberate [memory] fts_only = true mode — search falls back to the memories_fts FTS5 index over memory content, ranked by bm25. The index is maintained by triggers on the memories table, so nothing has to remember to keep it current, and it is rebuilt from memories when the schema migration runs. Queries that the index cannot serve (a substring landing inside a word, a query made only of punctuation) still fall through to substring matching, so the index only ever adds matches.

API Endpoint

# Semantic search API (GET with query params)
curl "http://127.0.0.1:4545/api/memory/search?q=find+information+about+AI&limit=10&threshold=0.8"

Knowledge Graph

Entity Management

# Add entity
librefang kg add-entity --type "person" --name "John Doe" --properties '{"role": "developer"}'

# List entities
librefang kg list-entities --type "person"

# Search entities
librefang kg search-entities "John"

Relationship Management

# Add relationship
librefang kg add-relation \
  --from "person:john" \
  --relation "works_at" \
  --to "company:acme"

# List relationships
librefang kg list-relations --entity "person:john"

# Query relationships
librefang kg query --from "person:john" --relation "works_at"

Graph Queries

# Query path
librefang kg path --from "person:alice" --to "company:acme"

# Query subgraph
librefang kg subgraph --entity "person:bob" --depth 2

Session Compaction

Auto Compaction

Automatic compaction when session message count reaches threshold. Compaction settings are managed internally by the runtime and are not exposed in config.toml. The default threshold is 80 messages, keeping the 20 most recent verbatim and summarizing the rest.

Manual Compaction

# Compact session
librefang session compact <session-id>

# Compact all sessions
librefang session compact-all

# View compaction status
librefang session compaction-status

Compaction Algorithm

  1. Keep Recent N Messages
  2. Extract Key Information
  3. Generate Summary
  4. Keep Tool Call History

Background Consolidation (Auto-Dream)

Session compaction keeps a single session from overflowing. Auto-dream is a complementary mechanism that works at a longer time horizon: once in a while it wakes an opted-in agent up, hands it a four-phase prompt (Orient / Gather / Consolidate / Prune), and lets the agent curate its own long-term memory across recent sessions — upserting durable insights via memory_semantic_add and retracting stale ones via memory_semantic_forget.

Enabling it

Two toggles must both be true:

# ~/.librefang/config.toml
[auto_dream]
enabled = true
min_hours = 24         # earliest a given agent can re-dream
min_sessions = 5       # require this much real activity since the last dream
# agent manifest (.toml)
auto_dream_enabled = true

# Optional per-agent overrides of the global thresholds
auto_dream_min_hours = 168      # weekly, for a quiet agent
auto_dream_min_sessions = 1     # after every session, for a chatty one

How a dream runs

  1. The primary trigger is the AgentLoopEnd hook — the moment an agent finishes a turn, the kernel evaluates its four gates in order: global enabled → time since last dream → session activity count → per-agent file lock. Any miss and the dream is skipped. A sparse backstop scheduler (default check_interval_secs = 86400 / 1 day) covers opted-in agents that never turn (e.g. channel bots waiting on inbound traffic).
  2. On a pass, the agent is invoked as a forked turn off the canonical session via kernel.run_forked_agent_streaming — same system prompt, tools, and message prefix as the parent turn so Anthropic's prompt cache hits. Fork turns don't persist their messages back to canonical session, so the user's conversation history isn't polluted by consolidation chatter. A runtime tool allowlist covering the key/value and memory_semantic_* tools is enforced at execute time, not request-build time, so the schema stays byte-identical to the parent and cache alignment holds. Prompt-injected dreams that try a non-memory tool get a synthetic error back and can't actually invoke anything outside the allowlist. memory_semantic_consolidate is in that allowlist but reaches a dream only for an agent whose manifest also sets allow_self_consolidation — so three switches, all defaulting to off, stand between a fresh install and an unattended whole-store merge.
  3. Streamed progress (phase, tool calls, memories touched, last turn preview, token/cost) is kept in a per-agent registry and surfaced via the status endpoint.
  4. On success the lock's mtime advances to "now" (driving the time gate); on failure or abort the mtime is rolled back so the next tick will retry.

Surfaces

  • Web Dashboard — Settings → Auto-Dream card renders one row per agent with an opt-in toggle, live status badge, "Dream now" and "Abort" buttons, and a progress preview while the dream is in flight.
  • TUI — librefang → Dashboard tab shows a compact DREAMS strip with per-agent status glyphs; hides itself when no agent has dreamed yet.
  • HTTP API — GET /api/auto-dream/status, POST /api/auto-dream/agents/{id}/trigger, POST /api/auto-dream/agents/{id}/abort, PUT /api/auto-dream/agents/{id}/enabled.
  • Audit trail — Every dream emits a DreamConsolidation audit event with phase, token usage, and USD cost.

Manual controls

  • POST /api/auto-dream/agents/{id}/trigger bypasses the time and session gates (still respects the lock and the opt-in flag).
  • POST /api/auto-dream/agents/{id}/abort cancels an in-flight manual dream and rolls the lock back so the time gate reopens.
  • Scheduled dreams run inline and cannot be interrupted individually — they'll hit timeout_secs (default 600s) or finish.

See [auto_dream] for every field, default, and manifest override, plus how the runtime tool allowlist is enforced.


External Vector Backend

Approximate-nearest-neighbour search is the one part of the memory subsystem that accepts an external implementation today, and the seam is the VectorStore trait. LibreFang keeps owning the SQLite row store, the extraction pipeline and every isolation guard; the backend you attach decides only which memory IDs are nearest to a query embedding.

Enabling the HTTP backend

[memory]
vector_backend = "http"                                        # "sqlite" (the default) keeps everything in-process
vector_store_url = "http://127.0.0.1:6333/collections/memories"

vector_backend = "http" without vector_store_url fails the boot rather than silently falling back to SQLite, and any value other than "http" / "sqlite" is rejected by name. The base URL is used verbatim with a trailing slash stripped, so paths are appended directly to it.

The HTTP contract

Your service answers four endpoints under that base URL. Responses are capped at 64 MiB, connect timeout is 10s and total request timeout is 30s.

MethodPathRequest bodyResponse body
POST/insert{ id, embedding, payload, metadata }ignored; any 2xx means stored
POST/search{ query_embedding, limit, filter? }[{ id, payload, score, metadata }]
DELETE/delete{ id }ignored; any 2xx means deleted
POST/get_embeddings{ ids }{ "<id>": [f32, …], … }

filter is the caller's MemoryFilter serialized as-is — agent, source, scope, minimum confidence, created-at bounds, metadata equality pairs and peer. Honouring it lets the backend prune before it ranks, which is the whole reason it is sent.

What the backend does not get to decide

Everything /search returns is treated as untrusted, because the search result is an ID list from a process LibreFang does not control.

  • IDs are hydrated from SQLite, never from the backend's payload. The payload and metadata a backend returns are not what reaches the prompt; the row LibreFang already has under that ID is.
  • The caller's MemoryFilter is re-applied after hydration. A backend that ignores the filter it was handed cannot widen the caller's scope: agent, scope, source, confidence floor, created-at bounds and metadata equality are all re-checked against the hydrated rows, and peer_id is enforced as a SQL predicate during hydration because MemoryFragment does not carry it.
  • Cross-chat (chat_scope) and cross-session (session_scope) isolation run after that, on the same post-recall path as the in-process backend, so attaching a backend cannot reopen either guard.

Failure behaviour

  • An ID that is not a UUID is dropped with a WARN and the rest of the result set is still hydrated. One malformed row costs that row, not the recall.
  • An ID that hydrates to nothing — deleted, unknown, or belonging to another peer — is dropped the same way.
  • A transport-level failure (connection refused, timeout, non-2xx, unparseable body) surfaces as an error from that recall. Callers differ in how they treat it: the prompt-assembly path logs and continues with whatever else it has, while a direct API search returns the error.

Implementing the trait in-process

vector_backend = "http" covers any backend you can put behind a small HTTP service, which is the supported path. If you would rather link a native client, implement VectorStore and attach it to the substrate at boot:

use librefang_types::memory::{MemoryFilter, VectorSearchResult, VectorStore};

#[async_trait]
impl VectorStore for QdrantVectorStore {
    async fn insert(&self, id: &str, embedding: &[f32], payload: &str,
                    metadata: HashMap<String, serde_json::Value>) -> LibreFangResult<()> { … }

    async fn search(&self, query_embedding: &[f32], limit: usize,
                    filter: Option<MemoryFilter>) -> LibreFangResult<Vec<VectorSearchResult>> { … }

    async fn delete(&self, id: &str) -> LibreFangResult<()> { … }

    async fn get_embeddings(&self, ids: &[&str]) -> LibreFangResult<HashMap<String, Vec<f32>>> { … }

    fn backend_name(&self) -> &str { "qdrant" }
}

search is called from a blocking bridge on the recall path, so it must not block the executor for long — the timeouts above exist for exactly that reason and a native client needs its own.


Usage Tracking

View Usage

# View usage statistics
librefang usage

# View Agent usage
librefang usage --agent <agent-id>

# View provider usage
librefang usage --provider

Cost Tracking

# View costs
librefang cost

# View costs by date range
librefang cost --from 2025-01-01 --to 2025-01-31

# Export report
librefang cost export --format csv

API Endpoints

KV Memory Operations

EndpointMethodDescription
/api/memory/searchGETSearch memory (query params: q, limit, threshold)
/api/memoryPOSTStore a memory entry
/api/memory/{id}GETGet a specific memory entry
/api/memory/{id}DELETEDelete a memory entry

Session Operations

EndpointMethodDescription
/api/memory/sessionsGETList sessions
/api/memory/sessions/{id}GETGet session details
/api/memory/sessions/{id}/compactPOSTCompact session

Knowledge Graph

The knowledge graph lives per-agent, not as a global REST surface. Use the per-agent endpoints:

EndpointMethodDescription
/api/memory/agents/{id}/relationsGETQuery the agent's knowledge graph (entities + relations)
/api/memory/agents/{id}/relationsPOSTStore new relations on the agent's graph

Agents can also build the graph from inside the loop with the knowledge_add_entity / knowledge_add_relation / knowledge_query tools. The earlier /api/memory/kg/* endpoints in older documentation never shipped — those paths return 404.

Proactive Memory

Proactive memory lets agents autonomously surface, consolidate, and recall long-term knowledge without explicit tool calls. Each entry is stored three ways simultaneously — semantic store (text + embedding), structured KV store (memory:{id}), and knowledge graph (extracted entity / relation triples) — so agents can recall a memory by similarity, by ID, or by traversing the graph from a known entity.

Entry schema

FieldTypePurpose
idUUIDUnique memory identifier
agent_idUUIDOwning agent — entries are scoped per agent, no cross-agent leakage
contenttextThe memory text/fact
sourceenumHow created: auto_memorize, manual_add, …
scopeenumLevel: user_memory, session_memory, agent_memory
confidencef64 (0.0–1.0)Relevance score; decays over time
metadataJSONFree-form KV; includes category (see below)
created_atRFC3339 timestampCreation time
accessed_atRFC3339 timestampLast access — drives decay
access_countintTimes this memory was recalled
deletedboolSoft-delete flag
embeddingvector (separate table)For cosine similarity search

Scopes

ScopeLifetimeUse
user_memoryPersistent across sessionsUser-level facts and preferences (may be shared across this user's agents).
session_memoryAuto-deleted after session_ttl_hours (default 24)Working notes for the current conversation.
agent_memoryPer-agent persistentAgent-learned behaviour and skills, isolated.

Categories

Categories live on metadata.category. The defaults extracted by auto_memorize are: communication_style, preference, expertise, work_style, project_context, personal_detail, frustration. The category list is configurable.

Auto-consolidation

Triggered automatically every 10 auto_memorize calls per agent (no explicit cron / no user action needed). The consolidator:

  1. Finds duplicate or near-duplicate memories using a tiered similarity ladder:
    • Substring containment (exact / superset / subset)
    • Vector cosine (when embeddings stored)
    • Jaccard word overlap (fallback)
  2. Keeps the most recently created of each duplicate cluster.
  3. Soft-deletes the rest and logs the merge count for audit.

The consolidator only scans the most recent 100 entries to keep the dedup pass O(n²)-safe on a hot loop. Run a manual POST /api/memory/agents/{id}/consolidate to scan the full history.

Decay

For memories not accessed for more than 1 day:

decayed_conf = original_conf × e^(-decay_rate × days_since_access)
final_conf   = min(decayed_conf × (1 + log2(access_count)), 1.0)

The decay sweep is rate-limited to once per hour (checked at the top of auto_retrieve). Defaults: confidence_decay_rate = 0.01 (very slow — halves in roughly 70 days), session_ttl_hours = 24. A memory is "stale" when (a) untouched for more than a day, or (b) its scope is session_memory and the TTL has elapsed.

Per-agent vs cross-agent

All entries except user_memory are strictly per-agent. Search, consolidate, and eviction all filter on agent_id — there is no path for one agent to read another's agent_memory or session_memory. The per-agent cap defaults to 1000 entries; when exceeded the oldest / lowest-confidence entries are evicted.

API endpoints

The proactive-memory endpoints live under two roots — global and per-agent. (KV memory at /api/memory/agents/{id}/kv/* is a separate subsystem; not listed here.)

Global / cross-agent:

MethodPathPurpose
GET/api/memoryList all proactive memories (paginated, optional ?category=)
POST/api/memoryAdd a memory entry
GET/api/memory/search?q=&limit=Semantic search across all entries
GET/api/memory/statsGlobal stats
GET/api/memory/configGet memory config
PATCH/api/memory/configUpdate memory config
POST/api/memory/cleanupManual session-TTL cleanup pass
POST/api/memory/decayManual confidence-decay pass
POST/api/memory/bulk-deleteDelete multiple entries
PUT/api/memory/items/{id}Update a single memory
DELETE/api/memory/items/{id}Delete a single memory
GET/api/memory/items/{id}/historyMemory edit history
GET/api/memory/user/{user_id}User-level memories across all agents

Per-agent:

MethodPathPurpose
GET/api/memory/agents/{id}List the agent's memories
DELETE/api/memory/agents/{id}Reset (clear all of) the agent's memories
GET/api/memory/agents/{id}/search?q=Agent-scoped semantic search
GET/api/memory/agents/{id}/statsPer-agent stats
DELETE/api/memory/agents/{id}/level/{level}Clear by scope (session, agent, user)
GET/api/memory/agents/{id}/duplicatesFind near-duplicates without deleting
POST/api/memory/agents/{id}/consolidateTrigger consolidation on the full history
GET/api/memory/agents/{id}/countMemory count
GET/api/memory/agents/{id}/relationsQuery the agent's knowledge graph
POST/api/memory/agents/{id}/relationsStore relations
GET/api/memory/agents/{id}/exportExport the agent's memories (JSON)
POST/api/memory/agents/{id}/importImport memories (JSON)

The earlier /api/memory/proactive/* namespace was a draft surface that never shipped — use the endpoints above. If your client still calls the old paths, expect 404.


Best Practices

  1. Regular Compaction - Prevent sessions from growing too large
  2. Use Tags - Easy organization and search
  3. Set Decay Rate - Control memory confidence
  4. Monitor Usage - Track costs and usage

Troubleshooting

# Rebuild vector index
librefang memory reindex

# Check index status
librefang memory index-status

Database Bloat

# Clean old data
librefang memory cleanup --older-than 30d

# Vacuum database
librefang memory vacuum

Memory Loss

# Check database integrity
librefang doctor

# Restore backup
librefang memory restore --backup <path>