Memory System
LibreFang's memory system provides persistent storage, semantic search, and knowledge graph functionality.
Table of Contents
- Overview
- Architecture
- Configuration
- Session Management
- Memory Operations
- Vector Search
- Knowledge Graph
- Session Compaction
- Background Consolidation (Auto-Dream)
- External Vector Backend
- Usage Tracking
- API Endpoints
- Best Practices
- Troubleshooting
Overview
The LibreFang memory system includes:
- SQLite Persistence - Structured KV storage
- Vector Embeddings - Semantic search capability
- Knowledge Graph - Entities and relationships
- Session Management - Cross-channel memory
- Session Compaction - Shrinks an active session when it grows past a threshold
- Auto-Dream - Optional, opt-in background consolidation that asks agents to periodically reflect on and curate their own long-term memory. See Background Consolidation (Auto-Dream).
- Usage Tracking - Cost and usage statistics
Tip: Memory data is stored in ~/.librefang/data/librefang.db by default, with configurable storage path.
Architecture
┌─────────────────────────────────────────┐
│ Agent Loop │
└─────────────────┬───────────────────────┘
│
▼
┌─────────────────────────────────────────┐
│ Memory Subsystem │
├─────────────────────────────────────────┤
│ ┌─────────┐ ┌─────────┐ ┌─────────┐ │
│ │ Session │ │ Vector │ │Knowledge│ │
│ │ Store │ │ Search │ │ Graph │ │
│ └─────────┘ └─────────┘ └─────────┘ │
├─────────────────────────────────────────┤
│ SQLite Database │
└─────────────────────────────────────────┘
Configuration
Basic Configuration
[memory]
decay_rate = 0.05
sqlite_path = "~/.librefang/data/librefang.db"
Advanced Configuration
[memory]
decay_rate = 0.05
sqlite_path = "~/.librefang/data/librefang.db"
vector_dimension = 1536
max_memory_items = 10000
auto_compact = true
| Field | Type | Default | Description |
|---|---|---|---|
decay_rate | Float | 0.05 | Memory confidence decay rate |
sqlite_path | String | ~/.librefang/data/librefang.db | Database path |
vector_dimension | Integer | 1536 | Vector dimension |
max_memory_items | Integer | 10000 | Maximum memory entries |
auto_compact | Boolean | true | Auto compaction |
Session Management
Create Session
# Create new session
librefang session create --name "research-project"
List Sessions
# List all sessions
librefang session list
Session Operations
# View session details
librefang session info <session-id>
# Delete session
librefang session delete <session-id>
# Compact session
librefang session compact <session-id>
# Export session
librefang session export <session-id> --format json
Memory Operations
Store Memory
# Store simple memory
librefang memory store <agent> "user:preference:theme" "dark"
# Store a key-value pair (positional: agent key value)
librefang memory store coder "project:version" "1.0"
# Equivalent long form using the set subcommand
librefang memory set coder "note:1" "Important note"
Search Memory
# Keyword search
librefang memory search "project"
# Vector search (semantic search)
librefang memory search --vector "find information about AI agents"
# Search with filters
librefang memory search "meeting" --tags "work" --limit 10
Memory Operations
# Read memory
librefang memory get <key>
# Update memory
librefang memory update <key> --value "new value"
# Delete memory
librefang memory delete <key>
# List all memories
librefang memory list --prefix "project:"
Agent-callable semantic memory
Semantic memory is not only something an agent is given — it is something an agent can ask.
The memory_semantic_* tools reach the embedding-backed memories table, which is the store the agent loop already recalls from automatically before each turn.
The three memory_store / memory_recall / memory_list tools are a separate, exact-key key/value store.
memory_recall matches a key character for character; it does no fuzzy, semantic, or substring matching.
Reaching for it to "search memory" returns a key miss, not an empty result.
| Tool | Effect |
|---|---|
memory_semantic_search(query, limit?, min_confidence?, min_similarity?) | Search remembered facts by meaning. Each result carries an id, and a similarity when embeddings ranked it. |
memory_semantic_add(content) | Record a durable fact. The extractor may distil, merge, or decline it; the result says which. |
memory_semantic_forget(memory_id) | Retract one memory so it stops being recalled. |
memory_semantic_stats() | Counts per level and category, plus whether the subsystem is switched on. |
memory_semantic_duplicates() | Report near-duplicate groups. Changes nothing. |
memory_semantic_consolidate() | Merge every near-duplicate group, keeping the newest of each and retracting the rest. Opt-in, see below. |
Only memory_semantic_search ships in every request; the rest are reachable through tool_search / tool_load.
All six are withheld when [proactive_memory] enabled = false, and gated a second time on the agent's declared memory_read / memory_write capability scopes.
Consolidation is opt-in, per agent
memory_semantic_consolidate is the only tool here whose reach is the agent's whole store rather than a row it named: it soft-deletes every member but the newest of each near-duplicate group, in one unattended call, with no undo the agent can reach.
It is therefore withheld unless the agent's own manifest asks for it.
# {workspace}/agent.toml — NOT config.toml (#5476)
[proactive_memory]
allow_self_consolidation = true
An explicit capabilities.tools entry does not stand in for this, and neither does tools = ["*"]: naming a tool grants reach, and this switch grants permission to delete rows the caller never named.
There is deliberately no kernel-global counterpart — a deployment-wide switch would arm destructive maintenance on agents nobody re-examined.
Leaving it off costs nothing an operator cannot recover.
memory_semantic_duplicates still reports every group the merge would have touched, and POST /api/memory/agents/{id}/consolidate still performs it with a human deciding when.
Asking for nothing rather than noise
Recall over-fetches candidates, ranks them by cosine similarity, and truncates to the top few. With no floor, a sparse store fills that top-k with whatever exists — irrelevant memories are promoted on merit and land in the prompt as if they were answers.
min_similarity sets the floor below which a memory is not recalled at all, which makes an empty recall possible:
# ~/.librefang/config.toml — deployment default, applies to automatic recall too
[proactive_memory]
min_similarity = 0.3
# {workspace}/agent.toml — per-agent override
[proactive_memory]
min_similarity = 0.45
Resolution is per-call argument → agent override → deployment default → no floor.
It is inert when no embedding provider is configured, because then nothing has been scored: distinct from min_confidence, which is decay-derived trust in a memory's content and says nothing about whether that memory answers this query.
Vector Search
Semantic Search
LibreFang supports semantic search with vector embeddings:
# Semantic search example
librefang memory search --vector "machine learning techniques for text classification"
Similarity Threshold
[memory]
similarity_threshold = 0.75
Full-text fallback
When no embedding provider is reachable — including the deliberate [memory] fts_only = true mode — search falls back to the memories_fts FTS5 index over memory content, ranked by bm25.
The index is maintained by triggers on the memories table, so nothing has to remember to keep it current, and it is rebuilt from memories when the schema migration runs.
Queries that the index cannot serve (a substring landing inside a word, a query made only of punctuation) still fall through to substring matching, so the index only ever adds matches.
API Endpoint
# Semantic search API (GET with query params)
curl "http://127.0.0.1:4545/api/memory/search?q=find+information+about+AI&limit=10&threshold=0.8"
Knowledge Graph
Entity Management
# Add entity
librefang kg add-entity --type "person" --name "John Doe" --properties '{"role": "developer"}'
# List entities
librefang kg list-entities --type "person"
# Search entities
librefang kg search-entities "John"
Relationship Management
# Add relationship
librefang kg add-relation \
--from "person:john" \
--relation "works_at" \
--to "company:acme"
# List relationships
librefang kg list-relations --entity "person:john"
# Query relationships
librefang kg query --from "person:john" --relation "works_at"
Graph Queries
# Query path
librefang kg path --from "person:alice" --to "company:acme"
# Query subgraph
librefang kg subgraph --entity "person:bob" --depth 2
Session Compaction
Auto Compaction
Automatic compaction when session message count reaches threshold. Compaction settings are managed internally by the runtime and are not exposed in config.toml. The default threshold is 80 messages, keeping the 20 most recent verbatim and summarizing the rest.
Manual Compaction
# Compact session
librefang session compact <session-id>
# Compact all sessions
librefang session compact-all
# View compaction status
librefang session compaction-status
Compaction Algorithm
- Keep Recent N Messages
- Extract Key Information
- Generate Summary
- Keep Tool Call History
Background Consolidation (Auto-Dream)
Session compaction keeps a single session from overflowing. Auto-dream is a complementary mechanism that works at a longer time horizon: once in a while it wakes an opted-in agent up, hands it a four-phase prompt (Orient / Gather / Consolidate / Prune), and lets the agent curate its own long-term memory across recent sessions — upserting durable insights via memory_semantic_add and retracting stale ones via memory_semantic_forget.
Disabled by default. Auto-dream spends real tokens on a recurring schedule, so both the global switch and every agent's opt-in flag are false unless you turn them on. See [auto_dream] for the full reference.
Enabling it
Two toggles must both be true:
# ~/.librefang/config.toml
[auto_dream]
enabled = true
min_hours = 24 # earliest a given agent can re-dream
min_sessions = 5 # require this much real activity since the last dream
# agent manifest (.toml)
auto_dream_enabled = true
# Optional per-agent overrides of the global thresholds
auto_dream_min_hours = 168 # weekly, for a quiet agent
auto_dream_min_sessions = 1 # after every session, for a chatty one
How a dream runs
- The primary trigger is the
AgentLoopEndhook — the moment an agent finishes a turn, the kernel evaluates its four gates in order: global enabled → time since last dream → session activity count → per-agent file lock. Any miss and the dream is skipped. A sparse backstop scheduler (defaultcheck_interval_secs = 86400/ 1 day) covers opted-in agents that never turn (e.g. channel bots waiting on inbound traffic). - On a pass, the agent is invoked as a forked turn off the canonical session via
kernel.run_forked_agent_streaming— same system prompt, tools, and message prefix as the parent turn so Anthropic's prompt cache hits. Fork turns don't persist their messages back to canonical session, so the user's conversation history isn't polluted by consolidation chatter. A runtime tool allowlist covering the key/value andmemory_semantic_*tools is enforced at execute time, not request-build time, so the schema stays byte-identical to the parent and cache alignment holds. Prompt-injected dreams that try a non-memory tool get a synthetic error back and can't actually invoke anything outside the allowlist.memory_semantic_consolidateis in that allowlist but reaches a dream only for an agent whose manifest also setsallow_self_consolidation— so three switches, all defaulting to off, stand between a fresh install and an unattended whole-store merge. - Streamed progress (phase, tool calls, memories touched, last turn preview, token/cost) is kept in a per-agent registry and surfaced via the status endpoint.
- On success the lock's mtime advances to "now" (driving the time gate); on failure or abort the mtime is rolled back so the next tick will retry.
Surfaces
- Web Dashboard — Settings → Auto-Dream card renders one row per agent with an opt-in toggle, live status badge, "Dream now" and "Abort" buttons, and a progress preview while the dream is in flight.
- TUI —
librefang→ Dashboard tab shows a compact DREAMS strip with per-agent status glyphs; hides itself when no agent has dreamed yet. - HTTP API —
GET /api/auto-dream/status,POST /api/auto-dream/agents/{id}/trigger,POST /api/auto-dream/agents/{id}/abort,PUT /api/auto-dream/agents/{id}/enabled. - Audit trail — Every dream emits a
DreamConsolidationaudit event with phase, token usage, and USD cost.
Manual controls
POST /api/auto-dream/agents/{id}/triggerbypasses the time and session gates (still respects the lock and the opt-in flag).POST /api/auto-dream/agents/{id}/abortcancels an in-flight manual dream and rolls the lock back so the time gate reopens.- Scheduled dreams run inline and cannot be interrupted individually — they'll hit
timeout_secs(default 600s) or finish.
See [auto_dream] for every field, default, and manifest override, plus how the runtime tool allowlist is enforced.
External Vector Backend
Approximate-nearest-neighbour search is the one part of the memory subsystem that accepts an external implementation today, and the seam is the VectorStore trait.
LibreFang keeps owning the SQLite row store, the extraction pipeline and every isolation guard; the backend you attach decides only which memory IDs are nearest to a query embedding.
Enabling the HTTP backend
[memory]
vector_backend = "http" # "sqlite" (the default) keeps everything in-process
vector_store_url = "http://127.0.0.1:6333/collections/memories"
vector_backend = "http" without vector_store_url fails the boot rather than silently falling back to SQLite, and any value other than "http" / "sqlite" is rejected by name.
The base URL is used verbatim with a trailing slash stripped, so paths are appended directly to it.
The HTTP contract
Your service answers four endpoints under that base URL. Responses are capped at 64 MiB, connect timeout is 10s and total request timeout is 30s.
| Method | Path | Request body | Response body |
|---|---|---|---|
| POST | /insert | { id, embedding, payload, metadata } | ignored; any 2xx means stored |
| POST | /search | { query_embedding, limit, filter? } | [{ id, payload, score, metadata }] |
| DELETE | /delete | { id } | ignored; any 2xx means deleted |
| POST | /get_embeddings | { ids } | { "<id>": [f32, …], … } |
filter is the caller's MemoryFilter serialized as-is — agent, source, scope, minimum confidence, created-at bounds, metadata equality pairs and peer.
Honouring it lets the backend prune before it ranks, which is the whole reason it is sent.
What the backend does not get to decide
Everything /search returns is treated as untrusted, because the search result is an ID list from a process LibreFang does not control.
- IDs are hydrated from SQLite, never from the backend's payload. The
payloadandmetadataa backend returns are not what reaches the prompt; the row LibreFang already has under that ID is. - The caller's
MemoryFilteris re-applied after hydration. A backend that ignores the filter it was handed cannot widen the caller's scope: agent, scope, source, confidence floor, created-at bounds and metadata equality are all re-checked against the hydrated rows, andpeer_idis enforced as a SQL predicate during hydration becauseMemoryFragmentdoes not carry it. - Cross-chat (
chat_scope) and cross-session (session_scope) isolation run after that, on the same post-recall path as the in-process backend, so attaching a backend cannot reopen either guard.
Failure behaviour
- An ID that is not a UUID is dropped with a
WARNand the rest of the result set is still hydrated. One malformed row costs that row, not the recall. - An ID that hydrates to nothing — deleted, unknown, or belonging to another peer — is dropped the same way.
- A transport-level failure (connection refused, timeout, non-2xx, unparseable body) surfaces as an error from that recall. Callers differ in how they treat it: the prompt-assembly path logs and continues with whatever else it has, while a direct API search returns the error.
Implementing the trait in-process
vector_backend = "http" covers any backend you can put behind a small HTTP service, which is the supported path.
If you would rather link a native client, implement VectorStore and attach it to the substrate at boot:
use librefang_types::memory::{MemoryFilter, VectorSearchResult, VectorStore};
#[async_trait]
impl VectorStore for QdrantVectorStore {
async fn insert(&self, id: &str, embedding: &[f32], payload: &str,
metadata: HashMap<String, serde_json::Value>) -> LibreFangResult<()> { … }
async fn search(&self, query_embedding: &[f32], limit: usize,
filter: Option<MemoryFilter>) -> LibreFangResult<Vec<VectorSearchResult>> { … }
async fn delete(&self, id: &str) -> LibreFangResult<()> { … }
async fn get_embeddings(&self, ids: &[&str]) -> LibreFangResult<HashMap<String, Vec<f32>>> { … }
fn backend_name(&self) -> &str { "qdrant" }
}
search is called from a blocking bridge on the recall path, so it must not block the executor for long — the timeouts above exist for exactly that reason and a native client needs its own.
Usage Tracking
View Usage
# View usage statistics
librefang usage
# View Agent usage
librefang usage --agent <agent-id>
# View provider usage
librefang usage --provider
Cost Tracking
# View costs
librefang cost
# View costs by date range
librefang cost --from 2025-01-01 --to 2025-01-31
# Export report
librefang cost export --format csv
API Endpoints
KV Memory Operations
| Endpoint | Method | Description |
|---|---|---|
/api/memory/search | GET | Search memory (query params: q, limit, threshold) |
/api/memory | POST | Store a memory entry |
/api/memory/{id} | GET | Get a specific memory entry |
/api/memory/{id} | DELETE | Delete a memory entry |
Session Operations
| Endpoint | Method | Description |
|---|---|---|
/api/memory/sessions | GET | List sessions |
/api/memory/sessions/{id} | GET | Get session details |
/api/memory/sessions/{id}/compact | POST | Compact session |
Knowledge Graph
The knowledge graph lives per-agent, not as a global REST surface. Use the per-agent endpoints:
| Endpoint | Method | Description |
|---|---|---|
/api/memory/agents/{id}/relations | GET | Query the agent's knowledge graph (entities + relations) |
/api/memory/agents/{id}/relations | POST | Store new relations on the agent's graph |
Agents can also build the graph from inside the loop with the knowledge_add_entity / knowledge_add_relation / knowledge_query tools. The earlier /api/memory/kg/* endpoints in older documentation never shipped — those paths return 404.
Proactive Memory
Proactive memory lets agents autonomously surface, consolidate, and recall long-term knowledge without explicit tool calls. Each entry is stored three ways simultaneously — semantic store (text + embedding), structured KV store (memory:{id}), and knowledge graph (extracted entity / relation triples) — so agents can recall a memory by similarity, by ID, or by traversing the graph from a known entity.
Entry schema
| Field | Type | Purpose |
|---|---|---|
id | UUID | Unique memory identifier |
agent_id | UUID | Owning agent — entries are scoped per agent, no cross-agent leakage |
content | text | The memory text/fact |
source | enum | How created: auto_memorize, manual_add, … |
scope | enum | Level: user_memory, session_memory, agent_memory |
confidence | f64 (0.0–1.0) | Relevance score; decays over time |
metadata | JSON | Free-form KV; includes category (see below) |
created_at | RFC3339 timestamp | Creation time |
accessed_at | RFC3339 timestamp | Last access — drives decay |
access_count | int | Times this memory was recalled |
deleted | bool | Soft-delete flag |
embedding | vector (separate table) | For cosine similarity search |
Scopes
| Scope | Lifetime | Use |
|---|---|---|
user_memory | Persistent across sessions | User-level facts and preferences (may be shared across this user's agents). |
session_memory | Auto-deleted after session_ttl_hours (default 24) | Working notes for the current conversation. |
agent_memory | Per-agent persistent | Agent-learned behaviour and skills, isolated. |
Categories
Categories live on metadata.category. The defaults extracted by auto_memorize are: communication_style, preference, expertise, work_style, project_context, personal_detail, frustration. The category list is configurable.
Auto-consolidation
Triggered automatically every 10 auto_memorize calls per agent (no explicit cron / no user action needed). The consolidator:
- Finds duplicate or near-duplicate memories using a tiered similarity ladder:
- Substring containment (exact / superset / subset)
- Vector cosine (when embeddings stored)
- Jaccard word overlap (fallback)
- Keeps the most recently created of each duplicate cluster.
- Soft-deletes the rest and logs the merge count for audit.
The consolidator only scans the most recent 100 entries to keep the dedup pass O(n²)-safe on a hot loop. Run a manual POST /api/memory/agents/{id}/consolidate to scan the full history.
Decay
For memories not accessed for more than 1 day:
decayed_conf = original_conf × e^(-decay_rate × days_since_access)
final_conf = min(decayed_conf × (1 + log2(access_count)), 1.0)
The decay sweep is rate-limited to once per hour (checked at the top of auto_retrieve). Defaults: confidence_decay_rate = 0.01 (very slow — halves in roughly 70 days), session_ttl_hours = 24. A memory is "stale" when (a) untouched for more than a day, or (b) its scope is session_memory and the TTL has elapsed.
Per-agent vs cross-agent
All entries except user_memory are strictly per-agent. Search, consolidate, and eviction all filter on agent_id — there is no path for one agent to read another's agent_memory or session_memory. The per-agent cap defaults to 1000 entries; when exceeded the oldest / lowest-confidence entries are evicted.
API endpoints
The proactive-memory endpoints live under two roots — global and per-agent. (KV memory at /api/memory/agents/{id}/kv/* is a separate subsystem; not listed here.)
Global / cross-agent:
| Method | Path | Purpose |
|---|---|---|
| GET | /api/memory | List all proactive memories (paginated, optional ?category=) |
| POST | /api/memory | Add a memory entry |
| GET | /api/memory/search?q=&limit= | Semantic search across all entries |
| GET | /api/memory/stats | Global stats |
| GET | /api/memory/config | Get memory config |
| PATCH | /api/memory/config | Update memory config |
| POST | /api/memory/cleanup | Manual session-TTL cleanup pass |
| POST | /api/memory/decay | Manual confidence-decay pass |
| POST | /api/memory/bulk-delete | Delete multiple entries |
| PUT | /api/memory/items/{id} | Update a single memory |
| DELETE | /api/memory/items/{id} | Delete a single memory |
| GET | /api/memory/items/{id}/history | Memory edit history |
| GET | /api/memory/user/{user_id} | User-level memories across all agents |
Per-agent:
| Method | Path | Purpose |
|---|---|---|
| GET | /api/memory/agents/{id} | List the agent's memories |
| DELETE | /api/memory/agents/{id} | Reset (clear all of) the agent's memories |
| GET | /api/memory/agents/{id}/search?q= | Agent-scoped semantic search |
| GET | /api/memory/agents/{id}/stats | Per-agent stats |
| DELETE | /api/memory/agents/{id}/level/{level} | Clear by scope (session, agent, user) |
| GET | /api/memory/agents/{id}/duplicates | Find near-duplicates without deleting |
| POST | /api/memory/agents/{id}/consolidate | Trigger consolidation on the full history |
| GET | /api/memory/agents/{id}/count | Memory count |
| GET | /api/memory/agents/{id}/relations | Query the agent's knowledge graph |
| POST | /api/memory/agents/{id}/relations | Store relations |
| GET | /api/memory/agents/{id}/export | Export the agent's memories (JSON) |
| POST | /api/memory/agents/{id}/import | Import memories (JSON) |
The earlier /api/memory/proactive/* namespace was a draft surface that never shipped — use the endpoints above. If your client still calls the old paths, expect 404.
Best Practices
- Regular Compaction - Prevent sessions from growing too large
- Use Tags - Easy organization and search
- Set Decay Rate - Control memory confidence
- Monitor Usage - Track costs and usage
Troubleshooting
Slow Memory Search
# Rebuild vector index
librefang memory reindex
# Check index status
librefang memory index-status
Database Bloat
# Clean old data
librefang memory cleanup --older-than 30d
# Vacuum database
librefang memory vacuum
Memory Loss
# Check database integrity
librefang doctor
# Restore backup
librefang memory restore --backup <path>