Hindsight: Agent Memory That Learns Over Time
GitHub Repo
MIT
October 2, 2026 at 09:19 AM
0 views

Hindsight: Agent Memory That Learns Over Time

@vectorize-ioProject Author

What Hindsight is

Hindsight is an agent memory system built by Vectorize. Its README frames the problem directly: most agent memory systems focus on recalling conversation history, while Hindsight is aimed at agents that learn, not just remember. It runs as a server (or embedded in a Python process) that agents write facts and experiences into, and later query for relevant context or for a reasoned answer over everything they have accumulated.

The project is MIT-licensed, written primarily in Python, and ships clients for Python, Node.js/TypeScript and Go plus a CLI and a REST API. It targets developers building conversational assistants that need per-user memory, and autonomous agents ("AI employees" in the README's words) that should change behavior based on feedback. The README is candid that for simple n8n-style workflows it "may be overkill".

How it works

Memories live in banks, isolated stores intended to be one "brain" per user, agent or project, with no cross-bank leakage. Hindsight organizes what goes into a bank into four memory types, which the README calls biomimetic data structures:

  • World facts: facts about the world.
  • Experiences: the agent's own experiences.
  • Observations: consolidated, evidence-backed beliefs formed from many memories.
  • Mental models: learned understanding synthesized from observations and facts.

Everything flows through three operations. retain uses an LLM to extract key facts, temporal data, entities and relationships from input, then normalizes them into canonical entities, time series and search indexes with sparse and dense vector representations. recall runs four retrieval strategies in parallel (semantic vector similarity, BM25 keyword matching, graph traversal over entity, temporal and causal links, and time-range filtering), merges results with reciprocal rank fusion, reranks them with a cross-encoder and trims them to a token budget. reflect performs a deeper analysis over existing memories to answer questions that need reasoning rather than lookup, shaped by per-bank disposition traits such as skepticism, literalism and empathy.

In the background, related facts are consolidated into observations that keep exact supporting quotes and a proof count, and are refined rather than overwritten when new evidence arrives. Mental models are standing answers to questions you define ("What are this user's preferences?") that Hindsight rewrites as the bank learns; reading one is a plain database read with no LLM call. Knowledge pages wrap mental models as wiki-like documents that can be projected to disk as Markdown.

Key features

  • Hybrid retrieval: semantic, keyword, graph and temporal recall fused with RRF and cross-encoder reranking.
  • 25+ LLM providers: hosted (OpenAI, Anthropic, Gemini, Groq, Bedrock, Vertex AI, DeepSeek and others), local (Ollama, LM Studio, llama.cpp), any OpenAI-compatible endpoint, and gateways like LiteLLM. Existing ChatGPT, Claude, Cursor and GitHub Copilot subscriptions can be used without an API key.
  • LLM wrapper: wrap_openai() and wrap_anthropic() recall memories before each call and retain the conversation after it.
  • Built-in MCP server: one endpoint per bank at /mcp/{bank_id}/, enabled by default, exposing retain, recall and reflect as tools.
  • Coding agent memory: a package that builds a per-repo bank from git history and past sessions for Claude Code, Codex CLI, Cursor CLI, opencode and others.
  • 60+ integrations: LangGraph, LlamaIndex, CrewAI, Pydantic AI, OpenAI Agents SDK, Google ADK, Vercel AI SDK, n8n, Dify and more.
  • Multilingual by default: facts stay in their original language and entities keep their native script.
  • Memory Defense: an opt-in per-bank policy that scans retained content for secrets and PII against 45 patterns and redacts or blocks matches.
  • Production tooling: hierarchical configuration (global, tenant, bank), Prometheus metrics, an admin CLI, and webhooks for lifecycle events.

Getting started

The recommended path is Docker:

export OPENAI_API_KEY=sk-xxx

docker run -it --pull always --name hindsight --restart unless-stopped -p 8888:8888 -p 9999:9999 \
  -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
  -v hindsight-data:/home/hindsight/.pg0 \
  ghcr.io/vectorize-io/hindsight:latest

The API listens on http://localhost:8888 and the UI on http://localhost:9999. Alternatively, install it bare metal:

pip install hindsight-api
export HINDSIGHT_API_LLM_API_KEY=sk-xxx

hindsight-api

Then connect a client and use the three operations:

pip install hindsight-client -U
from hindsight_client import Hindsight

client = Hindsight(base_url="http://localhost:8888")

# Retain: Store information
client.retain(bank_id="my-bank", content="Alice works at Google as a software engineer")

# Recall: Search memories
client.recall(bank_id="my-bank", query="What does Alice do?")

# Reflect: Generate disposition-aware response
client.reflect(bank_id="my-bank", query="Tell me about Alice")

For scripts and notebooks there is an embedded mode with no separate server (pip install hindsight-all -U), plus a Helm chart for Kubernetes and a Docker Compose setup for external PostgreSQL. To give CLI coding agents project memory:

npx @vectorize-io/hindsight-coding-agents install all          # every detected agent, wired natively

Use cases

  • Personalized chatbots: retain user inputs and tool calls with metadata that scopes memories to a user, then filter on recall.
  • Coding agents with project memory: let Claude Code or Codex start each session with knowledge of a repo's architecture, conventions and in-flight work.
  • Reflective task agents: the README's examples include a project-manager agent reflecting on risks, a sales agent reflecting on which outreach messages got responses, and a support agent spotting questions the docs do not answer.
  • Drop-in memory for existing apps: wrap an OpenAI or Anthropic client to add memory without restructuring code.
  • Shared memory across tools: expose a bank over MCP so several MCP clients read and write the same memory.

How it compares

The README positions Hindsight against RAG and knowledge-graph approaches to memory, claiming it "eliminates the shortcomings" of both, and links a dedicated RAG-vs-memory page in its docs. On benchmarks, the README claims state-of-the-art results on LongMemEval and says its numbers were independently reproduced by researchers at Virginia Tech's Sanghani Center and The Washington Post, while noting that other vendors' scores are self-reported. Live results are published on a separate benchmarks site, and the approach is described in an arXiv paper. These are vendor claims; teams evaluating agent memory should test against their own workloads. In practice, the closest comparison is other agent-memory layers that sit between an agent and a vector store; Hindsight's differentiators are the reflect operation, background consolidation into observations, and mental models that can be read without an LLM call.

Things to know before adopting

  • LLM costs: retain uses an LLM for extraction and reflect uses one for reasoning, so ingest volume translates directly into model calls.
  • Storage: production deployments use PostgreSQL with pgvector, or Oracle AI Database 23ai; the default Docker image uses an embedded database (pg0) stored in a volume.
  • Platform support: Linux, macOS and Windows are supported; on Intel Macs the README says to use the hindsight-all-slim package instead.
  • Hosted option: Hindsight Cloud is a usage-based managed service from Vectorize. Note that the LLM wrapper defaults to Hindsight Cloud unless you pass hindsight_api_url for a self-hosted server.
  • Commercial backing: the project is built by Vectorize.io, which also sells Cloud and Enterprise tiers; the open-source server is MIT-licensed.
  • Age: the repository is about a year old, so expect APIs and integrations to keep moving.

Project activity

As of October 2026 the repository has about 44,500 stars on GitHub. It was created on 2025-10-30, is written primarily in Python, and is released under the MIT license. Packages are published to PyPI (hindsight-api, hindsight-client) and npm (@vectorize-io/hindsight-client). Source code is at github.com/vectorize-io/hindsight and documentation is at hindsight.vectorize.io.

Enjoying this project?

Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.

Project
hindsight-agent-memory
Created
October 2
Last Updated
October 2, 2026 at 09:19 AM

Find more projects like this

One email a week: new and trending developer tools, fresh comparisons, and what shipped. Unsubscribe in one click.