One interface.
Every LLM backend.
SynapseKit is an async-first Python framework for RAG, agents, and graph workflows. Two hard dependencies. Plain Python you can read end to end.
Watch a request move through the graph.
Every step below is a plain Python node. Run it once per tab, then read the code that produced it.
Each node is a plain async function. Swap, extend, or step through any of them in a debugger.
What's new
The 2.x line is about trust and autonomy in production: provable agent behavior, self-managing memory, richer retrieval, local-first operation, and policy enforcement at the LLM boundary. All additive, no breaking changes.
Trust and verification
- Verifiable AgentsSigned, hash-chained audit trails with a standalone MATCH / DRIFT / UNVERIFIABLE verifier.
- GuardrailsPolicy middleware for any LLM call: block, redact, flag, or require human review, with HIPAA/GDPR/PCI-DSS rulepacks.
- NeuroSymbolicAgentThe LLM proposes constraints, a Z3, SymPy, MiniZinc, or Prolog solver checks them before you trust the answer.
- Orchestration evalDetects handoff loops, per-transfer context loss, and non-deterministic mis-routing across multi-agent runs.
Memory and retrieval
- Living MemoryAgents propose signed, diffable patches to their memory files instead of silently overwriting them.
- Property Graph RAG, WorldModelRAGVector search fused with graph traversal, plus a temporal knowledge graph with causal links.
- Personal Knowledge MeshLocal-first, incremental indexing across every project on your machine, with a CLI and MCP tools.
- Embeddings and reranker layerA provider-agnostic BaseEmbeddings contract across 9 hosted providers, plus a Reranker interface.
Agents
- AgentSwarmMarket-based routing across distributed agents: sealed-bid, Vickrey, English, and coalition auctions.
- SelfImprovingAgentEval-gated config evolution with signed patches and canary rollout. Every bad patch is blocked by the gate.
- EdgeRuntimeLocal-first inference with policy-gated cloud fallback and PII redaction before any data leaves the device.
- Dream Mode, Ambient daemonOffline reflection over past runs, and a background daemon that watches for moments to intervene.
- Code Archaeology agentReasons across a repo's history: as-of scoping, drift detection, and generated change narratives.
- Digital Twin, Hive Mode, Agent OS ShellA versioned profile of your voice for drafting in your style, multi-agent coordination, and a local agent shell.
Operations
- SynapseKit LiveA zero-dependency, real-time dashboard. Every LLM call, tool, retrieval, and cost streams to your browser.
- Official Docker imagesdocker pull ghcr.io/synapsekit/synapsekit: core and all-extras variants, published on every release.
- CAG/RAG routerRoutes between cache-augmented and retrieval-augmented generation, with a llama.cpp KV-cache backend.
- Signed agent marketplace, PC TwinEd25519-signed agent bundles with a hardened install flow, and a sandboxed environment for safe automation.
Six layers. Use one, use all.
Each layer is a normal Python import, not a hidden dependency. Take the vector store without the agent runtime, or the graph engine without the CLI.
Existing frameworks accumulate weight.
Dependency bloat, inconsistent async, and paid observability are the three complaints that come up most from teams evaluating alternatives.
Most frameworks pull in half of PyPI.
50+ dependencies for a 200 MB install is common. Every import is a surprise. SynapseKit needs only numpy and rank-bm25. Everything else is an optional extra.
Async gets bolted on, not designed in.
Partial async support is unpredictable: some methods await, others block the event loop without warning. SynapseKit is async/await native at every layer.
Cost tracking is usually a separate SaaS product.
Observability shouldn't require a subscription or an external agent. SynapseKit tracks cost, tokens, and latency out of the box, and streams it locally.
SynapseKit does the same in 10 lines.
Plain Python. No magic classes, no global state: functions you can read, debug, and extend.
Full async/await throughout, no sync/async mismatch
Token-level streaming from every provider
Swap model or provider in one line
Cost tracking on every call, no SaaS needed
No hidden chains. Every step is plain Python.
What is actually in the box.
46 LLM providers, 83 loaders, 32 vector stores, 56 tools, a guardrails layer, and a real-time dashboard. Every piece is plain Python you can read end to end.
Core architecture
Async by default, not bolted on.
Every public IO method is a coroutine. Blocking calls run through an executor. A CI gate checks this on every commit, so the async contract cannot regress silently.
Minimal footprint
Two hard dependencies.
numpy and rank-bm25. Every provider, loader, and store is an optional extra you install by name.
Output
Streaming is the default.
Token-level streaming across all 46 providers, not an opt-in mode.
Agents
ReAct or native function calling, 56 tools.
A ReAct loop that works with any LLM, or native function calling for OpenAI, Anthropic, Gemini, and Mistral. 56 built-in tools; write your own in five lines with @tool.
Orchestration
Graph workflows with typed state.
DAG-based async pipelines. Independent nodes run concurrently in waves. Conditional routing, fan-out/fan-in, human-in-the-loop, checkpointing, and Mermaid export.
Retrieval
32 vector stores, 9 embeddings providers, one interface.
From a zero-dependency in-memory store to managed cloud services, all behind VectorStore. BM25 reranking, property-graph RAG that fuses vector search with graph traversal, and federated retrieval that fans out to local and remote sources with score fusion.
Ecosystem
46 providers behind one interface.
OpenAI, Anthropic, Gemini, Ollama, Bedrock and 41 more, all implementing BaseLLM. Swap providers by changing a model string, not your call sites.
Policy enforcement
Guardrails at the LLM boundary.
Wrap any BaseLLM in a policy: block, redact, flag, or require human approval. Prompt-injection and jailbreak guards, PII redaction, HIPAA/GDPR/PCI rulepacks, and a signed audit trail.
Observability
Cost and latency, tracked automatically.
Per-call cost, tokens, and latency on every request. SynapseKit Live streams it to a local dashboard, no SaaS required.
Cost / query
$0.0012
Avg latency
1.34s
Tokens used
2.4M
Total spend
$2.87
Trust
Signed, hash-chained audit trails.
Every agent run can produce a replayable, Ed25519-signed log. A standalone verifier returns MATCH, DRIFT, or UNVERIFIABLE against pinned trusted keys.
Evaluation
EvalCI blocks regressions in CI.
A GitHub Action that runs your @eval_case suite on every pull request and fails the build on quality regression.
How we stack up.
SynapseKit, LangChain, and LlamaIndex, side by side.
| Feature | SynapseKit | LangChain | LlamaIndex |
|---|---|---|---|
| Hard dependencies | +2 | –50+ | –20+ |
| Install size | +~5 MB | –~200 MB+ | –~100 MB+ |
| Async-native | +Default | ~Partial | ~Partial |
| Streaming | +Default | ~Varies | ~Varies |
| Cost tracking | +Built-in | –SaaS add-on | –No |
| Evaluation (EvalCI) | +CLI + GitHub Action | –SaaS add-on | ~Built-in |
| Graph workflows | +Built-in | ~Separate package | –No |
| Agent federation | +Built-in | –No | –No |
| Guardrails middleware | +Built-in | –No | –No |
| Verifiable audit trails | +Signed, hash-chained | –No | –No |
| Agent memory backends | +10 built-in | ~Community plugins | ~Community plugins |
| Observability | +Prometheus + Grafana | –No | –No |
| Type safety | +Strict dataclasses | ~Partial | ~Partial |
| LLM providers | 46 | 38+ | 20+ |
| License | Apache 2.0 | MIT | MIT |
LangChain has more raw integrations and tutorials. SynapseKit optimizes for shipping and debugging a production LLM feature: readable code, predictable async behavior, and no surprise SaaS bill.
Three surfaces. One interface.
83
Loaders
32
Vector Stores
9
Embeddings & Reranker Providers
83 loaders: PDF, Word, YouTube, S3, Notion, HubSpot, BigQuery, Salesforce, Airtable, Obsidian, and more
32 vector stores: Chroma, Pinecone, Weaviate, pgvector, Redis, MongoDB Atlas, SQLiteVec, and more
Hybrid search: semantic vector search and multi-hop knowledge graph retrieval in one call
A dedicated embeddings and reranker layer, a provider-agnostic BaseEmbeddings contract across 9 hosted providers
Built-in RAG evaluation with cost and benefit tracking, alert sinks, and per-call scoring
Anthropic prompt caching via SmartContextManager, which cuts costs on repeated context
Ship with confidence.
Gate quality on every PR.
EvalCI is a GitHub Action that runs your LLM evaluation suite on every pull request, before anything merges. It catches regressions automatically, not manually.
- +Define eval cases with the @eval_case decorator
- +Compare every PR against a baseline model output
- +Block merges when quality drops below threshold
- +Track factual accuracy, faithfulness, cost per query
Your entire stack, already supported.
46 LLM providers behind one unified API. Swap without rewriting a line.
Load from anywhere. Get Documents everywhere.
Every loader returns the same Document object, whether it's a PDF, a YouTube video, a Salesforce export, or a BigQuery table. Your pipeline never needs to change.
Start local. Go prod. Zero rewrites.
Chroma for your laptop, Pinecone for production, pgvector for your existing Postgres, all behind one interface. Change one line, not your entire codebase.
One decorator. Real-world actions.
Decorate any function with @tool and your agent can call it. Browser automation, SQL queries, GitHub PRs, Slack messages: all wired up and production-tested.
Agents that remember across sessions.
Episodic memory stores what happened. Semantic memory stores what matters. Both work across SQLite, Redis, Postgres, and Firestore: start in-memory, scale out in one line.
Start in seconds.
Install only what you need. Extras are truly optional.
Full options: installation docs
Everything documented.
Nothing hidden.
All docsSynapseKit Live
Watch every LLM call, tool, retrieval, memory write, and cost stream to your browser in real time. Zero extra dependencies.
ReadRAG Guide
Pipelines, loaders, hybrid retrieval, vector stores, evaluation.
ReadAgents
Function calling, tool use, episodic memory, and AgentFederation across services.
ReadGraph Workflows
Compose pipelines as graphs. Branch, merge, loop, and run subgraphs in parallel.
ReadLLM Providers
ReasoningLLM, CostQualityRouter, streaming, structured output across all providers.
ReadEvalCI
LLM eval suites that run as a GitHub Action, so regressions get caught before merge, not in production.
ReadAPI Reference
Full reference for every public symbol, parameter, return type, and exception.
ReadQuestions
Answers to what people usually ask before adopting SynapseKit or switching from LangChain or LlamaIndex.
SynapseKit is an async-native, open-source Python framework for building LLM-powered applications. It provides RAG pipelines, ReAct agents, graph workflows, and AgentFederation with only 2 hard dependencies (numpy and rank-bm25). It supports 46 LLM providers, 83 document loaders, and 32 vector stores out of the box.