Copied to clipboard
v2.xGuardrails, orchestration eval, embeddings layer

One interface.
Every LLM backend.

SynapseKit is an async-first Python framework for RAG, agents, and graph workflows. Two hard dependencies. Plain Python you can read end to end.

pip install synapsekitView source
OpenAIAnthropicGeminiMistralOllama+ 41 moreBaseLLMone interfaceRAGRetrievalAgentsReAct, toolsGraphWorkflows
0
LLM providers
0
vector stores
0
data loaders
0
hard dependencies

Watch a request move through the graph.

Every step below is a plain Python node. Run it once per tab, then read the code that produced it.

rag_pipeline.py
done
User Querynatural languageLoader83 sourcesVector Store32 backendsLLM46 providersAnswerstreaming
output
source
from synapsekit import RAG

rag = RAG(model="gpt-4o-mini", api_key="sk-...")
await rag.add("Q1 2025 earnings report...", metadata={"source": "10-Q"})

async for token in rag.stream("What changed in Q1 2025?"):
    print(token, end="", flush=True)

Each node is a plain async function. Swap, extend, or step through any of them in a debugger.

What's new

The 2.x line is about trust and autonomy in production: provable agent behavior, self-managing memory, richer retrieval, local-first operation, and policy enforcement at the LLM boundary. All additive, no breaking changes.

Trust and verification

Memory and retrieval

Agents

Operations

Six layers. Use one, use all.

Each layer is a normal Python import, not a hidden dependency. Take the vector store without the agent runtime, or the graph engine without the CLI.

01
Application layer
Your code: scripts, FastAPI services, Jupyter notebooks
ScriptsFastAPINotebooks
02
Tools and CLI
CLI commands, PromptHub versioned prompts, 56 built-in tools
CLIPromptHub56 tools
03
Core orchestration
RAG pipelines, ReAct and function-calling agents, graph workflows, guardrails, AgentSwarm
RAGPipelineAgentStateGraphGuardrails
04
Data and memory
83 loaders, 32 vector stores, 9 embeddings/reranker providers, 10 memory backends
83 loaders32 stores10 backends
05
LLM providers
46 providers, one BaseLLM interface, streaming and structured output by default
OpenAIAnthropicGemini+43 more
06
Async core
numpy and rank-bm25 only. Fully async, sync wrappers available
async/awaitnumpyrank-bm25
46LLM providers
83Loaders
32Vector stores
56Tools
Read the architecture docs

Existing frameworks accumulate weight.

Dependency bloat, inconsistent async, and paid observability are the three complaints that come up most from teams evaluating alternatives.

Most frameworks pull in half of PyPI.

50+ dependencies for a 200 MB install is common. Every import is a surprise. SynapseKit needs only numpy and rank-bm25. Everything else is an optional extra.

Async gets bolted on, not designed in.

Partial async support is unpredictable: some methods await, others block the event loop without warning. SynapseKit is async/await native at every layer.

Cost tracking is usually a separate SaaS product.

Observability shouldn't require a subscription or an external agent. SynapseKit tracks cost, tokens, and latency out of the box, and streams it locally.

SynapseKit does the same in 10 lines.

Plain Python. No magic classes, no global state: functions you can read, debug, and extend.

agent_example.py
from synapsekit import agent, tool
 
@tool
def get_weather(city: str) -> str:
return f"Sunny, 22°C in {city}"
 
# One line to create a full agent
my_agent = agent(
model="gpt-4o-mini",
api_key="sk-...",
tools=[get_weather],
)
 
print(my_agent.run("What's the weather in Tokyo?"))
✓

Full async/await throughout, no sync/async mismatch

✓

Token-level streaming from every provider

✓

Swap model or provider in one line

✓

Cost tracking on every call, no SaaS needed

✓

No hidden chains. Every step is plain Python.

What is actually in the box.

46 LLM providers, 83 loaders, 32 vector stores, 56 tools, a guardrails layer, and a real-time dashboard. Every piece is plain Python you can read end to end.

Core architecture

Async by default, not bolted on.

Every public IO method is a coroutine. Blocking calls run through an executor. A CI gate checks this on every commit, so the async contract cannot regress silently.

# streaming is the default, not an add-on
async for token in llm.stream(prompt):
print(token, end="")

Minimal footprint

Two hard dependencies.

numpy and rank-bm25. Every provider, loader, and store is an optional extra you install by name.

0hard dependencies
SynapseKit2
LangChain50+
LlamaIndex20
numpyrank-bm25

Output

Streaming is the default.

Token-level streaming across all 46 providers, not an opt-in mode.

Agents

ReAct or native function calling, 56 tools.

A ReAct loop that works with any LLM, or native function calling for OpenAI, Anthropic, Gemini, and Mistral. 56 built-in tools; write your own in five lines with @tool.

Orchestration

Graph workflows with typed state.

DAG-based async pipelines. Independent nodes run concurrently in waves. Conditional routing, fan-out/fan-in, human-in-the-loop, checkpointing, and Mermaid export.

Retrieval

32 vector stores, 9 embeddings providers, one interface.

From a zero-dependency in-memory store to managed cloud services, all behind VectorStore. BM25 reranking, property-graph RAG that fuses vector search with graph traversal, and federated retrieval that fans out to local and remote sources with score fusion.

ChromaPineconeQdrantWeaviatepgvectorTurbopufferDeepLake+25 more

Ecosystem

46 providers behind one interface.

OpenAI, Anthropic, Gemini, Ollama, Bedrock and 41 more, all implementing BaseLLM. Swap providers by changing a model string, not your call sites.

OpenAIAnthropicGeminiOllamaBedrockMistralGroqTogetherDeepSeekCohereFireworksReplicateHuggingFacexAIvLLMLM StudioWriterNovitaAzureVertex+26 more

Policy enforcement

Guardrails at the LLM boundary.

Wrap any BaseLLM in a policy: block, redact, flag, or require human approval. Prompt-injection and jailbreak guards, PII redaction, HIPAA/GDPR/PCI rulepacks, and a signed audit trail.

block
reject and log
redact
strip PII, continue
flag
pass through, mark for review
require_human
hold for approval

Observability

Cost and latency, tracked automatically.

Per-call cost, tokens, and latency on every request. SynapseKit Live streams it to a local dashboard, no SaaS required.

Cost / query

$0.0012

Avg latency

1.34s

Tokens used

2.4M

Total spend

$2.87

Trust

Signed, hash-chained audit trails.

Every agent run can produce a replayable, Ed25519-signed log. A standalone verifier returns MATCH, DRIFT, or UNVERIFIABLE against pinned trusted keys.

Evaluation

EvalCI blocks regressions in CI.

A GitHub Action that runs your @eval_case suite on every pull request and fails the build on quality regression.

How we stack up.

SynapseKit, LangChain, and LlamaIndex, side by side.

FeatureSynapseKitLangChainLlamaIndex
Hard dependencies+2–50+–20+
Install size+~5 MB–~200 MB+–~100 MB+
Async-native+Default~Partial~Partial
Streaming+Default~Varies~Varies
Cost tracking+Built-in–SaaS add-on–No
Evaluation (EvalCI)+CLI + GitHub Action–SaaS add-on~Built-in
Graph workflows+Built-in~Separate package–No
Agent federation+Built-in–No–No
Guardrails middleware+Built-in–No–No
Verifiable audit trails+Signed, hash-chained–No–No
Agent memory backends+10 built-in~Community plugins~Community plugins
Observability+Prometheus + Grafana–No–No
Type safety+Strict dataclasses~Partial~Partial
LLM providers4638+20+
LicenseApache 2.0MITMIT

LangChain has more raw integrations and tutorials. SynapseKit optimizes for shipping and debugging a production LLM feature: readable code, predictable async behavior, and no surprise SaaS bill.

Three surfaces. One interface.

83

Loaders

32

Vector Stores

9

Embeddings & Reranker Providers

  • 83 loaders: PDF, Word, YouTube, S3, Notion, HubSpot, BigQuery, Salesforce, Airtable, Obsidian, and more

  • 32 vector stores: Chroma, Pinecone, Weaviate, pgvector, Redis, MongoDB Atlas, SQLiteVec, and more

  • Hybrid search: semantic vector search and multi-hop knowledge graph retrieval in one call

  • A dedicated embeddings and reranker layer, a provider-agnostic BaseEmbeddings contract across 9 hosted providers

  • Built-in RAG evaluation with cost and benefit tracking, alert sinks, and per-call scoring

  • Anthropic prompt caching via SmartContextManager, which cuts costs on repeated context

Ship with confidence.
Gate quality on every PR.

EvalCI is a GitHub Action that runs your LLM evaluation suite on every pull request, before anything merges. It catches regressions automatically, not manually.

  • +Define eval cases with the @eval_case decorator
  • +Compare every PR against a baseline model output
  • +Block merges when quality drops below threshold
  • +Track factual accuracy, faithfulness, cost per query
View EvalCI docs
evalci, PR #247
$ synapsekit eval run --suite tests/evals/
Running 24 eval cases against gpt-4o-mini.
Comparing against baseline (main@f3a9c1b).
PASS factual_accuracy 0.94 (+0.02)
PASS context_relevance 0.91 (+0.00)
PASS answer_faithfulness 0.88 (+0.03)
PASS cost_per_query $0.0012 (-8%)
24/24 passed, 0 regressions. GitHub PR: approved
▋

Your entire stack, already supported.

46 LLM providers behind one unified API. Swap without rewriting a line.

OpenAIAnthropicGoogle GeminiOllamaAWS BedrockCohereMistralxAI GrokTogether AIDeepSeekGroqReplicateHuggingFaceWriterNovitaLM StudioGPT4AllvLLMAzure OpenAIVertex AIFireworksPerplexityAnyscaleDeepInfraOpenRouterCerebrasAI21Cloudflare AIAleph AlphaVoyage AIMosaicMLPredibaseDatabricksOpenAIAnthropicGoogle GeminiOllamaAWS BedrockCohereMistralxAI GrokTogether AIDeepSeekGroqReplicateHuggingFaceWriterNovitaLM StudioGPT4AllvLLMAzure OpenAIVertex AIFireworksPerplexityAnyscaleDeepInfraOpenRouterCerebrasAI21Cloudflare AIAleph AlphaVoyage AIMosaicMLPredibaseDatabricks
Data Loaders83

Load from anywhere. Get Documents everywhere.

Every loader returns the same Document object, whether it's a PDF, a YouTube video, a Salesforce export, or a BigQuery table. Your pipeline never needs to change.

docs = await PDFLoader('report.pdf').load()
PDFPDF
YouTubeYouTube
S3S3
NotionNotion
HubSpotHubSpot
BigQueryBigQuery
SalesforceSalesforce
MongoDBMongoDB
Vector Stores32

Start local. Go prod. Zero rewrites.

Chroma for your laptop, Pinecone for production, pgvector for your existing Postgres, all behind one interface. Change one line, not your entire codebase.

store = ChromaVectorStore() # swap to Pinecone later
ChromaChroma
PineconePinecone
WeaviateWeaviate
PostgresPostgres
RedisRedis
MongoDBMongoDB
QdrantQdrant
OpenSearchOpenSearch
Agent Tools56

One decorator. Real-world actions.

Decorate any function with @tool and your agent can call it. Browser automation, SQL queries, GitHub PRs, Slack messages: all wired up and production-tested.

@tool async def query_db(sql: str) -> str: ...
GitHubGitHub
SlackSlack
StripeStripe
TwilioTwilio
JiraJira
NotionNotion
LinearLinear
AWS LambdaAWS Lambda
Memory Backends10

Agents that remember across sessions.

Episodic memory stores what happened. Semantic memory stores what matters. Both work across SQLite, Redis, Postgres, and Firestore: start in-memory, scale out in one line.

agent = Agent(memory=RedisMemory(url=REDIS_URL))
SQLiteSQLite
RedisRedis
PostgresPostgres
In-memoryIn-memory

Start in seconds.

Install only what you need. Extras are truly optional.

terminal
# OpenAIpip install synapsekit[openai]
# Anthropicpip install synapsekit[anthropic]
# Ollama (local)pip install synapsekit[ollama]
# Observabilitypip install synapsekit[observe]
# Everythingpip install synapsekit[all]

Full options: installation docs

Everything documented.
Nothing hidden.

All docs

Up in 5 minutes

Quickstart

Build your first RAG pipeline or agent. pip install, configure a provider, ship.

Read the guide
from synapsekit import RAGPipeline
pipeline = RAGPipeline(llm=llm, store=store)
result = await pipeline.query("How does X work?")
import synapsekit.live as live
Glass-box dashboard

SynapseKit Live

Watch every LLM call, tool, retrieval, memory write, and cost stream to your browser in real time. Zero extra dependencies.

Read
loader = PDFLoader("paper.pdf")
docs = await loader.load()
Retrieval-augmented generation

RAG Guide

Pipelines, loaders, hybrid retrieval, vector stores, evaluation.

Read
agent = Agent(llm=llm, tools=[search, sql])
result = await agent.run(task)
ReAct, tools, memory

Agents

Function calling, tool use, episodic memory, and AgentFederation across services.

Read
graph = Graph()
graph.add_edge(fetch, summarize)
DAG, parallel, conditional

Graph Workflows

Compose pipelines as graphs. Branch, merge, loop, and run subgraphs in parallel.

Read
llm = LLM(model="gpt-4o") # OpenAI
llm = LLM(model="claude-sonnet-5") # Anthropic
46 providers, one interface

LLM Providers

ReasoningLLM, CostQualityRouter, streaming, structured output across all providers.

Read
@eval_case
def test_summary_quality():
Quality gates on every PR

EvalCI

LLM eval suites that run as a GitHub Action, so regressions get caught before merge, not in production.

Read
# Auto-generated from source
# Searchable, versioned, always current
Every class, every method

API Reference

Full reference for every public symbol, parameter, return type, and exception.

Read

Questions

Answers to what people usually ask before adopting SynapseKit or switching from LangChain or LlamaIndex.

SynapseKit is an async-native, open-source Python framework for building LLM-powered applications. It provides RAG pipelines, ReAct agents, graph workflows, and AgentFederation with only 2 hard dependencies (numpy and rank-bm25). It supports 46 LLM providers, 83 document loaders, and 32 vector stores out of the box.