Search, RAG & memory, with a verdict on each.

Getting a model to answer from your own documents, and finding the right passage fast enough to put in a prompt.

The honest caveat for most small teams: Postgres with pgvector handles a few million vectors and you already run Postgres.

21 tracked3 we would adopt2 with a licence to check

Moving fastest this week: claude-mem (+3.6k), ragflow (+3.2k), langchain (+2.3k).

Adopt now

3

Mature enough to put in production this quarter.

Pilot it

12

Worth a timeboxed spike before you bet on it.

ragflow90k
infiniflow
Pilot it+3.2k this week

A document-understanding engine that parses messy PDFs and scans properly before answering questions about them.

Instead of: Paying a vendor per page for document extraction.

Apache-2.0Updated today
Pilot it
mem064k
mem0ai
Pilot it+2.3k this week

A memory layer that lets an AI assistant remember facts about a user across separate conversations.

Instead of: Stuffing the entire conversation history into every prompt and paying for the tokens.

Apache-2.0Updated 2 days ago
Pilot it
langchain145k
langchain-ai
Pilot it+2.3k this week

The most widely used Python toolkit for chaining model calls, tools and retrieval into an application.

Instead of: Writing your own provider adapters, retry logic and tool-calling loop.

MITUpdated today
Pilot it
cognee30k
topoteretes
Pilot it+774 this week

Self-hosted knowledge graph engine for AI agents

Instead of: Manual memory management or paid AI platforms

Apache-2.0Updated today
Pilot it
qdrant34k
qdrant
Pilot it+597 this week

Self-hosted vector database for storing and searching embeddings, built in Rust.

Instead of: Manual embedding storage or paid vector DBs like Pinecone when budget or control matters.

Apache-2.0Updated today
Pilot it
llama_index52k
run-llama
Pilot it+704 this week

A Python toolkit focused on the retrieval half of AI apps: loading documents, chunking, indexing and querying them.

Instead of: Writing your own chunking, embedding and reranking pipeline.

MITUpdated today
Pilot it
langchain4j13k
langchain4j
Pilot it+236 this week

Java library for LLM-powered apps

Instead of: Manual LLM integration or paid APIs

Apache-2.0Updated today
Pilot it
zvec16k
alibaba
Pilot it+221 this week

In-process vector database for search

Instead of: Paid vector databases or manual implementations

Apache-2.0Updated 2 days ago
Pilot it
txtai13k
neuml
Pilot it+148 this week

AI framework for semantic search and language models

Instead of: manual search and NLP workflows

Apache-2.0Updated 3 days ago
Pilot it
Flowise55k
FlowiseAI
Pilot it+374 this week

Build LLM chains and agents by dragging boxes on a canvas, then expose them as an API.

Instead of: A developer's time building the first version of an internal AI tool.

Custom / check the LICENSE fileLast push 17 days ago
Pilot it
LEANN13k
StarTrail-org
Pilot it+102 this week

Private RAG search with storage savings

Instead of: Paid search APIs or manual indexing

MITUpdated today
Pilot it
weaviate17k
weaviate
Pilot it+95 this week

Vector database for search and filtering

Instead of: Manual vector search implementations

BSD-3-ClauseUpdated today
Pilot it

Watch, or skip

6

Real software; just not where a small team's next hundred hours should go.

Search, RAG & memory, in short

What are the best open source RAG and vector search tools right now?
We would put anything-llm, PaddleOCR, meilisearch into production this quarter. anything-llm — A single application that turns a folder of your documents into a chat assistant, as a desktop app or a hosted install. The full list below ranks 21 by momentum, size and how recently they shipped.
Which of these can I use in a commercial product?
19 of the 21 carry a permissive licence (MIT, Apache-2.0, BSD). 2 do not — meilisearch (Custom / check the LICENSE file), Flowise (Custom / check the LICENSE file). Copyleft and source-available licences carry obligations when you ship them inside something you sell. Read the LICENSE file; this is a reading of the label, not legal advice.
How is this list ranked?
By real week-over-week star movement first, with a size floor so an established tool is not buried by a two-week-old project, and a hard penalty for anything that has stopped shipping. Verdicts are an editorial call for a team of 2-20 with limited engineering hours — not a judgement on the software, and not the same call we would make for a lab.

Which of these matters to your business?

Tell us what you sell in one sentence and we will hand you three specific moves — priced per month, with the arithmetic shown. Free, no signup.

Give me three moves

Stars, forks, licence and last-push data come from the public GitHub API. Verdicts are NoizeOff's editorial opinion for a team of 2–20, not advice from any project's maintainers, and not legal advice on licensing.

All categories · Adoption Radar · Home