GitHub - neuml/txtai: 💡 All-in-one open-source embeddings database for semantic search, LLM orchestration and language model workflows

All-in-one embeddings database

txtai is an all-in-one embeddings database for semantic search, LLM orchestration and language model workflows.

Embeddings databases are a union of vector indexes (sparse and dense), graph networks and relational databases. This enables vector search with SQL, topic modeling, graph analysis and more.

Embeddings databases can stand on their own and/or serve as a powerful knowledge source for large language model (LLM) prompts. Build autonomous agents, retrieval augmented generation (RAG) processes and multi-model workflows.

Summary of txtai features:

🔎 Vector search with SQL, object storage, topic modeling, graph analysis and multimodal indexing
📄 Create embeddings for text, documents, audio, images and video
💡 Pipelines powered by language models that run LLM prompts, question-answering, labeling, transcription, translation, summarization and more
↪️️ Workflows to join pipelines together and aggregate business logic. txtai processes can be simple microservices or multi-model workflows.
🤖 Agents that intelligently connect embeddings, pipelines, workflows and other agents together to autonomously solve complex problems
⚙️ Build with Python or YAML. API bindings available for JavaScript, Java, Rust and Go.
☁️ Run local or scale out with container orchestration

txtai is built with Python 3.9+, Hugging Face Transformers, Sentence Transformers and FastAPI. txtai is open-source under an Apache 2.0 license.

Interested in an easy and secure way to run hosted txtai applications? Then join the txtai.cloud preview to learn more.

Why txtai?

New vector databases, LLM frameworks and everything in between are sprouting up daily. Why build with txtai?

Up and running in minutes with pip or Docker

# Get started in a couple lines
import txtai

embeddings = txtai.Embeddings()
embeddings.index(["Correct", "Not what we hoped"])
embeddings.search("positive", 1)
#[(0, 0.29862046241760254)]

Built-in API makes it easy to develop applications using your programming language of choice

# app.yml
embeddings:
    path: sentence-transformers/all-MiniLM-L6-v2

CONFIG=app.yml uvicorn "txtai.api:app"
curl -X GET "http://localhost:8000/search?query=positive"

Run local - no need to ship data off to disparate remote services
Work with micromodels all the way up to large language models (LLMs)
Low footprint - install additional dependencies and scale up when needed
Learn by example - notebooks cover all available functionality

Use Cases

The following sections introduce common txtai use cases. A comprehensive set of over 60 example notebooks and applications are also available.

Semantic Search

Build semantic/similarity/vector/neural search applications.

Traditional search systems use keywords to find data. Semantic search has an understanding of natural language and identifies results that have the same meaning, not necessarily the same keywords.

Get started with the following examples.

Notebook	Description
Introducing txtai ▶️	Overview of the functionality provided by txtai
Similarity search with images	Embed images and text into the same space for search
Build a QA database	Question matching with semantic search
Semantic Graphs	Explore topics, data connectivity and run network analysis

LLM Orchestration

Autonomous agents, retrieval augmented generation (RAG), chat with your data, pipelines and workflows that interface with large language models (LLMs).

See below to learn more.

Notebook	Description
Prompt templates and task chains	Build model prompts and connect tasks together with workflows
Integrate LLM frameworks	Integrate llama.cpp, LiteLLM and custom generation frameworks
Build knowledge graphs with LLMs	Build knowledge graphs with LLM-driven entity extraction

Agents

Agents connect embeddings, pipelines, workflows and other agents together to autonomously solve complex problems.

txtai agents are built on top of the Transformers Agent framework. This supports all LLMs txtai supports (Hugging Face, llama.cpp, OpenAI / Claude / AWS Bedrock via LiteLLM).

Retrieval augmented generation

Retrieval augmented generation (RAG) reduces the risk of LLM hallucinations by constraining the output with a knowledge base as context. RAG is commonly used to "chat with your data".

A novel feature of txtai is that it can provide both an answer and source citation.

Notebook	Description
Build RAG pipelines with txtai	Guide on retrieval augmented generation including how to create citations
How RAG with txtai works	Create RAG processes, API services and Docker instances
Advanced RAG with graph path traversal	Graph path traversal to collect complex sets of data for advanced RAG
Speech to Speech RAG ▶️	Full cycle speech to speech workflow with RAG

Language Model Workflows

Language model workflows, also known as semantic workflows, connect language models together to build intelligent applications.

While LLMs are powerful, there are plenty of smaller, more specialized models that work better and faster for specific tasks. This includes models for extractive question-answering, automatic summarization, text-to-speech, transcription and translation.

Notebook	Description
Run pipeline workflows ▶️	Simple yet powerful constructs to efficiently process data
Building abstractive text summaries	Run abstractive text summarization
Transcribe audio to text	Convert audio files to text
Translate text between languages	Streamline machine translation and language detection

Installation

The easiest way to install is via pip and PyPI

pip install txtai

Python 3.9+ is supported. Using a Python virtual environment is recommended.

See the detailed install instructions for more information covering optional dependencies, environment specific prerequisites, installing from source, conda support and how to run with containers.

Model guide

See the table below for the current recommended models. These models all allow commercial use and offer a blend of speed and performance.

Component	Model(s)
Embeddings	all-MiniLM-L6-v2
Image Captions	BLIP
Labels - Zero Shot	BART-Large-MNLI
Labels - Fixed	Fine-tune with training pipeline
Large Language Model (LLM)	Mistral 7B OpenOrca
Summarization	DistilBART
Text-to-Speech	ESPnet JETS
Transcription	Whisper
Translation	OPUS Model Series

Models can be loaded as either a path from the Hugging Face Hub or a local directory. Model paths are optional, defaults are loaded when not specified. For tasks with no recommended model, txtai uses the default models as shown in the Hugging Face Tasks guide.

See the following links to learn more.

Powered by txtai

The following applications are powered by txtai.

Application	Description
txtchat	Retrieval Augmented Generation (RAG) powered search
paperai	Semantic search and workflows for medical/scientific papers
codequestion	Semantic search for developers
tldrstory	Semantic search for headlines and story text

In addition to this list, there are also many other open-source projects, published research and closed proprietary/commercial projects that have built on txtai in production.

Documentation

Full documentation on txtai including configuration settings for embeddings, pipelines, workflows, API and a FAQ with common questions/issues is available.

Contributing

For those who would like to contribute to txtai, please see this guide.

Name		Name	Last commit message	Last commit date
Latest commit History 1,453 Commits
.github/workflows		.github/workflows
docker		docker
docs		docs
examples		examples
src/python/txtai		src/python/txtai
test/python		test/python
.coveragerc		.coveragerc
.gitignore		.gitignore
.pre-commit-config.yaml		.pre-commit-config.yaml
.pylintrc		.pylintrc
CITATION.cff		CITATION.cff
LICENSE		LICENSE
Makefile		Makefile
README.md		README.md
apps.jpg		apps.jpg
demo.gif		demo.gif
logo.png		logo.png
mkdocs.yml		mkdocs.yml
pyproject.toml		pyproject.toml
setup.py		setup.py

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Why txtai?

Use Cases

Semantic Search

LLM Orchestration

Agents

Retrieval augmented generation

Language Model Workflows

Installation

Model guide

Powered by txtai

Further Reading

Documentation

Contributing

About

Releases 43

Packages

Used by 367

Contributors 16

Languages

License

neuml/txtai

Folders and files

Latest commit

History

Repository files navigation

Why txtai?

Use Cases

Semantic Search

LLM Orchestration

Agents

Retrieval augmented generation

Language Model Workflows

Installation

Model guide

Powered by txtai

Further Reading

Documentation

Contributing

About

Topics

Resources

License

Stars

Watchers

Forks

Releases 43

Packages 0

Used by 367

Contributors 16

Languages

Packages