LangChain is not the same thing as LlamaIndex, even though both get described as “frameworks for building LLM apps.” LangChain is a general-purpose orchestration layer — chains, agents, tools, memory, and connectors for arbitrary LLM workflows. LlamaIndex is a data framework — indexing, retrieval, and query engines for getting your own documents into an LLM’s context. Haystack is a pipeline framework — explicit, typed, composable nodes for search and retrieval-heavy applications, originally from deepset.
Those three definitions overlap enough that a newcomer can read all three landing pages and still not know which to install. This post approaches the choice the way you’d debug a misfiring integration: start from the symptom you’re seeing in your own planning, work back to the root cause in how each framework is designed, and apply the fix — which here means picking the right tool and knowing when to drop it.
Symptom: You can’t tell what problem you’re solving because every tutorial shows three different architectures
Cause: Framework marketing converges on the same diagram. A user prompt on the left, an LLM in the middle, a knowledge base on the right. That diagram hides the differences that matter, which are: where does state live, who controls the control flow, and how opinionated is the abstraction.
Fix: Classify your application before you pick a framework. There are three broad shapes:
- Workflow-heavy, retrieval-light. Multi-step reasoning, tool calls, conditional branching, multiple LLM calls in sequence. Example: an agent that reads a support ticket, classifies it, calls an API, drafts a response, and escalates if needed.
- Retrieval-heavy, workflow-light. The core operation is “answer a question using these documents.” Example: a policy Q&A bot over an internal wiki.
- Pipeline-heavy, mixed. A fixed sequence of processing steps — clean, split, embed, retrieve, rerank, generate, post-process — where each step is a named node you can inspect.
LangChain leans toward the first. LlamaIndex leans toward the second. Haystack leans toward the third. That is the whole decision in one sentence, and the rest of this post is how to turn it into code.
Symptom: You picked a framework and now everything you write is a wrapper around a wrapper around a prompt
Cause: Framework abstractions are useful until they obscure what is being sent to the model. LangChain in particular has historically exposed several overlapping ways to do the same thing (chains, LCEL, agents, tools as decorators, tools as classes), which makes it easy to stack abstractions until a two-line prompt becomes a fifteen-line construction and you lose track of what the actual API call looks like.
Fix: Prefer the minimal abstraction that solves your problem. All three frameworks support a “give me the raw prompt, give me the raw completion” mode. Start there, add abstraction only when you have a concrete reason.
Here is the same retrieval-augmented generation (RAG) question in each framework, kept deliberately minimal so you can compare the mental models.
LangChain (LCEL)
# pip install langchain langchain-openai langchain-community faiss-cpu
from langchain_community.vectorstores import FAISS
from langchain_openai import OpenAIEmbeddings, ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnablePassthrough
vectorstore = FAISS.from_texts(
["Refunds are processed within 5 business days.",
"Enterprise plans include SSO at no extra cost."],
embedding=OpenAIEmbeddings(model="text-embedding-3-small"),
)
retriever = vectorstore.as_retriever(search_kwargs={"k": 2})
prompt = ChatPromptTemplate.from_template(
"Answer using only the context below.\n\nContext:\n{context}\n\nQuestion: {question}"
)
llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)
chain = (
{"context": retriever, "question": RunnablePassthrough()}
| prompt
| llm
| StrOutputParser()
)
print(chain.invoke("How long do refunds take?"))
The | pipe syntax is LangChain Expression Language. Each component is a Runnable, and the whole chain is itself a Runnable — which means you can call .invoke, .batch, .stream, or .astream_events on it uniformly. That uniformity is LangChain’s real value proposition: anything that is a Runnable composes with anything else that is a Runnable.
LlamaIndex
# pip install llama-index llama-index-llms-openai llama-index-embeddings-openai
from llama_index.core import VectorStoreIndex, Document, Settings
from llama_index.llms.openai import OpenAI
from llama_index.embeddings.openai import OpenAIEmbeddings
Settings.llm = OpenAI(model="gpt-4o-mini", temperature=0)
Settings.embed_model = OpenAIEmbeddings(model="text-embedding-3-small")
documents = [
Document(text="Refunds are processed within 5 business days."),
Document(text="Enterprise plans include SSO at no extra cost."),
]
index = VectorStoreIndex.from_documents(documents)
query_engine = index.as_query_engine(similarity_top_k=2)
response = query_engine.query("How long do refunds take?")
print(response)
Notice how much is implicit. Settings is a global configuration object. from_documents handles chunking, embedding, and storage in one call. The query engine embeds a default prompt, a default retriever, and a default response synthesizer. That is LlamaIndex’s core bet: the common RAG path should require almost no configuration, and the uncommon path should be reachable by swapping named sub-components.
Haystack
# pip install haystack-ai
from haystack import Pipeline, Document
from haystack.components.embedders import OpenAIDocumentEmbedder, OpenAITextEmbedder
from haystack.components.writers import DocumentWriter
from haystack.components.retrievers.in_memory import InMemoryEmbeddingRetriever
from haystack.components.builders import PromptBuilder
from haystack.components.generators import OpenAIGenerator
from haystack.document_stores.in_memory import InMemoryDocumentStore
store = InMemoryDocumentStore()
pipeline = Pipeline()
pipeline.add_component("embed_docs", OpenAIDocumentEmbedder(model="text-embedding-3-small"))
pipeline.add_component("write", DocumentWriter(document_store=store))
pipeline.add_component("embed_query", OpenAITextEmbedder(model="text-embedding-3-small"))
pipeline.add_component("retriever", InMemoryEmbeddingRetriever(document_store=store, top_k=2))
pipeline.add_component("prompt", PromptBuilder(
template="Answer using only the context.\n\n{% for doc in documents %}{{ doc.content }}\n{% endfor %}\nQuestion: {{ question }}"
))
pipeline.add_component("llm", OpenAIGenerator(model="gpt-4o-mini"))
pipeline.connect("embed_docs.documents", "write.documents")
pipeline.connect("embed_query.embedding", "retriever.query_embedding")
pipeline.connect("retriever.documents", "prompt.documents")
pipeline.connect("prompt.prompt", "llm.prompt")
pipeline.run({"embed_docs": {"documents": [
Document(content="Refunds are processed within 5 business days."),
Document(content="Enterprise plans include SSO at no extra cost."),
]}})
result = pipeline.run({"embed_query": {"text": "How long do refunds take?"},
"prompt": {"question": "How long do refunds take?"}})
print(result["llm"]["replies"][0])
Haystack is verbose on purpose. Every connection is explicit, every node is named, and every input/output socket is typed. The payoff is that Haystack pipelines serialize to YAML and can be inspected, version-controlled, and visualized without executing them.
Symptom: You built a prototype in one framework and now you cannot add the feature you need
Cause: You picked the framework for the demo, not for the hard six months after the demo. The features that matter later — custom retrievers, rerankers, sub-question decomposition, streaming with tool calls, structured output validation, hybrid search — are distributed unevenly across the three.
Fix: Match the framework to the feature you know you will need, not the feature you have today. The table below summarizes where each framework’s gravity sits.
| Capability | LangChain | LlamaIndex | Haystack |
|---|---|---|---|
| Multi-step agent loops | Strong (LangGraph for stateful graphs) | Moderate (agent modules) | Moderate (agent components) |
| Document indexing & chunking | Basic | Strong, many node parsers | Strong, explicit components |
| Retrieval strategies | Basic retrievers, easy to swap | Extensive (auto-merging, recursive, hybrid) | Explicit, composable retrievers |
| Reranking | Available | Available, tight integration | Available, first-class node |
| Structured output | with_structured_output on chat models | Pydantic programs | Via generator + prompt |
| Pipeline serialization | Limited | Limited | YAML round-trip |
| Ecosystem breadth of integrations | Very large | Large for data sources | Moderate, focused on search |
| Opinionated? | Low | Medium | High |
If your app is “an agent that does things,” LangGraph (LangChain’s graph execution layer) is currently the most mature option of the three for stateful branching with checkpoints and time-travel debugging. If your app is “answer questions over a corpus,” LlamaIndex’s index and query abstractions are the shortest path. If your app is “a production search-and-answer service that a platform team needs to audit,” Haystack’s explicit pipelines age better.
Symptom: You want to switch frameworks midway and cannot tell what would break
Cause: Framework code is sticky. Once you have written prompt templates, retrievers, and evaluators against one framework’s types, porting means rewriting every integration point, not just the top-level orchestration.
Fix: Isolate the boundaries. Even inside a framework, keep three things framework-agnostic:
- Prompts as plain strings or files. All three frameworks accept a plain string. Do not use a framework-specific prompt DSL unless you need variable validation.
- Retrieval behind your own interface. Wrap the framework’s retriever in a function
def retrieve(query: str, k: int) -> list[dict]that returns dicts withtextandmetadata. Both LangChain and LlamaIndex can consume such a function via their custom-retriever constructs, and Haystack can wrap it in a@component. - LLM calls behind one client. If you use the OpenAI-compatible client directly for anything not covered by the framework, you have a fallback path when a framework’s model adapter lags behind an API change.
A worked migration path: suppose you prototyped in LlamaIndex and now need agentic control flow. Keep the LlamaIndex VectorStoreIndex as your retrieval backend, but call it from a LangGraph node:
# pip install langgraph llama-index langchain-openai
from langgraph.graph import StateGraph, END
from typing import TypedDict
from llama_index.core import VectorStoreIndex, Document
class State(TypedDict):
question: str
context: str
answer: str
index = VectorStoreIndex.from_documents([Document(text="Refunds take 5 days.")])
query_engine = index.as_query_engine()
def retrieve_node(state: State) -> State:
nodes = query_engine.retriever.retrieve(state["question"])
state["context"] = "\n".join(n.text for n in nodes)
return state
def answer_node(state: State) -> State:
state["answer"] = f"[draft answer using context: {state['context'][:40]}...]"
return state
graph = StateGraph(State)
graph.add_node("retrieve", retrieve_node)
graph.add_node("answer", answer_node)
graph.set_entry_point("retrieve")
graph.add_edge("retrieve", "answer")
graph.add_edge("answer", END)
app = graph.compile()
print(app.invoke({"question": "How long do refunds take?"}))
This hybrid is common in practice. The retrieval layer stays where it is best; the control flow moves to the framework that treats control flow as a first-class concern. You do not have to be all-in on one framework.
Symptom: The framework is eating your Python traceback and you cannot tell why the retrieval returned nothing
Cause: Abstractions that succeed loudly and fail quietly. A retriever that returns zero documents often produces an empty context, the LLM answers anyway, and the answer is confidently wrong. Haystack’s typed sockets catch this earlier because a node that receives an empty list can be asserted against; LangChain and LlamaIndex tend to pass empty context through.
Fix: Add explicit logging at each stage boundary, and pin your framework version in requirements.txt or pyproject.toml. All three ship breaking changes across minor versions more often than a normal library. A reproducible environment matters more here than in most Python work.
A useful pattern regardless of framework — wrap retrieval so an empty result becomes a hard failure you can see:
def retrieve_or_raise(retriever, query: str, min_docs: int = 1):
docs = retriever.invoke(query)
if len(docs) < min_docs:
raise RuntimeError(f"Retriever returned {len(docs)} docs for query: {query!r}")
return docs
Swap retriever.invoke for retriever.retrieve in LlamaIndex or your Haystack pipeline run call, and you get the same guard across frameworks.
Symptom: You are choosing between the three on a team with mixed experience levels
Cause: A framework that is easy to demo is not always easy to review. LangChain’s flexibility means two developers can solve the same problem in two structurally different ways, which makes code review harder. LlamaIndex’s implicit defaults are pleasant until a junior developer assumes the default prompt is fine for a regulated domain. Haystack’s explicitness is verbose but legible.
Fix: Match framework to team review culture.
- Solo developer or small team shipping fast: LlamaIndex. Fewer lines, faster time to a working RAG.
- Team building workflow-heavy features with branching and retries: LangChain + LangGraph. The graph model is reviewable once the team learns it.
- Platform team responsible for a search service other teams depend on: Haystack. The YAML pipeline and named nodes make change review mechanical.
When NOT to use any of these frameworks
If your application is a single prompt call, you do not need a framework. The OpenAI Python SDK, Anthropic SDK, or any provider SDK does the job with fewer dependencies and fewer upgrade surprises. Frameworks pay off when you have at least two of: retrieval over a corpus, multi-step control flow, multiple integrations, or a need to serialize and inspect the workflow. Below that threshold, they add surface area without adding capability.
Similarly, if you need maximum performance on a single retrieval pattern, calling the vector database’s own client directly — Pinecone, Weaviate, Qdrant, pgvector — with a hand-written prompt is often faster to debug and cheaper to run than routing through a framework. Frameworks are for composition, not for squeezing latency.
A short comparison table to keep
| Question | LangChain | LlamaIndex | Haystack |
|---|---|---|---|
| Best first project | Multi-tool agent | Document Q&A over your PDFs | Search pipeline with audit trail |
| Smallest working RAG | ~15 lines of LCEL | ~8 lines | ~25 lines |
| Serialization of flow | Code only | Code only | YAML export |
| Debug story | Callbacks, LangSmith | Callbacks, observability integrations | Component-level logging via pipeline tracing |
| Comfortable in production when | Control flow complexity grows | Retrieval quality is the hard problem | Reviewability and reproducibility are hard requirements |
None of the three is a wrong answer. The one that becomes a wrong answer is the one you picked because a tutorial used it and you never opened the abstraction to see what it was doing. Start minimal, add abstraction one layer at a time, and re-evaluate once a quarter — the framework landscape moves, and your second framework choice is usually better than your first.
🔗 Recommended Reading
- How to Evaluate LLM Outputs: A Beginner's Step-by-Step Tutorial to Testing and Scoring AI Responses
- LLM Guardrails for Beginners: A Step-by-Step Tutorial to Filtering Unsafe Outputs
- Building Your First AI Agent: A Beginner Tutorial with Python and the OpenAI API
- Function Calling in LLM APIs: A Beginner Tutorial for Connecting AI to Real Tools
- RAG Chunking Strategies for Beginners: A Step-by-Step Tutorial