Side-by-side comparison of building a Retrieval-Augmented Generation pipeline from scratch versus using the Mapbox Agents SDK.
| DIY | SDK | |
|---|---|---|
| Vector store | ~20 lines (cosine similarity by hand) | InMemoryVectorDB — zero |
| Embedding | ~5 lines | OpenAIEmbedding — one line |
| Retrieval | ~10 lines | RetrieveTool — ~15 lines (but agentic) |
| Generation | ~10 lines | ToolOrchestrator + ConversationManager |
| Total | ~65 lines | ~70 lines |
The raw line counts are similar. The difference is what each line does.
- Cosine similarity — implemented and tested in
InMemoryVectorDB - Embedding batching — providers handle chunking large inputs automatically
- Rate limiting — built into all providers via
TokenBucket - Agentic retrieval — the LLM decides when to retrieve and can call
retrievemultiple times for multi-hop questions (DIY always retrieves exactly once, before generation) - Conversation history —
ConversationManagerpersists turns; follow-up questions work out of the box - Swap without rewrites — switch
InMemoryVectorDB→QdrantAdapterorPineconeAdapter, orOpenAIEmbedding→CohereEmbedding, by changing one line
- OpenAI API key — for embeddings (all current providers are API-based)
- Anthropic API key — for the LLM
Coming soon:
TransformersEmbeddingwill add local/offline embedding via Transformers.js — no embedding API key required.
If your pipeline is truly fixed — always one retrieval, no conversation, no need to swap components — the DIY version is simpler. The SDK pays off when you need agentic retrieval, multi-turn conversations, or want to graduate from in-memory to a production vector store without rewriting everything.