Some time ago, we set out to build the best semantic code search platform we could: a RAG pipeline that gives LLM agents precise, citable evidence from real repositories instead of whatever grep happens to surface. The eventual solution was Air Context. We got it working, we got it into production, and we collected a lot of scar tissue along the way. In this series of posts, we’ll share the parts we wish someone had told us on day one.
Coding agents are undoubtedly the biggest technology leap for software development of our decade. However, as more and more development processes become agent-driven, the agent’s efficiency and the quality of the produced code become increasingly important. For large-scale code bases specifically, the agent would spend a great deal of time searching for the relevant pieces of code.
Attempting to locate the right code snippets, the agent will resort to traditional tools for code search such as keyword search and grep. These tools require the agent to know in advance which exact text to search for. This is where retrieval-augmented generation (RAG) comes into the picture. If we can index the source code in a way that captures its semantics and then allow the agent to retrieve the relevant pieces on demand using free text search, we create an interface that plays to the agent’s strengths.
Like many great ideas in the agentic era, a native, prototype implementation is extremely simple. A well-evaluated production grade solution most certainly is not. This first part of the series will cover the initial stages of the pipeline: parsing and chunking, where raw source files are divided into properly scoped units, and vectorization, where those units are transformed into a representation that supports semantic search.
Source: Hacker News · Summarized by HeadlinesBriefing