Building a RAG pipeline for semantic code search

(blog.jetbrains.com)

30 points | by saikatsg 2 hours ago

3 comments

  • keeda 25 minutes ago
    I think something like this would be key to improving the quality of coding agents. A very common issue (maybe the biggest one) observed by many people is that agents often produce a lot of duplicate and redundant code; multiple abstractions, methods, classes, data structures, etc. serving minor variations of the same purpose... sometimes within the same file!

    My theory is that this is due to a kind of "tunnel vision" these models have as they execute on a given task, because engineers new to a company do the same thing until they learn the "lay of the land" and figure out that similar problems have been solved elsewhere.

    In a past job my team owned the internal multi-repo codesearch tool, which was by far the most popular internal tool, and later another team added a similar semantic search capability. This was very exciting, but I left before I could see how well it worked out in real-life.

    Like, you'd do a keyword search and explore if you need some major piece of functionality that would require significant work, or whenever you encounter an abstraction whose code does not exist in your repo and you want to learn more about it. But when you're in the flow and inventing smaller abstractions, like a class or utility method, you don't necessarily think to search for it. Worse, even if you did, you could not search for it effectively because something similar may exist with slightly different naming or terminology or a typo that a keyword search would miss. Predictably, at scale you ended up with a dozen different implementations doing the same thing.

    Now however, you could automate this with agents. I suspect these days simply prompting an agent to look for any relevant code to reuse would actually work pretty well. But they would need to store the entire codebase in their context window (if it fits at all) to refer to it all the time, which would burn a ton of tokens AND reduce performance due to a heavily polluted context. Instead, something like this would be invaluable to provide as a tool / MCP to the agent so that it could locate relevant, reusable code during its planning phase. (Or maybe a post-codegen linter-like check, which IIRC some people have tried, but why fix when you could prevent?)

    I'm not sure if the exact method in TFA would work best, though; maybe a pipeline that generates comments/docs for each unit of functionality and then semantically indexes those, rather than structure-aware chunks of code itself?

  • duhhhhh1212 1 hour ago
    https://www.pangram.com/history/c901e80e-9cb7-46e4-bf10-7348...

    I don't want to say don't waste your time since the first half is human written. Questions for the authors: did y'all just get tired of writing and said "fuck it let's have the LLM finish the rest"? Or did one of you use LLM to write the last half and the other used their own words?

    • verdverm 1 hour ago
      This looks like the only kind of comment you make lately (to a Ai writing detector)

      1. detectors are unreliable, humans are "writing like ai" now (first we shape our tools, then they shape us)

      2. more people don't care as long as the content is quality

      3. there is a spectrum of ai-human writing and how people make use of the tools, some good, some NS;NT

      https://news.ycombinator.com/item?id=47089907

      ---

      While I am not without my critiques of the article (such as the lack of any measurements), I still find the ideas of interest (such as AST based chunking).

      • duhhhhh1212 48 minutes ago
        > detectors are unreliable

        If you can prove pangram is unreliable I will stop using it. By prove, I mean show their false positives and false negative rates are made up.

        > more people don’t care as long as the content is quality

        That’s your opinion. If you want to read AI generated content then there are wonderful sites that can produce that within seconds.

  • simianwords 1 hour ago
    Here we go again, the industry largely gave on up RAG. In fact I have hardly seen any case where grep doesn't work as well as RAG.