Rendered at 20:01:55 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
ryuuseijin 23 hours ago [-]
This is what I wanted to help with duplicate code detection, which is now a real problem with agentic programming since the agent tends to duplicate helper functions a lot in bigger codebases.
I ended up writing my own tool [1] that uses vector search, which works but can be unusably slow in large codebases.
ooh I'll give yours a shot - thanks for calling my project out on your readme!
I was fortunate to have a contributor push me to consider optimizations a few months back, and it definitely helped improve the speed/memory metrics.
I see the CLI knobs are somewhat similar ('min lines', 'thresholds'), however your 'search' is fascinating. I could see myself wondering "how much error handling is duplicated for HTTP responses?" and your tool would give me a start at an answer. Nice!
adityamishra241 15 hours ago [-]
Nice approach. AST-level similarity seems much more useful than simple text matching.
Game_Ender 22 hours ago [-]
Thank you for a README that is readable. It might have a bit of AI help but it’s concise and makes it clear how to use the tool.
anaqin 7 hours ago [-]
I dunno… I’m finding myself yearning for some emojis, a dedicated marketing web site, and a bunch of incomprehensible paragraphs that lack any soul.
91awebsi 4 hours ago [-]
Yes, a little AI on the main readme. I started the bones of the thing, but Claude has been behind the keyboard for much of the implementation.
bradleyy 1 days ago [-]
Thanks for posting this; I'm very interested in finding similar chunks of code in an AST-ish fashion, and using tree-sitter seems perfect.
91awebsi 4 hours ago [-]
I keep meaning to write a post for how this has worked for my projects, I think its been helpful.
For active projects I setup a weekly-ish 'code audit' agent: it runs treepeat to find opportunities to refactor and bumps up enforced coverage numbers. Over time for some projects I lower the '--similarity' threshold to more broadly find similar structured code (potentially more ambitious refactoring)
I ended up writing my own tool [1] that uses vector search, which works but can be unusably slow in large codebases.
I will give this a shot.
[1]: https://github.com/ninjaxtools/slopdex
I was fortunate to have a contributor push me to consider optimizations a few months back, and it definitely helped improve the speed/memory metrics.
I see the CLI knobs are somewhat similar ('min lines', 'thresholds'), however your 'search' is fascinating. I could see myself wondering "how much error handling is duplicated for HTTP responses?" and your tool would give me a start at an answer. Nice!
For active projects I setup a weekly-ish 'code audit' agent: it runs treepeat to find opportunities to refactor and bumps up enforced coverage numbers. Over time for some projects I lower the '--similarity' threshold to more broadly find similar structured code (potentially more ambitious refactoring)