7 comments

  • copemaxxxing 6 minutes ago
    People who build software factory are in for a disappointing future. Spoken as someone who uses AI day to day for job and personal projects, I've came across a few super gnarly bugs already, ranging from frontend React apps (yes, believe it or not, frontend is far from solved) to embedded that Claude can't solve. Those bugs are feature breaker. My primary skillset is fullstack/frontend leaning, not embedded. Granted, I might be out of depth in embedded/RTOS but I am very qualified to do frontend. Yet I found bugs in those areas that Claude Opus high can't even solve for hours.

    I pointed Claude at those for hours and hours, didn't fucking work, ended up creating more and more slop. Until I stepped in and fixed it in a few lines of code. Thankfully with all my AI assisted work, I took time to understand everything so I was able to debug and fix.

    The problem for software factory is, that up to a certain point, you'll encounter some nasty bugs that are really important for your features, but you can't ship it because while the AI deals with the 90%, turns out that last remaining 10% are the one that make or break your project. And you need to deal with that remaining 10%. But how can you deal with the remaining 10% if you don't understand that 90% or majority of it?

    Also, don't listen to influencers like Kent Dodds or Uncle Bob or Steve Yegge, who advocated for software factory walk away from the code kinda thing. They are already rich, and they sell teaching content, but most importantly, their time in this industry has passed. They are not in the weeds doing enterprise software development anymore.

    Their first and primary concern is trying to be relevant to sell courses. They don't care about you. Yes I used to listen to them (though not often), but now not anymore. The world doesn't get better or worse whether they choose to sip margarita on a beach or spin yet another "so many words to say absolutely nothing" video/blogs.

  • RajT88 1 hour ago
    So - apparently it's not fully self-hosted, since I don't see a GPU.

    I'm interested in hearing from folks who are hosting their own GPU to run coding models. So far my own results are... not great. Seems like frontier models are needed via the big providers?

    • primitivesuave 1 minute ago
      My favorite real-world example that I worked on: I created a YouTube documentary series about corruption in a small town in Illinois. This required downloading thousands of hours of government meeting videos from YouTube, transcribing + chunking + embedding, summarizing meeting segments, running a couple passes of validation, then searching for interesting storylines. I also scraped thousands of public documents which exposed campaign finance violations and some truly nefarious stuff going on behind the scenes.

      I did 90-95% of the work on this with a local GPU. I still ran Claude and Codex, but they were just writing code to orchestrate API calls against local models, local Whisper, etc. Doing the same with the public cloud would have cost thousands of dollars, instead it cost a few hundred dollars for a frontier lab subscription and tens of dollars for electricity.

      I recently got into the DGX Spark and just a month ago, got an AMD Ryzen developer platform. Both are incredibly useful tools for some upcoming projects I'm working on (also related to government corruption), but in my limited experience of trying to run opencode on them, the "good models" are still completely unusable because the system prompt alone requires a minute of thinking.

    • dang 57 minutes ago
      We've put almost back in the title above.

      (Submitted title was "A self hosted AI software factory")

      • jakelsaunders94 37 minutes ago
        Apologies, I was trying to choose between a snappy title and the whole story.
    • 0x457 1 hour ago
      I recently picked up R9700 to run Qwen 3.8-27B.

      1) I have knowledge DB that I query, spend a lot of time to get "query N models, stream to me all N results, let me pick which one" - only to find out that qwen beat all other contenders (those were picked before I got R9700, so the rest is =<8B parameters)

      2) Setup OpenCode to use it and gave it a few tasks:

         - first task (new feature to my MCP) took 40 minutes to complete with zero input from me, while it took 20 for sonnet-5 and sonnet-5 kept bugging me. Results are near identical.
       
         - second task (port a specific version of a package to my flake) it got stuck in a hilarious loop where model already been told what hash to use by nix itself, but it wanted to figure our how to get hash another way for some reason.
      
         - third task (another task, but much harder than first one with most of the discovery already done), kept doing discovery and running out of context, went through 3 compactions (256k context is what I can fit on R9700). Room got too hot, so I stopped it.
      
        3) virtual assistant like Hermes but my own: no notes, works great.
      
      
      I'm pretty sure codding issues are just harness and lack of memory that Claude Code already had. Pretty nice setup, similar to mine but I built my own lightweight PaaS that is highly specific to what I run.
    • nater5000 1 hour ago
      I have a similar setup to the OP (although I'm not sure if I'd call it a "software factory"), and I utilize a local model for some aspects of the setup. Specifically, I have a single RTX 3090 Ti with 24GB VRAM on my (main) home server, and I've been primarily running Qwen 3.6 35B A3B out of it via Ollama (I just switched to Qwen 3.8 27B, though, and I've also tried other Qwen models as well as Gemma models).

      The Qwen models are decent, but they don't come close to the full Claude experience I've come to expect. As such, I only use the local models for specific tasks where it makes sense to do so. Really the setup is that my Claude-powered agents are able to incorporate my local model into work it builds out. The agents can perform inference against the Ollama API as they see fit, and I encourage them to do so for tasks where (a) the low-level capacity of the local models make sense and/or (b) where costs can become a concern.

      It seems to work well when it comes into play (like having Claude drive a web browsing session but letting Qwen handle much of the actual browser interactions, image analysis, etc.). Still, Qwen just isn't smart enough (or fast enough on my machine) to handle anything agentic that isn't non-trivial.

    • yipinwong 48 minutes ago
      Check out some NIXOS communities. People they are doing fun crazy stuff, as NixOS is immutable, thus creates separate sandboxes (building/taking down) per agent, etc.

      I don't know most of the stuff there, but at least you can dive there.

  • codazoda 51 minutes ago
    I was recently inspired by another article here to start my own. My skills are written and tested, the factory has built the first test project, and I’m setting up the final machine to run it. Here’s my initial post about my motivation and early plans plus some follow-ups, and there are more to come.

    https://joeldare.com/creating-a-minimal-dark-factory

    • jakelsaunders94 40 minutes ago
      Just read the intro post, it’s a good read! I’ll read the rest when I get home from work. It’s nice to see others have had the same idea.
  • ramon156 40 minutes ago
    Can anyone that has OpenClaw/Hermex experience tell me what it's like working on this over OpenCode served?

    On bigger tasks OpenCode sometimes hangs. I'm not sure if that's on the provider side or on my network side, but i sometimes have to stop a subagent because its done but doesnt clean up. I wanted to try Pi, but now I'm wondering if Hermis is a better fit.

  • 100percentjake 1 hour ago
    Neat stuff. I use Hermes with the $20/mo ChatGPT Codex integration, myself, as well as Qwen running locally on a 64GB M1 Max Macbook Pro to help save my ever-quickly shrinking Codex credits. It's been wonderful for stuff like, say, giving Hermes an account on my Home Assistant server (after taking a backup) and having it make a bunch of configurations, or draw conclusions based on historical sensor data (trying to determine if my A/C is undersized by feeding Hermes all the capabilities of my system, size of house, and letting it pull room temp sensor data from HA is one project I did recently), is absolutely fantastic. It also spins me up little webapps for things like a very me-specific RSS reader, a "to do" checklist that references my email (and archives said emails when I mark an item off the list), and various little internal IT tools and report generators.

    The Hermes subreddit is a curious place where every second person has quick their $300k a year job and is making a living off of Hermes doing... something? They treat themselves as the CEO of a bunch of agentic employees and have AI generated infographics of their "Stack" (all hail the mighty Stack) and everybody stands in a circle and applauds the most convoluted Hermes setups you've ever seen with not a single word how any of this is supposedly making anybody any money.

    Are people developing like this? Are people asking money for something they one-shot instructions to an orchestration agent which delegated to fifteen other agents in Kanban and then spat out something that "works"? Every project I've ever had an LLM do a majority of the work for me has been strictly for my personal use; I'd never let anyone else use this stuff because it doesn't pass the vibe check. When a new frontier model comes out I'll pass the previous frontier model's work past the new model and let it tear it to shreds and see what improvements could be made for shits, giggles, and to waste a week's worth of tokens in the course of thirty minutes, but Reddit is overflowing with seeming non-coders who are passing this 100% organic slop off as sellable product?

    I need fewer morals.

    • jakelsaunders94 27 minutes ago
      That’s a great use of Hermes. Had it concluded if your A/C is undersized?

      Before this little experiment I generally used it as a very specific search engine for houses, ‘Find me a house in the country but close to amenities with a workshop barn but close enough on the train to Manchester’. It’s great for this!

      I’ve never been on the Hermes subreddit I’m gonna go check it out. I agree with the commenter below, it’s fine when I need no record my reps but don’t really care how. It writes tests but I don’t verify. It’s a manual ‘works or it doesn’t’ Thing. I’d never sell the code to anyone.

    • pmontra 49 minutes ago
      In my experience agentic sw development delivers if you are not too picky with the details or if you don't notice what's not done according to your instructions. If you want exactly what you had in mind, the last 10% might last a very long time and cost many tokens. More or less like working with humans.
  • pianopatrick 46 minutes ago
    Maybe I missed it but what AI model did you use for this?
    • ramon156 41 minutes ago
      > All on my home server, without another cloud infrastructure bill. The only ongoing cost specific to this experiment is a £20 Codex sub.
  • mempko 1 hour ago
    It seems at this stage hermes has the upper hand over OpenClaw in people building multi-agent systems. What about hermes do people like above OpenClaw. I'm building a completely different kind of "software factory" called Abject and I'm curious what people are valuing in Hermes above other systems.
    • tcdent 1 hour ago
      People assemble projects from a collection of buzzwords these days without understanding the underlying technology. Did he need to use tailscale to interact with the box remotely? No. But, it's what a google search will tell you to do.
    • aitchnyu 34 minutes ago
      Please share a link if its on the internet.