Show HN: Sokoban AI Solver

(mkornreich.me)

39 points | by enjoyyourlife 3 hours ago

14 comments

  • TimTheTinker 1 hour ago
    I love seeing the term "AI" used in the classic sense. Old AI is full of fascinating developments. Expert systems, A* search, genetic algorithms over S-expressions for creating arbitrary solutions, and SAT algorithms were once thought to be that which would eventually scale into AGI.

    I suspect that the next big AI breakthrough will result at least in part from constraining LLM decisions with old AI approaches. Frank Coyle presented the idea of ontologies constraining LLM output about a month ago: https://www.youtube.com/watch?v=Sir59K8ZDPU

    Going beyond that, I wonder if an agent could keep a running list of assumptions & known facts (with confidence levels/intervals), test them (actively & passively), update them when observations contradict them, and act based on them -- not merely as an emergent behavior, but as a provably correct (old AI based) algorithm embedded in the transformer architecture.

    • dietr1ch 1 hour ago
      > I suspect that the next big AI breakthrough will result at least in part from constraining LLM decisions with old AI approaches.

      AFAIK bridging deductive and inductive AI has been understood as the trick for "AGI" for a long time, probably even before it was called AGI.

      I really want this winter of deductive AI to be short. We need both sides and can't afford a long winter like the one inductive AI suffered.

  • epiccoleman 1 hour ago
    I'm kind of surprised to find myself enjoying this because I've had a certain hatred for box pushing games. (maybe it's trauma from the sliding blocks in Pokemon games, heh). I guess I'm getting over it (maybe it's happy memories from Baba Is You).

    Anyway, one thing that's fun here is that you can trigger the AI solve from any board state. So in particular on puzzle 12 I was interested to see that an initial push (to escape from the 'box' where you start) I'd written off as untenable turns out to be the optimal solution. Then of course it's fun to watch the solver tackle the initial conditions I solved under (and still beat my number of moves).

    Might be kind of fun to play with "pessimizing" the puzzle - like, how can you move blocks around to provide a maximally adversarial place to hit the "solve with AI" button? (obviously you don't get to count your initial moves around the board, or you could just move back and forth to get the most pessimum (thanks, Mel) solution.)

    Edit: Puzzle 14 feels odd. Super easy, why is it at 14? Maybe something tricky about it that I'm not seeing, perhaps the shape of the arena makes A* harder or something?

    Also, 15 is interesting and highlights a theme I'd noticed, which is that often the initial moves of a puzzle seem pretty locked in, and the place where the AI shaves moves off my solution in in some clever approach to the "stacking" of boxes onto the goals. I guess that seems kind of obvious when I write it out.

    Anyway, thanks for something to noodle on this morning!

  • xpct 1 hour ago
    Got me curious: is there some way to approximate solvability of a puzzle in a certain time frame, or is that completely intractable?

    Also, what counts as "complexity" in Sokoban puzzles. Does it plateau at a point, where board size/box count starts scaling the solving time more linearly?

    • jan_Inkepa 29 minutes ago
      There's computational complexity, then there's human complexity. I've thought a lot about this over the years (I've made a bunch of puzzle games, and puzzlescript, an engine/language for making grid-based puzzle games), and the only paper I've read that's made me think 'huh' was "Difficulty Rating of Sokoban Puzzle" by Jarušek and Pelánek ( https://www.fi.muni.cz/~xpelanek/publications/stairs2010-fin... ).

      While I have a feeling that subject 'difficulty' is necessarily a slippery concept, they focus on 'context switching' as a key element of difficulty. In sokoban terms - how often you have to alternate between pushing one box and pushing another. This too can be gamed/trivialized, but, when I used it as a heuristic is was very good at generating the most horrifically difficult levels, much moreso than just going for 'solution length'.

      On more general notions of complexity. In sokoban terms, the number of crates trumps everything else - for solvers I've written you quickly get exponential explosions with the number of crates. Nothing else really is significant.

      I've also been working on solvers for more general classes of these games (puzzlescript games) and it's surprising how powerful generic solvers still are. PuzzleScript+MIS https://dekeyser.ch/puzzlescriptmis/ (not by me) is one powerful tool that uses PuzzleScript as a basis. I've worked on speeding up the solver a bunch (not currently integrated), figuring out good general heuristics for different kinds of games ( https://github.com/increpare/puzzlescript-labs has various experiments in this direction, including a modded version of PS+MIS). It's a nice optimizaiton problem for focusing on making numbers go down - there are lots of games to test on.

      • xpct 4 minutes ago
        Wow, thank you for your input! The "alternation" proxy for difficulty is fascinating to me, doesn't feel like something I would have thought of right away. And, I guess I wasn't aware of how much thought goes into designing Sokobans :)

        I think the crates trumping other complexity metrics isn't entirely obvious to me. For problem 15 in the OP's post, author says it was too expensive to compute at runtime in the browser. From a human perspective, it's not apparent why, as a large part of the solution is very repetitive. It feels as if there should be a more condensed representation for iterating over problems like that one.

        If I may gauge your opinion on it, have you looked into MazeBench? It comes from LLM benchmarking circles, but seems to suggest a search space that's too difficult for LLMs, even with tools, to solve. Curious how much overlap the PS/MIS solvers would have with solving something like this.

    • FartyMcFarter 26 minutes ago
      According to Wikipedia Sokoban is NP-hard, which means there's no known polynomial time algorithm to solve it. It also means it's unlikely such an algorithm exists, as that would imply P=NP which is not believed to be the case.
    • yobbo 1 hour ago
      For games in general, one measure of complexity is branching factor. It means average number of possible actions or states at each turn. It is knowable.

      "Solvability" would mean number of turns to solve the game. It is known for some puzzles and can be found by brute force, otherwise you need to figure out a proof.

      • xpct 1 hour ago
        Thanks. Given a solver, could we extrapolate a problem's branching factor? For classic Sokoban, I'd guess it's on the lower side?
        • mightybyte 58 minutes ago
          I think there are two ways one could look at this. One is to make each move be a move of the player's location. If you do that, then the branching factor is obviously < 8. But there's a second way you could define a "move" for the purposes of a solver. And that would be to only consider pushes. In that case, the branching factor would be < 8*num_stones.

          In either case, I think when trying to assess complexity it might also be useful to consider the "narrowness" of the winning move sequence. Positions where the number of moves that win/make progress towards the goal is a small fraction of the number of available moves would arguably be harder or more complex than positions where a larger percentage of the moves win/make progress. In other words, finding a smaller needle and/or in a larger haystack makes the problem harder / more complex.

          • xpct 28 minutes ago
            Hmm. The push representation makes sense because solve progress is entirely dependent on it. And the movement state tree can be reduced to the push tree, which would only prune useless paths. The push tree can probably also be pruned for moves that leave to softlock, but I wonder whether it can be reduced to a different representation still. Push tree already requires us to maintain a mask of where we can move to, so it's not computationally free. I can imagine representing box pushes as every position we can push it to in the current setup, but that would also make it more computationally expensive.

            I feel like there's an interesting tradeoff of storing/computing cheap representations vs exploring a smaller tree.

  • npinsker 2 hours ago
    Intuitively, I feel like the final board might also be able to be tackled in browser, if you use WASM and speed up the solver.

    I wonder: maybe the state is overly compressed? Could it speed things up to store (boxes, [every position the keeper can reach without pushing]) rather than (boxes, representative keeper position), so we can reduce recomputation of the keeper walking around?

    I wonder: maybe A* is counterproductive, as obvious heuristics have traps? Maybe BFS is better?

    I wonder: the search doesn't actually "skip over" walking states, it just hides them in the processing of each element in the queue, so adding them to the queue might actually be faster?

    I wonder: are there any other simple pruning techniques that you could incorporate? Any learnings from state-of-the-art Sokoban solvers, like this one? -- https://ieee-cog.org/2020/papers/paper_44.pdf

    Many interesting questions... sadly, the webpage is written by AI, so there's zero discussion of these tradeoffs, future avenues, or rejected ideas, in favor of meaningless self-congratulatory copy about the "provable optimum" and silly claims like a bucket queue being allocation-free.

    • throwaway219450 51 minutes ago
      Showing the exploration would be nice. À la RedBlob tutorials, seeing the solver work is part of the fun. As is I have no intuition for where the algorithm would spend all its time and where it can easily rule out. 1GB of RAM for the final puzzle isn’t too bad for a browser demo if you warn the user and don’t run automatically (is the state space compressible?)
  • GPerson 3 hours ago
    “What runs here is a plain-JavaScript port of a native C++ optimal solver I wrote.”

    Seems to be AI in the older sense from 10 years ago?

    • dev_dan_2 2 hours ago
      Hmm, I would say even older than that (which, of course, is in no way intended to be a value statement of any kind, I like the website and the project, cool idea! :D).

      In 2015, https://en.wikipedia.org/wiki/AlphaGo came around and latest from there on, AI was associated heavily with NNs, deep learning and so on (but not with the transformer architecture which became popular later, the foundational paper itself was published in 2017: https://en.wikipedia.org/wiki/Attention_Is_All_You_Need).

      If you squint a little, the linked project is basically a https://en.wikipedia.org/wiki/A*_search_algorithm with optimized implementation, heuristics and so on. I also think that A* was associated with AI due to its use in path finding in early robotics - But I am not sure!

    • nairboon 2 hours ago
      No, that's still AI in today's sense, just not an LLM.
      • GPerson 1 hour ago
        On further reflection I actually now disagree that a an algorithm based puzzle solver was ever referred to as AI, even in the context of video games, in which AI refers to the behavior of NPCs.
  • amelius 1 hour ago
    Doesn't this break down quickly as the area increases? (Ironically, the complexity goes down as there are more squares to use).
  • cbondurant 2 hours ago
    While impressive that the optimal can be proven, I feel like the example puzzles here aren't ones that are particularly hard to find solutions for (when move count doesnt matter). I'd be interested to see at least one example that has a lot of tricky dead states that would act as traps.
    • conmod278 2 hours ago
      Imagine providing AI with ability to poke around a large bank of gridbased game problem instances. Ask it to solve them and learn from them and then generate new problem instances.
    • Retr0id 2 hours ago
      This was a coursework problem in my CS course, back in the day. For larger canvases, the state-space blows up and it gets slow/intractable to solve.
  • qbane 2 hours ago
    Compared to original sokoban game, the player's final position does not matter, and the number of boxes is strictly equal to the number of goal marks.
  • Sebastian_09 2 hours ago
    Fun game! Solver seems really smart. It would be great to disable double tap to zoom or make it slightly more adapted to phone screen sizes
  • k2xl 2 hours ago
    I wonder how this would do with Thinky.gg games (Pathology or Sokopath). Are you familiar with the site? There's a group of engineers working on various types of solvers in the thinky.gg discord too.
  • CatalystPz 2 hours ago
    pretty cool stuff, enjoyed it
  • mohamedkoubaa 2 hours ago
    Terms like AI used to mean something specific
  • j16sdiz 3 hours ago
    [dead]