Atlas: A World Model for Spatial Intelligence

(worldlabs.ai)

85 points | by johnsutor 3 hours ago

5 comments

  • brettdev 3 minutes ago
    A camera moving through a 3D space the world model understands is getting much closer to real robotics applications
  • modeless 1 hour ago
    This seems like by far the best model yet for reconstructing 3D spaces from sparse images. It looks like you could reconstruct your whole house with pretty good fidelity from a dozen or so images taken on your phone.

    They show it working with videos that have motion, but it seems like time is always frozen while the camera is moving, and they always return to a ground truth camera view before advancing time again. Maybe the temporal consistency isn't very good? This surprises me given how well it understands space. I guess modeling physics and time is the next step in the development of this kind of model.

  • thinkingkong 1 hour ago
    What exactly does world model mean? Ive seen it used so many times in so many ways to just describe SOTA anything its lost its meaning.
    • CSMastermind 1 hour ago
      It's an overloaded term for AI models that have spatial reasoning LLMs currently lack.

      Best definition I've heard is: AI systems that can build an internal map of their surroundings to anticipate what happens next and make decisions based on their predictions about the consequences the different actions they can take would have.

      There's a bunch of different approaches people are trying:

      - World labs (linked in this post) is going down the route of neural 3D representation work (NeRFs, 3D Gaussian Splatting)

      - Yann LeCun is pretty famously betting on JEPA architectures (check out the excellent Welch Labs videos for more)

      - Google is betting on generative video

      - Karl Friston was pursuing 'active interference,' which is just traditional RL techniques with different reward functions

    • KaiserPro 1 hour ago
      It means everything to everyone.

      However essentially a world model is something that has the understanding of 3d world and can generate novel view point given either text or image input.

      The use I have seen is for robotics. You feed in the current view and describe the action you want it to do, and then it plans the arm movements. (really useful for softbody manipulation.

      There are other meanings. but essentially a world model is able to reason in 3d, rather than text.

    • bluecalm 1 hour ago
      It means it builds internal representation of the world it understands (can do physics on/predict/modify) and then renders it.
  • monkeydust 53 minutes ago
    > For robotics, reconstruction is only half the job: as a simulated robot moves through space, Atlas also generates the RGB and depth data its sensors would observe along the way. The world and the robot's view of it come from the same model.

    Potentially very significant for accelerating the data flywheel challenge for robotics

  • doctorpangloss 44 minutes ago
    "Can reconstruct [scenes from Unreal Engine]"