Rust's derive often implies inline

(yossarian.net)

81 points | by woodruffw 3 days ago

5 comments

  • vlovich123 3 hours ago
    I feel like Debug should be lazily emitted altogether when it’s first used - just a special marker that’s never expanded since 99% of the Debug implementations aren’t used and having the rest marked #[cold] as inline is obviously wrong. Of course implementing it in practice sounds exceptionally difficult.

    That being said, even the justifying performance improvement PR was itself a mix of improvements and regressions

    • bloppe 50 minutes ago
      Lazily emitted from what? You would need information about the type at runtime to derive the implementation, and normally that information is not available at runtime. The debug implementation itself is probably within spitting distance of any other representation that would be sufficient to lazily emit the debug implementation.
      • vlovich123 1 minute ago
        You say runtime but that's ambiguous in this context. My point is that what gets emitted for Debug is just "type X lazily implements Debug" and you never even generate the equivalent Rust code for it until someone uses it and at that point you mark it with cold & inlineable if it's a small function.

        > within spitting distance of any other representation

        If marking Debug as #[inline] vs #[noinline] has an effect, it's clearly already true that there's a lot of time being spent on emitting the Debug trait eagerly and having it go through all the compiler stages.

      • Lvl999Noob 28 minutes ago
        You can do it at compile time. Or more likely, link time. If the implementation is ever actually used, it is kept in the binary. Otherwise it is used. The compiler can do the same, keeping all derives as just markers until it finds a place that actually uses it then firing off a background worker to compile the derive impl.
        • saghm 12 minutes ago
          My instinct is that this would end up being more costly than it's worth. You'd be essentially adding extra bookkeeping and logic for every single type that needs to remain alive as long as the type is still possible to reference.

          Moreover, how would you deal with downstream usage by dependents? Should I be able to make my dependency's type implement Debug (which is at least in spirit a violation of the orphan rule, and would make the bookkeeping/extra logic described above explode for any non-trivial dependency tree)?

          People already find the Rust compiler too slow. If there's budget for adding more expensive checks, I don't think I'd want it to be spent on something like this.

  • Sharlin 4 hours ago
    I’m fairly convinced that Debug should never be inlined. Display probably neither, the fmt machinery is heavy enough that not inlining is probably not a bottleneck even in serialization-heavy workloads. I’ve had to #[inline(never)] some of my own Debug/Display impls, shrinking the binary by tens of kilobytes (out of a few hundred, so relatively a significant reduction).
    • aw1621107 3 hours ago
      For what it's worth, according to the PR that added the annotation [0] doing so generally resulted in decreases in compile times and binary sizes on benchmarks. Furthermore, an additional experiment that avoided emitting the inline attribute on structs with >5 fields resulted in benchmark regressions compared to always emitting the attribute [1]. I'd guess this is one of those things which may help in aggregate but hurts for specific cases.

      That being said, one of the Rust devs indicated in the corresponding lobste.rs discussion [2] that they're open to revisiting/rebalancing things if they get enough bug reports indicating something is up, so it might not hurt to tag onto the bug report the author will (hopefully) eventually submit.

      [0]: https://github.com/rust-lang/rust/pull/117727

      [1]: https://github.com/rust-lang/rust/pull/118031

      [2]: https://lobste.rs/s/dldhpw/rust_s_derive_often_implies_inlin...

      • afdbcreid 3 hours ago
        A reasonable conjecture was raised on lobsters that this is because `#[inline]` makes actual codegen (LLVM IR and down from MIR) lazy, and most `Debug` impls are never used.
        • saghm 10 minutes ago
          That makes sense. If you're a library, it's just good manners to derive Debug on types that you don't have a good reason not to, but you have no way of knowing if anyone will actually use it in a given program.
        • infogulch 2 hours ago
          How much code is never used and compilation could be skipped entirely? Maybe applying a reachability pass to skip compiling unused code would be helpful.
          • saghm 8 minutes ago
            I've said for a while that I think that the "real" issue with compile times is that there's a ton of dead code getting compiled across the dependency tree. I personally blame Cargo features not being ergonomic enough, because on paper they're perfectly suited to solve this problem, but in practice the amount of boilerplate needed to use them for this is infeasible. I wrote a manifesto about this a while back: https://saghm.com/cargo-features-rust-compile-times/
          • mgsloan2 1 hour ago
            A cross-crate dead code analysis would mean that compilation of a crate now depends on information about its dependents. This would break reuse of compiled crates and cause recompiles when the analysis changes.

            Something does seem a little off about this, though. Ideally for this `Debug` case there would be an annotation that says "compile this lazily, don't inline". Maybe there doesn't even need to be a new annotation, just `#[inline] #[cold]`. Which looks pretty weird, but might work already.

      • Sharlin 3 hours ago
        Thanks, interesting!
  • scottlamb 2 hours ago
    I wonder if they ever considered a table-driven approach for `#[derive(Debug)]`, as `facet` [1] does. That would have been my first instinct for something this formulaic where binary size and compilation time matter more than execution speed. But my impression is facet hasn't quite realized its promise on those fronts, so maybe the table-driven approach in std was similarly tried and rejected.

    [1] https://crates.io/crates/facet

  • zamazan4ik 3 hours ago
    Or just try to avoid all of these optimization guesses by using Profile-Guided Optimization (PGO), that inserts/deletes all inlines based on actual application runtime profile.
    • woodruffw 3 hours ago
      The codebase in question (uv) uses PGO already. I suspect there isn’t a general way to guarantee that PGO ensures that only the “right” things get inlined.
  • api 3 hours ago
    There's a joke that the LLVM heuristic for whether to inline a function is "return true;" LLVM tends to inline aggressively.

    You can control this behavior with opt level "s" or "z" or "#[inline(never)]", but be aware that too little inlining can have large negative performance impacts.

    It's hard to get inlining exactly right without profile guided optimization.