13 comments

  • kelnos 52 minutes ago
    I agree. The frontier models are based on training data from tons of copyrighted work. Some of that work was obtained illegally, even. They could not exist without strip-mining the commons. The labs have no moral or ethical ownership to the end result, and others should feel free to treat any company-imposed restrictions on their use as invalid.

    I don't expect Tan's position to be based on any kind of real moral high ground, but his conclusion is correct.

    I love the "illicit distillation attacks" framing from the incumbents. There's nothing illicit. There's no attack. You just don't like it because it threatens your market position and business model.

    • torginus 0 minutes ago
      With the recent Navier-Stokes controversy, I think there's a credible suspicion that all your IP you run through these models will end up in these companies' possession. OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this.

      Obviously it's not possible to run a company whose value is predicated on its IP that uploads said IP to a third party which might get access to it.

      This could mean every potential serious customer would have no option but to seek alternatives to these online services.

    • giancarlostoro 3 minutes ago
      Abolish copyright and make it less ridiculous. Sampling music was never a thing that required royalties until the 1990s when I guess someone got angry that rappers were making money off their sampled music. Its insane to me. Make it illegal to transfer ownership of copyrighted work too, only the spouse or one single inheritor who isnt a company can have the rights transferred, after both die, the work enters public domain.

      LLMs should just pay a flat fee to use a specific book and thats it. Fees should be reasonable (not a million dollars per book), so long as the model doesnt spit out the entire book.

    • stymaar 2 minutes ago
      This. Distillation “attacks” are a made up concept. It's as if I claimed that Anthropic made a “training attack” when training on my internet writing.
    • mobelkh 21 minutes ago
      why can't I use the tokens i paid for anyway?
    • knollimar 20 minutes ago
      I'm sure they put some BS in their TOS
      • stymaar 5 minutes ago
        I'm also certain that they violated countless ToS when they scrapped the internet for training purpose.
  • TheJCDenton 1 hour ago
    > He also notes that the proprietary AI labs didn’t ask permission when they vacuumed up as much human knowledge as they could to train their models.

    I think this should desactivate the moral high ground from which Anthropic is trying to speak. That they would want to make distillation orderly IMHO is fair, but to make it illegal is very rich from any AI frontier lab, really.

    • Bluestein 38 minutes ago
      Also, as said elsewhere: "Lab" is rich here, for outfits that, facing these giant, energy swallowing black boxes have really no clue what's going on inside.-

      The moniker gives them an air of scientific, knowledgeable, tranquil, pro-social, pro bono work.-

      Of course they are entitled to kill off a few mice, or pillage the commons to forward their "lab" work.-

      • Den_VR 10 minutes ago
        “We don’t know what’s going on” is essentially marketing. Sure we don’t _know_ but we have intuitions about why, where, and how to make certain changes…
      • travisgriggs 14 minutes ago
        We also associate laboratories with evil scientists and Frankenstein and the like. I can just hear Boris Karloff (er Bobby Picket) uttering “I was working in the lab late one night. When my eyes beheld an eerie sight… … … …the monster mash”. If anything, I associate _uncertainty_ with labs. The result is never known up front, they’re a place of discovery.

        But I get your meaning. What should they be called instead? AI Sausage Factories maybe (cue Upton Sinclair?)?

        • Avicebron 7 minutes ago
          > What should they be called instead? AI Sausage Factories maybe (cue Upton Sinclair?)?

          That's actually great? Slaughterhouses killing off the collective genius of humanity and grinding it into a bland paste for mass consumption.

    • sobellian 56 minutes ago
      I reflected on this myself recently. Model distillation seems to be at least as fair a use as distilling a book.
      • causal 46 minutes ago
        More than fair if you consider that the tokens are paid for.
        • dathery 40 minutes ago
          Both labs even explicitly promise the customer owns the outputs. It feels like they want to have their cake (ensure enterprises don't get spooked away from using as many LLMs as possible) while eating it too (still arguing some level of control over the outputs).

          > Ownership of content. As between you and OpenAI, and to the extent permitted by applicable law, you (a) retain your ownership rights in Input and (b) own the Output. We hereby assign to you all our right, title, and interest, if any, in and to Output.

          https://openai.com/policies/terms-of-use/

          > As between the parties and to the extent permitted by applicable law, Anthropic agrees that Customer (a) retains all rights to its Inputs, and (b) owns its Outputs. Anthropic disclaims any rights it receives to the Customer Content under these Terms. Subject to Customer’s compliance with these Terms, Anthropic hereby assigns to Customer its right, title and interest (if any) in and to Outputs.

          https://www.anthropic.com/legal/commercial-terms

          Obviously there is some bad behavior going on in the distillation scene with gray-market token resellers but that is "just" normal fraud.

          • zenoprax 11 minutes ago
            > Both labs even explicitly promise the customer owns the outputs.

            > to the extent permitted by applicable law, you (a) retain your ownership rights in Input and (b) own the Output

            If the argument is that the model itself is under copyright protection then "as permitted by applicable law" would be doing some heavy lifting. Assuming that were true, given that locally-run LLMs exist, what would be illegal: the distillation itself or the provision of service of the distilled model?

        • bayindirh 15 minutes ago
          And LLM output can't be copyrighted, even.
    • toomuchtodo 58 minutes ago
      YC does better if its startups get open weight frontier benefits. Garry’s just advocating for his book, which is his job. Consider how much capital YC portfolio companies would have to burn until liquidity if they have to pay OpenAI and Anthropic, versus relying on open weight frontier capabilities.
      • SOLAR_FIELDS 41 minutes ago
        If someone proposes the right thing for selfish reasons, do we call that bad? Or do we call it proper incentive alignment?
  • layer8 1 minute ago
  • pton_xd 25 minutes ago
    Agreed! Allow US companies to innovate by creating an ecosystem of smaller, more efficient open weight models and it will be a net benefit for everyone. Distillation is a good thing.

    Preventing token-consumers from developing competing products should be litigated as anti-competitive behavior.

  • fmnxl 11 minutes ago
    If it were so easy why aren't the frontier labs doing it themselves?
    • layer8 3 minutes ago
      Distilled models are worse than the original, so you can’t fully compete. Also, if all frontier labs did that, there would be nothing left to distill from.
  • quicklywilliam 32 minutes ago
    I see it as analogous to companies building fiber in the public ROW during the last big infrastructure bubble. Under the Telecoms Act, these companies had to allow competitors to use their fiber at a fair price.

    Similarly, AI companies should be required to allow distillation at a fair price. Fair Use doesn’t make sense as a social contract if it only cuts one way!

    • jimnotgym 8 minutes ago
      But if they tried to set a fair price they would have to report how much money they are losing on each token sold. This might be bad for the real business of ai firms, hoovering up as much capital as they can
  • ViktorRay 23 minutes ago
    https://youtu.be/ZIaOBAjvc38

    Garry Tan and Sam Altman recently did this interview together. They seemed pretty friendly with each other during it. Wonder what Sam Altman would say about Tan advocating for OpenAI’s models to be distilled.

    Then again this is the same OpenAI that has gotten into legal trouble recently regarding Apple’s IP so who knows

  • re-thc 31 minutes ago
    There were comparisons and Muse Spark is so very similar to Fable / Opus... so...
  • brcmthrowaway 35 minutes ago
    Ozempic worked for him, nice.
  • Hikikomori 5 minutes ago
    Garry also goes to Thiels silicon valley church.
  • okasaki 14 minutes ago
    Like Gates saying there should be UBI, or Musk saying... well, whatever.

    They know it won't happen, so arguing for it is 'effectively free' and purely personal marketing.

    A bullshit game played by politicians and wannabes.

    • seanmcdirmid 11 minutes ago
      Gates probably honestly believes in UBI; the guy is practical to a fault but evil misleading genius he is not. I actually don’t see any better options than UBI long term.
  • zombiwoof 21 minutes ago
    [dead]
  • 9865322689965 21 minutes ago
    [dead]