• lurker@awful.systems
    link
    fedilink
    English
    arrow-up
    9
    ·
    2 days ago

    Someone pointed out that on some tests, Astra is actually a regression from previous models. Totally AGI guys

    • scruiser@awful.systems
      link
      fedilink
      English
      arrow-up
      3
      ·
      1 day ago

      Benchmaxxing for one set of benchmarks can actually degrade performance on other benchmarks. They’ve plateaued for a while, all they can do is scale inference compute up and down (at logarithmically poor rates of exchange) and trade performance in one area for performance in others

    • diz@awful.systems
      link
      fedilink
      English
      arrow-up
      6
      ·
      2 days ago

      Yeah sloppers I know went from “see it proved new math! that totally clears us of plagiarism accusations we get when we repost that someone ai generated a frogger and a WoW clone” to “tried to use it for coding and it didn’t do a good job”.

      The math stuff is mostly using Lean proof verifier and all that, by the way, the only use of LLM that they found that is actually legit because it doesn’t matter if its slopping, and you want your attempts randomized.