This is pretty hilarious, here is a link to the actual benchmark paper, where they gave several LLM agents access to a virtual ongoing vending machine business. Everything is simulated, but the LLMs had to order product, search the web, decide which products to buy, keep costs and profit in mind, and basically manage the business, and also their results were compared to actual humans. Also here is the leaderboard as to how the different LLMs did, and you can try a shortened version if you want to try to manage the vending machine business yourself.

  • 3DMVR@lemm.ee
    link
    fedilink
    English
    arrow-up
    0
    ·
    2 days ago

    As long as science fiction is used as training data they are gonna tweak like this, they shouldve never used hella books without authors permission