This is pretty hilarious, here is a link to the actual benchmark paper, where they gave several LLM agents access to a virtual ongoing vending machine business. Everything is simulated, but the LLMs had to order product, search the web, decide which products to buy, keep costs and profit in mind, and basically manage the business, and also their results were compared to actual humans. Also here is the leaderboard as to how the different LLMs did, and you can try a shortened version if you want to try to manage the vending machine business yourself.

  • Rhaedas@fedia.io
    link
    fedilink
    arrow-up
    0
    ·
    9 days ago

    A company would look at this and determine not that LLMs might have something going on that would be bad for long term business. They would see the bigger net dollar amount and figure that they just had to calculate when to “reset” the LLM. It’s just another IT problem where the solution isn’t to address the problem but to find a workaround that reduces cost while continuing operations.