11/ P5: Okay, so the first shocking thing about this table is how low even the best success rate is for atomic calls, which in theory should be “easy”. Wonder what’s going on here? GPT-4 is usually very good at precisely populating parameter values.Also, the success rate for full tasks for the open models are uselessly low. I question what the point is for even including them. Going from a 0% success rate to a 2% success rate is completely irrelevant.Also, the list of models is pretty outdated. Who the heck still uses davinci-002?? Where’s Mixtral? The January GPT-4-turbo? I know that didn’t come out that long ago but it can’t be that hard to rerun the checks with a new model string.
cited on: mixtral-8x7b
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.