I am asking the models to generate an image where fictional characters play chess or Texas Holdem. None of them can make a realistic chess position or poker game. Always something is off like too many pawns or too may cards, or some cards being ace-up when they shouldn't be.
Just ask LLM to write one on top of OpenRouter, AI SDK and Bun
To take your .md input file and save outputs as md files (or whatever you need)
Take https://github.com/T3-Content/auto-draftify as example
Yeah I’ve wondered about the same myself… My evals are also a pile of text snippets, as are some of my workflows. Thought I’d have a look to see what’s out there and found Promptfoo and Inspect AI. Haven’t tried either but will for my next round of evals
And it shouldn't be shared publicly so that the models won't learn about it accidentally :)