Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> It reduces prompt costs by half for those shared prefix tokens, but you have to pay $4.50/million tokens/hour to keep that cache warm - so probably not a useful optimization for most lower traffic applications

That's on a model with $3.5/1M input token cost, so half price on cached prefix tokens for $4.5/1M/hour breaks even at a little over 2.5 requests/hour using the cached prefix.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: