> Testing GPT-5, Claude, Gemini, Grok, and DeepSeek with $100K each over 8 month...

PTRFRLL · 2025-12-04T23:14:57 1764890097

> We were cautious to only run after each model’s training cutoff dates for the LLM models. That way we could be sure models couldn’t have memorized market outcomes.

stusmall · 2025-12-04T23:28:34 1764890914

Even if it is after the cut off date wouldn't the models be able to query external sources to get data that could positively impact them? If the returns were smaller I could reasonably believe it but beating the S&P500 returns by 4x+ strains credulity.

cheeseblubber · 2025-12-04T23:43:30 1764891810

We used the LLMs API and provided custom tools like a stock ticker tool that only gave stock price information for that date of backtest for the model. We did this for news apis, technical indicator apis etc. It took quite a long time to make sure that there weren't any data leakage. The whole process took us about a month or two to build out.

alchemist1e9 · 2025-12-05T00:14:58 1764893698

I have a hunch Grok model cutoff is not accurate and somehow it has updated weights though they still call it the same Grok model as the params and size are unchanged but they are incrementally training it in the background. Of course I don’t know this but it’s what I would do in their situation since ongoing incremental training could he a neat trick to improve their ongoing results against competitors, even if marginal. I also wouldn’t trust the models to honestly disclose their decision process either.

That said. This is a fascinating area of research and I do think LLM driven fundamental investing and trading has a future.

plufz · 2025-12-04T23:20:28 1764890428

I know very little about how the environment where they run these models look, but surely they have access to different tools like vector embeddings with more current data on various topics?

endtime · 2025-12-04T23:52:09 1764892329

If they could "see" the future and exploit that they'd probably have much higher returns.

plufz · 2025-12-05T08:06:15 1764921975

I would say that if these models independently could create such high returns all these companies would shut down the external access to the models and just have their own money making machine. :)

alchemist1e9 · 2025-12-05T00:16:00 1764893760

56% over 8 months with the constraints provided are pretty good results for Grok.

disconcision · 2025-12-04T23:44:53 1764891893

you can (via the api, or to a lesser degree through the setting in the web client) determine what tools if any a model can use

plufz · 2025-12-05T08:08:23 1764922103

But isn’t that more which MCP:s you can configure it to use? Do we have any idea which secret sauce stuff they have? Surely it’s not just a raw model that they are executing?

disconcision · 2025-12-04T23:55:25 1764892525

with the exception that it doesn't seem possible to fully disable this for grok 4

alchemist1e9 · 2025-12-05T00:16:24 1764893784

which is curiously the best model …

itake · 2025-12-04T23:14:59 1764890099

> We time segmented the APIs to make sure that the simulation isn’t leaking the future into the model’s context.

I wish they could explain what this actually means.

devmor · 2025-12-04T23:26:45 1764890805

It's a very silly way of saying that the data the LLMs had access to was presented in chronological order, so that for instance, when they were trading on stocks at the start of the 8 month window, the LLMs could not just query their APIs to see the data from the end of the 8 month window.

nullbound · 2025-12-04T23:25:57 1764890757

Overall, it does sound weird. On the one hand, assuming I properly I understand what they are saying is that they removed model's ability to cheat based on their specific training. And I do get that nuance ablation is a thing, but this is not what they are discussing there. They are only removing one avenue of the model to 'cheat'. For all we know, some that data may have been part of its training set already...

CPLX · 2025-12-04T23:14:11 1764890051

Not sure how sound the analysis is but they did apparently actually think of that.

joegibbs · 2025-12-04T23:23:46 1764890626

That's only if they're trained on data more recent than 8 months ago