They do. Internally, the gatekeepers of this data are mostly old-school chemists whose ideas of AI are stuck in the "maybe AI can do some property predictions better than our first principles FORTRAN 77 models, but I doubt it" mindset.
That is just false. The gatekeepers of this data see a new algorithm, get excited, try it, test it against the aforementioned model, and then keep the old model because it's just better /s.
Real talk now, I'm an "AI person" from big pharma, and we're quite up to date. topological neural networwks, diffusion models, QM neural potentials, large scale meta-learning, systematic active learning, we're doing it all. We also know that most of the time, a small bayesian GLM or random forest on run-of-the-mill descriptors actually works very well and fails predictably, which is important.
Data in pharma in particular tends to be sparse and shallow: a hundred datapoints clustered tightly in chemical space because that's what the process generates. Sharing the data can lead competitors to the precious IP you're protecting, which is why we also invest a lot in blind federated learning etc.
Anyway, the going's tough, but everyone is doing their best. nobody's dismissing AI at all... we're just more aware of the domain-specific pitfalls.
No. I'm talking about things like polymer rheological or film properties, or industrial scale reactor performance. That's not to mention one-offs like cloud point predictors, or specialized thermo packages.