I am biased because I work with Corin on tooling for chemistry/physics/biology.
TLDR: Building a strong tooling layer will probably become very important for increasingly powerful AI models. Crystallized intelligence is effective, and also lends itself to better interpretability (think CoT and sequences of actions performed by the model). No need to “reinvent the wheel” every time you do a new task.
This paragraph from this OpenAI blog about LifeSciBench, a life sciences reasoning benchmark, supports Corin’s thesis:
Performance remains much weaker on artifact-heavy, design-heavy, and operationally constrained scientific work. Namely, Design, Optimization, & Prediction remains one of the hardest workflows, with GPT‑Rosalind passrate at 30.7%; Analysis is similarly difficult at 30.3%.
Artifact use is a particularly clear gap. While GPT‑Rosalind performs better than GPT‑5.5 in artifact-heavy settings, its pass rate still drops from 45.1% on text-only tasks to 28.1% on tasks with artifacts or URLs. GPT‑5.5 shows the same pattern, dropping from 29.9% to 21.9%. A more detailed analysis confirms that frontier models struggle at extracting information from complex figures or large sequence files and integrating that information into the final answer.
vibing inverse NMR is probably too hard!!!!