Insights & Perspectives

Ideas for real-world impact.

Sarah Andrews Sarah Andrews

Stop Guessing About Your AI Maturity

Frontier AI labs are making massive leaps. There’s an explosion of research, accelerating in every industry. Then you look around at your own AI org, and progress feels... underwhelming. It’s easy to lay the blame on bad data, small budgets, or the wrong headcount, but the real problem is your org’s AI maturity — whether or not you understand the level you’ve reached, and what you need to work on next. We've built a tool to give you a clear picture of exactly these things — the AI Org Maturity Assessment.

Read More
Sarah Andrews Sarah Andrews

Why Research Agility Is the Most Important AI Metric You’re Not Tracking

ML infrastructure is painfully finicky, and using it entails managing a dozen irrelevant technical details just to do basic research. There are ideas that researchers don’t let themselves have, because engaging with them would mean weeks of infrastructure-wrangling before anyone gets to learn a thing. Sooner or later, your infrastructure starts setting your research agenda. We measure this as Research Agility; and it might be the most important metric in your organization.

Read More
Sarah Andrews Sarah Andrews

Data Is A Process, Not A Project

An uncomfortable claim: In pharma, biotech, and any serious applied AI, projects almost never fail because of the model. They fail because of the data. Everyone agrees that data matters, but not everyone is interested in funding the work needed to keep it useful. Data work has to run as a standing program, owned, scheduled, and budgeted, not as a one-time cleanup at the start of a project.

Read More
Sarah Andrews Sarah Andrews

Are My Evals Lying to Me?

For any model tackling any problem, the hardest part of evaluation is keeping the metric aligned with reality. An evaluation is just a measurement tool, and tools can be miscalibrated, biased, or simply pointed at the wrong thing. The urgent question is less: “Are my evals lying to me?” (They are, or they soon will be.) Instead, it’s: “When my evals do lie, how will I know — and what can I do about it?”

Read More

Hear Me Out: The Potential of Low-Latency Voice AI

Picture this: two users need advice on a health issue. One employs an AI text interface, resulting in a plan that leaves them feeling informed and empowered. A voice interface leads the other through a back-and-forth conversation with the AI; they feel cared for, supported. Same need, two very different experiences. All because of the interface.

Read More
Technical Deep Dives, Ethics/Guardrails Sarah Andrews Technical Deep Dives, Ethics/Guardrails Sarah Andrews

Leashing Your LLM: Practical and Cost-Saving Tips for Staying on Topic

The general nature of LLMs makes them inherently powerful but notoriously difficult to control. Operators must defend against unintended and potentially risky interactions. Our team has investigated many of the relatively nascent solutions out there for this issue; we share what we’ve learned in this post.

Read More

The Most Important Uses for LLMs Aren’t Chatbots

We love chatbots – ChatGPT and others in its class are amazing tools – but, as an AI consultancy with a long history of projects in the space before the current mania, we’re sensitive to the conflation of LLMs and chatbots. Many of the most exciting potential uses for LLMs have little to do with the chatbot interface, and we think those should get more attention.

Read More
Technical Deep Dives Sarah Andrews Technical Deep Dives Sarah Andrews

Engineer Better Research Results From a Solid Workbench

Treating the process of your work as important as the result will improve the quality of your results. A lot of focus gets put on building the right thing for customers, and rightfully so, but it’s important to remember that we have to first build our workbench. Whether we do that haphazardly or intentionally can have an enormous impact on the quality of our results.

Read More
Technical Deep Dives Sarah Andrews Technical Deep Dives Sarah Andrews

Testing Research Code: Is It Worth It?

Machine learning researchers often don’t write tests for their code. They’re not software engineers, and their code needs only to train a model or prove out an experiment. However, at Hop, we’ve found that adding certain kinds of tests can actually accelerate research and increase confidence in results through improving code quality and encouraging reuse.

Read More