Jev makes it easier to get a classifier running. And with the excitement that comes with new tools, it is easy to overlook what it takes to keep it doing its job.
Plenty of people talk about that kind of often invisible work, including the TypeSafe AI folks. But in the conversations around me, the tool still gets most of the attention. And I say that as someone who’s usually the first to get excited about these tools and this fascinating, terrifying Language Machine.
In 2016, we built and operated a classifier at Apple’s Beta Program to take over the toil of sorting feedback from customers running beta versions of the OS. Before that, during crunch season, more than 200 people categorized the incoming feedback and routed it to the right product team. Our early discussions covered embeddings, vectors, and cosine distance. We also had to define success, maintain labeled examples and holdouts, review mistakes, revise the categories, and work with the teams using the results. Production failures became cases to test against. Somebody had to own that loop. Back then we didn’t call it a loop. It was just work. :)
Golden datasets and prompt optimization help. For example, when working with DSPy.rb, I used a scoring rule that rewarded catching more cases at the cost of more false alarms. Someone still has to decide which mistakes matter, and whether a better score improves the work.
So you’re thinking of adding a classifier to your product. Do it. Just make sure the tool doesn’t end up deciding how your team works.