According to a recent LinkedIn post from SuperAnnotate, travel platform GetYourGuide has reportedly centralized its large language model evaluation workflows using a purpose-built tool. The post highlights commentary from GetYourGuide’s Director of Engineering, who suggests that moving away from spreadsheet-based evaluations reduced errors and improved reliability.
The company’s LinkedIn post describes four LLM evaluation types—chat, classification, extraction and free-form grading—being consolidated into a single schema-driven workflow with automatic routing and versioning by model, prompt and dataset. This structure appears to allow new use cases to integrate into an existing framework rather than launching separate ad hoc processes.
The post further suggests that the tooling places subject matter experts at the center of the evaluation design, with limited engineering overhead. For investors, this emphasis on scalable, expert-led evaluation could indicate that SuperAnnotate is positioning its platform as infrastructure for enterprises seeking robust, repeatable LLM testing and governance.
If adopted more broadly, such workflow tooling may enhance customer stickiness and support higher-value enterprise contracts, especially in sectors where AI reliability and auditability are commercially critical. The focus on reducing errors and standardizing evaluation across use cases could strengthen SuperAnnotate’s competitive stance in the AI data and tooling ecosystem, potentially supporting long-term revenue growth.

