Skip to content
Topic Hub
Prof Rod Avatar

Filed under Evaluation

1 entry and 3 thoughts filed under this tag.

Thoughts on Evaluation

A saved prompt is not a complete experiment record. The model, relevant settings and test inputs can all affect the result. When comparing an OpenAI workflow, record the exact model identifier used, the prompt version…

openaievaluation

Prompt tuning is hard to assess when the definition of “better” changes after every answer. Freeze a small acceptance set first. For a document-routing task, define the allowed destinations and label a collection of…

openaievaluation

A model can handle ordinary support tickets well and still fail on the exceptions that consume most of your team's time. Build a small evaluation set from representative, appropriately sanitized cases. Include a missing…

openaievaluation