Skip to content

Tag

#Testing

Evaluating LLM Output in CI

Retrieval-augmented generation demos beautifully and fails quietly. The gap between a notebook that answers questions about ten PDFs and a system that serves an entire organisation is not model quality — it is retrieval...

7 min read 13