llama-stack
llama-stack copied to clipboard
[rag evals][2/n] add more braintrust scoring fns for RAG eval
What does this PR do?
- add more braintrust scoring functions for RAG eval
- add tests for evaluating against context
Test Plan
pytest -v -s -m braintrust_scoring_together_inference scoring/test_scoring.py
Example Output
- https://gist.github.com/yanxi0830/2acf3b8b3e8132fda2a48b1f0a49711b
Sources
Please link relevant resources if necessary.
Before submitting
- [ ] This PR fixes a typo or improves the docs (you can dismiss the other checks if that's the case).
- [ ] Ran pre-commit to handle lint / formatting issues.
- [ ] Read the contributor guideline, Pull Request section?
- [ ] Updated relevant documentation.
- [ ] Wrote necessary unit or integration tests.