codekingpro/portable-devtools
114k
1"""**Evaluation** chains for grading LLM and Chain outputs.2 3This module contains off-the-shelf evaluation chains for grading the output of4LangChain primitives such as language models and chains.5 6**Loading an evaluator**7 8To load an evaluator, you can use the `load_evaluators <langchain.evaluation.loading.load_evaluators>` or9`load_evaluator <langchain.evaluation.loading.load_evaluator>` functions with the10names of the evaluators to load.11 12```python13from langchain_classic.evaluation import load_evaluator14 15evaluator = load_evaluator("qa")16evaluator.evaluate_strings(17 prediction="We sold more than 40,000 units last week",18 input="How many units did we sell last week?",19 reference="We sold 32,378 units",20)21```22 23The evaluator must be one of `EvaluatorType <langchain.evaluation.schema.EvaluatorType>`.24 25**Datasets**26 27To load one of the LangChain HuggingFace datasets, you can use the `load_dataset <langchain.evaluation.loading.load_dataset>` function with the28name of the dataset to load.29 30```python31from langchain_classic.evaluation import load_dataset32 33ds = load_dataset("llm-math")34```35 36**Some common use cases for evaluation include:**37 38- Grading the accuracy of a response against ground truth answers: `QAEvalChain <langchain.evaluation.qa.eval_chain.QAEvalChain>`39- Comparing the output of two models: `PairwiseStringEvalChain <langchain.evaluation.comparison.eval_chain.PairwiseStringEvalChain>` or `LabeledPairwiseStringEvalChain <langchain.evaluation.comparison.eval_chain.LabeledPairwiseStringEvalChain>` when there is additionally a reference label.40- Judging the efficacy of an agent's tool usage: `TrajectoryEvalChain <langchain.evaluation.agents.trajectory_eval_chain.TrajectoryEvalChain>`41- Checking whether an output complies with a set of criteria: `CriteriaEvalChain <langchain.evaluation.criteria.eval_chain.CriteriaEvalChain>` or `LabeledCriteriaEvalChain <langchain.evaluation.criteria.eval_chain.LabeledCriteriaEvalChain>` when there is additionally a reference label.42- Computing semantic difference between a prediction and reference: `EmbeddingDistanceEvalChain <langchain.evaluation.embedding_distance.base.EmbeddingDistanceEvalChain>` or between two predictions: `PairwiseEmbeddingDistanceEvalChain <langchain.evaluation.embedding_distance.base.PairwiseEmbeddingDistanceEvalChain>`43- Measuring the string distance between a prediction and reference `StringDistanceEvalChain <langchain.evaluation.string_distance.base.StringDistanceEvalChain>` or between two predictions `PairwiseStringDistanceEvalChain <langchain.evaluation.string_distance.base.PairwiseStringDistanceEvalChain>`44 45**Low-level API**46 47These evaluators implement one of the following interfaces:48 49- `StringEvaluator <langchain.evaluation.schema.StringEvaluator>`: Evaluate a prediction string against a reference label and/or input context.50- `PairwiseStringEvaluator <langchain.evaluation.schema.PairwiseStringEvaluator>`: Evaluate two prediction strings against each other. Useful for scoring preferences, measuring similarity between two chain or llm agents, or comparing outputs on similar inputs.51- `AgentTrajectoryEvaluator <langchain.evaluation.schema.AgentTrajectoryEvaluator>` Evaluate the full sequence of actions taken by an agent.52 53These interfaces enable easier composability and usage within a higher level evaluation framework.54 55""" # noqa: E50156 57from langchain_classic.evaluation.agents import TrajectoryEvalChain58from langchain_classic.evaluation.comparison import (59 LabeledPairwiseStringEvalChain,60 PairwiseStringEvalChain,61)62from langchain_classic.evaluation.criteria import (63 Criteria,64 CriteriaEvalChain,65 LabeledCriteriaEvalChain,66)67from langchain_classic.evaluation.embedding_distance import (68 EmbeddingDistance,69 EmbeddingDistanceEvalChain,70 PairwiseEmbeddingDistanceEvalChain,71)72from langchain_classic.evaluation.exact_match.base import ExactMatchStringEvaluator73from langchain_classic.evaluation.loading import (74 load_dataset,75 load_evaluator,76 load_evaluators,77)78from langchain_classic.evaluation.parsing.base import (79 JsonEqualityEvaluator,80 JsonValidityEvaluator,81)82from langchain_classic.evaluation.parsing.json_distance import JsonEditDistanceEvaluator83from langchain_classic.evaluation.parsing.json_schema import JsonSchemaEvaluator84from langchain_classic.evaluation.qa import (85 ContextQAEvalChain,86 CotQAEvalChain,87 QAEvalChain,88)89from langchain_classic.evaluation.regex_match.base import RegexMatchStringEvaluator90from langchain_classic.evaluation.schema import (91 AgentTrajectoryEvaluator,92 EvaluatorType,93 PairwiseStringEvaluator,94 StringEvaluator,95)96from langchain_classic.evaluation.scoring import (97 LabeledScoreStringEvalChain,98 ScoreStringEvalChain,99)100from langchain_classic.evaluation.string_distance import (101 PairwiseStringDistanceEvalChain,102 StringDistance,103 StringDistanceEvalChain,104)105 106__all__ = [107 "AgentTrajectoryEvaluator",108 "ContextQAEvalChain",109 "CotQAEvalChain",110 "Criteria",111 "CriteriaEvalChain",112 "EmbeddingDistance",113 "EmbeddingDistanceEvalChain",114 "EvaluatorType",115 "ExactMatchStringEvaluator",116 "JsonEditDistanceEvaluator",117 "JsonEqualityEvaluator",118 "JsonSchemaEvaluator",119 "JsonValidityEvaluator",120 "LabeledCriteriaEvalChain",121 "LabeledPairwiseStringEvalChain",122 "LabeledScoreStringEvalChain",123 "PairwiseEmbeddingDistanceEvalChain",124 "PairwiseStringDistanceEvalChain",125 "PairwiseStringEvalChain",126 "PairwiseStringEvaluator",127 "QAEvalChain",128 "RegexMatchStringEvaluator",129 "ScoreStringEvalChain",130 "StringDistance",131 "StringDistanceEvalChain",132 "StringEvaluator",133 "TrajectoryEvalChain",134 "load_dataset",135 "load_evaluator",136 "load_evaluators",137]138 