Team Ai
Datasetpublic

codekingpro/portable-devtools

sourceHugging Faceupdated 5mo agoView on Hugging Face
1likes14kdownloads
__init__.py138 linesDownload Raw Back to evaluation
1"""**Evaluation** chains for grading LLM and Chain outputs.2 3This module contains off-the-shelf evaluation chains for grading the output of4LangChain primitives such as language models and chains.5 6**Loading an evaluator**7 8To load an evaluator, you can use the `load_evaluators <langchain.evaluation.loading.load_evaluators>` or9`load_evaluator <langchain.evaluation.loading.load_evaluator>` functions with the10names of the evaluators to load.11 12```python13from langchain_classic.evaluation import load_evaluator14 15evaluator = load_evaluator("qa")16evaluator.evaluate_strings(17    prediction="We sold more than 40,000 units last week",18    input="How many units did we sell last week?",19    reference="We sold 32,378 units",20)21```22 23The evaluator must be one of `EvaluatorType <langchain.evaluation.schema.EvaluatorType>`.24 25**Datasets**26 27To load one of the LangChain HuggingFace datasets, you can use the `load_dataset <langchain.evaluation.loading.load_dataset>` function with the28name of the dataset to load.29 30```python31from langchain_classic.evaluation import load_dataset32 33ds = load_dataset("llm-math")34```35 36**Some common use cases for evaluation include:**37 38- Grading the accuracy of a response against ground truth answers: `QAEvalChain <langchain.evaluation.qa.eval_chain.QAEvalChain>`39- Comparing the output of two models: `PairwiseStringEvalChain <langchain.evaluation.comparison.eval_chain.PairwiseStringEvalChain>` or `LabeledPairwiseStringEvalChain <langchain.evaluation.comparison.eval_chain.LabeledPairwiseStringEvalChain>` when there is additionally a reference label.40- Judging the efficacy of an agent's tool usage: `TrajectoryEvalChain <langchain.evaluation.agents.trajectory_eval_chain.TrajectoryEvalChain>`41- Checking whether an output complies with a set of criteria: `CriteriaEvalChain <langchain.evaluation.criteria.eval_chain.CriteriaEvalChain>` or `LabeledCriteriaEvalChain <langchain.evaluation.criteria.eval_chain.LabeledCriteriaEvalChain>` when there is additionally a reference label.42- Computing semantic difference between a prediction and reference: `EmbeddingDistanceEvalChain <langchain.evaluation.embedding_distance.base.EmbeddingDistanceEvalChain>` or between two predictions: `PairwiseEmbeddingDistanceEvalChain <langchain.evaluation.embedding_distance.base.PairwiseEmbeddingDistanceEvalChain>`43- Measuring the string distance between a prediction and reference `StringDistanceEvalChain <langchain.evaluation.string_distance.base.StringDistanceEvalChain>` or between two predictions `PairwiseStringDistanceEvalChain <langchain.evaluation.string_distance.base.PairwiseStringDistanceEvalChain>`44 45**Low-level API**46 47These evaluators implement one of the following interfaces:48 49- `StringEvaluator <langchain.evaluation.schema.StringEvaluator>`: Evaluate a prediction string against a reference label and/or input context.50- `PairwiseStringEvaluator <langchain.evaluation.schema.PairwiseStringEvaluator>`: Evaluate two prediction strings against each other. Useful for scoring preferences, measuring similarity between two chain or llm agents, or comparing outputs on similar inputs.51- `AgentTrajectoryEvaluator <langchain.evaluation.schema.AgentTrajectoryEvaluator>` Evaluate the full sequence of actions taken by an agent.52 53These interfaces enable easier composability and usage within a higher level evaluation framework.54 55"""  # noqa: E50156 57from langchain_classic.evaluation.agents import TrajectoryEvalChain58from langchain_classic.evaluation.comparison import (59    LabeledPairwiseStringEvalChain,60    PairwiseStringEvalChain,61)62from langchain_classic.evaluation.criteria import (63    Criteria,64    CriteriaEvalChain,65    LabeledCriteriaEvalChain,66)67from langchain_classic.evaluation.embedding_distance import (68    EmbeddingDistance,69    EmbeddingDistanceEvalChain,70    PairwiseEmbeddingDistanceEvalChain,71)72from langchain_classic.evaluation.exact_match.base import ExactMatchStringEvaluator73from langchain_classic.evaluation.loading import (74    load_dataset,75    load_evaluator,76    load_evaluators,77)78from langchain_classic.evaluation.parsing.base import (79    JsonEqualityEvaluator,80    JsonValidityEvaluator,81)82from langchain_classic.evaluation.parsing.json_distance import JsonEditDistanceEvaluator83from langchain_classic.evaluation.parsing.json_schema import JsonSchemaEvaluator84from langchain_classic.evaluation.qa import (85    ContextQAEvalChain,86    CotQAEvalChain,87    QAEvalChain,88)89from langchain_classic.evaluation.regex_match.base import RegexMatchStringEvaluator90from langchain_classic.evaluation.schema import (91    AgentTrajectoryEvaluator,92    EvaluatorType,93    PairwiseStringEvaluator,94    StringEvaluator,95)96from langchain_classic.evaluation.scoring import (97    LabeledScoreStringEvalChain,98    ScoreStringEvalChain,99)100from langchain_classic.evaluation.string_distance import (101    PairwiseStringDistanceEvalChain,102    StringDistance,103    StringDistanceEvalChain,104)105 106__all__ = [107    "AgentTrajectoryEvaluator",108    "ContextQAEvalChain",109    "CotQAEvalChain",110    "Criteria",111    "CriteriaEvalChain",112    "EmbeddingDistance",113    "EmbeddingDistanceEvalChain",114    "EvaluatorType",115    "ExactMatchStringEvaluator",116    "JsonEditDistanceEvaluator",117    "JsonEqualityEvaluator",118    "JsonSchemaEvaluator",119    "JsonValidityEvaluator",120    "LabeledCriteriaEvalChain",121    "LabeledPairwiseStringEvalChain",122    "LabeledScoreStringEvalChain",123    "PairwiseEmbeddingDistanceEvalChain",124    "PairwiseStringDistanceEvalChain",125    "PairwiseStringEvalChain",126    "PairwiseStringEvaluator",127    "QAEvalChain",128    "RegexMatchStringEvaluator",129    "ScoreStringEvalChain",130    "StringDistance",131    "StringDistanceEvalChain",132    "StringEvaluator",133    "TrajectoryEvalChain",134    "load_dataset",135    "load_evaluator",136    "load_evaluators",137]138 
codekingpro/portable-devtools · Team Ai