Skip to main content
Github

❊ Info

These evaluators run a defined function on the response. How does it work A function evaluator runs a provided function along with the arguments for this function on the response and return whether the function passed or not. Required Args Your dataset must contain these fields:
  • response: The LLM generated response for the user query
Metrics
  • Passed: Boolean(True/False) value specifying whether the function passed or not.

▷ Run the function eval on a single datapoint


▷ Run the function eval on a dataset

  1. Load your data into a dictionary
  1. Run the evaluator on your dataset

Following are examples of the various function evaluators we support

Regex

Description: Checks if the response contains the regex pattern. Arguments:
  • pattern: str Pattern to search for.
Sample Code:

Contains Any

Description: Checks if the response contains any word from the list of keywords. Arguments:
  • keywords: List[str] List of keywords
  • case_sensitive: Optional[bool]. Defaults to False.
Sample Code:

Contains None

Description: Checks if the response does not contain any of the specified substrings. Arguments:
  • keywords: List of strings - keywords to check for absence in the context.
Sample Code:

Contains

Description:
Checks if the response contains the specified keyword.
Arguments:
  • keyword: string to check for presence in the response.
Sample Code:

ContainsAll

Description:
Checks if all the provided keywords are present in the response.
Arguments:
  • keywords: List[str] - The list of keywords to search for in the response.
  • case_sensitive: bool, optional - If True, the comparison is case-sensitive. Defaults to False.
Sample Code:

ContainsJson

Description:
Checks if the response contains a valid JSON.
Arguments:
  • None
Sample Code:

ContainsEmail

Description:
Checks if the response contains a valid email address.
Arguments:
  • None
Sample Code:

IsJson

Description:
Checks if the response is a valid JSON.
Arguments:
  • None
Sample Code:

IsEmail

Description:
Checks if the response is a valid email address.
Arguments:
  • None
Sample Code:
Description:
Checks if the response contains any links.
Arguments:
  • None
Sample Code:
Description:
Checks if the response contains valid links.
Arguments:
  • None
Sample Code:
Description:
Checks if the response does not contain any invalid links.
Arguments:
  • None
Sample Code:

ApiCall

Description:
Performs an API call to a specified endpoint and picks up the evaluation result from the response. This evaluator is useful when you want to run some complex or custom logic on the response.
Arguments:
  • url: string - API endpoint to call. Note that this API should accept POST request.
  • headers: dict - Headers to include in the API call.
  • payload: dict - Body to send with the API call. This payload will have the Response added to it.
Sample Code:
  • We expect the API response to be in JSON format with two keys namely result and reason. - The result key should contain the evaluation result which should be a boolean value. - The reason key should contain the reason for the evaluation result which should be a string. - The dataset should contain the response and optionally the query, context and expected_response to be passed to the API.

Equals

Description: Checks if the response is exactly equal to the specified string. Arguments:
  • expected_response: str String to compare the response with.
Sample Code:

StartsWith

Description: checks if the response starts with the specified substring. Arguments:
  • substring: str string to check at the start of the response.
Sample Code:

EndsWith

Description: checks if the response ends with the specified substring. Arguments:
  • substring: str string to check at the end of the response.
Sample Code:

LengthLessThan

Description: Checks if the length of the response is less than a maximum length. Arguments:
  • max_length: int the maximum allowable length for the response.
Sample Code:

LengthGreaterThan

Description: Checks if the length of the response is more than a minimum length. Arguments:
  • min_length: int the minimum allowable length for the response.
Sample Code:

Length Between

Description: Checks if the length of the response is between the minimum and maximum length. Arguments:
  • min_length: int the minimum allowable length for the response.
  • max_length: int the maximum allowable length for the response.
Sample Code:

One Line

Description: Checks if the response is a single line. Arguments:
  • None
Sample Code:

CustomCodeEval

Description: Runs a custom code as an evaluator. Arguments:
  • code: str Code to be executed. The code should contain a function named main which takes **kwargs as input and returns a boolean value.
Sample Code:
Read more about CustomCodeEval

JsonSchema

Description: Validates the JSON structure against a specified JSON schema. Arguments:
  • schema: str The JSON schema to validate against.
Sample Code:

JsonValidation

Description: Validates the value of a JSON field against a specified condition. Arguments: validations: list A list of validation rules. Each rule is a dictionary with the following keys: json_path: str The JSON path to the field to validate. validating_function: str The name of the validation function to use.
  • validations: list The validations list
Sample Code: