Repository navigation
[DOCS-15714] Add Node.js experiment setup instructions - #40594
mehulsonowal wants to merge 3 commits into
Conversation
|
Opened DOCS-15968 to assign a Docs writer and follow up with editorial review. |
domalessi
left a comment
There was a problem hiding this comment.
Left some feedback in line. Almost there!
| NODE_OPTIONS="--import dd-trace/initialize.mjs" node <YOUR_APP_ENTRYPOINT> | ||
| ``` | ||
|
|
||
| For more information, see the [Node.js tracer command-line setup](/llm_observability/instrument/sdk?tab=nodejs#command-line-setup). |
There was a problem hiding this comment.
Links to other pages should be defined reference style. Please apply this feedback to all inline links (inline anchor links to sections on the same page are fine).
| {{% /tab %}} | ||
| {{< /tabs >}} | ||
|
|
||
| ### APM Trace correlation |
There was a problem hiding this comment.
| ### APM Trace correlation | |
| ### APM trace correlation |
| text: "Debug and evaluate your AI app from your coding agent with Datadog Agent Observability" | ||
| --- | ||
|
|
||
| This guide describes how to set up and use Agent Observability experiments with the Python or Node.js SDK. For complete runnable Node.js examples, see the [Node.js experiments examples](https://github.com/DataDog/llm-observability/tree/main/experiments/nodejs). |
There was a problem hiding this comment.
Link should also be reference style
| This guide describes how to set up and use Agent Observability experiments with the Python or Node.js SDK. For complete runnable Node.js examples, see the [Node.js experiments examples](https://github.com/DataDog/llm-observability/tree/main/experiments/nodejs). | |
| This guide describes how to set up and use Agent Observability experiments with the Python or Node.js SDK. See [runnable Node.js experiments examples](https://github.com/DataDog/llm-observability/tree/main/experiments/nodejs). |
There was a problem hiding this comment.
I wonder if we should have a more featured/explicit call-out to the Setup docs. You can find it in Further reading, but a user might not readily spot it. The page feels fairly sparse and odd without a direct pointer to setup instructions, I think?
|
|
||
| ### 3. Define evaluators | ||
|
|
||
| Evaluators measure how well your model or agent performs on each record. Both SDKs support function-based evaluators and reusable class-based evaluators. |
There was a problem hiding this comment.
Clarify the distinction between local evaluators defined as functions or classes and managed remote evaluators, so the later remote-evaluator section does not appear to introduce an unexplained third type.
| Evaluators measure how well your model or agent performs on each record. Both SDKs support function-based evaluators and reusable class-based evaluators. | |
| Evaluators measure how well your model or agent performs on each record. Both SDKs support local evaluators defined as functions or classes. The Node.js SDK also supports managed remote evaluators, which run evaluations configured in Datadog. |
| {{% /tab %}} | ||
|
|
||
| {{% tab "Node.js" %}} | ||
| The Node.js SDK supports function and class-based evaluators. Return a Boolean, number, string, or JSON-serializable object. Use an object map when you want to assign evaluator names explicitly. |
There was a problem hiding this comment.
Label the first example as function-based and remove the repeated SDK overview. This also makes the Node.js tab parallel to the Python tab.
| The Node.js SDK supports function and class-based evaluators. Return a Boolean, number, string, or JSON-serializable object. Use an object map when you want to assign evaluator names explicitly. | |
| #### Function-based evaluators | |
| Return a Boolean, number, string, or JSON-serializable object. Use an object map to assign evaluator names explicitly. |
| {{% /tab %}} | ||
|
|
||
| {{% tab "Node.js" %}} | ||
| Summary evaluators are functions that receive the inputs, outputs, expected outputs, and an object containing the results from each evaluator. The optional fifth argument contains record metadata. |
There was a problem hiding this comment.
Add a heading for the function-based example, parallel to the Python tab and the class-based subsection that follows. Scope the explanation to functions rather than implying that all summary evaluators are functions.
| Summary evaluators are functions that receive the inputs, outputs, expected outputs, and an object containing the results from each evaluator. The optional fifth argument contains record metadata. | |
| #### Function-based summary evaluators | |
| Summary evaluator functions receive the inputs, outputs, expected outputs, and an object containing the results from each evaluator. The optional fifth argument contains record metadata. |
| {{% /tab %}} | ||
| {{< /tabs >}} | ||
|
|
||
| ### 5. Create and run the experiment. |
There was a problem hiding this comment.
Introduce the first examples so readers can distinguish creating and running an experiment from the execution options shown later. Also remove the period from the heading.
| ### 5. Create and run the experiment. | |
| ### 5. Create and run the experiment | |
| Create an experiment with your dataset, task, and evaluators, then run it and inspect the results: |
| {{% /tab %}} | ||
| {{< /tabs >}} | ||
|
|
||
| ### 6. Review your experiment results in Datadog. |
There was a problem hiding this comment.
| ### 6. Review your experiment results in Datadog. | |
| ### 6. Review your experiment results in Datadog |
| ```python | ||
| def num_exact_matches(inputs, outputs, expected_outputs, evaluators_results): | ||
| return evaluators_results["exact_match"].count(True) | ||
| ``` | ||
|
|
||
| Summary evaluator functions can take a list of any non-null type as `inputs` (string, number, Boolean, object, or array); `outputs` and `expected_outputs` can be lists of any type. `evaluators_results` is a dictionary of lists of results from evaluators, keyed by the name of the evaluator function. |
There was a problem hiding this comment.
| ```python | |
| def num_exact_matches(inputs, outputs, expected_outputs, evaluators_results): | |
| return evaluators_results["exact_match"].count(True) | |
| ``` | |
| Summary evaluator functions can take a list of any non-null type as `inputs` (string, number, Boolean, object, or array); `outputs` and `expected_outputs` can be lists of any type. `evaluators_results` is a dictionary of lists of results from evaluators, keyed by the name of the evaluator function. | |
| Summary evaluator functions can take a list of any non-null type as `inputs` (string, number, Boolean, object, or array); `outputs` and `expected_outputs` can be lists of any type. `evaluators_results` is a dictionary of lists of results from evaluators, keyed by the name of the evaluator function. | |
| ```python | |
| def num_exact_matches(inputs, outputs, expected_outputs, evaluators_results): | |
| return evaluators_results["exact_match"].count(True) | |
| ``` |
|
🔄 Datadog auto-retried 2 jobs - 0 passed on retry 🔗 Commit SHA: 3bc2cfb | Docs | View more details | Give us feedback! |
What does this PR do? What is the motivation?
Fixes DOCS-15714
Adds Node.js setup and usage instructions for Agent Observability experiments and reorganizes the detailed experiment documentation under Guides. This PR supersedes #39963 and contains the finalized documentation changes with all authored commits GitHub-verified under the
mehulsonowalaccount.Merge readiness
AI assistance
Documentation edits were prepared with Pi and validated with the repository hooks and whitespace checks.
Additional notes
The previous PR is closed in favor of this branch. All authored commits on this branch use the
mehulsonowalGitHub signing identity.