Skip to content

[DOCS-15714] Add Node.js experiment setup instructions - #40594

Open
mehulsonowal wants to merge 3 commits into
masterfrom
docs/nodejs-experiments-setup-mehulsonowal
Open

mehulsonowal wants to merge 3 commits into
masterfrom
docs/nodejs-experiments-setup-mehulsonowal

Conversation

@mehulsonowal

@mehulsonowal mehulsonowal commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

What does this PR do? What is the motivation?

Fixes DOCS-15714

Adds Node.js setup and usage instructions for Agent Observability experiments and reorganizes the detailed experiment documentation under Guides. This PR supersedes #39963 and contains the finalized documentation changes with all authored commits GitHub-verified under the mehulsonowal account.

Merge readiness

  • Ready for merge

AI assistance

Documentation edits were prepared with Pi and validated with the repository hooks and whitespace checks.

Additional notes

The previous PR is closed in favor of this branch. All authored commits on this branch use the mehulsonowal GitHub signing identity.

@mehulsonowal
mehulsonowal requested a review from a team as a code owner October 9, 2026 16:51
@github-actions github-actions Bot added Architecture Everything related to the Doc backend Guide Content impacting a guide labels Oct 9, 2026
@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Preview links (active after the build_preview check completes)

New or renamed files

Removed or renamed files (these should redirect)

Renamed files

Modified Files

@joepeeples joepeeples added the editorial review Waiting on a more in-depth review label Oct 9, 2026
@joepeeples

Copy link
Copy Markdown
Contributor

Opened DOCS-15968 to assign a Docs writer and follow up with editorial review.

@domalessi domalessi left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Left some feedback in line. Almost there!

NODE_OPTIONS="--import dd-trace/initialize.mjs" node <YOUR_APP_ENTRYPOINT>
```

For more information, see the [Node.js tracer command-line setup](/llm_observability/instrument/sdk?tab=nodejs#command-line-setup).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Links to other pages should be defined reference style. Please apply this feedback to all inline links (inline anchor links to sections on the same page are fine).

{{% /tab %}}
{{< /tabs >}}

### APM Trace correlation

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### APM Trace correlation
### APM trace correlation

text: "Debug and evaluate your AI app from your coding agent with Datadog Agent Observability"
---

This guide describes how to set up and use Agent Observability experiments with the Python or Node.js SDK. For complete runnable Node.js examples, see the [Node.js experiments examples](https://github.com/DataDog/llm-observability/tree/main/experiments/nodejs).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Link should also be reference style

Suggested change
This guide describes how to set up and use Agent Observability experiments with the Python or Node.js SDK. For complete runnable Node.js examples, see the [Node.js experiments examples](https://github.com/DataDog/llm-observability/tree/main/experiments/nodejs).
This guide describes how to set up and use Agent Observability experiments with the Python or Node.js SDK. See [runnable Node.js experiments examples](https://github.com/DataDog/llm-observability/tree/main/experiments/nodejs).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder if we should have a more featured/explicit call-out to the Setup docs. You can find it in Further reading, but a user might not readily spot it. The page feels fairly sparse and odd without a direct pointer to setup instructions, I think?


### 3. Define evaluators

Evaluators measure how well your model or agent performs on each record. Both SDKs support function-based evaluators and reusable class-based evaluators.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clarify the distinction between local evaluators defined as functions or classes and managed remote evaluators, so the later remote-evaluator section does not appear to introduce an unexplained third type.

Suggested change
Evaluators measure how well your model or agent performs on each record. Both SDKs support function-based evaluators and reusable class-based evaluators.
Evaluators measure how well your model or agent performs on each record. Both SDKs support local evaluators defined as functions or classes. The Node.js SDK also supports managed remote evaluators, which run evaluations configured in Datadog.

{{% /tab %}}

{{% tab "Node.js" %}}
The Node.js SDK supports function and class-based evaluators. Return a Boolean, number, string, or JSON-serializable object. Use an object map when you want to assign evaluator names explicitly.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Label the first example as function-based and remove the repeated SDK overview. This also makes the Node.js tab parallel to the Python tab.

Suggested change
The Node.js SDK supports function and class-based evaluators. Return a Boolean, number, string, or JSON-serializable object. Use an object map when you want to assign evaluator names explicitly.
#### Function-based evaluators
Return a Boolean, number, string, or JSON-serializable object. Use an object map to assign evaluator names explicitly.

{{% /tab %}}

{{% tab "Node.js" %}}
Summary evaluators are functions that receive the inputs, outputs, expected outputs, and an object containing the results from each evaluator. The optional fifth argument contains record metadata.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Add a heading for the function-based example, parallel to the Python tab and the class-based subsection that follows. Scope the explanation to functions rather than implying that all summary evaluators are functions.

Suggested change
Summary evaluators are functions that receive the inputs, outputs, expected outputs, and an object containing the results from each evaluator. The optional fifth argument contains record metadata.
#### Function-based summary evaluators
Summary evaluator functions receive the inputs, outputs, expected outputs, and an object containing the results from each evaluator. The optional fifth argument contains record metadata.

{{% /tab %}}
{{< /tabs >}}

### 5. Create and run the experiment.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Introduce the first examples so readers can distinguish creating and running an experiment from the execution options shown later. Also remove the period from the heading.

Suggested change
### 5. Create and run the experiment.
### 5. Create and run the experiment
Create an experiment with your dataset, task, and evaluators, then run it and inspect the results:

{{% /tab %}}
{{< /tabs >}}

### 6. Review your experiment results in Datadog.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### 6. Review your experiment results in Datadog.
### 6. Review your experiment results in Datadog

Comment on lines +624 to +629
```python
def num_exact_matches(inputs, outputs, expected_outputs, evaluators_results):
return evaluators_results["exact_match"].count(True)
```

Summary evaluator functions can take a list of any non-null type as `inputs` (string, number, Boolean, object, or array); `outputs` and `expected_outputs` can be lists of any type. `evaluators_results` is a dictionary of lists of results from evaluators, keyed by the name of the evaluator function.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
```python
def num_exact_matches(inputs, outputs, expected_outputs, evaluators_results):
return evaluators_results["exact_match"].count(True)
```
Summary evaluator functions can take a list of any non-null type as `inputs` (string, number, Boolean, object, or array); `outputs` and `expected_outputs` can be lists of any type. `evaluators_results` is a dictionary of lists of results from evaluators, keyed by the name of the evaluator function.
Summary evaluator functions can take a list of any non-null type as `inputs` (string, number, Boolean, object, or array); `outputs` and `expected_outputs` can be lists of any type. `evaluators_results` is a dictionary of lists of results from evaluators, keyed by the name of the evaluator function.
```python
def num_exact_matches(inputs, outputs, expected_outputs, evaluators_results):
return evaluators_results["exact_match"].count(True)
```

@datadog-official

datadog-official Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Tests

🔄 Datadog auto-retried 2 jobs - 0 passed on retry View in Datadog

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 3bc2cfb | Docs | View more details | Give us feedback!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Architecture Everything related to the Doc backend editorial review Waiting on a more in-depth review Guide Content impacting a guide

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants