Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
44 changes: 44 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
# Python
__pycache__/
*.py[cod]
*$py.class
*.so
.Python
build/
develop-eggs/
dist/
downloads/
eggs/
.eggs/
lib/
lib64/
parts/
sdist/
var/
wheels/
*.egg-info/
.installed.cfg
*.egg
MANIFEST

# Virtual environments
venv/
env/
ENV/
.venv

# IDEs
.vscode/
.idea/
*.swp
*.swo
*~

# OS
.DS_Store
Thumbs.db

# Test results
results/
output/
*.log
438 changes: 267 additions & 171 deletions README.md

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In my quick scan of the code and the readme, I'm having a hard time figuring out what the story is. It seems like examples are doing multiple things at once, instead of each illustrating a single concept. The readme descriptions of the example then doesn't enlighten the case much.

Below is my chat with Copilot about this:

===

What is this readme trying to show? The illustrative examples seem garbled and unclear and like each does multipel things -- just making them confusing. Are they confusing?

===

Yes — looking at this diff, the README is trying to do too much in each section, and it does make the examples harder to follow. A few concrete problems:

  1. Each "Try the X package" section mixes multiple execution modes without clearly separating them

For example, the CLI-only section shows, in sequence:

installing the package
running it as an installed shell command (my-dh-query)
then also importing the underlying function directly in Python (from my_dh_cli.cli import my_dh_query)
These are two different usage patterns (CLI invocation vs. programmatic call) bolted together under one "Try it" header, with only a short transition sentence ("The underlying function is also importable...") separating them. A reader skimming for "how do I just run this from the terminal" has to wade through Python import code too.

  1. The combined package section is the worst offender

It shows, back-to-back: two CLI commands (my-dh-toolkit-query, my-dh-toolkit-process), then a full server-startup snippet, then library imports, then calling the same underlying functions two different ways (from my_dh_toolkit.queries import ... and from my_dh_toolkit import my_dh_query, batch_process). That's four different ways to do roughly two things, all in one code block region. It's demonstrating flexibility, but at the cost of a clear "here's the one obvious way to use this" path.

  1. Repeated boilerplate obscures the actual point of each example

The Server(port=10000, jvm_args=["-Xmx4g"]).start() snippet is repeated near-verbatim in almost every section. That's necessary context, but it dilutes the examples — the reader has to mentally filter out the same 3 lines every time to find what's actually different about that example.

  1. Structural/informational overlap

There's now a "Packaging scenarios" section (when to use library-only/CLI-only/combined) and a separate "Execution patterns" section (entry points vs. module execution) and per-package "Usage" snippets at the top and "Try the X package" walkthroughs further down. That's four different places explaining overlapping concepts (when/how to use each package type), which is a lot for a reader to reconcile into one mental model.

Net assessment: The content itself is accurate and reasonably complete, but the presentation conflates "here's the simplest way to use this" with "here's every way you could possibly use this." A cleaner version would pick one primary usage pattern per package type for the main walkthrough (e.g., just the CLI command) and move the "you can also import the underlying function" variant into a clearly separate, secondary callout — rather than interleaving them as equally-weighted steps.

If you'd like, I can draft a tightened version of one of these sections (e.g., the combined package walkthrough) to show the difference concretely.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot reviw for the repo was set at a low / cheap level. I increased the level and reran it.

Large diffs are not rendered by default.

4 changes: 4 additions & 0 deletions data/batch/file1.csv
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
Name,Score,Value,Category
Kara,81,130,A
Liam,94,175,B
Mona,77,100,A
4 changes: 4 additions & 0 deletions data/batch/file2.csv
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
Name,Score,Value,Category
Nina,86,115,C
Omar,90,155,B
Pria,74,80,A
5 changes: 5 additions & 0 deletions data/batch/file3.csv
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
Name,Score,Value,Category
Quinn,93,165,C
Rosa,84,125,B
Sam,79,95,A
Tara,88,145,C
11 changes: 11 additions & 0 deletions data/sample.csv

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In reviewing this, I've had some confusion. This file looks like a concatenation of the prior files. I'm not sure if this is intentional, confusing, or just doesn't matter. I haven't read enough of the code to have a view. Just noting it here.

Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
Name,Score,Value,Category
Alice,85,120,A
Bob,92,150,B
Charlie,78,95,A
Diana,88,110,C
Eve,95,180,B
Frank,72,85,A
Grace,91,160,C
Henry,83,105,B
Iris,89,140,A
Jack,76,90,C
63 changes: 63 additions & 0 deletions my_dh_cli/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
# My Deephaven CLI

An example of packaging a Deephaven script as a command line tool. Installing this package creates one terminal command, `my-dh-query`. No library code is exposed; the package is used only through that command, so no Python needs to be written to use it.

The command is defined by the `[project.scripts]` entry point in [`pyproject.toml`](pyproject.toml):

```toml
[project.scripts]
my-dh-query = "my_dh_cli.cli:main"
```

## Installation

From the repository root:

```bash
pip install ./my_dh_cli
```

Or in editable mode for development:

```bash
pip install -e ./my_dh_cli
```

## Usage

Run the installed command on a CSV file. The command starts its own Deephaven server, so no separate setup is needed:

```bash
my-dh-query data/sample.csv --verbose
```

It reads the file, adds a `DoubleScore` computed column, and reports the number of rows processed.

The command binds its server to port 10000. If that port is already in use (for example, by Deephaven running in Docker), change the `port` value in [`cli.py`](src/my_dh_cli/cli.py).

During development, the package also runs without an entry point via [`__main__.py`](src/my_dh_cli/__main__.py):

```bash
python -m my_dh_cli data/sample.csv --verbose
```

## Command reference

### my-dh-query

Process a CSV file with Deephaven. The file must contain a `Score` column.

**Arguments:**

- `input_file` - Path to the CSV file to process.

**Options:**

- `--verbose, -v` - Enable verbose output.

## Requirements

- Python 3.9 or later
- Java 17 or later
- deephaven-server 0.35.0 or later (installed automatically as a dependency)
- Click 8.0.0 or later (installed automatically as a dependency)
22 changes: 22 additions & 0 deletions my_dh_cli/pyproject.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
[build-system]
requires = ["setuptools>=61.0", "wheel"]
build-backend = "setuptools.build_meta"

[project]
name = "my_dh_cli"
version = "0.1.0"
description = "Command line tool for data processing"
readme = "README.md"
requires-python = ">=3.9"
dependencies = [
# deephaven-server also provides the deephaven module (through its deephaven-core dependency).
"deephaven-server>=0.35.0",
# click implements the command line interface.
"click>=8.0.0",
]

[project.scripts]
my-dh-query = "my_dh_cli.cli:main"

[tool.setuptools.packages.find]
where = ["src"]
3 changes: 3 additions & 0 deletions my_dh_cli/src/my_dh_cli/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
"""Command line tool that processes a CSV file with Deephaven."""

__version__ = "0.1.0"
4 changes: 4 additions & 0 deletions my_dh_cli/src/my_dh_cli/__main__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
from my_dh_cli.cli import main

if __name__ == "__main__":
main()
61 changes: 61 additions & 0 deletions my_dh_cli/src/my_dh_cli/cli.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
import click


def my_dh_query(input_file: str, verbose: bool = False):
"""Read a CSV file and perform a simple query operation on the data."""
# Imported here, not at module level: deephaven requires a running server.
# The entry point starts the server first, then calls this function.
from deephaven import read_csv
from pathlib import Path

input_path = Path(input_file)

if not input_path.exists():
raise click.ClickException(f"Input file does not exist: '{input_path}'")
if not input_path.is_file():
raise click.ClickException(f"Input path is not a file: '{input_path}'")

if verbose:
click.echo(f"Processing {input_file}...")

try:
source = read_csv(input_file)
except Exception as e:
raise click.ClickException(f"Failed to read CSV file '{input_file}': {e}")

column_names = [col.name for col in source.columns]
if "Score" not in column_names:
raise click.ClickException(
f"File '{input_path.name}' is missing required column 'Score'. "
f"Available columns: {', '.join(column_names)}"
)

result = source.update(formulas=["DoubleScore = Score * 2"])
Comment thread
elijahpetty marked this conversation as resolved.

if verbose:
click.echo(f"Processed {result.size} rows")

return result


@click.command()
@click.argument("input_file", type=click.Path(exists=True))
@click.option("--verbose", "-v", is_flag=True, help="Enable verbose output")
def main(input_file: str, verbose: bool) -> None:
"""Process data with Deephaven."""
try:
from deephaven_server import Server

server = Server(port=10000, jvm_args=["-Xmx4g"])
server.start()
except Exception as e:
raise click.ClickException(
f"Failed to start Deephaven server on port 10000: {e}"
)

my_dh_query(input_file, verbose)
click.echo("Processing complete!")


if __name__ == "__main__":
main()
82 changes: 82 additions & 0 deletions my_dh_client/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,82 @@
# My Deephaven Client

An example of packaging a [`pydeephaven`](https://pypi.org/project/pydeephaven/) client program as a command line tool. Installing this package creates one terminal command, `my-dh-client`, which connects to a Deephaven server that is already running, uploads a CSV file to it, processes the data there, and binds the result under a name so it appears in the server's IDE.

Compare it with [`my_dh_cli`](../my_dh_cli/), which does similar work with an embedded server. The two differ in three ways:

- The dependency is `pydeephaven` rather than `deephaven-server`, so installing it does not pull in a JVM and the command starts quickly.
- The command does not start a server. It needs one to connect to, and many copies of the command can run against the same server at once.
- `pydeephaven` is imported at the top of `cli.py`. The embedded-server command has to delay its `deephaven` import until after the server starts; the client has no such constraint.

The command is defined by the `[project.scripts]` entry point in [`pyproject.toml`](pyproject.toml):

```toml
[project.scripts]
my-dh-client = "my_dh_client.cli:main"
```

## Installation

From the repository root:

```bash
pip install ./my_dh_client
```

Or in editable mode for development:

```bash
pip install -e ./my_dh_client
```

## Usage

> [!NOTE]
> A Deephaven server must already be running. By default the command connects to `localhost:10000` with anonymous authentication. See [Start a server for the client examples](../README.md#start-a-server-for-the-client-examples) in the repository README for ways to start one, and for connecting to a server that uses a pre-shared key.

Run the installed command on a CSV file:

```bash
my-dh-client data/sample.csv --verbose
```

It reads the file locally with pyarrow, uploads it to the server, adds a `DoubleScore` computed column there, and binds the result as a table named `sample` (the file's stem). Open the server's IDE at `http://localhost:10000` to see the table, or pick a different name with `--name`.

To connect to a different server, pass `--host` and `--port`. For a server that requires a token, pass `--auth-type` and put the token in the `DH_AUTH_TOKEN` environment variable so it stays out of your shell history:

```bash
DH_AUTH_TOKEN=my-secret-key my-dh-client data/sample.csv \
--host dh.example.com \
--auth-type io.deephaven.authentication.psk.PskAuthenticationHandler
```

During development, the package also runs without an entry point via [`__main__.py`](src/my_dh_client/__main__.py):

```bash
python -m my_dh_client data/sample.csv --verbose
```

## Command reference

### my-dh-client

Upload a CSV file to a running Deephaven server and process it there. The file must contain a `Score` column.

**Arguments:**

- `input_file` - Path to the CSV file to upload.

**Options:**

- `--host` - Deephaven server host. Default: `localhost`.
- `--port` - Deephaven server port. Default: `10000`.
- `--auth-type` - Authentication type. Default: `Anonymous`. For other types, set the token in the `DH_AUTH_TOKEN` environment variable.
- `--name` - Name to bind the result table under on the server. Default: the input file's stem.
- `--verbose, -v` - Enable verbose output.

## Requirements

- Python 3.9 or later
- pydeephaven 0.35.0 or later (installed automatically as a dependency)
- Click 8.0.0 or later (installed automatically as a dependency)
- A running Deephaven server to connect to. Java is not required on the client machine.
22 changes: 22 additions & 0 deletions my_dh_client/pyproject.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
[build-system]
requires = ["setuptools>=61.0", "wheel"]
build-backend = "setuptools.build_meta"

[project]
name = "my_dh_client"
version = "0.1.0"
description = "Command line client for a running Deephaven server"
readme = "README.md"
requires-python = ">=3.9"
dependencies = [
# pydeephaven is the Python client. It connects to a running server and does not start one.
"pydeephaven>=0.35.0",
# click implements the command line interface.
"click>=8.0.0",
]

[project.scripts]
my-dh-client = "my_dh_client.cli:main"

[tool.setuptools.packages.find]
where = ["src"]
3 changes: 3 additions & 0 deletions my_dh_client/src/my_dh_client/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
"""Command line client that uploads a CSV file to a running Deephaven server."""

__version__ = "0.1.0"
4 changes: 4 additions & 0 deletions my_dh_client/src/my_dh_client/__main__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
from my_dh_client.cli import main

if __name__ == "__main__":
main()
Loading