Repository navigation
Add examples #1
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: main
Are you sure you want to change the base?
Add examples #1
Changes from all commits
08e0360
2059ecb
1bfc5d3
6da0501
d5c749d
99946fb
339df38
f111a48
bc71cf1
dfab129
c8d8dff
d6fb791
7173ab4
1c589d6
211ca32
9f1800f
9064aef
b487cbd
acda4a6
333eac6
2730abd
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,44 @@ | ||
| # Python | ||
| __pycache__/ | ||
| *.py[cod] | ||
| *$py.class | ||
| *.so | ||
| .Python | ||
| build/ | ||
| develop-eggs/ | ||
| dist/ | ||
| downloads/ | ||
| eggs/ | ||
| .eggs/ | ||
| lib/ | ||
| lib64/ | ||
| parts/ | ||
| sdist/ | ||
| var/ | ||
| wheels/ | ||
| *.egg-info/ | ||
| .installed.cfg | ||
| *.egg | ||
| MANIFEST | ||
|
|
||
| # Virtual environments | ||
| venv/ | ||
| env/ | ||
| ENV/ | ||
| .venv | ||
|
|
||
| # IDEs | ||
| .vscode/ | ||
| .idea/ | ||
| *.swp | ||
| *.swo | ||
| *~ | ||
|
|
||
| # OS | ||
| .DS_Store | ||
| Thumbs.db | ||
|
|
||
| # Test results | ||
| results/ | ||
| output/ | ||
| *.log |
Large diffs are not rendered by default.
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,4 @@ | ||
| Name,Score,Value,Category | ||
| Kara,81,130,A | ||
| Liam,94,175,B | ||
| Mona,77,100,A |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,4 @@ | ||
| Name,Score,Value,Category | ||
| Nina,86,115,C | ||
| Omar,90,155,B | ||
| Pria,74,80,A |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,5 @@ | ||
| Name,Score,Value,Category | ||
| Quinn,93,165,C | ||
| Rosa,84,125,B | ||
| Sam,79,95,A | ||
| Tara,88,145,C |
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. In reviewing this, I've had some confusion. This file looks like a concatenation of the prior files. I'm not sure if this is intentional, confusing, or just doesn't matter. I haven't read enough of the code to have a view. Just noting it here. |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,11 @@ | ||
| Name,Score,Value,Category | ||
| Alice,85,120,A | ||
| Bob,92,150,B | ||
| Charlie,78,95,A | ||
| Diana,88,110,C | ||
| Eve,95,180,B | ||
| Frank,72,85,A | ||
| Grace,91,160,C | ||
| Henry,83,105,B | ||
| Iris,89,140,A | ||
| Jack,76,90,C |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,63 @@ | ||
| # My Deephaven CLI | ||
|
|
||
| An example of packaging a Deephaven script as a command line tool. Installing this package creates one terminal command, `my-dh-query`. No library code is exposed; the package is used only through that command, so no Python needs to be written to use it. | ||
|
|
||
| The command is defined by the `[project.scripts]` entry point in [`pyproject.toml`](pyproject.toml): | ||
|
|
||
| ```toml | ||
| [project.scripts] | ||
| my-dh-query = "my_dh_cli.cli:main" | ||
| ``` | ||
|
|
||
| ## Installation | ||
|
|
||
| From the repository root: | ||
|
|
||
| ```bash | ||
| pip install ./my_dh_cli | ||
| ``` | ||
|
|
||
| Or in editable mode for development: | ||
|
|
||
| ```bash | ||
| pip install -e ./my_dh_cli | ||
| ``` | ||
|
|
||
| ## Usage | ||
|
|
||
| Run the installed command on a CSV file. The command starts its own Deephaven server, so no separate setup is needed: | ||
|
|
||
| ```bash | ||
| my-dh-query data/sample.csv --verbose | ||
| ``` | ||
|
|
||
| It reads the file, adds a `DoubleScore` computed column, and reports the number of rows processed. | ||
|
|
||
| The command binds its server to port 10000. If that port is already in use (for example, by Deephaven running in Docker), change the `port` value in [`cli.py`](src/my_dh_cli/cli.py). | ||
|
|
||
| During development, the package also runs without an entry point via [`__main__.py`](src/my_dh_cli/__main__.py): | ||
|
|
||
| ```bash | ||
| python -m my_dh_cli data/sample.csv --verbose | ||
| ``` | ||
|
|
||
| ## Command reference | ||
|
|
||
| ### my-dh-query | ||
|
|
||
| Process a CSV file with Deephaven. The file must contain a `Score` column. | ||
|
|
||
| **Arguments:** | ||
|
|
||
| - `input_file` - Path to the CSV file to process. | ||
|
|
||
| **Options:** | ||
|
|
||
| - `--verbose, -v` - Enable verbose output. | ||
|
|
||
| ## Requirements | ||
|
|
||
| - Python 3.9 or later | ||
| - Java 17 or later | ||
| - deephaven-server 0.35.0 or later (installed automatically as a dependency) | ||
| - Click 8.0.0 or later (installed automatically as a dependency) |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,22 @@ | ||
| [build-system] | ||
| requires = ["setuptools>=61.0", "wheel"] | ||
| build-backend = "setuptools.build_meta" | ||
|
|
||
| [project] | ||
| name = "my_dh_cli" | ||
| version = "0.1.0" | ||
| description = "Command line tool for data processing" | ||
| readme = "README.md" | ||
| requires-python = ">=3.9" | ||
| dependencies = [ | ||
| # deephaven-server also provides the deephaven module (through its deephaven-core dependency). | ||
| "deephaven-server>=0.35.0", | ||
| # click implements the command line interface. | ||
| "click>=8.0.0", | ||
| ] | ||
|
|
||
| [project.scripts] | ||
| my-dh-query = "my_dh_cli.cli:main" | ||
|
|
||
| [tool.setuptools.packages.find] | ||
| where = ["src"] |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,3 @@ | ||
| """Command line tool that processes a CSV file with Deephaven.""" | ||
|
|
||
| __version__ = "0.1.0" |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,4 @@ | ||
| from my_dh_cli.cli import main | ||
|
|
||
| if __name__ == "__main__": | ||
| main() |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,61 @@ | ||
| import click | ||
|
|
||
|
|
||
| def my_dh_query(input_file: str, verbose: bool = False): | ||
| """Read a CSV file and perform a simple query operation on the data.""" | ||
| # Imported here, not at module level: deephaven requires a running server. | ||
| # The entry point starts the server first, then calls this function. | ||
| from deephaven import read_csv | ||
| from pathlib import Path | ||
|
|
||
| input_path = Path(input_file) | ||
|
|
||
| if not input_path.exists(): | ||
| raise click.ClickException(f"Input file does not exist: '{input_path}'") | ||
| if not input_path.is_file(): | ||
| raise click.ClickException(f"Input path is not a file: '{input_path}'") | ||
|
|
||
| if verbose: | ||
| click.echo(f"Processing {input_file}...") | ||
|
|
||
| try: | ||
| source = read_csv(input_file) | ||
| except Exception as e: | ||
| raise click.ClickException(f"Failed to read CSV file '{input_file}': {e}") | ||
|
|
||
| column_names = [col.name for col in source.columns] | ||
| if "Score" not in column_names: | ||
| raise click.ClickException( | ||
| f"File '{input_path.name}' is missing required column 'Score'. " | ||
| f"Available columns: {', '.join(column_names)}" | ||
| ) | ||
|
|
||
| result = source.update(formulas=["DoubleScore = Score * 2"]) | ||
|
elijahpetty marked this conversation as resolved.
|
||
|
|
||
| if verbose: | ||
| click.echo(f"Processed {result.size} rows") | ||
|
|
||
| return result | ||
|
|
||
|
|
||
| @click.command() | ||
| @click.argument("input_file", type=click.Path(exists=True)) | ||
| @click.option("--verbose", "-v", is_flag=True, help="Enable verbose output") | ||
| def main(input_file: str, verbose: bool) -> None: | ||
| """Process data with Deephaven.""" | ||
| try: | ||
| from deephaven_server import Server | ||
|
|
||
| server = Server(port=10000, jvm_args=["-Xmx4g"]) | ||
| server.start() | ||
| except Exception as e: | ||
| raise click.ClickException( | ||
| f"Failed to start Deephaven server on port 10000: {e}" | ||
| ) | ||
|
|
||
| my_dh_query(input_file, verbose) | ||
| click.echo("Processing complete!") | ||
|
|
||
|
|
||
| if __name__ == "__main__": | ||
| main() | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,82 @@ | ||
| # My Deephaven Client | ||
|
|
||
| An example of packaging a [`pydeephaven`](https://pypi.org/project/pydeephaven/) client program as a command line tool. Installing this package creates one terminal command, `my-dh-client`, which connects to a Deephaven server that is already running, uploads a CSV file to it, processes the data there, and binds the result under a name so it appears in the server's IDE. | ||
|
|
||
| Compare it with [`my_dh_cli`](../my_dh_cli/), which does similar work with an embedded server. The two differ in three ways: | ||
|
|
||
| - The dependency is `pydeephaven` rather than `deephaven-server`, so installing it does not pull in a JVM and the command starts quickly. | ||
| - The command does not start a server. It needs one to connect to, and many copies of the command can run against the same server at once. | ||
| - `pydeephaven` is imported at the top of `cli.py`. The embedded-server command has to delay its `deephaven` import until after the server starts; the client has no such constraint. | ||
|
|
||
| The command is defined by the `[project.scripts]` entry point in [`pyproject.toml`](pyproject.toml): | ||
|
|
||
| ```toml | ||
| [project.scripts] | ||
| my-dh-client = "my_dh_client.cli:main" | ||
| ``` | ||
|
|
||
| ## Installation | ||
|
|
||
| From the repository root: | ||
|
|
||
| ```bash | ||
| pip install ./my_dh_client | ||
| ``` | ||
|
|
||
| Or in editable mode for development: | ||
|
|
||
| ```bash | ||
| pip install -e ./my_dh_client | ||
| ``` | ||
|
|
||
| ## Usage | ||
|
|
||
| > [!NOTE] | ||
| > A Deephaven server must already be running. By default the command connects to `localhost:10000` with anonymous authentication. See [Start a server for the client examples](../README.md#start-a-server-for-the-client-examples) in the repository README for ways to start one, and for connecting to a server that uses a pre-shared key. | ||
|
|
||
| Run the installed command on a CSV file: | ||
|
|
||
| ```bash | ||
| my-dh-client data/sample.csv --verbose | ||
| ``` | ||
|
|
||
| It reads the file locally with pyarrow, uploads it to the server, adds a `DoubleScore` computed column there, and binds the result as a table named `sample` (the file's stem). Open the server's IDE at `http://localhost:10000` to see the table, or pick a different name with `--name`. | ||
|
|
||
| To connect to a different server, pass `--host` and `--port`. For a server that requires a token, pass `--auth-type` and put the token in the `DH_AUTH_TOKEN` environment variable so it stays out of your shell history: | ||
|
|
||
| ```bash | ||
| DH_AUTH_TOKEN=my-secret-key my-dh-client data/sample.csv \ | ||
| --host dh.example.com \ | ||
| --auth-type io.deephaven.authentication.psk.PskAuthenticationHandler | ||
| ``` | ||
|
|
||
| During development, the package also runs without an entry point via [`__main__.py`](src/my_dh_client/__main__.py): | ||
|
|
||
| ```bash | ||
| python -m my_dh_client data/sample.csv --verbose | ||
| ``` | ||
|
|
||
| ## Command reference | ||
|
|
||
| ### my-dh-client | ||
|
|
||
| Upload a CSV file to a running Deephaven server and process it there. The file must contain a `Score` column. | ||
|
|
||
| **Arguments:** | ||
|
|
||
| - `input_file` - Path to the CSV file to upload. | ||
|
|
||
| **Options:** | ||
|
|
||
| - `--host` - Deephaven server host. Default: `localhost`. | ||
| - `--port` - Deephaven server port. Default: `10000`. | ||
| - `--auth-type` - Authentication type. Default: `Anonymous`. For other types, set the token in the `DH_AUTH_TOKEN` environment variable. | ||
| - `--name` - Name to bind the result table under on the server. Default: the input file's stem. | ||
| - `--verbose, -v` - Enable verbose output. | ||
|
|
||
| ## Requirements | ||
|
|
||
| - Python 3.9 or later | ||
| - pydeephaven 0.35.0 or later (installed automatically as a dependency) | ||
| - Click 8.0.0 or later (installed automatically as a dependency) | ||
| - A running Deephaven server to connect to. Java is not required on the client machine. |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,22 @@ | ||
| [build-system] | ||
| requires = ["setuptools>=61.0", "wheel"] | ||
| build-backend = "setuptools.build_meta" | ||
|
|
||
| [project] | ||
| name = "my_dh_client" | ||
| version = "0.1.0" | ||
| description = "Command line client for a running Deephaven server" | ||
| readme = "README.md" | ||
| requires-python = ">=3.9" | ||
| dependencies = [ | ||
| # pydeephaven is the Python client. It connects to a running server and does not start one. | ||
| "pydeephaven>=0.35.0", | ||
| # click implements the command line interface. | ||
| "click>=8.0.0", | ||
| ] | ||
|
|
||
| [project.scripts] | ||
| my-dh-client = "my_dh_client.cli:main" | ||
|
|
||
| [tool.setuptools.packages.find] | ||
| where = ["src"] |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,3 @@ | ||
| """Command line client that uploads a CSV file to a running Deephaven server.""" | ||
|
|
||
| __version__ = "0.1.0" |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,4 @@ | ||
| from my_dh_client.cli import main | ||
|
|
||
| if __name__ == "__main__": | ||
| main() |
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
In my quick scan of the code and the readme, I'm having a hard time figuring out what the story is. It seems like examples are doing multiple things at once, instead of each illustrating a single concept. The readme descriptions of the example then doesn't enlighten the case much.
Below is my chat with Copilot about this:
===
What is this readme trying to show? The illustrative examples seem garbled and unclear and like each does multipel things -- just making them confusing. Are they confusing?
===
Yes — looking at this diff, the README is trying to do too much in each section, and it does make the examples harder to follow. A few concrete problems:
For example, the CLI-only section shows, in sequence:
installing the package
running it as an installed shell command (my-dh-query)
then also importing the underlying function directly in Python (from my_dh_cli.cli import my_dh_query)
These are two different usage patterns (CLI invocation vs. programmatic call) bolted together under one "Try it" header, with only a short transition sentence ("The underlying function is also importable...") separating them. A reader skimming for "how do I just run this from the terminal" has to wade through Python import code too.
It shows, back-to-back: two CLI commands (my-dh-toolkit-query, my-dh-toolkit-process), then a full server-startup snippet, then library imports, then calling the same underlying functions two different ways (from my_dh_toolkit.queries import ... and from my_dh_toolkit import my_dh_query, batch_process). That's four different ways to do roughly two things, all in one code block region. It's demonstrating flexibility, but at the cost of a clear "here's the one obvious way to use this" path.
The Server(port=10000, jvm_args=["-Xmx4g"]).start() snippet is repeated near-verbatim in almost every section. That's necessary context, but it dilutes the examples — the reader has to mentally filter out the same 3 lines every time to find what's actually different about that example.
There's now a "Packaging scenarios" section (when to use library-only/CLI-only/combined) and a separate "Execution patterns" section (entry points vs. module execution) and per-package "Usage" snippets at the top and "Try the X package" walkthroughs further down. That's four different places explaining overlapping concepts (when/how to use each package type), which is a lot for a reader to reconcile into one mental model.
Net assessment: The content itself is accurate and reasonably complete, but the presentation conflates "here's the simplest way to use this" with "here's every way you could possibly use this." A cleaner version would pick one primary usage pattern per package type for the main walkthrough (e.g., just the CLI command) and move the "you can also import the underlying function" variant into a clearly separate, secondary callout — rather than interleaving them as equally-weighted steps.
If you'd like, I can draft a tightened version of one of these sections (e.g., the combined package walkthrough) to show the difference concretely.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Copilot reviw for the repo was set at a low / cheap level. I increased the level and reran it.