Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion .github/workflows/pr.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -75,7 +75,10 @@ jobs:
run: |
echo "## Test Coverage" >> $GITHUB_STEP_SUMMARY

go install github.com/vladopajic/go-test-coverage/v2@latest
# Pinned: v2.20.0 requires Go 1.27 and this module builds with the
# Go of its go.mod (1.26, GOTOOLCHAIN=local). "@latest" broke the job
# the day that version was published. Raise it with the go directive.
go install github.com/vladopajic/go-test-coverage/v2@v2.19.0

# execute again to get the summary
echo "" >> $GITHUB_STEP_SUMMARY
Expand Down
5 changes: 4 additions & 1 deletion .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,10 @@ jobs:
run: |
echo "## Test Coverage" >> $GITHUB_STEP_SUMMARY

go install github.com/vladopajic/go-test-coverage/v2@latest
# Pinned: v2.20.0 requires Go 1.27 and this module builds with the
# Go of its go.mod (1.26, GOTOOLCHAIN=local). "@latest" broke the job
# the day that version was published. Raise it with the go directive.
go install github.com/vladopajic/go-test-coverage/v2@v2.19.0

# execute again to get the summary
echo "" >> $GITHUB_STEP_SUMMARY
Expand Down
6 changes: 4 additions & 2 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -103,15 +103,17 @@ A negated **predicate** is governed by its base operator, not by `OpNot`:
are allowed whenever `OpIn` / `OpLike` / `OpBetween` / `OpSimilarTo` /
`OpIsDistinctFrom` (respectively) are allowed.
- `OpNot` governs only the **standalone** logical `NOT`, e.g. `NOT (age > 30)`.
- `IS NOT DISTINCT FROM` is negated inside the `IS` predicate, not by a
leading `NOT`: `age NOT DISTINCT FROM 30` is refused.

```mermaid
flowchart TD
N["NOT appears in input"] --> K{"what follows?"}
K -->|"( expr ) or a comparison"| U["UnaryOperatorNode<br/>governed by OpNot"]
K -->|"IN / LIKE / BETWEEN /<br/>SIMILAR TO / DISTINCT FROM"| P["predicate node with IsNot=true<br/>governed by the base operator"]
K -->|"IN / LIKE / ILIKE /<br/>BETWEEN / SIMILAR TO"| P["predicate node with IsNot=true<br/>governed by the base operator"]

U -. "allow-list check" .-> gnot{"OpNot allowed?"}
P -. "allow-list check" .-> gbase{"OpIn / OpLike / OpBetween /<br/>OpSimilarTo / OpIsDistinctFrom allowed?"}
P -. "allow-list check" .-> gbase{"OpIn / OpLike / OpILike /<br/>OpBetween / OpSimilarTo allowed?"}
```

So allowing `OpEqual` but not `OpNot` accepts `a = 1` but rejects `NOT (a = 1)`.
Expand Down
54 changes: 54 additions & 0 deletions docs/filtering.md
Original file line number Diff line number Diff line change
Expand Up @@ -76,6 +76,50 @@ binds tighter than `OR`.
**Literals**: single-quoted strings (`''` escapes a quote), integers, floats,
and booleans (`TRUE`/`FALSE`/`YES`/`NO`).

### Numbers

A number is a decimal literal, with an optional minus sign **glued to it**:
`42`, `-1`, `3.14`, `-.5`, `1.`, `1e5`, `-1e-5`. The sign is part of the
literal wherever a literal is taken (`id IN (-1, 2)`, `age BETWEEN -5 AND 5`,
`n IS DISTINCT FROM -1`); the grammar has no arithmetic, so `- 1` (with a
space), `1-1`, `+1` and `-name` are refused.

The parser refuses what PostgreSQL refuses, so that an accepted expression
can be spliced into a `WHERE` clause:

| Input | Why it is refused |
| --- | --- |
| `id = 1AND name = 'x'` | a word glued to a number: PostgreSQL's "trailing junk after numeric literal". (`name = 'x'AND id = 1` is fine: a string ends at its quote.) |
| `price = 0x1p-2` | a hexadecimal float is a Go literal, not a SQL one |
| `id !=-1`, `name ~~-1`, `name ~~*-1`, `name !~-1` | PostgreSQL reads the longest operator, and one written with `!` or `~` may end in `-`: this is the unknown operator `!=-`. Write `id != -1`. (`id =-1`, `id <>-1` and `id >=-1` are fine.) |
| `id = --1`, `id = 1--` | `--` opens a SQL comment |
| `id = 1e` | an exponent with no digits |

It is also stricter than PostgreSQL on purpose: `0x10`, `0b101`, `0o17`,
`1_000` and `1_000.5` are numbers in PostgreSQL 16 and later and are refused
here. A number is written with decimal digits, a point and an exponent, and
nothing else.

### Keywords, and what bounds an expression

- **A keyword is ASCII, in any case.** `like`, `Like` and `LIKE` are the
keyword; `lıke` (a dotless `ı`) and `Iſ` (a long `ſ`) are not, although
Unicode upper-cases those two letters to `I` and `S`. PostgreSQL folds
case in ASCII only and reads them as names.
- **An expression does not start with a byte order mark** (U+FEFF).
- **Nesting is at most 100 levels deep**: parentheses inside parentheses,
`NOT` applied to `NOT`. The parser is recursive and the input decides how
deep it recurses; without a bound, a quarter of a megabyte of `(` ends the
process with a stack overflow, which no `recover` catches. A long flat
expression (`a = 1 AND b = 2 AND …`) is not nested and is not limited by
this. **Bound the length of what you hand to `Parse`** as well: the parser
reads its whole input before it answers.

**The parser does not know a column's type.** `name LIKE -1`, `id = 'x'` and a
bare `-1` are accepted, and PostgreSQL refuses each for its type, not for its
syntax. What this page promises is that an accepted expression is not a
syntax error; a caller still has to answer for a value of the wrong type.

## Nested (dot-notation) field names

Field names may contain dots, so you can allow-list and filter on nested paths:
Expand Down Expand Up @@ -109,6 +153,14 @@ flowchart LR
N -. "shorthands" .- SH["field ISNULL / NOTNULL<br/>→ IsNullNode"]
```

`IS` is the only way in. `id DISTINCT FROM 1` and `id NOT DISTINCT FROM 1`,
without `IS`, are not SQL and are refused; so is `active IS YES`. `YES` and
`NO` are this parser's spellings of a boolean **literal** and not truth values
of the `IS` test. They are not PostgreSQL's either: `active = YES` parses to a
`LiteralNode` whose `Value` is `true` and which renders as `true`, but the
text `YES` spliced into SQL is read as a column name. Write `TRUE`/`FALSE`
where the input itself is spliced.

## Worked examples

```go
Expand All @@ -124,12 +176,14 @@ flowchart LR
"age > 30"
"age <> 30" // not equal
"age != 30" // not equal (alias)
"balance < -10.5" // a sign glued to the number

// Pattern matching
"first_name LIKE 'J%'"
"first_name NOT LIKE 'J%'"
"first_name ILIKE 'j%'" // case-insensitive
"first_name NOT ILIKE 'j%'"
"first_name !~~ 'J%'" // NOT LIKE, as an operator ("NOT ~~" is not SQL)
"name SIMILAR TO 'J%n'" // SQL-standard regex
"name NOT SIMILAR TO 'J%n'"

Expand Down
55 changes: 55 additions & 0 deletions docs/migration.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,60 @@
# Migration Guide

## v1.0.2 → v1.0.3

`v1.0.3` makes the filter parser refuse what PostgreSQL refuses as a syntax
error. A caller that passed one of these inputs through now gets the parser's
error instead of the database's.

| Input | v1.0.2 | v1.0.3 | PostgreSQL 18 |
| --- | --- | --- | --- |
| `id DISTINCT FROM 1`, `id NOT DISTINCT FROM 1` | ✅ accepted | ❌ error | syntax error: write `IS [NOT] DISTINCT FROM` |
| `price = 0x1p-2` (a hexadecimal float) | ✅ accepted | ❌ error | trailing junk after numeric literal |
| `id=1AND name='x'` (a word glued to a number) | ✅ accepted | ❌ error | trailing junk after numeric literal |
| `active IS YES`, `active IS NOT NO` | ✅ accepted | ❌ error | syntax error: `IS` takes `TRUE`, `FALSE`, `UNKNOWN`, `NULL` |
| `name NOT ~~ 'x'`, `name NOT ~~* 'x'` | ✅ accepted | ❌ error | syntax error: write `NOT LIKE` or `!~~` |

Three more refusals, of input no client writes by accident:

| Input | v1.0.2 | v1.0.3 |
| --- | --- | --- |
| a keyword spelled with a non-ASCII letter (`name lıke 'x'`, `id Iſ NULL`, `name aſc` in a sort) | ✅ accepted | ❌ error: PostgreSQL folds case in ASCII only, and reads these as names |
| an expression that starts with a byte order mark (U+FEFF) | ✅ accepted, the mark dropped | ❌ error |
| an expression nested more than 100 levels deep (`((((…`, `NOT NOT NOT …`) | ✅ accepted, or **the process ended**: about 250 000 opening parentheses overflow the stack, a fatal error no `recover` catches | ❌ error |

The last one is a reason to upgrade whatever else you do, if the expression
comes from a request: bound its length before calling `Parse`, too.

**One input that ran is now refused:** a float written with `_` separators
(`price = 1_000.5`, `1e1_0`). v1.0.2 refused the integer form (`1_000`) and
accepted the float by accident of how each was converted; PostgreSQL 16 and
later run both. A number is now decimal digits, a point and an exponent, for
integers and floats alike. Write `1000.5`.

**One input that was refused is now accepted:** a negative number.

| Input | v1.0.2 | v1.0.3 |
| --- | --- | --- |
| `id = -1`, `id IN (-1, 2)`, `age BETWEEN -5 AND 5` | ❌ error | ✅ accepted |

The sign is part of the literal: `LiteralNode.Text` is `-1` and `Value` is
`int64(-1)`. It must be glued to the number, and not to an operator written
with `!` or `~`: `id !=-1` is refused, because PostgreSQL reads `!=-` as one
operator. The parser still does not know a column's type, so a negative
number is accepted wherever a literal is (`name LIKE -1`, a bare `-1`), as a
positive one already was; PostgreSQL refuses those for their type. See
[Filtering › Numbers](filtering.md#numbers).

**Error messages and tokens.** For an input that was and is refused, the text
can differ: `0x10`, `0b101`, `0o17` and `1_000` are `illegal token` (it was
`invalid integer`), and a word glued to a number is one illegal token named
whole (`error on field '1name': illegal token`). `Lexer` users: `-` followed
by a digit or a point is no longer its own illegal token but the start of an
`INT` or `FLOAT` whose `Value` carries the sign.

**Migration:** replace `x [NOT] DISTINCT FROM y` with `x IS [NOT] DISTINCT FROM
y`, and write numbers without `_`.

## v0.0.x → v0.1.0

`v0.1.0` fixes correctness bugs in the filter parser, adds the PostgreSQL 18
Expand Down
110 changes: 105 additions & 5 deletions filter_lexer.go
Original file line number Diff line number Diff line change
Expand Up @@ -36,8 +36,34 @@ func NewLexer(input string) *Lexer {
return l
}

// asciiUpper returns s in upper case when s is written in ASCII, and s as it
// is otherwise.
//
// It is how a keyword is recognised. SQL folds case in ASCII only, and
// strings.ToUpper does not: it maps the dotless 'ı' (U+0131) to 'I' and the
// long 'ſ' (U+017F) to 'S', so "lıke" and "Iſ" read as LIKE and IS here and
// are a syntax error in PostgreSQL. A word with a non-ASCII letter is never a
// keyword.
func asciiUpper(s string) string {
for i := 0; i < len(s); i++ {
if s[i] >= 0x80 {
return s
}
}

return strings.ToUpper(s)
}

// byteOrderMark is U+FEFF. text/scanner drops one at the start of its input
// without a token; PostgreSQL reads it as part of the first word.
const byteOrderMark = "\uFEFF"

// Parse reads all tokens from the scanner and buffers them.
func (l *Lexer) Parse() {
if strings.HasPrefix(l.input, byteOrderMark) {
l.tokens = append(l.tokens, Token{Type: TokenIllegal, Value: byteOrderMark})
}

for {
scanTok := l.s.Scan()
pos := l.s.Position
Expand All @@ -53,7 +79,7 @@ func (l *Lexer) Parse() {
if lit == "is" {
tok = TokenIdentifier
} else {
upperLit := strings.ToUpper(lit)
upperLit := asciiUpper(lit)
switch upperLit {
case "AND":
tok = TokenOperatorAnd
Expand Down Expand Up @@ -86,10 +112,18 @@ func (l *Lexer) Parse() {
tok = TokenIdentifier
}
}
case scanner.Int:
tok = TokenInt
case scanner.Float:
tok = TokenFloat
case scanner.Int, scanner.Float:
tok, lit = l.numberToken(scanTok, lit)
case '-':
// A minus sign is the sign of a numeric literal when it is glued
// to one ("-1", "-.5"). Anything else stays illegal: the grammar
// has no arithmetic, and "--" opens a SQL comment.
tok = TokenIllegal

if next := l.s.Peek(); (next == '.' || (next >= '0' && next <= '9')) && !l.gluedToOperator(pos) {
numTok := l.s.Scan()
tok, lit = l.numberToken(numTok, "-"+l.s.TokenText())
}
case scanner.String: // Built-in scanner string (double quotes) - treat as illegal for this SQL-like syntax
tok = TokenIllegal
// For the test cases, we need to ensure the token value matches the expected format
Expand Down Expand Up @@ -235,6 +269,72 @@ func (l *Lexer) Parse() {
}
}

// numberToken classifies a number read by text/scanner, whose text is lit.
//
// text/scanner reads Go numbers, and those are not SQL's. It takes a
// hexadecimal float ("0x1p-2"), prefixed integers and '_' separators, and it
// ends a number at the first character that cannot continue it, so "1AND" is
// the number 1 followed by the keyword AND. PostgreSQL has no hexadecimal
// float and refuses a number with a word glued to it ("trailing junk after
// numeric literal"), so both are illegal tokens here: an expression this
// package accepts must be one PostgreSQL can run.
func (l *Lexer) numberToken(scanTok rune, lit string) (TokenType, string) {
// A word glued to the number: consume it, so that the error names the
// whole of what was written.
if next := l.s.Peek(); next == '_' || unicode.IsLetter(next) {
l.s.Scan()

return TokenIllegal, lit + l.s.TokenText()
}

if !isDecimalNumber(lit) {
return TokenIllegal, lit
}

switch scanTok {
case scanner.Int:
return TokenInt, lit
case scanner.Float:
return TokenFloat, lit
}

// A sign followed by something that is not a number after all ("-.").
return TokenIllegal, lit
}

// isDecimalNumber reports whether lit is written with decimal digits, a
// decimal point, an exponent and signs only. Whether it is well formed ("1e"
// is not) is left to strconv, in the parser.
func isDecimalNumber(lit string) bool {
for _, r := range lit {
if (r < '0' || r > '9') && !strings.ContainsRune(".eE+-", r) {
return false
}
}

return true
}

// gluedToOperator reports whether the character at pos directly follows an
// operator written with '!' or '~'.
//
// PostgreSQL reads the longest operator it can, and an operator that contains
// '!' or '~' may end in '-': "a !=-1" is the unknown operator "!=-" applied to
// 1, not "a != -1". An operator made of '=', '<' and '>' alone cannot end in
// '-', so "a =-1" and "a <>-1" compare with minus one.
func (l *Lexer) gluedToOperator(pos scanner.Position) bool {
if len(l.tokens) == 0 {
return false
}

prev := l.tokens[len(l.tokens)-1]
if prev.Pos.Offset+len(prev.Value) != pos.Offset {
return false
}

return strings.ContainsAny(prev.Value, "!~") && strings.Trim(prev.Value, "!~*=<>") == ""
}

// Peek returns the next token without consuming it.
func (l *Lexer) Peek() Token {
if l.pos+1 >= len(l.tokens) {
Expand Down
8 changes: 4 additions & 4 deletions filter_nodes.go
Original file line number Diff line number Diff line change
Expand Up @@ -19,8 +19,8 @@ const (
NodeTypeGroup NodeType = "GROUP" // (expression) -> (name = "John" AND age > 30)
NodeTypeIsNull NodeType = "IS_NULL" // (field) -> name IS NULL
NodeTypeIsNotNull NodeType = "IS_NOT_NULL" // (field) -> name IS NOT NULL
NodeTypeDistinct NodeType = "DISTINCT" // (field) -> name DISTINCT
NodeTypeNotDistinct NodeType = "NOT_DISTINCT" // (field) -> name NOT DISTINCT
NodeTypeDistinct NodeType = "DISTINCT" // (field) -> name IS DISTINCT FROM
NodeTypeNotDistinct NodeType = "NOT_DISTINCT" // (field) -> name IS NOT DISTINCT FROM
NodeTypeBetween NodeType = "BETWEEN" // (field, lower, upper) -> age BETWEEN 30 AND 40
NodeTypeNotBetween NodeType = "NOT_BETWEEN" // (field, lower, upper) -> age NOT BETWEEN 30 AND 40
NodeTypeIn NodeType = "IN" // (field, values) -> name IN ("John", "Doe")
Expand Down Expand Up @@ -218,12 +218,12 @@ func (n *InNode) String() string {
}
func (n *InNode) Pos() scanner.Position { return n.pos }

// DistinctNode represents a DISTINCT FROM expression (e.g., name DISTINCT FROM 'John')
// DistinctNode represents an IS [NOT] DISTINCT FROM expression (e.g., name IS DISTINCT FROM 'John')
type DistinctNode struct {
baseNode
Field Node
Value Node // the value being compared against; nil when absent
IsNot bool // true for NOT DISTINCT FROM
IsNot bool // true for IS NOT DISTINCT FROM
}

func (n *DistinctNode) Type() NodeType {
Expand Down
Loading
Loading