golars

golars scripting language (.glr)

.glr files are a tiny, line-oriented language for pipeline-style DataFrame work. Every REPL command is a script statement.

.glr files are a tiny, line-oriented language for pipeline-style DataFrame work. Designed to feel like "your REPL session in a file", nothing more. Every REPL command you know is also a script statement.

# trades-daily.glr
load data/trades.csv          as trades
load data/symbols.csv         as symbols

use trades
filter volume > 100
groupby symbol amount:sum:total
join symbols on symbol
sort total desc
limit 10
show

Run it:

golars run trades-daily.glr       # one-shot

From inside the REPL:

golars » .source trades-daily.glr

Grammar

program      = { statement NL } ;
statement    = empty
             | comment
             | command ;
comment      = "#" { any-char-until-NL } ;
command      = [ "." ] identifier { arg } ;
arg          = identifier | number | string | operator | list | "as" | "on" ;
list         = "[" [ arg { "," arg } ] "]" ;
string       = '"' { any-char } '"' ;
identifier   = ( letter | "_" ) { letter | digit | "_" | "." | "-" | "/" | ":" } ;
number       = [ "-" ] digit { digit } [ "." digit { digit } ] ;
operator     = "==" | "!=" | "<=" | ">=" | "<" | ">" | "and" | "or" ;

That is the statement level. filter, with, select and groupby hand the rest of the line to the expression language.

Statements are line-terminated. A trailing \ (after any trailing whitespace) continues onto the next physical line, useful for long filter predicates:

filter salary > 100000 \
  and dept == "eng" \
  and tenure_years >= 2

The leading . on every command is optional: load foo.csv and .load foo.csv do the same thing. A # inside a "..." string is treated as a literal - only unquoted # starts a comment.

If you mis-spell a command golars tries to be helpful:

$ golars run typo.glr
typo.glr:3: unknown command "filtet" (did you mean "filter"?)

Statement reference

Commands are grouped the way .help prints them. The canonical list, with signatures and hover text, lives in script/spec.go; the REPL, the LSP, the Jupyter kernel and the editor grammars all read from it. Aliases are listed next to the command they stand for.

Three kinds of statement behave differently:

  • Pipeline statements add a lazy step to the focused frame. Nothing runs until something collects (show, head, save, ...).
  • Reshape statements have no lazy form. They collect the focus, transform it, and make the result the new focus.
  • Inspect and aggregate statements print something and leave the focus unchanged.

io

StatementWhat it does
load PATHFocus a new frame. Formats come from the extension: .csv, .tsv, .parquet/.pq, .arrow/.ipc, .json, .ndjson/.jsonl. Empty CSV fields read as null.
load PATH as NAMEStage a frame under NAME without touching the focus.
save PATH (alias write)Collect the focused pipeline and write it; the format comes from the extension.
scan_csv PATH [as NAME]Register a lazy CSV scan (push-down friendly). A .tsv path keeps the tab delimiter.
scan_parquet PATH [as NAME]Lazy Parquet scan.
scan_ipc PATH [as NAME] (alias scan_arrow)Lazy Arrow IPC scan.
scan_json PATH [as NAME]Lazy JSON scan.
scan_ndjson PATH [as NAME] (alias scan_jsonl)Lazy NDJSON scan.
scan_auto PATH [as NAME]Lazy scan with the reader inferred from the extension.

Without as NAME a scan replaces the focus and the file is opened when the pipeline is collected. With as NAME the scan is collected and staged, like load PATH as NAME.

frames

StatementWhat it does
use NAMESwitch focus to a clone of NAME. NAME stays staged so repeated use branches off the same base; the prior focus is discarded.
stash NAMECollect the focus and stage a copy under NAME; the focus continues from the snapshot.
framesList loaded frames. The focused one is marked *.
drop_frame NAMERelease NAME from the registry.

pipeline (lazy)

StatementWhat it does
select COL [, COL...]Project columns. Commas and spaces both separate plain names.
select ITEM, ITEM...Project expressions: each ITEM is a column, an expression, or name = expr, as in select name, year = dt.year(hired).
drop COL [, COL...]Drop columns.
filter PREDKeep rows where the boolean expression PRED holds. See predicates.
sort COL [asc|desc] [COL [asc|desc]]...Sort by one or more columns. Nulls sort first, as in polars.
limit NKeep the first N rows.
groupby KEYS AGG [AGG...]Group and aggregate. KEYS is comma-separated. AGG is col:op[:alias] (op: sum, mean/avg, min, max, count, null_count, first, last, median, std, var, n_unique) or name=expr, as in n=len(). Write an expression AGG without spaces or wrap it in parentheses.
group_by_dynamic TIME every DUR [period DUR] [offset DUR] [by KEYS] [closed C] [label L] AGG... (alias groupby_dynamic)Group rows into time windows of the sorted column TIME and aggregate. DUR is 30m, 1h, 1d, 1w, 1mo, 1y (or 3i for integer columns). AGG takes a bare column, as in amount:sum:total or n=amount.count().
join PATH|NAME on KEYS [TYPE] [suffix S]Join the focus with a staged frame or a file on KEYS, one column or a comma separated list present in both frames. TYPE is inner (default), left, right, full (or outer), semi, anti or cross, with polars semantics: full keeps separate key columns, semi and anti keep only the focus columns. Right columns whose name collides with a focus column get suffix S (default _right). The join runs eagerly and becomes the new focus.
join_asof PATH|NAME on KEY [by COLS] [backward|forward|nearest] [tolerance T]Join each row to the nearest row of the other frame by KEY: the last one at or before it (backward, the default), the first at or after it (forward), or the closest (nearest). by requires exact matches first; tolerance bounds the distance (a number, or a duration such as 2m). Both frames must be sorted by KEY.
with NAME = EXPRAdd a derived column. See expression language.
collectRun the pipeline and make the result the focus.
resetDiscard the pending lazy steps; keep the source.
reverseReverse the row order.
uniqueDrop duplicate rows across every column.
cast COL TYPECast COL to TYPE: i8 to i64, u8 to u64, f32, f64, bool, str, binary, date, datetime (or datetime[ms], [us], [ns]), duration, time, categorical.
fill_null VALUE (alias fillnull)Replace nulls across compatible columns with VALUE.
drop_null [COL...] (alias dropnull)Drop rows with nulls in any (or the listed) columns.
fill_nan VALUEReplace NaN with VALUE in every float column.
forward_fill [LIMIT] (alias ff)Forward-fill nulls per column (LIMIT 0 is unlimited). Leading nulls stay null.
backward_fill [LIMIT] (alias bf)Backward-fill nulls per column. Trailing nulls stay null.
rename OLD as NEWRename one column.
with_row_index NAME [OFFSET]Prepend an int64 row index.
sum_horizontal OUT [COL...]Append a row-wise sum column (nulls ignored).
mean_horizontal OUT [COL...]Row-wise mean.
min_horizontal OUT [COL...]Row-wise min.
max_horizontal OUT [COL...]Row-wise max.
all_horizontal OUT [COL...]Row-wise boolean AND.
any_horizontal OUT [COL...]Row-wise boolean OR.

reshape (replaces the focus)

StatementWhat it does
sample N [SEED]Keep N rows sampled uniformly without replacement (seed defaults to 42).
shuffle [SEED]Randomly reorder every row.
top_k K COLKeep the K rows with the largest values in COL.
bottom_k K COLKeep the K rows with the smallest values in COL.
transpose [HEADER_COL] [PREFIX]Transpose the focus (numeric/bool columns).
unpivot IDS [VALS] (alias melt)Wide-to-long reshape. IDS/VALS are comma-separated lists.
pivot INDEX ON VALUES [AGG]Long-to-wide pivot. AGG is first (default), sum, mean, min, max or count.
unnest COLProject the fields of a struct column as top-level columns.
explode COL[,COL...]Fan out each element of a list column into its own row. Several columns explode together.
to_dummies [COL...] [drop_first]One-hot encode the listed columns (all when none are given) into u8 columns named COL_VALUE.
upsample COL EVERYFill a sorted timestamp column at ns/us/ms/s/m/h/d/w intervals.

inspect

StatementWhat it does
show [N]Collect and print the first N rows (default 10). Same as head.
head [N]Collect and print the first N rows (default 10).
tail [N]Collect and print the last N rows (default 10).
schemaPrint column names and dtypes.
describe [COL...]count, null_count, mean, std, min, quartiles and max per column (or for the listed columns).
glimpse [N]Compact peek at the first N rows (default 5).
sizeEstimated Arrow byte size of the pipeline result.
null_count (alias null-count)Per-column null count as a one-row frame.
partition_by KEYSPrint the row count of each key combination.
ishow (alias browse)Open the focus in the interactive browse TUI. Quit with q to return.

aggregate

StatementWhat it does
sum COL / mean COL (alias avg) / min COL / max COL / median COL / std COLPrint one scalar for COL.
skew COL / kurtosis COLSkewness / excess kurtosis.
approx_n_unique COL (alias approx_nunique)HyperLogLog estimate of the distinct count.
corr COL1 COL2 / cov COL1 COL2Pearson correlation / sample covariance.
sum_all / mean_all / min_all / max_all / std_all / var_all / median_allOne row of per-column aggregates over every numeric column.
count_all / null_count_allOne row of per-column (null) counts.

plan

StatementWhat it does
explainPrint the logical plan, optimiser trace and optimised plan.
explain_tree (alias tree)The same report rendered as a box-drawn tree.
graph (alias show_graph)Plan tree with colour per node kind.
mermaidEmit the plan as a Mermaid flowchart. Pipe into mmdc for PNG/SVG.

session

StatementWhat it does
source PATHRun another .glr file inline.
help (aliases h, ?)Print the command reference.
exit (aliases quit, q)Quit the REPL. In golars run, stop the script without an error.
timingToggle per-statement timing.
infoRuntime info: Go version, heap, uptime, focus shape.
clearClear the screen.
pwd / ls [PATH] / cd [PATH]Working-directory helpers.

Expression language

with, filter, select and the name=expr form of groupby share one expression language. It reaches the golars expression API directly: every method on expr.Expr and on its namespaces is callable by its snake_case name, so new library functions show up in glr without parser changes.

Grammar

expr     := orExpr
orExpr   := andExpr ("or" andExpr)*
andExpr  := notExpr ("and" notExpr)*
notExpr  := "not" notExpr | cmpExpr
cmpExpr  := addExpr ( cmpOp addExpr | "is_null" | "is_not_null"
                    | strOp addExpr | ["not"] "in" list )?
cmpOp    := "==" | "!=" | "<" | "<=" | ">" | ">="
strOp    := "contains" | "starts_with" | "ends_with" | "like" | "not_like"
addExpr  := mulExpr (("+" | "-") mulExpr)*
mulExpr  := unary (("*" | "/" | "//" | "%") unary)*
unary    := "-" unary | power
power    := postfix ("**" unary)?
postfix  := primary ("." [ns "."] ident [args])*
primary  := literal | list | "(" expr ")" | when
          | ns "." ident args | ident args | ident
when     := "when" orExpr "then" orExpr ("when" orExpr "then" orExpr)*
            ("otherwise" orExpr)?
args     := "(" [arg ("," arg)*] ")"
arg      := ident "=" expr | expr
list     := "[" [expr ("," expr)*] "]"
ns       := "str" | "dt" | "list" | "arr" | "struct" | "name" | "bin" | "cat"

Bare identifiers are columns; col("my col") reaches names with spaces or odd characters. Keywords (and, or, not, in, when, ...) are case-insensitive.

Operators

glrMeaningGo
a + b, a - b, a * barithmeticAdd, Sub, Mul
a / btrue division; integers give f64, as in polarsDiv
a // bfloor divisionFloorDiv
a % bmodulo (sign follows the divisor, as in polars)Mod
a ** bpowerPowExpr
==, !=, <, <=, >, >=comparisonEq, Ne, ...
and, or, notboolean logicAnd, Or, Not
x in [1, 2], x not in [...]membershipIsIn
x is_null, x is_not_nullnull testsIsNull, IsNotNull
s contains "a", starts_with, ends_with, like, not_likestring testsStr().Contains(...), ...
when c then a [when c2 then b] otherwise dconditional; without otherwise the rest is nullWhen(c).Then(a).Otherwise(...)

Calls

Every function has two equivalent spellings: a method on the value, or a call whose first argument is the value.

with r1 = x.round(2)             # same as round(x, 2)
with y1 = ts.dt.year()           # same as dt.year(ts)
with n1 = str.len_chars(name)    # same as name.str.len_chars()
  • Namespaces: str, dt, list, arr, struct, name, bin, cat, matching Expr.Str(), Expr.Dt(), and so on. A column whose name is also a namespace (such as name) still works as a column: name.str.upper() is the column name.
  • Lists are written [a, b]: cut(x, [0, 10]), x in ["a", "b"], replace(x, [1, 2], [10, 20]).
  • Keyword arguments set options, as in polars: cut(x, [0, 10], labels=["lo", "mid", "hi"], left_closed=true), replace_strict(x, [1, 2], ["a", "b"], default="other"), str.to_date(s, "%d/%m/%Y", strict=false), rolling_mean_by(x, ts, "7d", min_periods=1), concat_str(a, b, separator="-").
  • Defaults follow polars where they exist: shift(x) shifts by 1, round(x) rounds to 0 places, rank(x) is average, date_range(a, b) steps one day.
  • Parentheses are optional for calls without arguments: price.sum.
  • sum("x") (a quoted name) aggregates the column x, for compatibility with older scripts.

Common functions

AreaFunctions
temporaldt.year, dt.month, dt.day, dt.weekday, dt.hour, dt.minute, dt.date, dt.truncate(ts, "1h"), dt.round(ts, "1d"), dt.offset_by(ts, "1mo"), dt.strftime(ts, "%Y-%m"), dt.epoch(ts, "ms"), dt.total_days(d), dt.replace_time_zone(ts, "UTC"), dt.convert_time_zone(ts, "Asia/Tokyo"), dt.month_start, dt.month_end
parsingstr.to_date(s, "%Y-%m-%d"), str.to_datetime(s, "%Y-%m-%d %H:%M", time_zone="UTC"), str.to_time(s, "%H:%M"), str.strptime(s, "date", "%Y-%m-%d"), str.to_integer(s), cast(x, "f64")
stringsstr.to_uppercase, str.to_lowercase, str.len_chars, str.contains, str.starts_with, str.replace, str.replace_all, str.slice(s, 0, 3), str.split(s, ","), str.strip, str.pad_start(s, 5, "0"), str.zfill, str.extract(s, "(\d+)"), str.count_matches, concat_str, format("{}-{}", a, b)
listslist.len, list.get(l, 0), list.first, list.last, list.contains(l, "a"), list.join(l, ","), list.sum, list.mean, list.sort, list.unique, list.eval(l, element() * 2)
structsstruct(a, b), struct.field(s, "a"), struct.rename_fields(s, "x", "y")
namesname.suffix(x, "_raw"), name.prefix(x, "p_"), name.to_uppercase(x)
binningcut(x, [0, 10, 100]), qcut(x, 4), qcut(x, [0.25, 0.75])
mappingreplace(x, [1], [10]), replace_strict(x, [1, 2], ["a", "b"], default="z"), coalesce(a, b), fill_null(x, v), fill_null(x, other), interpolate(x)
numericabs, round(x, 2), floor, ceil, sqrt, exp, log(x), log(x, 10), clip(x, 0, 1), sign, is_between(x, 1, 5), is_nan
windowsshift, diff, pct_change, cum_sum, cum_max, rank(x, "dense"), rolling_mean(x, 7, 1), rolling_mean_by(x, ts, "7d"), ewm_mean(x, 0.3), sum(x).over("k"), rle_id(x)
aggregatessum, mean, min, max, count, len(), median, std, var, quantile(x, 0.9), n_unique, first, last, mode
orderingsort(x, descending=true), sort_by(x, y), top_k(x, 3), top_k_by(x, y, 3), arg_max, reverse
constructorslit(0), col("name"), date(2024, 1, 31), datetime(2024, 1, 31, hour=9), duration(days=1), date_range(a, b, "1d"), int_range(0, 10), sum_horizontal(a, b), max_horizontal(a, b)

Functions that take a Go callback (map_batches, map_elements, fold) are not reachable from glr; use Go for those.

Examples

with signed_up = str.to_date(signup, "%d/%m/%Y")
with tenure_d  = dt.total_days(dt.date(ts) - signed_up)
with slot      = dt.strftime(dt.truncate(ts, "30m"), "%H:%M")
with ntags     = list.len(str.split(tags, ";"))
with size      = cut(amount, [20, 100], labels=["small", "medium", "large"])
with bucket    = amount // 50 * 50
with parity    = when amount % 2 == 0 then "even" otherwise "odd"
with trend     = amount.rolling_mean(7, 1)
with score     = coalesce(primary, backup).str.trim()

Predicates

filter takes any boolean expression. The classic clause forms read naturally, and binds tighter than or, and parentheses group:

filter age >= 21 and salary > 50000
filter symbol == "AAPL"
filter is_active and created_at > 1704067200000000
filter note is_null
filter name like "a%" or name contains "z"
filter dept in ["eng", "ops"] and not (age < 21)
filter dt.year(ts) == 2024 and str.len_chars(name) > 3

A bare boolean column is a predicate on its own: filter is_active. String values are double- or single-quoted.


Multi-source workflows

Scripts regularly need N frames. The as NAME / use NAME / .frames trio is the whole story: there's no hidden namespace:

# Stage every input up front. None of these promote themselves to
# focus, so we can read them in any order.
load data/trades.csv    as trades
load data/symbols.csv   as symbols
load data/users.csv     as users

# Work on one, stash it, work on the next.
use trades
filter volume > 100
groupby user_id amount:sum:total_bought
stash trade_totals

# `use` is non-consuming: trade_totals stays staged, and so does the
# original trades frame: we could `use trades` again to branch off a
# different filter.
use users
filter region == "US"
join trade_totals on user_id
join symbols on symbol
sort total_bought desc
show

stash is the "save into a variable" move: it materializes whatever lazy pipeline is on the focus and parks a copy under NAME so later use NAME gives you that snapshot. The focus itself keeps going from the snapshot, so the idiomatic branching pattern is:

load data/trades.csv
filter volume > 100
stash base

filter side == "buy"
stash buys

use base
filter side == "sell"
stash sells

use buys
join sells on symbol

When .join sees a name that exists in the frame registry, it consumes that frame (keeping it in the registry for reuse) instead of treating the argument as a path. Paths win only when no frame matches.

Anonymous load PATH

The short form load PATH (no as) is equivalent to use NAME where NAME is empty. It's the "single-frame script" ergonomic:

load data/trades.csv
filter volume > 100
show

No registry, no juggling: just pipe.


Transpile to Go

golars transpile SCRIPT.glr [-o OUT.go] [--package NAME] emits a standalone Go program that reproduces the pipeline through the lazy API. The generated source is piped through go/format and has its imports pruned by go/ast, so the output is always gofmt'd and free of unused imports.

golars transpile examples/script/pipeline.glr -o main.go --package main
go run main.go

Mapping:

.glrGo
load PATHgolars.ReadCSV(PATH, csv.WithNullValues(""))
load PATH as NAMEstashes the LazyFrame in an internal map for later use
use NAMEretargets focus onto the stashed frame
filter EXPR.Filter(EXPR)
with NAME = EXPR.WithColumns(EXPR.Alias("NAME"))
select A, n = EXPR.Select(expr.Col("A"), EXPR.Alias("n"))
groupby KEY COL:OP[:ALIAS] ....GroupBy(KEY).Agg(...)
group_by_dynamic T every D AGG....GroupByDynamic(T, dataframe.DynamicGroupOptions{Every: D}).Agg(...)
sort COL [desc].Sort(COL, desc)
sort A asc B desc.SortBy([]string{A, B}, []compute.SortOptions{...})
join_asof NAME on KEY ....JoinAsof(other, dataframe.AsofOptions{On: KEY, ...})
to_dummies, unnest, top_k, bottom_kCollect, the DataFrame method, then lazy.FromDataFrame
limit N.Limit(N)
head N.Limit(N) + .Collect + fmt.Println
join NAME on KEY [inner|left|cross].Join(other, []string{KEY}, dataframe.InnerJoin)
show.Head(10).Collect + fmt.Println
save PATHgolars.WriteCSV(df, PATH) (or matching writer)

Expressions are lowered to the same expr calls a hand-written program would use (dt.year(ts) becomes expr.Col("ts").Dt().Year()). REPL-only commands (.tree, .graph, .mermaid, .frames, ...) are skipped and listed in a header comment; other commands without a lowering emit a TODO(glr): comment so the file still compiles. If no show / head / collect / save appears in the script, transpile adds an implicit final Collect + fmt.Println so the generated binary prints something instead of exiting silently.

See examples/script/transpiled/ for a transpiled copy of every bundled .glr example.

Interop with code

Anything that implements script.Executor can host the language. cmd/golars is the reference, but the package ships a generic runner:

import "github.com/Gaurav-Gosain/golars/script"

r := script.Runner{
    Exec:  script.ExecutorFunc(func(line string) error { /* … */ return nil }),
    Trace: func(line string) { fmt.Println(">", line) },
    ContinueOnErr: true,
    ErrOut: os.Stderr,
}
if err := r.RunFile("pipeline.glr"); err != nil {
    log.Fatal(err)
}
  • Trace receives every normalised statement just before execution.
  • ContinueOnErr + ErrOut emits errors inline and keeps running.
  • script.Normalize(raw) is exported so third parties can apply the same parsing rules (comment stripping, leading . insertion).

Editor support

Tree-sitter grammar + highlight queries live at editors/tree-sitter-golars/. Install notes for Neovim (nvim-treesitter) and VS Code in that directory's README.

LSP

golars-lsp in Neovim: inlay hints showing frame shape after every .glr statement

cmd/golars-lsp is a minimal Language Server that ships:

  • Inline completions for commands, staged-frame names, file paths, and column names read from loaded CSV files.
  • Inlay hints showing each pipeline step's output shape - → 5 rows × 3 cols appears at the end of every shape-changing statement. Row counts propagate as upper bounds: limit N clamps to N, left joins preserve the left side's count, filters and inner joins mark rows ?.
  • Hover docs with signature + long description on any command token.
  • Diagnostics for unknown commands and files that don't resolve.

# ^? probe: live table previews

Drop # ^? on its own line anywhere in a script and the Neovim plugin renders the focused frame's current table as virtual text below the comment (via a golars --preview subprocess). This is the scripting equivalent of Twoslash/Quokka probes: a live peek at the data at that pipeline position:

load data/trades.csv
filter volume > 100
sort amount desc
limit 5
# ^?

The preview updates on save + debounced text changes; configure via require("golars").setup({ preview_cmd = { "/path/to/golars" }, preview_rows = 20, preview_timeout_ms = 3000 }). Set preview = false to disable.

golars --preview <path>

Invoke the preview pipeline from any editor, or manually via:

golars --preview path/to/script.glr
golars --preview path/to/script.glr --preview-rows 25

Runs the script silently (no banner, no trace, no success chrome) and prints exactly one rendered table: the focused pipeline's head. Exit code 0 on success, non-zero on script error (message on stderr).


What this language is NOT

  • No variables beyond the named-frame registry. If you need branching or reusable expressions, write a Go program that drives script.Runner with your own logic.
  • No control flow (no if, no loops). The idiom for conditional runs is shell scripting around golars run, or a Go host with an Executor that dispatches.
  • No user-defined functions. with NAME = EXPR covers derived columns; anything beyond that belongs in Go, where the full expression API is available.

The design target is "drop a day of REPL work into a file and have it run again tomorrow." Everything else is out of scope.

On this page