API Reference

The public surface is everything importable from nemdatatools. All date arguments accept naive datetimes, dates, or YYYY/MM/DD / YYYY/MM/DD HH:MM:SS strings, interpreted as NEM time (fixed UTC+10).

Fetching data

nemdatatools.fetch(table: str, start: str | datetime, end: str | datetime, regions: list[str] | None = None, cache: Cache | None = None) → DataFrame[source]

Fetch one table over a date range, stitching tiers as needed.

Parameters:
  • table – Curated table name (see nemdatatools.tables()).

  • start – Range start, inclusive, naive NEM time.

  • end – Range end, inclusive, naive NEM time.

  • regions – Optional NEM region filter, for tables with a region column.

  • cache – Download/parse cache; a default one under ~/.nemdatatools is used when omitted.

Returns:

Rows sorted by the table’s time column, which becomes the index. De-duplicated on the table’s identity columns where tiers overlap.

Raises:
  • AvailabilityGapError – If the range crosses a known hole in the table’s history (the message names the substitute table).

  • CoverageError – If part of the range cannot be served by any tier.

  • KeyError – If the table is not in the curated catalog.

nemdatatools.fetch_price_and_demand(start: str | datetime, end: str | datetime, regions: list[str] | None = None, cache: Cache | None = None) → DataFrame[source]

Fetch aggregated 5-minute price and demand by region.

Parameters:
  • start – Range start, inclusive, naive NEM time.

  • end – Range end, inclusive, naive NEM time.

  • regions – NEM regions to include; all five when omitted.

  • cache – Download cache; a default one is used when omitted.

Returns:

Columns REGION, TOTALDEMAND, RRP, PERIODTYPE indexed by SETTLEMENTDATE, sorted, covering [start, end].

nemdatatools.fetch_mmsdm_table(table: str, start: str | datetime, end: str | datetime, subdir: str = 'DATA', cache: Cache | None = None) → DataFrame[source]

Fetch any MMSDM table by name, without curated metadata.

Escape hatch for the ~236 archive tables outside the curated catalog: files are still discovered by listing-and-matching in both filename eras and all FILEnn parts are fetched, but no Reports stitching, gap checking, or de-duplication is applied. The time column is auto-detected for range filtering; rows are returned unfiltered when none is recognised.

Parameters:
  • table – MMSDM filename table token, e.g. "GENCONDATA".

  • start – Range start, inclusive, naive NEM time.

  • end – Range end, inclusive, naive NEM time.

  • subdir – Snapshot subdirectory (DATA, P5MIN_ALL_DATA or PREDISP_ALL_DATA).

  • cache – Download/parse cache; a default one is used when omitted.

Returns:

All rows of the table across the months covering the range.

Raises:

CoverageError – If no month in the range carries the table.

Discovery

nemdatatools.tables() → list[str][source]

List the curated table names accepted by fetch().

nemdatatools.availability(table: str) → dict[str, object][source]

Describe where and when a curated table is available.

Parameters:

table – Curated table name.

Returns:

A mapping with the table’s tier locations and known MMSDM gaps, suitable for printing.

nemdatatools.NEM_REGIONS = ('NSW1', 'QLD1', 'SA1', 'TAS1', 'VIC1')

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable’s items.

If the argument is a tuple, the return value is the same object.

Transforms

nemdatatools.resample(frame: DataFrame, rule: str, agg: str | dict[str, str] = 'mean', by: str | list[str] | None = None, trading_day: bool = False) → DataFrame[source]

Resample interval-ending AEMO data to a coarser interval.

Parameters:
  • frame – Data indexed by a DatetimeIndex of interval-ending timestamps (as returned by nemdatatools.fetch()).

  • rule – Target interval, e.g. "30min", "1h", "1D".

  • agg – Aggregation for numeric columns — a single function name, or a column-to-function mapping (non-mapped columns are dropped).

  • by – Identity column(s) to group by first, e.g. "REGIONID" or "DUID". Required when the frame carries more than one entity, otherwise their values would be averaged together.

  • trading_day – When aggregating to days or coarser, align buckets to the 04:00-04:00 NEM trading day instead of midnight.

Returns:

Aggregated rows labelled with interval-ending timestamps, grouped columns preserved as regular columns.

Raises:

ValueError – If the index is not a DatetimeIndex, or the frame has several entities and by was not given.

Caching

class nemdatatools.Cache(root: Path | str = PosixPath('/home/runner/.nemdatatools'), session: Session | None = None)[source]

Raw-download and parsed-table cache rooted at one directory.

download(url: str) → Path[source]

Return a local copy of url, downloading only when missing.

Parameters:

url – Absolute nemweb file URL.

Returns:

Path of the cached raw file.

Raises:

requests.HTTPError – If the download fails.

load_table(url: str, cid_key: tuple[str, str]) → DataFrame[source]

Fetch url and return one table from its C/I/D payload.

Results are memoized as Parquet per (payload, table); other tables found while parsing are memoized too, so multi-table payloads (e.g. TradingIS) parse once.

Parameters:
  • url – Absolute nemweb zip URL.

  • cid_key – (component, table) to extract.

Returns:

The table’s rows from this payload; empty when the payload carries no such segment.

Errors

exception nemdatatools.NemDataError[source]

Base class for nemdatatools errors.

exception nemdatatools.AvailabilityGapError[source]

The requested range crosses a known hole in a table’s history.

The message names the gap and the substitute table to use across it; callers who want partial data can re-request around the gap bounds.

exception nemdatatools.CoverageError[source]

No publication tier can serve part of the requested range.