Utility functions

Package providing utility functions for the user.

tctrack.utils.batching(tracker, input_files, interval, *, time_range, preprocessing=None, retrieve_data=None, tracker_inputs=['*'], config=None)[source]

Perform tracking in batches with optional steps for retrieval and preprocessing.

NetCDF input files for each batch period are put inside batch_[i] directories inside the output directory. After any preprocessing, these are then provided to the tracker’s run_tracker() method.

The outputs tracks are placed in the output directory with names tracks_[i].nc. If combine_outputs is True then these will be combined into a single tracks.nc file.

Parameters:
  • tracker (TCTracker) – The tracker object used to perform the tropical cyclone tracking.

  • input_files (str | Iterable[str | tuple[str, InputArgs]]) –

    Input netcdf files to use for each batch. These will be copied into the batch directory. The * wildcard can be used to match multiple files, in which case they will be combined into one file. A tuple can also be passed for each input file where the second value is a dictionary providing additional arguments. These can be:

    • batch_file: The filename for the file in the batch directory. Use None to not do so. By default it will use the same file name.

    • store: Keys for storing the fields in-memory for use in the preprocessing. If there are multiple fields these will be matched by netcdf variable name if possible, otherwise they must match the full number of fields, in which case they are matched positionally.

    • time_varying: Whether the fields should be subset in time. Default: True.

  • interval ({"month", "year"} | cf.TimeDuration) – The calendar interval of each batch. The final batch may be shorter to end at the end of time_range.

  • time_range (tuple[str, str]) – The start and end datetimes for all batches. Must in YYYY-MM-DD format. The end time is open (not inclusive).

  • preprocessing (Sequence[PreprocessStep] | None) –

    (optional) The list of preprocessing steps. These are each specified by a tuple.

    • The first entry in the tuple should be one of the preprocessing functions defined in tctrack.preprocessing. Alternatively, it can be a function that returns either a cf.Field, a list of cf.Field, or nothing. If it takes a field / fields as input this should be the first argument.

    • The second entry is a dictionary containing the arguments to pass to the function. String arguments can refer to the batch directory with %BATCH%.

    • The optional third entry is a dictionary that can take store and/or use keys which allows fields to be stored and passed from memory to avoid unnecessary file IO. store behaves the same as in input_files.

  • retrieve_data (Callable[[int, Path], None] | None) – (optional) A user-defined function that is called each iteration to retrieve data and put it in the batch directory. E.g. to download the data if it will not all fit on the filesystem in one go. The first argument of the function is the batch index, the second argument is the batch directory.

  • tracker_inputs (Iterable[str]) – (optional) A list of filenames from the batch directory to pass to the tracker. By default it uses all the files (using ["*"]).

  • config (BatchingConfig | None) –

    (optional) A dictionary of additional arguments.

    Option

    Description

    output_dir

    The location to save the outputs.
    Default: "tctrack_outputs".

    combine_outputs

    Whether to combine the outputs from each batch
    into a single tracks.nc file. Default: True.

    delete_batch_dirs

    Whether to delete the batch directories.
    Default: True.

    buffer_period

    timedelta for extra time at the end of each batch
    so tracks reaching a batch boundary are not cut off.
    Default: None.

    start_buffer_period

    timedelta for extra time included at the start
    of each batch. This ensures that any tracks starting
    before the batch period are removed. Defaults to
    one day, but is only used when buffer_period is set.

Return type:

None

Examples

Get tracks for each month in 1950 using tempest extremes. The input files matching “psl_*.nc” are loaded and then preprocessed to halve the latitude resolution and rename the netCDF variable for each month before running the tracker.

>>> from tctrack.utils import batching
>>> from tctrack import tempest_extremes as te
>>> from tctrack.preprocessing import subsample_field, set_nc_variable_name
>>> batching(
...     te.TETracker(),
...     [("psl_*.nc", {"store": "psl", "batch_file": None})],
...     interval="month",
...     time_range=("1950-01-01", "1951-01-01"),
...     preprocessing=[
...         (
...             subsample_field,
...             {"X": slice(0, None, 2)},
...             {"use": "psl", "store": "psl"},
...         ),
...         (
...             set_nc_variable_name,
...             {"field_name": "p", "output_file": "%BATCH%/p.nc"},
...             {"use": "psl"},
...         ),
...     ],
... )
tctrack.utils.load_tracker_metadata(filename)[source]

Read the TCTrack metadata from an output file, returns the parameters.

Parameters:

filename (str) – The filename of the output netcdf file containing TC tracks.

Return type:

tuple[type, list[TCTrackerParameters]]

Returns:

  • tracker_cls (type) – The relevant TCTracker type. This can be initialised by: tracker_cls(*parameters).

  • parameters (list[TCTrackerParameters]) – A list of parameter objects that duplicate the information in the metadata.

Examples

To read in the metadata from an output file, instantiate a tracker and reproduce the results:

>>> tracker_cls, parameters = load_tracker_metadata("tracks.nc")
>>> tracker = tracker_cls(*parameters)
>>> tracker.run_tracker("inputs.nc", "tracks_new.nc")
tctrack.utils.read_tracker_metadata(filename)[source]

Print the TCTrack metadata from an output file to stdout in a readable format.

Parameters:

filename (str) – The filename of the output netcdf file containing TC tracks.

Return type:

None