2023-03-04 · 5 min
sportsdataversedata: sportsdataverse data storage functions
sportsdataversedata is both a repository and an R package. As a repository it is the release store for the SportsDataverse: every processed dataset any load_*() function reads (play-by-play, schedules, rosters, and the rest) is published here as a tagged GitHub release. As a package it is the plumbing that writes those releases — save a data frame in the right formats, attach the metadata that lets a reader know when it was last refreshed, and upload it. It is not something most users install directly; it is what the producer side of every SportsDataverse *-data pipeline depends on. It sits underneath the rest of the R family, closer to infrastructure than to analysis.
Each dataset is produced by a pair of pipelines: a *-raw repository scrapes a source on a schedule, then dispatches to a *-data repository that cleans the result and calls into this package to publish it. wehoop-wnba-raw feeding wehoop-wnba-data is one example of the pair; cfbfastR-raw feeding cfbfastR-data is another. sportsdataversedata is the shared endpoint every one of those *-data repositories writes to.
Installation
sportsdataversedata is not on CRAN. Install the development version from GitHub with pak:
pak::pak("sportsdataverse/sportsdataverse-data")Uploading needs a GitHub token with permission to write releases on the target repository. sportsdataverse_save() defaults its .token argument to gh::gh_token(), so the usual gh package / GITHUB_PAT environment variable setup covers it; the gh_cli_* helpers instead shell out to the gh CLI directly and expect it to already be authenticated (gh_cli_available() checks that it's on the PATH and working).
Getting started
library(sportsdataversedata)
sportsdataverse_save(
data_frame = my_df,
file_name = "espn_cfb_pbp_2024",
sportsdataverse_type = "cfb play-by-play",
release_tag = "cfb_pbp",
pkg_function = "cfbfastR::load_cfb_pbp"
)sportsdataverse_save() attaches sportsdataverse_type and a timestamp as attributes on data_frame, writes it to a temporary directory in every format listed in file_types, and uploads the results to release_tag on the target repo. Nothing is returned to the console beyond the upload result; the release itself is what changes.
What's in the box
Nine exported names, split between the upload/save pair and the gh CLI wrappers underneath them:
sportsdataverse_save(data_frame, file_name, sportsdataverse_type, release_tag, pkg_function, .token = gh::gh_token(), file_types = c("rds", "csv", "parquet"), repo = "sportsdataverse/sportsdataverse-data")— the main entry point: tags a data frame with type and timestamp metadata, writes it out in each offile_types(any of"rds","csv","parquet","qs","csv.gz"), and uploads every file torelease_tag.sportsdataverse_upload(...)— the lower-level upload stepsportsdataverse_save()calls internally, for uploading pre-written files (files,tag,pkg_function,repo,overwrite = TRUE) without the save/format step.gh_cli_available()— checks whether theghCLI is installed and onPATH; the othergh_cli_*functions depend on it.gh_cli_release_upload(files, tag, ..., repo = "sportsdataverse/sportsdataverse-data", overwrite = TRUE)— uploads local files to a release tag viagh release upload, skipping any file path that doesn't exist rather than erroring outright.gh_cli_release_tags(repo = "sportsdataverse/sportsdataverse-data")— lists every release tag in a repository viagh release list.gh_cli_release_assets(tag, ..., repo = "sportsdataverse/sportsdataverse-data")— lists the assets attached to one release tag viagh release view.gh_cli_rate_limits(verbose = TRUE)— reports the current GitHub API rate limit viagh api rate_limit, with the reset time parsed to a readable timestamp..invoke_cli_command()/.cli_parse_json()— internal helpers that run aghCLI command and parse its JSON output; exported so other SportsDataverse pipelines can reuse them directly instead of duplicating the shell-out logic.
A second worked example, reading a release's assets back rather than writing them:
library(sportsdataversedata)
gh_cli_available()
tags <- gh_cli_release_tags(repo = "sportsdataverse/sportsdataverse-data")
head(tags)
assets <- gh_cli_release_assets(tag = "cfb_pbp", repo = "sportsdataverse/sportsdataverse-data")
assetsgh_cli_available() returns TRUE when the gh CLI is on the path and stops with an error (through cli::cli_abort()) when it is not, so wrap it in tryCatch() if you want a fallback branch instead of a stop. gh_cli_release_tags() returns every release tag the repository has, and gh_cli_release_assets() returns the asset listing for one of them — the same information a load_*() function's URL-building logic reads to find the current file for a given season or dataset.
Good to know
- The timestamp contract. Every
sportsdataverse_save()/sportsdataverse_upload()call writes atimestamp.json(and a plain-texttimestamp.txt) alongside the data assets, generated right before the upload begins rather than after, so it reflects when the release was actually pushed. Downstream, a release's "last updated" badge reads thattimestamp.jsonlive; a release with data assets but notimestamp.jsonpredates the timestamp being added, and shows its GitHub release date instead. A release tag with no data assets at all means that dataset's pipeline hasn't produced anything yet. - Formats.
sportsdataverse_save()defaults tords,csvandparquet; passfile_typesexplicitly to addqsorcsv.gz, or to narrow the set. - Repo default. Every function defaults
repoto"sportsdataverse/sportsdataverse-data", but accepts anyowner/repostring, so the same helpers back other SportsDataverse release repositories too. - Automation, not analysis. The repository's own README is mostly a live status dashboard — per-league scrape-and-process badges and a full release-by-release "last updated" listing — for maintainers checking whether a pipeline has gone stale, not a data catalog to browse for finding a dataset by hand.
- Python port.
sportsdataverse.releasein sportsdataverse-py is a port of this same plumbing — a GitHub-release publish/download layer plus a pure-Python byte-parity RDS writer — so a Python pipeline can write releases this package can read back in R, and vice versa. - Rate limits.
gh_cli_rate_limits(verbose = TRUE)reports the GitHub API's current rate limit and parses the reset time into a readable timestamp — worth checking before a large batch upload, since everygh_cli_release_upload()call spends part of that budget. - Missing files are skipped, not fatal.
gh_cli_release_upload()checksfile.exists()on every path first; if some are missing it warns and uploads the files that do exist rather than aborting the whole call. - License. The repository is licensed CC BY 4.0, not the MIT license most SportsDataverse code packages use — the data itself carries an attribution requirement separate from the code that produces it.
Related
sportsdataversedata is the release store behind every SportsDataverse load_*() function in cfbfastR, hoopR, wehoop, baseballr and fastRhockey, and it is installed alongside those by sportsdataverse (R). The same release-store role exists in Python as sportsdataverse.release, part of sportsdataverse-py.
Data and automation
sportsdataverse-data releases is where the assets land. The producer repos that publish through this package's helpers (or their Python port) are: cfbfastR-data, cfbfastR-cfb-data, cfbfastR-cfb-raw, ncaa-mfb-football-data, hoopR-mbb-data, hoopR-nba-data, hoopR-nba-stats-data, ncaa-mbb-hoops-data, wehoop-wbb-data, wehoop-wnba-data, wehoop-wnba-stats-data, ncaa-wbb-hoops-data, fastRhockey-nhl-data, fastRhockey-pwhl-data, baseballr-data. Each one's workflow status is on its Actions tab and in the nightly ecosystem status.
- Package checks:
- Cheat sheets for the rest of the family are at sportsdataverse.org/cheatsheets; this package does not have one.
- Ecosystem status — a nightly snapshot of every SportsDataverse repo: workflow conclusions, release-asset freshness, open PRs and issues.
Links
My role: author and maintainer. Part of the SportsDataverse — open sports data tooling for R, Python and JavaScript.