Skip to content

Storage Guide

This guide covers the most common storage workflows.

Initialize client

from gfslib.storage.client import StorageServices

svc = StorageServices("https://.../services/storage")
svc.set_api_key("your_api_key")

List files

short = svc.ls()
long = svc.ls_long()

# List inside a folder
folder = svc.ls(path="sync-tests", recursive=False, include_directories=True)

# Recursive listing including subfolders
tree = svc.ls_long(path="sync-tests", recursive=True, include_directories=True)

ls() and ls_long() both accept these optional parameters:

  • path: list a specific remote folder or prefix
  • recursive: include nested files and folders
  • include_directories: include directory entries in the response when supported by the server

Inspect a remote path

svc.exists("folder/sample.txt")
svc.is_file("folder/sample.txt")
svc.is_dir("folder")
info = svc.path_info("folder/sample.txt")

Missing paths return False from the boolean helpers. Authentication and server errors raise exceptions. path_info() returns the server's exists flag and, when present, the entry containing path, type, size, and modification time.

Upload a file

svc.upload("folder/sample.txt", "path/to/local/sample.txt")

Prefer upload_file(remote_path, local_path) when uploading a file and upload_text(remote_path, text) when uploading text. upload_file accepts pathlib.Path and raises FileNotFoundError for missing files. The existing upload method still treats strings that are not existing paths as text.

metadata() raises on HTTP errors and invalid JSON. Sync raises on failed uploads instead of reporting success. Low-level upload methods continue to return a response; call raise_for_status() when using them directly.

Upload multiple files in one request

from pathlib import Path

svc.upload_many({
    "inputs/table.csv": Path("./table.csv"),
    "inputs/image.tif": Path("./image.tif"),
})

upload_many() accepts remote paths mapped to local filenames. It hashes and streams one file at a time, with a configurable chunk_size (default 1 MiB), without loading the entire upload into memory. Paths are validated before the request. An empty mapping or a nonpositive chunk size raises ValueError.

The server's fileComparison policy controls replacement; file_comparison defaults to "TimeModified". A successful request does not imply every existing file was replaced. Source modification during transfer raises an error when detected. This operation is not transactional: earlier files may already exist on the server after a failure. No automatic retries are performed.

Download a file

# bytes in memory
content = svc.download("folder/sample.txt")

# write directly to disk
svc.download("folder/sample.txt", "downloaded/sample.txt")

# download multiple files into a folder
paths = [
    "folder/sample.txt",
    "folder/sub/other.txt",
]
svc.download(paths, dest="./downloads")

# skip paths that do not exist instead of failing the whole batch
svc.download(paths, dest="./downloads", ignore_missing=True)

# Preserve workspace-relative paths and verify SHA-256 before publishing files
svc.download(paths, dest="./downloads", preserve_paths=True,
             verify_sha=True, overwrite=False)

# Download folder contents, relative to the selected remote directory
svc.download_directory("folder", "./folder-copy", verify_sha=True)

Batch downloads retain their existing flat layout by default, but reject duplicate destination names. download_directory() preserves the tree and creates its destination folder. It defaults to overwrite=False; existing files raise FileExistsError. Single/batch downloads retain overwrite=True. Empty remote directories are not represented by the server's download stream.

Files are written to temporary siblings and published only after a complete transfer and successful checksum verification when requested. A failed file leaves an existing destination unchanged. Earlier completed files remain if a later file in a batch fails. verify_sha=True requires server hash metadata and cannot be combined with ignore_sha=True or partial (byte_range) downloads.

Metadata

meta = svc.metadata(["folder/sample.txt"], ignore_sha=False)
print(meta)

Create and remove directories

svc.mkdir("empty-folder")
svc.mkdir(["folder-one", "folder-two"])
results = svc.rmdir("empty-folder")
results = svc.rmdir(["folder-one", "folder-two"], recursive=True)

Both methods accept a string or an iterable of paths and return the server's per-path path, status, and error records. Inspect these records for partial failures even when the HTTP request succeeds. rmdir() defaults to recursive=False, so it does not remove a directory's contents.

Rename or move

svc.rename("uploads/old.txt", "archive/new.txt")
svc.rename("uploads/latest.txt", "archive/latest.txt", overwrite=True)

Renaming happens on the server. The default is overwrite=False; HTTP failures raise exceptions. The returned value is the server response.

Delete a single/multiple files/folders

svc.delete("folder/sample.txt")

svc.delete([
    "folder_one",
    "folder_two/sample.txt",
    "folder_two/sub/other.txt",
])

Sync local folder to remote

result = svc.sync_local_to_remote(
    local_dir="./data",
    remote_prefix="uploads",
    ignore_sha=False,
    dry_run=False,
)
print(result)

Plan and filter syncs

plan = svc.plan_sync("./data", remote_prefix="uploads",
                     include=["*.csv", "*.tif"], exclude=["scratch/*"])
for entry in plan:
    print(entry.path, entry.action, entry.reason)

svc.sync_local_to_remote("./data", "uploads", include=["*.csv"])
svc.sync_remote_to_local("./copy", "uploads", dry_run=True)
svc.sync_remote_to_local("./copy", "uploads", verify_sha=True)

plan_sync() returns SyncEntry objects without modifying local or remote files. Set direction="download" to plan remote-to-local transfers. Paths are relative to the selected directory. Patterns use case-sensitive shell matching on paths with forward slashes; * can match slashes. Exclusions take precedence and an empty include=[] selects nothing. These are not gitignore patterns.

Both sync methods return relative paths mapped to uploaded, downloaded, or skipped; planned transfers use uploaded (dry-run) or downloaded (dry-run). Dry runs do not create destination directories. Neither direction deletes files. ignore_sha=True transfers all selected files instead of comparing hashes. Remote-to-local sync defaults to replacing changed files; overwrite=False protects existing files. verify_sha=True checks downloaded content separately from the comparison used to plan transfers.