Migrating to Batch v2
The v2 Batch and Files API replaces single-file batch inputs with multi-file upload sessions and batches, improving speed, security and flexibility.
The core changes to the Batch API are related to how you get data in and results out: instead of uploading one file at a time
you can now upload one or more files into a session, keeping subfolder organizations and reference them when submitting a batch job as avra://uploads/<session_id>/file.ext.
What changed
| Concern | v1 (deprecated) | v2 |
|---|---|---|
| Upload | POST /v1/api/batch-inputs/upload — one file, referenced by UUID input_file_id |
POST /v2/api/files/upload — many files per session, integrity-checked with MD5 |
| Discover files | GET /v1/api/batch-inputs |
GET /v2/api/files with prefix / session / suffix filters |
| Create batch | input_file_id + optional inputs, webhook_url |
input.files[] (avra:// paths), optional reference_date for backtesting |
| Result | content_at URL on GET /v1/api/batches/{id}/result |
download_url + expires_at on GET /v2/api/batches/{id}/result |
Webhooks are now configurable only through our notifications tool, via dashboard (see Webhooks). Input and Result deletions and Input downloads are no longer available in self-service ways to ensure security. Reach out to Avra if you need to remove a file.
Accepted file formats
- Parquet — recommended for all inputs, and the format results are written in (
result.parquet). - CSV — supported for low-complexity schemas.
- JSONL — deprecated due to low compression. Migrate to Parquet.
Migration steps
Upload files into a session
Call POST /v2/api/files/upload with one target per file. Each target declares the destination file path, its content_type, and a base64-encoded md5 digest. You will receive one signed upload_url per target and a session expiration.
{
"targets": [
{ "file": "base/input1.csv", "content_type": "text/csv", "md5": "KPRxZBtAn5UfexFEKfKr7Q==" }
]
}For each target file, you will receive:
file: a reference for that file within Avra’s domain,avra://uploads/<session_id>/file.extcontent_type: a MIME type for that file, matching the Content-Type provided in request or, if not provided, inferred from the file extension.md5: exact MD5 provided in the request as a counter-proof to use.upload_url: a signed, short-lived URL to upload the described file, that expectsContent-TypeandContent-MD5headers matching the response values. A mismatched digest will be rejected by the storage.
{
"targets": [
{
"file": "avra://uploads/1784669990715/base/input1.csv",
"content_type": "text/csv",
"md5": "KPRxZBtAn5UfexFEKfKr7Q==",
"upload_url": "https://storage.googleapis.com/..."
}
],
"expiration": "2026-07-21T22:39:50Z"
}(Optional) List what you uploaded
Call GET /v2/api/files to enumerate uploaded files, optionally filtered by session, prefix, or suffix. Paths are returned in the avra:// scheme, ready to reference in a batch.
Create the batch
Call POST /v2/api/batches referencing the uploaded files. Provide model_id + model_version_id (or alias), an optional reference_date if backtesting, and the input definition.
{
"model_id": "00000000-0000-0000-0000-000000000000",
"model_version_id": "00000000-0000-0000-0000-000000000000",
"reference_date": "2023-01-01",
"input": {
"files": [
"avra://uploads/1784143129686/base/input1.csv"
]
}
}Avra returns the batch id and status so you can track the job.
Poll and download
Poll GET /v2/api/batches/{id} for status. When complete, call GET /v2/api/batches/{id}/result to receive a signed download_url that expires at expires_at.