Skip to content
FiveCord Docs

Data harvests

A data harvest is a ZIP archive of the current account’s data. FiveCord compiles it in the background, reports progress while it runs, and issues a temporary download URL once the archive is written.

Every route except Download data harvest archive is user-only. FiveCord rejects a bot or OAuth2 bearer credential with 403 ACCESS_DENIED, and an account that has suspicious activity flags with 403 ACCOUNT_SUSPICIOUS_ACTIVITY. No route here requires sudo mode. Delete current user’s messages takes the same filter shape, destroys the messages, and requires it.

Use status to track preparation of the archive.

ValueDescription
pendingNo processing attempt has started
processingAn attempt has started and has neither completed nor failed
completedThe archive was written and is ready for download
failedThe latest attempt recorded a failure

FiveCord retries the underlying job. A retry clears failed_at and error_message, sets started_at again, resets the reported progress to 0, and reports processing while it runs. On success it records completed_at, file_size, and a download deadline, and reports completed.

The acknowledgement returned by both creation operations.

FieldTypeDescription
harvest_idsnowflakeThe ID of the created harvest
status1stringThe harvest status at creation
created_at2ISO8601 timestampWhen the harvest was requested

1 Always pending

2 Recorded at the same moment the snowflake is generated, so ordering harvests by ID is the same as ordering them by created_at

The full state of one harvest. It extends the harvest creation object with progress and completion fields.

FieldTypeDescription
harvest_idsnowflakeThe ID of the harvest
statusstringThe derived harvest status
created_atISO8601 timestampWhen the harvest was requested
started_at1?ISO8601 timestampWhen the current processing attempt began, or null while the harvest is pending
completed_at?ISO8601 timestampWhen the archive was written, or null when it has not been
failed_at2?ISO8601 timestampWhen the latest attempt recorded a failure, or null when it has not
file_size3?stringThe archive size in bytes as a decimal string, or null before the archive was written
progress_percent4numberProgress from 0 through 100
progress_step5?stringThe harvest progress step for the current stage
error_message6?stringThe latest attempt’s failure description, or null when it has not failed
download_url_expires_at7?ISO8601 timestampWhen the archive’s download deadline elapses, or null before completion
expires_at8?ISO8601 timestampWhen the archive’s download deadline elapses, or null before completion

1 A new processing attempt replaces this timestamp, resets progress_percent to 0, and sets progress_step to Starting harvest

2 Cleared when a new attempt starts, so it is non-null only when the latest attempt failed

3 A decimal string because the value is 64-bit

4 Every recorded value is a whole number. It never decreases inside a single attempt, and completion sets it to exactly 100

5 Every record is created with the step Queued, so this member is never null

6 The string form of the underlying failure. It is diagnostic and is not stable for programmatic matching

7 Recorded when the archive is written, exactly seven days after that moment. Past it FiveCord issues no new download URL

8 A second name for download_url_expires_at. The stored value is the same

A short English label describing the current stage. Every value is fixed except the harvested-message label, which includes the message count.

ValueDescription
QueuedRecorded when the harvest record is created, with progress 0
Starting harvestRecorded when a processing attempt begins, with progress reset to 0
Harvesting messagesMessage collection has started, with progress 5
Harvested N messages1Message collection has finished, where N is the exact decimal count collected
Collecting user metadataAccount data is being collected, with progress 60
Downloading attachments and creating archiveThe ZIP archive is being written, with progress 65
CompletedThe archive is ready, with progress 100

1 Skipped when no messages are collected

One issued download URL and its stated expiry.

FieldTypeDescription
download_url1stringThe temporary archive download URL (1-2048 characters)
expires_at2ISO8601 timestampSeven days after this response was produced

1 Resolves to a single ZIP object served as application/zip

2 Exactly seven days after issuance, even when the archive’s own deadline falls earlier

A URL pointing at Download data harvest archive still stops resolving at the archive’s own deadline, so its stated expiry can outlast the point at which it stops working.

The filter chooses which messages a harvest collects. Every non-message section of the archive is included whatever the filter says.

FieldTypeDescription
scope?1stringWhich set of contexts the harvest targets, either selected or inaccessible_only (default selected)
include_dms?booleanWhether to include one-to-one direct messages the caller still has open (default true)
include_dms_closed?2booleanWhether to include one-to-one direct messages the caller has closed (default true)
include_group_dms?booleanWhether to include group DMs the caller is still a member of (default true)
include_guilds?booleanWhether to include text channels in guilds the caller is a member of (default true)
guild_filter_mode?3stringHow the guild list is read, either exclude or include_only (default exclude)
excluded_guild_ids?array[snowflake]The guilds to leave out in exclude mode (max 500, default empty)
included_guild_ids?array[snowflake]The only guilds to include in include_only mode (max 500, default empty)
start_date?4?ISO8601 timestampInclusive lower bound on message timestamps, or null for no lower bound
end_date?4?ISO8601 timestampExclusive upper bound on message timestamps, or null for no upper bound

1 The selected scope requires at least one of the include fields to be true, and a body disabling them all is rejected at path include_dms

2 Independent of include_dms, so setting include_dms false and include_dms_closed true targets closed conversations only

3 The guild lists are read only when include_guilds is true and the scope is selected, and the list that does not match the mode is ignored

4 Compared against the creation time embedded in the message snowflake, so editing an old message does not move it into a later window

Under the selected scope, each context has its own condition:

  • A one-to-one direct message is collected when the conversation is still open and include_dms is true, or when it has been closed and include_dms_closed is true.
  • A group DM is collected only while the caller is still a recipient and include_group_dms is true.
  • A guild channel is collected only while the caller is still a member, include_guilds is true, and the guild is admitted by the configured list mode.

The inaccessible_only scope selects messages in guilds the caller has left or been removed from and in group DMs the caller has left. It never selects one-to-one direct messages, and it ignores the toggles, the guild filter mode, and both guild lists.

Both scopes exclude the caller’s personal notes channel, and both exclude a message whose channel no longer resolves or is neither a direct message, a group DM, nor a guild channel.

Supplying both date bounds requires start_date strictly earlier than end_date. An equal or reversed pair is rejected at path end_date.

POST/v1/users/@me/harvest

Creates an unfiltered harvest of the account. Returns a harvest creation object on success.

The operation takes no request body and always returns before the archive exists, so the caller polls Get data harvest status or Get latest data harvest with the returned ID.

StatusBodyCondition
200harvest creation objectThe harvest request was accepted
404error responseThe authenticated account no longer resolves and the request returns UNKNOWN_USER
500error responseThe harvest record could not be written or the job could not be queued

A harvest record is created with a new snowflake, the request time as created_at, progress 0, the step Queued, and every other timestamp null. Creating a harvest removes no earlier record and enforces no ceiling on how many an account holds, so the route bucket is the only bound.

FiveCord prepares the archive in the background. It contains:

  • user.json, the account document, including connections with their identifiers, visibility settings and verification dates but no credentials
  • one channels/{channel_id}/messages.json file for each channel that contributed a message, with its messages ordered oldest first
  • payments/payment_history.json
  • integrations/oauth.json, the applications the account owns, with no entry for an application it only authorised
  • account/security.json
  • the account avatar and banner under assets/user/ when present

Attachment metadata appears with its message and has the attachment ID, filename, size, content type, CDN URL, and pixel dimensions. The CDN URL is a data package URL, so it does not expire and still reads its attachment after the archive’s own seven day deadline passes. Removing the secret that signed it is the only thing that ends it. The attachment files themselves are never included.

Every authored message is collected, with no ceiling on the count. Messages that no longer exist are omitted. Other read failures fail the attempt. After completion, FiveCord attempts to email a download URL if the account has an email address and the instance has email delivery enabled. That URL expires seven days after it was issued. Email failure does not prevent downloading a completed archive.

5 requests per 30 minutes for each authenticated user, on the user:data:harvest bucket shared with Request filtered data harvest.

POST/v1/users/@me/harvest/filtered

Creates a harvest whose message collection is constrained by a filter. Returns a harvest creation object on success.

The body is a harvest message filter object. Every field is optional, so an empty object is accepted and every default admits its context. An empty request body is read as an empty object.

StatusBodyCondition
200harvest creation objectThe filtered harvest request was accepted
400error responseA filter field failed validation, the selected scope enabled no context, a guild list exceeded 500 entries, or the date bounds were not strictly increasing
404error responseThe authenticated account no longer resolves and the request returns UNKNOWN_USER
500error responseThe harvest record could not be written or the job could not be queued

The side effects match Request data harvest, except that only messages selected by the filter are collected.

5 requests per 30 minutes for each authenticated user, on the user:data:harvest bucket shared with Request data harvest.

GET/v1/users/@me/harvest/latest

Returns the account’s most recently created harvest status object.

Returns the newest harvest regardless of status, including a failed harvest newer than a completed one.

StatusBodyCondition
200harvest status object | nullThe latest harvest state was returned, or null when none exists

40 requests per 10 seconds for each authenticated user, on the user:harvest:latest bucket.

GET/v1/users/@me/harvest/{harvestId}

Returns one harvest status object owned by the current account.

A harvest belonging to another account returns 404 UNKNOWN_HARVEST.

FiveCord removes no harvest record, so an ID that once resolved keeps resolving. A 404 means the ID was never issued to this account, and another account’s harvest is never distinguishable from one that never existed.

FieldTypeDescription
harvestIdsnowflakeThe ID of the harvest
StatusBodyCondition
200harvest status objectThe harvest state was returned
404error responseThe harvest is unknown or belongs to another account, returning UNKNOWN_HARVEST

40 requests per 10 seconds for each authenticated user, on the user:harvest:status bucket.

GET/v1/users/@me/harvest/{harvestId}/download

Issues a temporary download URL for one completed harvest. Returns a harvest download object on success.

FiveCord evaluates the rejection cases in a fixed order, each with its own code. An unknown or foreign harvest is 404 UNKNOWN_HARVEST. A harvest whose latest attempt failed is 400 HARVEST_FAILED. A harvest with no recorded completion or no stored archive is 400 HARVEST_NOT_READY. A harvest past its download deadline is 400 HARVEST_EXPIRED.

The URL takes one of the forms below. When the instance has presigned harvest downloads enabled, which is the default, it is an object storage presigned URL. Otherwise it points at Download data harvest archive on this API with a signed token query parameter.

FieldTypeDescription
harvestIdsnowflakeThe ID of the harvest
StatusBodyCondition
200harvest download objectA download URL was issued
400error responseThe harvest returns HARVEST_NOT_READY, HARVEST_FAILED, or HARVEST_EXPIRED
404error responseThe harvest is unknown or belongs to another account, returning UNKNOWN_HARVEST
500error responseA download URL could not be issued

10 requests per minute for each authenticated user, on the user:harvest:download bucket.

GET/v1/harvest-downloads/{harvestId}Unauthenticated

Streams the archive of one completed harvest.

The token query parameter issued by Get data harvest download URL authorises the download without a session. Use the complete issued URL and keep it secret.

The operation is active only while the instance has presigned harvest downloads disabled. That setting is enabled by default, so a default deployment answers every request here as 404 without inspecting the token.

FieldTypeDescription
harvestIdsnowflakeThe ID of the harvest
FieldTypeDescription
tokenstringSigned download token issued with the download URL
FieldTypeDescription
Range?stringStandard byte range, answered with 206 and a Content-Range header
StatusBodyCondition
200archive bytesThe complete archive was streamed
206archive bytesThe requested byte range was streamed
404Not FoundThe token is missing, invalid, or expired, the harvest is not downloadable, or the instance issues presigned URLs instead
500error responseObject storage could not be read

The 200 has Content-Type: application/zip, Content-Disposition1, Content-Length, Accept-Ranges: bytes, and Cache-Control: private, no-store. The 206 has the 200 headers plus Content-Range. The 404 has Content-Type: text/plain.

1 Content-Disposition is always attachment and names the file fluxer-data-{harvestId}.zip

The archive is read and streamed. The harvest is unchanged and the token is not consumed, so it stays usable until its own expiry.

60 requests per minute for each client IP address, or for the authenticated account when the request has a resolvable credential, on the user:harvest:download_file bucket.