Cloud storage, and why not just S3
Pulling a catalogue
This is the other half of “we pull, on a schedule, with whatever credentials the owner gave us”: for a bucket the owner already has, on any S3-compatible storage. No vendor is named anywhere in this: not in the schema, not in the connect form, not in this document beyond the phrase “S3-compatible” itself.
Connecting is verification, not prediction. The form asks for an endpoint, a bucket, an access key id and a secret, plus an optional folder prefix. Region and path-style-versus-virtual-hosted addressing sit behind an “advanced” disclosure, since both can usually be derived from the endpoint and only occasionally need a person to override them. Saving runs a real ListBucket and reports whatever came back, including the storage’s own error text, rather than trying to predict whether the details are right.
Read-only, always. The key is asked for and used as GetObject and ListBucket only. There is no code path anywhere that writes to, deletes from, or otherwise modifies somebody’s bucket. Filling it happens in the storage provider’s own tools, exactly as filling the bridge’s folder happens in the computer’s own file manager.
Cataloguing is a scheduled sweep, not a one-off. Every ten minutes, and once immediately after a bucket is connected, we list it (respecting the prefix) and match media by extension: .mp4, .m4v, .webm, .mov, .mp3 and .m4a are playable, and .mkv, .avi, .wmv and .flv are catalogued too and marked playable: false, for the reason given under Formats. We parse “Artist - Title” out of the filename the same way the bridge does (see below), and upsert the result through the same internal path that the sync endpoint uses, so there is exactly one place a catalogue lands, regardless of which direction it arrived from.
The filename is the title. The extension is dropped and underscores become spaces. If what is left contains a space, a hyphen and a space, Artist - Title, then what comes before the first one is the artist and what comes after is the title, and a title that contains the separator again keeps it. If it does not, there is no artist and the whole name is the title. A hyphen with no spaces around it is not a separator. Tags inside the file are not read. This is the same rule the bridge follows for a file on a disk, and it is described for people in a karaoke night.
Artwork follows the bridge’s own convention: a sibling image beside each file, same basename, any of .jpg, .jpeg, .png or .webp, matched without regard to case. Since the service can reach the bucket directly, the image is fetched once and stored as bytes, by the same path as a push to /providers/art (see Artwork), rather than as a link, which would expire in the catalogue long before anyone looked at it. An object’s own ETag is used to skip re-fetching an image that has not changed since the last sweep, and the same 1 MB limit applies.
The object key is the externalId. Unlike the bridge’s content fingerprint, this means renaming an object in the bucket loses any correction attached to it. The schema has nowhere else to keep a stable id for something we only ever list, never fingerprint ourselves.
Resolving media mints a link, not a token. Where a bridge’s address carries a bearer secret with no expiry, a cloud track resolves to a presigned GetObject link good for a few minutes: long enough for an ordinary song, short enough that a leaked link is not worth much. The Stage asks for this itself when a song is about to play, holds it, and asks again before it expires, or on resuming a song that was paused past its link’s lifetime.
canSeek is measured. On every sweep we ask the bucket for a range of one of its playable files and note whether it honoured it, rather than assuming. A bucket is never declared canAnalyseAudio.
Why this is not just S3
A recurring suggestion: drop all of the above and say “a provider is an S3-compatible bucket”. It is a good instinct — one protocol, off-the-shelf servers, a presigned link instead of a media token — and it was considered and rejected. The reasons, so it does not have to be relitigated from scratch:
The expensive half is unshared. Speaking S3 means verifying SigV4 on every request: canonical requests, signed headers, payload hashes, clock skew. That is security-critical code replacing a token comparison, and a client-side presigner, which we would write anyway to consume buckets, does not help with any of it.
It does not unify the thing that is actually split. The catalogue moves in opposite directions depending on whether we can reach you, and that is about reachability, not protocol. A bucket on a home NAS still cannot be polled from outside it. S3 changes nothing about that.
It widens what a provider exposes. Today the contract is one verb on one path. A real S3 endpoint invites ListObjects, which is a catalogue anyone on the network can walk.
It raises the floor for everyone else. “Serve one GET with range support” is twenty lines in any language, which is the entire point. “Implement S3” is not, and the contract exists so that a provider can be written by somebody who has never heard of us.
A bucket is still a perfectly good source. It just consumes this contract rather than replacing it, with the link minted rather than concatenated, as described above. Signing is pure computation, so that link can even point at an address we cannot ourselves reach: a NAS could have its catalogue pushed by something on its own network while media resolves through a presigned link the Stage dials directly. The certificate problem is the same one the bridge has, and has the same answer, described in A name and a certificate.