Initial implementation of youtube-automate

A DVR for YouTube subscriptions, laid out so Jellyfin presents each channel
as a show and each video as an episode. Cron-driven, idempotent, with a
public admin UI for a non-operator.

Verified end to end on susan against three real channels: PO tokens, h264
downloads, Jellyfin resolution from local NFOs with all providers disabled,
retention and tombstones.

Corrections to the original design handover (specs.md documents each with
the evidence, and specs.handover-original.md preserves the original):

- The format sort selected 360p. Ranking acodec above res makes `bv*` prefer
  the combined 360p stream, which carries AAC, over the 720p video-only
  stream whose acodec is none. vcodec now leads, so a video without h264 at
  720p yields h264 lower down rather than VP9 this hardware cannot transcode.
- yt-dlp now requires a JS runtime and the yt-dlp-ejs solver scripts, which
  only ship with the [default] extra. Without them the n challenge fails and
  the mweb formats disappear entirely.
- --flat-playlist carries no upload dates, so the specced client-side date
  filter for backfill was impossible. Backfill is RSS-first.
- skipped_old was terminal, so raising a channel's retention appeared to do
  nothing. Added an explicit rescan.
- is_upcoming premieres now defer and retry instead of being skipped forever.
- TubeArchivist is gone, so the media root and the tube.jihakuz.xyz vhost
  were both reclaimed; the latter still pointed at its dead port.

240 offline tests, no network and no real yt-dlp invocation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Tom Flux
2026-08-11 21:42:48 +01:00
co-authored by Claude Opus 5
commit 18bb2e420b
44 changed files with 7188 additions and 0 deletions
+656
View File
@@ -0,0 +1,656 @@
# `ytsubs` — implementation handover
**Target machine:** `susan`
**Audience:** a Claude Code agent implementing this from scratch
**Status:** design agreed, not started
---
## 1. What we're building
A self-hosted pipeline that keeps a rolling 14-day window of recent YouTube uploads from a
handful of subscribed channels, laid out on disk so Jellyfin presents each channel as a TV
show and each video as an episode. A small web UI lets a non-operator (the requester's
brother) subscribe, unsubscribe, and tweak settings.
The mental model is **a DVR for YouTube subscriptions**, not an archiver. Videos are
disposable. If a video ages out or a channel is unsubscribed, it goes away permanently.
### Explicitly in scope
- Poll subscribed channels for new uploads
- Download at ≤720p, h264/AAC, into a Jellyfin-friendly layout
- Generate Kodi-style `.nfo` metadata + artwork (no Jellyfin plugin)
- Delete videos older than the retention window
- Web admin: add/remove channels, edit settings
- Cron-driven, idempotent, monitored via Healthchecks
### Explicitly out of scope — do not build these
- **Cookie handling of any kind.** No browser extension, no cookie-receiving API, no
`cookies.txt`. See §3.
- Members-only, age-restricted, or private content (follows from the above)
- Migration from the existing TubeArchivist install
- Multi-user support, authentication beyond a single shared credential
- Transcoding of any kind (see §6 — this matters on this hardware)
---
## 2. Environment facts you need
- **susan**: Linux rack server, dual Xeon X58, 12 threads, **no AVX**. Old and slow.
Software video transcoding is effectively off the table.
- Already running on susan: Jellyfin, Docker, nginx (Let's Encrypt), Netdata,
Healthchecks (`hc.jihakuz.xyz`), `runitor` for cron wrapping.
- susan is on a **residential IP**. This matters — YouTube flags datacenter IPs far more
aggressively. Do not move any part of this to `victoria` (the Linode VPS).
- Tailnet: Headscale at `headscale.jihakuz.xyz`. The admin UI should bind to the tailnet
interface rather than being published publicly (see §9).
- `/disks/Plex/YouTube` is **already in use by TubeArchivist**. Do not touch it. Use a
fresh root (§5).
- Existing conventions to match: state under `/var/lib/<service>/`, deployed executables
under `/usr/local/bin/`, bare git repos under `/disks/git-repos/`.
### Proposed paths
| Purpose | Path |
|---|---|
| Source (bare repo) | `/disks/git-repos/youtube-automate.git` |
| Checkout | `/opt/youtube-automate` |
| Entry point | `/usr/local/bin/youtube-automate` |
| State DB | `/var/lib/youtube-automate/subs.db` |
| Lock file | `/var/lib/youtube-automate/run.lock` |
| Media root | `/disks/Plex/YouTube/` |
| Scratch/work dir | `/disks/Plex/YouTube/.ingest/` |
The work dir **must** be on the same filesystem as the media root so finished downloads
move into place with an atomic `rename()` rather than a copy. Jellyfin should be
configured to ignore `.ingest` (it's dot-prefixed, which Jellyfin skips by default —
verify).
---
## 3. YouTube access: PO tokens, not cookies
This was the main open question in design and the answer is that it collapses.
yt-dlp's current recommended setup for reliable downloads is the **mweb client plus a PO
Token Provider plugin** supplying tokens for GVS (Google Video Server) requests. No
account, no cookies, no browser. Cookies only unlock age-restricted / members-only /
private content, which is out of scope, and carry a real risk of the Google account being
banned.
### Setup
1. Run the provider as a Docker sidecar:
!! tom: needs a way of starting up when system reboots
```
docker run --name bgutil-provider -d --restart unless-stopped --init \
-p 127.0.0.1:4416:4416 brainicism/bgutil-ytdlp-pot-provider
```
2. Install the matching plugin into the same Python environment as yt-dlp:
!! tom: do this in uv venv
```
pip install bgutil-ytdlp-pot-provider
```
The plugin and server versions should match.
3. Verify before writing any other code:
```
yt-dlp -v 'https://www.youtube.com/watch?v=<some_id>' -F
```
You must see a line like
`[debug] [youtube] [pot] PO Token Providers: bgutil:http-1.x.x (external), ...`
If you don't, stop and fix this first — everything downstream depends on it.
4. `ytsubs` should health-check the provider before each download batch (a plain HTTP
GET against `http://127.0.0.1:4416/ping`) and fail loudly if it's down, rather than
silently accumulating 403s.
### Caveats to be aware of
- The bgutil README carries a warning that PO tokens no longer bypass the "Sign in to
confirm you're not a bot" interstitial in most cases. That is a **different failure
mode** (IP reputation) from the 403s on format URLs that this solves. Residential IP +
low request volume + sleep intervals should keep us clear of it. If it does start
firing, that's the point at which cookies get reconsidered — not before.
- **yt-dlp's YouTube handling churns every few weeks.** Client recommendations,
extractor-arg names, and which clients need tokens all move. Before implementing, read
the current `PO Token Guide` and `Extractors` pages on the yt-dlp wiki and adjust the
flags in §6 accordingly. Do not treat the flags in this document as authoritative —
treat them as the shape of the answer.
- Keep yt-dlp itself updated (weekly `pip install -U yt-dlp` under the same runitor
wrapper). A stale yt-dlp is the single most likely cause of "everything broke".
---
## 4. Discovery: RSS feeds
YouTube publishes an unauthenticated Atom feed per channel. Use the **undocumented
`UULF` playlist variant**, which returns long-form videos only — no Shorts, no
livestreams:
```
https://www.youtube.com/feeds/videos.xml?playlist_id=UULF<channel_id without the UC prefix>
```
So `UCabc123...` → `https://www.youtube.com/feeds/videos.xml?playlist_id=UULFabc123...`
This does most of the Shorts/livestream filtering for free, at the cheapest possible
point in the pipeline.
Related prefixes, for reference: `UU` all uploads, `UUSH` shorts, `UULV` livestreams,
`UUMF`/`UUMO` members-only.
### Robustness
These prefixes are undocumented and there have been reports through 2026 of intermittent
failures and missing entries in YouTube's native feeds. So:
- If the `UULF` feed 404s or returns zero entries, **fall back** to
`?channel_id=UC...` and record `discovery_source = 'uc_feed'` on the resulting rows.
- Rows discovered via the fallback path get the duration/live filter applied at download
time (§6). Rows from `UULF` don't need it.
- Track consecutive poll failures per channel and surface the count in the admin UI. Two
channels quietly failing for a month is the realistic failure mode here.
### Parsing
`xml.etree.ElementTree`, no dependencies. Namespaces:
```python
NS = {
"atom": "http://www.w3.org/2005/Atom",
"yt": "http://www.youtube.com/xml/schemas/2015",
"media": "http://search.yahoo.com/mrss/",
}
```
Per `atom:entry`: `yt:videoId`, `atom:title`, `atom:published`,
`media:group/media:description`. The feed does **not** carry duration.
### Backfill on subscribe
New subscriptions pull the last **7 days**. The RSS feed only returns ~15 items, which
can be under 7 days for a prolific channel. For the initial backfill only, use:
```
yt-dlp --flat-playlist --playlist-end 50 -J \
'https://www.youtube.com/playlist?list=UULF<...>'
```
and date-filter the entries client-side. Steady-state polling uses the RSS feed.
### Getting the subscription list in
There is no need for a Google account or a subscriptions RSS feed. Either the brother
pastes channel URLs into the admin page one at a time, or he does a one-off Google
Takeout export (`subscriptions.csv`) and we bulk-insert. With single-digit channels,
manual entry is fine — do not build a CSV importer unless asked.
---
## 5. On-disk layout
Media root: `/disks/Plex/YouTubeSubs/`
```
/disks/Plex/YouTubeSubs/
├── .work/ # scratch, ignored by Jellyfin
└── Some Channel/
├── tvshow.nfo
├── poster.jpg
├── fanart.jpg
└── Season 2026/
├── Some Channel - S2026E08110 - Video Title [dQw4w9WgXcQ].mp4
├── Some Channel - S2026E08110 - Video Title [dQw4w9WgXcQ].nfo
├── Some Channel - S2026E08110 - Video Title [dQw4w9WgXcQ]-thumb.jpg
├── Some Channel - S2026E08110 - Video Title [dQw4w9WgXcQ].en.srt
└── Some Channel - S2026E08110 - Video Title [dQw4w9WgXcQ].info.json
```
### Season / episode numbering
- **Season** = upload year (`2026`)
- **Episode** = `MMDD * 10 + ordinal_within_day`
So the first video uploaded on 11 August is `8110`, the second that day is `8111`, and
the first on 12 August is `8120`. This scheme sorts correctly across the whole year
(unlike appending a suffix to a bare `MMDD`) and tolerates up to 10 uploads per channel
per day. Clamp the ordinal at 9 and log a warning if exceeded.
The ordinal must be computed against **what's already in the DB for that channel+date**,
not against the current batch, so it's stable across runs.
### Filename sanitisation
Strip `/ \ : * ? " < > |`, collapse runs of whitespace, strip leading/trailing dots and
spaces, truncate the title component to 120 chars on a word boundary. The `[videoid]`
suffix guarantees uniqueness regardless.
Store the channel directory name in the DB (`channel.dir_name`) at subscribe time and
never recompute it — channels rename themselves and we don't want orphaned directories.
---
## 6. Download
### Format selection — read this bit carefully
susan has no AVX and cannot realistically transcode. YouTube serves 720p as VP9 or AV1 by
default. If a client can't direct-play those, Jellyfin will try to transcode and it will
be miserable. So we force h264 + AAC in mp4 at download time:
```
-f "bv*[height<=720]+ba/b[height<=720]"
-S "vcodec:h264,acodec:aac,res:720"
--merge-output-format mp4
```
If a video genuinely has no h264 rendition at ≤720p, `-S` will fall back to VP9 rather
than fail. Log those cases; they should be rare.
Do **not** use `--sponsorblock-remove` — it requires cutting and re-encoding. Use
`--sponsorblock-mark all` with `--embed-chapters`, which writes chapter markers only and
is free.
### Full invocation
```
yt-dlp \
--extractor-args "youtube:player_client=default,mweb" \
--extractor-args "youtubepot-bgutilhttp:base_url=http://127.0.0.1:4416" \
-f "bv*[height<=720]+ba/b[height<=720]" \
-S "vcodec:h264,acodec:aac,res:720" \
--merge-output-format mp4 \
--no-playlist \
--write-info-json \
--write-thumbnail --convert-thumbnails jpg \
--write-auto-subs --sub-langs "en.*" --convert-subs srt \
--sponsorblock-mark all --embed-chapters \
--retries 3 --fragment-retries 10 \
--sleep-requests 2 --sleep-interval 5 --max-sleep-interval 15 \
-P "/disks/Plex/YouTubeSubs/.work" \
-o "%(id)s.%(ext)s" \
"https://www.youtube.com/watch?v=<VIDEO_ID>"
```
Verify the `youtubepot-bgutilhttp:base_url` arg name against the current plugin README —
this is the kind of thing that gets renamed.
### Match filter (fallback-discovered videos only)
For rows with `discovery_source = 'uc_feed'`, append:
```
--match-filter "duration>?{min_duration_seconds} & live_status!=is_live & live_status!=is_upcoming & !was_live"
```
The `>?` form allows videos with unknown duration through rather than rejecting them.
`min_duration_seconds` defaults to 120 (Shorts can now run to 3 minutes, but so can
legitimate short videos — this is the tradeoff, and it's why `UULF` is the primary path).
A video rejected by the match filter should be recorded as `skipped_short` /
`skipped_live` and **never retried**.
### Concurrency and politeness
One download at a time. Single-digit channels over a 14-day window is a small workload;
there is no reason to be aggressive and every reason not to be.
### After a successful download
1. Read the `.info.json` from `.work/`
2. Compute season/episode/filename
3. Generate the episode `.nfo` (§7)
4. `os.rename()` all artefacts into the season directory
5. Update the DB row: `state='downloaded'`, `rel_path`, `downloaded_at`
6. After the whole batch, trigger a Jellyfin library refresh:
`POST {jellyfin_url}/Library/Refresh` with header `X-Emby-Token: {api_key}`
If any step 14 fails, clean up `.work/` for that video ID and mark `failed` with
`attempts += 1`. Give up after 5 attempts and surface it in the UI.
---
## 7. Metadata — NFO files, no Jellyfin plugin
Jellyfin reads Kodi-style NFO sidecars natively. Configure the library as **Shows**,
disable all internet metadata providers, enable *Prefer local metadata* and *Save
artwork/metadata into media folders*. Do not write or install a metadata provider plugin.
### `tvshow.nfo` (per channel, written at subscribe time, refreshed on title change)
```xml
<?xml version="1.0" encoding="utf-8"?>
<tvshow>
<title>Some Channel</title>
<plot>Channel description from yt-dlp.</plot>
<studio>YouTube</studio>
<uniqueid type="youtube" default="true">UCabc123...</uniqueid>
</tvshow>
```
### Episode `.nfo` (one per video, filename matches the media file)
```xml
<?xml version="1.0" encoding="utf-8"?>
<episodedetails>
<title>Video Title</title>
<showtitle>Some Channel</showtitle>
<season>2026</season>
<episode>8110</episode>
<plot>Video description.</plot>
<aired>2026-08-11</aired>
<runtime>12</runtime>
<studio>YouTube</studio>
<uniqueid type="youtube" default="true">dQw4w9WgXcQ</uniqueid>
</episodedetails>
```
`runtime` is in **minutes**. Escape all text content properly — video descriptions
contain everything.
### Artwork
At subscribe time, `yt-dlp --flat-playlist --playlist-items 0 -J <channel_url>` returns a
`thumbnails` array containing entries with `id` values like `avatar_uncropped` and
`banner_uncropped`. Download the avatar to `poster.jpg` and the banner to `fanart.jpg`.
Treat both as best-effort — if they're missing, carry on without them.
Per-episode thumbnails come from `--write-thumbnail`; rename to `<basename>-thumb.jpg`.
---
## 8. Data model
SQLite at `/var/lib/ytsubs/subs.db`. **Enable WAL mode** — the cron job and the web
server both write.
```sql
PRAGMA journal_mode = WAL;
PRAGMA foreign_keys = ON;
CREATE TABLE channel (
id INTEGER PRIMARY KEY,
channel_id TEXT NOT NULL UNIQUE, -- UC...
handle TEXT, -- @handle, informational
title TEXT NOT NULL,
description TEXT,
dir_name TEXT NOT NULL UNIQUE, -- sanitised, immutable after creation
added_at TEXT NOT NULL,
backfilled INTEGER NOT NULL DEFAULT 0,
last_polled_at TEXT,
last_poll_ok INTEGER,
consecutive_poll_failures INTEGER NOT NULL DEFAULT 0
);
CREATE TABLE video (
id INTEGER PRIMARY KEY,
video_id TEXT NOT NULL UNIQUE,
channel_pk INTEGER NOT NULL REFERENCES channel(id) ON DELETE CASCADE,
title TEXT,
upload_date TEXT, -- YYYY-MM-DD
duration INTEGER,
season INTEGER,
episode INTEGER,
state TEXT NOT NULL,
discovery_source TEXT NOT NULL, -- 'uulf_feed' | 'uc_feed' | 'backfill'
rel_path TEXT, -- relative to media root, NULL unless downloaded
size_bytes INTEGER,
attempts INTEGER NOT NULL DEFAULT 0,
last_error TEXT,
discovered_at TEXT NOT NULL,
downloaded_at TEXT,
deleted_at TEXT
);
CREATE INDEX idx_video_state ON video(state);
CREATE INDEX idx_video_upload_date ON video(upload_date);
CREATE TABLE setting (
key TEXT PRIMARY KEY,
value TEXT NOT NULL
);
```
### `video.state` values
| State | Meaning | Retried? |
|---|---|---|
| `pending` | discovered, queued | — |
| `downloading` | claimed by a worker | recovered on startup → `pending` |
| `downloaded` | on disk, `rel_path` set | — |
| `deleted` | aged out; tombstone prevents re-download | never |
| `skipped_short` | rejected by match filter | never |
| `skipped_live` | rejected by match filter | never |
| `skipped_old` | discovered but already outside the window | never |
| `failed` | download error | yes, up to 5 attempts |
The tombstone behaviour is the important part: a `deleted` row must never be
re-discovered and re-downloaded. This is why we keep our own state table rather than
using yt-dlp's `--download-archive`.
### Settings (all editable from the admin page)
Store as strings; provide typed accessors with defaults so a missing key never crashes.
| Key | Default | Notes |
|---|---|---|
| `retention_days` | `14` | delete videos with `upload_date` older than this |
| `backfill_days` | `7` | window pulled when a channel is added |
| `max_height` | `720` | |
| `min_duration_seconds` | `120` | fallback-path Shorts filter |
| `sponsorblock_mark` | `true` | |
| `write_subs` | `true` | |
| `sub_langs` | `en.*` | |
| `respect_jellyfin_watch_state` | `true` | §10 |
| `jellyfin_url` | `http://127.0.0.1:8096` | |
| `jellyfin_api_key` | *(empty)* | masked in the UI |
| `jellyfin_user_id` | *(empty)* | |
| `pot_provider_url` | `http://127.0.0.1:4416` | |
| `max_attempts` | `5` | |
| `disk_cap_gb` | `0` | `0` = disabled; if set, evict oldest downloaded videos |
Validate on save (integers parse, URLs well-formed) and re-render the form with an inline
error rather than 500-ing.
---
## 9. Entry point and scheduling
Single executable `ytsubs` with subcommands:
| Command | Purpose |
|---|---|
| `ytsubs run` | poll → download → reap. This is what cron calls. |
| `ytsubs poll [--channel ID]` | discovery only |
| `ytsubs download` | drain the pending queue |
| `ytsubs reap` | retention pass |
| `ytsubs serve` | admin HTTP server (systemd unit) |
| `ytsubs subscribe <url>` / `unsubscribe <id>` | CLI equivalents, useful for debugging |
| `ytsubs doctor` | check yt-dlp version, POT provider reachable, DB writable, media root writable, Jellyfin reachable |
`run` takes a **non-blocking `flock`** on `/var/lib/ytsubs/run.lock` and exits 0 silently
if already held. On startup, reset any `downloading` rows to `pending` (crash recovery).
### Cron
```
17 * * * * runitor -uuid <hc-uuid> -- /usr/local/bin/ytsubs run
```
Hourly is ample for single-digit channels. Register the check in Healthchecks with a
generous grace period.
### systemd unit for the admin server
Standard `simple` unit, `Restart=always`, running as the same unprivileged user that owns
the DB and media root. When the user subscribes to a channel via the web UI, the handler
should spawn `ytsubs run --channel <id>` detached so the backfill starts immediately
rather than waiting up to an hour.
### nginx
Bind the admin server to `127.0.0.1:8085`. Front it with nginx on the **tailnet
interface only** — `yt.jihakuz.xyz` resolving inside Headscale, HTTP basic auth on top.
Given the brother will be on the tailnet anyway, there's no reason to expose this
publicly, and it removes the need for a Let's Encrypt cert and any real CSRF story.
Confirm this before implementing; if it does need to be public, add proper CSRF tokens.
---
## 10. Retention and deletion
### Aging out (`ytsubs reap`)
Candidates: `state = 'downloaded'` and `upload_date < today - retention_days`.
Before deleting, if `respect_jellyfin_watch_state` is on:
1. `GET {jellyfin_url}/Items?userId={user_id}&recursive=true&includeItemTypes=Episode&fields=Path,UserData`
with header `X-Emby-Token: {api_key}` — one call per reap, build a `path → UserData` map.
2. **Skip** any candidate where `UserData.PlaybackPositionTicks > 0 and not UserData.Played`
(part-watched) or `UserData.IsFavorite` is true.
3. Apply a hard backstop: delete anyway once `upload_date` is older than
`retention_days * 2`, so a half-watched video doesn't live forever.
Nothing wrecks "set and forget" like a video disappearing at the 20-minute mark. This
check is worth the complexity.
Deletion removes the media file plus its `.nfo`, `-thumb.jpg`, `.srt`, and `.info.json`
siblings, then prunes the season directory if empty. Set `state='deleted'`,
`deleted_at`, and `rel_path=NULL`. **Keep the row** — it's the tombstone.
If the Jellyfin API is unreachable, skip the whole reap for that run and log it. Do not
delete blindly.
### Unsubscribe
Hard delete, as agreed: `shutil.rmtree()` the channel directory, then `DELETE FROM
channel` (cascades to `video`). Re-subscribing later starts from scratch and re-downloads
the last 7 days. Put a confirmation step in the UI — this is destructive and irreversible.
### Disk cap (optional, `disk_cap_gb`)
If enabled, after each reap, sum `size_bytes` over `downloaded` rows and evict oldest-by-
`upload_date` until under the cap, honouring the same watch-state protection.
---
## 11. Admin UI
Stdlib only — `http.server.ThreadingHTTPServer` + `BaseHTTPRequestHandler`. No Flask, no
FastAPI, no npm. Server-rendered HTML with a single embedded `<style>` block. The only
JavaScript should be a `confirm()` on the unsubscribe button.
### Routes
| Method | Path | Behaviour |
|---|---|---|
| `GET` | `/` | Channel list + add form + settings form |
| `POST` | `/channels` | Resolve URL → insert → spawn backfill → redirect to `/` |
| `POST` | `/channels/<id>/delete` | Confirm-guarded hard delete → redirect |
| `POST` | `/settings` | Validate + persist → redirect |
| `GET` | `/health` | JSON: yt-dlp version, POT provider up, last run time, queue depth |
All POSTs redirect (303) so refresh doesn't resubmit.
### Channel list should show, per channel
Title, `@handle`, video count on disk, disk usage, most recent upload date, last poll
time, and a warning badge if `consecutive_poll_failures > 2`.
### Channel URL resolution
Accept: `https://www.youtube.com/@handle`, `/channel/UC...`, `/c/name`, `/user/name`,
a bare `@handle`, or a bare `UC...` ID.
```
yt-dlp --flat-playlist --playlist-items 0 -J <url>
```
`--playlist-items 0` fetches channel metadata **without enumerating the uploads**, which
is the cheap way to do this. Pull `channel_id`, `channel`, `description`, `thumbnails`
from the result. If resolution fails or returns no `channel_id`, re-render the form with
an error — don't insert a half-formed row.
Reject duplicates on `channel_id` (not on the submitted URL — the same channel has many
URL forms).
---
## 12. Known gotchas
- **yt-dlp churn is the top operational risk.** Extractor args, client names, and PO
token requirements change frequently. Weekly `pip install -U yt-dlp` under runitor.
When something breaks, check the yt-dlp issue tracker before debugging our code.
- **The `UULF` prefix is undocumented** and could vanish. The `channel_id` fallback path
is not optional.
- **RSS feeds cap at ~15 entries.** Fine for hourly polling, insufficient for backfill —
hence the separate `--flat-playlist` path in §4.
- **`.work` must share a filesystem with the media root**, or every completed download
becomes a full copy.
- **Don't transcode.** If Jellyfin is transcoding these files, the format selection in §6
is wrong — fix it there, not by throwing hardware at it.
- **Descriptions are hostile input** for XML generation. Use `xml.etree`'s serialiser
rather than string-formatting the NFO by hand.
- **WAL mode is required** — two writers.
- Don't reuse `/disks/Plex/YouTube` (TubeArchivist's tree).
---
## 13. Acceptance criteria
Work through these in order; each is a real check, not a code-reading exercise.
1. `ytsubs doctor` passes on a clean install.
2. Subscribing to a real, active channel creates the directory, `tvshow.nfo`,
`poster.jpg`, and queues the last 7 days of videos.
3. At least one real video downloads end to end and lands in the right season directory
with a correct filename, NFO, thumbnail, and subtitle sidecar.
4. `ffprobe` on the downloaded file shows **h264 video and AAC audio**.
5. Jellyfin, after a library scan, shows the channel as a show, the year as a season, and
the video as an episode with the correct title, description, and air date — with all
internet metadata providers disabled.
6. The video **direct-plays** on a client with no transcoding (check the Jellyfin
dashboard's active-streams panel).
7. Re-running `ytsubs run` immediately downloads nothing and errors on nothing
(idempotency).
8. Manually backdating a video's `upload_date` past the retention window causes `reap` to
delete it and its sidecars, prune the empty directory, and leave a `deleted` tombstone.
9. Re-running `poll` after that does **not** re-download the deleted video.
10. Marking a video part-watched in Jellyfin protects it from reap (with
`respect_jellyfin_watch_state` on).
11. Unsubscribing removes the directory and all rows.
12. Killing `ytsubs run` mid-download leaves no orphan in `.work`, and the next run
recovers the `downloading` row to `pending`.
13. The admin page renders, adds, removes, and persists settings; a bad settings value
produces an inline error rather than a traceback.
14. Editing `retention_days` in the UI visibly changes reap behaviour on the next run.
---
## 14. Suggested build order
!! tom: we also wants tests through this
1. Skeleton: DB schema, migrations, settings accessors, `doctor`
2. POT provider + a hardcoded single-video download proving §3 and §6 work — **do this
before writing anything else of substance**
3. Channel resolution + `subscribe`/`unsubscribe` on the CLI
4. Discovery (`poll`), both feed paths
5. Download worker + NFO/artwork generation + move-into-place
6. Jellyfin library verification (acceptance criteria 5 and 6) — a checkpoint, not a step
7. `reap`, including the watch-state check
8. `run` orchestration, flock, crash recovery, cron + runitor + Healthchecks
9. Admin server
10. systemd unit, nginx on the tailnet, deploy scripts
Steps 2 and 6 are the two places where this design could turn out to be wrong. Hit them
early and report back rather than building on top of an unverified assumption.