Files
youtube-automate/specs.md
T
Tom FluxandClaude Opus 5 18bb2e420b Initial implementation of youtube-automate
A DVR for YouTube subscriptions, laid out so Jellyfin presents each channel
as a show and each video as an episode. Cron-driven, idempotent, with a
public admin UI for a non-operator.

Verified end to end on susan against three real channels: PO tokens, h264
downloads, Jellyfin resolution from local NFOs with all providers disabled,
retention and tombstones.

Corrections to the original design handover (specs.md documents each with
the evidence, and specs.handover-original.md preserves the original):

- The format sort selected 360p. Ranking acodec above res makes `bv*` prefer
  the combined 360p stream, which carries AAC, over the 720p video-only
  stream whose acodec is none. vcodec now leads, so a video without h264 at
  720p yields h264 lower down rather than VP9 this hardware cannot transcode.
- yt-dlp now requires a JS runtime and the yt-dlp-ejs solver scripts, which
  only ship with the [default] extra. Without them the n challenge fails and
  the mweb formats disappear entirely.
- --flat-playlist carries no upload dates, so the specced client-side date
  filter for backfill was impossible. Backfill is RSS-first.
- skipped_old was terminal, so raising a channel's retention appeared to do
  nothing. Added an explicit rescan.
- is_upcoming premieres now defer and retry instead of being skipped forever.
- TubeArchivist is gone, so the media root and the tube.jihakuz.xyz vhost
  were both reclaimed; the latter still pointed at its dead port.

240 offline tests, no network and no real yt-dlp invocation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 21:42:48 +01:00

48 KiB
Raw Blame History

youtube-automate — implementation spec

Target machine: susan Status: design agreed and verified against the live machine on 2026-08-11. Ready to build. Supersedes: specs.handover-original.md (the original design handover). Section 15 lists every change and why.


1. What we're building

A self-hosted pipeline that keeps a rolling retention window of recent YouTube uploads from a handful of subscribed channels, laid out on disk so Jellyfin presents each channel as a TV show and each video as an episode. A small web UI lets a non-operator (the requester's brother) subscribe, unsubscribe, and tweak settings.

The mental model is a DVR for YouTube subscriptions, not an archiver. Videos are disposable. If a video ages out or a channel is unsubscribed, it goes away permanently.

Explicitly in scope

  • Poll subscribed channels for new uploads
  • Download at ≤720p, h264/AAC, into a Jellyfin-friendly layout
  • Generate Kodi-style .nfo metadata + artwork (no Jellyfin plugin)
  • Delete videos older than the retention window
  • Web admin over public HTTPS: add/remove channels, edit settings
  • Cron-driven, idempotent, monitored via Healthchecks
  • Offline unit tests

Explicitly out of scope — do not build these

  • Cookie handling of any kind. No browser extension, no cookie-receiving API, no cookies.txt. See §3.
  • Members-only, age-restricted, or private content (follows from the above)
  • Transcoding of any kind (see §6 — this matters on this hardware)
  • Multi-user support. One shared credential, one admin UI.
  • Jellyfin watch-state protection (declined; see §10 and §15)
  • CSV / Takeout importers

2. Environment facts

All of the following was verified on the live machine on 2026-08-11.

  • susan: Debian 12, kernel 6.1.0-41. Dual Xeon X5675 (Westmere), 2×6 cores = 12 threads, 47 GB RAM. No AVX — confirmed absent from /proc/cpuinfo. Software video transcoding is effectively off the table.
  • Jellyfin 10.11.4, running as a native systemd service (User=jellyfin, Group=jellyfin, UMask=0022) on 0.0.0.0:8096, enabled at boot. Not containerised.
  • Docker running and enabled at boot, so --restart unless-stopped containers do come back after a reboot. No extra unit needed for the POT provider.
  • Healthchecks container on 127.0.0.1:8001, published at hc.jihakuz.xyz.
  • nginx active, Let's Encrypt via certbot, DNS at njal.la driven by ~/.local/bin/update-dns.sh (dynamic-DNS update keys, run every 15 min from cron).
  • runitor v1.4.1 at /usr/local/bin/runitor. Cron convention: entries live in susan's crontab with HC_API_URL=https://hc.jihakuz.xyz/ping set at the top and one UUID per job.
  • Tailscale 1.98.2, susan is 100.64.0.2. Headscale runs on the Linode, not here.
  • susan is on a residential IP. This matters — YouTube flags datacenter IPs far more aggressively. Do not move any part of this to victoria (the Linode VPS).
  • Toolchain: uv 0.11.23, Python 3.11.2, ffmpeg/ffprobe 5.1.7, sqlite 3.40.1, flock 2.38.1.
  • /usr/local/bin/yt-dlp is version 2023.11.16 — a standalone binary from Dec 2023. It is far too old to work against YouTube today. Nothing on the machine references it (checked /usr/local/bin, /var/lib/radio, ~/.local/bin, and all crontabs), so it is left alone and we use our own venv copy instead.
  • susan has no passwordless sudo. Anything touching /usr/local/bin, /etc/systemd/system or /etc/nginx must be run by the operator via deploy/deploy.sh.

TubeArchivist is gone

The original handover said /disks/Plex/YouTube was in use by TubeArchivist and must not be touched. This is no longer true. Verified: no TubeArchivist container exists (running or stopped), the directory is empty, and no Jellyfin library references it — the libraries are Audiobooks, Anime, Recordings, TV Shows, Music and Films. tube.jihakuz.xyz still has a DNS record in update-dns.sh but nothing serves it, so we reuse that hostname for the admin UI.

Paths

Purpose Path
Source (bare repo) /disks/git-repos/youtube-automate.git
Checkout /opt/youtube-automate
Entry point /usr/local/bin/youtube-automate
Virtualenv /var/lib/youtube-automate/venv
Deno (JS runtime, §3) /var/lib/youtube-automate/venv/bin/deno
State DB /var/lib/youtube-automate/subs.db
Lock file /var/lib/youtube-automate/run.lock
Media root /disks/Plex/YouTube/
Scratch/work dir /disks/Plex/YouTube/.work/

/disks is a single 5.5 TB ext4 volume (/dev/sdb2) with 1.3 TB free, so the work dir and the media root share a filesystem and finished downloads move into place with an atomic rename().

Estimated steady-state usage for the three test channels is 510 GB. Disk is a non-issue.

Ownership and permissions — do not skip this

The media tree convention on susan is susan:mediaserver, directories 0770, files 0664. /disks/Plex is 0770 with no setgid bit. Jellyfin reaches the tree only via its mediaserver supplementary group.

If the service runs as susan with the default umask 0022, new directories land susan:susan and only work by the o+rx bit — which breaks the moment the unit is hardened. Therefore:

  • systemd unit runs User=susan, Group=mediaserver, UMask=0002
  • the media root gets the setgid bit so new directories inherit mediaserver
  • cron entries that call the CLI directly go through sg mediaserver "...", matching the existing radio entries in susan's crontab

3. YouTube access: PO tokens, not cookies

Verified against the yt-dlp wiki on 2026-08-11: the current recommendation is still the mweb client plus a PO Token Provider plugin supplying tokens for GVS (Google Video Server) requests. No account, no cookies, no browser. Cookies only unlock age-restricted / members-only / private content, which is out of scope, and carry a real risk of the Google account being banned.

Setup

  1. Run the provider as a Docker sidecar. Docker is enabled at boot and unless-stopped survives a reboot, so no systemd unit is required:

    docker run --name bgutil-provider -d --restart unless-stopped --init \
      -p 127.0.0.1:4416:4416 brainicism/bgutil-ytdlp-pot-provider:1.3.1-deno
    

    Pin the tag. Floating latest on both halves is how you get a silent version skew.

  2. Build the venv. All four of these are required — see "JS challenge solving" below for why the last two are not optional:

    uv venv --python 3.11 /var/lib/youtube-automate/venv
    VIRTUAL_ENV=/var/lib/youtube-automate/venv uv pip install \
      'yt-dlp[default]' \
      'bgutil-ytdlp-pot-provider==1.3.1' \
      'curl-cffi<0.16'
    
    • yt-dlp[default] — the [default] extra is what pulls in yt-dlp-ejs, the challenge solver scripts. Plain pip install yt-dlp does not include it.
    • bgutil-ytdlp-pot-provider==1.3.1 — must match the container tag.
    • curl-cffi<0.16 — enables request impersonation. yt-dlp 2026.7.4 accepts 0.5.10 or 0.10.x0.15.x only; the current release (0.16.0) is rejected as "unsupported", and an unpinned install silently gets you that. Verify with yt-dlp --list-impersonate-targets, which must list Chrome/Safari targets rather than only (unavailable) lines.
  3. Install a JavaScript runtime. Deno is the recommended runtime and there is none on susan. Install it inside the venv's bin/ so it needs no root and no PATH changes beyond the venv directory the CLI already uses:

    curl -fsSL -o deno.zip \
      https://github.com/denoland/deno/releases/latest/download/deno-x86_64-unknown-linux-gnu.zip
    unzip -q deno.zip && install -m 0755 deno /var/lib/youtube-automate/venv/bin/deno
    

    Deno 2.9.5 (V8 15.0) runs fine on Westmere despite the absent AVX — verified. Minimum supported version is 2.3.0. yt-dlp locates it on PATH, so every yt-dlp invocation must run with /var/lib/youtube-automate/venv/bin prepended to PATH.

JS challenge solving is mandatory — this is new since the original handover

Without a JS runtime and yt-dlp-ejs, yt-dlp emits:

WARNING: n challenge solving failed: Some formats may be missing. Ensure you have a
supported JavaScript runtime and challenge solver script distribution installed.

That's the nsig throttling signature. Unsolved, formats go missing and downloads can be throttled. Measured on susan: without the runtime, 22 formats and no mweb formats at all (everything came from the android client); with Deno + yt-dlp-ejs, 29 formats including the 12 mweb ones, and a 44 MB download completed at 7.7 MiB/s with no throttling.

Confirm the good state in -v output — all three lines must be present:

[debug] JS runtimes: deno-2.9.5
[debug] [youtube] [jsc] JS Challenge Providers: ... deno ...
[youtube] [jsc:deno] Solving JS challenges using deno

A benign line you can ignore: [pot:bgutil:script-deno] Script path doesn't exist: /home/susan/bgutil-ytdlp-pot-provider/server/.... That is the script variant of the bgutil plugin probing for a local checkout we deliberately don't have; the http variant is the one in use and it reports as available.

Verification

Before writing any other code:

PATH=/var/lib/youtube-automate/venv/bin:$PATH yt-dlp -v '<any watch URL>' -F

You must see:

[debug] [youtube] [pot] PO Token Providers: bgutil:http-1.3.1 (external), ...
[youtube] [pot:bgutil:http] Generating a gvs PO Token for mweb client via bgutil HTTP server
[debug] [youtube] <id>: Retrieved a gvs PO Token for mweb client

If you don't, stop and fix this first — everything downstream depends on it.

doctor and every run health-check the provider with a plain GET against http://127.0.0.1:4416/ping before a download batch, and fail loudly if it's down rather than silently accumulating 403s. The endpoint is confirmed to exist and returns {"server_uptime": <n>, "version": "<s>"}doctor should also compare that version against the installed plugin version and warn on a mismatch.

The exact extractor-arg key youtubepot-bgutilhttp:base_url was confirmed against the current plugin README.

Status: proven on susan on 2026-08-11. Video cBwBMyJ_bxU downloaded end to end and ffprobe reports h264 (Main) 1280x720 @60fps + aac (LC) stereo in an mp4 container, with 3 SponsorBlock chapter marks embedded and no re-encoding. Files landed susan:mediaserver 0664 via the setgid bit and umask 0002.

Caveats

  • The bgutil README warns that PO tokens no longer bypass the "Sign in to confirm you're not a bot" interstitial in most cases. That is a different failure mode (IP reputation) from the 403s on format URLs that this solves. Residential IP + low request volume + sleep intervals should keep us clear of it. If it starts firing, that's the point at which cookies get reconsidered — not before.
  • yt-dlp's YouTube handling churns every few weeks. Client recommendations, extractor-arg names, and which clients need tokens all move. Treat the flags in §6 as the shape of the answer, not as gospel, and re-read the wiki when something breaks.
  • Keep yt-dlp updated weekly under runitor (§9). A stale yt-dlp is the single most likely cause of "everything broke" — the 2023 binary already on this box is the cautionary example.

4. Discovery: RSS feeds

YouTube publishes an unauthenticated Atom feed per channel. Use the undocumented UULF playlist variant, which returns long-form videos only — no Shorts, no livestreams:

https://www.youtube.com/feeds/videos.xml?playlist_id=UULF<channel_id without the UC prefix>

So UCabc123...https://www.youtube.com/feeds/videos.xml?playlist_id=UULFabc123...

Related prefixes: UU all uploads, UUSH shorts, UULV livestreams, UUMF/UUMO members-only.

The filtering is real — verified

Tested on 2026-08-11 against the three subscription targets:

Channel UULF UUSH UULV Shorts leaking into UULF
clabretro 200, 15 entries 200, 2 entries 200, 1 entry 0
EthosLab 200, 15 entries 404 404 n/a
DolanDarkest 200, 15 entries 404 404 n/a

clabretro is the useful case: it genuinely has 2 Shorts and 1 livestream, and none of them appear in the UULF feed. The primary discovery path does what it claims.

Note also that UUSH/UULV return HTTP 404 when a channel has none — that means "empty", not "broken". Only a UULF 404 (or an empty UULF result) triggers the fallback.

Robustness

These prefixes are undocumented and there have been reports of intermittent failures and missing entries in YouTube's native feeds. So:

  • If the UULF feed 404s or returns zero entries, fall back to ?channel_id=UC... and record discovery_source = 'uc_feed' on the resulting rows.
  • Rows discovered via the fallback path get the duration/live filter applied at download time (§6). Rows from UULF don't need it.
  • A skipped_short row from the fallback path is re-queued if a later UULF poll lists that video. UULF is authoritative about what is and isn't a Short, so this repairs the case where a transient feed outage would otherwise permanently drop a legitimate short-but-not-Short video. See §6 for why this matters concretely.
  • Track consecutive poll failures per channel and surface the count in the admin UI. Two channels quietly failing for a month is the realistic failure mode here.

Parsing

xml.etree.ElementTree, no dependencies. Namespaces:

NS = {
    "atom":  "http://www.w3.org/2005/Atom",
    "yt":    "http://www.youtube.com/xml/schemas/2015",
    "media": "http://search.yahoo.com/mrss/",
}

Per atom:entry: yt:videoId, atom:title, atom:published, media:group/media:description. The feed does not carry duration.

Backfill on subscribe

New subscriptions pull the last backfill_days (default 7).

--flat-playlist carries no upload dates. The original handover said to pull 50 entries with --flat-playlist -J and "date-filter the entries client-side". Verified against clabretro: every entry comes back with timestamp: None. There is no date to filter on. Entries do carry id, title, duration and live_status.

So the backfill is RSS-first:

  1. Read the UULF feed (falling back to channel_id), which is the only source that carries dates. It returns ~15 entries — enough to cover a 7-day window for any channel uploading less than twice a day, which is all three subscription targets.
  2. Only if the feed's oldest entry is still inside the backfill window might there be more, in which case extend via --flat-playlist --playlist-end 50 and resolve those extra dates one video at a time with --print "%(upload_date)s". The playlist is reverse-chronological, so stop at the first video outside the window.

In practice step 2 never runs for the current subscriptions. It exists so that raising backfill_days doesn't silently truncate at 15 videos.

skipped_old and rescan

A video discovered outside the window is skipped_old, which is terminal. That creates a trap: raising a channel's retention_days afterwards appears to do nothing, because every one of that channel's videos is already marked old. Observed with EthosLab — his most recent upload was 22 days before subscribe, so a 7-day backfill queued nothing at all for him.

rescan re-queues skipped_old rows whose upload_date now falls inside the channel's current effective retention window. It is an explicit action (poll --rescan, or a button in the admin UI), deliberately not something poll does silently — doing it on every poll would make backfill_days meaningless, since it would immediately re-queue everything the initial backfill had just decided to leave behind.

Tombstones are never touched: only skipped_old is eligible, never deleted.

Verified: setting EthosLab to 60 days and rescanning queued exactly the 3 videos inside that window and left the other 12 alone.

Getting the subscription list in

Manual entry via the admin page, one channel URL at a time. With single-digit channels that is enough. No CSV importer.


5. On-disk layout

Media root: /disks/Plex/YouTube/

/disks/Plex/YouTube/
├── .work/                                   # scratch, ignored by Jellyfin
│   └── .ignore                              # belt-and-braces, see below
└── Some Channel/
    ├── tvshow.nfo
    ├── poster.jpg
    ├── fanart.jpg
    └── Season 2026/
        ├── Some Channel - S2026E8110 - Video Title [dQw4w9WgXcQ].mp4
        ├── Some Channel - S2026E8110 - Video Title [dQw4w9WgXcQ].nfo
        ├── Some Channel - S2026E8110 - Video Title [dQw4w9WgXcQ]-thumb.jpg
        ├── Some Channel - S2026E8110 - Video Title [dQw4w9WgXcQ].en.srt
        └── Some Channel - S2026E8110 - Video Title [dQw4w9WgXcQ].info.json

The work dir is dot-prefixed (which Jellyfin skips) and contains an .ignore file, which Jellyfin also honours. Two independent mechanisms, because relying on undocumented hidden-folder behaviour for the directory holding partial downloads is not worth the risk.

Season / episode numbering

  • Season = upload year (2026)
  • Episode = MMDD * 10 + ordinal_within_day

So the first video uploaded on 11 August is 8110, the second that day is 8111, and the first on 12 August is 8120. This sorts correctly across the whole year — 1 Jan is 1010 and 31 Dec is 12310 — and tolerates up to 10 uploads per channel per day. Clamp the ordinal at 9 and log a warning if exceeded.

The ordinal must be computed against what's already in the DB for that channel+date, not against the current batch, so it's stable across runs.

Write the episode number unpadded (E8110), matching the <episode> value in the NFO exactly.

Filename sanitisation

Strip / \ : * ? " < > |, collapse runs of whitespace, strip leading/trailing dots and spaces, truncate the title component to 120 chars on a word boundary. The [videoid] suffix guarantees uniqueness regardless.

Store the channel directory name in the DB (channel.dir_name) at subscribe time and never recompute it — channels rename themselves and we don't want orphaned directories.


6. Download

Format selection — read this bit carefully

susan has no AVX and cannot realistically transcode. YouTube serves 720p as VP9 or AV1 by default. If a client can't direct-play those, Jellyfin will try to transcode and it will be miserable. So we force h264 + AAC in mp4 at download time:

-f "bv*[height<=720]+ba/b[height<=720]"
-S "vcodec:h264,res:720,acodec:aac"
--merge-output-format mp4

The sort field order matters and the original handover had it wrong. With -S "vcodec:h264,acodec:aac,res:720"acodec ranked above res — this selects the legacy combined 360p format 18 rather than 720p. bv* means "best video, possibly with audio", so ranking acodec:aac before res makes a combined 360p stream (which has AAC) outrank a 720p video-only stream (whose acodec is none). Verified empirically: the original ordering yields 18 | 640x360, the corrected ordering yields 298+140 | 1280x720 | avc1 | mp4a.

Keep vcodec first, ahead of res. If a video has no h264 at 720p but does at 480p, this picks h264 at 480p — which direct-plays — rather than 720p VP9, which would force a transcode. On hardware that cannot transcode, codec fidelity beats resolution. Log any case that still falls back to VP9/AV1; they should be rare.

Do not use --sponsorblock-remove — it requires cutting and re-encoding. Use --sponsorblock-mark all with --embed-chapters, which writes chapter markers only and is free.

Full invocation

yt-dlp \
  --extractor-args "youtube:player_client=default,mweb" \
  --extractor-args "youtubepot-bgutilhttp:base_url=http://127.0.0.1:4416" \
  -f "bv*[height<=720]+ba/b[height<=720]" \
  -S "vcodec:h264,res:720,acodec:aac" \
  --merge-output-format mp4 \
  --no-playlist \
  --write-info-json \
  --write-thumbnail --convert-thumbnails jpg \
  --write-subs --write-auto-subs --sub-langs "en.*" --convert-subs srt \
  --sponsorblock-mark all --embed-chapters \
  --retries 3 --fragment-retries 10 \
  --sleep-requests 2 --sleep-interval 5 --max-sleep-interval 15 \
  -P "/disks/Plex/YouTube/.work" \
  -o "%(id)s.%(ext)s" \
  "https://www.youtube.com/watch?v=<VIDEO_ID>"

--write-subs is included alongside --write-auto-subs so channels with real uploaded subtitles aren't silently skipped in favour of nothing.

Subtitle de-duplication. --sub-langs "en.*" matches both en and en-orig, so a typical video yields two sidecars (<id>.en.srt and <id>.en-orig.srt) which Jellyfin would present as two identical "English" tracks. Observed on the test download. In the move-into-place step, keep .en.srt if present and discard .en-orig.srt; if only en-orig exists, rename it to .en.srt.

Match filter (fallback-discovered videos only)

For rows with discovery_source = 'uc_feed', append:

--match-filter "duration>?{min_duration_seconds} & live_status!=?is_live & live_status!=?is_upcoming & !was_live"

The >? and !=? forms allow videos with unknown values through rather than rejecting them — an absent live_status should not disqualify a video any more than an absent duration does.

min_duration_seconds defaults to 120.

Why 120 is a knowingly tight default

DolanDarkest — one of the three subscription targets — uploads daily videos of 122167 seconds. One of them clears the filter by two seconds. Those videos arrive via UULF, which is exempt from the match filter, so in normal operation this is fine.

The risk is a transient UULF outage: rows then get discovery_source='uc_feed', the duration filter applies, and anything under two minutes would be dropped. This is why §4 re-queues a skipped_short row when a later UULF poll lists the video. With that repair in place the only residual loss case is a sub-120s video whose UULF feed never recovers inside the retention window, which would also be visible in the poll-failure counter.

A per-channel min_duration_seconds override was considered and declined; revisit if DolanDarkest starts posting sub-two-minute videos regularly.

Skip and defer semantics

Outcome State Retried?
Below min_duration_seconds skipped_short Only if a later UULF poll lists it (§4)
is_live or was_live skipped_live Never — livestreams are not wanted
is_upcoming deferred Yes — re-queued by later polls

is_upcoming must not be a permanent skip. A scheduled premiere becomes an ordinary downloadable video once it airs, so permanently rejecting it would silently lose content from any channel that uses premieres. Give up on a deferred row once it falls outside the retention window.

Concurrency and politeness

One download at a time. Single-digit channels over a 14-day window is a small workload; there is no reason to be aggressive and every reason not to be.

After a successful download

  1. Read the .info.json from .work/
  2. Compute season/episode/filename
  3. Generate the episode .nfo (§7)
  4. os.rename() all artefacts into the season directory
  5. Update the DB row: state='downloaded', rel_path, downloaded_at, size_bytes
  6. After the whole batch, trigger a Jellyfin library refresh (§7)

size_bytes is required — disk_cap_gb cannot work without it.

If any step 14 fails, clean up .work/ for that video ID and mark failed with attempts += 1. Give up after max_attempts (5) and surface it in the UI.


7. Metadata — NFO files, no Jellyfin plugin

Jellyfin reads Kodi-style NFO sidecars natively. The library is created as Shows with all internet metadata providers disabled, Prefer local metadata on, and Save artwork/metadata into media folders on. Do not write or install a metadata provider plugin.

The library is created programmatically via POST /Library/VirtualFolders using the stored admin API key, so the setup is reproducible rather than a hand-clicked state.

tvshow.nfo (per channel, written at subscribe time, refreshed on title change)

<?xml version="1.0" encoding="utf-8"?>
<tvshow>
  <title>Some Channel</title>
  <plot>Channel description from yt-dlp.</plot>
  <studio>YouTube</studio>
  <uniqueid type="youtube" default="true">UCabc123...</uniqueid>
</tvshow>

Episode .nfo (one per video, filename matches the media file)

<?xml version="1.0" encoding="utf-8"?>
<episodedetails>
  <title>Video Title</title>
  <showtitle>Some Channel</showtitle>
  <season>2026</season>
  <episode>8110</episode>
  <plot>Video description.</plot>
  <aired>2026-08-11</aired>
  <runtime>12</runtime>
  <studio>YouTube</studio>
  <uniqueid type="youtube" default="true">dQw4w9WgXcQ</uniqueid>
</episodedetails>

runtime is in minutes. Build these with xml.etree's serialiser, never string formatting — video descriptions are hostile input and contain everything.

Artwork

At subscribe time, yt-dlp --flat-playlist --playlist-items 0 -J <channel_url> returns a thumbnails array containing entries with id values like avatar_uncropped and banner_uncropped. Download the avatar to poster.jpg and the banner to fanart.jpg. Treat both as best-effort — if they're missing, carry on without them.

Per-episode thumbnails come from --write-thumbnail; rename to <basename>-thumb.jpg.

Library refresh

POST {jellyfin_url}/Library/Refresh with header X-Emby-Token: {api_key}, triggered after:

  • a download batch that produced at least one new file
  • a reap that deleted at least one file
  • an unsubscribe

Refreshing after deletions matters — without it Jellyfin shows ghost episodes until its own scheduled scan.


8. Data model

SQLite at /var/lib/youtube-automate/subs.db. WAL mode — the cron job and the web server both write.

PRAGMA journal_mode = WAL;
PRAGMA foreign_keys = ON;

CREATE TABLE channel (
  id                INTEGER PRIMARY KEY,
  channel_id        TEXT NOT NULL UNIQUE,      -- UC...
  handle            TEXT,                      -- @handle, informational
  title             TEXT NOT NULL,
  description       TEXT,
  dir_name          TEXT NOT NULL UNIQUE,      -- sanitised, immutable after creation
  added_at          TEXT NOT NULL,
  backfilled        INTEGER NOT NULL DEFAULT 0,
  retention_days    INTEGER,                   -- NULL = use global setting
  last_polled_at    TEXT,
  last_poll_ok      INTEGER,
  consecutive_poll_failures INTEGER NOT NULL DEFAULT 0
);

CREATE TABLE video (
  id                INTEGER PRIMARY KEY,
  video_id          TEXT NOT NULL UNIQUE,
  channel_pk        INTEGER NOT NULL REFERENCES channel(id) ON DELETE CASCADE,
  title             TEXT,
  upload_date       TEXT,                      -- YYYY-MM-DD
  duration          INTEGER,
  season            INTEGER,
  episode           INTEGER,
  state             TEXT NOT NULL,
  discovery_source  TEXT NOT NULL,             -- 'uulf_feed' | 'uc_feed' | 'backfill'
  rel_path          TEXT,                      -- relative to media root, NULL unless downloaded
  size_bytes        INTEGER,
  attempts          INTEGER NOT NULL DEFAULT 0,
  last_error        TEXT,
  discovered_at     TEXT NOT NULL,
  downloaded_at     TEXT,
  deleted_at        TEXT
);

CREATE INDEX idx_video_state ON video(state);
CREATE INDEX idx_video_upload_date ON video(upload_date);
CREATE INDEX idx_video_channel ON video(channel_pk);

CREATE TABLE setting (
  key   TEXT PRIMARY KEY,
  value TEXT NOT NULL
);

Schema changes are applied by an idempotent migration function keyed on PRAGMA user_version.

video.state values

State Meaning Retried?
pending discovered, queued
downloading claimed by a worker recovered on startup → pending
downloaded on disk, rel_path set
deleted aged out; tombstone prevents re-download never
deferred premiere/upcoming, not yet available yes, on later polls
skipped_short rejected by match filter only if later seen in UULF (§4)
skipped_live rejected by match filter never
skipped_old discovered but already outside the window never
failed download error yes, up to max_attempts

The tombstone behaviour is the important part: a deleted row must never be re-discovered and re-downloaded, and the skipped_short repair in §4 must never resurrect one. This is why we keep our own state table rather than using yt-dlp's --download-archive.

Settings

Store as strings; provide typed accessors with defaults so a missing key never crashes.

Key Default Notes
retention_days 14 global default; per-channel override wins
backfill_days 7 window pulled when a channel is added
max_height 720
min_duration_seconds 120 fallback-path Shorts filter
sponsorblock_mark true
write_subs true
sub_langs en.*
jellyfin_url http://127.0.0.1:8096
jellyfin_api_key (empty) masked in the UI; set via CLI
pot_provider_url http://127.0.0.1:4416
max_attempts 5
disk_cap_gb 0 0 = disabled; if set, evict oldest downloaded videos

Validate on save (integers parse, URLs well-formed) and re-render the form with an inline error rather than 500-ing.

Two further keys are not editable through the settings form and are never rendered: admin_password_hash (see §11) and session_secret (generated on first run).

No secret is committed to the repo. The Jellyfin API key and the admin password are set with CLI subcommands and live only in the DB.


9. Entry point and scheduling

/usr/local/bin/youtube-automate is a shim that execs the venv interpreter against the package. Subcommands:

Command Purpose
run [--channel ID] poll → download → reap. This is what cron calls.
poll [--channel ID] discovery only
download drain the pending queue
reap retention pass
serve admin HTTP server (systemd unit)
subscribe <url> / unsubscribe <id> CLI equivalents, useful for debugging
set-password prompt for and store the admin password (never on the command line)
set-jellyfin-key store the Jellyfin API key
setup-jellyfin-library create/verify the YouTube Shows library with the right options
doctor yt-dlp version, POT provider reachable, DB writable, media root writable, Jellyfin reachable, permissions sane

run takes a non-blocking flock on /var/lib/youtube-automate/run.lock and exits 0 silently if already held. On startup it resets any downloading rows to pending and clears matching orphans out of .work/ (crash recovery).

run accepts --channel so the web subscribe handler can spawn an immediate backfill rather than waiting up to an hour.

Cron (susan's crontab, alongside the existing entries)

17 * * * * runitor -uuid 41a4d61a-7743-43d9-9b5d-d37d536e4726 -- sg mediaserver "/usr/local/bin/youtube-automate run"
40 4 * * 1 runitor -uuid 721e4cf0-d796-48e7-a184-79d21e1ba373 -- /var/lib/youtube-automate/venv/bin/python -m uv pip install --upgrade yt-dlp

Hourly is ample for single-digit channels; both checks are registered in Healthchecks with generous grace periods. The weekly updater is monitored separately so a stale yt-dlp alerts distinctly from a failed run — §12 calls that the top operational risk.

systemd unit for the admin server

Type=simple, Restart=always, User=susan, Group=mediaserver, UMask=0002. Binds 127.0.0.1:8085 only.

nginx

tube.jihakuz.xyz127.0.0.1:8085, public, TLS via certbot. The DNS record and njal.la update key already exist in update-dns.sh, so no DNS work is needed.

The hostname was already claimed. A leftover TubeArchivist server block inside sites-available/jihakuz.xyz served tube.jihakuz.xyz, proxying to the now-dead 127.0.0.1:8003. nginx uses the first server block matching a name and that file loads before any standalone tube.jihakuz.xyz file, so installing a second vhost changed nothing and every request returned 502. The old block also already owns the Let's Encrypt certificate for the hostname, so the fix is to repoint it rather than duplicate it — deploy/fix-nginx-tube.sh does that, and removes the redundant standalone vhost.

The forwarding headers are not cosmetic: the app throttles failed logins per client address and reads X-Forwarded-For. Without it every attempt appears to come from nginx itself, so a single attacker would lock out everyone.

Because the UI is public, the app carries real CSRF tokens and treats every input as hostile (§11).


10. Retention and deletion

Aging out (reap)

Effective retention for a video = its channel's retention_days if set, else the global retention_days. Candidates: state = 'downloaded' and upload_date < today - effective.

Per-channel overrides exist because the subscription mix spans wildly different cadences. EthosLab uploads roughly every two to three weeks (20 Jul, 1 Jul, 19 Jun, 1 Jun, 11 May), so at a flat 14 days his show would sit empty in Jellyfin more often than not.

Deletion removes the media file plus its .nfo, -thumb.jpg, .srt and .info.json siblings, then prunes the season directory if it's empty, and the channel directory if that leaves it bare of seasons. Set state='deleted', deleted_at, and rel_path=NULL. Keep the row — it's the tombstone. Trigger a Jellyfin refresh if anything was deleted.

Reap is a purely local operation and does not require Jellyfin to be reachable. Watch-state protection was considered and declined (§15), so a video can in principle vanish mid-watch; the mitigation available is raising that channel's retention_days.

Unsubscribe

Hard delete: shutil.rmtree() the channel directory, then DELETE FROM channel (cascades to video). Re-subscribing later starts from scratch and re-downloads the backfill window. The UI requires a confirmation step — this is destructive and irreversible.

Disk cap (optional, disk_cap_gb)

If enabled, after each reap sum size_bytes over downloaded rows and evict oldest-by- upload_date until under the cap.


11. Admin UI

Stdlib only — http.server.ThreadingHTTPServer + BaseHTTPRequestHandler. No Flask, no FastAPI, no npm. Server-rendered HTML with a single embedded <style> block. The only JavaScript is a confirm() on the unsubscribe button.

Authentication

The UI is publicly reachable, so this is real auth, and it must not nag — the requirement is "log in once per device, effectively never again".

  • One shared password, stored as a scrypt hash (n=2**14, r=8, p=1, per-install salt) in setting.admin_password_hash. Set only via youtube-automate set-password, which prompts — never passed as an argument, never written to the repo.
  • Login form posts to /login. On success, set a session cookie containing an HMAC-signed token (session_secret, generated on first run), with attributes Secure; HttpOnly; SameSite=Lax; Path=/; Max-Age=31536000 — a 365-day expiry.
  • Every state-changing POST carries a CSRF token bound to the session and verified constant-time.
  • Login throttling: after 5 consecutive failures from an IP, reject for 60 seconds. In-memory is fine.
  • All comparisons of secrets use hmac.compare_digest.

Routes

Method Path Behaviour
GET /login login form
POST /login authenticate, set cookie, 303 → /
POST /logout clear cookie, 303 → /login
GET / channel list + add form + settings form
POST /channels resolve URL → insert → spawn backfill → 303 → /
POST /channels/<id>/retention set or clear the per-channel override
POST /channels/<id>/delete confirm-guarded hard delete → 303
POST /settings validate + persist → 303
POST /channels/<id>/rescan re-queue skipped_old inside the current window → 303
GET /health JSON: yt-dlp version, POT provider up, last run time, queue depth

All POSTs redirect (303) so refresh doesn't resubmit. Everything except /login requires a valid session — including /health, which the original spec left open. That was reasonable when the UI was going to be tailnet-only; on a public hostname it is gratuitous fingerprinting surface (yt-dlp version, queue depth, channel count) for no benefit, since the cron job is monitored by Healthchecks rather than by polling this endpoint.

Channel list should show, per channel

Title, @handle, video count on disk, disk usage, most recent upload date, last poll time, the effective retention (and whether it's an override), and a warning badge if consecutive_poll_failures > 2.

Channel URL resolution

Accept: https://www.youtube.com/@handle, /channel/UC..., /c/name, /user/name, a bare @handle, or a bare UC... ID.

yt-dlp --flat-playlist --playlist-items 0 -J <url>

--playlist-items 0 fetches channel metadata without enumerating the uploads, which is the cheap way to do this. Pull channel_id, channel, description, thumbnails from the result. If resolution fails or returns no channel_id, re-render the form with an error — don't insert a half-formed row.

Reject duplicates on channel_id, not on the submitted URL — the same channel has many URL forms.


12. Known gotchas

  • yt-dlp churn is the top operational risk. Extractor args, client names, and PO token requirements change frequently. Weekly update under runitor. When something breaks, check the yt-dlp issue tracker before debugging our code.
  • The UULF prefix is undocumented and could vanish. The channel_id fallback path is not optional. UUSH/UULV 404 means "none", not "broken".
  • RSS feeds cap at ~15 entries. Fine for hourly polling, insufficient for backfill — hence the separate --flat-playlist path in §4.
  • .work must share a filesystem with the media root, or every completed download becomes a full copy. Both are under /disks (/dev/sdb2).
  • Don't transcode. If Jellyfin is transcoding these files, the format selection in §6 is wrong — fix it there, not by throwing hardware at it.
  • Descriptions are hostile input for XML generation. Use xml.etree's serialiser.
  • WAL mode is required — two writers.
  • Group and umask matter (§2). Files that Jellyfin can't read are a silent failure.
  • No passwordless sudo — root steps go in deploy/deploy.sh for the operator to run.

13. Acceptance criteria

Work through these in order; each is a real check, not a code-reading exercise.

  1. doctor passes on a clean install.
  2. Subscribing to a real, active channel creates the directory, tvshow.nfo, poster.jpg, and queues the backfill window.
  3. At least one real video downloads end to end and lands in the right season directory with a correct filename, NFO, thumbnail, and subtitle sidecar.
  4. ffprobe on the downloaded file shows h264 video and AAC audio.
  5. Files and directories are readable by the jellyfin user (group mediaserver, 0664/0770).
  6. Jellyfin, after a library scan, shows the channel as a show, the year as a season, and the video as an episode with the correct title, description, and air date — with all internet metadata providers disabled.
  7. The video direct-plays on a client with no transcoding — the Jellyfin dashboard's active-streams panel must show neither video nor audio transcoding.
  8. Re-running run immediately downloads nothing and errors on nothing (idempotency).
  9. Manually backdating a video's upload_date past the retention window causes reap to delete it and its sidecars, prune the empty directory, and leave a deleted tombstone.
  10. Re-running poll after that does not re-download the deleted video.
  11. A per-channel retention_days override visibly changes which videos that channel's reap deletes, while other channels follow the global value.
  12. Unsubscribing removes the directory and all rows.
  13. Killing run mid-download leaves no orphan in .work, and the next run recovers the downloading row to pending.
  14. The admin page requires login, stays logged in across a browser restart, and rejects a POST with a missing or wrong CSRF token.
  15. The admin page renders, adds, removes, and persists settings; a bad settings value produces an inline error rather than a traceback.
  16. Editing retention_days in the UI visibly changes reap behaviour on the next run.
  17. pytest passes with no network access.

14. Build order

  1. Skeleton: package layout, DB schema + migrations, settings accessors, doctor
  2. venv + POT provider + a hardcoded single-video download proving §3 and §6 work — do this before writing anything else of substance
  3. Channel resolution + subscribe/unsubscribe on the CLI
  4. Discovery (poll), both feed paths, including the skipped_short repair
  5. Download worker + NFO/artwork generation + move-into-place
  6. Jellyfin library creation and verification (criteria 6 and 7) — a checkpoint, not a step
  7. reap, including per-channel overrides
  8. run orchestration, flock, crash recovery, cron + runitor
  9. Admin server, including auth and CSRF
  10. systemd unit, nginx vhost, certbot, deploy.sh

Steps 2 and 6 are the two places where this design could still turn out to be wrong. Hit them early and report back rather than building on top of an unverified assumption.

Tests

pytest, offline, no network and no real yt-dlp invocation. Fakes for yt-dlp and the Jellyfin API. Coverage targets:

  • filename sanitisation, including hostile titles and the 120-char truncation
  • season/episode numbering and the ordinal-within-day logic, including the clamp at 9
  • RSS parsing from fixture XML, both UULF and fallback shapes, including a 404
  • the skipped_shortpending repair, and proof it never resurrects a deleted tombstone
  • NFO generation with descriptions full of &, <, emoji and newlines
  • state machine transitions, including crash recovery and deferred
  • reap candidate selection with and without per-channel overrides
  • settings validation and typed accessors with missing keys
  • auth: scrypt round-trip, cookie signing, CSRF verification, throttling

15. Changes from the original handover

Area Change Why
Naming ytsubsyoutube-automate throughout The original contradicted itself; the bare repo, checkout and state dir already existed under youtube-automate
Media root /disks/Plex/YouTubeSubs/disks/Plex/YouTube; .ingest.work TubeArchivist is not installed, the directory is empty and no Jellyfin library references it — verified
Admin UI Tailnet-only + basic auth → public HTTPS at tube.jihakuz.xyz with app login, scrypt, 365-day signed cookie, CSRF Requester's brother has no tailnet device; basic auth re-prompts on some mobile browsers
Watch state Removed Declined. There are five Jellyfin users, so a single jellyfin_user_id would protect one person's progress; per-channel retention is the mitigation instead
Retention Added per-channel retention_days override EthosLab's 23 week cadence would leave his show empty at a flat 14 days
Shorts filter min_duration_seconds stays 120, but skipped_short is repaired on a later UULF sighting DolanDarkest posts 122167s videos; the repair removes the transient-outage data-loss path
Premieres is_upcoming now deferred and retried, not permanently skipped A premiere becomes downloadable once it airs
Subtitles Added --write-subs alongside --write-auto-subs Channels with real uploaded subs were being skipped
Match filter !=!=? on live_status An unknown live_status shouldn't disqualify a video, same logic as duration>?
Download Record size_bytes disk_cap_gb was unimplementable without it
Refresh Also refresh Jellyfin after reap and unsubscribe Otherwise ghost episodes linger
Episode format E08110E8110 Consistency with the <episode> value in the NFO
Permissions Added Group=mediaserver, UMask=0002, setgid on media root Original was silent on it; Jellyfin reads the tree only via that group
.work Added an .ignore file as well as the dot prefix Not worth relying on undocumented behaviour for the partial-download directory
Versions Pinned bgutil to 1.3.1 / 1.3.1-deno Floating latest on both halves invites version skew
yt-dlp Install in the venv; leave /usr/local/bin/yt-dlp (2023.11.16) alone Nothing on the box references it, so replacing it buys nothing and risks other jobs
run Accepts --channel §9 spawned run --channel but only defined the flag on poll
Monitoring Two Healthchecks UUIDs, one per job Matches the existing one-UUID-per-cron-entry convention on susan
Tests Added §14 test plan Requested

Found during build-order step 2 (the PO token proof)

These were discovered by actually running the pipeline, not by reading the doc.

Area Change Why
Format sort -S "vcodec:h264,acodec:aac,res:720"-S "vcodec:h264,res:720,acodec:aac" The original selects 360p. Ranking acodec above res makes the combined 360p format (which has AAC) outrank the 720p video-only stream (acodec=none) under bv*. Verified: original gives 18 | 640x360, corrected gives 298+140 | 1280x720
JS runtime Added Deno into the venv bin/, plus PATH requirement New yt-dlp requirement. Without it, n challenge solving failed — 22 formats and zero mweb formats vs 29 with it. Deno's V8 runs fine without AVX
yt-dlp extras yt-dlpyt-dlp[default] The [default] extra is what ships yt-dlp-ejs, the challenge solver scripts. Plain yt-dlp does not include it
Impersonation Added curl-cffi<0.16 Unpinned installs get 0.16.0, which yt-dlp rejects as "unsupported"; the supported range is 0.5.10 or 0.10.x0.15.x. Removes a per-download warning and helps against the bot-detection interstitial
Subtitles De-duplicate en vs en-orig on move --sub-langs "en.*" matches both, producing two identical English tracks in Jellyfin
doctor Also compare /ping reported version against the installed plugin version Version skew between container and plugin is silent otherwise

Found during the rest of the build

Area Change Why
Backfill --flat-playlist has no dates; RSS is now the primary source Verified: every flat-playlist entry returns timestamp: None, so the specced client-side date filter was impossible
skipped_old Added explicit rescan Otherwise raising a channel's retention appears to do nothing. EthosLab's 7-day backfill queued zero videos
/health Now requires a session The UI moved from tailnet-only to public
Reap Channel directory is kept when its last season is pruned Deleting it would make an active subscription vanish from Jellyfin and reappear later, and would bin the artwork
Premieres is_upcomingdeferred, retried A premiere becomes downloadable once it airs

Verified end to end on 2026-08-11

doctor passes clean. All three channels subscribe from three different URL forms, with tvshow.nfo and real JPEG artwork. Backfill queued the right windows. A real video downloaded to h264 Main 1280x720 / aac LC in mp4 with SponsorBlock chapters, landed as EthosLab - S2026E7010 - ... [j5jhOZvm5Vo].mp4 with all four sidecars and susan:mediaserver 0664. Jellyfin resolved it from the local NFO alone — correct title, season 2026, episode 7010, aired date, overview — and reports SupportsDirectPlay: true. Reap deleted all five artefacts, pruned the season directory, kept the channel artwork and left a tombstone that survives both poll and rescan. The lock refuses a second run instantly. 240 offline tests pass.