Skip to content

Repository files navigation

StreamETH Light

A read-only archive of the StreamETH video library. Browses organizations → events → sessions, plus a searchable "all videos" view, all served from a static JSON snapshot of the production database committed to data/.

  • Public sessions only (published: "public").
  • No backend at runtime — pages read data/*.json at build/request time.
  • Video playback resolves Livepeer playbackId to an HLS stream (https://livepeercdn.studio/hls/{playbackId}/index.m3u8), falling back to a raw videoUrl when present.

Refreshing the data

The production Mongo isn't reachable directly from a laptop — it only exists on the app's private Docker network on the VPS. scripts/export-db.sh automates the whole round trip: SSHes in, starts the mongodb container if it's stopped, runs the export inside a throwaway container on that network, copies the resulting JSON into data/, and puts the container back the way it found it.

cp scripts/.env.export.example scripts/.env.export  # fill in DB_PASSWORD
pnpm export-db

Search index

The homepage feed's search/filtering queries data/streameth.db — a SQLite database with an FTS5 full-text index (including transcripts) unifying StreamETH sessions and tracked YouTube videos into one videos table (see scripts/build-db.mjs, lib/videoDb.ts). It's generated from the committed JSON, not itself committed — pnpm dev/pnpm build run pnpm build-db automatically via predev/prebuild. Uses Node's built-in node:sqlite (Node 22.5+, no native compilation), so it runs anywhere the app's Node runtime does.

MCP server

/api/mcp is a read-only MCP endpoint (Streamable HTTP, app/api/mcp/route.ts) over the same search index, so AI agents can search the archive and read transcripts. Tools: search_videos, get_video, get_transcript (paged), list_channels, list_topics.

claude mcp add --transport http streameth https://<your-domain>/api/mcp

Signed-in users get a personal token on /connect ("Connect with MCP" in the top bar) and pass it as Authorization: Bearer smcp_…. A user with no tokens gets one created automatically on their visit, since the page is the only place a token can be shown; they can add more per app. Only a SHA-256 hash is stored (mcp_tokens, supabase/migrations/); /api/mcp checks it through the verify_mcp_token database function. Users can revoke tokens on the same page.

Apps that support OAuth (e.g. Claude.ai connectors) can connect without a token, as a signed-in account. Supabase Auth's OAuth 2.1 server is the authorization server: /.well-known/oauth-protected-resource/api/mcp points clients at it, the client registers itself (dynamic client registration), and the user signs in and approves on /oauth/consent. The MCP route then verifies the Supabase access token (lib/supabase/mcpAuth.ts).

On the hosted project, in Authentication → OAuth Server: enable it, set the authorization path to /oauth/consent, and allow dynamic client registration. The project's Site URL must be this app's domain, since Supabase builds the consent URL from it. Locally, supabase/config.toml already has these set.

Ask the archive (AI answers)

The home page's question box (components/AskBox.tsx) posts to /api/ask, which runs a model through OpenRouter (default deepseek/deepseek-v4-flash, lib/ask.ts) with one tool, search_archive: an any-term FTS5 search over the same index that returns the best-matching talks (transcripts first) with numbered passages. The model searches a few times, then writes an answer citing passages as [n]; the route streams search steps, sources and answer text as NDJSON, and the UI turns [n] into links to the talks. /?ask=<question> links ask on open.

With JEV_API_KEY set, each search's passages first go through TypeSafe AI's Jev decision model (lib/jev.ts): one yes/no question per passage — does it help answer the question? — and passages below the threshold are dropped before the answering model sees them. If Jev is unset, slow (>5s) or errors, all passages are kept.

Asking requires signing in (the boxes show for everyone; asking while signed out shows a sign-in prompt that returns to the question). /api/ask is rate limited through Supabase (consume_ask_quota, migration …03_ask_rate_limit.sql): per user per hour and site-wide per day. In production the route refuses to answer if the limit can't be checked.

Env var
OPENROUTER_API_KEY Required — Ask is disabled without it
OPENROUTER_MODEL Any OpenRouter model with tool calling (default deepseek/deepseek-v4-flash)
JEV_API_KEY Optional — BeatAPI key for the Jev relevance filter
JEV_MODEL Jev model (default jev-1.13; jev-1.13-free is limited to 1 request/min)
JEV_THRESHOLD Minimum relevance probability to keep a passage (default 0.3)
ASK_HOURLY_LIMIT Questions per user per hour (default 20)
ASK_DAILY_LIMIT Questions per day across the site (default 2000)

Weekly email digest

Visitors can subscribe on the home page to a Monday email of the past week's new talks, grouped by event (lib/digest.ts). It's double opt-in: the confirmation link opens /digest/confirm, where a button (not the page load, so link-scanning mail filters can't confirm) activates the subscription. Every email has an unsubscribe link and RFC 8058 one-click List-Unsubscribe headers. Subscribers live in digest_subscribers (migration …04_digest_subscribers.sql), readable only with the service-role key. A Vercel cron (vercel.json, Mondays 09:00 UTC) calls /api/digest/send; mail goes out through Resend.

Env var
SUPABASE_SERVICE_ROLE_KEY Server-only, for digest_subscribers
RESEND_API_KEY Resend API key
DIGEST_FROM Sender, e.g. StreamETH <digest@streameth.org> (domain verified in Resend)
CRON_SECRET Vercel sends it with cron requests; /api/digest/send rejects anything else

Watch analytics

The players report plays and watch time to /api/views (lib/viewTracking.ts): one row per playback in video_views (migration …05_video_views.sql), with seconds actually played (seeking excluded), the furthest point reached, and the viewer if signed in. mode is video for the watch page and audio for Listen mode. Per-video totals are in the video_view_stats view. Both are readable only with the service-role key (SQL editor or supabase CLI); it uses the same SUPABASE_SERVICE_ROLE_KEY as the digest.

Accounts (email code or Google, saved videos)

People sign in with an emailed one-time code (the same email has a magic link for the current browser) or with Google, through Supabase Auth (components/SignInForm.tsx). There are no passwords. Google and magic links land on /auth/callback, which swaps the code for a session cookie. Supabase (Postgres + Auth) is the one persistent, writable piece of an otherwise read-only/static app. Schema lives in supabase/migrations/; saved_videos rows are protected by row-level security so a user can only see/write their own. Accounts from the earlier wallet sign-in can no longer sign in, and their MCP tokens no longer verify.

Environment variables (.env.local, and the Vercel project):

Variable Required Purpose
NEXT_PUBLIC_SUPABASE_URL yes Supabase project URL
NEXT_PUBLIC_SUPABASE_ANON_KEY yes Supabase anon/publishable key
SEND_EMAIL_HOOK_SECRET production Verifies Supabase's Send Email hook calls (sign-in codes)

New environment (e.g. a fresh Supabase project): supabase link --project-ref <ref> then supabase db push to apply the migrations. Then in the dashboard:

  • Authentication → Sign In / Providers → Email: on, signups allowed.
  • Authentication → Hooks → Send Email (HTTPS): https://<domain>/api/auth/send-email. Supabase then hands every auth email to the app, which sends it through Resend with the digest's RESEND_API_KEY and DIGEST_FROM (lib/authEmail.ts). Put the hook's secret (v1,whsec_…) in SEND_EMAIL_HOOK_SECRET. Without the hook, Supabase's built-in mailer only delivers to the project's team.
  • Authentication → Sign In / Providers → Google: a Google Cloud OAuth client ID and secret, with https://<ref>.supabase.co/auth/v1/callback as its authorized redirect URI. The Google button appears once this is on.
  • Authentication → URL Configuration: site URL, plus https://<domain>/auth/callback (and preview domains) as redirect URLs.

Local stack: supabase start uses supabase/config.toml, which enables email codes (with supabase/templates/magic_link.html, through the local mail catcher rather than the hook); set SUPABASE_AUTH_EXTERNAL_GOOGLE_* and turn on [auth.external.google] to try Google locally.

Search engines and AI crawlers

Every video should be indexable by Google and quotable by AI answer engines:

  • /sitemap.xml lists channels, events, speakers and topics; /watch/sitemap/<n>.xml are video sitemaps covering every watch page (5,000 per file). /robots.txt lists them all and explicitly allows the major AI crawlers.
  • Watch pages carry VideoObject + BreadcrumbList JSON-LD, a canonical URL, and the full transcript in the server-rendered HTML. Videos without a cover image get a generated thumbnail at /watch/<id>/poster.png.
  • /llms.txt maps the archive for AI assistants, and /watch/<id>.md is a plain-markdown copy of any talk (metadata, description, transcript).
  • pnpm build runs scripts/check-seo.mjs afterwards and fails if any watch page or sitemap entry is missing what indexing needs.

Set NEXT_PUBLIC_SITE_URL to the production domain if it differs from Vercel's production URL — canonical URLs, sitemaps and JSON-LD are built from it. Set GOOGLE_SITE_VERIFICATION / BING_SITE_VERIFICATION to verify the site in Google Search Console / Bing Webmaster Tools, then submit /robots.txt's sitemaps there and watch the Video indexing report.

Development

pnpm install
pnpm dev

Deploy

Deployed on Vercel. data/*.json is committed to git, so the video archive itself needs no environment variables or database — data/streameth.db is rebuilt from it during pnpm build. Accounts need the env vars above set in the Vercel project.

License

MIT

Releases

Packages

Contributors

Languages