Persistent instruction set for this repository. Read this before doing any work here.
A self-hosted Grafana dashboard for GitHub repository Insights (Traffic: views, unique visitors, clones, unique cloners, referring sites, popular content; plus contributor/commit stats where useful).
This is not just a pretty traffic chart; it is a lead-generation / audience-insight tool. Design panels and the data model so the owner can answer:
$repository to switch a
single repo into every panel.$repository picker via the repos dimension table. Keep the
schema tall/normalized so new filters are just WHERE clauses, not new tables.Persist enough dimension detail (repo, day, metric, source/path, unique-vs-total, and repo attributes) that any of the above is a query, never a re-fetch. Never re-download traffic just to add a filter — repo attributes come from the cheap repo-list API.
LICENSE).README.md using
- [ ] / - [x] checklists. Update the checklist as part of the same change.Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>(repository, date, metric) and
upserted, so re-running a day never double-counts. Daily overlap of 1–2 days is a
feature (late-arriving counts get corrected), not a bug.gh CLI (already logged in as maifeeulasad).
Read data with gh api .... No PAT is configured or committed.~/.config/gh/hosts.yml. So mounting ~/.config/gh into a container does not
carry auth. The reliable bridge is to export a token at runtime:
GH_TOKEN="$(gh auth token)" and pass it into the container as an env var. The
container has its own gh (or plain curl) and uses that token.Design for a swap: put auth behind an interface so we can move from CLI to a PAT or GitHub App without touching call sites.
AuthProvider (interface) -> token() -> str
├─ CliAuthProvider (default) -> shells `gh auth token`
├─ EnvTokenProvider (provision)-> reads GH_TOKEN / PAT from env
└─ GitHubAppProvider (future) -> installation token
Only CliAuthProvider/EnvTokenProvider are needed initially; the others are stubs
with clear TODOs. Never hardcode or commit a token; env + .gitignore only.
gh api)gh api repos/{owner}/{repo}/traffic/views # total + unique views, 14d
gh api repos/{owner}/{repo}/traffic/clones # total + unique clones, 14d
gh api repos/{owner}/{repo}/traffic/popular/referrers # referring sites
gh api repos/{owner}/{repo}/traffic/popular/paths # popular content
gh api repos/{owner}/{repo}/stats/contributors # contributor stats
Traffic endpoints require push access to the repo (owner token satisfies this).
┌──────────────┐ daily ┌──────────────┐ SQL upsert ┌──────────────┐
│ collector │──────────► │ GitHubClient │──────────────►│ SQLite │
│ (scheduler) │ │ + AuthProvider│ │ data/traffic.db│
└──────────────┘ └──────────────┘ └──────┬───────┘
│ query (plugin)
┌──────▼───────┐
│ Grafana │
│ dashboards │
└──────────────┘
frser-sqlite-datasource plugin), not Postgres/Prometheus: our data is small, daily,
historical counts that need backfill + idempotent upserts. SQLite gives us SQL and
zero DB-server overhead — no extra container, just a mounted data/traffic.db. Schema
keyed by (repository, day, metric). If volume ever outgrows it, the DAO layer is the
only thing that changes (swap to Postgres).AuthProvider → GitHubClient →
domain models (TrafficPoint, Referrer, PopularPath, RepoMeta) → Storage/DAO
layer → SQLite. Scheduling starts as a simple loop; keep it swappable. The client
retries on rate-limit (403/429) and transient DNS/network errors; Storage sets
busy_timeout so a one-off --sync-repos can run alongside the collector.repos table: is_fork, visibility, is_archived) is synced from the
repo-list API (GET /user/repos) — a handful of cheap calls, never the traffic
endpoints. It powers the visibility/type/archived filters.grafana/provisioning/alerting/ as code: per-repo WoW change via
time series → reduce → threshold (frser labels series by repository).grafana/provisioning/), so the whole stack is reproducible from docker compose up.
Only two containers: grafana and collector, sharing the data/ volume.docker compose here: grafana and the collector, sharing
a SQLite named volume (dbdata). Prefer testing in-repo with compose.docker context (Docker Desktop socket); don’t hardcode a socket path.GH_TOKEN="$(gh auth token)" in the compose env (never commit it; use a
git-ignored .env). Document the one-liner in the README.REPO_DELAY adds politeness between repos. Order
repos owned-first then forks, newest push first.collect --once / --backfill) for manual
testing and on a daily schedule for production.README.md box.