# Publish Assistant Hugo-generated website that helps researchers: - identify important venues to publish in; - keep track of submission deadlines; - retrieve important papers for each issue. --- ## Overview The site is domain-driven: a researcher picks a research area (e.g. "edge and cloud systems") and the assistant builds a curated, ranked, up-to-date snapshot of where to publish, when to submit, and what to read. The output is a static Hugo site that can be rebuilt on demand. Multiple research topics live under the same Hugo instance as subsites (`/cloud-edge/`, `/embedded-ai/`, …). Each topic has its own venue list, deadlines, digests, and calendar. Layouts, theme, and JavaScript are shared. Work is split between **automated scripts** (data fetching, Hugo content generation) and **agent tasks** (domain curation, paper selection, deadline gap-filling). The README below documents both halves so the site can be kept fresh over time. --- ## Repository Layout ``` publish-assistant/ ├── .gitea/workflows/gh-pages.yaml # CI: renders the site, deploys to the gh-pages branch ├── site/ # Hugo project │ ├── hugo.toml │ ├── package.json # FullCalendar deps (resolved by Hugo's js.Build) │ ├── themes/PaperMod/ # PaperMod theme (git submodule — never edit in place) │ ├── assets/ │ │ ├── js/calendar.js # FullCalendar bootstrap, bundled by js.Build │ │ └── css/extended/custom.css # Theme overrides (auto-appended by PaperMod's head) │ ├── layouts/ │ │ ├── _default/calendar.html # FullCalendar layout │ │ ├── _default/digests.html # Digest list layout │ │ ├── _default/_markup/render-link.html # Rewrites root-relative links for subpath baseURL │ │ └── partials/ │ │ ├── header.html # Section-aware nav override │ │ └── extend_head.html # Calendar CSS injection │ ├── content/ # Generated — do not edit by hand │ │ ├── _index.md # Global landing page │ │ └── / # One subdir per topic, e.g. cloud-edge/ │ │ ├── _index.md # Topic home page │ │ ├── venues/ # Venue pages (generated) │ │ ├── calendar/ # Deadline calendar (generated) │ │ └── digests/ # Paper digests (generated) │ └── data/ │ ├── topics.yaml # Topic registry (nav + build loop) │ ├── rankings/ # Shared across all topics │ │ ├── icore.csv # ICORE conference rankings │ │ └── scimago.csv # SCImago journal rankings (manual download, committed) │ └── / # One subdir per topic, e.g. cloud-edge/ │ ├── venues.yaml # Master venue list ← primary edit target │ ├── deadlines.yaml # Deadline cache; manual entries preserved │ ├── best_papers.yaml # Best-paper awards │ └── papers/ │ ├── --candidates.yaml # Full paper list (pa-fetch-papers) │ └── --digest.yaml # Curated selection (agent, Task 3) ├── src/publish_assistant/ # Python package │ ├── fetch_icore.py │ ├── fetch_scimago.py │ ├── fetch_deadlines.py │ ├── fetch_papers.py │ ├── fetch_best_papers.py │ └── generate_content.py ├── pyproject.toml # uv project; CLI entry points ├── uv.lock ├── build.sh └── README.md ``` **Topic isolation**: each topic owns its `site/data//` directory (venues, deadlines, papers) and generates into `site/content//`. The shared `site/data/rankings/` CSVs are reused by every topic. The nav bar automatically shows topic-relative Venues / Calendar / Digests links when inside a topic, and lists all topics from `topics.yaml` on the root page. --- ## Setup ```bash uv sync # installs all dependencies + registers CLI tools npm --prefix site ci # FullCalendar packages; required for any hugo build # Available commands after sync: uv run pa-fetch-icore uv run pa-fetch-scimago uv run pa-fetch-deadlines uv run pa-fetch-best-papers uv run pa-fetch-papers --venue OSDI --year 2025 uv run pa-generate # Local dev server: ./build.sh --dev # or, to process one topic without re-fetching: ./build.sh --skip-rankings --topic cloud-edge --dev ``` `site/assets/js/calendar.js` imports `@fullcalendar/*`, and Hugo's `js.Build` resolves those from `site/node_modules/` at render time. That directory is gitignored, so a fresh clone needs `npm --prefix site ci` before `hugo` (or `build.sh`) will succeed. Cloning: the theme is a submodule — `git clone --recurse-submodules`, or `git submodule update --init --recursive` after the fact. --- ## Deployment `.gitea/workflows/gh-pages.yaml` runs on every push to `main` (and on manual dispatch): checkout with submodules → install Hugo (pinned via `HUGO_VERSION`) → `npm ci` in `site/` → `hugo --source site --minify --gc --baseURL "$SITE_BASE_URL"` → push `site/public/` to the orphan `gh-pages` branch with `peaceiris/actions-gh-pages`. **CI does not fetch data and does not run `build.sh`.** Everything under `site/content/` is generated locally by `./build.sh` and committed; CI only renders it. If generated content is not committed, the deployed site is stale. **Base URL**: the site is served from a subdirectory (`https://pub.sqrt.fr///` by default). Override with the repo variable `SITE_BASE_URL` (Settings → Actions → Variables) if the site moves. Because generated Markdown stores links as root-relative paths (`/cloud-edge/venues/…`), three layouts run them through `relURL` so they survive the subpath — see "Subpath-safe links" under Hugo Content Structure. --- ## Data Sources | Source | What it provides | Automatable? | Known issues | | --- | --- | --- | --- | | [ICORE](https://portal.core.edu.au/conf-ranks/) | Conference rankings (A*, A, B, C) | Yes | Pagination uses `javascript:jumpPage('N')` — handled in `fetch_icore.py` | | [SCImago](https://www.scimagojr.com/) | Journal quartiles, SJR, H-index | Blocked | Anti-bot returns HTML. A manually downloaded CSV is committed at `site/data/rankings/scimago.csv`; `pa-generate --scimago` matches journals by title. Inline `venues.yaml` values still win. | | [DBLP](https://dblp.org/) | Paper metadata by venue | Yes | Use `dblp_key` field in `venues.yaml` | | [OpenAlex](https://openalex.org/) | Papers, open-access links | Yes | Fallback when DBLP is thin | | [WikiCFP](http://wikicfp.com/) | Submission deadlines | Partially | See detailed notes below | | [Conference websites](.) | Authoritative deadlines | Partially | See Task 2 | | [jeffhuang.com/best_paper_awards/](https://jeffhuang.com/best_paper_awards/) | Best paper awards since 1996, ~32 venues | Yes | Manually maintained; run `pa-fetch-best-papers` annually | --- ## WikiCFP Integration — Detailed Notes WikiCFP is the primary deadline source but has several quirks that required workarounds: **HTML structure**: Detail pages use `` for row labels (not ``). The parser in `fetch_cfp_details()` specifically looks for `` + `` pairs. Do not revert to `find_all("td")` — it will find zero deadline rows. **Direct ID lookup**: Add `wikicfp_id: ""` to a venue entry in `venues.yaml` to skip the search and fetch that page directly. This avoids wrong matches on common acronyms. Verified IDs for the cloud-edge topic: | Venue | WikiCFP event ID | | | --- | --- | --- | | SOSP | 191399 | | | EuroSys | 186524 | superseded — `source: manual` (two cycles) | | SoCC | 191071 | superseded — `source: manual` (two cycles) | | Middleware | 190153 | superseded — `source: manual` (two cycles) | | IPDPS | 189093 | superseded — `source: manual` (2027 edition) | | HPDC | 191029 | | A venue whose `deadlines.yaml` entry says `source: manual` is skipped by the fetcher regardless of its `wikicfp_id`; the ID is kept so the lookup still works if the entry is ever handed back to WikiCFP. **Skipping search**: Set `wikicfp_id: false` to skip WikiCFP entirely for a venue (e.g., ATC, SC, SEC — where the search returns wrong events). Deadlines for these must be filled manually. **Conferences not on WikiCFP** (for systems/networking): OSDI, NSDI, USENIX ATC, SC, MobiSys, SEC. Use `wikicfp_id: false` for all of them. **Manual deadline entries**: Add entries with `source: manual` to `site/data//deadlines.yaml`. The fetcher preserves all `source: manual` entries across runs. Format: ```yaml ATC: source: manual event_dates: Nov 16-18, 2026 location: Hong Kong submission_deadline: Jun 10, 2026 notification: Sep 18, 2026 camera_ready: Oct 16, 2026 cfp_url: https://sigops.org/s/conferences/atc/2026/cfp.html ``` ### Multiple Submission Cycles Many systems conferences run more than one submission round per year, all feeding the same event. Recording only one round makes the site claim a conference has been missed when a later round is still open. Give such a venue a `cycles` list instead of top-level deadline fields: ```yaml EuroSys: source: manual event_dates: Apr 19-23, 2027 # event fields stay at the top level location: Rabat, Morocco cfp_url: https://2027.eurosys.org/cfp.html cycles: - name: Spring # free text, verbatim from the CFP abstract_deadline: May 7, 2026 submission_deadline: May 14, 2026 notification: Aug 21, 2026 camera_ready: Sep 25, 2026 - name: Fall abstract_deadline: Sep 17, 2026 submission_deadline: Sep 24, 2026 notification: Jan 29, 2027 camera_ready: Mar 5, 2027 ``` Cycle fields are `name`, `abstract_deadline`, `submission_deadline`, `notification`, `camera_ready`, plus an optional per-cycle `cfp_url` where a venue publishes a separate call per round. Everything else (`source`, `event_dates`, `location`, `cfp_url`) describes the event and stays at the top level. **Single-round venues need no change.** `load_deadlines()` folds a flat entry into a one-element `cycles` list on read, so the two shapes are interchangeable and only genuinely multi-round venues need the extra nesting. What the generator does with cycles: | Output | Behaviour | | --- | --- | | Venue page | Multi-cycle venues get a **Submission Cycles** table with every round; the next open one is bolded and marked `(next)`. Single-cycle venues keep the **Upcoming Deadline** block unchanged. | | Venue index | One row per venue, showing the next open cycle with its name appended to the paper deadline. | | Calendar table | One row per cycle, with a `Cycle` column, sorted by paper deadline. | | Calendar grid | One event set per cycle; titles read `EuroSys Fall — paper deadline`. Elapsed cycles are kept — the grid is time-indexed, so past rounds only show when you navigate back to their month. | | Front matter | `next_deadline` / `abstract_deadline` / `notification` / `camera_ready` reflect the **next** cycle. Multi-cycle venues additionally carry `cycles` and `next_cycle`. | "Next" means the earliest cycle whose paper deadline has not passed; if every round has elapsed, the last one, so a page shows the most recent round rather than going blank. **WikiCFP cannot express this.** It lists one deadline set per event, so a multi-round venue fetched from WikiCFP will silently record whichever round WikiCFP happens to hold. Research such venues from the CFP itself and mark them `source: manual` — otherwise the next fetcher run overwrites the extra rounds. SoCC, Middleware and IPDPS were converted this way even though they still carry `wikicfp_id` in `venues.yaml`. --- ## `venues.yaml` Schema ```yaml conferences: - acronym: SOSP full_name: "ACM Symposium on Operating Systems Principles" domain: [, operating-systems, distributed-systems] url: "https://sigops.org/s/conferences/sosp/" dblp_key: "conf/sosp" wikicfp_id: "191399" # direct lookup; omit to use search; false to skip entirely journals: - acronym: TPDS full_name: "IEEE Transactions on Parallel and Distributed Systems" domain: [, parallel-computing, distributed-systems] issn: "1045-9219" url: "https://www.computer.org/csdl/journal/td" dblp_key: "journals/tpds" submission_model: rolling scimago_quartile: Q1 # optional — falls back to rankings/scimago.csv by title scimago_sjr: "1.560" scimago_h_index: "131" ``` **Important**: Hugo reserves the front matter field `url` as a page URL override. `generate_content.py` maps `venues.yaml:url` → front matter field `homepage` to avoid this conflict. ### Journal Special Issues Journals with `submission_model: rolling` have no fixed deadlines, so WikiCFP yields nothing for them. Special issue CFPs are the only dated journal deadlines worth tracking. Add a `special_issues` list to the journal entry: ```yaml journals: - acronym: TPDS # ... existing fields ... special_issues: - title: "" guest_editors: ["", ""] cfp_url: "https://" abstract_deadline: "Aug 15, 2026" submission_deadline: "Sep 1, 2026" notification: "Jan 2027" ``` Every field must come from a published CFP. There is no fetcher for these, so nothing validates them — an unverified entry publishes a deadline that looks authoritative and names real people as editors. Record `cfp_url` pointing at the specific call, so any reader can check the entry against its source. Special issue deadlines appear in three places: - **Journal page** — a "Special Issues" table below the journal metadata - **Calendar page** — a "Journal Special Issues" table below the conference deadlines. The heading always renders; with no entries recorded it carries a line saying no call is open, so the page doesn't read as if journals went untracked - **FullCalendar grid** — events colored purple (`#7b2cbf`), titled `ACR SI — Title — deadline type`, carrying `cfp_url` in `extendedProps` Dates use the same `MMM D, YYYY` format as conference deadlines. `notification` additionally accepts month-only precision (`Jan 2027`), which lands on the first of that month in the calendar grid; the tables always print the raw string. These entries are `source: manual` by nature — WikiCFP doesn't cover journal special issues. Add them directly to `site/data//venues.yaml`. --- ## `topics.yaml` Schema ```yaml topics: - slug: cloud-edge title: "Edge and Cloud Systems" description: "Conferences and journals for edge computing, cloud systems, and distributed systems." - slug: embedded-ai title: "Embedded AI" description: "Conferences and journals for on-device ML, edge inference, and TinyML." ``` The `slug` must match the directory names under `site/data/` and `site/content/`. It also becomes the URL prefix (`/cloud-edge/`, `/embedded-ai/`). The `title` appears in the nav bar on the root page and on the topic's section home. --- ## Scripts All scripts are installed as CLI entry points by `uv sync`. ### `pa-fetch-icore` Downloads ICORE rankings. Shared across all topics — run once per build. ```bash uv run pa-fetch-icore # fetch all uv run pa-fetch-icore --query "distributed systems" ``` ### `pa-fetch-scimago` Downloads SCImago CSV. **Currently blocked by anti-bot.** Will raise a descriptive error if it receives HTML instead of CSV (`build.sh` treats this as a warning and continues on cached data). ```bash uv run pa-fetch-scimago --list-areas # show area codes uv run pa-fetch-scimago --area 1705 # networks ``` The working substitute is a manual download: scimagojr.com → Journal Rankings → pick an area → Download, saved as `site/data/rankings/scimago.csv` (semicolon-separated; the loader sniffs the delimiter). That file is committed and `pa-generate` reads it, so journals without inline `scimago_*` fields still get a quartile, SJR, and H-index. Matching is by `full_name` against the CSV `Title` column (exact, then substring), so a journal whose name differs from SCImago's still needs inline values. ### `pa-fetch-deadlines` Fetches submission deadlines from WikiCFP for a specific topic. Preserves `source: manual` entries. ```bash uv run pa-fetch-deadlines \ --venues site/data/cloud-edge/venues.yaml \ --output site/data/cloud-edge/deadlines.yaml ``` After running: check `deadlines.yaml` for the `missing:` list, then do Task 2. ### `pa-fetch-best-papers` Scrapes [jeffhuang.com/best_paper_awards/](https://jeffhuang.com/best_paper_awards/) and writes `best_papers.yaml`. Run once per year. ```bash uv run pa-fetch-best-papers ``` ### `pa-fetch-papers` Fetches paper lists from DBLP (with OpenAlex fallback). ```bash uv run pa-fetch-papers --venue OSDI --year 2024 uv run pa-fetch-papers --venue TPDS --year 2024 --source openalex ``` ### `pa-generate` Regenerates all Hugo content for one topic from its data files. Safe to re-run at any time. ```bash # Explicit (for a specific topic): uv run pa-generate \ --venues site/data/cloud-edge/venues.yaml \ --deadlines site/data/cloud-edge/deadlines.yaml \ --papers-dir site/data/cloud-edge/papers \ --content site/content/cloud-edge \ --base-path /cloud-edge # Defaults (cloud-edge): uv run pa-generate ``` `--base-path` prefixes all internal links in generated markdown (e.g. `/cloud-edge/venues/…`). It must match the topic slug in the URL. Preserved fields (never overwritten): `notes`, `deadline_source`. Stripped fields (removed on regen to avoid stale data): `url` (Hugo reserved), `papers`. ### `build.sh` Full pipeline orchestrator. Loops over all topics in `topics.yaml` by default. ```bash ./build.sh # full build, all topics ./build.sh --topic cloud-edge # one topic only ./build.sh --skip-rankings # skip icore/scimago fetches (use cached CSVs) ./build.sh --skip-deadlines # use cached deadlines.yaml ./build.sh --dev # hugo server instead of build ./build.sh --topic cloud-edge --dev # dev server, one topic ``` --- ## Hugo Content Structure ### Multi-topic routing Hugo treats `site/content//` as a section. All pages inside it are served under `//`. The nav bar partial (`site/layouts/partials/header.html`) detects `.Section` at render time: - **Inside a topic** (`cloud-edge`, `embedded-ai`, …): renders Venues / Calendar / Digests links relative to that section. - **At the root** (`/`): renders one link per topic from `site/data/topics.yaml`. To add a topic: populate `site/data/topics.yaml` + `site/data//venues.yaml`, then run `./build.sh --topic `. The new section appears in the global nav automatically. ### Subpath-safe links `pa-generate` writes root-relative links (`/cloud-edge/venues/…`). Those are correct when the site is served from a domain root, but they ignore the subdirectory in `baseURL` and 404 under `https://pub.sqrt.fr///`. Rather than teaching the generator about deployment, three layouts normalize at render time — each strips the leading slash so `relURL` will prepend the baseURL subpath (`relURL` leaves already-root-relative input alone): - **`layouts/_default/_markup/render-link.html`** — Hugo render hook applied to every Markdown link in generated bodies. Protocol-relative `//host/…` links are left as-is. - **`layouts/_default/calendar.html`** — rewrites the `url` of each event before `jsonify`, so calendar clicks land on the right page. - **`layouts/partials/header.html`** — nav links go through `relURL`, and the active-item check compares against `.RelPermalink`, which carries the same prefix. All of this is a no-op when the site is served from the domain root (the local dev server). ### Theme and styling `themes/PaperMod` is a pristine git submodule pinned to an upstream commit. **Never edit files inside it**: those changes cannot be committed from this repo, and CI checks the submodule out fresh, so they would silently vanish from the deployed site. Project CSS lives in `site/assets/css/extended/custom.css`, which PaperMod's `head.html` appends after its own stylesheet — currently a shorter `.first-entry` hero, since PaperMod reserves 320px/260px for a lead image the site does not use. Calendar-only CSS stays in `layouts/partials/extend_head.html`, gated on `layout == "calendar"`. ### Venues **`site/content//venues/_index.md`** — overview, links to conferences and journals. Generated. **`site/content//venues/conferences/_index.md`** — table of all conferences sorted by ICORE rank with deadlines, domains, and digest count. Generated. **`site/content//venues/journals/_index.md`** — table of all journals sorted by SCImago quartile. Generated. Each venue page body is **fully generated Markdown**. The body includes a metadata table (rank, domains, latest digest link), an upcoming deadline block, and previous-edition info blocks (paper count, topics, digest link). ### Calendar **`site/content//calendar/_index.md`** — uses `layout: calendar`. Events are embedded as JSON-ready YAML in the `events:` front matter field by `pa-generate`. Colors: orange = abstract deadline, red = paper deadline, blue = conference dates. ### Digests **`site/content//digests/-/index.md`** — leaf page (`index.md`, not `_index.md`). Body contains the full paper list with TL;DR and why-notable for each paper. Papers data lives in `site/data//papers/--digest.yaml`. --- ## Agent Prompts Ready-to-use prompts. Paste directly into Claude Code (or any agent) as a starting point. --- ### Bootstrap a new topic ``` Bootstrap a new publish-assistant topic for "" (slug: ). Steps: 1. Research the top 10–15 conferences and 5–10 journals for this domain. Use csrankings.org, the ICORE portal (portal.core.edu.au/conf-ranks/), and SCImago (scimagojr.com) for rankings. 2. For each conference: record acronym, full name, ICORE rank, official website URL for the upcoming edition, DBLP stream key (conf/), and WikiCFP event ID if you can find it (set wikicfp_id: false for short or ambiguous acronyms). 3. For each journal: record acronym, full name, ISSN, DBLP stream key (journals/), submission model (rolling / special issues). Quartile, SJR, and H-index come from site/data/rankings/scimago.csv when full_name matches the CSV Title; add them inline to venues.yaml only when it doesn't. 4. Create site/data//venues.yaml following the schema in the README. 5. Create site/data//papers/ (empty directory). 6. Add the topic to site/data/topics.yaml: - slug: title: "" description: "" 7. Create site/content//_index.md: --- title: "" description: "" draft: false --- 8. Run: ./build.sh --skip-rankings --topic (Use --skip-rankings to reuse cached ICORE data if already fresh.) 9. For any conferences under deadlines.yaml missing:, do the manual deadline task (Task 2 in the README). 10. Run: ./build.sh --skip-rankings --topic Then: hugo --source site --minify Confirm the site builds cleanly and //venues/ loads correctly. 11. Commit the generated site/content// tree along with the data files — CI renders committed content and never regenerates it. Checklist before finishing: - [ ] site/data//venues.yaml has all venues with correct dblp_key - [ ] wikicfp_id set or false on every conference - [ ] Every journal shows a quartile (from scimago.csv, or inline if unmatched) - [ ] site/data/topics.yaml updated - [ ] site/content//_index.md created - [ ] ./build.sh --skip-rankings --topic runs without errors - [ ] hugo --source site --minify succeeds - [ ] Generated content under site/content// committed ``` --- ### Add a venue to an existing topic ``` Add to the publish-assistant topic "". 1. Look up: - Full name and ICORE rank (conferences) or SCImago quartile + SJR + H-index (journals) - Official website URL for the upcoming edition - DBLP stream key at dblp.org - WikiCFP event ID, or note if this acronym is ambiguous (set wikicfp_id: false) 2. Append the entry to site/data//venues.yaml following the existing schema. 3. Run: uv run pa-fetch-deadlines \ --venues site/data//venues.yaml \ --output site/data//deadlines.yaml 4. If the conference appears under missing: in deadlines.yaml, find the deadline on the official CFP page and add a source: manual entry. 5. Run: uv run pa-generate (defaults to cloud-edge) or with explicit --venues / --content / --base-path flags for the target topic. 6. Confirm the new venue page appears at //venues/conferences// (or journals/) with correct metadata. For a journal, check that the quartile resolved from site/data/rankings/scimago.csv; if it came out blank, the full_name doesn't match the CSV Title — add scimago_* fields inline. 7. Commit the regenerated site/content// files together with the data changes; CI deploys committed content and does not regenerate it. ``` --- ### Build a digest for a venue + year ``` Build a paper digest for in the publish-assistant topic "". 1. Run: uv run pa-fetch-papers --venue --year This writes site/data//papers/--candidates.yaml. 2. Run: uv run pa-fetch-best-papers (Skip if site/data/best_papers.yaml already exists and is recent.) 3. Read the candidates file. Select 8–15 papers that are: - Methodologically novel (new algorithms, system designs, formal proofs) - Attracting community attention (highly cited if issue is ≥1 year old; well-known authors or top-venue co-publications if recent) - Representative of the breadth of the issue (avoid over-indexing on one subtheme) - Preferably open-access (arXiv, USENIX, ACM OpenTOC) Flag any paper in best_papers.yaml (include it; it's a strong signal). 4. For each selected paper write: - tldr: one sentence, the core technical contribution - why_notable: 1–2 sentences — novelty, impact, surprising result, or influential technique; what would make a program committee member recommend this paper to colleagues 5. Write site/data//papers/--digest.yaml: venue: year: date: "" tags: [<3–5 topic tags>] selected: - dblp_key: "..." title: "..." tldr: "..." why_notable: "..." 6. Run: uv run pa-generate (or with explicit flags for the topic) 7. Confirm the digest page at //digests/-/ renders correctly. ``` --- ### Annual cycle refresh for a topic ``` Refresh the publish-assistant topic "" for the new conference cycle. 1. For each conference in site/data//venues.yaml: a. Check whether the url field points to the upcoming edition (many venues use year-specific URLs like osdi26, 2027.eurosys.org, mobisys/2026/). Update any that point to past editions. b. Verify wikicfp_id still points to the upcoming edition by visiting http://wikicfp.com/cfp/servlet/event.showcfp?eventid=. If it points to a past event, search WikiCFP for the new edition. If the new event page doesn't exist yet, set wikicfp_id: false and add a source: manual entry to deadlines.yaml; restore the ID once it appears. 2. Run: uv run pa-fetch-deadlines \ --venues site/data//venues.yaml \ --output site/data//deadlines.yaml 3. For each entry under missing: in deadlines.yaml, visit the conference CFP page and add a source: manual entry to deadlines.yaml. 4. Run: ./build.sh --skip-rankings --topic Then confirm hugo --source site --minify succeeds and all venue pages show correct upcoming deadlines. Checklist: - [ ] All conference url fields updated to upcoming edition - [ ] All wikicfp_id values verified (or set to false + manual entry) - [ ] missing: list in deadlines.yaml is empty - [ ] Build succeeds, no broken links - [ ] Regenerated site/content// committed so CI deploys the new deadlines ``` --- ## Agent Tasks (Reference) ### Task 1 — Bootstrap a topic See the "Bootstrap a new topic" prompt above. The one-shot checklist is embedded in the prompt. ### Task 2 — Fill in missing deadlines **Trigger:** `pa-fetch-deadlines` lists conferences under `missing:`, or a deadline looks wrong. 1. For each missing conference, visit the official website's "Call for Papers" / "Important Dates" page. 2. Also check WikiCFP manually — if you find the right event ID, add `wikicfp_id` to `venues.yaml` so future runs fetch it automatically. 3. Extract: abstract deadline, paper deadline, notification, camera-ready, event dates, location. 4. Add a `source: manual` entry to `site/data//deadlines.yaml`. 5. Run `uv run pa-generate` (with explicit flags for the topic). **Conferences reliably NOT on WikiCFP** (systems/networking): - USENIX family: OSDI, NSDI, USENIX ATC, USENIX Security — use usenix.org - SC — use sc``.supercomputing.org - MobiSys — use sigmobile.org/mobisys/``/ - SEC — use acm-ieee-sec.org/``/ - Short/common acronyms (ATC, SEC) collide with unrelated events — always use `wikicfp_id: false` ### Task 3 — Build a digest See the "Build a digest for a venue + year" prompt above. **Digest YAML format:** ```yaml venue: OSDI year: 2024 date: "2024-07-10" tags: [llm-serving, distributed-systems, storage] selected: - dblp_key: "conf/osdi/ZhongLCHZL0024" title: "DistServe: Disaggregating Prefill and Decoding ..." tldr: "Separates prefill and decode onto different GPU pools, eliminating head-of-line blocking." why_notable: "Became one of the most-cited LLM systems papers of 2024; disaggregation is now standard in production inference stacks." ``` ### Task 4 — Refresh rankings 1. Run `uv run pa-fetch-icore` for fresh ICORE data. 2. For SCImago: download the CSV manually from [scimagojr.com](https://www.scimagojr.com/journalrank.php) and replace `site/data/rankings/scimago.csv` (commit it — CI does not fetch). The automated fetch is blocked. 3. For each journal in `venues.yaml` that carries inline `scimago_quartile` / `scimago_sjr`, compare against the new CSV and update if changed. Journals without inline values pick the new numbers up automatically on the next `pa-generate`. 4. Run `./build.sh --skip-deadlines` (or per-topic with `--topic `). ### Task 5 — Add a new venue mid-cycle See the "Add a venue to an existing topic" prompt above. ### Task 6 — Annual cycle refresh See the "Annual cycle refresh for a topic" prompt above. --- ## Known Gotchas (for agents picking this up) - **`--base-path` must match the topic slug**: `pa-generate --base-path /cloud-edge` prefixes all internal links in generated markdown. If you run `pa-generate` without this flag (old default was `/cloud-edge`), all links on the new topic would resolve to `/cloud-edge/...` instead. Always pass explicit `--base-path /` when generating for a non-default topic, or use `build.sh --topic ` which sets it automatically. - **Hugo `url` field**: reserved by Hugo to override the page URL. `venues.yaml` uses `url:` but `generate_content.py` maps it to `homepage:` in front matter. Never write `url:` in Hugo front matter via the generator. - **PaperMod renders body, not front matter**: all visible content must be in the Markdown body. Front matter is used only for metadata and Hugo taxonomy. If a page looks empty, check that `_conf_body()` / `_journal_body()` etc. are being called. - **`_index.md` vs `index.md`**: section pages use `_index.md` (list template), leaf pages use `index.md` (single template). Digest pages are `index.md` — using `_index.md` makes them section pages and breaks pagination. - **Rankings CSVs are in `site/data/rankings/`**: Hugo's data loader is configured to ignore `data/rankings/*.csv` via `ignoreFiles` in `hugo.toml`. If you move these files or add new CSVs, update `ignoreFiles` accordingly — Hugo cannot parse arbitrary CSV as a data map and will error on build. - **ICORE pagination**: the ICORE portal uses `javascript:jumpPage('N')` links, not standard `?page=N` URLs. `fetch_icore.py` handles this. If you get only 50 results instead of ~900, pagination is broken. - **SCImago blocked**: `pa-fetch-scimago` will raise a clear error if anti-bot HTML is returned. Refresh `site/data/rankings/scimago.csv` by hand, or add data inline to `venues.yaml`. - **Never edit `site/themes/PaperMod/`**: it is a submodule pinned to upstream, and CI checks it out fresh — edits there are not committed and disappear on deploy. Put CSS in `site/assets/css/extended/custom.css` and template overrides in `site/layouts/`. - **Generated content is committed, and CI never regenerates it**: the Gitea workflow only runs `hugo`. After editing any `site/data/` file, run `./build.sh` (or `pa-generate`) and commit the resulting `site/content/` diff, or the deployed site will not change. - **`npm ci` before `hugo`**: `js.Build` resolves `@fullcalendar/*` from `site/node_modules/`, which is gitignored. A fresh clone that skips it gets a build error on the calendar page. - **Root-relative links vs. subpath `baseURL`**: generated links start with `/`, which Hugo leaves untouched — they break when the site is served from a subdirectory. The render hook, calendar layout, and header partial each strip the leading slash and call `relURL`. If you add a layout that emits links from data, do the same; do not "fix" it by changing `--base-path`, which sets the topic prefix, not the deployment prefix. - **WikiCFP `` labels**: deadline detail pages use `` for label cells, not ``. The parser looks for `+` pairs. - **Calendar events**: `pa-generate` embeds events as YAML in the `events:` front matter field. The layout at `site/layouts/_default/calendar.html` reads `.Params.events` and initializes FullCalendar. Do not remove the `layout: calendar` front matter field. - **Deadline preservation**: never strip deadline fields from `deadlines.yaml` just because the submission window has closed. Remove an entry only once the conference has taken place. - **WikiCFP IDs are edition-specific**: each year's event gets a new ID. IDs verified at one point in time will be wrong once a new cycle begins. See Task 6 for the annual refresh checklist. --- ## Future Considerations - **iCal feed** — generate a `.ics` file from deadline data so researchers can subscribe from their calendar app. - **Email/RSS notifications** — alert when a deadline is within N weeks. - **Citation tracking** — periodically re-query OpenAlex for citation counts on digest papers. - **Best paper badge** — `pa-generate` could cross-reference `best_papers.yaml` with digest candidates and add a `best_paper_award: true` flag, then render a badge in `_digest_body()`. - **Automated WikiCFP ID discovery** — when a venue has no `wikicfp_id`, attempt a search and record the result in `venues.yaml` for future runs (reduces manual work when bootstrapping new topics). - **Cross-topic venue pages** — some venues (e.g., ASPLOS) span multiple topics. A future `cross_listed: [embedded-ai, cloud-edge]` field in `venues.yaml` could render the venue under multiple topic sections without duplicating data.