# Publish Assistant Hugo-generated website that helps researchers: - identify important venues to publish in; - keep track of submission deadlines; - retrieve important papers for each issue. --- ## Overview The site is domain-driven: a researcher picks a research area (e.g. "edge and cloud systems") and the assistant builds a curated, ranked, up-to-date snapshot of where to publish, when to submit, and what to read. The output is a static Hugo site that can be rebuilt on demand. Multiple research topics live under the same Hugo instance as subsites (`/cloud-edge/`, `/embedded-ai/`, …). Each topic has its own venue list, deadlines, digests, and calendar. Layouts, theme, and JavaScript are shared. Work is split between **automated scripts** (data fetching, Hugo content generation) and **agent tasks** (domain curation, paper selection, deadline gap-filling). The README below documents both halves so the site can be kept fresh over time. --- ## Repository Layout ``` publish-assistant/ ├── site/ # Hugo project │ ├── hugo.toml │ ├── themes/PaperMod/ # PaperMod theme (git submodule) │ ├── layouts/ │ │ ├── _default/calendar.html # FullCalendar layout │ │ ├── _default/digests.html # Digest list layout │ │ └── partials/ │ │ ├── header.html # Section-aware nav override │ │ └── extend_head.html # Calendar CSS injection │ ├── content/ # Generated — do not edit by hand │ │ ├── _index.md # Global landing page │ │ └── / # One subdir per topic, e.g. cloud-edge/ │ │ ├── _index.md # Topic home page │ │ ├── venues/ # Venue pages (generated) │ │ ├── calendar/ # Deadline calendar (generated) │ │ └── digests/ # Paper digests (generated) │ └── data/ │ ├── topics.yaml # Topic registry (nav + build loop) │ ├── rankings/ # Shared across all topics │ │ ├── icore.csv # ICORE conference rankings │ │ └── scimago.csv # SCImago journal rankings (optional) │ └── / # One subdir per topic, e.g. cloud-edge/ │ ├── venues.yaml # Master venue list ← primary edit target │ ├── deadlines.yaml # Deadline cache; manual entries preserved │ ├── best_papers.yaml # Best-paper awards │ └── papers/ │ ├── --candidates.yaml # Full paper list (pa-fetch-papers) │ └── --digest.yaml # Curated selection (agent, Task 3) ├── src/publish_assistant/ # Python package │ ├── fetch_icore.py │ ├── fetch_scimago.py │ ├── fetch_deadlines.py │ ├── fetch_papers.py │ ├── fetch_best_papers.py │ └── generate_content.py ├── pyproject.toml # uv project; CLI entry points ├── uv.lock ├── build.sh └── README.md ``` **Topic isolation**: each topic owns its `site/data//` directory (venues, deadlines, papers) and generates into `site/content//`. The shared `site/data/rankings/` CSVs are reused by every topic. The nav bar automatically shows topic-relative Venues / Calendar / Digests links when inside a topic, and lists all topics from `topics.yaml` on the root page. --- ## Setup ```bash uv sync # installs all dependencies + registers CLI tools # Available commands after sync: uv run pa-fetch-icore uv run pa-fetch-scimago uv run pa-fetch-deadlines uv run pa-fetch-best-papers uv run pa-fetch-papers --venue OSDI --year 2025 uv run pa-generate # Local dev server: ./build.sh --dev # or, to process one topic without re-fetching: ./build.sh --skip-rankings --topic cloud-edge --dev ``` --- ## Data Sources | Source | What it provides | Automatable? | Known issues | | --- | --- | --- | --- | | [ICORE](https://portal.core.edu.au/conf-ranks/) | Conference rankings (A*, A, B, C) | Yes | Pagination uses `javascript:jumpPage('N')` — handled in `fetch_icore.py` | | [SCImago](https://www.scimagojr.com/) | Journal quartiles, SJR, H-index | Blocked | Anti-bot returns HTML; add data manually to `venues.yaml` under `scimago_quartile` etc. | | [DBLP](https://dblp.org/) | Paper metadata by venue | Yes | Use `dblp_key` field in `venues.yaml` | | [OpenAlex](https://openalex.org/) | Papers, open-access links | Yes | Fallback when DBLP is thin | | [WikiCFP](http://wikicfp.com/) | Submission deadlines | Partially | See detailed notes below | | [Conference websites](.) | Authoritative deadlines | Partially | See Task 2 | | [jeffhuang.com/best_paper_awards/](https://jeffhuang.com/best_paper_awards/) | Best paper awards since 1996, ~32 venues | Yes | Manually maintained; run `pa-fetch-best-papers` annually | --- ## WikiCFP Integration — Detailed Notes WikiCFP is the primary deadline source but has several quirks that required workarounds: **HTML structure**: Detail pages use `` for row labels (not ``). The parser in `fetch_cfp_details()` specifically looks for `` + `` pairs. Do not revert to `find_all("td")` — it will find zero deadline rows. **Direct ID lookup**: Add `wikicfp_id: ""` to a venue entry in `venues.yaml` to skip the search and fetch that page directly. This avoids wrong matches on common acronyms. Verified IDs for the cloud-edge topic: | Venue | WikiCFP event ID | | --- | --- | | SOSP | 191399 | | EuroSys | 186524 | | SoCC | 191071 | | Middleware | 190153 | | IPDPS | 189093 | | HPDC | 191029 | **Skipping search**: Set `wikicfp_id: false` to skip WikiCFP entirely for a venue (e.g., ATC, SC, SEC — where the search returns wrong events). Deadlines for these must be filled manually. **Conferences not on WikiCFP** (for systems/networking): OSDI, NSDI, USENIX ATC, SC, MobiSys, SEC. Use `wikicfp_id: false` for all of them. **Manual deadline entries**: Add entries with `source: manual` to `site/data//deadlines.yaml`. The fetcher preserves all `source: manual` entries across runs. Format: ```yaml ATC: source: manual event_dates: Nov 16-18, 2026 location: Hong Kong submission_deadline: Jun 10, 2026 notification: Sep 18, 2026 camera_ready: Oct 16, 2026 cfp_url: https://sigops.org/s/conferences/atc/2026/cfp.html ``` --- ## `venues.yaml` Schema ```yaml conferences: - acronym: SOSP full_name: "ACM Symposium on Operating Systems Principles" domain: [, operating-systems, distributed-systems] url: "https://sigops.org/s/conferences/sosp/" dblp_key: "conf/sosp" wikicfp_id: "191399" # direct lookup; omit to use search; false to skip entirely journals: - acronym: TPDS full_name: "IEEE Transactions on Parallel and Distributed Systems" domain: [, parallel-computing, distributed-systems] issn: "1045-9219" url: "https://www.computer.org/csdl/journal/td" dblp_key: "journals/tpds" submission_model: rolling scimago_quartile: Q1 # add manually — SCImago CSV download is blocked scimago_sjr: "1.560" scimago_h_index: "131" ``` **Important**: Hugo reserves the front matter field `url` as a page URL override. `generate_content.py` maps `venues.yaml:url` → front matter field `homepage` to avoid this conflict. --- ## `topics.yaml` Schema ```yaml topics: - slug: cloud-edge title: "Edge and Cloud Systems" description: "Conferences and journals for edge computing, cloud systems, and distributed systems." - slug: embedded-ai title: "Embedded AI" description: "Conferences and journals for on-device ML, edge inference, and TinyML." ``` The `slug` must match the directory names under `site/data/` and `site/content/`. It also becomes the URL prefix (`/cloud-edge/`, `/embedded-ai/`). The `title` appears in the nav bar on the root page and on the topic's section home. --- ## Scripts All scripts are installed as CLI entry points by `uv sync`. ### `pa-fetch-icore` Downloads ICORE rankings. Shared across all topics — run once per build. ```bash uv run pa-fetch-icore # fetch all uv run pa-fetch-icore --query "distributed systems" ``` ### `pa-fetch-scimago` Downloads SCImago CSV. **Currently blocked by anti-bot.** Will raise a descriptive error if it receives HTML instead of CSV. Add journal data manually to `venues.yaml` instead. ```bash uv run pa-fetch-scimago --list-areas # show area codes uv run pa-fetch-scimago --area 1705 # networks ``` ### `pa-fetch-deadlines` Fetches submission deadlines from WikiCFP for a specific topic. Preserves `source: manual` entries. ```bash uv run pa-fetch-deadlines \ --venues site/data/cloud-edge/venues.yaml \ --output site/data/cloud-edge/deadlines.yaml ``` After running: check `deadlines.yaml` for the `missing:` list, then do Task 2. ### `pa-fetch-best-papers` Scrapes [jeffhuang.com/best_paper_awards/](https://jeffhuang.com/best_paper_awards/) and writes `best_papers.yaml`. Run once per year. ```bash uv run pa-fetch-best-papers ``` ### `pa-fetch-papers` Fetches paper lists from DBLP (with OpenAlex fallback). ```bash uv run pa-fetch-papers --venue OSDI --year 2024 uv run pa-fetch-papers --venue TPDS --year 2024 --source openalex ``` ### `pa-generate` Regenerates all Hugo content for one topic from its data files. Safe to re-run at any time. ```bash # Explicit (for a specific topic): uv run pa-generate \ --venues site/data/cloud-edge/venues.yaml \ --deadlines site/data/cloud-edge/deadlines.yaml \ --papers-dir site/data/cloud-edge/papers \ --content site/content/cloud-edge \ --base-path /cloud-edge # Defaults (cloud-edge): uv run pa-generate ``` `--base-path` prefixes all internal links in generated markdown (e.g. `/cloud-edge/venues/…`). It must match the topic slug in the URL. Preserved fields (never overwritten): `notes`, `deadline_source`. Stripped fields (removed on regen to avoid stale data): `url` (Hugo reserved), `papers`. ### `build.sh` Full pipeline orchestrator. Loops over all topics in `topics.yaml` by default. ```bash ./build.sh # full build, all topics ./build.sh --topic cloud-edge # one topic only ./build.sh --skip-rankings # skip icore/scimago fetches (use cached CSVs) ./build.sh --skip-deadlines # use cached deadlines.yaml ./build.sh --dev # hugo server instead of build ./build.sh --topic cloud-edge --dev # dev server, one topic ``` --- ## Hugo Content Structure ### Multi-topic routing Hugo treats `site/content//` as a section. All pages inside it are served under `//`. The nav bar partial (`site/layouts/partials/header.html`) detects `.Section` at render time: - **Inside a topic** (`cloud-edge`, `embedded-ai`, …): renders Venues / Calendar / Digests links relative to that section. - **At the root** (`/`): renders one link per topic from `site/data/topics.yaml`. To add a topic: populate `site/data/topics.yaml` + `site/data//venues.yaml`, then run `./build.sh --topic `. The new section appears in the global nav automatically. ### Venues **`site/content//venues/_index.md`** — overview, links to conferences and journals. Generated. **`site/content//venues/conferences/_index.md`** — table of all conferences sorted by ICORE rank with deadlines, domains, and digest count. Generated. **`site/content//venues/journals/_index.md`** — table of all journals sorted by SCImago quartile. Generated. Each venue page body is **fully generated Markdown**. The body includes a metadata table (rank, domains, latest digest link), an upcoming deadline block, and previous-edition info blocks (paper count, topics, digest link). ### Calendar **`site/content//calendar/_index.md`** — uses `layout: calendar`. Events are embedded as JSON-ready YAML in the `events:` front matter field by `pa-generate`. Colors: orange = abstract deadline, red = paper deadline, blue = conference dates. ### Digests **`site/content//digests/-/index.md`** — leaf page (`index.md`, not `_index.md`). Body contains the full paper list with TL;DR and why-notable for each paper. Papers data lives in `site/data//papers/--digest.yaml`. --- ## Agent Prompts Ready-to-use prompts. Paste directly into Claude Code (or any agent) as a starting point. --- ### Bootstrap a new topic ``` Bootstrap a new publish-assistant topic for "" (slug: ). Steps: 1. Research the top 10–15 conferences and 5–10 journals for this domain. Use csrankings.org, the ICORE portal (portal.core.edu.au/conf-ranks/), and SCImago (scimagojr.com) for rankings. 2. For each conference: record acronym, full name, ICORE rank, official website URL for the upcoming edition, DBLP stream key (conf/), and WikiCFP event ID if you can find it (set wikicfp_id: false for short or ambiguous acronyms). 3. For each journal: record acronym, full name, ISSN, SCImago quartile + SJR + H-index (add inline to venues.yaml — the CSV download is blocked), DBLP stream key (journals/), submission model (rolling / special issues). 4. Create site/data//venues.yaml following the schema in the README. 5. Create site/data//papers/ (empty directory). 6. Add the topic to site/data/topics.yaml: - slug: title: "" description: "" 7. Create site/content//_index.md: --- title: "" description: "" draft: false --- 8. Run: ./build.sh --skip-rankings --topic (Use --skip-rankings to reuse cached ICORE data if already fresh.) 9. For any conferences under deadlines.yaml missing:, do the manual deadline task (Task 2 in the README). 10. Run: ./build.sh --skip-rankings --topic Then: hugo --source site --minify Confirm the site builds cleanly and //venues/ loads correctly. Checklist before finishing: - [ ] site/data//venues.yaml has all venues with correct dblp_key - [ ] wikicfp_id set or false on every conference - [ ] SCImago data added inline for every journal - [ ] site/data/topics.yaml updated - [ ] site/content//_index.md created - [ ] ./build.sh --skip-rankings --topic runs without errors - [ ] hugo --source site --minify succeeds ``` --- ### Add a venue to an existing topic ``` Add to the publish-assistant topic "". 1. Look up: - Full name and ICORE rank (conferences) or SCImago quartile + SJR + H-index (journals) - Official website URL for the upcoming edition - DBLP stream key at dblp.org - WikiCFP event ID, or note if this acronym is ambiguous (set wikicfp_id: false) 2. Append the entry to site/data//venues.yaml following the existing schema. 3. Run: uv run pa-fetch-deadlines \ --venues site/data//venues.yaml \ --output site/data//deadlines.yaml 4. If the conference appears under missing: in deadlines.yaml, find the deadline on the official CFP page and add a source: manual entry. 5. Run: uv run pa-generate (defaults to cloud-edge) or with explicit --venues / --content / --base-path flags for the target topic. 6. Confirm the new venue page appears at //venues/conferences// (or journals/) with correct metadata. ``` --- ### Build a digest for a venue + year ``` Build a paper digest for in the publish-assistant topic "". 1. Run: uv run pa-fetch-papers --venue --year This writes site/data//papers/--candidates.yaml. 2. Run: uv run pa-fetch-best-papers (Skip if site/data/best_papers.yaml already exists and is recent.) 3. Read the candidates file. Select 8–15 papers that are: - Methodologically novel (new algorithms, system designs, formal proofs) - Attracting community attention (highly cited if issue is ≥1 year old; well-known authors or top-venue co-publications if recent) - Representative of the breadth of the issue (avoid over-indexing on one subtheme) - Preferably open-access (arXiv, USENIX, ACM OpenTOC) Flag any paper in best_papers.yaml (include it; it's a strong signal). 4. For each selected paper write: - tldr: one sentence, the core technical contribution - why_notable: 1–2 sentences — novelty, impact, surprising result, or influential technique; what would make a program committee member recommend this paper to colleagues 5. Write site/data//papers/--digest.yaml: venue: year: date: "" tags: [<3–5 topic tags>] selected: - dblp_key: "..." title: "..." tldr: "..." why_notable: "..." 6. Run: uv run pa-generate (or with explicit flags for the topic) 7. Confirm the digest page at //digests/-/ renders correctly. ``` --- ### Annual cycle refresh for a topic ``` Refresh the publish-assistant topic "" for the new conference cycle. 1. For each conference in site/data//venues.yaml: a. Check whether the url field points to the upcoming edition (many venues use year-specific URLs like osdi26, 2027.eurosys.org, mobisys/2026/). Update any that point to past editions. b. Verify wikicfp_id still points to the upcoming edition by visiting http://wikicfp.com/cfp/servlet/event.showcfp?eventid=. If it points to a past event, search WikiCFP for the new edition. If the new event page doesn't exist yet, set wikicfp_id: false and add a source: manual entry to deadlines.yaml; restore the ID once it appears. 2. Run: uv run pa-fetch-deadlines \ --venues site/data//venues.yaml \ --output site/data//deadlines.yaml 3. For each entry under missing: in deadlines.yaml, visit the conference CFP page and add a source: manual entry to deadlines.yaml. 4. Run: ./build.sh --skip-rankings --topic Then confirm hugo --source site --minify succeeds and all venue pages show correct upcoming deadlines. Checklist: - [ ] All conference url fields updated to upcoming edition - [ ] All wikicfp_id values verified (or set to false + manual entry) - [ ] missing: list in deadlines.yaml is empty - [ ] Build succeeds, no broken links ``` --- ## Agent Tasks (Reference) ### Task 1 — Bootstrap a topic See the "Bootstrap a new topic" prompt above. The one-shot checklist is embedded in the prompt. ### Task 2 — Fill in missing deadlines **Trigger:** `pa-fetch-deadlines` lists conferences under `missing:`, or a deadline looks wrong. 1. For each missing conference, visit the official website's "Call for Papers" / "Important Dates" page. 2. Also check WikiCFP manually — if you find the right event ID, add `wikicfp_id` to `venues.yaml` so future runs fetch it automatically. 3. Extract: abstract deadline, paper deadline, notification, camera-ready, event dates, location. 4. Add a `source: manual` entry to `site/data//deadlines.yaml`. 5. Run `uv run pa-generate` (with explicit flags for the topic). **Conferences reliably NOT on WikiCFP** (systems/networking): - USENIX family: OSDI, NSDI, USENIX ATC, USENIX Security — use usenix.org - SC — use sc``.supercomputing.org - MobiSys — use sigmobile.org/mobisys/``/ - SEC — use acm-ieee-sec.org/``/ - Short/common acronyms (ATC, SEC) collide with unrelated events — always use `wikicfp_id: false` ### Task 3 — Build a digest See the "Build a digest for a venue + year" prompt above. **Digest YAML format:** ```yaml venue: OSDI year: 2024 date: "2024-07-10" tags: [llm-serving, distributed-systems, storage] selected: - dblp_key: "conf/osdi/ZhongLCHZL0024" title: "DistServe: Disaggregating Prefill and Decoding ..." tldr: "Separates prefill and decode onto different GPU pools, eliminating head-of-line blocking." why_notable: "Became one of the most-cited LLM systems papers of 2024; disaggregation is now standard in production inference stacks." ``` ### Task 4 — Refresh rankings 1. Run `uv run pa-fetch-icore` for fresh ICORE data. 2. For SCImago: download the CSV manually from [scimagojr.com](https://www.scimagojr.com/journalrank.php) and place at `site/data/rankings/scimago.csv`. The automated fetch is blocked. 3. For each journal in `venues.yaml`, compare `scimago_quartile` / `scimago_sjr` against new CSV. Update inline values if changed. 4. Run `./build.sh --skip-deadlines` (or per-topic with `--topic `). ### Task 5 — Add a new venue mid-cycle See the "Add a venue to an existing topic" prompt above. ### Task 6 — Annual cycle refresh See the "Annual cycle refresh for a topic" prompt above. --- ## Known Gotchas (for agents picking this up) - **`--base-path` must match the topic slug**: `pa-generate --base-path /cloud-edge` prefixes all internal links in generated markdown. If you run `pa-generate` without this flag (old default was `/cloud-edge`), all links on the new topic would resolve to `/cloud-edge/...` instead. Always pass explicit `--base-path /` when generating for a non-default topic, or use `build.sh --topic ` which sets it automatically. - **Hugo `url` field**: reserved by Hugo to override the page URL. `venues.yaml` uses `url:` but `generate_content.py` maps it to `homepage:` in front matter. Never write `url:` in Hugo front matter via the generator. - **PaperMod renders body, not front matter**: all visible content must be in the Markdown body. Front matter is used only for metadata and Hugo taxonomy. If a page looks empty, check that `_conf_body()` / `_journal_body()` etc. are being called. - **`_index.md` vs `index.md`**: section pages use `_index.md` (list template), leaf pages use `index.md` (single template). Digest pages are `index.md` — using `_index.md` makes them section pages and breaks pagination. - **Rankings CSVs are in `site/data/rankings/`**: Hugo's data loader is configured to ignore `data/rankings/*.csv` via `ignoreFiles` in `hugo.toml`. If you move these files or add new CSVs, update `ignoreFiles` accordingly — Hugo cannot parse arbitrary CSV as a data map and will error on build. - **ICORE pagination**: the ICORE portal uses `javascript:jumpPage('N')` links, not standard `?page=N` URLs. `fetch_icore.py` handles this. If you get only 50 results instead of ~900, pagination is broken. - **SCImago blocked**: `pa-fetch-scimago` will raise a clear error if anti-bot HTML is returned. Add data inline to `venues.yaml` instead. - **WikiCFP `` labels**: deadline detail pages use `` for label cells, not ``. The parser looks for `+` pairs. - **Calendar events**: `pa-generate` embeds events as YAML in the `events:` front matter field. The layout at `site/layouts/_default/calendar.html` reads `.Params.events` and initializes FullCalendar. Do not remove the `layout: calendar` front matter field. - **Deadline preservation**: never strip deadline fields from `deadlines.yaml` just because the submission window has closed. Remove an entry only once the conference has taken place. - **WikiCFP IDs are edition-specific**: each year's event gets a new ID. IDs verified at one point in time will be wrong once a new cycle begins. See Task 6 for the annual refresh checklist. --- ## Future Considerations - **iCal feed** — generate a `.ics` file from deadline data so researchers can subscribe from their calendar app. - **Email/RSS notifications** — alert when a deadline is within N weeks. - **Citation tracking** — periodically re-query OpenAlex for citation counts on digest papers. - **Best paper badge** — `pa-generate` could cross-reference `best_papers.yaml` with digest candidates and add a `best_paper_award: true` flag, then render a badge in `_digest_body()`. - **Automated WikiCFP ID discovery** — when a venue has no `wikicfp_id`, attempt a search and record the result in `venues.yaml` for future runs (reduces manual work when bootstrapping new topics). - **Cross-topic venue pages** — some venues (e.g., ASPLOS) span multiple topics. A future `cross_listed: [embedded-ai, cloud-edge]` field in `venues.yaml` could render the venue under multiple topic sections without duplicating data.