Files
publish-assistant/README.md
khannurien 8484abea47 Initial commit
Hugo/PaperMod static site tracking 13 conferences and 7 journals for
  edge and cloud systems research. Includes FullCalendar deadline view,
  ICORE/SCImago rankings, DBLP paper digest pipeline, and Python
  fetch/generate scripts. PaperMod added as a git submodule.
2026-04-24 11:52:33 +00:00

408 lines
22 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Publish Assistant
Hugo-generated website that helps researchers:
- identify important venues to publish in;
- keep track of submission deadlines;
- retrieve important papers for each issue.
---
## Overview
The site is domain-driven: a researcher picks a research area (e.g. "edge and cloud systems") and the assistant builds a curated, ranked, up-to-date snapshot of where to publish, when to submit, and what to read. The output is a static Hugo site that can be rebuilt on demand.
Work is split between **automated scripts** (data fetching, Hugo content generation) and **agent tasks** (domain curation, paper selection, deadline gap-filling). The README below documents both halves so the site can be kept fresh over time.
---
## Repository Layout
```
publish-assistant/
├── site/ # Hugo project
│ ├── hugo.toml
│ ├── themes/PaperMod/ # PaperMod theme (git submodule)
│ ├── layouts/
│ │ ├── _default/calendar.html # FullCalendar layout override
│ │ └── partials/extend_head.html
│ ├── content/ # Generated — do not edit by hand
│ │ ├── venues/
│ │ │ ├── _index.md # Venues overview (generated)
│ │ │ ├── conferences/ # One subdir per tracked conference
│ │ │ └── journals/ # One subdir per tracked journal
│ │ ├── calendar/ # Aggregated deadline view + FullCalendar
│ │ └── digests/ # Per-issue paper digests
│ └── data/ # Structured data — edit these
│ ├── venues.yaml # Master venue list ← primary edit target
│ ├── deadlines.yaml # Deadline cache; manual entries preserved
│ ├── best_papers.yaml # Best-paper awards (written by pa-fetch-best-papers)
│ ├── rankings/
│ │ ├── icore.csv # ICORE rankings cache
│ │ └── scimago.csv # SCImago rankings cache (optional)
│ └── papers/
│ ├── <V>-<Y>-candidates.yaml # Full paper list (pa-fetch-papers)
│ └── <V>-<Y>-digest.yaml # Curated selection (agent, Task 3)
├── src/publish_assistant/ # Python package
│ ├── fetch_icore.py
│ ├── fetch_scimago.py
│ ├── fetch_deadlines.py
│ ├── fetch_papers.py
│ ├── fetch_best_papers.py
│ └── generate_content.py
├── pyproject.toml # uv project; CLI entry points
├── uv.lock
├── build.sh
└── README.md
```
---
## Setup
```bash
uv sync # installs all dependencies + registers CLI tools
# Available commands after sync:
uv run pa-fetch-icore
uv run pa-fetch-scimago
uv run pa-fetch-deadlines
uv run pa-fetch-best-papers
uv run pa-fetch-papers --venue OSDI --year 2025
uv run pa-generate
# Local dev server:
hugo server --source site
```
---
## Data Sources
| Source | What it provides | Automatable? | Known issues |
| ---------------------------------------------------------------------------- | ---------------------------------------- | ------------ | --------------------------------------------------------------------------------------- |
| [ICORE](https://portal.core.edu.au/conf-ranks/) | Conference rankings (A*, A, B, C) | Yes | Pagination uses `javascript:jumpPage('N')` — handled in `fetch_icore.py` |
| [SCImago](https://www.scimagojr.com/) | Journal quartiles, SJR, H-index | Blocked | Anti-bot returns HTML; add data manually to `venues.yaml` under `scimago_quartile` etc. |
| [DBLP](https://dblp.org/) | Paper metadata by venue | Yes | Use `dblp_key` field in `venues.yaml` |
| [OpenAlex](https://openalex.org/) | Papers, open-access links | Yes | Fallback when DBLP is thin |
| [WikiCFP](http://wikicfp.com/) | Submission deadlines | Partially | See detailed notes below |
| [Conference websites](.) | Authoritative deadlines | Partially | See Task 2 |
| [jeffhuang.com/best_paper_awards/](https://jeffhuang.com/best_paper_awards/) | Best paper awards since 1996, ~32 venues | Yes | Manually maintained; run `pa-fetch-best-papers` annually |
---
## WikiCFP Integration — Detailed Notes
WikiCFP is the primary deadline source but has several quirks that required workarounds:
**HTML structure**: Detail pages use `<th>` for row labels (not `<td>`). The parser in `fetch_cfp_details()` specifically looks for `<th>` + `<td>` pairs. Do not revert to `find_all("td")` — it will find zero deadline rows.
**Direct ID lookup**: Add `wikicfp_id: "<event_id>"` to a venue entry in `venues.yaml` to skip the search and fetch that page directly. This avoids wrong matches on common acronyms. Verified IDs for this domain:
| Venue | WikiCFP event ID |
| ---------- | ---------------- |
| SOSP | 191399 |
| EuroSys | 186524 |
| SoCC | 191071 |
| Middleware | 190153 |
| IPDPS | 189093 |
| HPDC | 191029 |
**Skipping search**: Set `wikicfp_id: false` to skip WikiCFP entirely for a venue (e.g., ATC, SC, SEC — where the search returns wrong events). Deadlines for these must be filled manually.
**Conferences not on WikiCFP** (for edge/cloud systems domain): OSDI, NSDI, USENIX ATC, SC, MobiSys, SEC. Use `wikicfp_id: false` for all of them.
**Manual deadline entries**: Add entries with `source: manual` to `site/data/deadlines.yaml`. The fetcher preserves all `source: manual` entries across runs. Format:
```yaml
ATC:
source: manual
event_dates: Nov 16-18, 2026
location: Hong Kong
submission_deadline: Jun 10, 2026
notification: Sep 18, 2026
camera_ready: Oct 16, 2026
cfp_url: https://sigops.org/s/conferences/atc/2026/cfp.html
```
---
## `venues.yaml` Schema
```yaml
conferences:
- acronym: SOSP
full_name: "ACM Symposium on Operating Systems Principles"
domain: [edge-and-cloud, operating-systems, distributed-systems]
url: "https://sigops.org/s/conferences/sosp/"
dblp_key: "conf/sosp"
wikicfp_id: "191399" # direct lookup; omit to use search; false to skip entirely
journals:
- acronym: TPDS
full_name: "IEEE Transactions on Parallel and Distributed Systems"
domain: [edge-and-cloud, parallel-computing, distributed-systems]
issn: "1045-9219"
url: "https://www.computer.org/csdl/journal/td"
dblp_key: "journals/tpds"
submission_model: rolling
scimago_quartile: Q1 # add manually — SCImago CSV download is blocked
scimago_sjr: "1.560"
scimago_h_index: "131"
```
**Important**: Hugo reserves the front matter field `url` as a page URL override. `generate_content.py` maps `venues.yaml:url` → front matter field `homepage` to avoid this conflict.
---
## Scripts
All scripts are installed as CLI entry points by `uv sync`.
### `pa-fetch-icore`
Downloads ICORE rankings. Handles the portal's JavaScript-based pagination.
```bash
uv run pa-fetch-icore # fetch all
uv run pa-fetch-icore --query "distributed systems"
```
### `pa-fetch-scimago`
Downloads SCImago CSV. **Currently blocked by anti-bot.** Will raise a descriptive error if it receives HTML instead of CSV. Add journal data manually to `venues.yaml` instead.
```bash
uv run pa-fetch-scimago --list-areas # show area codes
uv run pa-fetch-scimago --area 1705 # networks
```
### `pa-fetch-deadlines`
Fetches submission deadlines from WikiCFP. Preserves `source: manual` entries.
```bash
uv run pa-fetch-deadlines
```
After running: check `site/data/deadlines.yaml` for the `missing:` list, then do Task 2.
### `pa-fetch-best-papers`
Scrapes [jeffhuang.com/best_paper_awards/](https://jeffhuang.com/best_paper_awards/) and writes `site/data/best_papers.yaml`. Run once per year (the source is updated annually).
```bash
uv run pa-fetch-best-papers
```
### `pa-fetch-papers`
Fetches paper lists from DBLP (with OpenAlex fallback).
```bash
uv run pa-fetch-papers --venue OSDI --year 2024
uv run pa-fetch-papers --venue TPDS --year 2024 --source openalex
```
### `pa-generate`
Regenerates all Hugo content from data files. Safe to re-run at any time.
```bash
uv run pa-generate
```
Preserved fields (never overwritten): `notes`, `deadline_source`.
Stripped fields (removed on regen to avoid stale data): `url` (Hugo reserved), `papers`.
### `build.sh`
Full pipeline orchestrator.
```bash
./build.sh
./build.sh --skip-rankings # skip icore/scimago fetches
./build.sh --skip-deadlines # use cached deadlines.yaml
./build.sh --dev # hugo server instead of build
```
---
## Hugo Content Structure
### Venues
**`site/content/venues/_index.md`** — overview, links to conferences and journals. Generated.
**`site/content/venues/conferences/_index.md`** — table of all conferences sorted by ICORE rank with deadlines. Generated.
**`site/content/venues/journals/_index.md`** — table of all journals sorted by SCImago quartile. Generated.
Each venue page body is **fully generated Markdown** — PaperMod renders body content, not front matter fields. The body includes a metadata table and a deadline section.
### Calendar
**`site/content/calendar/_index.md`** — uses `layout: calendar`, which activates `site/layouts/_default/calendar.html`. The layout renders a FullCalendar (CDN) month grid above the deadline table. Events are embedded as a JSON-ready YAML list in the `events` front matter field. Colors: orange = abstract deadline, red = paper deadline, blue = conference dates.
### Digests
**`site/content/digests/<VENUE>-<YEAR>/index.md`** — uses `index.md` (not `_index.md`) to be a leaf page, not a section. Body contains the full paper list with TL;DR and why-notable for each paper. Papers data lives in `site/data/papers/<V>-<Y>-digest.yaml`; do not put it in front matter (it was stripped for causing empty pages with PaperMod).
---
## Agent Instructions
Tasks requiring agent involvement (domain knowledge, judgment, web research).
---
### Task 1 — Bootstrap a domain
**Trigger:** User asks to set up tracking for a new research domain.
**Steps:**
1. Ask for the domain name and any seed venues.
2. Research: identify top 1015 conferences and 510 journals. Use [csrankings.org](https://csrankings.org), [ICORE portal](https://portal.core.edu.au/conf-ranks/), and [SCImago](https://www.scimagojr.com/) for rankings.
3. For each conference: record acronym, full name, ICORE rank, official 2025/2026 website URL, DBLP stream key (`conf/<key>`).
4. For each journal: record acronym, full name, ISSN, SCImago quartile + SJR (add inline to `venues.yaml` — CSV download is blocked), DBLP stream key (`journals/<key>`), submission model (rolling / special issues).
5. Append all venues to `site/data/venues.yaml`.
6. For WikiCFP: add `wikicfp_id: "<id>"` if you can find the event page (search at wikicfp.com). Set `wikicfp_id: false` for conferences where search returns wrong matches (short/common acronyms are risky).
7. Run `uv run pa-fetch-icore` and `uv run pa-fetch-deadlines`.
8. For conferences not found by the deadline fetcher, do Task 2 immediately.
9. Run `uv run pa-generate` and `hugo --source site` to validate.
**One-shot checklist for agents:**
- [ ] `site/data/venues.yaml` populated with all venues
- [ ] `wikicfp_id` set or `wikicfp_id: false` on every conference
- [ ] SCImago data added inline to every journal entry
- [ ] `uv run pa-fetch-icore` succeeded (check `site/data/rankings/icore.csv`)
- [ ] `uv run pa-fetch-deadlines` ran (check `site/data/deadlines.yaml`)
- [ ] All conferences in `deadlines.yaml:missing` handled via Task 2
- [ ] `uv run pa-generate` ran cleanly
- [ ] `hugo --source site --minify` built without errors
---
### Task 2 — Fill in missing deadlines
**Trigger:** `pa-fetch-deadlines` lists conferences under `missing:`, or a deadline looks wrong.
**Steps:**
1. For each missing conference, visit the official website. Conferences typically have a "Call for Papers" page with an "Important Dates" section.
2. Also check [WikiCFP](http://wikicfp.com/cfp/servlet/tool.search?q=<ACRONYM>&year=f) manually — if you find the right event ID, add `wikicfp_id` to `venues.yaml` so future runs fetch it automatically.
3. Extract: abstract deadline, paper deadline, notification, camera-ready, event dates, location.
4. Add a `source: manual` entry to `site/data/deadlines.yaml`. This entry survives future `pa-fetch-deadlines` runs.
5. Run `uv run pa-generate` to propagate.
**Conferences reliably NOT on WikiCFP** (for systems/networking):
- USENIX family: OSDI, NSDI, USENIX ATC, USENIX Security — use usenix.org directly
- SC (Supercomputing) — use sc<YY>.supercomputing.org/program/papers/
- MobiSys — use sigmobile.org/mobisys/<YEAR>/
- SEC (Edge Computing) — use acm-ieee-sec.org/<YEAR>/
- Short or common acronyms (ATC, SEC) collide with unrelated events — always use `wikicfp_id: false` and fetch manually
---
### Task 3 — Build a digest for a conference issue
**Trigger:** User asks to build a digest for a specific venue + year.
**Steps:**
1. Run `uv run pa-fetch-papers --venue <ACRONYM> --year <YEAR>` to get the candidate pool (`site/data/papers/<V>-<Y>-candidates.yaml`).
2. Check `site/data/best_papers.yaml` (run `pa-fetch-best-papers` first if it doesn't exist). Papers with matching titles in the best-papers list should be included and flagged.
3. Select 815 papers that are:
- Methodologically novel (new algorithms, systems designs, formal proofs);
- Attracting community attention (highly cited if the issue is ≥ 1 year old; in top venues / co-authored by known researchers if recent);
- Representative of the breadth of the issue (avoid over-indexing on one subtheme);
- Preferably open-access (arXiv, USENIX, ACM OpenTOC).
4. For each selected paper, write:
- `tldr`: one sentence, the core technical contribution.
- `why_notable`: 12 sentences — novelty, impact, surprising result, or influential technique.
5. Write `site/data/papers/<V>-<Y>-digest.yaml` with a `selected:` list.
6. Run `uv run pa-generate` — the digest page is created at `site/content/digests/<V>-<Y>/index.md`.
**Digest YAML format:**
```yaml
venue: OSDI
year: 2024
date: "2024-07-10"
tags: [llm-serving, distributed-systems, storage]
selected:
- dblp_key: "conf/osdi/ZhongLCHZL0024"
title: "DistServe: Disaggregating Prefill and Decoding ..."
tldr: "Separates prefill and decode onto different GPU pools, eliminating head-of-line blocking."
why_notable: "Became one of the most-cited LLM systems papers of 2024; disaggregation is now standard in production inference stacks."
```
---
### Task 4 — Refresh rankings
**Trigger:** ICORE releases a new round (every 23 years); SCImago releases new data (annually, each spring).
**Steps:**
1. Run `uv run pa-fetch-icore` for fresh ICORE data.
2. For SCImago: download the CSV manually from [scimagojr.com](https://www.scimagojr.com/journalrank.php) (CSV download button on the rankings page) and place it at `site/data/rankings/scimago.csv`. The automated fetch is blocked.
3. Check `site/data/venues.yaml` — for journals, compare `scimago_quartile` / `scimago_sjr` against the new CSV. Update inline values if changed.
4. Run `uv run pa-generate`.
---
### Task 5 — Add a new venue mid-cycle
**Steps:**
1. Look up ICORE rank (conferences) or SCImago quartile (journals).
2. Find the DBLP stream key at [dblp.org](https://dblp.org).
3. Add to `site/data/venues.yaml`.
4. For conferences: look up WikiCFP event ID or set `wikicfp_id: false`.
5. Run `uv run pa-fetch-deadlines` + `uv run pa-generate`.
---
### Task 6 — Annual cycle refresh
**Trigger:** A new conference cycle begins (roughly each autumn/spring depending on the venue). Signs: event URLs return 404s, WikiCFP fetches pull wrong editions, or deadlines are over a year old.
**Steps:**
1. **Update edition URLs** — for each conference in `venues.yaml`, check whether `url` points to the upcoming edition. Many venues use year-specific URLs (`osdi26`, `2027.eurosys.org`, `mobisys/2026/`). Update these to the new edition. Generic/stable URLs (e.g., `sigops.org/s/conferences/sosp/`) do not need changing.
2. **Refresh WikiCFP IDs** — for each conference with a `wikicfp_id`, verify the ID still matches the upcoming edition by visiting `http://wikicfp.com/cfp/servlet/event.showcfp?eventid=<ID>`. If it points to a past event, search WikiCFP for the new edition and update the ID. If the new event page does not exist yet, set `wikicfp_id: false` temporarily and add a `source: manual` deadline entry; restore the ID once the page appears.
3. Run `uv run pa-fetch-deadlines` and check the `missing:` list. Fill gaps via Task 2.
4. Run `uv run pa-generate` and `hugo --source site --minify` to validate.
**One-shot checklist for agents:**
- [ ] All conference `url` fields in `venues.yaml` point to the upcoming edition
- [ ] All `wikicfp_id` values verified against the upcoming edition (or set to `false` with a manual entry)
- [ ] `uv run pa-fetch-deadlines` ran cleanly; `missing:` list is empty
- [ ] `uv run pa-generate` + `hugo --source site --minify` succeed
---
## What the Build Script Cannot Do
| Task | Why automation fails | Agent task |
| ----------------------------------- | ---------------------------- | ---------- |
| Initial venue curation | Requires domain knowledge | Task 1 |
| Fetching missing deadlines | No standard CFP structure | Task 2 |
| Selecting notable papers | Requires reading + judgment | Task 3 |
| Writing `why_notable` | Requires synthesis | Task 3 |
| Detecting meaningful rank changes | Requires domain context | Task 4 |
| Evaluating new venues for inclusion | Requires community awareness | Task 5 |
---
## Known Gotchas (for agents picking this up)
- **Hugo `url` field**: reserved by Hugo to override the page URL. `venues.yaml` uses `url:` but `generate_content.py` maps it to `homepage:` in front matter. Never write `url:` in Hugo front matter via the generator.
- **PaperMod renders body, not front matter**: all visible content must be in the Markdown body (after the second `---`). Front matter is used only for metadata and Hugo taxonomy. If a page looks empty, check that `_conf_body()` / `_journal_body()` etc. are being called.
- **`_index.md` vs `index.md`**: section pages use `_index.md` (list template), leaf pages use `index.md` (single template). Digest pages are `index.md` — using `_index.md` makes them section pages and breaks pagination.
- **ICORE pagination**: the ICORE portal uses `javascript:jumpPage('N')` links, not standard `?page=N` URLs. `fetch_icore.py` handles this. If you get only 50 results instead of ~900, pagination is broken.
- **SCImago blocked**: `pa-fetch-scimago` will raise a clear error if anti-bot HTML is returned. Add data inline to `venues.yaml` instead.
- **WikiCFP `<th>` labels**: deadline detail pages use `<th>` for label cells, not `<td>`. The parser looks for `<th>+<td>` pairs. A `find_all("td")` approach finds nothing.
- **Calendar events**: `pa-generate` embeds events as a YAML list in the `events:` front matter field of `content/calendar/_index.md`. The custom layout at `site/layouts/_default/calendar.html` reads `.Params.events` and initializes FullCalendar. Do not remove the `layout: calendar` front matter field.
- **FullCalendar CDN**: loaded from `cdn.jsdelivr.net`. The `extend_head.html` partial injects the CSS; the layout injects the JS. Both are conditional on `layout == "calendar"`.
- **Deadline preservation**: never strip deadline fields from `deadlines.yaml` just because the submission window has closed. Keep all fields (abstract deadline, submission deadline, notification, camera-ready) for every conference whose event date is still in the future. Remove an entry only once the conference has taken place. An agent doing a "refresh" or "cleanup" must not treat a past submission deadline as stale data worth deleting.
- **WikiCFP IDs are edition-specific**: each year's event gets a new WikiCFP event ID. The IDs in the table above are for specific editions and will be wrong once a new cycle begins. When `pa-fetch-deadlines` returns stale or mismatched data, check whether the `wikicfp_id` in `venues.yaml` still points to the upcoming edition. If the new event page doesn't exist on WikiCFP yet, set `wikicfp_id: false` and add a `source: manual` entry; update the ID once the new page appears. See Task 6 for the full annual refresh checklist.
---
## Future Considerations
- **iCal feed** — generate a `.ics` file from deadline data so researchers can subscribe from their calendar app.
- **Email/RSS notifications** — alert when a deadline is within N weeks.
- **Citation tracking** — periodically re-query OpenAlex for citation counts on digest papers.
- **Multi-domain support** — a single repo tracking multiple domains; run the pipeline per domain with a `--domain` flag.
- **Best paper badge** — `pa-generate` could cross-reference `best_papers.yaml` with digest candidates and add a `best_paper_award: true` flag, then render a badge in `_digest_body()`.
- **Automated WikiCFP ID discovery** — when a venue has no `wikicfp_id`, attempt a search and record the result in `venues.yaml` for future runs (reduces manual work when bootstrapping new domains).