Initial commit

Hugo/PaperMod static site tracking 13 conferences and 7 journals for
  edge and cloud systems research. Includes FullCalendar deadline view,
  ICORE/SCImago rankings, DBLP paper digest pipeline, and Python
  fetch/generate scripts. PaperMod added as a git submodule.
This commit is contained in:
khannurien
2026-04-24 11:52:33 +00:00
commit 8484abea47
57 changed files with 20183 additions and 0 deletions

407
README.md Normal file
View File

@@ -0,0 +1,407 @@
# Publish Assistant
Hugo-generated website that helps researchers:
- identify important venues to publish in;
- keep track of submission deadlines;
- retrieve important papers for each issue.
---
## Overview
The site is domain-driven: a researcher picks a research area (e.g. "edge and cloud systems") and the assistant builds a curated, ranked, up-to-date snapshot of where to publish, when to submit, and what to read. The output is a static Hugo site that can be rebuilt on demand.
Work is split between **automated scripts** (data fetching, Hugo content generation) and **agent tasks** (domain curation, paper selection, deadline gap-filling). The README below documents both halves so the site can be kept fresh over time.
---
## Repository Layout
```
publish-assistant/
├── site/ # Hugo project
│ ├── hugo.toml
│ ├── themes/PaperMod/ # PaperMod theme (git submodule)
│ ├── layouts/
│ │ ├── _default/calendar.html # FullCalendar layout override
│ │ └── partials/extend_head.html
│ ├── content/ # Generated — do not edit by hand
│ │ ├── venues/
│ │ │ ├── _index.md # Venues overview (generated)
│ │ │ ├── conferences/ # One subdir per tracked conference
│ │ │ └── journals/ # One subdir per tracked journal
│ │ ├── calendar/ # Aggregated deadline view + FullCalendar
│ │ └── digests/ # Per-issue paper digests
│ └── data/ # Structured data — edit these
│ ├── venues.yaml # Master venue list ← primary edit target
│ ├── deadlines.yaml # Deadline cache; manual entries preserved
│ ├── best_papers.yaml # Best-paper awards (written by pa-fetch-best-papers)
│ ├── rankings/
│ │ ├── icore.csv # ICORE rankings cache
│ │ └── scimago.csv # SCImago rankings cache (optional)
│ └── papers/
│ ├── <V>-<Y>-candidates.yaml # Full paper list (pa-fetch-papers)
│ └── <V>-<Y>-digest.yaml # Curated selection (agent, Task 3)
├── src/publish_assistant/ # Python package
│ ├── fetch_icore.py
│ ├── fetch_scimago.py
│ ├── fetch_deadlines.py
│ ├── fetch_papers.py
│ ├── fetch_best_papers.py
│ └── generate_content.py
├── pyproject.toml # uv project; CLI entry points
├── uv.lock
├── build.sh
└── README.md
```
---
## Setup
```bash
uv sync # installs all dependencies + registers CLI tools
# Available commands after sync:
uv run pa-fetch-icore
uv run pa-fetch-scimago
uv run pa-fetch-deadlines
uv run pa-fetch-best-papers
uv run pa-fetch-papers --venue OSDI --year 2025
uv run pa-generate
# Local dev server:
hugo server --source site
```
---
## Data Sources
| Source | What it provides | Automatable? | Known issues |
| ---------------------------------------------------------------------------- | ---------------------------------------- | ------------ | --------------------------------------------------------------------------------------- |
| [ICORE](https://portal.core.edu.au/conf-ranks/) | Conference rankings (A*, A, B, C) | Yes | Pagination uses `javascript:jumpPage('N')` — handled in `fetch_icore.py` |
| [SCImago](https://www.scimagojr.com/) | Journal quartiles, SJR, H-index | Blocked | Anti-bot returns HTML; add data manually to `venues.yaml` under `scimago_quartile` etc. |
| [DBLP](https://dblp.org/) | Paper metadata by venue | Yes | Use `dblp_key` field in `venues.yaml` |
| [OpenAlex](https://openalex.org/) | Papers, open-access links | Yes | Fallback when DBLP is thin |
| [WikiCFP](http://wikicfp.com/) | Submission deadlines | Partially | See detailed notes below |
| [Conference websites](.) | Authoritative deadlines | Partially | See Task 2 |
| [jeffhuang.com/best_paper_awards/](https://jeffhuang.com/best_paper_awards/) | Best paper awards since 1996, ~32 venues | Yes | Manually maintained; run `pa-fetch-best-papers` annually |
---
## WikiCFP Integration — Detailed Notes
WikiCFP is the primary deadline source but has several quirks that required workarounds:
**HTML structure**: Detail pages use `<th>` for row labels (not `<td>`). The parser in `fetch_cfp_details()` specifically looks for `<th>` + `<td>` pairs. Do not revert to `find_all("td")` — it will find zero deadline rows.
**Direct ID lookup**: Add `wikicfp_id: "<event_id>"` to a venue entry in `venues.yaml` to skip the search and fetch that page directly. This avoids wrong matches on common acronyms. Verified IDs for this domain:
| Venue | WikiCFP event ID |
| ---------- | ---------------- |
| SOSP | 191399 |
| EuroSys | 186524 |
| SoCC | 191071 |
| Middleware | 190153 |
| IPDPS | 189093 |
| HPDC | 191029 |
**Skipping search**: Set `wikicfp_id: false` to skip WikiCFP entirely for a venue (e.g., ATC, SC, SEC — where the search returns wrong events). Deadlines for these must be filled manually.
**Conferences not on WikiCFP** (for edge/cloud systems domain): OSDI, NSDI, USENIX ATC, SC, MobiSys, SEC. Use `wikicfp_id: false` for all of them.
**Manual deadline entries**: Add entries with `source: manual` to `site/data/deadlines.yaml`. The fetcher preserves all `source: manual` entries across runs. Format:
```yaml
ATC:
source: manual
event_dates: Nov 16-18, 2026
location: Hong Kong
submission_deadline: Jun 10, 2026
notification: Sep 18, 2026
camera_ready: Oct 16, 2026
cfp_url: https://sigops.org/s/conferences/atc/2026/cfp.html
```
---
## `venues.yaml` Schema
```yaml
conferences:
- acronym: SOSP
full_name: "ACM Symposium on Operating Systems Principles"
domain: [edge-and-cloud, operating-systems, distributed-systems]
url: "https://sigops.org/s/conferences/sosp/"
dblp_key: "conf/sosp"
wikicfp_id: "191399" # direct lookup; omit to use search; false to skip entirely
journals:
- acronym: TPDS
full_name: "IEEE Transactions on Parallel and Distributed Systems"
domain: [edge-and-cloud, parallel-computing, distributed-systems]
issn: "1045-9219"
url: "https://www.computer.org/csdl/journal/td"
dblp_key: "journals/tpds"
submission_model: rolling
scimago_quartile: Q1 # add manually — SCImago CSV download is blocked
scimago_sjr: "1.560"
scimago_h_index: "131"
```
**Important**: Hugo reserves the front matter field `url` as a page URL override. `generate_content.py` maps `venues.yaml:url` → front matter field `homepage` to avoid this conflict.
---
## Scripts
All scripts are installed as CLI entry points by `uv sync`.
### `pa-fetch-icore`
Downloads ICORE rankings. Handles the portal's JavaScript-based pagination.
```bash
uv run pa-fetch-icore # fetch all
uv run pa-fetch-icore --query "distributed systems"
```
### `pa-fetch-scimago`
Downloads SCImago CSV. **Currently blocked by anti-bot.** Will raise a descriptive error if it receives HTML instead of CSV. Add journal data manually to `venues.yaml` instead.
```bash
uv run pa-fetch-scimago --list-areas # show area codes
uv run pa-fetch-scimago --area 1705 # networks
```
### `pa-fetch-deadlines`
Fetches submission deadlines from WikiCFP. Preserves `source: manual` entries.
```bash
uv run pa-fetch-deadlines
```
After running: check `site/data/deadlines.yaml` for the `missing:` list, then do Task 2.
### `pa-fetch-best-papers`
Scrapes [jeffhuang.com/best_paper_awards/](https://jeffhuang.com/best_paper_awards/) and writes `site/data/best_papers.yaml`. Run once per year (the source is updated annually).
```bash
uv run pa-fetch-best-papers
```
### `pa-fetch-papers`
Fetches paper lists from DBLP (with OpenAlex fallback).
```bash
uv run pa-fetch-papers --venue OSDI --year 2024
uv run pa-fetch-papers --venue TPDS --year 2024 --source openalex
```
### `pa-generate`
Regenerates all Hugo content from data files. Safe to re-run at any time.
```bash
uv run pa-generate
```
Preserved fields (never overwritten): `notes`, `deadline_source`.
Stripped fields (removed on regen to avoid stale data): `url` (Hugo reserved), `papers`.
### `build.sh`
Full pipeline orchestrator.
```bash
./build.sh
./build.sh --skip-rankings # skip icore/scimago fetches
./build.sh --skip-deadlines # use cached deadlines.yaml
./build.sh --dev # hugo server instead of build
```
---
## Hugo Content Structure
### Venues
**`site/content/venues/_index.md`** — overview, links to conferences and journals. Generated.
**`site/content/venues/conferences/_index.md`** — table of all conferences sorted by ICORE rank with deadlines. Generated.
**`site/content/venues/journals/_index.md`** — table of all journals sorted by SCImago quartile. Generated.
Each venue page body is **fully generated Markdown** — PaperMod renders body content, not front matter fields. The body includes a metadata table and a deadline section.
### Calendar
**`site/content/calendar/_index.md`** — uses `layout: calendar`, which activates `site/layouts/_default/calendar.html`. The layout renders a FullCalendar (CDN) month grid above the deadline table. Events are embedded as a JSON-ready YAML list in the `events` front matter field. Colors: orange = abstract deadline, red = paper deadline, blue = conference dates.
### Digests
**`site/content/digests/<VENUE>-<YEAR>/index.md`** — uses `index.md` (not `_index.md`) to be a leaf page, not a section. Body contains the full paper list with TL;DR and why-notable for each paper. Papers data lives in `site/data/papers/<V>-<Y>-digest.yaml`; do not put it in front matter (it was stripped for causing empty pages with PaperMod).
---
## Agent Instructions
Tasks requiring agent involvement (domain knowledge, judgment, web research).
---
### Task 1 — Bootstrap a domain
**Trigger:** User asks to set up tracking for a new research domain.
**Steps:**
1. Ask for the domain name and any seed venues.
2. Research: identify top 1015 conferences and 510 journals. Use [csrankings.org](https://csrankings.org), [ICORE portal](https://portal.core.edu.au/conf-ranks/), and [SCImago](https://www.scimagojr.com/) for rankings.
3. For each conference: record acronym, full name, ICORE rank, official 2025/2026 website URL, DBLP stream key (`conf/<key>`).
4. For each journal: record acronym, full name, ISSN, SCImago quartile + SJR (add inline to `venues.yaml` — CSV download is blocked), DBLP stream key (`journals/<key>`), submission model (rolling / special issues).
5. Append all venues to `site/data/venues.yaml`.
6. For WikiCFP: add `wikicfp_id: "<id>"` if you can find the event page (search at wikicfp.com). Set `wikicfp_id: false` for conferences where search returns wrong matches (short/common acronyms are risky).
7. Run `uv run pa-fetch-icore` and `uv run pa-fetch-deadlines`.
8. For conferences not found by the deadline fetcher, do Task 2 immediately.
9. Run `uv run pa-generate` and `hugo --source site` to validate.
**One-shot checklist for agents:**
- [ ] `site/data/venues.yaml` populated with all venues
- [ ] `wikicfp_id` set or `wikicfp_id: false` on every conference
- [ ] SCImago data added inline to every journal entry
- [ ] `uv run pa-fetch-icore` succeeded (check `site/data/rankings/icore.csv`)
- [ ] `uv run pa-fetch-deadlines` ran (check `site/data/deadlines.yaml`)
- [ ] All conferences in `deadlines.yaml:missing` handled via Task 2
- [ ] `uv run pa-generate` ran cleanly
- [ ] `hugo --source site --minify` built without errors
---
### Task 2 — Fill in missing deadlines
**Trigger:** `pa-fetch-deadlines` lists conferences under `missing:`, or a deadline looks wrong.
**Steps:**
1. For each missing conference, visit the official website. Conferences typically have a "Call for Papers" page with an "Important Dates" section.
2. Also check [WikiCFP](http://wikicfp.com/cfp/servlet/tool.search?q=<ACRONYM>&year=f) manually — if you find the right event ID, add `wikicfp_id` to `venues.yaml` so future runs fetch it automatically.
3. Extract: abstract deadline, paper deadline, notification, camera-ready, event dates, location.
4. Add a `source: manual` entry to `site/data/deadlines.yaml`. This entry survives future `pa-fetch-deadlines` runs.
5. Run `uv run pa-generate` to propagate.
**Conferences reliably NOT on WikiCFP** (for systems/networking):
- USENIX family: OSDI, NSDI, USENIX ATC, USENIX Security — use usenix.org directly
- SC (Supercomputing) — use sc<YY>.supercomputing.org/program/papers/
- MobiSys — use sigmobile.org/mobisys/<YEAR>/
- SEC (Edge Computing) — use acm-ieee-sec.org/<YEAR>/
- Short or common acronyms (ATC, SEC) collide with unrelated events — always use `wikicfp_id: false` and fetch manually
---
### Task 3 — Build a digest for a conference issue
**Trigger:** User asks to build a digest for a specific venue + year.
**Steps:**
1. Run `uv run pa-fetch-papers --venue <ACRONYM> --year <YEAR>` to get the candidate pool (`site/data/papers/<V>-<Y>-candidates.yaml`).
2. Check `site/data/best_papers.yaml` (run `pa-fetch-best-papers` first if it doesn't exist). Papers with matching titles in the best-papers list should be included and flagged.
3. Select 815 papers that are:
- Methodologically novel (new algorithms, systems designs, formal proofs);
- Attracting community attention (highly cited if the issue is ≥ 1 year old; in top venues / co-authored by known researchers if recent);
- Representative of the breadth of the issue (avoid over-indexing on one subtheme);
- Preferably open-access (arXiv, USENIX, ACM OpenTOC).
4. For each selected paper, write:
- `tldr`: one sentence, the core technical contribution.
- `why_notable`: 12 sentences — novelty, impact, surprising result, or influential technique.
5. Write `site/data/papers/<V>-<Y>-digest.yaml` with a `selected:` list.
6. Run `uv run pa-generate` — the digest page is created at `site/content/digests/<V>-<Y>/index.md`.
**Digest YAML format:**
```yaml
venue: OSDI
year: 2024
date: "2024-07-10"
tags: [llm-serving, distributed-systems, storage]
selected:
- dblp_key: "conf/osdi/ZhongLCHZL0024"
title: "DistServe: Disaggregating Prefill and Decoding ..."
tldr: "Separates prefill and decode onto different GPU pools, eliminating head-of-line blocking."
why_notable: "Became one of the most-cited LLM systems papers of 2024; disaggregation is now standard in production inference stacks."
```
---
### Task 4 — Refresh rankings
**Trigger:** ICORE releases a new round (every 23 years); SCImago releases new data (annually, each spring).
**Steps:**
1. Run `uv run pa-fetch-icore` for fresh ICORE data.
2. For SCImago: download the CSV manually from [scimagojr.com](https://www.scimagojr.com/journalrank.php) (CSV download button on the rankings page) and place it at `site/data/rankings/scimago.csv`. The automated fetch is blocked.
3. Check `site/data/venues.yaml` — for journals, compare `scimago_quartile` / `scimago_sjr` against the new CSV. Update inline values if changed.
4. Run `uv run pa-generate`.
---
### Task 5 — Add a new venue mid-cycle
**Steps:**
1. Look up ICORE rank (conferences) or SCImago quartile (journals).
2. Find the DBLP stream key at [dblp.org](https://dblp.org).
3. Add to `site/data/venues.yaml`.
4. For conferences: look up WikiCFP event ID or set `wikicfp_id: false`.
5. Run `uv run pa-fetch-deadlines` + `uv run pa-generate`.
---
### Task 6 — Annual cycle refresh
**Trigger:** A new conference cycle begins (roughly each autumn/spring depending on the venue). Signs: event URLs return 404s, WikiCFP fetches pull wrong editions, or deadlines are over a year old.
**Steps:**
1. **Update edition URLs** — for each conference in `venues.yaml`, check whether `url` points to the upcoming edition. Many venues use year-specific URLs (`osdi26`, `2027.eurosys.org`, `mobisys/2026/`). Update these to the new edition. Generic/stable URLs (e.g., `sigops.org/s/conferences/sosp/`) do not need changing.
2. **Refresh WikiCFP IDs** — for each conference with a `wikicfp_id`, verify the ID still matches the upcoming edition by visiting `http://wikicfp.com/cfp/servlet/event.showcfp?eventid=<ID>`. If it points to a past event, search WikiCFP for the new edition and update the ID. If the new event page does not exist yet, set `wikicfp_id: false` temporarily and add a `source: manual` deadline entry; restore the ID once the page appears.
3. Run `uv run pa-fetch-deadlines` and check the `missing:` list. Fill gaps via Task 2.
4. Run `uv run pa-generate` and `hugo --source site --minify` to validate.
**One-shot checklist for agents:**
- [ ] All conference `url` fields in `venues.yaml` point to the upcoming edition
- [ ] All `wikicfp_id` values verified against the upcoming edition (or set to `false` with a manual entry)
- [ ] `uv run pa-fetch-deadlines` ran cleanly; `missing:` list is empty
- [ ] `uv run pa-generate` + `hugo --source site --minify` succeed
---
## What the Build Script Cannot Do
| Task | Why automation fails | Agent task |
| ----------------------------------- | ---------------------------- | ---------- |
| Initial venue curation | Requires domain knowledge | Task 1 |
| Fetching missing deadlines | No standard CFP structure | Task 2 |
| Selecting notable papers | Requires reading + judgment | Task 3 |
| Writing `why_notable` | Requires synthesis | Task 3 |
| Detecting meaningful rank changes | Requires domain context | Task 4 |
| Evaluating new venues for inclusion | Requires community awareness | Task 5 |
---
## Known Gotchas (for agents picking this up)
- **Hugo `url` field**: reserved by Hugo to override the page URL. `venues.yaml` uses `url:` but `generate_content.py` maps it to `homepage:` in front matter. Never write `url:` in Hugo front matter via the generator.
- **PaperMod renders body, not front matter**: all visible content must be in the Markdown body (after the second `---`). Front matter is used only for metadata and Hugo taxonomy. If a page looks empty, check that `_conf_body()` / `_journal_body()` etc. are being called.
- **`_index.md` vs `index.md`**: section pages use `_index.md` (list template), leaf pages use `index.md` (single template). Digest pages are `index.md` — using `_index.md` makes them section pages and breaks pagination.
- **ICORE pagination**: the ICORE portal uses `javascript:jumpPage('N')` links, not standard `?page=N` URLs. `fetch_icore.py` handles this. If you get only 50 results instead of ~900, pagination is broken.
- **SCImago blocked**: `pa-fetch-scimago` will raise a clear error if anti-bot HTML is returned. Add data inline to `venues.yaml` instead.
- **WikiCFP `<th>` labels**: deadline detail pages use `<th>` for label cells, not `<td>`. The parser looks for `<th>+<td>` pairs. A `find_all("td")` approach finds nothing.
- **Calendar events**: `pa-generate` embeds events as a YAML list in the `events:` front matter field of `content/calendar/_index.md`. The custom layout at `site/layouts/_default/calendar.html` reads `.Params.events` and initializes FullCalendar. Do not remove the `layout: calendar` front matter field.
- **FullCalendar CDN**: loaded from `cdn.jsdelivr.net`. The `extend_head.html` partial injects the CSS; the layout injects the JS. Both are conditional on `layout == "calendar"`.
- **Deadline preservation**: never strip deadline fields from `deadlines.yaml` just because the submission window has closed. Keep all fields (abstract deadline, submission deadline, notification, camera-ready) for every conference whose event date is still in the future. Remove an entry only once the conference has taken place. An agent doing a "refresh" or "cleanup" must not treat a past submission deadline as stale data worth deleting.
- **WikiCFP IDs are edition-specific**: each year's event gets a new WikiCFP event ID. The IDs in the table above are for specific editions and will be wrong once a new cycle begins. When `pa-fetch-deadlines` returns stale or mismatched data, check whether the `wikicfp_id` in `venues.yaml` still points to the upcoming edition. If the new event page doesn't exist on WikiCFP yet, set `wikicfp_id: false` and add a `source: manual` entry; update the ID once the new page appears. See Task 6 for the full annual refresh checklist.
---
## Future Considerations
- **iCal feed** — generate a `.ics` file from deadline data so researchers can subscribe from their calendar app.
- **Email/RSS notifications** — alert when a deadline is within N weeks.
- **Citation tracking** — periodically re-query OpenAlex for citation counts on digest papers.
- **Multi-domain support** — a single repo tracking multiple domains; run the pipeline per domain with a `--domain` flag.
- **Best paper badge** — `pa-generate` could cross-reference `best_papers.yaml` with digest candidates and add a `best_paper_award: true` flag, then render a badge in `_digest_body()`.
- **Automated WikiCFP ID discovery** — when a venue has no `wikicfp_id`, attempt a search and record the result in `venues.yaml` for future runs (reduces manual work when bootstrapping new domains).