speech-paper-daily-skill: The grumpy AI gatekeeper for speech research

A Markdown skill that turns arXiv monitoring into a checkpointed review pipeline, with a skeptical reviewer persona, recovery-safe caching, and doc publishing baked in.

8 min read · JusperLee/speech-paper-daily-skill

A stern customs checkpoint for tiny academic papers on a conveyor belt, with some papers stamped KEEP and others diverted into a reject chute. The scene explains that this repo is an opinionated filter for speech research, not a neutral feed reader.
The project's core move is editorial, not mechanical: it decides what deserves a human's attention.
Key Takeaways

It does not summarize papers. It ranks them

Most paper tools promise coverage. This one promises judgment. In speech-paper-daily-skill, the point is not to read every speech paper, but to sort the daily pile into the few that deserve a scientist's attention.

That is why the persona matters. The skill tells the agent to act like a skeptical senior reviewer, the kind who spots skin-deep novelty, weak experiments, and inflated claims quickly. A neutral voice would be nicer. It would also be less useful.

Why speech research needs a bouncer

Speech and audio research move fast, but not every new paper moves the field. A lot of the daily firehose is incremental, tightly scoped, or simply not relevant to the people who need to keep up. The repo is built around that asymmetry.

So the skill does not behave like a search engine. It behaves like a bouncer at a crowded venue, checking the list, ignoring the noise, and sending the wrong people back out the door. That is a sharper product idea than "paper summary" because it starts from a constraint: attention is scarce.

A researcher stands before a crushing wave of loose papers on one side, while the other side shows a small stamped tray of only a few selected papers. The contrast explains how the repo turns overflow into a narrow, trusted reading queue.
The project's value is not breadth. It is ruthless reduction.

The skill file is the product

The interesting part of this repository is that the product lives in skill.md. The file is not just a set of instructions. It is a workflow spec that defines source discovery, domain filtering, reviewer stance, file checkpoints, publication targets, and recovery behavior.

1. Discover papers from papers.cool
2. Fall back to arXiv when needed
3. Filter for speech and audio relevance
4. Review as a toxic senior reviewer
5. Save after every paper to /tmp/papers_YYYYMMDD/<index>_<arxiv_id>.md
6. Publish to Tencent Docs or Feishu
7. Resume from the last saved checkpoint if the run stops

That reads like a state machine because it is one. The prompt decides what enters, what gets rejected, where progress is persisted, and how the system resumes after interruption. In practice, that turns Markdown into orchestration.

The workflow is not a single pass. It is a recoverable pipeline with memory.

How the workflow survives real-world mess

The reliability story is where this repo stops feeling like a clever prompt and starts feeling like production engineering. The skill writes each paper to a temp file as it goes, so a crash or timeout does not erase the entire session. That matters when a daily pipeline is long, brittle, and full of API edges.

python3 -c 'import sys,json; print(json.dumps(sys.stdin.read()))'
sleep 1
path=/tmp/papers_20260330/03_asr1234.md

The small details tell the story. The skill includes fallback sourcing, deliberate pauses to avoid rate limits, and output escaping so Tencent Docs does not choke on content formatting. These are the sorts of fixes that usually live in a backend service, but here they are encoded in the prompt itself.

A close-up of file folders, a temp path label, and a taut thread running from a paper to a saved checkpoint. The scene explains how the repo preserves progress after each paper so the workflow can recover instead of restarting.
The hidden feature is not speed. It is survivability.

What it turns paper reading into

The endpoint is not a blog post and not a spreadsheet. It is a collaborative intelligence feed in Tencent Docs or Feishu, updated with a daily rhythm and written in a voice that already did the hard filtering. The reader gets a shortlist, not a firehose.

That changes the social contract of paper reading. Instead of asking a researcher to scan dozens of abstracts, the repo delivers a review trail with opinions attached. It is a quiet shift from aggregation to editorial control.

开源的每日 arXiv / OpenReview 论文推荐与 AI 阅读平台,支持 GitHub Actions 自动更新、GitHub Pages 部署与论文问答。

GitHub Repository Description, Project Documentation · ziwenhahaha/daily-paper-reader

How it compares

ProjectPrimary functionJudgment styleFailure handlingOutput destinationWhy it differs
speech-paper-daily-skillDaily speech paper reviewExplicitly skeptical and opinionatedCheckpointed temp files and resume logicTencent Docs or FeishuIt tries to decide what is worth reading, not just collect it.
daily-paper-readerGeneral paper recommendation and Q&AHelpful and broadPlatform-level automationGitHub Pages and GitHub ActionsIt is a fuller reading product, built for browsing and interaction.
arXiv Sanity / Hugging Face PapersPaper discovery and rankingMostly passive and user-drivenExternal to the reading loopWeb interfaceThey help you browse a stream. They do not act like a reviewer.

The contrast is blunt. The broader tools help you find papers. This repo tries to decide whether a paper deserves a slot on a busy human's calendar. That makes it feel less like an app and more like an editorial policy rendered as code.

Why the approach matters beyond speech

The domain does not have to be speech for this pattern to matter. Any field with high volume, weak signal, and tired experts can benefit from a prompt that encodes taste, reliability, and delivery in one place. The point is not to automate curiosity. It is to protect it.

That is the real lesson of speech-paper-daily-skill. Prompting is not just for tone or style when the model is already in the loop. It can become infrastructure when it tells an agent how to judge, how to persist, and how to recover.