speech-paper-daily-skill: The grumpy AI gatekeeper for speech research
A Markdown skill that turns arXiv monitoring into a checkpointed review pipeline, with a skeptical reviewer persona, recovery-safe caching, and doc publishing baked in.
- The repo's real invention is judgment, because it turns an AI into a skeptical gatekeeper instead of a passive summarizer.
- Its workflow matters as much as its wording, because checkpointing and fallback sourcing make the prompt behave like infrastructure.
- The harsh reviewer persona is a feature, because speech research needs a filter that can dismiss incremental work fast.
- The broader lesson is that prompts can encode operations, reliability, and delivery, not just tone.
It does not summarize papers. It ranks them
Most paper tools promise coverage. This one promises judgment. In speech-paper-daily-skill, the point is not to read every speech paper, but to sort the daily pile into the few that deserve a scientist's attention.
That is why the persona matters. The skill tells the agent to act like a skeptical senior reviewer, the kind who spots skin-deep novelty, weak experiments, and inflated claims quickly. A neutral voice would be nicer. It would also be less useful.
Why speech research needs a bouncer
Speech and audio research move fast, but not every new paper moves the field. A lot of the daily firehose is incremental, tightly scoped, or simply not relevant to the people who need to keep up. The repo is built around that asymmetry.
So the skill does not behave like a search engine. It behaves like a bouncer at a crowded venue, checking the list, ignoring the noise, and sending the wrong people back out the door. That is a sharper product idea than "paper summary" because it starts from a constraint: attention is scarce.
The skill file is the product
The interesting part of this repository is that the product lives in skill.md. The file is not just a set of instructions. It is a workflow spec that defines source discovery, domain filtering, reviewer stance, file checkpoints, publication targets, and recovery behavior.
1. Discover papers from papers.cool
2. Fall back to arXiv when needed
3. Filter for speech and audio relevance
4. Review as a toxic senior reviewer
5. Save after every paper to /tmp/papers_YYYYMMDD/<index>_<arxiv_id>.md
6. Publish to Tencent Docs or Feishu
7. Resume from the last saved checkpoint if the run stops
That reads like a state machine because it is one. The prompt decides what enters, what gets rejected, where progress is persisted, and how the system resumes after interruption. In practice, that turns Markdown into orchestration.
How the workflow survives real-world mess
The reliability story is where this repo stops feeling like a clever prompt and starts feeling like production engineering. The skill writes each paper to a temp file as it goes, so a crash or timeout does not erase the entire session. That matters when a daily pipeline is long, brittle, and full of API edges.
python3 -c 'import sys,json; print(json.dumps(sys.stdin.read()))'
sleep 1
path=/tmp/papers_20260330/03_asr1234.md
The small details tell the story. The skill includes fallback sourcing, deliberate pauses to avoid rate limits, and output escaping so Tencent Docs does not choke on content formatting. These are the sorts of fixes that usually live in a backend service, but here they are encoded in the prompt itself.
What it turns paper reading into
The endpoint is not a blog post and not a spreadsheet. It is a collaborative intelligence feed in Tencent Docs or Feishu, updated with a daily rhythm and written in a voice that already did the hard filtering. The reader gets a shortlist, not a firehose.
That changes the social contract of paper reading. Instead of asking a researcher to scan dozens of abstracts, the repo delivers a review trail with opinions attached. It is a quiet shift from aggregation to editorial control.
开源的每日 arXiv / OpenReview 论文推荐与 AI 阅读平台,支持 GitHub Actions 自动更新、GitHub Pages 部署与论文问答。
How it compares
| Project | Primary function | Judgment style | Failure handling | Output destination | Why it differs |
|---|---|---|---|---|---|
| speech-paper-daily-skill | Daily speech paper review | Explicitly skeptical and opinionated | Checkpointed temp files and resume logic | Tencent Docs or Feishu | It tries to decide what is worth reading, not just collect it. |
| daily-paper-reader | General paper recommendation and Q&A | Helpful and broad | Platform-level automation | GitHub Pages and GitHub Actions | It is a fuller reading product, built for browsing and interaction. |
| arXiv Sanity / Hugging Face Papers | Paper discovery and ranking | Mostly passive and user-driven | External to the reading loop | Web interface | They help you browse a stream. They do not act like a reviewer. |
The contrast is blunt. The broader tools help you find papers. This repo tries to decide whether a paper deserves a slot on a busy human's calendar. That makes it feel less like an app and more like an editorial policy rendered as code.
Why the approach matters beyond speech
The domain does not have to be speech for this pattern to matter. Any field with high volume, weak signal, and tired experts can benefit from a prompt that encodes taste, reliability, and delivery in one place. The point is not to automate curiosity. It is to protect it.
That is the real lesson of speech-paper-daily-skill. Prompting is not just for tone or style when the model is already in the loop. It can become infrastructure when it tells an agent how to judge, how to persist, and how to recover.