Contract and workflow
Read mdto spec podcast (or the adjacent SPEC.md in a checkout) and the
shared conventions before authoring or
repairing content. Outside a checkout, use the published spec.
This skill accompanies [email protected]. The spec and its fixtures govern conformance;
editorial advice below is non-normative and should adapt to the user's brief.
If the source declares another version, obtain its matching contract before changing it.
Use markdownto: [email protected] and exactly two speaker mappings with stable id, displayname, and distinct supported voice values. Use mdto podcast voices to inspect the
installed voice catalogue. Label every turn with a declared ID or exact display name;
bold labels are allowed. Keep IDs stable when changing display names.
Level-2 headings divide production sections and are not spoken. Write transitions into
the dialogue. Unlabeled prose, lists, tables, and code are not dialogue. Put non-spoken
notes in > [production:: note] blockquotes; a standalone [pause:: 750ms] adds exact
silence. Use supported style, language, and pronunciations frontmatter for direction,
not invented stage-direction syntax. Production must use native multi-speaker synthesis;
alternating independent single-voice clips does not implement this spec.
Choose the show's personality
Before drafting, choose a voice for this audience and this episode's purpose. Use the
user's brief and reference episodes first; do not assign a personality from the topic
alone. If the brief leaves room, make a reasonable choice and state it briefly with the
draft. Ask only when a missing choice would materially change the episode, such as
whether a religious discussion is devotional within a tradition or comparative education.
Make a compact editorial brief: who is listening and what they already know; the episode's
promise; the hosts' relationship and different contributions; the desired energy, warmth,
humor, and degree of formality; and how disagreement should sound. Turn adjectives into
writing decisions: “curious” means following up on an unanswered question; “skeptical”
means testing a claim with evidence; “playful” might mean a running callback rather than
a joke on every turn. A voice preset alone does not create a personality.
Possible directions, not genre rules:
| Context | A possible personality | How it changes the script |
|---|---|---|
| Sports | Two knowledgeable fans with affectionate rivalry | Brisk exchanges, specific plays, earned excitement, and predictions tested against evidence. A tactical explainer might instead be patient and analytical. |
| Business | Curious operators who challenge easy success stories | Concrete decisions, incentives, tradeoffs, and respectful skepticism; explain the mechanism behind the headline. A founder portrait can be warmer and more reflective. |
| Religious content | Reflective companions speaking from a specified tradition, or curious comparative educators | Choose the intended stance; make room for reflection and distinguish texts, interpretations, and personal belief. Match humor and disagreement to the audience rather than assuming solemnity or irreverence. |
Keep the chosen personality through explanations, transitions, and the ending. Let energy
change with the material without making the hosts unrecognizable. Do not invent a host's
credentials, lived experiences, or beliefs to sell the voice. Express concise performable
direction in the existing style field; keep a longer editorial brief outside the spoken
dialogue or in a supported production note. Do not add personality fields to the grammar.
For example, a business explainer could use:
style: Two curious colleagues working through a business puzzle; warm, conversational, lightly playful, with crisp questions and calm, evidence-led disagreement.Build the teaching before the dialogue
For educational episodes, establish a source-grounded content brief before drafting:
the central question, the explanation, the evidence that carries it, important uncertainty,
and the strongest counter-case. For a small request this can be a compact working outline;
a separate research document is useful for a substantial deep dive, not a required output.
Adapt this material into spoken dialogue. Do not introduce unsupported facts while making
it more entertaining. If multiple formats are requested, derive them from the same factual
base so the article, script, and visual treatment do not quietly disagree.
Calibrate to what this listener knows. A useful educational default, when the user gives
no other audience, is an intelligent, curious non-specialist. Teach the missing mechanism,
not every prerequisite: get to the live puzzle early and explain a specialized term in
one compact sentence when it unlocks the next piece of reasoning. Do not make an informed
listener sit through everyday definitions, or assume a beginner already knows the jargon.
Earn attention with a concrete way in, the central question, the payoff, and a brief sense
of the journey. Add a dated “why now” when a current development actually matters; evergreen
subjects do not need manufactured urgency. In a long-form episode, the opening minute can
usually carry this promise without a roll call of definitions or an elaborate show intro.
Plan an arc that can support the requested length: live question → the history or context
that explains it → how the mechanism works → concrete evidence → the strongest challenge
→ synthesis and what would change the conclusion. Adapt or omit acts that do not serve
the episode. Depth comes from examples, causes, consequences, and pressure tests, not
from padding a short explanation to an arbitrary hour.
Make the mechanism concrete before relying on abstract labels: who is doing what, what
problem they face, and what changes as a result. For a business, that might mean what it
sells, who buys it, why they choose it, and why they might leave. In other subjects, choose
an equally tangible example that lets the listener follow the causal chain.
Make each important number answer a named question. For example, a possession statistic
might help ask whether a team controlled the match, and a cash-flow measure might help
ask whether demand turns into cash. Explain the connection and its limits; do not read
out a table of figures. Give a plain-language rephrasing after a dense idea, and distinguish
observed evidence from interpretation. Date time-sensitive claims. For an analytical
question, land the conclusion the evidence supports, name the strongest counter-case, and
say what would change the view. A reflective or devotional episode may instead end with
an earned question, practice, or contemplation; do not impose a business-thesis ending.
Write a conversation with substance
For the educational cohost format, use two informed contributors: both teach, introduce
evidence, challenge assumptions, and carry parts of the explanation. Neither is merely
the expert's dictionary prompt or an agreement machine. Other relationships, such as an
interviewer and guest, may intentionally distribute the work differently. Each turn should
respond to the previous thought or move the inquiry forward; disagreement should clarify
a real tension rather than manufacture drama.
Give a host enough room to finish a thought. In a relaxed educational deep dive, many
turns may be around 40–90 words, with shorter questions, reactions, or interruptions only
where they earn a handoff. Use this as a diagnostic, not a quota: a lively sports exchange
or a reflective conversation can need another rhythm. Avoid both alternating miniature
essays and forced one-line ping-pong. Use natural contractions and occasional fragments
when appropriate to the language and personality; do not add filler to sound human.
Use callbacks to make earlier ideas do new work. Cut repeated agreement, repetitive
introductions, forced jokes, and recaps that only replay the last turn. Write transitions
into the dialogue because chapter headings are not spoken. Let the ending deliver the
opening promise in the show's chosen voice.
The episode must stand on its own for someone who cannot see the research. Attribute
consequential claims to real, findable sources in natural spoken language. Keep URLs,
raw citation tags, wikilinks, and references to an internal brief or knowledge base out
of spoken turns. Preserve detailed provenance in source notes or supported non-spoken
production notes. Do not invent quotations or claim certainty the sources do not provide.
Editorial pass before audio
Re-read the script as an editor, separately from drafting. Check that the opening earns
attention, explanations match the audience, both hosts do useful work, evidence supports
the conclusion, and the ending pays off the promise. Check that the chosen personality
is audible in the actual wording, not only declared in style.
Inspect spoken word count, turn lengths, and per-host word share when useful. In an equal
cohost show, roughly 40–60% per host is a useful imbalance warning; an interview need not
meet it. Numbers identify passages to review, not automatic rewrites. Read the opening,
one dense exchange, a transition, and the ending aloud; fix awkward phrasing and handoffs.
An audition can then test delivery when audio is requested and authorized. Do not mistake
parser success, rhythm statistics, or a mock render for a listening-quality review.
Check and deliver
Validate the file with mdto validate <path> before editing existing content and again
after changes. Read diagnostics against the spec's rationale; preserve meaning when
repairing syntax. Prefer spec-owned verbs for edits they support, checking their current--help; otherwise make a focused source edit, preserving unrelated content and IDs.
Render with mdto render <path> and inspect the result when available. Validation checks
structure, not editorial quality: review the content against the user's purpose as well.
Without the CLI, use the contract and examples and report that validation was not run.
Deliver the Markdown source and any requested artifacts, stating what was actually checked.
Before paid generation or audition, inspect the command's help and estimate the requested
scope. Use existing authorization when it covers the cost; otherwise show the estimate
and obtain approval before spending. Supply --yes only within that authorization.
A text draft does not require audio generation. Mock output checks the pipeline, not
spoken quality; never present it as a completed listening review.