VoiceDrop Home Developers 中文 Get it on iOS
Help Center

VoiceDrop Help Center

Speak it, it writes itself. This page covers every VoiceDrop feature — how to use it, the rules behind it, and the questions people ask most. From recording a voice memo to mining it into articles, editing, illustrating, sharing, and joining the community — all on one page.

🎙

Recording & Auto-Writing (Mining)

Record a quick voice memo, and VoiceDrop automatically transcribes it and "mines" it into one or more publish-ready articles.

This is the core VoiceDrop workflow: no typing, no organizing. Just say what's on your mind into your phone, and the app handles everything else — transcription, structuring, and turning it into a finished article.

On the "My Recordings" home screen, a red record button floats at the bottom. Tap it to enter the full-screen recording view and start recording. When you're done, tap Stop, and the recording is automatically queued for upload. Once uploaded, the server first transcribes your speech into text (Transcribing), then has AI mine your spoken words into a fully structured article (Mining) — with a title, an opening, and an ending, reading like an article you wrote yourself.

You don't have to watch any of this happen. Each recording moves through the statuses Uploading → Queued → Transcribing / Mining → Done in the list. Once done, tap it to hear the original audio, read the article, and publish with one tap. If no speech is detected in the recording, it's marked No Speech; if it's too fragmented to form a coherent piece, it shows "No article yet" — try again with a more complete thought.

How to use it

  1. Open the app and stay on the "My Recordings" tab at the bottom.
  2. Tap the red record button at the bottom. The screen switches to a full-screen recording view and recording starts immediately — "Recording" at the top, a large timer and a bouncing waveform in the middle. (Note: recording is tap-to-start, not press-and-hold. The "hold to speak" label on the red button is a different feature — holding the red button lets you give voice commands to the whole list, like "delete the second one".)
  3. To finish, tap the square Stop button in the middle. The recording view closes, you're back at the list, and the recording shows "Uploading".
  4. Want photos with your recording? Tap the light-colored camera icon on the right side of the recording view to snap a few shots. The photos upload together with the recording and get woven into the matching text by AI (taking photos doesn't interrupt recording).
  5. Then do nothing. The recording moves automatically from "Queued" through "Transcribing" and "Mining", and finally turns into a green "Done".
  6. Tap a "Done" recording to play back the original audio on top and read the mined article below. The ⋯ menu in the top right lets you "Publish to WeChat Drafts" or share.
  7. To edit the article: open the article view, hold the "Hold to Speak · Edit" talk bar at the bottom and speak your edit (like "delete the second paragraph"), then release — done.
  8. To have AI mine it again from scratch: in the list, long-press the "Done" badge on a finished recording and choose "Rewrite". It reuses the existing transcript and re-mines with the original logic (the output may differ from before).

Details & rules

  • How to start recording: tap the red record button on the home screen; the full-screen view starts recording automatically. Tap "Stop" to finish. Actual recording is tap-based, not press-and-hold.
  • How many articles per recording: the default is "merge aggressively, fewer but richer" — one voice memo usually produces just 1 article. Only when you clearly jump between unrelated topics does it split into 2–3 articles, each able to stand on its own. The cap is 3.
  • Only facts you actually said: the AI has a hard rule — it only uses content from the transcript, never invents or embellishes. Length follows the content: a few sentences can become a short piece; it never pads for word count.
  • Follow-up questions: after each article is written, the AI picks out the "thinnest" spots that only you would know and asks 1–3 short follow-up questions (real numbers, real names, judgments left unexplained). The questions are stored separately and never enter the article body; you can answer by voice and let it fill things in.
  • Status badges: Uploading (red) → Queued (amber) → Transcribing / Mining (spinner) → Done (green). No detected speech shows "No Speech"; blocked recordings show "Out of Credits" or "Recording Too Long".
  • When articles get written: the server processes automatically once an hour (and also triggers on upload). Transcribing a long recording may take several passes, so just check back a bit later.
  • Offline & retry: you can record without a connection. Recordings are saved to a local queue on your phone and upload automatically once you're back online. A failed upload retries automatically (up to 3 times, at 1.5s and 3s intervals); failed files stay safely on your phone and get retried later. The moment the network comes back, uploads resume automatically.
  • Recording length limit: a single recording can be up to 3 hours. Beyond that it's marked "Recording Too Long" and not processed.
  • Transcript too short: if the transcript is under 20 characters, it's marked "Too Short" and no article is written.
  • Credits (cost): new users get a one-time gift of 200 credits (valid for 1 year); 23 credits = ¥1. Transcription is billed by duration (about ¥0.8/hour); mining is billed by AI usage. When your balance hits 0 (out of credits), new recordings stop being processed until you top up.
  • Appending / editing logic: each recording maps to its own article(s) — there's no "append a new recording to an old article". To change an existing article, use the hold-to-speak edit bar; to redo it entirely, use "Rewrite" to mine again.
  • Recording from a tag page: if you start a recording from inside a tag page, the mined article gets that tag by default.

FAQ

How many articles can one recording produce?
The default is "merge aggressively, fewer but richer" — most recordings produce a single article that covers the topic well. Only when you clearly jump between several unrelated topics does it split into 2 or 3 articles, each with its own title, opening, and ending, publishable on its own. Three is the maximum.
Why does my recording still say "Queued", and how long until it's done?
After upload, recordings wait in a server queue: first transcription (Transcribing), then AI writing (Mining). The server processes automatically once an hour, and long recordings may take several passes to transcribe — so just wait a bit and check back later. The status will change to "Done" on its own.
Can I record without an internet connection?
Yes. Recordings are saved locally on your phone first and upload automatically once you're online — no manual steps. Failed uploads retry automatically, and failed files are never lost; they stay until the upload succeeds. The moment your connection comes back, the app pushes the backlog up automatically.
Why does my recording show "No Speech" or "No article yet"?
"No Speech" means no spoken voice was detected (e.g. pure background noise or an accidental tap). "No article yet" means speech was heard, but it was too fragmented to form a complete thought — the AI couldn't find anything that could stand as an article. Speaking your idea through more completely and re-recording usually fixes it. Also, a very short recording (transcript under 20 characters) is marked "Too Short".

Follow-up Questions

After an article is mined, the AI asks about its weakest spots — one or two details only you know. Hold to speak your answer, and the article grows richer with every reply.

You record a quick memo, and VoiceDrop mines it into a fully structured article. But recordings often leave out key information "only you know" — a real number, a person's name, a concrete scene, or a judgment you touched on but never unpacked. Fill those in, and the article truly stands up.

That's what follow-up questions are for. After each article takes shape, the AI finds its thinnest spots and asks 1 to 3 short, specific questions — like an editor pressing the author for details, not small talk. No typing needed: hold to speak your answer, and the information gets woven into the most relevant paragraphs automatically.

Follow-ups are a lightweight feature that's on by default. It stays out of the way — normally tucked away, showing only as a star with a number badge next to the talk bar on the recording detail page, reminding you how many questions are unanswered. Open it when you feel like answering; it never nags. Once you've answered or skipped them all, the star disappears.

How to use it

  1. Record as usual, wait for VoiceDrop to mine it into an article, and open the article's detail page.
  2. If the AI has questions for you, a star button lights up on the right of the "Hold to Speak · Edit" bar at the bottom. The small number badge in its corner is the count of unanswered questions.
  3. Tap the star, and the follow-up card expands from the bottom, wrapping around the talk bar: above it, "Question N/M" (current / total) and the question itself; below, a segmented progress bar.
  4. To answer, hold the talk bar (its label changes to "Hold to Speak · Answer"), speak your answer, and release to send. You don't need to echo the question — just state the facts you know.
  5. After you release, the AI weaves your answer into the most relevant paragraphs. The updated paragraphs highlight in yellow for a few seconds so you can see where the article grew, then it automatically flips to the next question.
  6. Don't want to answer one? Tap "Skip" in the card's top right corner to move on.
  7. When everything is answered or skipped, the card and star tuck away automatically, back to the normal talk bar.
  8. Want more questions? Just hold and say "ask me a few more" or "anything else you want to know" — it generates new questions on the spot and the star lights up again. (Note: while the follow-up card is open, anything you speak is treated as an answer. To ask for more questions, collapse the card first.)
  9. Don't want this feature at all? Turn off the "Follow-up after writing" switch in Settings (subtitle: AI asks a detail or two to thicken the article).

Details & rules

  • The AI asks 1 to 3 questions per article, all about specific information that can't be inferred from the recording — things only the author knows (real numbers, real names, real scenes, unexplained judgments). If nothing is worth asking, it doesn't ask.
  • Follow-up questions never enter the article body or version history. Publishing to WeChat, generating a share page, posting to the community, exporting to Xiaohongshu — every outlet is naturally free of these questions. Only you see them, inside the app.
  • Answers travel through exactly the same channel as regular voice edits — each answer is just an ordinary edit command, with no special waiting screen. Once sent, you get the familiar queue bubble and "editing" flow.
  • Per-question progress: green = answered, orange = current, gray = unanswered or skipped.
  • Unanswered questions expire automatically after 7 days — they don't hang around forever.
  • Follow-ups are per-article: each article has its own set of questions, and the star badge shows the unanswered count for the article you're viewing.
  • The yellow highlight after an answer just shows "what changed where" — it fades in a few seconds and doesn't block anything.

FAQ

Why do some articles get follow-up questions and others don't?
The AI only asks when the recording genuinely lacks key information "only you know" — and filling it in would make the article stronger. If a piece is already fairly complete with nothing worth asking, it stays quiet and the star never lights up.
Do I have to answer each question word by word?
No. Just hold to speak and state the facts you know, in your own words. The AI extracts the information from your answer and weaves it into the most relevant paragraphs — it won't copy the question or touch unrelated parts.
Will the follow-up questions get published with my article, where others can see them?
No. Follow-up questions live only in your own app — they never enter the article body or version history. WeChat publishing, share links, the community, Xiaohongshu — no outlet ever includes them.
How do I get the AI to ask me more questions?
First collapse the follow-up card (swipe it down, or finish the current batch), then hold to speak and say something like "ask me a few more". The AI adds 1 to 3 new questions on the spot and the star lights up again. Note that while the card is open, anything you speak counts as an answer — so collapse it first.
✏️

Editing Articles by Voice

Open a mined article, hold the talk bar at the bottom, and say what to change — the AI edits accordingly: body text, title, deleting paragraphs, swapping images, all of it.

A mined article isn't the end of the road. On any article's reading page, there's a persistent WeChat-style talk bar at the bottom — normally labeled Hold to Speak · Edit. Hold it, say what to change and how, release — and the AI understands your words and edits the article directly. No typing, no hunting for a cursor in a text box.

It edits the article you're currently reading. So you can point precisely, every line of the body is numbered "Line N / Image M" at the start — say "delete line 3", "change line 5 to…", "change the title", and the AI locates targets strictly by the numbers on screen, never miscounting. While you speak, the transcript shows live at the top of the screen, with locator phrases like "Line N / Image M" highlighted so you can confirm it heard the right line or image.

You can issue edits one after another. Each command joins a queue and runs in order — the one being executed shows a bouncing pencil, and the rest wait behind it. When an edit lands, the changed lines flash with a highlight and slowly fade, so you can see at a glance what changed where. Every voice edit is automatically saved as a new version, so if something goes wrong you can always undo.

How to use it

  1. Open a mined article's reading page. A persistent talk bar sits at the bottom, labeled Hold to Speak · Edit.
  2. Hold the talk bar and speak your edit, e.g. "delete line 2", "rewrite line 5 as…", "change the title to…", "add a paragraph after line 3 about…". Keep your finger down while speaking.
  3. As you speak, watch the dark bubble at the top of the screen — it shows the live transcript, with your "Line N / Image M" locators highlighted so you can confirm the target.
  4. When you're done, release to send (to abandon before releasing, swipe up to cancel as prompted). The command joins the queue and starts executing.
  5. Want several changes in a row? Don't wait — hold and speak the next one. New commands stack in the queue and run one by one. The one executing shows a bouncing pencil icon; queued ones show a little clock.
  6. When an edit finishes, the AI replies above the talk bar (on success, usually "Done"), and the changed lines flash briefly with a highlight.
  7. For more precise edits, you can long-press a paragraph — a small menu pops up with "Rewrite this paragraph", "Insert image", and more (plus a local "Copy"). Long-pressing a generated illustration lets you change that image's style. Your choice is handed to the same voice-edit queue.
  8. Not happy with an edit? Use the Undo / Redo buttons in the toolbar to step back or forward (they appear once there's more than one version).

Details & rules

  • It always edits the article you're looking at. The "Line N / Image M" numbers in the body exist only to align you and the AI on targets — they're never written into the article itself.
  • Commands you can give: delete a line, change a line, insert a paragraph after a line, change the title; you can also have it rewrite the whole piece. For images: swap an image or generate a new one. Deleting a paragraph means deleting the corresponding line N.
  • Merging two articles or deleting a whole article — cross-article operations — don't happen on the single-article page. Do them from the "My Recordings" list by holding the red button and speaking with article numbers (e.g. "merge article two and article three"). Merging saves a new article and keeps the originals; deleting asks you to confirm first.
  • Queued, serial execution: multiple commands don't run at once — they enter one queue and execute in order. The active one is highlighted; the rest wait. The server holds this queue and is the true source of authority.
  • Versions are kept: every voice edit automatically writes a new version (source tagged as agent), with a maximum of 10 versions kept — the oldest gets dropped beyond that. Undo / Redo only moves the "current version" pointer without writing a new version; if you undo and then make new edits, the undone "future" versions are discarded, just like git.
  • Feedback after edits: the AI leaves one line of reply above the talk bar — a bright dot icon on success, a red warning icon on error. The reply doesn't auto-dismiss until replaced by the next one or you tap elsewhere. Changed lines glow for a few seconds and fade.
  • Nothing is lost on disconnect / backgrounding / force-quit: commands you've spoken but the server hasn't confirmed are stored locally and resume automatically on reconnect. Stable command IDs deduplicate the resume, so the same edit never applies twice. Text-only commands survive even a force-quit; in the rare case an image edit is interrupted mid-flight, you'll need to say it again (image edits don't silently resume).
  • When a command fails: the reply turns into a red warning with a hint, and that command is removed from the queue — just rephrase more clearly and say it again. If the network dropped mid-way, the command stays in the queue and is sent again after auto-reconnect; no manual resend needed.

FAQ

Can I undo a bad edit?
Yes. Every voice edit is automatically saved as a new version, and the toolbar has Undo / Redo buttons (they appear once there's more than one version) — one tap steps back or forward. Up to 10 versions are kept. Note that if you undo and then make new edits, the undone newer versions are discarded, just like git.
Can I dictate several edits in one go?
Yes. Just hold and keep going — each command stacks into a queue and executes in order. The one being edited shows a bouncing pencil icon; the ones waiting show a little clock. No need to wait for one edit to finish before speaking the next.
How do I target an exact line or image?
Every line of the body is numbered "Line N / Image M" at the start — just say the number, like "delete line 3" or "restyle image 2". While you speak, your locator words are highlighted at the top of the screen so you can confirm the target. You can also long-press a paragraph and use the pop-up menu ("Rewrite this paragraph", etc.) for even more precision.
If the network drops mid-edit, or I quit the app, do I lose my changes?
No. Commands you've spoken but that haven't finished are stored locally and resume automatically on reconnect, with command IDs deduplicating so nothing applies twice. Text edits survive even a force-quit; only the rare case of an interrupted image edit needs to be spoken again.
🖼

Images

Take photos or import from your library inside VoiceDrop — let the AI write from what it sees, or generate cover images, cartoon explainers, and illustrations in various styles for existing articles.

VoiceDrop isn't just about voice — photos are material for the AI too. You can snap photos while recording to attach what's in front of you, or add photos to any finished article later, taken on the spot or picked from your library.

If a recording contains only photos and no speech, VoiceDrop doesn't discard it as "No Speech" — it automatically switches to photo mode: the AI observes the photos and writes a minimal photo essay for you. In other words, a single photo can become a short article on its own.

Photos can also flow the other way — you can have the AI draw for your article: long-press a paragraph to generate a WeChat Official Account cover image, or a cartoon explainer that helps readers grasp the article's structure at a glance. Long-press an existing illustration to redraw it in a different style — cartoon, watercolor, sketch, oil painting, film, or advertising. All drawing jobs join the same queue as voice edits; the AI processes them one by one, and the article updates automatically.

How to use it

Shoot while recording (photos during a recording)

  1. On the recording screen, tap the faint camera icon on the right to open the square camera.
  2. The viewfinder is square, with a rule-of-thirds grid. Tap the white shutter button to shoot; photos line up in a filmstrip at the bottom.
  3. The library icon in the bottom left picks photos from your system library (up to 9 at a time); the bottom right switches between front and rear cameras.
  4. To remove a photo, tap the ✕ in its top right corner.
  5. Tap "Done" in the top right when finished — the photos attach to the recording and upload with it.

Photos only, no speech — straight to an article

  1. Take or pick photos as above; you don't have to say a word.
  2. When the recording ends, the AI sees photos but no speech and automatically enters photo mode, observing the photos and writing a minimal photo essay for you.

Insert photos into an existing article

  1. Open an article and tap the "Insert photos" icon (film-strip style) in the top toolbar.
  2. The same square camera opens — shoot on the spot or pick from your library.
  3. Tap "Done"; the photos upload and the AI places each one near the paragraph it fits best.

Have the AI draw for your article (long-press a paragraph)

  1. In edit mode, long-press any paragraph to open the action menu.
  2. Choose "Insert image" — it has two options:
    • WeChat cover image: a 2.45:1 banner cover placed at the very top of the article, with a title distilled from the article (about 6–10 characters).
    • Cartoon explainer: a flat cartoon-style diagram inserted where it best aids understanding, making the article's structure clear at a glance.
  3. The menu also has "Rewrite this paragraph" (more concise / more casual / more formal / expand a bit) and "Copy".

Restyle an existing illustration (long-press the image)

  1. Long-press an AI-generated illustration in the article to open the "Image style" menu.
  2. Options: cartoon, advertising, watercolor, sketch, oil painting, film.
  3. Pick one, and the AI redraws the image in that style while keeping the composition and subject intact.

Details & rules

  • Library picker limit: up to 9 photos per pick.
  • Photo format: every photo is center-cropped to a square and scaled so the longest edge is at most 1080 pixels, saved as JPEG (under 900KB each). So illustrations in the article match exactly what you saw in the viewfinder — always 1:1.
  • Camera shots are also saved to your phone's photo library: the saved copy is the same square version as the viewfinder. Photos picked from the library are already there and aren't duplicated. (If library write permission is denied, they simply aren't saved — everything else still works.)
  • Where photos live: they upload to cloud storage under your own account, at paths like photos/<session-timestamp>/<second>-<random-suffix>.jpg. Each photo carries its own marker and is placed in the article by that marker, not by sequence number.
  • Photo-mode limits: articles generated from photos alone are minimal photo essays, and photo mode does not generate follow-up questions (there's no dictation to answer about).
  • Drawing shares one queue with voice edits: cover images, cartoon explainers, and restyle commands from the long-press menu all join the same edit queue. The AI processes them serially, showing "editing" while it works; the article updates automatically when done, with undo / redo available.
  • Cover vs. explainer conventions: WeChat cover images are fixed at a 2.45:1 banner; a cartoon explainer's aspect ratio follows the content (landscape, portrait, or square), aiming for "understood at a glance".
  • These menu items (rewrite, insert image, image style) are configured server-side and may change between releases.

FAQ

I only took photos and didn't say a word — can it still produce an article?
Yes. When VoiceDrop sees a recording with photos but no speech, it automatically switches to photo mode: the AI observes the photo content and writes a minimal photo essay for you, instead of discarding it as "No Speech".
How many photos can I pick from the library at once?
Up to 9 per pick. You can also shoot or pick in several rounds — everything lines up in the filmstrip at the bottom, and tapping "Done" uploads them together.
Do photos taken with the in-app camera get saved to my phone's photo library?
Yes. Camera shots are automatically saved as the same square version you saw in the viewfinder. Photos picked from the library are already there and won't be duplicated.
What's the difference between a "WeChat cover image" and a "cartoon explainer"?
The WeChat cover is a 2.45:1 banner placed at the very top of the article, with a short headline distilled from the title. The cartoon explainer is a flat cartoon diagram inserted mid-article, drawing out the article's structure, contrasts, or flow so it's understood at a glance. Both live in the "Insert image" menu after long-pressing a paragraph.
🖋

Writing Style

Teach VoiceDrop how you write, so every recording gets mined into an article that reads like you wrote it.

Your "writing style" is a style fingerprint VoiceDrop keeps for you — it records how you like to write: sentence rhythm and length, favorite words, how tight or loose your tone is, how you open and how you land — not what you write about or what you believe.

It has exactly one job, but a crucial one: every time your voice or photos get "mined" into an article, VoiceDrop folds this style into the AI's instructions, so the output tastes like you — not generic AI-speak.

Styles are stored versioned — every save adds a version, and you can switch or roll back anytime, or even mine the same recording with different versions to compare. There are three ways to get one: write or paste one yourself in Settings; use "Style Learning" to have the server distill one from articles you admire; or have Claude distill one from your published work. Once a style is in place, it applies to all future recordings — and older articles can be re-mined with any style version too.

How to use it

1. View and hand-write your own style

  1. Go to Settings and tap "Writing Style" in the first card (subtitle: "mimic this voice when writing").
  2. A full-page editor opens. If you have no style yet, it shows the hint: "No writing style yet. Paste a distilled style here, and mining will apply it to make articles sound more like you."
  3. Type or paste your style description into the box and tap "Save" in the top right. Every save adds a new version.
  4. A version bar at the top shows the current version number, style name, character count, and total versions. Tap it to expand the version dropdown and switch to any past version. Switching to an old version and saving without edits equals a rollback (no new version); editing the text and saving creates a new version.
  5. Tap "Cancel" in the top left to leave without saving.

2. Style Learning — distill a style from other people's articles

  1. From any app (a Safari page, a text selection, a PDF/Word/RTF document), use the system Share button and choose VoiceDrop.
  2. The "Style Dataset" panel appears: the top shows "N items collected · ~X characters", with the corpus so far and what this share added below. Web pages get their body text parsed automatically (parsing → collected; on failure it shows "link only", with a "Retry" button).
  3. Tap "Keep collecting" to close the panel and share in a few more pieces.
  4. When you have enough, tap the orange "Extract Writing Style". There's a "clear dataset after extraction" checkbox at the bottom (checked by default, so next time starts fresh).
  5. VoiceDrop distills in the background and jumps to "My Recordings" in the app. When done, an introduction article titled "Your writing style · name" appears.

3. Re-mine an old article in a different style (restyle)

  1. Open any finished article (reading page).
  2. Next to the date under the title, there's a style tag with a pencil icon (showing the current style version, e.g. "v8 style"). Tap it.
  3. The "Rewrite in a different style" panel opens: "Pick a style version and rewrite this article. The original stays; switch back anytime."
  4. Pick a style version and tap "Rewrite with vN" at the bottom. If this article has used that version before, it switches back instantly (free); otherwise it re-mines with that style and produces a new version.
  5. During the rewrite it shows "Rewriting in the new style…". The original dictation is untouched, and you can undo or switch versions anytime.

Details & rules

  • Style captures "how you write", not "how you think": the distiller extracts 9 dimensions — sentence rhythm, paragraph length, vocabulary, tone, argument structure, metaphors, emotional intensity, openings and endings, and "things you never do" — anchoring each with a few real quoted sentences, and names the style in 5 characters or fewer (shown as the first line of the version).
  • Style Learning has a word-count threshold: the corpus needs at least 300 characters of usable body text before extraction is allowed. Below that, the button is disabled with the hint "Not enough material yet (X characters total) — share a few articles with body text and collect 300+ characters first". Shares that carry only a title or link with no body (like book-title-only shares from reading apps) don't count.
  • Too few samples get shaky: with fewer than 3 pieces in the corpus, the distilled result is flagged "the fingerprint may be unstable" — it's easy to mistake one article's quirks for your signature.
  • Corpus limits: each sample keeps up to its first 4,000 characters, and one distillation feeds the model at most about 48,000 characters total — anything beyond is truncated in collection order.
  • Extraction runs on the mining pipeline: tapping "Extract Writing Style" actually uploads a silent placeholder recording that triggers the server-side miner — so you can watch progress in "My Recordings" like any recording, and retry on failure. More reliable than instant extraction.
  • It won't wreck anything when material is short: if the server finds the corpus under 300 characters, it does not touch your active style. Instead it gives you an explainer article — "Not enough samples, style unchanged" — listing what it received and how to add more. Your material stays in the dataset.
  • Versioning and rollback: styles are stored versioned just like articles, keeping roughly 10 versions of history; switch or roll back anytime from the version dropdown.
  • Multi-style comparison: the version dropdown has a "multi-style comparison" toggle, letting you check 2–3 versions (3 max). The idea: mining generates one article per style so you can flip between them at the top of the reading page.
  • Re-mining costs credits: restyling an old article runs a full mining pass and costs credits normally. If the article has already used that version, switching back is free — nothing is regenerated.

FAQ

I changed my style — will my previously mined articles change with it?
No. A new style only applies to new recordings; old articles stay as they are. To apply a new style to an old article, open it, tap the style tag next to the date, and use "Rewrite in a different style" to re-mine it with a chosen version — the original is untouched and you can switch back anytime.
How much material does Style Learning need before I can extract?
The corpus needs at least 300 characters of usable body text before "Extract Writing Style" is enabled; below that the button stays gray. The key is body text — select the article's actual text before sharing, or share a web link whose body can be parsed. Title-only or book-name-only shares (from some reading apps) can't build a real writing fingerprint.
Will extracting a new style overwrite my old one for good?
No. Every extraction adds a new version; all old versions remain. Go to Settings → Writing Style, tap the version bar at the top to expand the dropdown, and you can switch or roll back to any past version.
A shared web page says "parse failed · link only" — what do I do?
Some pages' body text can't be fetched. Tap "Retry" on that row; if it still fails, go back to the page, select the body text yourself, and share the text directly — it goes into the style dataset just the same.
🔗

Share & Export

Turn mined articles into public links, push them to your WeChat Official Account drafts with one tap, or export all recordings and articles — audio included — to your device.

Once VoiceDrop mines your voice memos into articles, there are several ways to get them out into the world: a public link for any single article that anyone can open, pushing an article to your own WeChat Official Account drafts, converting to Xiaohongshu (RED) copy with images, and a full export of all recordings and articles to your device.

Sharing for a single article lives in the ⋯ menu at the top right of the article reading page. Only recordings that have produced an article (status "Done") can be shared — a recording with no article has no body to share. Tapping "Share" first has the server issue a public page link for the article, in the form https://jianshuo.dev/voicedrop/<id>, then opens the system share sheet so you can send it to WeChat, X, or any app. In WeChat, the recipient sees a link card with the title and an image (the card image comes from the first photo in the article, falling back to the page default).

To let someone hear the original audio, use Export Data: it packs all your recordings' audio, mined articles, photos, and subtitles into one archive — a set of web pages you can open right in a browser, with a "Play recording" bar under every article.

If you run a WeChat Official Account, once it's configured, "Publish to WeChat Drafts" in the ⋯ menu sends the article straight into your account's draft box, where you finish layout and publish.

How to use it

Create and share a public link

  1. Open a recording's article reading page (it must be "Done").
  2. Tap in the top right and choose "Share".
  3. Wait a moment while the app requests a link from the server, then the system share sheet appears. The link carries the article you're currently reading (with multiple articles, it's the current one).
  4. Pick WeChat / X / Copy and send the jianshuo.dev/voicedrop/… link out. Sending to WeChat automatically produces a link card.

Publish to WeChat Official Account drafts

  1. First configure your account in Settings → WeChat Official Account (AppID / Secret, and add the server IP to your account's whitelist).
  2. Back in the article's menu, tap "Publish to WeChat Drafts" (if the article has been published before, it shows "Update WeChat Draft").
  3. Wait for the result — on success you'll see "Saved to drafts" or "Draft updated". Then go to the WeChat backend's draft box to lay out and publish.

Export all recordings and articles

  1. Open Settings and tap "Export Data" under "Sync & Storage" (subtitle: download all recordings and articles in one package).
  2. A progress sheet appears: Preparing → Exporting data (showing X / total) → Packing.
  3. When done, it shows "Export complete" and the file size. Tap "Share / Save locally" and use the system sheet to save to the Files app, AirDrop it, or send it to someone.

Convert to Xiaohongshu / post to the community

  • "Share to Xiaohongshu" in the ⋯ menu auto-generates Xiaohongshu copy (put on your clipboard), saves the images to your photo library, and opens the Xiaohongshu app — pick the images and paste the copy in its composer.
  • The "Visible in VD Community" switch posts the article to the in-app VD Community.

Details & rules

  • What a public link looks like: https://jianshuo.dev/voicedrop/<id>, where the id is a token generated by the server signing the article. Opening the link requires neither VoiceDrop nor a login — anyone can read it.
  • Only articles can be shared: links are issued for "Done" articles. For recordings without an article yet, the ⋯ menu can only share plain text — no public link is available.
  • Multiple articles: one recording may mine into several articles. Sharing carries the article you're currently reading, and the public page opens and previews with that one.
  • WeChat link cards: what's shared to WeChat is a plain link; WeChat fetches the page and builds the card itself. The card thumbnail prefers the first photo in the article, falling back to the page default. Sharing to X and similar platforms sends "title + body + link" as one block of text.
  • What's in the export: an archive (named like voicedrop-export-YYYY-MM-DD.zip) of web pages rather than Markdown/plain text — an overview page index.html (cards listing "X recordings, X articles", each with "Read article" and "Play recording" buttons), one HTML file per article (body, an audio player bar, an expandable "raw transcript", and the article's photos), plus the original audio (.m4a) and subtitles (.srt). Open index.html in any browser to read and listen.
  • WeChat goes to drafts, not straight to broadcast: articles land in your account's draft box; layout and the actual send still happen in the WeChat backend. Publishing an already-published article updates the same draft in place — no duplicates.
  • WeChat needs setup first: without configuration, or with a bad one (wrong AppID/Secret, IP not whitelisted), you'll be prompted to check the "WeChat Official Account" settings.

FAQ

Do people opening my shared link need VoiceDrop or an account?
No. The public link is an ordinary web page (under the jianshuo.dev/voicedrop/… domain) — anyone can open it in a browser or inside WeChat and read the article directly, no app install, no login.
Does the export come as Markdown or plain text?
Neither. The export is an archive of web pages (HTML) you can open directly in a browser, plus the original audio files (.m4a) and subtitles (.srt). Open the included index.html to browse every article and tap "Play recording" to hear the original audio.
How do I send my original recordings to someone?
Use Settings → Export Data. It packs all your recordings' audio along with the articles; in the exported pages, every article has an audio player bar underneath. Send the package to them, and they can listen in any browser.
If I tap "Publish to WeChat Drafts", does the article get broadcast immediately?
No. It only pushes the article into your account's draft box — success shows "Saved to drafts". The actual layout and broadcast still happen in the WeChat backend, done by you. Prerequisite: configure your account in Settings and add the server IP to the whitelist.
💬

VD Community (Share, Browse, Respond, Tip & Report)

Share your AI-mined articles with fellow creators — read each other's work, record responses, toss coins in encouragement, and withdraw or report content anytime.

Beyond your own WeChat Official Account, articles mined in VoiceDrop can be shared to the VD Community — a public space of creators who also love turning speech into writing. Share a finished article, and others can read it, like it, toss coins, or even record a response article; you can read theirs and encourage them back.

The community is right at the top of the main screen: two tabs, "My Recordings" and "VD Community" — tap the latter to browse what everyone has shared. The list is sorted newest-first by share time (with some ranking based on your interests); pull down to refresh.

Every shared article gets a public link (like /voicedrop/<share-id>) you can forward to WeChat, X, and elsewhere — recipients read it right in the browser, no app install needed.

The community follows a set of Community Guidelines with zero tolerance for objectionable content: sexual content, violence, hate, harassment, illegal content, self-harm. You can report inappropriate content or block users you don't want to see at any time; reported content is taken down immediately and goes to human review.

How to use it

Share an article to the community

  1. Open a recording that has produced an article and go to its detail page (only recordings with a mined article can be shared).
  2. Tap the menu in the top right and turn on the "Visible in VD Community" switch.
  3. If it's your first time posting, the "Community Guidelines" appear first — read them and tap "Agree & Publish".
  4. First-time sharing also requires Sign in with Apple (community posts need a real, accountable identity; the app prompts the login automatically).
  5. On success you'll see "Now visible in VD Community". If the article trips the sensitive-word filter, the share is rejected.

Browse and read

  1. Switch to the "VD Community" tab at the top of the main screen to see all shared cards; pull down to refresh.
  2. Tap any card to open the article. Shares with multiple articles show switchable title chips at the top.

Interacting on someone's article (detail page top bar and ⋯ menu)

  1. Tap ❤️ (heart) to like the article.
  2. Tap ⚡ (lightning circle) to "toss a coin" in encouragement.
  3. Tap → "Write a response" and just record a voice memo; when you stop, the AI mines it into an article, published automatically as a "response" attached under the original.
  4. Tap → "Share" to forward the article to WeChat, X, and more (a public link card is generated).

Withdraw a shared article (two ways, either works)

  1. Return to that recording's detail page, tap , and turn off "Visible in VD Community" — you'll see "Hidden from VD Community".
  2. Or find your own post in the "VD Community" list, choose "Remove from community", and confirm "Remove".
  3. Both only take it off the community — your original article is untouched, and you can share it again anytime.

Report / block

  1. On an article's detail page, tap → "Report" → "Report & take down" — the article is immediately removed from the community pending review.
  2. Tap → "Block this user" to stop seeing their community content; unblock anytime in "Settings → Blocked Users".

Details & rules

  • What can be shared: only recordings that have produced an article. What goes public is the finished article and the scene photos in its body — the original audio recording is never made public.
  • Community Guidelines (must agree once before your first post): you are responsible for what you publish and must have the right to publish it; sexual or explicit content, graphic violence, hate or discrimination, harassment or bullying, illegal content, and self-harm are strictly forbidden — zero tolerance.
  • Posting requires Sign in with Apple: sharing, withdrawing, coin-tossing — all "write" actions require an accountable Apple identity. If you're anonymous, the app guides you through login and then retries automatically.
  • How coin-tossing works: one toss gives the author 2 coins and you 0.5 coins, instantly converted to credits at the current coin rate. One toss per article per person; you can't tip your own article; tossing again shows "already tossed on this one".
  • Daily tipping cap: when the day's credit pool is exhausted, you'll see "today's credit pool is used up — come back tomorrow".
  • Responses: writing a response is just recording another memo that gets mined into an article, automatically linked to the original. The original shows "N responses" underneath, each openable.
  • What reporting does: reported content is removed from the community immediately (you won't see it again either) and handled by human review within 24 hours. Repeat or serious offenders get removed.
  • Blocking is local: blocking takes effect only on your own device (filtering by author nickname), the other person isn't notified, and you can undo it anytime in "Settings → Blocked Users".
  • Contact & complaints: to reach us about content, email [email protected].

FAQ

Why does sharing to the community require Sign in with Apple?
The community is a public space, and posting needs a real, accountable identity — so sharing, withdrawing, and coin-tossing all require Apple sign-in. The first time you try, the app automatically shows "Sign in with Apple" and retries your action after you log in.
If I withdraw from the community, is my original article deleted?
No. Withdrawing only takes the piece off the community — others can no longer see it there, but the original article in your app is completely unaffected, and you can share it again anytime.
What is "coin-tossing"? What do coins do?
Tossing a coin is a way to encourage an article you like. One toss gives the author 2 coins and you 0.5 coins, instantly converted to credits at the current coin rate. You can toss only once per article, and not on your own articles.
What do I do about inappropriate content?
On the article's detail page, tap ⋯ in the top right, choose "Report", then "Report & take down" — the piece is removed from the community immediately and reviewed by a human within 24 hours. If you simply don't want to see someone anymore, choose "Block this user"; you can undo it later in "Settings → Blocked Users".

Credits

Credits are VoiceDrop's billing currency — transcription, article mining, voice edits, and image generation all settle in credits.

Credits are VoiceDrop's unified billing unit. Everything AI-powered you do in the app — transcribing recordings, "mining" dictation into articles, editing articles by voice over and over, generating cartoon explainers — deducts from your credit balance by actual usage.

New users don't need to pay first. Signing up grants a one-time gift of 200 credits, valid for 365 days (one year) — plenty for everyday dictation and mining for a long while. You can keep earning credits through promotional gifts and community coin-tossing.

Your balance and every transaction are recorded on the Settings → Credits page, fully transparent: the big number on top is your remaining credits and roughly how many articles they can mine, with "Credit sources" and "History" sections below — the whole story at a glance. When your balance runs dry, credit-consuming actions are temporarily blocked with a prompt; everything already produced is unaffected.

How to use it

  1. Open Settings and tap "Credits".
  2. The dark card on top shows your remaining credits, with an "≈ N articles" estimate next to it (at roughly 9 credits per article), and a line below reading "total granted X · used Y".
  3. Below that, "Credit sources" groups everything you've received by source — like "sign-up gift", "promotional gift", "monthly allowance".
  4. Further down, "History" is the transaction ledger, newest first. Green "+" entries are income (sign-up gift, coins received); "−" entries are spending (mining, transcription, voice edits, image edits), each timestamped.
  5. In normal use, no action is needed — recordings upload, the system transcribes, mines, and bills automatically. You just talk.

Details & rules

  • New-user gift: 200 credits on sign-up, valid for 365 days.
  • Article mining: billed by the AI tokens actually consumed — roughly 9 credits per article (the app's "articles remaining" estimate uses 9 credits/article).
  • Transcription: billed by recording length at ¥0.8/hour (about 18 credits/hour).
  • Voice edits: billed by actual usage per edit; each article allows up to 100 edits, after which you'll see "this article has reached its edit limit (100)".
  • Image edits / cartoon explainers: a flat 1.8 credits per image. When short, you'll see "not enough credits — one image costs 1.8 credits, please top up".
  • Style distillation and Xiaohongshu copy: also deducted by actual AI usage.
  • Recordings cap at 3 hours — longer ones aren't processed (marked too long).
  • Community coin rewards: when you share an article to the community and someone tosses a coin, both sides receive credits (the author gets 2 coins per toss, the tosser 0.5, converted to credits at the current coin rate). These reward credits expire in 90 days. The daily pool is 2,000 credits, and a single coin converts to at most 200 credits. Repeated tosses from the same person to the same author taper off (2nd toss at 70%, 3rd and later at 50%) to prevent farming.
  • Referral rewards: none at the moment. The ways to earn credits are the sign-up gift, promotional gifts, and community coins.
  • Monthly subscription: the card shows ¥19.9/month, 200 credits per month, reset at month end, cancel anytime — but it's marked "coming soon" and not yet available; tapping it says it's still in development.
  • When credits run out: transcription/mining of new recordings is paused and flagged (they can be processed once you have credits again), and voice edits show "not enough credits to keep editing". Finished articles and history are never lost.
  • Pricing basis: internally, 23 credits = ¥1.

FAQ

How many credits do new users get, and how long do they last?
Sign-up grants 200 credits, valid for one year (365 days). An article costs about 9 credits, plus transcription at ¥0.8/hour — everyday dictation and mining will last you a long time.
How much does mining an article or generating an image cost?
Mining an article is billed by actual AI usage — roughly 9 credits. Generating a cartoon explainer or editing an image is a flat 1.8 credits. Transcription is billed separately by recording length (¥0.8/hour).
Besides topping up, how can I get credits for free?
Two ways: promotional gifts (handed out from time to time), and sharing articles to the community — when someone tosses a coin, both you and the tosser receive credit rewards (valid for 90 days). The monthly subscription is still in development and not yet available.
What happens when I run out of credits? Do I lose my articles?
Nothing is lost. When your balance runs dry, transcription and mining of new recordings are paused and flagged (they resume once you have credits), and voice edits show "not enough credits". All finished articles and your full history remain intact.
🛠

Developer API & Account Sign-in

VoiceDrop's account system (anonymous tokens / Sign in with Apple / 6+4 device pairing) and the open HTTP API for power users.

VoiceDrop's backend is a set of open APIs running on Cloudflare — your recordings, articles, photos, and writing style are all readable and writable with a single credential. If you just use the app normally to record and read articles, you can skip this whole section: login, tokens, and credits are all handled for you.

This section is for two kinds of people:

  • People who want to sign in on a computer, or approve a computer login from their phone — understand how VoiceDrop accounts work (anonymous identity, Sign in with Apple, 6+4 phone device pairing).
  • Power users who want to write their own scripts / clients against the API — VoiceDrop exposes articles, recordings, photos, mining, credits, community, and writing style all as HTTP endpoints; one token goes everywhere.

The data flow in one sentence: you upload an .m4a recording, the server mines it automatically (transcription first, then AI writing) into articles, and the client reads those articles back to share, publish to WeChat, or post to the community. Every step has a corresponding endpoint.

How to use it

1. How accounts work

VoiceDrop doesn't force registration. Its account system has three layers, all pointing to the same copy of your data:

  1. Anonymous identity (anon token): the first time you use the app, it locally generates a high-entropy token starting with anon_ — that's your identity. Your data lives in an isolated space that belongs only to you; nobody else can cross into it. You can copy this token from the app's Settings and use it to call the API elsewhere.
  2. Sign in with Apple: signing in with Apple in the app simply binds your Apple ID to your current anonymous identity — your data doesn't move; it's still the same copy. Once bound, you can recover the account on a new device with the same Apple ID. Posting to the community requires Apple sign-in first.
  3. 6+4 phone device pairing: for signing in to your VoiceDrop account on a computer. Instead of hand-copying the long token, approve it from your phone (see below).

2. Approve a computer login from your phone (6+4 device pairing)

This is for when you want to log in on a computer (say, with a command-line tool) without hand-copying the token. The computer plays the "new device", the phone plays the "old device" — pair once, and the identity lands safely on the computer:

  1. Open the phone app's Settings → Account, which shows a 6-digit hexadecimal short ID.
  2. Start pairing on the computer and enter that 6-digit code.
  3. The phone immediately shows a 4-digit numeric code (with an "It's me / Not me" confirmation).
  4. Type the 4-digit code back on the computer — pairing complete, and the identity is now stored locally on the computer.

Prerequisites: the phone is online, the app is in the foreground, and it's signed in to the target account.

3. Calling the API yourself

  1. First get a token: copy the anon token from the app's Settings, run the 6+4 pairing above, or exchange an Apple sign-in for a session.
  2. Send Authorization: Bearer <your token> on every request (the Files API also accepts a ?token= query parameter).
  3. Call the endpoints you need: list articles, read full text, upload recordings to trigger mining, check your credit balance, share, publish to WeChat, browse the community, and more.
  4. For the complete endpoint list, request examples, and response fields, see the official developer docs (linked under "Details" below).

Details & rules

Three backend services (Base URLs)

  • Files APIhttps://jianshuo.dev/files/api/: accounts, recording / article files, photos, sharing, WeChat, community, Apple sign-in. The vast majority of calls go here.
  • Agent Workerhttps://jianshuo.dev/agent/: triggering mining, live status push, voice editing, credit balance, 6+4 device pairing.
  • Reco Workerhttps://jianshuo.dev/reco/: community feed ranking and interaction reporting. Can be unplugged anytime; if it's down, the main flow is unaffected.

Three kinds of credentials

  • anon token: a high-entropy string of 20+ characters starting with anon_, generated once by the client and used long-term. Its scope (isolated data space) is users/anon-<hash>/.
  • session JWT: obtained via POST /files/api/auth/apple with an Apple identityToken, valid for 365 days. anon and session resolve to the same user and the same data.
  • 24-hour read-only temp token: issued by GET /files/api/token/articles, with expires_in of 86400 seconds. It can only list / download — for handing your article list read-only to external tools. The Agent and Reco services do not accept this read-only token.

Permission boundaries

  • Your scope decides whose data you can see. All relative paths are automatically prefixed with your scope — there's no crossing out (.. and absolute paths are rejected outright).
  • Most endpoints accept any valid token; only community writes (share / withdraw) require an Apple-signed session, otherwise you get 403 needs_apple_signin.
  • The only public, token-free endpoints are: public photos photo/*, the WeChat cover gallery asset/wechat-covers/*, the Apple sign-in exchange auth/apple, and share-page HTML.

Endpoint groups (full list in the developer docs)

  • Articles: list, read full text, write (versioned, with undo / redo), delete, sidecar flags (SRT / no-speech / out-of-credits), tags.
  • Recordings: list, upload (upload auto-triggers mining), download.
  • Photos: list, upload, download (private scoped, or public key).
  • Mining: POST /agent/mine/trigger processes pending recordings (the server also auto-sweeps every 6 hours).
  • Credits: GET /agent/usage/balance for the balance, GET /agent/usage/ledger for the transaction ledger (priced at 23 credits = ¥1).
  • Community: feed, read one, reply, share / withdraw (requires Apple sign-in), report.
  • Writing style: read / write (versioned), style-learning corpus collection, server-side distillation, re-mining a single article with a chosen style.

Hard limits on 6+4 device pairing

  • The pairing code is valid for 2 minutes (start over after a timeout).
  • The 4-digit code allows at most 5 attempts, with a remaining-attempts hint on each miss; exhausting them voids the pairing.
  • A 6-digit prefix matches at most 10 candidate accounts.
  • Tapping "Not me" on the phone cancels the pairing immediately.

Security note: what 6+4 pairing hands over is the account's full identity key, not a separately revocable sub-token. Whoever holds the local credential file holds full control of the account — guard it carefully, and never commit or sync it anywhere others can read.

Developer documentation links

  • Developer home: /voicedrop/en/developer/
  • HTTP API reference (complete endpoints + request examples): /voicedrop/en/developer/api.html
  • Claude Code command-line tool guide: /voicedrop/en/developer/wjs-voicedrop.html

FAQ

I don't write code — do I need to care about these tokens and APIs?
Not at all. If you record, read, and publish normally in the phone app, login and tokens are handled entirely for you. This section exists only for power users who want to sign in on a computer or write their own scripts against the API.
What's the difference between the anonymous identity (anon token) and Sign in with Apple? Will I end up with two copies of my data?
No. Signing in with Apple simply binds your Apple ID to your current anonymous identity — your data stays as the same single copy, nothing moves. The difference: an anonymous identity lives only on the device and can't be recovered if you switch phones; once bound to an Apple ID, you can recover it on a new device with the same Apple account. Also, posting to the community requires Apple sign-in.
The 4-digit code never appeared on my phone during 6+4 pairing — what now?
First check that the phone is online, the app is in the foreground, and it's signed in to the target account. Pairing has a 2-minute window from entering the 6-digit code to completion — after a timeout, start over. If the 6-digit code was mistyped or the phone was offline, the computer side reports no match; fix it and try again.
Is it safe to copy my token out to call the API?
Be careful. Both the anon token and what 6+4 pairing hands over are the account's full identity key — not a separately revocable sub-token. Whoever holds it holds full control of your account. Guard it well, and never commit it to a code repository or sync it anywhere others can read. If you only want to hand your articles read-only to an external tool, the 24-hour read-only temp token is the safer choice.