SONA SETTINGS Voices Telegram bot UA
SONA reference

Settings

Everything required for the bot and API: authentication, voices, pronunciation, subtitles, images, files and errors.

Key & quick start

The bot and API share the same credit balance.

Get an API key

  1. Open @sonapro_bot.
  2. Choose 🔌 API → Create key.
  3. Send sk_user_… in every request header.
X-API-Key: sk_user_…

Important: never put the key in frontend code, public GitHub or chat. Creating a new key in the bot revokes the previous one.

First synthesis

The flow is asynchronous: create a job_id, poll it and download audio after done.

API="https://api.sonapro.app"
KEY="sk_user_…"

curl -X POST "$API/v1/tts" \
  -H "X-API-Key: $KEY" -H "Content-Type: application/json" \
  -d '{"text":"Hello!","voice":"vc_xxxxxxxxxx","language":"en","model":"sona-3.6"}'

curl "$API/v1/tts/JOB_ID" -H "X-API-Key: $KEY"
curl "$API/v1/tts/JOB_ID/audio" -H "X-API-Key: $KEY" -o out.mp3

Endpoints

Base URL: https://api.sonapro.app.

MethodPathPurpose
POST/v1/ttsCreate speech → job_id
GET/v1/tts/{job_id}Job status
GET/v1/tts/{job_id}/audioDownload audio
GET/v1/tts/{job_id}/linkShort public link
GET/v1/tts/{job_id}/subtitles?format=srt|vtt|ass|json|zipSubtitles; requires subtitles:true
GET/v1/voices?language=en&collection=sona_v1Catalog; collection: library or sona_v1
POST/v1/voices/cloneClone a voice → tmpl:N
GET/POST/v1/voices/templatesList or create a template
PATCH/DEL/v1/voices/templates/{num}Edit or delete a template
GET/v1/emotionsCurrent emotion list
GET/POST/v1/pronunciationList or save a pronunciation rule
DEL/v1/pronunciation/{term}Delete a rule
GET/v1/meAccount, tariff and balance
POST/v1/imagesGenerate an image
POST/v1/images/editGenerate with image references
GET/v1/images/{image_id}Job status
GET/v1/images/{image_id}/fileDownload PNG/JPEG/WebP
POST/v1/videosCreate a Veo 3.1 clip — from text or from 1–3 images → video_id
GET/v1/videos/{video_id}Clip status
GET/v1/videos/{video_id}/fileDownload MP4
GET/v1/videos?limit=20Account clip history

Balance and expiry — GET /v1/me

The answer covers only the account the key belongs to. The bot's Profile shows the same figures. Dates are ISO 8601 in UTC.

FieldVoice-over credits and tariff
credits_remainingWhat can be spent now: every live pack plus the starter credits. Next to it: credits_total and credits_used.
credits_expiring
credits_expires_at
How many credits burn at the nearest deadline, and when. Only the nearest date: every pack has its own term. With no credits left in packs: 0 and null.
tariffCurrent tariff: start, pro, business, ultra or null.
tariff_expires_atWhen that tariff ends. A spent pack keeps its tariff until its term ends, so this date can be later than credits_expires_at and is set even when that one is null. After it, tariff drops to a smaller live pack or becomes null.
jobs_limitHow many voice-overs the account may run at once.
sub_active
sub_expires_at
The Ultra subscription and its end; false and null on other tariffs.
unlimitedtrue — an unlimited account: nothing is charged, so the balance never moves, and tariff_expires_at is null.

Nano and Veo plans are in product_billing.products.nano.plan and product_billing.products.video.plan; null without a plan.

FieldNano / Veo plan
daily_limitGenerations per day; null on Veo, no limit.
used_today
remaining_today
Used and left in the current day.
resets_atWhen the next day starts. Days count from the start of the plan, not from midnight.
expires_atEnd of the plan's current term.
active_untilEnd including a renewal already bought. Without one it equals expires_at.

product_billing.legacy_image_credits — previously purchased credits that can pay for Nano. They are already part of credits_remaining, so do not add them twice. payg — the pay-as-you-go wallet: balance_usd available, reserved_usd held for running jobs; the money does not expire. payg is absent when PAYG is off for the account.

curl https://api.sonapro.app/v1/me -H "X-API-Key: $SONA_API_KEY"

Response, shortened:

{
  "credits_remaining": 400000,
  "credits_expiring": 400000,
  "credits_expires_at": "2026-10-27T11:37:49+00:00",
  "tariff": "pro",
  "tariff_expires_at": "2026-10-27T11:37:49+00:00",
  "jobs_limit": 4,
  "product_billing": {"products": {"nano": {"plan": {
    "daily_limit": 600, "used_today": 2, "remaining_today": 598,
    "resets_at": "2026-09-28T11:37:49+00:00",
    "expires_at": "2026-10-27T11:37:49+00:00",
    "active_until": "2026-11-26T11:37:49+00:00"}}}},
  "payg": {"balance_usd": "5.000000", "reserved_usd": "0.000000"}
}

Voices & templates

Each template stores its own parameters.

ID formats and pricing

vc_…Library voice1 credit / character
sona_…SONA Voice v11.5 credits / character
tmpl:NPersonal template/clone1.5 credits / character

Build a separate first-party picker with GET /v1/voices?collection=sona_v1. Every item in the combined catalog includes collection as sona_v1 or library. Filters compose, for example ?collection=sona_v1&language=uk.

Pull the voice catalog

One request returns the whole catalog for a language — handy for building your own table of id + name + description.

curl "$API/v1/voices?language=en" \
  -H "X-API-Key: $KEY"

# response (trimmed)
{
  "count": 393,
  "voices": [
    {
      "id": "vc_88664ca3a4",
      "name": "Emma",
      "language": "en",
      "gender": "feminine",
      "tags": ["Conversational"],
      "description": "Approachable American female ideal for customer care and support.",
      "is_pro": false,
      "has_preview": true,
      "collection": "library"
    }
  ]
}

One voice, multiple templates

A template stores its own speed, volume and emotion. The same voice can therefore have Neutral, Angry and Calm presets. The bot asks for them when creating a template and lets you edit them from its template card.

POST /v1/voices/templates
{"engine":"sona","voice":"sona_xxxxxxxxxxxx","name":"Taras · angry",
 "speed":1.05,"volume":1.0,"emotion":"angry"}

PATCH /v1/voices/templates/7
{"emotion":"neutral","speed":1.0}

A parameter sent explicitly with a synthesis request overrides the template preset for that request only.

Cloning

POST /v1/voices/clone accepts multipart fields clip, name, language, plus optional speed, volume and emotion. Use 3–10 seconds of clean speech without music.

Speech parameters

Fields accepted by POST /v1/tts. An asterisk marks required fields.

FieldType / rangeMeaning
text *stringText to synthesize
voice *vc_…, sona_…, tmpl:NVoice or template
languageauto, uk, en…auto uses the voice language
modelsona-3.5 / sona-3.6Defaults to sona-3.6. sona-3.5 applies only when the request names it; the legacy names sona-pro, sona-fast and sona-hd also run as sona-3.6. The model chosen in the bot does not change the API default
normalizationauto / off / localeSona 3.6 spoken numbers, dates, times, currencies and abbreviations; e.g. en-IN
formatmp3Output format
bitrate32000–192000MP3 bitrate; default 96000
sample_rate8000–48000Sample rate; default 44100
speed0.6–1.5SONA/clone speed
volume0.5–2.0SONA/clone volume
emotionstringOne delivery style for the request
billing_sourcecredits / paygDefault: plan credits. Set payg explicitly for your USD wallet. No automatic fallback.
auto_stressbool, default trueInternal Cyrillic capital vowel → U+0301
subtitlesboolSRT/VTT/ASS and JSON timings
namestringFilename without extension

Template vs request emotion

Using tmpl:N without emotion applies the stored template emotion. Sending emotion:"angry" overrides it for this one request.

Pay per use (PAYG)

Top up from $5 and pay as you use speech. Sona costs $8 per million billable characters; clones and PRO voices cost $12 (×1.5). 10 concurrent tasks; funds do not expire. For example, $16 covers 2,000,000 ordinary characters.

Open /payg in Telegram to view your balance, top up and choose your speech payment mode. The API uses the same key; send billing_source:"payg" on each request. Telegram's preference does not change the API default. Images and video have their own plans and the same PAYG balance: $0.015 per image, $0.02 per clip; see the Image and Video generation tabs.

GET /v1/billing/payg returns available/reserved USD and history. Create a top-up with POST /v1/billing/checkout: {"product":"payg","amount_usd":"16.25","method":"mono"} (or crypto). Payment is credited after provider verification. An empty wallet returns 402; failed/canceled speech refunds the same wallet. The TTS billing receipt reports the amount and settlement.

Stress, pronunciation & emotions

A clone carries timbre; the model still reads the text. Pronunciation is therefore controlled through text.

Three levels of stress control

LevelExampleWhen to use
1. U+0301 markза́мок / замо́кFirst hint for an ambiguous word
2. Capital vowelзАмок / замОкOnly shorthand: SONA normalizes it to the same U+0301
3. Phonetic replacementТарас → та-РАСOnly after both hints fail and the replacement passes a listening test

U+0301 and a capital vowel are not independent fallbacks. The latter is merely converted to the former, so a model may ignore both.

Personal SONA dictionary: Cyrillic behavior

This is not a complete pronunciation lexicon or an automatic stress detector. It is an account's list of explicit exceptions: before synthesis SONA finds term in the transcript and literally substitutes replacement. You do not pass a dictionary ID to POST /v1/tts; saved rules are applied automatically.

PropertySONA behavior
CyrillicSupported in both term and replacement. Sounds-like guidance may use ordinary Ukrainian or Russian letters.
CaseBy default, Тарас matches Тарас, тарас and ТАРАС.
BoundariesOnly a complete word or complete phrase matches. Тарас does not modify Тараса or Тарасові.
InflectionsSave each required word form separately: Тараса → та-РА-са, Тарасові → та-РА-со-ві. These demonstrate the format; listen-test every replacement.
ScopeRules belong to the account, not one voice. They apply in the bot and API to all of its voices and templates.
OrderLonger phrases run before shorter words. The limit is 50 rules; term is limited to 60 and replacement to 160 characters.

Case-sensitive exceptions: the default is case_sensitive: false, preserving existing rules. With case_sensitive: true, US does not match us, and LaTeX does not match latex. Lowercase keys also accept initial capitalization: cat → cat / Cat, but not CAT. Supported by Sona 3.5 and 3.6. Example for POST /v1/pronunciation: {"term":"US","replacement":"United States","case_sensitive":true}. Overlapping spellings return 422; edit or delete the existing rule first. Use the exact term returned by the list endpoint to update or delete a particular variant.

Keep a dictionary term in its ordinary spelling inside the script. Do not also add U+0301 or a capital-vowel stress mark: the dictionary runs after stress normalization, so that altered spelling no longer equals the saved term. The substituted text is what synthesis, subtitles and character accounting receive.

# add or update one rule
curl -X POST https://api.sonapro.app/v1/pronunciation \
  -H "X-API-Key: $SONA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"term":"Тарас","replacement":"та-РАС"}'

# save inflected forms as separate rules
curl -X POST https://api.sonapro.app/v1/pronunciation \
  -H "X-API-Key: $SONA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"term":"Тараса","replacement":"та-РА-са"}'

# list every account rule
curl https://api.sonapro.app/v1/pronunciation \
  -H "X-API-Key: $SONA_API_KEY"

# delete a rule; percent-encode Cyrillic in the URL
curl -X DELETE https://api.sonapro.app/v1/pronunciation/%D0%A2%D0%B0%D1%80%D0%B0%D1%81 \
  -H "X-API-Key: $SONA_API_KEY"

Use it for names, brands, abbreviations and terms that are consistently mispronounced. If U+0301 solves a stress-only issue, a dictionary rule is unnecessary. Sounds-like spelling guides rather than guarantees pronunciation, so test a short take in the intended language first.

Verified rule: Тарас

For the Ukrainian Taras voice, the model said ТАрас and ignored the stress mark. The listen-tested та-РАС replacement produced the intended ТарАс.

curl -X POST https://api.sonapro.app/v1/pronunciation \
  -H "X-API-Key: $SONA_API_KEY" -H "Content-Type: application/json" \
  -d '{"term":"Тарас","replacement":"та-РАС"}'

The rule matches a whole word, case-insensitively, for every synthesis on this API account. Never syllabify an entire script or invent IPA; listen-test each exception first.

Prompt for Claude / ChatGPT

Put this prompt before a script. The AI will inspect the whole text—not only the word “Тарас”—and return SONA-ready text.

You edit text for SONA speech synthesis.

Analyze the complete text in context and prepare it for natural machine narration.

Rules:
1. Preserve meaning, style, facts and language.
2. Keep natural punctuation and paragraphs; they control pauses.
3. Find words that a TTS engine may pronounce incorrectly. Pay special attention to:
   - given names and surnames;
   - place names;
   - brands, products and company names;
   - abbreviations;
   - foreign words and loanwords;
   - rare terms and homographs.
4. If only the stress is ambiguous, place U+0301 after the stressed vowel:
   за́мок / замо́к, му́ка / мука́.
5. If a stress mark may be insufficient and you know the correct pronunciation with
   high confidence, rewrite ONLY the problematic word phonetically:
   - split it into readable parts with hyphens;
   - write the stressed syllable in UPPERCASE;
   - preserve the grammatical ending of the word form used in the text.
6. Do not limit the analysis to “Тарас”. Detect other risky words from context, but
   do not invent a pronunciation when uncertain.
7. Format examples:
   - Тарас → та-РАС
   - OpenAI → оупен-ей-АЙ
   - ChatGPT → чат-джі-пі-ТІ
   - Renault → ре-НО
8. Do not syllabify ordinary words, insert IPA, or rewrite the entire script
   phonetically. Preserve existing <emotion value="..."/> tags unchanged.
9. Return only the narration-ready text, without explanations, lists or comments.

My text:
"""
[PASTE TEXT]
"""

Tags in text

The speed, volume and emotion fields set one style for the whole request. The same three controls can be placed inside the text instead — plus a pause, letter-by-letter spelling and a laugh. Tags work on sona-3.5 and sona-3.6; [laughter] is documented by the provider for sona-3.6, the default model; do not rely on it with "model":"sona-3.5".

<speed ratio="0.8"/>0.6–1.5from here on
<volume ratio="0.5"/>0.5–2.0from here on
<emotion value="sad"/>list belowfrom here on
<break time="700ms"/>ms or s, up to 10 sone pause here
<spell>Bob</spell>—read letter by letter
[laughter]—one laugh here
<volume ratio="0.6"/><speed ratio="0.8"/>This part is softer and slower.<break time="1s"/>
<volume ratio="1"/><speed ratio="1"/>And this part is back to normal.
<emotion value="excited"/> You will not believe what I learned! [laughter]
<emotion value="sad"/> But then everything went wrong…

Uppercase, a comma decimal and a missing slash are repaired for you (<SPEED RATIO="0,8"> → <speed ratio="0.8"/>); broken markup returns 422 with an explanation instead of silently doing nothing. Tags we do not know (<b>, 5 < 10 > 3) stay plain text. Long text is split into parts and any active speed/volume/emotion is re-opened in each of them, so the effect does not stop mid-way.

emotion is a provider beta documented for English — on other languages the difference may not be audible; speed and volume are guidance, not exact multipliers. <break> splits the generation at that point, so several in a row sound unnatural. Tag characters count towards chars but never reach subtitles: those contain only what is actually spoken.

Full list · 58 emotions

neutral, happy, excited, enthusiastic, elated, euphoric, triumphant, amazed, surprised, flirtatious, curious, content, peaceful, serene, calm, grateful, affectionate, trust, sympathetic, anticipation, mysterious, angry, mad, outraged, frustrated, agitated, threatened, disgusted, contempt, envious, sarcastic, ironic, sad, dejected, melancholic, disappointed, hurt, guilty, bored, tired, rejected, nostalgic, wistful, apologetic, hesitant, insecure, confused, resigned, anxious, panicked, alarmed, scared, proud, confident, distant, skeptical, contemplative, determined

Image generation

Generate from text or from image references.

POST /v1/imagesJSON: prompt, model, size → image_id. Prefer: respond-async returns 202 immediately.
POST /v1/images/editmultipart repeated images: up to 10 files for Nano Banana 2 / Pro and GPT Image 2.5
GET /v1/images/{image_id}Job status
GET /v1/images/{image_id}/fileDownload file
GET /v1/images?limit=20Account history

Example

# Nano plan or previously purchased credits: billing_source may be omitted (auto)
curl -X POST https://api.sonapro.app/v1/images \
  -H "X-API-Key: $SONA_API_KEY" -H "Content-Type: application/json" \
  -H "Prefer: respond-async" \
  -d '{"prompt":"A cinematic Crimean coast at sunrise","model":"nano-banana-2","size":"1024x1024"}'

# PAYG balance ($0.015): billing_source:"payg" is required; without it auto answers 402
curl -X POST https://api.sonapro.app/v1/images \
  -H "X-API-Key: $SONA_API_KEY" -H "Content-Type: application/json" \
  -H "Prefer: respond-async" \
  -d '{"prompt":"A cinematic Crimean coast at sunrise","model":"nano-banana-2","size":"1024x1024","billing_source":"payg","max_amount_microusd":15000}'

# edit with repeated references (repeat -F images=@... up to the model limit; for PAYG add -F billing_source=payg)
curl -X POST https://api.sonapro.app/v1/images/edit \
  -H "X-API-Key: $SONA_API_KEY" -H "Prefer: respond-async" \
  -F 'prompt=Preserve the people and combine the composition' \
  -F 'model=nano-banana-pro' \
  -F 'images=@reference-1.jpg' -F 'images=@reference-2.png'

# GPT Image 2.5 through PAYG ($0.03); with previously purchased credits billing_source may be omitted
curl -X POST https://api.sonapro.app/v1/images \
  -H "X-API-Key: $SONA_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Prefer: respond-async" \
  -d '{"prompt":"Cinematic portrait in soft blue light","model":"gpt-image-2.5","size":"2048x1152","billing_source":"payg","max_amount_microusd":30000}'

# poll GET /v1/images/IMAGE_ID until status=done, then download download_url

Sizes (size)

sizeFormatNano Banana 2 / ProGPT Image 2.5
1024x10241:1, square1024×1024≈1254×1254
2048x115216:9, wide2048×1152≈1672×941
1152x20489:16, vertical1152×2048≈941×1672
1536x11524:3, landscape1536×1152≈1448×1086
1152x15363:4, portrait1152×1536≈1086×1448

The default is 1024x1024; any other value returns 422. Nano Banana returns a file of exactly the requested size. For GPT Image 2.5, size only picks the aspect ratio: the file arrives at the model's own resolution (~1.5 MP), never upscaled. POST /v1/images/edit takes the same size.

Models

ModelmodelReferencesPaid with
Nano Banana 2nano-banana-2up to 10Nano plan, PAYG or previously purchased credits
Nano Banana Pronano-banana-proup to 10the same as Nano Banana 2, sharing the plan's daily allowance
GPT Image 2.5gpt-image-2.5up to 10PAYG — $0.03, or previously purchased credits (at the purchase tariff)

Do not hardcode the list: GET /v1/me returns in image.model_options only the models this account may use, with their price (billing_options). model is required; there is no default.

How an image is paid for

  • Nano plan, 30 days: Start $27 — 600 images a day, 5 at once; Pro $45 — 1,200 a day, 7 at once; Ultra $90 — 3,500 a day, 14 at once. Nano Banana 2 and Pro count towards one daily allowance; unused allowance does not roll over.
  • PAYG — $0.015 per generation from the money balance, only when asked for explicitly: send billing_source:"payg" (a billing_source=payg field in multipart). Add max_amount_microusd to cap the price.
  • Previously purchased credits — 1,000 per Nano Banana 2 or Pro image.
  • GPT Image 2.5 — PAYG, $0.03 per image (billing_source:"payg"), or credits purchased before the switch to plans, at the purchase tariff (Start 2,500 · Pro 3,333 · Business 6,000 · Ultra 9,000); auto spends those first. Plans and credits bought after the switch do not pay for it: without old credits auto returns 402. size only picks the aspect ratio — the file arrives at the model's own resolution (~1.5 MP, e.g. 1672×941 for 16:9). The legacy ids gpt-image-2 and sona-image are accepted.

The default billing_source:"auto" uses the plan, and without one the previously purchased credits. It never spends PAYG money: with neither a plan nor credits the answer is 402. A failed generation refunds automatically. Images ship as the 1080 original exactly as Flow generates them, no upscale. quality=2k was retired and returns 422.

Video generation

Veo 3.1 · 720p · 8 s from text or references, currently 6 s from a first frame · one endpoint for text and images.

POST /v1/videosCreate a clip → 202 + video_id. JSON for a clip from text; multipart/form-data to start from 1–3 images
GET /v1/videos/{video_id}Status: queued → processing → done or error
GET /v1/videos/{video_id}/fileDownload the MP4 (H.264 + AAC)
GET /v1/videos?limit=20Account clip history

Parameters

FieldDefaultDescription
prompt *—What happens in the shot, up to 4000 characters
modelveo-3.1The one video model
aspect_ratio16:916:9 (landscape) or 9:16 (portrait)
images—Multipart only, repeated field, PNG/JPEG/WEBP up to 10 MB each. 1 image is the first frame; 2 are the first and last frames; 3 are references (a character, an object, a style). More than three returns 422
billing_sourceautoauto, plan or payg, as for images
max_amount_microusd—PAYG price ceiling; anything dearer is refused with 409 and nothing is charged

Always asynchronous: a clip takes one to several minutes to render, so POST answers 202 at once and you fetch the file from the status's download_url. A clip from text is 8 s. The render service picks the length of a clip from images: currently 6 s from a first frame (1–2 images) and 8 s from three references. So its duration_s is null until it is done. A finished clip's duration_s is the real length of the file.

Example

# 1. A clip from text
curl -X POST https://api.sonapro.app/v1/videos \
  -H "X-API-Key: $SONA_API_KEY" -H "Content-Type: application/json" \
  -d '{"prompt":"A paper boat drifting across a calm pond at sunrise","aspect_ratio":"16:9"}'
# → 202 {"video_id":"...","status":"queued","model":"veo-3.1","status_url":"/v1/videos/..."}

# 2. A clip that starts from your picture
curl -X POST https://api.sonapro.app/v1/videos \
  -H "X-API-Key: $SONA_API_KEY" \
  -F 'prompt=The camera slowly pushes in, wind moves the leaves' \
  -F 'aspect_ratio=9:16' \
  -F 'images=@first-frame.jpg'

# 3. Poll until done, then download the MP4
curl https://api.sonapro.app/v1/videos/VIDEO_ID -H "X-API-Key: $SONA_API_KEY"
curl https://api.sonapro.app/v1/videos/VIDEO_ID/file -H "X-API-Key: $SONA_API_KEY" -o clip.mp4

Access and billing

  • Video is available to every account. If a manager has switched Veo off for an account, POST /v1/videos returns 403 and GET /v1/me shows video.allowed:false.
  • Veo plan, 30 days, unlimited clips: Start $45 — 2 clips at once; Pro $95 — 6; Ultra $140 — 10. A manager enables the plan.
  • PAYG — $0.02 per clip from the money balance, only when asked for explicitly: send billing_source:"payg". To lock the price, add max_amount_microusd (20000).

GET /v1/me → video: allowed, models, aspect_ratios, duration_s, max_references and billing_options, this account's price. If a render fails, the status says error with the reason in detail, and the payment is returned automatically. error_code is content_refused when Google refused the request itself under its policies (refusal_reason: real_person, intellectual_property, minor, audio, sexual, violence, unsafe, prohibited, prompt or policy; retryable:false — change the prompt or the image), otherwise generation_failed (retryable:true). The same request (prompt + images) refused twice (three times when the reason is the audio) is not accepted for 24 h: POST answers 422 at once and nothing is charged.

Results & errors

Poll the status endpoint until a terminal state.

Job lifecycle

queued→processing→done
queued→processing→error

Poll every 1–2 seconds. Download audio, subtitles, images or video only after done.

HTTP codes

400 / 422Invalid input, unsupported voice, language or format
401Missing, invalid or revoked API key
402Insufficient credits, no active plan, or an empty PAYG balance
403The model or product is not enabled for this account (e.g. video before a manager enables it)
404Job, template or file not found
409The price is above the ceiling you sent (max_credits / max_amount_microusd); nothing was charged
429Rate limit, every concurrent job slot busy, or the plan's daily allowance spent; retry after a delay
503Engine is temporarily paused; show the returned maintenance message