CupAI logo

← Back to all models

OpenAI logo

gpt-live-transcribe

⁦OpenAI⁩ · ⁦gpt-live-transcribe⁩

GPT Live Transcribe — رونویسی زندهٔ گفتار به متن با تأخیر کم (Azure)

⁦32K⁩ tokens

Starting price

4,320

per minute of audio

Billing

Pay as you go

Service status

Active

About gpt-live-transcribe

GPT Live Transcribe — رونویسی زندهٔ گفتار به متن با تأخیر کم (Azure)

gpt-live-transcribe pricing

Billed by the minutes of audio

⁦4,320⁩

per minute of audio

⁦259,182⁩ per hour

One flat rate per minute of audio in the session — no separate output, cache or token charges.

Model id

gpt-live-transcribe

Provider

OpenAI

Context window

32K

Billing basis

Minutes of audio

Status

Active

Share model

https://cupai.ir/en/models/gpt-live-transcribe

Sample code and API for gpt-live-transcribe

Point the base URL at CupAI and use your own API key; the rest of the request stays exactly as it is in any compatible client.

# npm i -g wscat
wscat -c 'wss://api.cupai.ir/v1/realtime?model=gpt-live-transcribe' \
  -H 'Authorization: Bearer $CUPAI_API_KEY'

# then send, in order:
# {"type":"session.update","session":{"type":"realtime"}}
# {"type":"response.create"}

Replace CUPAI_API_KEY with your own key.

Create an API key

Frequently asked questions about gpt-live-transcribe

How is using gpt-live-transcribe charged?
This model is billed by the minute of audio, not by tokens: one flat per-minute rate covers the session, with no separate output, cache or token charges. The per-minute rate is in the price table on this page and is deducted from your wallet.
How do I connect to gpt-live-transcribe?
Point the base URL at CupAI, put your API key in the request header, and send the model id in the request body. Ready-made samples are on this page.
How do payments work?
Payments are in Toman and need no foreign credit card. Credit you buy can be spent on any available model.

Other models

How model pricing works

Every AI model strikes a different balance between speed, answer quality and cost. Flagship models suit complex reasoning, long-document analysis and precise code generation, while lighter models cost considerably less for summarising, classification and short answers. The list above shows each model's rate broken down by input token, output token and cache, so you can estimate the cost of your workload before you start.

A token is the smallest unit of text a model processes; roughly speaking a thousand tokens is about 750 English words. The cost of a request is the sum of its input and output tokens, so shortening your prompt and capping the response length reduces cost directly. Some models have a second price tier for very long inputs, which reprices the entire request once it crosses a stated threshold.