Dong Flash v2.3
The small, fast one.
Dong Flash is for short answers at high volume: chat widgets, classification, and any job where waiting costs more than being slightly wrong. In testing on a GTX 750 it produced a token every 40 seconds.* Faster hardware is planned.
*Testing consisted of one prompt. The prompt was "hello".
- Model id
abliterate-dong-flash-v2.3- Context
- 32K
- Input rate, per 1M tokens
- $0.10
- Output rate, per 1M tokens
- $0.30
Planned. This endpoint does not exist yet.
curl https://api.abliterate.app/v1/chat/completions \
-H "Authorization: Bearer ab-demo-not-a-real-key" \
-H "Content-Type: application/json" \
-d '{"model": "abliterate-dong-flash-v2.3", "messages": [{"role": "user", "content": "Say hello."}]}'