Deep research

Add deep_research: true to a normal completion. The proxy runs the same loop as Chat: it plans several searches, reads public pages, and then asks your model for a cited report.

curl https://api.proxy.gonka.gg/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk-your-api-key-here" \
  -d '{
    "model": "deepseek-ai/DeepSeek-V4-Flash-0731",
    "deep_research": true,
    "messages": [
      {"role": "user", "content": "What moved Brent crude prices in 2026?"}
    ]
  }'

The call blocks while the research runs, often a minute or two, then returns a normal chat completion. stream: true streams that report. It does not stream the search progress.

Timeouts

Set the client timeout to at least 5 minutes. Nothing is sent until the searches finish, including when stream is true, so the first byte can take one to two minutes. A 60-second default will fail before the report starts.

client = OpenAI(
    base_url="https://api.proxy.gonka.gg/v1",
    api_key="sk-your-api-key-here",
    timeout=300,
)
curl --max-time 300 https://api.proxy.gonka.gg/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk-your-api-key-here" \
  -d '{
    "model": "deepseek-ai/DeepSeek-V4-Flash-0731",
    "deep_research": true,
    "stream": true,
    "messages": [
      {"role": "user", "content": "What moved Brent crude prices in 2026?"}
    ]
  }'

SDK note

OpenAI SDKs may drop unknown keys. Put the flag in extra_body:

client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V4-Flash-0731",
    messages=[{"role": "user", "content": "What moved Brent crude prices in 2026?"}],
    extra_body={"deep_research": True},
)

Pseudo-tool

Clients that can only send a tools array may pass {"type":"deep_research"}. The proxy consumes it and does not forward that tool upstream.

Cost

You are billed for the tokens of the model you named: the planning calls and the final report, at that model's rate. Search itself is not a separate line item.

Jev judges which pages to keep. That use is internal to the loop, the same way Chat does it. It is not billed, and it does not require the 100,000,000-token eligibility that a direct Jev call requires. Do not send jev as the model. Jev does not write the report.

What it does not combine with

web_search on the same request is ignored. Deep research does its own searches. MCP tools on the same request are not called.