Loading model details...
Fetching the latest models and pricing from the API.
Loading model details...
Fetching the latest models and pricing from the API.
NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts model with 3B active parameters, distilled from Nemotron 3 Ultra and built for the high-volume execution layer of always-on agents - tool calls, result validation, and subagent delegation.
/ 1M input tokens
/ 1M output tokens
Cache pricing
/ 1M tokens
Implicit cache
from openai import OpenAI
# Initialize the OpenAI client with Qubrid base URL
client = OpenAI(
base_url="https://qubrid.com/v1",
api_key="QUBRID_API_KEY",
)
response = client.chat.completions.create(
model="nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4",
messages=[
{
"role": "user",
"content": "Explain the main benefits of using a chat completion API for text generation."
}
],
max_tokens=4096,
temperature=0.6,
top_p=1,
stream=False,
extra_body={
"enable_thinking": True,
}
)
print(response.choices[0].message.content)from openai import OpenAI
# Initialize the OpenAI client with Qubrid base URL
client = OpenAI(
base_url="https://qubrid.com/v1",
api_key="QUBRID_API_KEY",
)
response = client.chat.completions.create(
model="nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4",
messages=[
{
"role": "user",
"content": "Explain the main benefits of using a chat completion API for text generation."
}
],
max_tokens=4096,
temperature=0.6,
top_p=1,
stream=False,
extra_body={
"enable_thinking": True,
}
)
print(response.choices[0].message.content)Example response
A chat completion API provides a standard way to send conversational input and receive model-generated text in a single request. Key benefits include: • Interoperability: any client can use HTTP with JSON request and response bodies. • Flexibility: system prompts, user messages, and parameters such as temperature and max tokens are easy to configure. • Observability: responses typically include token usage fields for cost and performance tracking. A typical integration sends a POST request with the model name and messages array, then reads the assistant message from the first choice in the response.
from openai import OpenAI
# Initialize the OpenAI client with Qubrid base URL
client = OpenAI(
base_url="https://qubrid.com/v1",
api_key="QUBRID_API_KEY",
)
response = client.chat.completions.create(
model="nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4",
messages=[
{
"role": "user",
"content": "Explain the main benefits of using a chat completion API for text generation."
}
],
max_tokens=4096,
temperature=0.6,
top_p=1,
stream=False,
extra_body={
"enable_thinking": True,
}
)
print(response.choices[0].message.content)from openai import OpenAI
# Initialize the OpenAI client with Qubrid base URL
client = OpenAI(
base_url="https://qubrid.com/v1",
api_key="QUBRID_API_KEY",
)
response = client.chat.completions.create(
model="nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4",
messages=[
{
"role": "user",
"content": "Explain the main benefits of using a chat completion API for text generation."
}
],
max_tokens=4096,
temperature=0.6,
top_p=1,
stream=False,
extra_body={
"enable_thinking": True,
}
)
print(response.choices[0].message.content)Example response
A chat completion API provides a standard way to send conversational input and receive model-generated text in a single request. Key benefits include: • Interoperability: any client can use HTTP with JSON request and response bodies. • Flexibility: system prompts, user messages, and parameters such as temperature and max tokens are easy to configure. • Observability: responses typically include token usage fields for cost and performance tracking. A typical integration sends a POST request with the model name and messages array, then reads the assistant message from the first choice in the response.
Streaming supported • Function calling supported • See all examples in Playground Open in Playground for streaming, files, and all parameters.
Open in Playground