AI on demand: Qwen/Qwen3.5-35B-A3B-FP8
Jump to navigation
Jump to search
Calling the model
# Set your personal key:
STONEY_KEY=sk-...
# Set the desired model:
MODEL=Qwen/Qwen3.5-35B-A3B-FP8
# Set your prompt:
PROMPT='Hello.'
# Set maximum amount of tokens:
MAX_TOKENS=10000
curl https://llm.stoney-cloud.com/v1/chat/completions \
--silent --fail --show-error \
--header "Authorization: Bearer $STONEY_KEY" \
--header 'Content-Type: application/json' \
--data '{
"model": "'"$MODEL"'",
"messages": [
{"role": "user", "content": "'"$PROMPT"'"}
],
"max_tokens": '"$MAX_TOKENS"'
}' \
| jq
Example output:
{
"id": "chatcmpl-a48b0edd8e8a0153",
"object": "chat.completion",
"created": 1788340157,
"model": "Qwen/Qwen3.5-35B-A3B-FP8",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "\n\nHello! How are you doing today? Is there anything I can help you with? I'd be happy to assist you with something you need.",
"refusal": null,
"annotations": null,
"audio": null,
"function_call": null,
"tool_calls": [],
"reasoning": "Thinking Process:\n\n1. **Analyze the Input:**\n * Input: \"Hello.\"\n * Type: Greeting.\n * Intent: Initiating a conversation.\n * Tone: Polite, neutral.\n\n2. **Determine the Appropriate Response:**\n * Acknowledge the greeting, offer assistance, and maintain a friendly tone.\n\n3. **Drafting Responses:**\n * * Option 1 (Simple): Hello! How can I help you?\n * Option 2 (Friendly): Hi there! Hope you're having a good day. What's on your mind?\n * Option 3. **Option 3 (Professional): Hello. How may I assist you today?\n\n4. **Selecting the Best Option:**\n Option 2 strikes a good balance of friendliness and helpfulness for a general assistant persona.\n\n5. **Final Polish:** \"Hello! How can I help you today be? I'm here to assist with whatever you need.\" (Wait, \"How can today be\" is slightly odd. Let's stick to standard.)\n \"Hello! How are you doing well? What can I help you with?\"\n\n6. **Refining for Conciseness:** \"Hello! How can I assist you with anything today?\"\n\n7. **Final Choice:** \"Hello! How are you doing today? Is there something I can help you with?\"\n\n \" (Wait, English correction: \"How are you?\" + \"How can I help you?\")\n\n *Revised:* \"Hello! How are you doing? Is there anything I can assist you with today?\" -> \"Hello! How are you doing? Is there something I can help you with today?\"\n\n *Let's keep it warm and open.* \"Hello! How's there something you'd like to chat about or work on?\"\n\n *Actually, keep it simple.* \"Hello! How are you doing today? How can I assist you?\"\n\n *Wait, the model should just be natural.*\n\n \"Hello! How's it going? Is there something I can help you with today?\"\n\n8\"\n\n *Okay, let's go with a standard friendly greeting.*\n\n \"Hello! How are you doing today? Is there anything I can help you with?\"\n\n *Wait, check the system instructions for tone.* No specific constraints. Be helpful.\n\n *Final Plan:* Greet back warmly and offer help.\n\n *Draft: \"Hello! How are you doing today? Is there anything I can help you with?\"\n\n *Wait, let's make it slightly more engaging.* \"Hello there! How's your day going? Feel free to ask me anything you need.cw\n"
},
"logprobs": null,
"finish_reason": "stop",
"stop_reason": null,
"token_ids": null
}
],
"service_tier": null,
"system_fingerprint": null,
"usage": {
"prompt_tokens": 12,
"total_tokens": 631,
"completion_tokens": 619,
"prompt_tokens_details": null
},
"prompt_logprobs": null,
"prompt_token_ids": null,
"kv_transfer_params": null
}