Get API key

Uncensored text API for creative pipelines

Quickstart for Acestep pipelines

Connect your creative pipeline to an uncensored LLM using our OpenAI-compatible interface. This guide covers the essential steps to generate text, handle streaming, and manage your API key without content filters.

Prerequisites: Get Your API Key

Start by creating an account on the Get API key page. You can sign in with Google or use an email and password. No phone number is required. Once registered, you receive an API key immediately. New accounts also get $0.50 in trial credit valid for 7 days, with no credit card needed. This key authenticates all requests to the acestep API. Keep it secure, as it is the only credential needed for access.

Installation

The API follows the OpenAI chat-completions standard, so you can use existing SDKs. For Python, install the official client library. For Node.js, install the compatible package. These libraries handle authentication and request formatting automatically. You do not need custom wrappers. Just ensure your library supports the base_url override to point to our endpoint.

Basic Request: Generate Text

Send a POST request to /v1/chat/completions. Use the model ID uncensored. The API accepts standard parameters like temperature and top_p. It returns generated text based on your prompt. No embeddings or media generation endpoints are available. This is strictly for text pipelines.

curl https://api.acestep.cc/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

The response includes the generated content and token usage. If you do not set max_tokens, the output is capped at 2,048 tokens. The context window supports up to 64,000 tokens total for prompt and completion combined.

Streaming Responses

Enable streaming by setting stream: true in your request. The API returns data via Server-Sent Events (SSE). Each chunk contains partial text. The final chunk includes the full token usage statistics. This is ideal for real-time user interfaces. It reduces perceived latency for long outputs.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Handle stream errors gracefully. Network interruptions may cut off the response. Reassemble chunks on the client side if needed for final processing.

Function Calling

Supports tool definitions and tool_choice. Define your functions in the request. The model returns structured arguments for your tools. This works with standard JSON schemas. Use it to extract data or trigger actions in your pipeline.

from openai import OpenAI

client = OpenAI(base_url="https://api.acestep.cc/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Ensure your client library supports function calling. Pass the tool definitions in the tools array. The model decides when to call them based on the prompt context.

JSON Mode

Force structured output by setting response_format to {"type": "json_object"}. The model returns valid JSON instead of natural language. This is useful for parsing data or configuring systems. It does not validate your schema, so ensure your prompt instructs the model correctly.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.acestep.cc/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Combine JSON mode with function calling for robust structured data extraction. Always validate the output on the client side before processing.

Questions and answers

What happens if I exceed the rate limit?

You receive a 429 error. The limit is 300 requests per minute per key. You can also have 8 concurrent requests. Wait before retrying or distribute load across keys if needed.

Why do I get a 402 error?

Your account has insufficient credit. Payments are prepaid via crypto. Errors and refusals do not consume credit, so you can retry without charge. Top up via the billing page.

Is the content truly uncensored?

Yes, for lawful adult, fictional, and controversial topics. We do not refuse based on standard content filters. The only hard limit is no sexual content involving minors. Requests for that are refused.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.