Skip to content

Detect Caption API

Pricing

25 credits per request

Fixed cost regardless of video length or number of detected captions.

Overview

The Detect Caption API detects caption positions in videos by analyzing multiple frames to distinguish captions (changing text) from static signs (fixed text).

Domain: api.revidapi.com


Endpoint

POST https://revidapi.com/v1/detect-caption

Authentication: Required - Header x-api-key


Request

Headers

  • x-api-key: Required. Your API key for authentication.
  • Content-Type: Required. Must be application/json.

Body Parameters

Required Parameters

Parameter Type Description
video_url string Video URL to detect (http/https)

Optional Parameters

Parameter Type Default Description
caption_region string "bottom" bottom, top, center, full, top_bottom
sample_frames integer 5 Number of frames to analyze (3-10, recommended 5-6)
srt_text string - JSON string array of SRT [{text, ...}] to assist detection
webhook_url string - URL to receive result when complete
id string - Custom tracking identifier

caption_region

Value Description
bottom Caption at bottom (most common)
top Caption at top
center Caption in middle
full Full screen
top_bottom 50% top & 50% bottom

Response

Immediate Response (Task Created)

{
  "task_id": "8ffd7873-3272-4277-a4d2-83fd0c3731b4",
  "status": "pending",
  "message": "Detection task created",
  "type": "detect"
}

Task Status (GET /paid/get/job/status/{task_id})

✅ Completed

{
  "task_id": "8ffd7873-3272-4277-a4d2-83fd0c3731b4",
  "type": "detect",
  "status": "completed",
  "progress": 100,
  "message": "Found 1 captions, 0 static texts",
  "result": {
    "video_size": {"width": 576, "height": 1024},
    "captions": [{"x": 45, "y": 657, "w": 486, "h": 62, "text": "Sample caption", "occurrences": 6}],
    "static_text": [{"x": 50, "y": 50, "w": 200, "h": 30, "text": "Static sign", "occurrences": 5}],
    "caption_area": {"x": 45, "y": 657, "w": 486, "h": 62},
    "detected_region": {"x": 0, "y": 768, "w": 576, "h": 256}
  },
  "created_at": "2025-12-22T11:18:35.134597"
}

⏳ Processing

{
  "task_id": "8ffd7873-3272-4277-a4d2-83fd0c3731b4",
  "status": "processing",
  "progress": 50,
  "message": "Analyzing frame 3/5..."
}

❌ Failed

{
  "task_id": "8ffd7873-3272-4277-a4d2-83fd0c3731b4",
  "status": "failed",
  "progress": 30,
  "message": "Failed to download video"
}

Response Fields

result.captions[]

Field Type Description
x, y, w, h integer Coordinates (pixels)
text string Detected text
occurrences integer Occurrences across frames

result.caption_area

Combined region of all captions - Use for blur and subtitle

result.static_text[]

Fixed text (signs) - Ignore, do not blur


Example Requests

Minimal:

curl -X POST "https://revidapi.com/v1/detect-caption" \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"video_url": "https://example.com/video.mp4"}'

Full (working payload):

{
  "video_url": "https://example.com/video.mp4",
  "caption_region": "top_bottom",
  "sample_frames": 6
}

With SRT support:

{
  "video_url": "https://example.com/video.mp4",
  "caption_region": "top_bottom",
  "sample_frames": 6,
  "srt_text": "[{\"text\":\"Segment 1\"},{\"text\":\"Segment 2\"}]"
}


Usage Notes

  1. POST returns immediately with task_id (does not wait for processing)
  2. GET task status to retrieve results (poll in loop)
  3. caption_area is the combined region - use for blur/subtitle
  4. static_text are static signs - no need to blur

Workflow

POST /paid/detect-caption
    ↓
GET /paid/get/job/status/{task_id} (loop)
    ↓
Get caption_area → Use for blur/subtitle

Tips

  1. Most captions at bottom: Use caption_region: "bottom" as default
  2. Uncertain position: Use caption_region: "full" to search entire frame
  3. More frames: Increase sample_frames to 7-10 for more accurate detection (slower)
  4. Use caption_area: Always use result.caption_area (not captions[0]) for blur/subtitle

Common Issues

  1. No captions detected: Try different caption_region or increase sample_frames
  2. Wrong region detected: Adjust caption_region parameter
  3. Static text detected: Use caption_area which excludes static_text

Best Practices

  1. Use webhooks: Always use webhooks for better reliability
  2. Unique IDs: Provide unique id values for tracking
  3. Default settings: Use caption_region: "bottom" and sample_frames: 5 for most cases
  4. Use caption_area: Always use result.caption_area for subsequent blur/subtitle operations