Feeding clean web data to AI agents with MCP and tool calls

By Siftwright team7 min read

Language models are good at reasoning over text they're given and bad at fetching it themselves. Ask an agent what a conference talk said, what a competitor's pricing page looks like today, or what the news says about a company this morning, and it either refuses, guesses, or burns through a generic browsing tool that returns 50 KB of navigation menus.

The fix is to give the agent a few narrow, reliable tools that return small, structured results. This post covers two ways to do that:

  1. MCP (Model Context Protocol): connect an MCP client such as Claude, Cursor or VS Code to a hosted server and the tools appear automatically.
  2. Plain tool calling: define the tools yourself in your own agent and call a REST API from the handler.

We'll use three tools throughout: YouTube transcripts, page screenshots, and Google News search.

What makes a good agent tool

Before any setup, it's worth being clear about what we want, because it drives every choice below.

  • Narrow purpose. "Get the transcript of this video" is easier for a model to use correctly than "browse the web".
  • Structured output. JSON with predictable fields beats raw HTML. The model doesn't have to find the content inside the markup.
  • Small output. Every token a tool returns is a token of context the model has to read and you have to pay for. Trim aggressively.
  • Clear failures. "This video has no captions" is useful. A stack trace isn't.
  • Bounded cost. Agents loop. A tool that can be called 500 times by accident needs a spending ceiling.

Option 1: MCP with a hosted server

MCP is an open protocol that lets AI clients discover and call tools on a server. With a hosted MCP server there's nothing to run yourself: you add a URL to your client's config and the tools show up.

All three Siftwright tools are published as Actors on the Apify Store, and Apify runs a hosted MCP server at mcp.apify.com. You choose which Actors to expose with the tools query parameter:

{
  "mcpServers": {
    "apify": {
      "url": "https://mcp.apify.com?tools=siftwright/youtube-transcript-extractor,siftwright/pageframe-screenshots,siftwright/google-news-scraper",
      "headers": { "Authorization": "Bearer YOUR_APIFY_TOKEN" }
    }
  }
}

That block works in clients that use the common mcpServers config format. Where it goes depends on the client:

  • Claude (desktop and web): add the URL as a custom connector under Settings → Connectors and sign in to Apify with OAuth when prompted, instead of using a token header.
  • Cursor: Settings → MCP → Add new MCP server, or edit ~/.cursor/mcp.json.
  • VS Code: add the server via the "MCP: Add Server" command, which writes the config for you.
  • Claude Code: from a terminal, claude mcp add --transport http apify "https://mcp.apify.com?tools=siftwright/youtube-transcript-extractor,siftwright/pageframe-screenshots,siftwright/google-news-scraper" --header "Authorization: Bearer YOUR_APIFY_TOKEN".

Your Apify token is in Apify Console under Settings → API & Integrations. Apify also supports signing in with OAuth instead of pasting a token, if your client supports it.

Once connected, the client lists three tools:

  • siftwright--youtube-transcript-extractor
  • siftwright--pageframe-screenshots
  • siftwright--google-news-scraper

Each comes with its full input schema and description, so the model knows which arguments exist. Now you can simply ask: "Get the transcript of this talk and list the three main claims with timestamps", or "What's in the news about our company this week?"

Billing: when tools run through Apify's MCP server, they run in your Apify account and Apify bills you per result at the Actor prices ($3 per 1,000 transcripts, $1.50 per 1,000 captures or articles). Set a monthly usage limit in your Apify account settings so a runaway agent can't surprise you.

When MCP is the right choice: interactive work in a chat or coding client, prototyping, and teams that already use Apify.

Option 2: tool calling against a REST API

If you're building your own agent, calling a model API directly or using a framework, you usually define tools yourself: a name, a description, and a JSON Schema for the input. The model decides when to call a tool and with what arguments; your code runs it and sends back the result.

Most model APIs accept a very similar shape. Here's a definition for news search:

{
  "name": "google_news",
  "description": "Search Google News for recent articles. Use for questions about current events, companies or markets. Returns title, source, publish time and link.",
  "input_schema": {
    "type": "object",
    "properties": {
      "query": { "type": "string", "description": "Keywords, a company name, or a quoted phrase" },
      "maxItems": { "type": "integer", "minimum": 1, "maximum": 20, "description": "How many articles to return" }
    },
    "required": ["query"]
  }
}

(Some APIs call the schema field parameters instead of input_schema; the content is the same JSON Schema.)

And one for transcripts:

{
  "name": "youtube_transcript",
  "description": "Get the transcript of a YouTube video. Returns the title, channel, duration and the transcript text with timestamps.",
  "input_schema": {
    "type": "object",
    "properties": {
      "url": { "type": "string", "description": "YouTube video URL or 11-character video ID" }
    },
    "required": ["url"]
  }
}

The handlers

The handlers call the Siftwright API and, crucially, trim the output before it goes back to the model:

import os
import requests

BASE = "https://siftwright.com/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['SIFTWRIGHT_API_KEY']}"}


def _post(endpoint, payload):
    r = requests.post(f"{BASE}/{endpoint}", headers=HEADERS, json=payload, timeout=300)
    body = r.json()
    if not r.ok:
        return {"error": body["error"]["message"]}
    return body


def google_news(query, maxItems=10):
    body = _post("google-news", {"query": query, "maxItems": min(int(maxItems), 20)})
    if "error" in body:
        return body
    return [
        {"title": a["title"], "source": a.get("source"), "publishedAt": a.get("publishedAt"), "link": a["link"]}
        for a in body["results"] if a["status"] == "ok"
    ]


def youtube_transcript(url, max_chars=20000):
    body = _post("youtube-transcript", {"url": url, "outputFormats": ["segments"]})
    if "error" in body:
        return body
    video = body["results"][0]
    if video["status"] != "ok":
        return {"error": f"No transcript available ({video['status']})."}
    lines, size = [], 0
    for seg in video["segments"]:
        line = f"[{int(seg['start']) // 60}:{int(seg['start']) % 60:02d}] {seg['text']}"
        size += len(line)
        if size > max_chars:
            lines.append("[transcript truncated]")
            break
        lines.append(line)
    return {
        "title": video["title"],
        "channel": video.get("channelName"),
        "durationSeconds": video.get("durationSeconds"),
        "transcript": "\n".join(lines),
    }


TOOLS = {"google_news": google_news, "youtube_transcript": youtube_transcript}


def run_tool(name, arguments):
    try:
        return TOOLS[name](**arguments)
    except Exception as exc:
        return {"error": f"Tool failed: {exc}"}

A few things are doing real work here:

  • Field selection. The news handler returns four fields per article, not every field the API provides. Snippets are dropped because the titles carry most of the signal.
  • Timestamped lines. [3:25] ... is compact, and the model can quote timestamps back to the user.
  • A character budget. A two-hour podcast transcript can be 25,000 words. Truncating (or summarising in a first pass) protects your context window.
  • Errors as data. The model gets a short explanation it can relay instead of an exception.

The loop

Every agent framework has its own syntax, but the loop is always the same:

  1. Send the conversation and the tool definitions to the model.
  2. If the model asks to call a tool, run run_tool(name, arguments).
  3. Send the tool result back as a tool-result message.
  4. Repeat until the model answers without calling a tool.

Keep a hard cap on iterations (say 10) so a confused model can't loop forever.

Screenshots for multimodal models

Vision-capable models can look at images. A screenshot tool lets an agent check a layout after a deploy, read a chart, or see a competitor's pricing page as a person would:

def screenshot(url, mobile=False):
    body = _post("screenshot", {
        "url": url,
        "format": "jpeg",
        "fullPage": False,
        "viewportWidth": 390 if mobile else 1280,
        "viewportHeight": 844 if mobile else 800,
    })
    if "error" in body:
        return body
    shot = body["results"][0]
    if shot["status"] != "ok":
        return {"error": f"Capture failed ({shot['status']})."}
    return {"imageUrl": shot["fileUrl"]}

Viewport-only JPEGs keep images small. Depending on your model API, either pass the imageUrl as an image input, or download the bytes and send them base64-encoded.

Controlling cost

Agents are enthusiastic tool users, so put limits in more than one place:

  • In the handler: cap maxItems and batch sizes regardless of what the model asks for.
  • In the loop: cap tool calls per task.
  • At the provider: on a Siftwright plan, when your monthly results are used up, calls return HTTP 429 with quota_exceeded and nothing more is charged. On Apify, set a usage limit.

Only successful results are billed in both setups, so a model that retries a video without captions doesn't cost you anything for those attempts.

MCP or REST: which should you use?

Hosted MCP (Apify) REST tool calls (Siftwright API)
Setup Paste a config block Write tool definitions and handlers
Best for Chat and coding clients Your own agents and products
Output control Tool returns the Actor's full output You trim to exactly what the model needs
Billing Pay-per-event on your Apify account Your Siftwright monthly plan

Many teams use both: MCP while exploring in a chat client, REST handlers once a workflow becomes part of a product.

Summary

Give agents narrow tools that return small, structured JSON, handle failures as data, and put spending limits in the handler, the loop and the account. MCP gets you there in one config block; tool calling gives you full control over what the model sees.

See the web data for AI agents page for the current tool list, or run each tool in the live demo to see exactly what comes back.

Related guides

Get the next guide by email

Practical tutorials on transcripts, screenshots, news data and agent tooling. A couple of emails a month at most.