Google Scholar contains valuable data -- paper titles, paper URLs, SERP snippets, and more. Scraping this data directly means dealing with anti-bot detection, CAPTCHAs, IP rotation, and constantly breaking selectors. The Scavio API handles all of that and returns clean, structured JSON from a single POST request.
This tutorial shows you how to scrape Google Scholar using PHP and the Scavio API. By the end, you will have a working PHP script that fetches real-time Google Scholar data and parses the results.
Prerequisites
- PHP installed on your machine
- A Scavio API key (free tier includes 50 credits on signup -- no credit card required)
Step 1: Install Dependencies
cURL is built into PHP, so there is nothing to install.
# The cURL extension ships with PHPStep 2: Make Your First Google Scholar Search
Send a POST request to the Scavio Google Scholar API endpoint with your query. The API returns structured JSON with paper titles, paper URLs, SERP snippets, and more.
<?php
// Scavio has no Google Scholar endpoint, so citation counts and author lists are not
// available. This runs a Google web search - narrow it with site:arxiv.org or
// filetype:pdf.
$apiKey = "sk_live_your_key";
$query = "retrieval augmented generation 2024";
$ch = curl_init("https://api.scavio.dev/api/v2/google");
curl_setopt_array($ch, [
CURLOPT_RETURNTRANSFER => true,
CURLOPT_POST => true,
CURLOPT_HTTPHEADER => [
"Authorization: Bearer $apiKey",
"Content-Type: application/json",
],
CURLOPT_POSTFIELDS => json_encode(["query" => $query]),
]);
$response = curl_exec($ch);
curl_close($ch);
$data = json_decode($response, true);
print_r($data);Step 3: Example Response
The API returns structured JSON. Here is an example response for a Google Scholar search:
{
"search_parameters": { "q": "cold brew coffee", "hl": "en", "gl": "us" },
"organic_results": [
{
"position": 1,
"title": "how do you guys make cold brew? : r/Coffee",
"link": "https://www.reddit.com/r/Coffee/comments/oi7rm7/how_do_you_guys_make_cold_brew/",
"snippet": "i wanna learn how to make cold brew coffee but theres a lot of ways...",
"source": "Reddit"
}
],
"related_searches": [{ "query": "cold brew ratio", "link": "https://www.google.com/search?q=cold+brew+ratio" }],
"response_time": 2841,
"credits_used": 1,
"credits_remaining": 4821
}Every field is structured and typed -- no HTML parsing, no CSS selectors, no regex extraction. Your PHP code can access any field directly.
Step 4: Full Working Example
Here is a complete, runnable PHP script that searches Google Scholar and prints the results:
<?php
/**
* Search Google Scholar data with the Scavio API.
* POST /api/v2/google - rows come back under organic_results.
*/
// Scavio has no Google Scholar endpoint, so citation counts and author lists are not
// available. This runs a Google web search - narrow it with site:arxiv.org or
// filetype:pdf.
const API_URL = "https://api.scavio.dev/api/v2/google";
function search_google_scholar(string $query): array {
$apiKey = getenv("SCAVIO_API_KEY");
$ch = curl_init(API_URL);
curl_setopt_array($ch, [
CURLOPT_RETURNTRANSFER => true,
CURLOPT_POST => true,
CURLOPT_HTTPHEADER => [
"Authorization: Bearer $apiKey",
"Content-Type: application/json",
],
CURLOPT_POSTFIELDS => json_encode(["query" => $query]),
]);
$response = curl_exec($ch);
$httpCode = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);
if ($httpCode !== 200) {
throw new Exception("Scavio API error: $httpCode");
}
return json_decode($response, true);
}
$data = search_google_scholar("retrieval augmented generation 2024");
echo json_encode($data, JSON_PRETTY_PRINT);Why Use Scavio Instead of Scraping Google Scholar Directly?
- No proxy management. Direct scraping requires rotating proxies to avoid IP bans. Scavio handles all of this server-side.
- No CAPTCHA solving. Google Scholar aggressively blocks automated requests. Scavio returns clean data every time.
- Structured JSON output. No HTML parsing or CSS selector maintenance. Get typed, consistent data from every request.
- Multi-platform in one API. Search Google, Amazon, YouTube, and Walmart from the same API key with the same authentication pattern.
- Free tier included. 50 credits on signup with no credit card required. Each search costs 1 credit.