Tutorials
Collect certificates and reports every night
When an organization publishes results, each of your students gains certificates and reports: PDF files you probably want in your own system. In this tutorial you build a job that runs every night, finds what has been released since the last run, downloads it, and remembers what it has. It paces itself under the rate limit, and if it is interrupted it resumes where it stopped.
What the API gives you
- Only released documents. An organization releases certificates and reports when it publishes results. Before that they are neither listed nor downloadable. A report that is later withdrawn stops being listed and served too.
- Per organization, per student.
GET /v1/<organizationId>/certificate/<studentId>andGET /v1/<organizationId>/report/<studentId>list one student's documents on one organization.<studentId>is the core id you registered the student with. - Only after a first sign-in. Both lists need the organization's copy of the student, which the
student's first sign-in creates. Until then they answer
409 conflict. For this job that simply means "nothing to collect yet". - Paginated.
pagestarts at 1, andlimitdefaults to 20 with a maximum of 100. A largerlimitis lowered to 100 rather than refused. See Pagination. - Downloads are files, not JSON.
GET /v1/<organizationId>/certificate/download/<certificateId>andGET /v1/<organizationId>/report/download/<reportId>stream the file. Either takes the document's_idor itsshortId. - One answer for everything you cannot have. A download of a document that does not exist, is not
released, belongs to another account's student, or has no file answers the same
404 not_found,Not found!.
Plan the run
- List all your students once:
GET /v1/student, 100 per page. - List the organizations once:
GET /v1/organization. It returns the five organizations that run exams, which is where certificates and reports live. - For each organization, go through every student.
- For each student, list certificates, then reports. A
409on the certificates means the student has never signed in to that organization. Skip their reports too, which need the same copy, and move on.
Do not narrow the students by activatedPlatformsThisSeason. That list describes the current season,
while certificates and reports from earlier seasons stay listed. A student who is no longer entered on
an organization can still have results there.
5. For each listed document you do not have yet, download it, save it, and record it.
6. Record each finished student and organization pair, so an interrupted run can resume.
Budget the rate limit
Each account may make 100 requests per 60 seconds to each operation, counted across all your servers. Operations are counted separately, and one operation means one kind of request: listing certificates is one operation whichever organization or student is in the path.
| Request this job makes | Calls per run |
|---|---|
GET /v1/student | Your students ÷ 100 |
GET /v1/organization | 1 |
GET /v1/<organizationId>/certificate/<studentId> | One per student and organization pair (more if a student has over 100) |
GET /v1/<organizationId>/report/<studentId> | The same |
GET …/certificate/download/<certificateId> | One per new certificate |
GET …/report/download/<reportId> | One per new report |
The job below spaces calls to the same operation at least 650 ms apart, which keeps it at about 92
a minute per operation, under the limit. The two listings take turns, so they overlap. With 500 students and
five organizations, that is 2,500 student and organization pairs, and the listing pass takes about
2,500 × 0.65 s ≈ 27 minutes. A pair whose certificate list answers 409 costs one call, not two.
Downloads add 0.65 s per new file. The first run downloads everything
released so far; later runs fetch only what is new.
If you still get 429 too_many_requests, another part of your system is using the same operation. Wait
for the number of seconds in Retry-After; the helper does this for you. Every counted response also
carries X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset. See
Rate limits.
Note
Run the job on one machine only. The budget belongs to your account, so two copies of the job share it and each runs half as fast, while also downloading the same files twice.
Before you start
- Permissions:
student/readandorganization/readonmto, pluscertificate/readandreport/readon each organization you collect from. The job skips an organization where it lacks them. See Permissions. - The helper file
mainteam.mjsormainteam.phpfrom Register and apply. - A directory to write to.
Step 1: list your students
curl -s "https://api.main-team.org/v1/student?page=1&limit=100" \
-H "Authorization: Bearer $TOKEN"
{
"success": true,
"message": "Students fetched successfully.",
"data": [
{
"_id": "66e6a3f5c1d2b30012a4f9c1",
"username": "XXB1045",
"firstName": "Jane",
"lastName": "Doe",
"activatedPlatformsThisSeason": ["common"]
}
],
"pagination": { "page": 1, "limit": 100, "total": 512, "totalPages": 6 }
}
Keep requesting pages until page reaches totalPages. An empty list has totalPages 0. The list is
offset-paginated, so a student registered while you page through may shift others by one position. The
next night's run catches anything missed.
Step 2: list a student's certificates
curl -s "https://api.main-team.org/v1/<organizationId>/certificate/66e6a3f5c1d2b30012a4f9c1?limit=100" \
-H "Authorization: Bearer $TOKEN"
{
"success": true,
"message": "Certificates fetched successfully.",
"data": [
{
"_id": "66f1c0d2e3f4a5b6c7d8e9f0",
"application": "66e6a4b1c1d2b30012a4fa07",
"title": "Certificate of Achievement",
"shortId": "K7Q2M9X4TB",
"active": true,
"createdAt": "2026-12-02T08:00:00.000Z"
}
],
"pagination": { "page": 1, "limit": 100, "total": 1, "totalPages": 1 }
}
A certificate is listed whether it is tied to one of the student's applications or to the student
directly, and every listed certificate can be downloaded. shortId is a short uppercase code that
identifies the document as well as its _id does.
If the student has never signed in to this organization:
{
"error": {
"code": "conflict",
"message": "Student has never signed in to stem, so stem holds no record for them. Generate a sign-in link first with POST /:organizationId/auth/signin.",
"documentation_url": "https://hub.main-team.org/api/errors#conflict",
"request_id": "c1a7e9d2-4b3f-4e5a-8d6c-7f8e9a0b1c2d"
}
}
A student who is not yours answers 404 not_found, Student not found!.
Step 3: list a student's reports
curl -s "https://api.main-team.org/v1/<organizationId>/report/66e6a3f5c1d2b30012a4f9c1?limit=100" \
-H "Authorization: Bearer $TOKEN"
{
"success": true,
"message": "Reports fetched successfully.",
"data": [
{
"_id": "66f1c0d2e3f4a5b6c7d8ea01",
"application": "66e6a4b1c1d2b30012a4fa07",
"shortId": "P3W8R2N6ZD",
"isActive": true,
"isCanceled": false,
"createdAt": "2026-12-02T08:05:00.000Z"
}
],
"pagination": { "page": 1, "limit": 100, "total": 1, "totalPages": 1 }
}
Reports from every season are listed, not only the current one.
Step 4: download a file
curl -s -D - -o certificate.pdf \
https://api.main-team.org/v1/<organizationId>/certificate/download/66f1c0d2e3f4a5b6c7d8e9f0 \
-H "Authorization: Bearer $TOKEN"
HTTP/1.1 200 OK
Content-Type: application/pdf
Content-Length: 184213
Content-Disposition: attachment; filename="Ogrenci Sukru.pdf"; filename*=UTF-8''%C3%96%C4%9Frenci%20%C5%9E%C3%BCkr%C3%BC.pdf
X-Request-Id: 8e2f4a6c-1b3d-4f5e-9a7c-0d2e4f6a8b1c
How to handle the response:
- The body is the file. There is no JSON envelope.
Content-Typeis not alwaysapplication/pdf. It is whatever the file was stored with, andapplication/octet-streamwhen none was. Do not reject a download on its content type.- Take the name from
filename*. It holds the real name, percent-encoded as UTF-8, accents included.filenameis a plain-ASCII fallback for clients that cannot readfilename*. - Errors arrive before the file. A refusal comes as the usual JSON error body with an error status, before any file bytes.
- A transfer that stops early is a failure, not a short file. Compare the bytes you received with
Content-Lengthwhen it is present, and discard a partial file.
Step 5: remember what you have
The job keeps one JSON file next to the downloads:
{
"documents": {
"stem/certificate/66f1c0d2e3f4a5b6c7d8e9f0": {
"shortId": "K7Q2M9X4TB",
"file": "results/stem/66e6a3f5c1d2b30012a4f9c1/certificate-K7Q2M9X4TB.pdf",
"serverName": "Öğrenci Şükrü.pdf",
"at": "2026-12-03T02:15:41.220Z"
}
},
"run": {
"startedAt": "2026-12-03T02:15:00.004Z",
"finished": { "stem:66e6a3f5c1d2b30012a4f9c1": true }
}
}
documentsis keyed by organization, kind and_id, so a document is never downloaded twice, even if it appears on a later page or in a later run.runrecords which student and organization pairs the current run has finished. If the job stops halfway, the next start within 20 hours skips those pairs. A finished run clears it.- The file is rewritten atomically (write a temporary file, then rename) after each download, so a crash never leaves it half-written.
The complete job
Node.js
// collect-results.mjs: download newly released certificates and reports for all your students.
// Usage: node collect-results.mjs ./results
import { createWriteStream } from 'node:fs';
import { mkdir, readFile, rename, rm, stat, writeFile } from 'node:fs/promises';
import { dirname, extname, join } from 'node:path';
import { Readable } from 'node:stream';
import { pipeline } from 'node:stream/promises';
import { api, ApiError } from './mainteam.mjs';
const OUT_DIR = process.argv[2] ?? './results';
const STATE_FILE = join(OUT_DIR, 'state.json');
const MIN_INTERVAL_MS = 650; // per operation: about 92 calls a minute, under the 100 allowed
const RESUME_WINDOW_MS = 20 * 60 * 60 * 1000;
// ---- checkpoint -------------------------------------------------------------
async function loadState() {
try {
return JSON.parse(await readFile(STATE_FILE, 'utf8'));
} catch {
return { documents: {}, run: null };
}
}
async function saveState(state) {
await mkdir(OUT_DIR, { recursive: true });
await writeFile(`${STATE_FILE}.tmp`, JSON.stringify(state, null, 2));
await rename(`${STATE_FILE}.tmp`, STATE_FILE);
}
// ---- pacing: one lane per operation ---------------------------------------
const lastCall = new Map();
async function paced(route, send) {
const wait = (lastCall.get(route) ?? 0) + MIN_INTERVAL_MS - Date.now();
if (wait > 0) await new Promise((resolve) => setTimeout(resolve, wait));
lastCall.set(route, Date.now());
return send();
}
// ---- every item of a paginated list -----------------------------------------
async function* allPages(route, path) {
for (let page = 1; ; page++) {
const separator = path.includes('?') ? '&' : '?';
const res = await paced(route, () => api('GET', `${path}${separator}page=${page}&limit=100`));
yield* res.data;
if (page >= res.pagination.totalPages) return;
}
}
// ---- downloads ----------------------------------------------------------------
function serverFileName(header) {
const encoded = /filename\*=UTF-8''([^;]+)/i.exec(header ?? '');
if (encoded) return decodeURIComponent(encoded[1]);
const plain = /filename="([^"]*)"/i.exec(header ?? '');
return plain ? plain[1] : null;
}
async function download(route, path, target) {
const res = await paced(route, () => api('GET', path, undefined, { raw: true, timeoutMs: 300_000 }));
const serverName = serverFileName(res.headers.get('content-disposition'));
const file = target + (extname(serverName ?? '') || '.pdf');
const part = `${file}.part`;
await mkdir(dirname(file), { recursive: true });
try {
await pipeline(Readable.fromWeb(res.body), createWriteStream(part));
const expected = Number(res.headers.get('content-length'));
if (expected && (await stat(part)).size !== expected) throw new Error('truncated download');
await rename(part, file);
} catch (error) {
await rm(part, { force: true });
throw error;
}
return { file, serverName };
}
const describe = (error) =>
error instanceof ApiError ? `${error.status} ${error.code} (request ${error.requestId})` : error.message;
// ---- the run -------------------------------------------------------------------
async function main() {
const state = await loadState();
if (state.run && Date.now() - Date.parse(state.run.startedAt) < RESUME_WINDOW_MS) {
console.log(`Resuming the run started at ${state.run.startedAt}`);
} else {
state.run = { startedAt: new Date().toISOString(), finished: {} };
}
const students = [];
for await (const student of allPages('students', '/student')) students.push(student);
const organizations = (await api('GET', '/organization?limit=100')).data;
const counts = { downloaded: 0, known: 0, notYet: 0, failed: 0 };
for (const org of organizations) {
console.log(`${org.slug}: ${students.length} students`);
for (const student of students) {
const pair = `${org.slug}:${student._id}`;
if (state.run.finished[pair]) continue;
try {
for (const kind of ['certificate', 'report']) {
for await (const doc of allPages(`${kind}-list`, `/${org._id}/${kind}/${student._id}`)) {
const key = `${org.slug}/${kind}/${doc._id}`;
if (state.documents[key]) {
counts.known++;
continue;
}
try {
const target = join(OUT_DIR, org.slug, student._id, `${kind}-${doc.shortId ?? doc._id}`);
const { file, serverName } = await download(
`${kind}-download`,
`/${org._id}/${kind}/download/${doc._id}`,
target,
);
state.documents[key] = { shortId: doc.shortId, file, serverName, at: new Date().toISOString() };
await saveState(state);
counts.downloaded++;
console.log(` + ${key} -> ${file}`);
} catch (error) {
counts.failed++; // not recorded, so the next run tries again
console.warn(` ! ${key}: ${describe(error)}`);
}
}
}
} catch (error) {
if (error instanceof ApiError && error.status === 409) {
counts.notYet++; // never signed in to this organization: nothing to collect yet
} else if (error instanceof ApiError && error.status === 403) {
console.warn(` ! ${org.slug}: ${describe(error)}; skipping this organization`);
break;
} else {
throw error; // the next start resumes from the last finished pair
}
}
state.run.finished[pair] = true;
await saveState(state);
}
}
state.run = null;
await saveState(state);
console.log(
`Done: ${counts.downloaded} downloaded, ${counts.known} already had, ` +
`${counts.notYet} pairs without a sign-in, ${counts.failed} failed.`,
);
}
main().catch((error) => {
console.error(`Stopped: ${describe(error)}`);
process.exitCode = 1;
});
PHP
<?php
// collect-results.php: download newly released certificates and reports for all your students.
// Usage: php collect-results.php ./results
declare(strict_types=1);
require __DIR__ . '/mainteam.php';
const MIN_INTERVAL_US = 650000; // per operation: about 92 calls a minute, under the 100 allowed
const RESUME_WINDOW_S = 20 * 3600;
$outDir = rtrim($argv[1] ?? './results', '/');
$stateFile = "$outDir/state.json";
$api = MainTeam::fromEnv();
// ---- checkpoint -------------------------------------------------------------
function loadState(string $file): array
{
$text = @file_get_contents($file);
$state = $text === false ? null : json_decode($text, true);
return is_array($state) ? $state : ['documents' => [], 'run' => null];
}
function saveState(string $file, array $state): void
{
@mkdir(dirname($file), 0775, true);
file_put_contents("$file.tmp", json_encode($state, JSON_PRETTY_PRINT | JSON_UNESCAPED_SLASHES | JSON_UNESCAPED_UNICODE));
rename("$file.tmp", $file);
}
// ---- pacing: one lane per operation ---------------------------------------
function paced(string $route): void
{
static $last = [];
$wait = ($last[$route] ?? 0) + MIN_INTERVAL_US - (int) (microtime(true) * 1e6);
if ($wait > 0) {
usleep($wait);
}
$last[$route] = (int) (microtime(true) * 1e6);
}
// ---- every item of a paginated list -----------------------------------------
function allPages(MainTeam $api, string $route, string $path): Generator
{
for ($page = 1; ; $page++) {
paced($route);
$separator = str_contains($path, '?') ? '&' : '?';
$res = $api->call('GET', "$path{$separator}page=$page&limit=100");
foreach ($res['data'] as $item) {
yield $item;
}
if ($page >= $res['pagination']['totalPages']) {
return;
}
}
}
// ---- downloads ----------------------------------------------------------------
function serverFileName(string $header): ?string
{
if (preg_match("/filename\\*=UTF-8''([^;]+)/i", $header, $m)) {
return rawurldecode($m[1]);
}
if (preg_match('/filename="([^"]*)"/i', $header, $m)) {
return $m[1];
}
return null;
}
/** Streams one file to $target plus its extension. Returns [path, server name]. */
function download(MainTeam $api, string $route, string $path, string $target): array
{
@mkdir(dirname($target), 0775, true);
$part = "$target.part";
for ($attempt = 1; ; $attempt++) {
paced($route);
$headers = [];
$fh = fopen($part, 'wb');
$ch = curl_init(MainTeam::BASE_URL . $path);
curl_setopt_array($ch, [
CURLOPT_HTTPHEADER => ['Authorization: Bearer ' . $api->token(), 'X-Request-Id: ' . bin2hex(random_bytes(16))],
CURLOPT_FILE => $fh,
CURLOPT_CONNECTTIMEOUT => 5,
CURLOPT_TIMEOUT => 300,
CURLOPT_HEADERFUNCTION => function ($ch, string $line) use (&$headers): int {
$parts = explode(':', $line, 2);
if (count($parts) === 2) {
$headers[strtolower(trim($parts[0]))] = trim($parts[1]);
}
return strlen($line);
},
]);
$ok = curl_exec($ch);
$status = (int) curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
$error = curl_error($ch);
curl_close($ch);
fclose($fh);
clearstatcache(true, $part);
if ($ok === false) {
@unlink($part);
throw new RuntimeException("Download failed: $error");
}
if ($status === 429 && $attempt <= 3) {
@unlink($part);
sleep(max(1, (int) ($headers['retry-after'] ?? '1')));
continue;
}
if ($status !== 200) {
$body = json_decode((string) file_get_contents($part), true);
@unlink($part);
$err = is_array($body) ? ($body['error'] ?? []) : [];
throw new ApiError(
$status,
$err['code'] ?? 'unknown',
$err['message'] ?? "HTTP $status",
$err['request_id'] ?? ($headers['x-request-id'] ?? null),
);
}
if (isset($headers['content-length']) && filesize($part) !== (int) $headers['content-length']) {
@unlink($part);
throw new RuntimeException('Truncated download');
}
$serverName = serverFileName($headers['content-disposition'] ?? '');
$extension = pathinfo($serverName ?? '', PATHINFO_EXTENSION);
$file = $target . '.' . ($extension !== '' ? $extension : 'pdf');
rename($part, $file);
return [$file, $serverName];
}
}
function describe(Throwable $e): string
{
return $e instanceof ApiError ? "{$e->status} {$e->errorCode} (request {$e->requestId})" : $e->getMessage();
}
// ---- the run -------------------------------------------------------------------
try {
$state = loadState($stateFile);
if ($state['run'] !== null && time() - strtotime($state['run']['startedAt']) < RESUME_WINDOW_S) {
echo "Resuming the run started at {$state['run']['startedAt']}\n";
} else {
$state['run'] = ['startedAt' => gmdate('c'), 'finished' => []];
}
$students = [];
foreach (allPages($api, 'students', '/student') as $student) {
$students[] = $student;
}
$organizations = $api->call('GET', '/organization?limit=100')['data'];
$counts = ['downloaded' => 0, 'known' => 0, 'notYet' => 0, 'failed' => 0];
foreach ($organizations as $org) {
echo "{$org['slug']}: " . count($students) . " students\n";
foreach ($students as $student) {
$pair = "{$org['slug']}:{$student['_id']}";
if (!empty($state['run']['finished'][$pair])) {
continue;
}
try {
foreach (['certificate', 'report'] as $kind) {
foreach (allPages($api, "$kind-list", "/{$org['_id']}/$kind/{$student['_id']}") as $doc) {
$key = "{$org['slug']}/$kind/{$doc['_id']}";
if (isset($state['documents'][$key])) {
$counts['known']++;
continue;
}
try {
$label = $doc['shortId'] ?? $doc['_id'];
$target = "$outDir/{$org['slug']}/{$student['_id']}/$kind-$label";
[$file, $serverName] = download($api, "$kind-download", "/{$org['_id']}/$kind/download/{$doc['_id']}", $target);
$state['documents'][$key] = [
'shortId' => $doc['shortId'] ?? null,
'file' => $file,
'serverName' => $serverName,
'at' => gmdate('c'),
];
saveState($stateFile, $state);
$counts['downloaded']++;
echo " + $key -> $file\n";
} catch (Throwable $e) {
$counts['failed']++; // not recorded, so the next run tries again
echo " ! $key: " . describe($e) . "\n";
}
}
}
} catch (ApiError $e) {
if ($e->status === 409) {
$counts['notYet']++; // never signed in to this organization: nothing to collect yet
} elseif ($e->status === 403) {
echo " ! {$org['slug']}: " . describe($e) . "; skipping this organization\n";
break;
} else {
throw $e; // the next start resumes from the last finished pair
}
}
$state['run']['finished'][$pair] = true;
saveState($stateFile, $state);
}
}
$state['run'] = null;
saveState($stateFile, $state);
echo "Done: {$counts['downloaded']} downloaded, {$counts['known']} already had, "
. "{$counts['notYet']} pairs without a sign-in, {$counts['failed']} failed.\n";
} catch (Throwable $e) {
fwrite(STDERR, 'Stopped: ' . describe($e) . "\n");
exit(1);
}
Expected output
The first run after results are published:
$ node collect-results.mjs ./results
stem: 512 students
+ stem/certificate/66f1c0d2e3f4a5b6c7d8e9f0 -> results/stem/66e6a3f5c1d2b30012a4f9c1/certificate-K7Q2M9X4TB.pdf
+ stem/report/66f1c0d2e3f4a5b6c7d8ea01 -> results/stem/66e6a3f5c1d2b30012a4f9c1/report-P3W8R2N6ZD.pdf
! stem/report/66f1c0d2e3f4a5b6c7d8ea3c: 404 not_found (request 0a1b2c3d-4e5f-4a6b-8c7d-9e0f1a2b3c4d)
neo: 512 students
…
Done: 1284 downloaded, 0 already had, 1466 pairs without a sign-in, 1 failed.
The next night, with nothing new released:
Done: 0 downloaded, 1284 already had, 1463 pairs without a sign-in, 0 failed.
The failed report is not in the state file, so the next run tries it again. A 404 on a document that
was just listed usually means it was withdrawn or has no file yet.
Schedule it
Run the job once a night from cron or your scheduler, with the credentials in its environment. Cron does not load your shell profile, so read them from a file only the job's user can read:
# 02:15 every night
15 2 * * * cd /srv/results-job && . ./credentials.env && node collect-results.mjs /srv/results >> /var/log/results-job.log 2>&1
The job prints no secrets and no tokens, so its log is safe to keep.
Handling failures
| What happens | What the job does | What you do |
|---|---|---|
409 conflict on a list | Counts the pair as without a sign-in, and moves on | Nothing. Their documents appear after their first sign-in there |
403 forbidden on a list | Skips the organization for this run | Ask your operator for certificate/read and report/read there |
404 not_found on a download | Logs it, does not record it, tries again next run | Nothing, unless it persists. Then send the request_id to support |
429 too_many_requests | Waits for Retry-After seconds, up to three times | Check whether something else uses the same route at the same time |
| A truncated transfer | Deletes the partial file; tries again next run | Nothing |
401 unauthorized | Signs a fresh token and retries once, then stops | See Handle tokens in production |
5xx, or a network error on a list | Stops; the next start resumes at the last finished pair | Rerun later |
Next steps
- Certificates and reports: every rule for listing and downloading.
- Pagination and Rate limits: the numbers this job is built on.
- Identifiers: why every request here takes the core student id.
- Reference: listStudentCertificates, downloadCertificate, listStudentReports, downloadReport.