robots.ts
Generate /robots.txt from code: app/robots.ts returns the crawler rules and GioJS writes the file.
app/robots.ts
import type { MetadataRoute } from '@gio.js/core';
export default function robots(): MetadataRoute.Robots {
return {
rules: { userAgent: '*', allow: '/', disallow: ['/account', '/api'] },
sitemap: '/sitemap.xml',
};
}Reference
File name and location
app/robots.ts or app/robots.js, at the root of app/ only. It answers /robots.txt.
Exports
| Field | Type | Default | Description |
|---|---|---|---|
default (required) | Robots | (() => Robots | Promise<Robots>) | - | The rules, or a function returning them. The function gets no arguments. |
revalidate | number | false | 3600 | Seconds the output is cached. false keeps it until the next deploy, 0 generates it on every request. |
The Robots object
| Field | Type | Default | Description |
|---|---|---|---|
rules (required) | RobotsRule | RobotsRule[] | - | One block of directives per rule, in order. |
rules[].userAgent | string | string[] | '*' | One User-Agent: line per value. |
rules[].allow | string | string[] | - | One Allow: line per path. |
rules[].disallow | string | string[] | - | One Disallow: line per path. |
rules[].crawlDelay | number | - | A Crawl-delay: line; must be 0 or more. |
host | string | - | A Host: line after the rules. |
sitemap | string | string[] | - | One Sitemap: line each, at the end. Relative URLs resolve against GIO_SITE_URL. |
The types are MetadataRoute.Robots and RobotsRule, from @gio.js/core.
Response
200,Content-Type: text/plain; charset=utf-8. Blocks are separated by a blank line.- Only
GETandHEAD: other methods get405withAllow: GET, HEAD. - Cached by the server for
revalidateseconds, with anETagand no automaticCache-Control. - A value with a line break is refused (it could smuggle in a directive of its own), as is a negative
crawlDelayor a result that is not an object. Those, and a function that throws, answer500Internal Server Error (ref <digest>), never cached.
Examples
Different rules per crawler
app/robots.ts
import type { MetadataRoute } from '@gio.js/core';
export const revalidate = false;
export default function robots(): MetadataRoute.Robots {
return {
rules: [
{ userAgent: '*', allow: '/', disallow: ['/account', '/api'] },
{ userAgent: 'GPTBot', disallow: '/' },
],
sitemap: '/sitemap.xml',
};
}With GIO_SITE_URL=https://example.com, /robots.txt is:
text
User-Agent: *
Allow: /
Disallow: /account
Disallow: /api
User-Agent: GPTBot
Disallow: /
Sitemap: https://example.com/sitemap.xmlBlock crawlers outside production
app/robots.ts
import type { MetadataRoute } from '@gio.js/core';
export default function robots(): MetadataRoute.Robots {
if (process.env.DEPLOY_ENV !== 'production') {
return { rules: { userAgent: '*', disallow: '/' } };
}
return { rules: { userAgent: '*', allow: '/' }, sitemap: '/sitemap.xml' };
}DEPLOY_ENV here is a variable of your own, set per deployment (a staging server runs with NODE_ENV=production too).
Good to know
- A
public/robots.txtwins overapp/robots.ts: it is served first, the module never runs, and startup warns. A page orroute.tsat/robots.txtstops startup. gio exportwritesout/robots.txtfrom it. Without one (and withoutpublic/robots.txt), the export writesUser-agent: */Allow: /, plus aSitemap:line whenGIO_SITE_URLis set. The server writes nothing in that case:/robots.txtis a 404.- Rules only ask crawlers to stay away. Protect private pages with guards or authentication, not with
Disallow.
Related
- Metadata & SEO - the guide.
- sitemap.ts, manifest.ts, public/
metadata.robots- per-page<meta name="robots">.
Version history
| Version | Changes |
|---|---|
v0.1.0-beta.8 | Introduced: app/robots.ts serves /robots.txt, on the server, in standalone builds and in gio export. |
v0.1.0-beta.3 | gio export generates a default robots.txt. |