MCP server for coding agents¶
Coding agents such as Claude Code, Cursor or Claude Desktop can call skforecast-ai as a set of tools through the Model Context Protocol (MCP). The agent brings the language model: it reads your request, calls the tools and explains the results. skforecast-ai brings the forecasting: every decision (forecaster, estimator, lags, metric, cross-validation) is made by the same deterministic rules as in Python, and every result comes with the script that produced it.
Once the server is connected, you ask in plain language:
- "Forecast the next 12 months of
data/sales.csvand tell me how accurate it is." - "Compare the candidate models for
data/demand.csvand show me the leaderboard." - "Give me the script that produced that forecast."
The server needs Python 3.10 or newer. Most of the setups below start it with uv (uvx), which installs the package for you, so install uv first.
Install¶
Pick the tab of your agent. Every setup gives the server one directory, --allow-dir: it only reads CSV files inside it, so give it the narrowest directory that holds your data.
Claude Code installs the server and its skill as a plugin from the marketplace of this repository:
/plugin marketplace add skforecast/skforecast-ai
/plugin install skforecast-ai@skforecast-ai
The plugin starts uvx --from "skforecast-ai[mcp]==<version>" skforecast-ai mcp --allow-dir <your project>, where <version> is the version of the plugin (claude plugin list shows it).
The plugin can read every CSV file of your project
The allowed directory is the project you open in Claude Code. That includes an export of credentials or of personal data saved as .csv, whose values an error or a summary can quote.
To pass other options (another directory, --allow-model, HF_HUB_OFFLINE, or the backend of a foundation model with uvx --with), add the server by hand instead and disable the plugin: with both, two servers run. For every project (--scope user):
claude mcp add --scope user skforecast-ai -- uvx --from "skforecast-ai[mcp]" skforecast-ai mcp --allow-dir /absolute/path/to/project
In .cursor/mcp.json of the project (or ~/.cursor/mcp.json for every project):
{
"mcpServers": {
"skforecast-ai": {
"command": "uvx",
"args": [
"--from", "skforecast-ai[mcp]",
"skforecast-ai", "mcp", "--allow-dir", "/absolute/path/to/project"
]
}
}
}
In .vscode/mcp.json of the project (or run MCP: Open User Configuration from the Command Palette for every project). The key is servers, not mcpServers:
{
"servers": {
"skforecast-ai": {
"command": "uvx",
"args": [
"--from", "skforecast-ai[mcp]",
"skforecast-ai", "mcp", "--allow-dir", "/absolute/path/to/project"
]
}
}
}
In ~/.codex/config.toml:
[mcp_servers.skforecast-ai]
command = "uvx"
args = [
"--from", "skforecast-ai[mcp]",
"skforecast-ai", "mcp", "--allow-dir", "/absolute/path/to/project",
]
startup_timeout_sec = 120
tool_timeout_sec = 1800
The two timeouts are needed: by default Codex waits 10 seconds for a server to start, less than the first start of uvx takes, and 60 seconds for a tool to answer, less than a comparison can take.
Open Settings > Developer > Edit Config, which opens claude_desktop_config.json (~/Library/Application Support/Claude/ on macOS, %APPDATA%\Claude\ on Windows), add the server and restart the application:
{
"mcpServers": {
"skforecast-ai": {
"command": "uvx",
"args": [
"--from", "skforecast-ai[mcp]",
"skforecast-ai", "mcp", "--allow-dir", "/absolute/path/to/data"
]
}
}
}
Install the package with its mcp extra (also included in all) in a Python environment, or with pipx in an environment of its own:
pip install "skforecast-ai[mcp]"
pipx install "skforecast-ai[mcp]"
The client then starts the command skforecast-ai mcp --allow-dir /absolute/path/to/data, for example with "command": "skforecast-ai" in the JSON of the Cursor tab. If it does not find skforecast-ai, give the absolute path of the one in your environment (which skforecast-ai on Linux and macOS, where skforecast-ai on Windows), or start it with the Python of that environment: /path/to/python -m skforecast_ai mcp --allow-dir /absolute/path/to/data.
Copy the skill by hand: find it in your installation with
python -c "from importlib.resources import files; print(files('skforecast_ai') / 'mcp' / 'skills' / 'skforecast-ai-forecasting')"
and copy that folder into the skills directory of your agent (.claude/skills/ of your project, or ~/.claude/skills/ for every project, with Claude Code).
The path of --allow-dir must be absolute: the client starts the server from a working directory of its own choosing. The place and format of the configuration are those of each client, so check its documentation if they differ.
Warm up uvx before the first start
The first start is slow. uvx downloads and installs the package and its dependencies the first time, and Python loads them for the first time: about 30 seconds with a fast connection, more with a slow one. A client may give up on a server that takes that long to start: Claude Code waits 30 seconds by default. Run this once before connecting the agent, so later starts take a few seconds (the first one up to 10, the next ones about 2). With the plugin of Claude Code, use "skforecast-ai[mcp]==<version>":
uvx --from "skforecast-ai[mcp]" skforecast-ai --version
In Claude Code, the MCP_TIMEOUT environment variable, in milliseconds, makes it wait longer instead: MCP_TIMEOUT=120000 claude.
Install the skill. The skill teaches the agent how to use the tools. The plugin of Claude Code already includes it. For Cursor, VS Code, Codex and other agents, install it with the skills command:
npx skills add skforecast/skforecast-ai --skill skforecast-ai-forecasting
Without the skill the server still works: the most important of its rules also reach the agent in the instructions of the server.
Check the installation¶
Check that the agent sees the server before you ask for a forecast:
| Client | How to check |
|---|---|
| Claude Code | /mcp in a session, or claude mcp list in a terminal |
| Cursor | The MCP section of its settings lists the server and its tools |
| VS Code | MCP: List Servers in the Command Palette; Show Output opens the log of the server |
| Codex | /mcp in a session, or codex mcp list in a terminal |
| Claude Desktop | After the restart, the server appears among the connectors of the chat box ("+" button) |
Then ask something that needs the server, for example "Profile /absolute/path/to/project/data.csv with skforecast-ai". The agent should call the profile tool and answer with the frequency of the data and the recommended forecaster. If the server does not appear, see Troubleshooting.
What a session looks like¶
You ask, for example, "Forecast the next 12 months of /path/to/data/h2o.csv and tell me how accurate it is". The agent then calls:
profile(data_path="/path/to/data/h2o.csv", target="x"): the frequency, the series, the exogenous columns and the recommended forecaster, as a profile id. Every column other than the target, the date and the series ids is an exogenous variable, unlessexog_columnsnames the ones to use.plan(profile_id=..., steps=12): lags, window features, metric and preprocessing, as a plan id.create_cv(plan_id=...): the cross-validation strategy, with its cost (estimator fits, or inference windows for a foundation model, which is never trained). A notice warns when the backtest will be long.backtest(cv_id=...)and, to compare configurations,compare(cv_id=...).forecast(plan_id=...): the forecast of the next 12 months.
Each of these tools returns an id for the next ones, a plain-text summary (the one describe() gives in Python) and the warnings of the call. The summary of a backtest or a forecast carries its metrics and statistics of the predictions, and that of a comparison its leaderboard; the predictions themselves, row by row, go to CSV files in the output directory, listed in files. get_code returns the script that ran, which you can run yourself without the server. The API reference lists every tool, its arguments and its errors.
Long calls. compare sends a progress notification when each candidate starts and ends. Any long call (a backtest, a forecast, a candidate of compare) also sends one every 5 seconds while it runs, naming what runs and for how long ("ForecasterStats: running (35 s)"), so a client does not give up on a request that is still working, as long as it asked for progress. Cancelling a call waits for the candidate or the backtest in progress to end: a cancelled compare skips its remaining candidates, and a cancelled call registers nothing. Until it ends, only the read tools (get_code, get_failure, list_objects, describe_object) answer; the others wait their turn.
Server options¶
The agent starts the server itself, as a command, and talks to it over its standard input and output. --allow-dir is required: the server only reads CSV files inside that directory.
skforecast-ai mcp --allow-dir /path/to/data
| Option | Default | Meaning |
|---|---|---|
--allow-dir |
(required) | Directory the server may read, as an absolute path (an empty or relative value is rejected). Only absolute paths of .csv files inside it are accepted, also after resolving symbolic links. |
--output-dir |
a new temporary directory | Where the server writes predictions, metrics, leaderboards and long texts. It is also the working directory of the server, and it is kept when the server stops. Its path is logged when the server starts. |
--max-objects |
256 | Most objects (profiles, plans, results) the server keeps. |
--max-memory-mb |
1024 | Memory the objects may take. Beyond either limit, the least recently used objects are removed. |
--max-file-mb |
256 | Largest CSV file (data or future exogenous values) the server reads, checked on the size of the file before reading it. 0 for no limit. |
--allow-model |
(none) | Model ID prefix of a foundation model that the server does not run by default, for example google/timesfm-3.0. Repeat it for several. See Foundation models and licenses. |
When it starts, the server checks that it can write to the output directory and stops with an error otherwise. It writes its log to the standard error, one line per event: the directories it uses when it starts, the message and traceback of an unexpected error with its id, and one line when the client disconnects (it then exits with code 0, also in the middle of a call).
Options in the configuration of the client. Options of the server go at the end of args (for example "--allow-model", "google/timesfm-3.0" to run TimesFM 3.0 once you have accepted its license), and variables of its environment in env:
{
"mcpServers": {
"skforecast-ai": {
"command": "uvx",
"args": [
"--from", "skforecast-ai[mcp]",
"skforecast-ai", "mcp", "--allow-dir", "/absolute/path/to/project",
"--allow-model", "google/timesfm-3.0"
],
"env": {"HF_HUB_OFFLINE": "1"}
}
}
}
With claude mcp add, the variables go before --: claude mcp add --scope user skforecast-ai -e HF_HUB_OFFLINE=1 -- uvx .... In ~/.codex/config.toml, an env = { HF_HUB_OFFLINE = "1" } line under [mcp_servers.skforecast-ai].
Update. The plugin of Claude Code runs the version it was published with: update it with claude plugin update skforecast-ai@skforecast-ai. uvx keeps using the version it downloaded first: run uvx --refresh --from "skforecast-ai[mcp]" skforecast-ai --version to get the newest one, or pin the one you want in args ("skforecast-ai[mcp]==<version>"). With pip, pip install -U "skforecast-ai[mcp]".
What the agent sees¶
The data never travels whole to the agent, and the agent's language model sees what the agent reads:
- Summaries carry statistics (minimum, maximum, mean, standard deviation, missing values), dates, column names and series ids, the decisions and their explanations, the metrics, statistics of the predictions and the leaderboard of a comparison. Never rows of the data or of the predictions. The summary of a plan also names the data file its script reads; the other summaries do not name it.
- Messages of errors and warnings are forwarded as the library writes them. They can name columns and series ids and quote up to 5 values of the data (categories, dates). The server cuts an error message at 4,000 characters, its hint at 1,000 and each text of its
detailsat 500, and sends at most 20 warnings of 1,000 characters each (notices_omittedcounts the rest). An unexpected error (internal_error) carries only the type of the exception and an id: its message and traceback, which can quote a value, go to the log of the server (stderr) under that id. - The allowed directory. The instructions the server gives the agent when it connects name the absolute path of
--allow-dir, so the agent can find a file you name by a relative path without searching your file system. - Scripts (
get_code) name the path of the data file they read. Failures (get_failure) do not (the code that ran reads the data in memory), but they hold a traceback, which can quote values. - Files in the output directory hold rows (predictions, metrics). The agent reads them only if it opens them.
values_included is always false in the response of a tool that creates an object, as a reminder that no rows of the data or of the predictions were sent; the metrics and the leaderboard are in the summary.
Security¶
- Files. The server reads only absolute paths of
.csvfiles inside--allow-dir. A path outside it is rejected before the server looks at the file system, so the error does not say whether the file exists, and checked again after resolving symbolic links. URLs are rejected: download the file first. A file that changes between the profile and a later call, or during a call, is rejected (data_changed). - Code. The server never accepts a plan, a profile or a strategy as JSON, only ids and typed arguments, and checks every value the scripts use.
- Network. The server does not open network connections itself. Foundation models (
ForecasterFoundation) download their weights the first time they run, and their backend contacts the Hugging Face Hub on each run to check them, without sending data. SetHF_HUB_OFFLINE=1in the environment of the server to forbid any connection to the Hugging Face Hub (and use only models already downloaded). Through the server, foundation models only take theestimator_kwargsthat keep the data on your machine. - Working directory. The server runs in the output directory, so a library that writes files next to it (CatBoost writes
catboost_info/) does not write into your project.
Scripts run with your permissions
The scripts run in the process of the server, with the permissions of the user who started it: the server limits what the agent can read and pass, not what a script can do. Run it as a user without access to what the agent should not reach.
Foundation models and licenses¶
The server reads the license of each foundation model from skforecast. By default it only runs the models whose license does not restrict commercial use, whose weights are not gated and whose provider requires no account of its own. The others need --allow-model with their prefix, once you have read and accepted their license. So does a model for which skforecast gives no license information.
| Model | Model ID prefix | Runs by default | License |
|---|---|---|---|
| Chronos-2 (the default model) | autogluon/chronos-2, amazon/chronos-2 |
Yes | |
| TimesFM 2.5 | google/timesfm-2.5 |
Yes | |
| TabICL | soda-inria/tabicl |
Yes | |
| Nori | Synthefy/Nori |
Yes | |
| t0 | theforecastingcompany/t0 |
Yes | |
| TimesFM 3.0 | google/timesfm-3.0 |
With --allow-model |
Non-commercial |
| Moirai | Salesforce/moirai-2 |
With --allow-model |
CC-BY-NC-4.0 |
| TabPFN | priorlabs/tabpfn |
With --allow-model |
Non-commercial; Prior Labs requires an account and accepting its license |
| TS-ICL | taharnbl/TS-ICL |
With --allow-model |
Non-commercial |
Without the option, plan, refine_plan and compare reject those models with model_not_allowed, whose hint tells the agent which option to ask you for.
The first time a model whose weights are not in the local Hugging Face cache is used, a ModelDownloadNotice tells the agent, with the license that skforecast registers for models of that name (the server does not check that the repository exists); later plans and comparisons with the model carry that license in a ModelLicenseNotice. TabPFN keeps its weights in a cache of its own, so its notice only says that the server cannot tell whether they are downloaded.
What the agent does before it calls the server¶
The server only controls its own tools. Its instructions, its errors and its skill tell the agent never to copy a file into --allow-dir, never to write data for you (future values of exogenous variables, missing months) and to ask before writing a corrected copy. An agent can still try to do it with the tools of its client (a cp in the shell, reading a file and writing it again), also before the first call to the server, when no message of the server has reached it.
In the checks of this release a small model tried to copy a file from outside the directory in 1 of 3 sessions, and to write a corrected copy before it was asked to in 1 of 12; a larger one in none. In one of those sessions the client allowed the write: the copy had the missing months filled with values of the agent's own, described as an interpolation, and the forecast was made on it.
Keep the agent asking before it writes
What stops it is the permission to write of your client: keep the agent asking before it writes or runs shell commands in the project (the default of Claude Code), and read what it proposes to write. With a small model, open any corrected copy and compare it with your file before you trust a forecast made on it.
Troubleshooting¶
| Symptom | Cause | Fix |
|---|---|---|
| The server does not appear, or fails to start the first time | uvx is still downloading the package when the client gives up |
Run the warm up command of Install once, then restart the client. In Claude Code, MCP_TIMEOUT=120000 claude also gives it time. In Codex, set startup_timeout_sec |
The client cannot find uvx or skforecast-ai |
The client starts the server with a PATH that is not the one of your terminal |
Give the absolute path of the command (which uvx, which skforecast-ai) |
| The server stops as soon as it starts | --allow-dir is missing, relative or not a directory, or the output directory is not writable |
Give an absolute path; the log of the server (stderr, shown by the client) says which |
path_not_allowed |
The file is outside --allow-dir |
Copy the file there yourself, or restart the server with another --allow-dir |
invalid_path or url_not_allowed |
The path is relative, is not a .csv file or is a URL |
Give the absolute path of a CSV file; download a URL first |
model_not_allowed |
The license of the foundation model restricts its use | Read its license and add --allow-model with its prefix |
missing_dependency |
The backend of a foundation model is not installed where the server runs | Install the package the hint names, or add it with --with to uvx |
unknown_id |
The server restarted, or removed the object to stay within its limits | Ask the agent to profile the data again |
| A tool times out in Codex | A comparison takes longer than the default 60 seconds | Set tool_timeout_sec |
Two skforecast-ai servers in /mcp of Claude Code |
The plugin and a server added by hand both run | Disable one of them |
| Two server processes with Claude Desktop | Claude Desktop starts the server twice when it opens and keeps both processes | Nothing to fix: it talks to one of them, and both stop when the application quits |
To see why a server does not start, run its command in a terminal: it prints the error and stops, or waits for a client (stop it with Ctrl+C).
Errors¶
A failed call returns an error whose text is Error executing tool <name>: followed by a JSON object with code, message, field, hint and details. code is stable, so the agent can act on it: the error codes are those of the Python API plus the ones of the server (unknown_id, inconsistent_ids, invalid_path, path_not_allowed, url_not_allowed, data_changed, model_not_allowed, file_too_large). When a script fails, details.failure_id names its traceback and code, which get_failure returns.
Limits¶
- Calls run one at a time; a call waits for the previous one to end.
- The server reads CSV files of at most
--max-file-mb(256 MB by default), and a horizon (steps) longer than the longest series is rejected when the plan is built. - Ids live while the server runs. An id of a previous run, or of an object removed to stay within
--max-objectsand--max-memory-mb, givesunknown_idsaying which. - A summary, a script or a failure longer than 20,000 characters is cut in the response; the full text is written to the output directory.
- Foundation models (
ForecasterFoundation) need their backend package where the server runs: Chronos-2, the default, needschronos-forecasting(thefoundationextra). Without it,backtestandforecastanswermissing_dependency, whose hint gives thepip installcommand and the--withoption ofuvx. - The script of a forecast with future exogenous values reads them from
exog_future.csvin its working directory: copy the file there to run it.
The skill¶
The skill, SKILL.md, teaches the agent the workflow, how far to trust each result, the cost of a backtest, the format of dates, what each error code asks for and what reaches the agent. It is written for the agent that calls the server; the skills for the LLM are a different thing, the skforecast guides that ask() sends to its own model. It follows the Agent Skills standard.
The SKILL.md shipped with the package
---
name: skforecast-ai-forecasting
description: Forecast time series stored in CSV files with the tools of the skforecast-ai MCP server (profile, plan, create_cv, backtest, compare, forecast). Use when asked to forecast, backtest or compare forecasting models on tabular time series and the skforecast-ai server is connected. Also use it before answering what the server or you can see of the user's data (privacy) and what the server does not do.
---
# Forecasting with the skforecast-ai MCP server
The server runs a deterministic forecasting workflow built on skforecast.
Every decision (forecaster, estimator, lags, metric, cross-validation) comes
from rules, so the same inputs give the same results. You choose the inputs
and explain the results; the server decides and computes. Never invent a
number: every figure you report must come from a response or one of its
files. Do not derive one either: no percentage, difference or ratio that a
response does not give ("45% better" from a metric, a margin between two
candidates).
## Workflow
1. `profile(data_path, target, date_column?, series_id_column?,
exog_columns?)`: the absolute path of a CSV file inside the directory
the server may read, which its instructions name: build the path from
it when the user gives a relative one, without searching the file
system. `target` is one column, or a list of columns for
several series side by side; `series_id_column` names the column of
series ids when the series are stacked. Every other column is an
exogenous variable unless `exog_columns` names the ones to use (an
empty list for none); set it only when the user asks. Read the summary
and the `notices`: frequency, series, gaps, exogenous columns and the
recommended forecaster. Do not open the data file with your own tools
to look at it: `profile` gives its columns and statistics without rows.
When the user named the column to forecast, pass it as `target` at
once. Only when they did not, call `profile` without `target`: its
error lists the columns of the file, so never guess a target to see
them. Read the file only to locate a problem an error reports.
2. `plan(profile_id, steps, ...)`: `steps` is the horizon in observations
(12 for a year of monthly data), at most the length of the longest
series. When the user gives no horizon (or no target, and more than one
column could be it), ask; if you assume one, say so before the results. Leave the other arguments out to take the recommendation; set
them only when the user asks. `metric` (one metric, or a list whose
first one ranks) replaces the metric selected from the data, and only
the metrics given are computed. `use_exog: false` leaves the
exogenous columns out, so `forecast` needs no `exog_path`.
`differentiation` (usually 1, for a series with a trend) differences
the target before training; build the strategy of `create_cv` from
that plan, since a backtest needs the same order in both.
`calendar_features` (an empty list for none), `target_transformer`
(`StandardScaler` or `none`) and `dropna_from_series` replace the
rules of the machine learning forecasters.
3. Optionally `refine_plan(plan_id, overrides)` to change some decisions.
An omitted key keeps the value of the plan; every key but
`forecaster`, `estimator` and `steps` set to null goes back to the
default.
4. `create_cv(plan_id, ...)`: the backtesting strategy. Read `cost` before
running anything; above 50 estimator fits (2000 inference windows for a
foundation model) it already carries the `LongTrainingWarning` the
backtest would emit. A notice says when a
`backtest` of its plan would fail (a direct forecaster with `gap`, or
a first training window shorter than the forecaster needs): change
the strategy as the notice says before running it.
5. `backtest(cv_id, plan_id?)`: the accuracy of the plan of the strategy
over its folds. `plan_id` backtests another plan of the same profile on
the same folds.
6. Optionally `compare(cv_id, candidates?)`: several configurations on the
same folds, ranked by its `metric`, else by the metric chosen for the
plan of the strategy, else by the one selected from the data, with a
seasonal naive baseline. `links.best_plan_id` is the plan of the winner. Without
`candidates` it runs the forecasters the profile recommends for the
family of the data (with several series, ForecasterRecursiveMultiSeries
and ForecasterFoundation), or the estimators of the recommended
forecaster when that leaves one, without those above 500 estimator
fits. Without `interval` it uses the interval of the plan of the
strategy, so the winner keeps it (with an asymmetric interval there
is no baseline: it only takes symmetric ones, such as `[0.1, 0.9]`).
The candidates do not take `use_exog` from that plan: to compare
without exogenous variables, `profile` with `exog_columns: []`.
7. `forecast(plan_id, test_size?, exog_path?)`: the future. `exog_path` is
required when the plan uses exogenous variables: without a file of
future values from the user, ask for it, or build the plan again with
`use_exog: false` and say that they were left out. Never write those
values yourself, nor answer with a `test_size` evaluation instead. With
`test_size` it is a single hold-out evaluation instead, without
`exog_path`: pass the integer `steps` (the last `steps` observations) or
the ISO 8601 date the test set starts at. A fraction only works when it
gives exactly `steps` observations.
`get_code(object_id)` returns the Python script that ran (for a plan, the
one that would run; a profile has none), so the user can reproduce any
result without the server, and `requirements`, the packages to install
for it with the versions the server runs: name those, not others. Hand
the script as it is; if you change anything (the path of the data, a
comment), say what. `describe_object(object_id)` returns a response
again; `list_objects()` lists the ids.
## How far to trust a result
From most to least reliable:
1. A `compare` in which the winner beats the baseline: measured over the
`n_folds` of the strategy (at least 2) against a reference. If the
baseline wins, say so: the data may not be forecastable better than
repeating the last season. There is no baseline with several series,
when the target has missing values or dates, or with an asymmetric
interval (the summary says why):
then read the rows per series of `files.best_metrics` (`files.metrics`
of a backtest). A `mean_absolute_scaled_error` below 1 beats the
one-step naive forecast of that series, above 1 does worse. The summary
only gives the average, so the worst series is not in it: name it.
2. A `backtest`: measured over the same folds, but without a reference.
3. A `forecast` with `test_size`: one window of `steps` observations. It
can be lucky or unlucky; do not present it as the accuracy of the model,
nor as the forecast of the future: its dates are already in the data (a
`HoldoutEvaluationNotice` names them).
4. A `forecast` of the future: no measure of error at all. Report it with
the accuracy of the backtest or comparison of the same plan.
`mean_absolute_scaled_error` and `root_mean_squared_scaled_error` divide
the error by that of the one-step naive forecast (repeat the previous
value) on the training data, in every result and every row of a
leaderboard (a `MetricReferenceNotice` of a backtest or a forecast says
so). That reference is not a seasonal naive forecast nor the baseline of
`compare`: never report a value below 1 as beating either, nor turn it
into a percentage against them.
Prediction intervals are estimates: report them as such, and only from the
rows of `files.predictions` (the bounds of each step). If you have not read
that file, do not describe the interval: name the file. The summary gives
the minimum, maximum and mean of each bound, not a width, so no width and
no range around the point comes from it. Read `notices` before you report
anything: any notice can change what the result means (a data problem, a
warning of the plan, a long training), so tell the user about it.
## Cost
`create_cv` returns the `cost` of backtesting its plan: `n_folds`, `n_fits`
(trainings of the forecaster), `estimator_fits` and `inference_windows`. A
direct forecaster trains one estimator per step; ForecasterStats is
refitted in every fold whatever `refit` says; foundation models and the
baseline count 0 estimator fits. A foundation model is never trained: its
cost is `inference_windows`, one per series and fold (and it downloads its
weights the first time). `compare` runs every candidate on the same folds,
so it costs about the sum of theirs (its response reports the totals).
`compare_estimator_fits` and `compare_inference_windows` of `create_cv`
are those sums for a `compare` without `candidates`, which can be far more
than the plan (with `refit=true`, ForecasterDirect fits one estimator per
step and fold; with many series, the foundation model forecasts each one in
each fold); a `CompareCostNotice` says so (for a foundation plan, its own
`LongTrainingWarning` notice says it instead). `inference_windows` is an
upper bound: a series without data in a fold is not forecast in it. Above
50 estimator fits, or 2000 inference windows (added up over the foundation
candidates of a `compare`), a run gets a `LongTrainingWarning` notice and
can take minutes on a CPU; `compare` without `candidates` leaves out the
candidates above 500 estimator fits. Before an expensive run (a `backtest` or a `compare` above those
thresholds), stop and do not run it: tell the user the number of fits and
the cheaper strategies, an integer `refit` (retrain every n folds), fewer
folds (a larger `fold_stride` or a later `initial_train_size`) or
`refit=false` (train once, no help for ForecasterStats nor for a foundation
model). Run the expensive one only when the user chooses it in so many
words: asking to retrain regularly is not that choice.
Progress and cancellation: `compare` reports when each candidate starts
and ends, and any long call (a backtest, a forecast, a candidate) sends a
progress notification every 5 seconds while it runs, naming what runs and
for how long ("ForecasterStats: running (35 s)"). Cancelling a `compare`
waits for the candidate in progress to end and skips the rest; cancelling
another tool waits for it to end (the backtest or the forecast in
progress). Meanwhile only the read tools (`get_code`, `get_failure`,
`list_objects`, `describe_object`) answer: the others wait their turn.
## What a result says, and what the server does not do
Report what was measured, never why. A ranking says which candidate had the
lowest error over the folds, not what makes it better for the data: give
no cause, even hedged, for a ranking, a metric or the shape of a forecast
(a seasonal pattern, an event, too little data). Asked why, say that the
server does not measure it and restate the metric and its values.
The server does not search hyperparameters (`compare` runs the candidates
you list: call it that, not a grid search), detect anomalies, select
features or fill in missing values. Say so, offer what it does and stop
there. Do not do any of it another way in the same answer, by hand from
the files or with your own script, even labeled as outside the server:
that is for the user to ask once they know.
## Inputs
- Paths: absolute paths of `.csv` files inside the allowed directory,
which the instructions of the server name (also `details.allowed_dir`
of a path error). No URLs: download the file first. No relative paths
and no `~`. Never copy or move a file of the user into that directory
yourself: tell them it is outside, and that they can copy it there or
restart the server with another `--allow-dir`.
- Dates: ISO 8601 text, `"2012-01-01"`, where an argument takes one
(`initial_train_size` of `create_cv`, `test_size` of `forecast`). A count
is a number, never text: `"12"` is rejected.
- Arguments are strict: an unknown argument or a wrong type is an error,
never ignored. Metrics are the names of skforecast:
`mean_absolute_error`, `mean_squared_error`,
`mean_absolute_scaled_error`, and the others the schema lists.
- Messages of the library name the arguments of its Python API: `data` is
`data_path`, `exog` is `exog_path`, `profile`, `plan` and `cv` are the
ids `profile_id`, `plan_id` and `cv_id`, and `forecast()` or
`backtest()` are the tools `forecast` and `backtest`. A message can
also give advice that needs Python (read the file with pandas,
`dayfirst=True`): follow the `hint` of the error instead.
- Future exogenous values (`exog_path`): one row per date of the horizon
(and per series when they are stacked), with the date column of the data.
- Ids are valid while the server runs. After a restart, or when an id was
removed to keep the server within its limits, create the object again.
## Data problems
When the profile, a notice or an error shows a problem in the CSV file
(missing dates, rows without a target, a wrong date column, duplicated
dates, dates written in more than one format or day first, a series without
values, an exogenous column named like a lag or a window feature), tell the
user what it is and what it changes. Never change their file. Only if they
agree, write a corrected copy inside the allowed directory, under a new
name, and `profile` the copy; say what you changed. An error names the
first problem it finds, so the file can have others (the error of
repeated dates with different values also counts the identical repeated
rows and the missing dates): tell the user all of them when you ask, and
ask again before fixing a problem they have not agreed to.
## Foundation models
`ForecasterFoundation` forecasts with a pre-trained model, without
training. Its default model is Chronos-2 (`autogluon/chronos-2-small`).
Each of the others has its own license and size, so tell the user which
model, its license and that it downloads its weights before you choose
one; never switch models on your own. State a license only as a response
gives it: a `ModelLicenseNotice` or a `ModelDownloadNotice` of a plan
with a foundation model or of a comparison that ran one, or the message
of `model_not_allowed`. Through the server they only take
the `estimator_kwargs` `context_length`, `cross_learning`,
`point_estimate`, `max_horizon`, `add_calendar_features` and
`n_fourier_terms`. Models whose license
restricts commercial use, whose weights are gated or whose provider
requires an account (today the prefixes `google/timesfm-3.0`,
`Salesforce/moirai-2`, `priorlabs/tabpfn` and `taharnbl/TS-ICL`), and any
model for which skforecast gives no license information, only run when
the user started the server with `--allow-model PREFIX`; without it they
are `model_not_allowed`. A model
without its backend package installed where the server runs is
`missing_dependency`.
## Responses
- `summary`: plain text with the decisions, their explanations and
statistics. A backtest or a forecast adds its metrics and the minimum,
maximum and mean of the predictions; a comparison, its leaderboard. A
long summary is cut at 20,000 characters; the full text is in
`files.summary`.
- `values_included` is always false: no response holds rows of the data
or of the predictions. The summaries do carry the metrics and the
leaderboard; the rows (predictions, metrics per fold or series, the
whole leaderboard) are CSV files listed in `files`. Read them when you
need the values.
- `notices`: the warnings of the call, with their source (`data`, `plan`
or `runtime`); at most 20, and `notices_omitted` counts the rest. A
profile carries the problems of the data (`DataProfileWarning`: missing
dates, short series, missing values), and a plan carries them again with
its own warnings, so you see them where you decide.
- `links`: the ids an object was built from. `changeable`: the arguments
of `refine_plan` or `create_cv` that build a variant of it.
## Errors
An error arrives as `Error executing tool <name>: ` followed by a JSON
object `{code, message, field, hint, details}`. Act on `code` and `field`,
and follow `hint` when there is one:
| code | What to do |
|---|---|
| `invalid_argument` | Fix the argument named in `field`, as the message says. |
| `insufficient_data` | Ask for less, as the message says: a shorter horizon, fewer lags, or a first training set that leaves room for the folds (smaller) or for the window of the forecaster (a later `initial_train_size`). A target column without any value, or a series too short for the forecaster (the message names it), is also reported this way. |
| `data_not_found`, `invalid_path`, `url_not_allowed` | Pass the absolute path of a CSV file inside the allowed directory. |
| `path_not_allowed` | The file is outside the allowed directory (`details.allowed_dir`). Do not copy or move it yourself: tell the user, who can copy it there or restart the server with another `--allow-dir`. |
| `data_unreadable` | The file is not a CSV the server can read (empty, binary, not UTF-8, or rows with more fields than the header). Tell the user, as for the data problems above. |
| `file_too_large` | The file is larger than the server reads (`--max-file-mb`, 256 MB by default): pass a smaller file, or ask the user to raise the limit. |
| `data_changed` | The file changed: call `profile` again (or the tool again for an exogenous file). |
| `unknown_id` | Use an id from `list_objects`, or create the object again. |
| `inconsistent_ids` | Pass `backtest` a plan and a strategy built from the same profile. |
| `execution_failed`, `all_candidates_failed` | `get_failure(details.failure_id)` returns the traceback and the code. |
| `missing_dependency` | Tell the user which package to install; `hint` says how, for pip and for uvx. A foundation model without its backend fails this way before running (in `compare`, such a candidate fails and is ranked last). |
| `model_not_allowed` | Tell the user the license in the message; only if they accept it, ask them to restart the server with the `--allow-model` option of `hint`. |
| `internal_error` | Report it to the user with `details.error_id`, which finds the message in the log of the server; do not retry with the same inputs. |
A candidate of `compare` that fails is ranked last instead of failing the
call: `get_failure(comparison_id, candidate)` says why.
## Privacy
The server never sends rows of data in a response. Messages of errors and
warnings are forwarded as the library writes them: they can name columns
and series ids and quote up to 5 values of the data (categories, dates).
The server cuts a message at 4,000 characters, a hint at 1,000, each text
of `details` at 500 and each notice at 1,000. An unexpected error
(`internal_error`) carries only the type of the exception and an id
(`details.error_id`): its message, which can quote a value, goes to the log
of the server with that id. A failure (`get_failure`) holds a traceback,
which can quote values. The scripts of `get_code` and the summary of a
plan name the path of the data file; the other summaries and the failures
do not. Tell the user when they ask what you can see.
Foundation models download their weights from the Hugging Face Hub the
first time they run; a `ModelDownloadNotice` (source `plan`) says so, with
the license skforecast registers for models of that name, the first time a
model whose weights are not in the local cache is used. Otherwise a
`ModelLicenseNotice` gives that license. Later runs still
contact the Hub to check the cached weights, without sending data. The
user can forbid any connection to it by starting the server with
`HF_HUB_OFFLINE=1`.
See also¶
- MCP server API: every tool, its arguments, its files and its errors.
- Using the CLI: the
mcpcommand next to the other commands. - Skills for the LLM: the skforecast guides that
ask()sends to its model, not to a coding agent. - Agent Skills specification: the format of a
SKILL.mdfile.