MCP server¶
The MCP server serves the deterministic workflow of ForecastingAssistant to coding agents (Claude Code, Cursor, Claude Desktop and any other MCP client). It needs the mcp extra:
pip install "skforecast-ai[mcp]"
skforecast-ai mcp --allow-dir /path/to/data
Importing skforecast_ai never imports the server or the mcp package. The options of the command are in the CLI reference; how to connect an agent and give it the skill that ships with the package, in MCP server for coding agents.
Tools¶
Every tool that creates an object returns a ToolResult with its id; later tools take ids, never objects, so no plan, profile or cross-validation strategy is ever accepted as JSON.
| Tool | Takes | Returns |
|---|---|---|
profile |
data_path, target (without it, the error lists the columns of the file), date_column, series_id_column, exog_columns (the exogenous columns to use, [] for none; null for every other column) |
a profile id |
plan |
profile_id, steps, interval, forecaster, estimator, estimator_kwargs, lags, window_features, metric, use_exog, differentiation, calendar_features, target_transformer, dropna_from_series |
a plan id |
refine_plan |
plan_id, overrides (the keys of refine_plan(); an omitted key keeps the value of the plan, and every key but forecaster, estimator and steps set to null asks for the default) |
a new plan id |
create_cv |
plan_id and the arguments of create_cv() |
a cv id, with its cost: folds, trainings, estimator fits and the inference windows of a foundation model (at most one per series and fold), for its plan and for a compare without candidates |
backtest |
cv_id, plan_id (another plan of the same profile; null for the plan of the strategy) |
a backtest id; predictions and metrics in CSV files |
compare |
cv_id, candidates ([{name, config}], null for those of the profile), interval, metric, baseline |
a comparison id, the plan of the winner in links.best_plan_id; the leaderboard and the predictions and metrics of the winner in CSV files |
forecast |
plan_id, test_size (null to forecast the future), exog_path |
a forecast id; predictions (and metrics in evaluation) in CSV files |
get_code |
object_id, candidate (of a comparison) |
a CodeResult: the script that ran, or that would run for a plan, and requirements, the packages it needs with the versions installed where the server runs. A forecast with exog_path reads the future exogenous values from exog_future.csv in its working directory: copy the file there to run it |
get_failure |
object_id (details.failure_id of an error, or a comparison), candidate |
a FailureResult: the traceback and the code of a failed run |
list_objects |
kind |
an ObjectList |
describe_object |
object_id |
the ToolResult that created the object |
compare ranks the candidates by its metric; without one, by the metric chosen for the plan of the strategy (metric of plan or refine_plan), and otherwise by the metric selected from the data (compare() in Python, which takes no plan, uses the metric selected from the data). A candidate whose arguments do not match the schema of the tool (an unknown key, a forecaster that does not exist) or that names a foundation model the server does not run is rejected with the whole call, before anything runs. A candidate that passes those checks and whose configuration is not valid (an unknown estimator, arguments the forecaster cannot use) is ranked last with its error, as in Python, and get_failure returns why it failed. compare sends a progress notification when each candidate starts and ends (2 * completed + started over 2 * total). Every tool that runs the core also sends one every 5 seconds while its worker thread is busy, to a client that asked for progress: k notifications after an event of progress p send p + k / (k + 1) (from 0, with no total, for a call without events), so the values grow without reaching the next event, and the message names what runs and for how long ("ForecasterStats: running (35 s)"). A cancelled compare stops before its next candidate; any other cancelled call waits for its work to end. A cancelled call registers nothing. While a call runs, only get_code, get_failure, list_objects and describe_object answer; the other tools wait their turn.
Foundation models only take the estimator_kwargs context_length, cross_learning, point_estimate, max_horizon, add_calendar_features and n_fourier_terms through the server: other arguments of the adapters of skforecast can send the data to a remote service or download files. The server decides from the FoundationModelInfo of skforecast: a foundation model whose license restricts commercial use (commercial_use_restricted), whose weights are gated (requires_hf_auth) or whose provider requires its own account (requires_provider_auth), or for which skforecast does not give these facts, only runs when the server was started with --allow-model and a prefix of its model ID; otherwise plan, refine_plan and compare raise model_not_allowed. The weights of a foundation model are downloaded the first time it runs: the first time a plan or a candidate uses a model whose weights are not in the local Hugging Face cache (looked up under weights_repo_id), the response carries a ModelDownloadNotice (source plan) with its license. Every other plan with a foundation model, and every comparison that ran one, carries a ModelLicenseNotice with that license, so the agent never has to state it from memory. A model whose backend keeps its weights outside that cache (weights_in_hf_cache is false, as for TabPFN) gets the notice the first time it is used, saying that the server cannot tell whether they are downloaded. Set HF_HUB_OFFLINE=1 in the environment of the server to forbid downloads from the Hugging Face Hub.
Arguments are checked strictly: an unknown argument, a number written as text or a float for an integer is an error. Dates are ISO 8601 text.
The files of a response are CSV files with the index of the data: predictions and metrics of a backtest or a forecast, leaderboard, best_predictions and best_metrics of a comparison. A summary, a script or a failure too long for a response is written to a file as well.
The notices of a response are the warnings the call emitted, deduplicated, with their source: data (reading or profiling the data), plan (a warning the plan carries in plan.warnings) or runtime. Deprecation warnings go to the log of the server (stderr) instead. Besides those, a profile carries data_profile.warnings (category DataProfileWarning, source data), a plan carries them too, with any text of plan.warnings not emitted in the call (category PlanWarning), and create_cv carries the LongTrainingWarning that a backtest of the strategy will emit (above 50 estimator fits, or 2000 inference windows of a foundation model, one per series and fold). The server adds notices of its own, besides the ModelDownloadNotice and the ModelLicenseNotice of a foundation model: create_cv carries a CostNotice (source runtime) when the backtest of the strategy, or a compare without candidates on it, is above those thresholds, next to the warning of the library that states the cost: the notice tells the agent to stop, to tell the user the numbers and the cheaper strategies, and to run the expensive one only if they choose it; create_cv also carries a CompareCostNotice (source runtime) when a compare without candidates on the strategy would cost more than its plan; create_cv also carries a MissingValuesNotice (source data) when the backtest of its plan can fail on missing values of the target, next to the warning of the library that says so: the notice tells the agent that those values are the user's (ask before filling, dropping or writing any, also in a copy), that an estimator that accepts missing values avoids the error without touching the data, and to say it when it switches or when it forecasts without a backtest; a backtest or a forecast whose metrics include MASE or RMSSE carries a MetricReferenceNotice (source runtime) saying that they are scaled by the one-step naive forecast on the training data, not by a seasonal naive forecast nor by the baseline of compare; a backtest, a forecast or a comparison that computes MAPE carries a MetricUnitNotice saying that it is a fraction (1.245 is 124.5 %); a forecast run with test_size carries a HoldoutEvaluationNotice (source runtime) with the dates it predicts, which are already in the data, so it is not presented as the forecast of the future; a plan that uses exogenous variables carries a FutureExogNotice (source plan) saying that forecast needs their future values from the user, and a forecast of a plan that leaves the exogenous columns of the data out carries an ExogLeftOutNotice. An error of the library about the content of a file (a problem of the CSV in profile; missing values that a prediction reads, or final rows without a target, in backtest, compare and forecast; the file of future exogenous values, or its absence, in forecast) has a hint of the server that leaves the values to the user: the agent tells them and asks, and never writes, fills in or drops values on its own.
compare without interval computes the interval of the plan the strategy was built for (unlike compare() in Python, whose default is no interval), so the plan of the winner keeps it; to compare without one, build the strategy from a plan without interval. Other decisions of that plan (lags, use_exog...) do not reach the candidates: each one is planned from the profile with its own config. A candidate with a differentiation other than the one of the strategy runs on a copy of the strategy with its own order, and the summary says so. backtest with a plan_id whose differentiation order is not the one of the strategy is an invalid_argument. The seasonal naive baseline and ForecasterStats only compute symmetric intervals (lower + upper = 1): with an asymmetric one the comparison has no baseline (its summary says why), and a ForecasterStats candidate fails and is ranked last. backtest and forecast of a ForecasterFoundation plan whose backend package is not installed where the server runs raise missing_dependency before running anything; a candidate of compare with that model fails and is ranked last, as in Python.
Limits¶
- The server keeps at most 256 objects and about 1 GB of them (
--max-objects,--max-memory-mb); beyond that, the least recently used ones are removed and their ids raiseunknown_idsaying so. Objects built from a removed one keep working. - Ids do not survive a restart of the server: an id of a previous run raises
unknown_idsaying so. - A summary, a script or a failure longer than 20,000 characters is cut in the response, and the full text is written to the output directory (
files). - Calls run one at a time; a call waits for the previous one to end.
- CSV files larger than
--max-file-mb(256 MB by default) are rejected before being read, andplanandrefine_planreject astepslonger than the longest series of the profile (invalid_argument). - The output directory (
--output-dir, by default a new temporary directory that is kept when the server stops) is also the working directory of the server. - The scripts run in the process of the server, with the permissions of the user who started it. The server trusts the agent as much as that user: it limits what the agent can read (
--allow-dir) and pass, not what the scripts it builds can do. - The scripts of
get_codeand the summary of a plan (its script lists the file it reads) name the path of the data file; the other summaries, the failures and the rest of the responses do not.
Errors¶
A failure of a tool reaches the agent as an error result whose text is Error executing tool <name>: followed by a JSON object with code, message, field, hint and details. code is one of the codes of the core or one of the server:
code |
When |
|---|---|
unknown_id |
The id does not exist, was removed to stay within the limits of the server (details.removed), or comes from a previous run of the server. |
inconsistent_ids |
The plan and the cross-validation strategy passed to backtest come from different profiles. |
invalid_path |
The path is not absolute, does not end in .csv or holds a control character. |
path_not_allowed |
The path is outside the directory given to --allow-dir, also after resolving symbolic links. It is checked before looking at the file, so the error does not say whether a file outside exists. Its hint names the directory and tells the agent to ask the user to copy the file there or to change --allow-dir, not to copy it itself. |
url_not_allowed |
The path is a URL. Download the file into the allowed directory. |
file_too_large |
The CSV file (data or future exogenous values) is larger than --max-file-mb (256 MB by default; 0 for no limit). Its size is checked before reading it; details has the size and the limit. |
model_not_allowed |
A foundation model that --allow-model does not allow: its license restricts commercial use, its weights are gated, its provider requires an account, or skforecast gives no license information for it. details has its license and those facts; hint names the option to ask the user for. |
data_changed |
The CSV file changed since it was profiled, or the data or the exogenous file changed while the server was reading it. Nothing is registered; call profile again (or the tool again, for the exogenous file). The Python API profiles such data again and runs; the server asks for a new profile, so that every id keeps describing the file it was built from. |
When a script fails (execution_failed) or every candidate of a comparison fails (all_candidates_failed), details.failure_id names the full failure, which get_failure returns: it never goes in the error itself. The error names only the type of what failed when skforecast-ai did not raise it (an error of pandas or of the estimator can quote a value of the data), as an internal_error does. The failure that get_failure returns holds the whole message, so it can quote values of the data; its traceback names the files by module, without the directory they are installed in, and the code it holds reads the data in memory and does not name the path of the data file.
An error of profile about the content of the file (repeated dates with different values, series of different frequencies) gets a hint that leaves the fix to the user: the agent must tell them and ask before writing a corrected copy. The messages of the core are forwarded as they are: they can name columns, series ids and values of the data, such as categories or dates (at most 5 values each). An error that skforecast-ai did not raise itself is an internal_error with only its type and an id (details.error_type, details.error_id): its message and traceback, which can quote a value, are written to the log of the server (stderr) with that id. The server cuts a message to 4,000 characters, a hint to 1,000 and each text of details to 500, and a response carries at most 20 notices of 1,000 characters each (notices_omitted counts the rest).
skforecast_ai.mcp.create_server ¶
create_server(
allow_dir,
output_dir=None,
max_objects=DEFAULT_MAX_OBJECTS,
max_memory_mb=DEFAULT_MAX_MEMORY_MB,
allow_models=(),
max_file_mb=DEFAULT_MAX_FILE_MB,
)
Create the MCP server of skforecast-ai, without running it.
Useful to test the server in memory (mcp.Client(server)) or to run it
with another transport. run_server() creates it and serves it over
stdio, as the skforecast-ai mcp command does.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
allow_dir
|
(str, Path)
|
Directory the server may read data from. Only absolute paths of CSV files inside it are accepted, also after resolving symbolic links. |
required |
output_dir
|
(str, Path)
|
Directory of the files the server writes: predictions, metrics and leaderboards as CSV, and summaries, scripts and failures too long for a response. It is created if it does not exist. When None, a new temporary directory, kept when the server stops. |
None
|
max_objects
|
int
|
Most objects the server keeps; the least recently used ones are removed beyond it. |
256
|
max_memory_mb
|
int
|
Memory, in MB, the objects may take (an estimate); the least recently used ones are removed beyond it. |
1024
|
allow_models
|
iterable of str
|
Model ID prefixes of foundation models that the server may run
although their license restricts commercial use, their weights are
gated, their provider requires an account or skforecast gives no
license information ( |
()
|
max_file_mb
|
int
|
Largest CSV file (data or future exogenous values) the server reads, in MB, checked on the size of the file before reading it. 0 for no limit. |
256
|
Returns:
| Name | Type | Description |
|---|---|---|
server |
MCPServer
|
Server of the |
Source code in skforecast_ai/mcp/server.py
2807 2808 2809 2810 2811 2812 2813 2814 2815 2816 2817 2818 2819 2820 2821 2822 2823 2824 2825 2826 2827 2828 2829 2830 2831 2832 2833 2834 2835 2836 2837 2838 2839 2840 2841 2842 2843 2844 2845 2846 2847 2848 2849 2850 2851 2852 2853 2854 2855 2856 2857 2858 2859 2860 | |
skforecast_ai.mcp.run_server ¶
run_server(
allow_dir,
output_dir=None,
max_objects=DEFAULT_MAX_OBJECTS,
max_memory_mb=DEFAULT_MAX_MEMORY_MB,
allow_models=(),
max_file_mb=DEFAULT_MAX_FILE_MB,
)
Run the MCP server of skforecast-ai over stdio until the client closes.
The working directory of the process becomes output_dir, so a library
that writes files next to it (CatBoost writes catboost_info/) does not
write them into the project of the user. While it serves, the logger
skforecast_ai.mcp writes one plain line per record to the standard
error and does not pass its records to the root logger; both are
restored when it returns. When the client disconnects, also during a
call, it logs one line and returns instead of raising, and the standard
output (the closed pipe of the client) is pointed to the null device so
flushing it at exit does not fail again.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
allow_dir
|
(str, Path)
|
Directory the server may read data from. Only absolute paths of CSV files inside it are accepted, also after resolving symbolic links. |
required |
output_dir
|
(str, Path)
|
Directory of the files the server writes: predictions, metrics and leaderboards as CSV, and summaries, scripts and failures too long for a response. It is created if it does not exist. When None, a new temporary directory, kept when the server stops. |
None
|
max_objects
|
int
|
Most objects the server keeps. |
256
|
max_memory_mb
|
int
|
Memory, in MB, the objects may take (an estimate). |
1024
|
allow_models
|
iterable of str
|
Model ID prefixes of foundation models that the server may run although their license restricts commercial use, their weights are gated, their provider requires an account or skforecast gives no license information. |
()
|
max_file_mb
|
int
|
Largest CSV file the server reads, in MB; 0 for no limit. |
256
|
Returns:
| Type | Description |
|---|---|
None
|
|
Source code in skforecast_ai/mcp/server.py
2863 2864 2865 2866 2867 2868 2869 2870 2871 2872 2873 2874 2875 2876 2877 2878 2879 2880 2881 2882 2883 2884 2885 2886 2887 2888 2889 2890 2891 2892 2893 2894 2895 2896 2897 2898 2899 2900 2901 2902 2903 2904 2905 2906 2907 2908 2909 2910 2911 2912 2913 2914 2915 2916 2917 2918 2919 2920 2921 2922 2923 2924 2925 2926 2927 2928 2929 2930 2931 2932 2933 2934 2935 2936 2937 2938 2939 2940 2941 2942 2943 2944 2945 2946 2947 2948 2949 | |
skforecast_ai.mcp.models ¶
Classes:
| Name | Description |
|---|---|
ToolResult |
Envelope returned by the tools that create an object, and by |
ToolNotice |
A warning emitted while a tool call ran. |
CodeResult |
Script of an object, returned by the |
FailureResult |
Full description of a failure, returned by the |
ObjectInfo |
An object registered in the server, as listed by |
ObjectList |
Objects registered in the server, returned by |
Classes¶
ToolResult ¶
Bases: BaseModel
Envelope returned by the tools that create an object, and by
describe_object.
Attributes:
| Name | Type | Description |
|---|---|---|
id |
str
|
Id of the object, to pass to the tools that take it. |
kind |
str
|
Kind of the object: |
links |
dict
|
Ids of related objects: those it was built from ( |
summary |
str
|
Plain-text description of the object ( |
summary_truncated |
bool
|
Whether |
notices |
list of ToolNotice
|
Warnings the call emitted, the first 20 distinct ones. Deprecation warnings go to the log of the server instead, and the failures of the candidates of a comparison are in its result. |
notices_omitted |
int
|
Number of distinct warnings left out of |
files |
dict
|
Absolute paths of the files written for the object, by role:
|
values_included |
bool
|
Always False: the response holds no rows of the data or of the
predictions. The summary carries statistics of the predictions, the
metrics and the leaderboard of a comparison; the rows are in
|
cost |
(dict, None)
|
Cost of a cross-validation strategy, a backtest or a comparison:
|
changeable |
list of str
|
Arguments of |
ToolNotice ¶
Bases: BaseModel
A warning emitted while a tool call ran.
Attributes:
| Name | Type | Description |
|---|---|---|
source |
str
|
Where the warning comes from: |
category |
str
|
Class name of the warning (e.g. |
message |
str
|
Text of the warning, without the suggestion of skforecast on how to silence it, cut to 1,000 characters. |
count |
int
|
Number of times the call emitted it. |
CodeResult ¶
Bases: BaseModel
Script of an object, returned by the get_code tool.
Attributes:
| Name | Type | Description |
|---|---|---|
id |
str
|
Id of the object. |
kind |
str
|
Kind of the object. |
candidate |
(str, None)
|
Candidate of a comparison whose script it is. |
code |
str
|
Python code, cut to 20,000 characters. |
code_truncated |
bool
|
Whether |
requirements |
list of str
|
Packages to install to run the script outside the server: those of
the modules it imports and the backend of its foundation model,
each with the version installed where the server runs
( |
files |
dict
|
Absolute path of the file with the full code, when it was cut. |
FailureResult ¶
Bases: BaseModel
Full description of a failure, returned by the get_failure tool.
Attributes:
| Name | Type | Description |
|---|---|---|
id |
str
|
Id of the failure, or of the comparison whose candidate failed. |
candidate |
(str, None)
|
Candidate of the comparison that failed. |
text |
str
|
Error, line and statement that failed, traceback and code that ran, cut to 20,000 characters. Like the messages of the errors, it can quote values of the data. The code that ran reads the data in memory, so it does not name the path of the data file. |
text_truncated |
bool
|
Whether |
files |
dict
|
Absolute path of the file with the full text, when it was cut. |
ObjectInfo ¶
ObjectList ¶
Bases: BaseModel
Objects registered in the server, returned by list_objects.
Attributes:
| Name | Type | Description |
|---|---|---|
objects |
list of ObjectInfo
|
Registered objects, most recently used first. |
max_objects |
int
|
Most objects the server keeps; the least recently used ones are removed beyond it. |
max_memory_mb |
int
|
Memory, in MB, the objects may take; the least recently used ones are removed beyond it. |
removed |
int
|
Number of objects removed so far to stay within those limits. A tool that receives the id of a removed object says so. |