pandas 在Python中进行多TS统计预测时,数据框定义出错

aiazj4mn  于 2022-11-20  发布在  Python
关注(0)|答案(1)|浏览(130)

我试图在python中复制这段用于统计预测的代码,我遇到了一个奇怪的错误***“name 'forecasts' is not defined”***这很奇怪,因为我之前能够复制代码而没有任何错误。
这里与参考代码(在下面的链接中给出,以及我能够成功实现的代码)的区别在于,我没有使用训练集并提取过去6个月的数据进行评估,而是使用整个训练数据来创建统计预测。
例如:如果我的时间序列数据是9月22日之前的数据,我想将9月22日之前的全部数据作为统计模型的训练集,而之前的训练数据是3月22日之前的时间序列,其余6个月是测试数据。但现在出现了错误,我无法理解为什么逻辑是相同的?
附件为计算所用的简化数据框:

{'Key': {0: 65162552161356, 1: 65162552635756, 2: 65162552843456, 3: 65162552842856, 4: 65162552736856}, '2021-04-01': {0: 31, 1: 0, 2: 281, 3: 207, 4: 55}, '2021-05-01': {0: 25, 1: 0, 2: 72, 3: 104, 4: 6}, '2021-06-01': {0: 16, 1: 0, 2: 108, 3: 32, 4: 14}, '2021-07-01': {0: 8, 1: 0, 2: 107, 3: 78, 4: 10}, '2021-08-01': {0: 21, 1: 0, 2: 80, 3: 40, 4: 9}, '2021-09-01': {0: 24, 1: 0, 2: 40, 3: 73, 4: 3}, '2021-10-01': {0: 13, 1: 0, 2: 36, 3: 79, 4: 11}, '2021-11-01': {0: 59, 1: 0, 2: 65, 3: 139, 4: 14}, '2021-12-01': {0: 51, 1: 0, 2: 41, 3: 87, 4: 10}, '2022-01-01': {0: 2, 1: 0, 2: 43, 3: 47, 4: 6}, '2022-02-01': {0: 0, 1: 0, 2: 0, 3: 63, 4: 3}, '2022-03-01': {0: 0, 1: 0, 2: 16, 3: 76, 4: 18}, '2022-04-01': {0: 0, 1: 0, 2: 37, 3: 32, 4: 8}, '2022-05-01': {0: 0, 1: 0, 2: 106, 3: 96, 4: 40}, '2022-06-01': {0: 0, 1: 0, 2: 101, 3: 75, 4: 16}, '2022-07-01': {0: 0, 1: 0, 2: 60, 3: 46, 4: 14}, '2022-08-01': {0: 0, 1: 0, 2: 73, 3: 91, 4: 13}, '2022-09-01': {0: 0, 1: 0, 2: 19, 3: 17, 4: 2}}

以下是参考链接:https://towardsdatascience.com/time-series-forecasting-with-statistical-models-f08dcd1d24d1

import random
from itertools import product
from IPython.display import display, Markdown
from multiprocessing import cpu_count
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
from statsforecast import StatsForecast
from nixtlats.data.datasets.m4 import M4, M4Info
from statsforecast.models import (
    adida, 
    croston_classic, 
    croston_sba, 
    croston_optimized,
    historic_average,
    imapa,
    naive,
    random_walk_with_drift, 
    seasonal_exponential_smoothing,
    seasonal_naive, 
    seasonal_window_average,
    ses, 
    tsb,
    window_average
)
df = pd.read_excel ('C:/X/X/X/2.1 Demand_Data_Used.xlsx')
df['Key'] = df['Key'].astype(str)
df = pd.melt(df,id_vars='Key',value_vars=list(df.columns[1:]),var_name ='ds')
df.columns = df.columns.str.replace('Key', 'unique_id')
df.columns = df.columns.str.replace('value', 'y')
df["ds"] = pd.to_datetime(df["ds"],format='%Y-%m-%d')
df=df[["ds","unique_id","y"]]

df['unique_id'] = df['unique_id'].astype('object')
df = df.set_index('unique_id')
df.reset_index()

seasonality = 30 #Monthly data

models = [
    adida,
    croston_classic,
    croston_sba,
    croston_optimized,
    historic_average,
    imapa,
    naive,
    random_walk_with_drift,
    (seasonal_exponential_smoothing, seasonality, 0.2),
    (seasonal_naive, seasonality),
    (seasonal_window_average, seasonality, 2 * seasonality),
    (ses, 0.1),
    (tsb, 0.3, 0.2),
    (window_average, 2 * seasonality)
    ]

fcst = StatsForecast(df=df, models=models, freq='MS', n_jobs=cpu_count())
%time forecasts = fcst.forecast(6)
forecasts.reset_index()

forecasts = forecasts.reset_index().merge(df_test, how='left', on=['unique_id', 'ds'])
models = forecasts.drop(columns=['unique_id', 'ds', 'y']).columns.to_list()

附件是错误图像:

有谁能让我知道我做错了什么吗?我将非常感激。

bkhjykvo

bkhjykvo1#

问题的出现是因为Croston一家。我已经打开了一个issue来解决这个问题。在此期间,跳过那些型号是有效的。

models = [
    adida,
    #croston_classic,
    #croston_sba,
    #croston_optimized,
    historic_average,
    imapa,
    naive,
    random_walk_with_drift,
    (seasonal_exponential_smoothing, seasonality, 0.2),
    (seasonal_naive, seasonality),
    (seasonal_window_average, seasonality, 2 * seasonality),
    (ses, 0.1),
    (tsb, 0.3, 0.2),
    (window_average, 2 * seasonality)
    ]
fcst = StatsForecast(df=df, models=models, freq='MS', n_jobs=cpu_count())
fcst.forecast(6)

更新:
StatsForecast的最新版本修复了此问题。您可以使用以下命令使用它:

from statsforecast.models import CrostonClassic, CrostonSBA, CrostonOptimized

models = [
    CrostonClassic(),
    CrostonSBA(),
    CrostonOptimized()
]

fcst = StatsForecast(df=df, models=models, freq='MS', n_jobs=cpu_count())
fcst.forecast(6)

相关问题