使用pyarrow时无法加载libhdfs

rjzwgtxy  于 2021-06-01  发布在  Hadoop
关注(0)|答案(1)|浏览(1125)

我正试图通过pyarrow连接到hdfs,但它不起作用,因为 libhdfs 无法加载库。 libhdfs.so$HADOOP_HOME/lib/native 以及在 $ARROW_LIBHDFS_DIR .

print(os.environ['ARROW_LIBHDFS_DIR'])
fs = hdfs.connect()

bash-3.2$ ls $ARROW_LIBHDFS_DIR
examples        libhadoop.so.1.0.0  libhdfs.a       libnativetask.a
libhadoop.a     libhadooppipes.a    libhdfs.so      libnativetask.so
libhadoop.so        libhadooputils.a    libhdfs.so.0.0.0    libnativetask.so.1.0.0

我得到的错误是:

Traceback (most recent call last):
  File "wine-pred-ml.py", line 31, in <module>
    fs = hdfs.connect()
  File "/Users/PVZP/Library/Python/2.7/lib/python/site-packages/pyarrow/hdfs.py", line 183, in connect
    extra_conf=extra_conf)
  File "/Users/PVZP/Library/Python/2.7/lib/python/site-packages/pyarrow/hdfs.py", line 37, in __init__
    self._connect(host, port, user, kerb_ticket, driver, extra_conf)
  File "pyarrow/io-hdfs.pxi", line 89, in pyarrow.lib.HadoopFileSystem._connect
  File "pyarrow/error.pxi", line 83, in pyarrow.lib.check_status
pyarrow.lib.ArrowIOError: Unable to load libhdfs
hi3rlvi2

hi3rlvi21#

这解决了我的问题:

conda install libhdfs3 pyarrow

在script.py中:

import os
os.environ['ARROW_LIBHDFS_DIR'] = '/opt/cloudera/parcels/CDH/lib64/'

其中path是libhdfs3所在的目录——在我的例子中,这是cloudera托管lib的地方

相关问题