配置单元-内存不足异常-java堆空间

nhhxz33t  于 2021-05-29  发布在  Hadoop
关注(0)|答案(1)|浏览(338)

我正在Parquet文件(使用spark创建)上运行一个配置单元插入。配置单元插入正在使用partitioned by子句。但是当屏幕最后打印类似“loading partition{=xyz,=123,=}这样的消息时,一个java堆空间异常就出现了。

java.lang.OutOfMemoryError: Java heap space
         at java.util.HashMap.createEntry(HashMap.java:901)
         at java.util.HashMap.addEntry(HashMap.java:888)
         at java.util.HashMap.put(HashMap.java:509)
         at org.apache.hadoop.hive.metastore.api.Partition.<init>(Partition.java:229)
         at org.apache.hadoop.hive.metastore.HiveMetaStoreClient.deepCopy(HiveMetaStoreClient.java:1356)
         at org.apache.hadoop.hive.metastore.HiveMetaStoreClient.getPartitionWithAuthInfo(HiveMetaStoreClient.java:1003)
         at sun.reflect.GeneratedMethodAccessor11.invoke(Unknown Source)
         at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)
         at java.lang.reflect.Method.invoke(Method.java:606)
         at org.apache.hadoop.hive.metastore.RetryingMetaStoreClient.invoke(RetryingMetaStoreClient.java:89)
         at com.sun.proxy.$Proxy9.getPartitionWithAuthInfo(Unknown Source)
         at org.apache.hadoop.hive.ql.metadata.Hive.getPartition(Hive.java:1611)
         at org.apache.hadoop.hive.ql.metadata.Hive.getPartition(Hive.java:1565)
         at org.apache.hadoop.hive.ql.exec.StatsTask.getPartitionsList(StatsTask.java:403)
         at org.apache.hadoop.hive.ql.exec.StatsTask.aggregateStats(StatsTask.java:150)
         at org.apache.hadoop.hive.ql.exec.StatsTask.execute(StatsTask.java:117)
         at org.apache.hadoop.hive.ql.exec.Task.executeTask(Task.java:153)
         at org.apache.hadoop.hive.ql.exec.TaskRunner.runSequential(TaskRunner.java:85)
         at org.apache.hadoop.hive.ql.Driver.launchTask(Driver.java:1508)
         at org.apache.hadoop.hive.ql.Driver.execute(Driver.java:1275)
         at org.apache.hadoop.hive.ql.Driver.runInternal(Driver.java:1093)
         at org.apache.hadoop.hive.ql.Driver.run(Driver.java:916)
         at org.apache.hadoop.hive.ql.Driver.run(Driver.java:906)
         at org.apache.hadoop.hive.cli.CliDriver.processLocalCmd(CliDriver.java:268)
         at org.apache.hadoop.hive.cli.CliDriver.processCmd(CliDriver.java:220)
         at org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:423)
         at org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:359)
         at org.apache.hadoop.hive.cli.CliDriver.processReader(CliDriver.java:456)
         at org.apache.hadoop.hive.cli.CliDriver.processFile(CliDriver.java:466)
         at org.apache.hadoop.hive.cli.CliDriver.executeDriver(CliDriver.java:748)
         at org.apache.hadoop.hive.cli.CliDriver.run(CliDriver.java:686)
         at org.apache.hadoop.hive.cli.CliDriver.main(CliDriver.java:625)

我在运行作业时设置了以下属性,并尝试将值更改为更高和更低,但每次最后都会发现此错误。
切换的属性:

set mapred.map.tasks=100;
 set mapred.reduce.tasks=100;
 set mapreduce.map.java.opts=-Xmx4096m;
 set mapreduce.reduce.java.opts=-Xmx4096m;
 set hive.exec.max.dynamic.partitions.pernode=100000;
 set hive.exec.max.dynamic.partitions=100000;

请说明这里出了什么问题。配置单元版本为0.13。
Hive环境.sh

if [ "$SERVICE" = "cli" ]; then
   if [ -z "$DEBUG" ]; then
     export HADOOP_OPTS="$HADOOP_OPTS -XX:NewRatio=12 -Xms12288m -XX:MaxHeapFreeRatio=40 -XX:MinHeapFreeRatio=15 -XX:+UseParNewGC -XX:-UseGCOverheadLimit"
   else
     export HADOOP_OPTS="$HADOOP_OPTS -XX:NewRatio=12 -Xms12288m -XX:MaxHeapFreeRatio=40 -XX:MinHeapFreeRatio=15 -XX:-UseGCOverheadLimit"
   fi
 fi

# The heap size of the jvm stared by hive shell script can be controlled via:

# 

export HADOOP_HEAPSIZE=4096
wtzytmuj

wtzytmuj1#

可能与hive-10149有关。尝试设置 hive.optimize.sort.dynamic.partitiontrue .

相关问题