自动故障转移在hadoop中不起作用

zd287kbt  于 2021-05-29  发布在  Hadoop
关注(0)|答案(2)|浏览(517)

我正在尝试构建一个3节点群集(2个namenode(nn1,nn2)和1个datanode(dn1))。使用namenode webui,我可以查看nn1是活动的,nn2是备用的。但是,当我杀死活动的nn1时,备用nn2将不处于活动状态。请帮助我我做错了什么或者需要修改什么
nn1/etc/主机

  1. 127.0.0.1 localhost
  2. 192.168.10.153 nn1
  3. 192.168.10.154 dn1
  4. 192.168.10.155 nn2

nn2/etc/主机

  1. 127.0.0.1 localhost nn2
  2. 127.0.1.1 ubuntu
  3. # The following lines are desirable for IPv6 capable hosts
  4. ::1 ip6-localhost ip6-loopback
  5. fe00::0 ip6-localnet
  6. ff00::0 ip6-mcastprefix
  7. ff02::1 ip6-allnodes
  8. ff02::2 ip6-allrouters

core-site.xml(nn1,nn2)

  1. <configuration>
  2. <property>
  3. <name>fs.defaultFS</name>
  4. <value>hdfs://192.168.10.153:8020</value>
  5. </property>
  6. <property>
  7. <name>dfs.journalnode.edits.dir</name>
  8. <value>/usr/local/hadoop/hdfs/data/jn</value>
  9. </property>
  10. <property>
  11. <name>ha.zookeeper.quorum</name>
  12. <value>192.168.10.153:2181,192.168.10.155:2181,192.168.10.154:2181</value>
  13. </property>
  14. </configuration>

hdfs site.xml(nn1,nn2,dn1)

  1. <property>
  2. <name>dfs.replication</name>
  3. <value>1</value>
  4. </property>
  5. <property>
  6. <name>dfs.permissions</name>
  7. <value>false</value>
  8. </property>
  9. <property>
  10. <name>dfs.nameservices</name>
  11. <value>ha-cluster</value>
  12. </property>
  13. <property>
  14. <name>dfs.ha.namenodes.ha-cluster</name>
  15. <value>nn1,nn2</value>
  16. </property>
  17. <property>
  18. <name>dfs.namenode.rpc-address.ha-cluster.nn1</name>
  19. <value>192.168.10.153:9000</value>
  20. </property>
  21. <property>
  22. <name>dfs.namenode.rpc-address.ha-cluster.nn2</name>
  23. <value>192.168.10.155:9000</value>
  24. </property>
  25. <property>/usr/local/hadoop/hdfs/datanode</value>
  26. <name>dfs.namenode.http-address.ha-cluster.nn1</name>
  27. <value>192.168.10.153:50070</value>
  28. </property>
  29. <property>
  30. <name>dfs.namenode.http-address.ha-cluster.nn2</name>
  31. <value>192.168.10.155:50070</value>
  32. </property>
  33. <property>
  34. <name>dfs.namenode.shared.edits.dir</name>
  35. <value>qjournal://192.168.10.153:8485;192.168.10.155:8485;192.168.10.154:8485/ha-cluster</value>
  36. </property>
  37. <property>
  38. <name>dfs.client.failover.proxy.provider.ha-cluster</name>
  39. <value>org.apache.hadoop.hdfs.server.namenode.ha.ConfiguredFailoverProxyProvider</value>
  40. </property>
  41. <property>
  42. <name>dfs.ha.automatic-failover.enabled</name>
  43. <value>true</value>
  44. </property>
  45. <property>
  46. <name>ha.zookeeper.quorum</name>
  47. <value>192.168.10.153:2181,192.168.10.155:2181,192.168.10.154:2181</value>
  48. </property>
  49. <property>
  50. <name>dfs.ha.fencing.methods</name>
  51. <value>sshfence</value>
  52. </property>
  53. <property>
  54. <name>dfs.ha.fencing.ssh.private-key-files</name>
  55. <value>/home/ci/.ssh/id_rsa</value></property></configuration>

停止nn1(活动节点)时的日志:(zkfc nn1,nn2)(namenode nn1,nn2)https://pastebin.com/bwvfnanq

xqkwcwgp

xqkwcwgp1#

你提到的 <IP>:<port> 为了 fs.defaultFS 在core-site.xml中。所以当关闭活动namenode时,它不知道重定向到哪里。
为名称服务选择逻辑名称,例如“mycluster”。
然后更改hdfs-site.xml, dfs.namenode.http-address.[nameservice ID].[name node ID] -要侦听的每个namenode的完全限定http地址
对你来说,你必须付出代价
core-site.xml文件

  1. <property>
  2. <name>fs.defaultFS</name>
  3. <value>hdfs://myCluster</value>
  4. </property>

hdfs-site.xml文件

  1. <property>
  2. <name>dfs.namenode.rpc-address.myCluster.nn1</name>
  3. <value>192.168.10.153:9000</value>
  4. </property>

把说明书读清楚https://hadoop.apache.org/docs/r2.7.2/hadoop-project-dist/hadoop-hdfs/hdfshighavailabilitywithqjm.html
希望这对你有帮助。

展开查看全部
kse8i1jr

kse8i1jr2#

您必须寻找自动故障转移
https://stackoverflow.com/a/27272565/3496666
https://hadoop.apache.org/docs/r2.7.2/hadoop-project-dist/hadoop-hdfs/hdfshighavailabilitywithqjm.html

相关问题