tcp - Why isn't Hadoop implemented using MPI?

Question

Welcome To Ask or Share your Answers For Others

tcp - Why isn't Hadoop implemented using MPI?

asked Oct 6, 2021 in Technique[技术] by 深蓝 (71.8m points)

Correct me if I'm wrong, but my understanding is that Hadoop does not use MPI for communication between different nodes.

What are the technical reasons for this?

I could hazard a few guesses, but I do not know enough of how MPI is implemented "under the hood" to know whether or not I'm right.

Come to think of it, I'm not entirely familiar with Hadoop's internals either. I understand the framework at a conceptual level (map/combine/shuffle/reduce and how that works at a high level) but I don't know the nitty gritty implementation details. I've always assumed Hadoop was transmitting serialized data structures (perhaps GPBs) over a TCP connection, eg during the shuffle phase. Let me know if that's not true.

question from:https://stackoverflow.com/questions/4590674/why-isnt-hadoop-implemented-using-mpi

与恶龙缠斗过久,自身亦成为恶龙；凝视深渊过久,深渊将回以凝视…

517 views

1 Answer

深蓝 · Answer 1 · 2021-10-06T05:06:01+0000

One of the big features of Hadoop/map-reduce is the fault tolerance. Fault tolerance is not supported in most (any?) current MPI implementations. It is being thought about for future versions of OpenMPI.

Sandia labs has a version of map-reduce which uses MPI, but it lacks fault tolerance.

Categories

tcp - Why isn't Hadoop implemented using MPI?

Please log in or register to add a comment.

Please log in or register to answer this question.

1 Answer

Please log in or register to add a comment.

Just Browsing Browsing

Most popular tags