Abstract

Apache Hadoop Ozone provides storage for Hadoop-based and Cloud-native environments.

It provides Object Storage semantics (like S3) and scales up to billions of objects.

It provides S3, Hadoop Compatible File System, and CSI interfaces.

The Hadoop community believes that further development on Ozone can be done better as a separate project as discussed on the Hadoop email lists

History

The first beta release of Ozone was just released and it was agreed that the next release will be announced as GA release. It seems to be a good time before the first GA to make a decision about the future of the project

(Note: if you need more information about the history of the project, you can check the detailed version from the source repository)

Move to a separated, Apache TLP

 During the last years, Ozone became more and more independent both at the community and the code side. The separation has been suggested again and again (for example by Owen [2] and Vinod [3])

 From COMMUNITY point of view:

 From CODE point of view Ozone became more and more independent:

Current Status of the project

Ozone is already part of a Top Level Apache Project and already created multiple separated releases that are approved by the Hadoop PMC and operating according to the Apache guidelines without any reported issues. As such a project it's already passed the Apache project maturity model. Therefore, in this section, we ignore some obvious statements which are a usual part of the project adoptions (Ozone source code is already part of the Apache, and it's already governed by Apache PMC) and focus on what project did so far for building a stronger community.

Building a community is a continuous effort.  We are at the beginning of a journey and moving Ozone to a separated TLP is a very important step. Some of the current challenges:

PMC/Committers

Ozone as a new project requires an initial PMC and committer list. But first, we need to define how the lists are created/selected:

  1. PMC: As Ozone is a Hadoop subproject today, all the existing Hadoop PMCs with noticeable Ozone contributions are added to the initial list. (Definition of noticeable contribution: all the related GitHub / Jira content is downloaded, and we selected all the Hadoop PMCs with at least 30 comments AND/OR commits since the beginning of 2019.
  2. A discussion is started with these people about
    1. what are the important factors of being a committer / PMC.
    2. Who should be added to the initial list?
  3. Committer: similar to submarine Hadoop committers can get opt-in committer membership to the Ozone project (except PMC veto)

Some points which are named as an important factor of being PMC:

  1. Involvement in releases (being RM, validating and voting on releases, roadmap for future releases)
  2. Being involved constructively in design discussions, keeping the big picture, and project direction in mind.
  3. Investing in build/CI quality. Ensuring that contributors and committers have a solid infrastructure to develop the project.
  4. Responsiveness on security, trademark, copyright issues.
  5. Positive involvement in the community (mailing lists, raising committer candidates).
  6. Keeping an eye on what needs to go better in the project (documentation, test quality, wiki pages). A meta-view beyond regular contributions and releases.

It's also found especially important to include the user community to the project governance. End-users and adopters – who are actively helping with the projects with feedback during the design discussions – should be invited to the PMC (even without code contribution). (During the discussion they are called as "user-seats" in PMC)

The initial selection rules and PMC list is shared on the ozone-dev mailing list (people who are nominated in 2b are added explanatio) where additional methods are suggested (add everybody to the PMC who are Hadoop PMC and contributed at least 10 patches in this year) and accepted.

Proposed Chair:

  1. Sammi Chen [sammichen] (Hadoop PMC)

Proposed PMC (Hadoop PMC)

  1. Arpit Agarwal [arp] (member, Hadoop PMC)
  2. Shashikant Banerjee [shashikant] (Hadoop PMC)
  3. Li Cheng [licheng] (Hadoop committer)
  4. Dinesh Chitlangia [dineshc] (Hadoop committer)
  5. Clay Baenziger
  6. Attila Doroszlai [adoroszlai] (Hadoop committer)
  7. Junping Du [junping_du] (member, Hadoop PMC)
  8. Márton Elek [elek] (Hadoop PMC)
  9. Anu Engineer [aengineer] (Hadoop PMC)
  10. Uma Maheswara Rao G [umamahesh] (member, Hadoop PMC)
  11. Lokesh Jain [ljain] (Hadoop PMC)
  12. Hanisha Koneru [hanishakoneru] (Hadoop PMC)
  13. Yiqun Lin [yqlin] (Hadoop PMC)
  14. Siyao Meng [siyao] (Hadoop committer)
  15. Jitendra Nath Pandey [jitendra] (member, Hadoop PMC)
  16. Rakesh Radhakrishnan [rakeshr] (member, Hadoop PMC)
  17. Matt Sharp
  18. Mukul Kumar Singh [msingh] (Hadoop PMC)
  19. Tsz-wo Sze [szetszwo] (member, Hadoop PMC)
  20. Xiaoyu Yao [xyao] (Hadoop PMC)
  21. Nandakumar Vadivelu [nanda] (Hadoop PMC)
  22. Bharat Viswanadham [bharat] (Hadoop PMC)
  23. Siddharth Wagle [swagle] (Hadoop committer)
  24. Stephen O'Donnell [sodonnell] (Hadoop committer)
  25. Vivek Ratnavel Subramanian [vivekratnavel] (Hadoop committer)
  26. Aravindan Vijayan [avijayan] (Hadoop committer)

Proposed committer list

  1. Wei-Chiu Chuang [weichiu] (Hadoop committer)
  2. István Fajth
  3. Nilotpal Nandi [nilotpalnandi] (Hadoop committer)
  4. Yisheng Lien [yisheng] (Hadoop committer)
  5. Baoloong Mao (github.com/maobaolong)
  6. Neo Yang (github.com/cku328)
  7. WeiWei Yang (wwei) (Hadoop committer)
  8.  Jie Wang [runzhiwang]
  9. Xiang Zhang (github.com/iamabug)
  10. Micah Zhao (github.com/captainzmc)
  11. Masatake Iwasaki [iwasakims] (Hadoop committer)
  12. Prabhu Joseph [prabhujoseph] (Hadoop committer)
  13. Ayush Saxena [ayushsaxena] (Hadoop PMC)
  14. He Xiaoqiao [hexiaoqiao] (Hadoop committer)
  15. Surendra Singh Lilhore [surendralilhore] (Hadoop PMC)
  16. Vinayakumar B [vinayakumarb] (Hadoop PMC)
  17. Bibin A Chundatt [bibinchundatt] (Hadoop PMC) 

Required Resources