XetHub raises $7.5M for its Git-based data collaboration platform • TechCrunch

Seattle-based XetHub, a startup that makes it easy for enterprises to use Git for data management, announced today that it has raised $7.5 million in a seed round led by Madrona. The basic idea here is to allow developers to work with data the same way they work with code, including all the collaboration features that tools like Git allow. The team describes XetHub as a “collaborative storage platform for data management.”

The company was co-founded by Yucheng Low (CEO), Ajit Banerjee, and Rajat Arya. The team has years of experience working with large data platforms. In fact, Low previously co-founded ML startup Turi, where Arya was the first employee. Apple acquired the company in his 2016, allowing Low and Arya to work on different parts of his stack on Apple’s ML platform. For example, Arya leads the data platform team at Apple. It was also at Apple that the two met Banerjee, who previously worked at Inktomi, Amazon and Facebook. He also previously founded his two startups.

The XetHub repository view is designed for navigating and visualizing your data repositories while maintaining GitHub sensitivity. XetHub automatically summarizes common file formats (CSV) and supports custom visualizations.

While working on the data platform at Apple, the team realized that there was still a lot of room for improvement in the area of ​​data management.

“Not surprisingly, data is much more important than anything else. More than models, more than anything else,” Rowe told me. “Controlling where this data is stored and how this data is collaborated on is very basic. You’ll find it very similar to how you did it: version control and collaboration is done by copy and paste, sometimes there are more elaborate versions, but no one knows what you’re doing. If you want to avoid touching it, it’s still copy and paste in the end.”

Just as developers have moved to tools like Git to collaborate on source code, XetHub wants to be able to use these same familiar primitives to work with data.

“Our mindset is that for the first time, developers will be able to work with data in exactly the same way that they work with code,” says Low. He points out that the team aims to create tools that not only mimic a Git-like experience, but also preserve Git’s core user experience, including all the integrations developers are accustomed to. Did.

XetHub extends Git to support large files, maintaining full Git compatibility while providing efficient storage and transfer with data deduplication.

Currently, the service can handle repositories of up to 1 TB of data, with plans to scale this to 100 TB soon. Since few developers want to clone such large repositories, it is also possible for developers to mount these repositories and make them behave like local file systems. Also note that this tool is file format agnostic.

From a marketing perspective, the team is focused on the AI/ML team, but it’s clear that users can use XetHub to manage all kinds of data.

Xethub is now generally available with a free community edition that can be used to manage up to 20 GB of deduped storage. Low told me the company is already in talks with some enterprise customers, but the team isn’t ready to name them yet.

“Yucheng and the outstanding XetHub team have been innovating in machine learning for over a decade and then applying their skills at Apple, the most iconic consumer technology company. can collaborate with others to manipulate large datasets and build intelligent, generative applications.” “The development and deployment of these applications is constrained by legacy infrastructure and complex data workflows. XetHub addresses these pain points from a developer’s perspective.”

Source link

Leave a Reply

Your email address will not be published. Required fields are marked *