Registration is now open for a workshop on “Improving Data Mobility and Management for International Cosmology” to be held Feb. 10-11 at Lawrence Berkeley National Laboratory in California. The workshop, one in a series of Cross-Connects workshops, is sponsored the by the Dept. of Energy’s ESnet and Internet2.
Early registration is encouraged as attendance is limited and the past two workshops were filled and had waiting lists. Registration is $200 including breakfast, lunch and refreshments for both days. Visit the Cross-Connects Workshop website for more information.
Cosmology data sets are already reaching into the petabyte scale and this trend will only continue, if not accelerate. This data is produced from sources ranging from supercomputing centers—where large-scale cosmological modeling and simulations are performed—to telescopes that are producing data daily. The workshop is aimed at helping cosmologists and data managers who struggle with data workflow, especially as the need for real-time analysis of cosmic events increases.
Renowned cosmology experts Peter Nugent and Salman Habib will give keynote speeches at the workshop.
Nugent, Senior Scientist and Division Deputy for Science Engagement in the Computational Research Division at Lawrence Berkeley National Laboratory, will deliver a talk on “The Palomar Transient Factory” and how observational data in astrophysics, integrated with high-performance computing resources, benefits the discovery pipeline for science.
Habib, a member of the High Energy Physics and Mathematics and Computer Science Divisions at Argonne National Laboratory, a Senior Member of the Kavli Institute for Cosmological Physics at the University of Chicago, and a Senior Fellow in the Computation Institute, will give the second keynote on “Cosmological Simulations and the Data Big Crunch.”
As part of its look at things to expect in 2015, Popular Science magazine highlights ESnet’s new trans-Atlantic links which will have a combined capacity of 340 gigabits per second. The three 100 Gbps and one 40 Gbps connections are being tested and are expected to go live at the end of January.
ESnet, the Department of Energy’s (DOE’s) Energy Sciences Network, is deploying four new high-speed transatlantic links, giving researchers at America’s national laboratories and universities ultra-fast access to scientific data from the Large Hadron Collider (LHC) and other research sites in Europe.
ESnet’s transatlantic extension will deliver a total capacity of 340 gigabits-per-second (Gbps), and serve dozens of scientific collaborations. To maximize the resiliency of the new infrastructure, ESnet equipment in Europe will be interconnected by dedicated 100 Gbps links from the pan-European networking organization GÉANT.
Funded by the DOE’s Office of Science and managed by Lawrence Berkeley National Laboratory, ESnet provides advanced networking capabilities and tools to support U.S. national laboratories, experimental facilities and supercomputing centers.
Among the first to benefit from the network extension will be U.S. high energy physicists conducting research at the Large Hadron Collider (LHC), the world’s most powerful particle collider, located near Geneva, Switzerland. DOE’s Brookhaven National Laboratory and Fermi National Accelerator Laboratory—major U.S. computing centers for the LHC’s ATLAS and CMS experiments, respectively—will make use of the links as soon as they are tested and commissioned.
Data from the University of New Mexico’s Cancer Center next-generation genome sequencers is now flying across a 10 Gbps link to the university’s Center for Advanced Research Computing, thanks to the Science DMZ model pioneered by ESnet, according to an article posted by UNM.
According to the article, “This point-to-point connection is a first step toward establishing a campus-wide research network at UNM. The connection is based on the “Science DMZ” model formalized by the Department of Energy’s ESnet in 2010. The new link delivers a low-latency, high-bandwidth, unfiltered connection via UNM’s campus network.”
The article states that the new 10 Gbps link enables fast, reliable, and secure transfer of enormous genome sequence files from the UNM Cancer Center for analysis and subsequent data warehouse archiving. And the model may pave the way for greater research collaborations across the state.
“This project is part of UNM’s larger direction to collaborate across campuses and expand network infrastructure for research here and statewide,” said Chief Information Officer Gil Gonzales. UNM IT works closely with departments and Centers at UNM, and with research institutions throughout New Mexico, to provide production, commodity, and research network services.
We are pleased to announce three influential keynote speakers for the upcoming Focused Technical Workshop titled “Improving Data Mobility and Management for International Climate Science”, which will be hosted by the National Oceanic and Atmospheric Administration (NOAA) in Boulder, CO from July 14-16, 2014.
The first keynote will be delivered by NOAA’s Dr. Alexander “Sandy” MacDonald, Chief Science Advisor and Director of the Earth System Research Laboratory (ESRL), who is known for his influential work in weather forecasting and high performance computing at NOAA.
Also from NOAA’s Geophysical Fluid Dynamics Laboratory (GFDL) and Princeton University, Dr. V. Balaji, Head of the Modeling Systems Group, will share the importance of bridging the worlds between science and software with workshop attendees.
And finally, Eli Dart, a highly-acclaimed network engineer from the Department of Energy’s ESnet who is credited with co-developing the Science DMZ model, will wrap up the workshop with the final keynote focused on how to create cohesive strategies for data mobility across computer systems, networks and science environments.
Inspired by each of the keynote speaker’s integral roles in climate science, computing and network architectures, the workshop intends to spark lively, interactive discussions between the research and education (R&E) and climate science communities to build long-term relationships and create useful tools and resources for improved climate data transport and management.
Starting this January, the Earth System Grid Federation (ESGF) has started a new working group—the International Climate Network Working Group—to help set up and optimize network infrastructures for their climate data sites around the world. They need network connections that can deal with petabytes of modeling and observational data, which will traverse more than 13,000 miles of networks (more than half the circumference of the Earth!), spanning two oceans.
By the end of 2014, this working group will aim to obtain at least 4Gbps of data transfer throughput at five of their climate data centers at PCMDI/LLNL (US), NCI/ANU (AU), CEDA/SFTC (UK), DKRZ (DE), and KNMI (NE). This goal runs in parallel with the Enlighten Your Research Global international networking program award that ESGF received this last November 2013. This initiative is lead by Dean Williams of Lawrence Livermore National Lab and ESnet’s Science Engagement Team, along with collaborating international network organizations in Australia (AARnet), Germany (DFN), the Netherlands (SURFnet), and the UK (Janet). We are helping to shepherd ESGF’s project and working group to make sure all their climate sites get up and running at proficient network speeds for the future peta-scale climate data that is expected within the next 5 years.
As we work closely with ESGF to pave the way for climate science, we look forward to developing a new set of networking best practices to help future climate science collaborations. In all, we are excited to get this started and see their science move forward!
The Enlighten Your Research Global program will set up, optimize and/or troubleshoot 5 ESGF locations in different countries throughout 2014.
As a research and education network, one of ESnet’s accomplishments came to light at the end of 2013 during an SC13 demo in Denver, CO. Using ESnet’s 100 Gbps backbone network, NASA Goddard’s High End Computer Networking (HECN) Team achieved a record single host pair network data transfer rate of over 91 Gbps for a disk-to-disk file transfer. By close collaboration with ESnet, Mid-Atlantic Crossroads (MAX), Brocade, Northwestern University’s International Center for Advanced Internet Research (iCAIR), and the University of Chicago’s Laboratory for Advanced Computing (LAC), the HECN Team showcased the ability to support next generation data-intensive petascale science, focusing on achieving end-to-end 100 Gbps data flows (both disk-to-disk and memory-to-memory) across real-world cross-country 100 Gbps wide-area networks (WANs).
To achieve 91+ Gbps disk-to-disk network data transfer rate between a single pair of high performance RAID servers, this demo required a number of techniques working in concert to avoid any bottlenecks in the end-to-end transfer process. This required parallelization using multiple CPU cores, RAID controllers, 40G NICs, and network data streams; a buffered pipelined approach to each data stream, with sufficient buffering at each point in the pipeline to prevent data stalls, including application, disk I/O, network socket, NIC, and network switch buffering; a completely clean end-to-end 100G network path (provided by ESnet and MAX) to prevent TCP retransmissions; synchronization of CPU affinities for the application process and the disk and network NIC interrupts; and a suitable Linux kernel.
The success of the HECN Team SC13 demo proves that it is possible to effectively fully utilize real-world 100G networks to transfer and share large-scale datasets in support of petascale science, using Commercial Off-The-Shelf system, RAID, and network components, together with open source software.
View of the MyESnet Portal during the SC13 demo. Top visual shows a network topology graph with colors denoting the scale of traffic going over each part of the network. Bottom graph shows the total rate of network traffic vs. time across the backbone.Diagram of the SC13 demo connections.
The perfSONAR-PS project celebrated a milestone by surpassing 1000 deployed software instances in December of 2013. The perfSONAR software is designed to assist network operators and end users with the task of monitoring end-to-end performance, and assisting with debugging tasks in the event that problems arise. As the number of deployments grows, the software becomes more effective by offering better coverage across more network paths, including ESnet and the connectivity to Department of Energy and NSF funded resources. perfSONAR-PS is a joint effort between ESnet, Fermilab, Georgia Tech, Indiana University, Internet2, SLAC, and the University of Delaware.
Hurricane Sandy took a terrible toll this week, and our thoughts remain with the millions of people whose lives were impacted.
From the nightly news, we learned about critical systems that temporarily failed during the extreme weather, and others that continued to function. I’d like to share an example in the second category.
Here’s a screenshot from our public web portal, http://my.es.net, showing traffic to and from Brookhaven National Laboratory on Long Island, during the worst of the storm:
Network traffic to and from Long Island’s Brookhaven National Laboratory was not impacted by the devastating storm that struck the New York area this week.
Why does it matter that Brookhaven’s connection to ESnet, and to the broader Internet, remained functional during the disaster? It matters because modern science has entered an age of extreme-scale data. In more and more fields, scientific discovery depends on data mobility, as huge data sets from experiments or simulations move to computational and analysis facilities around the globe. Large-scale science instruments are now designed around the premise that high-speed research networks exist. In this kind of architecture, loss of network connectivity impairs scientific productivity. In the worst case, unique and vital data can be lost.
Much of the network traffic in these graphs can be attributed to ATLAS, one of the extraordinary experiments at CERN’s Large Hadron Collider, outside of Geneva. It’s been a great year at CERN, with announcement of a Higgs-like particle this summer. At ESnet, we were very happy that productivity of the ATLAS experiment was not affected by the massive storm.
Although the research networking community may have benefited from good fortune during storm, it’s important to recognize that a lot of careful planning and sound engineering – conducted over the course of many years – contributed to this outcome. Research networks like ESnet and Internet2, regional networks like NYSERNet, campus networking groups at Brookhaven and elsewhere, and exchange points like MAN LAN all work very hard to harden themselves against individual points of failure. As with all engineering, the devil is in the details: fiber diversity, elimination of shared fate in single conduits and data centers, location and specification of generators, availability of fuel, optical protection schemes, failover for dynamic circuits… the list goes on and on.
We won’t always be so fortunate, but the screenshot above is something our research networking community can take pride in. My thanks to the hundreds of people whose sound decisions made this good outcome possible.
ESnet and its collaborators successfully completed three days of demonstrating its End-to-End Circuit Service at Layer 2 (ECSEL) software at the Open Networking Summit held at Stanford a couple of weeks ago. Our goal is to build “zero-configuration circuits” to help science applications seamlessly use networks for optimized end-to-end data transport. ECSEL, developed in collaboration with NEC, Indiana University, and the University of Delaware builds on some exciting new conceptual thinking in networking.
Wrangling Big Data
To put ECSEL in context, the proliferating tide of scientific data flows – anticipated at 2 petabytes per second as planned large-scale experiments get in motion – is already challenging networks to be exponentially more efficient. Wide area networks have vastly increased bandwidth and enable flexible, distributed, scientific workflows that involve connecting multiple scientific labs to a supercomputing site, a university campus, or even a cloud data center.
Heavy network traffic to come
The increasing adoption of distributed, service-oriented computing means that resource and vendor independence for service delivery is a key priority for users. Users expect seamless end-to-end performance and want the ability to send data flows on demand, no matter how many domains and service providers are involved. The hitch is that even though the Wide Area Network (WAN) can have turbocharged bandwidth, at these exponentially increasing rates of network traffic even a small blockage in the network can seriously impair the flow of data, trapping users in a situation resembling commute conditions on sluggish California freeways. These scientific data transport challenges that we and other R&E networks face are just a taste of what the commercial world will encounter with the increasing popularity of cloud computing and service-driven cloud storage.
Abstracting a solution
One of the key feedback from application developers, scientists and end-users is that they do not want to deal with the complexity at the infrastructure level while still accomplishing their mission. At ESnet, we are exploring various ways to make networks work better for users. A couple of concepts could be game-changers, according to Open Network Summit conference presenter and Berkeley professor Scott Shenker: 1) using abstraction to manage network complexity, and 2) extracting and exposing simplicity out of the network. Shenker himself cites Barbara Liskov’s Turing Lecture as inspiration.
ECSEL is leveraging OSCARS and OpenFlow within the Software Defined Networking (SDN) paradigm to elegantly prevent end-to-end network traffic jams. OpenFlow is an open standard to allow application-driven manipulation of network flows. ECSEL is using OSCARS-controlled MPLS virtual circuits with OpenFlow to dynamically stitch together a seamless data plane delivering services over multi-domain constructs. ECSEL also provides an additional level of simplicity to the application, as it can discover host-network interconnection points as necessary, removing the requirement of applications being “statically configured” with their network end-point connections. It also enables stitching of the paths end-to-end, while allowing each administrative entity to set and enforce its own policies. ECSEL can be easily enhanced to enable users to verify end-to-end performance, and dynamically select application-specific protocol forwarding rules in each domain.
The OpenFlow capabilities, whether it be in an enterprise/campus or within the data center, were demonstrated with the help of NEC’s ProgrammableFlow Switch (PFS) and ProgrammableFlow Controller (PFC). We leveraged a special interface developed by them to program a virtual path from ingress to egress of the OpenFlow domain. ECSEL accessed this special interface programmatically when executing the end-to-end path stitching workflow.
Our anticipated next step is to develop ECSEL as an end-to-end service by making it an integral part of a scientific workflow. The ECSEL software will essentially act as an abstraction layer, where the host (or virtual machine) doesn’t need to know how it is connected to the network–the software layer does all the work for it, mapping out the optimum topologies to direct data flow and make the magic happen. To implement this, ECSEL is leveraging the modular architecture and code of the new release of OSCARS 0.6. Developing this demonstration yielded sufficient proof that well-architected and modular software with simple APIs, like OSCARS 0.6, can speed up the development of new network services, which in turn validates the value-proposition of SDN. But we are not the only ones who think that ECSEL virtual circuits show promise as a platform for spurring further innovation. Vendors such as Brocade and Juniper, as well as other network providers attending the demo were enthusiastic about the potential of ECSEL.
But we are just getting started. We will reprise the ECSEL demo at SC11 in Seattle, this time with a GridFTP application using Remote Direct Memory Access (RDMA) which has been modified to include the XSP (eXtensible Session Protocol) that acts as a signaling mechanism enabling the application to become “network aware.” XSP, conceived and developed by Martin Swany and Ezra Kissel of Indiana University and University of Delaware, can directly interact with advanced network services like OSCARS – making the creation of virtual circuits transparent to the end user. In addition, once the application is network aware, it can then make more efficient use of scalable transport mechanisms like RDMA for very large data transfers over high capacity connections.
We look forward to seeing you there and exchanging ideas. Until Seattle, any questions or proposals on working together on this or other solutions to the “Big Data Problem,” don’t hesitate to contact me.
–Inder Monga
imonga@es.net
ECSEL Collaborators:
Eric Pouyoul, Vertika Singh (summer intern), Brian Tierney: ESnet
You must be logged in to post a comment.