# Interline Technologies > We're a product and consulting business that helps organizations understand and improve transportation networks, digitally. Public Ghost content for AI and LLM tooling. This file includes a bounded export of public pages first, then recent public posts. Append `.md` to any post or page URL to get the content in Markdown (for example, `/example-post.md`). ## Pages ### About Our Team and Firm URL: https://www.interline.io/about/ Last updated: 2024-09-03T22:36:13.000Z We founded Interline Technologies in 2018 in the San Francisco Bay Area, and we continue to serve our customers from here, along with a team that's distributed across the United States. ## Principals #### Drew Dara-Abrams ****Drew Dara-Abrams Ph.D.** is a co-founder and Principal at Interline Technologies. At Interline, Drew manages the firm's consulting engagements and product development. Drew's previous experience spans both industry and academic labs in transportation and geography. His experience in industry includes serving as Head of Mobility Products for the Mapzen division of Samsung. While at Mapzen, he recruited and managed a team to build the Transitland open transit data platform, develop a worldwide routing/trip-planning engine, and build a global system to derive road traffic speeds from GPS data for The World Bank. Previously, Drew served as Chief Technology Officer for Kinnexxus, Inc. where he and his colleagues completed two successful Small Business Technology Transfer (SBIR) grants from the National Institutes of Health. He's developed and managed web applications serving upwards of millions of users, and has provided strategic, technical, and statistical consulting to major corporations, universities, and start-ups. Drew holds a Ph.D. in computational geography from the University of California, Santa Barbara, and co-authored two textbooks, **Supporting Web Servers* and **E-Commerce & Internet Law*, published by Prentice-Hall. #### Ian Rees ****Ian Rees Ph.D.** is a co-founder and Principal at Interline Technologies. At Interline, Ian leads engineering and operations efforts, building and maintaining open data pipelines to power publicly accessible trip planners, accessibility analysis, and data validation for Interline’s clients. Prior to Interline, Dr. Rees was a software engineer at Samsung’s Mapzen division and technical lead of the Transitland project, an effort to create an open, user-contributed, user-editable, fully validated repository of all public GTFS data in the world. Previous to Samsung, Ian held postdoctoral positions at Lawrence Berkeley National Laboratory and Baylor College of Medicine, where he administered a data center and developed an online platform for biological imagery analysis. Ian holds a B.S. in Biochemistry from the University of Houston, and a Ph.D. in Structural and Computational Biology from Baylor College of Medicine. Ian is proficient in several programming languages including Python, Ruby, and JavaScript, and is experienced building complex applications using PostgreSQL, Redis, Docker, Kubernetes, and all three of the major public clouds (Amazon Web Services, Google Cloud Platform, and Microsoft Azure). ## Associates and Partners Using both our physical location in Silicon Valley and relationships across the North American transportation sector, Interline has built a network of specialized associates and partners to serve our clients. Interline’s associates enable our firm to deliver large results to SaaS customers and our consulting clients. Interline’s partners organizations join our team to support critical projects: #### Garnet Consulting [****Garnet Consulting**](https://www.garnetconsultingpdx.com/?ref=interline.io) (Portland, OR) specializes in collaborative and inclusive transit technology, operations, and software consulting for rural, small urban, and intercity agencies and the organizations that serve them. #### DCR Design [****DCR Design**](https://dcrdesign.net/?ref=interline.io) (Southern California) is a cartographic services and information design firm with deep experience in public transit, active transportation, and geospatial analysis. #### Lab Zero Innovations [****Lab Zero Innovations**](https://www.labzero.com/?ref=interline.io) (San Francisco) combines user-centered design skills, agile management methods, and full-stack engineering capabilities. #### Trillium Solutions [****Trillium Solutions**](https://trilliumtransit.com/?ref=interline.io) (Portland, OR) specializes in the creation of transit data and creates GTFS feeds for hundreds of transit agencies around the United States. Interline enlists Trillium for engagements that involve custom GTFS creation and data analysis. ## Small Business Certifications Interline Technologies LLC is certified as a small business by the following organizations: - California Department of General Services - Los Angeles County Metropolitan Transportation Authority (LA Metro) ## Associations Interline collaborates with fellow companies, researchers, and institutions through participation in the following associations: - [MobilityData](https://mobilitydata.org/?ref=interline.io) as a dues-paying member company - [Urban Computing Foundation](https://uc.foundation/?ref=interline.io) with Drew Dara-Abrams serving on the foundation’s Technical Advisory Council - [OpenTripPlanner](https://docs.opentripplanner.org/en/latest/Governance/?ref=interline.io) with Drew Dara-Abrams serving on the Project Leadership Committee - [Transportation Research Board](http://www.trb.org/?ref=interline.io) with Drew Dara-Abrams serving as a member on the standing committee on transit data (AP090) and the Transit IDEA review panel ## Contact #### Learning how to use Transitland? If you're a developer, we recommend you start by signing up for a Transitland Free account and review the documentation. If you have questions, please post them to the [Transitland discussion board on GitHub](https://github.com/transitland/transitland/discussions?ref=interline.io). #### Need a quote for Transitland Enterprise? If your needs can't be met by the Transitland Free or Transitland Professional plans, please [contact us for a quote for Transitland Enterprise](https://opcrm.page.link/q5DT?ref=interline.io). #### Have another question? You're welcome to email our team at [info@interline.io](mailto:info@interline.io) 📬 To follow Interline, find us on [LinkedIn](https://www.linkedin.com/company/interline-io/?ref=interline.io), our [newsletter](http://eepurl.com/dmWHln?ref=interline.io), and our [blog](https://www.interline.io/blog). ### Intro to Transitland URL: https://www.interline.io/transitland/ Last updated: 2026-01-04T23:34:38.000Z Transitland is an open-data platform built on thousands of public-transit data feeds from around the world. We started Transitland in 2014 and continue to expand the platform. It’s now the largest and most feature-rich GTFS and GTFS Realtime aggregator. Transitland is the ideal solution for using data from many transit operators in web or mobile apps, maps, data visualizations, GIS analyses, and travel demand models. As a platform, Transitland both serves *data consumers* (such as app developers) and *data producers* (such as transit agencies and private mobility providers). ## Transitland for Data Consumers Are you developing an app, a visualization, a map, a plan, an analysis, or another type of digital experience that needs transit data as an input? Learn more about how you can use Transitland to consume transit data: #### APIs for Developers Transitland provides a range of APIs to flexibly query across its world-wide store of transit data. Results can be produced in a wide variety of formats, including JSON, GeoJSON, CSV, vector map tiles, and static images. [Learn more about Transitland APIs for developers](https://www.interline.io/transitland/apis-for-developers) #### Bulk Transit Data Interline produces bulk data exports from the Transitland platform that are customized for specific needs. Our clients use these bulk exports to power their own transportation analysis engines, real-estate assessments, routing engines, and other use-cases. [Learn more about custom bulk data exports](https://www.interline.io/transitland/custom-bulk-transit-data/) ## Transitland for Data Producers Learn more about how you can use Transitland to publish and enrich transit data: #### Agencies publishing GTFS Share your GTFS and GTFS Realtime feeds with Transitland and we’ll help disseminate your open data to an even wider audience. Agency staff may [use GitHub to add GTFS feed URLs to the Transitland Atlas](https://github.com/transitland/transitland-atlas?ref=interline.io#how-to-add-a-new-feed), or email our data team for assistance. [Email your feed URL or file to the Transitland data team](mailto:info@interline.io) #### Enriching GTFS data Our tooling is used by both Interline staff and clients to enrich GTFS feeds with GTFS Pathways data for stations and GTFS Fares-v2 data for fares and transfer discounts. [Learn more about Interline Station Editor](https://www.interline.io/saas-tools/station-editor/) ## Questions & Answers #### What is GTFS? GTFS stands for General Transit Feed Specification. It’s a set of machine-readable files in a standard format that can be produced by a transit operator and disseminated across the internet. The contents describe that operator’s stations and stop locations, route lines and shapes, schedules, fares, and related information. Trip-planning apps can ingest GTFS feeds from one or more operator to plan journeys for their users. Analysts also use GTFS feed to understand travel patterns and propose changes. #### How can we add our feeds to Transitland? Read the instructions to add a feed URL to the [Transitland Atlas](https://github.com/transitland/transitland-atlas?ref=interline.io) or [email us for assistance](mailto:info@interline.io). #### Is Transitland open-source software? Transitland is built atop open-source tools and its own components are also available as open-source code. See [github.com/transitland](https://github.com/transitland?ref=interline.io) for our public repositories. Interline offers the core Transitland libraries under a dual licensing model: open-source for use by all under the GPLv3 license and also available under a flexible commercial license from Interline. [Contact us](https://www.interline.io/about/#contact) for more information about flexible commercial licensing. ### Home URL: https://www.interline.io/home/ Last updated: 2026-04-27T18:47:35.000Z ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2026/04/_DSF0189.jpg) ## Hello We help organizations understand and improve transportation networks, digitally. Our systems support thousands of public transit agencies, intercity bus and rail operators, and shuttle systems across the world. Our datasets power applications across the planning, real-estate, healthcare, and insurance sectors, as well as academic research. # Build with transit data ### Transitland API Power transit apps, maps, and dashboards with live schedules, real-time arrivals, and operator data from thousands of agencies worldwide. [Explore the API →](https://www.interline.io/transitland/apis-for-developers/) [View pricing →](https://www.transit.land/plans-pricing/?ref=interline.io) ### Transitland Datasets Enrich your analysis with transit intelligence — bulk stop, route, and service data for GIS teams, planners, and researchers working at regional or national scale. [Browse Datasets →](https://www.transit.land/datasets/?ref=interline.io) ### Routing Platform Calculate multi-modal routes and travel times at scale — from embedding journey planning in an app to running regional accessibility studies. [Explore the Routing Platform →](https://www.interline.io/routing-platform/) ### Transitland Feed Archive Track how transit service has evolved over time with years of historical GTFS snapshots across thousands of feeds worldwide. [Learn more →](https://www.transit.land/documentation/datasets/?ref=interline.io) [View pricing →](https://www.transit.land/plans-pricing/?ref=interline.io) # Enrich and improve transit ### Station Editor Map station layouts, entrances, and accessible pathways so trip planners can route every rider through your facilities. [Learn more →](https://www.interline.io/saas-tools/station-editor/) ### Transfer Analyst Identify missed connections at multi-agency hubs and coordinate schedules to improve transfer reliability for riders. [Learn more →](https://www.interline.io/saas-tools/transfer-analyst/) ## Recently on Interline's blog ### Intro to Valhalla URL: https://www.interline.io/valhalla/ Last updated: 2024-09-04T00:05:10.000Z Valhalla is a multi-purpose routing engine originally created by Mapzen's mobility team, which was led by and included the founders of Interline. Now, Valhalla is also used by organizations like [Mapillary](https://www.mapillary.com/?ref=interline.io) (street imagery), [Mapbox](https://www.mapbox.com/?ref=interline.io) (mapping APIs), [Tesla](https://www.tesla.com/?ref=interline.io) (electric cars), and [Sidewalk Labs](https://sidewalklabs.com/?ref=interline.io) (urban real-estate development and operations). Valhalla is open-source (MIT license) and is maintained by [many contributors](https://github.com/valhalla/valhalla/graphs/contributors?ref=interline.io). Interline continues its involvement in open-source Valhalla, providing a variety of products and services to help organizations and individuals use Valhalla. ## What functions can Valhalla perform? ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/09/turn-by-turn-routes-1.png) #### Turn-by-turn directions Plan trips from Point A to Point B — and any number of other destinations. ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/09/time-distance-matrixes.png) #### Time/distance matrixes Rapidly calculate travel distances and times between many locations. ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/09/optimized-routes.png) #### Optimized routes Plan the shortest route to visit many locations (a.k.a. traveling salesman). ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/09/access-isochrones.png) #### Access isochrones Compute the boundaries of how far one can travel in 15 minutes, 30 minutes, and so on from an origin location. ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/09/location-to-map-matches.png) #### Location-to-map snapping Snap GPS probe data (or location data from another "noisy" source) with roadways and other elements of a transportation network. ## What makes Valhalla unique? **🌎 OpenStreetMap:** Valhalla is designed to work with the rich data model of [OpenStreetMap](https://www.openstreetmap.org/?ref=interline.io). It normalizes a wide range of tags into a consistent *roadway hierarchy*. It applies *country-specific* parameters to handle different laws (e.g., drive on right vs. on left) and default speed limits. **📜 Narrative guidance:** When generating turn-by-turn routes, Valhalla generates rich narrative guidance that's available in multiple languages. Without duplicative manuevers (no unnecessary instructions to "continue"). Ready for output as text-to-speech (TTS) on a smartphone. **🚌 Multimodal:** Valhalla supports many travel modes: auto, high-occupany auto, bicycle, motor scooter, and pedestrian. Valhalla combines perfectly with the [Transitland Routing API](https://www.transit.land/documentation/routing-api/?ref=interline.io) to also plan journeys by bus, train, and subway using the latest schedules from public transit agencies across the entire United States. **⚙️ Dynamic costing:** All of Valhalla's travel modes can be customized on the fly by changing costing parameters in each query. **Tiled data structure:** Valhalla generates and consumes its routing graph as tiles. This format allows an instance of Valhalla to scale to cover the entire planet, or to be specific to one metropolitan region. **📱 Embeddable:** Valhalla is primarily written in C++. It can be used for embedded applications in auto infotainment systems, on iOS, and on Android. ## How can I run Valhalla? #### Docker Valhalla is distributed as a Docker container. Easy to run on your Windows or Mac computer for testing. Deploy to GCP Kubernetes Engine, AWS Elastic Container Service, or another container scheduler. [Interline's Dockerfile for Valhalla](https://github.com/interline-io/valhalla-docker?ref=interline.io) Valhalla can also be installed on Linux using a [PPA package](https://github.com/valhalla/homebrew-valhalla?ref=interline.io). ## How can Interline help me to use Valhalla? We recommend **Valhalla Tilepacks** if you: - want to control your own Valhalla server(s) - know how to deploy Docker containers or PPA packages - don't want to run an entire Valhalla tile build pipeline - want to get started quickly [Download Valhalla Tilepacks](https://www.interline.io/valhalla/tilepacks/) We recommend **Interline's hosted global Valhalla APIs**, offered through RapidAPI.com, if you: - don't know how to manage servers - want to test with free API access (pay only for higher rates of API usage) - want to get started *very* quickly [Sign up through RapidAPI](https://rapidapi.com/interline-technologies-interline-technologies-default/api/interline-global-valhalla-navigation-and-routing-engine/?ref=interline.io) ### Transitland APIs for Developers URL: https://www.interline.io/transitland-apis-for-developers/ Last updated: 2024-09-03T23:24:41.000Z Transitland is an open-data platform built on thousands of public-transit data feeds from around the world. We started Transitland in 2014 and continue to expand the platform. It’s now the largest and most feature-rich GTFS and GTFS Realtime aggregator. Transitland is the ideal solution for using data from many transit operators in web or mobile apps, maps, data visualizations, GIS analyses, and travel demand models. 💡 Transitland APIs are ideal for powering clients that query for data as needed. If you are analyzing transit data at large scales, see also Interline’s [custom bulk transit data services](https://www.interline.io/transitland/custom-bulk-transit-data/). ## Features ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/09/transitland-operator-list.png) #### 2,000+ source feeds Transitland aggregates from thousand of GTFS and GTFS Realtime feeds around the world, fetching new feed versions, validating, and importing on a daily basis. All are welcome to add to the Transitland Atlas feed registry. [See the current list](https://www.transit.land/operators?ref=interline.io) ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/09/transitland-global-map.png) #### Global transit map Transitland includes buses, trains, subways, and ferries across over 50 countries. Pan and zoom across Transitland’s global transit map to view routes and stops, or use the v2 Vector Tile API to build your own. [Explore the Transitland global transit map](https://www.transit.land/map?ref=interline.io) ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/09/transitland-what-is-a-onestop-id.png) #### Unique Onestop IDs When using feeds from multiple operators, IDs often clash. Unexpected conflicts can also occur when using multiple versions over time from the same operator. Transitland solves this problem with the Onestop ID. Onestop IDs are globally unique. Each Onestop ID provides just enough information to look up a feed, an operator, a stop/station, or a route. Onestop IDs are used across the Transitland Atlas, the Transitland v1 and v2 APIs, and the Transitland website. Many developers and analysts also use Onestop IDs within their own apps and datasets to associate their records with Transitland. [Read Onestop ID documentation](https://www.transit.land/documentation/onestop-id-scheme/?ref=interline.io) ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/09/transitland-rest-api-docs.png) #### Powerful REST API The Transitland v2 platform features a new REST API that is fast, flexible, and reliable. Search across Transitland’s contents using a wide range of query options. Quickly page through results. The v2 REST API outputs JSON, GeoJSON, and static PNG images (to generate maps of operator service areas, routes, and stops. [Read Transitland REST API documentation](https://www.transit.land/documentation/rest-api/?ref=interline.io) ## Questions & Answers #### We have our own software for ingesting GTFS feeds. What advantages does Transitland provide? Most transit operators create their own GTFS feeds with little awareness of how their feeds may — or may not — work with other operators’ feeds. Transitland is designed specifically to support thousands of transit operators providing overlapping and interconnected service. Onestop IDs are globally unique and prevent data from different operators from clashing, even if operators use similar internal IDs. Transitland’s import workflows, database schemas, and APIs are all tuned to handle many large feeds in parallel. Interline staff and partners maintain the Transitland Atlas feed registry. If you are only working with one or two GTFS feeds, your own software may be simplest — if you aim to work with more, the full package of the Transitland platform is both easier and more powerful. #### Can we run our own copy of Transitland? Yes. We re-designed [Transitland v2](https://www.transit.land/news/2019/10/17/tlv2?ref=interline.io) from the ground up to be modular. This flexibility allows Transitland to be customized and re-deployed. Interline runs custom versions of the Transitland platform for clients as part of their GTFS production infrastructure, as well as for clients who use their Transitland deployment to aggregate and consume GTFS. [Contact us](https://www.interline.io/about/#contact) to learn about how we can customize and re-deploy Transitland for your organization’s needs. #### Can you answer my technical question about the Transitland APIs, software components, algorithms, data, or a related topic? First, please see the [Transitland documentation](https://www.transit.land/documentation?ref=interline.io). If you are on a paid plan, [sign into your Interline account](https://app.interline.io/users/sign%5Fin?ref=interline.io) to file a support ticket. If you are on a free plan, please use the [Transitland discussion board at GitHub](https://github.com/transitland/transitland/discussions?ref=interline.io), where we welcome questions (and answers) from all. ## Plans & Pricing Ready to use Transitland APIs for developers? View the [Transitland API plans and pricing](https://www.interline.io/transitland/plans-pricing/#transitland-apis). ### Consulting services by Interline URL: https://www.interline.io/consulting/ Last updated: 2024-09-03T23:13:36.000Z We integrate Interline’s [Transitland platform](https://www.interline.io/transitland/) and routing platform into a wide range of systems, including: - Traveler-information and journey-planning systems (especially for public mass transit and “shared” mobility) - Transportation-planning and accessibility-analysis tools - Transportation-network operation systems Our [team and partners](https://www.interline.io/about/) have experience consulting for large corporations, small start-ups, NGOs, planning firms, universities, and government agencies. Please review our offerings and [contact us](https://www.interline.io/about/#contact) for more information. ### Transitland Bulk Transit Data Exports URL: https://www.interline.io/transitland-custom-bulk-transit-data/ Last updated: 2024-11-29T21:56:18.000Z Interline produces bulk data exports from the [Transitland platform](https://www.interline.io/transitland/). We provide both standardized exports and exports that are customized for each client’s needs. Our clients use these bulk exports to power their own transportation analysis engines, real-estate assessments, routing engines, and other use-cases. ## Standardized exports Before considering about Interline’s customized data exports, review our standardized data export products: ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/09/transitland-stop-export-across-us.gif) #### All transit stops in the United States - over ****610,000** transit stop locations - includes stops for ****buses, trains, subways/metros, ferries, and cable cars** - focused on ****public transit agencies**; also including privately operated shuttles open to the public - updated on a ****daily** basis - Transitland Terms provides ****clear and consistent licensing and attribution** [Learn more about Interline's US stops data export](https://www.interline.io/transitland/all-us-stops-bulk-data/) ## Benefits of customized exports - **Use the largest open transit data archive**: Draw from Transitland’s catalog of thousands of GTFS feeds, archive of tens of thousands of GTFS feed versions, and archive of terabytes of GTFS Realtime data. - **Customize by geography**: Geographic coverage can be customized to focus on one metro region, state/province, country, or a combination. - **Customize by timeframe**: Temporal coverage can be current or historical. If we don’t have an exact match for your historical time period, Interline’s tooling can “lift and shift” GTFS service to provide approximate coverage. - **With GTFS entities plus advanced metrics**: Any of the metrics provided by Transitland APIs can be included in a bulk data export. We often add additional custom metrics as needed by clients. Focus on the analysis that makes your service or product unique, and turn to Interline to prepare the right ingredients. - **Depend on expert guidance** In addition to providing data exports, Interline staff share their expert guidance with our clients. We help our clients understand the possibilities, constraints, and options associated with each aspect of transit data. ## Export formats Interline can prepare custom bulk transit extracts in the following formats: - GTFS (merged or individual feeds) - JSON - CSV - XLSX (Excel spreadsheet) - GeoJSON - [GeoJSONL](https://www.interline.io/blog/geojsonl-extracts/) - GeoPackage - Shapefile - Protocol Buffer (for GTFS Realtime) ## Questions & Answers #### Can you create for us a single GTFS feed for an entire metropolitan region? Yes, Interline regularly prepares merged GTFS feeds for a variety of clients. See a [case study of how we do so on a daily basis for the Bay Area’s Metropolitan Transportation Commission](https://www.interline.io/blog/mtc-regional-gtfs-feed-release/). For each merged feed, we offer clients a set of options for how to handle entity identifiers (IDs), entity conflicts, and merging logic. #### We have internal GTFS data that needs to be kept private. Can you process it together with public feeds from Transitland to create a single export? Yes, Interline can process for public feeds from Transitland and private feeds supplied by our clients. #### How will we receive our transit data export? Interline provides unique, secure links for clients to download their bulk transit data exports. We can also publish bulk data to any of the three major public clouds (AWS, Google Cloud, Microsoft Azure) for direct use by our clients’ cloud accounts. #### What are the costs and turn-around times for exports? We custom create each bulk data export process, but are able to speed up the process by using the foundation of existing Transitland data and components. Please contact us for a quote. ## Request a Quote Please complete the following form for more information: [Request Quote for a Dataset](https://app.interline.io/contact%5Fforms/datasets?ref=interline.io) ### Support URL: https://www.interline.io/support/ Last updated: 2024-09-03T23:30:39.000Z If you are on a paid Transitland or routing platform plan, [sign into your Interline account](https://app.interline.io/users/sign%5Fin?ref=interline.io) to file a support ticket. If you are on a free Transitland plan, please use the [Transitland discussion board at GitHub](https://github.com/transitland/transitland/discussions?ref=interline.io), where we welcome questions (as well as helpful answers) from all. 📖 Please also check Interline's self-serve [documentation](https://www.interline.io/docs/) for answers to your questions. ### Documentation URL: https://www.interline.io/docs/ Last updated: 2024-11-01T21:09:28.000Z For documentation of current Interline services, see: - [Transitland APIs](https://www.transit.land/documentation/?ref=interline.io) - [OSM Extracts download service](https://app.interline.io/osm%5Fextracts/interactive%5Fview?ref=interline.io) - [PlanetUtils library](https://github.com/interline-io/planetutils/blob/master/README.md?ref=interline.io) - [Valhalla Tilepacks](https://www.interline.io/valhalla/tilepacks/) documentation is sent to subscribers via email - [Station Editor](https://www.interline.io/transitland/station-editor/) documentation is sent to subscribers via email The following documentation is out-of-date but retained for reference: - [HERE XYZ tutorials](https://www.interline.io/docs/here-xyz/) ### Valhalla Tilepacks URL: https://www.interline.io/valhalla-tilepacks/ Last updated: 2024-09-03T23:48:29.000Z ## Why Valhalla Tilepacks? [Valhalla](https://www.interline.io/valhalla/) is fully open-source and available for anyone to operate. In our experience helping a wide range of organizations operate Valhalla, we've found that it's simple for everyone to run Valhalla instances for routing requests — but it's difficult to run the pipeline that generates Valhalla's internal tile datasets. That's why Interline provides Valhalla Tilepacks for instant download. Interline runs a pipeline that every day combines together 40Gb of OpenStreetMap data with 1.6Tb of elevation to produce Valhalla tiles for the entire planet. You can download the entire planet or regional extracts of tiles to power your own Valhalla instances. No need to run or tune the pipeline. Update frequency is up to you: You can download once a day, weekly, monthly, or quarterly. 💡 Valhalla Tilepacks are generated in the [v3 tile format](https://github.com/valhalla/valhalla/blob/master/CHANGELOG.md?ref=interline.io#release-date-2018-11-21-valhalla-300), for us by Valhalla 3.0.0 and up. ## What data goes into build Valhalla Tilepacks? Valhalla Tilepacks are created daily from the following data sources: ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/09/osm-valhalla-connectivity.png) #### OpenStreetMap - approximately ****40 gigabytes** of OpenStreetMap data - includes roadways, paths, and turn restrictions - tags normalized to Valhalla's roadway hierarchy - updated daily from [Planet.osm diffs](https://wiki.openstreetmap.org/wiki/Planet.osm/diffs?ref=interline.io) using [Interline PlanetUtils](https://github.com/interline-io/planetutils?ref=interline.io) ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/09/mapzen-terrain-tile-example.jpg) #### Terrain Tiles - approximately ****1.6 terabytes** of terrain tiles, originally created by Mapzen - refreshed from [Terrain Tiles on AWS](https://aws.amazon.com/public-datasets/terrain/?ref=interline.io) public dataset using [Interline PlanetUtils](https://github.com/interline-io/planetutils?ref=interline.io) - used to add grade to OpenStreetMap roadways and paths ## How do I use Valhalla Tilepacks? Using Valhalla Tilepacks is as simple as: 1. Select your plan and [sign up](https://www.interline.io/valhalla/tilepacks/#pricing--sign-up). 2. [Install Interline PlanetUtils](https://github.com/interline-io/planetutils?ref=interline.io#installation) on your computer or your server. 3. [Request a new Valhalla Tilepack](https://github.com/interline-io/planetutils?ref=interline.io#valhalla%5Ftilepack%5Fdownload) using Interline PlanetUtils 4. Place the `tiles.tar` in your Valhalla server's data directory. We're available to answer any questions along the way! ## Pricing & Sign Up #### Daily Planet $624\*/ month - Global tiles built every day - Download at any frequency - Includes 30-minute Q&A call \* This price includes a 22% discount for prepayment of a 12-month contract. Prepay a 6-month contract and receive a 10% discount. Or if paying month-to-month the cost is $800 per month. [Sign up now](https://app.interline.io/products/valhalla%5Ftilepacks%5Fdaily%5Fplanet/orders/new?ref=interline.io) #### ****Don't want to set up your own Valhalla server?** Sign up for Interline's [Global Valhalla APIs](https://rapidapi.com/interline-technologies-interline-technologies-default/api/interline-global-valhalla-navigation-and-routing-engine/?ref=interline.io) through RapidAPI.com. All users are provided with a number of free requests per day and then may pay fractions of cents for each additional request they wish to make. #### Want to also plan trips by public transit? See the [Transitland Routing API](https://www.transit.land/documentation/routing-api/?ref=interline.io). ### Interline Routing Platform URL: https://www.interline.io/routing-platform/ Last updated: 2026-03-04T22:54:32.000Z The Interline Routing Platform combines together: - [Valhalla](https://www.transit.land/documentation/routing-platform/valhalla/?ref=interline.io) \- which we host globally for auto, bike, and pedestrian directions - [Transitland Routing API](https://www.transit.land/documentation/routing-platform/transitland-routing-api/?ref=interline.io) — which we host across the United States for bus, train, and subway directions 🚌 Previously Interline was more closely involved in the [OpenTripPlanner (OTP) project](https://www.opentripplanner.org/?ref=interline.io). As of 2024, have now focused our transit-routing efforts on our Transitland Routing API. Other individuals and organizations continue to support OTP. ### Transitland vs. alternative transit data sources and systems URL: https://www.interline.io/transitland-compare/ Last updated: 2025-07-22T22:25:50.000Z Compare Transitland’s powerful functionality with other mobility data processing systems and sources for GTFS, GTFS Realtime, and GBFS feeds. #### Compare Transitland vs. GTFS Data Exchange Learn about the history of GTFS Data Exchange and how its key functionality is now available through the Transitland platform. [See comparison](https://www.interline.io/transitland/compare/gtfs-data-exchange/) #### Compare Transitland vs. OpenMobilityData Learn about OpenMobilityData, the website originally known as TransitFeeds, and how it compares to Transitland. [See comparison](https://www.interline.io/transitland/compare/openmobilitydata) #### Compare Transitland vs. US National Transit Map Learn about the National Transit Map, a GIS dataset published by the US Bureau of Transportation Statistics, and how it compares to Transitland. [See comparison](https://www.interline.io/transitland/compare/us-national-transit-map) #### Compare Transitland vs. Esri Community Maps and ArcGIS Learn about Esri's recently announced transit initiative, how ArcGIS Pro can load some tables from a GTFS feed using the Network Analyst extension, and how these options compare to Transitland. [See comparison](https://www.interline.io/transitland/compare/esri-community-maps/) #### Compare Transitland vs. Google Maps Platform Learn about the Google Maps Platform Routes API, the Google Maps JavaScript API Transit Layer, and how these options compare to Transitland. [See comparison](https://www.interline.io/transitland/compare/google-maps-platform/) ### All bus, train, and subway stops across the United States URL: https://www.interline.io/transitland-all-us-stops-bulk-data/ Last updated: 2024-11-05T21:55:20.000Z Our most commonly requested bulk export from the Transitland platform is a dataset of all transit stops across the United States. This export includes information about the stop/station location itself, as well as the routes that serve that stop. Routes include bus, train, subway/metro, ferry, and cable car. Also included in the export is information about the transit agency that operates each route that serves the given stop/station location. ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/09/transitland-stop-export-across-us-1.gif) #### Each export includes: - over ****610,000** transit stop locations - includes stops for ****buses, trains, subways/metros, ferries, and cable cars** - focused on ****public transit agencies**; also including privately operated shuttles open to the public - updated on a ****daily** basis - [Transitland Terms](https://www.transit.land/terms?ref=interline.io) provides ****clear and consistent licensing and attribution** ## Use cases Interline’s clients use our standardized US transit stops data product for a wide range of purposes: - real-estate analytics - navigation applications - healthcare access ## Data sample Download a sample to explore in your own software. This sample includes 100 stops randomly selected from across the United States. The full set of columns are included for each. Please note that this sample may only be used to evaluate this service before purchasing a subscription. [transitland-bulk-stops-export-sampletransitland-bulk-stops-export-sample.csv28 KBdownload-circle](https://www.interline.io/content/files/2024/11/transitland-bulk-stops-export-sample.csv "Download") ## Data dictionary Here is an overview of the columns provided in the CSV sample. Columns about the feed from which the stop was sourced: - `feed_id` \- Transitland’s Onestop ID for the source feed. For more information open `https://www.transit.land/feeds/` - `feed_version_sha1` \- identifies the feed version from which the stop was sourced - `feed_version_fetched_at` \- datetime for when the feed version was fetched from the transit operator Columns about the stop: - `stop_onestop_id` \- Transitland’s ID for the stop location. If multiple agencies provide separate records for the same stop location, they will be merged together into a single Onestop ID, assuming lat/lon and name are very similar. For more information open `https://www.transit.land/stops/` - `stop_id` \- transit operator’s ID for the stop - `stop_name` \- name for the stop, provided by transit operator - `stop_desc` \- optional description for the stop, provided by the transit operator - `stop_lon` \- longitude coordinate for the stop location - `stop_lat` \- latitude coordinate for the stop location Between 1 and 5 agency and route combinations will be defined in a single stop row. At least 1 agency and route combination will serve the stop: - `agency_id_1` \- ID for the transit operator/agency providing service to the stop - agency\_name\_1 - name of the transit operator/agency providing service to the stop - `route_id_1` \- ID for the route providing service to the stop - `route_short_name_1` \- short name for the route providing service to the stop - `route_long_name_1` \- long name for the route providing service to the stop - `route_type_1` \- vehicle type of the route \[\* see below for a list of the different route vehicle types\] If additional agency and route combinations serve the same exact stop, additional columns will have values (using the same definitions as above): - `agency_id_2` - `agency_name_2` - `route_id_2` - `route_short_name_2` - `route_long_name_2` - `route_type_2` - `agency_id_3` - `agency_name_3` - `route_id_3` - `route_short_name_3` - `route_long_name_3` - `route_type_3` - `agency_id_4` - `agency_name_4` - `route_id_4` - `route_short_name_4` - `route_long_name_4` - `route_type_4` - `agency_id_5` - `agency_name_5` - `route_id_5` - `route_short_name_5` - `route_long_name_5` - `route_type_5` **Route types** These are the definitions for the enum integers provided in the `route_type_1`, `route_type_2`, `route_type_3`, `route_type_4`, and `route_type_5` columns. Defined by the GTFS static specification: [https://gtfs.org/reference/static#routestxt](https://gtfs.org/reference/static?ref=interline.io#routestxt) - `0` \- Tram, Streetcar, Light rail. Any light rail or street level system within a metropolitan area. - `1` \- Subway, Metro. Any underground rail system within a metropolitan area. - `2` \- Rail. Used for intercity or long-distance travel. - `3` \- Bus. Used for short- and long-distance bus routes. - `4` \- Ferry. Used for short- and long-distance boat service. - `5` \- Cable tram. Used for street-level rail cars where the cable runs beneath the vehicle, e.g., cable car in San Francisco. - `6` \- Aerial lift, suspended cable car (e.g., gondola lift, aerial tramway). Cable transport where cabins, cars, gondolas or open chairs are suspended by means of one or more cables. - `7` \- Funicular. Any rail system designed for steep inclines. - `11` \- Trolleybus. Electric buses that draw power from overhead wires using poles. - `12` \- Monorail. Railway in which the track consists of a single rail or a beam. ## How to purchase Email [info@interline.io](mailto:info@interline.io) to purchase 🌎 Need stops for the 50+ other countries covered by Transitland? Need data exports in an alternative format? Or have other requirements to discuss, see our [custom bulk transit data services](https://www.interline.io/transitland/custom-bulk-transit-data/). ### How does Transitland compare to GTFS Data Exchange? URL: https://www.interline.io/transitland-compare-gtfs-data-exchange/ Last updated: 2024-11-01T20:17:56.000Z When GTFS Data Exchange shut down in 2016, its creator recommended that its users and developers switch to Transitland. Learn how the two platforms for finding GTFS feeds compare. ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/09/gtfs-data-exchange-screenshot-2-1.png) ## Brief history of GTFS Data Exchange GTFS Data Exchange operated from 2008 to 2016 as a website “designed to help developers and transit agencies efficiently share and retrieve GTFS data.” Transit agencies provided their GTFS feeds by either registering a URL or uploading a feed version. Developers could browse the website for updates or subscribe to an RSS feed. GTFS Data Exchange was created by Jehiah Czebotar, along with assistance from Trillium Transit and Walk Score. GTFS Data Exchange grew to contain approximately 1,000 feeds and 13,000 feed versions. [In 2016](https://www.transit.land/news/2016/05/16/gtfs-data-exchange/?ref=interline.io), Czebotar closed the website to feed additions and updates. GTFS Data Exchange continues to be available in a read-only mode for users looking for older historical GTFS feed versions. The website is no longer maintained and has inconsistent service, so users should not depend upon it. GTFS Data Exchanges now points users to use Transitland or [TransitFeeds.com](https://www.interline.io/transitland/compare/openmobilitydata). ## When to use GTFS Data Exchange If you need to download historical feed versions for public transit agencies from 2008 to 2016, GTFS Data Exchange remains a useful archive. The website does have some occasional downtime, but when it is operating, users can still download old GTFS feed versions. For example, our team used historical feeds from both Transitland and GTFS Data Exchange [analyze levels of services operated by American transit agencies during and after the “Great Recession”](https://www.transit.land/news/2017/10/10/transitland-historical-feed-versions?ref=interline.io). Carole Turley Voulgaris and Charuvi Begwani at the Harvard Graduate School of Design used feed archives from GTFS Data Exchange, [OpenMobilityData](https://www.interline.io/transitland/compare/openmobilitydata), and Transitland to [analyze when transit agencies adopted GTFS and what characterized the agencies that were quickest to do so](https://findingspress.org/article/57722-predictors-of-early-adoption-of-the-general-transit-feed-specification?ref=interline.io). ## When to use Transitland Transitland regularly fetches GTFS feeds from transit agencies on an ongoing basis. Transitland has archived over 100,000 feed versions. Use the [Transitland website](https://www.transit.land/?ref=interline.io) or [Transitland v2 REST API](https://www.transit.land/documentation/rest-api/?ref=interline.io) to browse this archive of feed versions. Transitland also offers functionality that was never part of GTFS Data Exchange. For example, the [Transitland v2 REST API](https://www.transit.land/documentation/rest-api/?ref=interline.io) allows developers to use GTFS entities (such as stops, routes, and trips) without downloading and processing a GTFS feed. ## Transitland is reliable When announcing the shutdown of GTFS Data Exchange in 2016, its creator, Jehiah Czebotar, wrote: > [https://transit.land/](https://transit.land/?ref=interline.io) has steadily been growing as a reference spot for worldwide GTFS data, and the spot for transit agencies to connect with developers. It is also built to expose schedule data from a single consistent database which makes it even easier for developers to access schedule information. Most importantly though, Transitland has a corporate sponsor Mapzen, so developers can rely on it whereas gtfs-data-exchange has always been a side project, intermittently available and sometimes neglected. Transitland was sponsored by Mapzen from 2014 to 2017\. From 2018 to today, Transitland is sponsored by Interline Technologies. We are pleased to be able to provide Transitland users with reliability and continuity. Transitland continues to offer free plans for API access to hobbyists and academics. Interline also offers paid plans with support available to commercial users and organizations that require additional assistance. GTFS Data Exchange helped to establish an open culture for sharing GTFS transit data feeds among transit agencies and data consumers. Interline is glad to be able to help continue this tradition through the Transitland platform. ## Feature Comparison | | Coverage | Archived feeds | API to access GTFS entities? | Support available? | | ------------------ | -------------- | -------------- | ---------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------- | | GTFS Data Exchange | 2006 - 2016 | \~13,000 | ❌Only feed-level information available via RSS | ❌No longer maintained. Did not offer on-call support | | Transitland | 2015 - ongoing | 100,000+ | ✅Robust set of APIs for developers to access feed catalog, feed versions, GTFS entities, and GTFS Realtime information | ✅Community support on GitHub Discussions for free users and on-call support via email with Interline for paid users | 💡 Learn more about [Transitland APIs for Developers](https://www.interline.io/transitland/apis-for-developers/). ### Applied Research URL: https://www.interline.io/resources-applied-research/ Last updated: 2024-11-29T21:57:48.000Z ## Overview Interline’s [founders as well as many of our associates and partners](https://www.interline.io/about/#principals) come from a background in research. In addition to our primary focus on software platforms and tooling, Interline carries out targeted, applied research projects. ## GTFS Realtime Feed Validation and Data Quality Platform ![TR news article about open platform real-time transit data](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/tr-news-first-page.png) Interline and our partner lab the [Center for Urban Transportation Research at University of South Florida](https://www.cutr.usf.edu/?ref=interline.io) collaborated on a Transit IDEA Program grant from the US Transportation Research Board. Our goal has been to help transit agencies to validate and improve the quality of their GTFS Realtime feeds using the Transitland platform. [TR News article by Interline teamMarch-April 2021 issuetr-news-interline-article.pdf361 KBdownload-circle](https://www.interline.io/content/files/2024/11/tr-news-interline-article.pdf "Download") - The full final report can be downloaded from the [TRB website](http://www.trb.org/Main/Blurbs/181415.aspx?ref=interline.io). - Read about the development of the project in this [announcement blog post](https://www.interline.io/blog/transportation-research-board-funds-gtfs-realtime/) and this [conclusion blog post](https://www.interline.io/blog/trb-transit-idea-gtfs-realtime-report/). - Listen to a [Researching Transit podcast episode featuring our TRB project](http://publictransportresearchgroup.info/portfolio-item/rt14-dr-zhenliang-ma-harnessing-data-science-in-transit-operations-and-planning-2/?ref=interline.io). - Browse the [catalog of GTFS Realtime feeds added to the Transitland platform](https://www.transit.land/feeds?fetch%5Fspecs=gtfs-rt&feed%5Fspecs=gtfs-rt&ref=interline.io). - The California Department of Transportation (Caltrans) has adopted our report’s proposed list of the most important GTFS Realtime validation errors and warnings to include in the [California Minimum General Transit Feed Specification (GTFS) Guidelines](https://dot.ca.gov/cal-itp/california-transit-data-guidelines?ref=interline.io). ## Commuter Wallet for Transportation Demand Management ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/commuter-wallet-1.png) ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/commuter-wallet-2.png) ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/commuter-wallet-3.png) The Commuter Wallet is a software solution created to shift employees from driving alone to using their benefits for transit and other modes of travel. The Commuter Wallet enables employees to plan intermodal commutes (powered by Interline's [managed OpenTripPlanner platform](https://www.interline.io/opentripplanner)), to view benefits relevant to their commute plan, and to log their trips taken and benefits used. This process, available on employees’ smartphones and computers, replaces the need for HR forms, intranet pages, and other dispersed information sources about an employer’s commute benefits. Interline built the Commuter Wallet through a multi-step process of design, development, and deployment, with UX and UI support from our partner firm Lab Zero. This iterative approach allowed us to learn from all project stakeholders (staff commuters, transportation demand managers, and project managers at the Cities of Palo Alto, Mountain View, Menlo Park, and Cupertino) and to adapt to useful findings. [Contact us](https://www.interline.io/about/#contact) for a white paper describing the design, development, and deployment of the Commuter Wallet. ## Here XYZ Web Mapping and Analysis Tutorials ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/xyz-tutorial-2-animation-2.gif) Interline collaborated with HERE to demonstrate the many ways their XYZ developer platform is useful for working with open geodata. Users can [follow a set of tutorials](https://www.interline.io/docs/here-xyz/) to combine together OSM Extracts by Interline, Transitland APIs, and other open geodata APIs to create unique combinations and interactive visualizations. ## Transitland for Hobbyists and Academics Interline supports educational and non-profit groups with their own research efforts [by subsidizing their access to the Transitland platform](https://app.interline.io/contact%5Fforms/hobbyist%5Facademic?ref=interline.io). ### Resources URL: https://www.interline.io/resources/ Last updated: 2024-11-01T20:01:08.000Z We're pleased to offer the following complementary resources: - [Interline Blog](https://www.interline.io/blog/) - [Applied Research Reports](https://www.interline.io/resources-applied-research/) - [Transitland](https://www.transit.land/?ref=interline.io) website - [OSM Extracts](https://app.interline.io/osm%5Fextracts/interactive%5Fview?ref=interline.io) website And we're pleased to offer the following paid resources: - [Expert Consulting Services](https://www.interline.io/consulting/) ### How does Transitland compare to OpenMobilityData? URL: https://www.interline.io/transitland-compare-openmobilitydata/ Last updated: 2024-11-01T20:17:30.000Z OpenMobilityData is a website and API for browsing GTFS and GTFS Realtime feeds. The service was originally known as TransitFeeds. It is no longer maintained, but portions do remain available for users. Learn how OpenMobilityData compares to Transitland. ## Brief history of TransitFeeds and OpenMobilityData TransitFeeds.com was created in 2013 by the Australia-based app-developer Quentin Zervaas. In subsequent years, TransitFeeds grew to catalog GTFS and GTFS Realtime feeds for roughly 1,300 transit agencies. Users could browse GTFS feeds, as well as stops and routes, using the TransitFeeds.com website. While Transitland was focused at the time on providing a rich set of APIs, TransitFeeds provided an attractive and easy-to-use web interface. TransitFeed service also provided an API for browsing feed records, but no API for GTFS entity-level information. In 2019, the for-profit TransitScreen and non-profit MobilityData together acquired TransitFeeds.com from Crunchy Bagel, Quentin Zervaas’s app-development business. TransitFeeds.com was rebranded as OpenMobilityData. Both domain names continue to be in use, so users searching for GTFS feeds on Google may see similar results listed for both `TransitFeeds.com` and `OpenMobilityData.org`. More recently, MobilityData shifted its efforts to developing a new MobilityData Mobility Database service and announced that it would no longer develop OpenMobilityData. They wrote that “we commit to giving 6 months notice once the decision is finalized” to shut down the OpenMobilityData web application. # Is OpenMobilityData still functional? OpenMobilityData is partially functional as of 2023: - No feeds have been added to OpenMobilityData since January 2021. - January 2021 is probably also the last time that feed records were updated with current URLs. - Static GTFS feeds continue to be fetched for the existing feed records that still have working URLs; we see some feeds fetched as recently as December 2022. - The API is available but it is not possible for new users to sign up for an API key, as the “sign in with GitHub” button no longer works. - Users can browse existing feed records by browsing openmobilitydata.org - Users are directed to MobilityData’s Mobility Database catalogs for more recent links to GTFS feeds. OpenMobilityData continues to be a useful source for historical GTFS data between approximately 2015 and 2021. For example, Carole Turley Voulgaris and Charuvi Begwani at the Harvard Graduate School of Design used feed archives from OpenMobilityData, [GTFS Data Exchange](https://www.interline.io/transitland/compare/gtfs-data-exchange), and Transitland to [analyze when transit agencies adopted GTFS and what characterized the agencies that were quickest to do so](https://findingspress.org/article/57722-predictors-of-early-adoption-of-the-general-transit-feed-specification?ref=interline.io). ## How does OpenMobilityData license its data? OpenMobilityData does not provide information about the terms and conditions attached to each GTFS feed, and leaves it as an exercise to its users to figure out the license for each feed. Transitland provides information about each source feed’s license, terms, and conditions. Use the [Transitland v2 REST API to view feed license information and to filter results by license terms](https://www.transit.land/documentation/rest-api/feeds?ref=interline.io#feed-license-information). Users of Transitland APIs are also able to simplify their experience of working with multiple open datasets using the Transitland Terms. By including a single attribution to Transitland and a link to `https://www.transit.land/terms`, developers are able to meet the attribution requirement for all source feeds. Transitland’s information about feed-level licensing enables individual users, large organizations, and for-profit businesses to all have confidence in the open-data sources they are consuming. ## Feature Comparison | Platform | Data coverage | Web interface | API | Source feed licensing information? | | -------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------- | | **OpenMobilityData** | 2015 - 2021 | Attractive web interface for browsing GTFS feeds, routes, and stops | OpenMobilityData provided an API for its feed catalog, although new users may not sign up any more | No information about feed terms or conditions | | **Transitland** | 2015 - ongoing | Transitland was originally focused on APIs. Transitland Version 2 now also provides a comprehensive website for browsing transit data | Robust set of APIs for developers to access feed catalog, feed versions, GTFS entities, and GTFS Realtime information | Feed-level licensing info and comprehensive Transitland Terms | 💡 Learn more about [Transitland APIs for Developers](https://www.interline.io/transitland/apis-for-developers/). ### How does Transitland compare to the National Transit Map? URL: https://www.interline.io/transitland-compare-us-national-transit-map/ Last updated: 2026-04-23T18:05:09.000Z The US National Transit Map is a set of GIS layers in the National Transportation Atlas Database (NTAD). This page describes how it's produced, what it includes, and how it compares with Transitland Datasets and Transitland APIs. ## Brief History of the US National Transit Map Many public transit agencies across the United States have been producing GTFS data feeds for years. Some regional governments and state governments have also been involved in producing and disseminating GTFS data, on behalf of the transit agencies in their areas of interest. But the US federal government had no involvement in encouraging the production of GTFS feeds or helping to disseminate GTFS data. The US federal government became involved in GTFS data in 2016 when Secretary of Transportation Anthony Foxx issued [a “dear colleague” letter announcing](https://web.archive.org/web/20160412004924/http://gis.rita.dot.gov/Transit/downloads/DearColleague.pdf): > The solution is straightforward: a national repository of voluntarily provided, public domain GTFS feed data that is compiled into a common format with data from fixed route systems. ![two-page letter written on US Department of Transportation letterhead addressed 'dear colleague'](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/dear-colleague-gtfs-letter.png) [Read the entire "Dear Colleague" letter about GTFSFrom the Secretary of Transportation to transit agencies across the USDearColleague GTFS.pdf1 MBdownload-circle](https://www.interline.io/content/files/2025/07/DearColleague-GTFS.pdf "Download") Since that announcement, the Bureau of Transportation Statistics (BTS), within the US Department of Transportation (USDOT), has been tasked with creating and maintaining the National Transit Map. ## How does the National Transit Map collect GTFS feeds? The FTA began requiring National Transit Database (NTD) reporters to submit GTFS feed URLs as part of their annual reporting starting in Report Year 2023, following a [final notice](https://www.bts.gov/national-transit-map?ref=interline.io) issued on March 3, 2023\. Transit agencies submit URLs through their existing FTA Federal Access Control and Entry System (FACES) accounts. Licensing shifted with the requirement: data submitted voluntarily before the change was granted to USDOT under the Creative Commons Attribution 3.0 United States (CC-BY-3.0) license, while data submitted under the NTD requirement enters the public domain. BTS pulls the submitted feeds, extracts data, and publishes nationwide layers to the National Transportation Atlas Database (NTAD) several times a year. For background on the NTD reporting transition, see our blog post: [US National Transit Database releases data and requests more feedback](https://www.interline.io/blog/us-national-transit-database-releases-data-and-requests-more-feedback-2/). ## What data is available from the National Transit Map? The US Bureau of Transportation Statistics downloads agencies’ static GTFS feeds a few times each year and uses the source feeds to produce nation-wide geospatial layers in the National Transportation Atlas Database (NTAD). The collected GTFS source data is not released to the public. Three layers are available as part of the National Transit Map: - [**NTM Agencies**](https://geodata.bts.gov/maps/national-transit-map-agencies?ref=interline.io): point geometries for transit agencies' headquarters along with some metadata about each agency - [**NTM Stops**](https://geodata.bts.gov/maps/national-transit-map-stops?ref=interline.io): point geometries for transit stops and a subset of the metadata available about stops in the source GTFS feeds - [**NTM Routes**](https://geodata.bts.gov/maps/national-transit-map-routes?ref=interline.io): line geometries for transit routes and a subset of the metadata available about routes in the source GTFS feeds The three layers are served to the public using the BTS Open Data portal, which is powered by Esri ArcGIS Online. The portal enables basic browsing of each layer, as well as download for use in desktop GIS software or scripts. ## Comparison of transit stop coverage Both platforms assemble nationwide transit stop data from agency-produced GTFS feeds. NTM collects feeds that transit agencies submit through the FTA. Transitland sources its feeds through the open [Transitland Atlas](https://www.transit.land/documentation/atlas?ref=interline.io) feed registry. As of April 2026, raw stop counts are comparable. The key differences are update frequency, schedule data, and how each platform handles stop deduplication and non-stop GTFS location types. | Platform | Stops coverage | Update frequency | Schedule data | Formats | Licensing | Cost | | ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------- | ----------------------- | ---------------------------------------------------------------------------- | -------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **US National Transit Map**(Stops layer) | 680,275 rows (April 2026; includes \~4,000 GTFS entrances and pathway nodes alongside stops, platforms, and stations) | Approximately quarterly | Not included | CSV, KML, Shapefile, GeoJSON, File Geodatabase | Creative Commons Attribution 3.0 | Free | | **Transitland Datasets** | 643,257 US stops (April 2026 export; stops and stations, deduplicated across agencies via Transitland onestop\_id) | Approximately monthly | Per-stop departure counts for each day of the week, broken down by direction | CSV and [GeoJSONL](https://www.interline.io/blog/geojsonl-extracts/) | Source-feed open-data licenses, aggregated under the [Transitland Terms](https://www.transit.land/terms/?ref=interline.io) | Free for non-commercial use; [request a quote](https://app.interline.io/contact%5Fforms/enterprise%5Finquiry?products%5B%5D=datasets&ref=interline.io) for commercial pricing | | **Transitland v2 REST & GraphQL APIs** | Same underlying corpus as Datasets | Daily | Queryable via API | JSON, GeoJSON | Transitland Terms | Free and paid plans | ## Layering transit stops and routes in GIS Both the National Transit Map and Transitland provide options for adding transit stops and routes as layers to GIS projects. As the National Transit Map agency, stop, and route layers are hosted on ArcGIS Online, the interface for browsing and downloading the layers will be familiar to users of Esri tooling. Transitland provides Transitland v2 Vector Tiles for use as layers in a wide range of GIS tooling. See this [tutorial on how to add Transitland stop and route tiles as layers in QGIS](https://www.transit.land/news/2022/06/10/vector-map-tiles?ref=interline.io#use-transitland-v2-vector-tiles-api-in-qgis). Transitland Vector Tiles can also be used to [create web maps with libraries like Mapbox GL and Maplibre](https://www.transit.land/documentation/vector-tiles?ref=interline.io#mapbox-gl-example). ## When to use the National Transit Map and to use Transitland? **Reach for the National Transit Map when:** - You need transit layers alongside other modes in the National Transportation Atlas Database. - Your project is focused on the United States. - Quarterly or annual update cycles are sufficient for your analysis or report. - You're working in Esri tooling and want layers via the BTS ArcGIS Hub. - You're producing state- or national-scale maps where the NTM agency list fits your coverage needs. **Reach for Transitland Datasets when:** - You need schedule data (per-stop departure counts by day of week), not just stop locations. - You need richer metadata attached to stops and routes: route short and long names, agency names, vehicle types, and route colors. - You want a shorter update cycle than NTM's quarterly cadence. - You're doing non-commercial research (free) or commercial analysis (standardized pricing plans, with support). **Reach for the Transitland APIs when:** - You need trip planning analysis with up-to-date schedule data. - You need to query specific stops, routes, or feeds rather than downloading national exports. - You're combining static GTFS and GTFS Realtime feeds. - You're building web maps with MVT vector tiles. - You need international coverage beyond the US. 💡 Learn more about [Transitland Datasets](https://www.transit.land/datasets/?ref=interline.io). 💡 Learn more about [Transitland APIs for Developers](https://www.interline.io/transitland/apis-for-developers/). ### HERE XYZ URL: https://www.interline.io/docs-here-xyz/ Last updated: 2024-11-12T21:58:54.000Z 💡 ****2021** HERE has deprecated HERE XYZ and Studio 1.0\. See the [HERE Studio 2.0 documentation](https://developer.here.com/documentation/studio/dev%5Fguide/index.html?ref=interline.io) for more information. These tutorials may not describe Studio 2.0 functionality. 💡 ****October 1, 2018** Read [our blog post for an introduction to this collaboration](https://www.interline.io/blog/here-xyz-open-geo-data). Interline is collaborating with HERE to demonstrate the many ways their new [XYZ](http://explore.xyz.here.com/?ref=interline.io) developer platform is useful for working with open geodata. In these tutorials, follow along to combine together [OSM Extracts by Interline](https://www.interline.io/osm/extracts), [Transitland APIs](https://transit.land/?ref=interline.io), and other open geodata APIs to create unique combinations and interactive visualizations. ## Use Crowdsourced Data from Interline OSM Extracts in an XYZ Web Map *Skill level: Intermediate* ![web map animation](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/xyz-tutorial-1-animation-1.gif) Working with OpenStreetMap extracts typically requires desktop GIS applications or libraries to filter the many different types of features included in each extract. Those steps aren't required when you're using XYZ. Filter and display using just the HERE CLI and XYZ API. [Follow the tutorial](https://interline-io.github.io/here-xyz-tutorials/tutorial-1/?ref=interline.io#0) ## Build an Interactive Web Map of Subway Stations and Routes Showing How Long You'll Typically Wait *Skill level: Advanced* ![web map animation](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/xyz-tutorial-2-animation.gif) Transitland provides open public-transit data from around the world, queryable using an API. Use the HERE CLI and XYZ API to combine together many different Transitland API queries, to create a map of stop locations, route lines, and the average time you'll have to wait to travel on each. [Follow the tutorial](https://interline-io.github.io/here-xyz-tutorials/tutorial-2/?ref=interline.io) ### Terms of Service URL: https://www.interline.io/legal-terms/ Last updated: 2025-05-09T22:40:21.000Z **Effective as of May 9, 2025** This revision: - Adjusts section numbering - Revises Section 2 regarding Third Party Data - Revises Section 2 to cover attribution of OpenStreetMap and Transitland data - Adds Section 5 regarding Publicity - In Section 6 specifies Alameda County, California Please carefully read these Terms of Service ("Terms"), which govern your use of this website and the geographical data service ("Service") provided through this website and owned and operated by Interline Technologies LLC ("Interline," "we" or "us"). 1. Registration and Ordering - To access certain features or areas of the Service and Content, you may be required to provide personal information such as your name and email address as part of a registration or log-in process. In addition, certain features of the Service are only available to our registered users, and to access those areas of the Services you will be required to log in using your username and email password. By submitting an order form or proceeding with an order on this website ("Order") or clicking "accept," you indicate you have read, understand, accept and agree to be bound by these Terms on behalf of yourself as an individual and, if applicable, as an authorized representative of your company, association, agency or other legal entity. If you do not agree with these Terms, do not submit an Order or click "accept." These Terms are effective when you accept them. You acknowledge these Terms are a contract between you and Interline, even though it is electronic and is not physically signed. These Terms supersede any other agreements between you and Interline. - You represent you have the right to provide the information you provide when you register for the Service and it is accurate and complete. If you accept these Terms on behalf of a company, association, agency or other legal entity, you represent you have the authority to bind the entity to these Terms, in which case the terms "you," "your" or related terms refer to the entity. You are responsible for all activity occurring when the Service is accessed through your account, whether authorized or not. We are not responsible for any loss or damage arising from your failure to protect your login or account information. Information you provide to us when you set up an account with us, and any billing information that we, or our billing processors, collect from you when you purchase Services is protected by our Privacy Policy at , the most current terms of which are incorporated herein by reference. 2. Service Description, Access and Use - Service Description. The "Service" includes all data, text, images, sounds, videos, reports, and other content made available through this website ("Content"), and any training and support provided by Interline. - License. Subject to these Terms, Interline grants you a worldwide, non-exclusive, non-transferable and revocable license to access the Service and use the Content ordered through the Service, on a one time basis or subscription basis during a specified term as set forth in the applicable Order, solely for personal or internal business use to integrate the Content into your apps, software or devices, and create and distribute maps and routing outputs for use by your end users. Certain portions of the Content are provided under license from third parties, and are subject to copyright and other intellectual property rights owned or licensed by such third parties. We and our licensors have the right to enforce such rights as contractual rights pursuant to these Terms. - Limitations. You may not (a) license, sublicense, sell, resell, rent, lease, transfer, assign, distribute, time share, host or otherwise commercially exploit or make the Service or Content available to any third party on a standalone basis or in a competitive service, except as expressly permitted by these Terms; (b) modify, adapt or "hack" the Service or Content to falsely imply any sponsorship or association with Interline, or otherwise attempt to gain unauthorized access to the Service or its related systems or networks; (c) use the Service or Content in any unlawful manner, including but not limited to violation of any person's privacy rights, infringing any person's intellectual property rights, or sending spam or otherwise duplicative or unsolicited messages in violation of applicable law, (d) use the Service or Content in any manner that interferes with or disrupts the integrity or performance of the Service or Content and its components; (e) attempt to decipher, decompile, reverse engineer or otherwise discover the source code of any software making up the Service or Content; (f) use the Service or Content to post, upload, link to, send or store any content that is unlawful, racist, hateful, obscene, discriminatory, or contains any viruses, malware, Trojan horses, time bombs, or any other similar harmful software; or (g) or use the Service or Content to create, deliver training on, improve (directly or indirectly) or other a substantially similar product or service, or otherwise in violation of these Terms. - End User License Agreement. For your products and service, you must have users accept an enforceable an end-user license agreement that includes a disclaimer of all liability of Interline (or generally, your "licensors"). If you use location data in the Content to produce dynamic maps and other visualizations that can be used to provide directions and similar results for a variety of commercial and consumer applications, such as turn-by-turn route guidance and other routing ("Real Time Navigation"), then you must include the following notice in your end user license agreement: "YOUR USE OF THIS REAL TIME ROUTE GUIDANCE APPLICATION IS AT YOUR SOLE RISK. LOCATION DATA MAY NOT BE ACCURATE." - Third Party Data and Services. Depending on the type of Service you use, you may be required to use data and materials of third parties, such as but not limited to OpenStreetMap data, GTFS data feeds, GTFS Realtime data feeds, or GBFS data feeds, whether you obtain them separately or through us ("Third Party Data"). Your use of Third Party Data may also be governed by terms and conditions of those Third Party Data providers. For example, any data subject to an ODC Open Database License (ODbL) may be used for commercial purposes only according to its terms and provided you comply with its attribution and share-alike terms. This website and the Service may also contain links to, or otherwise may allow you to connect to and use certain third-party products, services or software under separate terms and conditions (collectively, "Other Services") in conjunction with Interline's Service. If you decide to access and use such Other Services, be advised your use is governed solely by the terms and conditions of such Other Services, and Interline does not endorse, is not responsible for, and make no representations as to such Other Services, their content or the manner in which they handle your data. Interline is not liable for any damage or loss caused or alleged to be caused by or in connection with your access or use of any such Other Services, or your reliance on the privacy practices or other policies of such Other Services. - Proprietary Rights. All rights, title and interest in and to the Service and Content, all improvements, modifications and derivative works, and all related intellectual property rights, will remain with and belong exclusively to Interline and its third-party licensors. We reserve the right to make changes to, or to suspend or discontinue (temporarily or permanently), the Service, Content or any portion of the Services and Content. You agree that we will not be liable to you or to any third party for any such modification, suspension or discontinuance. - Attribution. You agree to attribute Content by including the following notice on all maps and analyses you create: - When using OpenStreetMap data sourced through Interline Services: "© OpenStreetMap contributors." - When using the transit feeds and data sourced through Interline's Transitland Service: "Transitland" or the Transitland logo, with a link to [https://www.transit.land/terms](https://www.transit.land/terms?ref=interline.io) - Communications. By using the Service, you consent to receiving electronic communications from Interline. These electronic communications may include notices about applicable fees and charges, transactional information and other information concerning or related to the Service. These electronic communications are part of your relationship with Interline and you receive them as part of your purchase. You agree any notices, agreements, disclosures or other communications we send you electronically will satisfy any legal communication requirements, including such communications be in writing. You agree Interline may access your account information to respond to your requests and support the Service. We will not disclose your account data unless permitted by you, compelled by law, or pursuant to the terms of our Privacy Policy available at , the most current terms of which are incorporated herein by reference. - Payment. You agree Interline may bill you the applicable fees per your Order and associated taxes (if any), and charge your credit card or invoice you in accordance with the Order billing terms in effect at the time the fees are due and payable. Invoices are due 14 days after the date of invoice. All purchases are non-cancellable and all charges are non-refundable except as expressly set forth herein. Payments made after their due date will incur a daily simple interest from the original invoice due date at a rate equal to one percent (1%) per month or the maximum rate permitted by applicable law, whichever is lower. You shall pay all such interest and reasonable costs of collection, including but not limited to, reasonable attorneys' fees and court costs. If payment cannot be charged to your credit card or is not received for any reason, Interline may suspend or terminate your Order and your access to the Service. 3. Maintenance and Support - You are solely responsible for maintaining and supporting the products and services into which you incorporate the Content, and any complaints about your products and services, and any loss, liability, damages, fines, penalties, costs and expenses that you, a user, we or a third party may suffer arising out of the use or distribution of your products and services. - Interline will maintain and support the Service and Content in accordance with support terms set forth at . 4. Warranty Disclaimer; Limitation of Liability; Indemnification. - The Service and Content are provided on an "as is" and "as available" basis without any warranties of any kind to the fullest extent permitted by law. Interline expressly disclaims any and all warranties, whether express or implied, including without limitation, any warranty of title or non-infringement and the implied warranties of merchantability or fitness for a particular purpose. You agree that use of our Service or Content is at your own risk and acknowledge Interline does not warrant the Service and/or Content will be uninterrupted, secure, timely, or error or virus free. Although we try to make the Content accurate and current, we reserve the right to change or correct the Content at any time. We cannot and do not guarantee the correctness, currency, precision or completeness of the Content. - If you are dissatisfied with the Service or Content, your sole remedy is to cease using the Service. Neither Interline nor any of its directors, officers, employees, agents, licensors or suppliers, to the maximum extent permitted by law, shall be liable for any incidental, consequential, special, indirect, punitive, exemplary or similar damages, including without limitation, lost profits, savings or revenue, business interruption, the use or inability to use the Service and/or any Content, or any other loss, however caused and on any theory of liability, whether liability is asserted in contract, tort (including negligence), strict liability, or otherwise, in any way arising out of the Service or Content, even if advised of the possibility of such damage and notwithstanding the failure of the essential purpose of any remedy. Interline's total liability shall be limited to the lesser of (a) direct damages incurred by you or (b) fees received by Interline from you during the twelve (12) months immediately prior to the date on which the cause of action for such claim arose. - Some states do not allow the exclusion of implied warranties or limitations of liability for incidental or consequential damages, which means some of the above limitations may not apply to you. In these states, Interline's liability will be limited to the greatest extent permitted by applicable law. - Neither Interline nor any of its directors, officers, employees, agents, licensors or suppliers will be responsible or liable for any damages or losses arising out of, relating to or that result from (a) your use of the Content in high risk activities where the use or failure of the Service or Content could lead to death, personal injury or property damage, or (b) for your end users' use of the Content in your products and services. You agree to defend, indemnify, and hold harmless Interline from and against any claims, actions or demands, including, without limitation, reasonable legal and professional services fees, arising or resulting from your breach of these Terms, or your and your end users' access to, use, misuse or unlawful use of the Service or Content. Interline will provide you notice of any such claim, suit, or proceeding. Interline reserves the right to assume the exclusive defense and control of any matter which is subject to indemnification under this section, in which case you agree to cooperate with any reasonable requests to assist Interline's defense of such matter. 5. Publicity - Interline is pleased to have you a customer. You hereby grant us a worldwide, non-exclusive, royalty-free, fully paid-up, transferable and sublicensable license to use your trademarks, service marks, and logos for the purpose of identifying you as an Interline customer to promote and market our services. 6. General - Termination. You may terminate these Terms at any time by ceasing use of the Service but you will not receive any refund of prepaid fees for submitted Orders. Interline may terminate these Terms on 30 days' prior written notice to you if you breach any of these Terms and the breach remains uncured at the end of such 30-day period. On any termination of these Terms you will cease using the Service and Interline may cease providing Content you ordered. All accrued but unpaid amounts and Sections 1, 2, 4 and 6 will survive any termination of these Terms. We have the right to deny access to, and to suspend or terminate your access to the Services and Content at any time and for any reason, including without limitation for any violation by you of these Terms of Use. In the event that we suspend or terminate your access to and/or use of the Service or Content, you will continue to be bound by the Terms of Use that were in effect as of the date of your suspension or termination. - Assignment. These Terms shall bind and benefit the parties and their respective successors and assigns. - Amendment. Interline may amend these Terms from time to time, in which case the new Terms will supersede prior versions. It is your responsibility to review these terms from time to time; provided, Interline will provide registered active users of the Service with an email notice of the change. In the event you do not agree to any revised Terms, you must immediately cease using the Service and these Terms will terminate. - Severability. If any provision in these Terms is held by a court of competent jurisdiction to be unenforceable, such provision shall be modified by the court and interpreted so as to best accomplish the original provision to the fullest extent permitted by law, and the remaining provisions of these Terms shall remain in effect. - Governing Law. These Terms shall be governed by the laws of the State of California without regard to conflict of laws principles. you hereby expressly agree to submit to the exclusive personal jurisdiction of the federal and state courts of the State of California located in Alameda County, to resolve any dispute relating to these Terms or your access to or use of the Service. - Waiver. Interline's failure to enforce at any time any provision of these Terms does not constitute a waiver of provision or of any other provision of these Terms, which must be in writing signed by Interline. - Force Majeure. Interline shall not be responsible or liable by reason of any failure or delay in the performance of its obligations under these Terms because of any cause beyond its reasonable control. - Export. Certain Content and software components of the Service may be subject to U.S. export control and economic sanctions laws. If you are subject to U.S. laws, you agree to comply with all such laws and regulations as they relate to the software and Content, and access and use of the Service. You shall not access or use the Service if you are located in a country in which the U.S. Department of Commerce, Bureau of Industry and Security, has prohibited certain exports. A list of such prohibited jurisdictions can be found its website at [http://www.bis.doc.gov/index.php/regulations/export-administration-regulations-ear](http://www.bis.doc.gov/index.php/regulations/export-administration-regulations-ear?ref=interline.io). You shall also not provide access to the Service to any government, entity or individual located in such prohibited jurisdictions. ## Posts ### Riding transit to California's beaches URL: https://www.interline.io/blog/riding-transit-to-californias-beaches/ Last updated: 2026-04-23T17:12:35.000Z The California Coastal Commission recently updated [YourCoast](https://www.coastal.ca.gov/YourCoast/?ref=interline.io#/map), an interactive web map of over 1,600 public access points along the California coast. It's a genuinely useful tool: you can filter by restrooms, surfing, camping, dog-friendliness, stroller access, and more. You can also filter by car parking. ![Screenshot of the YourCoast web map with "Parking" toggled on, showing clusters of access points up and down the California coast.](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2026/04/yourmap-car-parking-amenity.png) What you *can't* filter by is public transit. The filter panel lists ten amenity types, but getting to the beach without a car isn't one of them. This is an unfortunate gap for an agency whose founding mandate under the landmark California Coastal Act is to "maximize public access" to the coast. ## California's coastal access data has just a bit of transit The Commission publishes its access point data as an [open dataset](https://california-coastal-commission-open-data-1-3-coastalcomm.hub.arcgis.com/datasets/e45f130ce101423ea1de943a69445ece%5F0/explore?ref=interline.io), which is great. Digging into the raw CSV, there *is* a `PUBTRANSP` column, but out of 1,631 access points, only 39 have a value of "Yes" and 10 have "No." The remaining 1,582 are blank. That field is currently too sparse to be useful. So, let's enrich the Coastal Commission's public access points using Transitland Datasets. ## Spatial join: Transitland Datasets + Coastal Commission open data Rather than calling an API for each of the 1,600+ access points, we'll use [Transitland Datasets](https://www.transit.land/documentation/datasets/data-dictionary?ref=interline.io), bulk exports of transit stop and route data that you can download and work with locally. What makes Transitland Datasets especially useful here is the **schedule summary data**: for every stop, the dataset includes departure counts broken down by day of week. This isn't just a list of stop locations. Rather, it tells you *how much service actually runs* at each stop, every day of the week. The plan: 1. **Load the Transitland stops dataset** and filter to California using the `adm1_iso` column 2. **Load the Coastal Commission's access points CSV** 3. **Search at two radii**: a half mile (\~805 m) for stops within walking distance, and three miles (\~4.8 km) for stops that are in the approximate area but might require a hike or a bike. 4. **Aggregate the results**, including which routes serve nearby stops and how often they run on weekdays and weekends ## Getting the data You'll need two files: - **Transitland stops dataset (US)**: Available through [Transitland Datasets](https://www.transit.land/datasets/?ref=interline.io) (which are free for non-commercial use). The stop-oriented CSV includes stop locations, route names, agency names, vehicle types, and scheduled departure counts for each day of the week. See the [data dictionary](https://www.transit.land/documentation/datasets/data-dictionary?ref=interline.io) for full column definitions. - **California Coastal Commission access points**: Download from their [ArcGIS Hub open data site](https://california-coastal-commission-open-data-1-3-coastalcomm.hub.arcgis.com/datasets/e45f130ce101423ea1de943a69445ece%5F0/explore?ref=interline.io) in CSV format. ## Python script using GeoPandas and uv Here's a Python script that performs the analysis using [GeoPandas](https://geopandas.org/?ref=interline.io) for the spatial join. It defines its dependencies inline, so if you have [uv](https://docs.astral.sh/uv/?ref=interline.io) installed, you can run it directly (with no need to set up a virtualenv or perform a pip install): ```python #!/usr/bin/env -S uv run --script # /// script # requires-python = ">=3.10" # dependencies = ["geopandas", "pandas"] # /// """ Find California coastal access points with nearby public transit by spatially joining Coastal Commission open data with a Transitland Dataset. Searches at two radii: 🚌 ½ mile: walking distance to a transit stop 🔍 3 miles: transit in the area, worth planning a trip Usage: uv run coastal_transit.py """ import sys import pandas as pd import geopandas as gpd WALK_M = 805 # ~0.5 miles NEAR_M = 4828 # ~3.0 miles def _clean(val): """Return empty string for NaN-ish values.""" s = str(val or "").strip() return "" if s in ("", "nan", "None") else s def load_stops(path): """Load Transitland stops CSV, filter to California, build route labels.""" print("Loading Transitland stops...", file=sys.stderr) df = pd.read_csv(path, dtype=str) ca = df[df["adm1_iso"] == "US-CA"].copy() ca["stop_lat"] = ca["stop_lat"].astype(float) ca["stop_lon"] = ca["stop_lon"].astype(float) ca["tue"] = pd.to_numeric(ca["departure_count_dow2"], errors="coerce").fillna(0).astype(int) ca["sat"] = pd.to_numeric(ca["departure_count_dow6"], errors="coerce").fillna(0).astype(int) ca["sun"] = pd.to_numeric(ca["departure_count_dow7"], errors="coerce").fillna(0).astype(int) def route_labels(row): labels = [] for i in range(1, 6): name = _clean(row.get(f"route_short_name_{i}")) name = name or _clean(row.get(f"route_long_name_{i}")) agency = _clean(row.get(f"agency_name_{i}")) if name and agency: labels.append(f"{name} ({agency})") return labels ca["routes"] = ca.apply(route_labels, axis=1) print(f" {len(ca):,} California stops", file=sys.stderr) return gpd.GeoDataFrame( ca, geometry=gpd.points_from_xy(ca.stop_lon, ca.stop_lat), crs="EPSG:4326" ) def load_access_points(path): df = pd.read_csv(path, encoding="utf-8-sig") df = df.dropna(subset=["LATITUDE", "LONGITUDE"]) print(f" {len(df):,} coastal access points", file=sys.stderr) return gpd.GeoDataFrame( df, geometry=gpd.points_from_xy(df.LONGITUDE, df.LATITUDE), crs="EPSG:4326" ) def spatial_join(stops, pts, radius_m): """Buffer access points by radius_m and find all stops within.""" buf = pts.copy() buf["geometry"] = buf.geometry.buffer(radius_m) joined = gpd.sjoin(stops, buf, predicate="within") return joined.groupby("index_right").apply( lambda g: pd.Series({ "stops": len(g), "tue": int(g["tue"].sum()), "sat": int(g["sat"].sum()), "sun": int(g["sun"].sum()), "routes": list(dict.fromkeys(r for rs in g["routes"] for r in rs)), }), include_groups=False, ) def main(): if len(sys.argv) < 3: print("Usage: uv run coastal_transit.py ") sys.exit(1) stops = load_stops(sys.argv[1]).to_crs("EPSG:3310") pts = load_access_points(sys.argv[2]).to_crs("EPSG:3310") walk = spatial_join(stops, pts, WALK_M) near = spatial_join(stops, pts, NEAR_M) # Build result: prefer ½-mile data; fall back to 3-mile rows = [] for idx, pt in pts.iterrows(): w = walk.loc[idx] if idx in walk.index else None n = near.loc[idx] if idx in near.index else None if w is not None and w["stops"] > 0: tier, data = "walk", w elif n is not None and n["stops"] > 0: tier, data = "near", n else: tier, data = "none", None rows.append({ "name": str(pt["Name"])[:49], "county": pt["COUNTY"], "tier": tier, "stops": int(data["stops"]) if data is not None else 0, "tue": int(data["tue"]) if data is not None else 0, "sat": int(data["sat"]) if data is not None else 0, "sun": int(data["sun"]) if data is not None else 0, "routes": data["routes"] if data is not None else [], }) result = pd.DataFrame(rows) # Print table print(f"\n{'Name':50s} {'County':15s} {'':4s} {'Stops':>5s} {'Tue':>5s} {'Sat':>5s} {'Sun':>5s} Routes") print("=" * 120) for _, r in result.iterrows(): icon = {"walk": "🚌", "near": "🔍", "none": " "}[r["tier"]] route_str = ", ".join(r["routes"][:4]) if len(r["routes"]) > 4: route_str += f" (+{len(r['routes']) - 4} more)" print( f"{r['name']:50s} {r['county']:15s} {icon:4s} " f"{r['stops']:>5} {r['tue']:>5} {r['sat']:>5} {r['sun']:>5} " f"{route_str or '-'}" ) walk_n = (result["tier"] == "walk").sum() near_n = (result["tier"] == "near").sum() none_n = (result["tier"] == "none").sum() print(f"\n{'='*60}") print(f"🚌 Within ½ mile (walkable): {walk_n:>5,}") print(f"🔍 Within 3 miles (hikeable): {near_n:>5,}") print(f" No transit within 3 miles: {none_n:>5,}") print(f" Total: {len(result):>5,}") if __name__ == "__main__": main() ``` Run it: ```bash uv run coastal_transit.py tl-dataset-US-stops.csv AccessPoints.csv ``` `uv` reads the inline dependency metadata, installs `geopandas` and `pandas` into an isolated environment, and runs the script. The spatial joins use [California Albers](https://epsg.io/3310?ref=interline.io) (EPSG:3310) for accurate meter-based distance calculations. ## Results: transit frequency to reach the California coast We ran that script against an April 2026 Transitland Dataset covering all US transit operators. Out of 1,625 coastal access points with coordinates: - 🚌 **913** have a transit stop within a half mile (more than half of all coastal access points in the state) - 🔍 **427** more have transit within three miles (likely hikeable, assuming there are appropriate pedestrian shoulders or trails) - The remaining **285** have no transit within three miles Here's a sample of results across all three tiers: | Name | County | Tier | Stops | Tue | Sat | Sun | Routes | | -------------------------------- | ----------- | ---- | ----- | ----- | --- | --- | ------------------------------------------------------------------ | | Santa Monica State Beach | Los Angeles | 🚌 | 10 | 732 | 497 | 497 | 2, 9, 43, 3 (Big Blue Bus) | | Santa Cruz Beach Boardwalk | Santa Cruz | 🚌 | 7 | 426 | 342 | 344 | 19B, 20, 19 (Santa Cruz METRO) +1 more | | Monterey Bay Aquarium | Monterey | 🚌 | 17 | 547 | 419 | 419 | 2, 1, A (Monterey-Salinas Transit) +2 more | | Half Moon Bay State Beach | San Mateo | 🔍 | 60 | 1,054 | 885 | 885 | 294, 15, 117, 18 (SamTrans) | | Huntington City Beach | Orange | 🚌 | 13 | 506 | 551 | 549 | 29, 29A, 25 (OCTA) +1 more | | Coronado City Beach | San Diego | 🚌 | 9 | 451 | 365 | 243 | 901, 904 (MTS) | | Beach Front Park (Crescent City) | Del Norte | 🚌 | 22 | 267 | 156 | 2 | 199, 1, 20, 2 (Redwood Coast Transit) +5 more | | Fort Bragg Coastal Parkland | Mendocino | 🚌 | 5 | 73 | 10 | 4 | 5, 65 (Mendocino Transit Authority) | | Point Reyes Hostel | Marin | 🔍 | 2 | 17 | 14 | 14 | 68 (Marin Transit) | | Peter Strauss Ranch | Los Angeles | 🔍 | 54 | 879 | 476 | 476 | 161 (Metro - Los Angeles), 423, 422 (LADOT), KS (Kanan Shuttle) | | Pelican State Beach | Del Norte | 🔍 | 4 | 12 | 8 | 0 | 20 (Redwood Coast Transit), Coastal Express (Curry Public Transit) | | Pfeiffer Big Sur State Park | Monterey | | 0 | 0 | 0 | 0 | | A few patterns to point out: - **Departure counts tell a story that proximity alone can't.** The table uses Tuesday as a representative weekday (as that day avoids most holidays). Santa Monica has 732 weekday and nearly 500 Sunday departures across 10 nearby stops, so you can show up any day without needing to closely follow a schedule. Fort Bragg has stops nearby, but 73 weekday departures drop to 10 on Saturday and 4 on Sunday. Crescent City's Beach Front Park has 267 weekday departures that drop to 156 on Saturday and just 2 on Sunday. Instead of asking *is there a bus stop nearby?* using Transitland Datasets, we can ask *is there a bus stop with service nearby?* (Note that these are raw departure counts summed across all nearby stops and both directions, so they represent *service intensity* in the area.) - **Some access points show zero transit even at three miles.** This is genuinely the case for remote stretches like the Lost Coast and Big Sur. (In fairness, Highway 1 can be washed out by weather, so you may not be traveling to some of those coastal access points by private auto either!) - **The three-mile radius catches real transit corridors where the stop isn't right at the beach.** Point Reyes Hostel, for example, has no stops within a half mile but 17 weekday and 14 Sunday departures within three miles. That's Marin Transit route 68 running along Sir Francis Drake Boulevard to Point Reyes Station village. In Southern California, Peter Strauss Ranch is another example: 879 weekday and 476 Sunday departures within three miles, from LA Metro 161 plus LADOT and Calabasas shuttles threading through the Santa Monica Mountains, but none of them drop you at the trailhead. The 🔍 tier flags these cases: transit exists in the corridor, and it's worth checking the [Transitland map](https://www.transit.land/map?ref=interline.io) or using the Transitland Routing API to see if the full journey is actually feasible. (Note that some of these sites aren't literally right on the water. They may be on bluffs or in ocean-fronting preserves, so a hike is part of the point.) - **Half Moon Bay State Beach is a good case for the "hikeable" tier.** It shows 60 stops and over 1,000 weekday departures within three miles (SamTrans routes 294, 15, 117, and 18) but the closest stops are on Highway 1, not at the beach entrance. A short walk or bike ride will take you from the bus to the beach. (That's a favorite spot of our Bay Area-based staff!) This is exactly the kind of analysis that the schedule summary columns in Transitland Datasets are designed for. Each stop row includes departure counts for every day of the week (`departure_count_dow1` through `departure_count_dow7`). Departure counts are also broken apart by direction, for finer-grained analysis, but that's beyond the scope of this blog post. 💡 ****How does this compare to other open transit data?** The [USDOT National Transit Map](https://www.interline.io/transitland/compare/us-national-transit-map/) provides stop locations and route associations, and is a great resource for **visually* mapping where transit exists. Transitland Datasets are unique in summarizing and embedding temporal schedule data as well. ## Going further: full trip planning with Transitland's routing API Finding nearby stops and counting departures tells you whether transit *exists* near a beach. It doesn't tell you whether you can actually *get there* from home in a reasonable amount of time. For that, you need a transit routing engine. The [Interline Routing Platform](https://www.interline.io/routing/) provides transit routing powered by the same GTFS feeds behind Transitland, combined with OpenStreetMap's pedestrian network via the Valhalla engine. You could use it to compute actual travel times from residential areas to coastal access points, accounting for real schedules, transfers, and the final walking leg from the last stop to the sand. Imagine a version of YourCoast where the filter panel included: *"Reachable by transit from my location in under 60 minutes."* Or imagine an equivalent for your own state or country with a rich dataset of destinations, ready to reach via public transit. The data and APIs to build that exist today. ## Resources - **Transitland Datasets**, bulk CSV/GeoJSON exports of stops, routes, and schedules: [transit.land/documentation/datasets/data-dictionary](https://www.transit.land/documentation/datasets/data-dictionary?ref=interline.io) - **Transitland map**, explore transit stops and routes visually: [transit.land/map](https://www.transit.land/map?ref=interline.io) - **Transitland REST API (stops endpoint)**, query individual stops by location: [transit.land/documentation/rest-api/stops](https://www.transit.land/documentation/rest-api/stops?ref=interline.io) - **Transitland API key signup**: [transit.land/documentation#signing-up-for-an-api-key](https://www.transit.land/documentation/index?ref=interline.io#signing-up-for-an-api-key) - **California Coastal Commission open data**: [ArcGIS Hub dataset](https://california-coastal-commission-open-data-1-3-coastalcomm.hub.arcgis.com/datasets/e45f130ce101423ea1de943a69445ece%5F0/explore?ref=interline.io) - **YourCoast web map**: [coastal.ca.gov/YourCoast](https://www.coastal.ca.gov/YourCoast/?ref=interline.io#/map) ### Explore GTFS-Pathways station layouts using Transitland URL: https://www.interline.io/blog/explore-gtfs-pathways-station-layouts-using-transitland/ Last updated: 2026-03-05T04:48:03.000Z [Interline Station Editor](https://www.interline.io/saas-station-editor/) has helped transit agencies and consultants build GTFS-Pathways data for major transit hubs across North America. Now we're also bringing the core of that functionality to Transitland, so anyone can explore and inspect pathways data, wherever it exists and however it was originally created. 💡 GTFS-Pathways data is useful to provide wheelchair-accessible navigation and to enrich wayfinding instructions for riders using smartphone apps. The GTFS-Pathways extension adds `pathways.txt` and `levels.txt` files, as well as stops with `location_type={3,4}` so GTFS can represent the physical layout of a transit station. In contrast with other indoor mapping data formats, GTFS-Pathways represents a topological "skeleton" of how riders can travel through a multi-level subway, train, or bus station. Find a pathways-enabled station in Transitland and open the Transitland Pathways Explorer to use the interactive Map View tab to view the structure of multi-level stations. For example, [WMATA's Metro Center](https://www.transit.land/stops/s-dqcjr1jyg8-metrocenterredlinetrack1platform/pathways?ref=interline.io) underground subway station: ![animation of using Transitland to explore pathways on the underground Metro Center subway station in Washington, D.C.](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2026/03/CleanShot-2026-03-04-at-15.04.22.gif) Use the Network Diagram tab to understand the topological organization of a station's levels and nodes. For example, the San Francisco Bay Area's [Millbrae Intermodal Station](https://www.transit.land/stops/s-9q8vzhbs2k-millbraebart/pathways?ref=interline.io), which is served by BART and Caltrain and bus agencies: ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2026/03/CleanShot-2026-03-04-at-20.23.16@2x.png) Use the Route Finder tab to test a journey through a station. For example, [MBTA's Downtown Crossing](https://www.transit.land/stops/s-drt2yyxdb8-downtowncrossing/pathways?ref=interline.io) underground subway station: ![using Transitland to find a route through MBTA's Downtown Crossing subway station in Boston](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2026/03/image-2.png) The Validation tab provides a high-level summary of potential issues to debug and improve within a station's GTFS-Pathways data. 💡 The Transitland Pathways Explorer is great for understanding what's publicly available. When you need to build it, fix it, or validate it to detailed production standards, that's what [Interline Station Editor](https://www.interline.io/saas-station-editor/) is for. To find pathways-enabled stations, use the Transitland global transit map and look for the notice that "*Station pathway points are in view*" ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2026/03/CleanShot-2026-03-04-at-15.12.40@2x.png) Click on (or near) a relevant stop point to view the stop details, and to then open the Transitland Pathways Explorer: ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2026/03/CleanShot-2026-03-04-at-16.00.08@2x.png) Transitland Pathways Explorer is available to all registered Transitland users, including those on the free plan. As more transit operators publish GTFS-Pathways data in their feeds, we look forward to this functionality being even more widely useful. ### Same data, different zip: How to identify unique GTFS feed versions URL: https://www.interline.io/blog/gtfs-checksum-versioning/ Last updated: 2026-02-11T22:45:59.000Z If you've ever downloaded a GTFS feed twice from the same URL and gotten two different zip files, you've encountered a surprisingly common problem in the transit data ecosystem. The data inside hasn't changed (the same stops, routes, and schedules) but the zip archive itself is different. Maybe the server regenerated it dynamically. Maybe the compression level changed. Maybe the file timestamps inside the archive shifted by a second. For Transitland, which archives and imports [thousands of GTFS feeds](https://www.transit.land/feeds?ref=interline.io) from around the world, this matters. We need to know: *has this feed actually changed, or is it just the same data in a new wrapper?* The answer lies in how we calculate checksums — and you can use the same tools we do. ## The problem with hashing zip files The naive approach to detecting changes is straightforward: compute a [SHA1 hash](https://en.wikipedia.org/wiki/SHA-1?ref=interline.io) of the zip file and compare it to what you had before. If the hashes differ, something changed. But zip archives are containers, and containers have metadata. A zip file's checksum can change for reasons that have nothing to do with the transit data inside: - **Dynamic generation**: Some transit agencies serve GTFS through vendor platforms that create the zip on the fly for each download. Same data, different zip bytes every time. - **Compression differences**: Repackaging a zip with a different compression level produces a different file. - **Timestamps and ordering**: The internal file modification times and the order of entries within the archive can vary between builds. If Transitland treated every new zip checksum as a new feed version, we'd be archiving and importing duplicate data constantly, wasting storage, processing time, and cluttering the archive with phantom "updates." ## Two checksums, two purposes Transitland calculates two SHA1 checksums for every [feed version](https://www.transit.land/documentation/concepts/static-gtfs-feed-versions?ref=interline.io): **Zip SHA1**: a conventional hash of the entire zip archive file. This is fragile by nature (it changes whenever the packaging changes), but it's useful as a unique identifier for the exact file you downloaded. Transitland uses this as the primary key for feed versions in its [REST API](https://www.transit.land/documentation/rest-api/feed%5Fversions?ref=interline.io) and website URLs. **Directory SHA1**: a content-aware hash that looks *inside* the zip, directly at the GTFS CSV files. This is the one that solves the deduplication problem. Two zip files with completely different Zip SHA1 values will share the same Directory SHA1 if the actual transit data is identical. When Transitland's [feed fetching service](https://www.transit.land/documentation/concepts/source-feeds?ref=interline.io) downloads a new copy of a feed, it checks both hashes against its archive. If either one matches an existing feed version, it knows the data isn't new and skips the import. ## Inside the Directory SHA1 The Directory SHA1 algorithm in [transitland-lib](https://www.transit.land/documentation/transitland-lib?ref=interline.io) is deliberately simple, but each design choice is intentional: 1. **Open the zip and enumerate its entries.** We look at what's inside the archive, not the archive itself. 2. **Filter to GTFS content files.** Only lowercase `.txt` files in the root directory of the archive are included — `stops.txt`, `routes.txt`, `trips.txt`, and so on. Directories, hidden files (starting with `.`), files in subdirectories, and non-`.txt` files are all excluded. 3. **Sort the files alphabetically by name.** This neutralizes any variation in the order files appear within the zip. 4. **Stream each file's bytes sequentially into a single SHA1 hash.** The contents of `agency.txt`, then `calendar.txt`, then `calendar_dates.txt`, and so on, all fed into one running hash computation. The result is a 40-character hex string that represents the actual transit data, independent of how it was packaged. Re-zip the same GTFS files with different compression? Same Directory SHA1\. Download the feed again tomorrow from a server that regenerates the zip? Same Directory SHA1, as long as the CSV contents haven't changed. Notably, the algorithm does *not* reorder rows within files or normalize field ordering within CSVs. It operates on the raw byte contents of each file. This is a pragmatic choice: it avoids the complexity and potential bugs of CSV parsing at the checksum stage, while still achieving the goal of being packaging-independent. 💡 The transitland-lib CLI does also include a `diff` command that can do deeper, row-level comparisons between two feed versions, but that's a topic for another blog post. ## Try it yourself The `transitland` command-line interface includes a `checksum` command that computes both hashes for any GTFS feed on your local machine. Here's a real example using [MBTA's static feed](https://www.transit.land/feeds/f-drt-mbta/?ref=interline.io): ``` $ wget https://cdn.mbta.com/MBTA_GTFS.zip $ transitland checksum MBTA_GTFS.zip Zip SHA1 (archive file): 6d768217c1e441cc13f313595856669d4c24c013 Directory SHA1 (feed contents): 34e21ac222935b7acdd395e5f6b36dc996da0d60 Find via Transitland website: https://www.transit.land/feed-versions/6d768217c1e441cc13f313595856669d4c24c013 Find via Transitland REST API: https://transit.land/api/v2/rest/feed_versions/6d768217c1e441cc13f313595856669d4c24c013?apikey=YOUR_API_KEY ``` The CLI gives you everything you need to cross-reference the local file against Transitland's global archive. Click the website link and you'll land directly on [that feed version's page](https://www.transit.land/feed-versions/6d768217c1e441cc13f313595856669d4c24c013?ref=interline.io), where you can see when it was fetched, what date range it covers, and explore its stops, routes, and agencies. The CLI's output can also be compared against another zip file you have locally. This is useful in a few scenarios: - **Agency staff** can verify that Transitland has successfully fetched and archived their latest feed. Publish an update, wait for the next fetch cycle, then run `transitland checksum` on the same file and check if the Zip SHA1 appears in the archive. - **Developers** responsible for legacy systems can ensure their pipelines are only emitting new GTFS feeds when the Directory SHA1 changes from its previously cached value. - **Analysts** who download GTFS feeds directly from agencies can confirm they're working with the same version that Transitland has processed, which is helpful when comparing your analysis against Transitland's API results. - **Researchers** working with historical GTFS data can use the checksum to determine exactly which feed version in Transitland corresponds to a file they've been handed. They can then use Transitland's additional metadata to inform their research. ## Installing the CLI The `transitland` CLI is a single binary. You can: - download it prebuilt for Linux and macOS from the [transitland-lib releases page](https://github.com/interline-io/transitland-lib/releases?ref=interline.io) on GitHub - install it [using the Homebrew package manager](https://github.com/interline-io/homebrew-transitland-lib?ref=interline.io) - build it from Golang source if you prefer ## A small detail that makes a big difference Checksumming transit feeds might sound like plumbing. It's the kind of plumbing that makes a global transit data platform reliable. Without content-aware hashing, Transitland would either miss real updates (by relying on unreliable zip hashes) or drown in false positives (by treating every re-downloaded zip as new data). The `transitland checksum` command also puts this capability in your hands. Whether you're an agency verifying that your feed updates are being picked up, or an analyst reconciling local data against the Transitland archive, you can generate the same fingerprints that Transitland uses internally. If you have questions or want to explore further, check out the [Transitland documentation](https://www.transit.land/documentation/?ref=interline.io) or browse the [transitland-lib source code](https://github.com/interline-io/transitland-lib?ref=interline.io) on GitHub. ### Interline at MobilityData's 2025 Workshop URL: https://www.interline.io/blog/interline-at-mobilitydatas-2025-workshop/ Last updated: 2025-10-02T17:58:42.000Z We're excited to be participating in this year’s [MobilityData Workshop in Vancouver](https://mobilitydata.org/2025-vancouver-workshop/?ref=interline.io). Interline has been a dues-paying member of MobilityData since its founding, and our team contributed to the early meetings that led to the organization’s creation. We’re proud to continue supporting MobilityData’s mission to advance open data standards for mobility around the world. ## Conference Kickoff: Data for Seamless Mobility During Major Events Interline Principal **Drew Dara-Abrams** will present as part of the conference kickoff session on *“Using Data to Power Seamless Mobility During Major Events.”* His talk will highlight how Interline is helping transit agencies prepare for the Olympic games and other major events with **GTFS-Pathways data, GTFS-Fares v2 data,** and **coordinated transfer schedules**. These tools and enriched data allow riders to more easily navigate complex station environments and plan complete door-to-door journeys. ## Showcase: Real-World Success Stories Interline Principal **Ian Rees** will present a **technical update on our** [**Transitland platform and APIs for developers**](https://www.interline.io/transitland/apis-for-developers/). Transitland is an open-data platform built on thousands of public-transit data feeds from around the world. We started Transitland in 2014 and continue to expand the platform. It’s now the largest and most feature-rich GTFS, GTFS Realtime, and GBFS aggregator. Transitland is the ideal solution for using data from many transit operators in web or mobile apps, maps, data visualizations, GIS analyses, and travel demand models. ## Join Us in Vancouver We’re looking forward to connecting with colleagues across the mobility data ecosystem, learning from others, and sharing how Interline’s work is supporting seamless mobility worldwide. Full details on the workshop program are available on the [MobilityData website](https://mobilitydata.org/2025-vancouver-workshop/?ref=interline.io). ### Easily inspect GTFS Realtime using Transitland's website or API URL: https://www.interline.io/blog/easily-inspect-gtfs-realtime-using-transitlands-website-or-api/ Last updated: 2025-07-22T03:45:24.000Z GTFS Realtime feeds provide live updates about vehicle positions, trip updates, and service alerts, but they can be challenging to inspect due to their Protocol Buffer format. This binary [format](https://protobuf.dev/?ref=interline.io) is highly efficient for data transmission but requires special tools to decode. You can't just open it up in a text editor or a web browser. Transitland now offers two convenient ways to parse and view GTFS Realtime feeds: through our web interface and via our REST API. ## Inspecting GTFS Realtime on the Transitland Website The Transitland website now provides an intuitive interface for viewing GTFS Realtime data directly in your browser. Let's walk through how to use this feature using Portland TriMet's GTFS Realtime feed as an example. ### Finding a GTFS Realtime Feed 1. Navigate to [www.transit.land](https://www.transit.land/?ref=interline.io) and go to the "Source Feeds" section 2. Look for feeds with the GTFS Realtime format designation 3. Click on any GTFS Realtime feed to view its details ![browsing Transitland website for GTFS Realtime feeds](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2025/07/image-1.png) For this example, we'll use the TriMet GTFS Realtime feed at [https://www.transit.land/feeds/f-trimet\~rt](https://www.transit.land/feeds/f-trimet~rt?ref=interline.io) ### Viewing Real-time Data On the feed page, you'll see comprehensive information about the feed including: - **Feed identification**: Onestop ID and source URLs - **Last fetch time**: When the data was most recently updated (Transitland typically refreshes GTFS Realtime feeds once per minute) - **Authorization details**: Any required API keys or parameters (Transitland keeps its own credentials private; to access the source feed directly, you'll have to sign up for your own credentials) ⏱️ [Learn more](https://www.transit.land/documentation/concepts/source-feeds/?ref=interline.io#gtfs-realtime-feed-fetching-and-caching) about how Transitland caches GTFS Realtime feeds. The most exciting feature is the "GTFS Realtime Feed Messages" section, which provides three ways to inspect each type of real-time data: ![screenshot of https://www.transit.land/feeds/f-trimet~rt](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2025/07/CleanShot-2025-07-21-at-13.25.59.png) ### Interactive JSON Viewer When you click "View as JSON" for any real-time endpoint, a modal window opens showing: - **Last fetched timestamp**: When the data was fetched from the source feed - **Formatted JSON data**: Human-readable data structure. Click the caret arrows to open or close nested sections. - **Download options**: Buttons to download the entire JSON document or copy it to clipboard ![Interactive JSON viewer for TriMet GTFS Realtime vehicle positions feed](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2025/07/image.png) The JSON viewer displays the data in a structured format that's easy to read and understand, showing all the fields defined in the [GTFS Realtime specification](https://gtfs.org/documentation/realtime/reference/?ref=interline.io). ## Accessing GTFS Realtime via the Transitland API For programmatic access or integration into your applications, the Transitland REST API provides direct access to Transitland's cached GTFS Realtime data. ### API Endpoint Structure The API uses the following endpoint pattern: ``` GET https://transit.land/api/v2/rest/feeds/{feed_key}/download_latest_rt/{rt_type}.{format} ``` Where: - `{feed_key}` is the feed's Onestop ID (e.g., `f-trimet~rt`) - `{rt_type}` is one of: - `vehicle_positions` for vehicle location data - `trip_updates` for arrival/departure time changes - `alerts` for service disruptions - `{format}` is either `json` or `pb` (raw Protocol Buffers) ### Example API Call To access TriMet's cached GTFS Realtime service alerts through the Transitland REST API in JSON format: ```bash curl https://transit.land/api/v2/rest/feeds/f-trimet~rt/download_latest_rt/alerts.json?apikey=YOUR_API_KEY ``` ### API Response Format When requesting JSON format, the API returns data in the standard GTFS Realtime JSON structure. For example: ```json { "header": { "gtfsRealtimeVersion": "2.0", "incrementality": "FULL_DATASET", "timestamp": "1753153829" }, "entity": [ { "alert": { "activePeriod": [ { "start": "1675370449" } ], "descriptionText": { "translation": [ { "text": "No service to the eastbound stop at SW Tualatin-Sherwood Rd & Langer Farms (Stop ID 13837) due to Tualatin-Sherwood Rd project." } ] }, "headerText": { "translation": [ { "text": "" } ] }, "informedEntity": [ { "routeId": "97" }, { "stopId": "13837" } ], "url": { "translation": [ { "text": "https://trimet.org/alerts/" } ] } }, "id": "164849" } ``` ### Finding Available GTFS Realtime Feeds To discover available GTFS Realtime feeds, use the [feeds REST endpoint](https://www.transit.land/documentation/rest-api/feeds?ref=interline.io) with the `spec=gtfs-rt` parameter: ```bash curl https://transit.land/api/v2/rest/feeds?spec=gtfs-rt&apikey=YOUR_API_KEY ``` You can also [filter by other criteria](https://www.transit.land/documentation/rest-api/feeds?ref=interline.io) to find feeds relevant to your needs. ## Benefits of Using Transitland for GTFS Realtime Transitland's website and API provide simplified access: - **No Protocol Buffer tools required**: View data directly in JSON format - **Real-time updates**: Access the latest data as it becomes available - **Standardized interface**: Consistent API across all feeds Transitland also offers comprehensive coverage: - **Multiple feed types**: Vehicle positions, trip updates, and service alerts (when available) - **Global coverage**: Access feeds from transit agencies worldwide 🔍 Don't see a GTFS Realtime feed available through Transitland? Help [add it the Transitland Atlas feed registry](https://www.transit.land/documentation/atlas?ref=interline.io). Alternatively, transit agency staff are welcome to email us GTFS Realtime feed URLs at [hello@transit.land](mailto:hello@transit.land). You may also share auth credentials, which we will keep private. ❌ Finally, please note that some agencies do not allow Transitland to redistribute their feeds in full. In these cases, GTFS Realtime feed export is not allowed through either the Transitland API or website. ## Getting Started - **Sign up**: Inspecting GTFS Realtime feeds requires a [free Interline account](https://app.interline.io/?ref=interline.io) (which helps us to limit concurrent requests and ensure prompt responses). - **For web browsing**: Visit [www.transit.land](https://www.transit.land/?ref=interline.io) and explore the Source Feeds section. - **For API access**: Get an API key by following the [sign-up instructions](https://www.transit.land/documentation/index?ref=interline.io#signing-up-for-an-api-key). Whether you're an analyst learning your way around GTFS data or an agency staffer trying to quickly debug your own systems, Transitland provides the tools you need to automatically parse and read GTFS Realtime feeds. The combination of our web interface and REST API makes it simple to explore real-time transit data without the complexity of Protocol Buffer conversion tools. ### Announcing transitland-lib v1.0.0 URL: https://www.interline.io/blog/announcing-transitland-lib-v1-0-0/ Last updated: 2025-03-28T19:56:25.000Z We're pleased to announce the release of transitland-lib version 1.0.0\. This software library, written in the [Go programming language](https://go.dev/?ref=interline.io), is the foundation of the Transitland platform and integral to a growing range of purpose-built applications. ## From experiment to production In 2019, the Interline team began exploring how Go's performance characteristics could serve as a new foundation for [Transitland's v2 platform](https://www.interline.io/blog/tlv2/). The goal was simple: process the largest and most complex GTFS feeds faster and more efficiently. After years of iteration and real-world testing, with hundreds of thousands of GTFS feeds processed, we've finally tagged a v1.0.0\. This release represents a stable production-ready library that powers critical transit data systems. ## Real-world uses transitland-lib can be found within many transit data platforms operated by Interline: - The canonical Transitland deployment at [www.transit.land](https://www.transit.land/?ref=interline.io) - [MTC's 511 Regional Feed for the San Francisco Bay Area](https://www.interline.io/blog/tag/511-sf-bay-regional-feed/) - A new transit analysis tool for public agencies across the states of Washington, Oregon, and California - A routing engine deployment for a business client that covers an entire country - The new [Transitland Routing API](https://www.transit.land/documentation/routing-api/?ref=interline.io) beta that currently provides journey planning across the entire United States This component is also in use by other organizations. For example, Portland TriMet uses transitland-lib to efficiently clip [Amtrak's country-wide feed](https://www.transit.land/feeds/f-9-amtrak~amtrakcalifornia~amtrakcharteredvehicle?ref=interline.io) to create a focused feed to ingest into their own [OpenTripPlanner routing engine for the greater Portland, Oregon, region](https://trimet.org/home/planner/?ref=interline.io). ## Built for handling many feeds transitland-lib provides: - Fast feed processing for both static GTFS and GTFS Realtime - Memory-efficient operations - Concurrent data handling - Robust error handling - Defining multiple input feeds using the [Distributed Mobility Feed Registry (DMFR) format](https://github.com/transitland/distributed-mobility-feed-registry?ref=interline.io) ## Performance in action Let's look at one example of transitland-lib's capabilities to process and validate massive GTFS feeds. This demonstration covers Switzerland's entire transit network, one of the largest GTFS feeds in the world ([view on transit.land](https://www.transit.land/feeds/f-u0-switzerland?ref=interline.io)): ```sh ✗ transitland validate https://data.opentransportdata.swiss/en/dataset/timetable-2025-gtfs2020/permalink ``` ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2025/03/CleanShot-2025-03-26-at-13.52.03.gif) In about one and a half minutes on a laptop, transitland-lib downloads and processes a GTFS feed including 460 separate agencies, over 90,000 stops, over 4,000 routes, over 1.2 million trips, and a complicated set of calendar and calendar date records. The library's ability to handle such large datasets reliably makes it an ideal choice for processing national-scale transit feeds or creating regional subsets of larger networks. ## Extensible architecture transitland-lib's extension framework enables support for both standard GTFS and [experimental extensions](https://gtfs.org/community/extensions/overview/?ref=interline.io). This flexibility has been crucial for: - Processing [MTC GTFS+ files](https://www.transitwiki.org/TransitWiki/images/e/e7/GTFS%2B%5FAdditional%5FFiles%5FFormat%5FVer%5F1.7.pdf?ref=interline.io) for the SF Bay Regional Feed - Supporting [GTFS Fares-v2](https://www.interline.io/blog/mtc-regional-gtfs-feed-fares-updates/) as it has iteratively evolved from proposal to adopted specification - Supporting the addition of [GTFS Pathways](https://www.interline.io/blog/mtc-regional-gtfs-feed-additions/#station-pathways-and-levels) using [Interline's Station Editor](https://www.interline.io/transitland/station-editor/) We update transitland-lib to track the latest static GTFS and GTFS Realtime specifications. To check which version of the specifications you are using, simply run the `transitland version` command: ```sh ✗ transitland version transitland-lib version: v1.0.0 transitland-lib commit: https://github.com/interline-io/transitland-lib/commit/bae91cd7f32c67ffe89965cab3636e22ddbce817 (time: 2025-03-11T18:47:50Z) GTFS specification version: https://github.com/google/transit/blob/11a49075c1f50d0130b934833b7eeb3fe518961c/gtfs/spec/en/reference.md GTFS Realtime specification version: https://github.com/google/transit/blob/7b9f229dfa0b539c3fcf461986638890024feb06/gtfs-realtime/proto/gtfs-realtime.proto ``` ## Validation and best practices GTFS feeds can be messy. Therefore, transitland-lib: - Validates required fields and data types - Ensures referential integrity between files - Checks that feeds meet many of the [best practices](https://gtfs.org/documentation/schedule/schedule-best-practices/?ref=interline.io) - Also supports validating GTFS Realtime feeds With both static GTFS and GTFS Realtime processed using the same library, we're now able to put into ongoing production-scale use [Interline's previous research with University of South Florida into GTFS Realtime validation](https://www.interline.io/blog/trb-transit-idea-gtfs-realtime-report/). 💡 While transitland-lib implements many of the same validation and best practice rules as [MobilityData's static GTFS validator](https://github.com/MobilityData/gtfs-validator?ref=interline.io), the purposes and advantages of each library are distinct. The MobilityData validator lends itself well to interactive use when people and organizations creating GTFS feeds want to check to ensure their static feed is ready for submission to a third-party consumer. The MobilityData validator outputs a report that is useful for reference and sharing. In contrast, transitland-lib is designed to efficiently process many feeds, to serve as a component within production-scale processing pipelines, and to optionally validate GTFS Realtime endpoints with the context of a static GTFS feed with associated schedules. transitland-lib also gives developers options for how to handle invalid data: - For the canonical Transitland deployment, we've configured transitland-lib to fail *soft*. When errors are limited to certain rows in a GTFS feed, those entities are filtered out. The rest of the feed's data is imported and made available to users. - For some client deployments, we've configured transitland-lib to fail *hard*. Any error in a source feed will halt the entire workflow and instantly notify staff to review the issue. Each application may have different goals and different tradeoffs for how to handle invalid data, and transitland-lib provides options to customize and control this behavior. ## Powerful filtering system transitland-lib's filtering system enables modifying and transforming GTFS feeds for specific project contexts, such as: - Producing date-specific GTFS feeds - Building regional routing engines with optimized data - Merging together many source feeds to create regional, state, or national scale feeds - Standardizing data across multiple agencies, with customized rules for namespacing entities Interline uses transitland-lib to [curate full-service data deliveries for our clients needing consistent GTFS data across states and countries](https://app.interline.io/contact%5Fforms/curated%5Fgtfs?ref=interline.io). ## Using transitland-lib If you use the Transitland website or the Transitland APIs, you're already using transitland-lib. Developers can visit [github.com/interline-io/transitland-lib](https://github.com/interline-io/transitland-lib?ref=interline.io) to learn more about transitland-lib as a command-line interface (CLI) or a library for programmatic usage. transitland-lib is available under a [dual license model](https://github.com/interline-io/transitland-lib/blob/main/LICENSE?ref=interline.io) enabling open-source use under the permissive GPLv3 license and also providing flexibility to Interline's clients under a customizable business license. We look forward to continuing to maintain transitland-lib for another six years as the GTFS and GTFS Realtime specifications continue to evolve and as the number and size of source feeds continues to grow. ### Transitland for nighttime URL: https://www.interline.io/blog/transitland-for-nighttime/ Last updated: 2025-02-25T22:23:06.000Z Our engineering team has recently been refactoring some of the internals of the Transitland website. The main goal has been to update various components and dependencies, to better support Interline's consulting clients who also use these components. As a side effect: the canonical Transitland website now has an optional dark mode. Go to [www.transit.land](https://www.transit.land/?ref=interline.io) and the website's styling will now respect your device's settings. So if your smartphone or laptop automatically switches to dark mode in the evening, when you access Transitland, you'll see a similarly dark background. Or tap the little sun or moon icon in the top bar to manually switch the appearance. We hope you enjoy Transitland — whatever time of day you're using it. ### Built with Transitland: Digital departures screen URL: https://www.interline.io/blog/built-with-transitland-instructables/ Last updated: 2025-01-10T17:42:13.000Z Want to build your own digital display to show when a bus, train, or ferry next departs? Here's a fun how-to guide from Instructables user @mpreston21: > In this project, I iterated on the concept of a subway clock and created a version that displays the current schedule of the NYC ferry— for those who ride in style! > > This product was created for my boyfriend who takes the ferry from the same stop every day to get to work. The design incorporates a perpetual scroll display that shares the next 3 departure times from his stop (North Williamsburg). He can glance over at the ferry times in the morning and determine which ferry to catch, and whether to walk, bike, or skateboard to catch his ferry. Using this clock reduces the likelihood that he will be distracted when looking up the schedule on the NYC Ferry's mobile app. It's a great way to start the day for anyone who relies on the ferry schedule with regularity :) Their instructions guide you through the process of creating all the relevant electronics: ![picture and list of supplies needed for the project, including an Arduino board, an LED RGB matrix, and wires](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2025/01/CleanShot-2025-01-10-at-09.36.59@2x.png) When it's time to program, the instructions describe how to query the [Transitland v2 REST API stop departures endpoint](https://www.transit.land/documentation/rest-api/departures?ref=interline.io): > In order to work with GTFS directly, you need to spin up a server to manage the rate/volume of data when you make your query. > > \[...\] > > after digging around on the internet, I found **Transit.Land**\-- an open data platform that collects GTFS and other open data feeds from transit providers around the world. Transitland makes all of this data queryable via **REST APIs**. There are feeds from [over 2,500 operators in over 55 countries](https://www.transit.land/operators?ref=interline.io), including the NYC ferry! > > I read through the Transit.Land documentation, and started to decode the the key/value pair data for [ferry stops](https://www.transit.land/operators/o-dr5r-nycferry?ref=interline.io#stops) and [ferry routes](https://www.transit.land/operators/o-dr5r-nycferry?ref=interline.io#routes) from the source feeds. I created an account on Interline.co to get a **secure API key**. Armed with my API key, I started to get to work on the program for my ESP32. Thanks to @mpreston21 for sharing their creation. See the [full set of instructions, code, and illustrations on the Instructables website](https://www.instructables.com/NYC-Ferry-Schedule-Ticker-Clock/?ref=interline.io). 💡 ****Want to share your own project you've created using Transitland APIs or datasets?** Please post to the ["show and tell" section](https://github.com/transitland/transitland/discussions/categories/show-and-tell?ref=interline.io) of the Transitland discussion board on GitHub. ### Interline at TRB AM 2025 URL: https://www.interline.io/blog/interline-at-trb-am-2025/ Last updated: 2025-01-03T18:37:33.000Z We'll be represented at the [US Transportation Research Board Annual Meeting](https://trb-annual-meeting.nationalacademies.org/?ref=interline.io) in Washington, D.C., by Interline principal Drew Dara-Abrams. He'll be presenting on Sunday, January 5, 2025 as part of the workshop on [Open-Source Data Repositories and Technologies](https://annualmeeting.mytrb.org/OnlineProgram/Details/22540?ref=interline.io): ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2025/01/CleanShot-2025-01-03-at-10.26.48@2x.png) All conference attendees are welcome to attend (not just members of the organizing committee on Statewide/National Transportation Data and Information Management). ### US National Transit Database releases data and requests more feedback URL: https://www.interline.io/blog/us-national-transit-database-releases-data-and-requests-more-feedback-2/ Last updated: 2025-01-06T20:02:03.000Z ****January 3, 2025**: FTA has extended the comment period to January 29 regarding [National Transit Database: Proposed Reporting Changes and Clarifications for Report Years 2025 and 2026](https://www.regulations.gov/document/FTA-2024-0013-0001?ref=interline.io). 📧 We've submitted the following public comment to the Federal Transit Administration regarding [National Transit Database: Proposed Reporting Changes and Clarifications for Report Years 2025 and 2026](https://www.regulations.gov/document/FTA-2024-0013-0001?ref=interline.io). 📖 For background, see earlier posts on the Interline blog: "[US Federal Transit Administration finalizes plan to collect GTFS URLs in US National Transit Database](https://www.interline.io/blog/us-ntd-reporting-gtfs-adopted/)" (2023) and "[US National Transit Database to collect GTFS URLs](https://www.interline.io/blog/us-ntd-reporting-gtfs/)" (2022). Interline Technologies is the maintainer of the [Transitland open data platform](https://www.transit.land/?ref=interline.io). Transitland aggregates approximately 1,000 static GTFS feeds from across the United States (as well as real-time feeds, plus feeds from international sources). To power the platform's APIs, our staff and external contributors maintain the [Transitland Atlas feed registry repository](https://github.com/transitland/transitland-atlas?ref=interline.io), which is publicly hosted on GitHub under a permissive license. This feed registry maintains lists of public GTFS feed URLs in addition to associated metadata (including US NTD IDs). The following comments regarding the proposed NTD reporting changes are based on our experience creating and operating these open-data systems since 2015. ## **Re Section B. Additional Data Within Publicly Hosted General Transit Feed Specification (GTFS) Datasets** Thank you to agencies and NTD for beginning to collect GTFS feeds from reporting agencies. The [2023 Annual Database General Transit Feed Specification (GTFS) Weblinks dataset](https://www.transit.dot.gov/ntd/data-product/2023-annual-database-general-transit-feed-specification-gtfs-weblinks?ref=interline.io) is a useful public release. 🙋 Transitland Atlas maintainers and contributors are reviewing the NTD weblinks dataset to look for new feed sources to add to the Transitland Atlas, as well as out-of-date references to update in Transitland Atlas. To help, see the [ntd-gtfs-weblinks projects subdirectory](https://github.com/transitland/transitland-atlas/tree/main/scripts/ntd-gtfs-weblinks?ref=interline.io). 🚌 ****For transit agencies currently unable to host their GTFS feed at a public URL**: We invite you to email a copy of your latest feed to our team at [hello@transit.land](mailto:hello@transit.land) To further increase the use of this weblinks dataset in future years, please guide as many agencies as possible to host their GTFS feeds at URLs that are public and stable. That is: - URLs - hosting their GTFS feed on an agency-controlled website (instead of submitting feed version as archives via email) - public - a URL that third-parties can download from (rather than a private submission to NTD) - stable - a URL that remains the same, while the GTFS feed archive itself is changed as the agency releases new schedules/versions (i.e., no dates in the file name) Finally, please keep in mind that it's best for agencies to be reporting to NTD the exact same GTFS feed versions that they share with third-party navigation apps and open-data aggregation platforms. This ensures that all data consumers are working from similar information, and makes it easier for data producers to focus their efforts. This is even more relevant for agencies that produce GTFS Realtime feeds as well (which need to match certain identifiers within associated static feeds). Our feedback regarding adding NTD IDs is also based on this goal of preventing agencies from having to produce separate feed files for NTD reporting purposes than they already produce for trip-planning and operational purposes. ## **Re Align `agency_id` and NTD ID Within GTFS File** We support the proposed goal of adding NTD IDs to GTFS feeds. NTD IDs are a useful "crosswalk" between the rider-facing data in a GTFS feed and the operational/planning/financial datasets in the NTD. However, we recommend that the means of adding NTD IDs to GTFS feeds be carefully considered. Interline strongly agrees with related feedback submitted from other commenters, including [MBTA](https://www.regulations.gov/comment/FTA-2024-0013-0003?ref=interline.io) and the [MobilityData](https://www.regulations.gov/comment/FTA-2024-0013-0006?ref=interline.io) non-profit (of which Interline is also a member). We recommend: 1. **NTD IDs should be added to GTFS feeds in a flexible manner that can map on to an entire feed, an agency within a feed, or a subset of routes within a feed.** This is necessary because rider-facing "brand names" of transit agencies do not always line up 1-to-1 with NTD reporting IDs, nor do NTD reporting IDs necessarily line up 1-to-1 with the internal means of organizing an agencies' data systems that produce records in agencies.txt 2. **NTD IDs be added to GTFS feeds in a way that does not conflict with existing data-producing systems.** This is toward the goal of agencies reporting their existing GTFS feeds to NTD, rather than creating "one off" feed versions that they customize and submit once a year to NTD. Based on our experience working with a wide range of transit agencies, we recommend that data-producers have the choice to either add an entirely new file or add columns to existing files. For some agencies, it may be simpler to add a single hard-coded file to their feed, while for other agencies, it may be simpler to add an extra custom column to the relevant files in their feed. 3. **NTD IDs be added to GTFS feeds in a way that does not conflict with existing data-consuming systems.** Consumers can choose to accept or ignore new columns and/or new files in GTFS feeds. This makes it simpler for them to "opt in" rather than to have to adapt to a change in the meaning of agency\_id. Our specific recommendation for your consideration: - Give reporters two approaches for adding NTD IDs to their feeds. - The first approach will be to add a file named `us-ntd-ids.txt` (or similar) to their feeds. This would be a CSV file in which each row maps an NTD ID to a chosen entity (e.g., an agency or a route). For a simple situation in which there is one NTD ID for a single agent in a feed, this file would only have one row (after the header row); for a complex situation, there may be multiple rows defining how different agencies and/or routes are mapped to different NTD IDs. - The second approach will be to add a `us_ntd_id` column (or similar) to existing `feed_info.txt`, `agencies.txt`, and/or `routes.txt` files. Data producers could add this column to just the files in which they want to list one or more NTD IDs. Our overall recommendation is to provide flexibility to data-producers in where and how they add US NTD IDs. This will lower the uncertainty of changing existing fields (such as the meaning of the `agency_id`). This will also decrease the technical trade-offs for agencies with complicated existing IT systems, which may limit how they can change or customize their GTFS feeds. This proposal does slightly increase the burden of data-consumers to extract the relevant NTD IDs from GTFS feeds and figure out how to apply them to feed contents. We believe that is a reasonable burden for data-consumers to take on. Transitland already does similar operations across the thousands of GTFS feeds it consumes. In exchange, the benefit of this approach is that it encourages agencies to fit NTD IDs into their production systems (rather than producing one-off feed versions just for NTD reporting). Finally, we recommend reviewing uptake of these two approaches by agencies after a year or two. Perhaps one approach will be fine and the requirement can be refined in the future. Still, we recommend beginning with reporting agencies having more initial flexibility. ## **Re `shapes.txt` File (Geospatial Drawing of Routes) as Part of GTFS Submission** We agree that recommending agencies add `shapes.txt` files for all routes/trips will be beneficial to data consumers, especially trip-planners, GIS analyses, and visualizations. However, in our experience, the quality of `shapes.txt` geometries can vary wildly. Imprecise and inaccurate shapes can often be worse than no shapes. Data consumers such as Transitland and trip planning apps already have to run automated checks to see whether our systems even want to ingest `shapes.txt` records for a given trip or to instead ignore them. The [shapes guidance webpage](https://gtfs.org/documentation/schedule/examples/shapes/?ref=interline.io) shared by MobilityData is worth sharing with agencies. It's also worth sharing with agencies and reiterating the value of the [GTFS "best practices."](https://gtfs.org/documentation/schedule/schedule-best-practices/?ref=interline.io) Thank you for accepting this feedback. ### How current is that schedule? URL: https://www.interline.io/blog/how-current-is-that-schedule/ Last updated: 2024-11-25T19:41:26.000Z 💡 This blog post is based on a presentation by Ian Rees and a panel discussion along with representatives from Google, Transit app, and MobilityData regarding "**How to Make Sure Transit Riders Get New Schedules from GTFS Static Datasets on Time?*" at the [2024 International Mobility Data Summit](https://mobilitydata.org/the-2024-international-mobility-data-summit-new/?ref=interline.io). [Transitland](https://www.interline.io/transitland/) monitors schedule data for thousands of public transit operators. This data is collected and archived — currently at a pace of tens of thousands of updates per year: ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/unnamed.png) The number of [feed versions](https://www.transit.land/documentation/concepts/static-gtfs-feed-versions/?ref=interline.io) in the Transitland archive continues to grow year by year. We rely on operators to set their own policies for updates to their schedule data, and the time period covered by that schedule data. Operators may publish data that contains scheduled service for the next week, the next month, or even the next year. In some instances, operators publish data that only contains service that begins at a future date. This is a critical, and often hidden, factor in how useful the data is to riders and other data consumers. ## The "7-day rule" Data that changes frequently or data that only contains schedules for the next few days struggles against an important constraint: Data consumers might require many hours or days before data updates can be processed and available. (For example, to rebuild a routing graph used in trip-planning software.) These constraints mean service changes may be in effect, on the ground, before these changes can be communicated back to riders, via apps and websites. Schedules that contain less than a week of future scheduled service can also force consumers to make (often faulty) assumptions about trips on future dates. Currently, the [GTFS Best Practices](https://gtfs.org/documentation/schedule/schedule-best-practices/?ref=interline.io) guide recommends publishing data for the current data plus at least 7 days in the future; we'll call this the "7-day rule." ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/unnamed--1--1.png) The percentage of feed versions in Transitland's archive by year (on the x-asis) that have 7 days, 30 days, or 90 days of forward-looking schedule coverage as of when they were fetched by Transitland. Fortunately, this is a mostly "good news" situation. ## Using Transitland's feed archive to assess update frequency Averaged across all transit operators in Transitland, about 95% of static schedule updates follow the 7-day rule. This number remains about 90% even when expanding the future service window to 30 days. It's also a good picture when disaggregating and looking at the individual operator level: The vast majority of operators have 100% of schedule updates following the 7-day rule. The most common exceptions are: 1. operators that have a tendency to publish data that begins in the future, with no scheduled service on the date of publication 2. operators that fail to update the values in `feed_info.txt` which explicitly set the range of valid service dates and remove guesswork and ambiguity for the data consumer. Even if these exceptions apply to comparatively few feeds, it's still unfortunate if they occur in feeds that are important to a data consumer's use-case (e.g., to provide up-to-date trip-planning for travelers on an important transit operator). ## Even more carefully characterizing updates This provides an optimistic baseline for the current state of GTFS updates. For an even more holistic picture, we also need to incorporate a few additional metrics: - Percentage of days for each operator where the data is "stale", that is, outside the scheduled service window for the last published data. - The number of operators that publish data on a very frequent basis that assumes consumers have a fast turn-around process, and the characteristics of these updates. These questions will be explored in a future blog post. 💡 ****The take-away from this blog post for transit data producers**: Please check to ensure that your GTFS feed always includes service for the current point-in-time when it is published, and include at least 7 days of forward-looking schedule coverage (if not 30 or 90 days). ### Giving equal prominence to buses and trains on the map URL: https://www.interline.io/blog/giving-equal-prominence-to-buses-and-trains-on-the-map/ Last updated: 2024-11-25T19:37:00.000Z Transitland’s global map of transit has an ambitious goal: to map buses, trains, subways, cable cars, ferries, and other modes of public transit across all urban and rural regions where we can source [standardized open-data feeds](https://www.transit.land/documentation/concepts/source-feeds/?ref=interline.io). While the data feeds may be standardized, every route/trip/stop record also represents the manifestation of local history, politics, and economics — factors that can vary widely from country to country, or metro region to metro region. When we design our map of transit to be consistently usable and understandable across the world, we also touch on questions unique to each place. Earlier this year, **Jeffrey Tumlin,** [director of the San Francisco Municipal Transportation Agency](https://www.sfmta.com/people/jeffrey-tumlin?ref=interline.io), took the time to discuss with our team some of the complicated questions raised by Transitland’s global map (and by similar transit navigation apps): - *Why are intercity trains, which only serve their riders a few times a day, often displayed more prominently than bus routes that actually provide much more frequent service?* - *When transit agencies offer a network of high-frequency bus service, how can riders identify this backbone to reach the nearest “entry point” and to plan transfers to most efficiently reach their ultimate destination?* - *Are colors a meaningful way to brand rail lines or high-frequency bus lines in a regional context, particularly if multiple overlapping agencies end up selecting the same colors for different lines?* The discussion was spirited: Tumlin shared how his current experience leading a major US transit agency and his previous experience consulting internationally highlights how much of an outlier the US is in terms of privileging rail transit over bus transit. Legacies of racism continue to shape the transit networks of today. Whether transit service is provided to riders by a bus with rubber wheels or a train on steel tracks still resonates in an almost cultural manner, aside from the practicalities of whether the route will help riders reach their destinations. The discussion was also optimistic: SFMTA is strengthening its network of bus routes with service of at least every 10 minutes. As frequencies increase further, riders are finding it even more convenient to quickly transfer from one high-frequency route to another. As a square of roughly 7 miles by 7 miles, San Francisco is now covered by a grid of high-frequency bus routes, enabling riders to travel “diagonally” with a single transfer with minimal wait time. SFMTA may stand out in the Bay Area in this respect, but it’s using patterns long common to other parts of the world and increasingly adopted in the dense cores of more American cities. 💡 Interline is currently engaged by the SF Bay Area's Metropolitan Transportation Commission to build a regional mapping data services platform to support transit cartography throughout the entire Bay Area. This platform will help to support MTC's [Regional Mapping and Wayfinding project](https://mtc.ca.gov/operations/transit-regional-network-management/regional-mapping-wayfinding?ref=interline.io), of which SFMTA is also a key stakeholder. All transit riders and residents of the greater Bay Area stand to gain from this initiative to design consistent maps and signage, as well as [related initiatives to better integrate transit service](https://mtc.ca.gov/operations/transit-regional-network-management/transit-fare-coordination-integration?ref=interline.io). Based on input from Tumlin, as well as from **SFMTA transit planner** [**Steve Boland**](http://calurbanist.com/?ref=interline.io), we’ve tuned the Transitland global map. ## Rail routes are now zoom-level dependent Rail lines continue to be more prominent at national or region-wide views. As you zoom into a city or neighborhood scale, rail lines now shrink to be equally sized with bus lines: ![screenshots of Transitland global transit map before and after changes to rail styling](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/Refinements-to-rail-line-width.png) This change reflects how rail lines are helpful for riders orienting across broader viewers and relevant for traveling longer distances, but for intra-urban travel, a bus may often be competitive if not better than a rail-based transit service. ## Bus routes are grouped by styling Previously, our map classified buses routes into two groups by their frequency: less frequent routes were styled with a lighter blue color, more frequent routes were styled with a darker and more prominent blue. Now our map groups buses into three different groups by their *headway* (the term that transit planners use to refer to how frequently a transit service operates): ![screenshots of Transitland global transit map before and after changes to bus frequency styling](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/Bus-frequency.png) In cities and metros with a network of high-frequency service, this change will help to emphasize that subset of fast and dependable service, while still also showing connections with less frequent bus services that provide coverage. ## Onward We’re working on some longer term improvements to the way bus and rail route shapes are handled within the Transitland software stack, and we look forward to sharing more in the future. It’s also always a pleasure to place the work we do with data and software within the broader context of public transit in the US. Thanks to SFMTA staff for a lively conversation and useful feedback! ### Turning off the Transitland v1 Datastore API URL: https://www.interline.io/blog/turning-off-the-transitland-v1-datastore-api/ Last updated: 2024-11-05T21:59:51.000Z 💡 ****October, 31, 2023**: We've now fully shut down the servers that respond to queries at `https://transit.land/api/v1`. Thank you to the many users who have migrated to using v2 APIs. For the handful who have yet to make the transition, we hope you'll read the following to learn more about the benefits of Transitland v2 APIs. We announced the [multi-step deprecation of the Transitland v1 API in a blog post in 2019](https://www.interline.io/blog/tlv2#gradual-deprecation-of-tlv1). In the intervening years, we've been greatly expanding the functionality, speed, and data coverage of the new v2 APIs. We've been pointing existing and new Transitland developers to the v2 APIs — particularly the v2 REST API — for a few years. The time has now come the Interline team to finally turn off the v1 Datastore API. On **October 31, 2023** we are planning to shut down the systems that handle queries to URLs to all endpoints that begin in `https://transit.land/api/v1` If you're in the subset of Transitland developers who continue to use the v1 APIs, please switch as soon as possible. We've emailed everyone with an API key that has queried the v1 APIs in the last 30 days to share this information. Most applications that used the v1 APIs should find the v2 REST API to be similar or better. The v2 REST API mixes together information from both static GTFS and GTFS Realtime feeds; it's highly performant; and we're about to announce some exciting new functionality that will only be available through the v2 endpoints. The same Transitland v1 API key will work with the v2 REST API. For more information on how to migrate an application or script to the v2 REST API: - See the [Transitland v2 REST API documentation](https://www.transit.land/documentation/rest-api/?ref=interline.io). - Post questions to the [Transitland discussions forum on GitHub](https://github.com/transitland/transitland/discussions?ref=interline.io) to be answered by Interline staff and fellow developers as time permits. - Sign up for a paid [Transitland Professional or Enterprise subscription](https://www.interline.io/transitland/plans-pricing/) to receive support via email by Interline staff. These plans also include the Transitland v2 GraphQL API, with even more functionality and flexibility. We appreciate having so many developers who've used the v1 APIs since 2015 to power such a wide range of applications and use-cases, and we look forward to fully focusing our efforts on further expanding v2 functionality and continuing its solid record of uptime and performance. Thanks to the developer-users been part of the Transitland platform over many years! ### US Federal Transit Administration finalizes plan to collect GTFS URLs in US National Transit Database URL: https://www.interline.io/blog/us-ntd-reporting-gtfs-adopted/ Last updated: 2024-11-05T00:09:55.000Z The US Bipartisan Infrastructure Law passed by Congress and signed into law by President Biden in 2021 was big, very big. While most people have focused on the large sums of funding that are now becoming available to improve American transportation infrastructure, we’ve focused on one of the details: The “BIL” requires the US Federal Transit Administration to collect data on American transit agencies’ “geographic service area coverage.” ![first page of FTA's notice and request for comments regarding additions to the NTD reporting process](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/bil-geographic-service-coverage-area-1.png) The US Bipartian Infrastructure Law, formally known as the [Infrastructure Investment and Jobs Act](https://www.congress.gov/bill/117th-congress/house-bill/3684/text?ref=interline.io), added a requirement that the US National Transit Database collect data on "geographic service area coverage" for American transit agencies receiving federal funds. ## Proposal To meet this requirement, the FTA issued a proposal on July 7, 2022 to collect static GTFS feeds from all American transit agencies that receive federal funds and operate fixed-route transit service. The [Interline team submitted feedback to FTA](https://www.interline.io/blog/us-ntd-reporting-gtfs) based on our experience operating the Transitland open transit data platform, which aggregates GTFS feeds from nearly a thousand transit operators in the US (and over 3,000 when including other countries). ## Adoption On March 3, 2023, the FTA released responses to feedback and formally adopted its proposal, with no changes made to the GTFS reporting requirement. For more information see [this notice in the US Federal Register](https://www.federalregister.gov/documents/2023/03/03/2023-04379/national-transit-database-reporting-changes-and-clarifications?ref=interline.io). ## Timing While the NTD reporting requirements have now been formally changed, agencies will have some time until they begin submitting their GTFS feeds. In response to a question from Interline about timing, the NTD team tells us that: > NTD will begin collecting GTFS data in Report Year 2023\. Agencies will submit 2023 data to the NTD beginning in October of 2023, and the associated RY23 data products containing GTFS data will be released in Fall of 2024. ## Data release When US National Transit Database Report Year 2023 data products are released, they will be available on the [NTD Data](https://www.transit.dot.gov/ntd/ntd-data?ref=interline.io) website. ## Improvements big and small The word that’s been most often said in recent years about the BIL is that it’s a “generational” amount of funding and policy change for American physical infrastructure and transportation networks. The collection and dissemination of URLs for GTFS feeds by the US Federal Transit Administration may not be as massive of a change in dollar amounts. Yet, this change still reflects a substantive development in the way that the US federal government engages with GTFS. We’re looking forward to seeing how FTA and other parts of the US Department of Transportation continue to promote and engage with transit data. ### Follow the Interline blog using RSS URL: https://www.interline.io/blog/interline-blog-rss-feed/ Last updated: 2024-11-12T22:10:09.000Z 💡 ****November 2024**: We've updated our RSS URL. For the latest updates from Interline, you can now follow our blog using an RSS feed reader. Just point your RSS feed reader to: Our [Transitland News & Updates](https://www.transit.land/news/?ref=interline.io) blog is also available by RSS feed at: [https://www.transit.land/feed.xml](https://www.transit.land/feed.xml?ref=interline.io) You may already be have the software for following RSS feeds without realizing it. For example, [Microsoft Outlook allows users to follow RSS feeds](https://support.microsoft.com/en-us/office/subscribe-to-an-rss-feed-73c6e717-7815-4594-98e5-81fa369e951c?ref=interline.io). Wikipedia [lists many more feed aggregators](https://en.wikipedia.org/wiki/Comparison%5Fof%5Ffeed%5Faggregators?ref=interline.io) that can be used to follow and read RSS feeds. ### Updated GTFS-Fares v2 for the SF Bay Area URL: https://www.interline.io/blog/mtc-regional-gtfs-feed-fares-updates/ Last updated: 2024-11-05T21:49:32.000Z 💡 ****April 20, 2023**: The GTFS-Fares v2 extension originally named `fare_containers.txt` has been renamed to `fare_media.txt`, as part its recent [adoption in the GTFS specification](https://github.com/google/transit/pull/355?ref=interline.io). In 2020, Interline and the Metropolitan Transportation Commission [released](https://www.interline.io/blog/mtc-regional-gtfs-feed-additions/#fares-and-transfer-discounts) a beta version of fares and transfer discounts for eight SF Bay Area transit agencies. Since then, we’ve been working together to improve data coverage and its schema. As of this August, the Regional GTFS Feed now includes fares and transfer discounts for over 30 Bay Area transit agencies. The updated dataset uses the adopted GTFS-Fares v2 “base implementation,” as well a few proposed additions to the GTFS-Fares v2 data schema. This data is now available for use, and our team is updating it on an ongoing basis. In this blog post we’ll cover: - [GTFS-Fares v2 in action](#gtfs-fares-v2-in-action) - [GTFS-Fares v2 base implementation and future additions to the specification](#gtfs-fares-v2-base-implementation-and-future-additions-to-the-specification) - [SF Bay Area transit agencies with GTFS-Fares v2 data](#sf-bay-area-transit-agencies-with-gtfs-fares-v2-data) - [Using GTFS-Fares v2 data for the Bay Area](#using-gtfs-fares-v2-data-for-the-bay-area) ## GTFS-Fares v2 in action If you have an Apple iPhone running [iOS 16](https://www.apple.com/ios/ios-16/?ref=interline.io), you can see the estimated costs of transit journeys. This data is sourced from the the Regional Feed. ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/ios-directions-with-fares-1.png) ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/ios-directions-with-fares-2-1.png) ![](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/ios-directions-with-fares-3.png) All of the above fares information is defined in the Regional GTFS Feed using the GTFS-Fares v2 specification. Here are extracts of the relevant GTFS files from the feeds for the itinerary displayed on the iPhone: ### fare\_products.txt | fare\_product\_id | fare\_product\_name | amount | currency | rider\_category\_id | fare\_media\_id | | ------------------- | --------------------- | ------ | -------- | ------------------- | --------------- | | AC:local:single | AC Transit local fare | 2.25 | USD | adult | clipper | | BA:matrix:FTVL-EMBR | generated | 2.15 | USD | adult | clipper | ### fare\_media.txt Updated | fare\_media\_id | fare\_media\_name | fare\_media\_type | | --------------- | ----------------- | ----------------- | | clipper | Clipper | 2 | ### fare\_leg\_rules.txt | leg\_group\_id | from\_area\_id | to\_area\_id | network\_id | fare\_product\_id | | -------------- | -------------- | ------------ | ----------- | ------------------- | | AC | AC:local | AC:local | AC | AC:local:single | | BA | FTVL | EMBR | BA | BA:matrix:FTVL-EMBR | ### stop\_areas.txt | area\_id | stop\_id | | -------- | -------- | | AC:local | 55227 | | AC:local | 55550 | | FTVL | FTVL | | EMBR | EMBR | Note that this data schema has changed since the examples in [our first release blog post](https://www.interline.io/blog/mtc-regional-gtfs-feed-additions/#an-example-of-fares-v2-in-practice). This new schema was adopted as part of the GTFS-Fares v2 “base implementation,” which we will now discuss in more detail. ## GTFS-Fares v2 base implementation and future additions to the specification Since the [original GTFS-Fares v2 proposal](https://bit.ly/gtfs-fares?ref=interline.io), many stakeholders in GTFS have provided additional input. In May, stakeholders voted to adopt the core of the refined GTFS-Fares v2 proposal as a [“base implementation”](https://github.com/google/transit/pull/286?ref=interline.io). The “base implementation” simplifies the organization of `fare_leg_rules.txt`, `fare_transfer_rules.txt`, and `fare_products.txt`. This is a breaking change, so consumers who implemented their GTFS processing software against the original proposal will need to update their tools. To fully capture the complexity of Bay Area transit fares and transfer discounts requires additional data tables and fields that were included in the original proposal but not in the “base implementation.” Here is an overview of all the GTFS-Fares v2 tables and columns/fields included in the Regional GTFS Feed: | GTFS file | Spec version(s) | Columns added for extensions | Documentation | | ------------------------- | --------------------------------------------- | ---------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- | | fare\_products.txt | "base" with additional columns for extensions | rider\_category\_id | [See gtfs.org reference](https://gtfs.org/schedule/reference/?ref=interline.io#fare%5Fproductstxt) | | fare\_leg\_rules.txt | "base" with additional columns for extensions | transfer\_only\* | [See gtfs.org reference](https://gtfs.org/schedule/reference/?ref=interline.io#fare%5Fleg%5Frulestxt) | | fare\_transfer\_rules.txt | "base" with additional columns for extensions | filter\_fare\_product\_id\* | [See gtfs.org reference](https://gtfs.org/schedule/reference/?ref=interline.io#fare%5Ftransfer%5Frulestxt) | | areas.txt | "base" | | [See gtfs.org reference](https://gtfs.org/schedule/reference/?ref=interline.io#areastxt) | | stop\_areas.txt | "base" | | [See gtfs.org reference](https://gtfs.org/schedule/reference/?ref=interline.io#stop%5Fareastxt) | | fare\_media.txt | "base" as of March 14, 2023 Updated | | [See gtfs.org reference](https://gtfs.org/schedule/reference/?ref=interline.io#fare%5Fmediatxt) | | rider\_categories.txt | extension | rider\_category\_id rider\_category\_name min\_age max\_age eligibility\_url | [See proposal Google Doc](https://docs.google.com/document/d/19j-f-wZ5C%5FkYXmkLBye1g42U-kvfSVgYLkkG5oyBauY/edit?ref=interline.io#heading=h.9gqyl4w7f6x1) | | routes.txt | extension | as\_route | [See proposal Google Doc](https://docs.google.com/document/d/19j-f-wZ5C%5FkYXmkLBye1g42U-kvfSVgYLkkG5oyBauY/edit?ref=interline.io#heading=h.v160xvt5oen5) | \* `fare_leg_rules.txt transfer_only` and `fare_transfer_rules.txt filter_fare_product_id` are two fields proposed for describing partial credit on inter-agency transfers. Thanks to the [Transit](https://transitapp.com/?ref=interline.io) app team for proposing these additions to make some of the transfer discounts in the Regional GTFS Feed more explicit for consuming software to parse. Interline and MTC will continue to produce these “extensions” to the “base” data schema, and we’ll continue to work with data consumers, [MobilityData](https://mobilitydata.org/?ref=interline.io), and other stakeholders to finalize their adoption into the GTFS spec. ## SF Bay Area transit agencies with GTFS-Fares v2 data Interline and MTC now produce fares, intra-agency transfer discounts, and inter-agency transfer discounts for the following agencies: | 511 operator code | Operator full name | | ----------------- | ----------------------------------------------------- | | 3D | Tri Delta Transit | | AC | AC Transit (including transbay routes) | | AF | Angel Island Tiburon Ferry | | AM | Capitol Corridor | | BA | Bay Area Rapid Transit (BART) | | CC | County Connection | | CE | Altamont Corridor Express (ACE) | | CM | Commute.org Shuttles | | CT | Caltrain | | DE | Dumbarton Express Consortium | | EM | Emery Go-Round | | FS | Fairfield and Suisun Transit | | GF | Golden Gate Ferry | | GG | Golden Gate Transit | | MA | Marin Transit | | MV | MVgo Mountain View | | PE | Petaluma Transit | | RV | Rio Vista Delta Breeze | | SA | Sonoma Marin Area Rail Transit | | SB | San Francisco Bay Ferry | | SC | Valley Transportation Authority (VTA) | | SF | San Francisco Municipal Transportation Agency (SFMTA) | | SI | San Francisco International Airport (SFO) | | SM | SamTrans | | SO | Sonoma County Transit | | SR | Santa Rosa CityBus | | SS | City of South San Francisco | | ST | SolTrans | | TD | Tideline Water Taxi | | UC | Union City Transit | | VC | Vacaville City Coach | | VN | VINE Transit | | WC | Western Contra Costa (WestCat) | | WH | Livermore Amador Valley Transit Authority (LAVTA) | Fares are included for Clipper Card, cash, and a wide variety of passes. ([Clipper Card](https://www.clippercard.com/?ref=interline.io) is the contact-less payment card and system operated by MTC and available on buses, trains, and ferries throughout the Bay Area.) Transfer discounts between routes operated by the same agency (intra-agency transfers) are included for all methods of payment. Transfer discounts between routes operated by different agencies (inter-agency transfers) are only included for Clipper Card. Our team is updating these fares and transfer discounts on a monthly basis. ## Using GTFS-Fares v2 data for the Bay Area Download the daily Regional Feed like so: 1. [Sign up for a 511 Open Data API token](https://511.org/open-data/token?ref=interline.io) 2. Download from `http://api.511.org/transit/datafeeds?api_key=[your_key]&operator_id=RG` We welcome questions sent to the [511SFBayDeveloperResources mailing list](https://groups.google.com/forum/?ref=interline.io#!forum/511sfbaydeveloperresources). # Acknowledgements Credit and many thanks to project team members including Ian Rees (Interline), Nome Dickerson (Garnet Consulting), Nisar Kapeel and Kapeel Daryani (MTC), and our partners at Bay Area transit agencies. ### US National Transit Database to collect GTFS URLs URL: https://www.interline.io/blog/us-ntd-reporting-gtfs/ Last updated: 2024-11-05T00:13:53.000Z **April 28, 2023**: The Federal Transit Administration has now adopted these new reporting requirements. For a brief overview, see our more recent [blog post describing when FTA will begin requiring US transit agencies to submit URLs for their GTFS feeds to the National Transit Database](https://www.interline.io/blog/us-ntd-reporting-gtfs-adopted/). American public transit agencies that receive federal funds are required to report information for inclusion in the [National Transit Database](https://www.transit.dot.gov/ntd?ref=interline.io). The Federal Transit Administration, which administers to NTD, is [proposing](https://www.regulations.gov/document/FTA-2022-0018-0001?ref=interline.io) that each agency will also provide a URL for a GTFS feed in their NTD reports. Submitted GTFS feeds will enter the public domain. This is great news for all producers and consumers of GTFS data. ![first page of FTA's notice and request for comments regarding additions to the NTD reporting process](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/us-ntd-federal-register-1.png) [Notice and request for comments](https://www.regulations.gov/document/FTA-2022-0018-0001?ref=interline.io) on NTD reporting requirements from the Federal Transit Administration. We’ve submitted feedback to the FTA based on our experience building and operating Transitland, and we’re sharing it here as well: > At Interline Technologies, we specialize in software services for GTFS, GTFS Realtime, and related data specifications. Our Transitland data platform ([www.transit.land](http://www.transit.land/?ref=interline.io)) aggregates open-data feeds from over 3,000 operators in over 50 countries, including 908 operators in the United States. We serve organizations that both produce and consume these data feeds. Our clients include public transit agencies, private mobility providers, planning firms, and academic research institutions. This is just a sampling of the many types of firms that make use of GTFS and related data feeds. > Based on our own experience, as well as the experience of our clients, we see much value in FTA’s proposed addition to the NTD reporting process to have agencies provide a public GTFS feed. Increasing the number and accessibility of openly licensed GTFS feeds with stable URLs from American transit agencies will have many benefits. In addition to being useful for transit trip planning (the best known use-case for GTFS), this will also assist with transportation planning, land-use planning, real-estate assessment, academic research, and local advocacy. Coverage of the largest and the smallest transit agencies is important for many of these endeavors, as is comprehensive coverage of both urban and rural operators. Using NTD reporting to promote the production and open sharing of GTFS data feeds will assist with all of these use-cases. > We are pleased to see that agencies will be encouraged to report a stable URL that is updated on their own servers. Transitland and similar data aggregators operate best when the “source of truth” for each feed is maintained on each agency’s servers and the aggregator fetches from that location on a regular basis. > We encourage FTA to consider multiple means of distributing the lists of GTFS feed URLs. While most NTD data is included in a single Excel file with multiple sheets, we would propose that FTA also disseminate a separate file that just lists GTFS feed URLs. Ideally this would be in the CSV format, be hosted at a stable URL, and include the associated NTD IDs for each GTFS feed URL. In addition to the NTD ID, we suggest also listing the agency’s ID used on the NTD website for agency profile pages. For example, the final portion of `https://www.transit.dot.gov/ntd/transit-agency-profiles/san-francisco-bay-area-rapid-transit-district` > FTA may consider posting a CSV listing the feed URLs, associated NTD IDs, and NTD website IDs to an open-data portal such as data.gov on an ongoing basis. This will be useful for a very wide range of users. > At present, our staff and external contributors maintain associations between GTFS feed URLs and NTD IDs in the Transitland Atlas repository on GitHub. It is open to public use and contributions at [https://github.com/transitland/transitland-atlas](https://github.com/transitland/transitland-atlas?ref=interline.io) While we believe this structure is likely more complex than is necessary for FTA and NTD’s purposes, we offer this as one example for how to provide information on many feeds, along with metadata about each operator and (when relevant) its US NTD ID. ![screenshot of BART operator detail page listing its US NTD ID](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/bart-operator-page-screenshot.png) US NTD ID listed on [Transitland's operator detail page for BART](https://www.transit.land/operators/o-9q9-bart?ref=interline.io). > Finally, we strongly support the proposed change that “upon publication, all GTFS data submitted to the NTD will enter the public domain.” Since 2014, our team has been working to improve the state of licensing for GTFS data. At our last company (Mapzen) we prepared a model license for GTFS feeds and discussed it in depth with USDOT OST staff. At Interline, we continue to assess the state of licenses and terms attached to public transit data. In the Transitland Atlas repository mentioned previously, our staff maintain metadata about the licenses attached to many of the feeds. This is important for our institutional clients, so they can be assured that we and they are making a good faith effort to follow the license and terms for each agency. However, all of this effort to track and review such a wide range of licenses is resource intensive. We do not think it is a good use of anyone’s time. Furthermore, we do not believe it benefits or protects transit agencies that produce GTFS data. GTFS feeds contain no privileged data, nor do they contain personally identifiable information (PII). Publishing a GTFS feed in the public domain does not entail giving up other protections such as a transit agency’s trademark of their brand name and logo. Transit agencies benefit when their GTFS feeds are useful widely by as many consumers as possible. We are pleased to see FTA proposing to simplify this overly complicated situation by instead requiring that GTFS feeds submitted to the NTD reporting process enter the public domain. > Thank you for taking these comments. You are welcome to reach us at [info@interline.io](mailto:info@interline.io) if we can provide any more information. ### Observed stop arrivals and departures through the SF Bay Area URL: https://www.interline.io/blog/mtc-regional-stop-observations/ Last updated: 2024-11-05T00:15:25.000Z When did the bus arrived at that stop yesterday? This isn’t a question that transit riders ask often — but it is a question that analysts and planners often do ask. Interline and the Metropolitan Transportation Commission have been archiving and processing the [Regional GTFS Realtime](https://www.interline.io/blog/mtc-regional-gtfs-realtime-feed/) feed in order to answer these questions. ## What are stop observations? “Stop observations” describes transit service as was delivered to riders: the arrival time of a vehicle at a designated stop, along with any additional contextual information available from GTFS and GTFS Realtime feeds. Stop observations are in contrast with the `stop_times.txt` file in the [daily Regional Feed](https://www.interline.io/blog/mtc-regional-gtfs-feed-release/), which uses static GTFS data to capture transit service as it is planned and scheduled. This is also in contrast with the [monthly Historic Regional Feed](https://www.interline.io/blog/mtc-regional-gtfs-feed-additions/#historical-feeds), which uses static GTFS data to capture transit service schedules, retrospectively, as they changed over an entire month (with a different schedule for each agency for each day). Stop observations can be used to better understand the real-world realities of transit service as it fluctuates from day to day and location to location. Stop observations can be used to estimate on-time performance (also known as schedule adherence or punctuality) and to inform potential schedule adjustments by service planners. Stop observations can also be used by navigation app developers to provide their users with contextual information, such as an estimated range of times when a route is likely to depart from a given stop. ## Measurement error and (im)precision Stop observations are inferred from public GTFS Realtime data feeds (trip updates and vehicle positions) and therefore, there are many potential sources for measurement error. Potential sources of error include GPS error on transit vehicles, misassignment of vehicles to “runs” of a given trip and route, imprecision or data errors in the transit agency’s computer-aided dispatch/automatic vehicle location (CAD/AVL) system, issues in the transit agency’s GTFS Realtime emitting system, issues in 511.org’s GTFS Realtime aggregation system, or issues in the [Regional Realtime GTFS Feed](https://www.interline.io/blog/mtc-regional-gtfs-realtime-feed/). We’ll discuss how stop observations are calculated later in this blog post, to help users better understand stop observations as measurements. When considering potential applications for stop observations, we ask users to keep the limitations of this data source in mind. ## What information is contained in a stop observation? We are distributing stop observations in the Historic Regional Feed. When users request a Historic Regional Feed with stop observations for a given year and month, a `stop_observations.txt` file will be included with all stop observations for that month. The `stop_observations.txt` file is a CSV file based on the [GTFS-Performance draft proposal](https://bit.ly/gtfs-performance?ref=interline.io). GTFS-Performance is an extension to the GTFS specification prepared by Remix, Swiftly, and their partners. The following table is based on the GTFS-Performance draft proposal and covers columns included in our implementation of `stops_observations.txt`, as well as columns that are in the proposal but not in our implementation: | Field Name | Details | Regional GTFS stop observations support | | ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------ | | trip\_id | Identifies a trip. When schedule\_relationship=DUPLICATED, then this references the trip that is copied, like GTFS Realtime TripUpdate.trip.trip\_id. | ✅ | | trip\_start\_time | See GTFS Realtime TripDescriptor.start\_time. Required to specify the start time of a trip that’s not defined in Static GTFS which occurs when using frequency-based service defined by frequencies.txt or when an ad-hoc added service based on an existing trip via TripUpdate.tripschedule\_relationship=DUPLICATED. | ✅ (when provided in agency source RT feed) | | stop\_id | Identifies a stop | ✅ (added by Interline for developer convenience) | | stop\_sequence | Identifies the order of stops for a particular trip. | ✅ | | schedule\_relationship | Valid options are: 0 - SCHEDULED 1 - SKIPPED 2 - NO\_DATA 3 - UNSCHEDULED 4 - CANCELED 5 - DUPLICATED 6 - MODIFIED | ✅ | | schedule\_adherence\_stop | Indicates if this observation is used for reporting purposes when a subset of stops are utilized, instead of all stops. | No | | trip\_start\_date | The start date of this trip instance in YYYYMMDD format to specify the service date of the corresponding trip. This should match the value in TripDescriptor.start\_date. | ✅ | | vehicle\_id | Unique identifier for the vehicle that served this stop. This should match the value provided in real-time via VehicleDescriptor.id. | ✅ (when provided in agency source RT feed) | | scheduled\_arrival\_time | GTFS static scheduled arrival time | ✅ (added by Interline for developer convenience) | | observed\_arrival\_time | Observed arrival time at a specific stop for a specific trip on a route. For times occurring after midnight on the service day, enter the time as a value greater than 24:00:00 in HH:MM:SS local time for the day on which the trip schedule begins. | ✅ | | observed\_departure\_time | Observed departure time at a specific stop for a specific trip on a route. For times occurring after midnight on the service day, enter the time as a value greater than 24:00:00 in HH:MM:SS local time for the day on which the trip schedule begins. | ✅ (when available) | | new\_departure\_time | Required when schedule\_relationship=DEPARTURE\_TIME\_MODIFIED. | No | | dwell\_time\_secs | Seconds spend boarding/alighting riders at the stop. | No | | scheduled\_dwell\_time\_secs | Seconds scheduled for boarding/alighting riders at the stop (when provided in agency source RT feed). | ✅ | | boardings | See [GTFS-ride](https://gtfsride.org/?ref=interline.io)'s board\_alight.boardings | No | | alightings | See [GTFS-ride](https://gtfsride.org/?ref=interline.io)'s board\_alight.alightings | No | | source | Valid options are: 0 - Manual 1 - Door Sensor 2 - Estimated using GPS/AVL data | No | | uncertainty | See StopTimeEvent.uncertainty. | ✅ | | occupancy\_status | See VehiclePosition.occupancy\_status. | ✅ (when provided in agency source RT feed) | | occupancy\_percentage | See VehiclePosition.occupancy\_percentage. | ✅ (when provided in agency source RT feed) | | occupancy\_type | See [GTFS-ride](https://gtfsride.org/?ref=interline.io)'s board\_alight.load\_type. | No | | occupancy\_source | See [GTFS-ride](https://gtfsride.org/?ref=interline.io)'s board\_alight.source. | No | ## How are stop observations calculated? Stop observations are derived retrospectively using an archive of the Regional GTFS Realtime Feed (updated every 15 seconds). Once a day, our workflow processes GTFS Realtime trip updates and vehicle positions from the archive, combined together with the day’s static GTFS daily regional feed. **Deriving stop observations from trip updates:** When an agency supplies GTFS Realtime trip updates, our workflow identifies the final `TripUpdate` sent before a vehicle reaches a given stop. From that record, our workflow extracts the most recent `StopTimeUpdate`. Here is an example showing three trip updates fetched from BART. The first two trip updates are before the train reaches Montgomery Station; the third trip update is after the train departs Montgomery Station. The portion of all of the records that is use to create a stop observation is marked with the orange box: ![illustration of three GTFS Realtime feed trip updates, showing the portion of stop time updates retained to create a stop observation record](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/trip-update-selection-1.png) **Deriving stop observations from vehicle positions**: When an agency supplies vehicle positions for a given route, our workflow also identifies the last update sent before the stop location. To do this, we need to consider the geometry of the route alignment from `shapes.txt` in the static RG feed. The workflow snaps the latitude/longitude point from each vehicle position to the route alignment, enabling it to determine from which update to take the `StopTimeUpdate` portion. Like many geometric algorithms, this approach can be stymied by complicated paths, such as looping routes that visit the same station multiple times. Our workflow can take advantage of stop sequence, when specified. Still, the algorithm can only do its best to estimate when the route alignment is complicated and/or the update frequency is infrequent. By being able to derive stop observations from two different parts of GTFS Realtime feeds, we have also been able to test and tune our workflow, to ensure that the two approaches agree when source feeds are of high quality. ## Related data specifications Our stop observations data model is based on the [GTFS-Performance](http://bit.ly/gtfs-performance?ref=interline.io) proposal. The GTFS-Performance proposal is, in turn, inspired by the larger [TCRP G-18](https://apps.trb.org/cmsfeed/TRBNetProjectDisplay.asp?ProjectID=4687&ref=interline.io) project. The GTFS-Performance `stop_visits.txt` is supposed to represent a subset of the TCRP G-18 `stop_visits.txt` file. Thank you to the creators of GTFS-Performance and to the G-18 research team. GTFS-Performance and TCRP G-18 go well beyond the functionality that we have implemented in our system for producing stop observations, to also address more internal operating concerns of transit agencies. We have intentionally implemented our system to create stop observations using only publicly available GTFS and GTFS Realtime inputs. ## Accessing stop observations Stop observations are now included in historical monthly RGs. To download a historical monthly RG including stop observations, make an HTTP request like so: ``` wget https://api.511.org/transit/datafeeds?operator_id=RG&historic=2022-03-so&api_key=xxx ``` Replace `2022-03` with a given year and month. Note the addition of `-so` to request the inclusion of stop observations. Note that `wget` is a Unix command that can be replaced by any other utility to download the URL; also note that `xxx` should be replaced by your own [511.org Open Data API key](https://511.org/open-data/token?ref=interline.io). Monthly RGs without stop observations are roughly 50Mb in size; with stop observations, the monthly RG will be roughly 500Mb to download as a compressed ZIP archive and then roughly 2Gb of uncompressed CSV files. Stop observations are available on a more frequent basis to staff at MTC and Bay Area transit agencies. Transit agencies may contact the 511 SF Bay Data Portal at [transitdata@511.org](mailto:transitdata@511.org) for more information. ## Using stop observations to understand transit Public-transit agencies can be hesitant to release this type of data. In part, it’s a challenge to assemble operations data in an ongoing and consistent manner. In part, it’s a risk that the general public will take operations data out of context. While we began this blog post by asking *When did the bus arrived at that stop yesterday?*, we hope that users of stop observations will ask an even wider range of questions of this dataset, especially questions that can explore the complexity of daily operations for over 20 agencies operating buses, trains, subways, and ferries across the Bay Area and questions that can help to improve the experience for riders using navigation apps and similar software. ### Downloading archived GTFS feed versions from Transitland URL: https://www.interline.io/blog/downloading-archived-gtfs-feed-versions-from-transitland/ Last updated: 2024-11-05T00:44:29.000Z 💡 ****July 19, 2022**: We appreciate questions and feedback we've received from a wide range of Transitland users about this change. Based on feedback, we are adding a Hobbyist/Academic plan with no cost. Users developing non-commercial applications may use this plan to access historical feed versions through the v2 REST API. Commercial users who need access to historical feed versions, as well as academic users who require support and/or access to the v2 GraphQL API, are directed to the Professional plan. See a comparison of functionality on the [Transitland plans and pricing](https://www.interline.io/transitland/plans-pricing/) page. (And one more update to this blog post: the count of feed version in the Transitland archive is out of date. Now there are 105,763 :) For roughly seven years, Transitland has been fetching and archiving static GTFS feeds from transit operators. We've also imported a selection of feed versions from other public archives, such as [GTFS Data Exchange](https://www.transit.land/news/2016/05/16/gtfs-data-exchange?ref=interline.io). Today Transitland contains 102,191 processed and archived feed versions. To date, we have allowed the public to freely download archived feed versions through the Transitland website. The [free tier for Transitland API access](https://www.interline.io/transitland/plans-pricing/) has also provided a generous allotment of queries, which can be used to download archived feed versions, in addition to querying into data contents imported from those feed versions. Some users are taking advantage of this openness by scraping the entire contents of Transitland's feed version archive. These mass downloads put load on Transitland servers and bandwidth, increasing our operating expenses. It also blurs distinctions between the API free tier (intended to be generous enough for hobbyist projects and pilot projects by companies) and the professional tier (intended to support production projects). We are adjusting access to Transitland's GTFS feed version archive to better balance these concerns. ## Downloading feed versions through the website This website now allows users to download the latest version of any feed (with the exception of those marked as having licenses that prevent redistribution). Click through to a feed page and you'll see a download icon in the top frow of the table of archived feed versions. For example, the [Caltrain feed](https://www.transit.land/feeds/f-9q9-caltrain?ref=interline.io): [](https://www.transit.land/feeds/f-9q9-caltrain?ref=interline.io) ## Downloading feed versions through the v2 REST API Developers can use the v2 REST API to download feed versions by two means: 1. Use [feed endpoint](https://www.transit.land/documentation/rest-api/feeds?ref=interline.io#downloading-source-gtfs) to download the latest feed version of a given feed. For example: `GET /api/v2/rest/feeds/f-9q9-caltrain/download_latest_feed_version` 2. Use the [feed versions endpoint](https://www.transit.land/documentation/rest-api/feed%5Fversions?ref=interline.io#downloading-source-gtfs) to specify a feed version by its unique SHA1 hash and download it. For example: `GET /api/v2/rest/feed_versions/8f99f69503edfa52e06ec5673e582d3a684c05ca/download` Starting in seven days (on Thursday, June 23, 2022), Transitland will only allow users on the [Professional or Enterprise plan](https://www.interline.io/transitland/plans-pricing/) to download feeds using the second of those options. (As of July 19, 2022, non-commercial users on the [Hobbyist/Academic plan](https://www.interline.io/transitland/plans-pricing/) are also being given this capability.) All registered developers, including those on the Free plan, will still be able to use the first of those options. ## Continuing to tune and improve We founded Transitland in 2014 to serve as "a community-edited data service aggregating transit networks across metropolitan and rural areas around the world." We've added, adjusted, and removed functionality over time. We've also supported Transitland's expenses and upkeep through a changing mixture of organizational sponsorship, charitable sponsorship, and customer revenue. We expect to continue to adjust Transitland's functionality and access policies over time to maintain and expand this platform as the best possible set of open transit and mobility APIs. Thanks to our users — both old and new, and on both free and paid plans — for making such good use of Transitland. ### Transitland Vector Tile Improvements URL: https://www.interline.io/blog/vector-map-tiles/ Last updated: 2024-11-05T00:54:29.000Z Transitland's [global transit map](https://www.transit.land/map?ref=interline.io) provides a worldwide view of public bus, train, subway, ferry, and gondola route lines and stop points around the world. Using the [v2 Vector Tiles API](https://www.transit.land/documentation/vector-tiles?ref=interline.io) developers can integrate Transitland routes and stops tiles into their own web maps. Since [releasing](https://www.interline.io/blog/transitland-v2-vector-map-tiles) the global transit map and v2 Vector Tiles API last year, we've added a variety of improvements in functionality and in rendering. ## Performance Improvements The Transitland global transit map and v2 Vector Tiles API are based on tiles in the [MVT (Mapbox Vector Tile) format](https://docs.mapbox.com/data/tilesets/guides/vector-tiles-standards/?ref=interline.io). We've made a number of changes to the data that's encoded into each MVT tile for route geometries. The changes are small but add up to a meaningful decrease in the size of the tiles and an increase in rendering performance. These improvements are especially useful at zoomed out (small scale) views: [](https://www.transit.land/map?ref=interline.io#4.93/55.95/5.4) ## Representative Route Geometries While some transit routes are simple, with all trips following the same line, other transit routes can be very complicated. Routes may have trips along certain lines at peak hours and along other lines at off-hours. Routes may branch and go in multiple directions depending upon the trip. Many different combinations are possible all under the same branding of a single route name. Transitland creates "representative" route geometries for every route. This geometry must balance competing goals: - show all of the different variants of the route (all of the possible trip geometries) - provide a single geometry that can be understood on its or on a map with many other routes - not be too large in file size When a source feed does not provide `shapes.txt` for a given trip, Transitland generates "stop-to-stop geometries" by drawing straight lines between stop locations. The representative geometry for each route is created using the following rules: 1. for each route direction: - select the longest shape - select any shape used for at least 20% of trips 2. from this collection, drop shapes with totally broken geometry, e.g. a point at (0,0) 3. drop shapes contained entirely within another selected shape 4. include generated stop-to-stop geometries only if no `shapes.txt` shapes are available 5. combine the selected shapes into a representative `MultiLineString` 6. the single most frequently used shape becomes the representative `LineString` ## Problematic Route Geometries Creating high quality `shapes.txt` records in GTFS feeds can be a challenge. Occasionally agencies do make mistakes. For example: - routes shapes that travel to [Null Island](https://blogs.loc.gov/maps/2016/04/the-geographical-oddity-of-null-island/?ref=interline.io) (include a 0,0 point) - routes shapes that zigzag between stop locations out of order - route shapes that have not been updated to reflect current stop locations and sequence Transitland's goal when importing from `shapes.txt` is the same as our overall goal with GTFS source feeds: We aim to import as many records as possible, while ignoring records that are invalid or that are of such low quality that they will not be usable. Transitland continues to import the overall feed, even if it must leave out some records. (Transitland's import strategy is in contrast with validators and tooling that are used by agencies when generating GTFS — for those types of use-cases, the user may wish to have the process fail an entire feed when one of its entities causes a problem. The operator will then fix the feed before repeating the validation process.) Transitland tags as problematic route geometries that fit the following criteria: - Shape length greater than 5,000km - Any individual line segment above 50km for `shapes.txt` geometries - Any individual line segment above 5km for generated stop-to-stop geometries - "Zigzag" shapes where points are in obviously incorrect order Problematic geometries are still included in the MVT vector tile data, but the global transit map hides these geometries by default. To see them, open *Options* and select *Show problematic route geometries*: [](https://www.transit.land/map?ref=interline.io#6.53/-36.665/144.322) ## Two tiers of feature detail A typical tile set contains over 150,000 routes! This can be expensive to load and render at low zoom levels, so some compromises are baked into the tiles. Vector tile zoom levels 8 and higher include the full `MultiLineString` representative geometries and all route features. Tile zoom levels between 0 and 7 include a single `LineString` geometry for each route, and on particularly dense tiles, a percentage of route features are dropped to help manage the file size of each tile (routes with high levels of service are weighted towards inclusion). Zoom levels 0 to 13 include line simplification appropriate for displaying at that zoom level; zoom level 14 includes the full, unsimplified geometry. | Zoom level | Feature(s) | Level of detail | | ---------- | -------------------- | --------------------------------------------------- | | 0 | single LineString | geometry simplified for display at given zoom level | | 1 | single LineString | geometry simplified for display at given zoom level | | 2 | single LineString | geometry simplified for display at given zoom level | | 3 | single LineString | geometry simplified for display at given zoom level | | 4 | single LineString | geometry simplified for display at given zoom level | | 5 | single LineString | geometry simplified for display at given zoom level | | 6 | single LineString | geometry simplified for display at given zoom level | | 7 | single LineString | geometry simplified for display at given zoom level | | 8 | full MultiLineString | geometry simplified for display at given zoom level | | 9 | full MultiLineString | geometry simplified for display at given zoom level | | 10 | full MultiLineString | geometry simplified for display at given zoom level | | 11 | full MultiLineString | geometry simplified for display at given zoom level | | 12 | full MultiLineString | geometry simplified for display at given zoom level | | 13 | full MultiLineString | geometry simplified for display at given zoom level | | 14 | full MultiLineString | full unsimplified geometry | ## Use Transitland v2 Vector Tiles API in QGIS Can you use Transitland vector tiles in desktop GIS software? Thanks to a curious user for [asking this question](https://github.com/transitland/transitland/discussions/355?ref=interline.io). Turns out the answer is yes. If you want to use Transitland route lines and stop locations as layers in your own GIS map, along with your other data sources, here are instructions for how to do so in QGIS: 1. If you have not already, download and install the [QGIS desktop GIS software package](https://qgis.org/?ref=interline.io). 2. If you have not already, [sign up for a Transitland API key](https://www.transit.land/documentation?ref=interline.io#signing-up-for-an-api-key). 3. From the menu, select *Layer > Add Layer > Add Vector Tile Layer...* 4. Press the *New* button and select *New Generic Connection...* 5. Enter "Transitland stops" or "Transitland routes" at the name (or whatever you wish) 6. For URL, enter `https://transit.land/api/v2/tiles/routes/tiles/{z}/{x}/{y}.pbf?apikey=xxx` or `https://transit.land/api/v2/tiles/stops/tiles/{z}/{x}/{y}.pbf?apikey=xxx` You will need to replace `xxx` with your own API key 7. Press the *OK* button 8. In the *Browser*side menu, you will now see the new source listed under *Vector Tiles*. Right click on it and select *Add Layer to Project* 9. You should now see the layer visualized. Repeat the steps to add the other layer, if you wish to have both stop points and route lines. 10. Customize the styling for one or both of the layers. ![screenshot of QGIS desktop GIS software displaying Transitland vector tiles](https://www.transit.land/images/vector-map-tiles/transitland-tiles-in-qgis.png) Please note that the [Transitland terms](https://www.transit.land/terms?ref=interline.io) require that you provide attribution on maps and other creations using output from the Transitland APIs. We look forward to seeing what GIS users, software developers, and data analysts create using these enhanced transit map tiles. We welcome all to post their creations on [Transitland's "show and tell" discussion board](https://github.com/transitland/transitland/discussions/categories/show-and-tell?ref=interline.io). ### Supporting the Mobility Data Interoperability Principles URL: https://www.interline.io/blog/mobility-data-interoperability-principles/ Last updated: 2024-11-12T22:01:05.000Z Interline is pleased to sign the newly released [Mobility Data Interoperability Principles](https://www.interoperablemobility.org/?ref=interline.io). Shared data specifications and practices are key to our services and products, as well as to the overarching goals of our customers. When software systems, datasets, and vendors “play well with others,” the benefits accrue to all participants. The principles are as follows: > All systems creating, modifying, or consuming mobility data should be interoperable.Interoperability should be achieved through the development, adoption, and widespread implementation of open standards that support the efficient exchange and portability of mobility data.Transit agencies and other mobility service providers should have access to tools that present high-quality mobility data accessibly, equitably, and in real time to assist travelers in meeting their mobility needs.Transit agencies, other mobility service providers, and travellers should be able to select the transportation technology components that best meet their needs.All individuals and the public should be empowered through high-quality, well-distributed mobility data to find, access, and utilize high-quality mobility options that meet their needs as they see fit, while maintaining their privacy.[Mobility Data Interoperability Principles v1](https://www.interoperablemobility.org/?ref=interline.io) released in October 2021. Also included is a [glossary](https://www.interoperablemobility.org/definitions/?ref=interline.io). This level of detail is useful for newcomers to the field, as well as for referencing the principles in formal documents like a public-sector RFP (request for proposals): > **Interoperability**: The ability for any mobility technology component to exchange data in an open standard or schema with other components in that mobility technology system. The measure of success will be putting these principles into practice. Interline had the opportunity, along with other vendors, to provide [feedback and suggestions](https://www.interoperablemobility.org/public%5Fdraft%5Freview/?ref=interline.io) based on our experiences building open-source and open-data transportation tooling, as well as our experiences helping our clients to integrate with proprietary and/or legacy systems. We thank the authors (all public-sector agencies) for using this feedback from vendors to refine the principles and supporting documentation. These principles are a useful articulation of the collaborative practices that have been growing throughout the geospatial and transportation tech industries for years. The principles are also — to be frank — a useful articulation of what our industries should avoid: software systems and business models based on silos, walled gardens, or excessive vertical integration. It will always be a balancing act to find the optimal mix of open versus closed attributes for a software system, not to mention an organization. These principles succeed in reminding us all of the important overall goals of “empowering transit agencies and other mobility service providers and transportation system managers to provide better service, improve the customer experience, and build systems that are equitable and sustainable.” ### SF Bay Area's Regional GTFS Realtime Feed URL: https://www.interline.io/blog/mtc-regional-gtfs-realtime-feed/ Last updated: 2024-11-01T22:21:17.000Z Last year, [Interline and the Metropolitan Transportation Commission released the Regional GTFS Feed for the San Francisco Bay Area](https://www.interline.io/blog/mtc-regional-gtfs-feed-release/) and also a [series of improvements to the static data including station pathways and fares](https://www.interline.io/blog/mtc-regional-gtfs-feed-additions/). We’re now pleased to announce a single consolidated Regional GTFS Realtime feed. Developers can combine the static Regional GTFS Feed and the Regional GTFS Realtime endpoints for comprehensive and up-to-date views of the Bay Area’s bus, train, and ferry services. Combining feeds is straightforward, as the same ID schemes are used for data entities across all feeds. ## 21 agencies included The Regional GTFS Realtime Feed currently aggregates data for the following Bay Area transit agencies: | Operator Name | Service Alerts | Trip Updates | Vehicle Positions | | ----------------------------------------------------- | -------------- | ------------ | ----------------- | | AC Transit | ✅ | ✅ | ✅ | | Bay Area Rapid Transit (BART) | ✅ | ✅ | \- | | Caltrain | ✅ | ✅ | ✅ | | County Connection | ✅ | ✅ | ✅ | | Dumbarton Express Consortium | ✅ | ✅ | ✅ | | Emery Go-Round | ✅ | ✅ | ✅ | | Fairfield and Suisun Transit | ✅ | ✅ | ✅ | | Golden Gate Transit (GGT) | ✅ | ✅ | ✅ | | Livermore Amador Valley Transit Authority (LAVA) | ✅ | ✅ | ✅ | | Marin Transit | ✅ | ✅ | ✅ | | Petaluma Transit | ✅ | ✅ | ✅ | | SamTrans | ✅ | ✅ | ✅ | | San Francisco Municipal Transportation Agency (SFMTA) | ✅ | ✅ | ✅ | | Santa Rosa CityBus | ✅ | ✅ | ✅ | | SolTrans | ✅ | ✅ | ✅ | | Sonoma County Transit | ✅ | ✅ | ✅ | | Sonoma Marin Area Rail Transit (SMART) | ✅ | ✅ | ✅ | | Tri Delta Transit | ✅ | ✅ | ✅ | | Union City Transit | ✅ | ✅ | ✅ | | VINE Transit | ✅ | ✅ | ✅ | | Valley Transportation Authority (VTA) | ✅ | ✅ | ✅ | MTC will be announcing additional agencies providing GTFS Realtime data in the future. ## Built using shared tooling and validation processes Like the static Regional GTFS Feed, the Regional GTFS Realtime Feed is build using Interline’s open-source tooling, including the [transitland-lib software library](https://github.com/interline-io/transitland-lib?ref=interline.io). While developing the Regional GTFS Realtime Feed, our teams wanted to ensure that the process of merging together all the individual input agency feeds into the single set of output endpoints did not introduce any additional data errors. We made us of the same GTFS Realtime validation tooling and test process as we recently described in a [Transportation Research Board report by Interline and our partners at University of South Florida](https://www.interline.io/blog/trb-transit-idea-gtfs-realtime-report/). ## How to consume the Regional GTFS Realtime Feed The Regional GTFS Realtime Feed is now available to all developers through the 511 Open Transit Data API. For instructions on how to access the endpoints and to sign up for an API key, see this [post to the 511 SF Bay Developer Resources mailing list](https://groups.google.com/g/511sfbaydeveloperresources/c/VlHlYyj4vRA/m/sLdfHXhdAgAJ?ref=interline.io). ### Use Transitland v2 to Search Across Operators, Feeds, and Routes URL: https://www.interline.io/blog/transitland-v2-search/ Last updated: 2024-11-01T22:52:16.000Z Over the last year (or two?) [we've been creating Transitland v2](https://www.interline.io/blog/tlv2), a major rewrite of the platform's internals. Transitland v2 brings performance improvements, the option to deploy any number of customized Transitland installations (in addition to the "canonical" Transitland that lives here at [transit.land](https://www.transit.land/?ref=interline.io)), and useful new feature additions. Here's one of our favorite new additions: search across all of the operators, feed records, and routes in Transitland. ![animated screenshot showing a user type 'Staten' in to the search field and click through to view the route page for the Staten Island Ferry in New York City](https://www.transit.land/images/tlv2-feature-announcements/tlv2-search-animation.gif) The new search functionality currently indexes millions of records and returns results quickly — very quickly! Try it yourself. Type into the search box in the upper right corner of this page and watch as the results are returned. Search can also be scoped. For example, open up the Chicago Transit Authority operator page and you can search within that operator's stops and routes: [](https://www.transit.land/operators/o-dp3-chicagotransitauthority?ref=interline.io#routes) Developers, we welcome you to start thinking of ways you might want to integrate Transitland's search functionality into your own apps, websites, and creations. Search will be included in the [upcoming Transitland v2 APIs](https://www.transit.land/documentation?ref=interline.io). Search is always complicated. There are questions of what to index, how to weight different parameters, and so on. For example, we are not yet including stops/stations in the global search at this point. We welcome feedback and bug reports. Are you searching Transitland only to find that we do not have data for your favorite transit operator? Help expand Transitland's coverage by contributing to the [Transitland Atlas](https://github.com/transitland/transitland-atlas?ref=interline.io), the open directory of feeds that powers Transitland's services. ### Transitland v2 Adds a Global Map and Vector Tiles URL: https://www.interline.io/blog/transitland-v2-vector-map-tiles/ Last updated: 2024-11-05T00:54:57.000Z 💡 ****June 10, 2022**: Here's a more recent blog post [describing further improvements and refinements to Transitland's global transit map and vector tile API](https://www.interline.io/blog/vector-map-tiles). Over the last year (or two?) [we've been creating Transitland v2](https://www.interline.io/blog/tlv2/), a major rewrite of the platform's internals. Transitland v2 brings performance improvements, the option to deploy any number of customized Transitland installations (in addition to the "canonical" Transitland that lives here at [transit.land](https://www.transit.land/?ref=interline.io)), and useful new feature additions. Here's one of our favorite new additions: a "slippy map" of transit routes and stops around the world. ![map showing transit routes around the world, zooming in to Los Angeles, panning to Amsterdam](https://www.transit.land/images/tlv2-feature-announcements/tlv2-map-animation.gif) Try the Transitland global transit map yourself at [https://www.transit.land/map](https://www.transit.land/map?ref=interline.io) Pan, zoom, and place your mouse cursor over features to learn more. Click to see more details about particular routes and stops. The Transitland global transit map powered by vector tiles in the [Mapbox Vector Tile format](https://docs.mapbox.com/vector-tiles/specification/?ref=interline.io). Any user of the [Transitland v2 Vector Tile API](https://www.transit.land/documentation/vector-tiles?ref=interline.io) can build these route line and stop location layers into their own "slippy maps" on the web, with their own styling and customizations. The tiles are always fresh — Transitland updates them with the latest data each day. Note that Transitland continues to provide GeoJSON through the [Datastore v1 API and the upcoming v2 APIs](https://www.transit.land/documentation?ref=interline.io). When creating smaller maps or static maps, it may be simpler to use GeoJSON. The power of MVT is when you need both wide coverage and precise detail. Not seeing public transit routes and stops where you expect to find them on the map? Help expand Transitland's coverage by contributing to the [Transitland Atlas](https://github.com/transitland/transitland-atlas?ref=interline.io), the open directory of feeds that powers Transitland's services. ### Our GTFS Realtime validation report published by the Transportation Research Board URL: https://www.interline.io/blog/trb-transit-idea-gtfs-realtime-report/ Last updated: 2024-11-05T00:53:15.000Z In 2019 and 2020, Interline and the [Center for Urban Transportation Research](https://www.cutr.usf.edu/?ref=interline.io) at University of South Florida collaborated on a [Transit IDEA Program grant](http://apps.trb.org/cmsfeed/TRBNetProjectDisplay.asp?ProjectID=4695&ref=interline.io) from the US Transportation Research Board. Our goal has been to help transit agencies to validate and improve the quality of their GTFS Realtime feeds. This first grant is now complete and TRB has published [the full report with our findings and proposed next steps](http://www.trb.org/Main/Blurbs/181415.aspx?ref=interline.io). ![TRB Transit IDEA T-93 report cover](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/trb-t-93-report-cover-1.png) The full report can be downloaded from the [TRB website](http://www.trb.org/Main/Blurbs/181415.aspx?ref=interline.io). ## Executive Summary Real-time transit information has been shown to have many benefits for transit riders and agencies, including shorter perceived and actual wait times, a more welcoming experience for new riders, and an increased feeling of safety, and increased ridership. Real-time transit data is, in comparison with many other potential operational or capital improvements to bus or rail service, an affordable means of increasing ridership. In the last few years, a real-time complement to the General Transit Feed Specification (GTFS) format, GTFS Realtime, has emerged, which enables transit agencies to share real-time predictions, vehicle positions, and service alert data in a standardized format. Despite its promise, adoption of GTFS Realtime by transit agencies has been hampered by readily available validation tools. ![technical architecture diagram showing GTFS Realtime feeds connecting transit agency data systems to consuming applications](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/gtfs-realtime-architecture-diagram.png) GTFS Realtime feeds (the blueish box on the right) are a critical connection between the complexity of transit agencies' data systems and consuming applications, like smartphone navigation apps. In this project, Interline Technologies LLC and the Center for Urban Transportation Research (CUTR) at the University of South Florida (USF) have created a prototype platform that makes GTFS Realtime validation tools readily available to, potentially, all transit agencies in North America. Our team is building upon two open-source projects: the [GTFS Realtime validator prototype](https://github.com/CUTR-at-USF/gtfs-realtime-validator/?ref=interline.io) and [Transitland](https://www.transit.land/?ref=interline.io), an open transit data platform. This project applies the open-source and open-data community models to the challenges of creating and improving GTFS Realtime data. **Stage 1: Build and Test GTFS Realtime Data Platform**: The research team combined the Transitland open data platform, the GTFS Realtime Validator, and a list of 162 GTFS Realtime feed endpoints (provided by a partner organization). The combined platform collects GTFS Realtime data from each feed, runs the validator process, and produces a report on any detected errors. Each report shows the counts of data entities, the percentage with errors, and a brief text description of any errors. Links take users to additional documentation about each error type. Some errors also provide further contextual information in maps and tables to assist users as they try to determine root causes. ![screenshots of the Transitland GTFS Realtime validation reports](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/transitland-validation-report-screenshots.png) Screenshots of example GTFS Realtime validation reports. **Left*: Listing errors (with red severity labels) and a warning (with a yellow severity label) detected in 24 hours of vehicle-position updates. **Right*: Viewing a specific issue to understand why a vehicle (the little red dot) is far from its assigned route (the blue line). **Stage 2: Testing and Expanding Catalogue**: The project team has tested the platform by preparing validation reports for seven public-transit agencies and reviewing the results in the platform user interface with agency staff members over video calls. In these user-testing sessions, the project team collected information from agency staff about how GTFS and GTFS Realtime data are currently created at each agency, known issues, and any open goals. After being given a tour through the platform and its interface, agency staff reviewed the reports for their own GTFS Realtime feeds. Agency staff were asked to provide input on both the specific quality checks and the overall presentation and approach used by the platform with a standardized question list. The research team has summarized notes from the user-testing sessions and used these results to inform preliminary plans for expanding the platform for a wider range of users in the future. **Product Pay-Off Potential**: By combining the open-source components of a GTFS Realtime validator with a catalog of GTFS Realtime feeds, hosted on Transitland’s cloud servers, this project will make the process of validating real-time data simple and accessible to agency staff from any computer with a web browser. As a result, GTFS Realtime data will improve in quality and availability. Transit riders will have a better experience (which has been linked to higher ridership), agency staff will provide better service with less effort and cost, and system vendors will provide a higher quality product. **Product Transfer**: To succeed, this platform will need to be usable by agency staff and to provide them with results that they can act upon, both within the context of their agencies and with their vendors. Therefore, the project team has involved agency stakeholders early and often in the project. The team recruited nine transit agencies to write letters of support as part of the Transit IDEA proposal, and seven agencies have participated in our user-testing process. These agencies represent a diverse range of rider population sizes, staff skill level, and location (urban and rural). The final report includes information on immediate plans to continue and expand Transitland’s catalog GTFS Realtime feeds and longer term plans to develop a GTFS Realtime certification process supported by the platform. ## Next Steps We invite everyone to browse Transitland’s [list of GTFS Realtime feeds](https://www.transit.land/feeds/?ref=interline.io). Please help us to maintain this list by contributing additions and edits to the [open Transitland Atlas feed registry](https://github.com/transitland/transitland-atlas?ref=interline.io). Do you work for a transit agency that publishes GTFS Realtime data? Please check to make sure it’s on Transitland. If not, please add it to the [open Transitland Atlas feed registry](https://github.com/transitland/transitland-atlas?ref=interline.io) or [contact us for assistance](mailto:info@interline.io). Do you work for a transit agency that has issues with GTFS Realtime data, or do you represent a group with a stake in the future of GTFS Realtime data? We’re looking to collaborate on opening this GTFS Realtime validation platform for self-serve use by agencies. [We welcome your interest](mailto:info@interline.io). ## Acknowledgements Research team members include: - [Ian Rees](https://www.linkedin.com/in/irees?ref=interline.io), Interline, Principal - [Sean Barbeau](https://www.linkedin.com/in/seanbarbeau/?ref=interline.io), Center for Urban Transportation Research (CUTR), University of South Florida (USF), Principal Mobile Software Architect for R&D Thanks to members of the T-93 expert review panel (ERP): - Patricia Collette - Stephen M. Stark - Aaron Antrim - Carol Schweiger - Murat Omay - Santosh Mishra Many thanks to Velvet Basemera-Fitzpatrick, Ph.D., Transit IDEA Program Manager. Thanks to the following transit agencies and staff members for participating in user testing: - Metropolitan Transportation Commission (San Francisco): Nisar Ahmed - Metro Transit (Minneapolis): Laura Matson, Mark DiPasquale, Gary Nyberg, Joey Reid, Juan Villanueva - Massachusetts Bay Transportation Authority (Boston): Paul Swartz, Jessie, Richards, Logan Nash - TriMet (Portland, Oregon): Mike Gilligan, Guy Tinat - Tompkins Consolidated Area Transit (Ithaca, New York): Tom Clavel, Matthew Yarrow - Washington Metropolitan Area Transit Authority (Washington, D.C.): Stephanie Jones, Richard Carman - Hillsborough Area Regional Transit (Tampa, Florida): Dexter Corbin, William Mozel, Tim Wictor - Rogue Valley Transportation District (Medford, Oregon): Melissa Lowry ### SF Bay Area's Regional GTFS Feed Expanded URL: https://www.interline.io/blog/mtc-regional-gtfs-feed-additions/ Last updated: 2024-11-05T00:45:36.000Z 💡 ****October 4, 2022**: MTC and Interline have released updated fares and transfer discounts for all Bay Area transit agencies using the adopted GTFS-Fares v2 "base implementation." See [this blog post for more information](https://www.interline.io/blog/mtc-regional-gtfs-feed-fares-updates/). 💡 ****September 10, 2021**: MTC and Interline have now released the Regional GTFS **Realtime* Feed, which provides up-to-the-minute service alerts, trip updates, and vehicle positions using the same ID scheme as the static Regional GTFS Feed. See [this blog post for an overview](https://www.interline.io/blog/mtc-regional-gtfs-realtime-feed/). Earlier this year, [Interline and the Metropolitan Transportation Commission released the Regional GTFS Feed for the San Francisco Bay Area](https://www.interline.io/blog/mtc-regional-gtfs-feed-release/). The Regional GTFS Feed is produced on a daily basis and made available through [511 SF Bay’s Open Data Portal](https://511.org/open-data?ref=interline.io) and its Open Transit Data API. We’re pleased to now share a series of additions to the Regional GTFS Feed that we’ve released together over recent months: - [**Historical Feeds**](#historical-feeds) which provide a retrospective look at an entire month of service - [**Station Pathways and Levels**](#station-pathways-and-levels) to provide richer wayfinding for all types of riders - [**Fares and Transfer Discounts**](#fares-and-transfer-discounts) \[beta\] to calculate the cost of transit journeys on and across the eight largest agencies # Historical Feeds Since we released the Regional GTFS Feed in January, much has changed across the Bay Area. Transit agencies and their dedicated staff have rapidly reduced and re-targeted transit service for those riders who perform essential work or depend upon transit for their necessary journeys. Using the newly released Historical Feeds component of the Regional GTFS Feed, we can visualize how transit service has changed throughout the Bay Area from January to July of 2020: ![Animated series of maps showing bus, train, and ferry route lines throughout the Bay Area with the lines given different weights to show which routes provide the most frequent service](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/headway-animation.gif) Buses, trains, and ferries throughout the San Francisco Bay Area color-coded by frequency: the reddest route lines are routes that operate at least every 10 minutes; medium hues of red are 10 - 20 minute headways; the lightest hue of red are routes with service less than once every 20 minutes. This animation is created by using the Historical Feeds for the months of January through July of 2020. Historical Regional Feed products are fully valid GTFS feeds, but they differ somewhat in their contents from the daily Regional Feed products. Read on to understand the process used to produce the Historical Regional Feed products and their key differences, or skip to the end of this section to [download Historical Feed products](#download-historical-feed-products). ## Slicing regional feeds Each day, the Regional Feed is produced from the versions of agency feeds on 511.org that provide the best view of service on that day. Each month, the Historical Regional Feed creation process takes these Regional Feeds and combines them together, taking one day of service from each feed, which we are calling a “slice.” For example: | Feed filename | Published | Contributes service slice for | | -------------------------------- | ---------- | ----------------------------- | | mtc-regional-feed-2020-04-24.zip | 2020-04-24 | 2020-04-24 | | mtc-regional-feed-2020-04-23.zip | 2020-04-23 | 2020-04-23 | | mtc-regional-feed-2020-04-22.zip | 2020-04-22 | 2020-04-22 | | mtc-regional-feed-2020-04-21.zip | 2020-04-21 | 2020-04-21 | | mtc-regional-feed-2020-04-20.zip | 2020-04-20 | 2020-04-20 | If the Regional Feed for a given day is missing, the closest previous day provides service. For instance, if 2020-04-22 was missing, the 2020-04-21 feed slice would cover both 2020-04-21 and 2020-04-22. ## Global entity copying Agencies, stops, and routes are considered “global”, and are handled using a simple ID-based merge with the most recent version winning. For example, if BART has a route with ID “OR-S” that is called “Richmond - Warm Springs”, but then later renames it to “Richmond to Warm Springs”, then the latter version will be used. ## Trip hashing, comparison, and copying Trips are more complicated and handled separately. A simple combining of all the trips and stop\_times in all of the input files can easily create a GTFS feed that is too large for practical use, especially given that programs like [OpenTripPlanner](https://www.interline.io/opentripplanner/) need to hold the entire schedule in memory. Therefore, duplicate copies of trips are detected using a hash based approach and only copied to the output once. This reduces the output size by approximately 90%. For example, here are three hypothetical versions of Trip ID “BA:2210503” from three consecutive days of input regional feeds. | Feed filename | Trip ID | Route ID | Service ID | Headsign | 1st stop | 2nd stop | n stops | Hash | | -------------- | ---------- | -------- | ----------------------- | ----------------------------- | --------- | --------- | ------------ | ------ | | 2020-04-24.zip | BA:2210503 | BA:OR-S | BA:Wkd\_BASE-Weekday-07 | Warm Springs/South Fremont | RICH 5:03 | DELN 5:07 | PLZA 5:10... | 8c4ecb | | 2020-04-23.zip | BA:2210503 | BA:OR-S | BA:Wkd\_BASE-Weekday-07 | Warm Springs/South Fremont | RICH 5:03 | DELN 5:07 | PLZA 5:10... | 8c4ecb | | 2020-04-22.zip | BA:2210503 | BA:OR-S | BA:Wkd\_BASE-Weekday-07 | Warm Springs to South Fremont | RICH 5:04 | DELN 5:06 | PLZA 5:12... | a4bf1a | For each of these, the hashing function takes into account all trip attributes, all the calendar attributes for that trip, and the full details of each entry in `stop_times.txt`. Any change in any field will result in a different hash. This allows us to directly compare trips between versions of the input feed. Above, all details and schedule for 2020-04-23 and 2020-04-24 match exactly, so these trips two will be considered identical. The trip for 2020-04-22 has some minor differences in name and schedule, so will generate a different hash, and be considered a different trip. As the historical feed merging program processes each input feed, it calculates the hash of each trip in the feed. If it has not seen a trip before, it copies it to the output and notes the hash for future use. If it has been seen before, it is not copied again. To prevent clashes, the original Trip IDs are appended with the trip hash (e.g. `BA:2210503` \-> `BA:2210503:8c4ecb`). The merging program then takes all trips in the input feed (both seen and unseen) and examines the calendars to see which are active for each day in this slice, and then creates `calendar_dates.txt` entries for each trip on each day where that trip is scheduled to run. The original service IDs are changed to be the same as the hash appended Trip ID, and the calendar is unrolled into a day-by-day format, but it works reliably. This hashing approach is resource efficient and allows us to create historical feeds of arbitrary duration while minimizing the output size. Example output `calendar_dates.txt`: | Service ID | Date | Exception Type | | ----------------- | ---------- | -------------- | | BA:2210503:8c4ecb | 2020-04-24 | 1 (Added) | | BA:2210503:8c4ecb | 2020-04-23 | 1 (Added) | | BA:2210503:a4bf1a | 2020-04-22 | 1 (Added) | In this way, the `8c4ecb` version of the trip is scheduled to run on the two days of input data where it was seen, and the `a4bf1a` version is scheduled to run on the other day. ## Differences between Regional and Historic feeds Historic Regional Feeds are equivalent to the original daily Regional Feeds in the stops, routes, and scheduled services they contain. Using a Historic will produce the same output in a routing engine or another type of analysis. Historic Feeds are different from Regional Feeds in their specific GTFS structure: - `calendars.txt` records are removed and rewritten in `calendar_dates.txt` - `trips.txt` records are hashed and compared (as described above) - IDs for global records are namespaced (as described above) These differences should not affect routing engine or similar types of analysis. However, keep these differences in mind if you are trying to use historical feeds to understand changes in GTFS data and its practices over time at Bay Area agencies. ## Download Historical Feed products To use the Historical Feed products: 1. [Sign up for a 511 Open Data API token](https://511.org/open-data/token?ref=interline.io) 2. Download from `http://api.511.org/transit/datafeeds?api_key=[your_key]&operator_id=RG&historic=YYYY-MM` (for example, to request May 2020: `historic=2020-05`) 3. At the start of each month, the Historic Feed for the last month is created and posted. You may download as many months as you wish and combine them to analyze as many months/quarters/years as you wish at once. As of this blog post, the following months are available for download (that is, you can you any of these values for the `historic` query parameter): - `2020-01` - `2020-02` - `2020-03` - `2020-04` - `2020-05` - `2020-06` - `2020-07` # Station Pathways and Levels The Regional Feed exists to both merge together individual agency GTFS feeds and to serve as a home for new GTFS data that describes cross-agency conditions. The daily Regional Feed now comes with layouts for 35 key transit stations across the Bay Area, each of which serves multiple agencies. Using the newly added `pathways.txt` and `levels.txt` files in GTFS, Interline and MTC are now able to provide more detailed information about how to transfer from one agency to another. This includes information on elevators, escalators, and routes that may not be accessible to those in wheelchairs. Pathways and levels information will equip trip planning apps to provide more helpful wayfinding information to their users, particular those who are new to the Bay Area or who have limited vision. ![Animated views of pathways and levels through the University Ave/Downtown Palo Alto station served by Caltrain regional rail and buses](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/university-ave-pathways.gif) Pathways and levels throughout the University Ave/Downtown Palo Alto station served by Caltrain regional rail on two platforms and VTA and SamTrans buses at bus bays. ## 35 Regional Transit Hubs Here is a list of the 35 stations currently in the Regional Feed: | Transit station/hub | County | | -------------------------------- | ------------- | | Embarcadero BART | San Francisco | | Montgomery BART | San Francisco | | Caltrain Station 4th & King | San Francisco | | Salesforce Transit Center | San Francisco | | 12th St Oakland City Center BART | Alameda | | El Cerrito Del Norte BART | Contra Costa | | 19TH St Oakland BART | Alameda | | Powell ST BART | San Francisco | | Civic Center BART | San Francisco | | Walnut Creek BART | Contra Costa | | Richmond BART/Amtrak | Contra Costa | | San Jose Diridon Station | Santa Clara | | Palo Alto Station | Santa Clara | | San Rafael Transit Center | Marin | | Millbrae BART | San Mateo | | Pleasant Hill BART | Contra Costa | | San Francisco Ferry Terminal | San Francisco | | Daly City BART | San Francisco | | Santa Rosa Transit Mall | Sonoma | | Union City BART | Alameda | | MacArthur BART | Alameda | | Dublin/ Pleasanton BART | Alameda | | Warm Springs/ South Fremont BART | Alameda | | Santa Clara Caltrain | Santa Clara | | Oakland Coliseum BART | Alameda | | SFO | San Francisco | | Fairfield Transportation Center | Solano | | OAK | Alameda | | Petaluma Transit Mall | Marin | | Vallejo Ferry Terminal | Solano | | Mountain View Station | Santa Clara | | Great America | Santa Clara | | Napa Intermodal | Napa | | SJC | Santa Clara | | Great Mall/Milpitas BART | Santa Clara | Note that just as the ongoing pandemic has made our work on the Historic Feed all the more relevant, our work on Station Pathways and Levels has had to adapt to current circumstances. Our staff have been working from home using aerial imagery, architectural drawings, station maps, and other materials. Interline’s Station Editor tool also works on tablets, and we look forward to again going out into the field to all of these stations to correct and improve details. We welcome questions and corrections sent to the [511SFBayDeveloperResources mailing list](https://groups.google.com/forum/?ref=interline.io#!forum/511sfbaydeveloperresources). ## Using Station Pathways and Levels Download the daily Regional Feed like so: 1. [Sign up for a 511 Open Data API token](https://511.org/open-data/token?ref=interline.io) 2. Download from `http://api.511.org/transit/datafeeds?api_key=[your_key]&operator_id=RG` 3. Look for the `pathways.txt` and `levels.txt` files inside the zip archive. 4. For more information, see the [static GTFS documentation](https://gtfs.org/reference/static?ref=interline.io#pathwaystxt) and the [GTFS-Pathways extension proposal document](https://bit.ly/gtfs-pathways?ref=interline.io). # Fares and Transfer Discounts To date, the Regional GTFS Feed has focused on how to plan a journey by transit. Now we’re curating and adding additional data to help riders understand how to pay for their journeys. The Bay Area’s Regional GTFS Feed is now the first in the world to include data using the GTFS Fares-v2 specification. This [newly proposed specification](https://bit.ly/gtfs-fares?ref=interline.io) allows us to model the many different fare products and discounts that are available to the Bay Area’s transit riders. For example, we can now account for how some agencies provide riders a discount when they pay by Clipper Card rather than by cash. ([Clipper Card](https://www.clippercard.com/?ref=interline.io) is the contact-less payment card and system operated by MTC and available on buses, trains, and ferries throughout the Bay Area.) We can also capture how riders can receive a discount when transferring from certain agencies to other agencies. Read on for a detailed example of how Fares-v2 work in practice, or skip ahead to [use the Fares-v2 beta data](https://www.interline.io/blog/mtc-regional-gtfs-feed-additions/#using-fares-and-transfer-discounts). ## An Example of Fares-v2 in Practice ![map of a journey by BART and AC Transit originating in San Francisco, transferring in Oakland, and ending in Alameda](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/bart-to-ac-transit-route.png) A journey by BART and AC Transit, from San Francisco to Oakland to Alameda, which we will use in the following example of how to calculate the total cost. (Credit: BART Trip Planner) Fares-v2 expands the traditional route- and zone-based GTFS fares model with several additional files, each focused on modeling a different part of a complex fare scheme. Imagine a rider making a transfer from BART to AC Transit; both BART and AC Transit have different prices for adult fares, eligible discount fares, and different prices when paying with cash or when using a Clipper Card; additionally, the discount applied when transferring from BART to AC Transit has different values and rules when using Clipper. Calculating the individual fares for each leg of this trip requires a description of any zones that apply to routes and stops (`fare_networks.txt` and `fare_areas.txt`), the costs and requirements of the base fare for each leg (`fare_leg_rules.txt`), any discount categories that might include the rider (`fare_profiles.txt`), and the type of payment used by the rider (`fare_containers.txt`). All of these factors are considered and then used to search for any transfer discounts that may apply for the trip (`fare_transfer_rules.txt`). Let’s work through the hypothetical trip above, first taking BART from Embarcadero Station to 12th St. in Oakland, and then riding an AC Local bus, 51A, to Alameda. The following is an except from `fare_leg_rules.txt` with some columns removed for brevity. This file contains sets of rules that are matched against a single leg of a trip, without regard to transfers. The columns `from_area_id`, `to_area_id`, `fare_container_id` and `fare_category_id` references values defined in the other files mentioned above. | order | leg\_group\_id | from\_area\_id | to\_area\_id | amount | fare\_container\_id | fare\_category\_id | notes | | ----- | -------------- | -------------- | ------------ | ------ | ------------------- | -------------------------- | --------------------------------------------- | | 90 | BA: | EMBR | 12TH | 1.35 | clipper | BA:Senior/Disabled Clipper | Embarcadero to 12th with Clipper and discount | | 100 | BA: | EMBR | 12TH | 3.70 | clipper | | Embarcadero to 12th with Clipper | | 110 | BA: | EMBR | 12TH | 4.20 | | | Embarcadero to 12th with cash surcharge | The BART leg would match the first the rules above based on `from_area_id` and `to_area_id`, with fares of $1.35, $3.70, and $4.20\. However, the first two rules are only available when using a Clipper Card (`fare_container_id`), and the first of these is only available to riders eligible for a discounted fare (`fare_category_id`). The `order` field is used to select the correct fare when multiple rules match the trip leg: any applicable discounted Clipper fare would match first, followed by adult Clipper fare, then finally the fare including additional charge for cash riders. In the event two (or more) rules match with the same order value, both are considered valid options to present to the rider. | order | leg\_group\_id | from\_area\_id | to\_area\_id | amount | fare\_container\_id | fare\_product\_id | fare\_category\_id | notes | | | ----- | -------------- | -------------- | ------------ | ------ | ------------------- | ----------------- | ------------------ | ---------------------------------------- | | | 100 | AC:local | AC:local | AC:local | 2.25 | clipper | | | Adult local fare with clipper $2.25 | | | 110 | AC:local | AC:local | AC:local | 2.5 | | | | Adult local fare $2.5 | | | 100 | AC:local | AC:local | AC:local | 1.12 | | | AC:senior | Discounted local fare $1.12 with Clipper | | | 100 | AC:local | AC:local | AC:local | 1.25 | | | AC:senior | Discounted local fare $1.25 | | | 100 | AC:local | AC:local | AC:local | 0 | | AC:local:monthly | | Free local fare with monthly pass | | The AC Transit fare rules are similar, with the addition of a few extra columns. The 51A is an East Bay local only route (`to_area_id` is assumed), and matches each rule, and the rules for discounts and cash are applied as above. The `fare_product_id` field describes additional rules that are available to users of certain prepaid products, such as day passes and monthly passes. In this case, a rider with an Adult Local 31 Day Pass ($84, described in `fare_products.txt`) would enjoy a free ride. More complex AC Transit fares use additional columns not pictured here (`fare_network_id`) and handle cases such as riding a Transbay bus and the different prices applied when riding within the East Bay vs. taking the trip all the way to the Salesforce Transit Center in downtown San Francisco. Any applicable transfers are calculated by reading the rules in `fare_transfer_rules.txt` and matching each leg of the trip against each subsequent leg of the trip. This can create quite complicated models, but let’s start by looking at our BART and AC Transit legs. The clever bit is that many `fare_leg_rules.txt` rules can share the same `leg_group_id`, which simplifies the number and types of transfer rules that must be defined. For example, all of the BART rules above use `BA:` and all of the AC Transit rules use `AC:local`. The table below is an except of `fare_transfer_rules.txt`. | order | from\_leg\_group\_id | to\_leg\_group\_id | fare\_container\_id | amount | duration\_limit | duration\_limit\_type | fare\_transfer\_type | spanning\_limit | notes | | ----- | -------------------- | ------------------ | ------------------- | ------ | --------------- | --------------------- | -------------------- | --------------- | ---------------------------------------- | | 100 | BA: | AC:local | clipper | \-0.5 | 90 | 2 | 1 | 2 | 1 credit of $-0.50 to AC w/in 90 minutes | | 110 | BA: | AC:local | | \-0.25 | | | 1 | 3 | Cash: 2 fare credits of $0.25 each | Two different discounts apply when taking an AC Transit leg after a BART leg. When using Clipper, a $0.50 (`amount=-0.5`, `fare_transfer_type=1`) discount is applied to the first AC Transit leg (`spanning_limit=2`, which means it can only match a subjourney of two legs). Alternatively, a paper transfer slip given to cash users contains two tabs, each of which applies a $0.25 discount to the AC Transit cash fare (`spanning_limit=3`, or a subjourney of up to three legs). Additionally, the Clipper discount expires within 90 minutes after tagging off BART (`duration_limit=90`, `duration_limit_type=2`). Because the `order` value is different, these two transfer options are mutually exclusive to Clipper and cash riders respectively. The combination of matching fare rules and fare transfer rules allows describing even complicated transfers, such as discounts to SamTrans local bus riders who hold a Caltrain 2 zone or higher monthly pass, the several categories of upgrade charges when transfering from an AC Transit Local to an AC Transit Transbay bus, and the many varied products and transfers that apply full or partial fare to trips on the Dumbarton Express. ## Agencies in the Beta Release Note that this is a “beta” release. The GTFS Fares-v2 specification is still being finalized. It does not fully capture some functionality, like how AC Transit and VTA use “fare capping” so that riders never have to pay more in one day than the cost of a day pass. We also expect that as trip planners and other apps begin to consume the Regional Feed’s fare data that we may need to revise some of the data or our schema. To our knowledge, Interline is the only organization to have a “rules engine” that can calculate journey costs using GTFS Fares-v2 data. We expect that as other organizations adopt the Interline fares rules engine or build their own equivalents, we may need to revisit some of the assumptions built into the engine. The beta release of data includes fares and transfer discounts within and between eight agencies: - BART - SFMTA - AC Transit - SamTrans - VTA - Caltrain - Golden Gate Transit (bus only; not ferries) - Dumbarton Express This list includes the seven largest agencies and a key connector across San Francisco Bay (Dumbarton Express). It should provide a broad enough sample of fares and transfer discounts to power a wide range of potential applications. This sample is also deep enough to inform analyses of fares, the cost of riding transit, and ideally even ways to reform fares and transfer discounts to produce more equitable outcomes for Bay Area transit riders. ## Using Fares and Transfer Discounts Download the daily Regional Feed like so: 1. [Sign up for a 511 Open Data API token](https://511.org/open-data/token?ref=interline.io) 2. Download from `http://api.511.org/transit/datafeeds?api_key=[your_key]&operator_id=RG` 3. For more information about the fares and transfer discount files in the GTFS feed, see the [GTFS-Fares v2 proposal document](https://bit.ly/gtfs-fares?ref=interline.io). Interline and MTC expect to improve both the Fares-v2 beta data and schemas as we learn more. We welcome questions and corrections sent to the [511SFBayDeveloperResources mailing list](https://groups.google.com/forum/?ref=interline.io#!forum/511sfbaydeveloperresources). # Public Transit in 2020 A final note that behind every piece of transit data and every multi-modal trip plan are real people who drive vehicles, serve riders, and maintain equipment and facilities. Thank you to all the front-line workers keeping the Bay Area’s public transit systems and riders moving safely! # Acknowledgements Credit and many thanks to project team members including Ian Rees and Ruth Miller (Interline), Nisar Kapeel and Kapeel Daryani (MTC), and our partners at Bay Area transit agencies. ### A Regional GTFS Feed for the San Francisco Bay Area URL: https://www.interline.io/blog/mtc-regional-gtfs-feed-release/ Last updated: 2024-11-05T00:23:35.000Z 💡 ****September 10, 2021**: MTC and Interline have now released the Regional GTFS **Realtime* Feed, which provides up-to-the-minute service alerts, trip updates, and vehicle positions using the same ID scheme as the static Regional GTFS Feed. See [this blog post for an overview](https://www.interline.io/blog/mtc-regional-gtfs-realtime-feed/). 💡 ****August 10, 2020**: MTC and Interline have released additional functionality in the Regional GTFS Feed. See [this more recent blog post](https://www.interline.io/blog/mtc-regional-gtfs-feed-additions/). Last year, [Interline began a “startup in residence” at the Metropolitan Transportation Commission](https://www.interline.io/blog/metropolitan-transportation-commission-selects-interline/). Now we’re pleased to publicly release our first creation together: the San Francisco Bay Area’s Regional GTFS Feed. Every day, Interline’s workflow processes 31 agency feeds and creates a single unified static GTFS feed. ![animated map of all transit routes in the Bay Area](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/mtc-regional-feed-animation.gif) Transit routes around the Bay Area, first by agency feed, then from the aggregated Regional GTFS Feed. ## Why a Regional Feed? People don’t stop their commutes or their travels at transit agency boundaries. Similarly, many applications require combining multiple agency feeds: trip planners, visualizations, data analyses, travel demand models, etc. ***The Regional GTFS Feed will provide data consumers an easier option.*** Less effort to ingest multiple feeds, more time toward your actual goals. ***The Regional GTFS Feed enables the addition of cross-agency data***: fares, transfer discounts, other transfer information, and transit hub layouts and accessibility information. These types of data reference IDs and entities from multiple agencies and are a natural fit within a consolidated feed. The workflow that creates the Regional GTFS Feed also identifies conflicts between agency feeds, data entities, and IDs. ***By validating agency feeds together, the workflow identifies data quality issues that are only apparent once individual feeds are combined.*** We’re sharing this information with MTC and agency staff to help further improve all feeds. Data consumers are still able to access individual feeds from agency websites and [511.org developer APIs](https://511.org/developers/list/apis/?ref=interline.io). The Regional GTFS Feed is produced using a “conservative” approach: data transformation rules are targeted and constrained. Our motto is *Merge. Don’t lose.* ## Future Plans In 2020, we’re continuing to work with MTC to further enrich the Regional GTFS Feed: - We’re ***mapping key transit hubs*** around the Bay Area, creating GTFS-Pathways data using Interline’s new Gotransit editing tools.[\*](https://www.interline.io/blog/mtc-regional-gtfs-feed-release/#footnote-1) This will help trip planning websites and apps to provide better directions to those who use wheelchairs or have other mobility limitations. - We’re ***cataloging fares, tickets, and transfer discounts around the Bay Area***, creating GTFS-Fares v2 data, also using Interline’s Gotransit editing tools.[\*\*](https://www.interline.io/blog/mtc-regional-gtfs-feed-release/#footnote-2) Transit fares and discounts across agencies is a problem of combinatorial complexity. The Regional GTFS Feed will not be able to precisely capture every possible fare, pass, and transfer discount. The goal is to capture sufficient detail to create realistic fare estimates for a variety of rider types riding and transferring between the agencies of highest ridership. - To start, the Regional GTFS Feed is updated once per day. In the future, we’ll also **create historical versions of the Regional GTFS Feed** that combine “slices in time” from different agency feed versions. While trip planning is best powered by the latest GTFS feed, many types of analyses and travel models instead require archival feed versions from the past. Please watch 511.org’s [Open Transit Data](https://511.org/open-data/transit?ref=interline.io) section for updates as these additions are released. ## Powered by Transitland The Regional GTFS Feed has much in common with [Transitland](https://transit.land/?ref=interline.io), our open data platform that aggregates GTFS data for over 2,500 agencies around the world. We’ve used many of the lessons we’ve learned from Transitland to create the Regional GTFS Feed workflow. In fact, the workflow is build using the core, open-source ingredients of [Transitland version 2](https://transit.land/news/2019/10/17/tlv2.html?ref=interline.io) (also known as “Tlv2”) The 30+ agency feeds are represented as a DMFR (a Distributed Mobility Feed Registry) and processed using the Gotransit library. The Bay Area’s Regional GTFS Feed workflow is, in effect, a Transitland of its own, which is exposed using a static GTFS export. ## To Use the Feed and Ask Questions To use the new Regional GTFS Feed: 1. [Sign up for a 511 Open Data API token](https://511.org/open-data/token?ref=interline.io) 2. Download from `http://api.511.org/transit/datafeeds?api_key=[your_key]&operator_id=RG` 3. The feed is updated daily; download it as frequently as your application needs updates. To provide feedback or ask questions, please visit the [511SFBayDeveloperResources](https://groups.google.com/forum/?ref=interline.io#!forum/511sfbaydeveloperresources) Google group. ## Acknowledgements ![Metropolitan Transportation Commission logo](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/mtc-logo.png) Thanks to our partners in MTC’s Technology Services division. They’ve been serving transit data through 511.org since before GTFS and GTFS Realtime specifications. It’s a pleasure learning from their experience and working together to take advantage of the latest advances in transit data. ![Startup in Residence logo](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/stir-banner-1.png) Thanks to also to [City Innovate](https://www.cityinnovate.com/?ref=interline.io) for managing the Startup in Residence (STIR) program. We encourage procurement officers and civic-oriented startups to learn more about the program. ## Notes \* For more on the GTFS-Pathways additions, see [this Google doc](http://bit.ly/gtfs-pathways?ref=interline.io). \*\* For more on the GTFS-Fares-v2 proposal, see [this Google doc](http://bit.ly/gtfs-fares?ref=interline.io). ### Even more geospatial tools supporting the "GeoJSONL" format URL: https://www.interline.io/blog/here-cli-support-geojsonl/ Last updated: 2024-11-05T22:18:20.000Z Last year we expanded our [OSM Extracts](https://www.interline.io/osm/extracts/) service to produce what we call GeoJSONL. This is a data format that goes by many names: newline-delimited GeoJSON, line-oriented GeoJSON, GeoJSONSeq, or GeoJSON Text Sequences. Whatever its name, the format enables effficient reading and efficient writing of [vector geo-data](https://en.wikipedia.org/wiki/GIS%5Ffile%5Fformats?ref=interline.io#Vector) features. For more background, see our [previous blog post on GeoJSONL](https://www.interline.io/blog/geojsonl-extracts/) and see this new [overview of the format’s many names and its defining properties](https://stevage.github.io/ndgeojson/?ref=interline.io). Over the last year, there’s been even more progress on tools for writing and reading GeoJSONL (or whatever your preferred name is for this format). ## GeoJSONL support in HERE XYZ HERE XYZ now supports GeoJSONL ingest. Don’t know what HERE XYZ is? Here’s how we summarized the platform in a [blog post](https://www.interline.io/blog/here-xyz-open-geo-data/) last year: > HERE XYZ provides the full set of components that "modern" geo developers have come to expect: GeoJSON input and output, "slippy map" libraries, a RESTful API, a JavaScript SDK, a cross-platform CLI, base maps and styles. As we've been experimenting with the HERE XYZ beta, we've found it provides some unique combinations that make it especially well suited to combining open geodata APIs — especially in large quantities. Using the latest version of the [HERE CLI](https://www.npmjs.com/package/@here/cli?ref=interline.io), you can parse and upload GeoJSONL files to the HERE XYZ platform. Here’s an update of our [“Use Crowdsourced Data from Interline OSM Extracts in an XYZ Web Map” tutorial](https://www.interline.io/docs/here-xyz/tutorials/tutorial-1/?index=docs) that uses the GeoJSONL format. Download a GeoJSONL extract from OSM Extracts: [![terminal showing usage of HERE CLI to upload OSM Extract in GeoJSONL format](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/download-geojsonl-extract-1.png)](https://www.interline.io/osm/extracts/) Upload to XYZ and set tags: ![screenshot of a terminal window](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/here-cli-upload-geojsonl.png) Previously you might hit memory limits on your computer if you tried to upload an extract for a big metro region. Now, thanks to GeoJSONL’s efficient parsing, you should be able to upload any extract of any size! ## The growing toolchain There’s a long history of tools that work with this format, and it’s been growing over the last year. Here are some more highlights: **Command line** On the command line, read `.geojsonl` and `.geojsons` files using GDAL’s [drv\_geojsonseq](https://gdal.org/drivers/vector/geojsonseq.html?ref=interline.io). **Desktop GIS** GDAL also powers [QGIS](https://www.qgis.org/?ref=interline.io), so now you can load `.geojsonl` and `.geojsons` files into our favorite desktop GIS sofware. **Parse in Python**: [Fiona](https://github.com/Toblerity/Fiona?ref=interline.io) is a Python library that wraps GDAL and provides a “neat and nimble” interface. The next release will include [support](https://github.com/Toblerity/Fiona/issues/720?ref=interline.io) for `drv_geojsonseq`.[\*](https://www.interline.io/blog/here-cli-supports-geojsonl/#footnote-1) **Statistics and data science in R** If you’re loading vector geo-data into R, try the [geojson](https://cran.r-project.org/package=geojson?ref=interline.io) package on CRAN. See these [instructions](https://github.com/ropensci/geojson?ref=interline.io#newline-delimited-geojson) for how to load a `.geojsonl` file directly from Interline OSM Extracts. **Convert from a shapefile** Install the [shapefile NPM package](https://www.npmjs.com/package/shapefile?ref=interline.io) and run `shp2json --newline-delimited` to convert a shapefile to GeoJSONL. See these [instructions](https://github.com/mbostock/shapefile?ref=interline.io#shp2json%5Fnewline%5Fdelimited). **Convert to vector tiles** [Tippecanoe](https://github.com/mapbox/tippecanoe/?ref=interline.io) can take GeoJSONL as input and turn it in to [MVT vector tiles](https://github.com/mapbox/awesome-vector-tiles?ref=interline.io). It also provides its own `tippecanoe-json-tool` command to convert a standard GeoJSON file to GeoJSONL format (see [these instructions](https://github.com/mapbox/tippecanoe/?ref=interline.io#tippecanoe-json-tool)). **Stream in to a web browser** Watch a GeoJSONL file stream into a web in this [interactive JavaScript notebook](https://observablehq.com/@bcamper/streaming-geojsonl?ref=interline.io).[†](https://www.interline.io/blog/here-cli-supports-geojsonl/#footnote-2) If you have an Interline OSM Extracts API key and an Observable account, try [this notebook](https://observablehq.com/@drewda/streaming-geojsonl-from-your-choice-of-openstreetmap-extr?ref=interline.io) to select from any of the 200+ cities and watch the stream arrive: [![Observable notebook streaming OpenStreetMap data in GeoJSONL format](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/streaming-geojsonl.gif)](https://observablehq.com/@bcamper/streaming-geojsonl?ref=interline.io) ## When to try GeoJSONL If you like GeoJSON but find your data files are too big to hold in memory, give one of these tools a try. If you need to run a bunch of vector geo-data through a [MapReduce](https://en.wikipedia.org/wiki/MapReduce?ref=interline.io) operation that, give this format a try. (It’s easy to divide a GeoJSONL file in to chunks before parsing its exact contents.) If you want to download [OSM Extracts](https://www.interline.io/osm/extracts/) from Interline and explore even the largest of extracts using HERE XYZ, give this a try. --- #### Notes \* Shoutout to [Sean Gillies](https://github.com/sgillies?ref=interline.io), the lead creator of Fiona and instigator of [RFC 8142 - GeoJSON Text Sequences](https://tools.ietf.org/html/rfc8142?ref=interline.io). We may have differing views on the [record separator character](https://en.wikipedia.org/wiki/Delimiter?ref=interline.io#ASCII%5Fdelimited%5Ftext). Regardless, his packages like Fiona and [Rasterio](https://rasterio.readthedocs.io/?ref=interline.io) are fabulous and we’re happy users of them. † Shoutout to [Brett Camper](http://vector.io/?ref=interline.io), the author of this Observable notebook and a fellow alum of [Mapzen](https://mapzen.com/?ref=interline.io). ### Transportation Research Board funds an open platform for transit agencies to improve the quality of their real-time data URL: https://www.interline.io/blog/transportation-research-board-funds-gtfs-realtime/ Last updated: 2024-11-05T00:46:32.000Z 💡 ****January 19, 2021**: Interline and CUTR@USF have now published a Transportation Research Board report on this project. See [this more recent blog post](https://www.interline.io/blog/trb-transit-idea-gtfs-realtime-report/). We’re excited to share that the [Transportation Research Board](http://www.trb.org/?ref=interline.io) has awarded a [Transit IDEA Program grant](http://apps.trb.org/cmsfeed/TRBNetProjectDisplay.asp?ProjectID=4695&ref=interline.io) to Interline. We’re using this unique opportunity to expand [Transitland](https://transit.land/?ref=interline.io) to support [GTFS Realtime](https://developers.google.com/transit/gtfs-realtime/?ref=interline.io) feeds. Our goal is for Transitland to provide a suite of validation tools for public-transit agency staff to validate the contents and quality of their real-time data feeds. ## High-quality real-time data is hard but worthwhile ![the problem statement](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/trb-idea-the-problem.png) Real-time transit information has many benefits to transit riders and agencies, including shorter perceived and actual wait times, a more welcoming experience for new riders, and an increased feeling of safety, and increased ridership. Real-time transit data is, in comparison with many other potential operational or capital improvements to bus or rail service, an affordable means of increasing ridership. In the last few years, a real-time complement to the General Transit Feed Specification (GTFS) format, GTFS Realtime, has emerged. Despite its promise, adoption of GTFS Realtime by transit agencies has been hampered by a lack of clear documentation and readily available validation tools. Hardware and software vendors for automatic vehicle location (AVL) systems also vary in the depth and quality of their support for the GTFS Realtime specification. As a result, transit agencies must invest significant time and effort to create and maintain high quality GTFS Realtime feeds. Furthermore, bad data has been shown to have a negative effect on ridership, the rider’s opinion of the agency, and the rider’s satisfaction with multimodal trip-planning apps. Thus, transit agencies must put even more effort toward ensuring that the GTFS Realtime data produced by their systems is of sufficient quality. ## Our Solution: GTFS Validation in Transitland In this project, Interline and the [Center for Urban Transportation Research](https://www.cutr.usf.edu/?ref=interline.io) are creating a prototype platform that makes GTFS Realtime validation tools readily available to, potentially, all transit agencies in North America. We will build upon two open-source projects: the [GTFS Realtime validator prototype](https://github.com/CUTR-at-USF/gtfs-realtime-validator?ref=interline.io) and Transitland, an open transit data platform. This project will apply the open-source and open-data community models to the challenges of creating and improving GTFS Realtime data. By combining the open-source components of a GTFS Realtime validator with a catalog of GTFS Realtime feeds, hosted on Transitland’s cloud servers, this project will make the process of validating real-time data simple and accessible to agency staff from any computer with a web browser. As a result, GTFS Realtime data will improve in quality and availability. Transit riders will have a better experience (which has been linked to higher ridership), agency staff will provide better service with less effort and cost, and system vendors will provide higher quality product. ## Our Partner: Center for Urban Transportation Research We’re glad to have on our project team an expert on GTFS Realtime: Sean Barbeau. Sean and his colleagues at the University of South Florida’s Center for Urban Transportation Research have been involved in GTFS, GTFS Realtime, and related transit data specifications and systems for many years. Most recently, they have created a GTFS Realtime validator prototype. We’ll be working together to test and deploy this validator tooling within the Transitland platform. For more about their preliminary findings about the quality of GTFS Realtime feeds, see [this recent summary of their TRB paper](https://medium.com/@sjbarbeau/introducing-the-gtfs-realtime-validator-e1aae3185439?ref=interline.io). Here are the most common errors and warnings their prototype validator found in a sample of GTFS Realtime feeds: ![most common errors found in a sample of GTFS Realtime feeds](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/validator-test-errors.png) Image Credit: Sean Barbeau, Transportation Research Board 2018 paper [18–05585 "Quality Control — Lessons Learned from the Deployment and Evaluation of GTFS-realtime Feeds"](https://trid.trb.org/view/1496848?ref=interline.io). ## About the Transit IDEA program Here’s an overview of the [Transit IDEA Program](http://www.trb.org/IDEAProgram/IDEATransit.aspx?ref=interline.io): > The Transit IDEA Program is part of the Transit Cooperative Research Program, a cooperative effort of the Federal Transit Administration (FTA), the Transportation Research Board (TRB), and the Transit Development Corporation (a nonprofit educational and research organization of the American Public Transportation Association). The program is funded by the FTA and is managed by TRB. > > The Transit IDEA Panel has established four high-priority focus areas to encourage proposals in the following areas:Increasing transit ridershipImproving transit safety, security, and emergency preparednessImproving transit capital and operating efficienciesProtecting the environment and promoting energy independence. IDEA stands for “Innovations Deserving Exploratory Analysis” — which we think is an apt description for the Transitland platform providing GTFS Realtime feed validation to agency staff! ## An open call to transit agency staff If your agency provides a GTFS Realtime feed, we welcome you to participate in the user-testing phase of this project. Please write us at [info@interline.io](mailto:info@interline.io) for more information. ### Metropolitan Transportation Commission and Interline to tackle regional transit data challenges URL: https://www.interline.io/blog/metropolitan-transportation-commission-selects-interline/ Last updated: 2024-11-05T21:56:18.000Z 💡 ****January 8, 2020**: MTC and Interline have released the Regional GTFS Feed. See [this more recent blog post](https://www.interline.io/blog/mtc-regional-gtfs-feed-release/). 💡 ****August 10, 2020**: MTC and Interline have released additional functionality in the Regional GTFS Feed. See [this more recent blog post](https://www.interline.io/blog/mtc-regional-gtfs-feed-additions/). We’re excited to share that the [Metropolitan Transportation Commission has selected Interline](https://mtc.ca.gov/whats-happening/news/mtc-work-startups-tackle-transportation-challenges?ref=interline.io) through a competitive process to participate in its 2019 Startup in Residence (STIR) program. ## Regional Transit Data for the Bay Area Interline and MTC will collaborate on the problem of regional transit data. The Bay Area has over 30 public-transit operators. Most agencies provide open-data in the [GTFS format](http://gtfs.org/?ref=interline.io), which is in turn used by many companies and developers to power navigation and transportation apps. However, to use all of these GTFS feeds together presents many technical challenges to both public agencies and private consumers. During the STIR residency, Interline will use the [Transitland](https://transit.land/?ref=interline.io) platform to help the MTC improve regional transit data distribution. Transitland currently aggregates GTFS feeds from approximately 2,500 transit operators around the world. Transitland users — including startups, large corporations, advocacy groups, and government agencies — currently query Transitland’s open APIs five million times per month. Transitland is powered by open-source software, and can be customized by Interline for particular agency needs. Transitland currently aggregates transit data for the San Francisco Bay Area from 36 operators: [![Transitland Feed Registry listing 36 operators in the San Francisco Bay Area](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/feed-registry-bay-area.png)](https://transit.land/feed-registry/?metro=San%20Francisco%20Bay%20Area&ref=interline.io) For a current list of Bay Area transit agencies that provide open GTFS data, see the [Transitland Feed Registry](https://transit.land/feed-registry/?metro=San%20Francisco%20Bay%20Area&ref=interline.io). ## Startup in Residence Program ![Start up in Residence program logo](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/stir-banner-2.png) The national STIR program is supported by the [City Innovate Foundation](https://www.cityinnovate.com/?ref=interline.io). “The Startup in Residence program is a model for civic innovation and national collaboration,” says Jay Nath, former Chief Innovation Officer for San Francisco and Executive Director for City Innovate. “This program is a unique opportunity for government agencies and startups to think creatively about how we can all work together to modernize government to benefit residents.” For more on STIR, see [City Innovate’s announcement of the 2019 program](https://medium.com/@CityInnovate/700-startups-compete-to-work-on-urban-challenges-through-startup-in-residence-55052fd781ea?ref=interline.io). ## Metropolitan Transportation Commission ![Metropolitan Transportation Commission (MTC) logo](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/mtc-logo-1.png) The [Metropolitan Transportation Commission](https://mtc.ca.gov/?ref=interline.io) (MTC) is the transportation planning, financing, and coordinating agency for the nine-county San Francisco Bay Area. MTC is based out of downtown San Francisco. MTC provides [511.org](https://511.org/?ref=interline.io) with transportation information for travelers around the Bay Area and [511 Developer Resources](https://511.org/developers/list/resources/?ref=interline.io) to power third-party apps and websites. ## Transitland and Interline ![Transitland logo](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/transitland_logo-1.png) Transitland is powered by open-source software and open data. Started in 2014 by Mapzen, [Transitland is now supported by Interline and a growing list of partners](https://transit.land/news/2019/01/08/transitland-continues.html?ref=interline.io). To follow Interline’s progress on this project, follow this blog or [subscribe to our newsletter](http://eepurl.com/dmWHln?ref=interline.io). Interline’s team of principals and specialists all live in the Bay Area and ride public transit regularly. We’re pleased to have this opportunity to collaborate with the MTC on improvements to our region and its transportation systems! ### Transitland Version 2 URL: https://www.interline.io/blog/tlv2/ Last updated: 2024-11-05T21:50:59.000Z 💡 ****March 4, 2021**: Here's a more recent blog post [demonstrating Transitland v2's new ability to search across transit operators, feeds, and routes](https://www.interline.io/blog/transitland-v2-search/). 💡 ****February 25, 2021**: Here's a more recent blog post [demonstrating Transitland v2's new global transit map and vector tiles](https://www.interline.io/blog/transitland-v2-vector-map-tiles/). It's been a quiet year on this blog, and we're overdue to share all the work that's been going on behind the scenes since the [last post](https://www.interline.io/blog/transitland-continues/). What first started as two simple, separate, little experiments has grown into the seeds of a major revision of the entire Transitland platform: Transitland Version 2. ## Tlv2 Functionality and Goals *Tlv2* (as we've come to call this new version) provides continuity on the most used parts of the Transitland platform. Tlv2 will still provide a canonical service at transit.land that includes a Feed Registry of open GTFS data feeds from around the world, a queryable API, and a range of user interfaces. Tlv2, like Tlv1, will be powered by open-source components. Tlv2's API will continue to be accessible without charge, and most of the Tlv1 API will continue to be available during the transition. (As before, we will enforce rate limits to ensure no one user hogs too many server resources; more on how rate limits will change below.) What we are changing in Tlv2 is in support of three goals: - Higher Performance - More Modular Architecture - Distributed Deployments ## Goal #1: Higher Performance Tlv1 successfully parses and imports GTFS feeds from thousands of transit agencies. However its "FeedEater" component slows down or halts on a handful of the very larger GTFS feeds — typically single feeds that cover entire countries. This has always been a constraint on Tlv1. Interline's new `transitland-lib` (tlib) library solves this problem. Transitland-lib is a new transit data library written in the [Go programming language](https://golang.org/?ref=interline.io), which provides finer-grained control of memory usage, as well as much easier patterns for concurrency and parallelization of complex tasks. Transitland-lib replaces replaces FeedEater, and will be used for feed validation, pre-processing, and importing data into the Tlv2 database. Transitland-lib is efficient enough to handle the largest feeds in a small memory footprint, and is fast enough to encourage interactive use; imports that took hours with FeedEater are completed in minutes. Transitland-lib began as an experiment by Ian Rees over the last winter holidays, and it now complete enough to power multiple Interline projects, from the San Francisco Bay Area Regional GTFS Feed (which merges \~30 feeds into one) to a transit-to-routing-engine pipeline that Interline built for another client (which covers transit service for an entire country!). We look forward to releasing transitland-lib under an open-source license shortly. Transitland's most demanding API users have also hit performance limitations over the years. As part of Tlv2, we'll be providing bulk data dumps that give users another option. Use the Tlv2 API to explore data or power a small web app; use Tlv2 data extracts to power a large analysis, a giant web app, or a country-wide routing engine. ## Goal #2: More Modular Architecture Tlv1 has succeeded in part due to its "monolith" architecture. FeedEater, the API, and administrative operations all live within the same [Transitland Datastore](https://transit.land/documentation/datastore/?ref=interline.io) codebase. This has allowed for efficient reuse of data models, helpers, and test code. The single Datastore codebase is also straightforward to deploy and configure for the canonical transit.land site. The Datastore codebase can be reused — some have in fact redeployed it for their own regional or internal GTFS management needs—but customizing and redeploying the Datastore is not straightforward. For Tlv2, we're breaking the Transitland Datastore into modular components. Some components will be in Go (data processing), some will remain in Ruby (general back-end) and JavaScript (user interfaces). The canonical transit.land deployment will continue to use all. Other deployments by Interline, our partners, or open-source users may only use a subset. This more modular architecture will also allow deployments that connect Tlv2 components to internal data pipelines, routing engines, and other tools that may not be relevant to the canonical transit.land deployment. While the software components of Tlv2 will be more modular, they will still be glued together using the data schemes developed for Transitland, including [Onestop IDs](https://transit.land/documentation/concepts/onestop-id-scheme?ref=interline.io) and a common feed format spec (described below). Additionally, we will be sharing our database schema in versions appropriate for PostgreSQL and SQLite to help simplify additional application development, including outside of the canonical transit.land. ## Goal #3: Distributed Deployments The Feed Registry at transit.land will continue to be our canonical list of GTFS feeds. However, some deployments of Tlv2 components may want to use just a subset of the Feed Registry or a private list of feeds. We also want to open up the canonical Feed Registry for more collaborative tending, beyond just staff from Interline and Trillium Transit. The second of the two experiments that has led to Tlv2 is called the [Distributed Mobility Feed Registry](https://github.com/transitland/distributed-mobility-feed-registry?ref=interline.io). DMFR (for short) is a standard way of representing a list of GTFS feeds as one or more JSON files. It also currently supports GTFS-RT, and we expect it to extend to GBFS (bike rental) and MDS (e-scooter rental) in the future. (Some history: This experiment actually started way back in the early tests of Transitland, before the fully built out Tlv1 monolith. It's still [archived](https://github.com/transitland/transitland-datastore?ref=interline.io) on GitHub. Now we've come full circle to revisit and build out this approach.) Soon the `transitland/distributed-mobility-feed-registry` repository on GitHub will become the source of truth for the Transitland Feed Registry of public open data feeds. We'll accept pull requests against the JSON files to add or edit feeds. This should substantially lower the burden for outside contributors to contribute new feeds and help manage existing records. However, this will be just be one repository of DMFR files. In effect we're blowing up the single Transitland Feed Registry and inviting the creation of many. Transitland-lib can manage a datastore using one or more DMFR files, meaning that the canonical transit.land, distributed instances of Tlv2, and temporary interactive sessions with transitland-lib on the command line can all share the same representation for mobility data feeds. At Interline, we've gone one step further and also use DMFR files to configure some of our routing engines and any other workflow that takes two or more mobility data feeds as input. We welcome other consumers or transformers of GTFS, GTFS-RT, and related data to try a DMFR file for their own needs. ## Tlv2 Architecture Diagram Now that we've introduced and discussed the new components of Tlv2, here's a diagram of how they connect: ![Tlv2 architecture diagram](https://www.transit.land/images/tlv2/tlv2-architecture-diagram.png) ## Gradual Deprecation of Tlv1 Tlv1 requires a fair amount of attention and resources to maintain. While developing Tlv2, we've already had to split our attention. As we shift our focus to deploying and supporting Tlv2, we'll be gradually deprecating Tlv1\. Don't worry — it won't be too quickly. **Step 1: Tlv2 API Preview** The initial release of Tlv2 will include a preview of a new [GraphQL](https://graphql.org/?ref=interline.io) API that provides flexible, structured access to transit data. All new data will be available through this API from the start. API access will require a key. Sign-up will be free. We will institute tight rate limits at first and increase them over time, as we better understand the type and complexity of queries that users run. **Step 2: Deprecate Tlv1 Schedule API** Immediately upon Tlv2 release, we will stop importing new [schedule data](https://transit.land/documentation/datastore/schedules.html?ref=interline.io) into Tlv1\. The `ScheduleStopPair` data model is the most expensive part of Tlv1 to run and maintain. We'll first stop imports, keeping the API available against the last imported version. We'll share tutorials on how to switch to the Tlv2 GraphQL API to access schedules. Then, after sufficient notice, we'll empty old Tlv1 schedule data and remove that API endpoint. **Step 3: Other Tlv1 Entity API Endpoints** While the Tlv2 GraphQL API will make a number of complex queries much simpler than with the current Tlv1 API, it may take a while for us to make all Tlv1 query options available in the Tlv2 API. During this time, we'll try to continue to power the Tlv1 APIs (with the exception of the schedules endpoint). We'll recommend that you build any new applications on the Tlv2 API, in case we do have to announce any specific incompatibilities during the Tlv1-to-Tlv2 API transition. We'll post updates to this site. Or watch [@transitland](https://twitter.com/transitland?ref=interline.io) and [@interline\_io](https://twitter.com/interline%5Fio?ref=interline.io) for updates on Twitter. ## How You Can Help Here are ways you can help with the rollout of Tlv2: - Provide your input on Tlv2 by taking our [survey and signing up for the beta tester list](https://docs.google.com/forms/d/e/1FAIpQLSdmhW28zcP1kBMBkmelAt7OQGqKjRcoILjsbZvGXLsGfmFNGw/viewform?ref=interline.io). - Comment on the [DMFR GitHub repository](https://funding.communitybridge.org/projects/transitland?ref=interline.io). Contribute edits once we migrate over the Feed Registry to DMFR files. - If you work for a GTFS producing transit operator or mobility provider, [contact Interline](mailto:info@interline.io) to learn more about "Tlv2 on premise" and regional GTFS feeds. It's consulting projects like this that help us serve both mobility providers and further the development of transitland-lib. - Contribute financially to the upkeep and growth of Tlv2 through the Linux Foundation's [CommunityBridge donation page for Transitland](https://funding.communitybridge.org/projects/transitland?ref=interline.io). These contributions go directly toward upkeep of the canonical transit.land. ### Transitland Continues in 2019 URL: https://www.interline.io/blog/transitland-continues/ Last updated: 2024-11-01T22:47:27.000Z Just like a garden, open-source and open-data projects require continual tending. During 2018, we quietly maintained the core of Transitland's code and data. We also worked to put in place new partnerships and structures to support Transitland after the [shuttering of Mapzen](https://mapzen.com/blog/shutdown/?ref=interline.io). With 2019 beginning, we're glad to share more widely these plans for Transitland's ongoing tending and growth as an open platform. ## Core Support and Maintenance We have a combination of organizations that are together supporting Transitland: - Transitland's open-source code is moving to the [Linux Foundation](https://www.linuxfoundation.org/?ref=interline.io), allowing code use and contributions by any and all. - Amazon Web Services continues to provide servers and access for Transitland's API users, thanks to awards from the [Earth on AWS Cloud Credits for Research Program](https://aws.amazon.com/earth/research-credits/?ref=interline.io). - [Interline Technologies](https://www.interline.io/) (the firm that I am now a part of) has been donating resources to maintain code and is now also offering professional services and support to the companies that depend on Transitland APIs. - [Trillium Solutions](https://trilliumtransit.com/?ref=interline.io), a team of experts at GTFS production, is helping to curate submissions to the Transitland Feed Registry and maintain existing records. Previously Mapzen was a single "keystone" supporter for Transitland. Now we have a growing list of "tentpole" supporters for Transitland's code, servers, and data. Together in 2018, we quietly shifted Transitland on to a new generation of infrastructure, fixed pressing bugs, and gradually increased feed coverage. Transitland now aggregates data for 2,466 transit operators across 53 countries! Each month, the API handles 5 million queries! ![bar chart of Transitland operators by country](https://www.transit.land/images/transitland-continues/transitland-operators-by-country-2019.png) *Operators records in Transitland, by* [*two-letter country code*](https://en.wikipedia.org/wiki/ISO%5F3166-1%5Falpha-2?ref=interline.io)*.* Now that Transitland again has a stable foundation of support, post Mapzen, we can return to feature development, major data additions, and community outreach. ## Growth and Development in 2019 We're working together with a series of partners and sponsors on new Transitland platform features and data additions in 2019\. The list of topics tentatively includes: - [GTFS-realtime support](https://www.interline.io/blog/transportation-research-board-funds-gtfs-realtime/) - feed validation and best-practices report cards - [feed merging](https://www.interline.io/blog/metropolitan-transportation-commission-selects-interline/) - "Transitland Tiles" for bulk download and analysis - website and documentation updates - an experimental, distributed replacement to the Transitland Feed Registry **Technical folks**: As you've likely guessed, many of these new experiments require a level of performance that Transitland's current mix of scripting languages cannot always provide. My colleague [Ian Rees](https://www.interline.io/about/), who is Interline's principal engineer, is leading an effort to re-architect key Transitland components in the high-performance Go programming language. Transitland's technical architecture has always had the goal of accessibility — now we're pursuing both accessibility *and* performance. **Professional users of Transitland APIs**: We've heard from many of you in 2018 that new functionality is nice, but that dependable service is even more important. Interline is now preparing professional services and support around the Transitland APIs. [Sign up for the Interline newsletter](http://eepurl.com/dmWHln?ref=interline.io) for related announcements or contact [info@interline.io](mailto:info@interline.io) with questions. Note that these commercial services supplement and do not change existing free access to the Transitland APIs. **All Transitland Community Members**: Thanks for sticking with us over the past year. Despite all the work happening behind the scenes, we weren't able to share full news of Transitland's progress publicly. We're excited to return to our previous cadance of public announcements and conversations. Watch this website for updates. We also welcome questions and participation from potential partners in the above topics and others — please write to [hello@transit.land](mailto:hello@transit.land). Here's to a new year of open transit data! ### Bringing Back the Spirit of Map Mashups: Combine Open Geodata APIs using HERE XYZ URL: https://www.interline.io/blog/here-xyz-open-geo-data/ Last updated: 2024-11-05T22:17:55.000Z At Interline, we spend much of our time with open geodata APIs. We create our own open geodata APIs (like [Transitland](https://transit.land/?ref=interline.io)) and we use others (for example, whenever we need to throw a bunch of GeoJSON on a base map). At their best, open geodata APIs are fun — you get to combine pieces together to create something new. Then again, public APIs can have many limitations — combine together a few and you can end up with the worst of all worlds: duplicate data, misregistered overlays, a slow user experience, incompatible data formats… With all these mixed experiences in our minds, we’ve been playing with the initial beta of [HERE XYZ](http://explore.xyz.here.com/?ref=interline.io). It’s a brand new developer service from one of the oldest names in geodata. HERE XYZ provides the full set of components that “modern” geo developers have come to expect: GeoJSON input and output, “slippy map” libraries, a RESTful API, a JavaScript SDK, a cross-platform CLI, base maps and styles. As we’ve been experimenting with the HERE XYZ beta, we’ve found it provides some unique combinations that make it especially well suited to combining open geodata APIs — especially in large quantities. Combining geodata is nothing new. [Ian McHarg](http://ladprofile.weebly.com/ian-mcharg-1920-2001.html?ref=interline.io), a prominent environmental designer, printed maps on transparent plastic sheets and overlaid them atop each other, inspiring the idea of layers in GIS ([geographic information systems](https://en.wikipedia.org/wiki/Geographic%5Finformation%5Fsystem?ref=interline.io)): ![photo of Ian McHarg](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/ian-mcharg.jpg) *Image credit: University of Pennsylvania* More recently, Google Maps brought us *mashups*. At first, it involved hacking the Google Maps JavaScript code in order to overlay [crimes in Chicago](https://www.nytimes.com/2005/12/11/magazine/doityourself-cartography.html?ref=interline.io) and [Craigslist house listings](http://www.housingmaps.com/?ref=interline.io): ![screenshot of old housingmaps.com](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/housingmaps2.png) *Image credit: housingmaps.com* Neither hacking nor reverse-engineering are required any more to combine multiple sources of geodata in a web browser. Companies like Google, Esri, Mapbox, Carto — also a now defunct start-up where we previously worked, called Mapzen — all offer accessible APIs for creating web maps. We’ve also been freed of many of the constraints on dataset size in 2005 mashups. Original mashups involved drawing points and other geographic features on top of base-map images, with a low limit on the number of these vector features that could be drawn at once. More recent techniques send both the base map and overlays as vector tiles. Instead of an image, a vector tile is effectively a square of data, which can be pre-generated by a server. Many vector tiles are then loaded and drawn by your web browser simultaneously. We’ve left behind the constraints of mashups — and it also feels like we’ve forgotten the fun of mashups. Vector tiles help greatly with performance, but often add complexity to the overall set of steps required to combine together existing and new data sources on a web map or another type of visualization. Our experiments with the HERE CLI and the XYZ API are bringing back that creative spirit of 2005 mashups. The XYZ platform generates vector tiles near instantly for any new imported data sources. The HERE CLI and XYZ API provide a range of ways to import GeoJSON files, connect to APIs, or stream in from other CLI tools. It’s the best of 2005 and 2018\. If you use the SDKs to style your web maps to look like the transparency overlays in Ian McHarg’s book *Design With Nature*, it’s the best of 1969 proto-GIS, too. We’re collaborating with the HERE XYZ team to turn our experiments into tutorials that demonstrate how Interline’s open geodata services can be combined, mixed, filtered, analyzed, and visualized. In the first tutorial, you’ll fetch and filter data from [Interline OSM Extracts](https://www.interline.io/osm/extracts): [![Honolulu web map animation](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/xyz-tutorial-1-animation-2.gif)](https://www.interline.io/docs/here-xyz/tutorials/tutorial-1-web-map/) In the second tutorial, you’ll query the [Transitland](https://transit.land/?ref=interline.io) API repeatedly to build up an analysis and a web map of just not subway locations but also the frequency of how often each subway line operates. (Especially when working with geodata for transportation applications, both space and time matter.) We’ll show you how to do this for the entire planet: [![web map animation zooming into Chicago, showing "L" routes](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/xyz-tutorial-2-animation-3.gif)](https://www.interline.io/docs/here-xyz/tutorials/tutorial-2-web-map/) The sample projects and tutorials take intermediate and advanced users on a tour through some of the unique features of XYZ and the HERE CLI: - filtering Interline OSM Extracts using OpenStreetMap tags - writing scripts that stream geodata from external APIs into an XYZ space using the HERE CLI - using multiple API queries to aggregate together properties about the same features (identified by unique IDs, such as [Transitland Onestop IDs](https://transit.land/documentation/onestop-id-scheme/?ref=interline.io)) - quickly building global tile sets We’re experimenting with additional features, for creating a new generation of mashups, and we look forward to sharing more soon. Now, get started with our [first two tutorials on combining HERE XYZ, Interline OSM Extracts, and Transitland](https://www.interline.io/docs-here-xyz/). ### GeoJSONL: An optimized format for large geographic datasets URL: https://www.interline.io/blog/geojsonl-extracts/ Last updated: 2024-11-05T00:47:26.000Z 💡 ****March 8, 2019**: We've published a follow-up blog post: [Even more geospatial tools supporting the "GeoJSONL" format](https://www.interline.io/blog/here-cli-supports-geojsonl/). Now you can stream GeoJSONL files up to HERE XYZ, down to a web browser, or between your own programs using a growing list of software libraries. [GeoJSON](http://geojson.org/?ref=interline.io) is a popular format for exchanging geographic data. It builds on the success and flexibility of JSON to describe structured data, and simplifies many of the more complicated aspects of traditional GIS file formats. Initially developed through an ad hoc process, GeoJSON is now formalized by [IETF RFC 7946](https://tools.ietf.org/html/rfc7946?ref=interline.io), with a [working group](https://datatracker.ietf.org/wg/geojson/charter/?ref=interline.io) to guide future additions. The universality of the underlying JSON format and the [WGS84](https://en.wikipedia.org/wiki/World%5FGeodetic%5FSystem?ref=interline.io) coordinate system make GeoJSON nearly ideal for web-based mapping projects. However, very large data sets are common in GIS, and the structure of a GeoJSON file generally requires the entire file to be read into memory and decoded all at once. A large file (e.g. 1 Gb) can easily exhaust the resources of a desktop computer. While streaming JSON parsers exist, they are generally not part of standard libraries and introduce yet another dependency and complexity. One potential solution to the problem of large JSON files is to restructure the file as an array of objects, rather than having a single root object. Because the JSON specification prohibits the newline `\n` character from being used in string literals without escaping, it opens the possibility of having one JSON object per line in a way that is compatible with both standard UNIX text-processing tools and can be read line-by-line with existing JSON parsers. This convention has arisen independently in many contexts, and is known as [Newline Delimited JSON](http://ndjson.org/?ref=interline.io) (`ndjson`), [JSON Lines](http://jsonlines.org/?ref=interline.io) (`jsonl`), or [GeoJSON Text Sequences](https://tools.ietf.org/html/rfc8142?ref=interline.io).[\*](#footnote-1) For GeoJSON specifically, the root level `FeatureCollection` object is replaced with a simple array of features, one per line. This file can then be read line-by-line, feature-by-feature, and easily integrated with other tools that use newline-delimited records such as [GNU parallel](https://www.gnu.org/software/parallel/?ref=interline.io). Newline-delimited GeoJSON (GeoJSONL) has been casually [proposed](https://macwright.org/2015/03/23/geojson-second-bite.html?ref=interline.io#streaming) several times before, and is natively supported by some tools such as the [Osmium](https://docs.osmcode.org/osmium/latest/osmium-export.html?ref=interline.io) `export` command, Mapbox’s [Tippecanoe](https://github.com/mapbox/tippecanoe?ref=interline.io), and [jq](https://stedolan.github.io/jq/?ref=interline.io), the Swiss-army knife of JSON processors. At Interline, our [OSM Extract service](https://www.interline.io/osm/extracts/) now provides these files as well as traditional GeoJSON and OSM PBF formats, with the goal of providing a simple basis for reading and filtering OSM extracts without the need for more complicated and specialized libraries. Here is an [example GeoJSONL](https://storage.googleapis.com/osm-extracts.interline.io/honolulu%5Fhawaii.geojsonl?ref=interline.io) file, which you can compare to [regular GeoJSON](https://storage.googleapis.com/osm-extracts.interline.io/honolulu%5Fhawaii.geojson?ref=interline.io). While the file sizes are nearly identical, reading and parsing the entire \~27 Mb GeoJSON file requires several seconds and \~240 Mb of memory. Iterating through the GeoJSONL file line-by-line takes a similar amount of time, but uses negligible amounts of memory: ```python import json with open('honolulu_hawaii.geojsonl') as f: for feature in f: print json.loads(feature) ``` The above Python snippet also demonstrates that no additional dependencies are required to read the GeoJSONL file; a line iterator and parsing each row separately works well. Still, [libraries are available](http://ndjson.org/libraries.html?ref=interline.io) for many languages that provide native support. As mentioned above, GeoJSONL also provides excellent input for traditional text-processing tools. The following (contrived) example runs `jq` to find the frequency of `name`s for features with the `highway` tag, without loading the entire file into memory: ```sh # cat honolulu_hawaii.geojsonl | jq 'select(.properties.highway!=null) | .properties.name' | sort | uniq -c | sort -n ...snip... 36 "State Highway 83" 49 "John A. Burns Freeway" 59 "Makakilo Drive" 106 "Farrington Highway" 107 "Kamehameha Highway" ``` Given the increasing complexity and importance of geographic data in nearly every domain, and with increasing file sizes, GeoJSONL is useful and timely. We’re glad to now support the format in [Interline OSM Extracts](https://www.interline.io/osm/extracts/) and we encourage other tools to offer native support when possible. --- #### Notes \* Minor differences exist between `ndjson` and `jsonl`. GeoJSON Text Sequences adds a record separator (RS) character, in addition to the newline used by `ndjson` and `jsonl`. Based on these existing options, we recommend the use of `ndjson`, which has a [simple spec document](https://github.com/ndjson/ndjson-spec?ref=interline.io) and allows the presence of blank lines. This supports the widest range of tools for parsing and consuming. We use `geojsonl` as the file extension, although this is just a matter of taste. ### Practices that help complex and collaborative open-source projects survive and thrive URL: https://www.interline.io/blog/open-source-best-practices-for-otp-summit/ Last updated: 2024-11-05T00:42:36.000Z At Interline, we specialize in deploying, creating, and managing the growth of open-source software. It’s less an individual matter of principle than a set of shared advantages: our clients build on each other’s advances, and we collaborate productively with a wide range of partners. However, the health of an open-source project is not a given — it requires ongoing attention and thought. In this blog post, we’ll share an overview of practices used by a wide variety of open-source projects: **Useful practices** - 👍 [Clean licenses](#clean-licenses) - 👍 [Technical steering committees](#technical-steering-committees) - 👍 [RFC (request for comments) processes](#rfc-request-for-comments-processes) - 👍 [Feature flags](#feature-flags) - 👍 [“Contrib” vs. core](#contrib-vs-core) - 👍 [Automated acceptance test and performance suites](#automated-acceptance-test-and-performance-suites) - 👍 [Automated dependency checks](#automated-dependency-checks) - 👍 [Release management](#release-management) - 👍 [Codes of conduct](#codes-of-conduct) **Unclear practices** - 🤷 [“The roadmap”](#the-roadmap) **Bad outcomes** - 👎 [“The eternal rewrite”](#the-eternal-rewrite) - 👎 [“The big switchover”](#the-big-switchover) ## OpenTripPlanner: An open-source project going on 10 years ![OpenTripPlanner logo](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/otp-logo.png) This blog post comes out of a presentation Interline gave at last month’s OpenTripPlanner Summit in Boston. None of the practices are particular to that project. However, some background will help explain how relevant many of these open-source practices are to OpenTripPlanner and similar open-source projects that have become critical — but also somewhat neglected — pieces of infrastructure for many organizations. The OpenTripPlanner routing engine (also known as OTP) [was created in 2009](http://docs.opentripplanner.org/en/latest/History/?ref=interline.io) by TriMet (the public-transit agency in Portland, Oregon), OpenPlans (a civic tech non-profit), and individual experts and enthusiasts. Since then, OTP has grown to be used by public-transit agencies, planning firms, and transportation researchers [around the world](http://docs.opentripplanner.org/en/latest/Deployments/?ref=interline.io). (Interline itself provides managed hosting of OTP for a variety of public-transit agencies around North America.) And yet the OTP project rests on an unstable foundation. Many individuals and organizations contribute to developing new features and expanding OTP’s reach, but few are involved directly in maintenance, coordination, and the other “housekeeping chores” of an open-source project. One outcome of last month’s OpenTripPlanner Summit is that more organizations will be collaborating on this “housekeeping.” Interline is joining the [OTP Project Leadership Committee](http://docs.opentripplanner.org/en/latest/Governance/?ref=interline.io), alongside Cambridge Systematics, Conveyal, TriMet, Ruter (the public-transit agency of Oslo), the University of South Florida, the PlannerStack Foundation, and the Helsinki Regional Transport Authority. Now, with this context in mind, let’s zoom out and consider a wide range of practices that complex and collaborative open-source projects use to grow and thrive. ## Clean licenses ### Choosing a (single) license At the core of an open-source project is its license. This is what legally enables use and contribution to the project by many individuals and organizations. The Open Source Initiative [catalogs hundreds of licenses and highlights a short list of the most popular options](https://opensource.org/licenses?ref=interline.io): [![screenshot of list of open-source licenses](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/os-licenses.png)](https://opensource.org/licenses?ref=interline.io) Different licenses can encourage different outcomes. OpenTripPlanner uses the LGPLv3 license, which requires that any organization directly customizing OTP code publicly share their changes. However, the LGPLv3 license does allow organizations to use OTP within the context of larger proprietary systems. ### Using dual licenses Some projects are released under dual licenses. For example, the Sidekiq project is available for free under LGPLv3\. Organizations that require more flexibility (for example, to customize code and keep their changes private), may pay the primary creator of Sidekiq for an alternative license: [![the Sidekiq project is available under two licenses, shown in this screenshot](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/sidekiq-dual-licenses.png)](https://sidekiq.org/?ref=interline.io) The dual license model can have trade-offs. It may provide resources to build a business around an open-source project, and it may provide a predictable means for buying support for large users. However, it can be off-putting or at least confusing to smaller users who assume “open-source = completely free.” Clear communication is important. Consider the case of the ExtJS project, which [switched to a dual license in 2008](https://en.wikipedia.org/wiki/Ext%5FJS?ref=interline.io#License%5Fhistory), alienating many of its users. ### Contributor license agreements The presence of a license file in an open-source code repository is a meaningful signal, but it isn’t always sufficent reassurance to businesses that the code is free of legal risk. Some open-source project sponsors ask all contributors to sign a [contributor agreement](http://contributoragreements.org/?ref=interline.io) before their code is merged into the open-source project. Typically, these agreements ask the “signer” to say that they personally created whatever code it is that they are contributing and that they are not knowingly contributing materials for which they don’t hold rights. Some contributor agreements go further, asking contributors to sign over intellectual property rights to the primary company that maintains the project. At a minimum, such an agreement provides all companies involved with some reassurance that contributors aren’t intentionally adding others’ code or private materials. The Linux Foundation no longer recommends broad contributor agreements; instead they offer the simple [Developer Certificate of Origin](https://developercertificate.org/?ref=interline.io). Open-source projects that do need to enforce one of these agreements can add a “bot” to GitHub or a similar source-code repository. Here’s [a bot to enforce the Developer Certificate of Origin](https://github.com/apps/dco?ref=interline.io) and [a bot to enforce a custom contributor license agreement](https://cla-assistant.io/?ref=interline.io): [![screenshot: the Developer Certificate of Origin is enforced using this GitHub app](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/dco-app.png)](https://github.com/apps/dco?ref=interline.io) Code licenses and contributor license agreements are necessary but not sufficient for an open-source project to succeed. They set the incentives and the bounds for collaboration and growth. ## Technical steering committees Licenses allow contribution, but who will accept the contributions? Many projects appoint a technical steering committee (or TSC) to make such decisions. For example, here are the [responsibilities of the Node.js TSC](https://github.com/nodejs/TSC?ref=interline.io#list-of-tsc-responsibilities): - **Managing code and documentation creation and changes** for the listed projects and resources - **Setting and maintaining standards** covering contributions of code, documentation and other materials - **Managing code and binary releases**: types, schedules, frequency, delivery mechanisms - Making decisions regarding dependencies of the Node.js Core project, - including what those dependencies are and how they are bundled with source code and releases - **Creating new repositories and projects** under the nodejs GitHub organization as required - **Setting overall technical direction** for the Node.js Core project, - including high-level goals and low-level specifics regarding features and functionality - **Setting and maintaining appropriate standards for community discourse** via the various mediums under TSC control - **Setting and maintaining governance rules** for the conduct and make-up of the TSC, Working Groups and other bodies within the TSC's domain The [history of the Node.js TSC](https://en.wikipedia.org/wiki/Node.js?ref=interline.io#History) is useful to consider. Development of Node.js had nominally been overseen by a single company, named Joyent. When outside contributors to the project found their contributions and their specific technical needs not addressed by Joyent staff, [they forked the Node.js code](https://www.javaworld.com/article/2855639/open-source-tools/qanda-why-io-js-decided-to-fork-node-js.html?ref=interline.io) (that is, they copied the code) and started a new effort. This new project, called io.js, was formed around a TSC and an “open” governance process. After much negotiation, Joyent and the io.js TSC combined their efforts under the newly formed Node.js Foundation. The code improvements made by this “splinter” effort found their way back into the Node.js code base. More importantly, the new governance structure continued. Consider another example: the OneBusAway project. It’s governed by [a steering committee](https://onebusaway.org/the-onebusaway-project/project-members/?ref=interline.io) that is structured to include a mix of representatives from the private sector, the public sector, universities, and independent individuals. As these two examples demonstrate, the primary challenge of a TSC is ensuring that it has an appropriate representation across the open-source project’s contributors and users. If a TSC has too few members or members that are only concerned with low-level software development concerns, the project will miss opportunities to learn from less technically sophisticated users. Alternatively, with a TSC that is too broad and includes individuals or organizations that are not deeply invested in the project, the risk is that the TSC will have many ideas but little action. ## RFC (request for comments) processes Technical steering committees (TSCs) can put thoughts into actions by various means. Let’s look at the RFC (that is, request for comments) process now. Later on, we’ll also consider the ”[roadmap](https://www.interline.io/blog/open-source-best-practices-for-otp-summit/#the-roadmap)“. The Internet and its protocols [were created and continue to mature using RFCs](https://en.wikipedia.org/wiki/Request%5Ffor%5FComments?ref=interline.io#History). RFCs are documents that put into words all of the ingredients of a proposed technical change. They are both speculative and precise. They are always pedantic (many RFCs begin by defining exactly what the authors mean by “must” and “must not”) and sometimes humorous (well, humorous in the eyes of fellow engineers). An RFC is a structured format for proposing new software development and honing plans based on feedback. The Rust programming language has [a well defined RFC process](http://rust-lang.github.io/rfcs/?ref=interline.io); it’s been used as a model by many other open-soure projects. For the Rust project, contributors must write an RFC for: - Making “Substantial changes” that are not bug fixes - Removing existing features - Changing interfaces between components RFCs are not required for: - Refactoring, reorganizing - “Additions that strictly improve objective, numerical quality criteria (warning removal, speedup, better platform coverage, more parallelism, trap more errors, etc.)” - “Additions only likely to be noticed by other developers-of-rust, invisible to users-of-rust.” When drafting a new RFC, contributors are asked to include: - Summary (one paragraph) - Motivation - “Guide-level explanation”: features, API input/output, example use-cases - “Reference-level explanation”: technical implementation details - Drawbacks: “why should we not do this?” - Rationale and alternatives - Why is this design the best in the space of possible designs? - What other designs have been considered? What is the rationale for not choosing them? - What is the impact of not doing this? - Prior art - Unresolved questions An RFC process front-loads the collaborative work of major changes or additions to an open-source project. That is, instead of waiting until a pull request arrives with lines of code changes, discussion begins early. Contributors are required to put in effort up-front to help others in the project community understand their proposed changes. The project’s TSC and its extended developer community are invited to provide input within the structure of the RFC. RFCs do introduce more overhead to a project. A developer can no longer just assume a code contribution speaks for itself; they do also have to put their ideas into words and arguments. However, the RFC process is often a net gain for open-source projects of a certain size and complexity. Any individual or organization is invited to contribute to the project’s growth — they are just asked to work within a structure. ## Feature flags Once an RFC has been “green lighted” by a TSC (or whatever approval critiera is used on a particular project), there are the questions of when and how to merge the related code changes and additions into the project. The sooner code is merged, the easier it will be for others to test and integrate with their own work-in-progress. The later code is merged, the more polished it may be. However, there is also the unfortunate reality of open-source development: The sooner code is merged, the more likely it will introduce bugs and needlessly expand future “surface area” requiring maintenance. The later code is merged, the more likely the developer will disappear, having lost interest in seeing their contribution reach fruition. One practice that strikes a balance between “too soon” and “too late” is to use feature flags. New additions are merged into a code base before they are ready for general use. Users can only access one of these new features by setting a configuration “flag” to opt-in. Feature flags provide a means to stage the introduction of major changes and additions. Code changes are merged early and often; functionality changes are hidden from most users until later. The feature-flag practice is also used within many web sites and applications. Facebook selectively rolls out feature changes to certain numbers and types of users, much like how junk food manufacturers test their new sodas and chips in small but representative markets. The [EmberJS project uses feature flags](https://guides.emberjs.com/v3.2.0/configuring-ember/feature-flags/?ref=interline.io) together with an RFC process. Each RFC typically turns into code hidden behind a feature flag. Advanced users can opt-in to the feature in early releases. After further testing and improvement, the feature flag will be removed and all users will use that code in future releases. (Later, we’ll also discuss [release management](https://www.interline.io/blog/open-source-best-practices-for-otp-summit/#release-management), which helps pace this process of introducing new functionality to a wider audience.) ## “Contrib” vs. core RFCs and feature flags work for contributions that are relevant to all users of a project, but what if a contribution is only relevant to some? An alternative approach is to divide an open-source project into two zones: “core” and “contrib” (that is, contributed). The standards for acceptance for changes to “core” are high — everyone depends on the stability and performance of these components. The standards for “contrib” are lower — these components may offer expanded functionality, but they come with fewer promises of current performance and future maintenance. “Contrib” is an area for experimentation, while “core” is more carefully managed through RFCs and similar types of structure. Here’s how the Drupal CMS differentiates between its core and contributed modules: [![screenshot of Drupal documentation](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/drupal-contrib-modules.png)](https://www.drupal.org/docs/7/understanding-drupal/general-concepts?ref=interline.io) Note how Drupal also refers to custom modules. In the case of OTP, the reason for our OpenTripPlanner Summit last month is that many potential code contributions to OTP have been accumulating in branches and forks, but few are being merged back into the main OTP repository’s master branch. In effect, these branches and forks are all custom modules — without a clear path to entry to a “contrib” or a “core.” To succeed, the core/contrib split does require certain conditions. First, the open-source project must be modular and expandable. That is, “contrib” code should be able to extend or modify aspects of “core” code. Otherwise, all changes would need to happen directly within “core.” (This is a challenge for many projects, OTP included.) Second, “contrib” is [a poor shorthand](https://blog.startifact.com/posts/against-contrib.html?ref=interline.io) for “code developed by non-core committers.” The core/contrib split is best used as a dividing line between the level of commitment a project has to certain functionality, rather than a dividing line between classes of developers. ## Automated acceptance test and performance suites Whether code is going behind a feature flag, into “core”, or into “contrib,” there should be a process for verifying its quality before it is merged. This is important in any type of software development effort; it’s especially important for open-source development, where contributions may be coming from developers of various skill levels, using various tools and development environments. Many open-source projects require that new code contributions also come with automated tests. These tests use sample data and conditions to demonstrate that new functionality behaves as expected. Automated tests help prevent regressions (a future change breaking current functionality). Tests also help document functionality, for other developers to understand a system’s goals. There are many types of software tests: [unit tests](https://en.wikipedia.org/wiki/Unit%5Ftesting?ref=interline.io), [integration tests](https://en.wikipedia.org/wiki/Integration%5Ftesting?ref=interline.io), [smoke tests](https://en.wikipedia.org/wiki/Smoke%5Ftesting%5F%28software%29?ref=interline.io), [GUI testing](https://en.wikipedia.org/wiki/Graphical%5Fuser%5Finterface%5Ftesting?ref=interline.io), [performance testing](https://en.wikipedia.org/wiki/Software%5Fperformance%5Ftesting?ref=interline.io#Testing%5Ftypes), security vulnerability scanning, and [accessibility testing](https://www.w3.org/wiki/Accessibility%5Ftesting?ref=interline.io). The specific test types will depend on the type of project. Different mixes of testing approaches are useful for command-line interface scripts, web applications, mobile applications, desktop GUIs, or embedded system modules. An open-source project should aim to run automated tests for two overall goals: acceptance and performance. An acceptance test suite tells developers that a proposed code change works as expected and doesn’t break existing functionality. A performance test suite warns developers when proposed code changes will have dramatic changes on speed or resource consumption. For some systems, a single set of tests can inform both acceptance and performance. For other systems, acceptance and performance are more easily tested as separate steps, before and after the piece of software is built or compiled. Acceptance and performance tests can be automatically run against every proposed code change by a build/CI/CD server. (CI stands for continuous integration; CD stands for continuous delivery.) We especially like [CircleCI](https://circleci.com/?ref=interline.io) at Interline, and will be blogging more about this soon. Acceptance and performance tests often produce a green ✅ or a red ❌ — that is, they give a “yes” or “no” answer to the question of whether a proposed code change is good to merge. However, good test suites also produce more detailed information for ongoing use by a TSC and the regular contributors to an open-source project. Here’s a code coverage report for the Transitland Datastore, the web application at the center of the [Transitland](https://transit.land/?ref=interline.io) open public-transit data platform. This report is generated by the acceptance test suite. We occasionally review it to ensure that acceptance tests are covering a growing portion of the code bases: ![screenshot of Transitland Datastore test code coverage](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/transitland-test-coverage.png) For programming languages and other heavily used software infrastructure, small differences can make major unexpected performance differences. The [RubyBench](https://rubybench.org/?ref=interline.io) project regularly evaluates hundred of actions in the Ruby scripting language and its most important packages. The same tests are run against new releases, enabling comparison with past versions. Developers can use the website to, hopefully, see performance increasing over time: [![screenshot of Ruby on Rails benchmark reports over many versions](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/rubybench.png)](https://rubybench.org/rails/rails/releases?result%5Ftype=activerecord/mysql2%5Fpluck&ref=interline.io) All software projects benefit from testing infrastructure. Testing infrastructure that is easy to access and use is all the more important for open-source software. Learning to read and write tests is an important step in the education of a software engineer. For some, an open-source project will be the first time they encounter auomated test infrastructure. Even for experienced developers, seeing their proposed change produce the red ❌ of a failing change may be the first occasion they have to dig deeply into an unfamiliar code base. Writing and running tests does not need to feel like flossing, a boring but necessary chore. When test suites are well structured and made easily accessible, testing helps improve everyone’s confidence in an open-source project and to reduce the risk of contributing. ## Automated dependency checks Most pieces of open-source software are built upon others. These are software dependencies. Sometimes, dependencies are just copies of code. Hopefully, dependencies are included using a package manager. Package managers make it simple to upgrade dependencies, when new versions are released. Just as open-source projects benefit from making automated test suites accessible to developers, projects can benefit from automating and making accessible the process of checking for new dependencies. Here’s an email I received a while back from a service called Gemnasium, reporting that a new version of a dependency is available and fixes a known security vulnerability: ![screenshot of a dependency security alert](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/gemnasium-update.png) [Gemnasium has now been acquired by Gitlab](https://docs.gitlab.com/ee/user/project/import/gemnasium.html?ref=interline.io), where they are beginning to offer a similar automated service. GitHub now also offers its own dependency updates. Some services, like [Greenkeeper](https://greenkeeper.io/?ref=interline.io), will even perform updates automatically, without waiting for a developer to do so on their own computer. Out-of-date dependencies are one of the main ways that software projects open themselves to security vulnerabilities. Closely monitoring dependencies to know when they need to be updated is an important but often forgotten task for open-source projects. Having an automated means of doing so can help turn this from a difficult chore to an easily shared test (similar to automated test suites). However, automated dependency checks do come with wrinkles: First, updates about dependencies with security vulnerabilities need to be acted upon quickly. Otherwise, a security alert is a glaring sign saying “come take advantage of this software!” Second, the process of updating dependencies relies on good release habits within those projects. How do you know what has changed within a new version of a dependency? How do you know if any of the changes will require your own project to change? Let’s consider release management next, since it’s through release management that open-source projects can provide stability to build atop each other. ## Release management The goal of release management is to produce a new version of a software package, ready for use. For commercial software, release management can be as grand an effort as the launch of a Hollywood movie. For open-source software, release management can often be an afterthought. As with all these other practices, having a structured and consistent approach to release management can be a great benefit to an open-source project. ### Release manager(s) Who performs the work of producing a new release? For every open-source project that has grown beyond a single core contributor, it’s worth naming a release manager or a release team. This person or people will deserve credit and thanks when the work of producing new releases goes well — and when producing a new release is a challenge, the release manager/team will need sufficient authority to make decisions to resolve problems. On the Kubernetes project, there is a “special interest group” (or SIG) focused on release management. [Their responsibilities](https://github.com/kubernetes/sig-release/blob/master/release-team/README.md?ref=interline.io) include: - Authority to build and publish releases at the published release date under the auspices of the CNCF - Authority to accept or reject cherrypick merge requests to the release branch - Authority to accept or reject PRs to the kubernetes/kubernetes master branch during code slush period - Changing the release timeline as needed if keeping it would materially impact the stability or quality of the release these responsibilities will continue to be discharged by SIG release through the Release Team. This charter grants SIG Release the following additional responsibilities - Authority to revert code changes which imperil the ability to produce a release by the communicated date or otherwise negatively impact the quality of the release (e.g. insufficient testing, lack of documentation) - Authority to guard code changes behind a feature flag which do not meet criteria for release - Authority to close the submit queue to any changes which do not target the release as part of enforcing the code slush period Note how the Kubernetes release team uses [feature flags](https://www.interline.io/blog/open-source-best-practices-for-otp-summit/#feature-flags) and other tools we’ve discussed previously: They act in a reactive manner, doing whatever is necessary to ship a release. This is in contrast with our previous discussion of [technical steering committees (TSCs)](https://www.interline.io/blog/open-source-best-practices-for-otp-summit/#technical-steering-committees) and how they can use tools like feature flags in a proactive manner to pace development. Ideally, an open-source project is able to have both proactive planning (to build up code and features that are worthy of release) and reactive release management (to ensure that releases do happen consistently and regularly). ### Release checklist A release manager/team should have a checklist that they follow each time they “cut” a release. In the short term, this checklist helps make the process of producing a new release easier to repeat. In the long term, this checklist makes it simpler to add new members to a release team and let burnt-out members depart, without losing the “secret” steps of how to produce a release. ### Release notes Once a release has been produced, how will users understand what’s included and what changed? Ideally there is a release notes document that includes this information. It can be as simple as a bulleted list (also known as a change-log) or as thorough as an illustrated guide. Each month, the Visual Studio Code project cuts a new release and includes illustrated release notes, which are also posted to [their blog](https://code.visualstudio.com/updates/v1%5F25?ref=interline.io). For the Transitland Datastore, the web application at the center of the [Transitland](https://transit.land/?ref=interline.io) open public-transit data platform, we use [a semi-automated process](https://github.com/github-changelog-generator/github-changelog-generator?ref=interline.io) to generate [a change-log from GitHub issues and pull requests](https://github.com/transitland/transitland-datastore/blob/master/CHANGELOG.md?ref=interline.io). The level of detail in release notes or change-log depends in large part on the audience. Visual Studio Code’s release notes are prepared for both end-users and plug-in developers, all of whom need detail. Transitland Datastore’s change-log is prepared for other developers, who are at least somewhat familiar with the existing project and its APIs. ### Semantic versioning Typically, each software release gets a version number. The number typically increases with each release. However, there are different ways of formatting and increasing version numbers. This can serve as a signal to users of what and how much changed within the software. [Semantic Versioning](https://semver.org/?ref=interline.io) (a.k.a. SemVer) is one approach to version numbers. Here’s an overview: Given a version number MAJOR.MINOR.PATCH, increment the: 1. MAJOR version when you make incompatible API changes, 2. MINOR version when you add functionality in a backwards-compatible manner, and 3. PATCH version when you make backwards-compatible bug fixes. Additional labels for pre-release and build metadata are available as extensions to the MAJOR.MINOR.PATCH format. This signals to users that, for example, upgrading from 2.4.1 to 2.5.0 won’t require them to make any changes to their own systems. However, upgrading from 2.4.1 to 3.0.0 should only be done after reading release notes and testing the new software; a function they depend upon may have changed. Semantic Versioning is a useful ingredient for [automated dependency checks](https://www.interline.io/blog/open-source-best-practices-for-otp-summit/#automated-dependency-checks). If you trust that your dependencies are faithfully following Semantic Versioning principles, you can automatically upgrade to new releases that just increment the “minor” or “patch” version number. Ideally, your project is also using [automated acceptance tests](https://www.interline.io/blog/open-source-best-practices-for-otp-summit/#automated-acceptance-test-and-performance-suites), in order to catch any unexpected side effects that arise from upgrading a dependency. ### LTS (long-term support) versioning Even if there are advantages to users in upgrading software as soon as possible, there are also costs. Some users will never be able to keep up with the latest version of a project, and may still have support needs and questions about their outdated versions. One way out of such situations is to label certain releases as having long-term support (or LTS). LTS releases will be maintained and support for a longer time than other releases. Bug fixes and security fixes made in future releases may also be “backported” so that they are also available to users of an LTS release. The Ubuntu operating system project [releases](https://wiki.ubuntu.com/LTS?ref=interline.io) new versions every 9 months. Every 2 years, one of these releases is marked as an LTS release. This LTS release is supported for 5 years. ### “Train” release cycles What should actually go into a release? That question is quite dependent on project specifics. Some projects actually try to avoid answering the question. Instead they set up a “train” release cycle, where releases happen on a consistent schedule. Whatever functionality is ready to release when the next date arrives is included in that release. Here’s an illustration of how the EmberJS framework releases beta and stable versions every 6 weeks: [![Ember.js release cycle website screenshot](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/ember-release-cycle.png)](https://www.emberjs.com/builds/?ref=interline.io) ## Codes of conduct Our discussion of [release management](https://www.interline.io/blog/open-source-best-practices-for-otp-summit/#release-management) hinted at some of the problems that can arise in open-source projects. For example: disagreement over features that make it into a release, or unexpected problems merging together code from multiple contributors. Our discussion of [technical steering committees](https://www.interline.io/blog/open-source-best-practices-for-otp-summit/#technical-steering-committees) also hinted at the challenges of many people working together on a complex technical project. In the short term, tempers can flare. In the long term, projects can become exclusive clubs, inadvertently sending signals to potential contributors that they are not welcome. These days, every open-source project should seriously consider adopting a code of conduct. It’s a proactive means of welcoming new contributors, and it’s a means of providing some “guardrails” for existing contributors to keep in mind. The following codes of conduct are widely used and referenced: - [Contributor Covenant](https://www.contributor-covenant.org/?ref=interline.io) - [Citizen Code of Conduct](http://citizencodeofconduct.org/?ref=interline.io) - [Geek Feminism community anti-harassment policy](http://geekfeminism.wikia.com/wiki/Community%5Fanti-harassment/Policy?ref=interline.io) ## “The roadmap” So far we’ve avoided mentioning the “roadmap.” A roadmap is a forward-looking plan, ordering the development of new functionality and assigning dates to future releases. Developers often find the process of writing roadmaps to be fun and inspiring — the focus is on the potential of tomorrow, rather than the bug fixes or the maintenance tasks of today. Users often find that reading roadmaps provides clarity and confidence — the roadmap tells them what further functionality they can expect, if they invest in committing to use a software package. However, a roadmap does not necessarily represent reality. Projects that are regularly unable to meet their self-imposed targets can lose whatever inspiration and clarity were generated by publishing a roadmap. By committing to a firm roadmap, open-source projects can also give up some of their room for flexibility that can be a useful advantage. Open-source projects are often more engineering-driven than proprietary projects. That is, developers and designers who are “closer to the code” are responsible for more decision making. This may provide more opportunities for unexpected optimizations and creativity during implementation. Roadmaps that are overly detailed can hinder opportunistic creation of new functionality and refactoring. In reaction to these shortcomings, some open-source projects never write roadmaps. For example, the ReactJS framework: [![screenshot of a ReactJS issue thread discussing the lack of a project roadmap](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/react-roadmap.png)](https://github.com/facebook/react/issues/11645?ref=interline.io) Some projects use [“train” release cycles](https://www.interline.io/blog/open-source-best-practices-for-otp-summit/#train-release-cycles) to effectively replace roadmaps. For example, on the Django project, whatever additions are completed by a certain date are considered part of the “roadmap” for a release: ![screenshot of the Django 2.1 roadmap webpage](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/django-2-1-roadmap.png) Finally, some projects find a mid-point. They may use a roadmap to define major themes for future releases, but avoid providing fine details about implementation and functionality that will likely be subject to change. They may use a roadmap to order development priorities and releases, but avoid attaching specific dates. They may use a roadmap to help divide duties among multiple organizations for release management and other maintenance tasks, but avoid listing any further goals for each release. When writing roadmaps, it’s too easy to be overconfident in the future of an open-source project. Before committing to a roadmap, consider if some of the other practices we’ve discussed could help achieve similar goals, with more room for flexibility along the way. After considering other practices, if there are still questions that remain that would be best answered by a roadmap-style plan, start with your roadmap only providing a “skeleton” for the most important aspects of the project to plan out in advance. ## “The eternal rewrite” We’ve discussed many positive practices of open-source projects. Let’s end by considering some unfortunate outcomes. First, consider how many open-source projects encounter problems during major version transitions. A 2.0 release promises to fix all current problems and provide great improvements. But 2.0 still hasn’t arrived, so users continue to work with and improve 1.0 in the meantime. Bug reports keep arriving on 1.0, distracting developers from completing 2.0\. It’s unclear what version new users should use, and there are no clear ways for new contributors to help. Unfortunately this is what’s happened to the ActiveModelSerializers project, with parallel 0.8, 0.9, and 0.10 versions in various states of usage and completion: [![screenshot of the ActiveModelSerializers project readme](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/ams-version-development.png)](https://github.com/rails-api/active%5Fmodel%5Fserializers?ref=interline.io) Consider “the eternal rewrite” situation. What of the practices that we discussed could help a project avoid this fate? What of the practices could help a project like this recover and move forward? (As they say in math textbooks, we’ll leave this as an exercise to the reader.) ## “The big switchover” Another unfortunate outcome to consider for a version change: “the big switchover.” That is, the new version is available, but major changes leave a large portion of the user community behind. For example, this is effectively a decade-long process for users and developers to switch from Python 2.x to 3.x: [![screenshot of a report on Python 2 and Python 3 usage levels over time](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/python-2-to-3.png)](https://www.webucator.com/blog/2016/03/still-using-python-2-it-is-time-to-upgrade/?ref=interline.io) For the Angular project, 1.x and 2.x were so different that some users and developers have never made the leap and continue on 1.x, as if it were a completely different project: [![Angularjs 1.0 and Angular 2 logos](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/angular-1-to-2-1.png)](https://www.concettolabs.com/blog/comparative-study-angular-js-1-angular-js-2/?ref=interline.io) “The big switchover” is a good problem to have. It means that a project is so widely used that it has entrenched users. Then again, it may also show a lack of applying some of the open-source practices that we discussed previously. How could a project have better engaged its users, its contributors, and its organizers in the process of building and migrating to a new version? ## A holistic view of open-source For a blog post about open-source software, we’ve talked surprisingly little about software engineering or about the business of software. That’s by design. There’s lots to read online about software engineering (The users of [Stackoverflow](https://stackoverflow.com/?ref=interline.io) will try to answer most any question you have). There’s also lots to read about the business of open-source software (even if it’s grand-standing when describing successes and similarly overconfident when judging failures). What is often missing are recipes for how to organize and scale open-source projects that have had enough success to get into trouble. That is, projects like OpenTripPlanner that have many users and interested contributors, but a lack of structure to take full advantage of that attention and channel it toward sustainable growth and maintenance. Software does still need to get written, and individuals and organizations do need to earn their income — however, those are necessary but not sufficient ingredients for project success. At Interline, we depend upon many open-source projects to succeed. Our clients do, too. That’s why we take this holistic view when evaluating, advising, and contributing to open-source projects. We think it helps improve the odds of success for an open-source project, beyond just code and cash. ### Producing 200 OpenStreetMap extracts in 35 minutes using a scalable data workflow URL: https://www.interline.io/blog/scaling-openstreetmap-data-workflows/ Last updated: 2024-11-05T00:43:33.000Z Interline now offers [OSM Extracts](https://www.interline.io/osm/extracts/), a service enabling software developers and GIS professionals to download chunks of OpenStreetMap data for 200 major cities and regions around the world. OSM Extracts is simple. It’s a 3 step process that you can run yourself. We’ll describe these steps, as well as how we run those 3 steps for all 200 extracts in parallel on Interline’s cloud infrastructure. ![screenshot of OSM Extracts website](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/osm-extracts-screenshot.png) A screenshot of [OSM Extracts by Interline](https://www.interline.io/osm/extracts/). ### Creating OSM Extracts in 3 steps Every day at Interline our servers download the latest updates from [OpenStreetMap](https://www.openstreetmap.org/?ref=interline.io), update our local copies of the planet file, and process this data to generate a variety of geospatial data products. Reliably updating a planet file requires a bit of legwork behind the scenes, which we have encapsulated as helper scripts in our [PlanetUtils](https://github.com/interline-io/planetutils?ref=interline.io) library. This simplifies the process down to 3 steps: 1. Run `osm_planet_update` to download the most recent OSM planet file (released weekly at [planet.openstreetmap.org](https://planet.openstreetmap.org/?ref=interline.io) and other mirrors) and then apply minutely “diffs” to update the planet file to the current point in time. 2. Fetch a copy of [cities.json](https://github.com/interline-io/osm-extracts/blob/master/cities.json?ref=interline.io) file from GitHub. This GeoJSON file contains polygons defining the boundaries of each extract, and is a continuation of the extract definitions created by over 100 contributors for the Mapzen Metro Extracts service. The file is open to additions and revisions by all. 3. Run `osm_planet_extract`, specifying the local copy of the cities.json file, to generate extracts for each of the cities/regions. Behind the scenes, PlanetUtils handles calls to `osmosis` and `osmconvert` to update the planet and generate the extracts. For more information on each of these commands, see the [PlanetUtils readme](https://github.com/interline-io/planetutils/blob/master/README.md?ref=interline.io), or install the package yourself using Homebrew or Docker. This workflow is simple but slow. Updating the planet and generating all 200 extracts on a single machine can take more than a day. In addition to the extracts, the updated planet is the starting point for additional data products, such as our [Valhalla Tilepacks](https://www.interline.io/valhalla/tilepacks/). Valhalla Tilepacks combine the OSM planet file with 1.6Tb of elevation data to produce downloadable data for organizations to run their own instances of the [Valhalla routing engine](https://www.interline.io/valhalla/). All of this heavy lifting requires a substantial mix of computing resources. Parallel workflows to the rescue! ### Managing our OSM Extracts workflow ![screenshot of Argo workflow dashboard](https://storage.ghost.io/c/57/10/5710b77b-3fc0-4094-a25f-875ff38c9df1/content/images/2024/11/argo-workflow-screenshot-1.png) OSM Extracts and Valhalla Tilepacks are all generated using an Argo workflow. This is screenshot from Argo's management dashboard. To simplify the image, only 2 extract jobs are shown. At Interline, we have a great appreciation for [Kubernetes](https://kubernetes.io/?ref=interline.io) which powers much of our infrastructure. Kubernetes is ideal for running highly available services and managing complex resources. We wanted our data workflow system to leverage this infrastructure and introduce as little additional complexity as possible. With these constraints, we found that [Argo Workflow Manager](https://applatix.com/open-source/argo/?ref=interline.io) was a great fit for our needs: - Each Argo workflow step is a [Kubernetes pod](https://kubernetes.io/docs/concepts/workloads/pods/pod/?ref=interline.io) and [Docker container](https://www.docker.com/what-container?ref=interline.io), simplifying resource requests and access to secrets (e.g., passwords and access keys for external services). - Workflows can be described as a directed graph, with an easy syntax for specifying dependencies between steps. - Argo provides a flexible method for specifying output files and parameters that can be fed into subsequent steps. - Parallel workflows are well supported, and one step can “fan out” to a number of parallel tasks (as you can see in the diagram above). Our Argo workflow resembles the simple 3-step process with a few key adjustments: 1. Update our local planet using `osm_planet_update` and copy to our cloud storage bucket. 2. Fetch the `cities.json` file, and divide the GeoJSON features into smaller chunks of about 8 extracts each. 3. Run each chunk as a parallel task to quickly generate all 200 extracts. Because each Argo task is defined as a Kubernetes pod, we have very fine-grained control over the resources allocated to each task. The Kubernetes scheduler is very effective at auto-scaling your cluster based, quickly increasing cluster size to run all tasks in parallel, and just as importantly, reducing the cluster size when the tasks are complete. Additionally, because some tasks are CPU intensive and others are memory intensive, we define multiple node pools to ensure a high utilization of cluster resources. Currently, we run about 25 tasks of 8 extracts each. With each task requesting an 8 CPU node, generating all 200 extracts in parallel takes about 35 minutes. This workflow also allows us to run multiple data pipelines together, sharing common outputs and reducing the amount of work that has to be duplicated. For instance, the updated planet file generated in the first step above is also used as the input to our Valhalla workflow. ### Next steps in the workflow Our strategy for working with OpenStreetMap and other data sources is “snout to tail” — that is, we look for opportunities to turn each piece of a giant data download into a useful “meal” for users. We’re working on additional inputs and outputs to Interline’s workflow. Tuning the workflow is also always a work in process. Each week, we find more ways to improve the performance of a step. We look forward to sharing more updates on our workflow in the future. In the meantime, we’d like to invite you to add a few next steps to your “professional workflow”: - If you work with OpenStreetMap data, sign up for the free [OSM Extracts](https://www.interline.io/osm/extracts/) developer preview. - If your organization needs to deploy a worldwide routing engine on premise or its in own cloud account, evaluate [Valhalla Tilepacks](https://www.interline.io/valhalla/tilepacks). - If your organization runs its own custom data workflow and needs help consider these or other workflow approaches, contact us to discuss Interline’s [consulting service](https://www.interline.io/consulting). We’re glad to share our experiences with Argo, Kubernetes, and related technologies.