# Ping Labz > Ping Labz helps you master IT certifications with hands-on Cisco networking labs, cybersecurity tutorials, and practical guides for CCNA, CCNP, and Security+. Public Ghost content for AI and LLM tooling. This file includes a bounded export of public pages first, then recent public posts. Append `.md` to any post or page URL to get the content in Markdown (for example, `/example-post.md`). ## Pages ### About PingLabz URL: https://www.pinglabz.com/about/ Last updated: 2026-08-02T05:47:56.000Z PingLabz is a networking education site for engineers, sysadmins, and IT professionals who want production-realistic Cisco content - not certification-mill regurgitation. Cluster pillars on the protocols that run real enterprise networks, hands-on labs with downloadable Cisco Modeling Labs topologies, and printable field-reference PDFs you can tape to your monitor. Everything written with one rule: if it would not survive a code review on a real Cisco IOS XE 17.x device, it does not ship here. ## What you will find here The site is built around four content types. Together they cover the topics a working network engineer hits in their first decade of practice. ### Cluster pillar guides Long-form complete-guide treatments of fourteen networking topics: [BGP](https://www.pinglabz.com/bgp/), [OSPF](https://www.pinglabz.com/ospf/), [VLANs and Layer 2 Switching](https://www.pinglabz.com/vlans-layer-2-switching/), [Spanning Tree](https://www.pinglabz.com/spanning-tree-protocol/), [802.1X and NAC](https://www.pinglabz.com/802-1x/), [Wireless and Catalyst 9800](https://www.pinglabz.com/wireless/), [SD-WAN](https://www.pinglabz.com/sd-wan/), [QoS](https://www.pinglabz.com/qos/), [EIGRP](https://www.pinglabz.com/eigrp/), [MPLS](https://www.pinglabz.com/mpls/), [FHRP (HSRP/VRRP/GLBP)](https://www.pinglabz.com/fhrp/), [IPv6](https://www.pinglabz.com/ipv6/), [GRE](https://www.pinglabz.com/gre/), and [Cisco ASA](https://www.pinglabz.com/cisco-asa/). Each pillar has 30 to 60 supporting deep-dive articles linked from the cluster index. ### The CCNA Labs library Sixty hands-on labs covering the entire Cisco CCNA 200-301 exam blueprint, each with a downloadable Cisco Modeling Labs (CML) topology .yaml file and a starter config bundle. Five labs are free preview; the rest are part of a PingLabz Pro membership at $19/month or $149/year. The labs are organized into five pillars matching the official CCNA exam domains: [Network Fundamentals](https://www.pinglabz.com/ccna-labs-network-fundamentals/), [Network Access](https://www.pinglabz.com/ccna-labs-network-access/), [IP Connectivity](https://www.pinglabz.com/ccna-labs-ip-connectivity/), [IP Services](https://www.pinglabz.com/ccna-labs-ip-services/), and [Security Fundamentals](https://www.pinglabz.com/ccna-labs-security-fundamentals/). ### Field-reference PDFs Multi-page printable cheat-sheets for the most-asked-about clusters. Three are live: [BGP Field Reference](https://www.pinglabz.com/bgp-cheatsheet/), [OSPF Field Reference](https://www.pinglabz.com/ospf-cheatsheet/), and the new [Cisco ASA Field Reference](https://www.pinglabz.com/asa-cheatsheet/). Each is a 9-page print-ready PDF with quick-reference tables, real lab captures from CML, troubleshooting decision trees, and copy-paste config templates. Free with email signup. ### Deep-dive articles Several hundred articles linked from the cluster pillars - each one a focused single-topic deep-dive (e.g. *"BGP path attributes"*, *"OSPF DR/BDR election"*, *"VLAN hopping defense"*). Use the cluster pillar pages above as the index; search is also available. ## What makes this different Three things separate PingLabz from the rest of the networking education space: - **Real captures, not synthetic.** Every `show` command output you read in a PingLabz article was captured from a working lab - usually a Cisco Modeling Labs IOL-XE 17.16 router. Where output is intentionally illustrative (not captured from the lab), it is clearly labeled. No fabricated CLI dressed up as evidence. - **Modern Cisco IOS XE 17.x syntax.** A lot of CCNA and CCNP material on the internet is from 2010-2015 and uses IOS 12.x / 15.x syntax. PingLabz uses current IOS XE syntax throughout, and notes legacy alternatives where they differ. - **Standardized lab IP scheme.** All labs across the site use the canonical [PingLabz lab IP scheme](https://www.pinglabz.com/lab-ip-scheme/): `10.255.0.x` loopbacks, `10.30.30.0/30` transit links, `10.20.0.0/24` LANs, `192.0.2.0/30` for inter-AS / public-internet simulation. The labs you do for OSPF use the same addresses as the labs you do for QoS or NAT. Concepts compound instead of resetting every time. ## Who runs PingLabz PingLabz is written by Alex, a working network engineer with hands-on time across Cisco enterprise gear, ISP environments, and modern multi-vendor data centers. The site is independent - no advertising, no sponsored content, no affiliate links. Revenue comes from PingLabz Pro memberships, which is what funds the writing and the Cisco Modeling Labs infrastructure that backs every lab. Site started in 2022 as a single-author networking blog. The 60-lab CCNA Labs library and the field-reference PDF series launched in 2026 alongside the Pro membership tier. ## How to engage - **Read for free.** Every cluster pillar and most deep-dive articles are open. No paywall on the theory content. - **Sign up free** at [the signup portal](https://www.pinglabz.com/membership/) for the field-reference PDFs and the newsletter. No card required for free signup. - **Subscribe to Pro** at $19/month or $149/year for the full 60-lab CCNA library, the downloadable CML topologies, the config bundles, and every future paid lab and PDF. - **Reach out** at [contact@pinglabz.com](mailto:contact@pinglabz.com) if you have questions, corrections, suggestions, or want to discuss content partnerships. See the [contact page](https://www.pinglabz.com/contact/) for details. ## The standards behind every post A few editorial rules I keep: - No synthetic or fabricated CLI captures presented as real. Where output is illustrative, it is labeled as such. - Every config example uses current Cisco IOS XE syntax (17.x). Legacy syntax is noted where it diverges. - Cross-links go to the most relevant prior article. The site is a connected graph, not a flat collection. - No em dashes (a personal style preference). No buzzword filler. No "in conclusion" paragraphs. - If you find an error, please [tell me](mailto:contact@pinglabz.com). Corrections ship within 48 hours. Thanks for reading. ### Contact URL: https://www.pinglabz.com/contact/ Last updated: 2026-08-02T05:47:56.000Z Get in touch about anything related to PingLabz - technical questions about a lab, billing or account issues, corrections, suggestions, or collaboration ideas. ## Primary contact **Email:** [contact@pinglabz.com](mailto:contact@pinglabz.com) I aim to respond within **2 business days**. Billing-related issues (refunds, subscription problems) usually get a same-day response. ## What you might be writing about Billing or refund The email on your PingLabz account + the date of charge Cannot access content The URL you are trying to reach + the email on your PingLabz account Technical error in a lab The lab URL, the step you are on, and what is going wrong (a screenshot of the CLI output helps a lot) Found a typo or factual error The URL + a short quote of the problem text. Corrections ship within 48 hours. Collaboration / partnership idea A short description of what you have in mind. Independent and ad-free, so most "affiliate" or "sponsored content" requests will be declined politely. Press, interviews, or speaking The publication, the deadline, and what angle you are after ## Faster routes for common issues ### Account access If you cannot sign in, try the password reset link at the bottom of the sign-in form first. Most "I cannot access content" issues are session-related and resolve with sign out + sign in. ### Cancel a subscription Sign in at [your account portal](https://www.pinglabz.com/#/portal/account) and click Cancel subscription. Your access continues until the end of the current paid period. No email needed. ### Refund request See the [Refund Policy](https://www.pinglabz.com/refunds/). First-time subscribers can email within 7 days of initial purchase for a full refund. After that, cancellation prevents future charges but does not refund prior periods. ### Privacy or data request See the [Privacy Policy](https://www.pinglabz.com/privacy-policy/). To delete your account or export your data, email [contact@pinglabz.com](mailto:contact@pinglabz.com) from the email address on your account. ## Stay in touch The fastest way to hear about new content is the PingLabz newsletter. [Sign up for free](https://www.pinglabz.com/signup/) \- one email per week at most, never shared, unsubscribe at any time. ### Privacy Policy URL: https://www.pinglabz.com/privacy-policy/ Last updated: 2026-06-13T20:09:49.000Z *Last updated: 11 May 2026* PingLabz ("we," "us," "our") operates www.pinglabz.com and publishes networking education content - written guides, hands-on labs for Cisco Modeling Labs, and downloadable field-reference PDFs. This Privacy Policy explains what data we collect, how we use it, and your rights. ## What we collect ### Account information When you create a free or Pro account, we collect: - Your email address (required) - Your name (optional) - Your account preferences (newsletter subscription, tier, etc.) ### Payment information When you subscribe to PingLabz Pro, payment is processed by Stripe. We do not store your credit card number, CVV, or full billing details. Stripe holds those. We receive only a customer reference and subscription status from Stripe. See [Stripe's privacy policy](https://stripe.com/privacy?ref=pinglabz.com) for how Stripe handles your payment data. ### Usage information Standard server logs (IP address, browser type, pages viewed) are retained temporarily for security and operational purposes. Our hosting platform uses session cookies to remember your login. We do not run third-party tracking or advertising scripts on the site. ## How we use it - Account management (signup, signin, password reset) - Transactional emails (welcome, payment receipts, subscription notifications, cancellation confirmations) - Newsletter emails (only if you have opted in) - Aggregate analytics to understand which content is read most - Compliance with legal obligations (financial records, tax) **We do not sell your data. We do not share your data with third parties for marketing purposes.** ## Services we use Stripe PurposePayment processing Data shared Email, name, payment details Mailgun Purpose Transactional email delivery Data sharedEmail, message content Self-hosted Ghost Purpose CMS and member database Data shared Account info, content access ## Your rights - **View your account.** Sign in at www.pinglabz.com and open your account portal. - **Update your information.** Edit your email and preferences from the account portal. - **Unsubscribe from emails.** Use the unsubscribe link in any newsletter, or update your preferences in the account portal. - **Delete your account.** Email contact@pinglabz.com or use the delete option in your portal. We will delete your account and associated personal data within 30 days, except where retention is legally required (financial records, etc.). - **Data portability.** Email contact@pinglabz.com to request an export of your account data. ## International users If you access PingLabz from outside the United States, you consent to your data being processed in the United States. We aim to comply with applicable data protection laws including the EU General Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA). If you have specific requests under those frameworks, email contact@pinglabz.com. ## Cookies We use a small number of strictly-necessary cookies for sign-in and session management. We do not use third-party advertising cookies. Browser settings allow you to block cookies, though signing in will not work without them. ## Children PingLabz content is intended for adults and professionals. We do not knowingly collect personal information from children under 16\. If you believe a child has provided information to us, email contact@pinglabz.com and we will remove the information. ## Data retention We retain account information for as long as your account is active, plus a reasonable period after closure for financial-record and tax-compliance purposes. We delete personal data sooner upon request, subject to legal requirements. ## Security We use industry-standard practices to protect your data: HTTPS for all traffic, encrypted storage of credentials, and access controls on the member database. No system is perfectly secure; if you believe your account has been compromised, email contact@pinglabz.com immediately. ## Updates to this policy We may update this Privacy Policy from time to time. Significant changes will be announced via email to active subscribers. The "Last updated" date at the top of this page reflects the current version. ## Contact Privacy questions or requests: **contact@pinglabz.com**. ### Terms of Service URL: https://www.pinglabz.com/terms-of-service/ Last updated: 2026-05-12T05:35:21.000Z *Last updated: 11 May 2026* These Terms of Service ("Terms") govern your use of www.pinglabz.com and the content and services PingLabz provides. By creating an account or accessing paid content, you agree to these Terms. ## About PingLabz PingLabz is a networking education website that publishes written guides, hands-on labs for Cisco Modeling Labs, downloadable field-reference PDFs, and related content for networking engineers, CCNA candidates, and IT professionals. ## Accounts You may create a free account or subscribe to a paid PingLabz Pro tier. You are responsible for keeping your account credentials secure. You agree to provide accurate information when you sign up. We may suspend or terminate accounts that violate these Terms. ## Subscriptions and billing PingLabz Pro is offered as a monthly or annual subscription. Subscriptions **auto-renew** at the end of each billing period unless you cancel. - You can cancel at any time from your account portal. - Cancellation prevents future charges and takes effect at the end of your current billing period. - Your access continues until the end of the current paid period. - Pricing changes may occur. Existing subscribers are grandfathered at their current rate until they cancel and resubscribe. See our [Refund Policy](https://www.pinglabz.com/refunds/) for refund eligibility. ## Acceptable use You agree NOT to: - Share your account credentials with others - Redistribute, resell, or republish PingLabz content without permission - Scrape content programmatically or via automated tools - Use PingLabz content to create competing courseware for resale - Misuse the labs (for example, attacking networks you do not own or have permission to test) - Attempt to bypass payment, access controls, or other security measures You MAY: - Use the content for personal study and professional development - Adapt configuration examples and lab .yaml files for your own use - Cite or quote PingLabz content with proper attribution - Print and share cheat-sheet PDFs with colleagues for non-commercial use ## Intellectual property All content on PingLabz, including text, diagrams, lab .yaml files, configurations, and PDFs, is copyright PingLabz unless otherwise noted. Configuration examples and lab files are licensed for personal use. Quoted command output from Cisco IOS XE and related Cisco products is property of Cisco Systems, Inc. and used for educational purposes. The PingLabz name and brand are trademarks of PingLabz. ## No warranty PingLabz content is provided "as is." We make reasonable efforts to ensure accuracy and currency but we do not guarantee: - That the content will help you pass any specific certification exam - That configurations will work in your specific environment without modification - That Cisco software syntax described will remain unchanged in future releases - Uninterrupted availability of the website or services You use the content at your own risk and judgement. ## Limitation of liability To the maximum extent permitted by law, PingLabz is not liable for any indirect, incidental, special, consequential, or punitive damages arising from your use of the site, content, or services. Our total liability to you is limited to the amount you have paid PingLabz in the twelve (12) months prior to the claim. This limitation applies whether the claim is based in contract, tort, statute, or any other legal theory. ## Termination We may terminate or suspend your account at any time, with or without notice, if you violate these Terms. Upon termination for a Terms violation, no refund is provided. You may close your account at any time by emailing contact@pinglabz.com or using the delete option in your account portal. ## Changes to these Terms We may update these Terms. Significant changes will be announced via email to active members. Your continued use of the site after such announcement constitutes acceptance of the updated Terms. If you do not agree to a change, you may close your account. ## Governing law These Terms are governed by the laws of the United States and the state in which PingLabz is operated, without regard to conflict-of-laws principles. Any dispute arising from these Terms will be resolved in the courts of that jurisdiction. ## Severability If any provision of these Terms is found unenforceable, the remaining provisions remain in full effect. ## Contact Questions about these Terms: **contact@pinglabz.com**. ### Disclaimer URL: https://www.pinglabz.com/disclaimer/ Last updated: 2026-05-12T05:35:22.000Z *Last updated: 11 May 2026* The information published on www.pinglabz.com is provided for general informational and educational purposes only. By using this website, you agree to the terms outlined in this disclaimer along with our [Terms of Service](https://www.pinglabz.com/terms-of-service/) and [Privacy Policy](https://www.pinglabz.com/privacy-policy/). ## Educational purpose only All tutorials, guides, hands-on labs, and downloadable materials on PingLabz are intended to support IT learning, career development, and certification preparation. We make reasonable efforts to keep the content accurate and current, but we do not guarantee that it is error-free or suitable for every situation. Use it at your own discretion. ## No professional advice PingLabz does not offer legal, financial, or professional IT consulting services. The content here is not a substitute for professional advice tailored to your specific environment. Before implementing any configuration in a production network, verify the steps in a lab environment and consult a qualified engineer where appropriate. ## Use at your own risk You are responsible for any actions you take based on PingLabz content. Lab configurations, command examples, and downloadable .yaml files are provided as starting points for your own learning. They have not been validated against every possible Cisco software version, platform, or production scenario. Always test in a lab before applying to a live network. ## No guarantee of certification success PingLabz content is designed to support CCNA and related certification preparation, but we do not guarantee that studying PingLabz material will result in passing any specific certification exam. Exam content, blueprints, and weighting may change at any time without notice. Always cross-reference with current official Cisco materials. ## Trademarks Cisco, Cisco IOS, IOS XE, Catalyst, ASA, ISE, Catalyst SD-WAN, and related names and logos are trademarks of Cisco Systems, Inc. Command output and screenshots reproduced from Cisco products are used for educational purposes under fair use. PingLabz is not affiliated with, endorsed by, or sponsored by Cisco Systems, Inc. Other product names, logos, and trademarks referenced on this site are the property of their respective owners. ## External links PingLabz content may include links to external websites for reference. We are not responsible for the content, accuracy, or privacy practices of those external sites. A link does not imply endorsement. ## Lab safety The lab content on PingLabz is designed for use in isolated lab environments (Cisco Modeling Labs Free, CML Personal, or equivalent virtual networking platforms). Do not run configurations or experiments against networks you do not own or do not have explicit permission to test. Misuse of network testing tools and configurations may violate applicable laws. ## Updates We may update this Disclaimer from time to time. The "Last updated" date at the top of the page reflects the current version. Continued use of the site after an update constitutes acceptance of the revised Disclaimer. ## Contact Questions about this Disclaimer: **contact@pinglabz.com**. ### OSPF Complete Guide: From Fundamentals to Enterprise Design URL: https://www.pinglabz.com/ospf/ Last updated: 2026-08-02T05:13:24.000Z OSPF (Open Shortest Path First) is the most widely deployed interior gateway protocol in enterprise networks, and the protocol most CCNP and CCIE candidates spend the longest time in the lab with. It is fast-converging, vendor-neutral, scales cleanly to thousands of routers when you design it right, and has just enough complexity (LSA types, area types, DR/BDR election, virtual links) to keep it interesting for a long career. This is the cluster overview for the full PingLabz OSPF series: 38 articles covering fundamentals, configuration, troubleshooting, internals, and enterprise design, all built on Cisco IOS XE 17.x. If you are studying for CCNA/CCNP/CCIE, designing a multi-area campus, or troubleshooting a stuck adjacency at 2 AM, start here. We will work through what OSPF is, how the protocol operates, the LSA types you actually need to know, and the configuration commands to bring up a working topology, with links into the deeper articles where you need them. Take the OSPF reference with you The free OSPF field-reference PDF: neighbor states, LSA types, and the show commands you will actually run. Delivered by email, no card required. [Get the free PDF](https://www.pinglabz.com/ospf-cheatsheet/) ## What OSPF Solves Inside a single autonomous system you need a routing protocol that can react to a link failure in well under a second, scale to thousands of routes without flooding the network with updates, and let multiple network operators express decisions about which links to prefer. RIP is too slow and does not scale. EIGRP is fast but proprietary (open since 2013, but adoption beyond Cisco is rare). BGP is too slow on purpose and the wrong abstraction for intra-domain routing. OSPF is what most networks reach for. It is: - **Link-state.** Every router floods its view of the local topology to every other router in the area, so all routers build the same map and run the same shortest-path computation independently. - **Standards-based** ([RFC 2328](https://www.rfc-editor.org/rfc/rfc2328?ref=pinglabz.com) for OSPFv2, RFC 5340 for OSPFv3). Vendor interop is genuinely good. - **Fast-converging.** Sub-second failover is achievable with the default timers and trivial with BFD. - **Hierarchical.** Areas let you contain LSA flooding and SPF computation to a sub-region, which is how OSPF scales past a few hundred routers. - **Cost-based.** The metric is a 16-bit integer derived from interface bandwidth, with a single tiebreaker (equal-cost multipath, ECMP, by default up to 4 paths and tunable). You will run OSPF underneath BGP on most production networks: OSPF for fast internal reachability, BGP at the edge for inter-AS policy. The two are complementary, not competitive. [What is OSPF? A Complete Guide to Open Shortest Path First](https://www.pinglabz.com/what-is-ospf-routing-protocol/) has the long-form intro. ## How OSPF Works (the 10,000-Foot View) OSPF runs directly on top of IP (protocol number 89, no TCP or UDP). It uses two multicast addresses on broadcast networks: 224.0.0.5 (AllSPFRouters) and 224.0.0.6 (AllDRouters). The protocol moves through three phases on every link it activates: 1. **Neighbor discovery.** Routers send Hello packets. If the Hello parameters match (area ID, hello/dead timers, subnet mask on broadcast links, authentication, MTU), the routers form a neighbor relationship. 2. **Database synchronization.** Once neighbors are at ExStart/Exchange, they swap Database Description (DBD) packets summarizing their LSDB, then request and exchange any missing LSAs via LSR/LSU/LSAck. 3. **SPF computation.** When the LSDB stabilizes, every router runs Dijkstra's SPF algorithm against its own copy and installs the resulting routes. Because every router in the area has an identical LSDB, every router computes the same topology. That is the link-state guarantee. The full mechanics are in [Introduction to OSPF: How It Works and Why It Matters](https://www.pinglabz.com/ospf-introduction/), with packet-level detail in [OSPF Packet Types Explained: Hello, DBD, LSR, LSU, LSAck](https://www.pinglabz.com/ospf-packet-types-explained/). ![Multi-area OSPF topology with two ABRs in Area 0 backbone connected by 10.0.0.0/30, R1 hanging off ABR1 in Area 1 (standard), R2 hanging off ABR2 in Area 2 (stub), each router labeled with router ID and DR/BDR priority](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/topology-2.png) Figure 1\. The reference topology used through this guide: Area 0 backbone with two ABRs, plus a standard area and a stub area hanging off it. ## OSPF Neighbor States: What show ip ospf neighbor Tells You Every OSPF adjacency walks through these states. If you see a neighbor stuck somewhere along the way, the state itself is the diagnostic clue: Down What's happeningNo Hellos received yet Stuck here means... L2 problem, OSPF not enabled, ACL Attempt What's happening NBMA only: trying to send unicast Hello Stuck here means... Manual neighbor config issue Init What's happening Hello received but our router ID is not yet in their Hello Stuck here means... One-way Hello, asymmetric ACL/filter 2-Way What's happening Bidirectional Hellos confirmed; this is the final state for non-DR/BDR pairs on broadcast links Stuck here means... Healthy on broadcast non-DR pairs ExStart What's happening Negotiating master/slave for DBD exchange Stuck here means... MTU mismatch (#1 cause) Exchange What's happening DBD packets being swapped Stuck here means... MTU mismatch, packet drop Loading What's happening Requesting missing LSAs via LSR Stuck here means...Rare; LSU loss Full What's happening LSDBs synchronized; healthy steady state Stuck here means...This is what you want Stuck-in-ExStart is so common it has its own article: [OSPF MTU Mismatch: Symptoms and Fixes](https://www.pinglabz.com/ospf-mtu-mismatch-troubleshooting/). The full state walkthrough with packet captures is in [OSPF Neighbor States Explained](https://www.pinglabz.com/ospf-neighbor-states-explained/). ![Eight OSPF neighbor states grouped into two phases: Discover Neighbor (Down, Attempt, Init, 2-Way) and Synchronize LSDB (ExStart, Exchange, Loading, Full). ExStart highlighted in red as where MTU mismatch lives, Full highlighted in teal as the target state](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/states.png) Figure 2\. The eight states split by what they're doing. ExStart is the MTU-mismatch graveyard; Full is where a healthy adjacency lives. ## LSA Types: The Heart of OSPF Internals OSPF carries topology information in Link-State Advertisements (LSAs). Different LSA types describe different scopes and propagate differently. Memorize the first six; the rest are special-case: Type 1: Router LSA Originated byEvery router ScopeSingle area Carries Router's own links and costs Type 2: Network LSA Originated byDR on broadcast/NBMA ScopeSingle area Carries Routers attached to the segment Type 3: Summary LSA (network) Originated byABR ScopeOther areas CarriesInter-area prefix Type 4: Summary LSA (ASBR) Originated byABR ScopeOther areas CarriesHow to reach an ASBR Type 5: External LSA Originated byASBR ScopeWhole AS (not stub) Carries Redistributed external routes Type 7: NSSA External LSA Originated byASBR in NSSA Scope NSSA only, then converted to Type 5 by ABR Carries External routes from inside an NSSA Type 9-11: Opaque LSAs Originated byVarious ScopeLink / area / AS Carries MPLS-TE, traffic engineering, segment routing The reason area types exist (stub, totally stubby, NSSA, totally NSSA) is to control which of these LSA types make it into the area, which is how you keep small areas small. The full reference, including how each LSA looks in `show ip ospf database`, is in [OSPF LSA Types Explained (Type 1-7)](https://www.pinglabz.com/ospf-lsa-types-explained/). ![OSPF LSA types reference: Type 1 Router LSA originated by every router with single-area scope, Type 2 Network LSA by DR on broadcast, Type 3 Summary by ABR for inter-area prefixes, Type 4 ASBR Summary, Type 5 External by ASBR scoped to whole AS except stubs, Type 7 NSSA External scoped to NSSA and converted to Type 5 at the ABR](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/lsa.png) Figure 3\. The six LSA types you actually have to know, and how far each one floods. ## Areas: Why OSPF Scales OSPF scales by splitting a network into areas. Every router in an area has an identical LSDB, but routers in different areas only see summarized inter-area information. Three rules govern area design: 1. Every multi-area OSPF deployment must have an Area 0 (the backbone). 2. All non-backbone areas must connect to Area 0, either directly via an Area Border Router (ABR), or indirectly through a virtual link (try to avoid these). 3. Inter-area traffic must transit Area 0\. The protocol does not support arbitrary area-to-area shortcuts. The four area flavors and what they filter: Standard Type 3 summary?Yes Type 4 summary?Yes Type 5 external?Yes Type 7 NSSA?No Stub Type 3 summary?Yes Type 4 summary?No Type 5 external?No Type 7 NSSA?No Totally stubby Type 3 summary? No (default route only) Type 4 summary?No Type 5 external?No Type 7 NSSA?No NSSA Type 3 summary?Yes Type 4 summary?No Type 5 external?No Type 7 NSSA? Yes (converted to T5 at ABR) Totally NSSA Type 3 summary? No (default route only) Type 4 summary?No Type 5 external?No Type 7 NSSA?Yes Use stub areas wherever you can; the smaller the LSDB the faster the SPF run. The full design walkthrough is in [OSPF Areas Explained: Why and How to Use Them](https://www.pinglabz.com/ospf-areas-explained/), configuration in [OSPF Stub Area Configuration](https://www.pinglabz.com/ospf-stub-area-configuration/), and the rare-but-needed [OSPF Virtual Links Configuration](https://www.pinglabz.com/ospf-virtual-link-configuration/) for backbone discontinuities. ![OSPF area type filtering matrix: Type 3 inter-area allowed in Standard, Stub, NSSA and replaced by default in Totally Stubby and Totally NSSA. Type 4 ASBR summary and Type 5 external blocked everywhere except Standard. Type 7 NSSA external allowed only in NSSA and Totally NSSA](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/areas.png) Figure 4\. What each area type filters. Smaller LSDB = faster SPF. Use the most restrictive area type the topology allows. ## DR/BDR Election on Multi-Access Links On a broadcast or NBMA segment with N routers, full-mesh adjacencies would require N(N-1)/2 sessions. Instead, OSPF elects one Designated Router (DR) and one Backup DR (BDR), and every other router only forms full adjacencies with those two. The DR generates the Type 2 LSA describing the segment. Election rules in order: 1. Highest OSPF priority (default 1, range 0-255; 0 means "never DR") 2. Highest router ID (which itself defaults to the highest loopback IP, falling back to highest physical interface IP at process start) The election is non-preemptive. If you bring up a router with priority 100 onto a segment that already has a DR, the existing DR stays put. To force a change, bounce the OSPF process or take the link down. [OSPF DR and BDR: What They Are and Why They Matter](https://www.pinglabz.com/ospf-dr-bdr-designated-router/) has the full election walkthrough. ![DR/BDR election on a five-router broadcast segment: R1 wins as DR with priority 100, R2 takes BDR with priority 50, R3, R4, and R5 stay as DROther with default priority 1. Adjacency matrix shows DROther to DR/BDR pairs go Full while DROther to DROther pairs stop at 2-Way, giving 7 full adjacencies instead of 10](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/drbdr.png) Figure 5\. The election rules on a multi-access segment, and the adjacency math: 5 routers, 7 full adjacencies (not 10), because DROthers only fully peer with the DR and BDR. On point-to-point links (the default network type for most modern WAN circuits), there is no DR election. Full adjacencies form directly between the two routers. ## OSPF Network Types: The One Setting Most Engineers Forget OSPF behaves differently depending on the network type assigned to each interface. The default depends on the interface (broadcast on Ethernet, point-to-point on serial, NBMA on Frame Relay), but you can override it: Broadcast DR election?Yes Hello/Dead10/40 Neighbor discoveryMulticast Use when Default Ethernet, multi-access Point-to-point DR election?No Hello/Dead10/40 Neighbor discoveryMulticast Use whenPoint-to-point links Point-to-multipoint DR election?No Hello/Dead30/120 Neighbor discoveryMulticast Use when Hub-and-spoke without DR overhead NBMA DR election?Yes Hello/Dead30/120 Neighbor discoveryManual neighbor config Use when Frame Relay full mesh (rare today) Loopback DR election?No Hello/Deadn/a Neighbor discoveryn/a Use when Default for loopback interfaces The single most useful trick: change a broadcast interface to point-to-point with `ip ospf network point-to-point`. It skips DR election and shaves a few seconds off neighbor formation. [OSPF Network Types Explained](https://www.pinglabz.com/ospf-network-types-explained/) covers when to do this and when not to. ## Configuration on Cisco IOS XE: Minimum Viable OSPF The smallest possible single-area OSPF config: ``` R1(config)# router ospf 1 R1(config-router)# router-id 1.1.1.1 R1(config-router)# network 10.0.0.0 0.0.255.255 area 0 R1(config-router)# passive-interface default R1(config-router)# no passive-interface GigabitEthernet0/0/1 ``` Three things to notice. First, the process ID (`1`) is locally significant only; you do not have to match it on neighbors. Second, the wildcard mask in the `network` statement is inverted from a regular subnet mask (`0.0.255.255` is the inverse of `/16`). Third, `passive-interface default` followed by selective `no passive-interface` is the safe pattern: it stops you from accidentally forming OSPF adjacencies on every interface in the network statement. You will also see the interface-based form, which is cleaner for multi-process or selective enablement: ``` R1(config)# interface GigabitEthernet0/0/1 R1(config-if)# ip ospf 1 area 0 ``` Once both sides are up, verification: ``` R1# show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 2.2.2.2 1 FULL/DR 00:00:36 10.0.12.2 GigabitEthernet0/0/1 ``` If you see anything other than `FULL/-` on a P2P link or `FULL/DR` or `FULL/BDR` on a broadcast link, you have a problem. Walk back up the neighbor states table. The full configuration walkthrough is in [How to Configure Single-Area OSPF on Cisco Routers](https://www.pinglabz.com/configure-single-area-ospf-cisco/) and [Configuring Multi-Area OSPF on Cisco Routers](https://www.pinglabz.com/configure-multi-area-ospf-cisco/). Other essentials: - [OSPF Passive Interfaces: When and How to Use Them](https://www.pinglabz.com/ospf-passive-interface/) \- the single most important hardening control - [OSPF Router ID: What It Is and How to Configure It](https://www.pinglabz.com/ospf-router-id-configuration-guide/) \- always set this manually - [OSPF Timers: Hello and Dead Intervals Explained](https://www.pinglabz.com/ospf-timers-hello-dead-intervals/) \- tune for sub-second failover with BFD - [OSPF Authentication Configuration (Plain Text and MD5)](https://www.pinglabz.com/ospf-authentication-cisco-configuration/) \- mandatory on any shared Layer 2 - [How to Advertise a Default Route in OSPF](https://www.pinglabz.com/advertise-default-route-ospf/) \- the classic `default-information originate` For IPv6, the protocol becomes OSPFv3: same SPF engine, redesigned plumbing. [OSPFv3 explained](https://www.pinglabz.com/ospfv3-explained-ipv6/) covers what actually changed (link-local next hops, the new LSA model, and RFC 5838 address families that carry IPv4 too), and the [OSPFv3 configuration guide](https://www.pinglabz.com/ospfv3-configuration-cisco-ios-xe/) is the step-by-step IOS XE build. The addressing fundamentals it assumes are in the [IPv6 complete guide](https://www.pinglabz.com/ipv6/). ## Metric and Cost: How OSPF Picks the Best Path OSPF uses cost (a 16-bit integer) as its only metric. Lower cost wins, and the SPF algorithm sums costs along the path. The default Cisco formula is: ``` cost = reference_bandwidth / interface_bandwidth ``` The default reference bandwidth is 100 Mbps, which means anything 100 Mbps or faster gets a cost of 1\. That is wrong on any modern network. Set the reference bandwidth high enough to differentiate your fastest link: ``` R1(config)# router ospf 1 R1(config-router)# auto-cost reference-bandwidth 100000 ! 100 Gbps ``` Set the same value on every router in the OSPF domain. [How OSPF Calculates Metric and Cost](https://www.pinglabz.com/ospf-metric-and-cost-calculation/) walks through the math and the gotchas. ![OSPF cost formula equals reference bandwidth divided by interface bandwidth. With the default 100 Mbps reference, every interface 100 Mbps or faster collapses to cost 1, hiding 1G, 10G, and 100G differences. Setting auto-cost reference-bandwidth 100000 (100 Gbps) restores meaningful cost values: 10000 for 10 Mbps, 1000 for 100 Mbps, 100 for 1G, 10 for 10G, 1 for 100G](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/cost.png) Figure 6\. Why the default reference bandwidth is wrong on every modern network: anything above 100 Mbps clamps to cost 1 and OSPF can't tell a 1G link from a 100G link. Get the OSPF Field Reference - 9 pages, free Everything in this pillar, on nine printable pages. State machine diagram, LSA types, troubleshooting decision tree, copy-paste IOS XE templates, and real lab captures from a Cisco Modeling Labs build. Free for PingLabz members - just sign up with your email. [Get the OSPF cheat-sheet](https://www.pinglabz.com/ospf-cheatsheet/) ## OSPF vs Other Routing Protocols Type OSPFLink-state EIGRPDistance-vector (DUAL) IS-ISLink-state BGPPath-vector Standards OSPFOpen (RFC 2328) EIGRP Open since 2013, Cisco-led IS-ISOpen (ISO 10589) BGPOpen (RFC 4271) Default AD (Cisco) OSPF110 EIGRP 90 internal / 170 external IS-IS115 BGP20 / 200 Convergence OSPFSub-second with tuning EIGRPSub-second (DUAL) IS-ISSub-second with tuning BGPSlow on purpose Metric OSPF Cost (bandwidth-derived) EIGRP Composite (bandwidth, delay) IS-ISCost (configurable) BGPPath attributes Hierarchy OSPF Areas with strict rules EIGRPNone native IS-ISTwo-level (L1/L2) BGPConfederations / RR Scope OSPFIntra-AS EIGRPIntra-AS IS-ISIntra-AS (huge ISPs) BGPInter-AS If you also run BGP (and most production networks do), see the BGP pillar at [BGP (Border Gateway Protocol): The Complete Guide](https://www.pinglabz.com/bgp/) for how the two coexist. The dedicated head-to-head is [BGP vs OSPF: When to Use Each Routing Protocol](https://www.pinglabz.com/bgp-vs-ospf/), and [OSPF vs Other Routing Protocols](https://www.pinglabz.com/ospf-vs-other-routing-protocols/) goes deeper on the trade-offs. The [EIGRP complete guide](https://www.pinglabz.com/eigrp/) covers the other Cisco IGP in the same depth. And in service-provider and large-enterprise cores, OSPF is most often the IGP running underneath [MPLS](https://www.pinglabz.com/mpls/), providing the loopback reachability that LDP and MP-BGP depend on. ## Redistribution: Bringing Routes In and Out of OSPF Every multi-protocol network needs redistribution somewhere: from connected/static into OSPF, between OSPF processes, between OSPF and BGP, between OSPF and EIGRP. Two things matter: 1. **Filter aggressively.** Redistribution defaults are dangerous (one ASBR can pull thousands of routes into OSPF Type 5 LSAs and explode the LSDB). Use route maps with prefix-list matches. 2. **Pick external metric type carefully.** Type 1 (E1) adds the OSPF cost to reach the ASBR; Type 2 (E2, default) does not. Use E1 inside a single AS to allow internal cost tiebreakers; E2 for routes coming in from outside. The full walkthrough with worked examples is in [OSPF Redistribution: How to Inject Routes from Other Protocols](https://www.pinglabz.com/ospf-redistribution-configuration/), and summarization at the ABR / ASBR boundary is in [OSPF Route Summarization: Strategy and Configuration](https://www.pinglabz.com/ospf-route-summarization-configuration/). ## OSPF Security and the Common Mistakes OSPF was not designed with hostile networks in mind, but the modern hardening checklist is short and well-understood: - **Authentication on every adjacency.** MD5 minimum, SHA where supported. Plain-text exists only for migration scenarios. - **Passive-interface default** followed by explicit `no passive-interface` on the interfaces that should peer. This is by far the most common control failure: a network statement that accidentally pulls in a customer-facing interface. - **Strict TTL check (GTSM)** with `ip ospf ttl-security hops 1` on point-to-point links to defeat off-link attackers. - **maxprefix-style** redistribution filtering to cap blast radius from a misconfigured ASBR. - **Stub or NSSA** on edge areas to limit the LSAs a compromised router can inject. The full hardening pattern is in [OSPF Design Best Practices for Enterprise Networks](https://www.pinglabz.com/ospf-design-best-practices-enterprise/). ## Troubleshooting: The Five Failures You Will See - [OSPF Neighbors Not Forming: Complete Troubleshooting Guide](https://www.pinglabz.com/troubleshoot-ospf-neighbors-not-forming/) - [Fixing OSPF Area Mismatch Issues](https://www.pinglabz.com/fix-ospf-area-mismatch/) - [OSPF MTU Mismatch: Symptoms and Fixes](https://www.pinglabz.com/ospf-mtu-mismatch-troubleshooting/) (the #1 cause of stuck-in-ExStart) - [OSPF Authentication Mismatch Troubleshooting](https://www.pinglabz.com/troubleshoot-ospf-authentication-mismatch/) - [Fixing Duplicate OSPF Router ID Issues](https://www.pinglabz.com/fix-duplicate-ospf-router-id/) - [OSPF Routes Not Appearing in Routing Table](https://www.pinglabz.com/ospf-routes-not-appearing-routing-table/) - [Common OSPF Passive Interface Mistakes](https://www.pinglabz.com/ospf-passive-interface/) - [OSPF Subnet Mask Mismatch: How to Troubleshoot and Fix](https://www.pinglabz.com/ospf-subnet-mask-mismatch/) Adjacency failures that appear only across tunnels are usually MTU problems in disguise - the [GRE complete guide](https://www.pinglabz.com/gre/) covers tunnel MTU, fragmentation, and MSS clamping in depth. ## The Full OSPF Cluster, in Reading Order ### Fundamentals 1\. [What is OSPF? A Complete Guide to Open Shortest Path First](https://www.pinglabz.com/what-is-ospf-routing-protocol/) 2\. [OSPF Key Terms and Concepts Every Network Engineer Should Know](https://www.pinglabz.com/ospf-key-terms-and-concepts/) 3\. [OSPF Neighbor States Explained](https://www.pinglabz.com/ospf-neighbor-states-explained/) 4\. [OSPF Areas Explained: Why and How to Use Them](https://www.pinglabz.com/ospf-areas-explained/) 5\. [OSPF DR and BDR: What They Are and Why They Matter](https://www.pinglabz.com/ospf-dr-bdr-designated-router/) 6\. [How OSPF Calculates Metric and Cost](https://www.pinglabz.com/ospf-metric-and-cost-calculation/) 7\. [OSPF Router ID: What It Is and How to Configure It](https://www.pinglabz.com/ospf-router-id-configuration-guide/) 8\. [OSPF Packet Types Explained](https://www.pinglabz.com/ospf-packet-types-explained/) 9\. [OSPF vs Other Routing Protocols](https://www.pinglabz.com/ospf-vs-other-routing-protocols/) ### Configuration 10\. [How to Configure Single-Area OSPF on Cisco Routers](https://www.pinglabz.com/configure-single-area-ospf-cisco/) 11\. [Configure Single Area OSPFv2: Complete Lab Guide](https://www.pinglabz.com/configure-ospfv2-cisco-routers/) 12\. [OSPF Passive Interfaces: When and How to Use Them](https://www.pinglabz.com/ospf-passive-interface/) 13\. [Interface-Based OSPF Configuration](https://www.pinglabz.com/interface-based-ospf-configuration/) 14\. [How to Advertise a Default Route in OSPF](https://www.pinglabz.com/advertise-default-route-ospf/) 15\. [Configuring Multi-Area OSPF on Cisco Routers](https://www.pinglabz.com/configure-multi-area-ospf-cisco/) 16\. [OSPF Authentication Configuration](https://www.pinglabz.com/ospf-authentication-cisco-configuration/) 17\. [OSPF Stub Area Configuration](https://www.pinglabz.com/ospf-stub-area-configuration/) 18\. [OSPF Virtual Links Configuration](https://www.pinglabz.com/ospf-virtual-link-configuration/) 19\. [OSPF Timers: Hello and Dead Intervals Explained](https://www.pinglabz.com/ospf-timers-hello-dead-intervals/) 20\. [OSPF Network Types Explained](https://www.pinglabz.com/ospf-network-types-explained/) 21\. [Cisco OSPF Configuration Guide: Step-by-Step Tutorial](https://www.pinglabz.com/configuring-ospf-cisco-routers-guide/) 22\. [Configuring OSPF Router IDs and Why They Matter](https://www.pinglabz.com/cisco-ospf-router-id-configuration-guide/) ### Troubleshooting 23\. [OSPF Neighbors Not Forming](https://www.pinglabz.com/troubleshoot-ospf-neighbors-not-forming/) 24\. [Fixing OSPF Area Mismatch Issues](https://www.pinglabz.com/fix-ospf-area-mismatch/) 25\. [OSPF MTU Mismatch](https://www.pinglabz.com/ospf-mtu-mismatch-troubleshooting/) 26\. [OSPF Authentication Mismatch](https://www.pinglabz.com/troubleshoot-ospf-authentication-mismatch/) 27\. [Fixing Duplicate OSPF Router ID Issues](https://www.pinglabz.com/fix-duplicate-ospf-router-id/) 28\. [OSPF Routes Not Appearing in Routing Table](https://www.pinglabz.com/ospf-routes-not-appearing-routing-table/) 29\. [Common OSPF Passive Interface Mistakes](https://www.pinglabz.com/ospf-passive-interface/) 30\. [OSPF Subnet Mask Mismatch](https://www.pinglabz.com/ospf-subnet-mask-mismatch/) ### Deep Dives 31\. [OSPF LSA Types Explained (Type 1-7)](https://www.pinglabz.com/ospf-lsa-types-explained/) 32\. [How OSPF SPF Algorithm and LSDB Work](https://www.pinglabz.com/ospf-spf-algorithm-lsdb-explained/) 33\. [OSPF Neighbor Relationships: The Foundation of OSPF](https://www.pinglabz.com/ospf-neighbor-relationships-guide/) 34\. [Understanding OSPF Terminology and Concepts](https://www.pinglabz.com/ospf-terminology-concepts-guide/) ### Design and Scaling 35\. [OSPF Route Summarization](https://www.pinglabz.com/ospf-route-summarization-configuration/) 36\. [OSPF Redistribution](https://www.pinglabz.com/ospf-redistribution-configuration/) 37\. [OSPF Design Best Practices for Enterprise Networks](https://www.pinglabz.com/ospf-design-best-practices-enterprise/) 38\. [OSPF Basics: How It Works and Why It Matters](https://www.pinglabz.com/ospf-introduction/) 39\. [OSPFv3 Explained: OSPF for IPv6 (and Address Families for IPv4)](https://www.pinglabz.com/ospfv3-explained-ipv6/) 40\. [OSPFv3 Configuration on Cisco IOS XE](https://www.pinglabz.com/ospfv3-configuration-cisco-ios-xe/) Hands-on OSPF - 5 CCNA labs included Configure OSPF single-area (free preview), multi-area with ABRs and inter-area routes, network types, DR/BDR election, and MD5 authentication on real Cisco IOS XE 17.16 routers. Downloadable CML topology .yaml + starter configs. Open the PingLabz CCNA Labs library to start. [Open the OSPF labs](https://www.pinglabz.com/ccna-labs-ip-connectivity/) ### More OSPF guides in this cluster 1\. [EIGRP vs OSPF: When to Use Each](https://www.pinglabz.com/eigrp-vs-ospf/) 2\. [OSPF Adjacency States: The 8-State FSM, Explained](https://www.pinglabz.com/ospf-adjacency-states/) 3\. [OSPF Cost: Reference Bandwidth, Manual Overrides, and Gotchas](https://www.pinglabz.com/ospf-cost/) 4\. [OSPF Link-State Advertisements: What an LSA Is](https://www.pinglabz.com/ospf-link-state-advertisement/) 5\. [OSPF Passive Interface: What It Does and Where to Use It](https://www.pinglabz.com/ospf-passive-interface/) 6\. [OSPF vs BGP Redistribution: Which Protocol Wins Which Tie](https://www.pinglabz.com/ospf-vs-bgp-redistribution/) 7\. [Routing Protocols Over GRE: OSPF, EIGRP, BGP](https://www.pinglabz.com/routing-protocols-over-gre/) ### Adjacency troubleshooting: the three ways it fails Arrive at these knowing only the symptom. Each one starts from the state you are staring at in `show ip ospf neighbor` and works backwards to the cause. [The neighbor is stuck in INIT](https://www.pinglabz.com/ospf-stuck-in-init-one-way-hello/) You can hear them, they cannot hear you. Reading the hello's neighbor list to find the direction that is broken. [The neighbor is stuck in EXSTART or EXCHANGE](https://www.pinglabz.com/ospf-mtu-mismatch-troubleshooting/) An MTU mismatch stalls the DBD exchange. Both fixes proven to FULL on a live lab, and why mtu-ignore is a workaround. [The adjacency keeps flapping](https://www.pinglabz.com/fix-duplicate-ospf-router-id/) Duplicate router IDs and dead-timer expiry, with the captured neighbor table showing the same ID twice. [Which LSA type carries which route](https://www.pinglabz.com/ospf-lsa-types-explained/) Every common LSA type mapped to real show ip ospf database output and the route code it produces. [What each area type actually filters](https://www.pinglabz.com/ospf-stub-area-configuration/) Stub, totally stubby and NSSA walked in order, proving in the LSDB which LSAs survive each one. ## Expert OSPF: Timers, Suppression, and Pathologies (CCIE level) The articles above build a correct OSPF network. The ones below are about the edges: the features that behave in non-obvious ways, the optimisations that quietly become outages, and the diagnostic habits that separate a professional from an expert. Every one is built on real Cisco IOS XE output from a six-router CML topology - a multi-area design with a normal area, an NSSA, an ASBR redistributing across a shared broadcast segment, and dual paths engineered so the pathologies actually appear. 1 [OSPF SPF and LSA throttling timers](https://www.pinglabz.com/ospf-spf-lsa-throttling-timers/) Exponential backoff, the LSA arrival constraint that will silently desync your database, and why the modern IOS XE defaults are already fast. 2 [OSPF prefix suppression: shrinking the routing table](https://www.pinglabz.com/ospf-prefix-suppression/) Delete the transit links nobody needs to reach - and understand the external route it can silently take down with it. 3 [The OSPF stub router: max-metric router-lsa](https://www.pinglabz.com/ospf-max-metric-stub-router/) Drain a router for maintenance without dropping a packet, and the one line every OSPF+BGP router should already have. 4 [The OSPF forwarding address](https://www.pinglabz.com/ospf-forwarding-address/) The five conditions, the hop it eliminates, and the silent RIB failure that no log message will ever tell you about. 5 [OSPF filtering: area range vs filter-list vs distribute-list vs summary-address](https://www.pinglabz.com/ospf-filtering-comparison/) Which mechanisms remove the LSA, which only touch your own RIB, and the Null0 discard route that black-holes a sloppy summary. 6 [OSPF path preference: O, O IA, E1, N1, E2, N2](https://www.pinglabz.com/ospf-path-preference-rules/) Route type is compared before metric. Proven with an N1 at cost 31 beating an E2 at cost 20 in a live lab. 7 [Expert OSPF troubleshooting: five broken scenarios, ticket style](https://www.pinglabz.com/expert-ospf-troubleshooting-scenarios/) Area type mismatch, a bad summary range, duplicate router-id, missing NSSA defaults, and a network-type mismatch - each broken for real. ## Studying for the CCIE? This cluster is part of the full CCNA to CCNP to CCIE Enterprise ladder on PingLabz, every rung built on real Cisco output. For expert-level depth across every EI v1.1 blueprint domain - and the four integration Super Labs - see the [CCIE Enterprise Infrastructure study hub](https://www.pinglabz.com/ccie-enterprise/). ## Frequently Asked Questions ### What does OSPF stand for? OSPF stands for Open Shortest Path First. It is a link-state interior gateway protocol defined in RFC 2328 (OSPFv2 for IPv4) and RFC 5340 (OSPFv3 for IPv6 and now IPv4 too). ### What protocol number does OSPF use? OSPF runs directly on top of IP using protocol number 89\. It does not use TCP or UDP. Hellos and most updates are sent to multicast 224.0.0.5 (AllSPFRouters) and DR/BDR-only traffic to 224.0.0.6 (AllDRouters). ### What is the administrative distance of OSPF? 110 on Cisco. Lower than RIP (120) and IS-IS (115), higher than EIGRP internal (90) and eBGP (20). The AD is used when multiple routing protocols offer routes to the same prefix; the protocol with the lowest AD wins. ### OSPF vs EIGRP, which one should I use? EIGRP converges slightly faster on small networks because of DUAL's local computation, but OSPF is the safer enterprise choice in 2026: it is genuinely vendor-neutral, scales further (multiple areas), and every certification track expects you to know it. Most CCIE candidates run both in the lab and OSPF in production. ### OSPF vs BGP, when do you use each? OSPF for fast intra-AS reachability. BGP for inter-AS policy and DFZ-scale prefix counts. You almost always run both: OSPF underneath BGP so the iBGP TCP sessions stay up and BGP NEXT\_HOPs resolve. See the [BGP pillar](https://www.pinglabz.com/bgp/) for the inter-AS half of the story. ### How many OSPF neighbor states are there? Eight: Down, Attempt, Init, 2-Way, ExStart, Exchange, Loading, Full. The first four are about discovering the neighbor; the last four are about synchronizing the LSDB. A healthy adjacency on a point-to-point link ends in Full; on a broadcast link, non-DR/BDR pairs stay at 2-Way and only the DR/BDR pair reaches Full. ## Key Takeaways If you take one thing away from this guide, make it this: OSPF rewards careful design at the area level. Every other concept (LSA types, area types, DR/BDR, network types) becomes obvious once you understand why areas exist. Memorize the neighbor states and the LSA types, set `passive-interface default` on every router, set the reference bandwidth on every router, and verify with `show ip ospf neighbor` after every change. Bookmark this page, work through the cluster articles in order, and lab every configuration. By the time you finish, you will be ready for any OSPF question a CCIE lab or a 3 AM ticket can throw at you. **Studying for the CCNA?** Test your OSPF knowledge on [PingLabz CCNA Flashcards](https://www.pinglabz.com/ccna-flashcards/) \- 200 free multiple-choice questions by topic, mixed, or a full mock exam, each with a plain-English explanation. ### References - [RFC 2328 - OSPF Version 2](https://www.rfc-editor.org/rfc/rfc2328?ref=pinglabz.com) - [Cisco OSPF technology documentation](https://www.cisco.com/c/en/us/tech/ip/open-shortest-path-first-ospf/index.html?ref=pinglabz.com) ### 802.1X Complete Guide: Port-Based Network Access Control URL: https://www.pinglabz.com/802-1x/ Last updated: 2026-08-01T19:48:25.000Z 802.1X is the IEEE standard for port-based network access control: the gatekeeper that decides whether a device gets on your network before a single user frame passes through. Combined with Cisco ISE as the policy server and a supplicant on the endpoint, it lets you enforce who and what connects to every wired port and wireless SSID, at scale, with central policy. This is the cluster overview for the full PingLabz 802.1X and Cisco ISE series: 31 articles covering fundamentals, EAP methods, switch configuration, host modes, dynamic VLANs, dACLs, troubleshooting, and enterprise rollout strategy. We will work through what 802.1X actually does, the three roles in the protocol, the EAP methods you need to choose between, and a minimum-viable IOS XE + ISE configuration, with links into the deeper articles where you need them. New labs and guides, in your inbox Every new PingLabz lab and deep-dive, built and verified on real Cisco IOS XE - free, straight to your inbox. [Join free](https://www.pinglabz.com/signup/) ## What 802.1X Solves Without 802.1X, any device that plugs into an Ethernet port or connects to your Wi-Fi gets Layer 2 access immediately. The only authentication that exists is whatever lives in firewalls and applications above. That model worked when networks were physically constrained (only employees could reach the cables). It does not work in 2026 when conference rooms have wall jacks, BYOD is universal, contractors come and go, and IoT devices outnumber humans. 802.1X moves authentication down to the port. A device cannot send any non-EAPOL traffic until it has been identified, authenticated, and authorized. The port enforcement is hardware-level: the switch's data-plane drops frames from unauthorized supplicants before they reach the network. The three things 802.1X gives you in production: - **Identity-based access.** Users authenticate with credentials, certificates, or both. Devices authenticate with MAC addresses (MAB) or certificates. The same port enforces different policies based on who connected. - **Dynamic policy.** ISE can return a VLAN, an ACL, an SGT, or a redirect URL based on attributes (group membership, posture, device type, time of day). The same physical port becomes guest, employee, or quarantine on demand. - **Zero-trust foundation.** Together with TrustSec/SGTs, 802.1X is how Cisco's macro-segmentation story works. Every flow has an identity attached. The intro is in [What Is 802.1X? Port-Based Network Access Control Explained](https://www.pinglabz.com/what-is-802-1x/). ## The Three Components: Supplicant, Authenticator, Authentication Server Supplicant What it is The endpoint requesting access What it does Provides credentials/certs to the authenticator Examples Windows native, NAM, macOS, Linux wpa\_supplicant, IP phones, printers Authenticator What it is The network access device What it does Enforces port state; relays EAP between supplicant and AAA server via RADIUS Examples Cisco Catalyst switch, C9800 wireless controller, ISR/ASR Authentication Server What it is The AAA / RADIUS server What it does Evaluates credentials against identity stores; returns Accept/Reject + attributes Examples Cisco ISE, FreeRADIUS, Microsoft NPS, Aruba ClearPass The protocol between supplicant and authenticator is EAPOL (EAP over LAN, Ethertype 0x888E for wired). Between authenticator and AAA server it is RADIUS (UDP 1812 for auth, 1813 for accounting). The switch is the bridge that translates between the two. [802.1X Components Explained](https://www.pinglabz.com/802-1x-components/) walks the architecture. ## The Authentication Flow, Step by Step 1. Port is in the unauthorized state. Only EAPOL frames pass. 2. Switch sends EAPOL-Identity-Request, OR supplicant sends EAPOL-Start. 3. Supplicant responds with EAPOL-Identity-Response (the username). 4. Switch wraps the response in a RADIUS Access-Request and sends it to ISE. 5. ISE selects an EAP method based on policy and starts the EAP exchange (PEAP, EAP-TLS, etc.). 6. Switch tunnels EAP-Request/Response messages between supplicant and ISE for the entire EAP conversation. The switch does not decrypt or inspect anything inside. 7. ISE evaluates final credentials, applies authorization policy, and returns a RADIUS Access-Accept (with optional VLAN, dACL, SGT, URL-Redirect) or Access-Reject. 8. Switch transitions the port to authorized state and applies the returned attributes. Data traffic now flows. The full flow with packet captures is in [802.1X Authentication Flow Step by Step: From EAPOL Start to RADIUS Accept](https://www.pinglabz.com/802-1x-authentication-flow-step-by-step/). The packet-level EAPOL detail is in [EAPOL Explained: How 802.1X Traffic Moves Over the Wire](https://www.pinglabz.com/eapol-explained/), and RADIUS specifics in [Understanding RADIUS in 802.1X Authentication](https://www.pinglabz.com/radius-in-802-1x/). What those eight steps actually look like on the wire is illuminating. Turn on `debug dot1x events` and `debug radius authentication` on the switch, plug a properly-configured PEAP supplicant into the port, and you get this: ``` ! Documented EAP-PEAP/MSCHAPv2 successful auth (Cisco IOS XE debug output) *Apr 11 13:00:01.123: dot1x-ev:[Gi0/1] Received EAPOL pkt (EAPOL_START) *Apr 11 13:00:01.124: dot1x-ev:[Gi0/1] Sending EAPOL packet (EAP_REQ/Identity) *Apr 11 13:00:01.156: dot1x-ev:[Gi0/1] Received EAPOL pkt (EAP_RESP/Identity, length=21) *Apr 11 13:00:01.157: RADIUS: Send Access-Request to 10.20.0.20:1812 id 1645/12 *Apr 11 13:00:01.160: RADIUS: Received from id 1645/12 10.20.0.20:1812, Access-Challenge *Apr 11 13:00:01.161: dot1x-ev:[Gi0/1] Sending EAPOL packet (EAP_REQ/PEAP) *Apr 11 13:00:01.190: dot1x-ev:[Gi0/1] Received EAPOL pkt (EAP_RESP/PEAP) *Apr 11 13:00:01.220: RADIUS: Send Access-Request to 10.20.0.20:1812 id 1645/13 *Apr 11 13:00:01.225: RADIUS: Received from id 1645/13 10.20.0.20:1812, Access-Challenge (PEAP TLS tunnel established, MSCHAPv2 inner exchange follows) *Apr 11 13:00:01.310: RADIUS: Send Access-Request to 10.20.0.20:1812 id 1645/16 *Apr 11 13:00:01.315: RADIUS: Received from id 1645/16 10.20.0.20:1812, Access-Accept *Apr 11 13:00:01.316: dot1x-ev:[Gi0/1] Sending EAPOL packet (EAP_SUCCESS) *Apr 11 13:00:01.317: %AUTHMGR-5-START: Starting 'dot1x' for client (aabb.cc00.6400) on Interface Gi0/1 *Apr 11 13:00:01.318: %DOT1X-5-SUCCESS: Authentication successful for client (aabb.cc00.6400) on Interface Gi0/1 *Apr 11 13:00:01.320: %AUTHMGR-5-SUCCESS: Authorization succeeded for client (aabb.cc00.6400) on Interface Gi0/1 ``` Notice that the switch sees the entire conversation but understands none of it. The Access-Request / Access-Challenge ping-pong between switch and RADIUS is opaque to the switch: each EAPOL frame from the supplicant gets wrapped, forwarded to RADIUS in an `EAP-Message` attribute, and the RADIUS response is unwrapped and forwarded back as EAPOL. Only the final `Access-Accept` tells the switch what to do, by which time the supplicant and RADIUS have already proven the user's identity inside the TLS tunnel that the switch could not see into. This is by design - the authenticator is deliberately a dumb relay so that adding new EAP methods does not require switch firmware updates. ## EAP Methods: PEAP, EAP-TLS, EAP-FAST, EAP-TTLS EAP is a framework, not a protocol. Inside the EAP wrapper, you pick a specific method. The choice is consequential and not as simple as "use the most secure one": EAP-TLS Server cert?Yes Client cert?Yes Credentials Mutual cert auth, no password Best for Highest security; corporate-managed devices with PKI PEAP-MSCHAPv2 Server cert?Yes Client cert?No Credentials Username/password inside TLS tunnel Best for Most common in mixed environments; weakest of the modern methods EAP-TTLS Server cert?Yes Client cert?No Credentials Username/password (or other) inside TLS tunnel Best for Like PEAP but more flexible inner methods; less common in Cisco environments EAP-FAST Server cert?Optional Client cert?Optional Credentials PAC-based shared secret Best for Cisco-led, PAC provisioning is the operational headache EAP-MD5 Server cert?No Client cert?No CredentialsPlain MD5 challenge Best forNever. Deprecated. The two real choices in 2026 are EAP-TLS (where you have the PKI in place) and PEAP-MSCHAPv2 (where you do not). EAP-TLS is genuinely more secure (no password material on the wire, even encrypted) but operationally heavier (you need a working enrollment story for every device). PEAP is the pragmatic default if your endpoints are domain-joined Windows machines. Detail in [How EAP Works in 802.1X: EAP Methods Compared](https://www.pinglabz.com/eap-methods-compared/). ## Cisco ISE: The Policy Engine Cisco ISE (Identity Services Engine) is the AAA server for the Cisco ecosystem. You can run 802.1X with FreeRADIUS or Microsoft NPS, but ISE is what almost every Cisco-shop deployment uses because of its identity store integration (AD, LDAP, certificates, internal users), policy granularity, and TrustSec/SGT integration. The ISE constructs you must understand: - **Network Device.** A switch or WLC registered with ISE, including the shared RADIUS secret. - **Identity Store.** Where ISE looks up usernames and certs. AD, internal users, certificate authentication profiles. - **Policy Set.** Conditions that select an authentication and authorization workflow. - **Authentication Policy.** "If MAB, use Internal Endpoints. If 802.1X, use Active Directory." - **Authorization Policy.** "If user is in Domain Users group AND device is compliant, return VLAN 10 + permit-all dACL." - **Authorization Profile.** The attribute bundle returned in the RADIUS Access-Accept (VLAN, dACL, SGT, URL). The end-to-end walkthrough that brings the components together is in [Introduction to Cisco ISE](https://www.pinglabz.com/cisco-ise-intro/), and the full wired deployment guide in [Cisco ISE 802.1x Wired Configuration: A Practical Step-by-Step Guide](https://www.pinglabz.com/cisco-ise-802-1x-wired-configuration-guide/). ## Minimum Viable Switch Configuration on IOS XE The smallest possible 802.1X port configuration: ``` ! Global Switch(config)# aaa new-model Switch(config)# aaa authentication dot1x default group radius Switch(config)# aaa authorization network default group radius Switch(config)# dot1x system-auth-control Switch(config)# radius server ISE-PSN-01 Switch(config-radius-server)# address ipv4 10.10.10.10 auth-port 1812 acct-port 1813 Switch(config-radius-server)# key 7 [secret] ! Port Switch(config)# interface GigabitEthernet1/0/1 Switch(config-if)# switchport mode access Switch(config-if)# switchport access vlan 10 Switch(config-if)# authentication port-control auto Switch(config-if)# authentication host-mode multi-domain Switch(config-if)# dot1x pae authenticator Switch(config-if)# mab Switch(config-if)# authentication order dot1x mab Switch(config-if)# authentication priority dot1x mab Switch(config-if)# spanning-tree portfast Switch(config-if)# spanning-tree bpduguard enable ``` This is the exact authenticator config running in the PingLabz 802.1X reference lab (an iosvl2 switch facing a wpa\_supplicant Alpine host on Gi0/1 and a FreeRADIUS server on Gi0/0). Pulled live with `show running-config`: ``` ! From the PingLabz lab (iosvl2 15.2 ADVENTERPRISEK9-M) aaa new-model aaa authentication dot1x default group radius aaa authorization network default group radius aaa accounting dot1x default start-stop group radius dot1x system-auth-control radius server PINGLABZ-RAD address ipv4 10.20.0.20 auth-port 1812 acct-port 1813 key PingLabzRadius! interface GigabitEthernet0/1 description To HOST1 supplicant switchport mode access authentication port-control auto authentication periodic authentication timer reauthenticate server mab dot1x pae authenticator dot1x timeout tx-period 10 spanning-tree portfast ``` Three things to notice. First, `port-control auto` is what actually enables enforcement; without it the switch attempts authentication but never blocks. Second, `mab` falls back to MAC Authentication Bypass for devices without a supplicant (printers, IP cameras, IoT). Third, `multi-domain` host mode allows one voice device + one data device per port, which is the dominant pattern for modern enterprises. Two quick checks confirm the global state is right before you touch a port: ``` SW1# show dot1x Sysauthcontrol Enabled Dot1x Protocol Version 3 SW1# show dot1x all summary Interface PAE Client Status -------------------------------------------------------- Gi0/1 AUTH none UNAUTHORIZED ``` `Sysauthcontrol Enabled` is the global `dot1x system-auth-control` command confirmed loaded. The summary shows Gi0/1 in the `UNAUTHORIZED` state with no client, which is what every dot1x-configured port looks like before a supplicant has been seen. The status will flip to `AUTHORIZED` in the same column the instant authentication succeeds. The single most useful per-session view (this is the illustrative output from an actual production-style authenticated session): ``` Switch# show authentication sessions interface Gi1/0/1 details Interface: GigabitEthernet1/0/1 MAC Address: 0050.b6c1.1a2b IPv4 Address: 10.10.10.55 User-Name: alice@example.com Status: Authorized Domain: DATA Oper host mode: multi-domain Oper control dir: both Session timeout: N/A Common Session ID: 0A0A0A1400000123ABCD1234 Acct Session ID: 0x0000000A Handle: 0x12000123 Current Policy: POLICY_Gi1/0/1 Method status list: Method State dot1x Authc Success ``` And the RADIUS server side of the picture. `show aaa servers` is the first command to run when authentication is failing intermittently or universally - the counters tell you immediately whether the switch is even talking to RADIUS: ``` SW1# show aaa servers RADIUS: id 1, priority 1, host 10.20.0.20, auth-port 1812, acct-port 1813 State: current UP, duration 376s, previous duration 0s Dead: total time 0s, count 0 Authen: request 4, timeouts 4, failover 0, retransmission 3 Response: accept 0, reject 0, challenge 0 Response: unexpected 0, server error 0, incorrect 0, time 0ms Transaction: success 0, failure 1 Author: request 0, timeouts 0, failover 0, retransmission 0 Response: accept 0, reject 0, challenge 0 Account: request 0, timeouts 0, failover 0, retransmission 0 Request: start 0, interim 0, stop 0 ``` The pattern above (this is real output from the lab during a deliberately broken RADIUS state) is the canonical "RADIUS server unreachable" signature: `State current UP` but `timeouts` equal to `request` count and zero accepts, rejects, or challenges. `State UP` just means the switch has not yet dead-marked the server based on consecutive timeouts; it does not mean RADIUS is actually answering. A healthy RADIUS server shows accepts (and occasional rejects) climbing alongside requests, with timeouts at zero or near zero. Per-feature deep dives in [Basic 802.1X Port Configuration](https://www.pinglabz.com/basic-802-1x-port-configuration-cisco-ios-xe/) and [Configuring Cisco ISE as a RADIUS Server for 802.1X](https://www.pinglabz.com/configuring-cisco-ise-radius-server-802-1x/). ## Host Modes: How Many Devices Per Port? single-host Behavior Strictest. One MAC; second MAC errdisables the port. Use for High-security ports; lab/test only in most environments multi-host Behavior One MAC authenticates, all others piggyback (no per-MAC enforcement). Use for Avoid; gives up most of the point of 802.1X multi-domain (MDA) Behavior One voice device + one data device, separately authenticated. Use for Standard enterprise default with IP phones multi-auth Behavior Each MAC authenticated independently. Use for Conference rooms, hubs behind a port, virtualized hosts Detail in [802.1X Authentication Host Modes](https://www.pinglabz.com/802-1x-host-modes-cisco-ios-xe/). Whichever host mode is in effect, the per-interface dot1x timers and limits are what govern how aggressively the port retries and times out. The PingLabz lab port (default settings except for an explicit `dot1x timeout tx-period 10`) reports: ``` SW1# show dot1x interface GigabitEthernet0/1 details Dot1x Info for GigabitEthernet0/1 ----------------------------------- PAE = AUTHENTICATOR QuietPeriod = 60 ServerTimeout = 0 SuppTimeout = 30 ReAuthMax = 2 MaxReq = 2 TxPeriod = 10 Dot1x Authenticator Client List Empty ``` `TxPeriod` is the interval between EAP-Request/Identity frames the switch sends when no supplicant has replied yet. `QuietPeriod` is how long the port waits after a failed authentication before retrying. `MaxReq` caps the EAP-Request retries before the switch declares the supplicant absent and moves to MAB or the fallback action. Tune these in production, especially `QuietPeriod`, which is the difference between "snappy reconnect after a printer power-cycle" and "user blames you for the network being slow." ## Dynamic VLANs, dACLs, and Other Authorization Goodies The point of 802.1X is not just yes/no but what-policy. Five common authorization outcomes: - **Dynamic VLAN assignment.** ISE returns Tunnel-Type=VLAN, Tunnel-Medium-Type=802, Tunnel-Private-Group-ID=NN, and the switch puts the port into VLAN NN. [Dynamic VLAN Assignment with 802.1X and Cisco ISE](https://www.pinglabz.com/dynamic-vlan-assignment-802-1x-cisco-ise/). - **Downloadable ACLs (dACLs).** ISE pushes a per-session ACL that the switch programs in hardware. Lets you express user-specific filtering without provisioning ACLs on every switch. [Downloadable ACLs (dACLs) with Cisco ISE and 802.1X](https://www.pinglabz.com/dacl-downloadable-acl-cisco-ise-802-1x/). - **Guest VLAN / Auth-Fail VLAN / Critical VLAN.** What happens when the supplicant fails (Auth-Fail VLAN), is absent (Guest VLAN), or RADIUS is unreachable (Critical VLAN). [Guest VLAN, Auth-Fail VLAN, and Critical VLAN in 802.1X](https://www.pinglabz.com/802-1x-guest-vlan-auth-fail-vlan-critical-vlan/). - **Web Authentication fallback.** If 802.1X and MAB both fail, redirect to a captive portal. Common for guest scenarios. [Web Authentication as a Fallback in 802.1X](https://www.pinglabz.com/web-authentication-fallback-802-1x-cisco-ise/). - **Change of Authorization (CoA).** ISE can push a new policy mid-session without making the supplicant reauthenticate from scratch. Critical for posture remediation and re-evaluation. [Change of Authorization (CoA) in 802.1X](https://www.pinglabz.com/802-1x-coa-configuration-cisco-ios-xe-ise/). ## 802.1X with IP Phones (Multi-Domain Authentication) The IP phone use case is so common it has its own host mode. The phone authenticates first (often via certificate) and lands in the voice domain. The PC behind the phone authenticates separately and lands in the data domain. The switch enforces VLAN segregation between them. Configuration plus a real-world walkthrough in [802.1X with IP Phones](https://www.pinglabz.com/802-1x-ip-phones-multi-domain-authentication/). ## Phased Deployment: Don't Turn 802.1X On in One Day The single biggest lesson from production 802.1X deployments: roll out in phases. Going from "no 802.1X" to "closed mode everywhere" overnight will lock out half your users on day one and you will be reverting at 9 AM. Monitor mode What it does Authenticates everyone, but allows access regardless of result. Logs to ISE. Use for Phase 1: discovery. Find all the unsupplicantable devices. Low-impact mode What it does Pre-auth ACL allows DHCP/DNS/PXE/etc. Failed auth still gets limited access. Use for Phase 2: production rollout with a safety net. Closed mode What it does No traffic until authenticated. Strictest. Use for Phase 3: target end-state for high-security segments. Detail in [Monitor Mode vs Low-Impact Mode vs Closed Mode](https://www.pinglabz.com/802-1x-deployment-modes/) and [Phased 802.1X Deployment Strategy for Enterprise Networks](https://www.pinglabz.com/802-1x-phased-deployment/). ## RADIUS Resilience and Scale If your single ISE PSN dies and Critical VLAN is not configured, your network goes dark. Two articles cover the production-grade design: - [RADIUS Redundancy and Failover in 802.1X Deployments](https://www.pinglabz.com/802-1x-radius-redundancy/) - [802.1X Scalability and High Availability Design for Large Enterprise Networks](https://www.pinglabz.com/802-1x-scalability-ha/) And for the macro-segmentation story (SGTs), [Cisco TrustSec and SGTs](https://www.pinglabz.com/802-1x-trustsec-sgt/). ## AAA and the Wider Security Picture 802.1X is one consumer of a RADIUS server, but the same AAA infrastructure secures device administration too. The protocol split is worth internalising: **RADIUS admits users** (this cluster, plus VPN and Wi-Fi), while **TACACS+ administers devices** (per-command authorization for engineers logging into switches). Both are configured, with a live authentication captured on the wire, in [AAA on Cisco IOS XE: TACACS+ vs RADIUS](https://www.pinglabz.com/aaa-tacacs-radius-device-administration/). 802.1X also sits inside a broader access-edge hardening story: control-plane policing, anti-spoofing, IPv6 first-hop security, MACsec on the access link, and management-plane lockdown. The [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/) ties them together, and [MACsec](https://www.pinglabz.com/macsec-802-1ae-explained/) is the natural encryption partner for an 802.1X-authenticated port. ## Troubleshooting: When Authentication Fails - [802.1X Authentication Failing: Where to Start Troubleshooting](https://www.pinglabz.com/802-1x-authentication-failing-troubleshooting/) - [Client Stuck in Unauthorized State](https://www.pinglabz.com/802-1x-client-stuck-unauthorized-state/) - [RADIUS Server Unreachable in 802.1X](https://www.pinglabz.com/radius-server-unreachable-802-1x-cisco-ios-xe/) - [Dynamic VLAN Assignment Not Working](https://www.pinglabz.com/802-1x-dynamic-vlan-assignment-not-working/) - [dACL Not Applying Correctly in 802.1X](https://www.pinglabz.com/dacl-not-applying-802-1x-troubleshooting/) - [Troubleshooting 802.1X with show authentication sessions and debug Commands](https://www.pinglabz.com/802-1x-troubleshooting-show-authentication-sessions-debug/) Universal first-step commands: ``` Switch# show authentication sessions interface Gi1/0/1 details Switch# show aaa servers Switch# debug radius authentication Switch# debug dot1x events ``` The debug output for a real failed authentication (wrong password, RADIUS returns Access-Reject) looks like this. Notice how the RADIUS exchange completes - the server is reachable - but the final response is a reject rather than an accept: ``` ! Documented EAP-PEAP failed auth (wrong password) *Apr 11 13:05:14.001: dot1x-ev:[Gi0/1] Received EAPOL pkt (EAPOL_START) *Apr 11 13:05:14.002: dot1x-ev:[Gi0/1] Sending EAPOL packet (EAP_REQ/Identity) *Apr 11 13:05:14.040: dot1x-ev:[Gi0/1] Received EAPOL pkt (EAP_RESP/Identity, length=21) *Apr 11 13:05:14.041: RADIUS: Send Access-Request to 10.20.0.20:1812 id 1645/22 *Apr 11 13:05:14.045: RADIUS: Received from id 1645/22 10.20.0.20:1812, Access-Challenge *Apr 11 13:05:14.090: RADIUS: Send Access-Request to 10.20.0.20:1812 id 1645/24 *Apr 11 13:05:14.094: RADIUS: Received from id 1645/24 10.20.0.20:1812, Access-Reject *Apr 11 13:05:14.095: dot1x-ev:[Gi0/1] Sending EAPOL packet (EAP_FAIL) *Apr 11 13:05:14.096: %DOT1X-5-FAIL: Authentication failed for client (aabb.cc00.6400) on Interface Gi0/1 *Apr 11 13:05:14.098: %AUTHMGR-5-FAIL: Authorization failed or unapplied for client (aabb.cc00.6400) on Interface Gi0/1 ``` Contrast with the empty-port case (real capture from the lab when a port has no supplicant on the other side at all): ``` SW1# show logging | include DOT1X *May 11 13:59:40.859: dot1x-ev:[Gi0/1] No DOT1X subblock found for port down *May 11 13:59:41.273: dot1x-ev:DOT1X Supplicant not enabled on GigabitEthernet0/1 ``` Two lines and silence. The switch is sending EAP-Request/Identity every `TxPeriod` seconds and getting nothing back, so there is no debug output to print. This is what "client stuck Unauthorized" looks like 99% of the time: not a misconfigured RADIUS, not a wrong cert - a missing or misconfigured supplicant on the endpoint. Symptom-to-cause translation: `show aaa servers` shows timeouts equal to requests, zero accepts/rejects Likely cause RADIUS unreachable (routing, ACL, RADIUS daemon down, wrong port, wrong source interface) First check Ping RADIUS from the switch sourced from the same SVI; check `ip radius source-interface`; check no ACL blocks UDP/1812 Requests have rejects but no accepts Likely cause Wrong shared secret, unknown user, or supplicant misconfigured (wrong EAP method, wrong identity) First check Check `show aaa servers` reject count vs ISE auth log; verify shared secret on both sides; verify supplicant credentials `show dot1x all summary` stays UNAUTHORIZED with no debug output Likely cause No supplicant on the endpoint, or supplicant disabled, or wrong EAPOL ethertype on a hop First check Capture on the port; confirm EAPOL Ethertype 0x888E appearing from the supplicant side Port flaps between AUTHORIZED and UNAUTHORIZED every few minutes Likely cause Reauth timer too aggressive, or session-timeout returned by RADIUS that the supplicant cannot meet First check Check `authentication periodic` and `authentication timer reauthenticate`; review Session-Timeout AV-pair in Access-Accept Authentication succeeds but the port is in the wrong VLAN Likely cause Dynamic VLAN attributes missing from Access-Accept, or VLAN does not exist on the switch First check Check ISE authz profile (Tunnel-Type, Tunnel-Medium-Type, Tunnel-Private-Group-ID); `show vlan brief` for VLAN existence One port works, the adjacent identical port does not Likely cause Per-port config drift, or the interface inherits from a different template First check Diff `show run interface` for both; check `access-session inherit` if IBNS 2.0 ## The Full 802.1X Cluster, in Reading Order ### Fundamentals 1\. [What Is 802.1X?](https://www.pinglabz.com/what-is-802-1x/) 2\. [802.1X Components Explained](https://www.pinglabz.com/802-1x-components/) 3\. [How EAP Works in 802.1X: EAP Methods Compared](https://www.pinglabz.com/eap-methods-compared/) 4\. [EAPOL Explained](https://www.pinglabz.com/eapol-explained/) 5\. [Understanding RADIUS in 802.1X](https://www.pinglabz.com/radius-in-802-1x/) 6\. [Introduction to Cisco ISE](https://www.pinglabz.com/cisco-ise-intro/) 7\. [802.1X Authentication Flow Step by Step](https://www.pinglabz.com/802-1x-authentication-flow-step-by-step/) ### End-to-End Walkthrough 8\. [Cisco ISE 802.1x Wired Configuration: A Practical Step-by-Step Guide](https://www.pinglabz.com/cisco-ise-802-1x-wired-configuration-guide/) ### Switch and RADIUS Setup 9\. [Basic 802.1X Port Configuration on Cisco IOS XE](https://www.pinglabz.com/basic-802-1x-port-configuration-cisco-ios-xe/) 10\. [Configuring Cisco ISE as a RADIUS Server](https://www.pinglabz.com/configuring-cisco-ise-radius-server-802-1x/) ### Authentication Methods 11\. [Configuring PEAP Authentication](https://www.pinglabz.com/configuring-peap-authentication-cisco-ise-ios-xe/) 12\. [Configuring EAP-TLS with Certificates](https://www.pinglabz.com/configuring-eap-tls-certificates-cisco-ise-ios-xe/) 13\. [MAC Authentication Bypass (MAB) Configuration](https://www.pinglabz.com/mab-configuration-cisco-ios-xe-ise/) ### Host Modes and Policy 14\. [802.1X Authentication Host Modes](https://www.pinglabz.com/802-1x-host-modes-cisco-ios-xe/) 15\. [Dynamic VLAN Assignment](https://www.pinglabz.com/dynamic-vlan-assignment-802-1x-cisco-ise/) 16\. [Guest VLAN, Auth-Fail VLAN, and Critical VLAN](https://www.pinglabz.com/802-1x-guest-vlan-auth-fail-vlan-critical-vlan/) 17\. [Downloadable ACLs (dACLs) with ISE](https://www.pinglabz.com/dacl-downloadable-acl-cisco-ise-802-1x/) ### Advanced Features 18\. [802.1X with IP Phones](https://www.pinglabz.com/802-1x-ip-phones-multi-domain-authentication/) 19\. [Web Authentication as a Fallback](https://www.pinglabz.com/web-authentication-fallback-802-1x-cisco-ise/) 20\. [Change of Authorization (CoA)](https://www.pinglabz.com/802-1x-coa-configuration-cisco-ios-xe-ise/) ### Troubleshooting 21\. [802.1X Authentication Failing](https://www.pinglabz.com/802-1x-authentication-failing-troubleshooting/) 22\. [Client Stuck in Unauthorized State](https://www.pinglabz.com/802-1x-client-stuck-unauthorized-state/) 23\. [RADIUS Server Unreachable](https://www.pinglabz.com/radius-server-unreachable-802-1x-cisco-ios-xe/) 24\. [Dynamic VLAN Assignment Not Working](https://www.pinglabz.com/802-1x-dynamic-vlan-assignment-not-working/) 25\. [dACL Not Applying Correctly](https://www.pinglabz.com/dacl-not-applying-802-1x-troubleshooting/) 26\. [Troubleshooting with show authentication sessions and debug](https://www.pinglabz.com/802-1x-troubleshooting-show-authentication-sessions-debug/) ### Deployment Strategy and Design 27\. [Monitor / Low-Impact / Closed Mode](https://www.pinglabz.com/802-1x-deployment-modes/) 28\. [Phased 802.1X Deployment Strategy](https://www.pinglabz.com/802-1x-phased-deployment/) 29\. [RADIUS Redundancy and Failover](https://www.pinglabz.com/802-1x-radius-redundancy/) 30\. [Cisco TrustSec and SGTs](https://www.pinglabz.com/802-1x-trustsec-sgt/) 31\. [802.1X Scalability and High Availability Design](https://www.pinglabz.com/802-1x-scalability-ha/) Hands-on 802.1X - configure the authenticator Configure the switch as 802.1X authenticator with port-control auto and dot1x pae authenticator. Real switch-side config + the show commands that matter. Plus seven other security labs in the cluster. Open the PingLabz CCNA Labs library. [Open the 802.1X lab](https://www.pinglabz.com/ccna-labs-security-fundamentals/) ### More 802.1X guides in this cluster 1\. [802.1X on Cisco Switches: Step-by-Step Configuration Guide](https://www.pinglabz.com/802-1x-configuration-on-cisco-switches-a-practical-lesson/) 2\. [How to Configure 802.1X and WPA3 Enterprise on the Cisco C9800](https://www.pinglabz.com/c9800-8021x-wpa3-configuration/) 3\. [Enable IEEE 802.1X Authentication on Windows 11 (Manual + Group Policy)](https://www.pinglabz.com/enable-ieee-802-1x-authentication-windows-11/) 4\. [What is 802.1X Authentication? End-to-End Flow with Real show Output](https://www.pinglabz.com/what-is-802-1x-authentication/) **Studying for CCIE Security?** 802.1X, MAB, and ISE policy are the backbone of secure network access on the lab blueprint. Take identity-based access control to expert depth with the network access control and visibility track in [the CCIE Security study hub](https://www.pinglabz.com/ccie-security/). ### Build it yourself: the open-source path Most 802.1X material assumes Cisco ISE, which most engineers cannot lab. These two build the same thing on open-source software you can run at home. [802.1X against your own RADIUS server](https://www.pinglabz.com/802-1x-freeradius-cisco-full-lab/) The full supplicant, authenticator and server chain on FreeRADIUS, with the real Access-Accept. No ISE licence needed. [Letting a printer on without a supplicant](https://www.pinglabz.com/mab-freeradius-cisco-lab/) MAC Authentication Bypass on FreeRADIUS, including the MAC format gotchas and an honest word on how weak MAB really is. ## Frequently Asked Questions ### What does 802.1X stand for? 802.1X is the IEEE standard number, not an acronym. The full name is "IEEE 802.1X-2020 Port-Based Network Access Control." It defines how an authenticator (switch or wireless access point) controls access to a network at the port level using EAP-based authentication. ### What is the difference between 802.1X and Cisco ISE? 802.1X is the IEEE standard defining the protocol. Cisco ISE is a product that acts as the authentication server (the AAA/RADIUS endpoint) in an 802.1X deployment. You can run 802.1X with FreeRADIUS, NPS, or any RADIUS-speaking server; ISE provides Cisco-specific features (TrustSec/SGT, posture, dACLs, profiling). ### What is MAB and when do I use it? MAB (MAC Authentication Bypass) authenticates a device by its MAC address when no 802.1X supplicant is present. It is the fallback for printers, IP cameras, IoT, and anything else without an EAP supplicant. Configure 802.1X first, MAB second, in the authentication order. [MAB Configuration](https://www.pinglabz.com/mab-configuration-cisco-ios-xe-ise/) has the syntax. ### EAP-TLS vs PEAP, which one should I use? EAP-TLS is more secure (mutual cert auth, no password material on the wire) but requires a working PKI for every endpoint. PEAP is operationally simpler (server cert + username/password) and works well with domain-joined Windows. Most enterprises use PEAP-MSCHAPv2 for Windows users and EAP-TLS for high-security devices and BYOD. ### What ports does 802.1X / RADIUS use? EAPOL uses Ethertype 0x888E directly on Layer 2 (no IP, no UDP). RADIUS auth between switch and ISE uses UDP 1812 (modern) or UDP 1645 (legacy). RADIUS accounting uses UDP 1813 (modern) or UDP 1646 (legacy). ### What is monitor mode and why is it important? Monitor mode is an 802.1X deployment mode where the switch authenticates every device but allows access regardless of result. It is the safe way to roll 802.1X out: you discover every printer, camera, and rogue device on your network without locking anyone out. Spend a couple of weeks in monitor mode before tightening to low-impact, then closed. ## Key Takeaways If you take one thing away from this guide, make it this: 802.1X is operationally complex but conceptually simple. Three roles (supplicant, authenticator, authentication server). Two protocols (EAPOL and RADIUS). One state machine on every port (unauthorized to authorized). Pick your EAP methods deliberately, deploy in phases, and configure Critical VLAN before you ever turn closed mode on. Bookmark this page, work through the cluster articles in order, and lab every change. By the time you finish you will be able to design a wired NAC rollout from scratch and troubleshoot one at 3 AM. **Studying for the CCNA?** Test your 802.1X knowledge on [PingLabz CCNA Flashcards](https://www.pinglabz.com/ccna-flashcards/) \- 200 free multiple-choice questions by topic, mixed, or a full mock exam, each with a plain-English explanation. ### References - [IEEE 802.1X-2020 - Port-Based Network Access Control](https://standards.ieee.org/ieee/802.1X/7345/?ref=pinglabz.com) - [RFC 3748 - Extensible Authentication Protocol (EAP)](https://www.rfc-editor.org/rfc/rfc3748?ref=pinglabz.com) - [Cisco Identity Services Engine (ISE)](https://www.cisco.com/c/en/us/products/security/identity-services-engine/index.html?ref=pinglabz.com) ### Spanning Tree Protocol (STP) Complete Guide: From Fundamentals to Enterprise Hardening URL: https://www.pinglabz.com/spanning-tree-protocol/ Last updated: 2026-07-12T10:04:04.000Z Spanning Tree Protocol (STP) is the protocol that keeps Layer 2 networks loop-free. It is also the protocol that takes networks down when it goes wrong. Every Cisco campus switch you touch runs some variant of it (PVST+, Rapid PVST+, MST), and a 30-second STP convergence at 11 AM on a workday will end your week. If you understand STP cold, you become the engineer the rest of the team calls when the network has flatlined. This is the cluster overview for the full PingLabz Spanning Tree series: 25 articles covering fundamentals, the variants, configuration, hardening features, troubleshooting, and enterprise design, all built on Cisco Catalyst switches. We will work through what STP solves, how the algorithm picks a root and elects ports, the variants you need to know in 2026, and the hardening features (PortFast, BPDU Guard, Root Guard, Loop Guard) that turn STP from a footgun into a stable foundation. New labs and guides, in your inbox Every new PingLabz lab and deep-dive, built and verified on real Cisco IOS XE - free, straight to your inbox. [Join free](https://www.pinglabz.com/signup/) ## What STP Solves Ethernet has no TTL. A frame placed onto a Layer 2 loop circulates forever, multiplying every time it hits a flooding decision. Within seconds a single loop saturates every link in the broadcast domain, MAC address tables thrash, CPUs pin, and the network is unusable. This is the broadcast storm. The bridge loop problem is unavoidable in any redundant Layer 2 design: if you have two paths between two switches for redundancy, you have a loop. STP's job is to detect those loops and put exactly one port per loop into a Blocking state, while keeping the other links available for instant failover if the active path dies. The trade-off STP makes: classic 802.1D takes 30-50 seconds to converge after a topology change. That was acceptable in 1998\. It is not acceptable today, which is why every modern network runs Rapid PVST+ or MST. Detail in [What Is Spanning Tree Protocol (STP)? The Bridge Loop Problem Explained](https://www.pinglabz.com/what-is-spanning-tree-protocol/). ## How STP Works (the 10,000-Foot View) STP runs through three phases that repeat whenever the topology changes: 1. **Elect a Root Bridge.** Every switch starts believing it is the root and sends BPDUs (Bridge Protocol Data Units) advertising its bridge ID. The switch with the lowest bridge ID wins. Bridge ID = priority (16-bit, default 32768) + system ID extension (the VLAN number) + base MAC address. Lower priority wins; if priorities tie (default everywhere), the switch with the lowest MAC wins, which is rarely what you want. 2. **Elect a Root Port on every non-root switch.** The Root Port is the one with the lowest cost path to the Root Bridge. Cost is bandwidth-derived (4 for 1 Gbps, 2 for 10 Gbps, etc., on the new long-cost scale). 3. **Elect a Designated Port on every segment.** The Designated Port forwards on a given segment; the others on that segment are blocked. Tiebreakers walk a list: lowest sender root path cost, then lowest sender bridge ID, then lowest sender port priority, then lowest sender port ID. Once the dust settles you have exactly one path from every switch to the root, with one Designated Port per segment, and any other ports either Root, Alternate, or Blocking. [How STP Works: Root Bridge Election, BPDUs, and the Spanning Tree Algorithm](https://www.pinglabz.com/how-stp-works-root-bridge-bpdus/) walks through it with diagrams. ## STP Port Roles Root Port (RP) What it does Best path to the Root Bridge from a non-root switch Forwards data?Yes Designated Port (DP) What it does Best path onto a segment; one per segment Forwards data?Yes Alternate Port What it does Backup path to the Root Bridge (RSTP only) Forwards data? No, but ready to take over Backup Port What it does Backup Designated Port on the same segment (RSTP only, rare) Forwards data?No Disabled What it does Manually shut down or otherwise inactive Forwards data?No The PingLabz STP Reference Lab makes the roles visible on a real switch. SW3 is a leaf in an L2 triangle (SW1 + SW2 + SW3); SW1 was set as the explicit Root Bridge for VLAN 10 with `spanning-tree vlan 10 priority 4096`. From SW3's perspective the three port roles appear at once - the lab is small enough that the algorithm output is unambiguous: ``` SW3#show spanning-tree vlan 10 VLAN0010 Spanning tree enabled protocol rstp Root ID Priority 4106 Address 5254.008b.e4d6 Cost 4 Port 1 (GigabitEthernet0/0) Bridge ID Priority 32778 (priority 32768 sys-id-ext 10) Address 5254.008c.a6e0 Interface Role Sts Cost Prio.Nbr Type ------------------- ---- --- --------- -------- -------------------------------- Gi0/0 Root FWD 4 128.1 P2p Gi0/1 Altn BLK 4 128.2 P2p Gi0/2 Desg FWD 4 128.3 P2p Edge ``` Gi0/0 is the Root Port toward SW1 in Forwarding state with cost 4 (the IEEE 802.1D default for a Gigabit link). Gi0/1 is the Alternate Port toward SW2, blocked because the direct path to SW1 wins. Gi0/2 is a Designated edge port - the host-facing access port that has PortFast enabled, marked Type "P2p Edge". The Bridge ID Priority 32778 = 32768 default + sys-id-ext 10 (the VLAN ID) - the PVST+ encoding that keeps each VLAN's bridge ID unique even when the configured priority is identical. Full reference in [STP Port Roles Explained](https://www.pinglabz.com/stp-port-roles-explained/). ## STP Port States A port walks through several states before it forwards data, and these are where the famous 30-second convergence comes from. Classic 802.1D port states: Disabled Timen/a Forwards data?No Learns MACs?No Sends BPDUs?No Blocking Time20s (Max Age) Forwards data?No Learns MACs?No Sends BPDUs?Listens only Listening Time15s (Forward Delay) Forwards data?No Learns MACs?No Sends BPDUs?Yes Learning Time15s (Forward Delay) Forwards data?No Learns MACs?Yes Sends BPDUs?Yes Forwarding Timeindefinite Forwards data?Yes Learns MACs?Yes Sends BPDUs?Yes Add it up: 20 + 15 + 15 = 50 seconds for a Blocking port to start forwarding. That is the 802.1D convergence time. RSTP collapses it to 1-3 typical states (Discarding, Learning, Forwarding) and uses proposal/agreement handshakes to skip the timers entirely on point-to-point links, achieving sub-second failover. The lab shows per-port state in `show spanning-tree interface ... detail` \- and on a trunked link, the same physical port has independent state per VLAN. On SW3's Gi0/0 we see VLAN 10 in *root forwarding* and VLAN 99 in *alternate blocking* on the same wire: ``` SW3#show spanning-tree interface Gi0/0 detail Port 1 (GigabitEthernet0/0) of VLAN0010 is root forwarding Port path cost 4, Port priority 128, Port Identifier 128.1. Designated root has priority 4106, address 5254.008b.e4d6 Designated bridge has priority 4106, address 5254.008b.e4d6 Designated port id is 128.2, designated path cost 0 Timers: message age 15, forward delay 0, hold 0 Number of transitions to forwarding state: 1 Link type is point-to-point by default BPDU: sent 2, received 94 Port 1 (GigabitEthernet0/0) of VLAN0099 is alternate blocking Port path cost 4, Port priority 128, Port Identifier 128.1. Designated root has priority 32867, address 5254.0019.72a1 Designated bridge has priority 32867, address 5254.008b.e4d6 Designated port id is 128.2, designated path cost 4 Timers: message age 16, forward delay 0, hold 0 Number of transitions to forwarding state: 0 Link type is point-to-point by default BPDU: sent 4, received 88 ``` Two things to read in that output. **Number of transitions to forwarding state** tells you how many topology events the port has been through; a stable port has 1 (initial) or 0 (never transitioned). **BPDU: sent N, received M** shows BPDU exchange asymmetry - the upstream root is sending most of the BPDUs. This is the canonical capture for diagnosing one-way BPDU flow (a common cause of Loop Guard fires). Detail in [STP Port States: Blocking, Listening, Learning, Forwarding, and Disabled](https://www.pinglabz.com/stp-port-states-explained/). ## STP Variants: Which One Are You Running? 802.1D STP StandardIEEE 1990 Per-VLAN?No (CST) Convergence30-50s Status in 2026Legacy; do not deploy PVST+ StandardCisco Per-VLAN?Yes Convergence30-50s Status in 2026Legacy; do not deploy Rapid PVST+ Standard Cisco (based on 802.1w) Per-VLAN?Yes ConvergenceSub-second Status in 2026 Default Cisco choice for typical campus MST (802.1s) StandardIEEE Per-VLAN? Multiple instances, mapped to VLANs ConvergenceSub-second Status in 2026 Best for large campus / many VLANs The decision is usually between Rapid PVST+ and MST. Rapid PVST+ runs an STP instance per VLAN (so 100 VLANs = 100 STP instances and 100 sets of BPDUs every 2 seconds). MST runs a small number of instances (1-16) and maps multiple VLANs to each, scaling much better. If your network has more than \~50 VLANs, MST is worth the configuration effort. The dedicated comparisons are in [RSTP: What Changed from 802.1D STP](https://www.pinglabz.com/rapid-spanning-tree-protocol-rstp/) and [STP vs RSTP: Convergence, Port Roles, and When to Switch](https://www.pinglabz.com/stp-vs-rstp/). [802.1D vs PVST+ vs Rapid PVST+ vs MST](https://www.pinglabz.com/stp-variants-compared-pvst-rstp-mst/) goes deeper. Configuration walkthroughs in [Configuring Rapid PVST+ on Cisco Catalyst Switches](https://www.pinglabz.com/configure-rapid-pvst-cisco/) and [Configuring Multiple Spanning Tree (MST) on Cisco Switches](https://www.pinglabz.com/configure-mst-cisco-switches/). ## BPDUs and Path Cost STP communicates via BPDUs sent every 2 seconds (Hello timer) by all switches. Two main types: Configuration BPDU (carries root, cost, sender bridge ID) and TCN BPDU (Topology Change Notification). The arrival of a TCN tells every switch in the network "something changed, age out old MAC entries faster than usual." Path cost on Cisco's modern long-cost scale: 10 Mbps Cost (long)2,000,000 Cost (short, legacy)100 100 Mbps Cost (long)200,000 Cost (short, legacy)19 1 Gbps Cost (long)20,000 Cost (short, legacy)4 10 Gbps Cost (long)2,000 Cost (short, legacy)2 100 Gbps Cost (long)200 Cost (short, legacy)1 1 Tbps Cost (long)20 Cost (short, legacy)1 (clamps) Modern catalysts default to long-cost in software 16+. The short scale clamps at 10 Gbps, which means 10 Gbps and 100 Gbps look identical to STP, which can cause unexpected blocking. Always use long-cost. [STP Path Cost and How Cisco Switches Calculate the Best Path](https://www.pinglabz.com/stp-path-cost-calculation/) explains. ## Configuring the Root Bridge: Don't Let the Default Win If you do not set bridge priorities explicitly, the switch with the oldest MAC address becomes the root. That is almost certainly the wrong switch (it is probably an access switch in a closet). Always pick the root deliberately. The correct pattern: pick your two strongest distribution switches. Make one of them the primary root and the other the secondary root, both for every VLAN you care about: ``` DistA(config)# spanning-tree vlan 1-4094 root primary DistB(config)# spanning-tree vlan 1-4094 root secondary ``` Cisco's `root primary` macro sets the bridge priority to 24576 (or 4096 less than the current root if there is already a primary). `root secondary` sets it to 28672\. Both are well below the default 32768, so they win the election and the rest of the network does not have to care. [How to Configure the STP Root Bridge on Cisco Switches](https://www.pinglabz.com/configure-stp-root-bridge-cisco/) has the full walkthrough. Verification: ``` SW3#show spanning-tree root Root Hello Max Fwd Vlan Root ID Cost Time Age Dly Root Port ---------------- -------------------- --------- ----- --- --- ------------ VLAN0010 4106 5254.008b.e4d6 4 2 20 15 Gi0/0 VLAN0020 4116 5254.008b.e4d6 4 2 20 15 Gi0/0 VLAN0099 32867 5254.0019.72a1 4 2 20 15 Gi0/1 ``` One line per VLAN. SW1 (5254.008b.e4d6) is root for VLAN 10 (priority 4106) and VLAN 20 (4116) because of the explicit configuration. For VLAN 99 no priority was configured, so the election fell to the lowest MAC (SW2). This per-VLAN priority is the PVST+ feature that classic 802.1D did not have - you can engineer root placement separately per VLAN so the same physical trunk can be Forwarding for half your VLANs and Blocking for the other half, distributing load over the redundant links. ## STP Hardening Features STP without hardening is dangerous in a way most engineers underestimate. Five features take it from a default that anyone can disrupt to a controlled, predictable protocol: PortFast Goes onHost ports (access) What it does Skips listening/learning, port forwards immediately What it prevents 30-second DHCP delays for end hosts BPDU Guard Goes on Host ports (with PortFast) What it does Errdisables port if any BPDU is received What it prevents Rogue switches plugged into user ports BPDU Filter Goes onHost ports (sometimes) What it does Suppresses BPDU sending and receiving What it prevents Use sparingly; can mask loops if misconfigured Root Guard Goes on Designated ports facing access switches What it does Errdisables port if a superior BPDU is received What it prevents An access switch becoming the root Loop Guard Goes on Root and Alternate ports on point-to-point links What it does Blocks port if BPDUs stop arriving What it prevents Unidirectional link failures that would silently transition Blocking to Forwarding The PingLabz default: every host port gets PortFast + BPDU Guard. Every distribution-to-access link gets Root Guard on the distribution side. Every point-to-point trunk gets Loop Guard. Verify the hardening posture and STP mode on any switch with `show spanning-tree summary`: ``` SW1#show spanning-tree summary Switch is in rapid-pvst mode Root bridge for: VLAN0010, VLAN0020 Extended system ID is enabled Portfast Default is disabled Portfast Edge BPDU Guard Default is disabled Portfast Edge BPDU Filter Default is disabled Loopguard Default is disabled PVST Simulation Default is enabled but inactive in rapid-pvst mode Bridge Assurance is enabled EtherChannel misconfig guard is enabled Configured Pathcost method used is short UplinkFast is disabled BackboneFast is disabled Name Blocking Listening Learning Forwarding STP Active ---------------------- -------- --------- -------- ---------- ---------- VLAN0010 0 0 0 4 4 VLAN0020 0 0 0 4 4 VLAN0099 1 0 0 2 3 3 vlans 1 0 0 10 11 ``` Three things this output tells you in five seconds. The mode is **rapid-pvst** (good). This switch is the **Root bridge for VLAN 10 and VLAN 20**, which matches the design - that's the diagnostic line a troubleshooting flow asks first. The **Blocking column** shows VLAN 99 has one blocked port and VLAN 10 / 20 have zero, which proves the loop is being broken correctly. The hardening defaults at the top are all *disabled* here because the lab is minimal - in production you would expect to see `Portfast Edge BPDU Guard Default is enabled` after running `spanning-tree portfast bpduguard default` globally. Detail in [PortFast Configuration](https://www.pinglabz.com/portfast-configuration-cisco-switches/), [BPDU Guard Configuration](https://www.pinglabz.com/bpdu-guard-configuration-cisco/), [Root Guard and Loop Guard](https://www.pinglabz.com/root-guard-loop-guard-cisco-configuration/), and the careful-use article [Configuring BPDU Filter on Cisco Switches](https://www.pinglabz.com/bpdu-filter-configuration-cisco/). ## STP and Other Layer 2 Protocols STP does not exist in isolation. Three interactions trip people up: - **STP and EtherChannel.** STP sees a port-channel as a single logical link. If you bundle two physical links, STP treats them as one and does not block either. [STP and EtherChannel: Spanning Tree Behavior with Port Channels](https://www.pinglabz.com/stp-etherchannel-port-channel-behavior/). - **STP and Trunking.** Each VLAN has its own STP instance under PVST+ / Rapid PVST+, so the same physical trunk can be Forwarding for VLAN 10 and Blocking for VLAN 20 (load distribution by manipulating per-VLAN priority). [STP and VLAN Trunking](https://www.pinglabz.com/stp-vlan-trunking-behavior/). - **STP and HSRP/VRRP.** The active first-hop redundancy gateway should be aligned with the primary root. Otherwise traffic from access switches climbs to the secondary root and crosses the inter-distribution trunk to reach the active HSRP. [Spanning Tree and First-Hop Redundancy](https://www.pinglabz.com/stp-fhrp-hsrp-vrrp-alignment/). Two design notes belong here. PVST+ runs one spanning-tree instance per VLAN, so your [VLAN design](https://www.pinglabz.com/vlans-layer-2-switching/) and your STP topology are the same decision made twice. And the STP root bridge should sit on the same switch as the active [FHRP gateway](https://www.pinglabz.com/fhrp/), or every off-VLAN packet trombones across the inter-switch link. ## Troubleshooting: When STP Goes Wrong STP failure modes split into three buckets: - [Troubleshooting STP Loops and Broadcast Storms](https://www.pinglabz.com/troubleshoot-stp-loops-broadcast-storms/) \- the catastrophic mode. CPU pinned, every interface light flickering, links saturated. Apply Storm Control as a defense-in-depth and find the root cause. - [Troubleshooting STP Root Bridge Issues](https://www.pinglabz.com/troubleshoot-stp-root-bridge-issues/) \- the wrong switch is root, or the topology has unexpectedly converged in a way you did not design. Almost always a missed Root Guard. - [Troubleshooting Errdisable and STP Guard Features](https://www.pinglabz.com/troubleshoot-errdisable-stp-guard/) \- a port has been shut down by BPDU Guard or Root Guard. The symptom looks like a dead host; the cause is correct protection working as designed. - [Troubleshooting STP Convergence Problems and Slow Failover](https://www.pinglabz.com/troubleshoot-stp-convergence-slow-failover/) \- failover is taking too long. Almost always means classic STP / PVST+ instead of Rapid PVST+, or a misconfigured port type (link-type point-to-point not set on a P2P link). Universal first commands when STP looks wrong: ``` Switch# show spanning-tree vlan 10 Switch# show spanning-tree summary Switch# show spanning-tree inconsistentports Switch# show interfaces status err-disabled ``` The lab reproduces a controlled topology change so you can watch RSTP fail over before-and-after. Starting from the steady state (SW3 uses Gi0/0 as the Root Port toward SW1 with cost 4, Gi0/1 sits as the Alternate Port), shutting down SW3's Gi0/0 forces RSTP to promote the Alternate to Root - and because RSTP keeps the alternate already in a ready-to-forward state, the switch never goes through Listening or Learning: ``` SW3(config)#interface GigabitEthernet0/0 SW3(config-if)#shutdown ! kill the current Root Port SW3#show spanning-tree vlan 10 VLAN0010 Spanning tree enabled protocol rstp Root ID Priority 4106 Address 5254.008b.e4d6 Cost 8 <-- was 4, now 8 (2-hop path) Port 2 (GigabitEthernet0/1) <-- was Gi0/0, now Gi0/1 Interface Role Sts Cost Prio.Nbr Type ------------------- ---- --- --------- -------- -------------------------------- Gi0/1 Root FWD 4 128.2 P2p <-- was Altn BLK Gi0/2 Desg FWD 4 128.3 P2p Edge ``` Gi0/1 went from Alternate Blocking to Root Forwarding without passing through Listening or Learning - the RSTP shortcut on point-to-point links. The total path cost to root rose from 4 to 8 because the new path is SW3 -> SW2 -> SW1 (two hops) instead of SW3 -> SW1 directly. The Root Bridge identity is unchanged (still SW1, MAC 5254.008b.e4d6); only the path changed. `no shutdown` on Gi0/0 reverses the failover within seconds because the direct path has cost 4 and wins the next election. Reference of every show/debug you need is in [STP Toolkit Reference](https://www.pinglabz.com/stp-show-debug-commands-reference/). ## Design and the Hardening Checklist - [STP Design Best Practices for Enterprise Campus Networks](https://www.pinglabz.com/stp-design-best-practices-enterprise/) - [STP in Multi-Layer Campus Designs](https://www.pinglabz.com/stp-multi-layer-campus-design/) - [STP Configuration Checklist: Hardening Spanning Tree Before Go-Live](https://www.pinglabz.com/stp-hardening-checklist-go-live/) ## The Full STP Cluster, in Reading Order ### Fundamentals 1\. [What Is Spanning Tree Protocol (STP)?](https://www.pinglabz.com/what-is-spanning-tree-protocol/) 2\. [How STP Works: Root Bridge Election, BPDUs, and the Spanning Tree Algorithm](https://www.pinglabz.com/how-stp-works-root-bridge-bpdus/) 3\. [STP Port Roles Explained](https://www.pinglabz.com/stp-port-roles-explained/) 4\. [STP Port States](https://www.pinglabz.com/stp-port-states-explained/) 5\. [Understanding STP Timers](https://www.pinglabz.com/stp-timers-hello-forward-delay-max-age/) 6\. [STP Path Cost](https://www.pinglabz.com/stp-path-cost-calculation/) ### STP Variants 7\. [802.1D vs PVST+ vs Rapid PVST+ vs MST](https://www.pinglabz.com/stp-variants-compared-pvst-rstp-mst/) ### Configuration 8\. [How to Configure the STP Root Bridge on Cisco Switches](https://www.pinglabz.com/configure-stp-root-bridge-cisco/) 9\. [Configuring Rapid PVST+ on Cisco Catalyst Switches](https://www.pinglabz.com/configure-rapid-pvst-cisco/) 10\. [PortFast Configuration on Cisco Switches](https://www.pinglabz.com/portfast-configuration-cisco-switches/) 11\. [BPDU Guard Configuration](https://www.pinglabz.com/bpdu-guard-configuration-cisco/) 12\. [Root Guard and Loop Guard](https://www.pinglabz.com/root-guard-loop-guard-cisco-configuration/) 13\. [Configuring BPDU Filter on Cisco Switches](https://www.pinglabz.com/bpdu-filter-configuration-cisco/) 14\. [Configuring Multiple Spanning Tree (MST)](https://www.pinglabz.com/configure-mst-cisco-switches/) ### STP with Other Technologies 15\. [STP and EtherChannel](https://www.pinglabz.com/stp-etherchannel-port-channel-behavior/) 16\. [STP and VLAN Trunking](https://www.pinglabz.com/stp-vlan-trunking-behavior/) 17\. [Spanning Tree and First-Hop Redundancy](https://www.pinglabz.com/stp-fhrp-hsrp-vrrp-alignment/) ### Troubleshooting 18\. [Troubleshooting STP Loops and Broadcast Storms](https://www.pinglabz.com/troubleshoot-stp-loops-broadcast-storms/) 19\. [Troubleshooting STP Root Bridge Issues](https://www.pinglabz.com/troubleshoot-stp-root-bridge-issues/) 20\. [Troubleshooting Errdisable and STP Guard Features](https://www.pinglabz.com/troubleshoot-errdisable-stp-guard/) 21\. [Troubleshooting STP Convergence Problems and Slow Failover](https://www.pinglabz.com/troubleshoot-stp-convergence-slow-failover/) ### Design and Best Practices 22\. [STP Design Best Practices for Enterprise Campus Networks](https://www.pinglabz.com/stp-design-best-practices-enterprise/) 23\. [STP in Multi-Layer Campus Designs](https://www.pinglabz.com/stp-multi-layer-campus-design/) ### Reference and Checklists 24\. [STP Toolkit Reference](https://www.pinglabz.com/stp-show-debug-commands-reference/) 25\. [STP Configuration Checklist](https://www.pinglabz.com/stp-hardening-checklist-go-live/) Hands-on STP - Rapid-PVST, PortFast + BPDU Guard, Root Guard Configure Rapid-PVST root election on three IOSvL2 switches, then layer in PortFast + BPDU Guard on access ports and Root Guard on uplinks. Real captures of port roles and BPDU Guard err-disable events. Open the PingLabz CCNA Labs library. [Open the STP labs](https://www.pinglabz.com/ccna-labs-network-access/) ### More STP guides in this cluster 1\. [Cisco PVST Guide: Spanning Tree for Every VLAN](https://www.pinglabz.com/cisco-pvst-guide-spanning-tree-made-simple/) 2\. [STP and EtherChannel: When They Collide and Who Wins](https://www.pinglabz.com/etherchannel-spanning-tree/) 3\. [MSTP: Multiple Spanning Tree Protocol Explained (802.1s)](https://www.pinglabz.com/multiple-spanning-tree-protocol/) 4\. [Per-VLAN Spanning Tree (PVST+ and Rapid-PVST+) Explained](https://www.pinglabz.com/per-vlan-spanning-tree/) 5\. [RSTP: Rapid Spanning Tree Protocol Explained (802.1w)](https://www.pinglabz.com/rapid-spanning-tree-protocol/) 6\. [Spanning Tree PortFast: Faster Access Ports, Safely](https://www.pinglabz.com/spanning-tree-portfast/) ## Expert Layer 2: MST/PVST, UDLD, and the L2 Ticket Gauntlet (CCIE level) These articles close switching on PingLabz. Built on real Cisco output from a four-switch CML square - an MST region meeting a Rapid-PVST access layer, with a router-on-a-stick and the full L2 security stack - they cover the interop seams, the physical faults spanning tree cannot see, and the hardening features that separate a lab switch from a production one. Where a data-plane hardware feature could not be faithfully reproduced on the virtual switch, we say so and show the real config and operational state rather than staging fake output. 1 [MST and PVST+ interoperation: the boundary, the CIST, and the gotchas](https://www.pinglabz.com/mst-pvst-interoperation/) How an MST region presents itself to Rapid-PVST, and the PVST simulation inconsistency that blocks a whole boundary link. 2 [UDLD: detecting unidirectional links before they loop](https://www.pinglabz.com/udld-unidirectional-link-detection/) The one physical fault STP is blind to, the echo mechanism that catches it, and the copper gotcha that leaves it disabled. 3 [Storm control: stopping broadcast floods at the port](https://www.pinglabz.com/storm-control-configuration/) Cap broadcast, multicast and unknown-unicast per port, and know where to shut versus where to only rate-limit. 4 [DHCP snooping in depth: bindings, Option 82, and trusted ports](https://www.pinglabz.com/dhcp-snooping-in-depth/) The binding table every other L2 security feature is built on, and how to keep it alive across a reload. 5 [Dynamic ARP Inspection and IP Source Guard](https://www.pinglabz.com/dynamic-arp-inspection-ip-source-guard/) Turn the snooping binding table into active defence against ARP and IP spoofing - and handle the static hosts that break it. 6 [Switch administration: SDM templates, errdisable recovery, and CAM aging](https://www.pinglabz.com/switch-administration-sdm-errdisable/) The unglamorous features that decide whether a switch quietly runs or mysteriously breaks. 7 [Expert Layer 2 troubleshooting: five broken scenarios, ticket style](https://www.pinglabz.com/expert-layer-2-troubleshooting-scenarios/) Native VLAN mismatch, an MST region split, errdisable, a dead router-on-a-stick, and DAI vs a static host - each broken for real. ## Studying for the CCIE? This cluster is part of the full CCNA to CCNP to CCIE Enterprise ladder on PingLabz, every rung built on real Cisco output. For expert-level depth across every EI v1.1 blueprint domain - and the four integration Super Labs - see the [CCIE Enterprise Infrastructure study hub](https://www.pinglabz.com/ccie-enterprise/). ## Frequently Asked Questions ### What does STP stand for? STP stands for Spanning Tree Protocol, originally defined in IEEE 802.1D (1990). It is named after the graph-theory concept of a spanning tree: a subset of edges that connects every vertex without forming a cycle. STP runs that algorithm at switch level to keep Layer 2 networks loop-free. ### How many STP port states are there? Five in classic 802.1D: Disabled, Blocking, Listening, Learning, Forwarding. RSTP collapses to three: Discarding, Learning, Forwarding. The Disabled state is administrative only. ### What is RSTP and why does it matter? RSTP (Rapid Spanning Tree Protocol, 802.1w) is the 2001 update to STP that achieves sub-second convergence by replacing the classic timer-based state machine with proposal/agreement handshakes on point-to-point links. Cisco's Rapid PVST+ is RSTP run per-VLAN. Every modern campus should run Rapid PVST+ or MST, not classic 802.1D / PVST+. ### When should I use MST instead of Rapid PVST+? When you have more than about 50 VLANs and want to reduce control-plane overhead and switch CPU. MST runs a small number of STP instances (typically 1-16) and maps groups of VLANs to each, instead of running one instance per VLAN. The trade-off is configuration complexity (every switch in an MST region must have identical region config) and load-balancing granularity. ### What is the default Cisco bridge priority? 32768\. Plus the system ID extension (the VLAN number for PVST+/Rapid PVST+, the instance number for MST). You should never leave it at default on a switch you want to be root or want to keep from being root; set it explicitly with `spanning-tree vlan X root primary` or `spanning-tree vlan X priority N`. ### Should I enable PortFast on every port? On host (access) ports, yes. PortFast lets the port skip listening/learning so DHCP works in seconds rather than half a minute. On trunk ports, no. PortFast on a trunk that connects to another switch can cause loops during convergence. Always pair PortFast with BPDU Guard so the port is protected if someone connects a switch to it. ## Key Takeaways If you take one thing away from this guide, make it this: STP is one of the few protocols where defaults will hurt you. Pick your roots deliberately, run Rapid PVST+ or MST (not classic STP), and apply the hardening features uniformly. PortFast plus BPDU Guard on host ports. Root Guard at distribution-facing-access. Loop Guard on point-to-point trunks. Bookmark this page, work through the cluster articles in order, and lab every change. Spanning Tree is unforgiving, but it is also predictable once you understand it. **Studying for the CCNA?** Test your Spanning Tree knowledge on [PingLabz CCNA Flashcards](https://www.pinglabz.com/ccna-flashcards/) \- 200 free multiple-choice questions by topic, mixed, or a full mock exam, each with a plain-English explanation. ### References - [IEEE 802.1D - MAC Bridges (incorporates RSTP, formerly 802.1w)](https://standards.ieee.org/ieee/802.1D/3387/?ref=pinglabz.com) - [Cisco Spanning Tree Protocol technology documentation](https://www.cisco.com/c/en/us/tech/lan-switching/spanning-tree-protocol/index.html?ref=pinglabz.com) ### BGP (Border Gateway Protocol) URL: https://www.pinglabz.com/bgp/ Last updated: 2026-08-01T19:48:22.000Z Border Gateway Protocol (BGP) is the path-vector routing protocol that exchanges reachability information between autonomous systems, running over TCP port 179\. It is what holds the internet together: every AS learns where every prefix lives through BGP updates. Increasingly, it also runs large enterprise and data center networks - BGP is no longer just a service-provider protocol. This guide is the cluster overview for the full PingLabz BGP series: 30 articles covering fundamentals, configuration, traffic engineering, troubleshooting, security, and design, all built on a consistent Cisco IOS XE 17.x lab. If you are studying for CCNP ENCOR/ENARSI or CCIE Enterprise Infrastructure, or you just inherited a BGP-running network and need to come up to speed quickly, start here. We will work through what BGP is, how it operates, and how its attributes drive path selection, with links into the deeper articles where you need them. Take the BGP reference with you The free BGP field-reference PDF: path attributes, the 13-step best-path order, and the show commands that matter. Delivered by email, no card required. [Get the free PDF](https://www.pinglabz.com/bgp-cheatsheet/) ## What BGP Solves Interior gateway protocols (OSPF, EIGRP, IS-IS) are excellent at finding the shortest path inside a single administrative domain. They converge fast, they react to topology changes in sub-second timeframes, and they assume the routers running them are all owned and trusted by the same operator. That trust assumption breaks the moment you cross an organizational boundary. When AT&T's network needs to talk to Verizon's network, neither side is willing to flood link-state databases or run a shared SPF computation across the other operator's infrastructure. They need a protocol that can: - Exchange reachability information (which prefixes are reachable through whom) without exposing internal topology - Express rich policy about which routes to accept, prefer, or advertise, since peering relationships are commercial - Scale to a routing table that today exceeds 1,000,000 IPv4 prefixes in the default-free zone (DFZ) - Be slow on purpose: convergence trades raw speed for stability, because instability at internet scale is catastrophic BGP is that protocol. It is a path-vector protocol (it carries the AS-level path each route has traversed), it runs over TCP port 179 for reliable delivery, and it is fundamentally policy-driven rather than metric-driven. There is no "shortest" route in BGP, only the route that wins the best-path algorithm based on the attributes operators have chosen to set. ## How BGP Works (the 10,000-Foot View) Two BGP routers form a peering relationship, called a neighbor or session. The session runs over TCP/179, which means the two routers must already have IP reachability to each other (BGP itself does not discover neighbors the way OSPF does on a multi-access link). You configure each peer explicitly. Once the TCP session is up, the two peers walk through the BGP finite state machine. The neighbor states, in order, are: 1. **Idle** \- waiting to start 2. **Connect** \- trying to complete the TCP three-way handshake 3. **Active** \- TCP failed once and is being retried (counterintuitively, "Active" is a problem state, not a healthy one) 4. **OpenSent** \- we sent our OPEN message, waiting for theirs 5. **OpenConfirm** \- their OPEN looked good, waiting for KEEPALIVE 6. **Established** \- the session is up and updates can be exchanged If you ever see a peer stuck in Active, that is your TCP layer telling you something: bad neighbor IP, ACL blocking 179, mismatched source interfaces, or no route to the peer. The full state-by-state diagnostic is in [BGP Neighbor States: The 6-State FSM and How to Diagnose It](https://www.pinglabz.com/bgp-neighbor-states/). The full mechanics are covered in [How BGP Works: Peers, Updates, and the Finite State Machine](https://www.pinglabz.com/how-bgp-works/), with the four message types broken out in [BGP Message Types: OPEN, UPDATE, KEEPALIVE, and NOTIFICATION](https://www.pinglabz.com/bgp-message-types/). Once Established, peers exchange UPDATE messages containing prefixes (NLRI) and the path attributes that describe each prefix. Those updates feed into per-peer Adj-RIB-In tables, get evaluated against inbound policy, run through the best-path algorithm, and the winners are pushed into the main BGP table and (if they are best in the IP routing table too) the FIB. [How BGP Advertises Routes: NLRI, Withdrawn Routes, and the Adj-RIB](https://www.pinglabz.com/bgp-route-advertising/) walks through that pipeline. ## BGP vs IGPs: When You Reach for Each One of the most common interview questions for working network engineers is: when do you use BGP versus an IGP like OSPF? The honest answer is that they solve different problems and you almost always end up running both. The dedicated comparison is in [BGP vs OSPF: When to Use Each Routing Protocol](https://www.pinglabz.com/bgp-vs-ospf/); the short version below. Routing scope BGP Inter-AS (and increasingly intra-DC and SD-WAN overlays) OSPF / EIGRP / IS-IS Intra-AS (single administrative domain) Algorithm BGP Path-vector (AS\_PATH carried with each route) OSPF / EIGRP / IS-ISLink-state / DUAL Convergence BGP Slow on purpose (seconds to minutes) OSPF / EIGRP / IS-IS Fast (sub-second with tuning) Policy expressiveness BGP Extreme (route maps, communities, MED, LP, weight) OSPF / EIGRP / IS-ISLimited Default Cisco AD BGP20 (eBGP) / 200 (iBGP) OSPF / EIGRP / IS-IS 110 (OSPF) / 90 (EIGRP) / 115 (IS-IS) Typical prefix count BGP Hundreds of thousands to \~1M (DFZ) OSPF / EIGRP / IS-IS1,000 to 10,000 Transport BGP TCP/179 (reliable, point-to-point) OSPF / EIGRP / IS-IS IP protocol numbers, often multicast Authentication BGPMD5, TCP-AO, GTSM OSPF / EIGRP / IS-ISMD5, SHA In practice, you run an IGP underneath BGP to provide reachability between iBGP peers (so the TCP/179 sessions stay up and BGP NEXT\_HOPs resolve), and you run BGP at the edge to talk to other ASes. In a data center, BGP is now often used as the underlay too (BGP unnumbered, EVPN), but the IGP-versus-BGP split still applies conceptually. For the interior side of this decision, the same lab-driven treatment exists for both major IGPs: the [OSPF complete guide](https://www.pinglabz.com/ospf/) and the [EIGRP complete guide](https://www.pinglabz.com/eigrp/). ## eBGP vs iBGP The same protocol behaves differently depending on whether the two peers are in the same AS or different ASes. [eBGP vs iBGP: What's the Difference and When to Use Each](https://www.pinglabz.com/ebgp-vs-ibgp/) covers the full picture, but here is the short version: Default AD eBGP (different ASes)20 iBGP (same AS)200 TTL eBGP (different ASes) 1 by default (peers must be directly connected, override with `ebgp-multihop`) iBGP (same AS)255 AS\_PATH update eBGP (different ASes) Local AS prepended on advertise iBGP (same AS) Not modified between iBGP peers NEXT\_HOP update eBGP (different ASes) Rewritten to advertising router by default iBGP (same AS) Preserved (this is why `next-hop-self` exists) Topology requirement eBGP (different ASes) None beyond reachability iBGP (same AS) Full mesh, or use route reflectors / confederations Two iBGP-specific gotchas trip up almost everyone the first time: 1. **The NEXT\_HOP problem.** When an eBGP peer advertises a route into your AS, the NEXT\_HOP attribute is the eBGP peer's IP. Your iBGP peers then receive that route, but the NEXT\_HOP is unchanged, and they may have no route to it. The fix is [BGP Next-Hop-Self](https://www.pinglabz.com/bgp-next-hop-self/) on your edge router so internal peers see the advertising router as the NEXT\_HOP. 2. **The full-mesh problem.** iBGP routes received from one iBGP peer are not re-advertised to other iBGP peers (loop prevention). With *n* iBGP routers, that means you need *n(n-1)/2* sessions for everyone to learn everything. At 50 routers, that is 1,225 sessions. You do not want to manage that. The two ways out are [BGP Route Reflectors](https://www.pinglabz.com/bgp-route-reflectors/) (most common) and [BGP Confederations](https://www.pinglabz.com/bgp-confederations/). ## Path Attributes: The Heart of BGP BGP does not pick the shortest path. It picks the best path according to a 13-step algorithm that walks down a list of attributes in priority order. Understanding those attributes is the whole game; once you know them, BGP troubleshooting and traffic engineering become tractable. WEIGHT Type Cisco-proprietary, local only Affects directionOutbound Tiebreaker ruleHighest wins Set by Local router (not advertised to peers) LOCAL\_PREF Type Well-known discretionary Affects directionOutbound Tiebreaker ruleHighest wins Set by iBGP peers within an AS AS\_PATH length TypeWell-known mandatory Affects directionOutbound Tiebreaker ruleShortest wins Set by Each AS prepends on advertise ORIGIN TypeWell-known mandatory Affects directionOutbound Tiebreaker rule IGP < EGP < Incomplete Set by How the route entered BGP MED Type Optional non-transitive Affects direction Inbound (a hint to neighbors) Tiebreaker ruleLowest wins Set by Neighbor AS exit-point preference NEXT\_HOP TypeWell-known mandatory Affects directionForwarding Tiebreaker rule Must be reachable; eBGP-learned preferred Set byAdvertising router COMMUNITY TypeOptional transitive Affects directionPolicy tag Tiebreaker ruleMark-and-match Set byAny router in the path ORIGINATOR\_ID, CLUSTER\_LIST Type Optional non-transitive Affects directionLoop prevention Tiebreaker rulen/a Set byRoute reflectors The deep dive lives in [BGP Path Attributes Explained: Well-Known vs Optional](https://www.pinglabz.com/bgp-path-attributes/). The key thing to internalize: WEIGHT and LOCAL\_PREF influence what *your* AS chooses to do with traffic going out (outbound path selection from your perspective). MED and AS-Path Prepending are how you try to influence what *other* ASes do with traffic coming in. The latter is much harder, because remote operators are free to ignore your hints. ## Best Path Selection: The 13-Step Tiebreaker When BGP sees two or more paths to the same prefix, it walks this list in order and stops at the first step that picks a winner. Memorize the top 5 (the rest are edge cases): 1. Prefer the path with the highest **WEIGHT** (Cisco only, local) 2. Prefer the path with the highest **LOCAL\_PREF** 3. Prefer the path that was **locally originated** (network, redistribute, aggregate) 4. Prefer the path with the **shortest AS\_PATH** 5. Prefer the path with the lowest **ORIGIN** code (IGP < EGP < Incomplete) 6. Prefer the path with the lowest **MED** (only between paths from the same neighboring AS, by default) 7. Prefer **eBGP** over iBGP 8. Prefer the path with the lowest **IGP cost to the NEXT\_HOP** 9. Determine if multiple paths require multipath installation (ECMP) 10. For eBGP, prefer the **oldest path** (stability heuristic) 11. Prefer the path with the lowest **router ID** 12. Prefer the path with the shortest **CLUSTER\_LIST** 13. Prefer the path from the lowest **neighbor IP** If two paths still tie at step 13, BGP can install both with `maximum-paths` for ECMP load-sharing. Full algorithm walkthrough with worked examples lives in [BGP Best Path Selection Algorithm](https://www.pinglabz.com/bgp-best-path-selection/). Get the BGP Best-Path Selection cheat-sheet All 13 tiebreakers on a single printable page, with the gotcha for each rule. Free for members. [Download the cheat-sheet](https://www.pinglabz.com/bgp-cheatsheet/) ## Configuration on Cisco IOS XE: Minimum Viable BGP Here is the smallest possible eBGP setup. R1 in AS 65001 peering with R2 in AS 65002: ``` R1(config)# router bgp 65001 R1(config-router)# bgp router-id 1.1.1.1 R1(config-router)# neighbor 10.0.12.2 remote-as 65002 R1(config-router)# address-family ipv4 R1(config-router-af)# neighbor 10.0.12.2 activate R1(config-router-af)# network 192.168.1.0 mask 255.255.255.0 R1(config-router-af)# exit-address-family ``` The `network` statement does not enable BGP on an interface (that is an OSPF habit you want to unlearn). It tells BGP "if a route exactly matching 192.168.1.0/24 is in the routing table, originate it into BGP." Once both sides are configured, you should see: ``` R1# show ip bgp summary BGP router identifier 1.1.1.1, local AS number 65001 BGP table version is 3, main routing table version 3 2 network entries using 496 bytes of memory Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 10.0.12.2 4 65002 14 14 3 0 0 00:09:42 1 ``` The **State/PfxRcd** column is the diagnostic gold mine. A number means Established and that is how many prefixes the peer is advertising you. A word (Idle, Active, OpenSent) means the session is not up and you have a problem. A working session sending zero prefixes is also a problem (filter mismatch, no `activate`, etc.). The full configuration walkthrough, including iBGP, peer groups, and templates, is in [How to Configure eBGP Neighbors on Cisco IOS XE](https://www.pinglabz.com/configure-ebgp-neighbors/) and [How to Configure iBGP Neighbors and Full Mesh](https://www.pinglabz.com/configure-ibgp-neighbors/). The same neighbor and address-family model extends to IPv6 unicast - for the addressing side, start with the [IPv6 complete guide](https://www.pinglabz.com/ipv6/). ## Route Filtering and Policy You should never run BGP without filters on a public-facing session. The default Cisco behavior is "advertise everything you have" and "accept everything you are sent" - which is how route leaks happen. Three tools cover 95% of filtering needs: Prefix list Best for Filtering by network/prefix-length Match granularity Exact prefix or range (`ge`/`le`) AS-path access list Best for Filtering by AS\_PATH regex (e.g. "any path through AS 65010") Match granularity Regex on AS\_PATH string Route map Best for Combining matches and setting attributes (LP, MED, communities) Match granularity Anything, plus the ability to mutate Route maps are the universal lever. [BGP Route Maps: Matching, Setting, and Manipulating Paths](https://www.pinglabz.com/bgp-route-maps/) covers the syntax and the six most common production patterns. Prefix-list-only filtering is faster to read and lives in [BGP Route Filtering with Prefix Lists](https://www.pinglabz.com/bgp-route-filtering-prefix-lists/). And once you have multiple peering points, communities become the cleanest way to tag routes once and have policy elsewhere react to the tag - see [BGP Communities: Standard, Extended, and Large](https://www.pinglabz.com/bgp-communities/). ## Traffic Engineering with BGP BGP traffic engineering breaks into two halves: influencing the traffic *leaving* your AS (outbound, easy), and influencing the traffic *entering* your AS (inbound, hard). **Outbound**: you control your own routers, so you can use any attribute you like. WEIGHT (highest wins, Cisco-only, local-only) is the bluntest tool. [Cisco BGP Weight Attribute](https://www.pinglabz.com/bgp-weight-attribute/) covers when this is appropriate. [Manipulating BGP Local Preference for Inbound Path Selection](https://www.pinglabz.com/bgp-local-preference/) is the iBGP-friendly equivalent. **Inbound**: you have to convince other ASes to do what you want. The two main levers are MED (a hint to your direct neighbor about which of your links to prefer) and AS-Path Prepending (artificially making one path look longer so neighbors avoid it). MED only works between paths from the same neighboring AS by default, so it is useful for primary/backup at one peer but not for multi-vendor steering. AS-Path Prepending is broader but unreliable - [BGP AS-Path Prepending: When It Works and When It Doesn't](https://www.pinglabz.com/bgp-as-path-prepending/) explains why three prepends often have no effect (other operators may have a higher-priority LOCAL\_PREF policy that ignores AS-Path entirely). Two more useful TE tools sit in the same cluster: [BGP Route Aggregation and Summarization](https://www.pinglabz.com/bgp-aggregation-summarization/) for keeping your advertised footprint small, and [Originating a Default Route in BGP](https://www.pinglabz.com/bgp-default-route/) for downstream simplification. And when the physical path will not cooperate, BGP policy is often paired with [GRE tunnels](https://www.pinglabz.com/gre/) to build the path you actually want and route over it. ## Scaling iBGP: Route Reflectors and Confederations The iBGP full-mesh problem (mentioned earlier) is the thing that pushes designs toward route reflectors. With route reflectors, a small number of central routers (RRs) re-advertise iBGP routes to their clients, breaking the no-iBGP-readvertisement rule in a controlled way. You go from *O(n²)* sessions to *O(n)*. Confederations achieve a similar scaling result by splitting a single AS into sub-ASes (members) that run eBGP between each other internally and present a single AS number externally. Confederations are common in older service provider designs and rare in modern enterprise. [BGP Route Reflectors: Eliminating the Full Mesh](https://www.pinglabz.com/bgp-route-reflectors/) and [BGP Confederations: An Alternative to Route Reflectors](https://www.pinglabz.com/bgp-confederations/) have the configurations and trade-offs. Route reflectors are also the backbone of MPLS L3VPN, where MP-BGP carries VPNv4 customer routes between PEs - the [MPLS complete guide](https://www.pinglabz.com/mpls/) covers that architecture end to end. The same MP-BGP machinery also powers data-centre fabrics. **BGP EVPN** is an MP-BGP address family (`l2vpn evpn`) that advertises MAC and IP reachability between VXLAN tunnel endpoints, replacing the old flood-and-learn model. If you understand VPNv4 route distribution, EVPN uses the identical RD/RT construct for a MAC-and-host payload. See [VXLAN with BGP EVPN](https://www.pinglabz.com/vxlan-bgp-evpn-explained/) and the [Network Virtualization cluster](https://www.pinglabz.com/network-virtualization/). The VPNv4 address family is where BGP stops being an internet routing protocol and starts being a VPN transport. A customer prefix gets an 8-byte route distinguisher prepended (making it globally unique even when two customers both use 10.20.20.0/24), a route target attached as an extended community (deciding which VRFs import it), and an MPLS VPN label. The PE-to-PE session that carries all of this is ordinary iBGP with one extra address family. If you want to see it built line by line on real hardware, with the labelled traceroute at the end, read [MPLS L3VPN configuration step by step: VRF, RD, RT, and MP-BGP](https://www.pinglabz.com/mpls-l3vpn-configuration-step-by-step/). Two BGP-specific traps live in there and bite everyone once: `send-community extended` is mandatory (without it the route targets are stripped and the receiving PE accepts zero prefixes), and `as-override` is required whenever a customer reuses the same AS number at multiple sites, or BGP's own loop prevention will discard their routes. Both are dissected in [Troubleshooting MPLS L3VPN](https://www.pinglabz.com/troubleshooting-mpls-l3vpn/) and [PE-CE routing protocols in MPLS L3VPN](https://www.pinglabz.com/mpls-l3vpn-pe-ce-routing/). ## BGP Security and the State of the Internet BGP was designed in 1989\. Its security model was "we know everyone running BGP." That is no longer true, and the internet has paid for it many times. The 2008 Pakistan-YouTube hijack, the 2019 Cloudflare-Verizon route leak, and the 2022 KlaySwap incident are all in the textbooks because BGP did exactly what it was told. The protocol itself is not broken, but it trusts whoever you peer with. The modern hardening stack you should care about: - **Strict ingress filters** on every eBGP peer (prefix lists + AS-path filters). This is the single most effective control. Most route leaks are filter mistakes, not protocol attacks. - **RPKI Route Origin Validation**. RPKI lets prefix holders publish signed Route Origin Authorizations (ROAs), and your routers validate received routes against them. [BGP RPKI and Route Origin Validation on IOS XE](https://www.pinglabz.com/bgp-rpki-route-origin-validation/) walks through the IOS XE configuration. - **Session authentication** with MD5 or, on newer platforms, TCP-AO. Plus GTSM (TTL security) on eBGP sessions to defeat off-link attackers. [BGP Session Authentication: MD5 and TCP-AO](https://www.pinglabz.com/bgp-authentication/). - **maxprefix limits** on every peer to cap blast radius if a peer leaks the DFZ at you. The full picture, including the difference between hijacks and leaks and how to detect them, is in [BGP Route Leaks and Hijacks: Detection and Prevention](https://www.pinglabz.com/bgp-route-leaks-hijacks/). ## Troubleshooting: The Three Failures You Will See Most BGP problems fall into one of three buckets: the session is not coming up, the session is up but routes are not flowing, or the session is up and routes are flowing but BGP is choosing the wrong path. The PingLabz troubleshooting cluster has one article per bucket: - [Troubleshooting BGP Neighbor Adjacency Issues](https://www.pinglabz.com/troubleshoot-bgp-neighbor/) \- sessions stuck in Active or Idle - [Troubleshooting BGP Route Advertisement Problems](https://www.pinglabz.com/troubleshoot-bgp-route-advertisement/) \- routes missing from the table - [Why Is BGP Choosing the Wrong Path?](https://www.pinglabz.com/troubleshoot-bgp-path-selection/) \- 13-step algorithm in real outputs - [BGP Convergence: Timers, BFD, and Reducing Failover Time](https://www.pinglabz.com/bgp-convergence-troubleshooting/) (the BFD mechanism itself, timed live against default timers with `fall-over bfd` captures, is covered in [BFD: Sub-Second Failure Detection for OSPF, EIGRP, and BGP](https://www.pinglabz.com/bfd-bidirectional-forwarding-detection/)) - when sessions are healthy but failover is too slow ## Design and the Full BGP Cluster Once you have the fundamentals, the next question is how to lay BGP out across a real network. Two design articles cover the dominant patterns: - [BGP Design for Enterprise Networks: Dual-Homed and Multi-Homed](https://www.pinglabz.com/bgp-design-enterprise/) - [BGP Design for Service Providers: Full Tables, Peering, and Transit](https://www.pinglabz.com/bgp-design-service-provider/) The full 30-article cluster, in reading order: ### Fundamentals 1\. [What Is BGP? The Protocol That Runs the Internet](https://www.pinglabz.com/what-is-bgp/) 2\. [How BGP Works: Peers, Updates, and the Finite State Machine](https://www.pinglabz.com/how-bgp-works/) 3\. [eBGP vs iBGP: What's the Difference and When to Use Each](https://www.pinglabz.com/ebgp-vs-ibgp/) 4\. [BGP Message Types: OPEN, UPDATE, KEEPALIVE, and NOTIFICATION](https://www.pinglabz.com/bgp-message-types/) 5\. [BGP Path Attributes Explained: Well-Known vs Optional](https://www.pinglabz.com/bgp-path-attributes/) 6\. [BGP Best Path Selection Algorithm: All 13 Steps](https://www.pinglabz.com/bgp-best-path-selection/) 7\. [How BGP Advertises Routes: NLRI, Withdrawn Routes, and the Adj-RIB](https://www.pinglabz.com/bgp-route-advertising/) ### Configuration 8\. [How to Configure eBGP Neighbors on Cisco IOS XE](https://www.pinglabz.com/configure-ebgp-neighbors/) 9\. [How to Configure iBGP Neighbors and Full Mesh](https://www.pinglabz.com/configure-ibgp-neighbors/) 10\. [BGP Next-Hop-Self: When and Why You Need It](https://www.pinglabz.com/bgp-next-hop-self/) 11\. [BGP Route Reflectors: Eliminating the Full Mesh](https://www.pinglabz.com/bgp-route-reflectors/) 12\. [BGP Confederations: An Alternative to Route Reflectors](https://www.pinglabz.com/bgp-confederations/) 13\. [BGP Route Filtering with Prefix Lists](https://www.pinglabz.com/bgp-route-filtering-prefix-lists/) 14\. [BGP Route Maps: Matching, Setting, and Manipulating Paths](https://www.pinglabz.com/bgp-route-maps/) 15\. [BGP Communities: Standard, Extended, and Large](https://www.pinglabz.com/bgp-communities/) ### Traffic Engineering 16\. [Manipulating BGP Local Preference for Inbound Path Selection](https://www.pinglabz.com/bgp-local-preference/) 17\. [BGP MED (Multi-Exit Discriminator): Influencing Inbound Traffic](https://www.pinglabz.com/bgp-med/) 18\. [BGP AS-Path Prepending: When It Works and When It Doesn't](https://www.pinglabz.com/bgp-as-path-prepending/) 19\. [Cisco BGP Weight Attribute: Local Path Preference](https://www.pinglabz.com/bgp-weight-attribute/) 20\. [BGP Route Aggregation and Summarization](https://www.pinglabz.com/bgp-aggregation-summarization/) 21\. [Originating a Default Route in BGP](https://www.pinglabz.com/bgp-default-route/) 22\. [BGP Session Authentication: MD5 and TCP-AO](https://www.pinglabz.com/bgp-authentication/) ### Troubleshooting 23\. [Troubleshooting BGP Neighbor Adjacency Issues](https://www.pinglabz.com/troubleshoot-bgp-neighbor/) 24\. [Troubleshooting BGP Route Advertisement Problems](https://www.pinglabz.com/troubleshoot-bgp-route-advertisement/) 25\. [Why Is BGP Choosing the Wrong Path? Troubleshooting Best Path](https://www.pinglabz.com/troubleshoot-bgp-path-selection/) 26\. [BGP Convergence: Timers, BFD, and Reducing Failover Time](https://www.pinglabz.com/bgp-convergence-troubleshooting/) ### Security 27\. [BGP Route Leaks and Hijacks: Detection and Prevention](https://www.pinglabz.com/bgp-route-leaks-hijacks/) 28\. [BGP RPKI and Route Origin Validation on IOS XE](https://www.pinglabz.com/bgp-rpki-route-origin-validation/) ### Cross-Cluster and Deep Dives A. [BGP vs OSPF: When to Use Each Routing Protocol](https://www.pinglabz.com/bgp-vs-ospf/) B. [BGP Neighbor States: The 6-State FSM and How to Diagnose It](https://www.pinglabz.com/bgp-neighbor-states/) C. [MP-BGP: Multiprotocol BGP Address Families Explained](https://www.pinglabz.com/mp-bgp-multiprotocol-bgp/) ### Design and Architecture 29\. [BGP Design for Enterprise Networks: Dual-Homed and Multi-Homed](https://www.pinglabz.com/bgp-design-enterprise/) 30\. [BGP Design for Service Providers: Full Tables, Peering, and Transit](https://www.pinglabz.com/bgp-design-service-provider/) Solid on the CCNA foundations? - the PingLabz CCNA Labs library BGP is CCNP+ territory. The CCNA foundations - subnetting, OSPF, EIGRP, FHRP, ACLs, AAA, NAT, QoS - are covered in the 60-lab PingLabz CCNA Labs library with real captures from Cisco IOS XE 17.16\. If your CCNA fundamentals feel rusty, open the PingLabz CCNA Labs library to refresh. [Open the CCNA labs library](https://www.pinglabz.com/ccna-labs-network-fundamentals/) ### More BGP guides in this cluster 1\. [BGP Configuration on Cisco IOS XE: eBGP and iBGP](https://www.pinglabz.com/bgp-configuration/) 2\. [BGP Looking Glass: What It Is, Public Servers, and Hosting Your Own](https://www.pinglabz.com/bgp-looking-glass/) 3\. [BGP Weight: The Cisco-Only Path Attribute (and When to Use It)](https://www.pinglabz.com/bgp-weight/) 4\. [MPLS L3VPN with MP-BGP and VPNv4](https://www.pinglabz.com/mpls-l3vpn/) 4b. [MPLS L3VPN Configuration Step by Step: VRF, RD, RT, and MP-BGP](https://www.pinglabz.com/mpls-l3vpn-configuration-step-by-step/) 4c. [RD vs RT: The Most Confused Concepts in MPLS L3VPN](https://www.pinglabz.com/mpls-rd-vs-rt-explained/) 5\. [OSPF vs BGP Redistribution: Which Protocol Wins Which Tie](https://www.pinglabz.com/ospf-vs-bgp-redistribution/) 6\. [Routing Protocols Over GRE: OSPF, EIGRP, BGP](https://www.pinglabz.com/routing-protocols-over-gre/) ### Operational BGP: when it does not do what you expected [The session will not come up](https://www.pinglabz.com/bgp-neighbor-stuck-idle-active-connect/) Organised by the FSM state you are looking at right now. Active is a failure state, not a healthy one. [The prefix is in BGP but not in the RIB](https://www.pinglabz.com/bgp-route-not-in-routing-table/) Next-hop reachability, lost best-path selection, and reading the status codes most people skim past. [Which path wins, and why](https://www.pinglabz.com/bgp-best-path-selection/) The full decision order walked on captured output, with weight and local preference each proven to flip the winner. [Tag a prefix, match it, act on it](https://www.pinglabz.com/bgp-communities/) The full mechanics on IOS XE, including the send-community requirement everyone forgets. [Suppressing a prefix that keeps flapping](https://www.pinglabz.com/bgp-route-dampening/) Penalties, half-life decay and reuse limits, plus an honest account of why the technique fell out of favour. ## Expert BGP: Policy, Scale, and IPv6 Transport (CCIE level) Everything above gets you a working, safe BGP network. The articles below are the layer that separates a professional deployment from an expert one: policy that scales, features that fail in non-obvious ways, and the diagnostic habits that turn a symptom into a cause in minutes. Every one is built on real Cisco IOS XE output from a six-router CML topology - a dual-homed enterprise (AS 65001) sitting between two providers, with an IPv4-only MPLS core carrying IPv6. 1 [BGP conditional advertisement with advertise-map and non-exist-map](https://www.pinglabz.com/bgp-conditional-advertisement/) Withhold a backup prefix entirely until the primary path disappears. The only BGP mechanism that decides *whether* to announce based on live routing state. 2 [Outbound Route Filtering (ORF): pushing your prefix-list to the neighbor](https://www.pinglabz.com/bgp-outbound-route-filtering-orf/) Stop paying CPU and bandwidth to discard routes you never wanted. The capability negotiation, and the command that proves your filter arrived. 3 [BGP route dampening: penalties, half-lives, and the RIPE reversal](https://www.pinglabz.com/bgp-route-dampening/) How the penalty algorithm works, why default timers turn a 30-second outage into an hour, and the RIPE-580 timers to use instead. 4 [Designing BGP policy with communities, including a working RTBH](https://www.pinglabz.com/bgp-community-policy-design/) Origin, scope and blackhole tags as an addressing plan for policy - plus a remote-triggered black hole verified down to the CEF entry. 5 [BGP multipath and load sharing across eBGP, iBGP and eiBGP](https://www.pinglabz.com/bgp-multipath-load-sharing/) The equality rules for a second path, when to reach for multipath-relax, and the next-hop rewrite that hides a missing next-hop-self. 6 [MP-BGP for IPv6 and 6PE: IPv6 across an IPv4-only MPLS core](https://www.pinglabz.com/ipv6-bgp-6pe/) send-label, IPv4-mapped next-hops, and an IPv6 traceroute reporting an MPLS label from a router that has never heard of MPLS. 7 [Expert BGP troubleshooting: five broken scenarios, ticket style](https://www.pinglabz.com/expert-bgp-troubleshooting-scenarios/) update-source, next-hop-self, the route-map implicit deny, AS-path loops, and IPv6 activation - each broken for real, with the diagnostic trail. ## OMP: the SD-WAN analogue of BGP Engineers who know BGP inevitably ask what its Catalyst SD-WAN equivalent is, and the answer is OMP - the Overlay Management Protocol. It runs between each WAN Edge and the Controllers (which act as route reflectors), carrying routes, TLOCs and policy, with a best-path algorithm that mirrors the BGP one. A TLOC is essentially the BGP next-hop enriched with transport identity: [OMP deep dive: routes, TLOCs, service routes, and path selection](https://www.pinglabz.com/omp-deep-dive/). If you know BGP, you already know most of OMP. ## Studying for the CCIE? This cluster is part of the full CCNA to CCNP to CCIE Enterprise ladder on PingLabz, every rung built on real Cisco output. For expert-level depth across every EI v1.1 blueprint domain - and the four integration Super Labs - see the [CCIE Enterprise Infrastructure study hub](https://www.pinglabz.com/ccie-enterprise/). ## Frequently Asked Questions ### What does BGP stand for? BGP stands for Border Gateway Protocol. It is defined in [RFC 4271](https://www.rfc-editor.org/rfc/rfc4271?ref=pinglabz.com) (current version, BGP-4) and is the de facto inter-domain routing protocol of the internet. ### What port does BGP use, and why TCP? BGP uses TCP port 179\. TCP gives BGP reliable, ordered delivery and built-in congestion control, which is what you want when you are exchanging hundreds of thousands of prefixes between routers that may be hundreds of milliseconds apart. UDP would force BGP to re-implement reliability itself, and getting that right is harder than just using TCP. ### What is the difference between BGP and OSPF? OSPF is an interior gateway protocol designed for fast convergence inside a single administrative domain. BGP is an exterior gateway protocol designed for policy-driven routing between administrative domains. OSPF picks paths by metric (lowest cost wins). BGP picks paths by attributes that operators set explicitly. You almost always run both: an IGP for internal reachability, BGP at the edge. ### What is the administrative distance of BGP on Cisco? 20 for eBGP-learned routes, 200 for iBGP-learned routes. The high iBGP value is intentional: it ensures that if you also run an IGP, internal routes from the IGP win over iBGP for traffic that has an internal destination, which is usually what you want. ### What is an AS number, and what are the private ranges? An AS (autonomous system) number identifies a routing domain in BGP. Public AS numbers are assigned by the regional internet registries (RIRs) and are globally unique. The private ranges are 64512-65534 for 2-byte ASNs and 4200000000-4294967294 for 4-byte ASNs - useful for internal BGP designs (data center, SD-WAN overlay, multi-tenancy) where the AS numbers never escape your network. ### Is BGP secure? Not by default. BGP trusts whoever you peer with, which is why route leaks and hijacks happen. The modern hardening stack is: strict ingress filtering, RPKI Route Origin Validation, session authentication (MD5 or TCP-AO), GTSM, and maxprefix limits per neighbor. None of these are optional on a public-facing session in 2026. ## Key Takeaways If you take one thing away from this guide, make it this: BGP is a policy-driven protocol, not a metric-driven one. Once you internalize that, the 13-step best-path algorithm stops looking arbitrary and starts looking like a list of policy levers in priority order. Every other piece (configuration, troubleshooting, traffic engineering, security) flows from there. The protocol itself is small (four message types, a handful of well-known attributes). What makes BGP hard is operating it in a network you do not control end-to-end - which is exactly why it is the protocol that runs the internet. Bookmark this page, work through the cluster articles in the order above, and run every configuration in a lab. By the time you finish, you will be ready for any BGP question a CCIE lab or a 3 AM ticket can throw at you. **Studying for the CCNA?** Test your BGP knowledge on [PingLabz CCNA Flashcards](https://www.pinglabz.com/ccna-flashcards/) \- 200 free multiple-choice questions by topic, mixed, or a full mock exam, each with a plain-English explanation. ### References - [RFC 4271 - A Border Gateway Protocol 4 (BGP-4)](https://www.rfc-editor.org/rfc/rfc4271?ref=pinglabz.com) - [Cisco BGP technology documentation](https://www.cisco.com/c/en/us/tech/ip/border-gateway-protocol-bgp/index.html?ref=pinglabz.com) ### VLANs & Layer 2 Switching URL: https://www.pinglabz.com/vlans-layer-2-switching/ Last updated: 2026-08-01T19:48:27.000Z VLANs (Virtual LANs) are the foundation of every modern switched network. They let you carve a single physical switching fabric into multiple isolated broadcast domains, which is the only sane way to scale Ethernet past a few hundred hosts. If you are studying for CCNA, designing a campus network, or troubleshooting an inter-VLAN routing issue, VLANs are the layer 2 concept you must own cold. This is the cluster overview for the full PingLabz VLAN and Layer 2 switching series: 26 articles covering fundamentals, configuration, inter-VLAN routing, EtherChannel, troubleshooting, and campus design. Every article is written for engineers working with Cisco Catalyst switches on IOS XE. Real configs, real `show` commands, real troubleshooting. Start here, then drill into the deeper articles where you need them. New labs and guides, in your inbox Every new PingLabz lab and deep-dive, built and verified on real Cisco IOS XE - free, straight to your inbox. [Join free](https://www.pinglabz.com/signup/) ## What a VLAN Actually Is A VLAN is a logical broadcast domain. Two hosts on the same VLAN behave as if they were on the same physical Ethernet segment, even if they are connected to different switches in different buildings. Two hosts on different VLANs cannot reach each other at Layer 2 at all; they need Layer 3 routing to communicate. The problem VLANs solve: a flat Layer 2 network is fragile. One misbehaving host floods broadcasts to every other host. Spanning Tree gets harder to reason about. Security policy lives only in firewalls at the edge, with nothing in the middle. VLANs let you draw boundaries inside the switching fabric, which gives you four immediate wins: - **Smaller broadcast domains** \- a broadcast storm on one VLAN does not affect the others - **Logical grouping** \- hosts can share a VLAN regardless of physical location (same building, same room, even same closet) - **Security boundaries** \- inter-VLAN traffic must traverse a router or L3 switch where you can apply ACLs and inspection - **Reusable IP space** \- each VLAN gets its own subnet, so you can plan the address space once and re-use the pattern across sites The full intro is in [What Is a VLAN? Virtual LANs Explained for Network Engineers](https://www.pinglabz.com/what-is-a-vlan/), with the broadcast-domain mechanics in [How VLANs Work: Tagging, Broadcast Domains, and Frame Forwarding](https://www.pinglabz.com/how-vlans-work/). ## How VLANs Work on the Wire A switch tracks which VLAN each port belongs to in its CAM table. When a frame arrives on an access port, the switch tags it internally with the VLAN ID, looks up the destination MAC, and forwards only to ports in the same VLAN. If the destination MAC is unknown, the flood is also limited to the same VLAN. So far so simple. The complication arrives when frames need to traverse between switches: how does the second switch know which VLAN a frame belonged to? The answer is 802.1Q trunking. On a trunk port, the switch inserts a 4-byte 802.1Q tag into the Ethernet header carrying the VLAN ID, the receiving switch strips it back off, and the VLAN identity is preserved across the link. The 802.1Q tag has four fields: TPID (always 0x8100), PCP (3-bit priority for QoS), DEI (drop-eligible indicator), and the VLAN ID itself (12 bits, so theoretical range 0-4095, but 0 and 4095 are reserved, leaving 1-4094 usable). Frame size grows from 1518 bytes to 1522 bytes, which is why you sometimes see "baby giant" support on switch ports. The byte-by-byte deep dive is in [802.1Q VLAN Tag Explained: The 4 Bytes That Make Trunking Work](https://www.pinglabz.com/802-1q-vlan-tag-explained/). ## Port Modes: Access, Trunk, Dynamic Every Cisco switch port operates in one of three modes: Access CarriesOne VLAN TaggingUntagged Use for End hosts (PCs, printers, IP phones via voice VLAN) Trunk CarriesMultiple VLANs Tagging Tagged with 802.1Q (except native VLAN) Use for Switch-to-switch links, switch-to-router links Dynamic auto/desirable CarriesNegotiates via DTP TaggingNegotiates via DTP Use forDon't (security risk) Always set ports explicitly. Dynamic Trunking Protocol (DTP) is enabled by default on every Catalyst port and is a documented attack vector (a malicious host can negotiate trunking, see all VLANs, and pivot - see [VLAN Hopping Attacks: Switch Spoofing and Double Tagging](https://www.pinglabz.com/vlan-hopping-attacks/) for the full attack walkthrough). The PingLabz hardening pattern: `switchport mode access` on every host port and `switchport nonegotiate` on every trunk to disable DTP. The per-port view in `show interfaces ... switchport` spells out the mode and the VLAN assignments. Compare an access port (with a voice VLAN) against a trunk port from the same lab switch: ``` SW1#show interfaces Gi0/2 switchport Name: Gi0/2 Switchport: Enabled Administrative Mode: static access Operational Mode: down Administrative Trunking Encapsulation: negotiate Negotiation of Trunking: Off Access Mode VLAN: 10 (DATA) Trunking Native Mode VLAN: 1 (default) Voice VLAN: 20 (VOICE) ... SW1#show interfaces Gi0/0 switchport Administrative Mode: trunk Operational Mode: trunk Access Mode VLAN: 1 (default) Trunking Native Mode VLAN: 99 (NATIVE) Administrative Native VLAN tagging: enabled Voice VLAN: none Administrative private-vlan trunk encapsulation: dot1q Trunking VLANs Enabled: 10,20,99 ``` Three things to read in those two outputs. **Administrative Mode** is the configured intent (*static access* on Gi0/2, *trunk* on Gi0/0); "static" means DTP will not talk the port into anything else. **Access Mode VLAN: 10 (DATA)** on Gi0/2 plus **Voice VLAN: 20 (VOICE)** is the IP-phone-on-the-same-cable pattern - the phone tags voice VLAN 20, the PC behind it sends untagged frames that fall into VLAN 10\. **Trunking VLANs Enabled: 10,20,99** on Gi0/0 is the explicit allowed-VLAN list - the default would be "ALL", which is the wrong production default. Detail in [DTP: How It Works and Why You Should Disable It](https://www.pinglabz.com/dtp-dynamic-trunking-protocol/) and [VLAN Access Ports and Switchport Modes](https://www.pinglabz.com/vlan-access-ports-switchport-modes/). ## 802.1Q Trunking and the Native VLAN Quirk The full mental model for tagged vs untagged frames (and the cross-vendor vocabulary) is in [Tagged vs Untagged VLANs: How Trunks and Access Ports Really Work](https://www.pinglabz.com/tagged-vs-untagged-vlans/). On a trunk port, every VLAN gets tagged with 802.1Q except one: the native VLAN. The native VLAN is sent untagged, which exists for backwards compatibility with hubs and other devices that do not understand 802.1Q. By default, the native VLAN is VLAN 1. This default is dangerous for two reasons: 1. **VLAN 1 is everywhere.** CDP, VTP, PAgP, DTP all use VLAN 1\. Leaving the native VLAN as VLAN 1 means you are mixing control plane and user data on the same untagged segment. 2. **Native VLAN hopping attacks.** If an attacker is in the native VLAN, they can inject double-tagged frames that the first switch strips one tag from, then forwards to the next switch with the inner tag intact - landing in any VLAN they target. [VLAN Security Hardening](https://www.pinglabz.com/vlan-security-best-practices/) covers this and the complete set of L2 controls you need in production. The fix is simple: change the native VLAN to a dedicated unused VLAN (e.g. VLAN 999) and never let a host port be in it. [Native VLAN Configuration and Security on Cisco Switches](https://www.pinglabz.com/change-native-vlan-cisco-switch/) walks through the change. Trunk configuration itself is straightforward: ``` Switch(config)# interface GigabitEthernet1/0/24 Switch(config-if)# switchport mode trunk Switch(config-if)# switchport trunk encapsulation dot1q Switch(config-if)# switchport trunk native vlan 999 Switch(config-if)# switchport trunk allowed vlan 10,20,30,40 Switch(config-if)# switchport nonegotiate ``` Notice the `allowed vlan` list. By default trunks carry every VLAN that exists on the switch, which is rarely what you want; explicitly listing VLANs reduces the broadcast spread and limits attack surface. The trunk inventory in `show interfaces trunk` consolidates every trunk on the switch into one capture, including the per-trunk allowed VLAN list and what is actually forwarding through STP: ``` SW1#show interfaces trunk Port Mode Encapsulation Status Native vlan Gi0/0 on 802.1q trunking 99 Gi0/1 on 802.1q trunking 99 Po1 on 802.1q trunking 99 Port Vlans allowed on trunk Gi0/0 10,20,99 Gi0/1 10,20,99 Po1 10,20,99 Port Vlans allowed and active in management domain Gi0/0 10,20,99 Gi0/1 10,20,99 Po1 10,20,99 Port Vlans in spanning tree forwarding state and not pruned Gi0/0 10,20,99 Gi0/1 10,20,99 Po1 10,20 ``` Three things to notice. Mode `on` means unconditional trunk (no DTP negotiation, the production default). Native VLAN 99 is set explicitly across all three trunks, not left at VLAN 1\. The fourth table is the most useful diagnostic: VLAN 99 appears in the allowed list for Po1 but NOT in its STP-forwarding column, because STP is breaking a loop on the native VLAN by blocking Po1 for VLAN 99 specifically. Reading those four tables against each other immediately tells you whether a trunk problem is allowed-list, native-VLAN, or STP-pruning. [Configuring 802.1Q Trunks on Cisco Catalyst Switches](https://www.pinglabz.com/cisco-8021q-trunking-lab-guide/) has the full pattern. ## Basic VLAN Configuration on Cisco IOS XE The minimum to create a VLAN and assign a host port: ``` Switch(config)# vlan 10 Switch(config-vlan)# name USERS Switch(config-vlan)# exit Switch(config)# interface GigabitEthernet1/0/1 Switch(config-if)# switchport mode access Switch(config-if)# switchport access vlan 10 Switch(config-if)# switchport nonegotiate Switch(config-if)# spanning-tree portfast Switch(config-if)# spanning-tree bpduguard enable ``` The last two lines are not optional in production. PortFast skips the listening/learning states for end-host ports (so PCs DHCP cleanly within seconds rather than 30+ seconds), and BPDU Guard error-disables the port immediately if anyone plugs a switch into it. See the STP cluster at [Spanning Tree Protocol (STP)](https://www.pinglabz.com/spanning-tree-protocol/) for the why. Verification on the lab switch: ``` SW1#show vlan brief VLAN Name Status Ports ---- -------------------------------- --------- ------------------------------- 1 default active 10 DATA active Gi0/2 20 VOICE active Gi0/2 99 NATIVE active 1002 fddi-default act/unsup 1003 token-ring-default act/unsup 1004 fddinet-default act/unsup 1005 trnet-default act/unsup ``` VLAN 1 (default) always exists - you cannot delete it. VLAN 10/20/99 were created by config. Gi0/2 appears under **both** VLAN 10 and VLAN 20 because of the voice-VLAN configuration (data on VLAN 10, voice on VLAN 20, same physical port). Trunk ports do NOT appear in this list - trunks are not "assigned" to a VLAN, they carry many. The 1002-1005 entries are the FDDI/Token Ring reserved range from the original 1990s spec; they cannot be used for production VLANs. Detail and edge cases in [Configuring VLANs on Cisco Catalyst Switches](https://www.pinglabz.com/configuring-standard-vlans-on-catalyst-switches/) and [VLAN Naming, Ranges, and Management Best Practices on Cisco IOS XE](https://www.pinglabz.com/master-cisco-vlan-configuration-step-by-step-guide/). ## Inter-VLAN Routing: Three Ways to Cross VLAN Boundaries VLANs by themselves only isolate; they do not communicate. To let one VLAN reach another you need a Layer 3 device. The three options, in order of modern preference: L3 Switch with SVIs Where the routing happens Hardware ASIC on the switch Pros Wire-rate, scales to many VLANs, modern default Cons Requires L3 license / hardware Router-on-a-stick Where the routing happens External router with subinterfaces over a single trunk Pros Cheap, works on any router Cons Bottleneck on the trunk; software routing External router with multiple physical interfaces Where the routing happens One physical link per VLAN ProsConceptually simple Cons Doesn't scale past a few VLANs The L3 switch SVI pattern is what almost every modern campus uses. You configure a Switch Virtual Interface (SVI) for each VLAN with the gateway IP, enable IP routing on the switch, and the ASIC routes between VLANs at line rate: ``` Switch(config)# ip routing Switch(config)# interface vlan 10 Switch(config-if)# ip address 10.10.10.1 255.255.255.0 Switch(config-if)# no shutdown Switch(config)# interface vlan 20 Switch(config-if)# ip address 10.10.20.1 255.255.255.0 Switch(config-if)# no shutdown ``` The full walkthrough for each approach lives in [Configuring SVIs for Inter-VLAN Routing](https://www.pinglabz.com/cisco-vlan-ip-configuration-guide-step-by-step/), [Inter-VLAN Routing with Router-on-a-Stick](https://www.pinglabz.com/inter-vlan-routing-router-on-a-stick/), and [Inter-VLAN Routing on a Layer 3 Switch](https://www.pinglabz.com/inter-vlan-routing-layer-3-switch/). Whichever routing method you pick, the gateway itself should never be a single point of failure - pair it with HSRP, VRRP, or GLBP from the [FHRP complete guide](https://www.pinglabz.com/fhrp/). ## Specialized VLANs: Voice, Private, Management Beyond plain data VLANs, three specialized variants come up constantly: **Voice VLANs.** Cisco IP phones tag their voice traffic with one VLAN ID and pass through PC traffic untagged on a different VLAN. The same physical port carries both. [Configuring Voice VLANs on Cisco Switches for IP Phones](https://www.pinglabz.com/voice-vlan-cisco-configuration/) covers the syntax (`switchport voice vlan`) and the QoS implications. **Private VLANs (PVLANs).** Sometimes you need hosts in the same subnet to be isolated from each other (think shared hosting, hotel guest networks, certain DMZs). PVLANs subdivide a VLAN into Primary, Isolated, and Community sub-domains, all sharing one Layer 3 gateway. [Private VLANs on Cisco Catalyst Switches](https://www.pinglabz.com/private-vlans-cisco-configuration/) walks through the configuration. **Management VLAN.** The VLAN your network management traffic (SSH, SNMP, syslog, NTP) lives in. Should always be a dedicated VLAN, never VLAN 1, and accessible only from your jump hosts. Set the SVI in this VLAN as the source for management protocols. ## VTP: Use With Caution VLAN Trunking Protocol (VTP) propagates VLAN database changes across switches to save you typing. It also has a famous failure mode: a switch with a higher VTP revision number plugged into a VTP domain wipes every other switch's VLAN database to match its own, and entire networks have gone down because someone connected a lab switch to production. The PingLabz default: VTP transparent mode on every switch (each switch maintains its own VLAN database, propagates VTP advertisements but ignores them). VTP version 3 with controlled primary/secondary servers is acceptable in tightly-managed environments. [VTP Configuration on Cisco Switches](https://www.pinglabz.com/master-configuring-vtp-clients-and-servers-cisco-switches/) has the full pattern, including the rollover-incident-prevention drill. ## EtherChannel: Bundling Links for Redundancy and Bandwidth EtherChannel (also called LAG, Link Aggregation, or port-channel) bundles 2-8 physical links into a single logical link. It is essential for two reasons: it defeats the Layer 2 redundancy paradox (multiple links between the same two switches without Spanning Tree blocking some), and it scales bandwidth linearly without requiring faster physical interfaces. Three control protocols: LACP StandardIEEE 802.3ad Modesactive, passive When to use Default modern choice; vendor-neutral PAgP StandardCisco-proprietary Modesdesirable, auto When to use Cisco-only environments; legacy Static (on) StandardNone Modeson When to use Avoid; no negotiation means misconfig fails silently LACP is the modern default. Both ends should be in `active`, with `passive` only used on one side intentionally. Load balancing is configured separately and matters for actual throughput; the default `src-dst-ip` usually works but is worth checking for high-throughput links. The lab has a single channel-group between SW1 and SW2 running LACP active on both ends; `show etherchannel summary` is the one-line health check: ``` SW1#show etherchannel summary Flags: D - down P - bundled in port-channel I - stand-alone s - suspended H - Hot-standby (LACP only) R - Layer3 S - Layer2 U - in use N - not in use, no aggregation f - failed to allocate aggregator ... Number of channel-groups in use: 1 Number of aggregators: 1 Group Port-channel Protocol Ports ------+-------------+-----------+----------------------------------------------- 1 Po1(SU) LACP Gi0/3(P) ``` Two flags do the work. **Po1(SU)** means the port-channel is *in use* and *Layer 2* \- the bundle is forwarding and is a switchport (not a routed link). **Gi0/3(P)** means the physical port is *bundled in the port-channel*. Protocol LACP confirms dynamic negotiation; if the peer had been misconfigured as `mode on` the flag would shift to *s* (suspended) and the channel would never form - a common interop failure that this exact capture catches in five seconds. [EtherChannel Fundamentals](https://www.pinglabz.com/etherchannel-configuration-on-cisco-switches/) and [Configuring EtherChannel with LACP on Cisco Catalyst Switches](https://www.pinglabz.com/configure-lacp-etherchannel-cisco-ios/) have the full configurations. ## Layer 2 Security: The VLAN Hardening Checklist Layer 2 attacks bypass every firewall above them. The minimum production hardening: - **Disable DTP everywhere.** `switchport nonegotiate` on every port that is not specifically engineered for trunking negotiation. - **Move the native VLAN** away from VLAN 1 to a dedicated, unused VLAN. Disable hosts in this VLAN. - **Prune VLANs from trunks.** Use `switchport trunk allowed vlan` with an explicit list, not the default. - **BPDU Guard on every host port.** Combined with PortFast, this prevents an attacker from connecting a rogue switch. - **Storm Control** on host ports to limit broadcast/multicast flood from a misbehaving NIC. - **DHCP Snooping** \+ **Dynamic ARP Inspection** to block rogue DHCP servers and ARP spoofing. - **Port Security** to limit MAC addresses per port (with caution on voice + PC ports, which legitimately need 2). - **Disable VLAN 1.** No SVI, no host ports, no trunks carrying it. The full hardening pattern with configurations is in [VLAN Security Hardening: Protecting Your Layer 2 Network](https://www.pinglabz.com/vlan-security-best-practices/). ### L2 Security: The Switch Is an Attack Surface The switch is not neutral plumbing, it is a live attack surface. ARP poisoning, intra-VLAN lateral movement, and host-to-host pivots all play out below the firewall, on the same VLAN, where nothing above Layer 2 can see them. Cisco IOS XE answers in layers. Start with the [DHCP snooping binding table](https://www.pinglabz.com/dhcp-snooping-in-depth/), promote it to active defence with [Dynamic ARP Inspection and IP Source Guard](https://www.pinglabz.com/dynamic-arp-inspection-ip-source-guard/), then cover the statically-addressed hosts DAI would otherwise break using [static ARP ACLs](https://www.pinglabz.com/static-arp-acl-inspection/). Filter east-west traffic inside a single VLAN with [VLAN ACLs (VACLs)](https://www.pinglabz.com/vlan-acls-vacl-configuration/), isolate hosts that share a subnet with [private VLANs](https://www.pinglabz.com/private-vlans-explained/), and cap floods with [storm control](https://www.pinglabz.com/storm-control-configuration/). The capstone that ties every control together is the [Layer 2 attack surface hardening checklist](https://www.pinglabz.com/layer-2-attack-surface-hardening/). ## Campus VLAN Design The dominant pattern for campus networks is access-distribution-core. VLANs live at the access layer, get extended (or terminated) at the distribution layer, and the core routes between distribution blocks. Two key design choices: 1. **Where do you terminate Layer 2?** Modern best practice: terminate at the access layer (each access switch is its own L3 device, no VLANs span between access switches). This eliminates Spanning Tree as a critical failure mode and lets you use ECMP routing for redundancy. 2. **How do you address?** One VLAN per access switch (small subnet, /24 or smaller) is cleaner than VLANs spanning multiple closets. Voice and data on the same port using voice VLAN. The full design pattern with diagrams and worked examples is in [VLAN Design for Campus Networks: From Access to Core](https://www.pinglabz.com/vlan-design-campus-network/). ## Layer 2 Security Cross-References The switching features here (port security, DHCP snooping, DAI) are the IPv4 access-edge defences. Their IPv6 counterparts, and the device-level controls that protect the switches themselves, live in the [Infrastructure Security cluster](https://www.pinglabz.com/infrastructure-security/): [IPv6 RA Guard and DHCPv6 Guard](https://www.pinglabz.com/ipv6-ra-guard-dhcpv6-guard/) (the IPv6 parallel to DHCP snooping), [MACsec](https://www.pinglabz.com/macsec-802-1ae-explained/) for link-layer encryption, and [management-plane hardening](https://www.pinglabz.com/hardening-management-plane-ios-xe/) for the switch CLI itself. ## Beyond VLANs: VXLAN and the Overlay VLANs run out of room at two points: the 4094-ID ceiling, and the scale limits of large flat Layer 2\. **VXLAN** is the answer to both, wrapping the Ethernet frame in UDP and routing it across a fabric with a 24-bit segment ID (16 million segments). It does not replace VLANs; it carries them (as L2VNIs) beyond the reach of a physical switch. See [VXLAN deep dive](https://www.pinglabz.com/vxlan-deep-dive/) for the byte-level encapsulation, and [VRF, VLAN, VXLAN, LISP: choosing the right segmentation layer](https://www.pinglabz.com/vrf-vlan-vxlan-lisp-segmentation/) for when to reach for which. The full overlay story is in the [Network Virtualization cluster](https://www.pinglabz.com/network-virtualization/). ## Troubleshooting: Where VLAN Problems Hide - [Troubleshooting VLAN and Trunk Problems on Cisco Switches](https://www.pinglabz.com/troubleshoot-vlan-trunk-issues/) \- allowed list mismatches, native VLAN mismatches, and DTP misfires - [Troubleshooting Inter-VLAN Routing](https://www.pinglabz.com/troubleshoot-inter-vlan-routing/) \- missing SVIs, ACLs, ARP issues - [Troubleshooting SVI Up/Down Issues](https://www.pinglabz.com/fixing-cisco-vlan-interface-down/) \- the SVI that won't come up The one obscure-but-useful diagnostic worth knowing about is `show vlan internal usage`. Catalyst switches use the extended VLAN range (1006-4094) for internal allocations when you configure routed ports, L3 SVIs, or certain hardware features - and if you later try to configure a user VLAN on a number the platform has silently grabbed, the VLAN will refuse to come up with an unhelpful error. On the small lab switch the table is empty because no L3 features are configured: ``` SW1#show vlan internal usage VLAN Usage ---- -------------------- ``` On a production Catalyst 9000 with several L3 routed ports and SVIs, this same command shows the dynamic VLAN-to-feature mappings. The diagnostic move: if VLAN N refuses to come up and the error does not explain why, this is the command that explains it. ## The Full VLAN Cluster, in Reading Order ### Fundamentals 1\. [What Is a VLAN? Virtual LANs Explained for Network Engineers](https://www.pinglabz.com/what-is-a-vlan/) 2\. [How VLANs Work: Tagging, Broadcast Domains, and Frame Forwarding](https://www.pinglabz.com/how-vlans-work/) 3\. [VLAN Access Ports and Switchport Modes](https://www.pinglabz.com/vlan-access-ports-switchport-modes/) 4\. [VLAN Trunking Explained: 802.1Q, Allowed Lists, and Trunk Negotiation](https://www.pinglabz.com/understanding-vlan-trunking-a-network-admins-guide-to-cisco-trunks/) 5\. [DTP (Dynamic Trunking Protocol)](https://www.pinglabz.com/dtp-dynamic-trunking-protocol/) ### VLAN Configuration 6\. [Configuring VLANs on Cisco Catalyst Switches](https://www.pinglabz.com/configuring-standard-vlans-on-catalyst-switches/) 7\. [VLAN Naming, Ranges, and Management Best Practices](https://www.pinglabz.com/master-cisco-vlan-configuration-step-by-step-guide/) 8\. [Configuring 802.1Q Trunks on Cisco Catalyst Switches](https://www.pinglabz.com/cisco-8021q-trunking-lab-guide/) 9\. [Native VLAN Configuration and Security](https://www.pinglabz.com/change-native-vlan-cisco-switch/) 10\. [VTP Configuration on Cisco Switches](https://www.pinglabz.com/master-configuring-vtp-clients-and-servers-cisco-switches/) ### SVIs and Inter-VLAN Routing 11\. [Configuring SVIs for Inter-VLAN Routing](https://www.pinglabz.com/cisco-vlan-ip-configuration-guide-step-by-step/) 12\. [Inter-VLAN Routing with Router-on-a-Stick](https://www.pinglabz.com/inter-vlan-routing-router-on-a-stick/) 13\. [Inter-VLAN Routing on a Layer 3 Switch](https://www.pinglabz.com/inter-vlan-routing-layer-3-switch/) ### Specialized VLANs 14\. [Configuring Voice VLANs on Cisco Switches for IP Phones](https://www.pinglabz.com/voice-vlan-cisco-configuration/) 15\. [Private VLANs on Cisco Catalyst Switches](https://www.pinglabz.com/private-vlans-cisco-configuration/) ### EtherChannel 16\. [EtherChannel Fundamentals](https://www.pinglabz.com/etherchannel-configuration-on-cisco-switches/) 17\. [Configuring EtherChannel with LACP](https://www.pinglabz.com/configure-lacp-etherchannel-cisco-ios/) ### Troubleshooting 18\. [Troubleshooting VLAN and Trunk Problems](https://www.pinglabz.com/troubleshoot-vlan-trunk-issues/) 19\. [Troubleshooting Inter-VLAN Routing](https://www.pinglabz.com/troubleshoot-inter-vlan-routing/) 20\. [Troubleshooting SVI Up/Down Issues](https://www.pinglabz.com/fixing-cisco-vlan-interface-down/) ### Design and Security 21\. [VLAN Security Hardening](https://www.pinglabz.com/vlan-security-best-practices/) 22\. [Static ARP ACLs: Inspecting ARP Without DHCP Snooping](https://www.pinglabz.com/static-arp-acl-inspection/) 23\. [VLAN ACLs (VACLs): Filtering Traffic Inside a VLAN](https://www.pinglabz.com/vlan-acls-vacl-configuration/) 24\. [Private VLANs: Isolating Hosts That Share a Subnet](https://www.pinglabz.com/private-vlans-explained/) 25\. [The Layer 2 Attack Surface: A Hardening Checklist That Actually Holds](https://www.pinglabz.com/layer-2-attack-surface-hardening/) 26\. [VLAN Design for Campus Networks](https://www.pinglabz.com/vlan-design-campus-network/) Hands-on VLANs - 14 CCNA Network Access labs Configure VLANs, 802.1Q trunks, VTP, DTP, voice VLAN, EtherChannel (LACP + PAgP), and inter-VLAN routing on three Cisco IOSvL2 switches. VLANs+Trunks+VTP lab is free preview. Open the PingLabz CCNA Labs library. [Open the labs](https://www.pinglabz.com/ccna-labs-network-access/) ### More VLAN guides in this cluster 1\. [Understanding Collision Domains](https://www.pinglabz.com/collision-domain-guide/) 2\. [STP and EtherChannel: When They Collide and Who Wins](https://www.pinglabz.com/etherchannel-spanning-tree/) 3\. [Understanding Ethernet LAN Fundamentals for CCNA](https://www.pinglabz.com/ethernet-lan-fundamentals-ccna/) 4\. [Mastering ‘Show Interface Status’ in Cisco: Top 5 Essential Tips](https://www.pinglabz.com/how-to-show-interface-status-in-cisco-switch/) 5\. [Inter-VLAN Routing: SVI vs Router-on-a-Stick (with Real IOS XE Config)](https://www.pinglabz.com/inter-vlan-routing/) 6\. [Native VLAN Explained: Untagged Traffic and VLAN Hopping](https://www.pinglabz.com/native-vlan/) 7\. [Per-VLAN Spanning Tree (PVST+ and Rapid-PVST+) Explained](https://www.pinglabz.com/per-vlan-spanning-tree/) 8\. [Local Area Network Basics: Understanding Modern LANs](https://www.pinglabz.com/understanding-modern-lans/) 9\. [VLAN, Subnet, and Broadcast Domain: What's the Difference?](https://www.pinglabz.com/vlan-subnet-broadcast-domain-difference/) 10\. [VTP (VLAN Trunking Protocol): v1, v2, v3 and Why It's Dangerous](https://www.pinglabz.com/vlan-trunking-protocol/) 11\. [VLAN vs VXLAN: The L2 Overlay, Demystified](https://www.pinglabz.com/vlan-vs-vxlan/) ### When the switch port stops behaving [The port went err-disabled](https://www.pinglabz.com/errdisable-recovery-cisco/) The full cause list, how to identify which one fired, and when auto-recovery is a genuinely bad idea. [CDP is warning about a native VLAN mismatch](https://www.pinglabz.com/native-vlan-mismatch-troubleshooting/) Reading the two VLAN numbers out of the log line, what actually happens to the traffic, and why the warning disappears if you turn CDP off while the problem does not. ## Expert Layer 2: MST/PVST, UDLD, and the L2 Ticket Gauntlet (CCIE level) These articles close switching on PingLabz. Built on real Cisco output from a four-switch CML square - an MST region meeting a Rapid-PVST access layer, with a router-on-a-stick and the full L2 security stack - they cover the interop seams, the physical faults spanning tree cannot see, and the hardening features that separate a lab switch from a production one. Where a data-plane hardware feature could not be faithfully reproduced on the virtual switch, we say so and show the real config and operational state rather than staging fake output. 1 [MST and PVST+ interoperation: the boundary, the CIST, and the gotchas](https://www.pinglabz.com/mst-pvst-interoperation/) How an MST region presents itself to Rapid-PVST, and the PVST simulation inconsistency that blocks a whole boundary link. 2 [UDLD: detecting unidirectional links before they loop](https://www.pinglabz.com/udld-unidirectional-link-detection/) The one physical fault STP is blind to, the echo mechanism that catches it, and the copper gotcha that leaves it disabled. 3 [Storm control: stopping broadcast floods at the port](https://www.pinglabz.com/storm-control-configuration/) Cap broadcast, multicast and unknown-unicast per port, and know where to shut versus where to only rate-limit. 4 [DHCP snooping in depth: bindings, Option 82, and trusted ports](https://www.pinglabz.com/dhcp-snooping-in-depth/) The binding table every other L2 security feature is built on, and how to keep it alive across a reload. 5 [Dynamic ARP Inspection and IP Source Guard](https://www.pinglabz.com/dynamic-arp-inspection-ip-source-guard/) Turn the snooping binding table into active defence against ARP and IP spoofing - and handle the static hosts that break it. 6 [Switch administration: SDM templates, errdisable recovery, and CAM aging](https://www.pinglabz.com/switch-administration-sdm-errdisable/) The unglamorous features that decide whether a switch quietly runs or mysteriously breaks. 7 [Expert Layer 2 troubleshooting: five broken scenarios, ticket style](https://www.pinglabz.com/expert-layer-2-troubleshooting-scenarios/) Native VLAN mismatch, an MST region split, errdisable, a dead router-on-a-stick, and DAI vs a static host - each broken for real. ## Studying for the CCIE? This cluster is part of the full CCNA to CCNP to CCIE Enterprise ladder on PingLabz, every rung built on real Cisco output. For expert-level depth across every EI v1.1 blueprint domain - and the four integration Super Labs - see the [CCIE Enterprise Infrastructure study hub](https://www.pinglabz.com/ccie-enterprise/). **Studying for CCIE Security?** Layer 2 attacks and the VLAN hardening checklist map directly onto the network security domain. Take switch-port security and Layer 2 threat mitigation to expert depth with the infrastructure security track in [the CCIE Security study hub](https://www.pinglabz.com/ccie-security/). ## Frequently Asked Questions ### What does VLAN stand for? VLAN stands for Virtual Local Area Network. It is a logical broadcast domain that can span multiple physical switches. ### How many VLANs can I have? The 802.1Q tag uses 12 bits for the VLAN ID, giving a theoretical range of 0-4095\. VLANs 0 and 4095 are reserved by the standard. Cisco extends-range VLANs (1006-4094) require VTP transparent mode or VTP v3\. Most networks use IDs 2-1001 (the "normal range"). ### Is a VLAN the same as a subnet? No, but in a typical design they map one-to-one. A VLAN is a Layer 2 broadcast domain. A subnet is a Layer 3 IP network. You can technically run multiple subnets on the same VLAN (secondary IPs on an SVI) or one subnet across multiple VLANs (with bridging), but neither is recommended. ### What is the difference between an access port and a trunk port? An access port carries traffic for one VLAN, untagged. End hosts (PCs, printers, IP phones) connect to access ports. A trunk port carries traffic for multiple VLANs, with each frame tagged with its VLAN ID using 802.1Q (except the native VLAN). Switch-to-switch and switch-to-router links are trunks. ### What is a native VLAN? The native VLAN is the one VLAN on a trunk port whose frames are sent untagged. It exists for backwards compatibility. By default it is VLAN 1, but you should change it to a dedicated unused VLAN for security; mismatched native VLANs across a trunk also cause CDP/STP errors and possible double-tagging vulnerabilities. ### Can VLANs be attacked? Yes. Two main attacks: switch spoofing (a malicious host negotiates DTP and becomes a trunk, gaining access to all VLANs) and double tagging (an attacker injects a frame with two 802.1Q tags so the first switch strips one and forwards to the inner-tagged VLAN). Both are mitigated by disabling DTP, moving the native VLAN, and pruning trunk allowed lists. See [VLAN Security Hardening](https://www.pinglabz.com/vlan-security-best-practices/). ## Key Takeaways If you take one thing away from this guide, make it this: VLANs are how you scale Ethernet, but they only buy you isolation if you configure them defensively. Disable DTP. Move the native VLAN. Prune trunks. Use BPDU Guard and PortFast on host ports. Pick one inter-VLAN routing approach and apply it consistently. Bookmark this page, work through the cluster articles in order, and run every configuration in a lab. By the time you finish, you will be ready for any VLAN question a CCNA/CCNP exam or a 3 AM ticket can throw at you. **Studying for the CCNA?** Test your VLAN and Layer 2 switching knowledge on [PingLabz CCNA Flashcards](https://www.pinglabz.com/ccna-flashcards/) \- 200 free multiple-choice questions by topic, mixed, or a full mock exam, each with a plain-English explanation. ### References - [IEEE 802.1Q - Bridges and Bridged Networks](https://standards.ieee.org/ieee/802.1Q/10323/?ref=pinglabz.com) - [Cisco VLAN and VTP technology documentation](https://www.cisco.com/c/en/us/tech/lan-switching/virtual-lans-vlan-trunking-protocol-vlans-vtp/index.html?ref=pinglabz.com) ### Cisco Wireless Complete Guide: Catalyst 9800 Fundamentals, Configuration & Troubleshooting URL: https://www.pinglabz.com/wireless/ Last updated: 2026-07-04T22:53:18.000Z The Cisco Catalyst 9800 is Cisco's IOS XE-based wireless LAN controller - the platform that replaced AireOS at the center of enterprise Wi-Fi. It terminates CAPWAP tunnels from access points, applies configuration through a tags-and-profiles model, and scales from a 2-AP branch to a 6,000-AP campus. If you are migrating off AireOS, standing up your first C9800-CL in vCenter, or chasing a client that will not roam cleanly between APs, this is the cluster overview. This is the cluster overview for the full PingLabz Cisco wireless series: 41 articles covering platform architecture, the AP join process, configuration of WLANs and security, FlexConnect, RRM, mobility, troubleshooting, and SD-Access wireless integration. We will work through what makes the C9800 different, the configuration model that catches AireOS migrants off-guard, the roaming and RF concepts that define real-world performance, and the troubleshooting commands you will use most. New labs and guides, in your inbox Every new PingLabz lab and deep-dive, built and verified on real Cisco IOS XE - free, straight to your inbox. [Join free](https://www.pinglabz.com/signup/) ## Why Cisco Replaced AireOS The Cisco AireOS controllers (the 5500 / 8500 / vWLC line) carried wireless from 802.11n through 802.11ac. They were stable, well-understood, and had a configuration model built around fifteen years of accumulated wireless features. The Catalyst 9800 line is built on Cisco IOS XE. That single architectural choice drives most of what is different about the C9800: the same software platform as Catalyst 9000 switches, the same model-driven programmability stack (NETCONF/YANG, gRPC telemetry, RESTCONF), the same upgrade and HA story (ISSU, SSO), and the same configuration grammar. Operationally, a C9800 looks more like a switch than a legacy WLC. The reasons Cisco prioritized this transition: - **Programmability.** AireOS was hard to automate. IOS XE has a first-class NETCONF/YANG interface and streaming telemetry. - **Wi-Fi 6 / 6E / 7 features** ship on the C9800; many do not on AireOS at all. - **Deployment flexibility.** Physical (9800-40, 9800-80), virtual (9800-CL on ESXi/KVM/Hyper-V), cloud (AWS/Azure), and embedded on Catalyst 9000 switches all run the same image. - **SSO with sub-second failover** instead of AireOS's HA-SSO that still left clients to reauth. The platform overview lives in [Cisco Catalyst 9800 Series Wireless Controllers: The Complete Guide](https://www.pinglabz.com/cisco-catalyst-9800/) and [Platform Overview and Models Compared](https://www.pinglabz.com/cisco-catalyst-9800-overview/). The AireOS-vs-C9800 differences are in [C9800 vs AireOS: What Changed and Why It Matters](https://www.pinglabz.com/c9800-vs-aireos/). ## The Configuration Model: Tags, Profiles, Policies The single biggest stumbling block for AireOS migrants is the C9800 configuration model. AireOS used a flat WLAN-to-AP-group mapping. The C9800 uses three layers of tags that map APs to per-AP behavior: Policy Profile Carries VLAN, ACL, QoS, AAA, session timers Maps Network behavior of an SSID WLAN Profile Carries SSID, security (WPA2/WPA3), 802.1X parameters Maps The over-the-air broadcast Policy Tag Carries Pairs of (WLAN, Policy Profile) Maps Which SSIDs are broadcast and how RF Profile Carries 2.4 / 5 / 6 GHz radio settings, channel/power, DCA Maps Per-band radio behavior RF Tag Carries References RF Profiles for each band MapsPer-AP RF behavior Site Tag Carries AP Join Profile, Local-vs-Flex switching, country code Maps Per-site AP-to-controller binding AP Carries (referenced by all three tags above) Mapsn/a Every AP gets a Policy Tag, an RF Tag, and a Site Tag. Most environments build a small set (say, three or four of each) and assign them via AP location. [The C9800 Configuration Model: Tags, Profiles, and Policies Explained](https://www.pinglabz.com/c9800-configuration-model/) walks through the design pattern with examples. ## The AP Join Process An AP joining the controller follows a deterministic sequence. When something is wrong, the symptom is "AP not joining" and the debug walks through these stages: 1. **Discovery.** AP finds candidate WLCs via DHCP option 43, DNS (CISCO-CAPWAP-CONTROLLER.), broadcast, or static config. 2. **Selection.** AP picks the best WLC from the candidates. 3. **DTLS handshake.** AP and WLC mutually authenticate via certificates (this is where untrusted certs cause failures). 4. **Join Request / Response.** AP requests join; WLC accepts. 5. **Configuration.** AP downloads its image (if needed) and per-AP config from the WLC. 6. **Run state.** AP starts CAPWAP-tunneling client traffic. The full flow with packet captures is in [C9800 AP Join Process: Step-by-Step Explained](https://www.pinglabz.com/c9800-ap-join-process/), and the protocol-level transport detail in [CAPWAP Explained: How the C9800 Controls Access Points](https://www.pinglabz.com/capwap-explained/). ## Deployment Models 9800-40 / 9800-80 Form factorPhysical appliance Use for Large enterprise on-prem; up to 6000 APs 9800-L Form factor Smaller physical appliance Use for Mid-market; up to 250 APs 9800-CL Form factor Virtual (ESXi/KVM/Hyper-V/AWS/Azure) Use for Most modern deployments; private cloud Embedded on Cat 9k Form factor Software on a Catalyst 9000 switch Use for Branch / small site without dedicated WLC hardware The 9800-CL has become the dominant choice for new deployments because it virtualizes cleanly, scales by VM size, and matches the rest of an organization's compute lifecycle. [C9800 Deployment Models](https://www.pinglabz.com/c9800-deployment-models/) covers the trade-offs, and [C9800 Licensing](https://www.pinglabz.com/c9800-licensing/) covers Smart Licensing, DNA Advantage, and Network Advantage. ## Minimum Viable C9800 Configuration The initial setup wizard takes you through hostname, management IP, country, NTP, and credentials. After that, the smallest useful configuration is one WLAN with WPA3 security tied to one Policy Profile and one Policy Tag: ``` WLC(config)# wlan CORP 1 CORP WLC(config-wlan)# security wpa wpa3 WLC(config-wlan)# security wpa akm sae WLC(config-wlan)# no shutdown WLC(config-wlan)# exit WLC(config)# wireless profile policy CORP-POLICY WLC(config-wireless-policy)# vlan 100 WLC(config-wireless-policy)# no shutdown WLC(config-wireless-policy)# exit WLC(config)# wireless tag policy CORP-PT WLC(config-policy-tag)# wlan CORP policy CORP-POLICY WLC(config-policy-tag)# exit WLC(config)# ap 1234.5678.90ab WLC(config-ap-tag)# policy-tag CORP-PT WLC(config-ap-tag)# site-tag default-site-tag WLC(config-ap-tag)# rf-tag default-rf-tag ``` Verification: ``` WLC# show wireless tag policy summary WLC# show wireless tag policy detailed CORP-PT WLC# show ap summary WLC# show wireless client summary ``` End-to-end walkthrough in [C9800 Initial Setup: Step-by-Step Configuration Guide](https://www.pinglabz.com/c9800-initial-setup/) and [How to Configure WLANs on the Cisco Catalyst 9800](https://www.pinglabz.com/c9800-wlan-configuration/). ## Wireless Security in 2026: WPA3, iPSK, Enhanced Open WPA3 is the modern baseline. Three flavors matter: - **WPA3-Personal (SAE).** Replaces WPA2-PSK. Resistant to offline dictionary attacks via Simultaneous Authentication of Equals. - **WPA3-Enterprise.** 192-bit security mode for high-security environments; still uses 802.1X / RADIUS underneath. - **Enhanced Open (OWE).** Replaces open SSIDs (think guest, public). Encrypts the air without authentication. For corporate WLANs you almost always want WPA3-Enterprise with 802.1X (cross-link to the [802.1X pillar](https://www.pinglabz.com/802-1x/)). Configuration in [How to Configure 802.1X and WPA3 Enterprise on the Cisco C9800](https://www.pinglabz.com/c9800-8021x-wpa3-configuration/). Detail on the security primitives in [C9800 Wireless Security Deep Dive](https://www.pinglabz.com/c9800-wireless-security/). iPSK (identity-PSK) deserves a callout: it lets you have one SSID with one PSK per group of devices, which is how IoT / printer / camera fleets are typically onboarded. WPA3 transition mode allows mixed WPA2/WPA3 clients during migration. [C9800 Web Authentication and Captive Portal](https://www.pinglabz.com/c9800-web-authentication/) covers the guest pattern. ## FlexConnect: Local Switching at the Branch By default, an AP tunnels every client frame back to the WLC for switching ("local mode" or "central switching"). For a branch with a slow WAN, that is wasteful: a printer on the branch LAN that wants to talk to a PC on the same branch LAN should not need to round-trip to the controller. FlexConnect mode lets the AP switch traffic locally and only send control-plane CAPWAP messages back to the WLC. WAN failures still leave the AP usable. The trade-off: every AP needs its own VLAN configuration, FlexConnect ACLs are per-AP, and some features (like SSO state-table propagation) work differently. The decision walkthrough is in [C9800 FlexConnect vs Local Mode: How to Choose](https://www.pinglabz.com/c9800-flexconnect-vs-local-mode/), and configuration in [C9800 FlexConnect Configuration: Deployment, Switching, and ACLs](https://www.pinglabz.com/c9800-flexconnect-configuration/). At branches where the WAN is an [SD-WAN overlay](https://www.pinglabz.com/sd-wan/), FlexConnect local switching keeps user traffic out of the tunnels and off the central controller. ## High Availability: SSO and N+1 Two HA models on the C9800: SSO (Stateful Switchover) How it works Active + Standby pair sharing state via redundancy port Failover time Sub-second; clients do not reauth Use for Single site / metro pair N+1 How it works One backup WLC for multiple primary WLCs; APs failover via HA SKU Failover time Tens of seconds; clients reconnect Use for Multi-site or geo-redundant SSO is the modern default within a site. N+1 covers the cross-site scenario. [C9800 High Availability (SSO) Configuration Guide](https://www.pinglabz.com/c9800-sso-configuration/) and [C9800 N+1 Redundancy Configuration](https://www.pinglabz.com/c9800-n-plus-1-redundancy/) cover both. ## RRM and RF Design Wireless performance is rarely a configuration problem; it is usually an RF problem. The C9800's Radio Resource Management (RRM) suite handles the dynamic part: DCA picks channels, TPC tunes power, CHDM responds to coverage holes, and FRA manages the 5-GHz / 6-GHz radio split on tri-radio APs. None of that compensates for a bad site survey. The RRM concepts and tuning patterns are in [C9800 Radio Resource Management (RRM) Deep Dive](https://www.pinglabz.com/c9800-rrm/), and the design fundamentals (channel planning, AP placement, capacity) in [C9800 RF Design: Site Survey, Channel Planning, and Power Tuning](https://www.pinglabz.com/c9800-rf-design/). ## Fast Roaming: 802.11r, OKC, PMKID Caching A client roaming between APs has to reauthenticate, exchange new keys, and reconfigure its DHCP state. Done naively that is a 1-3 second blackout, which is fatal for voice and video. Three mechanisms accelerate it: - **802.11r Fast Transition (FT).** Pre-authenticates the client to neighbor APs over the wire so the new session key exchange happens in advance. - **OKC (Opportunistic Key Caching).** Cisco-proprietary; pre-shares the PMK with neighbor APs so the four-way handshake is shorter. - **PMKID Caching.** The client remembers a previous AP's PMK and presents it on rejoin to skip full 802.1X. For voice clients you turn 802.11r on; for legacy mixed-vintage clients you sometimes leave it off because some old supplicants do not handle FT well. Detail in [C9800 Client Roaming Deep Dive](https://www.pinglabz.com/c9800-client-roaming/), and inter-WLC roaming via mobility groups in [C9800 Mobility and Inter-Controller Roaming Explained](https://www.pinglabz.com/c9800-mobility/). Roaming performance matters most for voice, and it only pays off end to end - pair it with the marking and queueing from the [QoS complete guide](https://www.pinglabz.com/qos/). ## Wi-Fi 6, 6E, and the 6 GHz Band Wi-Fi 6 (802.11ax) brought OFDMA, MU-MIMO uplink, BSS coloring, and TWT. Wi-Fi 6E added the 6 GHz band (1200 MHz of new spectrum in most regions), which gives you clean channels with no DFS surprises and high client density. Wi-Fi 7 (802.11be) is now landing on Cisco APs in mid-2026 with multi-link operation and 320-MHz channels. The C9800 supports all of this; the constraint is usually AP and client capability. Detail in [Wi-Fi 6 and Wi-Fi 6E on the Cisco Catalyst 9800](https://www.pinglabz.com/c9800-wifi6-wifi6e/). ## Rogue Detection, WIPS, and Spectrum The C9800 ships with rogue detection (APs your APs see but were not deployed by you), WIPS (active rogue containment), client exclusion (lockout for repeated auth failures), and CleanAir (spectrum analysis to detect non-Wi-Fi interference). Rogue detection should be on everywhere; rogue containment is jurisdiction-sensitive (FCC in the US allows it; many other countries do not). [C9800 Rogue AP Detection, WIPS, and Client Exclusion](https://www.pinglabz.com/c9800-rogue-detection-wips/) and [C9800 CleanAir and Spectrum Intelligence Explained](https://www.pinglabz.com/c9800-cleanair/). ## Programmability and Telemetry The C9800 supports NETCONF, RESTCONF, and gRPC streaming telemetry on top of the IOS XE foundation. If you are building Ansible playbooks, Terraform providers, or feeding Grafana from the controller, this is a different world from AireOS's SNMP-only lineage. [C9800 NETCONF and RESTCONF: Automation and Programmability](https://www.pinglabz.com/c9800-netconf-restconf/) and [C9800 Model-Driven Telemetry](https://www.pinglabz.com/c9800-model-driven-telemetry/). ## Troubleshooting: The Five Failures You Will See - [C9800 AP Not Joining](https://www.pinglabz.com/c9800-ap-join-troubleshooting/) - [C9800 Client Connectivity Troubleshooting](https://www.pinglabz.com/c9800-client-troubleshooting/) - [C9800 RADIUS Authentication Failures](https://www.pinglabz.com/c9800-radius-troubleshooting/) - [C9800 Roaming Issues](https://www.pinglabz.com/c9800-roaming-troubleshooting/) - [C9800 High Availability Troubleshooting](https://www.pinglabz.com/c9800-ha-troubleshooting/) - [C9800 RF Troubleshooting](https://www.pinglabz.com/c9800-rf-troubleshooting/) Universal first commands: ``` WLC# show ap summary WLC# show wireless client summary WLC# show wireless mobility summary WLC# show ap config general ``` Reference of every show/debug in [C9800 Show Commands](https://www.pinglabz.com/c9800-show-commands/) and [C9800 Debug Commands Reference](https://www.pinglabz.com/c9800-debug-commands/). ## The Full Wireless Cluster, in Reading Order ### Start With the Pillar 1\. [Cisco Catalyst 9800 Series Wireless Controllers: The Complete Guide](https://www.pinglabz.com/cisco-catalyst-9800/) ### Wireless Fundamentals 2\. [Platform Overview and Models Compared](https://www.pinglabz.com/cisco-catalyst-9800-overview/) 3\. [C9800 Hardware and Software Architecture](https://www.pinglabz.com/c9800-architecture/) 4\. [The C9800 Configuration Model](https://www.pinglabz.com/c9800-configuration-model/) 5\. [C9800 vs AireOS](https://www.pinglabz.com/c9800-vs-aireos/) 6\. [C9800 Deployment Models](https://www.pinglabz.com/c9800-deployment-models/) 7\. [CAPWAP Explained](https://www.pinglabz.com/capwap-explained/) 8\. [C9800 AP Join Process](https://www.pinglabz.com/c9800-ap-join-process/) 9\. [C9800 Licensing](https://www.pinglabz.com/c9800-licensing/) ### Configuration Guides 10\. [C9800 Initial Setup](https://www.pinglabz.com/c9800-initial-setup/) 11\. [How to Configure WLANs](https://www.pinglabz.com/c9800-wlan-configuration/) 12\. [C9800 FlexConnect Configuration](https://www.pinglabz.com/c9800-flexconnect-configuration/) 13\. [C9800 RADIUS and AAA Configuration](https://www.pinglabz.com/c9800-radius-aaa-configuration/) 14\. [802.1X and WPA3 Enterprise Configuration](https://www.pinglabz.com/c9800-8021x-wpa3-configuration/) 15\. [C9800 Web Authentication and Captive Portal](https://www.pinglabz.com/c9800-web-authentication/) 16\. [C9800 QoS Configuration](https://www.pinglabz.com/c9800-qos-configuration/) 17\. [C9800 SSO Configuration](https://www.pinglabz.com/c9800-sso-configuration/) 18\. [C9800 N+1 Redundancy](https://www.pinglabz.com/c9800-n-plus-1-redundancy/) 19\. [C9800 Multicast and mDNS Gateway](https://www.pinglabz.com/c9800-multicast-mdns-configuration/) ### Troubleshooting 20\. [C9800 AP Not Joining](https://www.pinglabz.com/c9800-ap-join-troubleshooting/) 21\. [C9800 Client Connectivity Troubleshooting](https://www.pinglabz.com/c9800-client-troubleshooting/) 22\. [C9800 RADIUS Authentication Failures](https://www.pinglabz.com/c9800-radius-troubleshooting/) 23\. [C9800 Roaming Issues](https://www.pinglabz.com/c9800-roaming-troubleshooting/) 24\. [C9800 HA Troubleshooting](https://www.pinglabz.com/c9800-ha-troubleshooting/) 25\. [C9800 RF Troubleshooting](https://www.pinglabz.com/c9800-rf-troubleshooting/) 26\. [C9800 Debug Commands Reference](https://www.pinglabz.com/c9800-debug-commands/) 27\. [C9800 Show Commands](https://www.pinglabz.com/c9800-show-commands/) ### Deep Dives 28\. [C9800 Client Roaming Deep Dive](https://www.pinglabz.com/c9800-client-roaming/) 29\. [C9800 Mobility and Inter-Controller Roaming](https://www.pinglabz.com/c9800-mobility/) 30\. [C9800 Radio Resource Management (RRM) Deep Dive](https://www.pinglabz.com/c9800-rrm/) 31\. [Wi-Fi 6 and Wi-Fi 6E](https://www.pinglabz.com/c9800-wifi6-wifi6e/) 32\. [C9800 Wireless Security Deep Dive](https://www.pinglabz.com/c9800-wireless-security/) 33\. [C9800 Rogue AP Detection, WIPS](https://www.pinglabz.com/c9800-rogue-detection-wips/) 34\. [C9800 CleanAir and Spectrum Intelligence](https://www.pinglabz.com/c9800-cleanair/) 35\. [C9800 Model-Driven Telemetry](https://www.pinglabz.com/c9800-model-driven-telemetry/) 36\. [C9800 NETCONF and RESTCONF](https://www.pinglabz.com/c9800-netconf-restconf/) ### Design and Best Practices 37\. [Cisco C9800 Design Best Practices](https://www.pinglabz.com/c9800-design-best-practices/) 38\. [C9800 RF Design](https://www.pinglabz.com/c9800-rf-design/) 39\. [C9800 FlexConnect vs Local Mode](https://www.pinglabz.com/c9800-flexconnect-vs-local-mode/) 40\. [C9800 Fabric Mode and SD-Access Wireless](https://www.pinglabz.com/c9800-fabric-sda-wireless/) 41\. [Backing Up, Restoring, and Upgrading the C9800](https://www.pinglabz.com/c9800-backup-upgrade/) Wireless in the CCNA Labs library Wireless architecture and WPA2 vs WPA3 are covered in two concept labs in the CCNA Labs library. Full hands-on wireless (WLC + APs + clients) requires CML Personal - the labs explain why and provide the configuration templates. Open the PingLabz CCNA Labs library. [Open the wireless labs](https://www.pinglabz.com/ccna-labs-network-access/) ### More wireless guides in this cluster 1\. [Catalyst 9800-CL Day 0 Configuration, Line by Line (9800 Series Part 2)](https://www.pinglabz.com/catalyst-9800-cl-day-0-configuration/) 2\. [Build a Catalyst 9800-CL Wireless Lab in Cisco Modeling Labs (9800 Series Part 1)](https://www.pinglabz.com/catalyst-9800-cl-lab-cml/) 3\. [Catalyst 9800 Wireless Series: Lab Files & Downloads](https://www.pinglabz.com/catalyst-9800-lab-files/) 4\. [Cisco 9800-CL Wireless Controller Setup Made Easy](https://www.pinglabz.com/cisco-9800cl-initial-setup-cli-guide/) 5\. [Introduction to Cisco Catalyst 9800 Wireless Controllers](https://www.pinglabz.com/cisco-catalyst-9800-wireless-controller-introduction/) 6\. [Cisco Catalyst 9800 Wireless Controller Models Explained](https://www.pinglabz.com/cisco-catalyst-9800-wireless-controller-models/) 7\. [Cisco Wireless LAN Controller: What It Does and How It Works](https://www.pinglabz.com/cisco-wireless-lan-controller/) 8\. [Wi-Fi 7 for Cisco Shops: Upgrade Now or Wait?](https://www.pinglabz.com/wi-fi-7-for-cisco-shops-upgrade-now-or-wait/) 9\. [Cisco Wireless Access Point Modes Explained](https://www.pinglabz.com/wireless-access-point-modes/) ## Frequently Asked Questions ### What is the difference between the 9800-40, 9800-80, and 9800-L? Capacity. The 9800-80 is the largest (up to 6,000 APs / 64,000 clients), 9800-40 mid-range (up to 2,000 APs), 9800-L is the entry physical model (up to 250 APs). The 9800-CL is a virtualized controller with size scaling by VM resources. ### Should I deploy 9800-CL or a physical 9800? 9800-CL for most modern deployments. It runs in your existing virtualization environment, scales by VM size, has the same feature parity as the physical models, and avoids hardware lifecycle management. Physical appliances make sense when you need predictable hardware throughput, when virtualization is not available, or for very large 4,000+ AP sites where appliance hardware is more cost-effective. ### Do I have to migrate from AireOS? Yes, eventually. Cisco has set end-of-life and end-of-support dates for the AireOS platforms; new wireless features (Wi-Fi 6E, Wi-Fi 7, modern security primitives) are landing only on the C9800\. Plan migration projects with care: the configuration model is genuinely different, and a lift-and-shift will not work. ### How does 802.1X work on wireless? The same way it works on wired (cross-link to [802.1X pillar](https://www.pinglabz.com/802-1x/)): supplicant on the client, authenticator is the WLC + AP, RADIUS server (Cisco ISE in most enterprises). The over-the-air transport is EAPOL inside the 802.11 association. All the EAP method choices (PEAP, EAP-TLS, EAP-FAST) are the same. ### Should I use WPA3-only or WPA3 transition mode? WPA3-only when every client supports WPA3 (modern Windows, macOS, iOS, Android, recent Linux). WPA3 transition mode (also called WPA2/WPA3 mixed) when you have older clients that do not support WPA3\. Transition mode is operationally simpler during migration but slightly less secure than WPA3-only. ### Should I always enable 802.11r? Yes for voice and real-time application clients. No for legacy mixed-vintage fleets that include very old supplicants which mishandle FT. The compromise pattern: have a separate voice SSID with 802.11r on, and a data SSID with 802.11r off (or in mixed-mode if the controller supports it). ## Key Takeaways If you take one thing away from this guide, make it this: the C9800 is an IOS XE device that happens to terminate APs. Embrace that. The configuration model (tags, profiles, policies) is more verbose than AireOS but expresses real production needs that AireOS hacked around. The programmability story is genuinely better. The HA and RF tools are first-class. Bookmark this page, work through the cluster articles in order, and run every configuration on a 9800-CL VM before you touch production. By the time you finish you will be ready to design and operate enterprise wireless on the modern Cisco platform. **Studying for the CCNA?** Test your wireless knowledge on [PingLabz CCNA Flashcards](https://www.pinglabz.com/ccna-flashcards/) \- 200 free multiple-choice questions by topic, mixed, or a full mock exam, each with a plain-English explanation. ### References - [IEEE 802.11 Working Group (Wi-Fi standards)](https://www.ieee802.org/11/?ref=pinglabz.com) - [Cisco Catalyst 9800 Series documentation](https://www.cisco.com/c/en/us/support/wireless/catalyst-9800-series-wireless-controllers/series.html?ref=pinglabz.com) ### SD-WAN: The Complete Guide for Network Engineers URL: https://www.pinglabz.com/sd-wan/ Last updated: 2026-07-12T10:04:07.000Z SD-WAN (Software-Defined Wide Area Network) is the platform that replaced the dedicated-MPLS-circuit model of enterprise WAN. It uses commodity internet links plus encrypted overlay tunnels, a centralized controller, and per-application path policy to deliver a WAN that is cheaper, more flexible, and operationally aware of what your applications actually need. By 2026 it is the dominant enterprise WAN architecture, and "we run SD-WAN" is the default answer in every infrastructure design conversation. This is the cluster overview for the full PingLabz SD-WAN series: architecture, the Cisco Catalyst SD-WAN (Viptela) platform, vendor comparisons, security and the SASE convergence, deployment models, and the operational realities of running SD-WAN past day 200\. We will work through what SD-WAN actually is, the architecture every implementation shares, the four-component Cisco model that dominates enterprise deployments, and where SD-WAN starts to shade into SASE. If you are evaluating SD-WAN, designing a migration, or operating a deployed platform, start here. New labs and guides, in your inbox Every new PingLabz lab and deep-dive, built and verified on real Cisco IOS XE - free, straight to your inbox. [Join free](https://www.pinglabz.com/signup/) ## What SD-WAN Solves The traditional enterprise WAN model was built on dedicated MPLS circuits between branch sites and a central data center. Three things broke that model in the 2010s: - **SaaS and the public cloud.** Your applications stopped living in the data center. Backhauling Office 365, Salesforce, AWS, and Zoom traffic across the corporate WAN, through the central security stack, and back out the internet edge added 30-100ms of latency to every user transaction for no business reason. - **Cost asymmetry.** A 100 Mbps MPLS circuit cost roughly the same as a 1 Gbps internet link in most metros. The price-per-Mbps gap kept widening. - **Operational rigidity.** Adding a branch took weeks (carrier provisioning) and policy changes meant tickets, config drift, and the occasional 3 AM incident. The control plane was distributed across hundreds of manually-configured routers. SD-WAN's pitch was simple: replace expensive MPLS with cheaper broadband, steer traffic intelligently per-application, and centralize the control plane so policy changes propagate automatically. The first wave of deployments delivered measurable cost savings (30-50 percent reductions were typical) and the path-steering story genuinely solved problems that static routing could not. By 2026 the conversation has shifted. Path steering is table stakes across every vendor. The competitive question now is operational depth: how the platform handles template management at scale, how it surfaces policy drift, how cleanly it integrates with cloud security, and whether the observability story holds up when a user says "Salesforce is slow" and the controller shows everything green. See [Cisco SD-WAN in 2026: The Real Value Is Day-2 Operations](https://www.pinglabz.com/cisco-sd-wan-in-2026-the-real-value-is-day-2-operations/) for the full argument on day-2 operations being the durable differentiator. ## How SD-WAN Works (the 10,000-Foot View) Every SD-WAN platform shares the same architectural pattern. The differences between Cisco, Fortinet, VMware, Versa, Palo Alto, and Aruba are about the implementation; the model is consistent. 1. **Decoupled control and data planes.** A central controller computes policy and pushes it to the edge devices. The edge devices forward packets according to the policy without needing to consult the controller per-packet (which would not scale). 2. **Encrypted overlay tunnels.** Edge devices form IPsec tunnels to each other (or to a cloud gateway) over whatever transport happens to be available - MPLS, broadband internet, LTE/5G, satellite. The overlay is the WAN; the underlay is just transport. 3. **Multiple transports per site.** Branches typically have two or more transports (e.g. one MPLS plus one broadband, or two broadband plus LTE for redundancy). The SD-WAN edge device has tunnels over each, and chooses which to use per-application based on policy. 4. **Application-aware steering.** The edge identifies application traffic (DPI, DNS, IP/port heuristics, sometimes pre-built application signatures) and applies policy: "voice goes over the lowest-jitter path", "Salesforce goes direct to the internet from this branch", "everything else goes back to HQ". 5. **Centralized orchestration.** Policy, templates, certificates, and software upgrades are managed from a central pane (vManage, FortiManager, VeloCloud Orchestrator, Versa Director, etc.) and pushed to all edges. The result is a WAN that adapts to transport conditions in seconds, treats different applications differently without per-router policy, and can be onboarded at a new branch with zero-touch provisioning rather than days of configuration. Application-aware routing chooses the path, but it does not schedule the queue - you still need real [QoS marking and queueing](https://www.pinglabz.com/qos/) on the underlay interfaces. ## SD-WAN vs MPLS The most common framing question for new evaluators. Both can carry your enterprise traffic; the trade-offs are real. Cost per Mbps SD-WAN Low (commodity broadband) MPLS High (dedicated carrier circuit) Bandwidth available SD-WAN1 Gbps+ at low cost MPLS Typically 50-200 Mbps per branch QoS guarantee SD-WAN Best-effort over public internet (mitigated by multi-link, FEC, jitter buffers) MPLS Per-class hard QoS (carrier honors DSCP) Latency consistency SD-WAN Variable (internet path) MPLS Predictable (dedicated path) Provisioning time SD-WAN Days (broadband install) MPLS Weeks to months (carrier provisioning) Cloud / SaaS reach SD-WAN Direct internet break-out from branch MPLS Backhaul to HQ then out (or expensive cloud-direct) Policy granularity SD-WAN Per-application, dynamic MPLSPer-class, static Vendor lock SD-WAN Single SD-WAN vendor + multiple transport vendors MPLS Single carrier per region Encryption SD-WANIPsec by default MPLS Optional, customer-deployed The honest answer in 2026: most enterprises run hybrid. Critical real-time traffic (voice, certain regulated workflows) keeps an MPLS leg with hard QoS guarantees. Everything else runs over broadband-plus-broadband (sometimes broadband-plus-LTE) with SD-WAN steering. The cost story works because the marginal cost of dropping from MPLS-everywhere to one MPLS link plus broadband redundancy is huge and the user-experience improvement on cloud apps is significant. ## SD-WAN Architecture: Three Planes Every SD-WAN deployment has three logical planes. Understanding which component lives in which plane is what unlocks the operational picture: Management Job Configuration, monitoring, troubleshooting UI Cisco component vManage / SD-WAN Manager Fortinet component FortiManager + FortiAnalyzer VMware componentVeloCloud Orchestrator Control Job Compute and distribute routing/policy state Cisco componentvSmart controllers Fortinet componentFortiGate + iBGP VMware component VeloCloud Gateway / SD-WAN Hub Data Job Forward customer packets through the overlay Cisco component WAN Edge (cEdge / vEdge) Fortinet componentFortiGate edge VMware componentVeloCloud Edge Orchestration (auth) Job Bring up new edges, distribute certificates Cisco componentvBond orchestrator Fortinet componentFortiManager fabric VMware componentVeloCloud Orchestrator The Cisco model deliberately separates these into four distinct components because each has different scale and HA requirements. vManage is heavy (database, UI, REST API) and runs as a single virtual instance per region with cluster HA. vSmart controllers are lightweight and you typically run two or three per region for redundancy. WAN Edges are at every branch and number in the thousands. vBond is small and stateless. Other vendors collapse the planes more aggressively. Fortinet runs everything inside FortiGate appliances with FortiManager as the orchestrator. The control-plane work happens via iBGP between fabric members. This is operationally simpler if you are already a Fortinet shop and more complex if you are not. ## Cisco Catalyst SD-WAN (formerly Viptela) Cisco bought Viptela in 2017 and rebranded the platform several times: Viptela SD-WAN, then Cisco SD-WAN, now Catalyst SD-WAN. The underlying architecture is the same throughout. It is the dominant enterprise SD-WAN platform in 2026, especially in shops that already run Cisco for switching and routing. The four-component model: vManage (now SD-WAN Manager) Form factor VM (3 or 6 nodes for HA) Job Management UI, REST API, template engine, monitoring backend vSmart controllers Form factor VM (2-3 per region for HA) Job OMP routing protocol; distributes routes and policy to WAN Edges vBond orchestrator Form factor VM (typically internet-facing, 2 for HA) Job Authenticates new WAN Edges joining the fabric; relays initial certs WAN Edge (cEdge or vEdge) Form factor Hardware appliance or VM at every site Job Data plane; forms IPsec tunnels; enforces policy The OMP (Overlay Management Protocol) is the BGP-derived control protocol vSmarts use to distribute routes and policy. If you understand BGP path attributes and route maps, OMP makes immediate sense; the same mental model applies. See the [BGP cluster pillar](https://www.pinglabz.com/bgp/) for the foundational BGP concepts that OMP builds on. cEdge is the IOS-XE-based WAN Edge (Catalyst 8000, ISR 4000). vEdge is the older Viptela-OS-based hardware. Cisco has been migrating customers off vEdge and onto cEdge for several years; cEdge is the future and most new deployments default to it. For the deep dive on Cisco SD-WAN architecture, components, and the OMP control plane, see (article forthcoming in this cluster). ## Vendor Landscape The SD-WAN market consolidated around six vendors in 2026, plus a long tail of smaller players. Quick orientation: Cisco Platform Catalyst SD-WAN (Viptela), Meraki MX Strength Largest installed base; deep Cisco integration; Catalyst supports complex policy Trade-off Two distinct platforms (Catalyst and Meraki); operational complexity at scale VMware Platform VeloCloud (now part of Broadcom) Strength Strong cloud and SaaS optimization; mature multi-tenant model Trade-off Broadcom acquisition has unsettled some customers; pricing changes Fortinet Platform Secure SD-WAN (FortiGate-based) Strength Native security integration; cost-effective for Fortinet shops Trade-off Operational model collapses control and data; less visibility for pure SD-WAN ops Versa Networks PlatformVersa SASE / Concerto Strength Single platform for SD-WAN and full SASE stack Trade-off Smaller installed base; learning curve for Cisco-trained operators Palo Alto Platform Prisma SD-WAN (CloudGenix) Strength App-defined policies; cloud-native control plane Trade-off Newer entrant; ecosystem still maturing HPE Aruba Platform EdgeConnect (Silver Peak) Strength Strong path conditioning (FEC, packet ordering); excellent for poor links Trade-off Smaller share; less integrated with broader Aruba switching at the operational layer For most enterprises the vendor decision is heavily shaped by what is already in the network. Cisco shops pick Catalyst SD-WAN. Fortinet shops pick FortiGate Secure SD-WAN. Mixed-vendor shops evaluate Versa or Palo Alto on greenfield deployments. Few customers run two SD-WAN vendors simultaneously. ## Security and the SASE Convergence SD-WAN by itself only solves the WAN routing problem. It does not solve the security problem of what to do with the internet-direct traffic now leaving every branch. That gap is what SASE (Secure Access Service Edge, Gartner's term, pronounced "sassy") fills. The SASE model bundles SD-WAN with cloud-delivered security services: SWG (Secure Web Gateway), CASB (Cloud Access Security Broker), ZTNA (Zero Trust Network Access), DLP (Data Loss Prevention), and FWaaS (Firewall-as-a-Service). The branch SD-WAN edge sends internet-bound traffic to the SASE provider's nearest cloud PoP for inspection rather than backhauling to HQ. By 2026 the SD-WAN and SASE markets are effectively merging: - **Cisco** integrates Catalyst SD-WAN with Cisco Umbrella (cloud SWG/DNS) and Cisco Secure Access (ZTNA). - **VMware** ties VeloCloud SD-WAN to Symantec Cloud SWG and Lookout CASB through Broadcom integrations. - **Palo Alto** sells Prisma SASE as a unified platform (Prisma SD-WAN + Prisma Access). - **Versa**, **Fortinet**, and **Cato Networks** all sell single-vendor SASE. Practical implication: when you evaluate SD-WAN in 2026, you are also evaluating SASE. The choices are tightly coupled. A best-of-breed approach (one vendor for SD-WAN, another for cloud security) is operationally heavier than picking a single SASE platform but gives you better individual products. For PingLabz coverage of how 802.1X / NAC and SD-WAN interact for branch micro-segmentation, see the [802.1X cluster](https://www.pinglabz.com/802-1x/). ## Deployment Models The four common branch deployment patterns: 1. **Hybrid (MPLS + broadband).** One MPLS leg for hard-QoS traffic (voice, regulated apps), one or two broadband legs for everything else. Most common pattern in established enterprises migrating from MPLS-only. Used to be the recommendation; now usually the migration end-state. 2. **Internet-only (dual broadband + LTE).** No MPLS at all. Two broadband links from different ISPs plus LTE for tertiary backup. Cost-optimal; the SD-WAN's path conditioning handles the variability. Common for new branches and SaaS-heavy organizations. 3. **Cloud-on-ramp.** The branch SD-WAN edge connects directly to the cloud provider's network (AWS Cloud WAN, Azure Virtual WAN, Google Cloud Network Connectivity Center). Bypasses both the legacy WAN and the public internet for cloud-bound traffic. 4. **Multi-cloud direct.** Multiple cloud-on-ramps to different providers. The SD-WAN edge holds policy about which cloud serves which application and routes accordingly. Headquarters and data center sites tend to deploy SD-WAN as a redundant pair of edges with active-active or active-standby pairing. Hub-and-spoke topologies persist where compliance or latency requires central inspection; full-mesh is more common for collaborative organizations where branches need to talk to each other. At the branch, the SD-WAN edge usually lands in the same closet as the wireless stack - the [Cisco wireless complete guide](https://www.pinglabz.com/wireless/) covers that side. ## When SD-WAN Makes Sense (and When It Doesn't) SD-WAN is the right answer when: - You have multiple branch sites with significant cloud or SaaS usage. - Your MPLS contract is up for renewal and the price per Mbps is no longer competitive. - You need application-aware policy that static routing cannot express. - You want zero-touch provisioning for new branches. - You are evaluating SASE or already buying cloud security and want unified management. SD-WAN is overkill when: - You have one or two sites with simple connectivity needs (a static dual-WAN router is cheaper). - Your applications are entirely on-premises and well-served by MPLS. - Compliance requires every flow to traverse a specific carrier-managed path. - You do not have the operational maturity to run a centralized control plane (the operations cost of SD-WAN done badly is higher than legacy WAN done well). ## SD-WAN and Campus/DC Fabrics SD-WAN virtualizes the WAN; the campus and data centre have their own fabric story built on overlays. [VXLAN](https://www.pinglabz.com/vxlan-deep-dive/) and [BGP EVPN](https://www.pinglabz.com/vxlan-bgp-evpn-explained/) run the data centre, and [Cisco SD-Access](https://www.pinglabz.com/sd-access-architecture/) (LISP + VXLAN + Catalyst Center) runs the campus. All three share the overlay idea SD-WAN applies to the WAN, decoupling the logical network from the physical transport. The full picture is in the [Network Virtualization cluster](https://www.pinglabz.com/network-virtualization/). ## Automating Catalyst SD-WAN SD-WAN Manager (formerly vManage) exposes the whole fabric through a REST API, so you can pull per-tunnel loss, latency, and jitter across every site, read policies, and drive templates programmatically. It is covered from the automation angle in [the SD-WAN Manager API article](https://www.pinglabz.com/sd-wan-manager-api/), part of the [Network Automation cluster](https://www.pinglabz.com/network-automation/). ## SD-WAN Deep Dives in This Cluster The articles in this cluster, in reading order: 1. [SD-WAN vs MPLS: When Each Wins in 2026](https://www.pinglabz.com/sd-wan-vs-mpls/) 2. [SD-WAN Architecture: Control, Data, and Management Planes Explained](https://www.pinglabz.com/sd-wan-architecture/) 3. [Cisco Catalyst SD-WAN Architecture: vManage, vSmart, vBond, WAN Edge](https://www.pinglabz.com/cisco-sd-wan-architecture/) 4. [Cisco vManage / SD-WAN Manager: The Control Plane Walkthrough](https://www.pinglabz.com/cisco-vmanage/) 5. [SD-WAN Deployment Models: Hybrid, Internet-Only, and Cloud On-Ramp](https://www.pinglabz.com/sd-wan-deployment-models/) 6. [SD-WAN Security and the SASE Convergence](https://www.pinglabz.com/sd-wan-security/) 7. [Cisco SD-WAN in 2026: The Real Value Is Day-2 Operations](https://www.pinglabz.com/cisco-sd-wan-in-2026-the-real-value-is-day-2-operations/) This list grows as new articles are published. Check back for vendor-specific deep dives, configuration walkthroughs, and troubleshooting references. Solid on the CCNA foundations? - the PingLabz CCNA Labs library SD-WAN sits on top of CCNA-level routing, QoS, and security. Refresh those foundations in the 60-lab PingLabz CCNA Labs library: OSPF, EIGRP, ACLs, NAT, QoS, IPsec concepts. All with real captures from Cisco IOS XE 17.16\. Open the PingLabz CCNA Labs library. [Open the CCNA labs library](https://www.pinglabz.com/ccna-labs-network-fundamentals/) ### More SD-WAN guides in this cluster 1\. [Cisco Catalyst SD-WAN: The Rebrand, the Architecture, OMP](https://www.pinglabz.com/cisco-catalyst-sd-wan/) 2\. [MPLS L3VPN vs SD-WAN: When to Migrate](https://www.pinglabz.com/mpls-vs-sd-wan-migration/) 3\. [QoS in Modern Networks: SD-WAN, Cloud, and Application-Aware Steering](https://www.pinglabz.com/qos-sd-wan-modern-networks/) 4\. [SD-WAN: What Is It? The 5-Minute Primer](https://www.pinglabz.com/sd-wan-what-is-it/) ## Advanced SD-WAN: OMP, Policy, and Templates (CCIE level) These articles close the Catalyst SD-WAN half of the CCIE Software-Defined Infrastructure domain. They cover provisioning, the OMP control plane, the three policy types, DIA, and TLOC extension - grounded in Cisco's current 20.x CLI. Because a full Manager/Validator/Controller stack with certificate onboarding cannot be driven through a device CLI session, the command output here is presented as documented reference from Cisco's 20.x documentation, clearly labelled - never staged as a home-lab capture. 1 [Catalyst SD-WAN templates: device, feature, and CLI compared](https://www.pinglabz.com/sd-wan-templates-device-feature-cli/) The three provisioning models and which to use on a current version - configuration groups are the answer. 2 [OMP deep dive: routes, TLOCs, service routes, and path selection](https://www.pinglabz.com/omp-deep-dive/) SD-WAN's BGP. Routes point at TLOCs, TLOCs resolve to tunnels - the whole forwarding model, with the full BGP analogy. 3 [SD-WAN centralized control policy: shaping the overlay](https://www.pinglabz.com/sd-wan-centralized-control-policy/) One document at the Controller shapes what every edge learns - path preference, topology, and the direction gotcha. 4 [SD-WAN data policy and application-aware routing (AAR)](https://www.pinglabz.com/sd-wan-data-policy-aar/) Steer each app onto the path that meets its SLA, measured live via BFD - the headline reason SD-WAN replaced traditional WAN. 5 [SD-WAN localized policy: ACLs and route policies at the edge](https://www.pinglabz.com/sd-wan-localized-policy/) The per-device treatment - ACLs, QoS, route filtering - and the complete map of which policy type does what. 6 [Direct Internet Access (DIA) in Catalyst SD-WAN](https://www.pinglabz.com/sd-wan-direct-internet-access-dia/) Break out branch internet locally instead of backhauling - and the branch-edge security question it forces. 7 [TLOC extension: sharing transports between edge routers](https://www.pinglabz.com/sd-wan-tloc-extension/) True dual-router redundancy without duplicating every circuit - let one router use its partner's transport. ## Studying for the CCIE? This cluster is part of the full CCNA to CCNP to CCIE Enterprise ladder on PingLabz, every rung built on real Cisco output. For expert-level depth across every EI v1.1 blueprint domain - and the four integration Super Labs - see the [CCIE Enterprise Infrastructure study hub](https://www.pinglabz.com/ccie-enterprise/). ## Frequently Asked Questions ### What does SD-WAN stand for? SD-WAN stands for Software-Defined Wide Area Network. The "software-defined" part means the control plane (policy, routing, application steering) is decoupled from the data plane (the boxes forwarding packets) and centralized. The "wide area network" part means it connects geographically distributed sites - branches to HQ, branches to each other, branches to cloud. ### What is the difference between SD-WAN and SDN? SDN (Software-Defined Networking) is the broader architectural concept: decouple the control plane from the data plane and centralize control. It applies anywhere - data centers, campuses, WANs. SD-WAN is the specific application of SDN principles to the wide area network. Most modern SD-WAN platforms borrow heavily from SDN architectural patterns but add WAN-specific features (per-application path steering, IPsec overlays, cloud on-ramps). ### How much does SD-WAN actually save vs MPLS? The first wave of SD-WAN deployments (2016-2020) commonly reported 30-50 percent WAN cost reduction by replacing MPLS with broadband. The savings narrative is more nuanced now. Pure cost replacement is real where MPLS contracts have not been renegotiated; total-cost-of-ownership analyses that include the SD-WAN platform license, edge appliances, and operational uplift often show smaller net savings (10-30 percent). The bigger 2026 value is operational - faster branch provisioning, better cloud user experience, and the SASE convergence story. ### Does SD-WAN completely replace MPLS? Not always. Most 2026 enterprise deployments are hybrid: one MPLS leg for hard-QoS-required traffic (voice, certain regulated workloads) plus broadband legs for everything else. Pure internet-only SD-WAN works well for SaaS-heavy organizations and new branches but is less common in established enterprises with legacy real-time applications. The trend is toward fewer MPLS circuits, not zero. ### Is SD-WAN the same as a VPN? No. SD-WAN typically uses IPsec tunnels, which are technically VPN tunnels, but SD-WAN adds the centralized control plane, application-aware steering, and orchestration layer that a plain VPN does not have. Site-to-site IPsec VPN is a primitive SD-WAN can use; SD-WAN as a category is the platform that builds on top of those tunnels. ### SD-WAN vs SASE - what is the difference? SD-WAN is the WAN routing and steering layer. SASE adds cloud-delivered security services (SWG, CASB, ZTNA, DLP, FWaaS) that inspect the traffic SD-WAN is steering. By 2026 most vendors sell unified SD-WAN+SASE platforms; the distinction matters mostly when you are evaluating products from different vendors. SASE is roughly "SD-WAN plus the security stack you'd otherwise run separately." ### Cisco Catalyst SD-WAN vs Meraki MX - which one? Both are Cisco. Catalyst SD-WAN (formerly Viptela) is the enterprise-grade platform: complex policy, advanced traffic engineering, full vManage/vSmart/vBond/WAN Edge architecture, suited to large or complex deployments. Meraki MX is the cloud-managed simpler platform: fewer knobs, faster to deploy, suited to small and mid-sized deployments where ease of use beats policy expressiveness. Cisco typically positions Catalyst SD-WAN for larger enterprises and Meraki for SMB and lean-IT use cases. ## Key Takeaways SD-WAN is now the default enterprise WAN architecture. The case for it has broadened beyond cost replacement to include operational simplicity, cloud-direct user experience, and the SASE convergence story. Whatever vendor you pick, the architecture is consistent: centralized control plane, encrypted overlay tunnels over commodity transports, application-aware steering, and zero-touch branch provisioning. If you are evaluating SD-WAN in 2026, do not let the path-steering pitch dominate the conversation. It is table stakes. Push the conversation toward day-200 operations: template management at scale, policy drift detection, cloud-security integration, and observability when something goes wrong. Bookmark this page, work through the cluster articles in order, and lab every architecture decision in a controlled environment. The platforms are mature; the operational habits required to run them well are still being built. ### References - [Cisco SD-WAN solutions documentation](https://www.cisco.com/c/en/us/solutions/enterprise-networks/sd-wan/index.html?ref=pinglabz.com) ### QoS (Quality of Service): The Complete Guide for Cisco Engineers URL: https://www.pinglabz.com/qos/ Last updated: 2026-07-12T03:19:17.000Z QoS (Quality of Service) is the set of network mechanisms that decide which packets win when there is not enough bandwidth, buffer space, or scheduling priority to accommodate every flow. Without QoS, every packet is best-effort - the network treats your VoIP call the same as a Windows update download. With QoS, the network knows that voice packets matter more than backup traffic and behaves accordingly. This is the cluster overview for the full PingLabz QoS series: classification and marking, the Cisco MQC (Modular QoS CLI), DSCP and IP precedence, queueing/policing/shaping, voice and video QoS, and how QoS works in modern SD-WAN and wireless deployments. We will work through what QoS actually does, the four-step model every Cisco QoS implementation follows, the marking standards you must understand, and the trust-boundary discipline that separates a working QoS deployment from a misconfigured one. If you are studying for CCNP/CCIE, designing a QoS rollout, or troubleshooting why voice quality went bad after the last firmware push, start here. New labs and guides, in your inbox Every new PingLabz lab and deep-dive, built and verified on real Cisco IOS XE - free, straight to your inbox. [Join free](https://www.pinglabz.com/signup/) ## What QoS Solves Every network has finite resources: link bandwidth, switch buffers, scheduler slots, ingress and egress queues. When demand exceeds those resources (transient bursts, sustained congestion, or scheduled events like the 9 AM video-call rush), packets must be dropped or delayed. The question is which packets. Without QoS, the answer is "whichever packet happened to arrive when the queue was full." The network is fair in a literal sense and useless in a practical sense: voice gets dropped equally with backup traffic, and your CFO's video call breaks while a server pulls a 50-GB OS image from a CDN. QoS gives you four levers to bias the outcome: - **Classification.** Identify what kind of traffic each packet belongs to. Is this VoIP? Is this Salesforce? Is this YouTube? - **Marking.** Apply a label (DSCP, CoS, MPLS EXP) so downstream devices can act on the classification without redoing the work. - **Queueing and Scheduling.** Decide the order in which packets leave a congested interface. Voice goes first; bulk backups go last. - **Policing and Shaping.** Decide which packets to drop or delay when the offered load exceeds an agreed rate. Combined, these four mechanisms let you express policy like "voice gets priority queueing with no rate limit; business-critical apps get a guaranteed 30 percent of the link; bulk traffic gets the rest with no guarantees." That policy compiles into per-interface configurations on every router and switch in the path. ## The Four-Step QoS Model Every Cisco QoS deployment follows the same four-step pattern, sometimes called the QoS toolset: Ingress to the QoS domain (typically access port) Step1\. Classify Cisco mechanism class-map matching ACL, NBAR, DSCP, CoS, etc. Same place as classification Step2\. Mark Cisco mechanism policy-map with set dscp / set cos / set mpls experimental Egress on every congested interface Step3\. Queue and schedule Cisco mechanism policy-map with priority/bandwidth/fair-queue Either ingress (policing) or egress (shaping) Step4\. Police or shape Cisco mechanism policy-map with police / shape The discipline: classify and mark once at the network edge, then trust those marks throughout the QoS domain. Re-classifying at every hop is wasteful (CPU, memory) and error-prone. The trust boundary is where QoS-capable infrastructure starts; everything inside trusts the markings, everything outside is suspect. The Cisco implementation of this four-step model uses the MQC (Modular QoS CLI). MQC has three constructs: `class-map` (defines classification), `policy-map` (defines what to do with each class), and `service-policy` (applies a policy to an interface). See (article forthcoming) for the deep dive. ## Marking: DSCP, CoS, IP Precedence, MPLS EXP The whole point of marking is to label a packet at the edge so every downstream device can treat it consistently without re-classifying. Four marking fields exist depending on layer and protocol: IP Precedence Where it lives IPv4 ToS field, top 3 bits Bits3 (8 values) Used at Layer 3; legacy, mostly replaced by DSCP DSCP (Differentiated Services Code Point) Where it lives IPv4 ToS / IPv6 Traffic Class field, top 6 bits Bits6 (64 values) Used at Layer 3; modern standard CoS (802.1p Priority) Where it lives802.1Q tag, PCP field Bits3 (8 values) Used at Layer 2 (only on tagged frames) MPLS EXP (Traffic Class) Where it lives MPLS label, EXP/TC field Bits3 (8 values) Used atMPLS-labeled traffic DSCP is the dominant marking in modern networks. Its 6-bit value space gives you 64 possible classes, but in practice you use a small standardized set: EF (Expedited Forwarding) Decimal46 ClassVoice Typical use RTP voice payload; lowest jitter, lowest loss CS5 Decimal40 ClassBroadcast video Typical use One-way streaming video AF41 Decimal34 ClassVideo conferencing Typical use Two-way real-time video (Zoom, Teams) AF31 Decimal26 ClassMultimedia streaming Typical use One-way audio/video streaming AF21 Decimal18 ClassTransactional apps Typical use Salesforce, ERP, low-latency business apps AF11 Decimal10 ClassBulk data Typical useEmail, file transfers CS1 Decimal8 ClassScavenger Typical use Lower than best-effort; backups, peer-to-peer BE (Best Effort, Default) Decimal0 ClassDefault Typical useEverything else The IETF's RFC 4594 ("Configuration Guidelines for DiffServ Service Classes") defines the 12-class model that most enterprises base their QoS designs on. Cisco's published "QoS Best Practices" maps that to 8-class and 4-class simplifications for smaller deployments. The marking step itself is short. Define a class-map that picks out the traffic, then a policy-map that sets the DSCP value, then attach the policy as `service-policy input` on the access port closest to the source: ``` ! Identify voice traffic at the edge (typically by VLAN ! or DSCP coming from a trusted IP phone) class-map match-any VOICE-EDGE match access-group name VOICE-RTP ! Mark it to DSCP EF so every downstream device ! knows it is voice policy-map MARK-VOICE class VOICE-EDGE set dscp ef class class-default set dscp default ! Apply at ingress to the access port interface Vlan20 service-policy input MARK-VOICE ``` Once the packet is marked, every router and switch from here to the WAN edge acts on the DSCP value without re-classifying. That single edge marking is the entire point of the marking step. For the byte-level explanation of DSCP and IP Precedence, including how the 6 bits map to per-hop behaviors, see (article forthcoming). ## Trust Boundary: The Single Most Important Concept If you take one thing away from QoS, make it this: classify and mark once at a controlled edge, then trust those marks everywhere inside the QoS domain. The trust boundary is the perimeter where QoS-trusted infrastructure starts. Inside, every device respects the marking on incoming packets and applies policy accordingly. Outside (host PCs, IP phones, BYOD devices), markings are suspect and must be re-validated at ingress. Common trust boundary placements: Cisco IP phone behind PC Trust the phone's voice VLAN markings; re-mark or zero everything from the data VLAN Trusted application server Trust DSCP set by the application BYOD / guest device Never trust; mark all traffic to BE or scavenger Inter-switch trunk inside the QoS domain Trust DSCP and CoS WAN edge to ISP Trust outbound markings (your own); re-mark or zero inbound (ISP's) The classic failure: trusting markings from a host PC. A user's machine can mark every packet as DSCP EF (voice) and starve the actual voice traffic. Always re-mark or police at the access port for any device whose markings you cannot vouch for. ## Queueing: How Cisco Decides What Leaves Next When an interface is congested, multiple packets are waiting to leave. The queueing scheduler decides the order. Cisco supports several scheduling algorithms: FIFO (First-In-First-Out) BehaviorOne queue, no priority Use for Default on uncongested links; never on production WAN Priority Queueing (PQ) Behavior 4 queues; higher always served first; can starve lower Use for Legacy; rarely used today Weighted Fair Queueing (WFQ) Behavior Per-flow queues with proportional service Use for Default on slow serial interfaces; rarely tuned Class-Based Weighted Fair Queueing (CBWFQ) Behavior Per-class queues; configured bandwidth guarantees Use for Modern default for non-real-time traffic Low-Latency Queueing (LLQ) Behavior CBWFQ plus a strict-priority queue with policer Use for Modern default when voice/video share with data LLQ is the dominant queueing strategy in 2026\. It gives voice a strict priority queue (lowest latency, lowest jitter) but applies a built-in policer to prevent the voice queue from starving everything else if voice traffic explodes (e.g. a misbehaving SIP gateway). The remaining bandwidth is divided among the other classes proportionally via CBWFQ. A minimum-viable LLQ + CBWFQ policy on a WAN-facing interface looks like this. The class-maps match on DSCP at the egress; the policy-map is the new piece: ``` ! Match the DSCP markings at the WAN egress class-map match-any VOICE match dscp ef class-map match-any VIDEO match dscp af41 ! LLQ + CBWFQ: voice in priority queue, video and ! default with guaranteed bandwidth percentages policy-map WAN-OUT class VOICE priority percent 20 class VIDEO bandwidth percent 30 class class-default bandwidth percent 25 ! Apply at egress on the WAN interface interface Ethernet0/1 description WAN to provider bandwidth 100000 service-policy output WAN-OUT ``` This policy reserves 20 percent of the configured 100 Mbps interface bandwidth for voice (strict priority), 30 percent as a CBWFQ minimum for video, and 25 percent for everything else. The unallocated 25 percent is held in reserve for control-plane traffic. Use `show policy-map interface` to confirm the policy attached and is classifying correctly, covered in the verification section below. For the configuration walkthrough of LLQ, CBWFQ, and how priority+bandwidth statements interact, see (article forthcoming). ## Policing vs Shaping Both policing and shaping limit traffic to a configured rate. They differ in what happens to the excess: Policing Excess trafficDropped (or re-marked) Where appliedIngress or egress Use for Hard rate limits; SLA enforcement Shaping Excess trafficBuffered and delayed Where appliedEgress only Use for Smoothing bursts to fit downstream link The classic shaping use case: your branch has a 50 Mbps Metro Ethernet handoff but the carrier rate-limits you to 20 Mbps. Without shaping, your switch sends bursts of 50 Mbps and the carrier drops the excess. With egress shaping at 20 Mbps, the switch buffers the burst and feeds the link at the rate the carrier expects. No drops, no packet loss. The classic policing use case: an SLA contract says "you can send 100 Mbps; anything above gets dropped." The ISP polices at ingress; your egress shaping ensures you do not get policed in the first place. Both can re-mark instead of drop. A common pattern: police the scavenger class to a percentage of the link, and re-mark traffic that exceeds the cap to the same class but with a higher drop precedence. The Cisco MQC syntax for both is short. Policing drops the excess; shaping buffers and delays it: ``` ! Policer: cap at 10 Mbps, drop excess policy-map LIMIT-IN class class-default police 10000000 ! Shaper: cap at 20 Mbps, queue and pace excess policy-map SHAPE-OUT class class-default shape average 20000000 ``` The policer is one configured line (`police` with a rate in bits per second) plus optional burst sizes and exceed/violate actions. The shaper is one line (`shape average`) plus optional buffer-tuning knobs. Most production designs combine the two: shape outbound to fit the carrier's policed rate, so your traffic never gets policed in the first place. ## Voice and Video QoS The dominant real-time use case. Voice (VoIP / RTP) has tight requirements: - **Latency < 150ms** end-to-end (one-way; ITU-T G.114) - **Jitter < 30ms** - **Packet loss < 1 percent** (preferably much lower) The standard pattern: classify voice traffic at the access port (typically by voice VLAN), mark to DSCP EF, place into LLQ priority queue at every egress interface in the path, and reserve enough bandwidth so the priority queue never has to drop. For a 1 Gbps WAN link supporting 100 simultaneous VoIP calls (each \~80 kbps including overhead), you reserve 8 Mbps for the priority queue. The MQC implementation of LLQ is short. Define a priority queue for voice and let video and class-default share the rest with bandwidth-percent guarantees: ``` policy-map WAN-OUT class VOICE priority percent 20 ! strict priority + built-in policer class VIDEO bandwidth percent 30 ! CBWFQ guaranteed minimum class class-default bandwidth percent 25 ! CBWFQ for everything else ``` The `priority percent 20` line is what makes this LLQ rather than plain CBWFQ. It sets up a strict-priority queue for the VOICE class but caps it at 20 percent of the interface bandwidth so misbehaving voice traffic cannot starve video or data. The `bandwidth percent` lines are CBWFQ minimums: classes can use more bandwidth when available, but are guaranteed at least their configured share under congestion. Video conferencing (Zoom, Teams, Webex) has slightly looser requirements but is more bandwidth-hungry. Mark to AF41 or CS4 depending on your design; place in a guaranteed-bandwidth class with WRED for graceful degradation under congestion. For the full configuration walkthrough including IP phone trust, voice VLAN, LLQ tuning, and verification commands, see (article forthcoming). ## Verifying QoS The single most useful command for verifying that a QoS policy is doing what you designed is `show policy-map interface`. It reports per-class packet and byte counters, the queue depth, total drops, and (for LLQ) the priority-queue policer state. After driving 200 voice-marked, 200 video-marked, and 200 unmarked test packets through the lab's `WAN-OUT` policy on Cisco IOS XE 17.16, the populated output looks like this: ``` R1# show policy-map interface Ethernet0/1 Ethernet0/1 Service-policy output: WAN-OUT queue stats for all priority classes: Queueing queue limit 64 packets (queue depth/total drops/no-buffer drops) 0/131/0 (pkts output/bytes output) 69/69966 Class-map: VOICE (match-any) 200 packets, 202800 bytes 5 minute offered rate 0000 bps, drop rate 0000 bps Match: dscp ef (46) 200 packets, 202800 bytes Priority: 20% (20000 kbps), burst bytes 500000, b/w exceed drops: 0 Class-map: VIDEO (match-any) 200 packets, 282800 bytes 5 minute offered rate 3000 bps, drop rate 0000 bps Match: dscp af41 (34) 200 packets, 282800 bytes Queueing queue limit 64 packets (queue depth/total drops/no-buffer drops) 0/131/0 (pkts output/bytes output) 69/97566 bandwidth 30% (30000 kbps) Class-map: class-default (match-any) 241 packets, 309892 bytes 5 minute offered rate 5000 bps, drop rate 2000 bps Match: any Queueing queue limit 64 packets (queue depth/total drops/no-buffer drops) 0/131/0 (pkts output/bytes output) 110/111558 bandwidth 25% (25000 kbps) ``` Read this from the top: - **Service-policy output: WAN-OUT** confirms the policy attached. If the line is missing, the `service-policy output` statement is not on the interface. - **Class-map: VOICE ... Match: dscp ef (46)** proves classification is matching. The 200 packets matching DSCP EF were the 200 voice-marked test pings. - **Priority: 20% (20000 kbps), b/w exceed drops: 0** shows the LLQ policer is configured for 20 Mbps and the priority queue never hit its cap. The 131 drops in the priority queue stats above are queue-depth tail drops, not policer drops. - **VIDEO and class-default** show the same per-class accounting plus the configured `bandwidth percent` in absolute terms (30 percent of the 100 Mbps interface = 30000 kbps). Change the interface `bandwidth` statement and these absolute numbers change with it. - **5 minute drop rate** is the rolling drop-rate counter. Non-zero means the class is being shed faster than the queue can drain - the canary for "bump this class' bandwidth or police harder upstream." The contrast between a fresh attach (zero everywhere) and these populated counters is the whole point of QoS verification. Counters that move prove the policy is reaching the dataplane; counters that match expected packet counts prove classification is matching the right traffic; non-zero drops in the wrong class catch a misconfigured trust boundary or starving CBWFQ allocation before users complain. Pair with `show class-map`, `show policy-map`, and `show running-config interface` to cover the "is this thing configured?" questions on any QoS rollout. ## QoS in Modern Networks: SD-WAN and Wireless Two modern contexts where QoS is implemented differently from the legacy MPLS WAN model: **SD-WAN.** Application-aware steering replaces a lot of legacy QoS complexity. Instead of marking traffic and trusting the WAN to honor markings, the SD-WAN edge identifies applications via DPI and steers them across multiple transports based on real-time SLA measurements. QoS still exists - voice still gets a priority queue at each WAN egress - but the policy is expressed in terms of applications, not DSCP values, and the SD-WAN handles per-tunnel SLA monitoring. See the [SD-WAN cluster pillar](https://www.pinglabz.com/sd-wan/) for the architecture and the [SD-WAN architecture article](https://www.pinglabz.com/sd-wan-architecture/) for the per-tunnel SLA model. **Wireless.** Wireless QoS uses 802.11e WMM (Wi-Fi Multimedia) with four access categories: Voice, Video, Best Effort, Background. The Catalyst 9800 maps DSCP values from incoming wired traffic into WMM categories on the air, and CoS-marked frames from wireless clients into DSCP for forwarding upstream. Auto QoS and AVC (Application Visibility and Control) handle most of the configuration automatically. See [C9800 QoS Configuration: Auto QoS, DSCP Mapping, and Wireless Profiles](https://www.pinglabz.com/c9800-qos-configuration/) for the wireless-specific walkthrough. The wireless half of this story - AVC, metal QoS profiles, and per-SSID policing on the Catalyst 9800 - lives in the [Cisco wireless complete guide](https://www.pinglabz.com/wireless/). ## CCNP Depth: Advanced Queueing and Architecture The concepts above are the QoS foundation; the CCNP goes deeper into the queueing mechanics and pairs QoS with enterprise design. Four QoS deep-dives, all with real captures from a congested link: [LLQ and CBWFQ on Cisco IOS XE](https://www.pinglabz.com/llq-cbwfq-cisco-ios-xe/) builds the full queueing policy and proves it under real congestion: on a 5 Mbps choke, the bulk transfer lost 162 packets while the voice priority class lost zero (`b/w exceed drops: 0`). [WRED and congestion avoidance](https://www.pinglabz.com/wred-congestion-avoidance/) explains why you drop packets early and at random to defeat TCP global synchronization. [Hierarchical QoS](https://www.pinglabz.com/hierarchical-qos-shaping/) solves the case a flat policy cannot (a 100 Mbps interface delivering a 20 Mbps service) by shaping in a parent and queueing in a nested child. And [interpreting QoS configurations](https://www.pinglabz.com/interpreting-qos-configurations-encor/) is the exam skill of reading a policy and its output and saying whether it works. ## Enterprise Architecture and High Availability The ENCOR architecture domain pairs QoS with design. [Enterprise campus design](https://www.pinglabz.com/enterprise-campus-design-ccnp/) covers access/distribution/core, two-tier vs three-tier, the Layer 2/Layer 3 boundary, and the spine-leaf direction. [High availability techniques](https://www.pinglabz.com/sso-nsf-graceful-restart/) covers SSO, NSF, and graceful restart (with a real OSPF NSF helper capture) so a supervisor failure does not take the network down. And [on-prem vs cloud network design](https://www.pinglabz.com/on-prem-vs-cloud-network-design/) covers what changes, and what stays the same, when the workloads move to a VPC. ## QoS Deep Dives in This Cluster The articles in this cluster, in reading order: 1. [DSCP and IP Precedence Explained Byte by Byte](https://www.pinglabz.com/dscp-ip-precedence-explained/) 2. [Cisco MQC (Modular QoS CLI): The Operator's Walkthrough](https://www.pinglabz.com/cisco-mqc/) 3. [QoS Classification, Marking, and Trust Boundaries](https://www.pinglabz.com/qos-classification-marking-trust-boundary/) 4. [QoS Queueing, Policing, and Shaping Compared](https://www.pinglabz.com/qos-queueing-policing-shaping/) 5. [Voice and Video QoS on Cisco IOS XE](https://www.pinglabz.com/voice-video-qos-cisco/) 6. [QoS in Modern Networks: SD-WAN, Cloud, and Application-Aware Steering](https://www.pinglabz.com/qos-sd-wan-modern-networks/) 7. [C9800 QoS Configuration: Auto QoS, DSCP Mapping, and Wireless Profiles](https://www.pinglabz.com/c9800-qos-configuration/) 8. [LLQ and CBWFQ on Cisco IOS XE: Building the Queueing Policy](https://www.pinglabz.com/llq-cbwfq-cisco-ios-xe/) 9. [WRED and Congestion Avoidance: Dropping Packets on Purpose](https://www.pinglabz.com/wred-congestion-avoidance/) 10. [Hierarchical QoS: Shaping Parents and Queueing Children](https://www.pinglabz.com/hierarchical-qos-shaping/) 11. [Interpreting QoS Configurations: An ENCOR Exam Skill Guide](https://www.pinglabz.com/interpreting-qos-configurations-encor/) 12. [Enterprise Campus Design: 2-Tier vs 3-Tier at CCNP Depth](https://www.pinglabz.com/enterprise-campus-design-ccnp/) 13. [High Availability Techniques: SSO, NSF, and Graceful Restart](https://www.pinglabz.com/sso-nsf-graceful-restart/) 14. [On-Prem vs Cloud Network Design: What ENCOR Wants You to Know](https://www.pinglabz.com/on-prem-vs-cloud-network-design/) This list grows as new articles are published. Check back for vendor-specific deep dives, configuration walkthroughs, and troubleshooting references. Hands-on QoS - classification, marking, LLQ + CBWFQ Configure QoS class-maps for VOICE (DSCP EF) and VIDEO (AF41), then apply LLQ + CBWFQ on a Cisco WAN egress interface. Per-class packet counters from `show policy-map interface`. Open the PingLabz CCNA Labs library. [Open the QoS labs](https://www.pinglabz.com/ccna-labs-ip-services/) ### More QoS guides in this cluster 1\. [QoS for VoIP Across an MPLS WAN](https://www.pinglabz.com/qos-for-voip/) 2\. [QoS on a Router: A Practical Cisco IOS XE Walkthrough](https://www.pinglabz.com/qos-on-a-router/) 3\. [Configuring Voice VLANs on Cisco Switches for IP Phones](https://www.pinglabz.com/voice-vlan-cisco-configuration/) ## Frequently Asked Questions ### What does QoS stand for? QoS stands for Quality of Service. It refers to the set of network mechanisms (classification, marking, queueing, policing, shaping) that control how a network treats different kinds of traffic when resources are constrained. The same acronym is used loosely in other contexts (MQTT QoS levels, Kubernetes QoS classes, Queen of Spades cultural references), but in networking it means the IP/Ethernet QoS toolkit covered here. ### Do I need QoS on a network that is not congested? Probably not on the data plane. If your links are routinely under 50 percent utilization and you have no real-time applications (no VoIP, no video conferencing, no industrial control systems), best-effort works fine. You should still mark traffic at the edge (it costs nothing and lets the network be ready when it does become congested), but elaborate queueing and shaping is overkill. You almost certainly do need QoS on WAN edges, wireless, and any link prone to bursting (cloud egress, sub-1 Gbps WAN, congested interfaces during business hours). ### What is the difference between DSCP and CoS? DSCP is a Layer 3 marking in the IP header (6 bits, 64 values) and is preserved end-to-end across IP routing. CoS is a Layer 2 marking in the 802.1Q tag (3 bits, 8 values) and only exists on tagged Ethernet frames; it is stripped at every routed hop. Most networks classify and mark DSCP at the edge; CoS is only relevant inside switched Layer 2 segments. See the [802.1Q VLAN Tag Explained](https://www.pinglabz.com/802-1q-vlan-tag-explained/) for where CoS lives in the frame. ### What is LLQ and how does it relate to CBWFQ? LLQ (Low-Latency Queueing) is CBWFQ (Class-Based Weighted Fair Queueing) plus a strict-priority queue with a built-in policer. The priority queue is for voice and other real-time traffic; everything else is in CBWFQ classes with configured bandwidth guarantees. The built-in policer prevents the priority queue from starving the rest of the policy if voice traffic explodes. LLQ is the modern Cisco default for any link sharing voice with data. ### When should I use policing vs shaping? Policing drops or re-marks excess traffic; use it for hard rate limits and SLA enforcement. Shaping buffers and delays excess traffic; use it for smoothing bursts to fit a downstream link. Common pattern: shape egress to match the carrier's policed rate so you do not get policed in the first place. Detail in the queueing/policing/shaping article. ### Does QoS still matter in cloud? Differently. The cloud provider does not honor your DSCP markings inside their network (you do not own that path). What matters in cloud-bound traffic is QoS at your WAN edge (so voice and video get priority leaving your network) and at the cloud-on-ramp (Direct Connect, ExpressRoute) where you have control. Once the packet enters AWS or Azure, your markings are irrelevant. SD-WAN cloud on-ramps re-establish per-tunnel QoS over those links. ### Where should I mark traffic? As close to the source as possible, on a device you trust. The classic pattern: classify and mark at the access switch port for end hosts, at the WAN edge router for inbound flows from outside, and at the application server for traffic the application can mark itself. Re-marking elsewhere should be limited to scavenger-class enforcement at the trust boundary. ## Key Takeaways If you take one thing away from this guide, make it this: QoS is mostly about discipline at the edge. Classify and mark once at a controlled trust boundary, trust those markings everywhere inside the QoS domain, and apply queueing/policing/shaping at congestion points. The protocol mechanics (DSCP values, LLQ, MQC) are simpler than the operational discipline of getting the trust boundary right and keeping policy consistent across hundreds of devices. Bookmark this page, work through the cluster articles in reading order, and lab every change. QoS is the kind of network design where small misconfigurations cause subtle problems that only show up under load - exactly when you cannot afford them. The cluster articles will give you the patterns; the discipline of edge marking plus consistent end-to-end policy is what makes them work in production. **Studying for the CCNA?** Test your QoS knowledge on [PingLabz CCNA Flashcards](https://www.pinglabz.com/ccna-flashcards/) \- 200 free multiple-choice questions by topic, mixed, or a full mock exam, each with a plain-English explanation. ### References - [RFC 2474 - Definition of the Differentiated Services Field (DSCP)](https://www.rfc-editor.org/rfc/rfc2474?ref=pinglabz.com) - [RFC 3246 - An Expedited Forwarding PHB](https://www.rfc-editor.org/rfc/rfc3246?ref=pinglabz.com) - [Cisco QoS technology documentation](https://www.cisco.com/c/en/us/tech/quality-of-service-qos/index.html?ref=pinglabz.com) ### EIGRP (Enhanced Interior Gateway Routing Protocol): The Complete Guide URL: https://www.pinglabz.com/eigrp/ Last updated: 2026-07-12T10:04:02.000Z EIGRP (Enhanced Interior Gateway Routing Protocol) is the routing protocol that occupies the awkward middle ground between OSPF and BGP. It is fast like OSPF but converges differently. It is policy-rich like BGP but limited to a single AS. It was Cisco-proprietary for 26 years, opened to standards in 2013, and is still found in 2026 Cisco-shop networks where someone built a deployment in 2008 and it still works. This is the cluster overview for the full PingLabz EIGRP series: fundamentals, the DUAL algorithm that makes EIGRP convergence unique, the K values and composite metric, configuration on Cisco IOS XE, stub routing, and how EIGRP compares to OSPF and BGP. We will work through what EIGRP actually is, the DUAL state machine that drives loop-free convergence, the metric calculation that has confused generations of CCNP candidates, and the configuration patterns that work in production. Every capture on this page comes from a real four-router IOS XE lab (R1-R2-R3-R4 in a partial mesh, with an R2-R4 backup link so DUAL can be observed end-to-end). New labs and guides, in your inbox Every new PingLabz lab and deep-dive, built and verified on real Cisco IOS XE - free, straight to your inbox. [Join free](https://www.pinglabz.com/signup/) ## What EIGRP Is EIGRP is an advanced distance-vector protocol (some call it a hybrid protocol because it has link-state-like features). Unlike pure distance-vector protocols (RIP, IGRP), EIGRP does not periodically broadcast its full routing table. Instead it sends incremental updates only when topology changes, uses the Diffusing Update Algorithm (DUAL) to compute loop-free paths, and maintains neighbor relationships via Hello packets. The defining characteristics: - **Cisco-led, now open.** EIGRP was Cisco-proprietary from 1992 until 2013 when Cisco published the basic protocol as [RFC 7868](https://www.rfc-editor.org/rfc/rfc7868?ref=pinglabz.com). Adoption beyond Cisco remains rare; it is Cisco's IGP for Cisco shops. - **Fast convergence.** Sub-second on direct failure when a feasible successor exists. The DUAL algorithm pre-computes alternates so failover does not require recomputation. - **Composite metric.** Unlike OSPF (cost only) or RIP (hop count), EIGRP combines bandwidth, delay, and (configurably) load and reliability into a single 64-bit metric. The default uses bandwidth and delay only; the K values control the formula. - **Classless and hierarchical.** Supports VLSM and route summarization at any boundary, manual or automatic. - **Reliable transport.** RTP (Reliable Transport Protocol) provides ordered delivery for control-plane updates, similar to TCP but more efficient for small frequent messages. - **Multicast-based discovery.** Hellos and updates use 224.0.0.10 (IPv4) or FF02::A (IPv6). For the historical context and why EIGRP exists alongside OSPF and BGP, see [BGP vs OSPF](https://www.pinglabz.com/bgp-vs-ospf/) and the upcoming EIGRP vs OSPF comparison article. ## How EIGRP Works (the 10,000-Foot View) EIGRP forms neighbor relationships, exchanges full routing information once at session establishment, then sends incremental updates only when topology changes. Three protocol phases: 1. **Neighbor discovery.** Routers send Hello packets to multicast 224.0.0.10\. If parameters match (AS number, K values, authentication), neighbors form. EIGRP validates parameters strictly; mismatches mean no adjacency. 2. **Initial topology exchange.** Once neighbors form, they exchange full Update packets containing all known routes. Updates use RTP for reliable delivery. 3. **Steady state with DUAL.** All routes are stored in the topology table. Best paths (and feasible successors) are installed. Topology changes trigger DUAL recomputation, which uses the topology table to find loop-free alternates without polling the network. The genius of DUAL: when a primary path fails and a feasible successor exists in the topology table, the router fails over instantly using the cached alternate. No queries to neighbors needed. This is what makes EIGRP convergence so fast on healthy networks. ## DUAL: The Algorithm That Makes EIGRP Different DUAL (Diffusing Update Algorithm) is EIGRP's loop-prevention and convergence machinery. The four key concepts: Successor The neighbor offering the best (lowest-metric) path to a destination. The route via this neighbor is installed in the routing table. Feasible Successor (FS) A backup neighbor whose Reported Distance is less than our current Feasible Distance. Pre-computed; usable instantly if successor fails. With [variance](https://www.pinglabz.com/eigrp-variance-unequal-cost-load-balancing/), feasible successors can even carry traffic simultaneously. Feasible Distance (FD) The lowest metric we have ever seen for this route. Set when route is first installed; only updated downward. Reported Distance (RD) The metric the neighbor reports to us. The "distance from the neighbor's perspective." The Feasibility Condition: a route is a Feasible Successor if its Reported Distance is strictly less than the current Feasible Distance. RD < FD. This guarantees loop-freedom because the neighbor must already have a shorter path to the destination than ours. When the successor fails: - **If a Feasible Successor exists:** install it as the new successor immediately. Sub-second convergence. The route never enters Active state. - **If no Feasible Successor exists:** the route enters Active state. The router queries neighbors for alternative paths. Slower convergence; can take seconds. If queries time out (Stuck-In-Active), the neighbor is declared dead - the mechanics and the design fixes are in the [query process and SIA deep dive](https://www.pinglabz.com/eigrp-query-process-stuck-in-active/). The differentiator capture: a real DUAL state transition. R3 in the lab uses R2 as the successor for several prefixes (including 10.255.0.2/32, R2's loopback). Killing R3's Ethernet0/0 (the link to R2) makes EIGRP mark every R2-reached destination Active, query R4, receive the reply, and re-install the routes via R4 - the whole thing in roughly 23 milliseconds: ``` R3(config)#interface Ethernet0/0 R3(config-if)#shutdown ! kill the successor ! debug eigrp fsm output (captured to logging buffer): *May 11 06:18:02.245: DUAL: AS(100) Dest 10.255.0.2/32 entering active state for tid 0. *May 11 06:18:02.245: EIGRP-IPv4(100): Set reply-status table. Count is 1. *May 11 06:18:02.245: EIGRP-IPv4(100): Not doing split horizon *May 11 06:18:02.268: EIGRP-IPv4(100): dest(10.30.31.0/30) active *May 11 06:18:02.268: EIGRP-IPv4(100): rcvreply: 10.30.31.0/30 via 10.30.32.2 metric 262144000/196608000 *May 11 06:18:02.268: EIGRP-IPv4(100): reply count is 1 *May 11 06:18:02.268: DUAL: AS(100) Clearing handle 1, count now 0 *May 11 06:18:02.268: EIGRP-IPv4(100): Find FS for dest 10.30.31.0/30 ... found *May 11 06:18:02.268: DUAL: AS(100) RT installed 10.30.31.0/30 via 10.30.32.2 *May 11 06:18:02.269: EIGRP-IPv4(100): rcvreply: 10.255.0.1/32 via 10.30.32.2 metric 262225920/196689920 *May 11 06:18:02.269: DUAL: AS(100) RT installed 10.255.0.1/32 via 10.30.32.2 *May 11 06:18:02.269: EIGRP-IPv4(100): rcvreply: 10.255.0.2/32 via 10.30.32.2 metric 196689920/131153920 *May 11 06:18:02.269: DUAL: AS(100) RT installed 10.255.0.2/32 via 10.30.32.2 ``` Five things are happening in that trace. The route enters Active state. EIGRP sets a reply-status table for the remaining neighbor (R4 via 10.30.32.2) and sends a Query. R4 replies with its own known path. EIGRP runs *Find FS* to confirm the new path satisfies the feasibility condition. The new successor is installed and the prefix returns to Passive. No other IGP makes the state transition this visible - OSPF and IS-IS just flood an LSA / LSP and every router re-runs Dijkstra in its own head. For the deep-dive on DUAL with worked examples, see [EIGRP DUAL Algorithm Deep Dive](https://www.pinglabz.com/eigrp-dual-algorithm/). ## The EIGRP Composite Metric EIGRP's metric is the part most engineers find confusing. It uses a 64-bit composite calculated from up to five inputs (bandwidth, delay, load, reliability, MTU), weighted by configurable K values. The classic IGRP/EIGRP metric formula (default K1=1, K3=1, others=0): ``` metric = 256 * (10^7 / min_bandwidth + sum_delay) ``` Where: - **min\_bandwidth** is the minimum interface bandwidth along the path, in kbps - **sum\_delay** is the sum of all interface delays along the path, in tens of microseconds The 64-bit "wide metric" introduced in IOS 15.x uses the same formula but with larger ranges and units, supporting modern interface speeds without saturation. The full 64-bit format, picosecond delay units, and rib-scale are decoded in the [EIGRP wide metrics deep dive](https://www.pinglabz.com/eigrp-wide-metrics/). K values let you weight different inputs: 1 K valueK1 ControlsBandwidth weight 0 K valueK2 ControlsBandwidth/load weight 1 K valueK3 ControlsDelay weight 0 K valueK4 ControlsReliability weight 0 K valueK5 ControlsReliability weight Critical: all routers in the same EIGRP AS must use the same K values or neighbor relationships will not form. Cisco strongly recommends never changing the defaults. The K values are visible in `show ip protocols` on any EIGRP router: ``` R3#show ip protocols Routing Protocol is "eigrp 100" EIGRP-IPv4 VR(PINGLABZ) Address-Family Protocol for AS(100) Metric weight K1=1, K2=0, K3=1, K4=0, K5=0 K6=0 Metric rib-scale 128 Metric version 64bit Soft SIA disabled NSF-aware route hold timer is 240 Router-ID: 10.255.0.3 Topology : 0 (base) Active Timer: 3 min Distance: internal 90 external 170 Maximum path: 4 Maximum hopcount 100 Maximum metric variance 1 Total Prefix Count: 8 Total Redist Count: 0 Automatic Summarization: disabled ``` Two things are worth noticing in that output: `Metric version 64bit` confirms named-mode EIGRP is using the wide metric, and the `Distance: internal 90 external 170` line is the EIGRP administrative distance you will spot in any RIB tie-break against another protocol. *Maximum metric variance 1* is the unequal-cost load-balancing knob - leaving it at 1 means EIGRP installs equal-cost paths only; raising it to 2 would install any path whose metric is up to 2x the successor. For the full math walkthrough including the wide-metric variants and how to verify metrics in show output, see [EIGRP Metric and K Values Explained](https://www.pinglabz.com/eigrp-metric-k-values/). ## Neighbor States and Adjacency Requirements Five things must match for an EIGRP neighbor relationship to form (and when one does not, the [adjacency troubleshooting guide](https://www.pinglabz.com/troubleshooting-eigrp-neighbor-adjacencies/) shows the exact failure output for each): Same AS number The number after `router eigrp X` Same K values Default works everywhere; never change Authentication match If MD5/SHA configured, both ends must agree - config for both is in the [EIGRP authentication guide](https://www.pinglabz.com/eigrp-authentication-md5-sha/) Subnet must match Both ends on same primary subnet (rare for two routers to disagree but happens with secondary IPs) Hello/Hold timers compatible Hello defaults: 5s on Ethernet, 60s on slow NBMA. Mismatched is OK as long as Hold > Hello on both sides Verify with `show ip eigrp neighbors`. From the lab, R3 has two neighbors (R2 via Et0/0 and R4 via Et0/1): ``` R3#show ip eigrp neighbors EIGRP-IPv4 VR(PINGLABZ) Address-Family Neighbors for AS(100) H Address Interface Hold Uptime SRTT RTO Q Seq (sec) (ms) Cnt Num 1 10.30.32.2 Et0/1 12 00:01:01 2 100 0 6 0 10.30.31.1 Et0/0 11 00:01:07 1263 5000 0 13 ``` The H column is a DUAL-internal handle. Hold is the seconds remaining before the neighbor is declared dead (default 15 on broadcast media, refreshed by every Hello). SRTT is the smoothed round-trip time used by EIGRP's reliable transport (RTP). The Q (Queue) column matters: a non-zero queue means EIGRP is buffering updates because the neighbor has not acknowledged. Persistent non-zero queue indicates RTP issues - link congestion, ACL filtering, or a dying neighbor. ## Configuration on Cisco IOS XE The minimum EIGRP configuration: ``` R1(config)# router eigrp 100 R1(config-router)# network 10.0.0.0 0.0.255.255 R1(config-router)# passive-interface default R1(config-router)# no passive-interface GigabitEthernet0/0/0 R1(config-router)# no auto-summary ``` Three things to notice. First, the AS number (100) must match across all neighbors in the same EIGRP domain. Second, the wildcard mask in `network` is inverted from a regular subnet mask. Third, `no auto-summary` is mandatory in modern networks - the legacy auto-summarization at classful boundaries causes problems with VLSM and is enabled by default in old code. Once the process is up, the usual next steps are [route summarization and default-route origination](https://www.pinglabz.com/eigrp-summarization-default-routes/) and [route filtering with distribute lists](https://www.pinglabz.com/eigrp-route-filtering-distribute-lists/); routes that later go missing are usually one of those two doing its job, per the [missing-routes troubleshooting guide](https://www.pinglabz.com/troubleshooting-eigrp-missing-routes/). The named-mode configuration (preferred for new deployments since IOS 15.x): ``` R1(config)# router eigrp PINGLABZ R1(config-router)# address-family ipv4 unicast autonomous-system 100 R1(config-router-af)# af-interface default R1(config-router-af-interface)# passive-interface R1(config-router-af-interface)# exit-af-interface R1(config-router-af)# af-interface Ethernet0/0 R1(config-router-af-interface)# no passive-interface R1(config-router-af-interface)# exit-af-interface R1(config-router-af)# network 10.30.30.0 0.0.0.3 R1(config-router-af)# network 10.255.0.1 0.0.0.0 R1(config-router-af)# eigrp router-id 10.255.0.1 R1(config-router-af)# exit-address-family ``` Named mode separates IPv4 and IPv6 cleanly, scales better for multi-AS deployments, and is what new CCNP labs expect. After convergence, the topology table shows the successors for every learned prefix: ``` R3#show ip eigrp topology EIGRP-IPv4 VR(PINGLABZ) Topology Table for AS(100)/ID(10.255.0.3) Codes: P - Passive, A - Active, U - Update, Q - Query, R - Reply, r - reply Status, s - sia Status P 10.255.0.4/32, 1 successors, FD is 131153920 via 10.30.32.2 (131153920/163840), Ethernet0/1 P 10.255.0.1/32, 1 successors, FD is 196689920 via 10.30.31.1 (196689920/131153920), Ethernet0/0 P 10.30.33.0/30, 2 successors, FD is 196608000 via 10.30.31.1 (196608000/131072000), Ethernet0/0 via 10.30.32.2 (196608000/131072000), Ethernet0/1 P 10.30.30.0/30, 1 successors, FD is 196608000 via 10.30.31.1 (196608000/131072000), Ethernet0/0 P 10.30.32.0/30, 1 successors, FD is 131072000 via Connected, Ethernet0/1 P 10.30.31.0/30, 1 successors, FD is 131072000 via Connected, Ethernet0/0 P 10.255.0.2/32, 1 successors, FD is 131153920 via 10.30.31.1 (131153920/163840), Ethernet0/0 P 10.255.0.3/32, 1 successors, FD is 163840 via Connected, Loopback0 ``` Every line is `P` (Passive) - the steady state for a destination, meaning DUAL has finished all calculations and the route is installed. The FD (Feasible Distance) and the metric tuple `(total/RD)` are exactly what DUAL feeds into the feasibility condition. The `show ip eigrp topology all-links` variant shows every path EIGRP has heard about, not just the successors - that is the table DUAL consults when picking feasible successors. For the full walkthrough, see [EIGRP Configuration on Cisco IOS XE](https://www.pinglabz.com/eigrp-configuration-cisco/). Named mode also carries EIGRP into IPv6 through address families - the [IPv6 complete guide](https://www.pinglabz.com/ipv6/) covers the addressing side. ## Stub Routing EIGRP stub routing optimizes hub-and-spoke topologies by limiting what the spoke advertises and stopping the hub from querying the spoke during DUAL active states (queries and their failure mode, [stuck-in-active](https://www.pinglabz.com/eigrp-query-process-stuck-in-active/), are covered separately). The result: faster hub convergence and less control-plane churn. ``` ! On a spoke router router eigrp 100 eigrp stub connected summary ``` Stub options control what the spoke can announce: `connected`, `summary`, `static`, `redistributed`, or `receive-only`. The dominant pattern for branch routers in hub-and-spoke designs is `eigrp stub connected summary`. Detail in [EIGRP Stub Routing](https://www.pinglabz.com/eigrp-stub-routing/). Stub configuration is standard practice on DMVPN spokes, where EIGRP runs over multipoint [GRE tunnels](https://www.pinglabz.com/gre/) and an un-stubbed spoke can accidentally become transit for the whole WAN. ## EIGRP vs OSPF vs BGP Type EIGRP Advanced distance-vector (DUAL) OSPF Link-state (Dijkstra SPF) BGPPath-vector Standards EIGRP RFC 7868 (basic), Cisco-led OSPFRFC 2328 (open) BGPRFC 4271 (open) Default Cisco AD EIGRP 90 internal / 170 external OSPF110 BGP20 / 200 Convergence EIGRP Sub-second when FS exists; seconds for queries OSPFSub-second with tuning BGPSlow on purpose Metric EIGRP Composite (bandwidth, delay, load, reliability) OSPF Cost (bandwidth-derived) BGP 13-step best-path with attributes Hierarchy EIGRP None native; stub feature for hub-and-spoke OSPF Strict areas with backbone rule BGP Confederations / route reflectors Scope EIGRPIntra-AS OSPFIntra-AS BGPInter-AS Vendor EIGRPCisco-led OSPFUniversal BGPUniversal For the full comparison see [EIGRP vs OSPF: When to Use Each](https://www.pinglabz.com/eigrp-vs-ospf/) and the cross-cluster [BGP vs OSPF](https://www.pinglabz.com/bgp-vs-ospf/) piece. Both comparisons run deeper than one section: see the [OSPF complete guide](https://www.pinglabz.com/ospf/) and the [BGP complete guide](https://www.pinglabz.com/bgp/) for the full treatment of each protocol. In MPLS L3VPN environments, EIGRP also survives at the WAN edge as a PE-CE protocol - context in the [MPLS complete guide](https://www.pinglabz.com/mpls/). ## EIGRP Deep Dives in This Cluster Sixteen articles, in reading order. Foundations first, then configuration and policy, then the operational material. ### Foundations 1. [The DUAL algorithm: successors, feasible successors, FD and RD](https://www.pinglabz.com/eigrp-dual-algorithm/) 2. [The EIGRP composite metric and K values](https://www.pinglabz.com/eigrp-metric-k-values/) 3. [64-bit wide metrics in named mode](https://www.pinglabz.com/eigrp-wide-metrics/) 4. [EIGRP administrative distance: internal 90, external 170](https://www.pinglabz.com/eigrp-administrative-distance/) ### Configuration and Policy 1. [EIGRP configuration on IOS XE: classic and named mode](https://www.pinglabz.com/eigrp-configuration-cisco/) 2. [Route summarization, the Null0 discard route, and default routes](https://www.pinglabz.com/eigrp-summarization-default-routes/) 3. [Unequal-cost load balancing with variance](https://www.pinglabz.com/eigrp-variance-unequal-cost-load-balancing/) 4. [Authentication with MD5 key chains and HMAC-SHA-256](https://www.pinglabz.com/eigrp-authentication-md5-sha/) 5. [Route filtering with distribute lists, prefix lists, and route maps](https://www.pinglabz.com/eigrp-route-filtering-distribute-lists/) 6. [Stub routing for hub-and-spoke WANs](https://www.pinglabz.com/eigrp-stub-routing/) ### Convergence and Scale 1. [The 5 neighbor requirements that must match](https://www.pinglabz.com/eigrp-neighbor-requirements/) 2. [The query process and stuck-in-active (SIA)](https://www.pinglabz.com/eigrp-query-process-stuck-in-active/) ### Troubleshooting 1. [Troubleshooting neighbor adjacencies that will not form](https://www.pinglabz.com/troubleshooting-eigrp-neighbor-adjacencies/) 2. [Troubleshooting route advertisement and missing routes](https://www.pinglabz.com/troubleshooting-eigrp-missing-routes/) ### Design and Comparisons 1. [EIGRP vs OSPF: when to use each](https://www.pinglabz.com/eigrp-vs-ospf/) 2. [Running EIGRP, OSPF, and BGP over GRE tunnels](https://www.pinglabz.com/routing-protocols-over-gre/) 3. [Routing over DMVPN: split horizon, next-hop-self, and OSPF network types on an mGRE overlay](https://www.pinglabz.com/routing-over-dmvpn-eigrp-ospf/) Hands-on EIGRP - 2 CCNA labs included Configure EIGRP named-mode AS 100, then watch DUAL converge in real time with Feasible Successors. Real `show ip eigrp topology` output with FD/AD breakdowns. Part of the 14-lab CCNA IP Connectivity cluster. Open the PingLabz CCNA Labs library. [Open the EIGRP labs](https://www.pinglabz.com/ccna-labs-ip-connectivity/) ## Expert EIGRP: SoO, Leak Maps, DMVPN and Integration (CCIE level) EIGRP at professional depth is DUAL, feasible successors and a clean metric. EIGRP at expert depth is what happens when you put it over a multipoint tunnel, hand it a dual-homed site with no loop-prevention marker, and ask it to share a network with two other protocols. These articles are built on real Cisco IOS XE output from a dual-hub DMVPN in CML, with an OSPF domain and a BGP peer hanging off the hub so the redistribution scenarios are genuine. 1 [EIGRP over DMVPN: multi-hub design and the split-horizon problem](https://www.pinglabz.com/eigrp-over-dmvpn-multi-hub/) The two commands that make spoke-to-spoke routing work on an mGRE hub, and what breaks when either is missing. 2 [EIGRP Site of Origin: loop prevention for dual-homed sites](https://www.pinglabz.com/eigrp-site-of-origin-soo/) EIGRP has no AS-path, so a site will happily learn its own routes back. What SoO does, and what actually works in the global table. 3 [EIGRP summarization with leak maps](https://www.pinglabz.com/eigrp-summary-leak-map/) Send the summary, leak the specifics that matter, and understand the AD 5 discard route that can black-hole your own traffic. 4 [EIGRP offset lists: surgical metric manipulation](https://www.pinglabz.com/eigrp-offset-lists/) Make one hub primary with a single command on a single router - exact arithmetic, no guessing with bandwidth and delay. 5 [Multi-protocol redistribution: the CCIE scenarios that break networks](https://www.pinglabz.com/multi-protocol-redistribution-ccie-scenarios/) Route feedback and the AD race between EIGRP external (170) and OSPF (110), reproduced for real and then fixed with tags. 6 [IGP migration: EIGRP to OSPF without an outage](https://www.pinglabz.com/eigrp-to-ospf-migration/) Ships in the night. Run both, flip the administrative distance, verify, remove. Zero packet loss and one-command rollback. ## Studying for the CCIE? This cluster is part of the full CCNA to CCNP to CCIE Enterprise ladder on PingLabz, every rung built on real Cisco output. For expert-level depth across every EI v1.1 blueprint domain - and the four integration Super Labs - see the [CCIE Enterprise Infrastructure study hub](https://www.pinglabz.com/ccie-enterprise/). ## Frequently Asked Questions ### What does EIGRP stand for? EIGRP stands for Enhanced Interior Gateway Routing Protocol. It is the successor to IGRP (Interior Gateway Routing Protocol), Cisco's earlier distance-vector protocol from the 1980s. EIGRP added classless support, VLSM, the DUAL algorithm, and route summarization at any boundary. ### What is the administrative distance of EIGRP? 90 for internal EIGRP routes (learned from EIGRP neighbors in the same AS) and 170 for external EIGRP routes (redistributed from other protocols). Lower than OSPF (110) and IS-IS (115), higher than directly connected (0) and static (1). ### What protocol number does EIGRP use? EIGRP runs directly on top of IP using protocol number 88\. It does not use TCP or UDP. It uses multicast 224.0.0.10 (IPv4) or FF02::A (IPv6) for Hellos and unsolicited updates. ### EIGRP vs OSPF, which one should I use? OSPF for vendor neutrality and CCNP/CCIE expectations. EIGRP for Cisco-only environments where convergence speed is paramount and the hub-and-spoke stub feature simplifies design. In practice, many enterprises run OSPF (the safer enterprise default in 2026) and EIGRP shows up where someone deployed it years ago and it still works. ### Is EIGRP still Cisco-only? The basic protocol was opened to RFC 7868 in 2013\. Some non-Cisco implementations exist (Open EIGRP for Linux, partial support in some appliances). In production, EIGRP is essentially a Cisco-only protocol; deploying it for vendor-interop is not recommended. ### What is an EIGRP AS number? An AS number identifies the EIGRP routing domain. All routers in the same EIGRP AS exchange routes; routers in different ASes do not form neighbors. The AS number is locally significant in EIGRP (unlike BGP, which uses globally unique ASNs). Range is 1-65535. ## Key Takeaways EIGRP is the routing protocol that exists in the gap between OSPF (link-state, vendor-neutral, hierarchical) and BGP (path-vector, inter-AS, policy-rich). Its DUAL algorithm gives it sub-second convergence on healthy networks, its composite metric handles diverse interface speeds, and its stub feature simplifies hub-and-spoke designs. The cost is Cisco-affiliation: deploying EIGRP commits you to Cisco for IGP across that domain. If you take one thing away from this guide, make it the DUAL feasibility condition: RD < FD guarantees loop-freedom and lets EIGRP fail over instantly without re-querying neighbors. Master DUAL and the rest of EIGRP is mechanics. Bookmark this page, work through the cluster articles in order, and lab every change. **Studying for the CCNA?** Test your EIGRP knowledge on [PingLabz CCNA Flashcards](https://www.pinglabz.com/ccna-flashcards/) \- 200 free multiple-choice questions by topic, mixed, or a full mock exam, each with a plain-English explanation. ### References - [RFC 7868 - Cisco's Enhanced Interior Gateway Routing Protocol (EIGRP)](https://www.rfc-editor.org/rfc/rfc7868?ref=pinglabz.com) - [Cisco EIGRP technology documentation](https://www.cisco.com/c/en/us/tech/ip/enhanced-interior-gateway-routing-protocol-eigrp/index.html?ref=pinglabz.com) ### MPLS (Multiprotocol Label Switching): The Complete Guide URL: https://www.pinglabz.com/mpls/ Last updated: 2026-07-12T10:04:02.000Z MPLS (Multiprotocol Label Switching) is the protocol that runs the inside of every major service provider's network and a substantial fraction of large enterprise WANs. It uses fixed-length labels instead of IP lookup to forward packets, decouples the forwarding plane from the routing plane, and serves as the substrate for VPNs (L3VPN, L2VPN), traffic engineering (RSVP-TE), and now segment routing. It is also the protocol that SD-WAN was supposed to replace, except most production networks still run both. This is the cluster overview for the full PingLabz MPLS series: labels and the label stack, LDP for label distribution, MPLS L3VPN with MP-BGP and VPNv4, traffic engineering, and the segment routing successor story. We will work through what MPLS actually does, how labels move packets through the network, the L3VPN model that made MPLS a service-provider standard, and the modern segment-routing direction. If you are studying for CCIE Service Provider, designing an MPLS deployment, or trying to understand what your carrier is selling you, start here. Every capture on this page comes from a real five-router IOS XE lab (CE1 - PE1 - P - PE2 - CE2) running OSPF as the IGP, LDP across the core, and MP-iBGP carrying VPNv4 between the two PEs. New labs and guides, in your inbox Every new PingLabz lab and deep-dive, built and verified on real Cisco IOS XE - free, straight to your inbox. [Join free](https://www.pinglabz.com/signup/) ## What MPLS Solves Two problems drove MPLS in the late 1990s: 1. **IP forwarding was slow.** Routers had to perform a longest-prefix-match lookup against the full routing table for every packet. CPUs of the era struggled at gigabit speeds. MPLS replaced this with a fixed-length label lookup, which is dramatically faster. 2. **Traffic engineering was impossible with pure IP.** IP routing picks the shortest path; you cannot easily say "send this customer's traffic over the western backbone, this other customer's traffic over the eastern backbone." MPLS enabled explicit path control. The hardware-acceleration story for IP eventually caught up - modern routers do longest-prefix-match in silicon at terabit speeds - but by then MPLS had become the substrate for an even more important capability: VPNs. MPLS L3VPN lets a service provider carry many customers' overlapping IP address spaces over one shared backbone, with each customer seeing only their own routes and addresses. That capability is what made MPLS a service-provider standard. By 2026 the original speed argument is moot, traffic engineering has spawned segment routing as a successor, and SD-WAN has eaten into MPLS's enterprise WAN role. But MPLS still runs underneath most carrier networks and most of the BGP-VPN deployments enterprises buy from carriers. ## How MPLS Works (the 10,000-Foot View) An MPLS network has three router roles: Customer Edge AcronymCE Job The customer's router; runs IP, no MPLS Provider Edge AcronymPE Job The MPLS network's edge; pushes/pops labels; runs IP-facing-customer and MPLS-facing-core Provider Core AcronymP Job Core MPLS router; only label-switches; never sees customer IP A packet's journey: 1. The customer's CE router sends an IP packet to the PE router. 2. The PE looks up the destination, decides which Label Switched Path (LSP) to use, pushes one or more labels onto the packet, and forwards. 3. Each P router along the path looks at the outermost label, swaps it for the next hop's expected label (label switching), and forwards. 4. The egress PE pops the label(s) and forwards the original packet (or the inner packet of an L2VPN) to the destination CE. This is fundamentally different from IP forwarding. The P routers do not look at IP at all - they only see the MPLS label and forward based on that. The label-switched path is determined when the LSP is set up, not per-packet. This is what enables traffic engineering, and what makes the same backbone usable for L3VPNs, L2VPNs, traffic engineering, and IP transit simultaneously. None of this label switching works without an IGP already in place - most commonly [OSPF](https://www.pinglabz.com/ospf/) \- providing the loopback-to-loopback reachability that LDP builds on. ## MPLS Labels and the Label Stack An MPLS label is a 32-bit shim header inserted between the data-link layer header (e.g. Ethernet) and the IP header. The format: ``` +-----------------+-----+---+-------------+ | Label | EXP | S | TTL | | 20 bits | 3 | 1 | 8 bits | +-----------------+-----+---+-------------+ ``` Label Bits20 Purpose The label value (0-1048575); locally significant per LSR EXP / Traffic Class Bits3 Purpose QoS priority; equivalent to DSCP top 3 bits S (Bottom of Stack) Bits1 Purpose 1 = bottom label; 0 = more labels follow TTL Bits8 Purpose Hop count, decremented at each LSR Multiple labels can be stacked. A typical L3VPN packet has two labels: an outer "transport" label that gets it across the MPLS core, and an inner "VPN" label that identifies which customer VPN it belongs to. The S bit on the inner label is 1; the outer label has S=0. The lab P router (the one in the middle of the core) shows the label-swap that drives every MPLS diagram. Local label in, outgoing label out, prefix that label maps to: ``` P#show mpls forwarding-table Local Outgoing Prefix Bytes Label Outgoing Next Hop Label Label or Tunnel Id Switched interface 16 Pop Label 10.255.0.1/32 1548 Et0/0 10.30.30.1 17 Pop Label 10.255.0.2/32 1494 Et0/1 10.30.31.2 ``` P assigns label 16 for PE1's loopback (10.255.0.1/32) and label 17 for PE2's loopback (10.255.0.2/32). When traffic carrying one of those labels arrives, P pops the transport label and forwards the inner packet on - **penultimate-hop popping (PHP)**. PHP saves the egress PE one label lookup; the imp-null binding that triggers it is the default behaviour for the protocol's owner of a connected prefix or loopback. "Bytes Switched" is the running counter of bytes forwarded against each label entry. After the lab traceroute (below) you will see this counter incrementing. For the full byte-level walkthrough, see [MPLS Labels Explained](https://www.pinglabz.com/mpls-labels-explained/). ## LDP: Label Distribution Protocol Labels do not appear by magic. Some protocol must distribute them so each router knows which label to use for which destination. Three main label distribution protocols exist: LDP (Label Distribution Protocol) Use for IP-driven label assignment for unicast Status in 2026Dominant for IP/MPLS RSVP-TE Use for Traffic-engineered LSPs with bandwidth reservation Status in 2026 Dominant for TE; declining as Segment Routing takes over BGP-LU (BGP Labeled Unicast) Use for Inter-AS label distribution Status in 2026 Common in service provider Option B/C designs Segment Routing Use for Source routing without per-LSP signaling Status in 2026 Rising; replaces LDP and RSVP-TE in modern deployments LDP is the workhorse. Every PE and P router runs LDP, builds a session with each neighbor, and exchanges label mappings: "for prefix X, I will use label Y." The forwarding state derives from the IP routing table - LDP simply assigns labels for each prefix in the IGP and shares those mappings. Enabling LDP on Cisco IOS XE is two interface-level lines plus a global declaration of the protocol: ``` ! global mpls label protocol ldp ! ! per core-facing interface interface Ethernet0/1 ip address 10.30.30.1 255.255.255.252 ip ospf 1 area 0 mpls ip ``` Run that on every P and PE core-facing interface. LDP discovers neighbors using hello messages on UDP 646, then opens a TCP session on the same port to exchange label mappings. The session is established between the highest IP on the router and the neighbor (often the loopback in production), not necessarily the IP on the link. The lab P router has two LDP neighbors (PE1 and PE2). Both sessions are Oper, with the peer's loopback as the LDP Ident: ``` P#show mpls ldp neighbor Peer LDP Ident: 10.255.0.1:0; Local LDP Ident 10.255.0.5:0 TCP connection: 10.255.0.1.646 - 10.255.0.5.37099 State: Oper; Msgs sent/rcvd: 12/12; Downstream Up time: 00:04:04 LDP discovery sources: Ethernet0/0, Src IP addr: 10.30.30.1 Addresses bound to peer LDP Ident: 10.30.30.1 10.255.0.1 Peer LDP Ident: 10.255.0.2:0; Local LDP Ident 10.255.0.5:0 TCP connection: 10.255.0.2.646 - 10.255.0.5.33181 State: Oper; Msgs sent/rcvd: 12/12; Downstream Up time: 00:03:59 LDP discovery sources: Ethernet0/1, Src IP addr: 10.30.31.2 Addresses bound to peer LDP Ident: 10.30.31.2 10.255.0.2 ``` Each Peer LDP Ident is a router-id:label-space-id pair (the label space is almost always 0 for unicast). "Downstream" is the label-distribution method - the downstream router advertises labels for prefixes it is the next hop for. "Addresses bound" is the set of interface IPs the peer announced in its LDP address message, used to map next-hop IPs back to the right LSP. The Label Information Base (LIB) on PE1 shows the labels PE1 chose locally and the labels its neighbour (P) is willing to receive: ``` PE1#show mpls ldp bindings lib entry: 10.30.30.0/30, rev 4 local binding: label: imp-null remote binding: lsr: 10.255.0.5:0, label: imp-null lib entry: 10.30.31.0/30, rev 8 local binding: label: 17 remote binding: lsr: 10.255.0.5:0, label: imp-null lib entry: 10.255.0.1/32, rev 2 local binding: label: imp-null remote binding: lsr: 10.255.0.5:0, label: 16 lib entry: 10.255.0.2/32, rev 10 local binding: label: 18 remote binding: lsr: 10.255.0.5:0, label: 17 lib entry: 10.255.0.5/32, rev 6 local binding: label: 16 remote binding: lsr: 10.255.0.5:0, label: imp-null ``` Read it as: "for prefix X, I (PE1) will use label A locally; my peer (P) tells me to send packets to X with label B." The `imp-null` entries are PHP signals - "do not bother imposing my label on packets for this prefix; pop the previous label and send native IP." For the loopback of P (10.255.0.5/32), P is the owner, so P says imp-null and PE1 will pop the label before sending. For LDP fundamentals, configuration, and verification, see [LDP and MPLS Label Distribution](https://www.pinglabz.com/ldp-mpls-label-distribution/). ## MPLS L3VPN: The Service That Made MPLS Successful The L3VPN (Layer 3 VPN) model is what every enterprise customer of "MPLS service" actually buys. The carrier runs an MPLS backbone and offers each customer a private routing instance (VRF) with private label space. Customer routes never mix; customer A and customer B can both use 10.0.0.0/24 without conflict. The VRF construct itself does not require MPLS at all: [VRF-Lite on Cisco IOS XE](https://www.pinglabz.com/vrf-lite-configuration-cisco-ios-xe/) builds the same isolated routing tables standalone (with proven isolation and static route leaking), and is the best on-ramp to everything in this section. L3VPN adds the machinery that makes VRFs scale across a backbone. The mechanism (RFC 4364): 1. Each customer gets a VRF on the PE. 2. The customer's IPv4 routes are converted to VPNv4 routes by prepending an 8-byte Route Distinguisher (RD) - the same prefix gets a globally unique 12-byte VPNv4 representation. 3. VPNv4 routes are exchanged between PEs via MP-BGP. Route Targets (RTs) attached as extended communities determine which VRFs import which routes. 4. Each VPN route gets a per-VPN label. Two-label stack on the wire: outer transport label (LDP) plus inner VPN label. 5. The egress PE pops both labels and forwards the original IPv4 packet into the customer's VRF. Each of those five steps has its own failure mode and its own set of commands. The full build, in order, with real output at every stage, is in [MPLS L3VPN configuration step by step: VRF, RD, RT, and MP-BGP](https://www.pinglabz.com/mpls-l3vpn-configuration-step-by-step/). The two values that cause the most confusion get their own article: [route distinguisher vs route target](https://www.pinglabz.com/mpls-rd-vs-rt-explained/), proven with the same prefix appearing twice in one BGP table. And the choice of how the PE learns customer routes in the first place - static, OSPF, EIGRP, or eBGP - is compared side by side in [PE-CE routing protocols in MPLS L3VPN](https://www.pinglabz.com/mpls-l3vpn-pe-ce-routing/). The PE configuration is the heart of MPLS-VPN. Here is the working PE1 config from the lab - VRF definition with RD and import/export Route Targets, the customer-facing interface in the VRF, the core-facing interface in the global table with LDP, OSPF as the IGP for loopback reachability, and MP-iBGP to the other PE for VPNv4: ``` ! IGP for the core (loopback reachability) router ospf 1 router-id 10.255.0.1 network 0.0.0.0 255.255.255.255 area 0 ! ! Core-facing interface: IP + IGP + LDP interface Ethernet0/1 ip address 10.30.30.1 255.255.255.252 ip ospf 1 area 0 mpls ip ! ! Customer VRF definition vrf definition CUSTOMER rd 65000:1 route-target export 65000:1 route-target import 65000:1 address-family ipv4 exit-address-family ! ! Customer-facing interface lives in the VRF interface Ethernet0/0 vrf forwarding CUSTOMER ip address 10.40.10.1 255.255.255.252 ! ! MP-iBGP to PE2 for VPNv4 routes router bgp 65000 bgp router-id 10.255.0.1 no bgp default ipv4-unicast neighbor 10.255.0.2 remote-as 65000 neighbor 10.255.0.2 update-source Loopback0 ! address-family vpnv4 neighbor 10.255.0.2 activate neighbor 10.255.0.2 send-community both exit-address-family ! address-family ipv4 vrf CUSTOMER redistribute connected exit-address-family ``` The `no bgp default ipv4-unicast` stops the BGP session from negotiating the regular IPv4 unicast address-family - we only want VPNv4 between the PEs. `send-community both` ensures the extended communities (the Route Targets) get propagated, which is how the receiving PE knows which VRFs to import the prefix into. After OSPF, LDP, and MP-iBGP converge, the VRF shows the local interface and the matched RD: ``` PE1#show ip vrf Name Default RD Interfaces CUSTOMER 65000:1 Et0/0 PE1#show ip vrf detail CUSTOMER VRF CUSTOMER (VRF Id = 1); default RD 65000:1 Interfaces: Et0/0 Address family ipv4 unicast (Table ID = 0x1): Export VPN route-target communities RT:65000:1 Import VPN route-target communities RT:65000:1 VRF label allocation mode: per-prefix ``` The MP-iBGP VPNv4 session to PE2 is up and exchanging routes. One prefix received: the connected /30 PE2 redistributed from its own CUSTOMER VRF. ``` PE1#show ip bgp vpnv4 all summary BGP router identifier 10.255.0.1, local AS number 65000 BGP table version is 4, main routing table version 4 2 network entries using 528 bytes of memory 2 path entries using 272 bytes of memory Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 10.255.0.2 4 65000 5 4 4 0 0 00:01:37 1 ``` The CUSTOMER routing table on PE1 now has the imported PE2 prefix, marked `B` for BGP-learned, with the next-hop being PE2's loopback (10.255.0.2). The next-hop is not the PE-CE link - it is the egress PE's loopback. MPLS resolves the egress PE loopback via the IGP and LDP, which is how the packet finds its way across the core. ``` PE1#show ip route vrf CUSTOMER Routing Table: CUSTOMER Gateway of last resort is not set 10.0.0.0/8 is variably subnetted, 3 subnets, 2 masks C 10.40.10.0/30 is directly connected, Ethernet0/0 L 10.40.10.1/32 is directly connected, Ethernet0/0 B 10.40.20.0/30 [200/0] via 10.255.0.2, 00:00:42 ``` The BGP entry shows the VPN label PE2 advertised for that prefix. The local label is `nolabel` (PE1 is not the egress for this prefix); the out label is 19 - the label PE2 expects on packets arriving in the CUSTOMER VRF for 10.40.20.0/30: ``` PE1#show ip bgp vpnv4 vrf CUSTOMER 10.40.20.0/30 BGP routing table entry for 65000:1:10.40.20.0/30, version 4 Paths: (1 available, best #1, table CUSTOMER) Local 10.255.0.2 (metric 21) (via default) from 10.255.0.2 (10.255.0.2) Origin incomplete, metric 0, localpref 100, valid, internal, best Extended Community: RT:65000:1 mpls labels in/out nolabel/19 ``` The credibility moment: a traceroute from PE1 into the VRF surfaces both labels in the stack. Hop 1 (P) returns a TTL-exceeded that carries the original label stack as a header copy, so the labels show up in plain text in the IOS traceroute output: ``` PE1#traceroute vrf CUSTOMER 10.40.20.2 Type escape sequence to abort. Tracing the route to 10.40.20.2 VRF info: (vrf in name/id, vrf out name/id) 1 10.30.30.2 [MPLS: Labels 17/19 Exp 0] 2 msec 2 msec 2 msec 2 10.40.20.1 2 msec 3 msec 2 msec 3 10.40.20.2 4 msec * 4 msec ``` Hop 1 is P. The packet arrived at P with the label stack **17/19**: outer label 17 is what P expects on packets bound for PE2's loopback (the transport label, learned via LDP); inner label 19 is what PE2 expects on packets for 10.40.20.0/30 in the CUSTOMER VRF (the VPN label, learned via MP-iBGP). Exp 0 is the EXP/Traffic Class bits, zero because we did not mark anything. Hop 2 is PE2 inside the VRF; PE2 popped both labels and forwarded native IP into the customer-facing interface. Hop 3 is CE2. From the customer's side, a CE1 traceroute to CE2's interface address sees the same label stack. The MPLS layer is not hidden from the customer - on this Cisco platform, by default, the egress PE writes the label stack into the ICMP TTL-exceeded reply, and the customer sees the carrier's labels: ``` CE1#traceroute 10.40.20.2 1 10.40.10.1 2 msec 2 msec 1 msec 2 10.30.30.2 [MPLS: Labels 17/19 Exp 0] 4 msec 3 msec 3 msec 3 10.40.20.1 3 msec 3 msec 3 msec 4 10.40.20.2 4 msec * 6 msec ``` To hide the labels from the customer, carriers usually run `no mpls ip propagate-ttl forwarded` on the PEs, which makes the customer traceroute show the carrier as a single hop and the labels disappear from view. The PingLabz lab leaves the default in place so the labels are visible for teaching. End-to-end reachability is the proof that the whole stack works: ``` CE1#ping 10.40.20.2 source 10.40.10.2 repeat 5 Type escape sequence to abort. Sending 5, 100-byte ICMP Echos to 10.40.20.2, timeout is 2 seconds: Packet sent with a source address of 10.40.10.2 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 3/3/5 ms ``` This is where the BGP cluster connects directly. MP-BGP carries VPNv4 routes via address-family vpnv4 unicast. Route distinguishers and route targets are MPLS-VPN concepts that live in BGP. See the [MP-BGP article](https://www.pinglabz.com/mp-bgp-multiprotocol-bgp/) for the BGP side and [MPLS L3VPN with MP-BGP](https://www.pinglabz.com/mpls-l3vpn/) for the MPLS side. The control plane here is MP-BGP - if BGP is not yet second nature, the [BGP complete guide](https://www.pinglabz.com/bgp/) is the prerequisite read. On the PE-CE edge you will also meet [EIGRP](https://www.pinglabz.com/eigrp/) as a customer-facing routing protocol. ## L2VPN: Carrying Layer 2 Over MPLS L3VPN carries IP routes. L2VPN carries Layer 2 frames over the same MPLS backbone. Two main flavors: VPWS (Virtual Private Wire Service / EoMPLS) TopologyPoint-to-point Use for Replacing leased lines; "pseudowire" for one customer link VPLS (Virtual Private LAN Service) Topology Multipoint (one big virtual LAN) Use for Multi-site Layer 2 connectivity for one customer EVPN over MPLS Topology Multipoint with control-plane MAC learning Use for Modern replacement for VPLS EVPN is rapidly displacing VPLS for new deployments because of its control-plane MAC learning and active-active multihoming features. See the BGP cluster's [MP-BGP article](https://www.pinglabz.com/mp-bgp-multiprotocol-bgp/) for EVPN context. ## Traffic Engineering and Segment Routing RSVP-TE was the original MPLS traffic engineering protocol. Operators would specify constraints (bandwidth, link metrics, explicit paths) and RSVP-TE would signal LSPs through the network meeting those constraints. It worked but the per-LSP state was operationally heavy. Segment Routing (SR) is the modern direction. Instead of signaling LSPs, the source encodes the path as a list of segments (labels) in the packet itself. The MPLS data plane is unchanged - SR uses the same label format - but the control plane is dramatically simpler. No per-LSP state in the network; only segment-to-prefix mappings advertised by the IGP. RSVP-TE Control planePer-LSP signaling State per LSP State on every router along path Complexity High operational overhead Segment Routing (SR-MPLS) Control plane IGP-distributed segments State per LSP None per-LSP; only prefix-to-segment mappings ComplexityMuch simpler SRv6 Control plane IPv6-based segments (no MPLS) State per LSPNone Complexity Simplest; replaces MPLS data plane entirely Modern service-provider deployments are migrating from LDP+RSVP-TE to SR-MPLS, with SRv6 being the longer-term direction for greenfield IPv6-only networks. IPv6 rides this same architecture via 6PE/6VPE over the IPv4 label-switched core - the [IPv6 complete guide](https://www.pinglabz.com/ipv6/) covers the addressing fundamentals underneath. ## MPLS vs SD-WAN The most common framing question for enterprise customers is whether to renew MPLS or migrate to SD-WAN. The honest answer in 2026 is hybrid - run SD-WAN over multiple transports including MPLS for hard-QoS-required flows. See the dedicated [SD-WAN vs MPLS](https://www.pinglabz.com/sd-wan-vs-mpls/) article for the cost analysis and migration patterns. The PingLabz position: SD-WAN is the routing layer, MPLS is one of the transports. ## Troubleshooting MPLS: The Two Failure Families MPLS failures split cleanly into two groups, and knowing which one you are in saves most of the time. **Label-plane failures.** The IGP is converged, loopback pings succeed, and traffic still blackholes, because LDP is not distributing labels. The fingerprint is unmistakable once you have seen it: `show ip cef vrf CUST-A 10.20.20.0` returning `unusable: no label` while `show ip ospf neighbor` sits happily at FULL. Walk it in [Troubleshooting MPLS: LDP neighbors, label bindings, and broken LSPs](https://www.pinglabz.com/troubleshooting-mpls-ldp/). **VPN-plane failures.** The labels are fine and the route itself has gone missing somewhere between the customer's routing table and the far CE. Missing redistribution into MP-BGP, a missing route-target import, or RTs stripped in transit because `send-community extended` was never configured. Each one is silent, and each one has a specific command that names it. See [Troubleshooting MPLS L3VPN: where VPN routes go missing](https://www.pinglabz.com/troubleshooting-mpls-l3vpn/). ## The Full MPLS Cluster, in Reading Order ### Fundamentals 1. [MPLS label format, label stacking, and penultimate hop popping](https://www.pinglabz.com/mpls-labels-explained/) 2. [LDP: how routers agree on which label means what](https://www.pinglabz.com/ldp-mpls-label-distribution/) 3. [Configuring MPLS on Cisco IOS XE](https://www.pinglabz.com/cisco-mpls-configuration/) 4. [MP-BGP address families, including VPNv4](https://www.pinglabz.com/mp-bgp-multiprotocol-bgp/) ### MPLS L3VPN 1. [MPLS L3VPN with MP-BGP and VPNv4: the architecture](https://www.pinglabz.com/mpls-l3vpn/) 2. [Configuring an L3VPN step by step: VRF, RD, RT, and MP-BGP](https://www.pinglabz.com/mpls-l3vpn-configuration-step-by-step/) 3. [Route distinguisher vs route target, proven in one BGP table](https://www.pinglabz.com/mpls-rd-vs-rt-explained/) 4. [PE-CE routing: static, OSPF, EIGRP, and eBGP compared](https://www.pinglabz.com/mpls-l3vpn-pe-ce-routing/) 5. [VRF-Lite: the same VRF construct without MPLS](https://www.pinglabz.com/vrf-lite-configuration-cisco-ios-xe/) ### Troubleshooting 1. [LDP neighbors, label bindings, and broken LSPs](https://www.pinglabz.com/troubleshooting-mpls-ldp/) 2. [Where VPN routes go missing: RTs, redistribution, and import filters](https://www.pinglabz.com/troubleshooting-mpls-l3vpn/) ### Traffic Engineering, Segment Routing, and the SD-WAN Question 1. [MPLS traffic engineering explained](https://www.pinglabz.com/mpls-traffic-engineering/) 2. [Segment routing (SR-MPLS): replacing LDP and RSVP-TE](https://www.pinglabz.com/segment-routing-mpls/) 3. [QoS for VoIP across an MPLS WAN](https://www.pinglabz.com/qos-for-voip/) 4. [MPLS L3VPN vs SD-WAN: when to migrate](https://www.pinglabz.com/mpls-vs-sd-wan-migration/) 5. [SD-WAN vs MPLS: when each wins in 2026](https://www.pinglabz.com/sd-wan-vs-mpls/) Solid on the CCNA foundations? - the PingLabz CCNA Labs library MPLS-VPN is CCNP+. The CCNA prerequisites - OSPF, EIGRP, IPv6, ACLs, basic NAT, GRE - are all covered hands-on in the 60-lab PingLabz CCNA Labs library with real Cisco IOS XE captures. Open the PingLabz CCNA Labs library to refresh the foundations. [Open the CCNA labs library](https://www.pinglabz.com/ccna-labs-network-fundamentals/) ## IPv6 over an IPv4 MPLS core (6PE) You do not have to dual-stack the core to deliver IPv6\. 6PE carries IPv6 prefixes across an untouched IPv4 MPLS core by attaching an MPLS label to each prefix over MP-BGP, using an IPv4-mapped IPv6 next-hop that the receiving PE resolves through the existing LDP label-switched path. The P routers never learn an IPv6 route. Full walkthrough with real labels and an end-to-end IPv6 traceroute: [MP-BGP for IPv6 and 6PE explained](https://www.pinglabz.com/ipv6-bgp-6pe/). ## Expert Transport: 6VPE, PE-CE BGP, and Dual-Hub DMVPN (CCIE level) These articles close the CCIE transport domain. Built on real Cisco IOS XE output from two CML labs - a five-router MPLS L3VPN carrying VPNv4 and VPNv6, and a dual-hub IPsec DMVPN - they cover the label stacks, the BGP PE-CE edge cases, and the failover behaviour that separates a working transport from a resilient one. 1 [MPLS VPNv6 and 6VPE: IPv6 L3VPN over an IPv4 core](https://www.pinglabz.com/mpls-vpnv6-6vpe/) The two-label stack that carries IPv6 L3VPN across an unchanged IPv4 MPLS core, proven in an IPv6 traceroute. 2 [BGP as the PE-CE protocol: as-override, allowas-in, and SoO](https://www.pinglabz.com/bgp-pe-ce-as-override-allowas-in/) Why two customer sites sharing an AS number cannot reach each other, and the three tools that fix it. 3 [Dual-hub DMVPN: redundancy designs that actually fail over](https://www.pinglabz.com/dual-hub-dmvpn-design/) The NHRP holdtime trap that black-holes spoke-to-spoke traffic for hours after a hub failure - and how to fix it. 4 [DMVPN with IKEv1 vs IKEv2: configuration and migration](https://www.pinglabz.com/dmvpn-ikev1-vs-ikev2/) Both crypto suites on the same cloud, real SA output side by side, and the one-line profile-swap migration. 5 [MPLS and DMVPN together: choosing and combining transports](https://www.pinglabz.com/mpls-vs-dmvpn-enterprise-transport/) A purchased SLA-backed service vs a self-built encrypted overlay, the hybrid designs, and where SD-WAN fits. 6 [Expert transport troubleshooting: MPLS and DMVPN ticket scenarios](https://www.pinglabz.com/expert-transport-troubleshooting/) Five faults isolated layer by layer - label stack, AS-path loop, spoke-to-spoke, the NHRP trap, and an IPsec bounce. ## Studying for the CCIE? This cluster is part of the full CCNA to CCNP to CCIE Enterprise ladder on PingLabz, every rung built on real Cisco output. For expert-level depth across every EI v1.1 blueprint domain - and the four integration Super Labs - see the [CCIE Enterprise Infrastructure study hub](https://www.pinglabz.com/ccie-enterprise/). ## Frequently Asked Questions ### What does MPLS stand for? MPLS stands for Multiprotocol Label Switching. The "multiprotocol" part means it can carry IPv4, IPv6, Ethernet frames, ATM cells, or any other Layer 3 protocol. The "label switching" part means it forwards based on labels rather than IP destination addresses. ### What OSI layer is MPLS? Layer 2.5\. It sits between Layer 2 (the data link, e.g. Ethernet) and Layer 3 (IP). The MPLS shim header inserts between the Ethernet header and the IP header. This is sometimes called the "shim layer." ### What is the advantage of MPLS over plain IP? Three big ones: faster forwarding via fixed-length label lookup (mattered more in the 1990s than today), traffic engineering via explicit paths, and VPN services (L3VPN, L2VPN) over a shared backbone. The forwarding speed advantage is largely gone; the VPN and TE capabilities are why MPLS is still dominant in service-provider networks. ### What is the difference between MPLS and MPLS L3VPN? MPLS is the underlying label-switching protocol. MPLS L3VPN is a service built on top of MPLS that uses MP-BGP to distribute customer routes (in VPNv4 format) across the MPLS backbone, allowing many customers' overlapping IP address spaces to coexist on one provider network. When enterprises buy "MPLS" from a carrier, they almost always buy MPLS L3VPN. ### Is SD-WAN replacing MPLS? Partially. SD-WAN replaces MPLS as the enterprise WAN routing layer, but still uses MPLS as one of the transports for hard-QoS-required flows. Most 2026 enterprise deployments are hybrid: SD-WAN over a mix of MPLS, broadband, and LTE. Pure SD-WAN-internet-only is common at smaller branches. ### What is segment routing and is it replacing MPLS? Segment Routing is a modern source-routing approach where the source router specifies the path as a list of segments (labels). SR-MPLS uses the existing MPLS data plane with a simpler control plane. SRv6 uses IPv6 instead of MPLS. SR is replacing LDP and RSVP-TE in many service-provider networks; SRv6 may eventually replace MPLS itself, but that is a long migration. ## Key Takeaways MPLS is the protocol that runs the inside of every major service provider network and most of the BGP-VPN services enterprises buy. Labels (32 bits, stackable), LDP (label distribution), and MP-BGP (for VPN routes) together make L3VPN possible. Segment Routing is the modern direction; LDP and RSVP-TE are gradually being phased out in service-provider networks. If you take one thing away from this guide, make it the two-label stack model for L3VPN: outer transport label gets the packet across the MPLS core, inner VPN label tells the egress PE which customer VRF to deliver to. Master that and the rest of MPLS-VPN follows. Bookmark this page, work through the cluster articles in order, and lab every concept. **Studying for the CCNA?** Test your MPLS knowledge on [PingLabz CCNA Flashcards](https://www.pinglabz.com/ccna-flashcards/) \- 200 free multiple-choice questions by topic, mixed, or a full mock exam, each with a plain-English explanation. ### References - [RFC 3031 - Multiprotocol Label Switching Architecture](https://www.rfc-editor.org/rfc/rfc3031?ref=pinglabz.com) - [Cisco MPLS technology documentation](https://www.cisco.com/c/en/us/tech/multiprotocol-label-switching-mpls/index.html?ref=pinglabz.com) ### FHRP (First Hop Redundancy Protocol): The Complete Guide URL: https://www.pinglabz.com/fhrp/ Last updated: 2026-08-01T19:48:30.000Z FHRP (First Hop Redundancy Protocol) is the family of Layer 3 protocols that lets a group of routers share a virtual IP and MAC address so end hosts have a single, stable default gateway even when the underlying router fails. The three protocols in the family - HSRP (Cisco), VRRP (IETF standard), and GLBP (Cisco) - all solve the same problem with slightly different mechanics. Without FHRP, end host default gateways become single points of failure; with FHRP, the failure of a primary router becomes a sub-second event invisible to users. This is the cluster overview for the full PingLabz FHRP series: the protocol family, HSRP, VRRP, GLBP, and the design patterns that integrate FHRP with STP, Layer 2 VLANs, and the rest of a Cisco campus design. We will work through what FHRP solves, the three protocols in the family and how they differ, the configuration patterns, real Cisco IOS XE captures from a lab, and the alignment with STP that is critical for clean failover. New labs and guides, in your inbox Every new PingLabz lab and deep-dive, built and verified on real Cisco IOS XE - free, straight to your inbox. [Join free](https://www.pinglabz.com/signup/) ## What FHRP Solves End hosts use a default gateway IP for any traffic destined outside their local subnet. Traditionally that gateway is a single router. If the router goes down, the gateway becomes unreachable, and the host has no path off the subnet. The user sees connectivity loss until DHCP renews to a new gateway (which never happens automatically) or someone reconfigures the host. FHRP solves this by letting a group of routers share a virtual IP/MAC. The hosts use the virtual IP as their default gateway. The router currently active for that virtual IP forwards traffic; if it fails, another router takes over the virtual IP/MAC and traffic continues with sub-second interruption. The hosts never know anything changed. The three FHRP protocols differ in details but share the architecture: a group of routers, a virtual IP and MAC, election of an active forwarder, and failover when the active fails. ## The Three Protocols HSRP (v1) VendorCisco RFC2281 (informational) Default Hello3 sec Default Hold10 sec Active election Highest priority; tiebreak by IP HSRP (v2) VendorCisco RFC2281 Default Hello 3 sec (sub-second tunable) Default Hold10 sec Active electionSame as v1 VRRP (v2) VendorIETF standard RFC3768 / 5798 Default Hello1 sec Default Hold3 sec Active election Highest priority; tiebreak by IP VRRPv3 VendorIETF standard RFC5798 Default Hello 1 sec (sub-second tunable) Default Hold3 sec Active election Same as v2 + IPv6 support GLBP VendorCisco RFCNone (proprietary) Default Hello3 sec Default Hold10 sec Active election AVG (Active Virtual Gateway) elects; AVFs (Active Virtual Forwarders) load-balance HSRP and VRRP do active/standby - one router forwards, the others wait. GLBP does active/active - multiple routers forward simultaneously, sharing the load via different virtual MAC addresses for the same virtual IP. ## HSRP: Cisco's Default HSRP (Hot Standby Router Protocol) is Cisco's proprietary FHRP and the dominant protocol in Cisco-only campuses. The version 1 group number is 0-255 (limited); version 2 supports group 0-4095 and is the modern default. HSRP states each router walks through: Initial HSRP just started; not yet sending hellos Learn Waiting to learn the virtual IP from a Hello (rare; typically configured) Listen Heard from active and standby; not active or standby itself Speak Sending hellos; participating in active/standby election Standby Backup; will take over if active fails Active Currently forwarding traffic for the virtual IP Healthy steady state: one active, one standby, others in listen. Configuration is interface-level: a group number, a virtual IP, an optional priority, and preemption so the higher-priority router reclaims active state after a reboot. ``` ! R1 - intended active gateway interface Ethernet0/0 ip address 10.20.0.1 255.255.255.0 standby version 2 standby 10 ip 10.20.0.254 standby 10 priority 110 standby 10 preempt standby 10 name HSRP-PROD ! ! R2 - intended standby (default priority 100, no preempt) interface Ethernet0/0 ip address 10.20.0.2 255.255.255.0 standby version 2 standby 10 ip 10.20.0.254 standby 10 name HSRP-PROD ``` On R1 with the priority + preempt configuration above, HSRP comes up Active. The full state shows the v2 virtual MAC, the standby's identity, hello/hold timers, and the configured priority: ``` R1# show standby Ethernet0/0 - Group 10 (version 2) State is Active 2 state changes, last state change 00:01:06 Virtual IP address is 10.20.0.254 Active virtual MAC address is 0000.0c9f.f00a (MAC In Use) Local virtual MAC address is 0000.0c9f.f00a (v2 default) Hello time 3 sec, hold time 10 sec Next hello sent in 1.040 secs Preemption enabled Active router is local Standby router is 10.20.0.2, priority 100 (expires in 9.504 sec) Priority 110 (configured 110) Group name is "HSRP-PROD" (cfgd) FLAGS: 1/1 ``` Two details worth noting. The Active virtual MAC is `0000.0c9f.f00a`, not the burned-in MAC of R1's Ethernet0/0\. End hosts ARP for 10.20.0.254 and get this MAC, which means R2 can take over without any host needing to re-ARP. And the v2 MAC format is `0000.0c9f.fXXX`, not the v1 `0000.0c07.acXX`; group 10 ends up as `...f00a` because 10 in hex is `0a`. For the full configuration walkthrough including priority manipulation, interface tracking, and authentication, see [HSRP High Availability: Configure Cisco HSRP Step-by-Step](https://www.pinglabz.com/implementing-hsrp-for-high-availability-a-complete-guide-for-network-engineers/). ## VRRP: The Open Standard VRRP (Virtual Router Redundancy Protocol) is the IETF standard. [RFC 5798](https://www.rfc-editor.org/rfc/rfc5798?ref=pinglabz.com) defines VRRPv3 for both IPv4 and IPv6\. Functionally similar to HSRP with a few differences: - Master/Backup terminology (instead of Active/Standby) - Default master is the router whose interface IP matches the virtual IP (a feature that can simplify designs but causes confusion when not understood) - Faster default timers (1-second Hello, 3-second hold) - Vendor-neutral - works between Cisco, Juniper, Arista, Nokia, etc. The classic VRRPv2 configuration on Cisco IOS XE looks almost identical to HSRP at the interface level: ``` ! Global - force classic VRRPv2 syntax (defaults to v3 on IOS XE 17.x) fhrp version vrrp v2 ! ! R1 interface Ethernet0/0 vrrp 20 ip 10.20.0.253 vrrp 20 priority 110 vrrp 20 description VRRP-PROD ! ! R2 interface Ethernet0/0 vrrp 20 ip 10.20.0.253 vrrp 20 priority 100 vrrp 20 description VRRP-PROD ``` R1 wins the Master election and starts answering ARP for the virtual IP with the VRRP-style virtual MAC: ``` R1# show vrrp Ethernet0/0 - Group 20 VRRP-PROD State is Master Virtual IP address is 10.20.0.253 Virtual MAC address is 0000.5e00.0114 Advertisement interval is 1.000 sec Preemption enabled Priority is 110 Master Router is 10.20.0.1 (local), priority is 110 Master Advertisement interval is 1.000 sec Master Down interval is 3.570 sec FLAGS: 1/1 ``` Notice the virtual MAC format: `0000.5e00.0114`. The `0000.5e00.01XX` prefix is reserved by IANA for VRRP, and `14` is the hex of the group number (20). The 1-second advertisement interval and 3-second Master Down interval also show through - VRRP fails over faster than HSRP by default. For the full walkthrough including VRRPv3 syntax and IPv6, see [VRRP Explained](https://www.pinglabz.com/vrrp-explained/). ## GLBP: Active/Active Load Balancing GLBP (Gateway Load Balancing Protocol) is Cisco-proprietary and unique among FHRPs in that it actively load-balances across multiple routers. One router is the AVG (Active Virtual Gateway) which manages the protocol; up to four AVFs (Active Virtual Forwarders) actually forward traffic, each with its own virtual MAC. How it works: end hosts ARP for the virtual IP. The AVG responds with one of four virtual MACs, rotating across hosts. Each virtual MAC is owned by a different AVF, so different hosts naturally use different routers. Failover is per-AVF; if one AVF dies, the AVG redirects its virtual MAC to another AVF. The benefit: utilization across redundant routers instead of one router idle as standby. The cost: more complex troubleshooting and Cisco-only. Configuration is similar to HSRP but with the `glbp` keyword: ``` ! R1 - intended AVG interface Ethernet0/0 glbp 30 ip 10.20.0.252 glbp 30 priority 110 glbp 30 preempt glbp 30 name GLBP-PROD ! ! R2 - secondary AVF (default priority, no preempt) interface Ethernet0/0 glbp 30 ip 10.20.0.252 glbp 30 name GLBP-PROD ``` What separates GLBP from HSRP and VRRP shows up clearly in `show glbp brief`. Where the other two protocols show one row per group, GLBP shows the AVG row plus one row per forwarder: ``` R1# show glbp brief Interface Grp Fwd Pri State Address Active router Standby router Et0/0 30 - 110 Active 10.20.0.252 local 10.20.0.2 Et0/0 30 1 - Active 0007.b400.1e01 local - Et0/0 30 2 - Listen 0007.b400.1e02 10.20.0.2 - ``` R1 is both the AVG (top row) and the AVF for Forwarder 1 (MAC `0007.b400.1e01`). Forwarder 2 (`0007.b400.1e02`) is owned by R2, which is why R1 shows it as Listen. From R2's perspective the same group looks inverted: ``` R2# show glbp brief Interface Grp Fwd Pri State Address Active router Standby router Et0/0 30 - 100 Standby 10.20.0.252 10.20.0.1 local Et0/0 30 1 - Listen 0007.b400.1e01 10.20.0.1 - Et0/0 30 2 - Active 0007.b400.1e02 local - ``` R2 is the AVG Standby (it will take over the AVG role if R1 dies), but R2 is simultaneously the active AVF for Forwarder 2\. Both routers are forwarding traffic at the same time, just for different hosts. The GLBP MAC format is `0007.b400.XXYY` where `XX` is the group in hex (`1e` \= 30) and `YY` is the forwarder number. For the full walkthrough including weighting, load-balancing algorithms, and AVF priority, see [GLBP for Active/Active Load Balancing](https://www.pinglabz.com/glbp-load-balancing/). ## FHRP Design with STP and VLANs FHRP does not exist in isolation. In typical campus designs the active FHRP gateway should align with the STP root bridge, otherwise traffic from access switches takes a suboptimal path through the network: up to one distribution switch, across the inter-distribution trunk to the other, down to the actual gateway. The pattern: per-VLAN, set the same router to be both the STP root bridge and the FHRP active gateway. For VLAN load balancing, alternate which distribution switch is root and active for each VLAN. With Rapid PVST+ this is straightforward; with MST it requires alignment of MST instances to FHRP groups. Detail in [Spanning Tree and First-Hop Redundancy: Aligning STP with HSRP/VRRP](https://www.pinglabz.com/stp-fhrp-hsrp-vrrp-alignment/) in the STP cluster. The two layers underneath this design each have their own complete guide: [VLANs and Layer 2 switching](https://www.pinglabz.com/vlans-layer-2-switching/) for the broadcast-domain layout, and [spanning tree](https://www.pinglabz.com/spanning-tree-protocol/) for which links actually forward. ## Verification: Proving the Virtual MAC The single most important thing to internalize about FHRP is that the virtual IP is an L2 abstraction, not an L3 route. The active router answers ARP for the virtual IP with a virtual MAC; it does not install the virtual IP as a /32 in the routing table. The `show ip route` output on R1 (the HSRP/VRRP/GLBP active for all three groups) makes this obvious: ``` R1# show ip route 10.0.0.0/8 is variably subnetted, 3 subnets, 2 masks C 10.20.0.0/24 is directly connected, Ethernet0/0 L 10.20.0.1/32 is directly connected, Ethernet0/0 C 10.255.0.1/32 is directly connected, Loopback0 ``` None of `10.20.0.252`, `10.20.0.253`, or `10.20.0.254` appear. The router knows the subnet they live in but does not own them as routed addresses. They are pure ARP destinations. From the perspective of any host on the LAN that has ARPed for all three virtual IPs, the three protocols announce three different virtual MAC ranges. R2 in this lab has pinged each virtual IP once, so its ARP table holds the full picture: ``` R2# show ip arp Protocol Address Age (min) Hardware Addr Type Interface Internet 10.20.0.1 3 aabb.cc00.0b00 ARPA Ethernet0/0 Internet 10.20.0.2 - aabb.cc00.0c00 ARPA Ethernet0/0 Internet 10.20.0.252 - 0007.b400.1e02 ARPA Ethernet0/0 Internet 10.20.0.253 0 0000.5e00.0114 ARPA Ethernet0/0 Internet 10.20.0.254 0 0000.0c9f.f00a ARPA Ethernet0/0 ``` Three virtual IPs, three different vendor prefixes. `0000.0c9f.fXXX` for HSRPv2 (Cisco OUI `0000.0c` meets the HSRP-specific `9f.f` body). `0000.5e00.01XX` for VRRP (an IANA-assigned range). `0007.b400.XXYY` for GLBP (Cisco again, but a different sub-range, with the forwarder number baked into the last byte). None of these MAC addresses belongs to a physical interface on either router. When R1 fails, R2 starts answering ARP for the same virtual IPs with the same virtual MACs, and traffic continues. End hosts never re-ARP. This is also why FHRPs work over arbitrary L2 topologies (a single switch, a stack, a VSS, a vPC, or a long L2 stretch) without help from the routing protocol. The protocol that makes the world resilient at the first hop is, mechanically, a small ARP trick. ## HSRP vs VRRP vs GLBP Vendor HSRPCisco VRRPOpen standard GLBPCisco Active routers HSRP1 (others standby) VRRP1 (others backup) GLBPUp to 4 (AVG + 3 AVFs) Load balancing HSRP Per-VLAN (different active per VLAN) VRRPPer-VLAN GLBPPer-host (within VLAN) Default hello HSRP3 seconds VRRP1 second GLBP3 seconds Multicast address HSRP 224.0.0.2 (v1) / 224.0.0.102 (v2) VRRP224.0.0.18 GLBP224.0.0.102 Virtual MAC range HSRP 0000.0c07.acXX (v1) / 0000.0c9f.fXXX (v2) VRRP0000.5e00.01XX GLBP0007.b400.XXYY Authentication HSRPMD5, plain text VRRP None (v3); plain text/MD5 (v2) GLBPMD5, plain text Use case HSRP Cisco-only campus default VRRP Multi-vendor environments GLBP Cisco shops wanting load balancing without per-VLAN role assignment The three brief summaries side by side on R1 make the differences land. Same interface, three different protocols, three different group numbers: ``` R1# show standby brief Interface Grp Pri P State Active Standby Virtual IP Et0/0 10 110 P Active local 10.20.0.2 10.20.0.254 R1# show vrrp brief Interface Grp Pri Time Own Pre State Master addr Group addr Et0/0 20 110 3570 Y Master 10.20.0.1 10.20.0.253 R1# show glbp brief Interface Grp Fwd Pri State Address Active router Standby router Et0/0 30 - 110 Active 10.20.0.252 local 10.20.0.2 Et0/0 30 1 - Active 0007.b400.1e01 local - Et0/0 30 2 - Listen 0007.b400.1e02 10.20.0.2 - ``` HSRP and VRRP show one row per group (one active forwarder). GLBP shows three: the AVG row plus one row per AVF. That extra structure is the entire point of GLBP - traffic from a single LAN can be load-balanced across two routers without splitting VLANs. For the full comparison see [HSRP vs VRRP vs GLBP](https://www.pinglabz.com/hsrp-vs-vrrp-vs-glbp/). ## High Availability Beyond the First Hop FHRP protects the default gateway; the platform itself has its own high-availability story. [SSO, NSF, and graceful restart](https://www.pinglabz.com/sso-nsf-graceful-restart/) keep traffic flowing through a device when a supervisor fails or software is upgraded - SSO preserves the forwarding table, NSF keeps forwarding while the control plane restarts, and graceful restart stops neighbors from rerouting. Together with FHRP and [BFD](https://www.pinglabz.com/bfd-bidirectional-forwarding-detection/), they form the resilience layer of an [enterprise campus design](https://www.pinglabz.com/enterprise-campus-design-ccnp/). ## FHRP Deep Dives in This Cluster 1. [HSRP High Availability: Configure Cisco HSRP Step-by-Step](https://www.pinglabz.com/implementing-hsrp-for-high-availability-a-complete-guide-for-network-engineers/) 2. [VRRP Explained: The Vendor-Neutral FHRP](https://www.pinglabz.com/vrrp-explained/) 3. [GLBP for Active/Active Load Balancing](https://www.pinglabz.com/glbp-load-balancing/) 4. [HSRP vs VRRP vs GLBP: Choosing the Right FHRP](https://www.pinglabz.com/hsrp-vs-vrrp-vs-glbp/) 5. [Spanning Tree and First-Hop Redundancy: Aligning STP with HSRP/VRRP](https://www.pinglabz.com/stp-fhrp-hsrp-vrrp-alignment/) ### More FHRP guides in this cluster 1\. [HSRP States: The 6-State Machine with show standby Output](https://www.pinglabz.com/hsrp-state/) ## Tracking: The Failure FHRP Cannot See On Its Own An FHRP group protects against the gateway *dying*. It does nothing about the gateway that is alive, still winning the election, and has lost its uplink. Hosts keep ARPing for the virtual IP, the active router keeps answering, and every packet goes into a black hole. The group is healthy and the site is down. The fix is object tracking: attach a tracked object to the group and decrement priority when the uplink fails, so the standby takes over as active gateway. `standby 1 track 1 decrement 30` is the whole idea, and the tracked object can watch an interface, a route, or (much better) an [IP SLA probe](https://www.pinglabz.com/ip-sla-cisco-ios-xe/) that actually tests the path rather than the link. The tracking mechanism itself, including the failover captured live and the recursion trap that makes routers oscillate, is covered in [Object tracking and reliable static routes with IP SLA](https://www.pinglabz.com/object-tracking-reliable-static-routes/). The same tracked object drives HSRP priority, a floating static, or a PBR next-hop. ### Running and breaking a first-hop gateway [Both routers think they are Active](https://www.pinglabz.com/hsrp-troubleshooting-flapping-dual-active/) Dual-active means the hellos are not crossing. Captured side by side, with the ACL removed and the FSM transitioning back. [Choosing between the three protocols](https://www.pinglabz.com/hsrp-vs-vrrp-vs-glbp/) Every claim backed by captured output: the virtual MAC each one uses, VRRP's measured timer advantage, and GLBP genuinely forwarding on both routers. ## FAQ ### What does FHRP stand for? FHRP stands for First Hop Redundancy Protocol. It refers to the family of protocols (HSRP, VRRP, GLBP) that provide gateway redundancy for end hosts. ### What is the difference between HSRP and VRRP? Functionally similar. HSRP is Cisco-proprietary; VRRP is the IETF standard. HSRP has slightly slower default timers (3-second Hello vs 1-second); VRRP has slightly faster failover. In a Cisco-only environment, HSRP is the dominant choice; in mixed-vendor, VRRP is required. ### What makes GLBP different from HSRP and VRRP? HSRP and VRRP have one active router per group, with others on standby. GLBP has up to four active forwarders simultaneously, each with its own virtual MAC. ARP responses from the AVG distribute hosts across the AVFs, providing per-host load balancing without per-VLAN role tuning. ### What is HSRP priority and how does it work? HSRP priority is a 0-255 value (default 100). Higher priority wins the active election. With preemption enabled, a higher-priority router that comes up will take over from a lower-priority active. Track interface state to lower priority when uplinks fail, so the standby takes over. ### What is the virtual MAC for HSRP? HSRPv1 uses `0000.0c07.acXX` where `XX` is the group number in hex. HSRPv2 uses `0000.0c9f.fXXX` where `XXX` is the 12-bit group in hex (since v2 supports groups 0-4095, which need 3 hex digits). End hosts ARP for the virtual IP and receive this MAC; they use it as the destination for all upstream traffic. Group 10 under v2, for example, resolves to `0000.0c9f.f00a`. Hands-on FHRP - HSRP, VRRP, and GLBP labs Configure all three first-hop redundancy protocols on real Cisco IOS XE routers. Virtual MAC formats, priority manipulation, preemption, and GLBP active-active load balancing. Part of the 14-lab CCNA IP Connectivity cluster. Open the PingLabz CCNA Labs library. [Open the FHRP labs](https://www.pinglabz.com/ccna-labs-ip-connectivity/) ## Key Takeaways FHRP is mandatory for any production network where users care about gateway uptime. Pick HSRP for Cisco-only environments, VRRP for multi-vendor, GLBP when you want active/active load balancing without per-VLAN tuning. Whichever you pick, align the active gateway with the STP root bridge per VLAN to avoid suboptimal traffic paths, and remember that the virtual IP is an ARP trick rather than a routing-table entry - which is why FHRP failover is sub-second and host-invisible. Bookmark this page, work through the cluster articles in order, and lab every failover. **Studying for the CCNA?** Test your FHRP knowledge on [PingLabz CCNA Flashcards](https://www.pinglabz.com/ccna-flashcards/) \- 200 free multiple-choice questions by topic, mixed, or a full mock exam, each with a plain-English explanation. ### References - [RFC 5798 - Virtual Router Redundancy Protocol (VRRP) Version 3](https://www.rfc-editor.org/rfc/rfc5798?ref=pinglabz.com) - [Cisco HSRP technology documentation](https://www.cisco.com/c/en/us/tech/ip/hot-standby-router-protocol-hsrp/index.html?ref=pinglabz.com) ### IPv6: The Complete Guide for Network Engineers URL: https://www.pinglabz.com/ipv6/ Last updated: 2026-08-01T19:48:29.000Z IPv6 (Internet Protocol version 6) is the successor to IPv4 - a 128-bit address space (vs IPv4's 32-bit), a streamlined header, native autoconfiguration, and the protocol the internet has been "about to migrate to" for thirty years. The migration is real now: most major content providers, mobile networks, and cloud platforms support IPv6, and an increasing percentage of internet traffic is genuinely v6\. For network engineers in 2026, IPv6 is no longer optional knowledge. This is the cluster overview for the full PingLabz IPv6 series: the 128-bit address format, address types (link-local, global, unique-local), the simplified IPv6 header, ICMPv6 and Neighbor Discovery (the ARP replacement), SLAAC autoconfiguration, IPv6 routing protocols (OSPFv3, BGP for IPv6, EIGRP for IPv6, VRRPv3), and Cisco IOS XE configuration patterns. We will work through what IPv6 is, how it differs from IPv4, why the address format looks the way it does, and the configuration patterns that work in production. Every capture on this page comes from the dual-stack OSPFv2+OSPFv3 PingLabz reference lab - real IOS XE 17.16 output, not synthetic examples. New labs and guides, in your inbox Every new PingLabz lab and deep-dive, built and verified on real Cisco IOS XE - free, straight to your inbox. [Join free](https://www.pinglabz.com/signup/) ## Why IPv6 Exists IPv4's 32-bit address space provides 4.3 billion unique addresses. That sounded enormous in 1981; it is utterly inadequate now. The IANA exhausted its IPv4 pool in 2011; regional internet registries (ARIN, RIPE, APNIC, LACNIC, AFRINIC) ran out of fresh allocations between 2015 and 2020\. IPv4 transfers and reclamations now drive any new allocation. IPv6's 128-bit address space provides 340 undecillion (3.4 x 10^38) addresses. Enough that every grain of sand on Earth could have its own subnet. The address shortage is solved permanently. But IPv6 is not just bigger addresses. The protocol designers used the redesign as an opportunity to fix several IPv4 design issues: - **Streamlined header.** Fewer fields; fixed length (40 bytes vs IPv4's variable 20-60 bytes). Routers process IPv6 packets faster. - **No header checksum.** Layer 2 (Ethernet) and Layer 4 (TCP/UDP) already have checksums; the IPv6 header drops it to save processing. - **No fragmentation by routers.** Only the source can fragment; routers signal too-big errors via ICMPv6 PTB (Packet Too Big). - **Native autoconfiguration.** SLAAC (Stateless Address Autoconfiguration) lets hosts assign their own IPv6 address without DHCP. - **Mandatory IPsec support.** Originally required (now recommended); every IPv6 stack includes IPsec primitives. - **Built-in mobility, multicast, and extensibility.** Mobile IPv6, MLD multicast, extension headers for new features without changing the base header. The result is a protocol that scales to internet+IoT addresses and forwards faster than IPv4 on equivalent hardware. ## The IPv6 Address Format An IPv6 address is 128 bits, written as eight groups of four hexadecimal digits separated by colons: ``` 2001:0db8:85a3:0000:0000:8a2e:0370:7334 ``` Two compression rules make addresses readable: 1. **Leading zeros in each group can be dropped.** 0db8 becomes db8; 0000 becomes 0. 2. **One run of consecutive all-zero groups can be replaced with ::.** Only one :: per address (otherwise ambiguous). So the address above compresses to `2001:db8:85a3::8a2e:370:7334`. The /64 prefix length is the convention for hosts on a network. The first 64 bits identify the network; the last 64 bits identify the host. This 64/64 split is what enables SLAAC and Neighbor Discovery efficiencies. The real `show ipv6 interface` output makes this concrete. On the PingLabz OSPF Reference Lab, R1's Ethernet0/0 has both a link-local FE80::/10 address and an EUI-64-derived global address from the configured /64 prefix: ``` R1#show ipv6 interface Ethernet0/0 Ethernet0/0 is up, line protocol is up IPv6 is enabled, link-local address is FE80::A8BB:CCFF:FE00:300 No Virtual link-local address(es): Description: To SW1 (Area 0 broadcast LAN) Global unicast address(es): 2001:DB8:20:0:A8BB:CCFF:FE00:300, subnet is 2001:DB8:20::/64 [EUI] Joined group address(es): FF02::1 FF02::2 FF02::5 FF02::1:FF00:300 MTU is 1500 bytes ND DAD is enabled, number of DAD attempts: 1 ND reachable time is 30000 milliseconds (using 30000) ND router advertisements are sent every 200 seconds ND router advertisements live for 1800 seconds Hosts use stateless autoconfig for addresses. ``` Two things are happening on the same interface. There is a link-local address (FE80::A8BB:CCFF:FE00:300) - mandatory, automatic, never routed off-segment. And there is a global unicast address built from the configured prefix 2001:DB8:20::/64 plus the EUI-64 interface ID A8BB:CCFF:FE00:300, which is derived from R1's MAC address by inserting FF:FE in the middle and flipping the U/L bit. The `[EUI]` tag tells you exactly how the address was generated. The "Joined group address(es)" list shows the multicast groups this interface listens to: `FF02::1` all-nodes, `FF02::2` all-routers, `FF02::5` all-OSPFv3-routers, and `FF02::1:FF00:300` the solicited-node multicast that replaces ARP. Multicast is fundamental to IPv6 - it is how the protocol replaced broadcast. For the byte-level walkthrough including the compression rules and worked examples, see [IPv6 Address Format Explained](https://www.pinglabz.com/ipv6-address-format/). ## IPv6 Address Types 2000::/3 TypeGlobal Unicast ScopeGlobally routable Use Equivalent to IPv4 public addresses FC00::/7 (commonly FD00::/8) TypeUnique Local (ULA) Scope Site-local; not internet-routable Use Like IPv4 RFC 1918 private addresses FE80::/10 TypeLink-Local Scope Link-only; never crosses a router Use Required on every IPv6 interface; used for ND, OSPF, EIGRP FF00::/8 TypeMulticast ScopeGroup communication Use Replaces IPv4 broadcast for many functions ::1/128 TypeLoopback ScopeHost-local Use Equivalent to 127.0.0.1 ::/128 TypeUnspecified ScopeNone Use "No address yet" (e.g. during DAD) Two key observations: every IPv6 interface has a link-local address (FE80::/10) automatically; routing protocols use these as next-hops. Global unicast addresses are not strictly required for inside-AS forwarding; OSPFv3 and EIGRP for IPv6 work entirely over link-locals. ## The IPv6 Header The IPv6 header is 40 bytes, fixed length: ``` +-Version (4)-+-Traffic Class (8)-+-Flow Label (20)-----+ +-Payload Length (16)-+-Next Header (8)-+-Hop Limit (8)-+ +-Source Address (128 bits)----------------------------+ +-Destination Address (128 bits)-----------------------+ ``` Version Size4 bits PurposeAlways 6 for IPv6 Traffic Class Size8 bits Purpose QoS marking (DSCP+ECN, like IPv4 ToS) Flow Label Size20 bits Purpose Per-flow identifier (rarely used in practice) Payload Length Size16 bits Purpose Length of payload in bytes (excludes header) Next Header Size8 bits Purpose Type of next header (TCP=6, UDP=17, ICMPv6=58, etc.) - replaces IPv4's Protocol field Hop Limit Size8 bits Purpose TTL equivalent; decremented at each hop Source Address Size128 bits PurposeSource IPv6 address Destination Address Size128 bits Purpose Destination IPv6 address What is NOT in the IPv6 header (compared to IPv4): - Header checksum (gone) - Header length (fixed at 40) - Identification, flags, fragment offset (only in extension headers if fragmenting) - Options (replaced by extension headers, optional) The result: more compact, faster to parse, easier to hardware-accelerate. ## Neighbor Discovery: The ARP Replacement IPv6 does not use ARP. Address resolution uses ICMPv6 Neighbor Discovery (NDP, RFC 4861). Five message types matter: Router Solicitation (RS) ICMPv6 Type133 Purpose Host asks routers to send Router Advertisements Router Advertisement (RA) ICMPv6 Type134 Purpose Router announces its presence and prefix info; sent periodically and in response to RS Neighbor Solicitation (NS) ICMPv6 Type135 Purpose "What is the MAC for this IPv6 address?" - replaces ARP request Neighbor Advertisement (NA) ICMPv6 Type136 Purpose "My MAC for this IPv6 is X" - replaces ARP reply Redirect ICMPv6 Type137 Purpose Tell a host about a better next-hop NDP is a strictly link-local protocol; messages have hop limit 255 and are dropped if any router decrements the hop limit (which would happen if they crossed a router). This is the GTSM-like protection against off-link attacks. The neighbor table is the IPv6 equivalent of the ARP table. On the lab's ABR, three neighbors are visible - all by their link-local addresses, not their global ones: ``` R2#show ipv6 neighbors IPv6 Address Age Link-layer Addr State Interface FE80::A8BB:CCFF:FE00:300 0 aabb.cc00.0300 STALE Et0/0 FE80::A8BB:CCFF:FE00:400 0 aabb.cc00.0400 STALE Et0/0 FE80::A8BB:CCFF:FE00:500 0 aabb.cc00.0500 STALE Et0/1 ``` ND uses a state machine very different from ARP. STALE means "I last saw this neighbor more than 30 seconds ago and I have not actively forwarded traffic through it since" - which is normal for the steady state. Active forwarding triggers a transition through DELAY and PROBE; a confirmed-fresh entry sits in REACHABLE. The neighbor entries are always link-local because that is what ND learns from NS/NA exchanges on the segment. ## SLAAC: Stateless Address Autoconfiguration SLAAC is the IPv6 mechanism that lets hosts pick their own global address without DHCP. The flow: 1. Host comes up; auto-assigns a link-local FE80:: address using EUI-64 or a random interface ID 2. Host sends a Router Solicitation 3. Router responds with Router Advertisement carrying the network's /64 prefix 4. Host concatenates the prefix with its own interface ID to form a global address (e.g. prefix 2001:db8:1::/64 + interface ID becomes 2001:db8:1:0:abcd:ef01:2345:6789) 5. Host runs Duplicate Address Detection (DAD) by sending an NS for its tentative address; if no response, it adopts the address SLAAC is enabled by default on most IPv6 networks. DHCPv6 (a separate protocol) is used for stateful address assignment when the operator wants control, or for sending DNS server addresses (which RAs can also do via RDNSS option). The SLAAC capture from the lab. R3's Et0/0 is configured with `ipv6 address autoconfig` instead of a static address. It heard the unsolicited RAs from R1 and R2, learned the on-link prefix 2001:DB8:20::/64, built an EUI-64 suffix from its own MAC, and assigned itself a global address - complete with the lifetimes the RA advertised: ``` R3#show ipv6 interface Ethernet0/0 Ethernet0/0 is up, line protocol is up IPv6 is enabled, link-local address is FE80::A8BB:CCFF:FE00:400 Description: To SW1 (Area 0 broadcast LAN) Stateless address autoconfig enabled Global unicast address(es): 2001:DB8:20:0:A8BB:CCFF:FE00:400, subnet is 2001:DB8:20::/64 [EUI/CAL/PRE] valid lifetime 2591929 preferred lifetime 604729 Joined group address(es): FF02::1 FF02::2 FF02::5 FF02::6 FF02::1:FF00:400 MTU is 1500 bytes Hosts use stateless autoconfig for addresses. ``` The differences from R1's interface are exactly what SLAAC adds: the line "Stateless address autoconfig enabled" appears, the `[EUI/CAL/PRE]` tag has Calendar-valid and Preferred flags, and the address has explicit valid/preferred lifetimes (here 30 days valid, 7 days preferred - typical RA defaults). When those lifetimes expire, R3 will renew them via subsequent RAs or generate a new address. The forwarding table sees the SLAAC-installed on-link prefix as a route code `NDp`: ``` R3#show ipv6 route NDp 2001:DB8:20::/64 [2/0] via Ethernet0/0, directly connected ``` The `NDp` code (Neighbor Discovery Prefix) only appears when the prefix was learned via SLAAC. A statically-configured prefix would show as `C` (Connected). ## IPv6 Routing Protocols OSPFv3 (RFC 5340) ProtocolOSPF Notes Separate protocol from OSPFv2; carries IPv6 routes natively, plus IPv4 in modern code EIGRP for IPv6 ProtocolEIGRP Notes Same DUAL algorithm; IPv6 address-family in named-mode EIGRP MP-BGP with address-family ipv6 unicast ProtocolBGP Notes Same BGP, different AFI/SAFI RIPng (RFC 2080) ProtocolRIPng Notes RIP for IPv6; rare in production IS-IS multi-topology ProtocolIS-IS Notes IS-IS carries IPv6 prefixes natively OSPFv3 is the dominant IPv6 IGP in enterprise networks. Conceptually it is identical to OSPFv2 (same areas, same DR/BDR election on broadcast media, same SPF), but it runs as a separate protocol instance and uses link-local addresses for everything. The neighbor table on the dual-stack ABR makes both halves visible side by side: ``` R2#show ipv6 ospf neighbor OSPFv3 Router with ID (10.255.0.2) (Process ID 100) Neighbor ID Pri State Dead Time Interface ID Interface 10.255.0.1 1 FULL/DROTHER 00:00:35 1 Ethernet0/0 10.255.0.3 1 FULL/DR 00:00:35 1 Ethernet0/0 10.255.0.4 1 FULL/DR 00:00:38 1 Ethernet0/1 ``` OSPFv3 uses the IPv4 router-id (10.255.0.2) for the neighbor identifier - there is no separate v6 router-id space. State columns mirror OSPFv2: FULL is the steady state; DR/BDR/DROTHER apply on broadcast media. The IPv6 routing table on the same router shows OSPFv3-learned routes. Note that the next-hops are link-local addresses, not the global ones - that is the OSPFv3 convention: ``` R2#show ipv6 route ospf IPv6 Routing Table - default - 9 entries O 2001:DB8:255::1/128 [110/10] via FE80::A8BB:CCFF:FE00:300, Ethernet0/0 O 2001:DB8:255::3/128 [110/10] via FE80::A8BB:CCFF:FE00:400, Ethernet0/0 O 2001:DB8:255::4/128 [110/10] via FE80::A8BB:CCFF:FE00:500, Ethernet0/1 ``` Three loopback /128s, all reached through link-local next-hops on either the LAN (Et0/0) or the P2P link to R4 (Et0/1). End-to-end reachability across the dual-stack lab is the proof that everything is wired correctly: ``` R3#ping 2001:db8:255::4 Type escape sequence to abort. Sending 5, 100-byte ICMP Echos to 2001:DB8:255::4, timeout is 2 seconds: !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 2/3/8 ms ``` For BGP IPv6, see the [BGP cluster pillar](https://www.pinglabz.com/bgp/) and [MP-BGP article](https://www.pinglabz.com/mp-bgp-multiprotocol-bgp/). For OSPFv3, see the [OSPF cluster pillar](https://www.pinglabz.com/ospf/). Both IGPs have complete guides of their own: [EIGRP](https://www.pinglabz.com/eigrp/) and [OSPF](https://www.pinglabz.com/ospf/). In MPLS cores, IPv6 typically rides 6PE/6VPE over the IPv4 label-switched path - see the [MPLS complete guide](https://www.pinglabz.com/mpls/). ## Cisco IOS XE Configuration Enabling IPv6 globally and on an interface: ``` ! Global enablement ipv6 unicast-routing ipv6 cef ! Per-interface (static) interface Ethernet0/0 ipv6 address 2001:db8:20::2/64 ! Per-interface (EUI-64 generated suffix) interface Ethernet0/0 ipv6 address 2001:db8:20::/64 eui-64 ! Per-interface (SLAAC client - learns prefix from RAs) interface Ethernet0/0 ipv6 address autoconfig ! Link-local only, no global (useful for transit links between routers) interface Ethernet0/1 ipv6 enable ``` Adding OSPFv3 to make the lab routable is two extra lines per interface plus a protocol-level router-id. From the dual-stack PingLabz reference lab, R2 (the ABR) carries both v4 OSPFv2 and v6 OSPFv3 with parallel area structures - this is the production pattern for enterprises migrating gradually to IPv6: ``` ! Add to each router after the IPv4 OSPF is already working ipv6 router ospf 100 router-id 10.255.0.2 ! Reuse the IPv4 router-id; no separate v6 ID ! Per-interface enrollment in OSPFv3 interface Ethernet0/0 ipv6 address 2001:db8:20::2/64 ipv6 ospf 100 area 0 interface Ethernet0/1 ipv6 address 2001:db8:30::1/64 ipv6 ospf 100 area 30 ! Different area = ABR interface Loopback0 ipv6 address 2001:db8:255::2/128 ipv6 ospf 100 area 0 ``` That's it. OSPFv3 picks up the configured areas, forms adjacencies on the link-local addresses, and exchanges LSAs. The DR/BDR election runs on broadcast media exactly as OSPFv2 does, but you read it from `show ipv6 ospf interface` rather than `show ip ospf interface`: ``` R3#show ipv6 ospf interface Ethernet0/0 Ethernet0/0 is up, line protocol is up Link Local Address FE80::A8BB:CCFF:FE00:400, Interface ID 1 Area 0, Process ID 100, Instance ID 0, Router ID 10.255.0.3 Network Type BROADCAST, Cost: 10 Transmit Delay is 1 sec, State DR, Priority 1 Designated Router (ID) 10.255.0.3, local address FE80::A8BB:CCFF:FE00:400 Backup Designated router (ID) 10.255.0.2, local address FE80::A8BB:CCFF:FE00:200 Timer intervals configured, Hello 10, Dead 40, Wait 40, Retransmit 5 ``` R3 is DR on this segment (it tied on default priority and won on highest router-id). The local address recorded for the DR is the link-local FE80:: form, not the global - again, OSPFv3 cares about link-locals. Quick verification commands: ``` ! Show all IPv6-enabled interfaces Router# show ipv6 interface brief ! Inspect a specific interface (link-local + global + multicast groups) Router# show ipv6 interface Ethernet0/0 ! Show the IPv6 routing table Router# show ipv6 route ! Show the neighbor (ARP-replacement) table Router# show ipv6 neighbors ``` ## DHCPv6 and How Hosts Actually Get Addresses The biggest mental shift from IPv4: an IPv6 host does not need a DHCP server. The router advertises a prefix in a Router Advertisement and the host builds its own global address from it (SLAAC). DHCPv6 sits on top of that as an option, not a requirement. Two bits in the RA arbitrate. The **M (managed)** flag says "get your address from DHCPv6." The **O (other)** flag says "get your DNS and domain from DHCPv6." Those two bits produce the three deployment models: pure SLAAC, stateless DHCPv6, and stateful DHCPv6\. And a trap that catches everyone: setting M does not stop SLAAC, so unless you also mark the prefix `no-autoconfig`, hosts end up holding two addresses from two mechanisms. Note also that **DHCPv6 never provides a default gateway**. That always comes from the RA, which means you cannot run DHCPv6 without RAs. Full walkthrough, including DUIDs (DHCPv6 does not key on MAC addresses) and a working relay: [DHCPv6 explained: stateful, stateless, SLAAC, and the relay](https://www.pinglabz.com/dhcpv6-stateful-stateless-explained/). ## IPv6 First-Hop Security IPv6 hosts autoconfigure from unauthenticated Router Advertisements, and any device on the segment can send one. That makes the access edge more exposed than IPv4, and it needs switch-side defences. [RA Guard and DHCPv6 Guard](https://www.pinglabz.com/ipv6-ra-guard-dhcpv6-guard/) stop rogue gateways and rogue DHCPv6 servers (proven in the lab by flooding rogue RAs and confirming the victim never learned the rogue prefix), and [ND Inspection and Source Guard](https://www.pinglabz.com/ipv6-nd-inspection-source-guard/) build the device-tracking binding table and enforce it against address theft and spoofing. Both are part of the [Infrastructure Security cluster](https://www.pinglabz.com/infrastructure-security/). ## IPv6 Deep Dives in This Cluster 1. [IPv6 Address Format Explained Byte by Byte](https://www.pinglabz.com/ipv6-address-format/) 2. [IPv4 vs IPv6: The Real Differences](https://www.pinglabz.com/ipv4-vs-ipv6/) 3. [IPv6 Address Types: Link-Local, Global, Unique Local](https://www.pinglabz.com/ipv6-address-types/) 4. [IPv6 Configuration on Cisco IOS XE](https://www.pinglabz.com/ipv6-cisco-configuration/) 5. [IPv6 Header Format Explained](https://www.pinglabz.com/ipv6-header-explained/) 6. [EIGRP for IPv6: Classic and Named Mode on IOS XE](https://www.pinglabz.com/eigrp-for-ipv6/) 7. [IS-IS for IPv6: Multi-Topology and the Single-Topology Trap](https://www.pinglabz.com/is-is-for-ipv6/) 8. [IPv6 over IPv4 Tunnels: Manual, GRE, and 6to4 Compared](https://www.pinglabz.com/ipv6-over-ipv4-tunnels/) 9. [NAT64 Explained: Stateful, Stateless, and the DNS64 Half of the Story](https://www.pinglabz.com/nat64-explained/) 10. [IPv6 Transition Mechanisms: Which One, and When](https://www.pinglabz.com/ipv6-transition-mechanisms/) ### Neighbor Discovery and SLAAC This cluster covered addressing, routing and DHCPv6 in depth but had nothing on the protocol that actually glues an IPv6 segment together. That gap is now filled. [The host is stuck at link-local](https://www.pinglabz.com/ipv6-neighbor-discovery-troubleshooting/) Why no global address formed, reading the RA flags, the neighbor cache states, and why blocking ICMPv6 breaks IPv6 in a way blocking ICMP never did for IPv4. ## FAQ ### What does IPv6 stand for? IPv6 is Internet Protocol version 6, the successor to IPv4 (Internet Protocol version 4). The intermediate version 5 was an experimental streaming protocol that was never deployed; 6 was assigned to the next-generation IP project. ### How much bigger is IPv6 than IPv4? 340 undecillion (3.4 x 10^38) IPv6 addresses vs 4.3 billion IPv4\. Roughly 10^29 times bigger. Enough for every device on Earth to have many addresses each, and every grain of sand on Earth to have its own /64 subnet. ### Is IPv6 actually deployed? Yes. As of 2026, Google reports that around 45-50 percent of users connect via IPv6\. Major content (Google, Facebook, Netflix, AWS, Azure) all support IPv6 natively. Mobile carriers in many countries default to IPv6-only for cellular data. Enterprise IPv6 adoption lags consumer/mobile but is increasing. ### Should I disable IPv6 on my hosts? No. Microsoft, Apple, and the IETF strongly advise against disabling IPv6\. Disabled IPv6 causes issues with newer applications and breaks privacy extensions. If you don't want IPv6 routing, leave the protocol enabled but don't deploy IPv6 routing in your network. ### Does IPv6 use ARP? No. IPv6 uses ICMPv6 Neighbor Discovery (NDP) instead of ARP. NDP serves the same purpose (mapping IPv6 addresses to MAC addresses) but uses ICMPv6 messages over multicast rather than ARP's broadcast model. This is more efficient and integrates with the rest of IPv6 (RAs, etc.). ### What are the private IPv6 addresses? Unique Local Addresses (ULAs) in the FC00::/7 range, typically FD00::/8 in practice (the second half of the ULA range). Equivalent to IPv4's RFC 1918 private addresses but with a 40-bit pseudo-random Global ID to prevent collisions if networks are merged. Hands-on IPv6 - configure addressing and routing Configure IPv6 with manual and EUI-64 derivation, observe SLAAC, and build end-to-end IPv6 reachability with static routes using link-local next-hops. Real `show ipv6 neighbors` \+ ICMPv6 ping output. Open the PingLabz CCNA Labs library. [Open the IPv6 labs](https://www.pinglabz.com/ccna-labs-network-fundamentals/) ## Key Takeaways IPv6 is no longer optional knowledge for network engineers in 2026\. The address space is 128-bit (340 undecillion addresses), the header is 40 bytes fixed, ARP is replaced by Neighbor Discovery, and SLAAC handles host autoconfiguration without DHCP. Routing protocols (OSPFv3, MP-BGP for IPv6, EIGRP for IPv6, VRRPv3) handle IPv6 with the same conceptual models as their IPv4 counterparts. Master the address format, the address types, and Neighbor Discovery, and the rest of IPv6 follows naturally. Bookmark this page, work through the cluster articles in order, and lab everything in a controlled environment - IPv6 troubleshooting habits are different from IPv4 and need practice. **Studying for the CCNA?** Test your IPv6 knowledge on [PingLabz CCNA Flashcards](https://www.pinglabz.com/ccna-flashcards/) \- 200 free multiple-choice questions by topic, mixed, or a full mock exam, each with a plain-English explanation. ### References - [RFC 4291 - IP Version 6 Addressing Architecture](https://www.rfc-editor.org/rfc/rfc4291?ref=pinglabz.com) - [RFC 8200 - Internet Protocol, Version 6 (IPv6) Specification](https://www.rfc-editor.org/rfc/rfc8200?ref=pinglabz.com) - [Cisco IPv6 technology documentation](https://www.cisco.com/c/en/us/tech/ip/ip-version-6-ipv6/index.html?ref=pinglabz.com) ## IPv6 across a service provider core (6PE) Service providers with a mature IPv4 MPLS core rarely want to rebuild it just to deliver IPv6\. 6PE lets them carry IPv6 prefixes across that core inside MPLS labels, signalled by MP-BGP, without the P routers ever running IPv6\. It is the most common way IPv6 transit actually reaches customers today: [IPv6 over BGP: MP-BGP for IPv6 and 6PE explained](https://www.pinglabz.com/ipv6-bgp-6pe/). For BGP policy at expert depth, see the [BGP pillar guide](https://www.pinglabz.com/bgp/). ## IPv6 L3VPN over an IPv4 core (6VPE) A service provider with a mature IPv4 MPLS L3VPN does not want to rebuild it to sell IPv6 VPN service. 6VPE carries IPv6 L3VPN across that unchanged IPv4 core using the VPNv6 address family and a two-label stack - the per-VRF-isolated sibling of 6PE: [MPLS VPNv6 and 6VPE explained](https://www.pinglabz.com/mpls-vpnv6-6vpe/). See also [6PE](https://www.pinglabz.com/ipv6-bgp-6pe/) for the global-table version. ## IPv6 edge services: prefix delegation, general prefix, NPTv6 IPv6 numbers the network edge far more elegantly than IPv4 ever did. DHCPv6 prefix delegation hands a whole subnet to a downstream router automatically; a general prefix lets you renumber an entire site by changing one value; and NPTv6 translates prefixes statelessly when you genuinely need it, without the baggage of NAT: [IPv6 services closure](https://www.pinglabz.com/ipv6-nptv6-dhcpv6-pd/). ## IPv6 IGPs, Tunnels, and Translation Three pieces round out the cluster. On the IGP side, IPv6 has more than OSPFv3: [EIGRP for IPv6 in classic and named mode](https://www.pinglabz.com/eigrp-for-ipv6/) runs the same DUAL algorithm over an IPv6 address-family, and [IS-IS for IPv6 with multi-topology](https://www.pinglabz.com/is-is-for-ipv6/) keeps v4 and v6 in separate topologies to avoid the single-topology trap that black-holes traffic. When two IPv6 islands are separated by an IPv4 core, [IPv6-over-IPv4 tunneling with manual, GRE, and 6to4 tunnels](https://www.pinglabz.com/ipv6-over-ipv4-tunnels/) carries v6 across it. And when an IPv6-only host must reach an IPv4 service, [NAT64 and DNS64 translation](https://www.pinglabz.com/nat64-explained/) bridges the two protocols. If you are weighing the options, [which IPv6 transition mechanism to use, and when](https://www.pinglabz.com/ipv6-transition-mechanisms/) maps each one to the scenario it actually fits. ## Studying for the CCIE? This cluster is part of the full CCNA to CCNP to CCIE Enterprise ladder on PingLabz, every rung built on real Cisco output. For expert-level depth across every EI v1.1 blueprint domain - and the four integration Super Labs - see the [CCIE Enterprise Infrastructure study hub](https://www.pinglabz.com/ccie-enterprise/). **Studying for CCIE Security?** IPv6 addressing, Neighbor Discovery, and first-hop security appear across the lab blueprint. Take IPv6 threat mitigation and hardening to expert depth with the infrastructure security track in [the CCIE Security study hub](https://www.pinglabz.com/ccie-security/). ### GRE Tunnels: The Complete Guide for Network Engineers URL: https://www.pinglabz.com/gre/ Last updated: 2026-07-11T19:14:31.000Z GRE (Generic Routing Encapsulation) is the simplest tunnel protocol in the network engineer's toolkit. You take any kind of network packet, wrap it inside an IP packet, and send it across an IP network as if you had a direct cable between the two endpoints. That is GRE in one sentence. [RFC 2784](https://www.rfc-editor.org/rfc/rfc2784?ref=pinglabz.com) (with RFC 2890 adding key and sequence-number extensions) defines the format, Cisco invented it in the early 1990s, and three decades later it still underpins more enterprise overlays than any other tunnel technology because it is everywhere, it is well-understood, and it does exactly one thing well. This is the cluster overview for the full PingLabz GRE series: how GRE encapsulation works, the packet format, configuration on Cisco IOS XE, routing protocols over GRE, keepalives, the MTU and fragmentation problem, GRE over IPsec, multipoint GRE / DMVPN, and troubleshooting. We will walk through the basics, then the parts of GRE that catch engineers out in production, then the operational realities. If you are configuring your first GRE tunnel, designing a hub-and-spoke overlay, or chasing a recursive-routing flap at 2 AM, start here. New labs and guides, in your inbox Every new PingLabz lab and deep-dive, built and verified on real Cisco IOS XE - free, straight to your inbox. [Join free](https://www.pinglabz.com/signup/) ## What GRE Solves GRE was built to carry network protocols that the underlying IP network does not natively understand. In the early 1990s that meant tunneling IPX, AppleTalk, and DECnet across an IP backbone. In 2026 the use cases are different but the core value is the same: GRE lets you treat an IP path between two routers as a virtual point-to-point link. That virtual link gives you four things you cannot easily get from plain IP routing: - **Multicast over an internet path.** OSPF and EIGRP rely on multicast hellos, and the public internet does not forward multicast. Wrap them in GRE and the multicast becomes a unicast IP packet the internet is happy to deliver. - **Non-IP payload over IP.** IPv6 over IPv4 (6in4), MPLS over IP backbones, IS-IS over IP - GRE carries any protocol the receiver knows how to decode. - **A predictable Layer 3 next-hop.** The other end of the tunnel is always one IP hop away no matter how many physical hops the underlay takes, which simplifies recursive routes, traffic engineering, and policy-based routing. - **Decoupled overlay from underlay.** Tunnel endpoints can change carriers, IP addresses, or physical paths without touching the overlay routing protocol. GRE is the duct tape of network engineering, and it never wears out. What GRE deliberately does not do is encrypt. There is no authentication, no confidentiality, no integrity protection in the GRE header. If you put GRE on the public internet, anyone who can capture the packets can read them. That is why production deployments almost always pair GRE with IPsec, which is the subject of [GRE over IPsec: When and How to Combine Them](https://www.pinglabz.com/gre-over-ipsec/). ![Four problems GRE solves: multicast over the internet (OSPF/EIGRP hellos in unicast wrappers), non-IP payload over IP (IPv6-in-IPv4 6in4, MPLS over IP, IS-IS over IP), predictable Layer 3 next-hop, and decoupled overlay from underlay](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/what-gre-solves.png) The four jobs GRE actually does well. Notice none of them is encryption. ## How GRE Encapsulation Works (the 10,000-Foot View) A GRE tunnel is just two routers that have agreed to wrap traffic for each other. There are no sessions, no handshakes, no negotiation. Each router has a tunnel interface configured with a `tunnel source` (the local underlay IP), a `tunnel destination` (the remote underlay IP), and a `tunnel mode` (almost always `gre ip` for plain GRE over IPv4). Anything routed into that tunnel interface gets encapsulated and sent to the destination. The encapsulation steps in order: 1. A packet is routed into the tunnel interface by the IP routing table. The tunnel interface looks like any other interface to the routing process. 2. The router prepends a GRE header (4 bytes minimum, more if optional fields are enabled) describing the protocol type of the inner packet. 3. The router prepends an outer IP header. The outer source is the tunnel source IP, the outer destination is the tunnel destination IP, and the protocol field is set to 47 (the IANA-assigned GRE protocol number). 4. The packet is forwarded out the underlay interface that the routing table says reaches the tunnel destination. From the underlay's perspective it is a plain IP unicast packet bound for some destination on the public internet. 5. The remote router receives the packet, sees protocol 47, strips the outer IP and GRE headers, and routes the inner packet according to its own routing table. That is it. There is no per-packet state on either end. If a packet is dropped, GRE does not know and does not care; the inner protocol (TCP, the routing protocol, whatever) handles loss recovery. This statelessness is GRE's strength (no fancy failure modes) and its weakness (no native liveness check, which is why [GRE keepalives](https://www.pinglabz.com/gre-tunnel-keepalives/) were added later). ![Five-step GRE encapsulation flow on R1: packet routed into Tunnel0, prepend 4-byte GRE header with proto type, prepend 20-byte outer IP with proto 47, forward via underlay route to R2, R2 strips headers and routes the inner packet. R1 marked encapsulate, R2 marked decapsulate, with stateless-by-design callout](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/encap-flow.png) The five steps the local router runs on every packet entering the tunnel. No sessions, no handshake, no per-packet state. ## GRE Packet Format The wire format of a GRE-encapsulated packet, from outermost to innermost: Ethernet LayerOuter L2 Size (typical)14 bytes Notes Underlay frame header (varies by underlay) IP (proto 47) LayerOuter L3 Size (typical)20 bytes (IPv4) Notes Source = tunnel source IP, dest = tunnel destination IP GRE header LayerGRE Size (typical)4 bytes (minimum) Notes Protocol type, flags. Up to 16 bytes if key + sequence enabled IP / IPv6 / other LayerInner L3 Size (typical)20 bytes (inner IPv4) Notes The original packet that was routed into the tunnel TCP / UDP / etc LayerPayload Size (typical)variable NotesUser data Total overhead added per packet is 24 bytes for plain GRE over IPv4 (20-byte outer IPv4 plus 4-byte GRE). That is the reason the canonical Cisco MTU recommendation for a 1500-byte underlay is 1476 bytes inside the tunnel. Add IPsec encryption on top of GRE and you lose another 50-60 bytes depending on the cipher and mode, which is why MTU-and-MSS becomes its own deep dive at [GRE MTU and Fragmentation: Fixing Tunnel Packet Loss](https://www.pinglabz.com/gre-tunnel-mtu/). ![GRE packet format on the wire from outermost to innermost: 14-byte Ethernet, 20-byte outer IP (proto 47), 4-byte GRE, 20-byte inner IP, then payload. 24-byte GRE overhead annotated. MTU math: 1500 underlay minus 24 equals 1476 tunnel MTU. With IPsec: add 50-60 bytes in tunnel mode, ip mtu 1400 and adjust-mss 1360. Over IPv6: 40 outer plus 4 GRE equals 44-byte overhead](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/packet-format.png) The wire format. The 24-byte GRE overhead is the math behind the canonical 1476 tunnel MTU and the 1400 production setting. ## GRE Header Fields The 4-byte mandatory GRE header has five flag bits and a Protocol Type field. The flags decide whether optional fields follow: Checksum (C) Bits1 Purpose If set, GRE includes a 16-bit checksum over header + payload (rarely enabled in practice) Key (K) Bits1 Purpose If set, GRE includes a 32-bit Key field used to demultiplex multiple tunnels with same source / destination, or as a tag for mGRE / DMVPN Sequence (S) Bits1 Purpose If set, GRE includes a 32-bit sequence number for in-order delivery (rarely enabled) Version Bits3 Purpose 0 for standard GRE; 1 for PPTP (the old Microsoft VPN) Protocol Type Bits16 Purpose EtherType of the inner protocol. 0x0800 for IPv4, 0x86DD for IPv6, 0x8847 for MPLS, etc. The Protocol Type field is what makes GRE generic. By advertising the EtherType of the inner protocol, GRE can carry anything the receiver knows how to decode. Most of the time you are looking at 0x0800 (IPv4) or 0x86DD (IPv6) and the rest of the flags are zero. The Key field becomes important when you build multipoint GRE for DMVPN; see [mGRE and DMVPN Introduction](https://www.pinglabz.com/mgre-dmvpn-introduction/) for how the same tunnel source can serve many destinations. ![The 4-byte GRE header at the bit level: bits 0-2 are C/K/S flags (checksum, key, sequence), bits 3-12 are 9-bit reserved0 must-be-zero, bits 13-15 are 3-bit Version (0 standard, 1 PPTP), bits 16-31 are 16-bit Protocol Type carrying the EtherType. C is rarely enabled, K is used to demux mGRE/DMVPN, S is rarely enabled, Version is 0 for standard GRE, Protocol Type uses 0x0800 for IPv4, 0x86DD for IPv6, 0x8847 for MPLS](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/header-bits.png) The 4-byte mandatory header. Most fields are usually zero; the Protocol Type is what makes GRE generic. ## Configuration on Cisco IOS XE: Minimum Viable GRE The smallest working GRE tunnel between two Cisco routers: ``` ! ---- R1 ---- interface Tunnel0 ip address 10.0.0.1 255.255.255.252 tunnel source 198.51.100.1 tunnel destination 203.0.113.1 tunnel mode gre ip ! ! ---- R2 ---- interface Tunnel0 ip address 10.0.0.2 255.255.255.252 tunnel source 203.0.113.1 tunnel destination 198.51.100.1 tunnel mode gre ip ``` Two interface stanzas, four lines of relevant config each. The tunnel source and destination must be IPs that the underlay routing table can reach (typically internet-routable addresses on each end, or any reachable IP if the tunnel rides over a private WAN). The 10.0.0.0/30 inside the tunnel is the overlay; pick anything you like as long as both ends agree. Verification on R1: ``` R1#show interface Tunnel0 Tunnel0 is up, line protocol is up Hardware is Tunnel Description: GRE tunnel to R3 (source Lo0, dest 10.255.0.3) Internet address is 10.30.30.1/30 MTU 17916 bytes, BW 100 Kbit/sec, DLY 50000 usec, reliability 255/255, txload 1/255, rxload 1/255 Encapsulation TUNNEL, loopback not set Keepalive set (5 sec), retries 3 Tunnel linestate evaluation up Tunnel source 10.255.0.1 (Loopback0), destination 10.255.0.3 Tunnel Subblocks: src-track: Tunnel0 source tracking subblock associated with Loopback0 Set of tunnels with source Loopback0, 1 member, on interface Tunnel protocol/transport GRE/IP Key disabled, sequencing disabled Checksumming of packets disabled Tunnel TTL 255, Fast tunneling enabled Tunnel transport MTU 1476 bytes Tunnel state info: State change to up : May 11 2026 06:32:49 ``` The output is from the PingLabz GRE Reference Lab where R1's Tunnel0 is sourced from Loopback0 (10.255.0.1) and destined for R3's Loopback0 (10.255.0.3). The `Tunnel source 10.255.0.1 (Loopback0)` line is the production-style flavour - sourcing from a loopback rather than a physical interface means the tunnel survives any single underlay interface failure as long as the underlay routes to that loopback. `Tunnel transport MTU 1476` is Cisco doing the math: 1500 underlay minus 24 GRE+IP overhead. `Tunnel linestate evaluation up` is the modern IOS XE per-tunnel state tracker confirming the destination is reachable and the tunnel is forwarding. Note that line-protocol-up does not actually mean the remote end is processing traffic - a GRE tunnel will show up / up as long as the underlay route to the destination is reachable, even if the remote router is offline. That is exactly the problem keepalives were invented to solve. Full lab walkthrough at [GRE Tunnel Configuration: Step-by-Step Cisco IOS-XE Lab](https://www.pinglabz.com/gre-tunnel-configuration-cisco/). Sourcing Tunnel0 from a loopback (as the lab does) means the underlay needs a route to that loopback. In the lab R1 has a static route to R3's loopback via the transit router; in production this is normally the IGP's job. The IP routing table on R1 makes the resolution chain explicit: ``` R1#show ip route 10.0.0.0/8 is variably subnetted, 4 subnets, 2 masks C 10.30.30.0/30 is directly connected, Tunnel0 L 10.30.30.1/32 is directly connected, Tunnel0 C 10.255.0.1/32 is directly connected, Loopback0 S 10.255.0.3/32 [1/0] via 192.0.2.2 192.0.2.0/24 is variably subnetted, 2 subnets, 2 masks C 192.0.2.0/30 is directly connected, Ethernet0/0 L 192.0.2.1/32 is directly connected, Ethernet0/0 ``` The `S 10.255.0.3/32 [1/0] via 192.0.2.2` line is the underlay path that resolves the tunnel destination. Take it away or let the inside-the-tunnel routing protocol be the only source of that route, and you create the recursive-routing scenario covered at the end of this guide. ## Routing Over GRE The reason most GRE tunnels exist is to carry a routing protocol. Plain static routes do not need GRE; if you have a static route to the remote network via the underlay, you do not need a tunnel. GRE earns its keep when you want OSPF, EIGRP, or BGP to form a neighbor relationship across an underlay that does not natively support it. The pattern is the same regardless of routing protocol: 1. Bring up the GRE tunnel. 2. Put the tunnel interface in the routing protocol's interface set (`network` statement, `ip ospf` command, EIGRP `network` command, or BGP `neighbor` using the tunnel-overlay IP). 3. Watch the neighbor come up. From the routing protocol's perspective the tunnel looks like any other point-to-point link. OSPF over GRE is the most common case because OSPF needs multicast (224.0.0.5 / 224.0.0.6) for hellos and DR election. The internet does not forward multicast; GRE wraps the multicast in unicast and the problem disappears. EIGRP behaves the same way. iBGP over GRE is common in hub-and-spoke designs where the overlay carries the iBGP mesh. Worked configs and gotchas (recursive routing, OSPF network types, EIGRP delay tweaking) are at [Routing Protocols Over GRE: OSPF, EIGRP, BGP](https://www.pinglabz.com/routing-protocols-over-gre/). OSPF over the lab's tunnel comes up FULL with R3 (the only neighbor on the tunnel), and the neighbor address is R3's Tunnel0 IP (10.30.30.2), not its loopback - because OSPF talks to the IP on the interface it's running on: ``` R1#show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 10.255.0.3 0 FULL/ - 00:00:36 10.30.30.2 Tunnel0 ``` Priority 0 and state suffix `-` are the OSPF point-to-point network-type defaults; there is no DR election on a Tunnel0 P2P link. A ping across the overlay sourced from R1's loopback confirms end-to-end reachability over GRE: ``` R1#ping 10.255.0.3 source 10.255.0.1 Type escape sequence to abort. Sending 5, 100-byte ICMP Echos to 10.255.0.3, timeout is 2 seconds: Packet sent with a source address of 10.255.0.1 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 2/2/3 ms ``` One thing to plan for: any routing protocol you run inside the tunnel must not use the tunnel to learn the route to the tunnel destination itself. That is recursive routing, the tunnel will flap, and the symptom is "GRE tunnel keeps going up and down every few minutes." Use a static route to the tunnel destination via the underlay, or filter the underlay's IP out of the inside-the-tunnel routing protocol. The lab below shows what happens when you forget. Each protocol has its own tunnel quirks and its own complete guide: [OSPF](https://www.pinglabz.com/ospf/) (network types and MTU), [EIGRP](https://www.pinglabz.com/eigrp/) (DMVPN and stub spokes), and [BGP](https://www.pinglabz.com/bgp/) (eBGP multihop over tunnels). ## GRE Keepalives Vanilla GRE has no liveness check. The tunnel interface stays up / up as long as the local router can route to the configured tunnel destination. If the remote router crashes, is rebooted, or has a software fault that stops it processing GRE, the local end has no way to know and traffic black-holes. Cisco's GRE keepalives (a Cisco extension, not in the original RFC 2784 but widely adopted by other vendors) fix this by periodically sending a small GRE-encapsulated packet whose inner IP is addressed back to the sender. The remote router does not have to do anything special; it just routes the inner packet according to its routing table, which sends it back through the tunnel to the sender. If keepalives stop returning, the local end declares the tunnel down. ``` interface Tunnel0 keepalive 10 3 ``` That tells the router: send a keepalive every 10 seconds, and declare the tunnel down after 3 missed responses (so 30 seconds of no response). Defaults are 10 seconds and 3 retries. Keepalives are unidirectional - each end runs them independently - and they have a subtle interaction with IPsec that catches engineers out, which is the subject of the dedicated [GRE Tunnel Keepalives Explained](https://www.pinglabz.com/gre-tunnel-keepalives/) article. A few seconds of `debug tunnel keepalive` on R1 shows exactly what travels each way. Note the trick: R1 builds the keepalive with INNER source 10.255.0.3 (the remote loopback) and INNER dest 10.255.0.1 (its own loopback), so when R3 receives and decapsulates, it sees an ordinary IP packet from 10.255.0.3 to 10.255.0.1 and forwards it back through the tunnel: ``` R1#debug tunnel keepalive Tunnel keepalive debugging is on *May 11 06:35:34.539: Tunnel0: sending keepalive, 10.255.0.3->10.255.0.1 (len=24 ttl=255), counter=1 *May 11 06:35:34.542: Tunnel0: keepalive received, 10.255.0.3->10.255.0.1 (len=24 ttl=252), resetting counter *May 11 06:35:39.540: Tunnel0: sending keepalive, 10.255.0.3->10.255.0.1 (len=24 ttl=255), counter=1 *May 11 06:35:39.543: Tunnel0: keepalive received, 10.255.0.3->10.255.0.1 (len=24 ttl=252), resetting counter *May 11 06:35:44.540: Tunnel0: sending keepalive, 10.255.0.3->10.255.0.1 (len=24 ttl=255), counter=1 *May 11 06:35:44.544: Tunnel0: keepalive received, 10.255.0.3->10.255.0.1 (len=24 ttl=252), resetting counter R1#undebug all ``` The TTL falls from 255 (R1 originates) to 252 (R1 receives back). That is exactly the 3-hop round-trip across the underlay (R1 -> R2 -> R3, decap, forward back across the same path). The counter resets to zero on every successful return - if three consecutive sends miss their return, the tunnel goes down with a "tunnel keepalive failed" log. ## MTU, Fragmentation, and MSS Clamping The single most common operational problem with GRE tunnels is MTU. Every encapsulation adds bytes, and when the resulting packet is too big to fit in the underlay path MTU, something must give: either the packet is fragmented, or it is dropped with an ICMP "fragmentation needed" message, or, in the most painful case, it is silently dropped by an underlay device that filters ICMP. The math for plain GRE over IPv4 with a 1500-byte underlay: Underlay MTU (Ethernet) 1500 Outer IPv4 header 20 GRE header 4 Maximum inner IP packet 1476 Add IPsec on top in tunnel mode and you lose another 50-60 bytes. The two-line fix that prevents 90 percent of GRE MTU pain on Cisco: ``` interface Tunnel0 ip mtu 1400 ip tcp adjust-mss 1360 ``` `ip mtu 1400` tells the tunnel interface to fragment IP packets larger than 1400 bytes before encapsulation, leaving headroom for IPsec overhead. `ip tcp adjust-mss 1360` rewrites the MSS value in TCP SYN packets so endpoints negotiate a segment size that fits inside the tunnel without requiring fragmentation in the first place. Full breakdown of the math, when to set what value, and how to debug the symptoms (large pings fail, small ones succeed; HTTPS works for some sites but not others) is at [GRE MTU and Fragmentation](https://www.pinglabz.com/gre-tunnel-mtu/). ## Security: GRE Has None - Use IPsec GRE was designed for trusted networks. The header has no authentication, no integrity check (the optional checksum is integrity in the weakest sense), and no encryption. Anyone with packet-capture access on the underlay path can read every byte of every inner packet. For lab and private-WAN use that is fine. For anything traversing the public internet it is not. The standard pattern is GRE inside IPsec: GRE provides the encapsulation and routing-protocol multicast support, IPsec provides confidentiality, integrity, and authentication. Modern IOS XE deployments use IPsec profiles attached to the tunnel interface (the `tunnel protection ipsec profile` command), which is cleaner than the older crypto map approach. The full configuration walkthrough is at [GRE over IPsec: When and How to Combine Them](https://www.pinglabz.com/gre-over-ipsec/). For the framework on when to use plain GRE, GRE-over-IPsec, IPsec alone, or something else (VXLAN, WireGuard, SD-WAN overlays), see [GRE vs IPsec vs GRE-over-IPsec: Which Tunnel Type?](https://www.pinglabz.com/gre-vs-ipsec/). ## Multipoint GRE and DMVPN Standard GRE is point-to-point: one tunnel source, one tunnel destination, one neighbor. That works for two sites. For 50 branches connecting to two hubs, you do not want 100 manually configured GRE tunnels. Multipoint GRE (mGRE, `tunnel mode gre multipoint`) lets a single tunnel interface have many remote destinations. Combined with NHRP (Next Hop Resolution Protocol) for dynamic mapping and IPsec for encryption, mGRE becomes DMVPN (Dynamic Multipoint VPN), the Cisco standard for hub-and-spoke and spoke-to-spoke overlays. DMVPN Phase 1 is hub-and-spoke only (spoke-to-spoke goes through the hub). Phase 2 enables direct spoke-to-spoke after a hub-mediated NHRP resolution. Phase 3 (the modern default) uses NHRP shortcut switching for efficient direct paths. DMVPN now has its own full cluster, anchored by [DMVPN: The Complete Guide](https://www.pinglabz.com/dmvpn/). Start with the [mGRE and DMVPN introduction](https://www.pinglabz.com/mgre-dmvpn-introduction/) here in the GRE cluster, then continue into the deep material: [DMVPN: The Complete Guide](https://www.pinglabz.com/dmvpn/) \- the cluster pillar: architecture, phases, config, and troubleshooting in one place. [DMVPN Explained](https://www.pinglabz.com/dmvpn-explained/) \- how mGRE, NHRP, and routing interlock, traced packet by packet. [NHRP Deep Dive](https://www.pinglabz.com/nhrp-deep-dive/) \- registration, resolution, and redirects with real debugs. [Phase 1 vs Phase 2 vs Phase 3](https://www.pinglabz.com/dmvpn-phase-1-2-3-differences/) \- one lab migrated live through all three phases. [DMVPN Phase 3 Configuration on IOS XE](https://www.pinglabz.com/dmvpn-phase-3-configuration/) \- the complete working build with every verification step. [Routing Over DMVPN](https://www.pinglabz.com/routing-over-dmvpn-eigrp-ospf/) \- EIGRP and OSPF design choices compared on the same lab. [Securing DMVPN with IPsec (IKEv2)](https://www.pinglabz.com/dmvpn-ipsec-profiles-ikev2/) \- profiles, transform sets, and live SA verification. [Troubleshooting DMVPN](https://www.pinglabz.com/troubleshooting-dmvpn/) \- three failures broken on purpose and diagnosed. ## GRE Across Vendors GRE is a standard and every major vendor implements it. Cisco gets the most coverage in this guide because of its CLI ubiquity, but the wire format is the same everywhere. Quick orientation: Cisco IOS XE `interface Tunnel0` \+ `tunnel mode gre ip` Linux `ip tunnel add gre0 mode gre ...` Juniper Junos `set interfaces gr-0/0/0 unit 0 tunnel ...` Arista EOS `interface TunnelN` \+ `tunnel mode gre` MikroTik RouterOS `/interface gre add ...` Configuration syntax differs, behavior of optional fields (Key, Sequence) differs, and keepalive interop between Cisco and non-Cisco kit can be flaky because keepalives are a Cisco extension. The Linux deep dive is at [GRE on Linux: ip tunnel add Commands](https://www.pinglabz.com/gre-on-linux/). ## Troubleshooting: The Common Failures The five GRE failure modes you will actually see in production: - **Recursive routing.** The route to the tunnel destination is learned through the tunnel itself. The tunnel comes up, the routing protocol learns a better path to the underlay through the tunnel, the tunnel goes down, the routing protocol forgets, and the cycle repeats. Symptom: tunnel flaps every few seconds. Fix: static route to the tunnel destination via the underlay, or distribute-list the underlay IP out of the tunnel-side routing protocol. - **MTU drops.** Small packets work, large packets do not. Pings work, HTTPS to certain sites stalls. Fix: `ip mtu 1400` and `ip tcp adjust-mss 1360` on the tunnel interface. - **Keepalive flap with IPsec.** GRE keepalives stop returning when the IPsec tunnel rekeys, or when one direction's IPsec SA dies. Fix: enable IPsec DPD (Dead Peer Detection) and check keepalive intervals against rekey timers. - **Underlay reachability lost.** The local end can no longer route to the tunnel destination IP. Tunnel interface goes down even though "show running-config" looks correct. Fix: check the underlay routing table with `show ip route `. - **ACL blocking GRE.** A stateful firewall in the path drops protocol 47 because it does not match TCP, UDP, or ICMP. Fix: add an explicit permit for IP protocol 47 between tunnel endpoints. This catches a lot of greenfield deployments where the underlay firewall was provisioned without knowledge of the GRE plan. Recursive routing is the failure mode every GRE engineer remembers, and the lab can reproduce it on cue. Starting from the working topology - tunnel up, OSPF FULL across it, an underlay static route to the tunnel destination - simply remove the underlay static. R1's only remaining path to 10.255.0.3 is the OSPF route learned over Tunnel0 itself, and IOS detects the loop within a few seconds: ``` R1(config)#no ip route 10.255.0.3 255.255.255.255 192.0.2.2 ! ~30 seconds later, syslog reports: *May 11 06:36:16.894: %ADJ-5-PARENT: Midchain parent maintenance for IP midchain out of Tunnel0 - looped chain attempting to stack *May 11 06:36:19.540: %TUN-5-RECURDOWN: Tunnel0 temporarily disabled due to recursive routing *May 11 06:36:19.540: %LINEPROTO-5-UPDOWN: Line protocol on Interface Tunnel0, changed state to down *May 11 06:36:19.542: %OSPF-5-ADJCHG: Process 1, Nbr 10.255.0.3 on Tunnel0 from FULL to DOWN, Neighbor Down: Interface down or detached R1#show interface Tunnel0 | include line protocol|linestate|destination Tunnel0 is up, line protocol is down Tunnel linestate evaluation down - no output interface Tunnel source 10.255.0.1 (Loopback0), destination 10.255.0.3 ``` The `%TUN-5-RECURDOWN: Tunnel0 temporarily disabled due to recursive routing` line is the canonical "GRE blew up" log every engineer recognises. The supporting `%ADJ-5-PARENT ... looped chain attempting to stack` message is the CEF adjacency subsystem detecting the recursion at the forwarding-plane level - it spotted that resolving the tunnel's next-hop would require traversing the tunnel itself. Restoring the underlay static reverses the failure within seconds: the chain unwinds, the tunnel comes back up, and OSPF re-converges. The full debug walkthrough with `debug tunnel`, `debug ip packet detail`, and packet captures of the problem cases is at [GRE Tunnel Troubleshooting Guide](https://www.pinglabz.com/gre-tunnel-troubleshooting/). ![Five GRE production failure modes with symptom and fix for each: recursive routing (tunnel flaps every few seconds, fix with static route to tunnel destination via underlay); MTU drops (pings work, HTTPS stalls on some sites, fix with ip mtu 1400 and ip tcp adjust-mss 1360); keepalive flap with IPsec (keepalives stop returning at IPsec rekey, fix by enabling DPD and checking rekey vs keepalive timers); underlay reachability lost (tunnel down even though running-config looks fine, check show ip route for the tunnel destination); ACL blocking GRE (works in lab, fails in production, permit IP protocol 47 between endpoints)](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/failure-modes.png) The five GRE failure modes that cover the vast majority of real-world tickets. Symptom on top, fix at the bottom of each card. ## The Full GRE Cluster, in Reading Order ![GRE cluster reading-order map: the GRE Tunnels Complete Guide pillar at the centre, with nine articles connected as spokes - Tunnel Configuration on Cisco IOS XE, GRE over IPsec, Tunnel Keepalives, MTU and Fragmentation, GRE vs IPsec which tunnel, Routing Protocols over GRE, mGRE and DMVPN intro, Troubleshooting Guide, and GRE on Linux](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/cluster-map.png) The full cluster at a glance. Pillar at the centre, nine deep-dive articles as spokes, dashed lines for cross-cluster cousins. The articles in this cluster, in the order they make sense to read: 1. [GRE Tunnel Configuration: Step-by-Step Cisco IOS-XE Lab](https://www.pinglabz.com/gre-tunnel-configuration-cisco/) 2. [GRE over IPsec: When and How to Combine Them](https://www.pinglabz.com/gre-over-ipsec/) 3. [GRE Tunnel Keepalives Explained](https://www.pinglabz.com/gre-tunnel-keepalives/) 4. [GRE MTU and Fragmentation: Fixing Tunnel Packet Loss](https://www.pinglabz.com/gre-tunnel-mtu/) 5. [GRE vs IPsec vs GRE-over-IPsec: Which Tunnel Type?](https://www.pinglabz.com/gre-vs-ipsec/) 6. [Routing Protocols Over GRE: OSPF, EIGRP, BGP](https://www.pinglabz.com/routing-protocols-over-gre/) 7. [mGRE and DMVPN Introduction](https://www.pinglabz.com/mgre-dmvpn-introduction/) 8. [GRE Tunnel Troubleshooting Guide](https://www.pinglabz.com/gre-tunnel-troubleshooting/) 9. [GRE on Linux: ip tunnel add Commands](https://www.pinglabz.com/gre-on-linux/) This list grows as the cluster expands. Bookmark this page; the cluster is the canonical PingLabz reference for GRE. Hands-on GRE - build a real tunnel Configure a GRE Tunnel0 between two Cisco routers sourced from loopbacks, with R2 as a transparent transit. Capture the single-hop traceroute that defines an overlay. Part of the 12-lab CCNA Network Fundamentals cluster. Open the PingLabz CCNA Labs library. [Open the GRE lab](https://www.pinglabz.com/ccna-labs-network-fundamentals/) ## Frequently Asked Questions ### What is the GRE protocol number? GRE is IP protocol 47, assigned by IANA. That is the value in the Protocol field of the outer IP header. GRE is not a TCP or UDP protocol, so it has no port number. If a firewall in the path filters by TCP/UDP port and does not have an explicit allow for IP protocol 47, GRE packets get dropped at that hop. This is the single most common reason a GRE tunnel works in lab and fails in production. ### Is a GRE tunnel secure? No. GRE has no encryption, no authentication, and no integrity protection. Anyone who can capture packets on the underlay path can read every byte of the inner traffic. That is why production deployments wrap GRE inside IPsec (or run an alternative like WireGuard or an SD-WAN overlay). GRE alone is appropriate for trusted networks (private MPLS, internal data center) but not for the public internet. ### Is GRE a VPN? GRE is a tunnel, which is one of the components a VPN is built from, but GRE alone is not a VPN in the security sense. A VPN, as the term is normally used, implies confidentiality (encryption) and authentication. GRE provides neither. GRE-over-IPsec is a VPN. Plain GRE is a tunnel that happens to look topologically like a VPN. ### GRE vs IPsec - which one should I use? If you need encryption and only need to carry unicast IP, use IPsec alone (in tunnel mode). If you need to carry multicast (so OSPF / EIGRP can form neighbors over the tunnel) or you need to carry non-IP protocols, use GRE. If you need encryption and multicast support, use GRE over IPsec. The decision tree is at [GRE vs IPsec vs GRE-over-IPsec: Which Tunnel Type?](https://www.pinglabz.com/gre-vs-ipsec/). ### Is GRE a Cisco proprietary protocol? No. Cisco invented GRE and was the first vendor to ship it, but the protocol was standardized as RFC 1701 in 1994 and updated as RFC 2784 in 2000\. RFC 2890 added optional Key and Sequence Number extensions. Every major networking vendor (Juniper, Arista, Linux, FortiGate, MikroTik, Palo Alto) implements GRE as an interoperable standard. Some specific extensions, like Cisco GRE keepalives, are vendor-specific and may not interoperate cleanly across vendors. ### How much overhead does GRE add to each packet? Plain GRE over IPv4 adds 24 bytes per packet: 20 bytes of outer IPv4 header plus 4 bytes of GRE header. Enabling Key, Sequence, or Checksum adds 4 bytes per field. GRE over IPv6 has 40 bytes of outer IPv6 header plus 4 bytes of GRE for 44 bytes total. Add IPsec in tunnel mode and you lose another 50-60 bytes depending on the cipher. That is why `ip mtu 1400` on the tunnel interface is the safe default - it leaves headroom for any encryption you add later. ## Key Takeaways GRE is the simplest possible tunnel: an outer IP header, a 4-byte GRE header, and your original packet. It carries multicast over unicast underlays, runs routing protocols across paths that would otherwise not support them, and gives you a clean Layer 3 next-hop between two endpoints regardless of how complicated the underlay is. It has been a network-engineer staple for three decades because it does one thing well, almost never breaks, and works the same way on every vendor. If you take one thing away from this guide, take this: GRE without IPsec is a private-network tool, GRE with IPsec is the production tunnel of choice, and the operational pain of GRE always reduces to either MTU or recursive routing. Set `ip mtu 1400`, run `keepalive 10 3`, route the tunnel destination statically through the underlay, and you have eliminated 80 percent of the failure modes before you have a single user complaint. Bookmark this page, work through the cluster articles in order, and lab every design decision before you ship it. GRE is simple. Production GRE done well is the result of caring about the parts that look simple. **Studying for the CCNA?** Test your GRE knowledge on [PingLabz CCNA Flashcards](https://www.pinglabz.com/ccna-flashcards/) \- 200 free multiple-choice questions by topic, mixed, or a full mock exam, each with a plain-English explanation. ### References - [RFC 2784 - Generic Routing Encapsulation (GRE)](https://www.rfc-editor.org/rfc/rfc2784?ref=pinglabz.com) - [Cisco GRE technology documentation](https://www.cisco.com/c/en/us/tech/ip/generic-routing-encapsulation-gre/index.html?ref=pinglabz.com) ### Blog URL: https://www.pinglabz.com/blog/ Last updated: 2026-05-09T03:40:33.000Z _No content available._ ### Topics URL: https://www.pinglabz.com/tags/ Last updated: 2026-08-02T22:39:39.000Z _No content available._ ## Posts ### Native VLAN Mismatch: Read the CDP Log, Fix the Trunk URL: https://www.pinglabz.com/native-vlan-mismatch-troubleshooting/ Last updated: 2026-08-01T19:26:51.000Z Very few Cisco log messages hand you the interface and both sides of the misconfiguration in a single line. `%CDP-4-NATIVE_VLAN_MISMATCH` does, which is why it is at once the easiest Layer 2 fault to fix and one of the most commonly ignored. It is severity 4, it never takes the trunk down, and it repeats forever, so it gets filtered out of the syslog view and forgotten. That is a mistake, because of what is happening underneath the warning. A native VLAN mismatch does not black-hole a link in any obvious way. It quietly stitches two VLANs together across the trunk, giving you an unauthorised Layer 2 path between two broadcast domains that your firewall and your router ACLs never see. This article is the symptom-to-fix path: what the message says, what is happening to frames, why the warning can vanish while the fault stays, and how to close it out. If you are still assembling the bigger picture of [how VLANs and 802.1Q trunks carry traffic across a switched network](https://www.pinglabz.com/vlans-layer-2-switching/), start there and come back. One sentence of orientation and no more: the native VLAN is the single VLAN whose frames cross an 802.1Q trunk untagged, and if you want the full treatment of [why untagged frames exist on a tagged link at all](https://www.pinglabz.com/native-vlan/), that is a separate read. Everything below was captured on a live lab, two `ioll2-xe` switches running IOS XE 17.18.2 in Cisco Modeling Labs, trunked back to back with deliberately mismatched native VLANs. ## The message, and how to read the two VLAN numbers Here is the real thing, straight off the console of SW1\. The timestamp has been trimmed from the front of each line, which is why the leading asterisk is left dangling: ``` SW1: *CDP-4-NATIVE_VLAN_MISMATCH: Native VLAN mismatch discovered on Ethernet0/0 (1), with SW2 Ethernet0/0 (99). *CDP-4-NATIVE_VLAN_MISMATCH: Native VLAN mismatch discovered on Ethernet0/0 (1), with SW2 Ethernet0/0 (99). (repeats at the CDP interval: 19:56:21, 19:57:21, 19:58:18, 19:59:17, 20:00:15, 20:01:13 ...) ``` The message is self-diagnosing. Read it as four fields: - `on Ethernet0/0` is the **local** interface, on the switch that printed the line. - `(1)` immediately after it is the **local** native VLAN. That is your side. - `with SW2 Ethernet0/0` is the neighbour device name and its interface, learned from CDP. - `(99)` is the **neighbour's** native VLAN. So SW1 is running native VLAN 1 on Et0/0 and SW2 is running native VLAN 99 on the other end of the same cable. The parenthesised numbers are the entire investigation, and the only question left is which of the two your standard says it should be. Two operational details. Both switches log it, each with its own number first, so a central syslog shows two complementary messages for what is one fault. And look at the timestamps: roughly every 60 seconds, the CDP advertisement interval, forever. This is not a transient you can wait out. In a noisy buffer, `show logging | include NATIVE_VLAN_MISMATCH` pulls it straight out. ## What is actually broken: two VLANs bridged into one The warning is cosmetic. The forwarding behaviour is not. On an 802.1Q trunk every frame carries a VLAN tag except the ones in the native VLAN, which are sent bare. The receiving switch has no tag to read, so it applies its own rule: an untagged frame arriving on a trunk belongs to *my* native VLAN. When the two ends disagree, that rule silently rewrites VLAN membership at the trunk boundary. In the lab above, a frame in VLAN 1 leaves SW1 untagged. SW2 receives it, sees no tag, and places it into VLAN 99\. In the return direction, a VLAN 99 frame leaves SW2 untagged and SW1 drops it into VLAN 1\. The tagged VLANs on that trunk are unaffected and keep working perfectly, which is exactly why this failure is so good at hiding. Nothing goes down. What you get instead is: - **Leakage.** VLAN 1 on SW1 and VLAN 99 on SW2 are now one broadcast domain, mixing ARP, DHCP, broadcasts and unicast flooding. If either was your management or quarantine VLAN, it is no longer isolated. - **Black-holing.** A host in the native VLAN cannot reach its gateway across the trunk, because its frames arrive in the wrong VLAN and the gateway SVI never hears them. It presents as "one VLAN cannot reach anything past the trunk" while every other VLAN is fine. - **Address confusion.** Two subnets sharing one broadcast domain gives you ARP entries that will not resolve and DHCP offers from a scope that has no business being there. Spanning tree can react too, depending on the mode. Per-VLAN spanning tree sends the native VLAN's BPDUs untagged, so they land in the wrong instance on the far end, and PVST+ can spot the port-VLAN-ID inconsistency and block the port for the affected VLANs to contain the damage (MST shares one BPDU per region and will not raise the same per-VLAN complaint). If a trunk misbehaves for one VLAN and behaves for the rest, [how spanning tree builds a separate topology per VLAN](https://www.pinglabz.com/spanning-tree-protocol/) explains why the blast radius is VLAN-scoped rather than link-scoped. In this lab run, the only thing that complained on the console was CDP. ## The trap: CDP detects it, CDP does not cause it This is the part that catches people, and it is the reason to treat the message as a symptom rather than as the problem itself. The mismatch is a pure configuration disagreement between two `switchport trunk native vlan` statements, and it exists whether or not anything notices. CDP carries the native VLAN in its advertisements, so a Cisco switch can compare what its neighbour claims against its own setting and raise the alarm. That detection is a courtesy, not an enforcement mechanism, which means: - Run `no cdp run` globally, or `no cdp enable` on that interface, and the messages stop immediately while the two VLANs stay bridged. Hardening guides routinely recommend disabling CDP at the edge, so plenty of networks have already switched off their own smoke detector. - Trunk to a non-Cisco switch, a hypervisor uplink, or a firewall doing 802.1Q, and there is no CDP relationship to do the comparison. IOS will not raise an equivalent alarm from LLDP, so mismatches on those links get found by users, not by logs. - Silence after a change is not proof of a fix. It is proof that nothing is currently telling you. Verify the configuration, not the absence of a log line. ## The fix Confirm both ends first. The mismatch message already gave you the two numbers, but on a real trunk you want to see the running state, not just CDP's opinion of it: ``` show interfaces trunk ! compare the "Native vlan" column on both ends show cdp neighbors detail ! CDP reports the neighbor's native VLAN show run interface Ethernet0/0 | include native ``` Then pick the value your standard mandates (which for most networks is a dedicated, unused VLAN and not VLAN 1) and set it on the end that is wrong: ``` interface Ethernet0/0 switchport trunk native vlan 99 ``` Two things to get right while you do it. The VLAN you choose must exist on both switches and be in the trunk's allowed list, otherwise you have swapped a mismatch for a pruned native VLAN, which is a different and more confusing failure. And if you are working remotely, know your management path before you type, because between changing the first end and the second you are the one causing the mismatch. Standardising the native VLAN across an estate is a design decision rather than an incident action, so [moving the native VLAN off VLAN 1 without dropping the trunk](https://www.pinglabz.com/change-native-vlan-cisco-switch/) is worth planning properly rather than doing at 2am. The other legitimate fix is to stop having an untagged VLAN at all: ``` vlan dot1q tag native ``` This tags native VLAN traffic on egress instead of sending it bare, removing the ambiguity the whole problem depends on. It is global, and only safe if you apply it on **both** ends and every device sharing those trunks. Applying it on one side converts a mismatch you can see into an asymmetry you cannot. ## Verify 1. `show interfaces trunk` on both switches, and the Native vlan column reads the same number. This is the check that actually matters. 2. Wait out two CDP intervals, about two minutes, and confirm no new NATIVE\_VLAN\_MISMATCH entries. Note the timestamp of the last occurrence before you change anything so you can tell old messages from new. 3. Test a host in the native VLAN across the trunk, gateway included. That is the traffic that was broken, so that is the traffic that proves it is fixed. ## The same primitive an attacker uses on purpose A native VLAN mismatch is an accidental version of a documented attack. In a double-tagging VLAN hop, the attacker sits on an access port whose VLAN matches the trunk's native VLAN and sends a frame carrying two 802.1Q tags. The first switch strips the outer tag on the way onto the trunk, because that VLAN goes untagged, and the second switch reads the inner tag and delivers the frame into a VLAN the attacker was never allowed to touch. The attacker is deliberately engineering the exact condition your trunk has stumbled into: an untagged frame that means one thing on one switch and something else on the other. Treat the log line accordingly. It is a security finding as much as a connectivity one, and it belongs with the other Layer 2 trust failures covered in [how double tagging and switch spoofing let a host cross VLANs](https://www.pinglabz.com/vlan-hopping-attacks/). If two VLANs are bridged at a trunk, your segmentation model is wrong on that path whether or not anyone is currently abusing it. ## What this was captured on PlatformCisco Modeling Labs, ioll2-xe switches SoftwareIOS XE 17.18.2 TopologySW1 Et0/0 to SW2 Et0/0, 802.1Q trunk SW1 native VLAN1 (left at the default) SW2 native VLAN99 DetectionCDP, logged on both switches every \~60s CapturedBroken state, via native console logging ## Gotchas **On IOL-L2 you need the encapsulation line first.** `ioll2-xe` does accept `switchport trunk encapsulation dot1q`, and needs it before `switchport mode trunk` will take. On modern hardware that command is often gone because dot1q is the only option, so lab and production muscle memory differ here. **Everyone misreads which number is theirs.** The first parenthesised VLAN belongs to the switch that printed the line, not to the neighbour. Fix the wrong end and you have simply moved the mismatch. **Only the broken state was captured in this run.** The `ioll2-xe` image has no EEM and the PyATS console path was unavailable, so the post-fix silence was not machine-driven here. Run the verification steps above rather than trusting anyone's screenshot of a quiet log, this one included. **Severity 4 gets filtered.** In an aggregated syslog carrying thousands of events a minute, a level-4 message that repeats every 60 seconds forever is exactly the kind of thing that ends up in a suppression rule. Alert on this mnemonic explicitly. ## Key Takeaways - `%CDP-4-NATIVE_VLAN_MISMATCH` names both interfaces and both native VLANs. The first number in parentheses is the local switch's, the second is the neighbour's. - The damage is not the log noise. Untagged frames get re-homed at the trunk boundary, so two VLANs share one broadcast domain while every tagged VLAN keeps working normally. - CDP only detects the condition. Disable CDP, or trunk to a non-Cisco device, and the alarm goes away while the bridged VLANs remain. - Fix it with a matching `switchport trunk native vlan` on both ends, or `vlan dot1q tag native` on every device sharing those trunks. - Verify with `show interfaces trunk` on both sides plus a traffic test in the native VLAN. Silence in the log is not verification. - Treat it as a security event: a mismatch is the accidental form of the double-tagging VLAN hop and defeats segmentation the same way. Native VLAN mismatches sit in the small group of Layer 2 faults that are trivially fixable and disproportionately dangerous, alongside VTP revision accidents and unpruned trunks. For the rest of that group, and the order to learn them in, work through the [Layer 2 switching and trunking guide](https://www.pinglabz.com/vlans-layer-2-switching/). ### Errdisable Recovery on Cisco: Every Cause and How to Bring the Port Back URL: https://www.pinglabz.com/errdisable-recovery-cisco/ Last updated: 2026-08-01T19:26:50.000Z An err-disabled port is a switch telling you it made a decision. The link is fine, the cable is fine, the far end is happily transmitting, and the switch has decided that whatever is happening on that port is bad enough that no traffic should pass. Nothing comes back on its own either, because on a default Cisco switch nothing is configured to bring it back. That surprises people constantly: the fault cleared hours ago and the port is still down. This article is about two things. First, the cause taxonomy: everything that can put a port into err-disable on a modern IOS XE switch, and how to work out in about ten seconds which one fired. Second, the recovery mechanism itself: what `errdisable recovery cause` and the recovery interval actually do, what they do not do, and when leaving auto-recovery switched off is the right engineering call. If port behaviour at Layer 2 is still settling for you, the pillar on [how VLANs and Layer 2 switching work on a Cisco switch](https://www.pinglabz.com/vlans-layer-2-switching/) is the background this assumes. Everything below was captured on `ioll2-xe` switches running IOS XE 17.18.2 in Cisco Modeling Labs, using two separate breaks: a port-security violation on SW1 Ethernet0/3 driven by a deliberately wrong static MAC, and a BPDU guard trip on SW1 Ethernet0/0 from a neighbouring switch. The structured output was pulled with pyATS and Genie 26.6\. No output on this page was typed by hand. ## What this article covers, and what two neighbours cover instead Err-disable sits at a junction of three topics, so it is worth drawing the lines before we start. If your port went down because of a spanning-tree protection feature, the detail on [how BPDU guard, root guard and loop guard decide to shut a port](https://www.pinglabz.com/troubleshoot-errdisable-stp-guard/) belongs to that article, and it is the right place to go for how to design those guards in the first place. If you arrived from the platform administration side, [how SDM templates change what a switch can hold](https://www.pinglabz.com/switch-administration-sdm-errdisable/) covers that angle. What stays here is the part neither of those owns: the complete cause list, identifying which one fired on the port in front of you, the mechanics of the recovery timer, and whether to arm it. ## Err-disable is a state, not an error counter When a protection feature trips, the port manager moves the interface into err-disable and the port goes down in both directions. From the interface itself it looks like this: ``` SW1# show interfaces Ethernet0/3 | include line protocol|reset Ethernet0/3 is down, line protocol is down (err-disabled) 0 output errors, 0 collisions, 1 interface resets ``` That parenthetical `(err-disabled)` is the whole distinction. A cable fault gives you `down/down` with no qualifier. A shut port gives you `administratively down`. The qualifier means software took the port down on purpose, which also means no amount of re-seating fibre or swapping patch leads is going to help you. The port also stays down indefinitely. It is not a hold-down and not a penalty that decays. Absent configuration to the contrary, it is permanent. ## The full cause list is on the box You do not need to memorise which features can err-disable a port, because the switch will tell you. `show errdisable recovery` lists every cause the running image supports, and on IOS XE 17.18.2 that is 29 of them: ``` SW1# show errdisable recovery ErrDisable Reason Timer Status ----------------- -------------- arp-inspection Disabled bpduguard Disabled channel-misconfig Disabled dhcp-rate-limit Disabled dtp-flap Disabled evpn-mh-core-isolation Disabled gbic-invalid Disabled inline-power Disabled l2ptguard Disabled link-flap Disabled mac-limit Disabled link-monitor-failure Disabled loopback Disabled loopdetect Disabled oam-remote-failure Disabled pagp-flap Disabled port-mode-failure Disabled pppoe-ia-rate-limit Disabled psecure-violation Disabled security-violation Disabled sfp-config-mismatch Disabled storm-control Disabled udld Disabled unicast-flood Disabled vmps Disabled psp Disabled dual-active-recovery Disabled evc-lite input mapping fa Disabled mrp-miscabling Disabled Timer interval: 300 seconds Interfaces that will be enabled at the next timeout: ``` Read that output twice. Every single cause says `Disabled`, and that is the factory default. The 300 second timer interval is real but it is not doing anything, because no cause is armed to use it, and the empty list at the bottom confirms nothing is queued. This is the most common err-disable surprise in production: engineers assume a five minute self-heal exists, and it does not until somebody configures it. Twenty-nine causes is a lot to reason about individually, so group them by what they protect against. The family tells you whether auto-recovery is even a sensible idea. Security enforcement Causespsecure-violation, arp-inspection, dhcp-rate-limit TriggerA device broke a policy Auto-recoverUsually no Loop protection Causesbpduguard, loopdetect, loopback, l2ptguard TriggerTopology is not what you declared Auto-recoverNo Negotiation mismatch Causeschannel-misconfig, dtp-flap, pagp-flap, port-mode-failure TriggerTwo ends configured differently Auto-recoverPointless until fixed Physical and transceiver Causeslink-flap, gbic-invalid, sfp-config-mismatch, inline-power, udld TriggerHardware or cabling Auto-recoverOften sensible Rate and volume Causesstorm-control, unicast-flood, mac-limit, pppoe-ia-rate-limit TriggerA threshold was crossed Auto-recoverYes, with a long interval Your platform's list is authoritative, not the list in any article including this one. `evpn-mh-core-isolation`, `mrp-miscabling` and `psp` are present on 17.18.2 and absent from older images, and `evc-lite input mapping fa` is a truncated label rather than a typo. ## Identify which cause fired Here is a real port-security break. SW1 Ethernet0/3 faces a router. Port security is configured with `maximum 1` and a static MAC of `0000.dead.beef`, which is not the router's MAC, so the first frame the router sends is a violation. Before the break, the switch is clean: ``` SW1# show interfaces status Port Name Status Vlan Duplex Speed Type Et0/0 connected 1 full auto 10/100/1000BaseTX Et0/1 connected 1 full auto 10/100/1000BaseTX Et0/2 connected 1 full auto 10/100/1000BaseTX Et0/3 connected 1 full auto 10/100/1000BaseTX ``` One OSPF hello later: ``` SW1# show interfaces status Port Name Status Vlan Duplex Speed Type Et0/0 connected 1 full auto 10/100/1000BaseTX Et0/1 connected 1 full auto 10/100/1000BaseTX Et0/2 connected 1 full auto 10/100/1000BaseTX Et0/3 err-disabled 1 full auto 10/100/1000BaseTX <-- here ``` `show interfaces status` tells you which port, but not why. Add the filter and you get the reason column, which is the single most useful command on this page: ``` SW1# show interfaces status err-disabled Port Name Status Reason Err-disabled Vlans Et0/3 err-disabled psecure-violation ``` The word in the Reason column is not descriptive text. It is the exact keyword you feed to `errdisable recovery cause`. That mapping is the reason the command is worth running before you touch anything else. The syslog says the same thing at the moment it happens, in a different shape. From the BPDU guard break on the other lab port: ``` *Jul 20 13:02:52.871: %SPANTREE-2-BLOCK_BPDUGUARD: Received BPDU from bridge aabb.cc00.9600 on port Et0/0 with BPDU Guard enabled. Disabling port. *Jul 20 13:02:52.871: %PM-4-ERR_DISABLE: bpduguard error detected on Et0/0, putting Et0/0 in err-disable state *Jul 20 13:02:53.871: %LINEPROTO-5-UPDOWN: Line protocol on Interface Ethernet0/0, changed state to down *Jul 20 13:02:54.872: %LINK-3-UPDOWN: Interface Ethernet0/0, changed state to down ``` Two lines carry the diagnosis. The feature-specific message (`%SPANTREE-2-BLOCK_BPDUGUARD` here, and a port security equivalent for the other break) tells you what was detected and often names the offender, in this case the bridge MAC `aabb.cc00.9600`. Then `%PM-4-ERR_DISABLE` from the port manager confirms the action, and the word immediately before "error detected" is once again your recovery cause keyword. Reason column in show interfaces status err-disabledExact recovery cause keyword Word before "error detected" in %PM-4-ERR\_DISABLEExact recovery cause keyword Feature message above itNames the offending device or frame Timer Status in show errdisable recoveryWhether it will ever come back on its own Once you know the family, confirm it at the feature. For port security that means the per-interface counters, which give you a second independent confirmation and, usefully, the MAC that caused it: ``` SW1# show port-security interface Ethernet0/3 Port Security : Enabled Port Status : Secure-shutdown Violation Mode : Shutdown Maximum MAC Addresses : 1 Total MAC Addresses : 1 Configured MAC Addresses : 1 Sticky MAC Addresses : 0 Last Source Address:Vlan : aabb.cc00.dd00:1 Security Violation Count : 1 ``` `Secure-shutdown` is port security's own name for the same condition, and `Last Source Address` is the MAC that tripped it. In the lab that is the router's own address, which is the point: the violating device is frequently something legitimate that simply is not the MAC you pinned. In production this is where you find out somebody swapped a NIC. If you monitor rather than eyeball, all of this parses cleanly. Genie turns the same two commands into assertable structure, so a check can look for `parsed['interfaces']['Ethernet0/3']['status'] == 'err-disabled'` and read the reason as a value: ``` { "interfaces": { "Ethernet0/3": { "reason": "psecure-violation", "status": "err-disabled" } } } ``` ## Recovery path one: bounce it yourself Manual recovery is two commands and it always works: ``` SW1(config)# interface Ethernet0/3 SW1(config-if)# shutdown SW1(config-if)# no shutdown ``` The `shutdown` is not optional. `no shutdown` alone does not clear err-disable, because the port is not administratively down, so you have to move it into admin-down first to give the port manager a state to leave. Do the config fix before the bounce, or you are just going to watch it trip again. ## Recovery path two: arm the timer Auto-recovery is per cause, plus one global interval: ``` SW1(config)# errdisable recovery cause psecure-violation SW1(config)# errdisable recovery interval 30 ``` Re-run the show command and the table changes in two places. The cause flips to `Enabled`, the interval updates, and a queue appears at the bottom listing ports waiting to be brought back: ``` SW1# show errdisable recovery ErrDisable Reason Timer Status ----------------- -------------- bpduguard Disabled link-flap Disabled psecure-violation Enabled <-- armed storm-control Disabled udld Disabled Timer interval: 30 seconds Interfaces that will be enabled at the next timeout: Interface Errdisable reason Time left(sec) --------- ----------------- -------------- Et0/3 psecure-violation 299 ``` Look at the last two lines against the interval above them. The header says 30 seconds, the queued port says 299 seconds left. That is not a bug and it matters operationally: the countdown for Et0/3 was armed when the cause was enabled, off the old 300 second default, and changing `errdisable recovery interval` does not re-arm a port that is already queued. If you are in an outage and you shorten the interval expecting the port back in half a minute, you will be waiting nearly five. Bounce the port manually instead, and let the new interval apply to the next trip. The interval itself is exact once it is in effect. On the BPDU guard port, configured with `interval 30` from the start, the recovery attempts land at 13:03:22, 13:03:52, 13:04:22 and 13:04:52\. Thirty seconds apart, every time. ## When auto-recovery is the wrong answer Here is what those BPDU guard recovery attempts actually looked like, with the neighbouring switch still connected and still sending BPDUs: ``` *Jul 20 13:03:22.860: %PM-4-ERR_RECOVER: Attempting to recover from bpduguard err-disable state on Et0/0 *Jul 20 13:03:22.867: %SPANTREE-2-BLOCK_BPDUGUARD: Received BPDU ... on port Et0/0 ... Disabling port. *Jul 20 13:03:22.867: %PM-4-ERR_DISABLE: bpduguard error detected on Et0/0, putting Et0/0 in err-disable state *Jul 20 13:03:52.858: %PM-4-ERR_RECOVER: Attempting to recover from bpduguard err-disable state on Et0/0 ``` Recovered at 13:03:22.860, err-disabled again at 13:03:22.867\. Seven milliseconds of uptime, then straight back down, and the cycle repeats forever. Auto-recovery did exactly what it was told and achieved nothing, because it re-enables a port without knowing why the port went down. That is the whole argument in one capture. The recovery timer is a retry loop, not a repair. It is useful when the condition is genuinely transient (a storm that has passed, a link that flapped during a UPS transfer, an SFP that settled) and it is actively harmful when the condition is a device or a config that has not changed. Now apply that to `psecure-violation` specifically, because it is the cause people most often arm and most often should not. Port security exists to stop an unauthorised device using a port. If you enable auto-recovery on it with a five minute interval, you have told the switch to hand that unauthorised device a fresh attempt every five minutes, indefinitely, with nobody ever being asked about it. The security control still logs, but it no longer denies. A port that stays down until someone looks at it is not a failure of the design, it is the design. There is a second cost, and it is the one that bites at 3am. A port cycling every 30 or 300 seconds generates continuous link up and down events, which drive MAC table churn and spanning-tree topology changes across the domain. A permanently down port is a clean, quiet, single alarm. A flapping port is a moving target that fills your logs and, on a switch with any topology at all, makes downstream [ports run the listening and learning timers](https://www.pinglabz.com/stp-port-states-explained/) over and over. Flapping is often worse for the network than being down. A workable default: arm auto-recovery for `link-flap`, `storm-control`, `udld` and `inline-power`, keep the interval generous (the 300 second default is generous for a reason), and leave `psecure-violation`, `bpduguard` and the negotiation mismatch causes disabled so a person has to make a decision. ## What this was captured on Two breaks on CML, both on `ioll2-xe` switches running IOS XE 17.18.2\. The BPDU guard trip used SW1 Ethernet0/0 configured as an access port with `spanning-tree portfast` and `spanning-tree bpduguard enable`, cabled to SW2 running default spanning tree, with `errdisable recovery cause bpduguard` and `interval 30` in place to capture the flap loop. The port-security trip used SW1 Ethernet0/3 facing a router, with `switchport port-security`, `maximum 1`, `violation shutdown` and a static `mac-address 0000.dead.beef` that deliberately does not match the router. Structured output came from pyATS and Genie 26.6 on Python 3.13 against the live device. ## Gotchas - **Nothing auto-recovers by default.** Every cause shows `Disabled` on a fresh switch. The 300 second timer interval in the output is real but inert until you arm at least one cause. - **Changing the interval does not re-arm a queued port.** The capture shows `Timer interval: 30 seconds` alongside a port with 299 seconds left, because it was queued under the old default. - **Auto-recovery re-trips within milliseconds if the cause persists.** Recovered at 13:03:22.860, err-disabled at 13:03:22.867. - **`no shutdown` alone does not clear err-disable.** You need the `shutdown` first. - **The Genie parser key is misleading.** `show errdisable recovery` parses the global interval into a field named `bpduguard_timeout_recovery`, regardless of which cause is armed. If you assert on that key expecting a BPDU guard specific value, you are reading the global timer. - **`ioll2-xe` has no EEM.** `event manager` is rejected as invalid input, so on-box scripted capture of err-disable events is not available on that CML node type. Use console syslog or drive it from pyATS off-box. ## Key takeaways - Err-disable is a deliberate software state, shown as `(err-disabled)` on the interface line, and it is permanent by default. - `show interfaces status err-disabled` gives you the port and the reason, and that reason word is the exact keyword for `errdisable recovery cause`. - `show errdisable recovery` is the authoritative cause list for your image (29 causes on IOS XE 17.18.2) and tells you whether anything will come back on its own. - Manual recovery is `shutdown` then `no shutdown`, after the root cause is fixed, not before. - The recovery timer is a retry loop, not a repair. With the cause still present it produced a port that flapped every 30 seconds indefinitely. - Auto-recovery on `psecure-violation` converts a security control into a rate limiter on the attacker. Leave it off and let a human clear it. Err-disable is one of the places where the switch is being more careful than the person configuring it, and the recovery mechanism is the point where you decide how much of that care to keep. For where port security, guard features and STP all sit relative to each other, the [Layer 2 switching fundamentals every switch feature builds on](https://www.pinglabz.com/vlans-layer-2-switching/) ties the cluster together. ### MAB with FreeRADIUS on Cisco: MAC Authentication Bypass Without ISE URL: https://www.pinglabz.com/mab-freeradius-cisco-lab/ Last updated: 2026-08-01T19:26:50.000Z Every 802.1X rollout stalls on the same list of devices: the label printer in the warehouse, the ceiling cameras, the badge readers at the door, the ancient PLC nobody is allowed to touch. None of them run a supplicant, none of them will ever send an EAPOL frame, and on a port with `authentication port-control auto` they are simply dead. MAC Authentication Bypass is the escape hatch, and if you understand [how 802.1X port authentication works on Cisco switches](https://www.pinglabz.com/802-1x/) you are already 90 percent of the way to understanding MAB, because MAB is what happens when that process gives up. There is already a detailed article on this site covering [MAB with Cisco ISE, endpoint identity groups and profiling](https://www.pinglabz.com/mab-configuration-cisco-ios-xe-ise/). If you have ISE, read that one instead: profiling and endpoint groups are genuinely better tooling than anything below. This article owns the other path, the one most engineers actually have available on a Tuesday afternoon, which is FreeRADIUS on a Debian box and a switch. No licence, no appliance, no ISE node. Everything here was built and captured that way. The captures come from a CML lab: a Catalyst 9000v (`cat9000v-uadp`) running IOS XE 17.18.02 as the authenticator, FreeRADIUS 3.2.7 on a Debian 13 VM at 192.168.99.100 running in the foreground with `freeradius -X`, and a headless `net-tools` container as the endpoint with no supplicant at all. The payoff is a real `Method: mab / Authc Success` on a real switchport, and the RADIUS transaction that produced it, printed from both ends. ## MAB is not 802.1X, it is PAP with a MAC in it The single most useful thing to internalise about MAB is that it does not use EAP. 802.1X is an EAP transport: the supplicant and the RADIUS server run a full EAP method between them and the switch just relays EAPOL frames. MAB has no supplicant to talk to, so the switch fabricates an ordinary username and password request out of the source MAC address it saw on the wire, and sends that. Plain PAP. You do not have to take that on faith, because the server prints its own decision. This is the policy trail from `freeradius -X` as it handles a MAB request: ``` (4) Received Access-Request Id 176 ... User-Name = "525400031ace" (4) [preprocess] = ok (4) suffix: No '@' in User-Name = "525400031ace", looking up realm NULL (4) eap: No EAP-Message, not doing EAP <-- MAB is NOT EAP: bare PAP, no supplicant (4) files: users: Matched entry 525400031ace at line 222 (4) Found Auth-Type = PAP (4) pap: Comparing with "known good" Cleartext-Password ``` `eap: No EAP-Message, not doing EAP` followed immediately by `Found Auth-Type = PAP` is the whole distinction in two lines. For contrast, the same server handling a genuine 802.1X supplicant (driven with `eapol_test`, since the lab endpoint cannot do EAP) negotiates a full tunnelled method: ``` CTRL-EVENT-EAP-STARTED EAP authentication started CTRL-EVENT-EAP-PROPOSED-METHOD vendor=0 method=4 -> NAK CTRL-EVENT-EAP-PROPOSED-METHOD vendor=0 method=25 CTRL-EVENT-EAP-METHOD EAP vendor 0 method 25 (PEAP) selected EAP: Status notification: remote certificate verification (param=success) CTRL-EVENT-EAP-SUCCESS EAP authentication completed successfully MPPE keys OK: 1 mismatch: 0 SUCCESS ``` Same server, same users file, two completely different authentication paths. PEAP builds a TLS tunnel, verifies a server certificate and derives keying material. MAB compares a string to another string. Hold that thought for the security section at the end. ## The FreeRADIUS side: two files This is where the open-source path is genuinely simpler than the ISE path. There is no endpoint database, no identity group, no profiling engine. There is `clients.conf`, which says who is allowed to ask, and `users`, which says who is allowed in. ``` ! ---- /etc/freeradius/3.0/clients.conf : who may send RADIUS ---- client pinglabz-lab-subnet { ipaddr = 192.168.99.0/24 secret = PingLabzRAD123 nas_type = cisco shortname = pinglabz-lab } ``` That secret has to match the switch's `key` byte for byte. A mismatch shows up server-side as `Received packet from ... with invalid Message-Authenticator`, which is the most common failure in any RADIUS build and worth recognising on sight. If your switch is not getting answers at all, the diagnostic path for a [RADIUS server that the switch thinks is dead](https://www.pinglabz.com/radius-server-unreachable-802-1x-cisco-ios-xe/) is the same whether the server is ISE or FreeRADIUS. Then the identity itself. For MAB, the identity is the MAC address, used as both username and password: ``` ! ---- /etc/freeradius/3.0/users : the identities ---- ! 802.1X supplicant: alice Cleartext-Password := "PingLabz1x" Reply-Message = "RES-0004 dot1x accept for alice" ! MAB - Cisco sends the endpoint MAC as BOTH username and password, ! lowercase, no separators (default 'ietf' MAB format on IOS): 525400031ace Cleartext-Password := "525400031ace" Reply-Message = "RES-0015 MAB accept for H1 by MAC" ``` That is the entire MAB "policy". One line per device you are willing to let on. It is crude next to ISE, and it is also the reason you can stand this up in ten minutes. Before touching the switch, prove the server in isolation with `radtest`, passing the MAC as both credentials exactly as the switch will: ``` ! Command: radtest 525400031ace 525400031ace 127.0.0.1 0 testing123 Sent Access-Request Id 176 ... User-Name = "525400031ace" User-Password = "525400031ace" Received Access-Accept Id 176 from 127.0.0.1:1812 to 127.0.0.1:59405 length 73 Reply-Message = "RES-0015 MAB accept for H1 by MAC" ``` ## The MAC format gotcha that bites everyone Look closely at that username: `525400031ace`. The endpoint's MAC is `52:54:00:03:1a:ce`. Cisco's default MAB attribute format on IOS XE (the `ietf` format) is lowercase hex with no separators at all, and it is sent as the User-Name and the User-Password. If your `users` entry says `52:54:00:03:1a:ce`, or `5254.0003.1ace`, or `525400031ACE`, the file module simply does not match it and the server rejects. You will get an Access-Reject with no obvious reason, because nothing is wrong except capitalisation. What makes this genuinely confusing is that a single MAB Access-Request carries the same MAC in two different formats at once. Here is the request the switch actually sent for the live endpoint, MAC `52:54:00:f0:93:48`: ``` (2) Received Access-Request Id 3 from 192.168.99.12:61802 to 192.168.99.100:1812 length 291 (2) User-Name = "525400f09348" (2) User-Password = "525400f09348" (2) Service-Type = Call-Check <-- the MAB service-type (2) Cisco-AVPair = "service-type=Call Check" (2) Cisco-AVPair = "method=mab" <-- switch declares it is MAB (2) Cisco-AVPair = "audit-session-id=0C63A8C00000000E969AF179" (2) NAS-IP-Address = 192.168.99.12 (2) NAS-Port-Id = "GigabitEthernet1/0/2" (2) Calling-Station-Id = "52-54-00-F0-93-48" <-- the endpoint MAC (2) Called-Station-Id = "52-54-00-13-E9-16" <-- the switch port MAC ``` The User-Name is lowercase and unseparated. The Calling-Station-Id, in the very same packet, is uppercase and hyphen-separated. People copy the MAC out of a debug, paste it into the users file, and pick the wrong one of the two. Match the **User-Name** format, always, because that is the attribute the `files` module keys on. (Other vendors default to different formats, and IOS XE can be told to send others, which is exactly why "the MAC format" is a support ticket generator across every NAC product.) Two other things in that packet are worth knowing by heart: `Service-Type = Call-Check` and the Cisco AVPair `method=mab`. Those are the on-wire fingerprint of MAB. If you ever need to write a server-side policy that treats MAB differently from 802.1X, Call-Check is what you key on. ## The switch side: MAB is a fallback, and that is why it feels slow On the port, MAB is one command. The interesting part is the method order around it: ``` interface GigabitEthernet1/0/2 switchport mode access authentication port-control auto authentication order dot1x mab ! try 802.1X first, then MAB authentication priority dot1x mab authentication host-mode single-host authentication violation restrict dot1x pae authenticator dot1x timeout tx-period 5 ! how long each 802.1X probe waits dot1x max-reauth-req 2 ! probes before falling back to MAB mab spanning-tree portfast ``` `authentication order dot1x mab` means the switch tries 802.1X first every single time, even on a port where you know full well the device is a printer. It sends EAP-Request/Identity, waits `tx-period` seconds, retransmits, waits again, and only when the silence has gone on for `tx-period` multiplied by `max-reauth-req` does it give up and start MAB. With the values above that is about 15 seconds; with the IOS defaults it is closer to 30\. This is the answer to the most common MAB complaint, which is "the camera takes half a minute to come online after a reboot". Nothing is broken. You are watching 802.1X time out. `authentication priority dot1x mab` is the other half and it is a security control, not a timer. It says that if an EAPOL frame ever shows up on a port that MAB has already authorised, 802.1X wins and preempts the MAB session. Without it, someone who unplugs the printer and plugs in a laptop inherits the printer's authorisation. If a port is dedicated to a non-supplicant device forever, you can flip to `authentication order mab dot1x` and cut the wait to a couple of seconds. Keep `priority dot1x mab` when you do. ## The payoff: Method mab, Authc Success With the endpoint plugged into Gi1/0/2 and its MAC in the users file, this is what the switch reports. This is the output the whole article is built toward, and it tells the story in four lines: ``` SW9K# show authentication sessions interface GigabitEthernet1/0/2 details Interface: GigabitEthernet1/0/2 MAC Address: 5254.00f0.9348 User-Name: 525400f09348 Status: Authorized Domain: DATA Oper host mode: single-host Common Session ID: 0C63A8C00000000E969AF179 Method status list: Method State dot1x Stopped <-- 802.1X tried first, no supplicant, timed out mab Authc Success <-- fell back to MAB, RADIUS said yes ``` The method status list is the part to read. `dot1x Stopped` then `mab Authc Success` is `authentication order dot1x mab` doing exactly what it was told: 802.1X was attempted, the endpoint stayed silent, MAB took over and the RADIUS server said yes. If you are ever handed a port that will not come up, this list is the first thing to look at, and the wider technique for [reading show authentication sessions when a port stays unauthorized](https://www.pinglabz.com/802-1x-troubleshooting-show-authentication-sessions-debug/) applies unchanged to MAB. Server-side, the same transaction completes like this: ``` (2) eap: No EAP-Message, not doing EAP (2) files: users: Matched entry 525400f09348 at line 226 (2) Found Auth-Type = PAP (2) pap: User authenticated successfully (2) Sent Access-Accept Id 3 from 192.168.99.100:1812 to 192.168.99.12:61802 length 83 (2) Reply-Message = "RES-0015 live MAB accept for H9 on cat9000v" ``` And once the port is authorised, the switch pins the endpoint's MAC: ``` SW9K# show mac address-table interface GigabitEthernet1/0/2 1 5254.00f0.9348 STATIC Gi1/0/2 <-- pinned STATIC once authorized ``` A ping from the endpoint to the switch SVI came back 3 of 3 after authorisation. Before it, the port was unauthorised and the endpoint was cut off entirely. That is the control working. ## Which switch image actually implements MAB This cost more lab time than the entire FreeRADIUS build, so take it as given. Three common virtual switch images, three different answers, and only one of them is usable: ioll2-xe CLI acceptedNo Behaviour% Invalid input VerdictHonest rejection iosvl2 CLI acceptedYes BehaviourNever authenticates VerdictSilent no-op, avoid cat9000v-uadp CLI acceptedYes BehaviourFull authenticator VerdictUse this one The `iosvl2` case is the nasty one, because the configuration is accepted into the running config without a murmur and then nothing happens. No Access-Request ever leaves the box. `debug dot1x events` during an interface bounce explains why: ``` dot1x-ev:[Gi0/1] Interface state changed to UP dot1x-ev:DOT1X Supplicant not enabled on GigabitEthernet0/1 dot1x-ev:[Gi0/1] No DOT1X subblock found for port down ``` Build on `cat9000v-uadp`, or on real 9200/9300 hardware. If you want the full authenticator build (AAA, the RADIUS server block, the global `dot1x system-auth-control`) rather than just the MAB delta shown above, the [complete 802.1X and FreeRADIUS lab build](https://www.pinglabz.com/802-1x-freeradius-cisco-full-lab/) is the prerequisite for this one and covers the base configuration line by line. ## Gotchas from the build - **The port boots err-disabled with a security violation.** The default dot1x violation action is shutdown, so the very first frame from an unauthorised endpoint on a `single-host` port can kill the port before MAB ever runs: `%PM-4-ERR_DISABLE: security-violation error detected on Gi1/0/2`. Add `authentication violation restrict` (and `errdisable recovery cause security-violation`), then bounce the port. - **Wrong MAC format, silent reject.** Lowercase, no separators, matching the User-Name in the Access-Request. Not the Calling-Station-Id form. - **The first interface on a Catalyst 9000v is not a switchport.** Gi0/0 is a management port in `Mgmt-vrf`. Cabling your external connector there (the obvious first-slot wiring) puts your RADIUS path in the wrong VRF and nothing reaches the server. Cable a real `Gi1/0/x` switchport instead. - **After `aaa new-model`, your priv-15 user is not priv-15.** Add `aaa authorization exec default local` or SSH sessions land at `SW>` and automation fails to enter enable mode. On the virtual 9000v specifically, do not set an `enable secret` at all if you are driving the console with PyATS. - **Do not skip the `radtest` step.** Proving the users entry locally before you involve the switch turns a two-variable problem (server policy, switch config) into two one-variable problems. ## The honest part: MAB is inventory control, not authentication A MAC address is not a secret. It is printed on a sticker on the back of the device, it is broadcast in every frame the device sends, and any laptop can be made to claim it in one command. The captures above show why that matters at the protocol level: 802.1X ran PEAP, built a TLS tunnel, verified a certificate and derived keys. MAB compared `525400f09348` to `525400f09348`. Anyone who can read a sticker or run a sniffer for thirty seconds has the credential. So do not describe MAB to your security team as authentication, because it is not. It is an allow list of known hardware, and its real value is that it forces you to have an inventory at all. Treat every MAB port as a port that a determined attacker owns, and design around that: put MAB endpoints in their own VLAN, give them a restrictive dACL or port ACL, keep `authentication priority dot1x mab` so a supplicant preempts a spoofed session, and use `single-host` mode where the device genuinely never changes. Layer 2 controls matter more here than anywhere else, so pair it with the protections that stop a host from [poisoning ARP or claiming an IP that is not its own](https://www.pinglabz.com/dynamic-arp-inspection-ip-source-guard/). And if you have ISE, its profiling engine at least cross-checks that the thing claiming to be a printer behaves like a printer, which bare FreeRADIUS cannot do. ## Key Takeaways - MAB carries no EAP. The switch turns the source MAC into a plain PAP username and password, which is exactly why a device with no supplicant can pass it, and `eap: No EAP-Message, not doing EAP` in `freeradius -X` is the proof. - The username format is lowercase hex with no separators (`525400f09348`). The Calling-Station-Id in the same packet is uppercase and hyphenated. Match the User-Name. - `authentication order dot1x mab` means 802.1X is always tried first, so a MAB endpoint waits `tx-period` times `max-reauth-req` seconds before it is authorised. That delay is the design, not a fault. - `Method: mab / Authc Success` after `dot1x Stopped` in `show authentication sessions` is the confirmation that fallback fired, and `Service-Type = Call-Check` plus `method=mab` is the same event seen from the RADIUS server. - Lab it on `cat9000v-uadp` or real hardware. `ioll2-xe` rejects the commands and `iosvl2` accepts them and does nothing. - MAB is an allow list of known hardware, not authentication. Scope its access accordingly, and see the rest of the [wired port authentication cluster](https://www.pinglabz.com/802-1x/) for the controls that go around it. ### 802.1X with FreeRADIUS on Cisco: The Full Lab, No ISE Required URL: https://www.pinglabz.com/802-1x-freeradius-cisco-full-lab/ Last updated: 2026-08-01T19:26:49.000Z Almost every 802.1X tutorial on the internet opens the same way: "first, log in to ISE." That is fine if you work somewhere that already owns a Cisco ISE deployment and can spare a policy set to experiment in. It is useless if you are studying at your kitchen table, because ISE is a licensed appliance with an appetite for RAM that no home lab is going to satisfy. The result is that the one security feature running on practically every enterprise access port stays theoretical for the people who most need to practise it. You do not need ISE. 802.1X is an IEEE standard and RADIUS is an open protocol, and FreeRADIUS speaks both perfectly well. This article builds the whole chain from nothing: FreeRADIUS 3.2.7 on Debian 13, a Catalyst 9000v running IOS-XE 17.18.02 as the authenticator, and a real endpoint on a real access port. Every command works, and every block of output below was copied off that lab, including the moment the port flips to `Authorized`. If you want the theory of [how port-based network access control fits together](https://www.pinglabz.com/802-1x/) first, the pillar covers it; this piece is the build. One warning before you spend an evening on it. If you lab this in CML, the switch image you pick decides whether any of it works, and two of the three obvious choices fail. That section comes first, deliberately. ## Read this before you pick a switch image 802.1X is not a config you can practise on whatever L2 node you happen to have in your topology. I tested the same interface configuration on three CML images and got three different behaviours, only one of which is any use. `ioll2-xe`, the everyday IOL-XE switch most people reach for, simply does not have the port authenticator. The global `dot1x system-auth-control` command is accepted, which is what makes this so misleading, but every interface-level command is thrown out: ``` SW(config-if)# authentication port-control auto ^ % Invalid input detected at '^' marker. SW(config-if)# dot1x pae authenticator ^ % Invalid input detected at '^' marker. SW(config-if)# mab ^ % Invalid input detected at '^' marker. ``` `iosvl2` is worse, because it lies. The exact commands IOL-XE rejected are all accepted here without a murmur, they show up in `show running-config`, and nothing whatsoever happens. Wire a live endpoint to the port, bounce it, and you get this: ``` SW# show dot1x all summary Interface PAE Client Status Gi0/1 AUTH none UNAUTHORIZED SW# show authentication sessions interface Gi0/1 No sessions match supplied criteria. ! debug dot1x events during the interface bounce: dot1x-ev:[Gi0/1] Interface state changed to UP dot1x-ev:DOT1X Supplicant not enabled on GigabitEthernet0/1 dot1x-ev:[Gi0/1] No DOT1X subblock found for port down ``` The RADIUS server never receives a single Access-Request from that switch. The CLI parses into the running config but the authenticator state machine was never implemented, which is a silent no-op and considerably nastier than the honest rejection IOL-XE gives you. You will assume your shared secret is wrong, or your users file, or your supplicant, and you will be wrong about all three. `cat9000v-uadp` is the one that works. It accepts the full stack, saves it, and the boot log confirms the machinery actually came up: ``` %AAA-6-METHOD_LIST_STATE: authen mlist ... of DOT1X service ... current state is : ALIVE %AAA-6-RADIUS_SERVER_KEY_UPDATE: Radius server:PLZ key is updated %RADIUS-4-NON_TLS_SERVER_CONFIGURED: RADIUS server PLZ is configured without tls/dtls ``` ioll2-xe Interface CLIRejected AuthenticatorNone Access-Requests sentZero VerdictFails loudly iosvl2 Interface CLIAccepted AuthenticatorNever starts Access-Requests sentZero VerdictFails silently cat9000v-uadp Interface CLIAccepted AuthenticatorRuns Access-Requests sentYes VerdictUse this one Use `cat9000v-uadp`, or real hardware, or a 9200/9300 if you can borrow one. Nothing else in a typical CML refplat set will authenticate a port. And note the interface naming, because it matters later: the Catalyst 9000v has real switchports at `GigabitEthernet1/0/1` through `1/0/24`, not IOL's `Ethernet0/x`. ## The three roles, and which box is which 802.1X only has three participants and every troubleshooting session goes faster if you keep them straight. - **Supplicant**: software on the endpoint that speaks EAP over LAN. In this lab that is `eapol_test` from wpa\_supplicant 2.10 on the Debian VM. On your desk it is the native Windows or macOS 802.1X client. - **Authenticator**: the switch. It is a relay and a gate, nothing more. It wraps the endpoint's EAP frames into RADIUS packets, forwards them, and enforces whatever the server answers. It does not know or check any credentials itself. - **Authentication server**: RADIUS. Here, FreeRADIUS on Debian. This is where identity actually lives, and it is the box ISE would otherwise be. Two protocols on two different wires. EAP over LAN (EAPOL) runs between the endpoint and the switch, and RADIUS runs between the switch and the server. That split is why a shared-secret problem and a credentials problem look completely different when you go looking for them. If the vocabulary here is new, the [basic dot1x port configuration on a Cisco switch](https://www.pinglabz.com/802-1x-configuration-on-cisco-switches-a-practical-lesson/) is the shorter concept primer to read alongside this. ## Build the RADIUS server Install `freeradius` and `freeradius-utils` on Debian, then stop the service so you can run the daemon in the foreground where it prints every packet and every policy decision. That debug output is the single most useful diagnostic tool in the whole exercise. ``` sudo systemctl stop freeradius # release UDP 1812/1813 sudo freeradius -X # debug mode, full protocol decode ss -lunp | grep 181 # expect 0.0.0.0:1812 and 0.0.0.0:1813 ``` Note the binary name. On Debian 13 it is `/usr/sbin/freeradius`, and the `radiusd -X` command you will find in every older guide does not exist. Same daemon, different packaging. Next, tell the server which NAS devices are allowed to talk to it. Anything not listed in `clients.conf` gets ignored outright, which is a failure mode that produces total silence rather than a reject. In `/etc/freeradius/3.0/clients.conf`: ``` client pinglabz-lab-subnet { ipaddr = 192.168.99.0/24 secret = PingLabzRAD123 nas_type = cisco shortname = pinglabz-lab } ``` A subnet is convenient in a lab, but in production you list each switch by address. The `secret` has to match the switch's `key` character for character. Then the identities, in `/etc/freeradius/3.0/users`. Two entries, because this lab proves two paths: ``` ! 802.1X supplicant alice Cleartext-Password := "PingLabz1x" Reply-Message = "RES-0004 dot1x accept for alice" ! MAB: Cisco sends the MAC as BOTH username and password, ! lowercase, no separators (default IETF format on IOS) 525400031ace Cleartext-Password := "525400031ace" Reply-Message = "RES-0015 MAB accept for H1 by MAC" ``` Flat-file identities are a lab shortcut, obviously. Point the `files` module at LDAP or SQL later; the switch side does not change at all when you do, which is rather the point of RADIUS. ## Prove the server before you touch a switch Do not wire anything up yet. `radtest` lets you generate Access-Requests locally and confirm the server's decisions in isolation, so that when the switch does start sending packets you already know the server half is sound. Four commands cover every decision the deployment needs to get right. ``` radtest alice PingLabz1x 127.0.0.1 0 testing123 Sent Access-Request Id 181 from 0.0.0.0:35149 to 127.0.0.1:1812 length 75 User-Name = "alice" User-Password = "PingLabz1x" Received Access-Accept Id 181 from 127.0.0.1:1812 to 127.0.0.1:35149 length 71 Reply-Message = "RES-0004 dot1x accept for alice" ``` That is the packet a switch needs to see before it will open a port. Now break it, twice, in two different ways: ``` radtest alice WrongPassword 127.0.0.1 0 testing123 Received Access-Reject Id 48 from 127.0.0.1:1812 to 127.0.0.1:39339 length 71 radtest ghost whatever 127.0.0.1 0 testing123 Received Access-Reject Id 213 from 127.0.0.1:1812 to 127.0.0.1:56137 length 38 ``` Both are rejects, but look at the lengths. The wrong-password reject is 71 bytes and the unknown-user reject is 38\. The difference is the `Reply-Message`: FreeRADIUS sets it during the `authorize` stage, which happens *before* the password is checked, so a user who exists but fails the password still gets their reply attribute echoed back on the reject. The ghost user matched no entry at all, so there is nothing to echo. Read the packet type, never the reply message. Plenty of people have stared at `Reply-Message = "... accept for alice"` inside an Access-Reject and concluded the server was broken. ## Configure the switch The authenticator config splits into three parts: AAA, the RADIUS server definition, and the port. Global first: ``` aaa new-model aaa authentication login default local aaa authorization exec default local ! or your priv-15 user drops to priv-1 aaa authentication dot1x default group radius aaa authorization network default group radius ! dot1x system-auth-control ! the global on-switch for 802.1X ! radius server PLZ address ipv4 192.168.99.100 auth-port 1812 acct-port 1813 key PingLabzRAD123 ! ip radius source-interface Vlan1 radius-server attribute 6 on-for-login-auth radius-server attribute 8 include-in-access-req ``` `ip radius source-interface` is not optional housekeeping. The source address of the Access-Request is what FreeRADIUS matches against `clients.conf`, and if the switch sources from an interface you did not authorise, the server drops the packet without answering. Pin it to the SVI you actually listed. Then the port itself: ``` interface GigabitEthernet1/0/2 switchport mode access authentication order dot1x mab authentication priority dot1x mab authentication port-control auto authentication violation restrict mab dot1x pae authenticator dot1x timeout tx-period 5 spanning-tree portfast ``` `authentication port-control auto` is the line that arms the port. Without it you have a fully configured authenticator that never challenges anybody. `authentication order dot1x mab` says try 802.1X first and fall back to MAC Authentication Bypass, which is what handles printers and cameras that have no supplicant. `dot1x timeout tx-period 5` shortens the EAPOL retry interval so the fallback happens in seconds instead of a minute and a half, which makes a lab far less tedious. ## A real EAP exchange, end to end 802.1X *is* EAP, and it helps to watch a full handshake once before you start reading switch counters. The lab endpoint had no supplicant, so I drove a real PEAP/MSCHAPv2 conversation at FreeRADIUS with `eapol_test` instead, using the `alice` identity: ``` eapol_test -c eap.conf -a 127.0.0.1 -s testing123 -r0 CTRL-EVENT-EAP-STARTED EAP authentication started CTRL-EVENT-EAP-PROPOSED-METHOD vendor=0 method=4 -> NAK CTRL-EVENT-EAP-PROPOSED-METHOD vendor=0 method=25 CTRL-EVENT-EAP-METHOD EAP vendor 0 method 25 (PEAP) selected CTRL-EVENT-EAP-PEER-CERT depth=0 subject='/CN=j-deb-vm' EAP: Status notification: remote certificate verification (param=success) CTRL-EVENT-EAP-SUCCESS EAP authentication completed successfully MPPE keys OK: 1 mismatch: 0 SUCCESS ``` Read it as a negotiation. The server proposed method 4 (EAP-MD5), the supplicant NAKed it, and both sides settled on method 25 (PEAP). A TLS tunnel came up, the supplicant verified the server certificate, MSCHAPv2 ran inside the tunnel, and the exchange ended with EAP-SUCCESS and derived MPPE keys. On a wired port the switch relays exactly these EAP messages between EAPOL frames and RADIUS attributes without inspecting them. This is also where an EAP method mismatch bites in the real world: if the endpoint only offers EAP-TLS and the server has no certificate configuration for it, you get a failed negotiation, not a rejected password, and the two look nothing alike in the logs. ## The port goes Authorized Here is the moment the whole exercise exists for. A net-tools endpoint (MAC `52:54:00:f0:93:48`) plugged into `Gi1/0/2` on the Catalyst 9000v, with FreeRADIUS running in debug on the Debian VM: ``` SW9K# show authentication sessions interface GigabitEthernet1/0/2 details Interface: GigabitEthernet1/0/2 MAC Address: 5254.00f0.9348 User-Name: 525400f09348 Status: Authorized Domain: DATA Oper host mode: single-host Common Session ID: 0C63A8C00000000E969AF179 Method status list: Method State dot1x Stopped <-- 802.1X tried first, nobody answered mab Authc Success <-- fell back to MAB, RADIUS said yes ``` That method list is the entire story of a mixed access port. 802.1X was attempted, the endpoint had no supplicant to answer with, the authenticator gave up after its retries, and MAB took over. This is `authentication order dot1x mab` behaving exactly as designed, and it is why the fallback exists at all: [what happens when the supplicant never answers](https://www.pinglabz.com/mab-freeradius-cisco-lab/) is a whole topic of its own. On the server side, this is the switch's own Access-Request, not a `radtest` simulation: ``` (2) Received Access-Request Id 3 from 192.168.99.12:61802 to 192.168.99.100:1812 length 291 (2) User-Name = "525400f09348" (2) User-Password = "525400f09348" (2) Service-Type = Call-Check (2) Cisco-AVPair = "service-type=Call Check" (2) Cisco-AVPair = "method=mab" (2) Cisco-AVPair = "audit-session-id=0C63A8C00000000E969AF179" (2) NAS-IP-Address = 192.168.99.12 (2) NAS-Port-Id = "GigabitEthernet1/0/2" (2) Calling-Station-Id = "52-54-00-F0-93-48" (2) Called-Station-Id = "52-54-00-13-E9-16" (2) eap: No EAP-Message, not doing EAP (2) files: users: Matched entry 525400f09348 at line 226 (2) Found Auth-Type = PAP (2) pap: User authenticated successfully (2) Sent Access-Accept Id 3 from 192.168.99.100:1812 to 192.168.99.12:61802 length 83 (2) Reply-Message = "RES-0015 live MAB accept for H9 on cat9000v" ``` Everything you need for troubleshooting is in that request. `NAS-IP-Address` tells you which switch asked, `NAS-Port-Id` names the exact interface, `Calling-Station-Id` is the endpoint MAC and `Called-Station-Id` is the switch port's own. And one line separates the two authentication paths cleanly: `eap: No EAP-Message, not doing EAP`, followed by `Found Auth-Type = PAP`. MAB is a bare PAP request the switch fabricates from the source MAC, with the MAC as both username and password. A genuine 802.1X request carries an `EAP-Message` attribute and `Service-Type = Framed` instead. Same server, same users file, two entirely different conversations, and knowing which one you are looking at saves an hour. Once authorized, the switch pins the learned address: ``` SW9K# show mac address-table interface GigabitEthernet1/0/2 1 5254.00f0.9348 STATIC Gi1/0/2 ``` And traffic passes, which it did not before. A ping from the endpoint to the switch SVI at 192.168.99.12 returned 3 packets transmitted, 3 received, 0% packet loss. Before authorization the port was unauthorized in single-host mode and the endpoint was cut off completely. Worth remembering that MAC-based authorization is only as trustworthy as the MAC, so pair it with [stopping spoofed ARP and source addresses on an authorized port](https://www.pinglabz.com/dynamic-arp-inspection-ip-source-guard/) before you rely on it for anything sensitive. ## What this was captured on Authentication serverFreeRADIUS 3.2.7 on Debian 13 (trixie), 192.168.99.100 AuthenticatorCatalyst 9000v (cat9000v-uadp), IOS-XE 17.18.02, SVI Vlan1 192.168.99.12 Supplicanteapol\_test (wpa\_supplicant 2.10) for EAP; a net-tools endpoint for the port Wiringbridge1 to Gi1/0/1 (RADIUS path), endpoint to Gi1/0/2 (dot1x + MAB) Also testedioll2-xe and iosvl2, both unusable for 802.1X ## The three ways this breaks **The shared secret does not match.** This is the most common 802.1X-to-RADIUS failure, and it is invisible from the switch: you get no reply, the port times out, and nothing suggests a key problem. On the server, running in debug, a mismatched key shows up as a complaint about an invalid Message-Authenticator on the received packet. Diagnose it from the server side, always. And check the switch's source address against `clients.conf` at the same time, because an unlisted NAS produces the same symptom of total silence for a completely different reason. **The supplicant never answers.** Very common and usually not a fault. The endpoint has no 802.1X client, or the client is disabled, or a user never entered credentials. The authenticator retries EAPOL-Identity-Request on its `tx-period`, gives up, and you see `dot1x Stopped` in the method list. What happens next depends entirely on your config: with MAB configured, the port falls back and authorizes by MAC; without it, the port stays shut. That is precisely why MAB exists, and if you run ISE rather than FreeRADIUS the [MAB config for IOS-XE against ISE](https://www.pinglabz.com/mab-configuration-cisco-ios-xe-ise/) covers the same fallback with a different policy engine. **The port stays unauthorized.** Work outward from the switch. `show authentication sessions interface X details` tells you whether a session exists at all; no session means the authenticator never started, which on a lab image usually means you picked the wrong platform. A session that reaches the server and comes back rejected is a credentials or EAP-method problem, and the FreeRADIUS debug trail names it exactly. A session where the switch sends and nothing comes back is a transport or shared-secret problem. Three symptoms, three separate places to look, and the debug output distinguishes them instantly. ## Gotchas that cost me time - **The `aaa new-model` privilege trap.** After `aaa new-model`, a local user with `privilege 15` gets dropped to privilege 1 on VTY unless you also configure `aaa authorization exec default local`. Symptom: SSH lands you at `SW>` instead of `SW#`, and automation dies with "Failed to enter enable mode". It looks exactly like a credentials problem and is not. This bit both lab switches. - **The dot1x port boots err-disabled.** The default violation action is shutdown, so `Gi1/0/2` came up with `%PM-4-ERR_DISABLE: security-violation error detected on Gi1/0/2` before anything was even plugged in. `authentication violation restrict` plus `errdisable recovery cause security-violation` fixes it, then shut/no-shut to recover the port. - **Cable a switchport, not the management port.** The Catalyst 9000v's first data interface is `Gi0/0`, a dedicated management port in `Mgmt-vrf`. Wire your external connector there (the obvious first-interface choice) and the bridge lands on the management VRF, so nothing reaches your SVI or your access port. Use a `Gi1/0/x` switchport instead. This wasted two lab rebuilds. - **Skip `enable secret` on cat9000v in CML.** Setting one makes PyATS-driven console access fail with "Bad Password". With no enable secret and `aaa authorization exec default local` in place, admin lands straight at a priv-15 prompt. - **It is `freeradius -X`, not `radiusd -X`,** on Debian 13\. Ten minutes lost to a command that every tutorial written before 2020 still tells you to run. ## Key takeaways - 802.1X does not require ISE. FreeRADIUS plus a Catalyst switch gives you a complete, working authenticator chain you can build at home in an evening. - In CML, only `cat9000v-uadp` implements the port authenticator. `ioll2-xe` rejects the CLI outright and `iosvl2` accepts it and silently does nothing, which is the trap that wastes people's time. - Prove the server with `radtest` before you configure a single switch. It cleanly separates "my identities are wrong" from "my switch cannot reach RADIUS". - Read the RADIUS packet type, not the reply message: a `Reply-Message` set during authorize is echoed on an Access-Reject too. - One debug line separates the two paths. `eap: No EAP-Message, not doing EAP` plus `Auth-Type = PAP` means MAB; an `EAP-Message` attribute means real 802.1X. - `dot1x Stopped` followed by `mab Authc Success` in the method list is a silent supplicant handled correctly, not a failure. The rest of the picture (host modes, guest and critical VLANs, dynamic VLAN assignment, RADIUS redundancy and phased rollout) builds on exactly this chain, just with more policy at the server end. Work through [the full 802.1X cluster](https://www.pinglabz.com/802-1x/) once you have this lab authorizing a port, because every one of those features is easier to reason about after you have watched an Access-Accept arrive and a port flip state. ### NAT and PAT Troubleshooting on Cisco IOS XE: An Empty Translation Table URL: https://www.pinglabz.com/nat-pat-troubleshooting-cisco-ios-xe/ Last updated: 2026-08-01T18:57:58.000Z The ticket says the 192.168.1.0/24 subnet cannot reach the internet. You log into the border router and the configuration looks correct: there is an `ip nat inside source list 1 interface Ethernet0/1 overload` line, access list 1 permits the right subnet, there is a default route out. Nothing is obviously wrong. And nothing is being translated. This article starts at that moment, not before it. It is not a configuration walkthrough. It assumes you already have NAT configured, that it is not working, and that you now have to work backwards from a broken translation to the reason. If you are still building the configuration, the prerequisite is separate: go and [set up inside source NAT and PAT overload from scratch](https://www.pinglabz.com/nat-cisco-ios-xe/) first, then come back here when it refuses to translate. For where address translation sits alongside DHCP, NTP and the rest, the [IP services configuration and troubleshooting](https://www.pinglabz.com/ip-services/) pillar is the map. Everything below was captured on a live CML lab running `iol-xe` on IOS XE 17.18.2: R1 as the inside host, R2 doing PAT, R3 as the outside world holding 8.8.8.8, with the break baked in deliberately and repaired on the box while the captures ran. ## One command splits the problem in two Before you read a single line of the NAT configuration, run `show ip nat translations` while the failing traffic is actually flowing. That output puts you in one of exactly two worlds, and the two worlds have almost nothing in common. **No entry at all.** The router is not translating this flow. Either the packet is not arriving, or it is arriving and NAT is declining to act on it. Nothing downstream of the router matters yet, because nothing downstream has been given a chance to fail. **An entry exists, and traffic still fails.** NAT is doing its job. The translation was built, the packet left with a public source, and the failure is past the translation point: return routing, a filter on the far side, or an asymmetric path that bypasses the NAT router on the way back. Most wasted hours come from skipping that split and rereading the NAT configuration regardless of which world you are in. Here is world one, captured on R2 with R1 pinging 8.8.8.8 from 192.168.1.1 every fifteen seconds: ``` R2# show ip nat translations (empty) R2# show ip nat statistics Total active translations: 0 (0 static, 0 dynamic; 0 extended) Hits: 0 Misses: 0 <-- traffic is flowing and NAT has seen none of it ``` Note what makes that reading trustworthy: the pings were live at the time. An empty table on an idle router tells you nothing at all. ## No entry: the traffic never matched Three things must be true before IOS XE builds a translation. The packet arrives on an interface marked `ip nat inside`, it leaves via one marked `ip nat outside`, and it matches the ACL, pool or route-map named in the `ip nat inside source` statement. Miss any one and you get silence, not an error. Check them in that order, because that is the order of how often they are the culprit. ### Check the interface roles first, and check them backwards The single most common cause of a NAT that never fires is the inside and outside tags: one missing, or the two on the wrong interfaces. It happens on renumbering, on interface swaps, on a config restore, and above all when someone adds a second WAN link and tags the new interface without touching the old one. Do not read this from `show running-config` top to bottom. Ask each interface directly: ``` show run interface Ethernet0/0 | include nat show run interface Ethernet0/1 | include nat ``` You want `ip nat inside` on the interface facing your private hosts and `ip nat outside` on the interface facing the internet, and you want to say out loud which physical interface each of those actually is. Reversed tags are the cruel version, because the configuration looks complete: both keywords present, both on interfaces, and nothing translated, because IOS XE translates on the inside-to-outside path and you have told it the internet is your inside network. In the lab the break was the simpler version of this: R2's Ethernet0/0, the interface facing R1, was never marked. One line repaired it. ``` R2(config)# interface Ethernet0/0 R2(config-if)# ip nat inside ``` A related trap: if the inside interface lives in a VRF, a plain global NAT statement will not pick that traffic up at all, so everything above looks correct while nothing translates. That is a different rule set, covered in [translating traffic that lives inside a VRF](https://www.pinglabz.com/vrf-aware-nat-ios-xe/). Grep the interface for a `vrf forwarding` line before going further down this branch. ### Then check what selects the traffic If both roles are correct, the next suspect is the selector. Pull the ACL and compare it to the source address of the traffic that is actually failing, not to the subnet you believe your users are on: ``` show ip access-lists 1 show run | include ip nat inside source ``` Things that bite here: the ACL permits 192.168.1.0 0.0.0.255 but the host sits on a secondary subnet outside that range; the wildcard mask is wrong by one bit; an edit has put a `deny` above the `permit`; or the NAT statement references list 1 while somebody has been carefully maintaining list 101\. With a route-map instead of an ACL, remember a route-map whose `match` clause points at a non-existent ACL matches nothing while looking perfectly reasonable in the config. The fast confirmation is the ACL's own hit counters. Run `show ip access-lists 1` twice, thirty seconds apart, with the failing traffic running. If the match count does not move, the traffic is not reaching that ACL and you are still in the interface-role branch. ### Then confirm the packet is even arriving If the roles are right, the ACL is matching and the table is still empty, stop assuming the packet reaches the router. Check the inside host's route to the destination and confirm the NAT router is genuinely the next hop, not something the host learned from a competing default. Being able to [read the routing table and work out which next hop actually wins](https://www.pinglabz.com/how-to-read-cisco-routing-table/) settles that in seconds, for both halves of the path. Interface counters close it out. If input packets on the inside interface are not incrementing while the host is pinging, the problem is upstream of NAT and you have been troubleshooting the wrong device. ## An entry exists and it still fails Now the other world. `show ip nat translations` lists your flow, the inside global address is what you expect, and the user still reports failure. NAT has done everything asked of it, so rereading its configuration cannot help you. Three things to check, in order. **Return routing to the inside global address.** The outside world has to be able to reach whatever you translated to. Translating to an interface address on a link the upstream device already routes for is usually fine. Translating to a pool address on a subnet nobody upstream knows about means the request leaves and the reply dies out there with no route home. On the lab topology R3 can reach 203.0.113.1 only because that address sits on the directly connected link. **A filter on the far side.** An ACL on the outside interface, an upstream firewall, or a policy on the destination itself. Because the far side sees the translated source, this is also where you discover a permit list that was written against the private address and never updated for the public one. The tell is translation counters climbing while nothing comes back, which is exactly what a silently discarded return packet looks like. **Asymmetry.** If the return path does not cross the same NAT router, that router's translation table is never consulted and the reply reaches the host with a public destination address it does not own. Redundant edge routers with independent NAT tables produce this reliably, and it presents as intermittent failure rather than total failure, which is what makes it unpleasant to chase. ## Hits, misses, and the fixed state `show ip nat statistics` is the second half of the diagnosis and it is routinely ignored. Hits count packets that matched an existing translation, misses count packets that needed a new one built. Zero for both under live traffic is the clean signature of a NAT configuration gap rather than a routing problem, because the router is not even attempting to translate. Here is the same lab after the one-line repair, with the pings still running: ``` R2# show ip nat translations Pro Inside global Inside local Outside local Outside global icmp 203.0.113.1:1024 192.168.1.1:9 8.8.8.8:9 8.8.8.8:1024 icmp 203.0.113.1:1025 192.168.1.1:10 8.8.8.8:10 8.8.8.8:1025 icmp 203.0.113.1:1026 192.168.1.1:11 8.8.8.8:11 8.8.8.8:1026 ... R2# show ip nat statistics Total active translations: 5 (0 static, 5 dynamic; 5 extended) Hits: 240 Misses: 0 ``` Read the columns. Inside local is the real private address, inside global is what the outside world sees, and here one public address is shared across many port numbers. That is overload, which is what PAT means, and why the entries are marked extended. For a known-good build to compare a suspect configuration against, [a working PAT overload lab you can rebuild in an evening](https://www.pinglabz.com/ccna-lab-ips-05-nat-overload-pat/) gives you one. Note the five active translations for one pinging host. Each ICMP echo gets its own entry keyed on the ICMP identifier, so a busy host inflates that count fast. Never read a large translation count as evidence of a large number of users. ## What this was captured on CML, three `iol-xe` nodes on IOS XE 17.18.2\. R1 is the inside host with Loopback0 at 192.168.1.1/24 and Ethernet0/0 at 10.0.12.1/30, defaulting to R2\. R2 is the PAT router: Ethernet0/0 at 10.0.12.2 facing inside, Ethernet0/1 at 203.0.113.1 facing outside, running `ip nat inside source list 1 interface Ethernet0/1 overload` with access list 1 permitting 192.168.1.0/24, and a static route to 8.8.8.8 via R3\. R3 is the outside, with Loopback0 as 8.8.8.8/32. The break was that R2's Ethernet0/0 was missing `ip nat inside`. Everything else was correct, which is what makes it a fair reproduction of the real ticket. Captures were taken on-box with EEM applets writing `show` output to syslog while a second applet generated the pings, so every command ran against genuinely live traffic. ## Gotchas **ICMP translations age out fast, so ping and show together.** This one produces phantom bugs. You ping, it fails, you run `show ip nat translations` a minute later, the table is empty and you conclude NAT is broken when the entry existed and simply timed out. Keep a repeating ping running in another window, or script the ping and the show as one action, and only then trust an empty table. **Zero hits and zero misses is a specific signal.** Rising misses with no successful translations points at a pool or an ACL problem. Zero of both means the packet is never being offered to NAT at all, which sends you straight to the interface roles rather than into the selector. **Both roles or nothing.** Marking only `ip nat outside` is the classic half-configuration, and it fails completely silently. There is no log message, no error at the CLI and no counter anywhere that says "I would have translated this if you had told me which side was inside". **Clearing translations is a diagnostic, not a fix.** `clear ip nat translation *` tears down stale state and can genuinely resolve a stuck entry, but if the table repopulates and the same flow fails again, you have only learned that the entries were not stale. ## Key takeaways - Run `show ip nat translations` while the failing traffic is live. No entry and an existing entry are two different investigations. - No entry sends you to the interface roles first. Missing or reversed `ip nat inside` and `ip nat outside` is the most common cause and it fails silently. - If the roles are right, check the ACL, pool or route-map that selects the traffic, and confirm its hit counters are moving. - An entry that exists means NAT worked. Look downstream: return routing to the inside global address, a filter on the far side, or an asymmetric return path. - `show ip nat statistics` at 0 hits and 0 misses under live traffic is the signature of a NAT configuration gap, not a routing problem. - ICMP entries expire quickly, so pair the ping with the show command or you will chase a table that emptied itself. NAT sits at the boundary where a routing problem, a filtering problem and a translation problem all look identical from the user's desk, which is why a fixed order of checks beats intuition. The [other edge services that fail this quietly](https://www.pinglabz.com/ip-services/) follow the same pattern: read the state table first, then decide which half of the path you are actually in. ### HSRP Troubleshooting: Both Routers Active and Flapping States URL: https://www.pinglabz.com/hsrp-troubleshooting-flapping-dual-active/ Last updated: 2026-08-01T18:57:58.000Z HSRP fails in two shapes that look nothing alike from the CLI. Either both routers are permanently convinced they are the Active gateway, or the pair cannot stop swapping roles. Both come back to the same question: can these two routers actually hear each other? The trap in dual-active is that each router looks completely healthy on its own. Active state, owns the virtual IP, interface up. Nothing says "broken" until you put the peer's output next to it. This article walks a real capture from a lab built to break that way, then covers flapping, the version and authentication mismatches that produce the same silence, and the preempt behaviour that inverts your design intent without ever logging an error. If the election mechanics are still fuzzy, the [guide to how first-hop redundancy protocols pick a gateway](https://www.pinglabz.com/fhrp/) is the background this assumes. Everything below was captured on two `iol-xe` routers running IOS XE 17.18.2 in Cisco Modeling Labs, sharing one Ethernet segment, with an inbound ACL on one router dropping HSRP hellos in one direction. No output on this page was typed by hand. ## The symptom: two routers, both Active, no Standby Here is the broken state, held long enough to prove it was not a transient: ``` R1# show standby brief Interface Grp Pri P State Active Standby Virtual IP Et0/0 1 110 P Active local unknown 10.0.0.254 R2# show standby brief Interface Grp Pri P State Active Standby Virtual IP Et0/0 1 100 P Active local unknown 10.0.0.254 ``` Both routers report `State = Active` and `Active = local`, both report `Standby = unknown`, and both advertise the same virtual IP. That `unknown` column is the whole diagnosis. A healthy HSRP group always knows its partner, so a blank there says this router has never received a single hello from the other side. Notice too that priority is doing nothing: R1 sits at 110 and R2 at 100, and yet both are Active. Priority is only meaningful between routers that can talk to each other. On the wire this is worse than an outage. Two routers answer ARP for the same virtual MAC, so the upstream switch relearns that MAC on whichever port sent the last frame and flips back and forth. Host traffic lands on one router or the other at random, and you get intermittent loss plus MAC flap messages that look like an L2 problem. ## Why a silent peer promotes itself HSRP is a hello-driven election with no tie-breaker outside the hello stream. A router coming up runs through Initial, Listen and Speak, and every decision on the way rests purely on the hellos it receives. Hear a superior Active router and you settle into Standby or Listen. Hear nothing and you conclude there is no gateway here, so you take the job. That is correct behaviour when the peer genuinely is dead. The problem is that "the peer is dead" and "I cannot receive the peer's packets" are indistinguishable from inside the protocol, so both routers reach the same conclusion at the same time. The state name tells you which half of the conversation is missing, so the breakdown of [what each HSRP state means and what moves a router between them](https://www.pinglabz.com/hsrp-state/) is worth keeping beside the CLI. HSRP hellos ride UDP 1985, sent to `224.0.0.2` in version 1 and `224.0.0.102` in version 2\. Anything that stops that traffic either way produces the capture above. ## What actually stops the hellos The search space is small and splits into two families: the hellos are not delivered at all, or they are delivered and rejected. The tells differ, so work out which family you are in first. Inbound ACL on the segment FamilyNot delivered TellSilence, no logs Checkshow access-lists VLAN missing from a trunk FamilyNot delivered TellPeer IP unreachable Checkshow int trunk STP blocking the only path FamilyNot delivered TellFollows a topology change Checkshow spanning-tree vlan Version 1 against version 2 FamilyWrong multicast group TellSilence, no logs Checkshow standby Authentication mismatch FamilyDelivered, rejected Tell%HSRP-4-BADAUTH Checkshow standby Et0/0 1 Group or VIP mismatch FamilyDelivered, ignored Tell%HSRP-4-DIFFVIP1 Checkshow run int | i standby A version mismatch is the cruellest, because it produces perfect silence: the two routers subscribe to different multicast groups, so neither ever sees a malformed packet to complain about. Version 2 also uses a different virtual MAC range and allows groups up to 4095, which is why a group above 255 will not configure under version 1\. An authentication mismatch is the friendliest, because the packets arrive and the receiver logs `%HSRP-4-BADAUTH` (version 1 defaults to the plaintext string `cisco`, so a router with a key string silently rejects one without). The group and VIP case splits in two. Same group with different virtual IPs gives you `%HSRP-4-DIFFVIP1`, which names the problem for you. Different groups with the same virtual IP is the quiet one: each router builds its own group with its own virtual MAC (the group number is encoded in the last byte), and both answer ARP for the same address with no log at all. Take the layer 2 causes seriously in a campus. A blocked STP port between two distribution switches, or a VLAN left out of `switchport trunk allowed vlan`, cuts the HSRP conversation as cleanly as an ACL does. That is why [the STP root bridge and the HSRP active router should be the same box for a given VLAN](https://www.pinglabz.com/stp-fhrp-hsrp-vrrp-alignment/): if they are not, a topology change is far more likely to isolate the pair. ## The fix, and the state machine that proves it The lab break was an inbound ACL on R2 dropping HSRP from R1\. Removing it recovered the pair in real time: ``` R2(config)# interface Ethernet0/0 R2(config-if)# no ip access-group BLOCK-HSRP in *HSRP-5-STATECHANGE: Ethernet0/0 Grp 1 state Active -> Speak *HSRP-5-STATECHANGE: Ethernet0/0 Grp 1 state Speak -> Standby ``` Those two lines are the protocol working as designed. The instant R2 received a hello from a router with priority 110, higher than its own 100, it gave up the Active role and dropped to Speak, where it advertises itself as a candidate, then settled into Standby as the best remaining candidate. No reload, no `clear`, and R1 was never touched. Confirm on both routers, never just one: ``` R2# show standby brief Interface Grp Pri P State Active Standby Virtual IP Et0/0 1 100 P Standby 10.0.0.1 local 10.0.0.254 R1# show standby brief Et0/0 1 110 P Active local 10.0.0.2 10.0.0.254 ``` The important change is not the state names. It is that each router now names the other: R2 lists `Active = 10.0.0.1`, R1 lists `Standby = 10.0.0.2`. Those peer addresses are your proof that hellos cross both ways. If either reads `unknown`, you still have a one-way path. ## Preempt is not "the highest priority wins" Look at the `P` column above. Both routers have preempt configured, which is why the higher-priority R1 holds the Active role. Take that `P` away and the behaviour changes in a way that catches people out. Priority does not decide who is Active. It decides who wins an election, and an election only happens when there is no Active router, or when a router with preempt enabled hears an Active router of lower priority than its own. Without `standby 1 preempt`, a router that is already Active is never displaced by a better candidate. It stays Active until it dies. So with no preempt, the first router to boot keeps the Active role permanently, even if it is the one you gave the lower priority. Reload your primary and the secondary takes over correctly, but when the primary returns it hears an Active hello, decides the segment already has a gateway, and quietly becomes Standby. Your config still reads `priority 110` on the box you think of as primary, and nothing is logged, because nothing is wrong. If you want a defined primary you must configure preempt, so [setting HSRP priority and preempt on an interface](https://www.pinglabz.com/hsrp-configuration/) is worth getting right once. ## Flapping is the same fault at a different duty cycle Stable dual-active means hellos never arrive. Flapping means they arrive sometimes, or that something keeps re-running the election. Repeated `%HSRP-5-STATECHANGE` messages are the symptom, and the log matters more than the current state, because by the time you type `show standby brief` the pair may look perfect. Three causes account for most of it. First, a tracked object bouncing: every time it goes down the router's priority drops and the peer preempts, then when it recovers the priority rises and the original router preempts back, so one marginal uplink gives you two failovers per event. This is why [driving the tracked object from an IP SLA probe](https://www.pinglabz.com/ip-sla-cisco-ios-xe/) beats tracking a raw interface: an SLA lets you set thresholds and require consecutive failures before your gateway moves. Second, preempt with no delay: a rebooted router brings its interfaces up and preempts before its routing table has converged, becoming the gateway for a segment it cannot yet forward for. Use `standby 1 preempt delay minimum 90` on top of an interface-level `standby delay minimum 30 reload 60`. Third, intermittent hello loss. The default 3 second hello and 10 second hold time mean three lost hellos trigger a failover, so sub-second timers also tighten your tolerance for a busy control plane. Check interface counters and CPU before blaming HSRP. ## Clearing HSRP config without making it worse One command trap can turn a small problem into a dual-active outage while you are troubleshooting. `no standby version 2` does not remove an HSRP group. Version is an interface-level property, so negating it reverts the interface to version 1 and leaves the group, its priority, its virtual IP and its tracking fully configured. You have not cleaned anything up, you have moved that interface onto the version 1 multicast address while the peer is still on version 2. To remove a group, negate the group: `no standby 1` takes the whole group off the interface, and `no standby 1 ip` removes only the virtual IP. Verify with `show run interface` rather than assuming, and note that changing version deliberately bounces every group on that interface. ## What this was captured on Two `iol-xe` routers in Cisco Modeling Labs running IOS XE 17.18.2 on one Ethernet segment, `10.0.0.0/24`. R1 is `10.0.0.1` with HSRP group 1, priority 110 and preempt; R2 is `10.0.0.2` with priority 100 and preempt; virtual IP `10.0.0.254`. The fault was an inbound ACL on R2's Ethernet0/0 denying UDP 1985 to `224.0.0.2` and `224.0.0.102` from R1, which stops R2 hearing R1 while leaving the rest of the segment normal. Both `show standby brief` outputs were collected simultaneously by on-box EEM applets writing to syslog, which is why the side-by-side comparison above is honest rather than two snapshots a minute apart. ## Gotchas - **One router's output cannot show you dual-active.** R1 alone looks like a perfectly healthy Active gateway. You need both tables, ideally captured at the same moment. - **An ACL that permits everything else passes every test you run.** Pings between the router interfaces succeed and the segment looks clean while HSRP is dead. Test the actual traffic, UDP 1985 to a multicast address. - **A version mismatch logs nothing.** Authentication and virtual IP mismatches generate `%HSRP-4-BADAUTH` and `%HSRP-4-DIFFVIP1`, but a version mismatch is a clean, quiet, total failure. Check it explicitly. ## Key takeaways - Dual-active means the hellos are not crossing. Both routers show Active with `Standby = unknown`, and that second column is what gives it away. - HSRP hellos are UDP 1985 to `224.0.0.2` for version 1 and `224.0.0.102` for version 2\. Confirm that traffic passes both ways before touching any config. - Split the causes into "not delivered" (ACL, missing VLAN, STP block, version mismatch) and "delivered but rejected" (authentication, group or VIP mismatch). The second family logs something, the first usually does not. - Without preempt, the first router to boot keeps the Active role forever, even at a lower priority. No election runs while an Active router already exists. - Flapping is normally a tracked object bouncing or preempt firing before a rebooted router is ready. Use IP SLA thresholds and `standby preempt delay minimum`, not tighter timers. - `no standby version 2` reverts the interface to version 1, it does not remove the group. Use `no standby ` to delete it. Dual-active is one of the rare faults where the fix is almost never in the protocol configuration. The HSRP config on both routers in this lab was correct the whole time. For the wider picture, including [how HSRP differs from VRRP and GLBP when choosing a protocol](https://www.pinglabz.com/hsrp-vs-vrrp-vs-glbp/), work through the [complete first-hop redundancy configuration and troubleshooting guide](https://www.pinglabz.com/fhrp/). ### Routing Loop Troubleshooting: Detect, Break, and Prevent Them URL: https://www.pinglabz.com/routing-loops-detect-and-fix/ Last updated: 2026-08-01T18:57:57.000Z A routing loop is one of the few faults where the symptom hands you the diagnosis. Traffic to one prefix dies, everything else on the box is fine, and a traceroute comes back with the same two or three addresses repeating down the whole column until it gives up at hop 30\. That pattern is not ambiguous. Two routers each believe the other is the way to the destination, and your packets are passed back and forth between them until their TTL runs out. Every block of output below came off a two-router CML lab running Cisco IOS XE 17.18.2 on `iol-xe` nodes, break baked into the startup config: R1 and R2 back to back on 10.0.12.0/30, each holding a static for 172.16.99.0/24 pointing at the other. For the machinery underneath, the guide to [how a router picks a route and forwards on it](https://www.pinglabz.com/ip-routing/) covers the routing table behaviour this article assumes. ## The Symptom Signature Three things point at a loop, and the first is worth more than the other two put together. **The traceroute address column repeats.** This is the tell. Here is the lab, broken: ``` R1# traceroute 172.16.99.1 numeric timeout 1 probe 1 Tracing the route to 172.16.99.1 1 10.0.12.2 3 msec 2 10.0.12.1 2 msec 3 10.0.12.2 2 msec 4 10.0.12.1 2 msec ... (alternating 10.0.12.2 / 10.0.12.1) ... 29 10.0.12.2 13 msec 30 10.0.12.1 13 msec ``` Two addresses, alternating, all the way to the maximum TTL. Note `numeric` and `probe 1`: reverse DNS on thirty hops turns a four second traceroute into a long wait, and one probe per hop keeps the address column a single entry wide so the repeat is obvious at a glance. **You get TTL exceeded replies from a router, not the destination.** A user reporting "TTL expired in transit" is reporting a loop until proven otherwise, and the source address on that ICMP message is one of the looping routers. **A link sits at high utilisation carrying nothing legitimate.** Each packet crosses the looping link many times before it dies, so the loop is a traffic multiplier. A host starting at TTL 128 gets roughly sixty round trips per packet, and a trickle of traffic to a dead prefix presents as tens of megabits between the two routers. ## Why the Network Does Not Melt (and a Layer 2 Loop Does) People conflate routing loops with switching loops. They behave nothing alike, and the difference is one field in the IPv4 header. Every router that forwards a packet decrements its TTL by one, and a router that decrements it to zero discards the packet and sends an ICMP Time Exceeded message back to the source. A looping packet therefore has a hard, bounded lifetime and the damage is linear: one packet per second entering a two-router loop at TTL 255 becomes roughly 127 packets per second on that link. A fixed multiplier that never grows. An Ethernet frame has no TTL field, so a broadcast entering a Layer 2 loop is replicated at every switch on every port, each copy replicated again, and the segment saturates in seconds. That is why STP exists and why nothing equivalent is needed at Layer 3. The same mechanism is what traceroute exploits deliberately. It sends probes with TTL 1, then 2, then 3, and reads the source address of each Time Exceeded reply. When the routing table is looping, traceroute is printing the routers in the cycle in order, over and over. The alternating column is not an artefact of the tool. It is a faithful map of the path. ## What This Was Captured On PlatformCML, 2 x iol-xe nodes, Cisco IOS XE 17.18.2 LinkR1 to R2 back to back, 10.0.12.0/30 Target prefix172.16.99.0/24, which exists nowhere in the lab The breakR1 static to 10.0.12.2, R2 static to 10.0.12.1 Capture methodOn-box EEM running traceroute, output to syslog on R1 ## Finding Where the Loop Closes The traceroute gives you the members of the cycle. Now find the one routing table entry that is wrong. Walk the routers it named and ask each the same question: ``` show ip route 172.16.99.1 show ip cef 172.16.99.1 ``` The loop closes at the pair where A's entry names B as the next hop and B's names A. On a longer cycle, look for the router whose next hop points backwards along the path the packet arrived on. Check `show ip cef` as well as `show ip route`: the RIB is what the control plane decided, the FIB is what the linecard is doing, and the rare disagreement is the one worth catching. Then read that entry's route code and administrative distance, which name the cause: a static shows as S, a redistributed route as external (O E2, D EX), and a summary as a shorter prefix than the one you asked for. ## The Four Causes, in Order of Frequency ### 1\. Mutual redistribution without filtering By a wide margin the most common cause in production, and the only one here that turns up in networks nobody misconfigured on purpose. It needs two ingredients: two routing protocols, and two or more routers redistributing between them in both directions. A prefix native to OSPF is redistributed into EIGRP at router A, crosses the EIGRP domain to router B, and is redistributed back into OSPF as an external. Some router now has a choice between the real route and the round-trip copy, and if the copy wins on metric or administrative distance, it starts forwarding toward the EIGRP domain for a prefix that lives in OSPF. The structural fact worth holding on to: a single redistribution point cannot loop, because a route has nowhere to come back from. Two or more can. Anyone adding a second redistribution router to a single-point design is one route-map away from this, which is why [what breaks when two protocols redistribute into each other](https://www.pinglabz.com/troubleshooting-route-redistribution/) is worth reading first. ### 2\. A static route whose next hop routes the packet back The lab's loop, and the branch and edge classic. You point a static at a next hop, that next hop has no more specific route for the destination, so it falls back on its default, and its default points at you. The production version usually involves an internet edge. Your default points at the ISP, the ISP has a static for your address block pointing back at you, and somebody assigns a prefix inside that block that is not routed internally. Traffic to it goes out, comes straight back, and you trade the packet until TTL runs out. Nobody configured a loop and both statics are individually correct. Recursive statics earn the same suspicion: your router checks it can resolve the next hop and installs the route, and whether that next hop knows what to do with the destination is not something it can verify. ### 3\. Administrative distance manipulation Administrative distance is locally significant. That single sentence is the whole cause. Change the AD of a route to make one box prefer the path you want and you have changed one router's opinion and nobody else's. If the neighbour still prefers the old path, the two now disagree about direction, and traffic caught between them loops. Floating statics are the usual vehicle, along with per-protocol `distance` commands used during a migration. The migration case is worse, because it is applied router by router across a maintenance window and the loop lives for as long as the network is half changed. If you are going to touch AD, get clear on [how a router decides which routing source to trust](https://www.pinglabz.com/administrative-distance/), and change it symmetrically or not at all. ### 4\. Summarisation hiding a more specific route You advertise 10.1.0.0/16 upstream, but only 10.1.1.0/24 and 10.1.2.0/24 exist behind you. Traffic to 10.1.9.9 is drawn to you by the summary, you have no matching specific, so you fall back on your default, which points at the router that just sent you the packet. It sends it back because of your summary. This one surprises people, because summarisation is supposed to be good practice and nothing in the config looks wrong. Being fluent in [how longest prefix match picks the winning entry](https://www.pinglabz.com/understanding-ip-routing/) helps, since the failure is entirely a consequence of the summary being the longest match available for an address with no real home. ## Breaking the Loop The fix is to make one router stop forwarding the packet back. In the lab, repoint R2's static at `Null0`, a discard route: ``` R2(config)# no ip route 172.16.99.0 255.255.255.0 10.0.12.1 R2(config)# ip route 172.16.99.0 255.255.255.0 Null0 ``` Traceroute from R1 straight afterwards: ``` R1# traceroute 172.16.99.1 numeric timeout 1 probe 1 Tracing the route to 172.16.99.1 1 10.0.12.2 3 msec 2 10.0.12.2 !H <-- !H = destination host unreachable (R2 dropped it), loop gone ``` Two lines instead of thirty. Be precise about what the `!H` proves: the packet reached R2 and R2 *dropped* it rather than forwarding it back. An unreachable code is a healthy result for a prefix that genuinely does not exist. (You may see `!N` instead, depending on platform; either means dropped.) Treat the Null0 as stopping the bleeding, not as the fix. It ends the CPU and bandwidth damage in one command, but the reason the two routers disagreed is still in the config. ## Preventing the Next One **Tag routes on redistribution and deny the tag on the way back.** This is the answer to cause number one and the only one that scales. Set a tag when you inject routes into a protocol, then match and deny that tag on every route-map redistributing the other way. A route that originated in your OSPF domain cannot get back in, however many redistribution points you add later, and you never maintain a prefix list. The mechanics are in [stopping redistributed routes from being re-injected using tags](https://www.pinglabz.com/redistribution-loop-prevention-route-tags/). **Filter explicitly at the redistribution point when tags are not available.** Prefix lists on the route-map work, but every new subnet is a change to every filter, and the first one somebody forgets brings the loop back. Use them only where the prefix set is static. **Install a discard route for every summary you advertise.** A Null0 for the aggregate means traffic to non-existent specifics inside it is dropped by you instead of bounced upstream. EIGRP and OSPF do this when you use their own summarisation commands; summaries built by advertising a static or by redistribution do not get one. **Keep redistribution to one point where the topology allows.** One boundary cannot loop, and when redundancy forces you to two, the tagging goes in with the second router, not after the first outage. For a dual-homed site that could receive its own routes back from the core, [tagging routes with the site they came from](https://www.pinglabz.com/eigrp-site-of-origin-soo/) applies the same idea per site rather than per protocol. ## Common Mistakes and Gotchas - **Assuming the pattern is always two alternating addresses.** A three or four router loop cycles through three or four addresses in sequence. Look for repetition, not specifically for A/B/A/B. - **Only testing one direction.** Loops are frequently asymmetric. Traffic one way loops, the return path is clean, and the far end insists the link is fine because their pings work. Traceroute from both ends. - **Treating an unreachable as a failure.** After the fix the traceroute still does not reach the destination, because the destination does not exist. `!H` at hop 2 is the success condition here. - **Leaving the Null0 in and closing the ticket.** The discard route makes the symptom vanish completely, which is why it gets forgotten. Six months later somebody deploys that prefix for real and it is unreachable for reasons nobody can find. ## Key Takeaways - A traceroute whose hop addresses repeat in a cycle up to the maximum TTL is a routing loop. There is no other common explanation. - TTL is why a routing loop degrades a link instead of destroying a network. Ethernet frames have no equivalent field, which is why a switching loop is far worse. - Find it with `show ip route` and `show ip cef` for the destination on each router the traceroute named, looking for the pair pointing at each other. - Ranked by real frequency: mutual redistribution without filtering, a static whose next hop routes the packet back, asymmetric AD changes, and a summary with no matching specific. - A `Null0` discard route breaks the loop immediately and `!H` proves the packet is dropped rather than forwarded, but it is triage, not a fix. - Prevent the next one with route tags on redistribution, filtering at the boundary, and a discard route behind every summary you advertise. Loops are one of the small set of faults where the diagnostic is faster than the theory. For the rest of the material on route selection and forwarding behaviour underneath this one, start from the [routing troubleshooting and configuration hub](https://www.pinglabz.com/ip-routing/). ### CEF Load Balancing and Polarization: Why One Link Gets All the Traffic URL: https://www.pinglabz.com/ecmp-cef-load-balancing-polarization/ Last updated: 2026-08-01T18:57:57.000Z The ticket always reads the same way. Two links, same bandwidth, same protocol, same metric, both up. One is running at 400 Mbps and the other is doing 30\. Somebody escalates it as a load-balancing fault, and it is not one. The misconception is that equal-cost multipath means packets alternate between links. It does not. The routing table installs multiple next-hops, but CEF picks one of them per flow and keeps that flow there for as long as the path exists. A single TCP session will never use both links. Scale that up across a tiered topology where every router hashes the same fields with the same algorithm and you get polarization: entire layers agreeing on the same side and leaving the other side idle. This article proves the per-flow behaviour on real output, then explains the polarization mechanism that follows from it. For the wider context, start with [how a router decides where to send a packet](https://www.pinglabz.com/ip-routing/). Everything captured below came from CML using `iol-xe` nodes running IOS XE 17.18.2, on 2026-07-20. ## What the router installs when two paths tie Equal-cost multipath needs three things to line up: the same routing protocol, the same administrative distance, and the same metric. Miss any one of them and you get a single best path, because [the router ranks the source before it ever compares metrics](https://www.pinglabz.com/administrative-distance/). When they do line up, the prefix is installed with more than one Routing Descriptor Block. ``` R1# show ip route 8.8.8.0 Routing entry for 8.8.8.0/24 Known via "ospf 1", distance 110, metric 11, type intra area Routing Descriptor Blocks: 10.0.13.2, from 3.3.3.3, via Ethernet0/1 Route metric is 11, traffic share count is 1 * 10.0.12.2, from 2.2.2.2, via Ethernet0/0 Route metric is 11, traffic share count is 1 ``` Two descriptor blocks, both `Route metric is 11`, both `traffic share count is 1`. That share count is the ratio in which the router intends to distribute traffic, and 1 to 1 is what equal-cost gives you. The asterisk marks the next-hop used for a packet the router itself originates, which is worth knowing because it makes engineers think one path is preferred. It is not. If you are unsure which fields in that output mean what, [reading the anatomy of a show ip route entry](https://www.pinglabz.com/how-to-read-cisco-routing-table/) is worth the detour. The RIB is only half the story. Forwarding happens in CEF, so check the FIB too: ``` R1# show ip cef 8.8.8.0 255.255.255.0 8.8.8.0/24 nexthop 10.0.12.2 Ethernet0/0 nexthop 10.0.13.2 Ethernet0/1 ``` Two descriptor blocks in the RIB plus two next-hops in the FIB is the definitive answer to "is this prefix load-sharing?". If the RIB shows two and CEF shows one, you have a FIB problem, not a routing problem, and that is a completely different investigation. ## CEF picks a path per flow, not per packet Here is the command that ends most of these arguments. `show ip cef exact-route` takes a source and a destination and tells you exactly which physical interface that pair will use, right now, deterministically: ``` R1# show ip cef exact-route 1.1.1.1 8.8.8.10 1.1.1.1 -> 8.8.8.10 => IP adj out of Ethernet0/1, addr 10.0.13.2 ``` One source, one destination, one interface. Ethernet0/0 is installed, available, and completely irrelevant to this flow. Run the command again and you get the same answer, because CEF is not choosing at random, it is hashing the source and destination address and using the result to index into a load-sharing table. Same input, same output, every time. That is the mechanism behind the ticket. Your 400 Mbps backup job is one source talking to one destination, so it is one hash result, so it is one link. Adding a second 1 Gbps circuit did not make that transfer faster and never could. Only more flows can use more links. Layer 4 ports are not in the hash unless you explicitly add them with `ip cef load-sharing algorithm include-ports source destination`, and platform support for that varies. This matters more than it sounds: if a host pair opens sixteen parallel TCP sessions to move data faster, all sixteen still hash identically under the default algorithm, because the hashed fields are the same for all of them. Adding ports is what turns those sixteen sessions into sixteen independent hash results. ## Per-packet exists and you almost certainly do not want it IOS will do true round-robin if you ask it to, per interface, with `ip load-sharing per-packet`. It gives you genuinely even byte distribution across the links. It also delivers packets out of order whenever the two paths have even slightly different latency, which they always do. Out-of-order delivery is not cosmetic. TCP treats three duplicate ACKs as a loss signal and halves its congestion window, so a link pair that is dropping nothing still looks like a lossy path to every session crossing it. Add that per-packet forwarding is frequently punted rather than done in hardware, and the feature costs you throughput to buy a statistic nobody needs. The one defensible case is very low speed parallel links carrying traffic that does not care about ordering. Outside of that, leave it alone. ## Hashing distributes flows, not bytes This is the part that gets skipped, and it is the reason "fixing" the hash so often changes nothing. CEF distributes flows evenly. It has no idea how many bytes each flow will carry. With two links and four flows, a perfect hash puts two flows on each side. If one of those four is a database replication stream and the other three are SSH sessions, your links are 99 percent uneven and the load balancing is working correctly. You need a large number of similarly-sized flows before hash distribution starts to look like even utilisation, and plenty of real links carry a handful of elephant flows and a long tail of nothing. So before you tune hash algorithms, count your flows. If the answer is "six, and one of them is the backup", no algorithm will help you. The fix is at the application or the design layer, not in CEF. ## Polarization: when every router makes the same decision Now scale that deterministic hash across a tiered topology. An access router hashes a flow and sends it left. The flow arrives at a distribution router that has its own pair of equal-cost paths and runs the same hash function over the same source and destination fields. Nothing about the packet changed in transit, so the hash produces the same relative result and the distribution layer sends it left again. Each downstream router therefore only ever sees the subset of flows that hashed to one side, and half of its own uplinks go unused. That is CEF polarization. It is not a bug in the hash, it is what happens when every device in a path runs an identical deterministic function on identical input, which was exactly the behaviour of the original fixed CEF hashing algorithm. The fix is to make the input or the function differ per device. Modern IOS and IOS XE default to the universal algorithm, which mixes a router-specific unique ID into the hash so that two routers presented with the same flow do not necessarily reach the same conclusion: ``` R1(config)# ip cef load-sharing algorithm universal ``` You can append a hex ID to that command to set the seed explicitly, which guarantees adjacent layers differ rather than trusting them to differ by accident. Other options exist depending on platform, including `tunnel` for topologies where the outer header is nearly identical on every flow, and `include-ports` to widen the hash input. The principle is the same in all of them: vary either the algorithm or the fields per layer so the same flow does not get the same answer twice. If this pattern feels familiar, it should. [The way an EtherChannel picks a member link](https://www.pinglabz.com/etherchannel-configuration-on-cisco-switches/) is the same problem one layer down. A bundle hashes frames to members, a single flow rides a single member, and a bundle with a non-power-of-two member count distributes hash buckets unevenly before any traffic arrives. Diagnose both the same way: count flows first, look at the hash input second. ## When the paths are not equal Everything above assumes the metrics tie. Two knobs change that. `maximum-paths` caps how many equal-cost paths the protocol will install. Check what yours is set to with `show ip protocols | include maximum path`, because the default varies by platform and release. Setting it to 1 disables ECMP for that protocol entirely, which is a useful diagnostic in itself: ``` R1(config)# router ospf 1 R1(config-router)# maximum-paths 1 R1# show ip route 8.8.8.0 Routing entry for 8.8.8.0/24 Known via "ospf 1", distance 110, metric 11, type intra area Routing Descriptor Blocks: * 10.0.12.2, from 2.2.2.2, via Ethernet0/0 <-- only ONE path now R1# show ip cef 8.8.8.0 255.255.255.0 8.8.8.0/24 nexthop 10.0.12.2 Ethernet0/0 ``` One descriptor block, one CEF next-hop. The second path is still in the OSPF database and still valid, it just is not installed. Nothing about the topology changed, only the RIB install limit. The second knob is unequal-cost sharing, and this is where the protocol matters. OSPF, IS-IS and BGP install equal-cost paths only. EIGRP can install paths with different metrics using `variance`, subject to the feasibility condition, and it adjusts the traffic share counts to match the metric ratio. That is a protocol feature layered on top of everything described here, and it is covered separately in [how to share traffic across paths with different metrics](https://www.pinglabz.com/eigrp-variance-unequal-cost-load-balancing/). Keep the two ideas apart: variance decides which paths get installed and in what ratio, while everything in this article is platform-level forwarding behaviour that happens after the install and applies identically no matter which protocol put the paths there. Set variance and the polarization mechanism does not go away, it just operates on unevenly weighted buckets. ## What this was captured on PlatformCML iol-xe, IOS XE 17.18.2 R1 (the router under test)Et0/0 10.0.12.1, Et0/1 10.0.13.1 R210.0.12.2, Lo1 in 8.8.8.0/24 R310.0.13.2, Lo1 in 8.8.8.0/24 Contested prefix8.8.8.0/24, OSPF area 0, metric 11 both ways Capture methodOn-box EEM applet, output to syslog, verbatim R2 and R3 each advertise the same prefix into OSPF area 0 at identical cost, so R1 has no reason to prefer either. That symmetry is what makes the exact-route result meaningful: the flow is not landing on Ethernet0/1 because that path is better, it is landing there because the hash said so. ## Gotchas - **A loopback advertised into OSPF appears as a /32, not as the subnet you configured.** If you build this lab and your `8.8.8.0/24` never shows up, that is why. Put `ip ospf network point-to-point` on the loopback and OSPF advertises the real mask. - **A /30 has exactly two host addresses.** The first build of this lab put R3 on 10.0.13.3/30, which is the broadcast address. No adjacency formed, only one path was installed, and the whole thing looked like an ECMP failure. It was an addressing typo. Before you debug load sharing, confirm both neighbours are actually up. - **The asterisk in `show ip route` is not a preference marker.** With two equal descriptor blocks it just tags which next-hop the router itself would use for locally originated traffic. Engineers regularly read it as "this is the active path" and conclude ECMP is broken. - **Do not test load sharing with ping from the router.** Router-sourced traffic does not necessarily follow the CEF path a transit flow would take. Use `show ip cef exact-route` with the real source and destination instead. - **Interface counters lie about load sharing.** Uneven byte counts across two links are consistent with a healthy hash and a small number of flows. Counters tell you nothing until you know the flow count. ## Key Takeaways - More than one Routing Descriptor Block plus more than one CEF next-hop for the same prefix is the definitive test for whether a prefix is load sharing at all. - CEF hashes each flow to exactly one path, by default over source and destination IP. `show ip cef exact-route` proves it and settles the "why is my transfer not using both links" question in one command. - Per-packet load sharing gives even bytes and causes reordering, which TCP reads as loss. Reach for it almost never. - Hashing distributes flows, not bytes. A handful of large flows will never balance evenly no matter which algorithm you choose. - Polarization happens when every router in a tier runs the same hash over the same fields and reaches the same conclusion. Vary the algorithm, the seed, or the fields per layer to break it. - `maximum-paths` controls how many equal paths get installed and disables ECMP at 1\. Unequal-cost sharing is an EIGRP-only feature and sits on top of, not instead of, the forwarding behaviour described here. The pattern worth carrying away is that load sharing decisions happen in two separate places: the routing protocol decides which paths get installed, and CEF decides which installed path each flow uses. Almost every load-balancing complaint is really a question about the second one. For the rest of the path a packet takes from route lookup to egress, the [complete IP routing reference](https://www.pinglabz.com/ip-routing/) covers the surrounding pieces. ### IPv6 Neighbor Discovery Troubleshooting: Stuck at Link-Local URL: https://www.pinglabz.com/ipv6-neighbor-discovery-troubleshooting/ Last updated: 2026-08-01T18:57:56.000Z An IPv6 host sitting on nothing but an `FE80::` address is not broken the way an IPv4 host with no DHCP lease is broken. There is no server to check, no scope, no relay to blame. In IPv6 the address and the default gateway both arrive in a Router Advertisement, an ICMPv6 packet sent by a router on the segment, and if that packet is not sent or not delivered, the host stays link-local forever and tells you almost nothing about why. Everything below came off a two-router CML lab on Cisco IOS XE 17.18.2 (`iol-xe` nodes): R1 with `ipv6 unicast-routing` and a /64 on Ethernet0/0, R2 running `ipv6 address autoconfig`, and the break baked into R1's startup config as `ipv6 nd ra suppress all`. For the surrounding context, the [IPv6 addressing, routing and troubleshooting guide](https://www.pinglabz.com/ipv6/) maps the cluster in reading order. Short version: if the host has no global address, stop looking at the host. `show ipv6 routers` on the client answers the question in one line, and the answer is nearly always on the router side. ## The Five Messages Neighbor Discovery Runs On Neighbor Discovery (RFC 4861) is not a protocol sitting next to IPv6, it is part of how IPv6 works. It absorbs jobs that IPv4 split across ARP, ICMP redirects, router discovery and DHCP, using five ICMPv6 message types over link-local addresses and multicast. Router Solicitation (RS) ICMPv6 type133 Sent byHosts, to FF02::2 JobAsk for an RA now Router Advertisement (RA) ICMPv6 type134 Sent byRouters, to FF02::1 JobPrefix, flags, MTU, gateway Neighbor Solicitation (NS) ICMPv6 type135 Sent byAny node JobResolve L2, probe, run DAD Neighbor Advertisement (NA) ICMPv6 type136 Sent byAny node JobAnswer an NS, or announce Redirect ICMPv6 type137 Sent byRouters, to the host JobUse a better first hop The split that matters is RS/RA versus NS/NA. RS and RA are *configuration*: what prefix am I on, who is my default router. NS and NA are *reachability*: what MAC sits behind this address, and is it still answering. Different failures, different commands, so decide which half you are in first. ## What This Was Captured On PlatformCML, 2 x iol-xe nodes, Cisco IOS XE 17.18.2 LinkR1 Et0/0 to R2 Et0/0, back to back R1 (the router)ipv6 unicast-routing, 2001:DB8:12::1/64 R2 (the client)ipv6 enable plus ipv6 address autoconfig The breakR1 boots with ipv6 nd ra suppress all Also present10.0.12.0/24 IPv4, dual-stack (see Gotchas) ## The Symptom: Link-Local and Nothing Else R2 two minutes after boot, autoconfig configured, R1 silent: ``` R2# show ipv6 interface brief Ethernet0/0 [up/up] FE80::A8BB:CCFF:FE00:D900 ``` Up/up, IPv6 enabled, link-local present (which proves the stack is running and DAD passed on it). What is missing is any address from `2001:DB8:12::/64`. The detail says why: ``` R2# show ipv6 interface Ethernet0/0 Ethernet0/0 is up, line protocol is up IPv6 is enabled, link-local address is FE80::A8BB:CCFF:FE00:D900 Stateless address autoconfig enabled No global unicast address is configured Joined group address(es): FF02::1 (all-nodes) FF02::2 (all-routers... solicited) FF02::1:FF00:D900 (solicited-node multicast for DAD/ND) ND DAD is enabled, number of DAD attempts: 1 Hosts use stateless autoconfig for addresses. ``` **"Stateless address autoconfig enabled"** plus **"No global unicast address is configured"** is the diagnostic pair: the client is doing its job, waiting for a prefix that never arrives. Note FF02::1:FF00:D900, the solicited-node group derived from the low 24 bits of its own address, which is how an NS reaches one node instead of shouting at the whole segment the way ARP does. ## show ipv6 routers Answers It in One Line Run this first on any host stuck at link-local. It lists the routers whose RAs this node actually received, and what those RAs said: ``` R2# show ipv6 routers R2# ``` Nothing. No RA has been heard on this segment, from anybody, ever. That one empty result moves the investigation off the client and onto the router, in about three seconds. On Linux the equivalent is `rdisc6 eth0` or a capture filtered on `icmp6 and ip6[40] == 134`. Know the timing before you call it empty. A host sends an RS the moment its interface comes up and a router answers within a second or so; after that IOS sends unsolicited RAs every 200 seconds. Empty a full minute after link-up is a real finding, not impatience. To watch the exchange live, `debug ipv6 nd` shows RS, RA, NS and NA in both directions (lab only). ## Root Cause: The Router Is Not Advertising - **No `ipv6 unicast-routing`.** The classic. The box accepts an IPv6 address, answers pings on it and looks fully configured, but until unicast routing is on globally it is a *host*: no RAs, no transit forwarding. Nothing in `show ipv6 interface` flags it, which is why people burn an hour on the client. - **RAs suppressed on the interface.** `ipv6 nd ra suppress all` is legitimate on links where you do not want hosts autoconfiguring, and it is what this lab boots with. It also arrives via a template and gets forgotten. - **The RA is sent but not delivered.** An ACL, an RA Guard policy on the access switch, or an SVI in the wrong VLAN. The router thinks it did its part and the counters agree. - **The RA arrives but says the wrong thing.** That is the flags, below. The fix here was one line on R1, `no ipv6 nd ra suppress all`. Within a minute, R2, with no change to its own config: ``` R2# show ipv6 interface brief Ethernet0/0 [up/up] FE80::A8BB:CCFF:FE00:D900 2001:DB8:12:0:A8BB:CCFF:FE00:D900 <-- SLAAC global address ``` That is SLAAC in one screen. The client was never misconfigured; it got a routable address the moment a router started telling it what prefix it was on. ## Reading the RA: the A, M and O Flags ``` R2# show ipv6 routers Router FE80::A8BB:CCFF:FE00:D800 on Ethernet0/0, last update 0 min Hops 64, Lifetime 1800 sec, AddrFlag=0, OtherFlag=0, MTU=1500 HomeAgentFlag=0, Preference=Medium Reachable time 0 (unspecified), Retransmit time 0 (unspecified) Prefix 2001:DB8:12::/64 onlink autoconfig Valid lifetime 2592000, preferred lifetime 604800 ``` Three fields decide how a host gets addressed. **AddrFlag** is the M (Managed) bit: 1 means "get your address from DHCPv6". **OtherFlag** is O: 1 means "the address is your problem, but take DNS and other options from DHCPv6". `autoconfig` on the prefix line is the A flag in the Prefix Information option, the permission to build an address from that prefix. Both flags are 0 here, so this is pure SLAAC with no DHCPv6 involved. The failure mode is a mismatch between what the router advertises and what the client can act on. M=1 with no DHCPv6 server or relay anywhere leaves hosts with no address at all. A prefix with A cleared and M cleared says "this prefix is on-link, but do not build an address from it", a legal RA that produces a host with a gateway and no source address. If you are choosing between the models, [the difference between stateful and stateless DHCPv6](https://www.pinglabz.com/dhcpv6-stateful-stateless-explained/) is exactly those two bits. One hard constraint: SLAAC only works with a /64, because the interface identifier is defined as 64 bits. ## How the Address Got Built, and DAD The host bits R2 chose, `A8BB:CCFF:FE00:D900`, are the same host bits as its link-local. That is a modified EUI-64 identifier built from the MAC `aabb.cc00.d900`: split the MAC in half, insert `FFFE`, flip the seventh bit of the first byte (which is why AA becomes A8). If that transform is still fuzzy, [building an EUI-64 interface ID by hand](https://www.pinglabz.com/ccna-lab-nf-05-ipv6-addressing-and-eui64/) makes it stick. Real hosts no longer do this: Windows and current Linux use opaque per-prefix identifiers (RFC 7217) plus rotating temporary addresses, so expect several global addresses that do not correlate to a MAC. Before any address is usable it must pass **Duplicate Address Detection**: the node sends an NS for its own tentative address, sourced from `::`, and if anyone answers with an NA, IOS marks the address DUPLICATE and refuses to use it. A DAD failure on the *link-local* address disables IPv6 on that interface entirely, which is a spectacular outage to get from one cloned VM. Note `number of DAD attempts: 1` above; where a port takes a moment to start forwarding, one attempt can finish before anyone could have answered, so a real duplicate slips through. `ipv6 nd dad attempts 3` is cheap insurance. The RA also installs routing state: ``` R2# show ipv6 route NDp 2001:DB8:12::/64 [2/0] via Ethernet0/0, directly connected L 2001:DB8:12:0:A8BB:CCFF:FE00:D900/128 [0/0] via Ethernet0/0, receive ``` `NDp` is the ND-prefix route, learned from the RA rather than from a configured address, and it is a useful tell: NDp routes on a router you did not expect to be autoconfiguring mean an interface is taking directions from whatever else is on that segment. ## The Neighbor Cache and Its Five States Once addressing works, the other half of ND takes over: mapping IPv6 addresses to MACs. `show ipv6 neighbors` is the ARP table equivalent, except it exposes a state machine per entry. INCMPNS sent, no NA back yet. Stuck here means nothing is answering. REACHConfirmed reachable in the last 30 seconds. The only state that means good now. STALEKnown but unconfirmed, because nothing has needed it. Normal, not a fault. DELAYTraffic just used a STALE entry. Waiting a few seconds for upper-layer proof. PROBESending unicast NS on a timer. Repeated PROBE then deletion means it is gone. Most entries in a healthy cache are STALE and that is nothing to fix, because IPv6 confirms reachability from upper-layer progress rather than by re-resolving on a timer. An entry parked in INCMP is an incomplete ARP entry by another name: your solicitation goes out, nothing comes back. Check the L2 path, then check what is filtering ICMPv6. ## Do Not Filter ICMPv6 the Way You Filtered ICMP Blocking ICMP inbound on IPv4 is rude and breaks Path MTU Discovery, but the network keeps working. Do the same to ICMPv6 and you have not hardened the segment, you have unplugged it. Address resolution is ICMPv6\. Router discovery is ICMPv6\. DAD is ICMPv6\. PMTUD is *mandatory* in IPv6 because routers never fragment, so dropping Packet Too Big (type 2) gives you the classic failure where pings work and TLS handshakes hang. A blanket `deny icmp any any` in an IPv6 ACL stops hosts getting addresses, stops them resolving each other and stops large flows completing, all at once, and it looks entirely reasonable to whoever wrote it. Permit types 133 to 137, plus 2, 1, 3 and 4, before the deny. The same discipline applied more broadly is [how ACLs on router interfaces actually treat protocol traffic](https://www.pinglabz.com/infrastructure-security/). Routing protocols get caught by this too: [OSPFv3 peers over link-local addresses](https://www.pinglabz.com/ospfv3-explained-ipv6/), so a filter that interferes with ND tends to take the IGP with it. ## An Address but No Default Gateway The other failure shape is a host that autoconfigures perfectly and still cannot reach anything off-link. The gateway comes from exactly one place: the source of an RA whose **Router Lifetime is non-zero**. Above that is `Lifetime 1800 sec`, so R1's link-local `FE80::A8BB:CCFF:FE00:D800` becomes R2's default router. Three things break it: - **Router Lifetime 0.** `ipv6 nd ra lifetime 0` means "here is prefix information, but do not use me as a default router". Legitimate in some designs, catastrophic if it is the only router on the link. - **DHCPv6 without RAs.** DHCPv6 has no default gateway option. None. A stateful server hands out an address and DNS and still leaves the host with no route off-link, because that only ever comes from an RA. - **RAs filtered after the fact.** RA Guard or a later policy lets the first RA through and drops the refreshes. The lifetime expires, the default route vanishes, and off-link connectivity dies 30 minutes after nobody changed anything. `show ipv6 routers` plus the default route in `show ipv6 route` separate them. Router listed with a non-zero lifetime but no default route means look at the host; router not listed at all puts you back on the RA problem. ## Gotchas - **An empty `show ipv6 neighbors` is not a failure.** In this lab it stayed empty after SLAAC completed, because no data traffic had been exchanged. NS/NA resolution happens on demand, so ping the neighbor first, then read the cache. - **Link-local presence proves nothing about routing.** Every IPv6-enabled interface has one, with or without `ipv6 unicast-routing`, with or without an RA. It only says the stack is up and DAD passed. - **An IPv6-only startup config can fail to load on iol-xe.** A day-0 config with no IPv4 anywhere fell into the setup dialog and never applied; adding an IPv4 address per interface fixed it. If a CML node boots empty, check that before debugging IPv6. - **RA suppression is per-interface, unicast-routing is global.** Identical symptom from the client, two different places to fix, so check both. ## Key Takeaways - IPv6 hosts get their prefix *and* their default gateway from Router Advertisements. No RA, no global address, whatever the host is configured to do. - `show ipv6 routers` is the first command: empty means no RA was heard and the problem is upstream, populated means read the flags. - AddrFlag is M, OtherFlag is O, and the prefix line's `autoconfig` is A. Those three bits decide SLAAC, DHCPv6, or nothing. - Neighbor cache states are a state machine, not a health score. STALE is normal, REACH is confirmed, INCMP means your NS is unanswered. - Filtering ICMPv6 wholesale breaks addressing, resolution and PMTUD at once. Permit types 133 to 137 and type 2 before any deny. - DHCPv6 never supplies a default gateway, so a stateful deployment still needs a router sending RAs. Most IPv6 troubleshooting starts and ends at Neighbor Discovery, because addressing, gateway selection, address resolution and duplicate detection all ride the same five ICMPv6 messages. Once you can read an RA, the rest stops being mysterious: [what you need to run IPv6 in production](https://www.pinglabz.com/ipv6/) covers the addressing, transition and first-hop security material built on top of it. ### Detecting Nmap Scans: The View From the Defender's Side URL: https://www.pinglabz.com/detecting-nmap-scans-cisco-blue-team/ Last updated: 2026-08-01T18:57:56.000Z Most nmap writing, including most of what is on this site, is written from behind the keyboard that launches the scan. You learn which flag does what, how to read `open` versus `filtered`, and how to get through a firewall that would rather you did not. That is half the picture. The other half is what your routers, your switches and your packet captures record while all of that is happening, and whether anyone is actually looking. If you want the attacker-side material first, the guide to [running nmap scans against network gear](https://www.pinglabz.com/nmap/) covers it in depth. This article turns the camera around. Every scan below was run for real from a Kali box, and every detection artefact below was pulled off the target device and off the wire immediately afterwards. The point is not to sell you an IDS. The point is that a plain Cisco router, with no extra licensing and no sensor appliance, already records enough to spot a port scan, and that most engineers never look at that data because they think ACLs are only for blocking. Captures come from a CML lab: Kali running nmap 7.99 at 192.168.99.101, a target router R2 at 192.168.99.2 (`iol-xe`, IOS XE 17.18.2) with a logging ACL inbound on Et0/0, and a Debian VM at 192.168.99.100 running `tcpdump` on the same segment. Everything was captured on 2026-07-24. ## The lab behind these captures Three machines on one broadcast domain, which is deliberate. A scan crossing a firewall gets sanitised on the way; on the local segment you see the raw behaviour, including the ARP that a `-sn` sweep leans on. Kali is bridged into CML through the external connector, so this is a real Linux host attacking a real IOS XE image. AttackerKali 192.168.99.101 (eth1), nmap 7.99 Target and sensor 1R2 192.168.99.2, iol-xe, IOS XE 17.18.2 Sensor 2Debian VM 192.168.99.100, tcpdump on ens224 Detection configACL SCAN-DETECT applied inbound on Et0/0 Scan window27 seconds, four scan types ## From the attacker's seat, this is boring Start with what the operator sees, because the contrast is the whole argument. Two hundred TCP ports, half-open SYN scan, with `--reason` so nmap tells you why it decided what it decided: ``` $ sudo nmap -sS -p 1-200 --reason 192.168.99.2 Starting Nmap 7.99 ( https://nmap.org ) at 2026-07-24 16:01 -0700 Nmap scan report for 192.168.99.2 Host is up, received arp-response (0.0055s latency). Not shown: 199 closed tcp ports (reset) PORT STATE SERVICE REASON 22/tcp open ssh syn-ack ttl 255 MAC Address: AA:BB:CC:00:DC:00 (Unknown) Nmap done: 1 IP address (1 host up) scanned in 4.05 seconds ``` Four seconds. One open port. No error, no timeout, no sign that anything unusual happened. A UDP scan of four common ports came back just as fast, every port `closed` because R2 answered each probe with an ICMP port-unreachable at TTL 255\. A `-sn` sweep of the whole /24 found six live hosts in 2.93 seconds. An `-O` fingerprint guessed "Cisco 1921 router (IOS 15.1)" at 96 percent, wrong in the details and right about the thing that matters. The mechanics of why a lone SYN is enough to [tell an open port from a closed one](https://www.pinglabz.com/nmap-port-scanning/) are a separate story; the point here is that all four scans cost 27 seconds of wall clock and produced nothing that would make an attacker nervous. ## Sensor one: the ACL you already have The detection surface on R2 is not a feature you buy. It is the `log` keyword on an ACE, and the key insight is that `log` works on a **permit** as happily as it does on a deny. That turns an access list from a filter into an instrument. Traffic passes normally, but the first packet of each flow generates a `%SEC-6-IPACCESSLOGP` message and the ACE hit counter increments for every packet. ``` ip access-list extended SCAN-DETECT permit tcp any any eq 22 log permit tcp any host 192.168.99.2 log permit udp any any log permit icmp any any log permit ip any any ! interface Ethernet0/0 ip access-group SCAN-DETECT in ! service timestamps log datetime msec logging buffered 200000 debugging ``` The final `permit ip any any` with no `log` is the escape valve: everything uninteresting still passes without filling your buffer. Now read the counters after the scan window: ``` R2# show ip access-lists SCAN-DETECT Extended IP access list SCAN-DETECT 10 permit tcp any any eq 22 log (257 matches) 20 permit tcp any any eq telnet log (2 matches) 30 permit tcp any any eq www log (2 matches) 40 permit tcp any any eq 443 log (1 match) 50 permit tcp any host 192.168.99.2 log (1420 matches) <-- the sweep 60 permit udp any any log (26 matches) 70 permit icmp any any log (10 matches) 80 permit ip any any (17 matches) ``` That is the fingerprint, and you can see it without a single extra tool. Line 50, the catch-all for anything aimed at R2's own address, absorbed 1420 hits. The named-service lines that represent real traffic saw single or double digits: two telnet, two HTTP, one HTTPS. Normal traffic to a router is lopsided in the opposite direction, concentrated on the handful of ports that actually run services. When your generic line outweighs your service lines by two orders of magnitude, something enumerated you. The UDP and ICMP counters tell the rest of the story. Twenty six UDP hits covers the four-port `-sU` scan plus background broadcast noise, and the ten ICMP hits are the discovery probes from the `-sn` sweep hitting R2's interface. Different [nmap scan types land on different ACEs](https://www.pinglabz.com/nmap-scan-types/), which means a well-structured ACL gives you a crude protocol breakdown of the scan for free. ## The syslog fan-out: one source, many destination ports Counters tell you something happened. Syslog tells you who and what. This is `show logging` filtered to the ACL messages, trimmed to the interesting stretch: ``` R2# show logging | include IPACCESSLOG *Jul 24 23:01:26.597: %SEC-6-IPACCESSLOGP: list SCAN-DETECT permitted tcp 192.168.99.101(51831) -> 192.168.99.2(199), 1 packet *Jul 24 23:01:27.708: %SEC-6-IPACCESSLOGP: list SCAN-DETECT permitted tcp 192.168.99.101(51833) -> 192.168.99.2(85), 1 packet *Jul 24 23:01:28.814: %SEC-6-IPACCESSLOGP: list SCAN-DETECT permitted tcp 192.168.99.101(51835) -> 192.168.99.2(65), 1 packet *Jul 24 23:01:29.920: %SEC-6-IPACCESSLOGP: list SCAN-DETECT permitted tcp 192.168.99.101(51837) -> 192.168.99.2(79), 1 packet *Jul 24 23:01:31.968: %SEC-6-IPACCESSLOGP: list SCAN-DETECT permitted udp 192.168.99.101(47800) -> 192.168.99.2(123), 1 packet *Jul 24 23:01:33.073: %SEC-6-IPACCESSLOGP: list SCAN-DETECT permitted udp 192.168.99.101(47802) -> 192.168.99.2(67), 1 packet *Jul 24 23:01:34.178: %SEC-6-IPACCESSLOGP: list SCAN-DETECT permitted udp 192.168.99.101(47804) -> 192.168.99.2(161), 1 packet *Jul 24 23:01:38.190: %SEC-6-IPACCESSLOGP: list SCAN-DETECT permitted tcp 192.168.99.101(46826) -> 192.168.99.2(3306), 1 packet *Jul 24 23:01:39.312: %SEC-6-IPACCESSLOGP: list SCAN-DETECT permitted tcp 192.168.99.101(46828) -> 192.168.99.2(10010), 1 packet *Jul 24 23:01:40.417: %SEC-6-IPACCESSLOGP: list SCAN-DETECT permitted tcp 192.168.99.101(46830) -> 192.168.99.2(9101), 1 packet ``` Read the destination ports in order: 199, 85, 65, 79, then 3306, 10010, 9101\. There is no application on earth whose client talks to that set of ports in eleven seconds. Legitimate traffic converges on a small number of destination ports and spreads across many source ports; a scan does the opposite, marching through destination ports while the source port ticks up by two each time. That inversion is the single most reliable behavioural signature of a scan, and it is what every commercial detection engine is ultimately measuring, whatever they call it in the datasheet. Note that `1 packet` closes every line. A normal flow eventually logs a higher count as it continues. One packet per flow, hundreds of flows, one source, is enumeration, not communication. ## The rate-limit line nobody reads Buried in the same output is a message that is easy to dismiss as housekeeping noise. It is not. ``` *Jul 24 23:01:41.434: %SEC-6-IPACCESSLOGP: list SCAN-DETECT permitted tcp 192.168.99.101(58846) -> 192.168.99.2(22), 1 packet *Jul 24 23:01:41.920: %SEC-6-IPACCESSLOGRL: access-list logging rate-limited or missed 1423 packets *Jul 24 23:01:43.146: %SEC-6-IPACCESSLOGP: list SCAN-DETECT permitted tcp 192.168.99.101(59098) -> 192.168.99.2(22), 1 packet ``` IOS caps how fast it will generate ACL log messages, because building a syslog message is a process-switched operation and an attacker who can make you log fast enough can hurt the control plane. When the cap is hit, IOS emits `%SEC-6-IPACCESSLOGRL` and tells you how many packets it gave up on. Here that number is 1423, which is essentially the whole scan. Two conclusions follow, and they pull in opposite directions. The first is that `IPACCESSLOGRL` is itself a high-quality detection signal. Nothing in a healthy network generates 1423 new flows in a burst against one router. If that message appears in your syslog server and you have no maintenance window running, you have your alert (arguably a better one than any individual `IPACCESSLOGP` line, because it is a single unambiguous event rather than a pattern you have to correlate). The second is more sobering. Syslog undercounted the scan by more than a thousand packets: the ACE counter said 1420 hits, syslog showed roughly twenty messages. If your detection rule says "alert when I see fifty scan-like flows from one source", the box will never send you fifty. That is the honest argument for NetFlow or a proper IDS on top of ACL logging. ACL logging tells you a scan happened; it will not tell you how big it was. ## The scan on the wire The router's view is a summary. The capture is ground truth, and it is where the scan type becomes unmistakable. This is `tcpdump` on the Debian VM, filtered to the attacker, while a fresh `-sS` ran from Kali: ``` $ sudo tcpdump -i ens224 -nn -c 40 -w res0030-synscan.pcap 'tcp and host 192.168.99.101' -Z root tcpdump: listening on ens224, link-type EN10MB (Ethernet), snapshot length 262144 bytes 40 packets captured 201 packets received by filter 0 packets dropped by kernel $ tcpdump -nn -r res0030-synscan.pcap | head -20 16:03:10.968095 IP 192.168.99.101.47066 > 192.168.99.2.53: Flags [S], seq 3251516137, win 1024, options [mss 1460], length 0 16:03:10.968140 IP 192.168.99.101.47066 > 192.168.99.2.111: Flags [S], seq 3251516137, win 1024, options [mss 1460], length 0 16:03:10.968231 IP 192.168.99.101.47066 > 192.168.99.2.21: Flags [S], seq 3251516137, win 1024, options [mss 1460], length 0 16:03:10.968271 IP 192.168.99.101.47066 > 192.168.99.2.22: Flags [S], seq 3251516137, win 1024, options [mss 1460], length 0 16:03:10.968306 IP 192.168.99.101.47066 > 192.168.99.2.80: Flags [S], seq 3251516137, win 1024, options [mss 1460], length 0 16:03:10.972170 IP 192.168.99.2.53 > 192.168.99.101.47066: Flags [R.], seq 0, ack 3251516138, win 0, length 0 16:03:10.972248 IP 192.168.99.2.111 > 192.168.99.101.47066: Flags [R.], seq 0, ack 3251516138, win 0, length 0 16:03:10.972554 IP 192.168.99.2.21 > 192.168.99.101.47066: Flags [R.], seq 0, ack 3251516138, win 0, length 0 16:03:10.973236 IP 192.168.99.2.22 > 192.168.99.101.47066: Flags [S.], seq 3976364105, ack 3251516138, win 65535, options [mss 1460,sackOK,eol], length 0 16:03:10.973355 IP 192.168.99.101.47066 > 192.168.99.2.22: Flags [R], seq 3251516138, win 0, length 0 <-- half-open teardown ``` Five things in that capture identify a SYN scan beyond argument. All the probes come from one source port (47066) and carry an identical sequence number, because nmap is not tracking state per flow. Every probe is a bare `[S]` with `length 0` and a tiny 1024-byte window. The destination ports jump around with no application logic. Closed ports answer `[R.]`, so a wall of RSTs back to one host is as diagnostic as the SYNs going out. And the one port that answers `[S.]`, port 22, immediately gets an `[R]` from the attacker instead of the ACK that would complete the handshake. That refusal to finish the handshake is what makes `-sS` half-open, and it is why the scan never appears in the SSH daemon's own logs. Counting unique destination ports from that source in the 40-frame capture gives 29\. Twenty nine destinations from one source in a fraction of a second is not a threshold anyone needs to argue about. If you are not comfortable reading flag notation at speed, the [practical tcpdump guide for network engineers](https://www.pinglabz.com/tcpdump-for-network-engineers/) covers the filter syntax, and there is a broader walkthrough of [reading a capture to find the anomaly](https://www.pinglabz.com/packet-analysis/) in the packet analysis material. ## Signatures by scan type Here is how the four scans run in this lab map to what a defender can observe. Everything in this card was seen in the captures above. \-sS SYN scan On the wireLone \[S\], one source port, no ACK On the routerGeneric TCP ACE spikes (1420) Tell\[R\] back after a \[S.\] \-sU UDP scan On the wireICMP port-unreach per closed port On the routerUDP ACE, 26 matches Telludp flows to 53, 67, 123, 161 \-sn ping sweep On the wireARP on-segment, ICMP off-segment On the routerICMP ACE, 10 matches Tell256 addresses probed, 6 answered \-O fingerprint On the wire1000 ports plus odd probe flags On the routerLargest counter jump of the four Tell999 closed ports plus a UDP probe ## The slow scan problem, honestly Everything above assumes default timing. Nmap ran at its normal pace, dumped 1420 packets at one router in a few seconds, and lit up every counter on the box. Change one flag and most of that goes away. Timing templates from `-T0` through `-T5` control probe parallelism and inter-probe delay, and the slow end of that range exists specifically to fall under detection thresholds; the article on [tuning nmap timing and parallelism](https://www.pinglabz.com/nmap-timing-performance/) covers what each one changes. This lab did not capture a slow scan, so treat what follows as reasoning rather than measurement. The ACE hit counter survives. It is cumulative and it does not care whether 1420 packets arrived in four seconds or four days. If you baseline `show ip access-lists` weekly and diff it, a slow scan still shows up as a generic-line counter that grew when nothing else did. That is genuinely one of the few cheap defences against slow enumeration, and almost nobody does it. The packet-rate signals do not survive. `%SEC-6-IPACCESSLOGRL` never fires, because a probe every few seconds never outruns the logger. The per-flow messages still appear, but scattered across days of unrelated log volume, which turns finding them into a correlation problem rather than a reading problem. The tcpdump signature is technically identical (same lone SYNs, same half-open teardown) but you would never catch it in a 40-frame capture. Be honest with yourself about this. A patient attacker probing twenty ports a day from a residential IP is not going to be caught by an ACL counter and a tired engineer. Catching that requires stateful flow records with long retention and something doing the grouping for you. What the router gives you for free is excellent coverage of the loud, fast, automated scanning that makes up the overwhelming majority of what actually hits your perimeter. ## The arms race runs both directions Notice that every detection above keys on the source IP address. That is the weak joint, and attackers know it. Decoy scanning (`-D`) interleaves your real probes with probes bearing forged source addresses, so the defender's log fills with a dozen plausible-looking scanners and no way to tell which one is holding the keyboard. Fragmentation, source-port spoofing and idle scanning attack the same assumption from other angles; the full catalogue is in the piece on [getting a scan past a firewall undetected](https://www.pinglabz.com/nmap-firewall-evasion/). The same `%SEC-6-IPACCESSLOG` surface has been captured recording forged source addresses generated with Scapy, and it records them exactly as faithfully as it records real ones, because the ACL is reading a header field and the header field is a lie. That is a boundary on the technique, not a flaw in it. Source-based logging tells you enumeration is happening on your network; it does not reliably tell you who is doing it. That second question needs correlation the router cannot do alone: switch port mapping, upstream flow records, or a sensor that tracks conversation state. ## Gotchas - **Counters and syslog will disagree, and syslog is the liar.** The ACE showed 1420 matches while syslog showed about twenty messages plus one rate-limit warning. Always size your alerting off counters or flow data, never off a message count. - **Log on permits, not just denies.** Most people only put `log` on the implicit deny, which means they only see traffic their policy already stopped. The interesting traffic is the traffic you are allowing. - **Leave an unlogged catch-all at the bottom.** Without a final `permit ip any any` carrying no `log`, every packet on the interface generates logging and the router spends its day building syslog messages instead of forwarding. - **ACL logging is process-switched.** That is precisely why IOS rate-limits it. Do not respond to the rate-limit message by raising the cap on a production edge router without understanding what you are asking the CPU to do. - **Read counters twice.** Separate reads of the same ACL during and after the scan gave different totals on the SSH line, because background management traffic keeps ticking. Diff two timestamped reads rather than trusting one snapshot. - **A quiet SSH log does not mean nobody knocked.** The SYN scan found port 22 open and never completed a handshake, so `sshd` has nothing to report. Application logs are blind to half-open scanning by design. ## Key Takeaways - A `permit ... log` ACL turns any IOS XE device into a scan sensor with no extra licensing. The signature is a lopsided ACE hit count: 1420 matches on a generic line against single digits on the real service lines. - The behavioural tell in syslog is one source address reaching many different destination ports in a short window, one packet per flow. Real traffic inverts that ratio. - `%SEC-6-IPACCESSLOGRL` is an alert in its own right. It reported 1423 missed packets here, which means the burst was real and that syslog alone will always undercount it. - In a capture, a SYN scan is unmistakable: identical sequence numbers from a single source port, bare SYNs with a 1024-byte window, RSTs coming back from closed ports, and an `[R]` from the attacker on the one port that answered. - Slow scans defeat every rate-based signal but not the cumulative counter. Diffing `show ip access-lists` over time is one of the few cheap ways to see enumeration spread across days. - Source-based detection is defeated by decoys and spoofing, which is why this belongs alongside the rest of your [Cisco device hardening controls](https://www.pinglabz.com/infrastructure-security/) rather than standing in for them. ### Scapy vs Cisco Layer 2 Defenses: Crafting and Spoofing Packets URL: https://www.pinglabz.com/scapy-packet-crafting-spoofing-cisco/ Last updated: 2026-08-01T18:57:55.000Z If you defend switches, you need to know exactly what an attacker can put on the wire, because "the network trusts what the packet says" is the assumption every Layer 2 attack is built on. Scapy is the fastest way to prove that assumption is wrong: it lets you set every field in a frame by hand and send it, which is why red teams reach for it and why you should understand it before you sign off on a switch config. This article walks the attack and then walks the defence, and the defence is the point. On a lab switch with DHCP snooping and Dynamic ARP Inspection turned on, the same forged ARP that poisons an undefended host gets dropped at the port and the switch log names the attacker's real MAC. That is the moment worth reading for, and it sits at the centre of the [infrastructure security controls that stop Layer 2 spoofing](https://www.pinglabz.com/infrastructure-security/). Everything below was captured on lab equipment built only to be attacked. Nothing here is novel tradecraft. The value is the side-by-side: craft the packet, watch it work against an unprotected segment, then watch a correctly configured switch refuse it. Reproduce the defence and you can trust it in production. The run was Scapy 2.7.01 on Kali (Python 3), attacking a Debian 13 VM (kernel 6.12) and a Cisco IOS-XE 17.18.2 router and switch inside Cisco Modeling Labs. Full topology is in the "What this was captured on" note near the end. ## Craft a packet and put it on the wire Scapy builds a frame as a stack of layers you divide together with `/`. There is no template and no default you cannot override. Here is a hand-built ICMP echo carrying a custom payload, inspected before it is ever serialised, then sent with `sr1()` (send one packet at layer 3, wait for one reply): ``` from scapy.all import * conf.iface = "eth1" pkt = IP(dst="192.168.99.2")/ICMP()/Raw(load=b"PINGLABZ-RES0008") pkt.summary() pkt.show() ans = sr1(pkt, timeout=3) ``` The `.show()` output is the whole point of Scapy: every field is yours, and the ones left as `None` (length, checksum) are computed at send time so you do not have to. ``` ! Command: pkt.summary() IP / ICMP 192.168.99.101 > 192.168.99.2 echo-request 0 / Raw ! Command: pkt.show() (the layered structure Scapy will serialize) ###[ IP ]### version = 4 ttl = 64 proto = icmp src = 192.168.99.101 dst = 192.168.99.2 ###[ ICMP ]### type = echo-request code = 0 ###[ Raw ]### load = b'PINGLABZ-RES0008' ! Command: ans = sr1(pkt, timeout=3) # send layer-3, get one reply ! ans.summary(): IP / ICMP 192.168.99.2 > 192.168.99.101 echo-reply 0 / Raw / Padding ``` A frame you assembled field by field got a real echo-reply from the router. That is the primitive. Everything below is the same call with fields set to values that are not true. If you want to see these frames arrive from the receiving side rather than the sender's, this is where a capture on the target pays off - our companion walkthrough on [reading crafted frames with tcpdump](https://www.pinglabz.com/tcpdump-for-network-engineers/) shows the same packets landing on the wire. ## Forge the source IP and watch the router log the lie The source address in an IP header is whatever the sender types. A router forwarding a packet has no way to check it, so Scapy can stamp any source you like. To prove the router accepts the lie, R2 runs an inbound ACL with `permit ip any any log`, so it records the source of everything that arrives. We send from four addresses that do not exist on the segment: ``` send(IP(src="10.66.66.66", dst="192.168.99.2")/ICMP()/Raw(b"SPOOFED-BY-SCAPY"), count=3) send(IP(src="172.16.240.5", dst="192.168.99.2")/ICMP()/Raw(b"SPOOFED-BY-SCAPY"), count=3) send(IP(src="192.0.2.13", dst="192.168.99.2")/ICMP()/Raw(b"SPOOFED-BY-SCAPY"), count=3) send(IP(src="198.51.100.7", dst="192.168.99.2")/TCP(dport=22, flags="S"), count=2) ``` The router's own log, filtered to just those forged sources, shows all four accepted and recorded exactly as Scapy stamped them: ``` R2# show logging | include 10.66.66.66|172.16.240.5|192.0.2.13|198.51.100.7 %SEC-6-IPACCESSLOGP: list SCAN-DETECT permitted tcp 198.51.100.7(20) -> 192.168.99.2(22), 2 packets %SEC-6-IPACCESSLOGDP: list SCAN-DETECT permitted icmp 192.0.2.13 -> 192.168.99.2 (8/0), 3 packets %SEC-6-IPACCESSLOGDP: list SCAN-DETECT permitted icmp 10.66.66.66 -> 192.168.99.2 (8/0), 3 packets %SEC-6-IPACCESSLOGDP: list SCAN-DETECT permitted icmp 172.16.240.5 -> 192.168.99.2 (8/0), 3 packets ``` None of those hosts exist. The router logged them anyway, because the source IP is data, not identity. This is the single-screenshot argument for why you cannot make trust decisions on source address alone, and why [the controls that drop packets with impossible source addresses](https://www.pinglabz.com/anti-spoofing-acls-urpf/) (uRPF and anti-spoofing ACLs) exist. The router cannot tell truth from forgery in the header; it can only check the source against where the packet actually arrived from. ## ARP cache poisoning, and why the textbook one-liner fails ARP has no authentication at all, which is what makes it the classic Layer 2 attack. The idea is to send an unsolicited "192.168.99.2 is-at my MAC" reply so the victim sends R2's traffic to the attacker instead. In Scapy that is one line: ``` m = get_if_hwaddr("eth1") # the attacker's own MAC arp = ARP(op=2, psrc="192.168.99.2", hwsrc=m, pdst="192.168.99.100") sendp(Ether(dst="ff:ff:ff:ff:ff:ff")/arp, count=5, iface="eth1") ``` Here is the first surprise, and it is the opposite of what most ARP-spoofing tutorials imply. A single unsolicited reply, fired over a cache entry that is already `REACHABLE`, did not overwrite it. Modern Linux (this was Debian 13, even with `arp_accept=1`) refuses to replace a live entry on the strength of one gratuitous reply. The famous one-liner does nothing against a host that is currently talking to the real gateway. The realistic attack is a continuous flood. With Kali spraying roughly three poison replies a second, and the victim forced to re-resolve as a natural cache expiry would, we sampled the victim's ARP table every two seconds: ``` ! victim (Debian VM), watching the entry for R2: t+2s: 192.168.99.2 lladdr 00:0c:29:6f:c4:ac STALE <-- attacker's MAC (MITM window) t+4s: 192.168.99.2 lladdr aa:bb:cc:00:dc:00 STALE <-- real R2 reclaims it t+6s: 192.168.99.2 lladdr aa:bb:cc:00:dc:00 STALE ... t+24s: 192.168.99.2 lladdr aa:bb:cc:00:dc:00 STALE ! Samples pointing at the ATTACKER (00:0c:29:6f:c4:ac): 1 of 12 ! Samples pointing at the REAL R2 (aa:bb:cc:00:dc:00): 11 of 12 ``` The entry flaps. Every time it holds the attacker's MAC the victim's traffic to R2 is delivered to Kali, which is a working man-in-the-middle window; but the real R2 is right there answering too, so it keeps reclaiming the slot. One of twelve samples pointed at the attacker in this run. That is the honest result on a live segment, and it explains why real MITM tools spray many times a second and also actively suppress the real host: winning the race once is easy, holding it is not. Do not trust a demo that shows a clean, permanent takeover from a single packet; on a segment where the real gateway is answering, that is not what happens. The lesson for a defender is that the fragility of the attack is not your defence. An attacker who sprays fast enough and silences the gateway will hold the window. You need a control that never lets the forged reply reach the victim in the first place. That control lives on the switch. ## The defence: Dynamic ARP Inspection drops it cold This is the whole reason to run the attack. Dynamic ARP Inspection (DAI) is the blue-team answer to everything above, and this article is the attack half of a pair: the defence is documented in full in [the Dynamic ARP Inspection and IP Source Guard configuration guide](https://www.pinglabz.com/dynamic-arp-inspection-ip-source-guard/). DAI intercepts every ARP on an untrusted port and checks the sender IP-to-MAC claim against a trusted binding. If the claim does not match, the ARP is dropped before it ever reaches the victim. On the lab switch (an `ioll2-xe`) we bound the two legitimate hosts in a static ARP ACL and made the router uplinks trusted: ``` arp access-list PLZ-ARP-ACL permit ip host 192.168.99.2 mac host aabb.cc00.dc00 ! the real R2 binding permit ip host 192.168.99.100 mac host 000c.29b1.cc47 ! the VM, so it keeps working ! interface Ethernet0/1 ip arp inspection trust ! router/uplink ports = trusted ip arp inspection vlan 1 ip arp inspection filter PLZ-ARP-ACL vlan 1 ip arp inspection validate src-mac ip ``` In production you would drive those bindings from the DHCP snooping table rather than a static ACL, which is why DHCP snooping is DAI's prerequisite - the [DHCP snooping and DAI lab](https://www.pinglabz.com/ccna-lab-sec-05-dhcp-snooping-and-dai/) walks that pairing end to end. With DAI live, we fired the exact same forged ARP flood from Kali (20 replies this time) and read the switch. The statistics counter tells the story in two columns: ``` SW1# show ip arp inspection statistics vlan 1 Vlan Forwarded Dropped DHCP Drops ACL Drops ---- --------- ------- ---------- --------- 1 0 44 44 0 ``` Zero forged ARPs forwarded. Forty-four dropped. And DAI did not just drop the packets silently - it logged them, with the attacker's real MAC: ``` SW1# show logging | include SW_DAI %SW_DAI-4-DHCP_SNOOPING_DENY: 5 Invalid ARPs (Res) on Et0/0, vlan 1. ([000c.296f.c4ac/192.168.99.2/0000.0000.0000/192.168.99.100/...]) ``` Read that log line the way the switch did. Sender MAC `000c.296f.c4ac` (Kali) claimed to be sender IP `192.168.99.2` (R2), addressed to `192.168.99.100` (the VM). DAI compared the claimed pairing `192.168.99.2 -> 000c.296f.c4ac` against the ACL, which binds `192.168.99.2 -> aabb.cc00.dc00`, saw the mismatch, dropped the frame, and recorded the attacker's true hardware address in the process. That is the exact packet Scapy built, defeated at the port and attributed to its source. Two details make this a clean win rather than a blunt instrument. First, the interface trust state is what makes it surgical: ``` SW1# show ip arp inspection interfaces Interface Trust State Rate (pps) Burst Interval Et0/0 Untrusted 15 1 Et0/1 Trusted None N/A Et0/2 Trusted None N/A Et0/3 Trusted None N/A ``` Second, and this is the part that separates a good control from a self-inflicted outage: the legitimate host never noticed. Because the VM's own binding was in the ACL, its ARP passed inspection throughout. During the same attack window, the VM flushed its cache, re-resolved, and pinged R2 with zero loss, landing on the real R2 MAC. DAI dropped only the forgery and left the honest traffic alone. A defence that also breaks the users it protects does not survive contact with a change board; this one does not. ## What this was captured on AttackerKali, Scapy 2.7.01 on Python 3, eth1 = 192.168.99.101, MAC 00:0c:29:6f:c4:ac Victim hostDebian 13 VM (kernel 6.12), 192.168.99.100 on ens224 RouterR2, IOS-XE 17.18.2 (iol-xe), 192.168.99.2, real MAC aabb.cc00.dc00 Switch / defenderSW1, IOS-XE 17.18.2 (ioll2-xe), DHCP snooping + Dynamic ARP Inspection PlatformCisco Modeling Labs, Kali bridged into the lab via bridge1 Everything shown here is a controlled lab. Running these techniques against a network you do not own is illegal, and running the ARP flood against a production segment will cause an outage whether or not it "works". Build the topology, break it, defend it, and keep it inside the lab. ## Common mistakes and gotchas **Expecting the one-shot poison to stick.** A single gratuitous ARP does not overwrite a live `REACHABLE` entry on modern Linux, even with `arp_accept=1`. Tutorials that show a clean takeover from one packet are testing against an empty or expired cache. Against a host actively talking to the gateway, you need a sustained flood, and even then you only win intermittently. **Reading the flood result as a permanent takeover.** Our honest sample was one in twelve pointing at the attacker. The real gateway keeps reclaiming its slot, so the man-in-the-middle window flickers. If your defensive testing shows the entry flapping rather than pinning, that is correct behaviour, not a failed capture. **Forgetting the trusted-port config on DAI.** If you enable DAI on a VLAN and leave the router uplinks untrusted, the switch will start dropping the router's own legitimate ARP and you will take down the segment you meant to protect. Uplinks and known infrastructure ports must be `ip arp inspection trust`. **Leaving legitimate hosts out of the binding source.** DAI drops anything it cannot validate. Static ARP ACLs are fine for a couple of servers, but for user ports you want DHCP snooping populating the bindings automatically, or every host that renews a lease will get its ARP dropped. **Shared-bridge artifacts in a lab.** In our capture, a few DAI drop lines named MACs from a separate 802.1X lab whose ARP crossed the shared external bridge. That is a lab-only side effect of bridging multiple topologies onto one segment; on a production switch the shared bridge does not exist. Do not mistake cross-lab noise for a DAI error. **Treating a MAC as identity.** The same spoofability that makes ARP poisoning possible is why MAC-based controls are weak. A forged source MAC is trivial for Scapy, which is the same reason a spoofable Layer 2 identity undermines MAC-based access schemes and why [Layer 2 attacks like VLAN hopping](https://www.pinglabz.com/vlan-hopping-attacks/) deserve the same switch-side hardening. ## Key takeaways - Scapy lets you set every field in a frame, so any header value the network trusts (source IP, ARP sender, MAC) can be forged in one line. Treat headers as claims, not facts. - A router logs a spoofed source IP verbatim because it cannot verify it. Source address is not identity; uRPF and anti-spoofing ACLs are the answer. - The textbook one-shot ARP poison fails against a live cache on modern Linux. A sustained flood wins the race only intermittently while the real gateway keeps answering. - Dynamic ARP Inspection dropped 44 of 44 forged ARPs, forwarded zero, and logged the attacker's true MAC. It is the clean, attributable win against ARP spoofing. - A correctly scoped DAI deployment is surgical: the legitimate host kept pinging with zero loss because its binding was trusted. Trusted uplinks and a real binding source are what keep it from becoming an outage. The attack half of this is old news; the defence is what you take to work. Craft the packet so you understand the claim the network is being asked to trust, then put a control at the switch that checks the claim. If you are hardening a switching layer, treat DAI and IP Source Guard as baseline rather than optional, and fold packet captures into the routine so you can see the forgery arrive and confirm it was dropped - the [packet analysis workflow](https://www.pinglabz.com/packet-analysis/) makes that a repeatable check. For the wider set of switch-side controls that turn "the network trusts what the packet says" from a liability into a logged, dropped event, work through [the rest of the infrastructure hardening series](https://www.pinglabz.com/infrastructure-security/). ### tcpdump for Network Engineers: A Practical Reference URL: https://www.pinglabz.com/tcpdump-for-network-engineers/ Last updated: 2026-08-01T18:57:55.000Z Every other troubleshooting tool on a Linux box is telling you what it believes. `ss` believes a socket is established, the application log believes it sent the request, the router believes it has a neighbour. tcpdump is the only one that tells you what actually left the NIC. When the device claims one thing and the network is doing another, the capture settles it in about eight seconds. This is a working reference for **tcpdump for network engineers**, not a first-day walkthrough: interface selection, BPF filter syntax, reading the output line, writing and re-filtering pcap files, snaplen, and the ring-buffer flags that let a capture run overnight without filling a disk. Everything below was captured on a Debian 13 VM running tcpdump 4.99.5, bridged onto the same segment as three `iol-xe` routers running IOS-XE 17.18.2 in CML. It is one slice of [the packet analysis toolkit a network engineer actually needs](https://www.pinglabz.com/packet-analysis/), and it is the slice that works everywhere, because tcpdump is on the box already. Two things in here are not in the usual tutorials: why a rotating capture stops after exactly one file even under `sudo`, and a telnet login to a Cisco router with the username and password readable in the capture, character by character. ## What this was captured on The capture host is a Linux VM with a second NIC bridged into a virtual topology, so the routers on the wire are real IOS-XE control planes generating real OSPF traffic. Timestamps drift between blocks because each block is a separate run. Capture hostDebian 13 (trixie), kernel 6.12.96, tcpdump 4.99.5 Capture NICens224, 192.168.99.100/24, bridged into CML On the wireR1/R2/R3 at 192.168.99.1-3, iol-xe, IOS-XE 17.18.2 Background trafficOSPF 1 area 0, three speakers, Hellos every 10 s TopologyOne flat broadcast domain, no routing between VM and routers Three OSPF speakers on the segment is deliberate: the wire is never quiet, so no command here has the "nothing is happening" problem. Put a routing protocol on your lab segment and you have a free traffic generator. ## Permissions come before filters The number one tcpdump failure is not a filter mistake. It is this: ``` ! Command: tcpdump -i ens224 -c 1 # as user j, no sudo tcpdump: ens224: You don't have permission to perform this capture on that device (socket: Operation not permitted) ``` Opening a packet socket is privileged, and on Debian 13 the binary carries neither route around that: ``` ! Command: which tcpdump; ls -l $(which tcpdump) /usr/bin/tcpdump -rwxr-xr-x 1 root root 1273880 Feb 9 2025 /usr/bin/tcpdump ! Command: getcap $(which tcpdump) # file capabilities, if any <-- empty: no cap_net_raw either ``` No setuid bit, no `cap_net_raw`. So you run under `sudo`, or you grant the capability yourself with `setcap cap_net_raw,cap_net_admin=eip /usr/bin/tcpdump` (which has consequences, since anyone who can then execute it reads every frame on the box). Remember this, because tcpdump does something with those privileges later that surprises people. ## Pick the interface, and prove you picked the right one `-D` lists what is capturable, and it is the fastest way to avoid the second-most-common mistake, which is capturing on the management NIC and wondering where the traffic went: ``` ! Command: tcpdump -D # list capture interfaces 1.ens192 [Up, Running, Connected] 2.ens224 [Up, Running, Connected] <-- the NIC on the lab segment 3.any (Pseudo-device that captures on all interfaces) [Up, Running] 4.lo [Up, Running, Loopback] 5.bluetooth-monitor (Bluetooth Linux Monitor) [Wireless] 6.nflog (Linux netfilter log (NFLOG) interface) [none] 7.nfqueue (Linux netfilter queue (NFQUEUE) interface) [none] ``` `-i any` looks like the safe answer and mostly is not: it gives you a Linux cooked-capture link type instead of Ethernet, so no Ethernet header and no VLAN tag. Name the physical interface. Be honest about where the host sits, too: tcpdump only sees frames delivered to that NIC, so its own traffic, broadcast, multicast, and whatever a SPAN session feeds it. If you cannot get a Linux host onto the segment, [capture on the router itself with EPC](https://www.pinglabz.com/embedded-packet-capture-ios-xe/) is the Cisco-side equivalent, and on a firewall [the ASA has its own capture engine](https://www.pinglabz.com/cisco-asa-packet-capture/) with the same idea and different syntax. ## Bound the capture, and use -nn every time Two flags belong in muscle memory. `-c N` stops after N packets so the command always terminates, which matters when you are running tcpdump over the same SSH session you are capturing. `-nn` disables name resolution for addresses and ports. Compare a default run with the same run under `-nn`: ``` ! Command: tcpdump -i ens224 -c 8 listening on ens224, link-type EN10MB (Ethernet), snapshot length 262144 bytes 14:55:47.957527 IP 192.168.99.1 > ospf-all.mcast.net: OSPFv2, Hello, length 84 14:55:50.603060 IP 192.168.99.3 > ospf-all.mcast.net: OSPFv2, Hello, length 84 14:55:50.883184 IP 192.168.99.2 > ospf-all.mcast.net: OSPFv2, Hello, length 84 8 packets captured 9 packets received by filter 0 packets dropped by kernel ! Command: tcpdump -i ens224 -nn -c 8 14:56:14.914958 IP 0.0.0.0.68 > 255.255.255.255.67: BOOTP/DHCP, Request from 00:0c:29:b1:cc:47, length 301 14:56:16.913287 IP 192.168.99.1 > 224.0.0.5: OSPFv2, Hello, length 84 14:56:18.903784 IP 192.168.99.3 > 224.0.0.5: OSPFv2, Hello, length 84 14:56:19.218986 IP 192.168.99.2 > 224.0.0.5: OSPFv2, Hello, length 84 8 packets captured 8 packets received by filter 0 packets dropped by kernel ``` `ospf-all.mcast.net` is prettier and useless. You wanted `224.0.0.5`, because that is the number you will grep for. Worse, every unresolved address triggers a reverse DNS lookup, which slows the capture and injects your own DNS queries into the capture you are trying to read (enough to cause drops on a busy box). One `-n` kills host names, the second kills port names, and there is essentially never a reason not to use both. Add `-v` or `-vv` when you want the protocol decoded rather than summarised. tcpdump knows more protocols than people expect: ``` ! Command: tcpdump -i ens224 -nn -v -c 3 'proto ospf' 14:56:37.758339 IP (tos 0xc0, ttl 1, id 1866, offset 0, flags [none], proto OSPF (89), length 104) 192.168.99.3 > 224.0.0.5: OSPFv2, Hello, length 84 [len 52] Router-ID 3.3.3.3, Backbone Area, Authentication Type: none (0) Options [External, LLS] Hello Timer 10s, Dead Timer 40s, Mask 255.255.255.0, Priority 1 Designated Router 192.168.99.3, Backup Designated Router 192.168.99.2 Neighbor List: 1.1.1.1 2.2.2.2 ``` That single Hello gives you the timers, area, mask, DR and BDR, and this router's neighbour list. Every classic adjacency-mismatch cause, without touching the router. ## Reading the output line The default one-line format is the same shape for every IP packet: ``` 14:56:58.677331 IP 192.168.99.100.48124 > 192.168.99.1.22: Flags [S], seq 1351658335, win 64240, options [mss 1460,sackOK,TS val 3102991609 ecr 0,nop,wscale 7], length 0 ``` Left to right: microsecond timestamp (host clock, so NTP matters if you are correlating with anything else), the protocol tcpdump decoded, then `source-address.source-port > destination-address.destination-port`. That last octet-looking field is the port, which trips people up until they notice it is a fifth number. After the colon comes protocol detail, then `length`, which is payload bytes and not frame bytes (a pure ACK is `length 0`). The counters at the end of a live run are worth reading rather than skipping. `packets captured` is what tcpdump printed or wrote, `packets received by filter` is what the BPF matched, and `packets dropped by kernel` is the one that invalidates the capture: anything above zero means the kernel ring buffer overflowed and you are missing frames. ## BPF filter syntax The filter is the last argument, and it compiles into a kernel-side BPF program, which is why a good filter costs almost nothing. It is built from qualifiers in three families: type (`host`, `net`, `port`, `portrange`), direction (`src`, `dst`), and protocol (`ip`, `ip6`, `arp`, `tcp`, `udp`, `icmp`, `ether`, or `proto ospf` for anything with no keyword of its own). host MatchesOne address, either direction Examplehost 192.168.99.3 With directionsrc host 192.168.99.100 net MatchesA prefix, either direction Examplenet 192.168.99.0/24 NoteMask or CIDR both accepted port MatchesTCP or UDP port, both ends Exampletcp port 22 Rangedst portrange 33434-33534 proto MatchesIP protocol number or name Exampleproto ospf Shorthandsicmp, tcp, udp, arp Combine them with `and`, `or` and `not`. Two things catch people out. `not` binds tighter than `and`, so parenthesise whenever you negate more than one term, and quote the whole filter so the shell does not eat them. And qualifiers do not carry forward the way they look: `src host A or B` is not what you meant, so write `src host A or src host B`. The most useful filter on a lab segment is often the negative one: ``` ! Command: tcpdump -nn -r res0017-lab-edge.pcap 'not proto ospf and not port 22' 14:57:56.051646 IP 192.168.99.100 > 192.168.99.1: ICMP echo request, id 8, seq 1, length 64 14:57:56.055154 IP 192.168.99.1 > 192.168.99.100: ICMP echo reply, id 8, seq 1, length 64 14:57:58.064454 IP 192.168.99.100 > 192.168.99.3: ICMP echo request, id 9, seq 1, length 64 14:57:58.067156 IP 192.168.99.3 > 192.168.99.100: ICMP echo reply, id 9, seq 1, length 64 ``` ## ARP and ICMP, with the Ethernet header on `-e` prints the link-layer header, which turns an ARP exchange from an abstraction into something you can read straight down. Clear the neighbour entry first so the resolution happens in front of you: ``` ! sudo ip neigh del 192.168.99.2 dev ens224 ! Command: tcpdump -i ens224 -nn -e -c 6 'arp or icmp' 14:56:49.319348 00:0c:29:b1:cc:47 > ff:ff:ff:ff:ff:ff, ethertype ARP (0x0806), length 42: Request who-has 192.168.99.2 tell 192.168.99.100, length 28 14:56:49.323275 aa:bb:cc:00:dc:00 > 00:0c:29:b1:cc:47, ethertype ARP (0x0806), length 60: Reply 192.168.99.2 is-at aa:bb:cc:00:dc:00, length 46 14:56:49.323306 00:0c:29:b1:cc:47 > aa:bb:cc:00:dc:00, ethertype IPv4 (0x0800), length 98: 192.168.99.100 > 192.168.99.2: ICMP echo request, id 6, seq 1, length 64 14:56:49.325668 aa:bb:cc:00:dc:00 > 00:0c:29:b1:cc:47, ethertype IPv4 (0x0800), length 98: 192.168.99.2 > 192.168.99.100: ICMP echo reply, id 6, seq 1, length 64 6 packets captured 6 packets received by filter 0 packets dropped by kernel ``` Broadcast request to `ff:ff:ff:ff:ff:ff`, unicast reply straight back, then the echo request goes out with the learned MAC as its destination. The host's own view agrees afterwards (`192.168.99.2 lladdr aa:bb:cc:00:dc:00 REACHABLE`), and that correlation is the trick: a capture that agrees with the control plane means you can stop looking at Layer 2\. There is more on what the ICMP itself is doing in [what ping is actually putting on the wire](https://www.pinglabz.com/ping/). Note also that `6 packets received by filter` equals `6 packets captured`: the BPF narrowed this in the kernel, not at the display layer. On a 10 Gb link that is the difference between a usable capture and a lost one. ## A TCP handshake, and filtering on flags Six packets of an SSH connection to R1 covers the entire opening of a TCP session: ``` ! Command: tcpdump -i ens224 -nn -c 6 'tcp and host 192.168.99.1 and port 22' 14:56:58.677331 IP 192.168.99.100.48124 > 192.168.99.1.22: Flags [S], seq 1351658335, win 64240, options [mss 1460,sackOK,TS val 3102991609 ecr 0,nop,wscale 7], length 0 14:56:58.681261 IP 192.168.99.1.22 > 192.168.99.100.48124: Flags [S.], seq 894975465, ack 1351658336, win 65535, options [mss 1460,sackOK,nop,nop,wscale 2,eol], length 0 14:56:58.681343 IP 192.168.99.100.48124 > 192.168.99.1.22: Flags [.], ack 1, win 502, length 0 14:56:58.682650 IP 192.168.99.100.48124 > 192.168.99.1.22: Flags [P.], seq 1:42, ack 1, win 502, length 41: SSH: SSH-2.0-OpenSSH_10.0p2 Debian-7+deb13u4 14:56:58.684111 IP 192.168.99.1.22 > 192.168.99.100.48124: Flags [P.], seq 1:20, ack 1, win 32768, length 19: SSH: SSH-2.0-Cisco-1.25 14:56:58.684164 IP 192.168.99.100.48124 > 192.168.99.1.22: Flags [.], ack 20, win 502, length 0 ``` Four milliseconds for the SYN-ACK, and both version banners in the clear before encryption starts (which is how a scanner fingerprints your gear without ever authenticating: `SSH-2.0-Cisco-1.25` is a giveaway). The dot in the flag notation is ACK, so `[S.]` is SYN plus ACK and `[.]` is a bare ACK. tcpdump switches to relative sequence numbers after the first packet in each direction, which is why the SYN shows an absolute `seq` and the rest count from 1. \[S\]SYN, connection opener \[S.\]SYN plus ACK, the server agreeing \[.\]Bare ACK, usually length 0 \[P.\]PSH plus ACK, data with it \[R.\]RST plus ACK, refused or torn down \[F.\]FIN plus ACK, graceful close Once you can read the flags you can filter on them, with the byte-offset syntax BPF exposes when keywords run out. `tcp[tcpflags]` is the flags byte, and the named constants save remembering bit positions: ``` ! Command: tcpdump -i ens224 -nn -vv -c 1 \ ! 'tcp[tcpflags] & tcp-syn != 0 and tcp[tcpflags] & tcp-ack = 0' 14:57:13.608571 IP (tos 0xb8, ttl 64, id 59487, offset 0, flags [DF], proto TCP (6), length 60) 192.168.99.100.45308 > 192.168.99.2.22: Flags [S], cksum 0x47e6 (incorrect -> 0xcdf1), seq 2288597127, win 64240, options [mss 1460,sackOK,TS val 800626090 ecr 0,nop,wscale 7], length 0 1 packet captured 1 packet received by filter ``` SYN set and ACK clear means connection openers only, SYN-ACKs excluded. That filter answers "who is initiating connections here", which catches a service that keeps reconnecting, and it is the defender's view of [what a SYN scan looks like from the far end](https://www.pinglabz.com/nmap-port-scanning/). If you want to build the detection out on the network gear rather than the host, [catching scans from the Cisco side](https://www.pinglabz.com/detecting-nmap-scans-cisco-blue-team/) takes the same idea to ACLs and NetFlow. The `cksum 0x47e6 (incorrect -> 0xcdf1)` is not a bug, it is checksum offload: the NIC fills the checksum in after tcpdump has seen the frame, so locally originated packets nearly always look wrong. Ignore it on transmit, take it seriously on receive. A refused connection is just as readable. Here is R1's HTTP port answering a SYN with a reset: ``` ! Command: tcpdump -i ens224 -nn -c 4 'tcp port 80' 14:57:16.946378 IP 192.168.99.100.47292 > 192.168.99.1.80: Flags [S], seq 2550366896, win 64240, options [mss 1460,sackOK,TS val 2988725570 ecr 0,nop,wscale 7], length 0 14:57:16.950538 IP 192.168.99.1.80 > 192.168.99.100.47292: Flags [R.], seq 0, ack 2550366897, win 0, length 0 2 packets captured ``` SYN answered by RST means nothing is listening, which is diagnostically different from a SYN with no answer at all (that is a drop: check ACLs, routing, firewall). R1's own `show ip http server status` reported the server Enabled on port 80\. The device's opinion and the wire disagreed, and the wire wins. ## Write a pcap, then filter it on read `-w` writes raw frames to a file instead of printing them, so nothing appears on the terminal while it runs, and `-s` sets the snaplen, the bytes kept per frame: ``` ! Command: tcpdump -i ens224 -nn -s 0 -w res0017-lab-edge.pcap -c 40 listening on ens224, link-type EN10MB (Ethernet), snapshot length 262144 bytes 40 packets captured 44 packets received by filter 0 packets dropped by kernel ! Command: file res0017-lab-edge.pcap res0017-lab-edge.pcap: pcap capture file, microsecond ts (little-endian) - version 2.4 (Ethernet, capture length 262144) ``` Modern tcpdump defaults to a 262144-byte snaplen, so `-s 0` ("whole frame") is redundant on a current box and essential on an old one, where the historical default of 68 or 96 bytes truncates every payload you were trying to read. Keep the habit. Set a small snaplen deliberately only for headers-only capture off a very high-rate link. That is a standard libpcap v2.4 Ethernet capture, so Wireshark opens it natively and so does tcpdump itself with `-r`. Which is the workflow that matters: capture wide once, then ask the file different questions afterwards. ``` ! Command: tcpdump -nn -r res0017-lab-edge.pcap 'host 192.168.99.3' reading from file res0017-lab-edge.pcap, link-type EN10MB (Ethernet), snapshot length 262144 14:57:58.064454 IP 192.168.99.100 > 192.168.99.3: ICMP echo request, id 9, seq 1, length 64 14:57:58.067156 IP 192.168.99.3 > 192.168.99.100: ICMP echo reply, id 9, seq 1, length 64 14:57:59.066514 IP 192.168.99.100 > 192.168.99.3: ICMP echo request, id 9, seq 2, length 64 14:57:59.070107 IP 192.168.99.3 > 192.168.99.100: ICMP echo reply, id 9, seq 2, length 64 14:58:00.068480 IP 192.168.99.100 > 192.168.99.3: ICMP echo request, id 9, seq 3, length 64 14:58:00.072044 IP 192.168.99.3 > 192.168.99.100: ICMP echo reply, id 9, seq 3, length 64 ``` Six frames out of forty, one target, no re-capture. The same file answering six questions, with the real counts: no filter (all frames)40 net 192.168.99.0/24 and icmp12 src host 192.168.99.10019 port 80 or port 2227 greater 100 (frames over 100 bytes)11 ether multicast1 A filter you got wrong at capture time cannot be undone. A filter you got wrong at read time costs one more command. Capture wider than you think you need, then narrow on read. `-q` is the fastest first look at an unfamiliar file, one short line per packet with the protocol detail stripped. ## Long captures: -G, -W, and the privilege drop that breaks them An intermittent problem needs a capture that runs for hours without filling the disk. `-G N` starts a new file every N seconds, `-W N` caps the number of files, and together they make tcpdump self-terminating, which is what you want in a script or a cron job. Except that on Debian it does this: ``` ! Command: sudo tcpdump -i ens224 -nn -w 'rot-%Y%m%d-%H%M%S.pcap' -G 10 -W 2 listening on ens224, link-type EN10MB (Ethernet), snapshot length 262144 bytes tcpdump: /home/j/plzcap/ev3/rot-20260724-150201.pcap: Permission denied ! Command: ls -l rot-*.pcap -rw-r--r-- 1 tcpdump tcpdump 426 Jul 24 15:02 rot-20260724-150149.pcap ``` The whole command ran under `sudo` and it still failed. Look at the owner of the one file that did get written: `tcpdump:tcpdump`, not root. tcpdump opens the capture socket as root, then drops privileges to the unprivileged `tcpdump` user (Debian's packaging default). The first file is opened before the drop; every rotation after it happens as a user that cannot write into your home directory. Two fixes, both captured: ``` ! Fix 1 - keep the privileges with -Z root ! Command: sudo tcpdump -i ens224 -nn -w 'zrot-%Y%m%d-%H%M%S.pcap' -G 10 -W 2 -Z root listening on ens224, link-type EN10MB (Ethernet), snapshot length 262144 bytes Maximum file limit reached: 2 8 packets captured 8 packets received by filter 0 packets dropped by kernel ! Command: ls -l zrot-*.pcap -rw-r--r-- 1 root root 650 Jul 24 15:02 zrot-20260724-150201.pcap -rw-r--r-- 1 root root 560 Jul 24 15:02 zrot-20260724-150211.pcap ``` Two files ten seconds apart, and tcpdump exited on its own after the second. Fix 2 is to write somewhere the `tcpdump` user can reach, such as `/tmp`, with no `-Z` at all. That is the better habit on a shared box, since `-Z root` leaves the capture files owned by root. The other half of a safe long capture is size, not time. `-C 100` rotates at 100 MB per file, and `-C` with `-W` gives you a true ring buffer that overwrites the oldest file and never grows. `-G` without `-W` rotates forever and fills the disk, which is a bad outcome on a jump host. If you are correlating a capture with load you generated deliberately, [driving the link with iperf3](https://www.pinglabz.com/iperf/) makes the timing far easier to line up. ## Payload: -A, -X, and a router password on the wire `-A` renders the TCP payload as ASCII, `-X` puts hex beside it. Both are useless on encrypted traffic and devastating on anything that is not. Here is a telnet login to R3, read back from a pcap with `tcpdump -nn -A -r telnet.pcap 'tcp port 23 and greater 60'`: ``` 15:04:00.955095 IP 192.168.99.3.23 > 192.168.99.100.56194: Flags [P.], seq 13:55, ack 1, win 32768, length 42 E..R;?....7...c...cd.....$......P....... User Access Verification Username: 15:04:02.919204 IP 192.168.99.100.56194 > 192.168.99.3.23: Flags [P.], seq 45:51, ack 88, win 502, length 6 E...U.@.@..r..cd..c........F.$..P...G...admin 15:04:02.938662 IP 192.168.99.3.23 > 192.168.99.100.56194: Flags [P.], seq 93:105, ack 51, win 32755, length 12 E..4;P....7...c...cd.....$."...LP...o!.. Password: 15:04:04.924441 IP 192.168.99.100.56194 > 192.168.99.3.23: Flags [P.], seq 51:60, ack 105, win 502, length 9 E..1U.@.@..k..cd..c........L.$..P...G...Cisco123 ``` No decryption, no tooling, no privileged position: just a host on the same segment. The username `admin` and the password `Cisco123` sit in the payload as text, and because telnet echoes every keystroke the characters appear a second time in the server's replies, one per packet. tcpdump decodes the option negotiation inline too, including the part where the client asks for encryption and the router declines: ``` 15:04:00.954681 IP 192.168.99.3.23 > 192.168.99.100.56194: Flags [P.], seq 1:13, ack 1, win 32768, length 12 [telnet WILL ECHO, WILL SUPPRESS GO AHEAD, DO TERMINAL TYPE, DO NAWS] 15:04:00.967629 IP 192.168.99.3.23 > 192.168.99.100.56194: Flags [P.], seq 55:58, ack 31, win 32760, length 3 [telnet WONT ENCRYPT] ``` Compare that with the SSH session earlier, where the only readable thing was a version banner. This is the argument for `transport input ssh` in one capture, and it generalises: any management protocol without transport encryption, SNMPv2c and unauthenticated syslog included, is readable by anyone who can capture on the path. If you are building the case for cleaning that up, [hardening the management plane end to end](https://www.pinglabz.com/infrastructure-security/) is where the rest of it lives. Two notes. `greater 60` drops the pure ACKs so payload frames are not buried. And `-A` on a session carrying no payload prints header bytes as garbage rather than an error, so pointing it at a refused connection gives nonsense: the capture is telling you there was never any payload. ## Gotchas - **The rotation trap is a permissions problem, not a bug.** First rotated file succeeds, second fails with `Permission denied` even under sudo. Check the owner of the first file: `tcpdump:tcpdump` means you hit the privilege drop. Use `-Z root` or write where that user can reach. - **`ip[8] < 64` is noisier than the tutorials admit.** The classic "find traceroute" recipe matches the TTL byte, and on a router segment it fires on OSPF Hellos first, because IOS sends them with TTL 1\. The same filter did catch a real traceroute probe on another run (`192.168.99.100.50906 > 192.168.99.3.33434: UDP, length 32`), but tighten it to `udp and ip[8] < 64 and dst portrange 33434-33534`, or `icmp and ip[8] < 64` for the Windows flavour. - **Capturing on the wrong NIC.** On a multi-homed host `-D` takes two seconds and saves ten minutes of staring at an empty capture of the management LAN. - **Incorrect checksums on your own transmitted packets are normal.** Checksum offload means the NIC fills them in after tcpdump has seen the frame. - **Non-zero `packets dropped by kernel` invalidates the capture.** Tighten the filter, raise the buffer with `-B`, or write to faster storage before concluding anything about missing traffic. - **Never run tcpdump unbounded on a remote session.** Use `-c`, or `-G` plus `-W`. If you drive it from an automation harness that cuts commands off after a timeout, detach it with `setsid nohup ./capture.sh > /tmp/cap.log 2>&1 < /dev/null &`; plain `nohup` with an ampersand is not enough on its own. ## Key Takeaways - Permissions first: a stock Debian binary has no setuid bit and no `cap_net_raw`, so `You don't have permission to perform this capture` is a sudo problem, not a filter problem. - `-nn -i -c N` is the shape of every ad hoc capture. Names cost speed and inject your own DNS into the capture. - BPF filters run in the kernel, so `received by filter` equalling `captured` proves you narrowed at the right layer. Reach for `tcp[tcpflags]` and `icmp[icmptype]` when the keywords run out. - Capture wide with `-w` and full snaplen, then ask the file different questions with `-r`. A wrong capture-time filter is unrecoverable. - Bound long captures on both axes: `-G` plus `-W` for time, `-C` plus `-W` for size, and remember the privilege drop that kills the second rotated file. - `-A` against telnet shows the username and password as text, which is the strongest argument for encrypted management you will produce in ten seconds. tcpdump reads. When you need to write the packet as well, [crafting frames field by field with Scapy](https://www.pinglabz.com/scapy-packet-crafting-spoofing-cisco/) is the other half of the same skill, and both sit alongside the on-device capture tools in [the full guide to capturing and analysing traffic on Cisco and Linux](https://www.pinglabz.com/packet-analysis/). Build the flat lab segment with a routing protocol for background noise and every command on this page is reproducible in about five minutes. ### BGP Route Not in the Routing Table: Why It Is Not Installed URL: https://www.pinglabz.com/bgp-route-not-in-routing-table/ Last updated: 2026-08-01T18:36:14.000Z You run `show ip bgp` and the prefix is right there. You run `show ip route` for the same prefix and get `% Network not in table`. No neighbor has flapped, the session is Established, the path looks perfectly healthy, and traffic is still going somewhere else. That gap is not a bug and almost never a mystery. BGP keeps every path it is sent, hands the routing table exactly one path per prefix, and the RIB gets the final say on whether it accepts. A prefix that sits in BGP and never reaches the RIB failed at one of those two steps, and the leading characters on each line of `show ip bgp` tell you which. Most engineers skim straight past them. For the wider context on [how BGP chooses and installs paths](https://www.pinglabz.com/bgp/), the pillar covers the attribute machinery this article assumes. One boundary first, because these two problems get conflated constantly. If your neighbor never receives a prefix at all, that is the sending side, covered in [why a router is not advertising a prefix to its neighbor](https://www.pinglabz.com/troubleshoot-bgp-route-advertisement/). This article owns the receiving side: it arrived, you can see it, and it will not install. Everything below came off a live CML lab of three `iol-xe` routers on IOS XE 17.18.2, R1 in AS 65001 speaking eBGP to R2 in AS 65002, R2 speaking iBGP to R3. ## Read the status codes before you read anything else Every prefix line in `show ip bgp` starts with a two or three character status field, and that field is the entire diagnosis. It is explained in a legend at the top of the output that everyone reads once, in their first week, and scrolls past forever after. \*Valid. The path passed sanity checks and BGP is willing to consider it. \>Best. This is the one path BGP offers to the RIB and advertises to peers. iLearned from an iBGP peer. Installs at AD 200, not 20. rRIB-failure. BGP picked this path and the routing table refused it. sSuppressed by aggregation. A summary is being advertised instead. d / hDampened, or history. The prefix has been flapping and is being penalised. Two of those matter more than the rest here. A line with `*` and no `>` means the path is valid but is not the one BGP is offering the RIB, so the failure happened inside BGP. A line with `r` means the opposite: BGP selected a best path, tried to install it, and the RIB pushed back. Completely different fix lists, and the difference is one character. (Do not confuse the `i` in the status field with the trailing origin code at the end of the same line, which is also frequently `i`. Same letter, different column, unrelated meanings.) ## Cause 1: the next hop is not reachable This is the one. If you take a single reflex from this article, make it "check the next hop first", because unresolved next hops account for more not-installed BGP routes than everything else combined. Here is the broken state, R3 holding `100.100.100.0/24` from its iBGP peer R2: ``` R3# show ip bgp 100.100.100.0/24 BGP routing table entry for 100.100.100.0/24, version 0 Paths: (1 available, no best path) <-- one path, and BGP picked none of them Flag: 0x8100 Not advertised to any peer Refresh Epoch 1 65001 10.0.12.1 (inaccessible) from 10.0.23.2 (2.2.2.2) Origin IGP, metric 0, localpref 100, valid, internal rx pathid: 0, tx pathid: 0 R3# show ip route 100.100.100.0 % Network not in table ``` Three things there do all the work: the next hop reads `10.0.12.1 (inaccessible)`, the summary says `no best path` despite one path being available, and the attribute line says `valid, internal`. That last one trips people up, because the path is flagged valid and BGP still refuses to select it. Validity is about the update being well formed, selection is a separate step, and RFC 4271's decision process disqualifies any path whose next hop cannot be resolved in the RIB before the tie-breakers ever run. BGP does not discover topology, it advertises reachability and leans recursively on the local routing table to reach the next hop. R3 has no route to `10.0.12.1`, so everything stops there. Why is the next hop unreachable? `10.0.12.1` is R1's address on the R1 to R2 link, a subnet R3 has no route to. R2 passed the prefix from eBGP into iBGP and left the next hop untouched, which is specified behaviour rather than a bug (the assumption being that everyone inside the AS can reach the AS boundary). If your edge links are not in the IGP, that assumption is false and every downstream speaker inherits a path it cannot use. Confirm it in one command: take the next hop out of the BGP entry and run `show ip route 10.0.12.1`. If that returns `% Network not in table`, you are done diagnosing, with no need to look at attributes or filters at all. ## The fix, and what an installed path looks like Two honest options: advertise the edge subnet into your IGP so every internal router can resolve it, or have the border router rewrite the next hop to itself. The second is the standard answer, and the entire reason `next-hop-self` exists: ``` router bgp 65002 neighbor 10.0.23.1 next-hop-self ! clear ip bgp 10.0.23.1 soft out ``` Applied on R2, that rewrites the next hop to R2's own iBGP-facing address, which R3 reaches over a directly connected /30, and the soft clear pushes the updated attributes without tearing the session down. R3 immediately afterwards: ``` R3# show ip bgp 100.100.100.0/24 BGP routing table entry for 100.100.100.0/24, version 2 Paths: (1 available, best #1, table default) 65001 10.0.23.2 from 10.0.23.2 (2.2.2.2) Origin IGP, metric 0, localpref 100, valid, internal, best R3# show ip route 100.100.100.0 Routing entry for 100.100.100.0/24 Known via "bgp 65002", distance 200, metric 0 Tag 65001, type internal * 10.0.23.2, from 10.0.23.2, 00:00:17 ago ``` Same prefix, same peer, same AS path. The only change is a next hop the local RIB can resolve, and the path goes from `no best path` to `best #1` and lands in the routing table. Note `distance 200` and `type internal`, because that becomes cause five. If the fields in that entry are not second nature yet, ten minutes on [what each part of a Cisco routing table entry actually means](https://www.pinglabz.com/how-to-read-cisco-routing-table/) pays for itself, since half of BGP troubleshooting is really RIB reading. ## Cause 2: valid, but not best If the prefix has several paths and one carries `>`, the others are working as designed. BGP is not a load balancer by default: it picks a single winner, and only that winner is offered to the RIB and re-advertised. Every other path sits there as a warm standby. This becomes a real complaint when the winner is not the one you wanted, which produces the same phone call even though the prefix is in the RIB, just via a next hop nobody expected. Work down [the order BGP compares paths in](https://www.pinglabz.com/bgp-best-path-selection/) until you find the step that decided it, remembering that weight and local preference sit near the top, so a stray inbound route-map beats any amount of AS path engineering further down. If you want two paths installed, that is `maximum-paths`: opt-in, fussy about attribute equality, and not a fallback for a path that lost. ## Cause 3: it never reached this router in the first place Look again at one line in the broken capture: `Not advertised to any peer`. A path that is not best is not propagated, so R3 not only failed to install the prefix, it declined to pass it on, and a single unresolved next hop silently blackholes every device behind it. The related trap is iBGP split horizon. A route learned from an iBGP peer is never advertised to another iBGP peer, full stop, because there is no AS path loop prevention inside an AS (which is why iBGP needs a full mesh, or route reflectors to fake one). From the far end that looks identical to a filtering problem. Tell them apart at the middle router: prefix present there and absent downstream means split horizon or a non-best path, not a filter. Check the session is genuinely up while you are there, because a peer that keeps resetting shows prefixes appearing and vanishing, which reads as an install problem when it is really [a neighbor that is not staying in Established](https://www.pinglabz.com/bgp-neighbor-stuck-idle-active-connect/). ## Cause 4: synchronisation, and why it is probably not your problem The old synchronisation rule said a router must not use or advertise an iBGP-learned route until the same prefix appeared in its IGP, so that non-BGP routers in the middle of the AS would not black-hole transit traffic. It has been off by default since IOS 12.2(8)T and stays off on modern IOS XE, so on any current box this is not your problem. It earns a mention only because it appears in CCNP question banks and still lurks in configs copied from a 2003 template. ## Cause 5: something with a better administrative distance already owns the prefix This is what the `r` code exists for, and it is the most misread symptom in the list. RIB-failure does not mean BGP failed. It means BGP selected a best path, offered it to the routing table, and the routing table declined because it already holds that prefix from a source it trusts more. The numbers make it obvious. eBGP installs at AD 20, beating OSPF at 110 and EIGRP at 90\. iBGP installs at AD 200, losing to essentially every IGP and to any static route. Look back at the working capture: `distance 200`. Had that prefix also been present via OSPF, OSPF would have won and the BGP path would be flagged `r` instead. The asymmetry is deliberate, and the full picture of [which routing source wins when two protocols offer the same prefix](https://www.pinglabz.com/administrative-distance/) is worth keeping to hand. Reach for `show ip bgp rib-failure`, which lists every prefix BGP could not install and why. Note the sting in the tail: a RIB-failed route is still advertised to your peers, so you are telling the world you can reach a prefix while forwarding it by a different protocol's idea of the topology. ## Cause 6: inbound filtering, which looks like the same problem If the prefix is not in `show ip bgp` at all, the update was either never sent or discarded on arrival. Prove which by comparing what the peer sent against what you kept. With `neighbor x.x.x.x soft-reconfiguration inbound` configured, `show ip bgp neighbors x.x.x.x received-routes` gives the pre-policy view and `show ip bgp neighbors x.x.x.x routes` the post-policy view. A prefix in the first and not the second is your own prefix-list, distribute-list, or route-map eating it. The subtler version does belong here. An inbound route-map can leave a prefix visible but unusable, most cruelly with `set ip next-hop` pointing at an address the router cannot resolve, reproducing the `(inaccessible)` symptom from cause one on a router where the peering is fine and the neighbor is blameless. If the next hop is not an address you expected, read your own inbound policy before you call the other side. ## What this was captured on Three `iol-xe` nodes in Cisco Modeling Labs on IOS XE 17.18.2\. R1 in AS 65001 originates `100.100.100.0/24` from a loopback and peers eBGP with R2 over `10.0.12.0/30`. R2 in AS 65002 peers iBGP with R3 over `10.0.23.0/30` and starts without `next-hop-self`, which is the entire break, and R3 has no route to the edge link. Output was collected on-box by an EEM applet writing `show` results to syslog, read back through the CML console log API, so everything above is verbatim. ## Gotchas - **"valid" does not mean installable.** The capture shows `valid, internal` alongside `no best path` in one output. Validity and selection are different gates, with next-hop resolution between them. - **`(inaccessible)` ends the investigation.** Read the next-hop line before the attribute line, every time. - **eBGP next hops pass into iBGP unchanged.** If your edge links are not in the IGP, the symptom lands on routers that have no idea the eBGP session exists. - **A non-best path is not advertised onward.** One router's next-hop failure shows up downstream as a total absence of the prefix. - **Use a soft clear.** `clear ip bgp soft out` re-sends updates without dropping the session. A hard `clear ip bgp *` on a production edge is a resume-generating event. ## Key takeaways - The BGP table and the RIB are separate structures. A prefix can live in one and not the other, permanently, with nothing logging an error. - Read the status characters first. No `>` means BGP did not select the path, an `r` means the RIB refused it. Different problems, different fixes. - Check next-hop reachability before anything else. `show ip route ` answers the most common cause in one command, and `next-hop-self` is the standard fix. - An installed iBGP route carries AD 200, so any IGP holding the same prefix wins and pushes BGP into RIB-failure while you keep advertising the prefix to peers. - If the prefix is not in `show ip bgp` at all, that is the other problem, and the diagnosis moves to the sending side and your inbound policy. Next hop, then best path, then RIB. Three checks in that order resolve almost every instance of this, and all three are quicker than reading configuration. The rest of the [BGP troubleshooting and design series](https://www.pinglabz.com/bgp/) covers the attributes and neighbor states either side of this one. ### BGP Neighbor Stuck in Active, Idle or Connect: Decode the State URL: https://www.pinglabz.com/bgp-neighbor-stuck-idle-active-connect/ Last updated: 2026-08-01T18:36:13.000Z You have a BGP session that will not come up, and the right-hand column of `show ip bgp summary` is showing you a word instead of a number. That word is the whole reason you are here: Idle, Active or Connect, sitting there and not moving. This article is organised around that word, because the state you are staring at is the fastest way to narrow the search, and because two of the three names mean roughly the opposite of what they sound like. Start with the naming, since it trips up almost everyone. **Active does not mean the session is active.** It means the router is actively trying to open a TCP connection and failing, which is a broken state. **Idle does not mean waiting politely.** It is where the FSM falls back to when something has gone wrong, and where a session starts before it has any reason to try. Only Established is good. If you want the formal machine, the [full walk through every BGP neighbor state from Idle to Established](https://www.pinglabz.com/bgp-neighbor-states/) covers the transitions in order; this page is about what to do when one of them will not advance. For the wider protocol picture, the [complete BGP configuration and troubleshooting guide](https://www.pinglabz.com/bgp/) maps the whole cluster in reading order. Everything below was captured on live routers: two `iol-xe` nodes running IOS XE 17.18.2 in Cisco Modeling Labs, back to back over 10.0.12.0/30, peering iBGP in AS 65001 across their loopbacks. One of the three labs produced a finding that changes how you should read the state column, so read to the end before you change any configuration. ## First, read the summary correctly The `State/PfxRcd` column is doing double duty, and that is the source of a lot of wasted time. It shows either a state name or a prefix count, and which one you get is itself the diagnosis: - **A word** (Idle, Connect, Active, OpenSent, OpenConfirm) means the session is *not* up. There is nothing to receive prefixes into. - **A number** means the session is Established and that is how many prefixes the peer has sent you. A session showing `0` is up and healthy, it just has nothing to advertise yet. Here is the broken lab, held stable across samples from 06:34:33 to 06:37:03 so there was no argument about it being transient: ``` R1# show ip bgp summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 2.2.2.2 4 65001 0 0 1 0 0 never Idle ``` Three numbers on that line matter as much as the state. `MsgSent 0` and `MsgRcvd 0` mean not a single BGP message has ever crossed, so this is not a policy, capability or authentication problem (those all need a TCP session first). And `Up/Down never` means the session has never been Established in the lifetime of this process, which rules out flapping and a peer that reset you. ## What each state is actually claiming Each state is a statement about how far the router got before it stopped, and each one points at a different half of the problem. Idle Router isNot trying Usual causeNo route to peer First checkshow ip route peer Connect Router isWaiting on TCP Usual causeSYN unanswered First checkPeer BGP running Active Router isRetrying, failing Usual causeTCP 179 blocked First checkACLs, update-source OpenSent / OpenConfirm Router isNegotiating Usual causeAS or RID mismatch First checkLast NOTIFICATION The line to draw is between the first three and the last two. Idle, Connect and Active all sit below the TCP layer, so the problem is transport: routing, filtering, source address, TTL. OpenSent and OpenConfirm mean TCP succeeded and the routers are arguing about the OPEN message, so the problem is configuration: a wrong `remote-as`, a duplicate router ID, an authentication mismatch. Those two halves need completely different investigations, and knowing [why BGP runs over a TCP session in the first place](https://www.pinglabz.com/how-bgp-works/) makes the split obvious rather than arbitrary. Connect deserves one note because it is the state people see least. A router in Connect has sent a SYN and is sitting on the ConnectRetry timer waiting for the handshake, so catching a session there usually means the far end is silently discarding the SYN or its BGP process is not listening. Connect is short lived, and the FSM drops back to Active or Idle when the retry timer expires, which is why you rarely find a session parked in it. ## Case one: Idle because there is no route to the neighbor This is the classic, and in the lab it is the iBGP-over-loopbacks trap. R1 peers with 2.2.2.2, which is R2's Loopback0, but nothing advertises the loopbacks into an IGP and there is no `update-source` statement. Here is what the router says when you ask it properly: ``` R1# show ip bgp neighbors 2.2.2.2 BGP state = Idle, down for never Last reset never No active TCP connection R1# show ip route 2.2.2.2 % Network not in table ``` That is the diagnosis in two lines. `% Network not in table` means R1 has no idea how to reach the address it was told to peer with, so it never attempts the TCP connection, which is exactly what `No active TCP connection` reports. Nothing is wrong with the BGP configuration. The routing table is the problem, and BGP is the thing complaining about it. Peering over loopbacks needs two things and people routinely supply only one. You need a route to the peer's loopback, normally from the IGP, and you need `update-source Loopback0` so your packets are sourced from the address the peer expects. Miss the route and you get what is above. Miss the update-source and your SYN arrives from the physical interface address, the peer does not recognise it as a configured neighbor, and refuses the connection. Both look identical in the summary. The [correct way to build an iBGP peering across loopbacks](https://www.pinglabz.com/configure-ibgp-neighbors/) is worth a re-read if you inherited the config rather than writing it. ## The fix, and the proof that it was the fix The repair is the two things above, applied together: ``` router ospf 1 network 2.2.2.2 0.0.0.0 area 0 ! router bgp 65001 neighbor 2.2.2.2 update-source Loopback0 ``` What makes this convincing is not that the session came up, but the order in which it came up. In the fixed lab the routers booted with BGP already configured and OSPF not yet converged, so the session started out looking exactly like the broken lab: ``` *Jul 20 06:38:45: (still Idle at boot - OSPF not converged yet) R1# show ip bgp summary 2.2.2.2 4 65001 0 0 1 0 0 never Idle R1# show ip route 2.2.2.2 % Network not in table ``` Then OSPF finished loading and BGP came up 353 milliseconds later, unprompted: ``` *Jul 20 06:39:22.595: %OSPF-5-ADJCHG: Process 1, Nbr 2.2.2.2 on Ethernet0/0 from LOADING to FULL, Loading Done *Jul 20 06:39:22.948: %BGP-5-ADJCHANGE: neighbor 2.2.2.2 Up ``` That gap is the causal link. BGP was not reconfigured, cleared or touched. The instant the IGP installed 2.2.2.2/32 in the RIB, the FSM walked Idle to Connect to OpenSent to OpenConfirm to Established on its own. If your session is waiting on [an OSPF adjacency reaching FULL](https://www.pinglabz.com/ospf/), fix the IGP and stop looking at BGP. ``` R1# show ip bgp summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 2.2.2.2 4 65001 2 2 1 0 0 00:00:12 0 R1# show ip bgp neighbors 2.2.2.2 BGP state = Established, up for 00:00:13 R1# show ip route 2.2.2.2 Routing entry for 2.2.2.2/32 Known via "ospf 1", distance 110, metric 11, type intra area * 10.0.12.2, from 2.2.2.2, 00:00:14 ago, via Ethernet0/0 ``` Note the `0` in the State/PfxRcd column. That is a healthy session receiving zero prefixes, not a broken one. If you have got this far and your peer is up but nothing is in the RIB, that is a different article: [why a BGP route is missing from the routing table](https://www.pinglabz.com/bgp-route-not-in-routing-table/) is the failure you hit next. ## Case two: the route is there and the session still will not come up The third lab is the one that changes how you should read the state column. Both routers have `update-source Loopback0`, OSPF has converged and the peer loopback is in the RIB. The only difference is an inbound ACL on R2: ``` ip access-list extended BLOCKBGP deny tcp host 1.1.1.1 host 2.2.2.2 eq bgp deny tcp host 1.1.1.1 eq bgp host 2.2.2.2 permit ip any any ! interface Ethernet0/0 ip access-group BLOCKBGP in ``` That `permit ip any any` at the bottom is why this is nasty in production. Pings work, traceroute works, OSPF hits FULL. Only TCP 179 between those two addresses disappears, and it disappears silently, with no ICMP unreachable to tell the sender anything. Here is what R1 reported, held for over four minutes of samples: ``` R1# show ip bgp summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 2.2.2.2 4 65001 0 0 1 0 0 never Idle R1# show ip bgp neighbors 2.2.2.2 BGP state = Idle, down for never Connections established 0; dropped 0 Last reset never No active TCP connection R1# show ip route 2.2.2.2 Routing entry for 2.2.2.2/32 Known via "ospf 1", distance 110, metric 11, type intra area * 10.0.12.2, from 2.2.2.2, via Ethernet0/0 ``` Read the state, then read the route. The state is **Idle**, identical to case one, and the route is present. That combination is the finding: **on this IOS XE 17.18 build, a session with TCP 179 blocked reported Idle, not Active, for the entire test.** Every article you have read (including the textbook, and including the state table earlier on this page) says a router failing to complete TCP shows Active. Classic IOS usually does. This build did not. So the state name narrows the search but it does not identify the cause. What identifies the cause is the pair of outputs, not either one alone. ## The two-command triage that actually separates the causes Run these two on the router showing the stuck state, and read them together: Stuck state + "% Network not in table"No route to the peer. Fix the IGP or the update-source. Stuck state + route present + "No active TCP connection"TCP 179 is being dropped. Hunt the ACL, firewall or ttl-security. Route present + TCP connected + OpenSentTransport is fine. Read the last NOTIFICATION for the real error. The commands themselves are unglamorous: ``` show ip bgp summary show ip bgp neighbors ! read "BGP state" and the TCP connection line show ip route ! is there a route at all? ping source Loopback0 ! test from the address BGP will actually use ``` That last one matters more than it looks. A plain `ping 2.2.2.2` sources from the outgoing interface and can succeed while the BGP session, sourced from the loopback, is dropped by a filter matching the loopback address. And if the route is present but TCP will not open, the filter is not always on your side: in this lab the ACL was on the *peer's* inbound interface and R1 had no way to see it. Check both ends, and check [how an inbound ACL is evaluated on IOS XE](https://www.pinglabz.com/advanced-acls-cisco-ios-xe/) before you assume a rule that permits everything else must be innocent. ## What this was captured on R1 and R2 are `iol-xe` nodes running IOS XE 17.18.2 in CML, connected back to back on `Ethernet0/0` in 10.0.12.0/30, peering iBGP in AS 65001 between Loopback0 addresses 1.1.1.1 and 2.2.2.2\. Three separate labs were built: one with no IGP and no update-source (stuck Idle), one with OSPF area 0 advertising the loopbacks plus `update-source Loopback0` (Established), and one identical to the second except for the inbound `BLOCKBGP` ACL on R2. One note on method. The PyATS console path was unavailable during this run, so `show` output was captured by an on-box EEM applet logging to syslog, read back through the CML console-log API. The text is verbatim device output either way, and the repeated sampling is why the stuck states are described as stable for minutes rather than as a snapshot. ## Gotchas - **The state name is a hint, not a verdict.** The same word covered two completely different root causes on the same software version. Pair the state with `show ip route ` every single time. - **Active is a failure state.** If you report "the neighbor is Active" in a change call and someone relaxes, correct them. The router is failing and retrying. - **A number in State/PfxRcd is good news, including zero.** A peer showing `0` is Established. Do not troubleshoot a working session because the column looks empty. - **MsgSent 0 rules out an entire class of causes.** No messages sent means no TCP session, so it cannot be an AS mismatch, a password mismatch or a capability problem (those all need OPEN messages to have crossed). The [order in which BGP messages are exchanged](https://www.pinglabz.com/bgp-message-types/) tells you which failures are still possible at a given state. - **Ping is not proof.** An ACL that denies only TCP 179 and permits everything else leaves the link passing every test a reasonable engineer runs first. - **Check the far end.** The blocking filter, the missing neighbor statement and the wrong `remote-as` are invisible from the router showing the symptom. - **If the session did come up once, this is not your article.** A session that reaches Established and then drops is a hold-timer, max-prefix or link-quality problem. The [systematic checklist for a BGP adjacency that flaps or gets rejected](https://www.pinglabz.com/troubleshoot-bgp-neighbor/) covers flapping, NOTIFICATION decoding and authentication in detail. ## Key takeaways - A word in `State/PfxRcd` means the session is down. A number, even `0`, means it is Established. - Idle, Connect and Active are pre-TCP states, so the cause is transport: routing, filtering, source address or TTL. OpenSent and OpenConfirm mean TCP worked and the OPEN was rejected. - Active is a broken state despite the name, and Idle is where the FSM falls back to when it has given up. - The state alone does not identify the cause. On IOS XE 17.18.2 both "no route to peer" and "TCP 179 blocked by an ACL" presented as Idle with `No active TCP connection` and `MsgSent 0`. - `show ip route ` is the command that separates them. No route means fix the IGP or the update-source. Route present means hunt the filter, at both ends of the link. - For iBGP over loopbacks you need the IGP route *and* `update-source Loopback0`. In the lab BGP came up 353 ms after OSPF reached FULL, with no BGP change at all. The habit worth building is to stop treating the state column as an answer and start treating it as the first of two data points. Collect the state, the route and the TCP connection line together, and most stuck sessions resolve in under two minutes. When the peering is up and you move on to why the prefixes are not where you expect them, the rest of the series, from [peering design through path attributes and best-path selection](https://www.pinglabz.com/bgp/), picks up from there. ### OSPF Stuck in INIT: Diagnosing the One-Way Hello URL: https://www.pinglabz.com/ospf-stuck-in-init-one-way-hello/ Last updated: 2026-08-01T18:36:13.000Z An OSPF neighbor sitting in INIT is one of the few troubleshooting states that hands you the answer for free, if you know how to read it. INIT is not "still negotiating" and it is not a slow start. It is a router telling you something very specific: **I am receiving your hellos, and you are not receiving mine.** The adjacency is broken in exactly one direction, and that single fact eliminates most of the causes people go hunting for. The mistake almost everyone makes is to stare at the INIT side. That router is the healthy one. The interesting device is the other end, and the diagnostic move that solves this in about ninety seconds is to run `show ip ospf neighbor` on *both* routers and compare the two tables. If you are still building your mental model of [how OSPF forms adjacencies](https://www.pinglabz.com/ospf/), the state machine is the part worth internalising, because every neighbor problem you will ever meet announces itself by getting stuck in one particular state. Everything below was captured on a live lab: two `iol-xe` routers running IOS XE 17.18.2 in Cisco Modeling Labs, back to back over a /30, with an inbound ACL deliberately dropping OSPF in one direction. No output on this page was typed by hand. ## The symptom: one router sees a neighbor, the other sees nothing Here is the broken state, held stable for about two minutes so there was no doubt it was a steady condition rather than a transient: ``` R1# show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 2.2.2.2 1 INIT/DROTHER 00:00:32 10.0.12.2 Ethernet0/0 R2# show ip ospf neighbor (no output - R2 has no OSPF neighbor at all) ``` Look at what those two tables say together. R1 knows R2 exists. It has R2's router ID, its interface address, and a dead timer that is counting down and being reset, which means hellos from R2 are arriving on schedule. R2, meanwhile, has an empty neighbor table. Not an error, not a partially formed neighbor. Nothing. That combination has exactly one meaning. Hellos are flowing from R2 to R1 and are not flowing from R1 to R2\. Everything else in this problem is downstream of that sentence. ## What INIT actually means An OSPF hello packet carries a neighbor list: the router IDs of every neighbor this router has recently heard on that segment. RFC 2328 section 10.5 defines the receive logic in a way that is easy to state and easy to forget. When a router receives a hello, it checks whether its own router ID appears in that hello's neighbor list. If it does, the two routers have proved bidirectional reachability and the neighbor moves to 2WAY. If it does not, the neighbor stops at INIT. So INIT is not a timing artefact and it does not clear on its own. R1 is receiving hellos from R2 and R1's router ID (1.1.1.1) is not in them, because R2 has never heard from R1 and therefore has nothing to list. R1 records the neighbor, marks it INIT, and waits forever. The state [progresses from DOWN through 2WAY, EXSTART and LOADING to FULL](https://www.pinglabz.com/ospf-adjacency-states/) only after that mutual acknowledgement happens, so a one-way hello path stops the whole machine at the first gate. Two details in that R1 output are worth noticing. The dead time of `00:00:32` is healthy and refreshing, which tells you the reverse direction is fine and rules out a [hello and dead interval mismatch](https://www.pinglabz.com/ospf-timers-hello-dead-intervals/) on the path you can see. And the `DROTHER` role attached to the state is meaningless here, because DR election only runs once neighbors reach 2WAY. Do not read anything into it. ## Reading the direction, then reading the cause The asymmetry is the diagnosis, and it is what separates this failure from every other adjacency problem. Most neighbor mismatches are symmetric in their symptoms. An area ID mismatch, a subnet mask mismatch, an MTU mismatch, mismatched timers: in all of those, both routers receive each other's packets and both have something to complain about, whether that is a log message or a neighbor stuck in EXSTART. Both sides *see something*. A completely empty neighbor table on one end while the other end holds a stable INIT is different. That is a packet delivery problem, not a parameter problem. Something between R1's OSPF process and R2's OSPF process is eating packets in one direction only, and your search space is now the list of things that can be applied to a single direction: Inbound ACL on the silent router BlocksIP proto 89 Check withshow access-lists TellPings still work Passive interface on one end BlocksHellos sent Check withshow ip ospf interface TellPassive router sees peer Auth applied on one end only BlocksHellos discarded Check withdebug ip ospf adj TellMismatch logs one side Broken multicast in one direction Blocks224.0.0.5 Check withdebug ip ospf hello TellUnicast ping is fine Two of those deserve a note. Authentication is the one that fools people, because it can present as either symmetric or one-way depending on how the mismatch was created: if one side has authentication configured and the other has none, the configured side silently discards the unauthenticated hellos while its own hellos are quite happily accepted by the peer that is not checking. That is a genuine one-way hello with no packet filter anywhere, and it is worth reading up on [how OSPF authentication mismatches present in the logs](https://www.pinglabz.com/troubleshoot-ospf-authentication-mismatch/) before you go looking for an ACL that does not exist. The [passive-interface case](https://www.pinglabz.com/ospf-passive-interface/) is the inverse shape and the easiest to spot once you know it. A passive interface still listens, so the passive router is the one that ends up with the neighbor in INIT while the other router, which never hears a hello at all, sits with an empty table. ## What this was captured on R1 and R2 are `iol-xe` nodes on IOS XE 17.18.2 in CML, connected back to back on `Ethernet0/0` in 10.0.12.0/30, both in OSPF area 0, with router IDs 1.1.1.1 and 2.2.2.2\. The break was baked into R2 before the routers came up: ``` ! Applied inbound on R2 Ethernet0/0 ip access-list extended OSPF-BLOCK deny ospf host 10.0.12.1 any permit ip any any ``` That `permit ip any any` is the whole reason this failure is nasty in production. Only IP protocol 89 from one source is dropped, so ARP resolves, pings succeed both ways, the interface is up/up with no errors, and the link passes every test a reasonable engineer runs before deciding the problem must be "an OSPF thing". If you are auditing filters on a transit link, it is worth knowing [how an inbound ACL is evaluated on IOS XE](https://www.pinglabz.com/advanced-acls-cisco-ios-xe/) and which control-plane traffic an over-broad rule can quietly swallow. One note on method, for honesty: `show ip ospf neighbor` was logged to syslog by an on-box EEM applet and read back through the CML console-log API, because the PyATS console path was unavailable during this run. The text is verbatim device output either way. ## The fix and the proof Removing the filter is one line on the interface: ``` R2(config)# interface Ethernet0/0 R2(config-if)# no ip access-group OSPF-BLOCK in ``` The recovery is immediate, and this is the part to pay attention to, because it is what confirms the diagnosis rather than merely ending the outage. Five seconds after the ACL came off, both routers logged the transition, within a millisecond of each other: ``` *Jul 20 12:51:42: %SYS-5-CONFIG_I: Configured from console by on vty1 (EEM:FIX) FIX-DONE removed OSPF-BLOCK inbound ACL; expect FULL *Jul 20 12:51:47.620: %OSPF-5-ADJCHG: Process 1, Nbr 2.2.2.2 on Ethernet0/0 from LOADING to FULL, Loading Done (R1) *Jul 20 12:51:47.621: %OSPF-5-ADJCHG: Process 1, Nbr 1.1.1.1 on Ethernet0/0 from LOADING to FULL, Loading Done (R2) ``` R1 did not need to be touched, restarted, or cleared. It had been holding a valid neighbor entry the entire time, waiting for one hello that listed its router ID. The moment R2 could hear R1, R2 started including 1.1.1.1 in its own hellos, R1 moved to 2WAY on the next one it received, and the rest of the state machine ran to completion in under a second. Both tables agree afterwards: ``` R1# show ip ospf neighbor 2.2.2.2 1 FULL/DR 00:00:38 10.0.12.2 Ethernet0/0 R2# show ip ospf neighbor 1.1.1.1 1 FULL/BDR 00:00:37 10.0.12.1 Ethernet0/0 ``` Note that the DR election ran only now, and R2 won it. During the whole INIT period there was no election at all, which is why the earlier `DROTHER` label was noise. If you want the full breakdown of [what each OSPF neighbor state means and what stalls it](https://www.pinglabz.com/ospf-neighbor-states-explained/), that is the reference to keep open next to the CLI. ## Gotchas - **Do not troubleshoot the INIT router.** The router showing INIT is working correctly. Its peer, the one with the empty neighbor table, is where the filter, the passive interface or the authentication statement lives. - **An empty `show ip ospf neighbor` is not "OSPF is down".** It is easy to read blank output as a dead process and start re-checking `network` statements. Confirm the interface is actually in OSPF with `show ip ospf interface` before you rewrite any configuration. - **Ping proves nothing about OSPF.** A filter that drops protocol 89 and permits everything else leaves the link looking flawless to every generic connectivity test. - **Symmetric symptoms mean a different problem.** If both routers can see each other and are stuck together, you are not looking at a one-way hello. Stuck in EXSTART or EXCHANGE on both ends points at [an OSPF MTU mismatch](https://www.pinglabz.com/ospf-mtu-mismatch-troubleshooting/), and neighbors that keep reaching FULL and then dropping usually mean [two routers sharing a router ID](https://www.pinglabz.com/fix-duplicate-ospf-router-id/). - **Prove the direction before you change anything.** `debug ip ospf hello` on both routers, or [an on-box packet capture on IOS XE](https://www.pinglabz.com/embedded-packet-capture-ios-xe/), will show you hellos arriving on one router and never on the other. That is a two-minute check that saves you from a maintenance window spent editing OSPF config that was never wrong. ## Key takeaways - INIT means "I hear you, you do not hear me". The neighbor's hello arrived, but it did not list your router ID, so the adjacency cannot advance to 2WAY. - Run `show ip ospf neighbor` on both ends. INIT on one side plus an empty table on the other is a one-way packet delivery problem, not a parameter mismatch. - The silent router is the broken one. Look for an inbound ACL dropping IP protocol 89, a passive interface, authentication configured on only one end, or multicast to 224.0.0.5 failing in one direction. - An ACL that ends in `permit ip any any` will pass every ping and traceroute you throw at it while still killing the adjacency. - Recovery is automatic and near instant once hellos flow both ways. In the lab both routers hit FULL five seconds after the filter was removed, with no clear or reload on either side. - DR and BDR roles shown alongside INIT are meaningless, because election does not run until neighbors reach 2WAY. One-way hellos are the cleanest failure in the adjacency family precisely because the evidence is asymmetric, and asymmetry always points at direction. Get in the habit of collecting both neighbor tables before you form a theory. For the rest of the protocol, from area design to LSA behaviour and the other ways adjacencies fail, work through the [complete OSPF configuration and troubleshooting guide](https://www.pinglabz.com/ospf/). ### Netmiko vs NAPALM vs pyATS: Three Ways to Diff a Cisco Change URL: https://www.pinglabz.com/netmiko-vs-napalm-vs-pyats/ Last updated: 2026-08-01T17:24:30.000Z You push a change to three routers, the script scrolls past without a single `%` line, and everything looks fine. Now answer the only question that matters at 2am: did that run actually change anything, and is the network doing what you intended? Most automation arguments about Netmiko, NAPALM and pyATS are really arguments about that one question, because each of the three answers it at a completely different layer. This is not a feature table copied from three README files. All three tools were run against the same three routers, on the same segment, on the same afternoon, from the same Debian box. That is what makes the comparison worth reading, and it is the reason this sits in the middle of [how the Cisco automation stack actually fits together](https://www.pinglabz.com/network-automation/) rather than off to one side. The lab is a CML topology called VM Automation Edge: R1, R2 and R3 are `iol-xe` nodes running IOS-XE 17.18.2 with OSPF 1 area 0, plus an `ioll2-xe` switch, all bridged onto `192.168.99.0/24` with a Debian 13 VM sitting on the same broadcast domain at `192.168.99.100`. Short version of the finding: Netmiko answers at the session layer, NAPALM answers at the configuration layer, and pyATS Genie answers at the operational state layer. They are three rungs on a ladder, not three competitors. ## Netmiko: the session layer Netmiko is an SSH session with the sharp edges filed off. It logs in, disables the pager, sends what you tell it to send, and hands you back a string. That is genuinely useful, and for reading a device it is all you need. Push a loopback with `send_config_set()` and you get the whole config session echoed back: ``` ! Command: conn.send_config_set(config_commands) configure terminal Enter configuration commands, one per line. End with CNTL/Z. R1(config)#interface Loopback14 R1(config-if)# description PingLabz RES-0014 pushed by netmiko R1(config-if)# ip address 14.14.14.1 255.255.255.255 R1(config-if)#end R1# ``` The absence of a `%` line is how you know every command was accepted. Now run the exact same list a second time, with nothing changed in between: ``` ! Command: conn.send_config_set(config_commands) # second, identical run configure terminal Enter configuration commands, one per line. End with CNTL/Z. R1(config)#interface Loopback14 R1(config-if)# description PingLabz RES-0014 pushed by netmiko R1(config-if)# ip address 14.14.14.1 255.255.255.255 R1(config-if)#end R1# ``` Byte for byte the same output. IOS re-accepted every line, cheerfully, because the IOS CLI is imperative and does not care that the interface already exists. The return value of `send_config_set()` tells you precisely nothing about whether the device moved. Netmiko has no concept of desired state, so the only way to prove "no change" is to grab `show running-config` before and after and diff it yourself with `difflib`: ``` ! Command: difflib.unified_diff(show run before 2nd push, show run after 2nd push) ``` For contrast, that same diff after the *first* push is not empty at all: ``` --- running-config BEFORE 1st run +++ running-config AFTER 1st run @@ -20,2 +20,5 @@ ip ospf 1 area 0 +interface Loopback14 + description PingLabz RES-0014 pushed by netmiko + ip address 14.14.14.1 255.255.255.255 interface Ethernet0/0 ``` So Netmiko can tell you what changed, but only because you built the diff engine in Python and pointed it at scraped text. That distinction is the whole reason to reach for [driving Cisco SSH sessions from Python one command at a time](https://www.pinglabz.com/netmiko-ssh-automation-cisco/) when you need the flexibility, and to reach for something else when you need a guarantee. Netmiko does have `use_textfsm=True`, which turns the CLI table into a list of dicts and gets you out of regex territory (the same reason [structured data beats screen scraping](https://www.pinglabz.com/ccna-auto-04-json-structured-data/) everywhere else in automation), but parsing output is not the same thing as knowing whether you changed the box. ## NAPALM: the configuration layer NAPALM moves the question one layer up. You hand it a candidate configuration, and before anything is applied it asks the device what the difference would be. Driven from a Nornir inventory with `napalm_configure(..., dry_run=True)`, that is this: ``` napalm_configure DRY RUN******************************************************** * R1 ** changed : True ********************************************************* vvvv napalm_configure DRY RUN ** changed : True vvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvv INFO +interface Loopback1 + description NORNIR-MANAGED pinglabz-lab + ip address 10.99.1.1 255.255.255.0 + ip ospf 1 area 0 +snmp-server location PingLabz-Lab-Rack1 +snmp-server contact netops@pinglabz.com ^^^^ END napalm_configure DRY RUN ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ``` The important detail is where that diff came from. It was not computed in Python. NAPALM copied the candidate to `unix:candidate_config` over SCP and ran `show archive config incremental-diffs` on the router. The device generated its own diff. That is why it is trustworthy, and it is also why `archive` has to be enabled on the box before any of this works. Commit it (`dry_run=False`) and all three hosts come back `changed=True` with the same block of `+` lines. Then run the identical task again: ``` napalm_configure RE-RUN (identical input)*************************************** * R1 ** changed : False ******************************************************** vvvv napalm_configure RE-RUN (identical input) ** changed : False vvvvvvvvvvvvvv INFO ^^^^ END napalm_configure RE-RUN (identical input) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ! Command: python -> (result[host][0].changed, len(result[host][0].diff)) R1 changed=False diff_len=0 diff='' R2 changed=False diff_len=0 diff='' R3 changed=False diff_len=0 diff='' ``` Nothing between the `vvvv` and `^^^^` markers, because there was nothing to say. NAPALM asked the device for the incremental diff, got an empty string back, and skipped the commit entirely. That is real idempotence, and note how different it is from the Netmiko result: not "the commands were re-accepted without error", but "no configuration commands were sent at all". Same three routers, same segment, same day. If you have ever compared [how declarative config management tools differ from a script that just runs commands](https://www.pinglabz.com/ansible-vs-puppet-vs-chef/), this is that argument reduced to two evidence files. The other thing you get for free is a rendered-per-host inventory. `snmp_location` was defined once on the `ios_edge` group and resolved on all three devices, while each router got its own `mgmt_loopback` address, which is the practical case for [keeping a real inventory instead of a list of IP addresses](https://www.pinglabz.com/nornir-napalm-cisco-config-management/). ## pyATS and Genie: the operational state layer NAPALM will happily tell you a change was applied. It will not tell you the interface came up. That is a different question, and it is the one pyATS answers. Genie's `learn()` is not a parser for one command. It runs several `show` commands and merges them into a feature model, so `dev.learn('interface')` gives you what the device is currently doing. Snapshot before a change: ``` ! Command: pre = dev.learn('interface') ! Command: python -> pre.info['Loopback14'] { "bandwidth": 8000000, "description": "PingLabz RES-0014 pushed by netmiko", "enabled": true, "oper_status": "up" } ``` Change the description and shut the interface, snapshot again, and feed both into `Diff()`. The naive version is a mess, and that mess is the most useful thing in this whole article: ``` ! Command: d = Diff(pre.info, post.info); d.findDiff(); print(d) Ethernet0/0: counters: - in_octets: 179470 + in_octets: 191412 - in_pkts: 2396 + in_pkts: 2585 - out_octets: 328888 + out_octets: 366986 Loopback14: - description: PingLabz RES-0014 pushed by netmiko + description: CHANGED BY PYATS RES-0002 - enabled: True + enabled: False ``` The change you made is in there, buried under interface counters that move every second (three OSPF speakers on one segment guarantee traffic). A pre and post test written on that raw diff fails one hundred percent of the time for reasons that have nothing to do with the change under test. Exclude the volatile keys and you get the money shot: ``` ! Command: Diff(pre.info, post.info, exclude=['counters', 'accounting', 'rate', 'last_change', 'in_rate', 'out_rate']) Loopback14: - description: PingLabz RES-0014 pushed by netmiko + description: CHANGED BY PYATS RES-0002 - enabled: True + enabled: False - oper_status: up + oper_status: down ``` `oper_status` is the field neither of the other two tools has. Netmiko diffs config text you scraped. NAPALM diffs a candidate configuration. Only pyATS catches "the config applied cleanly but the interface never came up", because it is comparing measured state, not intent. Then roll the change back and run the same diff a third time: ``` ! Command: Diff(pre.info, rolled_back.info, exclude=['counters', 'accounting', 'rate', 'last_change', 'in_rate', 'out_rate']) ``` An empty diff there is not a boring result. It is the proof that the rollback was complete, measured against the state you actually recorded rather than an assumption that `no shutdown` did what it says. That is the shape of every useful [pre and post change validation test](https://www.pinglabz.com/pyats-genie-network-testing/): learn, change, learn, diff, and expect empty when you undo it. ## The three side by side Netmiko 4.5.0 AbstractsThe SSH session Reports change asNothing. You diff it Needs installednetmiko, ntc-templates CML gotchafast\_cli, silent TextFSM keys Reach for it whenThe one-off nothing else does NAPALM 5.1.0 + Nornir 3.5.0 AbstractsThe configuration Reports change aschanged + device diff Needs installednornir, nornir\_napalm, napalm CML gotchadest\_file\_system: "unix:" Reach for it whenPushing intent to many boxes pyATS + Genie 26.6 AbstractsOperational state Reports change asDiff() of learned models Needs installedpyats\[full\], testbed.yaml CML gotchassh\_config legacy algorithms Reach for it whenValidating a change landed ## What this was captured on One CML lab, one L2 segment, no routing between the automation host and the devices. R1, R2 and R3 are `iol-xe` on IOS-XE 17.18.2 at `192.168.99.1-3`, SW1 is `ioll2-xe` at `.4`, and the Debian 13 VM sits at `192.168.99.100` on an interface bridged into CML. Login is `admin` at privilege 15 with no enable secret, which is why `find_prompt()` returns `R1#` and not `R1>`. Netmiko 4.5.0 was the Debian system package on Python 3.13.5; Nornir 3.5.0, nornir\_napalm 0.5.0, NAPALM 5.1.0 and pyATS 26.6 all installed from wheels into a venv on the same interpreter. ## Gotchas that stop each tool before it starts - **NAPALM on IOL will not even reach the diff.** `iol-xe` has no `flash:` filesystem (`show file systems` lists `nvram:`, `unix:`, `disk0:`, `disk1:`, `system:`), and the napalm-ios driver defaults `dest_file_system="flash:"`. Every `load_merge_candidate()` fails until the inventory sets `optional_args: {dest_file_system: "unix:"}`. You also need `ip scp server enable` for the transfer and `archive` with `path unix:archive` for the diff engine itself. - **pyATS dies on SSH negotiation, not on pyATS.** A 2026 OpenSSH client will not talk to a 2020-era Cisco image. The fix is a `~/.ssh/config` stanza for `192.168.99.*` adding `KexAlgorithms +diffie-hellman-group14-sha1`, `HostKeyAlgorithms +ssh-rsa` and `PubkeyAcceptedAlgorithms +ssh-rsa`. It applies to Netmiko and NAPALM too. - **TextFSM fails silently.** The ntc-templates template for `show ip interface brief` names its columns `intf`, `ipaddr`, `status` and `proto`, not `interface` and `ip_address`. Guess wrong and every row prints `None` with no exception raised. Print `result[0].keys()` once before you write the loop. - **Netmiko 4.x defaults to `fast_cli=True`.** Over a bridged path to IOL that trimmed delay factor causes truncated reads. Every capture here used `fast_cli=False`. - **Genie `Diff()` without `exclude=` is unusable on a live segment.** Counters, accounting and rate keys move constantly. Exclude them or your test never passes. - **`napalm_get` forwards every kwarg to every getter.** `napalm_get(getters=['config','interfaces'], retrieve='running')` raises `TypeError: IOSDriver.get_interfaces() got an unexpected keyword argument 'retrieve'`. Getters with options need their own task. ## How the three fit in one pipeline Nobody picks one. A real change workflow uses NAPALM (usually through Nornir) to render intent from an inventory and push it, so the device generates the diff and a re-run is a no-op. It uses pyATS to learn the relevant feature model before and after and assert the operational diff contains exactly the intended change and nothing else. And it keeps Netmiko around for the awkward command a structured tool will not do: a vendor-specific `show` nothing has a parser for, a password recovery step, a box that only speaks Telnet. The same instinct applies to [running Python directly on the switch](https://www.pinglabz.com/on-box-python-ios-xe/) when the automation host cannot reach it at all. If you want the honest test of whether your pipeline is doing anything real, run it twice. A pipeline that cannot tell you the second run changed nothing is not automation, it is a fast way to type. ## Key Takeaways - Netmiko has no concept of desired state. An identical second `send_config_set()` produces an identical happy session echo, and only a `difflib` diff of `show running-config` proves the box did not move. - NAPALM's diff is generated by the device (`show archive config incremental-diffs`), so an identical second run returns `changed=False` with `diff=''` and sends no commands at all. - Genie `learn()` plus `Diff()` compares what the network is doing, including `oper_status`, which is the only one of the three that catches a change that applied but did not take effect. - An empty diff after a rollback is a result, not a non-event. It is the proof the rollback was complete. - Each tool has one setup detail that will stop you cold: `fast_cli` and TextFSM key names, `dest_file_system: "unix:"`, and the legacy SSH algorithm stanza. - Use all three. Push with NAPALM or Nornir, validate with pyATS, keep Netmiko for the one-off, and follow the rest of [the Cisco automation series](https://www.pinglabz.com/network-automation/) for the layers around them. ### pyATS and Genie: Diff Cisco Operational State, Not Config URL: https://www.pinglabz.com/pyats-genie-network-testing/ Last updated: 2026-08-01T17:24:30.000Z You pushed a change at 02:00\. The config looks right in `show run`. So why is half the branch black? Because config text is not state. The line went in and the interface never came up, or an adjacency reset and never came back, and nothing you diffed would have told you. Config diffing answers "did my text land". It does not answer "is this device still doing what it was doing an hour ago". That second question is what Cisco pyATS and Genie were built for, and it is the part of a [Python-based network automation toolchain](https://www.pinglabz.com/network-automation/) most people skip past on their way to writing yet another config pusher. Genie snapshots a device's live operational state, lets you make a change, snapshots again, and hands you a structured diff of what actually moved. Everything below was captured on real gear: a Debian 13 VM running pyATS 26.6, Genie 26.6 and Unicon 26.6 on Python 3.13.5, driving three `iol-xe` routers and an `ioll2-xe` switch on IOS-XE 17.18.2 in Cisco Modeling Labs. ## Installing pyATS in 2026: the old warning is dead Nearly every pyATS tutorial older than about a year opens with a warning about the install. Pin your Python version, expect C extensions to compile, do not even try the newest interpreter. Ignore all of it, because it is out of date. `pip install "pyats[full]"` now pulls cp313 wheels, and on Python 3.13.5 the whole stack (Genie, Unicon, the parser libraries, the traffic-generator bindings) installed with no compiler present and no errors. The only prerequisite was the venv module: ``` sudo apt-get install -y python3.13-venv python3 -m venv ~/venvs/plz ~/venvs/plz/bin/pip install "pyats[full]" ~/venvs/plz/bin/pyats version check ``` That reports pyats 26.6, genie 26.6 and unicon 26.6\. If you put pyATS off because you remembered the install being a fight, the fight is over. ## The SSH problem that stops everyone before line one Here is the failure you will actually hit, and it has nothing to do with pyATS. A modern OpenSSH client will not talk to a 2020-era Cisco image. It dropped SHA-1 host keys and the older key exchange groups years ago, and IOS is still offering exactly those. Your first `.connect()` dies with `unicon.core.errors.EOF` or `Host key verification failed`, and because the traceback comes out of Unicon, most people conclude pyATS is broken and go do something else. It is not broken. Your SSH client is refusing to negotiate. Fix it once in `~/.ssh/config`: ``` Host 192.168.99.* StrictHostKeyChecking no UserKnownHostsFile /dev/null KexAlgorithms +diffie-hellman-group14-sha1 HostKeyAlgorithms +ssh-rsa PubkeyAcceptedAlgorithms +ssh-rsa PubkeyAuthentication no ``` Three things worth understanding rather than pasting blindly. The `+` prefix appends to the client's default list instead of replacing it, so you are widening what you accept, not downgrading everything you connect to. Scoping the stanza to your lab subnet (not `Host *`) keeps that widening off production and internet SSH. And `PubkeyAuthentication no` matters because if your agent has keys loaded, the client burns its authentication attempts offering them and IOS drops the session before you reach the password prompt. The same stanza is what makes netmiko and NAPALM work from that box, so if you are building a [netmiko SSH automation script](https://www.pinglabz.com/netmiko-ssh-automation-cisco/) alongside this, you only do it once. ## The testbed file: your inventory, in YAML pyATS does not take a list of IPs. It takes a testbed: a YAML file describing devices, what OS they run, and how to reach them. The `os:` key is the important one, because it tells Unicon which plugin to load and therefore how to drive the CLI prompt state machine. You never write expect logic. ``` testbed: name: PingLabz-Research-VM-Automation-Edge credentials: default: username: admin password: Cisco123 devices: R1: os: iosxe type: router platform: iol connections: cli: protocol: ssh ip: 192.168.99.1 ``` Sanity-check it with `pyats validate testbed testbed.yaml` before you write any Python. Loading and connecting is then three lines, and `learn_hostname=True` makes Unicon read the real hostname off the prompt instead of trusting your YAML label, which is how you catch drifted inventory: ``` from genie.testbed import load tb = load("testbed.yaml") r1 = tb.devices["R1"] r1.connect(log_stdout=False, learn_hostname=True) ``` Each connected device then carries its resolved attributes: ``` DEV os platform type connected hostname learned R1 iosxe iol router True R1 SW1 iosxe iol switch True SW1 ``` ## parse(): show output becomes JSON you can index Two methods do the reading. `execute()` returns raw text exactly as the device printed it. `parse()` runs the same command and returns a dictionary. Compare them on `show ip interface brief`: ``` Interface IP-Address OK? Method Status Protocol Ethernet0/0 192.168.99.1 YES NVRAM up up Ethernet0/1 unassigned YES NVRAM administratively down down ``` ``` ! Command: dev.parse('show ip interface brief') # Genie parser { "interface": { "Ethernet0/0": { "interface_is_ok": "YES", "ip_address": "192.168.99.1", "method": "NVRAM", "protocol": "up", "status": "up" }, "Ethernet0/1": { "ip_address": "unassigned", "status": "administratively down" } } } ``` Look at the shape, because this is where Genie parts company with the TextFSM and ntc-templates world. TextFSM gives you a flat list of rows to loop over. Genie gives you a dictionary keyed by the thing itself, so you index straight in with `result['interface']['Ethernet0/0']['status']`, and every parser has a published schema, so the keys are a contract rather than something you discover by printing. If [turning CLI output into structured data you can query](https://www.pinglabz.com/ccna-auto-04-json-structured-data/) is new to you, that contrast is the whole point. A health check becomes a comprehension: filter that dict for `status == 'up'` and you get back `['Ethernet0/0', 'Loopback0', 'Loopback1', 'Loopback14']`. Protocol state works the same way. `show ip ospf neighbor` parses into neighbours keyed by router ID: ``` ! Command: dev.parse('show ip ospf neighbor') { "interfaces": { "Ethernet0/0": { "neighbors": { "2.2.2.2": { "address": "192.168.99.2", "dead_time": "00:00:31", "priority": 1, "state": "FULL/BDR" }, "3.3.3.3": { "address": "192.168.99.3", "state": "FULL/DR" } } } } } ``` Now "assert every neighbour is FULL" is a real test, not a regex against a screen scrape. `show version` parses the same way, handing you `version`, `uptime`, `chassis_sn` and `last_reload_reason` as fields. ## learn(): a whole feature, not one command `parse()` is one command in, one dict out. `learn()` runs several show commands, merges the results, and returns a model of an entire feature: a snapshot of what a protocol is doing. `r1.learn("ospf").info` comes back structured by VRF, address family, instance and area, and inside each area sits the full link-state database broken out by LSA type: ``` ! Command: dev.learn('ospf') -> .info { "vrf": { "default": { "address_family": { "ipv4": { "instance": { "1": { "areas": { "0.0.0.0": { "area_type": "normal", "database": { "lsa_types": { "1": { "lsa_type": 1, "lsas": { "1.1.1.1 1.1.1.1": { "adv_router": "1.1.1.1", "lsa_id": "1.1.1.1", ... ``` That is the LSDB as a dictionary, so you can assert about your topology directly against it. Only useful if you can read it, though, so if `lsa_types` keyed by number is not immediately obvious, get the [difference between a Type 1 and a Type 5 LSA](https://www.pinglabz.com/ospf-lsa-types-explained/) straight first. `learn('interface')` is flatter and more immediately useful: ``` ! Command: python -> {i: v.get('enabled') for i,v in .info.items()} { "Ethernet0/0": true, "Ethernet0/1": false, "Loopback0": true, "Loopback14": true } ``` The models are vendor-neutral by design: that same call returns the same structure on NX-OS or IOS-XR with the platform-specific parsing hidden underneath. That portability is the entire argument for models over per-command parsing in a mixed estate. ## The killer feature: learn, change, learn, Diff Here is the sequence that justifies the toolchain. Snapshot state, change something, snapshot again, diff the two. ``` from genie.utils.diff import Diff pre = r1.learn("interface") r1.configure(["interface Loopback14", " description CHANGED BY PYATS RES-0002", " shutdown"]) post = r1.learn("interface") ``` Before and after for that interface, straight out of the model: ``` ! Command: python -> pre.info['Loopback14'] { "description": "PingLabz RES-0014 pushed by netmiko", "enabled": true, "oper_status": "up" } ! Command: python -> post.info['Loopback14'] { "description": "CHANGED BY PYATS RES-0002", "enabled": false, "oper_status": "down" } ``` ### The counter-noise trap Now diff them naively and watch it fall over: ``` ! Command: d = Diff(pre.info, post.info); d.findDiff(); print(d) Ethernet0/0: counters: - in_octets: 179470 + in_octets: 191412 - out_errors: 0 + out_errors: 6 rate: - in_rate: 1000 + in_rate: 3000 Loopback14: - description: PingLabz RES-0014 pushed by netmiko + description: CHANGED BY PYATS RES-0002 - enabled: True + enabled: False ``` Your change is in there, buried under interface counters that moved because time passed. Octets, packet counts, spanning-tree bytes, computed rates: all of it climbs every second on any link carrying traffic, and with three OSPF speakers flooding hellos onto this segment there is no quiet moment. A pre/post test written against that raw diff fails one hundred percent of the time, for reasons unrelated to the change. This is the most common reason people try `Diff()` once and abandon it. The fix is the `exclude=` argument: name the volatile keys and they are dropped from both sides before comparison. Same snapshots, same change, and now the output is only the change: ``` ! Command: Diff(pre.info, post.info, exclude=['counters', 'accounting', 'rate', 'last_change', 'in_rate', 'out_rate']) Loopback14: - description: PingLabz RES-0014 pushed by netmiko + description: CHANGED BY PYATS RES-0002 - enabled: True + enabled: False - oper_status: up + oper_status: down ``` Note what those three lines are. `description` is the config you sent. `enabled` and `oper_status` are the device reporting the consequence: it went administratively down and the line protocol followed. Nothing you diff in a config file reports that second half. ### The empty diff is the proof Roll it back and diff `pre` against a fresh snapshot, same exclusions: ``` ! Command: dev.configure(['interface Loopback14',' no shutdown',' description PingLabz RES-0014 pushed by netmiko']) ! Command: Diff(pre.info, rolled_back.info, exclude=[...]) ``` Empty. That is a stronger claim than it looks. You have not merely verified that the rollback commands were accepted, which is all a config comparison tells you. You have verified that measured operational state is identical to what you recorded before you touched anything. "I typed `no shutdown` and it did not error" and "the interface is up and behaving as before" are different statements, and only one ends the change window. ## Configuration diffs versus operational diffs This distinction decides which tool you reach for, and it is why serious shops run more than one. netmiko DiffsText you scraped StructureYou build it Config landed?Yes Never came up?No NAPALM DiffsCandidate config StructureDevice-generated Config landed?Yes, before commit Never came up?No pyATS + Genie DiffsOperational state StructurePublished schemas Config landed?Yes Never came up?Yes, only this one NAPALM diffs configuration. pyATS Genie diffs operational state. Neither replaces the other, so use both: [let NAPALM show you the candidate config diff](https://www.pinglabz.com/nornir-napalm-cisco-config-management/) and gate the commit on it, then let Genie prove the network still works afterwards. For all three compared side by side on these same devices, see the [netmiko versus NAPALM versus pyATS breakdown](https://www.pinglabz.com/netmiko-vs-napalm-vs-pyats/). ## Parsing switch tables The same parsers work against Layer 2, where the ugliest column formats live. On the `ioll2-xe` switch, `show interfaces status` parses cleanly, and the parser normalises the CLI's `Et0/0` into `Ethernet0/0` so keys match across commands: ``` ! Command: sw1.parse('show interfaces status') { "interfaces": { "Ethernet0/0": { "duplex_code": "full", "port_speed": "auto", "status": "connected", "type": "10/100/1000BaseTX", "vlan": "1" } } } ``` `show errdisable recovery`, a two-column wall of reasons, becomes booleans plus the timer. `show vlan`, three stacked tables in one output, comes back keyed by VLAN ID with member ports as a list: ``` ! Command: sw1.parse('show vlan') { "vlans": { "1": { "interfaces": ["Ethernet0/0", "Ethernet0/1", "..."], "name": "default", "said": 100001, "state": "active", "vlan_id": "1" } } } ``` The important behaviour: where no parser exists for a command, Genie raises rather than returning something half-parsed. That is the correct failure mode. You want a loud exception, not a dictionary quietly missing the key your test asserts on. ## What this was captured on Automation hostDebian 13 VM, Python 3.13.5 in a venv Toolchainpyats 26.6, genie 26.6, unicon 26.6 TargetsR1/R2/R3 iol-xe, SW1 ioll2-xe, IOS-XE 17.18.2 TopologyOne flat segment, 192.168.99.0/24, OSPF 1 area 0 Data pathVM NIC bridged into CML via an external connector The switch between the CML external connector and the routers is what makes this zero-routing: the VM and all three routers share one broadcast domain, so the VM just ARPs for them. It also means the segment carries OSPF hellos continuously, which is why the counter noise above was so aggressive. ## Gotchas - **Do not probe with a question mark.** `r1.execute("show aaa ?")` raises `SubCommandFailure('... Incomplete command')` even though the help text prints on screen, because Unicon wants a settled prompt and context-sensitive help never gives it one. Test feature support by running the real command and catching `SchemaEmptyParserError`. - **The first connect to a freshly booted node can exceed 60 seconds and time out.** A retry against the same node succeeded immediately, so retry before assuming the device is broken. - **Genie's schema is nested, not row-based.** Coming from ntc-templates you will reflexively loop over rows. There are no rows. Index by key, and read the parser's schema page rather than guessing. - **Always exclude counters before you diff.** `counters`, `accounting`, `rate`, `in_rate`, `out_rate` and `last_change` covers the interface model. Other models have their own volatile fields (anything with a timer or a hit count), so build the list from your first noisy diff. - **Diff the whole model, not just the object you touched.** The point of a state diff is catching side effects you did not expect. Narrow it to `Loopback14` first and you have thrown away the reason you ran it. - **Watch the credentials block.** These routers run privilege 15 with no enable secret, so pyATS lands straight in exec mode. If yours need an enable password, it goes in the testbed credentials as a separate entry. ## Key takeaways - The install objection is retired. `pip install "pyats[full]"` ships cp313 wheels and installs clean on Python 3.13 with no compiler. - Your first blocker is SSH, not pyATS. Modern OpenSSH will not negotiate with 2020-era Cisco images until `KexAlgorithms +diffie-hellman-group14-sha1` and `HostKeyAlgorithms +ssh-rsa` are in `~/.ssh/config`. - `parse()` turns one command into a schema-backed dictionary. `learn()` merges several commands into a vendor-neutral model of a whole feature, which is what you snapshot. - `Diff()` without `exclude=` is useless on a live link, because packet counters move between the two learns whether or not you changed anything. - NAPALM diffs configuration, pyATS Genie diffs operational state. Only the state diff catches "the config applied but the interface never came up", which is why you run both. - An empty diff after rollback is the strongest proof of a clean back-out there is, because it is measured rather than assumed. Once you have a diff you trust, wrap it in AEtest so post-change validation runs as pass/fail in CI instead of a human squinting at output. The rest of the [tooling that gets you from manual show commands to tested changes](https://www.pinglabz.com/network-automation/) covers what sits around it, and [running Python directly on IOS-XE](https://www.pinglabz.com/on-box-python-ios-xe/) is the same idea with no automation host in the middle. ### Nornir and NAPALM: Cisco Config Management That Is Actually Idempotent URL: https://www.pinglabz.com/nornir-napalm-cisco-config-management/ Last updated: 2026-08-01T17:24:29.000Z Almost every Python automation story on a Cisco network starts the same way. You write a netmiko script, it works against one router, you wrap it in a `for` loop, and suddenly you own a hand-rolled inventory, a hand-rolled thread pool, and a config push that has no idea whether it changed anything. That last part is the one that actually hurts in production. Run the script twice and the output looks identical both times, and the only way to find out whether the second run did anything is to go and diff the running-config yourself. Nornir solves the inventory and the concurrency. NAPALM solves the "did anything change" problem, and it solves it by asking the router instead of guessing in Python. This article walks the entire loop against real gear: an inventory with group inheritance, a read across three devices at once, a dry-run diff, a commit that reports `changed=True`, and then the same run again reporting `changed=False` with an empty diff. If you are still working out [how the pieces of a network automation stack fit together](https://www.pinglabz.com/network-automation/), this is the layer where scripts stop being scripts and start being config management. Everything below is copied from a capture run on 2026-07-24: a Debian 13 VM running Python 3.13.5 in a venv, driving three `iol-xe` routers on IOS-XE 17.18.2 inside CML. Versions matter here, so they are all listed further down. ## What netmiko gives you, and what it does not netmiko is a session library. It opens an SSH channel, handles the pager and the prompt detection, sends lines, and gives you back a string. That is genuinely the hard part of talking to IOS, and [driving a single Cisco device over SSH from Python](https://www.pinglabz.com/netmiko-ssh-automation-cisco/) is still the right starting point for anyone learning this. What netmiko does not give you is any of the surrounding machinery. There is no concept of a device group, so credentials get pasted into every script. There is no concept of per-device variables, so the loopback address for R2 lives in a dictionary at the top of the file. And critically, there is no concept of state: `send_config_set()` returns whatever IOS echoed back, which is the same text whether the commands changed something or were no-ops. That is exactly the gap the declarative tools were built to close. If you have looked at [where declarative configuration management tools actually differ](https://www.pinglabz.com/ansible-vs-puppet-vs-chef/), the Nornir plus NAPALM combination sits in the same conceptual box, except the inventory is Python objects and the tasks are Python functions, so you never hit the wall where you need real logic and the DSL will not let you have it. ## The inventory is the point Nornir's `SimpleInventory` plugin reads two YAML files. `hosts.yaml` holds what is unique to each device, `groups.yaml` holds what is shared. Here is the host file used for the capture, verbatim: ``` R1: hostname: 192.168.99.1 groups: - ios_edge data: site: pinglabz-lab role: edge mgmt_loopback: 10.99.1.1 R2: hostname: 192.168.99.2 groups: - ios_edge data: site: pinglabz-lab role: edge mgmt_loopback: 10.99.2.1 R3: hostname: 192.168.99.3 groups: - ios_edge data: site: pinglabz-lab role: edge mgmt_loopback: 10.99.3.1 ``` Note that nothing in there is a credential and nothing in there is a platform. Those live once, on the group: ``` ios_edge: platform: ios username: admin password: Cisco123 data: snmp_location: PingLabz-Lab-Rack1 connection_options: napalm: extras: optional_args: dest_file_system: "unix:" netmiko: extras: fast_cli: false ``` `platform: ios` is worth pausing on. These routers run IOS-XE 17.18.2, but the NAPALM driver name is still `ios`. There is no `iosxe` driver, and setting the platform to something NAPALM does not recognise is one of the most common ways to get a connection failure that looks like an auth problem but is not. The `connection_options` block is where the whole thing becomes usable on lab gear, and it gets its own section below. First, the payoff of having an inventory at all. Loading it and printing the resolved attributes per host: ``` ! Command: python -> InitNornir(config_file='nornir/config.yaml') ! Command: python -> nr.inventory.hosts {'R1': Host: R1, 'R2': Host: R2, 'R3': Host: R3} ! Command: python -> per-host resolved attributes (group inheritance in action) HOST hostname platform username mgmt_loopback snmp_location R1 192.168.99.1 ios admin 10.99.1.1 PingLabz-Lab-Rack1 R2 192.168.99.2 ios admin 10.99.2.1 PingLabz-Lab-Rack1 R3 192.168.99.3 ios admin 10.99.3.1 PingLabz-Lab-Rack1 ! Command: python -> nr.filter(role='edge').inventory.hosts.keys() ['R1', 'R2', 'R3'] ``` `snmp_location` is written once, on the group, and resolves on all three hosts. `mgmt_loopback` is per host. That single table is the argument for keeping an inventory instead of a list of IP addresses, and `nr.filter()` is how a forty-device inventory becomes a three-device run without editing a file (change the filter, not the source of truth). The third file, `config.yaml`, wires it together and sets the runner: ``` inventory: plugin: SimpleInventory options: host_file: /home/j/plzcap/nornir/hosts.yaml group_file: /home/j/plzcap/nornir/groups.yaml runner: plugin: threaded options: num_workers: 5 logging: enabled: false ``` Five workers, threaded. You did not write a thread pool, and you will not write one. ## The one line that makes NAPALM work on iol-xe This is the thing that stops most people dead about five minutes into their first lab, so it gets its own heading. NAPALM's IOS driver does not type your config into a terminal. It writes the candidate configuration to a file on the device, transfers it over SCP, and then asks IOS to diff and merge it. The destination it writes to defaults to `flash:`, which is correct on almost every physical Catalyst or ISR you will ever touch. `iol-xe` nodes in CML have no `flash:` filesystem. `show file systems` on IOL lists `nvram:`, `unix:` (the 2 GB default), `disk0:`, `disk1:` and `system:`, and that is the complete list. Every `load_merge_candidate()` call dies before it ever produces a diff, and the traceback points at file operations rather than at anything you did wrong. The fix is four lines of inventory: ``` connection_options: napalm: extras: optional_args: dest_file_system: "unix:" ``` Two more things have to be true on the device itself, and neither is on by default: ``` ip scp server enable ! napalm ships the candidate config over SCP ! archive ! napalm's merge diff is computed BY THE ROUTER path unix:archive ! 'flash:archive' on a real box maximum 5 ``` Without `ip scp server enable`, the file transfer fails outright. Without `archive`, `show archive` answers `Archive feature not enabled` and there is no diff engine on the box for NAPALM to interrogate. Check all three before you blame the library: `show ip ssh`, `show archive`, and `show file systems` to find the writable disk. ## Reading three devices at once With the inventory loaded, a multi-device read is one line. `napalm_get` takes a list of getters and returns the same normalised structure for every platform: ``` from nornir import InitNornir from nornir_utils.plugins.functions import print_result from nornir_napalm.plugins.tasks import napalm_get, napalm_configure nr = InitNornir(config_file="nornir/config.yaml") r = nr.run(task=napalm_get, getters=["facts", "interfaces"]) print_result(r) ``` Trimmed to R1 (R2 and R3 returned the identical shape): ``` napalm_get********************************************************************** * R1 ** changed : False ******************************************************** vvvv napalm_get ** changed : False vvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvv INFO { 'facts': { 'fqdn': 'R1.pinglabz.lab', 'hostname': 'R1', 'interface_list': [ 'Ethernet0/0', 'Ethernet0/1', 'Ethernet0/2', 'Ethernet0/3', 'Loopback0', 'Loopback14'], 'model': 'Unknown', 'os_version': 'Linux Software (X86_64BI_LINUX-ADVENTERPRISEK9-M), ' 'Version 17.18.2, RELEASE SOFTWARE (fc3)', 'serial_number': '2040027', 'uptime': 480.0, 'vendor': 'Cisco'}, 'interfaces': { 'Ethernet0/0': { 'description': '', 'is_enabled': True, 'is_up': True, 'last_flapped': -1.0, 'mac_address': 'AA:BB:CC:00:DB:00', 'mtu': 1500, 'speed': 10.0}, 'Loopback14': { 'description': 'PingLabz RES-0014 pushed by ' 'netmiko', 'is_enabled': True, 'is_up': True, 'last_flapped': -1.0, 'mac_address': '', 'mtu': 1514, 'speed': 8000.0}}} ^^^^ END napalm_get ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ``` Two details in there are worth reading closely. `Loopback14` carries a description left behind by an earlier netmiko run against the same three routers, which is a small live demonstration of why you want a tool that can tell you what is on the box rather than what you think you put there. And `'model': 'Unknown'` is honest: IOL does not report a chassis model, so NAPALM returns the key with nothing useful in it rather than inventing a value. Same for `'mac_address': ''` and `'last_flapped': -1.0` on loopbacks. Normalised does not mean populated. Collapsing the same result set down to what you would actually put in a report: ``` HOST vendor model os_version fqdn uptime_s R1 Cisco Unknown Linux Soft R1.pinglabz.lab 480.0 R2 Cisco Unknown Linux Soft R2.pinglabz.lab 480.0 R3 Cisco Unknown Linux Soft R3.pinglabz.lab 480.0 R1 up=['Ethernet0/0', 'Loopback0', 'Loopback14'] R2 up=['Ethernet0/0', 'Loopback0', 'Loopback14'] R3 up=['Ethernet0/0', 'Loopback0', 'Loopback14'] ! failed hosts: {} (empty dict = every device answered) ``` `failed_hosts` being an empty dict is the check you want in every script you write. Nornir does not raise on a device failure, it records it and carries on with the rest, so a run that "worked" can quietly have skipped half your fleet if you never look. ## The diff comes from the router, not from Python This is the architectural difference between NAPALM and a session library, and it is worth being precise about it. The config is rendered per host from inventory data. A task function builds it, pulling `site` and `mgmt_loopback` from the host and `snmp_location` from the group: ``` def build(task): h = task.host return "\n".join([ "interface Loopback1", " description NORNIR-MANAGED %s" % h["site"], " ip address %s 255.255.255.0" % h["mgmt_loopback"], " ip ospf 1 area 0", "!", "snmp-server location %s" % h["snmp_location"], "snmp-server contact netops@pinglabz.com", ]) def cfg(task): return napalm_configure(task, configuration=build(task), replace=False, dry_run=True) print_result(nr.run(task=cfg, name="napalm_configure DRY RUN")) ``` Which renders, for R1: ``` ! ---- R1 ---- interface Loopback1 description NORNIR-MANAGED pinglabz-lab ip address 10.99.1.1 255.255.255.0 ip ospf 1 area 0 ! snmp-server location PingLabz-Lab-Rack1 snmp-server contact netops@pinglabz.com ``` And the dry run returns this: ``` napalm_configure DRY RUN******************************************************** * R1 ** changed : True ********************************************************* vvvv napalm_configure DRY RUN ** changed : True vvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvv INFO +interface Loopback1 + description NORNIR-MANAGED pinglabz-lab + ip address 10.99.1.1 255.255.255.0 + ip ospf 1 area 0 +snmp-server location PingLabz-Lab-Rack1 +snmp-server contact netops@pinglabz.com ^^^^ END napalm_configure DRY RUN ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ * R2 ** changed : True ********************************************************* vvvv napalm_configure DRY RUN ** changed : True vvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvv INFO +interface Loopback1 + description NORNIR-MANAGED pinglabz-lab + ip address 10.99.2.1 255.255.255.0 + ip ospf 1 area 0 +snmp-server location PingLabz-Lab-Rack1 +snmp-server contact netops@pinglabz.com ^^^^ END napalm_configure DRY RUN ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ``` Those `+` lines were not produced by Python comparing two strings. NAPALM copied the candidate to `unix:candidate_config` over SCP and then ran `show archive config incremental-diffs` on the router. The Contextual Configuration Diff Utility in IOS-XE, the same engine behind `configure replace` and config rollback, is what decided which lines are new. For a full replace instead of a merge the driver asks for `show archive config differences` instead. That distinction is why the output is trustworthy. A Python-side diff has to model IOS parser behaviour: submode context, command ordering, the fact that `ip ospf 1 area 0` under an interface is not the same statement as anywhere else, defaults that get suppressed from `show run`. The device already knows all of that, because it is the thing that parses the config. With `dry_run=True`, NAPALM discards the candidate instead of committing it, so you get the router's own answer to "what would this change" without changing anything. ## Commit, then run it again The commit is the same call with `dry_run=False`: ``` napalm_configure COMMIT********************************************************* * R1 ** changed : True ********************************************************* vvvv napalm_configure COMMIT ** changed : True vvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvv INFO +interface Loopback1 + description NORNIR-MANAGED pinglabz-lab + ip address 10.99.1.1 255.255.255.0 + ip ospf 1 area 0 +snmp-server location PingLabz-Lab-Rack1 +snmp-server contact netops@pinglabz.com ^^^^ END napalm_configure COMMIT ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ! Command: python -> result[host][0].changed R1 changed=True R2 changed=True R3 changed=True ``` Now the part this whole article exists for. Same script, same inventory, same rendered config, run a second time immediately afterwards: ``` napalm_configure RE-RUN (identical input)*************************************** * R1 ** changed : False ******************************************************** vvvv napalm_configure RE-RUN (identical input) ** changed : False vvvvvvvvvvvvvv INFO ^^^^ END napalm_configure RE-RUN (identical input) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ * R2 ** changed : False ******************************************************** vvvv napalm_configure RE-RUN (identical input) ** changed : False vvvvvvvvvvvvvv INFO ^^^^ END napalm_configure RE-RUN (identical input) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ * R3 ** changed : False ******************************************************** vvvv napalm_configure RE-RUN (identical input) ** changed : False vvvvvvvvvvvvvv INFO ^^^^ END napalm_configure RE-RUN (identical input) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ! Command: python -> (result[host][0].changed, len(result[host][0].diff)) R1 changed=False diff_len=0 diff='' R2 changed=False diff_len=0 diff='' R3 changed=False diff_len=0 diff='' ! First run: changed={'R1': True, 'R2': True, 'R3': True} ``` Read what is between the `vvvv` and `^^^^` markers on the second run: nothing. No diff, because the router had nothing to report. And because the diff came back empty, NAPALM skipped the commit entirely. That is the distinction people gloss over when they say a script is idempotent. Re-sending `snmp-server location PingLabz-Lab-Rack1` to a box that already has it is harmless, and a lot of automation is called idempotent on exactly that basis. It is not the same claim. Here, no configuration commands were sent at all, and the tool reported that fact in a boolean you can branch on. Second identical run with netmiko Commands sentAll of them, again Return valueSame text as run 1 Who computes the diffYou do, afterwards Preview before changeNot available Second identical run with napalm\_configure Commands sentNone Return valuechanged=False, diff='' Who computes the diffThe router, first Preview before changedry\_run=True ## Verify on the box, not in the return value A tool reporting success is not the same as the network being right. Pull the running config back and look: ``` ! ---- R1 ---- ! Command: napalm_get(getters=['config'], retrieve='running') -> Loopback1 stanza interface Loopback1 description NORNIR-MANAGED pinglabz-lab ip address 10.99.1.1 255.255.255.0 ip ospf 1 area 0 ! Command: python -> [l for l in running if l.startswith('snmp-server')] snmp-server location PingLabz-Lab-Rack1 snmp-server contact netops@pinglabz.com ! Command: napalm_get(getters=['interfaces']) -> ['Loopback1'] {'is_enabled': True, 'is_up': True, 'description': 'NORNIR-MANAGED pinglabz-lab', 'mac_address': '', 'last_flapped': -1.0, 'mtu': 1514, 'speed': 8000.0} ! ---- R3 ---- ! Command: napalm_get(getters=['config'], retrieve='running') -> Loopback1 stanza interface Loopback1 description NORNIR-MANAGED pinglabz-lab ip address 10.99.3.1 255.255.255.0 ip ospf 1 area 0 ! Command: python -> [l for l in running if l.startswith('snmp-server')] snmp-server location PingLabz-Lab-Rack1 snmp-server contact netops@pinglabz.com ``` R1 got `10.99.1.1`, R3 got `10.99.3.1`, each rendered from that host's own inventory data, and both got the same `snmp-server` lines because those came from the group. The interface is `is_up: True` with the description NAPALM was told to apply. The devices are running exactly what the inventory says they should be running, which is the entire claim configuration management makes. One trap in that snippet: it is two separate task runs, not one. `napalm_get` forwards every extra keyword argument to every getter you asked for, so combining a getter that takes options with one that does not blows up: ``` ! NOTE: napalm_get passes every extra kwarg to EVERY getter, so ! napalm_get(getters=['config','interfaces'], retrieve='running') ! fails with: TypeError: IOSDriver.get_interfaces() got an unexpected keyword ! argument 'retrieve'. Getters that take options must be run in their own task. ``` Verification like this is where the third leg of the stack picks up. Pulling config back and eyeballing it is fine for one loopback; when you want structured pass or fail assertions across a fleet, [turning device state into an actual test suite](https://www.pinglabz.com/pyats-genie-network-testing/) is the tool for that job. And if you want the short version of which of the three libraries to reach for and when, there is [a straight comparison of the three Python options](https://www.pinglabz.com/netmiko-vs-napalm-vs-pyats/). ## What this was captured on Runner hostDebian 13 (trixie) VM, 192.168.99.100 on ens224 Python3.13.5 in a venv Toolchainnornir 3.5.0, nornir\_napalm 0.5.0, nornir-utils 0.2.0, napalm 5.1.0, netmiko 4.7.0 TargetsR1/R2/R3, CML iol-xe, IOS-XE 17.18.2, 192.168.99.1-3 TopologyVM to CML external connector (bridge1) to an ioll2-xe switch to all three routers, one flat /24 Device prepRSA keys, ip ssh version 2, ip scp server enable, archive path unix:archive The flat layer 2 design is deliberate. The VM and all three routers sit in one broadcast domain, so there is no routing and no static routes between the automation host and the gear it manages, which removes an entire class of "is it my script or is it my network" ambiguity from the capture. ## Gotchas - **No `flash:` on `iol-xe`.** NAPALM's IOS driver defaults `dest_file_system` to `flash:`, and IOL does not have one. Set `dest_file_system: "unix:"` in `connection_options` or nothing that writes config will work. Run `show file systems` on any unfamiliar platform before you assume. - **`archive` is off by default.** `show archive` returning `Archive feature not enabled` means NAPALM has no diff engine to query, because the diff is `show archive config incremental-diffs` running on the router. - **`ip scp server enable` is not optional.** The candidate config goes over SCP. If you cannot enable the SCP server, NAPALM's `inline_transfer` optional argument is the alternative path. - **`napalm_get` sprays kwargs at every getter.** `getters=['config','interfaces']` with `retrieve='running'` raises `TypeError` on `get_interfaces()`. Split option-taking getters into their own task. - **Normalised does not mean populated.** IOL returns `'model': 'Unknown'`, empty MAC addresses on loopbacks and `last_flapped: -1.0`. NAPALM gives you a consistent key set across vendors, not a guarantee that every key has a value. - **The driver is `ios`, not `iosxe`.** IOS-XE 17.x uses the `ios` driver. A wrong platform string in `groups.yaml` surfaces as a connection error that reads like a credentials problem. - **Python 3.13 is a non-event, but Debian 13 is.** nornir, napalm and netmiko all installed cleanly from wheels on 3.13.5\. The only friction was that Debian 13 does not ship `python3.13-venv`, so `python3 -m venv` fails with an `ensurepip is not available` error until you install it. - **Check `failed_hosts`.** Nornir records failures rather than raising them. A run that printed happily can still have skipped devices, and `dict(result.failed_hosts)` being `{}` is the only thing that proves otherwise. ## Key takeaways - The inventory, not the SSH session, is what turns a script into config management. Group inheritance meant `snmp_location` was written once and resolved on all three hosts, while `mgmt_loopback` stayed per device. - NAPALM's diff is generated by the router (`show archive config incremental-diffs`), not by Python comparing strings. That is why `dry_run=True` is a preview you can actually trust before a change window. - A second identical run returned `changed=False` with `diff_len=0` on all three devices and sent no configuration commands at all. That is real idempotence, not "the commands were harmlessly re-accepted". - On CML `iol-xe` nodes, `dest_file_system: "unix:"` plus `archive` and `ip scp server enable` on the device are the three prerequisites. Miss any one and NAPALM fails before it reaches the interesting part. - Always verify on the box. The running config on R1, R2 and R3 matched the intended template with each device's own address, which is the only proof that matters. Nornir and NAPALM are one layer of a stack, not the whole thing. netmiko still underpins the session handling, on-box tooling has its place when you want logic running [on the switch itself rather than on a management server](https://www.pinglabz.com/on-box-python-ios-xe/), and a test framework covers the validation side. The [full automation cluster, in reading order](https://www.pinglabz.com/network-automation/), walks through where each of them belongs. ### Netmiko Tutorial for Cisco: SSH Automation That Proves It Worked URL: https://www.pinglabz.com/netmiko-ssh-automation-cisco/ Last updated: 2026-08-01T17:24:29.000Z You have a four-line change to make on forty routers, and the only management interface everyone agrees works is SSH. That is what Netmiko is for: a thin wrapper around Paramiko that knows how Cisco prompts behave, disables the pager for you, waits for the trailing `#` before handing anything back, and otherwise stays out of the way. This walkthrough is a **netmiko tutorial for Cisco** devices built entirely from a live capture: Netmiko 4.5.0 on a Debian 13 VM driving three `iol-xe` routers running IOS-XE 17.18.2 in a CML lab, all sitting on one bridged Layer 2 segment. Every block of device output below is copied from that run. If you are still deciding where scripted SSH sits next to the declarative tools, the [wider Python automation stack for Cisco gear](https://www.pinglabz.com/network-automation/) covers the neighbours. The part most tutorials skip is the part that matters in a change window: Netmiko has no idea whether your push changed anything. It will replay the same four lines a hundred times and print the same cheerful echo every time. Proving "nothing moved" is your job. ## What this was captured on The path is a real Linux host bridged into a real virtual network, which is why OSPF timers and route ages drift between blocks. Automation hostDebian 13 (trixie), Python 3.13.5, 192.168.99.100 on ens224 Librarynetmiko 4.5.0 with ntc-templates and textfsm DevicesR1/R2/R3, iol-xe, IOS-XE 17.18.2, 192.168.99.1-3 TopologyVM and all three routers on one broadcast domain, OSPF 1 area 0 Loginadmin / privilege 15 local user, no enable secret Device side, SSH needs this before any of it works, and one line cannot be staged in a startup config: ``` username admin privilege 15 secret Cisco123 ip domain name pinglabz.lab crypto key generate rsa modulus 2048 ip ssh version 2 line vty 0 4 login local transport input ssh ``` `crypto key generate rsa` has to run after the box boots, which is why automated lab builds all end up with a post-boot step. (On IOL, IOS-XE calls the command deprecated and then generates the key anyway.) ## Connect: ConnectHandler, and the missing enable secret The whole connection is a dictionary. `device_type` is the only field with any magic in it, because it selects the platform class that knows what a Cisco prompt looks like and which pager command to send. ``` from netmiko import ConnectHandler dev = { "device_type": "cisco_ios", "host": "192.168.99.1", "username": "admin", "password": "Cisco123", "fast_cli": False, } conn = ConnectHandler(**dev) print(conn.find_prompt()) ``` ``` R1# ``` That prompt is worth staring at. It is `R1#`, not `R1>`. Because the local user is `privilege 15`, Netmiko lands straight in exec mode and never answers an enable prompt, so there is no `secret=` key in the dict. Most tutorials show one, which is fine until someone copies the dict onto a box where the user is privilege 1 and cannot work out why config mode fails. (Lower privilege, add `"secret"` and call `conn.enable()` first.) `fast_cli=False` is deliberate. Netmiko 4.x defaults it to `True`, which trims the internal delay factors for speed; over a bridged path to virtual IOL nodes that made reads flakier, so it was turned off for stable captures. On real hardware, leave the default alone until it bites you. If you have not run Netmiko before, start with the CCNA lab: [your first netmiko show command against a router](https://www.pinglabz.com/ccna-auto-02-netmiko-show-commands/), then [pushing your first config from a script](https://www.pinglabz.com/ccna-auto-03-netmiko-config-push/). Everything below assumes the connection already works. ## Read: send\_command returns the whole thing, unpaged `send_command()` writes the string, then reads until it sees the prompt again. You get complete output, with no `--More--` in it, because Netmiko issued `terminal length 0` at login on your behalf. ``` conn.send_command("show version | include Software|uptime is") conn.send_command("show ip ospf neighbor") ``` ``` Cisco IOS Software [IOSXE], Linux Software (X86_64BI_LINUX-ADVENTERPRISEK9-M), Version 17.18.2, RELEASE SOFTWARE (fc3) R1 uptime is 4 minutes Neighbor ID Pri State Dead Time Address Interface 2.2.2.2 1 FULL/BDR 00:00:30 192.168.99.2 Ethernet0/0 3.3.3.3 1 FULL/DR 00:00:39 192.168.99.3 Ethernet0/0 ``` What you have there is a Python string containing a screen. That is screen-scraping, not automation. Everyone's first script reaches for a regex against that string, and everyone's first script breaks the first time a column shifts. ## Parse: use\_textfsm, and the key names nobody guesses right Add one keyword argument and Netmiko runs the output through TextFSM using the bundled ntc-templates, handing you a list of dictionaries instead of a screen. ``` rows = conn.send_command("show ip interface brief", use_textfsm=True) ``` ``` [{'intf': 'Ethernet0/0', 'ipaddr': '192.168.99.1', 'status': 'up', 'proto': 'up'}, {'intf': 'Ethernet0/1', 'ipaddr': 'unassigned', 'status': 'administratively down', 'proto': 'down'}, {'intf': 'Loopback14', 'ipaddr': '14.14.14.1', 'status': 'up', 'proto': 'up'}] ``` Now read those keys again. They are `intf`, `ipaddr`, `status` and `proto`. They are **not** `interface` and `ip_address`, which is what any reasonable person would guess, and what the first version of this capture script did guess. Here is why that matters more than it sounds. TextFSM does not raise. A dictionary lookup with `.get()`, an f-string over a missing key, a template that matches zero rows: all of them produce `None` or an empty list and keep going. The script ran to completion, printed a tidy report full of `None`, and exited zero. Nothing in the traceback told anyone, because there was no traceback. The fix costs one line, once, per command you have not parsed before: ``` sorted(rows[0].keys()) ``` ``` ['intf', 'ipaddr', 'proto', 'status'] ``` Do that before you write the loop, not after the report looks wrong. The other reliable source is the template itself: the `Value` lines at the top of `cisco_ios_show_ip_interface_brief.textfsm` in the ntc-templates repo are the definitive key list. Once the keys are right, a health check stops being a regex problem and becomes a list comprehension: ``` nbrs = conn.send_command("show ip ospf neighbor", use_textfsm=True) [n for n in nbrs if not n["state"].startswith("FULL")] ``` ``` [] ``` An empty list is the entire check. Every adjacency is FULL, with no column counting and no `include` filter that silently stops matching after an upgrade. `show version` parses the same way, into `version`, `hostname`, `uptime` and `serial`, which is the shape you want for an inventory. If that data is heading somewhere else, [handling JSON from network devices in Python](https://www.pinglabz.com/python-json-encor-exam/) picks up there, and [running Python on the switch itself](https://www.pinglabz.com/on-box-python-ios-xe/) is the other end of the same idea. ## Write: send\_config\_set builds the config session for you Pass a list of lines. Netmiko wraps them in `configure terminal` and `end`, sends them one at a time, and returns the whole session echo as a string. ``` cfg = [ "interface Loopback14", " description PingLabz RES-0014 pushed by netmiko", " ip address 14.14.14.1 255.255.255.255", ] print(conn.send_config_set(cfg)) ``` ``` configure terminal Enter configuration commands, one per line. End with CNTL/Z. R1(config)#interface Loopback14 R1(config-if)# description PingLabz RES-0014 pushed by netmiko R1(config-if)# ip address 14.14.14.1 255.255.255.255 R1(config-if)#end R1# ``` The absence of a line starting with `%` is how you know every command was accepted, and scanning the returned string for `"%"` is a real check worth automating. It is also about the only signal that return value gives you. Netmiko will not save for you, so add the write yourself: ``` conn.send_command("write memory") ``` ``` Building configuration... [OK] ``` ## Prove it: Netmiko cannot tell you whether anything changed This is what separates a demo from a change process. Push the identical list a second time, nothing altered in between. ``` configure terminal Enter configuration commands, one per line. End with CNTL/Z. R1(config)#interface Loopback14 R1(config-if)# description PingLabz RES-0014 pushed by netmiko R1(config-if)# ip address 14.14.14.1 255.255.255.255 R1(config-if)#end R1# ``` Byte for byte the same output as the first run. IOS re-accepted every line without a murmur, because the CLI is imperative: you are not describing a desired end state, you are typing commands, and typing a command that is already true is not an error. The string Netmiko handed back says nothing about whether the device moved. The only way to know is to diff the device's own state around the operation: ``` before = conn.send_command("show running-config") conn.send_config_set(cfg) after = conn.send_command("show running-config") import difflib list(difflib.unified_diff(before.splitlines(), after.splitlines())) ``` ``` ``` Empty list, zero diff lines. That is the proof, and producing it was your job, not the library's. For contrast, the same diff around the *first* push, when something genuinely did change: ``` --- running-config BEFORE 1st run +++ running-config AFTER 1st run @@ -20,2 +20,5 @@ ip ospf 1 area 0 +interface Loopback14 + description PingLabz RES-0014 pushed by netmiko + ip address 14.14.14.1 255.255.255.255 interface Ethernet0/0 ``` Hold that next to a declarative tool. NAPALM builds a candidate config, asks the device for the diff, and returns `changed=False` with an empty diff when there is nothing to do, before it commits. That guarantee is the whole reason to move up a layer: [netmiko gives you the session, Nornir gives you the inventory](https://www.pinglabz.com/nornir-napalm-cisco-config-management/) and NAPALM gives you the diff. ## Scale: one loop, three devices, one independent witness This is why anyone tolerates screen-scraping at all. One inventory dictionary, one credential set, one loop: ``` INV = { "R1": {"host": "192.168.99.1", "ip": "14.14.14.1"}, "R2": {"host": "192.168.99.2", "ip": "14.14.14.2"}, "R3": {"host": "192.168.99.3", "ip": "14.14.14.3"}, } for name, d in INV.items(): c = ConnectHandler(host=d["host"], **BASE) c.send_config_set([ "interface Loopback14", " ip address %s 255.255.255.255" % d["ip"], " ip ospf 1 area 0", ]) c.send_command("write memory") c.disconnect() ``` Each device confirms its own line: ``` ! R2 Loopback14 14.14.14.2 YES manual up up ! R3 Loopback14 14.14.14.3 YES manual up up ``` Except that proof is weak. You wrote to R2, then read from R2, over the same session with the same library. If anything in that path were lying to you (a stale buffer, a prompt-matching bug, a mocked connection someone left in), you would not find out this way. So put a witness in the loop. Each pushed loopback carries `ip ospf 1 area 0`, so R2 and R3 have to advertise it and R1, never told about any of it, has to learn it independently: ``` conn.send_command("show ip route ospf | include 14.14.14") ``` ``` O 14.14.14.2 [110/11] via 192.168.99.2, 00:00:06, Ethernet0/0 O 14.14.14.3 [110/11] via 192.168.99.3, 00:00:03, Ethernet0/0 ``` R1 learned both prefixes through OSPF. A routing protocol on a device you never touched now agrees that your change landed on two real, running boxes. That is a categorically stronger claim than re-reading the device you just wrote to, and it costs one extra `show`. Verify from the other side of the change wherever a protocol will do the verifying for you, which is the habit [testing a network with pyATS and Genie](https://www.pinglabz.com/pyats-genie-network-testing/) formalises. ## The session log: what actually goes over the wire Add `"session_log": "session.txt"` to the device dict and Netmiko writes every byte it sent and received to a file. It is the first thing to reach for when a script hangs, because it shows you exactly which prompt Netmiko was waiting for when it gave up. ``` R1# R1#terminal width 511 R1#terminal length 0 R1# R1#show version | include Software|uptime is Cisco IOS Software [IOSXE], Linux Software (X86_64BI_LINUX-ADVENTERPRISEK9-M), Version 17.18.2, RELEASE SOFTWARE (fc3) R1 uptime is 4 minutes R1# R1#exit ``` Notice the two lines you never wrote: `terminal width 511` and `terminal length 0`. Netmiko sends those itself at login, which is why your output never contains `--More--` and never wraps mid-column. It is also a reminder that this is a normal interactive session, landing in your AAA accounting and syslog exactly like a human typing. Automation over SSH is not a side channel. ## Gotchas that cost real time - **TextFSM fails silently.** Wrong key name, no exception, `None` everywhere, exit code zero. Print `result[0].keys()` once per new command, or read the template's `Value` lines. - **There is no desired state.** The return value of `send_config_set()` is a transcript, not a change report. Identical pushes look identical. Diff `show running-config`, or move to a tool that does it for you. - **send\_config\_set does not save.** Your change survives exactly until the next reload unless you follow it with `write memory`. - **Privilege level decides whether you need a secret.** With `privilege 15`, `find_prompt()` gives you `R1#` and no `secret=` is needed. Anything lower and you need `enable()` before config mode. - **fast\_cli is on by default in Netmiko 4.x.** Fast is good until it is not. On slow or virtual gear, `fast_cli=False` buys stability at the cost of a couple of seconds per device. - **Older IOS images will refuse your SSH client outright.** Modern Paramiko and OpenSSH dropped the key exchange and host key algorithms ancient boxes still offer, so you get a negotiation failure that looks nothing like an auth failure. Re-enabling those algorithms is a deliberate [weakening of management plane security](https://www.pinglabz.com/infrastructure-security/), so scope it per host and put the image upgrade on the plan. - **Watch which Netmiko you are actually running.** The Debian system package here was 4.5.0; a fresh virtualenv on the same host pulled 4.7.0\. Both worked, but when behaviour differs between your laptop and the jump host, that is usually why. ## Key takeaways - Netmiko gives you a reliable interactive SSH session with the pager disabled and prompt handling solved. That is the whole product, and it is a good one. - `use_textfsm=True` is the line between screen-scraping and automation, but verify the key names once before you trust them, because a wrong key is silent. - `send_config_set()` tells you commands were accepted (no `%` lines) and nothing more. It cannot tell you whether the device changed. - Diffing `show running-config` around the push is the only honest proof of "no change" with Netmiko. If you need that routinely, [choosing between netmiko, NAPALM and pyATS](https://www.pinglabz.com/netmiko-vs-napalm-vs-pyats/) is the next decision. - Verify from somewhere you did not write to. OSPF proved these pushes landed on R2 and R3 far better than re-reading R2 and R3 did. - Start scripted SSH here, then work outward through the rest of the [Cisco automation tooling walkthroughs](https://www.pinglabz.com/network-automation/) when your inventory outgrows a for loop. ### IPsec over GRE vs GRE over IPsec URL: https://www.pinglabz.com/ipsec-over-gre-tunnel/ Last updated: 2026-07-20T13:00:00.000Z "IPsec over GRE" and "GRE over IPsec" sound like the same thing said two ways. They are not. They are two different designs with the encapsulation layers stacked in the opposite order, and the order changes what gets encrypted, what is visible on the wire, and which one you should actually deploy. This post explains IPsec over GRE specifically, contrasts it cleanly with the more common reverse, and shows when each belongs. For the cluster overview, see the [GRE complete guide](https://www.pinglabz.com/gre/). ## Why two protocols at all GRE and IPsec each do half a job, and a secure site-to-site tunnel needs both halves. - **GRE** can encapsulate anything - multicast, broadcast, non-IP protocols - which means it can carry a routing protocol like OSPF or EIGRP across the tunnel. But GRE provides *no encryption*. A GRE packet is readable by anyone on the path. - **IPsec** encrypts and authenticates, but standard IPsec only protects *unicast IP*. It will not, on its own, carry the multicast that routing protocols depend on. Pair them and you get both: GRE carries the routing protocol, IPsec provides the encryption. The only question is which one wraps the other - and that is the whole "IPsec over GRE" versus "GRE over IPsec" question. ## What "over" means here The protocol that comes *after* "over" is the outer, transporting layer. The protocol before "over" rides inside it. - **GRE over IPsec** \- GRE is the passenger, IPsec is the carrier. GRE encapsulates the original packet first, then IPsec encrypts the resulting GRE packet. IPsec is the outermost header. - **IPsec over GRE** \- IPsec is the passenger, GRE is the carrier. Traffic is protected by IPsec first, then the GRE tunnel transports the IPsec-protected traffic. GRE is the outermost header. ## IPsec over GRE: how it is built In IPsec over GRE, you build the GRE tunnel first as a normal tunnel interface between the two sites. Then you apply IPsec to traffic *traversing* that GRE tunnel - the crypto policy sits inside the GRE relationship rather than wrapping the whole thing. The encapsulation order for a packet is: original packet, then IPsec protection applied to selected traffic, then GRE encapsulation, then the physical transport header. The GRE header ends up on the outside, exposed. That exposed GRE header is the defining characteristic. Anyone watching the path sees a GRE tunnel between the two endpoint addresses. They cannot read the IPsec-protected payload, but the tunnel itself, its endpoints, and any traffic in the GRE tunnel that was *not* selected for IPsec are all visible. ## GRE over IPsec: the reverse, and the default choice In GRE over IPsec, the GRE tunnel is encapsulated *inside* IPsec. IPsec is configured to match and protect the GRE traffic (IP protocol 47) between the two endpoints. Because IPsec wraps the entire GRE packet, **everything** inside the GRE tunnel is encrypted - the user data and the routing-protocol traffic alike - and the GRE header itself is hidden under the IPsec header. This is the design almost every modern deployment uses, and DMVPN is built on this model. The companion article on [GRE over IPsec](https://www.pinglabz.com/gre-over-ipsec/) covers it in full. ## Side by side Outer header on the wire IPsec over GREGRE GRE over IPsecIPsec What gets encrypted IPsec over GRE Only traffic selected by the IPsec policy, inside the tunnel GRE over IPsec The entire GRE tunnel - all traffic in it Routing protocol traffic encrypted? IPsec over GRE Only if explicitly selected GRE over IPsec Yes, automatically - it is inside the GRE that IPsec wraps GRE tunnel visible on the path? IPsec over GRE Yes - GRE header is exposed GRE over IPsec No - hidden under IPsec Typical use IPsec over GRE Niche - encrypt only a subset of tunnel traffic GRE over IPsec The standard secure site-to-site design ## When would you use IPsec over GRE? Honestly, rarely. It exists for the narrow case where a single GRE tunnel must carry a mix of traffic and you want to encrypt only part of it - leaving the rest as plain GRE for cost, performance, or policy reasons. It can also appear when an existing GRE tunnel topology is already in production and IPsec is being layered onto specific flows without rebuilding the tunnel design. For a brand-new secure site-to-site link, IPsec over GRE is the wrong default. It leaves the GRE header exposed, and it makes it easy to forget to protect the routing-protocol traffic - which means your OSPF or EIGRP adjacency could be crossing the internet in clear text while you believe the tunnel is "encrypted." GRE over IPsec avoids both problems by encrypting the whole GRE tunnel as one unit. ## Configuration shape GRE over IPsec - the recommended design - protects the GRE flow itself. The crypto policy matches GRE between the two physical endpoints: ``` ! GRE over IPsec - IPsec protects the GRE traffic (protocol 47) crypto ipsec profile PINGLABZ_PROFILE set transform-set PINGLABZ_TS ! interface Tunnel0 ip address 10.30.30.1 255.255.255.252 tunnel source GigabitEthernet0/0 tunnel destination 192.0.2.2 tunnel protection ipsec profile PINGLABZ_PROFILE ``` The single line `tunnel protection ipsec profile` is what makes everything entering Tunnel0 - including the routing protocol - encrypted. That one command is why GRE over IPsec is both simpler and safer, and why it is the design to reach for unless you have a specific reason not to. ## Common gotchas Routing protocol adjacency crossing the link in clear text An IPsec-over-GRE design where the routing traffic was never selected by the IPsec policy. Use GRE over IPsec. Tunnel up but no encryption happening The IPsec policy is not matching the traffic, or `tunnel protection` is missing on the tunnel interface. Fragmentation and MTU drops Both layers add overhead. Set `ip mtu` and `ip tcp adjust-mss` on the tunnel interface. Confusion over which design is deployed Check the outer header in a capture: IPsec (ESP) outermost is GRE over IPsec; GRE outermost is IPsec over GRE. Adjacency forms but encrypted traffic fails An ACL on the path is blocking ESP (IP protocol 50) or UDP 500/4500 for IKE. ## Key takeaways IPsec over GRE and GRE over IPsec stack the same two protocols in opposite order, and the order matters. The protocol after "over" is the outer carrier. IPsec over GRE puts GRE on the outside, exposes the GRE header, and encrypts only the traffic explicitly selected inside the tunnel - which makes it easy to accidentally leave routing-protocol traffic unencrypted. GRE over IPsec puts IPsec on the outside, hides the GRE header, and encrypts the entire GRE tunnel including the routing protocol, in one line of config. IPsec over GRE has narrow niche uses; for any new secure site-to-site tunnel, GRE over IPsec is the correct default. For the GRE cluster, see the [GRE pillar](https://www.pinglabz.com/gre/). ### Catalyst 9800-CL in CML: What You Can and Can't Actually Lab URL: https://www.pinglabz.com/catalyst-9800-cl-cml-limitations/ Last updated: 2026-07-20T05:01:16.000Z If you have tried to build a wireless lab in Cisco Modeling Labs, you have probably hit the same wall everyone hits: you drop a Catalyst 9800-CL on the canvas, configure it perfectly, and `show ap summary` stubbornly reports zero access points forever. Nothing is wrong with your configuration. The platform simply cannot do what you are asking. This article documents exactly what a 9800-CL can and cannot do inside CML, with real output captured from a controller running IOS XE 17.18.2\. Every command result below came off an actual lab controller, including the failures. If you are planning study time around a virtual wireless lab, read this first so you spend it on the things that actually work. For the broader configuration and design material, start at the [Cisco wireless cluster guide](https://www.pinglabz.com/wireless/). ## The Short Version Console is VNC-only at first Even with `platform console serial`, the first boot only talks on VNC. Serial works after day-0 is saved and the node restarts. No AP can ever join CML ships no access point image. The 2.10 `wireless-ap` node is a Linux hostapd box and does not speak CAPWAP. Default sizing blocks local mode 2 vCPU and 6 GB puts the node in Ultra-Low profile, which cannot tunnel AP traffic. FlexConnect works in every profile. HA parses but does nothing SSO and N+1 commands are accepted in full. Redundancy never becomes operational. ## The Console Will Not Talk to You on First Boot This is the first thing that stops people, and it looks like a broken image rather than expected behaviour. You start the node, watch the serial console, and the output dies partway through boot: ``` Both links down, not waiting for other chassis Chassis number is 1 ``` After that, silence. Nothing else ever appears on the serial console, and any automation you point at it (PyATS, Netmiko, an MCP connector) fails with a bare timeout rather than an authentication error, which sends you hunting for credentials problems that do not exist. Cisco documents the cause: the 9800-CL uses the VNC console initially, and *"even with the command `platform console serial` in the bootstrap configuration (which is included in the default node definition), the first boot will only output the full data and accept CLI input on the VNC console. Once the 9800-CL has booted and saved its day-0 configuration, you can restart the node. On subsequent boots, the node will use the serial console."* The practical fix, if you want a scriptable controller, is to stop the node, wipe it, inject a day-0 configuration that includes both the console directive and your credentials, and start it again: ``` hostname WLC1 ! platform console serial ! username admin privilege 15 secret Cisco@123 ! line con 0 exec-timeout 0 0 privilege level 15 logging synchronous line vty 0 4 login local privilege level 15 transport input ssh ``` After the wipe the node factory-resets itself once before applying the config: ``` Chassis 1 reloading, reason - Reload command %PMAN-5-EXITACTION: R0/0: pvp: Process manager is exiting: process exit with reload chassis code %IOSXEBOOT-4-FACTORY_RESET: (rp/0): This was not selected via cli. Rebooting/Shutting Down like normal ``` Budget roughly twelve minutes from wipe to a usable prompt. From the second boot onward the serial console works normally and automation attaches cleanly: ``` WLC1#show version | include Cisco IOS Software|uptime is Cisco IOS Software [IOSXE], C9800-CL Software (C9800-CL-K9_IOSXE), Version 17.18.2, RELEASE SOFTWARE (fc3) WLC1 uptime is 13 minutes ``` ## No Access Point Will Ever Join This is the limitation that matters most, and the one people burn the most hours on. Cisco states it directly in the 9800-CL platform notes: *"While CML does not offer images for access points (APs) or wireless clients, the 9800-CL can be used to manage physical access points via bridge external connectivity. Even without physical APs, you can use the 9800-CL to learn about wireless LAN configuration, creating RF profiles, etc."* There is no virtual lightweight AP in CML. There never has been. ### The CML 2.10 wireless nodes do not change this CML 2.10 added beta Wi-Fi simulation with two new node types, `wireless-ap` and `wireless-client`, and it is easy to assume these are the missing piece. They are not. Their node definition is an Ubuntu 24.04 cloud-init virtual machine, and the AP role is provided by `hostapd`: ``` write_files: - path: /home/cisco/hostapd.conf content: | interface=wlan0 driver=nl80211 ssid=openap channel=6 auth_algs=1 wpa=0 ``` That is a standalone Linux software access point. It speaks 802.11 to a `wireless-client` node running `wpa_supplicant`, which is genuinely useful for demonstrating association, beacons, and the WPA handshake. It has no CAPWAP implementation and no concept of a controller, so it cannot register to a Catalyst WLC. The result is permanent, no matter how correct your DHCP option 43 is: ``` WLC1#show ap summary Number of APs: 0 WLC1#show wireless stats ap join summary Number of APs: 0 Base MAC Ethernet MAC AP Name IP Address Status Last Failure Phase Last Disconnect Reason ----------------------------------------------------------------------------------------------------------------- ``` Note what is *not* in that second output: there is no failure phase and no disconnect reason, because no join was ever attempted. If you were troubleshooting a real AP that could not join, you would see entries here with a failure phase. An empty table is the signature of an AP that never spoke CAPWAP at all. The DHCP side works perfectly, which adds to the confusion. The AP node takes a lease from the lab DHCP server without complaint: ``` EDGE-RTR1#show ip dhcp binding IP address Client-ID/ Lease expiration Type State Interface 10.10.10.22 ff32.39f9.b500.0200. Jul 20 2026 04:43 AM Automatic Selecting Ethernet0/0.10 ``` If you need genuine AP join and client output, your options are a physical access point bridged into the lab through an external connector, or real hardware. There is no software path. ## The Default Node Size Blocks Local Mode Even once you attach a physical AP, the default node sizing constrains what deployment model you can run. Cisco puts it this way: *"The 9800-CL node definition uses 2 vCPUs and 6 GB of RAM by default, which puts it in Ultra-Low Profile. This profile does not support local mode AP deployments (i.e., you cannot have AP traffic tunneled over CAPWAP to the controller). If you increase the resource allocation of your CAT 9800-CL node to 4 vCPUs to 8 GB of memory, it will then support local mode deployments. In all modes, you can do FlexConnect deployments."* You can confirm which profile you are in from the controller itself. A maximum DRAM figure just under 6 GB means Ultra-Low: ``` WLC1#show platform resources **State Acronym: H - Healthy, W - Warning, C - Critical Resource Usage Max Warning Critical State ---------------------------------------------------------------------------------------------------- RP0 (ok, active) H Control Processor 54.11% 100% 80% 90% H DRAM 2133MB(36%) 5879MB 88% 93% H ``` The practical consequence: if you are labbing [FlexConnect versus local mode](https://www.pinglabz.com/c9800-flexconnect-vs-local-mode/), FlexConnect is available to you on any node size, and local mode requires bumping the node to 4 vCPU and 8 GB first. Plan your host memory accordingly, because a 9800-CL at 8 GB alongside a few IOL nodes adds up quickly. ## High Availability Configures Cleanly and Does Nothing This one is genuinely dangerous for study purposes, because the CLI gives you no indication that anything is wrong. Every SSO and N+1 command is accepted. The configuration appears in the running config. Verification commands return plausible-looking output. But redundancy never becomes operational. Cisco's note is unambiguous: *"high availability (HA) features of the physical Catalyst 9800 controllers are not supported by the 9800-CL node. The configuration options for HA are present, but they do not work."* Here is the proof from a live controller. Read the two mode lines together: ``` WLC1#show redundancy states my state = 13 -ACTIVE peer state = 1 -DISABLED Mode = Simplex Unit = Primary Unit ID = 1 Redundancy Mode (Operational) = Non-redundant Redundancy Mode (Configured) = sso Redundancy State = Non Redundant Manual Swact = disabled (system is simplex (no peer unit)) Communications = Down Reason: Simplex mode client count = 130 client_notification_TMR = 30000 milliseconds Gateway Monitoring = Enabled Gateway monitoring interval = 8 secs ``` `Configured` says `sso`. `Operational` says `Non-redundant`. That gap is the entire story. The chassis view confirms a single unit with no peer address: ``` WLC1#show chassis Chassis/Stack Mac Address : 0001.0202.aabb - Local Mac Address H/W Current Chassis# Role Mac Address Priority Version State IP ------------------------------------------------------------------------------------- *1 Active 0001.0202.aabb 1 V02 Ready 0.0.0.0 ``` If you are preparing for an exam or a production cutover, practise the [SSO configuration syntax](https://www.pinglabz.com/c9800-sso-configuration/) here by all means, but do not expect to observe a switchover. The failover behaviour, RMI health signalling, and client session preservation all need real hardware. ## What You Can Actually Lab None of this makes a virtual 9800 useless. The entire control and configuration plane is real, and that is where most of the learning curve lives. In a single session on this controller we exercised all of the following with genuine output: - The full tag and profile model: WLAN profiles, policy profiles, policy tags, site tags, and RF tags - WLAN security across WPA2-PSK, WPA3-SAE, and WPA3 Enhanced Open (OWE) - Wireless management interface assignment and country and regulatory domain configuration - Mobility group settings, DTLS cipher lists, and mobility domain identifiers - Radio network defaults for 2.4 GHz, 5 GHz, and 6 GHz - RF profile creation and the 802.11 network enable and disable workflow - AAA and RADIUS server configuration syntax - The web UI, which is fully functional at `https://` What you cannot lab is anything downstream of an AP joining: AP join troubleshooting, client onboarding, roaming, RRM in action, rogue detection, and RF troubleshooting all need radios. ## Three Gotchas Worth Knowing Before You Start ### Country code configuration is order-dependent Setting the country while the radio networks are up is rejected outright: ``` WLC1(config)#ap country US % 802.11bg/802.11a Network must be disabled ``` You have to shut both bands, set the country, and re-enable them. Each step prompts for confirmation: ``` ap dot11 24ghz shutdown ap dot11 5ghz shutdown ap country US no ap dot11 24ghz shutdown no ap dot11 5ghz shutdown ``` ### Adaptive FT silently blocks WPA3 and OWE This one is not in most guides. A WPA3-SAE or Enhanced Open WLAN will refuse to come up while adaptive fast transition is enabled, and the error only appears when you try to enable the WLAN: ``` WLC1(config-wlan)#no shutdown % node-(1):dbm:wireless:FT adaptive is not allowed with SAE/SAE-EXT-KEY AKM in WPA3 only WLAN ``` ``` WLC1(config-wlan)#no shutdown % node-(1):dbm:wireless:FT must be disabled in OWE WLAN ``` The fix is a single line, `no security ft adaptive`, but you will not find it by reading the WLAN configuration back. ### Removing WPA2 turns off PMF behind your back When you strip WPA2 from a WLAN on the way to a WPA3-only configuration, protected management frames are disabled as a side effect, and the controller tells you only once: ``` WLC1(config-wlan)#no security wpa wpa2 PMF is now disabled. ``` Since WPA3 requires PMF, you must re-apply `security pmf mandatory` afterward or the WLAN will not come up. Related reading: [Catalyst 9800 wireless security configuration](https://www.pinglabz.com/c9800-wireless-security/). ## Bonus: the 6 GHz Band Tells You Exactly What It Wants One genuinely pleasant surprise is how explicit the controller is about 6 GHz requirements. On a WPA2-PSK WLAN, the 6 GHz band reports itself down along with the reason: ``` Operational State of Radio Bands 2.4GHz : UP 5GHz : UP 6GHz : DOWN (Required config: Enable WPA3 & dot11ax) ``` Configure WPA3-SAE with PMF required on the same controller and the band comes up: ``` WLC1#show wlan id 2 | include Operational State|Security-6GHz|SAE|PMF Support Operational State of Radio Bands : All Bands Operational Security-2.4GHz/5GHz SAE PWE Method : Hash to Element, Hunting and Pecking(H2E-HNP) PMF Support : Required Security-6GHz SAE PWE Method : Hash to Element(H2E) PMF Support : Required ``` Note the SAE password element method differs by band. The 6 GHz band is Hash to Element only, while 2.4 and 5 GHz permit H2E with Hunting and Pecking for backward compatibility. That detail matters when you are debugging why a client associates on 5 GHz but not 6 GHz. More on this in the [Wi-Fi 6 and 6E guide](https://www.pinglabz.com/c9800-wifi6-wifi6e/). ## One More Trap: the External Connector If you attach an external connector to your wireless management VLAN in the hope of bridging a physical AP later, be careful about leaving it up before that AP exists. In this lab the connector bridged the management VLAN straight onto the physical home network, and the access switch learned twenty-eight foreign MAC addresses on a single port: ``` ACCESS-SW1#show mac address-table vlan 10 10 000c.295b.368b DYNAMIC Et0/1 10 187f.888c.9209 DYNAMIC Et0/1 10 1cb3.c902.f42f DYNAMIC Et0/1 ... (28 total on Et0/1) 10 5254.00f1.2139 DYNAMIC Et0/2 Total Mac Addresses for this criterion: 32 ``` Worse, the upstream router on that network was also offering DHCP, so lab clients were racing two DHCP servers and leases were landing unpredictably. Shut the connector until you genuinely need it. ## Key Takeaways - The 9800-CL in CML is a control-plane lab. Configuration, syntax, tags, profiles, and security modes are all real and worth your time. - No access point can join a 9800-CL in CML because no virtual AP image exists. The CML 2.10 `wireless-ap` node is a Linux hostapd device and does not speak CAPWAP. - Expect the serial console to be dead on first boot. Save a day-0 configuration and restart, and it behaves from then on. - The default 2 vCPU and 6 GB sizing is Ultra-Low profile and cannot do local mode. FlexConnect works everywhere. Local mode needs 4 vCPU and 8 GB. - High availability configuration is accepted but never becomes operational. Check `Redundancy Mode (Operational)`, not the configured value. - If you need AP join, client, roaming, or RRM output, plan for a physical AP bridged in through an external connector. For the full cluster covering configuration, security, design, and troubleshooting, see the [Cisco wireless guide](https://www.pinglabz.com/wireless/). ### References - [Cisco Modeling Labs: Catalyst 9800-CL platform notes](https://developer.cisco.com/docs/modeling-labs/catalyst-9800-cl/?ref=pinglabz.com) - [Cisco Modeling Labs 2.10 release notes](https://developer.cisco.com/docs/modeling-labs/cml-release-notes/?ref=pinglabz.com) ### dnswalk: Audit a DNS Zone for Consistency URL: https://www.pinglabz.com/dnswalk/ Last updated: 2026-07-18T17:52:33.000Z dnswalk is the DNS enumeration tool that defenders should love as much as attackers do. It is a DNS debugger: it transfers a zone and then checks the database for internal consistency and correctness - missing PTR records, mismatched addresses, bad delegations, and more. Where the other tools in this cluster extract information, dnswalk audits it. This article is part of the PingLabz [DNS enumeration guide](https://www.pinglabz.com/dns-enumeration/), and the capture below is real output from Kali in our lab. ## What dnswalk does dnswalk performs a zone transfer of a domain and then runs a battery of consistency checks against the records it receives. It flags forward records with no matching reverse (PTR) record, reverse records that do not match their forward counterpart, CNAMEs pointing at nonexistent names, missing or inconsistent delegations, and other structural problems. It is a Perl script built on the Net::DNS library, and it has one quirk worth remembering: the domain argument must end with a trailing dot. Because it depends on being able to transfer the zone, dnswalk is at its most useful where you legitimately can - auditing your own domains. It ships with Kali; elsewhere, `sudo apt install dnswalk`. ## Auditing the lab zone Run with recursion (`-r`) and debug (`-d`) against the fully-qualified domain (note the trailing dot): ``` root@kali:~# dnswalk -r -d pinglabz.lab. Checking pinglabz.lab. SOA=ns1.pinglabz.lab contact=hostmaster.pinglabz.lab. WARN: gw.pinglabz.lab A 192.168.99.1: no PTR record WARN: ns1.pinglabz.lab A 192.168.99.100: no PTR record Getting zone transfer of pinglabz.lab. from ns1.pinglabz.lab...done. 0 failures, 2 warnings, 0 errors. ``` dnswalk transferred the zone from ns1, then reported its findings. The two warnings are real problems it caught: gw and ns1 both have A records in the 192.168.99.0/24 range, but the lab only defined a reverse zone for 10.0.53.0/24, so those two hosts have no matching PTR record. That is exactly the kind of consistency gap dnswalk exists to find. The summary line - `0 failures, 2 warnings, 0 errors` \- is the report card for the zone. For an attacker, dnswalk doubles as another way to dump a zone when transfers are allowed. For an administrator, it is a quick way to sanity-check a zone file after edits: forward-and-reverse mismatches, dangling CNAMEs, and delegation slips are precisely the errors that cause hard-to-diagnose resolution problems in production. ## The checks and flags that matter trailing dot The domain argument must end with a "." (pinglabz.lab.) or dnswalk refuses to run. \-r Recurse into subdomains, walking delegations below the target zone. \-d Debug output - show progress and the checks as they run. \-F Perform "facist" (strict) checking for spurious and questionable records. \-a Turn on warning for duplicate A records. WARN / FAIL / ERROR Severity levels in the output - warnings are style issues, failures and errors are real breakage. ## Where dnswalk fits dnswalk sits at the overlap of offense and defense. As a recon tool it is a zone-transfer client with extra analysis, so if AXFR is open it dumps and critiques the zone in one go. But its real home is on the administrator's side: run it after every zone change to catch missing PTRs, mismatches, and dangling records before they cause outages. It depends entirely on being able to transfer the zone, which - if you have locked `allow-transfer` down properly - means you run it from a host that *is* allowed, not from the open internet. ## The defender's view This is the tool to add to your own DNS hygiene routine. Schedule dnswalk against your zones (from an authorised secondary) and treat its failures and errors as a punch list. But remember what makes it work: a zone transfer. If dnswalk can pull your zone from an *unauthorised* host, that is the finding - restrict `allow-transfer` to your secondaries. The complete hardening checklist is in the [DNS enumeration pillar](https://www.pinglabz.com/dns-enumeration/). ## Related tools dnswalk is the auditor's counterpart to the extractors [dnsrecon](https://www.pinglabz.com/dnsrecon/) and [dnsenum](https://www.pinglabz.com/dnsenum/). For subdomain discovery when transfers are blocked, use [dnsmap](https://www.pinglabz.com/dnsmap/) or [massdns](https://www.pinglabz.com/massdns/), and to understand delegation use [dnstracer](https://www.pinglabz.com/dnstracer/). Once your recon or audit is done, scan live hosts with the [Nmap cluster](https://www.pinglabz.com/nmap/). ## Key takeaways dnswalk transfers a zone and audits it for consistency - missing PTRs, forward/reverse mismatches, dangling CNAMEs, and delegation errors. In the lab it correctly flagged two hosts with no reverse record and summarised the zone as 0 failures, 2 warnings, 0 errors. Remember the trailing dot on the domain. It is as valuable to defenders auditing their own zones as it is to attackers dumping an open one - and either way, the fact that it needs a zone transfer to work is the reminder to lock those transfers down. *Only run dnswalk against domains you own or are explicitly authorised to test. The output here is from a private lab.* ### dnstracer: Follow the DNS Delegation Chain URL: https://www.pinglabz.com/dnstracer/ Last updated: 2026-07-18T17:52:32.000Z Most DNS enumeration tools ask a server for records. dnstracer asks a different question: *where does this answer actually come from?* It determines where a given nameserver gets its information for a hostname and follows the chain of DNS servers back to the authoritative source. That makes it the odd one out in the enumeration toolkit - less about dumping records, more about understanding delegation and trust. This article is part of the PingLabz [DNS enumeration guide](https://www.pinglabz.com/dns-enumeration/), and the capture below is real output from Kali in our lab. ## What dnstracer does DNS is a hierarchy. When you resolve a name on the public internet, the query walks down a delegation chain: the root servers point at the TLD servers, which point at the domain's authoritative servers, which give the final answer. dnstracer makes that chain visible. Starting from a server you choose (or the root servers), it follows the referrals down until it reaches the server that answers authoritatively, showing each hop and whether the answer was authoritative or cached. For a recon analyst, that is useful for mapping who controls a domain's DNS, spotting third-party DNS providers, and understanding delegation before deciding which servers to probe. dnstracer is in the Kali repos: `sudo apt install dnstracer`. ## Tracing in the lab In an isolated lab there is no root hierarchy - the BIND9 server is directly authoritative for the zone - so dnstracer resolves to a single authoritative hop. Pointing it at the lab server with `-s`, enabling the overview (`-o`) and asking for an A record shows exactly that: ``` root@kali:~# dnstracer -s 192.168.99.100 -r 2 -o -q A www.pinglabz.lab Tracing to www.pinglabz.lab[a] via 192.168.99.100, maximum of 2 retries 192.168.99.100 (192.168.99.100) Got authoritative answer 192.168.99.100 (192.168.99.100) www.pinglabz.lab -> 10.0.53.10 ``` The key phrase is `Got authoritative answer`. dnstracer confirms that 192.168.99.100 is not just relaying a cached record - it is the authority for this name, and it maps www.pinglabz.lab to 10.0.53.10\. Against a real public domain the same command produces a multi-line tree: the root servers, the TLD servers, and finally the domain's authoritative nameservers, with each hop labelled by whether it gave an authoritative answer or a referral. That tree is dnstracer's real value - it shows the delegation path, not just the final IP. ## The options that matter \-s Server for the initial request. Use "." to start from the root servers (A.ROOT-SERVERS.NET). \-q Query type for the requests (default A). Trace MX, NS, or other record types. \-o Print an overview of received answers at the end - the clean summary view. \-r Number of retries per request (default 3). Lower it to move faster on a reliable network. \-v Verbose - dump the full DNS header and packet detail for each hop. \-4 Do not query IPv6 servers - handy on IPv4-only networks. ## Where dnstracer fits dnstracer is a specialist, not an all-rounder. It will not brute force subdomains or dump a zone. Its job is delegation and provenance: confirming which server is authoritative, spotting inconsistencies where different servers give different answers, and mapping who actually runs a domain's DNS. In practice you use it alongside the enumeration tools - [dnsrecon](https://www.pinglabz.com/dnsrecon/) and [dnsenum](https://www.pinglabz.com/dnsenum/) tell you what the records are, and dnstracer tells you where they come from. It is also handy for debugging your own delegation when a name resolves differently than you expect. ## The defender's view dnstracer does not extract anything sensitive - it walks public delegation that has to be visible for DNS to work. But it is a useful diagnostic on your own side: use it to confirm your nameservers all give authoritative answers and agree with each other, and to verify a delegation change has propagated as intended. Inconsistent authoritative answers between your nameservers are a real misconfiguration that dnstracer surfaces quickly. General DNS hardening lives in the [DNS enumeration pillar](https://www.pinglabz.com/dns-enumeration/). ## Related tools Pair dnstracer with the record-dumping tools [dnsrecon](https://www.pinglabz.com/dnsrecon/) and [dnsenum](https://www.pinglabz.com/dnsenum/). For subdomain discovery use [dnsmap](https://www.pinglabz.com/dnsmap/) or [massdns](https://www.pinglabz.com/massdns/), and to audit a zone's contents use [dnswalk](https://www.pinglabz.com/dnswalk/). When recon is complete, scan the hosts you found with the [Nmap cluster](https://www.pinglabz.com/nmap/). ## Key takeaways dnstracer answers the provenance question in DNS enumeration: it follows the delegation chain from a chosen server (or the root) down to the authoritative source, showing each hop and whether it answered authoritatively. It is a specialist for mapping who controls a domain's DNS and for spotting inconsistent or misdelegated answers, not a record dumper or brute forcer. Use it beside dnsrecon and dnsenum, and use it on your own domains to confirm your nameservers are authoritative and in agreement. *Only run dnstracer against domains you own or are explicitly authorised to test. The output here is from a private lab.* ### massdns: High-Performance DNS Enumeration at Scale URL: https://www.pinglabz.com/massdns/ Last updated: 2026-07-18T17:52:31.000Z Subdomain brute forcing is only useful if you can go big, and going big means going fast. massdns is the tool that makes internet-scale DNS enumeration practical: a high-performance stub resolver built to resolve millions - even billions - of names by spreading queries across many public resolvers at once. Without any special tuning it will push over 350,000 names per second. This article is part of the PingLabz [DNS enumeration guide](https://www.pinglabz.com/dns-enumeration/), and the capture below is real output from Kali in our lab. ## What massdns does massdns is a stub resolver, not a brute forcer on its own. You feed it two things: a list of names to resolve and a list of DNS resolvers to spread the load across. It fires the queries out concurrently, tracks which ones succeed, and writes the answers in whatever output format you choose. In a modern subdomain-discovery pipeline it is the resolution engine: a tool like a wordlist generator or permutation tool produces candidate names, and massdns resolves them at speed. It is the workhorse behind many bug-bounty recon workflows. massdns is in the Kali repos: `sudo apt install massdns`. It is also trivially built from source. ## Running massdns against the lab You need a resolvers file (here, just the lab's BIND9 server) and a names file (candidate FQDNs). Then point massdns at both, choose record type A and simple output (`-o S`), and write to a file: ``` root@kali:~# massdns -r resolvers.txt -t A -o S -w massdns.out names.txt Concurrency: 10000 Processed queries: 34 Received packets: 34 Progress: 100.00% (00 h 00 min 00 sec / 00 h 00 min 00 sec) Current incoming rate: 7071 pps, average: 7071 pps Current success rate: 7071 pps, average: 7071 pps Finished total: 34, success: 34 (100.00%) root@kali:~# sort massdns.out admin.pinglabz.lab. A 10.0.53.99 backup.pinglabz.lab. A 10.0.53.60 db.pinglabz.lab. A 10.0.53.70 dc01.pinglabz.lab. A 10.0.53.5 jenkins.pinglabz.lab. A 10.0.53.91 mail.pinglabz.lab. A 10.0.53.20 router.pinglabz.lab. CNAME gw.pinglabz.lab. smtp.pinglabz.lab. CNAME mail.pinglabz.lab. vpn.pinglabz.lab. A 10.0.53.30 www.pinglabz.lab. A 10.0.53.10 ``` The live status screen is the massdns signature: concurrency, queries processed, and a real-time packets-per-second rate. Even in this tiny lab it hit 7,071 pps at 100% success. Scale the names file to millions and point it at a large resolver list and that rate is what lets it finish in minutes instead of days. The output resolves CNAME chains too, so you see that router is an alias for gw and smtp for mail. ## The options that matter \-r Resolvers file - the list of DNS servers to spread queries across. The heart of the tool. \-t Record type to resolve (default A). Use AAAA, CNAME, TXT, etc. as needed. \-o Output format: S (simple), F (full), L (domain list), J (ndjson) with sub-flags for detail. \-w Write results to a file instead of stdout. \-s Hashmap size - concurrent lookups (default 10000). Tune to your bandwidth and resolvers. \--norecurse Send non-recursive queries - useful for DNS cache snooping and resolver testing. One practical warning: massdns is only as reliable as its resolver list. A list full of dead or poisoned public resolvers produces false positives and misses. Recon practitioners maintain curated, regularly validated resolver lists for exactly this reason. Because massdns just resolves what you give it, the quality of your candidate names (from a good wordlist or a permutation tool) determines what you actually find. ## The defender's view massdns does not attack your server so much as hammer it, and a flood of sequential lookups for guessed names is a detectable pattern. Response Rate Limiting and query-log analysis on your authoritative servers will surface it. The deeper defence is the same as for any brute forcing: make sure a resolved name only ever returns a public address, so speed buys the attacker nothing but a list of things they were already meant to see. The full checklist is in the [DNS enumeration pillar](https://www.pinglabz.com/dns-enumeration/). ## Related tools massdns is the scale version of what [dnsmap](https://www.pinglabz.com/dnsmap/) does slowly. Generate the standard records and try zone transfers first with [dnsrecon](https://www.pinglabz.com/dnsrecon/) and [dnsenum](https://www.pinglabz.com/dnsenum/) \- if AXFR works, you do not need to brute force at all. To audit your own zone for errors, use [dnswalk](https://www.pinglabz.com/dnswalk/). Feed the live hosts massdns finds into the [Nmap cluster](https://www.pinglabz.com/nmap/) for port scanning. ## Key takeaways massdns is the speed tool of DNS enumeration - a stub resolver that spreads a huge list of candidate names across many resolvers to resolve hundreds of thousands per second. Give it a names file and a resolvers file, pick a record type and output format, and let it run. Its results are only as good as your resolver list and your candidate names, so curate both. It is the scale-up from dnsmap, and the defence against it is the same: keep internal addresses out of public zones and monitor for brute-force patterns. *Only run massdns against domains and resolvers you own or are explicitly authorised to test. The output here is from a private lab.* ### dnsmap: DNS Subdomain Brute Forcing on Kali URL: https://www.pinglabz.com/dnsmap/ Last updated: 2026-07-18T17:52:31.000Z When a nameserver refuses zone transfers - which, these days, is most of them - subdomain brute forcing is how you enumerate a domain. dnsmap is the tool built for exactly that: point it at a domain and it grinds through a wordlist of common hostnames, keeping every one that resolves. It is small, fast, needs no root, and has a habit of flagging when a domain leaks internal addresses. This article is part of the PingLabz [DNS enumeration guide](https://www.pinglabz.com/dns-enumeration/), and the output below is real, captured from Kali against a BIND9 server in our lab. ## What dnsmap does dnsmap scans a domain for subdomains using brute force. It ships with a built-in wordlist of around 1,000 common names (ns1, mail, smtp, firewall, vpn and the like) in English and Spanish, or you can supply your own with `-w`. For every name that resolves, it prints the hostname and IP, and results can be saved as plain text or CSV for later processing. Crucially, dnsmap uses the system resolver, so whatever nameserver is in `/etc/resolv.conf` is who it asks. As the author notes, dnsmap is deliberately meant for the enumeration phase of an assessment - and it is especially useful precisely because zone transfers are so rarely allowed publicly anymore. It comes with Kali; elsewhere, `sudo apt install dnsmap`. Run it as a normal user, not root. ## Brute forcing the lab domain With the resolver pointed at the lab's BIND9 server, dnsmap chews through a wordlist against `pinglabz.lab`: ``` root@kali:~# dnsmap pinglabz.lab -w /home/j/dns-recon/subs.txt dnsmap 0.36 - DNS Network Mapper [+] searching (sub)domains for pinglabz.lab using /home/j/dns-recon/subs.txt [+] using maximum random delay of 10 millisecond(s) between requests www.pinglabz.lab IP address #1: 10.0.53.10 [+] warning: internal IP address disclosed vpn.pinglabz.lab IP address #1: 10.0.53.30 [+] warning: internal IP address disclosed jenkins.pinglabz.lab IP address #1: 10.0.53.91 [+] warning: internal IP address disclosed admin.pinglabz.lab IP address #1: 10.0.53.99 [+] warning: internal IP address disclosed [+] 22 (sub)domains and 22 IP address(es) found [+] 22 internal IP address(es) disclosed [+] completion time: 1 second(s) ``` The standout feature is that `warning: internal IP address disclosed` line. dnsmap recognises RFC 1918 addresses and calls them out - because a public DNS record pointing at 10.0.53.x tells an attacker the internal network layout without ever touching the internal network. Here it found 22 hosts and flagged all 22 as internal-address disclosures. That summary at the end (22 subdomains, 22 internal IPs disclosed, one second) is the whole tool in three lines. ## The options that matter \-w Use an external wordlist instead of the built-in one. Bigger, targeted lists find more. \-r Save results to a plain-text file (or a directory for an auto-named file). \-c Save results in CSV format for further processing. \-d Maximum random delay between queries. Raise it to be gentler on bandwidth or stealthier. \-i Ignore specific IPs in results - useful for filtering out wildcard false positives. dnsmap-bulk A companion script that runs dnsmap across a whole file of domains in bulk. ## Where dnsmap fits dnsmap does one thing - subdomain brute forcing - and does it simply. It has no zone-transfer logic and no record-type enumeration, so it is not a replacement for [dnsrecon](https://www.pinglabz.com/dnsrecon/) or [dnsenum](https://www.pinglabz.com/dnsenum/); it is the focused brute forcer you reach for when those tools' AXFR attempts come back empty. Its one real limitation is speed: it queries sequentially, so against a large wordlist it is slow. When that matters, step up to [massdns](https://www.pinglabz.com/massdns/), which does the same job at tens of thousands of names per second. ## The defender's view dnsmap's internal-IP warnings are your to-do list. If your public DNS resolves names to RFC 1918 addresses, move those records into an internal-only view with split-horizon DNS. You cannot stop brute forcing outright - each query looks like a normal lookup - but you can make sure a successful guess only ever returns a public address, and you can rate-limit and log to spot the burst of sequential queries that brute forcing produces. See the hardening checklist in the [DNS enumeration pillar](https://www.pinglabz.com/dns-enumeration/). ## Related tools Use dnsmap after [dnsrecon](https://www.pinglabz.com/dnsrecon/) and [dnsenum](https://www.pinglabz.com/dnsenum/) when zone transfers are blocked. For the same brute force at scale, use [massdns](https://www.pinglabz.com/massdns/). To audit your own zone, run [dnswalk](https://www.pinglabz.com/dnswalk/). Once you have live hostnames, scan them with the [Nmap cluster](https://www.pinglabz.com/nmap/). ## Key takeaways dnsmap is the go-to subdomain brute forcer for when zone transfers are locked down. It runs a wordlist against a domain using the system resolver, keeps every name that resolves, and flags any internal (RFC 1918) addresses it uncovers. It is simple and needs no root, but it is single-threaded and slow on large lists - reach for massdns when you need scale. For defenders, its internal-IP warnings are a direct signal to adopt split-horizon DNS. *Only run dnsmap against domains you own or are explicitly authorised to test. The output here is from a private lab.* ### dnsenum: Enumerate a Domain's DNS End to End URL: https://www.pinglabz.com/dnsenum/ Last updated: 2026-07-18T17:52:30.000Z dnsenum is the classic DNS enumeration tool - a multithreaded Perl script that has been in every pentester's kit for over a decade. Point it at a domain and it gathers the host address, the nameservers, the mail servers, attempts a zone transfer on each nameserver, brute forces subdomains, and can even work out and reverse-sweep the surrounding network ranges, all in one report. This article is part of the PingLabz [DNS enumeration guide](https://www.pinglabz.com/dns-enumeration/), and every capture below is real output from a Kali box against a BIND9 server in our lab. ## What dnsenum does The point of dnsenum is to gather as much information about a domain as possible in a single run. It performs, in order: get the host's A record, get the nameservers, get the MX records, run AXFR against each nameserver and grab BIND versions, scrape Google for extra subdomains, brute force subdomains from a file (with optional recursion on any subdomain that has its own NS records), calculate class-C network ranges, run whois on them, and reverse-lookup the results. It is deliberately thorough. dnsenum ships with Kali. Elsewhere: `sudo apt install dnsenum`. ## Running dnsenum against the lab A focused run points dnsenum at a specific server with `--dnsserver`, supplies a brute-force wordlist with `-f`, and skips the noisy reverse-lookup phase with `--noreverse`: ``` root@kali:~# dnsenum --dnsserver 192.168.99.100 --nocolor --noreverse -f /home/j/dns-recon/subs.txt pinglabz.lab dnsenum VERSION:1.3.1 ----- pinglabz.lab ----- Host's addresses: __________________ pinglabz.lab. 604800 IN A 10.0.53.10 Name Servers: ______________ ns1.pinglabz.lab. 604800 IN A 192.168.99.100 ns2.pinglabz.lab. 604800 IN A 10.0.53.11 Mail (MX) Servers: ___________________ mail.pinglabz.lab. 604800 IN A 10.0.53.20 mail2.pinglabz.lab. 604800 IN A 10.0.53.21 Trying Zone Transfers and getting Bind Versions: _________________________________________________ Trying Zone Transfer for pinglabz.lab on ns2.pinglabz.lab ... Trying Zone Transfer for pinglabz.lab on ns1.pinglabz.lab ... pinglabz.lab. 604800 IN SOA ( ... pinglabz.lab. 604800 IN NS ns1.pinglabz.lab. pinglabz.lab. 604800 IN MX 10 _kerberos._tcp.pinglabz.lab. 604800 IN SRV 0 _ldap._tcp.pinglabz.lab. 604800 IN SRV 0 admin.pinglabz.lab. 604800 IN A 10.0.53.99 dc01.pinglabz.lab. 604800 IN A 10.0.53.5 jenkins.pinglabz.lab. 604800 IN A 10.0.53.91 vpn.pinglabz.lab. 604800 IN A 10.0.53.30 ``` dnsenum lays the report out in labelled sections, which is why people still love it: host address, nameservers, and mail servers up top, then it tries a zone transfer on each nameserver. ns2 (the dead address) returns nothing; ns1 hands over the full zone, including the SRV records that reveal LDAP and Kerberos - a strong hint that Active Directory sits behind this domain. ## The options that matter \--dnsserver Use this specific server for A, NS and MX queries. Point it straight at the target's nameserver. \-f Wordlist for subdomain brute forcing. Takes priority over the default dns.txt. \-r / --recursion Recursively brute force any discovered subdomain that has its own NS record. \--enum Shortcut for --threads 5 -s 15 -w (threads, Google scrape, whois). Convenient but noisy. \--noreverse Skip reverse lookups. Faster and quieter when you only care about forward records. \-o Write XML output, importable into tools like MagicTree. ## dnsenum vs dnsrecon The two tools cover almost the same ground, and both belong in your kit. dnsenum's report layout (labelled sections, BIND version grab, automatic network-range discovery and reverse sweeps) reads beautifully and is great for a first look. [dnsrecon](https://www.pinglabz.com/dnsrecon/) is more granular, actively maintained in Python, and exports structured XML/CSV/JSON that fits into automation. Many testers run dnsenum for the readable summary and dnsrecon when they need machine-readable output or a specific enumeration type. ## The defender's view dnsenum tries a zone transfer on every nameserver it finds, so the same fix applies: restrict `allow-transfer` to your secondaries. The tool also grabs BIND version strings, so consider setting `version "none";` in your BIND options to stop advertising your software version. The network-range discovery and reverse sweeps only work when your PTR records are populated for public ranges - another reason internal addressing does not belong in public zones. The full checklist is in the [DNS enumeration pillar](https://www.pinglabz.com/dns-enumeration/). ## Related tools For structured output and SRV-focused enumeration, pair dnsenum with [dnsrecon](https://www.pinglabz.com/dnsrecon/). When you need heavier brute forcing, use [dnsmap](https://www.pinglabz.com/dnsmap/) or [massdns](https://www.pinglabz.com/massdns/). To audit a transferred zone for errors, run [dnswalk](https://www.pinglabz.com/dnswalk/). When recon turns up live hosts, move to scanning them with the [Nmap cluster](https://www.pinglabz.com/nmap/). ## Key takeaways dnsenum is the readable all-in-one DNS enumerator: host addresses, nameservers, mail servers, per-nameserver zone transfers, subdomain brute forcing, and network-range discovery in one labelled report. Point it at the target's nameserver with `--dnsserver`, add a wordlist with `-f`, and use `--noreverse` to stay quiet. It overlaps with dnsrecon by design - run both, and defend against both by restricting zone transfers and hiding your BIND version. *Only run dnsenum against domains you own or are explicitly authorised to test. The output here is from a private lab.* ### dnsrecon: The Complete DNS Enumeration Tool Guide URL: https://www.pinglabz.com/dnsrecon/ Last updated: 2026-07-18T17:52:29.000Z If you learn one DNS enumeration tool, make it dnsrecon. It is the Swiss Army knife of the category: a single, actively maintained Python script that pulls standard records, attempts zone transfers against every nameserver, enumerates SRV records, brute forces subdomains, does reverse lookups, and exports clean XML, CSV, or JSON. This article is part of the PingLabz [DNS enumeration guide](https://www.pinglabz.com/dns-enumeration/), and every command below was run from Kali against a real BIND9 server in a lab we control. ## What dnsrecon does dnsrecon (the "DNS Recon" tool) rolls the entire enumeration workflow into one command with a `-t` type switch. It can check all NS records for zone transfers, enumerate general records (MX, SOA, NS, A, AAAA, SPF, TXT), perform SRV record enumeration, expand top-level domains, check for wildcard resolution, brute force subdomains from a wordlist, and do PTR lookups across a CIDR range. That breadth is why most testers reach for it first. It comes preinstalled on Kali. Elsewhere: `sudo apt install dnsrecon`, or `pip install dnsrecon` on any system with Python 3.6+. ## The lab target The target for this cluster is a BIND9 server at 192.168.99.100, authoritative for the internal domain `pinglabz.lab`. The zone lists two nameservers (ns1 is live, ns2 is a dead address), a pair of mail servers, SPF and verification TXT records, SRV records for LDAP and Kerberos, an IPv6 host, and a spread of internal A records. It deliberately allows zone transfers to the attacker so you can see the worst case. Full topology is in the [pillar guide](https://www.pinglabz.com/dns-enumeration/). ## Standard enumeration and the zone transfer The highest-value single command is a standard enumeration with a zone-transfer attempt (`-a`), pointed at a specific nameserver with `-n`: ``` root@kali:~# dnsrecon -n 192.168.99.100 -d pinglabz.lab -a [*] Checking for Zone Transfer for pinglabz.lab name servers [*] Resolving SOA Record [*] SOA ns1.pinglabz.lab 192.168.99.100 [*] Resolving NS Records [*] NS Servers found: [*] NS ns2.pinglabz.lab 10.0.53.11 [*] NS ns1.pinglabz.lab 192.168.99.100 [*] Trying NS server 10.0.53.11 [-] Zone Transfer Failed for 10.0.53.11! [-] Port 53 TCP is being filtered [*] Trying NS server 192.168.99.100 [+] 192.168.99.100 Has port 53 TCP Open [+] Zone Transfer was successful!! [*] NS ns1.pinglabz.lab 192.168.99.100 [*] TXT v=spf1 mx a:mail.pinglabz.lab -all [*] TXT plz-verification=dns-recon-lab-2026 [*] AAAA ipv6host.pinglabz.lab 2001:db8:53::10 [*] A admin.pinglabz.lab 10.0.53.99 [*] A backup.pinglabz.lab 10.0.53.60 [*] A dc01.pinglabz.lab 10.0.53.5 [*] A jenkins.pinglabz.lab 10.0.53.91 [*] A mail.pinglabz.lab 10.0.53.20 [*] A vpn.pinglabz.lab 10.0.53.30 ``` Notice how dnsrecon walks every nameserver in the NS set. The dead ns2 refuses (port 53 TCP filtered), then ns1 answers and the whole zone spills out - mail servers, the domain controller, a Jenkins box, the VPN gateway, an IPv6 host, and every internal 10.0.53.x address. That is the entire value of the tool in ten seconds. ## Brute forcing when the transfer is blocked On a properly configured target, that AXFR fails on every nameserver. Then you switch to brute force with `-t brt` and a wordlist: ``` root@kali:~# dnsrecon -n 192.168.99.100 -d pinglabz.lab -D /home/j/dns-recon/subs.txt -t brt [*] Performing host and subdomain brute force against pinglabz.lab... [*] A mail2.pinglabz.lab 10.0.53.21 [*] A www.pinglabz.lab 10.0.53.10 [*] A vpn.pinglabz.lab 10.0.53.30 [*] A intranet.pinglabz.lab 10.0.53.50 [*] A db.pinglabz.lab 10.0.53.70 [*] CNAME smtp.pinglabz.lab mail.pinglabz.lab [*] A jenkins.pinglabz.lab 10.0.53.91 [*] A dc01.pinglabz.lab 10.0.53.5 [*] CNAME web.pinglabz.lab www.pinglabz.lab [*] CNAME router.pinglabz.lab gw.pinglabz.lab [*] 25 Records Found ``` Brute force resolves CNAME aliases too, so you see the relationships (smtp is really mail, web is really www). It only finds names in your wordlist, so quality matters - the wordlists in `/usr/share/wordlists/` and SecLists are the usual starting points. ## The options that matter \-d DOMAIN The target domain. Required for everything. \-n NS\_SERVER Query a specific nameserver instead of the target's SOA. Essential when testing one server directly. \-a Attempt AXFR on every nameserver alongside standard enumeration. The high-value flag. \-t TYPE Enumeration type: std, brt (brute), srv, axfr, rvl (reverse), zonewalk, and more. \-D DICTIONARY Wordlist for brute forcing subdomains and hosts. \-x / -c / -j Export found records to XML, CSV, or JSON for reporting and pipelines. A common one-liner for a real engagement combines standard enumeration, a wordlist, and XML output: `dnsrecon -d target.com -D /usr/share/wordlists/dnsmap.txt -t std --xml dnsrecon.xml`. dnsrecon also checks the target's BIND version and tests for NXDOMAIN hijacking unless you disable those with `--disable_check_bindversion` and `--disable_check_nxdomain`. ## The defender's view The scariest line in the output above is "Zone Transfer was successful". Lock `allow-transfer` down to your known secondaries by IP and TSIG key, and that entire capture becomes a row of failures. Run dnsrecon against your own domain periodically - if `-a` ever succeeds from an unauthorised host, you have a finding to fix. See the hardening checklist in the [DNS enumeration pillar](https://www.pinglabz.com/dns-enumeration/). ## Related tools dnsrecon overlaps with [dnsenum](https://www.pinglabz.com/dnsenum/), which produces a similar all-in-one report in Perl. For heavy subdomain brute forcing, hand off to [dnsmap](https://www.pinglabz.com/dnsmap/) or, at scale, [massdns](https://www.pinglabz.com/massdns/). To audit a zone you have transferred, use [dnswalk](https://www.pinglabz.com/dnswalk/). Once recon is done and you have live hosts, move on to port scanning with the [Nmap cluster](https://www.pinglabz.com/nmap/). ## Key takeaways dnsrecon is the first DNS enumeration tool to learn because it does the whole workflow in one script: standard records, zone transfers across every nameserver, SRV enumeration, and subdomain brute forcing, with clean export formats. Lead with `-a` to catch a leaky server, fall back to `-t brt` with a good wordlist when transfers are blocked, and export to XML or JSON for reporting. Then run it against your own domains, because the same command an attacker uses is the fastest way to find out what your DNS is giving away. *Only run dnsrecon against domains you own or are explicitly authorised to test. The output here is from a private lab.* ### IPv6 Address Examples: Every Type, Worked Through URL: https://www.pinglabz.com/ipv6-address-examples/ Last updated: 2026-07-17T13:00:00.000Z IPv6 addresses look intimidating until you have seen enough real ones to recognize the patterns. The first hextet usually tells you everything: what kind of address it is and what it is for. This post is a worked reference - one concrete example of every IPv6 address type you will actually encounter, with the address broken down field by field, so the patterns stick. For the cluster overview, see the [IPv6 complete guide](https://www.pinglabz.com/ipv6/). ## The format, briefly An IPv6 address is 128 bits, written as eight groups of four hex digits (hextets) separated by colons. Two compression rules make them shorter: - **Drop leading zeros** in any hextet: `0db8` becomes `db8`, `0000` becomes `0`. - **Replace one run of all-zero hextets with `::`** \- but only once per address, otherwise it would be ambiguous. So `2001:0db8:0000:0000:0000:0000:0000:0001` compresses to `2001:db8::1`. Every example below is shown compressed. ## Global unicast - the public, routable address A global unicast address (GUA) is the IPv6 equivalent of a public IPv4 address: globally unique, internet-routable. Current GUAs are assigned out of **2000::/3**, which in practice means an address starting with 2 or 3. Worked example: `2001:db8:acad:10::5` Global routing prefix Bits/48 Value here2001:db8:acad Meaning Assigned to the organization by an ISP/RIR Subnet ID Bits16 Value here0010 Meaning Which subnet inside the org - here subnet 10 Interface ID Bits/64 Value here::5 Meaning The host portion - identifies the device The almost-universal split is a /64 for every subnet: 64 bits of network, 64 bits of interface ID. Note `2001:db8::/32` is the documentation range (reserved for examples like this one), so you will see it constantly in guides and never in the wild. ## Link-local - exists on every interface, never routed Every IPv6-enabled interface automatically has a link-local address. It is valid only on its own link, is never routed off that link, and is what neighbor discovery and routing-protocol hellos use. Link-local addresses come from **fe80::/10**, so they start with `fe80`. Worked example: `fe80::1ff:fe23:4567:890a` The `fe80` prefix is the giveaway. The interface ID is often EUI-64 derived (you can spot the tell-tale `fffe` inserted in the middle) or randomized. Because link-local addresses are not unique beyond the link, you often must specify which interface you mean - `ping fe80::1%GigabitEthernet0/0`. ## Unique local - private addressing Unique local addresses (ULAs) are the IPv6 counterpart to RFC 1918 private IPv4 space. They are routable inside an organization but not on the global internet. ULAs come from **fc00::/7**, and in practice you see **fd00::/8** because the locally-assigned bit is set. Worked example: `fd12:3456:789a:1::20` Prefix Valuefd MeaningULA, locally assigned Global ID Value12:3456:789a Meaning A 40-bit pseudo-random value - generate it, do not pick a tidy one Subnet ID Value0001 MeaningSubnet 1 Interface ID Value::20 MeaningThe host ## Multicast - one-to-many IPv6 has no broadcast. Its job is done by multicast, and every multicast address starts with **ff00::/8** \- the leading `ff` is the marker. The character after `ff` encodes flags and scope (1 = node, 2 = link, 5 = site). ff02::1 NameAll-nodes Used by Every IPv6 host on the link ff02::2 NameAll-routers Used by Every IPv6 router on the link ff02::5 NameAll OSPFv3 routers Used byOSPFv3 ff02::1:ff00:0/104 NameSolicited-node Used by Neighbor discovery (the IPv6 replacement for ARP) The solicited-node multicast address is worth knowing: a host joins the group `ff02::1:ff` plus the last 24 bits of its unicast address, so neighbor discovery can target it without flooding the whole link the way ARP does. ## The two special single addresses Loopback Address::1 Meaning The host itself - the IPv6 equivalent of 127.0.0.1 Unspecified Address:: Meaning "No address yet" - a source address before one is assigned Both are mostly zeros, which is why they compress so dramatically. `::1` is 127 zero bits and a one; `::` is all 128 bits zero. ## A compression worked example Compression trips people up most, so one full example end to end. Take the address `2001:0db8:0000:0000:0000:00a3:0000:8a2e`. - **Drop leading zeros** in each hextet: `2001:db8:0:0:0:a3:0:8a2e`. - **Collapse the longest run of zero hextets** with `::`. The longest run is the three zeros in the middle, so: `2001:db8::a3:0:8a2e`. Note what you cannot do: the single `0` hextet after `a3` stays as `0`, because `::` has already been used and an address may contain it only once. If two zero runs are equal length, collapse the first. Getting this right matters because `2001:db8::a3:0:8a2e` and the fully-expanded form are the same address - tools accept both, and you will see both in the wild. ## Anycast, briefly One type that has no special prefix: **anycast**. An anycast address is an ordinary unicast address (usually a global unicast) deliberately assigned to more than one device. A packet sent to it is delivered to the *nearest* instance by routing metric. There is no flag in the address that marks it as anycast - it looks like any other unicast address, and only the configuration on multiple devices makes it anycast. It is used for things like distributing a DNS resolver across many sites behind one address. ## Reading the first hextet at a glance 2 or 3 (2000::/3) Global unicast - routable fe80 (fe80::/10) Link-local - this link only fd (fd00::/8) Unique local - private ff (ff00::/8) Multicast ::1 / :: Loopback / unspecified ## Common gotchas Using `::` twice in one address Only one `::` is allowed - two would be ambiguous about how many zero hextets each represents. Expecting a broadcast address IPv6 has none. `ff02::1` (all-nodes multicast) is the closest equivalent. Pinging a link-local without an interface Link-local addresses are not link-unique. Append the zone ID: `fe80::1%Gi0/0`. Subnetting smaller than /64 SLAAC and EUI-64 assume a /64\. Going smaller breaks host autoconfiguration. Hand-picking a tidy ULA global ID The global ID is meant to be pseudo-random so two ULA networks rarely collide when merged. ## Key takeaways Every IPv6 address type has a recognizable opening. Global unicast starts with 2 or 3 and is the public, routable address, split as routing prefix, subnet ID, and a /64 interface ID. Link-local starts with `fe80`, exists automatically on every interface, and never leaves its link. Unique local starts with `fd` and is private addressing. Multicast starts with `ff` and replaces broadcast, with `ff02::1` and `ff02::2` as the everyday all-nodes and all-routers groups. `::1` is loopback and `::` is unspecified. Learn to read the first hextet and an IPv6 address tells you what it is before you read the rest. For the IPv6 cluster, see the [IPv6 pillar](https://www.pinglabz.com/ipv6/). ### HSRP Configuration on Cisco IOS XE URL: https://www.pinglabz.com/hsrp-configuration/ Last updated: 2026-08-01T19:35:30.000Z HSRP configuration is short - a working setup is two lines per router - but the short config hides the decisions that actually matter. Which router is active, whether it stays active after a reboot, and what happens when its uplink fails are all controlled by commands people skip. This post is a configuration walkthrough on Cisco IOS XE that covers the minimum config and the four settings that make HSRP behave in production. For the cluster overview, see the [FHRP complete guide](https://www.pinglabz.com/fhrp/). ## What you are building HSRP gives a LAN a **virtual gateway**. Two (or more) routers share a virtual IP address and a virtual MAC address. Hosts point their default gateway at that virtual IP. One router is **active** and answers for the virtual IP and MAC; the other is **standby**, waiting. If the active router fails, the standby takes over the same virtual IP and MAC, and hosts never notice - their gateway address did not change. So the configuration job is: pick a virtual IP, decide which router should normally own it, and decide how it should react to failures. ## Minimum viable config Two distribution-layer routers, R1 and R2, on VLAN 10 (subnet 10.20.0.0/24). The virtual gateway will be 10.20.0.1\. Each router has its own real address on the SVI. ``` ! R1 interface Vlan10 ip address 10.20.0.2 255.255.255.0 standby version 2 standby 10 ip 10.20.0.1 ! R2 interface Vlan10 ip address 10.20.0.3 255.255.255.0 standby version 2 standby 10 ip 10.20.0.1 ``` That is a complete, working HSRP setup. The `10` is the HSRP group number, the `ip 10.20.0.1` is the virtual gateway. Hosts on VLAN 10 use 10.20.0.1 as their default gateway. Without anything else, the router with the higher IP address becomes active (R2 here) - which is rarely the router you actually wanted, hence the next section. ## Setting 1: priority - choose the active router HSRP priority ranges 0-255, default 100\. The highest priority becomes active. To make R1 the intended active router, give it a priority above the default: ``` ! R1 interface Vlan10 standby 10 priority 110 ``` Leave R2 at the default 100\. R1's 110 beats it. Always set priority explicitly - relying on the default means the active router is decided by whichever IP address happens to be higher, which is not a design decision. ## Setting 2: preempt - the one everyone forgets Here is the trap. You set R1 to priority 110\. R1 reboots. While it is down, R2 (priority 100) becomes active - correct. R1 comes back up. Does R1 take active back? **By default, no.** HSRP is non-preemptive. A higher-priority router that comes online *after* a lower-priority router is already active will sit in standby. R1 has priority 110 and is still standby behind R2's 100\. Your carefully chosen "primary" is now the backup, indefinitely. The fix is `preempt`. It tells a router to take the active role back as soon as it can, if its priority is higher: ``` ! R1 interface Vlan10 standby 10 preempt ``` Configure `preempt` on the router you want to be primary. Priority chooses the active router; **preempt** is what makes that choice survive a reboot. The pair is incomplete without it. Preempt has a second reputation, though. Switch it on without thinking and a router that is still settling after a reboot grabs the gateway role straight back, and the pair starts trading it: [what preempt is really promising](https://www.pinglabz.com/hsrp-troubleshooting-flapping-dual-active/) is worth reading alongside this, because it is not simply "the highest priority wins". ## Setting 3: interface tracking - fail over for the right reason HSRP, on its own, only fails over if the active router stops sending hellos - it is watching the LAN side. It does not know if the active router's *upstream* link is down. Without tracking, R1 can keep happily serving as the gateway while its WAN uplink is dead, black-holing every packet. Object tracking fixes this. You track the uplink, and HSRP decrements the router's priority when that uplink goes down: ``` ! R1 track 1 interface GigabitEthernet0/0 line-protocol ! interface Vlan10 standby 10 priority 110 standby 10 preempt standby 10 track 1 decrement 20 ``` If Gi0/0 fails, R1's priority drops from 110 to 90\. R2 at 100 now outranks it, and (with preempt) takes active - failing over because the path to the rest of the network actually broke, not just the router. The decrement must be large enough to push the priority below the peer's: 110 minus 20 is 90, below R2's 100\. Get the arithmetic right. ## Setting 4: version and timers Use **HSRP version 2**. Version 1 supports only group numbers 0-255 and uses a coarser timer; version 2 supports groups up to 4095, has a distinct virtual MAC range, and supports millisecond timers and IPv6\. Both ends of a group must run the same version. Default timers are a 3-second hello and a 10-second hold, so failover takes up to 10 seconds. For faster failover, tighten them - but keep both routers identical: ``` interface Vlan10 standby 10 timers msec 250 msec 800 ``` ## Verifying ``` R1# show standby brief P indicates configured to preempt. | Interface Grp Pri P State Active Standby Virtual IP Vl10 10 110 P Active local 10.20.0.3 10.20.0.1 ``` Confirm three things: **State** is Active on the router you intended, the **P** flag is present (preempt configured), and the **Virtual IP** matches what hosts use as their gateway. For the full picture including tracking, use `show standby`. ## Common gotchas The "primary" router stays standby after a reboot No `preempt`. Priority alone does not reclaim the active role. Both routers think they are active They are not exchanging hellos - group number mismatch, version mismatch, or a Layer 2 problem between them. Gateway keeps working but traffic black-holes No interface tracking. The active router's uplink is down but HSRP does not know. Tracking configured but failover never happens The decrement is too small to drop priority below the peer's. Recheck the arithmetic. Hosts lose the gateway intermittently Timer mismatch between the two routers, or HSRP version mismatch. Make both identical. ## Key takeaways HSRP configuration is `standby ip ` on each router, and that alone gives a working virtual gateway. But the production-grade setup needs four more decisions: set `priority` explicitly to choose the active router rather than letting IP addresses decide; add `preempt` so that choice survives a reboot - the single most-forgotten command; configure interface `track` with a decrement large enough to fail over when the uplink dies, not just when the router dies; and run HSRP version 2 with matching timers on both ends. Verify with `show standby brief` and confirm the State, the P flag, and the virtual IP. For the FHRP cluster, see the [FHRP pillar](https://www.pinglabz.com/fhrp/). ### IPv6 Transition Mechanisms: Which One, and When URL: https://www.pinglabz.com/ipv6-transition-mechanisms/ Last updated: 2026-07-13T14:54:45.000Z There is no single "IPv6 migration button." Instead there are three transition mechanisms, and the skill is knowing which one to reach for in a given situation. Dual-stack, tunneling, and translation each solve a different shape of problem, and picking the wrong one turns a clean migration into a fragile mess. This article is the decision guide that ties together the hands-on work from the rest of the series into one framework you can apply on real networks. It is part of the PingLabz [IPv6 routing and services guide](https://www.pinglabz.com/ipv6/) and connects to the broader [IP routing fundamentals](https://www.pinglabz.com/ip-routing/). ## Three mechanisms, three different problems Before choosing, be clear about what each mechanism actually does. **Dual-stack** runs IPv4 and IPv6 in parallel on the same devices and links, so every host and router speaks both protocols natively. **Tunneling** wraps IPv6 packets inside IPv4 (or the reverse) so that traffic of one protocol can cross a network that only forwards the other. **Translation** (NAT64 with DNS64) converts between the two protocols at a gateway so a host that speaks only one can talk to a host that speaks only the other. Those are three genuinely different operations: running both, carrying one inside the other, and converting one into the other. Match the operation to your constraint and the decision makes itself. ## Dual-stack: the preferred end state Dual-stack is the goal, not a bridge. When a device is dual-stacked, it has an IPv4 address and an IPv6 address on its interfaces, participates in both an IPv4 and an IPv6 routing table, and simply uses whichever protocol a given destination needs. Nothing is encapsulated, nothing is translated, and no application breaks, because both protocols are genuinely present end to end. Applications increasingly prefer IPv6 when it is available (the "happy eyeballs" behaviour in modern clients) and fall back to IPv4 seamlessly. Every router in the lab for this series is dual-stacked, which is what let us run IPv6 IGPs, build tunnels, and configure NAT64 all on the same devices. That is the point: dual-stack is the foundation you build the other mechanisms on top of. The only real costs are operational: you run two protocols, so you maintain two sets of routing, two sets of ACLs, and two sets of security policy, and you must not let the IPv6 side lag behind the IPv4 side in hardening (an under-secured IPv6 stack on a dual-stack device is a classic blind spot). Choose dual-stack whenever you control both ends and can enable IPv6 everywhere along the path. It is the destination; the other two mechanisms exist only to get you there or to cover the gaps you cannot yet close. ## Tunnels: carrying IPv6 islands over an IPv4 core Sometimes you have IPv6 running at the edges but a core in the middle that is IPv4-only and cannot be upgraded yet (a provider network, a legacy WAN, a segment behind someone else's change freeze). Tunneling bridges those IPv6 islands by encapsulating their traffic in IPv4 so the core forwards it as ordinary IPv4, oblivious to the IPv6 payload inside. Both endpoints still speak native IPv6, so from the application's perspective nothing changed; only the transport across the gap differs. We proved a manual IPv6-over-IPv4 tunnel in the [tunnels article](https://www.pinglabz.com/ipv6-over-ipv4-tunnels/), carrying real IPv6 pings across an IPv4-only path at 100 percent success. The practical guidance from that work carries into the decision here: use a point-to-point tunnel (manual or GRE) when you have a known pair of endpoints, prefer GRE when you want to run an IPv6 IGP across the tunnel because GRE carries the multicast that OSPFv3 and EIGRP for IPv6 need, and avoid 6to4, which is automatic but deprecated. Remember the golden rule of tunnels: the underlay must be able to reach both endpoint addresses, and an up/up tunnel that will not pass traffic is almost always an underlay routing or MTU problem, not an IPv6 one. Tunnels are a bridge, deliberately temporary, that you keep until you can dual-stack the core and take them out. ## Translation: an IPv6-only side reaching IPv4-only content The hardest case is when one side genuinely does not speak the other protocol at all. An IPv6-only client network needs to reach an IPv4-only server, and there is no shared protocol to dual-stack and no way to tunnel because the client has no IPv4 stack to encapsulate. This is where translation, NAT64 paired with DNS64, is the answer. NAT64 rewrites IPv6 packets as IPv4 (and back) at a gateway, and DNS64 synthesises the IPv6 address the client uses to trigger that translation. We walked the full mechanism in the [NAT64 article](https://www.pinglabz.com/nat64-explained/), including the synthesized-address math and the honest note that the data-plane translation did not complete on the control-plane simulator we used. For the decision here, the key point is positioning: translation is a last resort, not a design goal. It is stateful, it introduces a choke point, and it quietly breaks applications that embed IP literals in their payloads. Reach for NAT64 and DNS64 specifically when you have an IPv6-only client network that must reach the legacy IPv4 internet or IPv4-only internal systems, and you cannot make those systems dual-stack. In every other situation, one of the first two mechanisms is a better fit. ## The decision tree Put the three together and the choice reduces to a short set of questions about what you control. ### Dual-stack Use whenYou control both ends and can enable v6 everywhere RoleThe goal / end state Breaks appsNo CostRun two protocols ### Tunnel Use whenv6 islands over a v4-only core you cannot upgrade RoleA bridge Breaks appsNo (both ends native v6) CostOverhead, MTU, underlay ### Translation (NAT64+DNS64) Use whenIPv6-only client to IPv4-only server RoleLast resort Breaks appsSome (IP literals) CostStateful choke point Read the tree as a sequence of questions. Can you control both ends and enable IPv6 along the whole path? Then dual-stack, and you are done. Do you have IPv6 islands separated by an IPv4-only core you cannot change? Then tunnel across it (GRE if you want dynamic routing, manual for a simple static pair). Is one side genuinely IPv6-only and the other genuinely IPv4-only, with no way to add the missing protocol? Then, and only then, translate with NAT64 and DNS64\. Dual-stack is the goal; tunnels and translation are the bridges you use until you get there. ## How these combine in a real migration In practice you rarely pick just one. A realistic migration dual-stacks everything you can reach, tunnels IPv6 across the parts of the core you cannot yet upgrade, and stands up NAT64 for the pockets that are IPv6-only but still need IPv4 content. Over time the tunnels come out as the core gets dual-stacked, and the NAT64 footprint shrinks as more content becomes reachable over native IPv6\. The direction of travel is always toward dual-stack (and eventually IPv6-only with less and less need for translation), with tunnels and translation as scaffolding you remove as the building takes shape. Thinking of them as scaffolding rather than as permanent fixtures keeps your design honest and keeps you from leaving fragile translation gateways in the path longer than necessary. ## Security across the transition Each mechanism carries its own security consideration, and the transition period is exactly when blind spots appear. On dual-stack, the biggest risk is an unhardened IPv6 stack: teams pour years into IPv4 ACLs and firewall rules and then bring up IPv6 with permissive or missing policy, handing attackers a wide-open second door into the same hosts. Treat the IPv6 control and data plane as a first-class citizen, matching every IPv4 protection with an IPv6 equivalent. On tunnels, remember that encapsulation can smuggle IPv6 straight through IPv4 inspection points that do not understand protocol 41 or GRE; protect any tunnel that crosses untrusted ground with IPsec, and filter rogue tunnels at the edge so users cannot build their own unsanctioned paths out of your network. On translation, the NAT64 gateway is a stateful choke point and therefore a denial-of-service target: its state table can be exhausted, so size it and rate-limit accordingly, and log translations for the same accountability reasons you log IPv4 NAT. ## Common mistakes when choosing a mechanism The most frequent error is reaching for translation when dual-stack was available. NAT64 is genuinely necessary only when one side cannot run the other protocol at all; if you could have enabled IPv6 on both ends, translation just added a fragile, app-breaking choke point for no reason. The second common mistake is treating tunnels as permanent. A tunnel is scaffolding for an IPv4-only core you have not upgraded yet, and every tunnel you leave in place is overhead, an MTU liability, and one more thing to troubleshoot; plan its removal from the day you build it. The third mistake is forgetting DNS: NAT64 without DNS64 does not work in practice because clients have no way to learn the synthesized address, and dual-stack deployments frequently trip over DNS resolvers that return IPv6 records before the IPv6 path is actually ready, causing timeouts that look like network faults. Get DNS behaviour right and half of your transition surprises disappear. Match the mechanism to the constraint, keep dual-stack as the destination, and treat everything else as temporary, and the migration stays under control. ## Key Takeaways - There are three IPv6 transition mechanisms, each for a different problem: dual-stack (run both), tunneling (carry one inside the other), translation (convert one into the other). - **Dual-stack** is the preferred end state. Use it wherever you control both ends and can enable IPv6 along the whole path; nothing breaks because both protocols are genuinely present. - **Tunnels** carry IPv6 islands across an IPv4-only core you cannot upgrade. Prefer GRE for dynamic routing, manual for a static pair, and avoid deprecated 6to4\. Both ends still speak native IPv6. - **Translation (NAT64 + DNS64)** is the last resort, for an IPv6-only client reaching IPv4-only content. It is stateful and breaks some applications, so use it only when you cannot dual-stack or tunnel. - The decision tree: control both ends and can enable v6, dual-stack; v6 islands over a v4 core, tunnel; IPv6-only to IPv4-only, NAT64 and DNS64\. Dual-stack is the goal, the others are bridges you remove over time. This closes the IPv6 routing and translation series. Revisit the deep-dives on [IPv6-over-IPv4 tunnels](https://www.pinglabz.com/ipv6-over-ipv4-tunnels/) and [NAT64](https://www.pinglabz.com/nat64-explained/), or head back to the [IPv6 guide](https://www.pinglabz.com/ipv6/) for the full cluster. ### NAT64 Explained: Stateful, Stateless, and the DNS64 Half of the Story URL: https://www.pinglabz.com/nat64-explained/ Last updated: 2026-07-13T14:54:45.000Z Dual-stack and tunnels both assume IPv6 exists on both ends of a conversation. NAT64 solves the harder case: an IPv6-only client that must reach a server which only speaks IPv4\. It does this by translating between the two protocols and, cleverly, by encoding the IPv4 server's address inside an IPv6 address so the client never has to know it is talking to a legacy host. This article explains how NAT64 works, walks the full stateful configuration on Cisco IOS XE 17.18, does the address-synthesis math by hand, and explains the DNS64 half that makes the whole thing usable. It also gives you an honest platform finding from the lab. It is part of the PingLabz [IPv6 routing and services guide](https://www.pinglabz.com/ipv6/). ## The problem NAT64 exists to solve Imagine you have finally built an IPv6-only client network, no IPv4 addresses to hand out, no dual-stack to maintain. That is the clean end state everyone wants. The catch is that large parts of the internet, and plenty of internal legacy systems, still run IPv4-only. Your IPv6-only client has no way to put a packet on the wire toward `10.20.20.80`, because it has no IPv4 stack and no IPv4 address. NAT64 sits at the boundary and translates: it takes IPv6 packets from the client, rewrites them as IPv4 packets toward the real server, and translates the replies back. The client stays purely IPv6; the server stays purely IPv4; the NAT64 gateway does the shapeshifting in the middle. ## The synthesized address: the whole trick NAT64's central idea is that an IPv4 address can be embedded inside an IPv6 address. You reserve an IPv6 `/96` prefix for NAT64, and the last 32 bits of that prefix hold the IPv4 address of the target server. On the lab the NAT64 prefix is `2001:DB8:64::/96`, and the IPv4-only server is `10.20.20.80`. Convert each octet of the IPv4 address to hex: ``` 10 . 20 . 20 . 80 0x0a 0x14 0x14 0x50 -> 0x0a141450 ``` Now append those 32 bits to the `/96` prefix. The IPv6-only client reaches the IPv4 server at: ``` 2001:db8:64::0a14:1450 (written by IOS as 2001:DB8:64::A14:1450) ``` That single address is the entire trick of NAT64\. The client sends a normal IPv6 packet to `2001:db8:64::0a14:1450`. The NAT64 gateway recognises the packet is destined for its own `/96` prefix, pulls the last 32 bits out (`0x0a141450`), reconstructs the IPv4 address `10.20.20.80`, and forwards an IPv4 packet to the real server. To the client it was pure IPv6; to the server it was pure IPv4\. The `/96` length is deliberate: 128 minus 96 is exactly 32 bits, precisely enough to hold one IPv4 address. ## The full stateful configuration Here is the complete stateful NAT64 configuration from EDGE1 on the lab. Ethernet0/0 faces the IPv6-only client side, Ethernet0/1 faces the IPv4 side: ``` nat64 prefix stateful 2001:DB8:64::/96 nat64 v4 pool PLZ-POOL 203.0.113.1 203.0.113.10 ipv6 access-list PLZ-V6-CLIENTS permit ipv6 2001:DB8:99:1::/64 any nat64 v6v4 list PLZ-V6-CLIENTS pool PLZ-POOL overload interface Ethernet0/0 nat64 enable interface Ethernet0/1 nat64 enable ``` Each piece has a job. `nat64 prefix stateful 2001:DB8:64::/96` defines the translation prefix, the one whose last 32 bits carry the embedded IPv4 address. `nat64 v4 pool PLZ-POOL 203.0.113.1 203.0.113.10` creates a pool of real IPv4 addresses that the gateway will use as the *source* of the translated packets (the server needs an IPv4 source to reply to). The `ipv6 access-list` selects which IPv6 clients are allowed to be translated, here the `2001:DB8:99:1::/64` client subnet. `nat64 v6v4 list ... pool ... overload` ties it together: translate IPv6-to-IPv4 for traffic matching the ACL, using the pool, with `overload` meaning PAT (many clients share the pool via port translation, exactly like IPv4 NAT overload). Finally, `nat64 enable` must be applied on both the ingress and egress interfaces, translation only happens on interfaces where it is switched on. ## The infrastructure comes up correctly With that configuration applied, the control plane builds everything it should. First, the NAT64 prefix gets a route, owned by the NAT64 virtual interface: ``` EDGE1#show ipv6 route 2001:DB8:64::A14:1450 Routing entry for 2001:DB8:64::/96 Known via "static", ... ::100.0.0.1, NVI0 ``` Look at what this confirms. A lookup for the synthesized address `2001:DB8:64::A14:1450` resolves to the `2001:DB8:64::/96` route, and that route points at `NVI0`, the NAT64 Virtual Interface. NVI0 is the internal pseudo-interface that owns the NAT64 prefix and hands matching packets to the translation engine. This is exactly the routing state you want: any IPv6 packet aimed at the prefix is steered into NAT64 rather than dropped. The prefix, the pool, the ACL, and this NVI0 route are all correctly in place. ## An honest platform finding: the data plane did not translate Now the honest part, because credibility matters more than a tidy story. With the IPv6-only host sourcing traffic correctly from `2001:db8:99:1::100` toward the synthesized address, the actual IPv6-to-IPv4 translation did **not** complete on this platform. The statistics stayed empty: ``` EDGE1#show nat64 statistics global Total active translations: 0 ... Packets translated (IPv6 -> IPv4) Stateful: 0 ``` Zero active translations, zero packets translated. This is not a configuration error, the config, the pool, the ACL, the prefix, and the NVI0 route are all present and correct. It is a platform limitation. The lab runs on IOL-XE (IOS on Linux), which is a control-plane-focused simulation: it faithfully accepts configuration and shows control-plane state, but several data-plane features do not fully forward in it. On this same platform, uRPF drop counters and WRED behave the same way, they configure and report state but do not exercise the full forwarding path. NAT64 translation falls into that category here. So the responsible thing is to be clear: we proved the configuration and the routing infrastructure, and we did the address-synthesis math, but we did not capture a translated packet, and we are not going to pretend we did. A completed end-to-end NAT64 translation is best demonstrated on a full data-plane platform such as the Catalyst 8000V or physical hardware. Everything you can verify without forwarding, we verified; the one thing that needs a real data plane, we are honest about. ## The DNS64 half of the story NAT64 by itself has a usability gap: how does the client know to send its packet to `2001:db8:64::0a14:1450` in the first place? A user types a hostname, and DNS is supposed to return an address. The IPv4-only server has an `A` record (`10.20.20.80`) but no `AAAA` record, so an IPv6-only client doing a normal `AAAA` lookup gets nothing back. This is where **DNS64** completes the design. DNS64 is a special resolver behaviour. When an IPv6-only client asks for the `AAAA` record of an IPv4-only host, the DNS64 resolver notices there is no real `AAAA`, so it looks up the `A` record instead, takes the returned IPv4 address, and *synthesises* an `AAAA` record on the fly by embedding that IPv4 address into the NAT64 `/96` prefix. For our server it would synthesise exactly `2001:db8:64::0a14:1450` and hand it back. The client just did an ordinary DNS lookup and received an ordinary-looking IPv6 address; it has no idea any translation is involved. It sends its packet to that address, the packet hits the NAT64 prefix, and the gateway translates. DNS64 and NAT64 are two halves of one mechanism: DNS64 manufactures the synthesized address at lookup time, NAT64 acts on it at forwarding time. Deploy one without the other and the design does not work in practice, because clients would have no way to discover the synthesized address on their own. ## Stateful vs stateless NAT64 NAT64 comes in two flavours, and the distinction matters for design. ### Stateful NAT64 MappingMany IPv6 to few IPv4 IPv4 poolYes (PAT/overload) Keeps stateYes (per flow) Needs DNS64Yes Best forMany clients, few IPv4 addresses ### Stateless NAT64 Mapping1:1 IPv6 to IPv4 IPv4 poolNo Keeps stateNo ConstraintIPv4-translatable addresses only Best forPredictable 1:1, servers **Stateful** NAT64 (the configuration above) maps many IPv6 clients onto a small pool of IPv4 addresses using port translation, exactly like classic IPv4 NAT overload. It keeps per-flow state and conserves IPv4 addresses, which is the realistic case for an IPv6-only client network reaching the IPv4 internet. It requires DNS64 to hand clients the synthesized address. **Stateless** NAT64 maps addresses one-to-one with no pool and no state table, deriving each IPv4 address algorithmically from the IPv6 address. It scales beautifully because there is no state to track, but it only works with IPv6 addresses that are built to be IPv4-translatable, and it does not conserve IPv4 space (every translated host needs its own IPv4 address). Stateless suits a bank of servers with a fixed, predictable mapping; stateful suits a crowd of clients sharing scarce IPv4. ## Verifying NAT64, in the order that finds problems fastest When you deploy NAT64 and it does not work, check the pieces in dependency order rather than poking at random. First, confirm the prefix route exists and points at NVI0 with `show ipv6 route` on the synthesized address, as in the capture above; if that route is missing, the NAT64 prefix was never installed and nothing downstream can work. Second, confirm `nat64 enable` is on both the IPv6-facing and IPv4-facing interfaces, forgetting one side is the most common configuration mistake. Third, confirm the ACL actually matches your client source; a client outside the permitted subnet is silently ignored. Fourth, confirm the pool has free addresses and that the IPv4 side can route back to the pool range (the server replies to a pool address, so that address must be reachable from the server's perspective). Only after all four are green do you look at `show nat64 statistics` and `show nat64 translations` to watch flows build. If the infrastructure is correct and translations still read zero, and you are on a control-plane simulator, you have hit the same platform limit this lab did, move to hardware or a data-plane-capable image to see live translation. ## What NAT64 breaks, and why it is a last resort NAT64 works well for the common case of a client fetching content over HTTP, HTTPS, or other protocols that carry no embedded IP addresses. It struggles with anything that puts an IP literal inside the payload. Protocols that embed addresses in their application data (classic FTP in active mode, some SIP and peer-to-peer signalling, and any application that hands a bare IPv4 literal to the user) do not get those embedded addresses translated, because NAT64 only rewrites the packet headers, not arbitrary application payloads. An IPv6-only client handed an IPv4 literal by such an application has no way to reach it, since there is no synthesized address for a literal that never went through DNS64\. This is the fundamental reason NAT64 is positioned as a last resort rather than a design goal: it is stateful, it is a translation choke point, and it quietly breaks a meaningful minority of applications. Dual-stack breaks nothing because both protocols are genuinely present; NAT64 is what you reach for only when you cannot dual-stack the client side and still must reach IPv4-only content. ## Where NAT64 fits among the transition tools It helps to place NAT64 next to its neighbors. Dual-stack runs IPv4 and IPv6 side by side and is the preferred end state; it has no translation and no tunnel, so nothing breaks, but it requires you to enable IPv6 everywhere and keep running IPv4 too. Tunnels carry IPv6 islands across an IPv4-only core you cannot upgrade; both ends still speak IPv6, so applications are unaffected, but you have taken on encapsulation overhead and MTU concerns. NAT64 is different in kind from both: it is the only one of the three that lets a host which does not speak IPv4 at all reach a host which does not speak IPv6 at all. That unique capability is also its cost, because bridging two protocols that were never designed to interoperate is inherently lossy. Use dual-stack where you can, tunnels where you must cross a hostile core, and NAT64 only where one side is genuinely single-protocol and cannot be changed. ## Key Takeaways - NAT64 lets an IPv6-only client reach an IPv4-only server by translating between the protocols at a gateway. The client stays pure IPv6, the server stays pure IPv4. - The synthesized address is the core concept: a NAT64 `/96` prefix with the server's IPv4 embedded in the last 32 bits. `10.20.20.80` becomes `0x0a141450`, so the client targets `2001:db8:64::a14:1450`. - DNS64 is the other half: it synthesises exactly that `AAAA` record when an IPv6-only client looks up an IPv4-only host, so the client discovers the synthesized address with a normal DNS query. NAT64 without DNS64 is not usable in practice. - Stateful NAT64 uses an IPv4 pool with overload for many-to-few mappings and needs DNS64; stateless NAT64 is 1:1 with no pool, no state, but only for IPv4-translatable addresses. - **Honest lab note:** the config, pool, ACL, prefix, and NVI0 route were all correct, but the data-plane translation did not complete on IOL-XE 17.18 (0 translations), consistent with it being a control-plane simulation. Demonstrate a completed NAT64 on Catalyst 8000V or hardware. NAT64 is the last resort when one side is IPv6-only and the other is IPv4-only. The next article ties dual-stack, tunnels, and translation into a single decision framework. For the full picture, return to the [IPv6 guide](https://www.pinglabz.com/ipv6/). ### IPv6 over IPv4 Tunnels: Manual, GRE, and 6to4 Compared URL: https://www.pinglabz.com/ipv6-over-ipv4-tunnels/ Last updated: 2026-07-13T14:54:44.000Z Sometimes you have IPv6 islands that need to talk to each other, but the network in between is IPv4-only and you cannot change it (a provider core, a legacy WAN, a merged network you do not fully control). The answer is to wrap IPv6 packets inside IPv4 and carry them across the gap. This article configures a manual IPv6-over-IPv4 tunnel on real Cisco IOS XE 17.18, proves it with a real ping across an IPv4-only core, and compares the three tunnel types you need to know: manual, GRE, and 6to4\. It is part of the PingLabz [IPv6 routing and services guide](https://www.pinglabz.com/ipv6/) and builds on the [GRE tunneling cluster](https://www.pinglabz.com/gre/). ## The idea: IPv6 as the payload, IPv4 as the transport Every tunnel is encapsulation. An IPv6-over-IPv4 tunnel takes a complete IPv6 packet and puts an IPv4 header in front of it. To the IPv4 core, the packet is just IPv4 traffic between two tunnel endpoints; the IPv6 payload is opaque and irrelevant. At the far end, the router strips the IPv4 header and hands the original IPv6 packet to its IPv6 stack. The two tunnel endpoints behave as if they share a direct IPv6 link, even though many IPv4 hops separate them. The crucial mental split is between the **underlay** (the IPv4 network doing the actual forwarding) and the **overlay** (the IPv6 adjacency riding on top). The overlay only works if the underlay can deliver packets between the two endpoint addresses. Hold on to that distinction, because it is the source of the single most common tunnel failure, covered at the end. ## Configuring a manual tunnel A manual (also called configured or static) tunnel is point-to-point: you explicitly set both a source and a destination, so it carries traffic between exactly two known endpoints. Here is the working configuration from EDGE1 on the lab, tunneling to BR1 across the IPv4 core: ``` interface Tunnel0 ipv6 address 2001:DB8:99:AAAA::1/64 tunnel source Loopback0 tunnel destination 10.255.0.3 tunnel mode ipv6ip ``` Four lines do the work. The `ipv6 address` gives the tunnel interface an IPv6 identity, so the two ends form an IPv6 subnet (here `2001:DB8:99:AAAA::/64`). `tunnel source Loopback0` and `tunnel destination 10.255.0.3` define the IPv4 endpoints of the underlay: the local loopback and the remote router's loopback. Using loopbacks rather than physical interfaces is deliberate, a loopback stays up as long as any path to the router exists, so the tunnel does not flap when a single physical link goes down. The line that sets the encapsulation is `tunnel mode ipv6ip`: this is native IPv6-in-IPv4, which uses IP protocol number 41 with no additional tunnel header. It is the leanest possible encapsulation, and it is what "manual IPv6 tunnel" means on IOS. ## What the tunnel interface actually shows Once the underlay can reach both endpoints, the tunnel comes up. Here is the real interface state: ``` EDGE1#show interface Tunnel0 Tunnel0 is up, line protocol is up Tunnel source 10.255.0.1 (Loopback0), destination 10.255.0.3 Tunnel protocol/transport IPv6/IP Tunnel transport MTU 1480 bytes ``` Read the two lines that matter. `Tunnel protocol/transport IPv6/IP` confirms the tunnel is carrying IPv6 as its protocol directly inside IPv4 as its transport, which is exactly what `tunnel mode ipv6ip` requests (protocol 41, no GRE). And `Tunnel transport MTU 1480 bytes` tells you the overhead: the underlay uses a standard 1500-byte MTU, the IPv4 header consumes 20 bytes, and the remaining 1480 bytes are available to the IPv6 payload. There is no GRE header here, so you lose only the 20 bytes of the IPv4 header. Remember that 1480 figure; it changes with GRE. ## Proving it: IPv6 across an IPv4-only path The whole point of the tunnel is that IPv6 reachability now exists where the underlay speaks only IPv4\. The proof is a real ping from EDGE1 to BR1's tunnel address: ``` EDGE1#ping 2001:DB8:99:AAAA::2 repeat 4 !!!! Success rate is 100 percent (4/4), round-trip min/avg/max = 3/3/4 ms ``` Four exclamation marks, 100 percent success. Those IPv6 echo requests were encapsulated in IPv4 at EDGE1, forwarded hop by hop across the IPv4 core (which never saw an IPv6 packet), decapsulated at BR1, answered, and returned the same way. From the IPv6 point of view, EDGE1 and BR1 are directly connected neighbors on the `2001:DB8:99:AAAA::/64` subnet. From the core's point of view, this was ordinary protocol-41 IPv4 traffic between two loopbacks. ## The three tunnel types compared Manual is one of three encapsulation styles you should recognise. They differ by exactly one configuration line and by what they can carry. ### Manual (ipv6ip) Mode`tunnel mode ipv6ip` TopologyPoint-to-point EncapsulationProtocol 41, no header MTU1480 MulticastNo StatusProven here ### GRE (gre ip) Mode`tunnel mode gre ip` TopologyPoint-to-point EncapsulationGRE over IPv4 MTU1476 (4 bytes more) MulticastYes (runs an IGP) StatusOne-line change ### 6to4 (ipv6ip 6to4) Mode`tunnel mode ipv6ip 6to4` TopologyAutomatic / multipoint Address2002::/16 from source IP DestinationNone (derived) MulticastNo StatusDeprecated The practical differences: **manual** is the lean default when you have a known pair of endpoints and only need unicast IPv6\. **GRE** costs four extra bytes of overhead (MTU drops from 1480 to 1476) but it can carry multicast, which means you can run an IPv6 IGP such as OSPFv3 or EIGRP for IPv6 over the tunnel and let it exchange routes dynamically. That single capability is why GRE is often the better choice for anything beyond a static pair. **6to4** is automatic and multipoint: it derives a 2002::/16 IPv6 prefix from the tunnel's IPv4 source address and needs no explicit destination, which sounds convenient but comes with real operational problems around relay reliability. It is deprecated and being retired, so treat it as exam knowledge and legacy documentation, not a design you would deploy today. ## The gotcha that will waste your afternoon: underlay reachability Here is the failure mode that catches everyone. A tunnel interface can show `up/up` and still not pass a single packet. The reason is the underlay. Because the tunnel is encapsulated in IPv4 and sent between the two endpoint addresses, those two addresses must be **mutually reachable in the IPv4 routing table**. On the lab, the tunnel stayed down until CORE1 had IPv4 routes to both loopbacks; without underlay reachability, there is nowhere for the encapsulated packets to go. What makes this deceptive is that the line protocol on a manual tunnel can come up based on the local source being valid, independent of whether the remote endpoint is actually reachable. So you see `up/up`, assume the tunnel is healthy, and then the ping fails. When that happens, do not stare at the IPv6 configuration. Check two things in order. First, can the local endpoint ping the remote endpoint over IPv4 (`ping 10.255.0.3` from EDGE1)? If not, it is an underlay routing problem: fix the IPv4 route to the far loopback. Second, if IPv4 reachability is fine but large packets or real applications fail while small pings succeed, suspect MTU and fragmentation: the encapsulation ate into the payload budget (1480 for manual, 1476 for GRE) and something in the path is not handling the reduced MTU or is dropping the fragments. The rule to remember: an up/up tunnel that will not pass traffic is almost always an underlay routing or MTU problem, not an overlay one. ## Why GRE wins when you need dynamic routing The single most useful practical difference between manual and GRE tunnels is multicast, and the reason it matters is dynamic routing protocols. OSPFv3 and EIGRP for IPv6 both rely on multicast to discover neighbors and exchange updates. A manual `ipv6ip` tunnel carries unicast only, so you cannot run an IGP across it; you are limited to static IPv6 routes pointing at the tunnel. That is fine for a single pair of endpoints with a handful of prefixes, but it does not scale and it does not react to failures. Switch the same tunnel to `tunnel mode gre ip` and multicast rides across, so you can enable OSPFv3 or EIGRP for IPv6 on the tunnel interface and let the two sides learn each other's IPv6 prefixes dynamically. Now, if a downstream link changes, the IGP reconverges automatically instead of waiting for you to edit static routes. The cost is four bytes of overhead and a slightly lower MTU (1476 instead of 1480). For anything beyond a static point-to-point pair, that trade is almost always worth it, which is why GRE, not manual, is the workhorse for IPv6 island interconnect in production. ## MTU and fragmentation, in a bit more detail Tunnels shrink the usable payload, and that has real consequences for applications. With a 1500-byte physical MTU, a manual tunnel leaves 1480 bytes for the IPv6 packet and GRE leaves 1476\. A host that still tries to send 1500-byte IPv6 packets will produce frames that do not fit once encapsulated. IPv6 does not allow routers to fragment in transit (unlike IPv4), so the tunnel source must either fragment the IPv4-encapsulated packet or rely on Path MTU Discovery to tell the origin host to send smaller packets. When PMTUD is broken (often because an overly aggressive firewall drops the ICMPv6 "packet too big" messages), you get the classic symptom: small pings succeed, but web pages hang and large transfers stall. The fix is usually to lower the TCP MSS on the tunnel interface with `ipv6 tcp adjust-mss 1440` (or similar), which clamps TCP segment sizes so they fit inside the tunnel without relying on PMTUD. Keep this in mind whenever a freshly built tunnel pings clean but real traffic misbehaves. ## A security note before you deploy A tunnel is a hole through your perimeter by design: it carries IPv6 traffic that your IPv4 firewalls may not inspect. Because native `ipv6ip` and GRE tunnels are unauthenticated and unencrypted, anyone who can reach the tunnel destination over IPv4 and spoof the source could, in principle, inject traffic. For a tunnel that crosses any untrusted network, protect the underlay with IPsec (for example, a GRE-over-IPsec design) so the encapsulated IPv6 is both authenticated and encrypted. On a fully private, trusted core, a plain tunnel is acceptable, but make the decision consciously rather than by default. ## Key Takeaways - IPv6-over-IPv4 tunnels encapsulate whole IPv6 packets inside IPv4 so IPv6 islands can cross an IPv4-only core. Keep the underlay (IPv4 transport) and overlay (IPv6 adjacency) distinct in your head. - A manual tunnel uses `tunnel mode ipv6ip` (protocol 41, no extra header, MTU 1480) and is point-to-point between two explicit endpoints. It carries unicast only, no multicast. - GRE (`tunnel mode gre ip`) is one line different, costs 4 bytes (MTU 1476), and crucially carries multicast, so you can run an IPv6 IGP over it. 6to4 (`tunnel mode ipv6ip 6to4`) is automatic and multipoint but deprecated. - `show interface Tunnel0` confirms the encapsulation (`Tunnel protocol/transport IPv6/IP`) and the transport MTU. A real repeat ping proved 100 percent IPv6 delivery across an IPv4-only path. - An `up/up` tunnel that will not ping is almost always an underlay problem: the two endpoint IPv4 addresses must be mutually reachable in the IPv4 table, and MTU/fragmentation is the second thing to check. Tunnels bridge IPv6 islands over an IPv4 core. When the problem is instead an IPv6-only client reaching IPv4-only content, you need translation, which is the NAT64 article next. For the full picture, return to the [IPv6 guide](https://www.pinglabz.com/ipv6/). ### IS-IS for IPv6: Multi-Topology and the Single-Topology Trap URL: https://www.pinglabz.com/is-is-for-ipv6/ Last updated: 2026-07-13T14:54:44.000Z IS-IS is the quietest way to break IPv6 in your network. It will happily form an adjacency, show every neighbor as up, and route IPv4 perfectly, while dropping your IPv6 traffic into a black hole on any link where IPv6 is not fully enabled. The reason is a default almost nobody thinks about: by default IS-IS runs a **single topology**, and IPv6 is forced to follow the IPv4 shortest-path tree whether or not IPv6 is actually usable along that path. This article shows the fix, multi-topology IS-IS, on real Cisco IOS XE 17.18, with the actual adjacency and the IPv6-only prefix it learns. It is part of the PingLabz [IPv6 routing and services guide](https://www.pinglabz.com/ipv6/). ## How IS-IS carries IPv6 in the first place IS-IS was designed to be protocol-agnostic. It runs directly over the data link (not inside IP), and it advertises reachability using TLVs (Type-Length-Value structures) inside its link-state PDUs. Adding a new network-layer protocol means defining new TLVs, not redesigning the protocol. IPv6 support arrived exactly this way: two new TLVs, one for IPv6 reachability and one for the IPv6 interface address, plus a Network Layer Protocol ID (NLPID) that advertises "I speak IPv6" to neighbors. Because IS-IS carries IPv4 and IPv6 in the same link-state database by default, the naive expectation is that turning on IPv6 "just works." And in a lab where every link is dual-stacked, it appears to. The problem only surfaces when your topology is not uniform, which in the real world it almost always is. ## The single-topology default: the trap Here is the mechanism that catches people. In its default mode, IS-IS computes **one** shortest-path-first (SPF) tree, based on the IPv4 metrics, and then uses that same tree to forward both IPv4 and IPv6\. There is one topology, one SPF result, and IPv6 is a passenger. Now picture a network where one link in the path is IPv4-only (perhaps an older segment nobody has migrated, or a link where IPv6 was simply never configured). IS-IS still considers that link part of the single topology because it carries IPv4\. The single SPF calculation may well choose a path that crosses that IPv4-only link because it is the shortest for IPv4\. IPv6 packets are then handed to a next hop that cannot forward them, because IPv6 is not enabled on that link. The result is a black hole: the routing table looks healthy, the adjacencies are up, and IPv6 traffic silently dies. Nothing in `show isis neighbors` hints at the problem, because the adjacency itself is perfectly fine. This is the single-topology trap, and it is exactly the kind of failure that eats an afternoon before someone thinks to question the topology mode. ## Multi-topology: giving IPv6 its own SPF The fix is **multi-topology IS-IS** (often abbreviated MT). With multi-topology enabled, IS-IS maintains and computes a *separate* topology and a *separate* SPF for IPv6\. Each link can advertise its own IPv6 metric, and the IPv6 SPF only considers links that are actually IPv6-capable. An IPv4-only link simply does not appear in the IPv6 topology, so the IPv6 SPF routes around it. IPv4 and IPv6 now make independent forwarding decisions, which is what you want the instant your two address families do not overlay perfectly. This is the same idea OSPF solves by running two completely separate processes (OSPFv2 for IPv4, OSPFv3 for IPv6). IS-IS keeps everything in one protocol instance and one adjacency, and simply computes two SPFs. That efficiency is a big part of why large service providers favour IS-IS: one protocol, one set of adjacencies, correct independent routing for both address families. ## The configuration Here is the working multi-topology configuration from CORE1 facing BR1 on the lab, IOS XE 17.18: ``` router isis PLZ net 49.0001.0000.0000.0002.00 metric-style wide address-family ipv6 multi-topology interface Ethernet0/2 ipv6 router isis PLZ isis network point-to-point ``` Walk through it line by line. The `net` (Network Entity Title) is the router's IS-IS identity: area `49.0001`, system ID `0000.0000.0002`, and the NSAP selector `00`. `metric-style wide` switches IS-IS to wide metrics (more on why that is mandatory below). Under `address-family ipv6`, the `multi-topology` command is the whole point of the exercise: it tells IS-IS to compute a dedicated IPv6 SPF. On the interface, `ipv6 router isis PLZ` enables IPv6 IS-IS on that link, and `isis network point-to-point` tells IS-IS to treat the link as a point-to-point circuit, which skips the DIS (Designated Intermediate System) election you do not need on a link with only two routers. ## The adjacency: one neighbor, both levels With both ends configured, the adjacency forms. Note the level: ``` CORE1#show isis neighbors Tag PLZ: System Id Type Interface State Holdtime Circuit Id BR1 L1L2 Et0/2 UP 27 02 ``` The neighbor BR1 is up as type `L1L2`, meaning the router participates in both Level 1 (intra-area) and Level 2 (inter-area) routing, which is the IOS default until you constrain it. The state is `UP` and the holdtime counts down normally. This is the deceptive part: this output looks identical whether or not multi-topology is enabled and whether or not IPv6 is actually forwarding. The adjacency is a link-layer relationship; it tells you nothing about whether IPv6 SPF is correct. That is precisely why the single-topology trap is so easy to miss, the neighbor table lies by omission. ## The IPv6-only prefix, learned correctly The proof that IPv6 is genuinely routing lives in the IPv6 routing table. On the lab, BR1 hosts an IPv6-only branch LAN, and CORE1 learns it through IS-IS: ``` CORE1#show ipv6 route isis I1 2001:DB8:99:66::/64 [115/20] I1 2001:DB8:99:255::3/128 [115/20] ``` The `I1` code is IS-IS Level 1\. The prefix `2001:DB8:99:66::/64` is the IPv6-only branch LAN, learned across the point-to-point link and installed with the standard IS-IS administrative distance of 115 and a metric of 20\. The second entry is BR1's loopback. Both arrived through the IPv6 topology, which is exactly what multi-topology gives you. In a single-topology network with a non-uniform path, those same prefixes could be present in the table yet point at a next hop that cannot forward IPv6, which is why "the route is in the table" is not by itself proof that IPv6 works end to end. ## metric-style wide is not optional This is the requirement people skip and then spend an hour debugging. IS-IS multi-topology **does not work with narrow metrics**. The original IS-IS metric style uses a 6-bit interface metric (values 0 to 63) and cannot carry the extended TLVs that multi-topology depends on. Multi-topology information is transported inside the wide-metric TLVs, so `metric-style wide` is a hard prerequisite, not a tuning preference. If you configure `multi-topology` under the IPv6 address family but leave the router on narrow metrics, MT simply will not engage, and you are quietly back in the single-topology trap you were trying to escape. Wide metrics are what you want regardless of IPv6\. Narrow metrics cap a single interface at 63 and a full path at 1023, which is far too coarse for modern networks. Wide metrics give a 24-bit interface metric and a 32-bit path metric, so link costs actually reflect bandwidth differences. Enable `metric-style wide` on every IS-IS router in the domain, consistently, before you turn on multi-topology. ## A note on point-to-point circuits The `isis network point-to-point` command on the interface is a small but worthwhile optimisation. On a link with exactly two routers, the default broadcast behaviour still elects a DIS and generates pseudonode LSPs, which is pure overhead when there can never be a third router on the segment. Declaring the circuit point-to-point skips the DIS election, reduces LSP flooding, and speeds up adjacency formation. Use it on any genuine point-to-point link (including routed Ethernet between two devices), and leave the default on true multi-access segments. ## Migrating an existing IS-IS domain to multi-topology You rarely deploy multi-topology on a greenfield network; you usually retrofit it onto a running IS-IS domain that started life IPv4-only. The order of operations matters. Do not flip one router to multi-topology while its neighbors are still single-topology, because a mismatch in topology capability between neighbors can drop the IPv6 routes you are trying to protect. IOS provides a `transition` keyword (`multi-topology transition`) precisely for this: it advertises both the old single-topology and the new multi-topology TLVs at the same time, so routers that have been converted and routers that have not can coexist while you work through the domain. Once every router is running multi-topology, you remove the `transition` keyword and the network is fully on MT. The safe sequence is: enable `metric-style wide` everywhere first, then add `multi-topology transition` everywhere, then remove `transition` once the whole domain is converted. Plan the metric-style change carefully in a live network. Because narrow and wide metrics are advertised differently, mixing them across a domain during the change can cause temporary suboptimal routing. Many operators schedule the `metric-style wide` rollout as its own maintenance step, confirm the IPv4 topology is stable on wide metrics, and only then begin the multi-topology work. Rushing both changes together is how a routine IPv6 enablement turns into an outage. ## Verifying that multi-topology actually engaged Because the neighbor table cannot tell you whether MT is working, lean on the IPv6-specific and topology-specific show commands. `show ipv6 route isis` confirms the IPv6 prefixes are present with the expected `I1` or `I2` codes and sane metrics, as in the capture above. `show isis ipv6 topology` displays the dedicated IPv6 SPF result, which is the direct evidence that IS-IS is computing a separate tree for IPv6 rather than reusing the IPv4 one. And `show isis database detail` lets you inspect the TLVs in the link-state PDUs, where you can confirm the multi-topology reachability TLVs are actually present. If those TLVs are missing, the usual cause is that `metric-style wide` was never applied, or was applied on some routers but not others. The final and least deniable test is always a real IPv6 ping or traceroute end to end, because a black hole caused by an IPv4-only link only reveals itself in the data plane, never in the adjacency table. ## Single-topology vs multi-topology at a glance ### Single-topology (default) SPF treesOne (IPv4-based) IPv6 pathFollows the IPv4 SPF IPv4-only link in pathBlackholes IPv6 Safe whenEvery link is dual-stacked ### Multi-topology (MT) SPF treesTwo (independent v4 and v6) IPv6 pathOwn IPv6 SPF IPv4-only link in pathRouted around Requires`metric-style wide` ## Key Takeaways - IS-IS carries IPv6 with dedicated TLVs inside its existing link-state PDUs, so one protocol instance and one adjacency handle both address families. - By default IS-IS runs a **single topology**: it computes one IPv4-based SPF and forces IPv6 to follow it. An IPv4-only link in the chosen path black-holes IPv6 while every adjacency still shows up. - `multi-topology` under `address-family ipv6` gives IPv6 its own independent SPF, so it routes around links where IPv6 is not enabled. - `metric-style wide` is a **hard requirement** for multi-topology. Narrow metrics cannot carry MT information, and MT silently fails to engage without it. - The adjacency table (`show isis neighbors` showing `L1L2 UP`) tells you nothing about IPv6 correctness. Confirm IPv6 with `show ipv6 route isis` and by testing real end-to-end reachability. Multi-topology fixes IPv6 where you run IS-IS everywhere. When you cannot enable IPv6 on the core at all, you tunnel over it instead, which is the next article. For the full picture of IPv6 routing and services, return to the [IPv6 guide](https://www.pinglabz.com/ipv6/). ### EIGRP for IPv6: Classic and Named Mode on IOS XE URL: https://www.pinglabz.com/eigrp-for-ipv6/ Last updated: 2026-07-13T14:54:43.000Z EIGRP did not just bolt IPv6 support onto the existing process. On Cisco IOS XE, EIGRP for IPv6 is a separate protocol instance with its own topology table, its own neighbors, and a configuration model that trips up engineers who expect it to look like IPv4 EIGRP. There is no `network` statement, the next hop is always a link-local address, and (depending on your IOS version) the process may or may not come up shut down. This article walks the classic interface-based configuration on real IOS XE 17.18, shows the actual adjacency and routes, and clears up the one gotcha that has confused CCNP and CCIE candidates for years. It is part of the PingLabz [IPv6 routing and services guide](https://www.pinglabz.com/ipv6/), and it builds directly on the [EIGRP cluster](https://www.pinglabz.com/eigrp/). ## Why EIGRP for IPv6 is its own protocol EIGRP for IPv4 and EIGRP for IPv6 share the same DUAL (Diffusing Update Algorithm) core, the same composite metric, and the same convergence behaviour. What differs is the plumbing. IPv6 does not use the `network` command to select which interfaces run the protocol. Instead, you enable EIGRP directly on each interface. That single design decision changes how you read a router's configuration: to know which interfaces are participating, you look at the interfaces themselves, not at a list of network statements under the routing process. The other structural difference is the next hop. IPv6 interior gateway protocols always form adjacencies and install routes using link-local addresses (the `FE80::/10` range that every IPv6 interface auto-generates). You will never see a global unicast next hop for an EIGRP-learned IPv6 route. This is not a quirk, it is by design: link-local addresses are guaranteed to exist and are guaranteed to be reachable on the local link, which is exactly what a next hop needs to be. ## Classic mode: the interface-based configuration Classic mode is the original configuration style and the one you still meet most often in production and in older documentation. Here is the working configuration from EDGE1 facing CORE1 on the lab, IOS XE 17.18.02: ``` ipv6 router eigrp 100 interface Ethernet0/1 ipv6 eigrp 100 interface Loopback0 ipv6 eigrp 100 ``` Read that carefully. The `ipv6 router eigrp 100` line creates the process for autonomous system 100\. But nothing under that process tells it which interfaces to use. The activation happens per interface with `ipv6 eigrp 100`. Enable it on the interfaces you want to advertise and form adjacencies over, and leave it off everywhere else. There is no `network 2001:db8::/64` equivalent, and there is no wildcard mask to fuss over. If an interface is not carrying EIGRP, it is because you did not put `ipv6 eigrp 100` on it. ## The adjacency and the routes: what you actually see Once both ends are enabled, the neighbor comes up. Notice the address the router lists for its neighbor: ``` EDGE1#show ipv6 eigrp neighbors EIGRP-IPv6 Neighbors for AS(100) H Address Interface Hold Uptime SRTT RTO Q Seq 0 Link-local address: Et0/1 14 00:00:17 1 100 0 3 FE80::A8BB:CCFF:FE00:7710 ``` The neighbor is identified by its link-local address `FE80::A8BB:CCFF:FE00:7710`, not by any global address you configured. That is the rule for every IPv6 IGP, and it is the first thing to internalise. Now look at the route EIGRP installs: ``` EDGE1#show ipv6 route eigrp D 2001:DB8:99:255::2/128 [90/409600] via FE80::A8BB:CCFF:FE00:7710, Ethernet0/1 ``` The learned prefix is a global unicast address (CORE1's loopback), but the `via` next hop is again the link-local address out Ethernet0/1\. The administrative distance is 90 (internal EIGRP, unchanged from IPv4) and the composite metric is 409600\. The `D` code is EIGRP, exactly as it is for IPv4\. Nothing about DUAL changed; only the address family did. The topology table confirms the feasible successors and feasible distances: ``` EDGE1#show ipv6 eigrp topology P 2001:DB8:99:255::1/128, 1 successors, FD is 128256 P 2001:DB8:99:12::/64, 1 successors, FD is 281600 P 2001:DB8:99:255::2/128, 1 successors, FD is 409600 ``` Each prefix is in the passive (`P`) state with one successor and a feasible distance. This is the same topology-table structure you already know from IPv4 EIGRP, which is the whole point: if you understand DUAL, you already understand the routing logic here. Only the enablement and next-hop conventions are new. ## The honest gotcha: does the process come up shut down? Here is the piece of folklore that has cost people hours. The classic teaching, repeated in study guides for years, is that `ipv6 router eigrp` comes up administratively **shut down** and that you must issue `no shutdown` under the process before any adjacency will form. For a long time on classic IOS that was genuinely true, and forgetting the `no shutdown` was a rite of passage. On this lab, running IOS XE 17.18, that behaviour did **not** reproduce. The adjacency came up immediately after enabling the interfaces, with no `no shutdown` required under the process. So the honest guidance is: do not assume either way. The shutdown-by-default behaviour was real on older IOS and may still appear on some platforms and versions, so verify it on the exact code you are running. And if a classic EIGRP-for-IPv6 adjacency stubbornly refuses to form, checking for a shut-down process is still the very first thing to rule out: ``` EDGE1#show ipv6 protocols ``` If `show ipv6 protocols` shows the EIGRP process is shut down, add `no shutdown` under `ipv6 router eigrp 100` and the neighbors will spring to life. If it shows the process active and you still have no neighbor, move on to the usual suspects: mismatched AS numbers, an interface that is missing the `ipv6 eigrp` command, an ACL blocking the multicast, or a mismatched authentication key. The lesson is not that the gotcha is fake; it is that it is version-dependent, so you confirm rather than assume. ## Named mode: the modern configuration Classic mode still works and still appears everywhere, but Cisco's modern recommendation is **named mode** EIGRP, which places every address family under a single named process. The IPv6 configuration lives inside an address-family block: ``` router eigrp PLZ address-family ipv6 unicast autonomous-system 100 af-interface Ethernet0/1 exit-af-interface exit-address-family ``` Named mode gives you two concrete advantages. First, it consolidates IPv4 and IPv6 configuration (plus any VRFs) under one `router eigrp NAME` hierarchy, so the whole EIGRP posture of the router lives in one place instead of scattered across `ipv6 router eigrp` and interface commands. Second, named mode enables **wide metrics** by default. Classic EIGRP scales its metric to a 32-bit value, which starts to lose resolution above roughly a gigabit of interface bandwidth. Wide metrics use a 64-bit metric so that 10G, 40G, and 100G links are distinguished properly instead of all flattening to the same value. On any modern high-speed network, that resolution matters for correct path selection. For a deeper look at how the composite metric and DUAL feasibility work, see the wider [EIGRP cluster](https://www.pinglabz.com/eigrp/), which covers the IPv4 mechanics that carry over unchanged into IPv6. ## What carries over unchanged from IPv4 EIGRP Once you get past the enablement and next-hop conventions, almost everything you already know about EIGRP applies without modification. The hello and hold timers default to the same values (5 seconds hello, 15 seconds hold on high-bandwidth broadcast and point-to-point links). Feasibility and the feasible-distance test that prevent routing loops are identical. Summarisation, stub routing, and the split-horizon rules all behave the same way, you just configure them in the IPv6 context. Authentication is still available, and on any production EIGRP-for-IPv6 deployment you should turn it on so that a rogue device on a segment cannot inject an adjacency. The composite metric still weighs bandwidth and delay by default (with reliability, load, and MTU available but unused unless you change the K-values). Because delay is part of the metric and delay is configured per interface in tens of microseconds, the same `delay` knob you use to influence IPv4 EIGRP path selection influences IPv6 path selection too, as long as both address families run over the same interface. The mental model to carry away is simple: EIGRP for IPv6 is EIGRP, with an IPv6 address family and a per-interface enablement model bolted on in place of the `network` command. ## Verifying and troubleshooting Three commands answer most questions. `show ipv6 eigrp neighbors` tells you whether the adjacency exists and over which interface. `show ipv6 route eigrp` tells you which prefixes made it into the routing table and confirms the link-local next hop. `show ipv6 protocols` tells you the process is running (and, critically, whether it is shut down) along with the interfaces and any redistribution. When a neighbor will not form, work that list top to bottom: process shut down, AS mismatch, missing `ipv6 eigrp` on an interface, an ACL dropping the EIGRP multicast, an authentication mismatch, or an MTU mismatch on the link. The order matters because the cheapest checks (process state and AS number) catch the most common mistakes. ## Classic vs named at a glance ### Classic mode Process`ipv6 router eigrp 100` Enable per interface`ipv6 eigrp 100` Wide metricsNo (32-bit) Best forLegacy configs you inherit ### Named mode Process`router eigrp PLZ` Enable`address-family ipv6 unicast autonomous-system 100` Wide metricsYes (64-bit) Best forNew builds, high-speed links, VRFs ## Key Takeaways - EIGRP for IPv6 is a separate protocol instance that reuses DUAL, the composite metric, and administrative distance 90 from IPv4 EIGRP. - There is no `network` statement. You enable classic mode per interface with `ipv6 eigrp `, so interfaces, not network commands, tell you what is participating. - The next hop is **always** a link-local `FE80::` address, both in `show ipv6 eigrp neighbors` and in `show ipv6 route eigrp`. Learned prefixes are global, next hops are link-local. - The classic "process comes up shut down, needs `no shutdown`" gotcha was real on older IOS but did not reproduce on IOS XE 17.18\. Verify on your version, and check `show ipv6 protocols` for a shut-down process as your first troubleshooting step. - Named mode (`router eigrp NAME` with an `address-family ipv6` block) is the modern configuration and enables 64-bit wide metrics, which you want on any high-speed network. Next in the series: how IS-IS handles IPv6 and the single-topology trap that quietly blackholes it. For the full picture of IPv6 routing and services, return to the [IPv6 guide](https://www.pinglabz.com/ipv6/). ### The Layer 2 Attack Surface: A Hardening Checklist That Actually Holds URL: https://www.pinglabz.com/layer-2-attack-surface-hardening/ Last updated: 2026-07-13T14:21:09.000Z Plug a laptop into a switchport in most networks and, by default, it is trusted completely. It can flood the CAM table, hand out DHCP leases, poison ARP caches, spoof its source IP, talk to every other host in the VLAN, and even elect itself the spanning-tree root. The default Layer 2 posture is wide open, because Ethernet was built for a world where physical access meant you were allowed. Every control in this checklist closes one specific door in that open house. This is the roundup for the PingLabz [VLANs and Layer 2 switching](https://www.pinglabz.com/vlans-layer-2-switching/) cluster and the [network infrastructure security](https://www.pinglabz.com/infrastructure-security/) guide: each control mapped to the exact attack it stops, and, wherever we could, a way to prove it on a real switch. Several of the proofs below were captured live on Cisco IOL-XE in CML. ## Why the default is dangerous A switch does its job by trusting what it hears. It learns MAC addresses from whatever frames arrive, it forwards DHCP from anyone, it caches ARP from any reply, and it runs spanning tree with any device that speaks BPDUs. None of that involves a check on whether the sender should be doing it. That is not a flaw so much as the absence of a feature, and the features that add the checks are exactly the hardening controls below. The goal of L2 hardening is to replace blanket trust with specific, verifiable trust: this port may have this MAC, that server is the only DHCP source, this host owns this IP-and-MAC pair, these ports may not talk to each other. Assemble the full set and the wide-open port becomes a port that only does what you intended. ## The control-to-attack map Each card below names a control, the attack it defeats, and how you would prove it works. "Proven here" means the drop was captured live in the CML lab behind this cluster; the rest link to the deep-dive that walks the configuration and verification. Port security Stops: CAM table overflow / MAC flooding How: Limit the number of MACs learned per port; violation shuts or restricts the port Prove it: Flood MACs past the limit, watch the port err-disable DHCP snooping Stops: Rogue DHCP server / DHCP starvation How: Trust only the real server's port; drop server-side DHCP from untrusted ports; build the bindings table Prove it: Offer a lease from an untrusted port, watch it get dropped DAI + static ARP ACL Stops: ARP poisoning / man-in-the-middle How: Validate every ARP against snooping bindings, or against a static IP-to-MAC ACL where there is no DHCP Proven here: 8 spoofed ARPs dropped IP Source Guard Stops: IP spoofing from a host How: Bind IP + MAC + port from the snooping table; drop traffic whose source does not match Prove it: Spoof a source IP, watch the frame get filtered at the port VLAN ACL (VACL) Stops: Lateral movement inside a subnet How: Filter frames as they are bridged inside the VLAN, where a router ACL never sees them Proven here: intra-VLAN ping 0% to 100% loss Private VLANs Stops: Host-to-host pivot in a shared segment How: Isolated secondary VLAN: members reach only the promiscuous uplink, never each other Config proven here: isolated + community association up BPDU Guard / Root Guard Stops: Rogue switch / STP root takeover How: BPDU Guard err-disables an access port that receives a BPDU; Root Guard blocks a port that tries to become root Prove it: Send a BPDU into a guarded port, watch it err-disable Storm control Stops: Broadcast / multicast / unknown-unicast storms How: Rate-limit each traffic class per port; drop or shut the port above threshold Prove it: Generate a broadcast flood, watch the port throttle or err-disable ## Port security stops CAM overflow A switch learns MAC addresses into its CAM table by watching source addresses. Fill that table with thousands of fake MACs and the switch runs out of room and fails open, flooding unknown-destination frames out every port - which is exactly what an attacker wants, because it turns the switch into a hub they can sniff. Port security caps how many MACs a port may learn and reacts (protect, restrict, or shutdown) when the cap is exceeded. It is the first control every access port should have. The mechanics of the table it protects are covered in [how switches learn MAC addresses](https://www.pinglabz.com/how-switches-learn-mac-addresses/), and the broader per-port hardening it belongs to is in [VLAN security hardening](https://www.pinglabz.com/vlan-security-best-practices/). ## DHCP snooping stops rogue DHCP Any host on a VLAN can answer a DHCP request. A rogue server can hand a victim a bad default gateway (itself) and become a man-in-the-middle, or exhaust the real server's pool (starvation) so nothing gets an address. DHCP snooping designates the real server's port as trusted and drops server-side DHCP messages (OFFER, ACK) arriving on any untrusted port. As a bonus, it builds the IP-to-MAC-to-port bindings table that DAI and IP Source Guard depend on. The full mechanism, including Option 82 and trusted ports, is in [DHCP snooping in depth](https://www.pinglabz.com/dhcp-snooping-in-depth/). ## DAI and static ARP ACLs stop ARP poisoning (proven) ARP has no authentication, so any host can claim any IP. Dynamic ARP Inspection validates each ARP against the snooping bindings, and where there is no DHCP, a static ARP ACL supplies the legitimate IP-to-MAC pairs by hand. This is the one attack we ran end to end. HOST2 fired eight forged ARP replies claiming HOST1's IP from HOST2's MAC, and DAI dropped all eight: ``` SW1#show ip arp inspection statistics vlan 120 Vlan Forwarded Dropped DHCP Drops ACL Drops ---- --------- ------- ---------- --------- 120 2 8 8 0 ``` Eight spoofed ARPs in, eight dropped, the man-in-the-middle dead on arrival. The full walkthrough, including the `static` keyword and the 15 pps rate-limit, is in [static ARP ACLs: inspecting ARP without DHCP snooping](https://www.pinglabz.com/static-arp-acl-inspection/). For the DHCP-backed version alongside IP Source Guard, see [Dynamic ARP Inspection and IP Source Guard](https://www.pinglabz.com/dynamic-arp-inspection-ip-source-guard/). ## IP Source Guard stops IP spoofing Even with ARP under control, a host can still forge its source IP to impersonate another machine or dodge an ACL. IP Source Guard uses the snooping bindings to pin each port to a specific IP-and-MAC and drops any frame whose source does not match. It is the natural partner to DHCP snooping and DAI, closing the source-address door those two leave ajar. The configuration and verification live in the same [L2 security stack article](https://www.pinglabz.com/dynamic-arp-inspection-ip-source-guard/). ## VACLs stop lateral movement (proven) A router ACL only sees traffic that routes between subnets, so it can never filter two hosts in the same VLAN. When an attacker compromises one host and pivots sideways to its neighbours, that traffic never touches the router. A VACL filters frames as they are bridged inside the VLAN, so it can. We proved it: the same intra-VLAN ping went from 0% loss to 100% loss the moment the VACL was applied, while the gateway ping still worked, showing it was the VACL and not a broken host. The before-and-after and the access-map config are in [VLAN ACLs (VACLs): filtering traffic inside a VLAN](https://www.pinglabz.com/vlan-acls-vacl-configuration/). ## Private VLANs stop the host-to-host pivot Where a VACL filters specific flows, a private VLAN isolates hosts wholesale. Put every server in an isolated secondary VLAN and each one reaches the gateway but none can reach its neighbours, so a compromised server cannot pivot across the segment - all without a subnet per host. The isolated-plus-community association was built and confirmed on IOL-XE in this cluster. The full model, config, and the DMZ use case are in [private VLANs: isolating hosts that share a subnet](https://www.pinglabz.com/private-vlans-explained/). ## BPDU Guard and Root Guard stop STP takeover Spanning tree trusts any device that speaks BPDUs, so a rogue switch (or a laptop running STP) can advertise a superior bridge ID, win the root election, and drag traffic through itself. BPDU Guard err-disables an access port the instant it receives any BPDU, on the logic that end hosts should never send them. Root Guard is gentler: it lets a port participate in STP but blocks it if it ever tries to become root, protecting your root placement on links to other switches you do not fully control. Both belong on the edge, and both are covered in the [Spanning Tree Protocol complete guide](https://www.pinglabz.com/spanning-tree-protocol/). The related switch-spoofing and double-tagging tricks are in [VLAN hopping attacks explained](https://www.pinglabz.com/vlan-hopping-attacks/). ## Storm control stops the flood A broadcast, multicast, or unknown-unicast storm (from a loop, a misbehaving NIC, or a deliberate flood) can saturate a segment and take the whole VLAN down. Storm control rate-limits each of those traffic classes per port and drops the excess or err-disables the port once it crosses a threshold, containing the storm at the point it enters. The thresholds and configuration are in [storm control: stopping broadcast floods at the port](https://www.pinglabz.com/storm-control-configuration/). ## The order to turn them on Sequence matters, because some controls depend on others. DHCP snooping comes first on any VLAN with DHCP clients, since its bindings table is the foundation that both DAI and IP Source Guard read from - enable them before snooping has learned anything and you will drop legitimate traffic. Port security, BPDU Guard, and storm control are independent and can go on every access port from day one; they need no shared state. DAI and IP Source Guard come next, once snooping is populated (or, for static segments, once your static ARP ACL is written). VACLs and private VLANs are design decisions layered on top, applied where the topology calls for intra-subnet isolation. A sane rollout is therefore: harden every access port with port security, BPDU Guard, and storm control; enable DHCP snooping and let it learn; turn up DAI and IP Source Guard against those bindings; then add VACLs or private VLANs where servers share a subnet. ## Testing each control safely in a lab The reason this cluster leans on CML captures is that every one of these controls is easy to misconfigure in a way that looks fine until it is attacked. The fix is to test the failure, not just the config. For each control, generate the exact attack it is meant to stop and confirm the drop: flood MACs past the port-security limit and watch the err-disable; offer a DHCP lease from an untrusted port and watch snooping discard it; fire forged ARPs with a tool like `nping` and read `show ip arp inspection statistics`; ping a same-VLAN neighbour before and after a VACL; send a BPDU into a guarded port and watch it shut. A control you have only ever seen in the running-config is a control you are trusting on faith. A control whose drop you have watched in show output is one you know works. That is the whole philosophy behind the proofs in this cluster, and it is why the [static ARP ACL](https://www.pinglabz.com/static-arp-acl-inspection/) and [VACL](https://www.pinglabz.com/vlan-acls-vacl-configuration/) articles show the attack and the counter, not just the syntax. ## Building the layered posture No single control is the answer; the point is that each one closes a different door, and an attacker only needs one door left open. A hardened access port typically carries port security, DHCP snooping (with DAI and IP Source Guard riding on its bindings), BPDU Guard, and storm control all at once, with VACLs or private VLANs adding intra-subnet isolation where the design calls for it. Turn them on together and the wide-open default becomes a port that trusts only what you have explicitly allowed. Work through the deep-dives linked above, and use the [VLANs and Layer 2 switching](https://www.pinglabz.com/vlans-layer-2-switching/) cluster and the [network infrastructure security](https://www.pinglabz.com/infrastructure-security/) guide as your map for the rest. ## Key Takeaways - The default L2 posture is wide open: a plugged-in port trusts every MAC, DHCP offer, ARP reply, source IP, and BPDU it hears. Each control replaces blanket trust with a specific, verifiable check. - Port security stops CAM overflow, DHCP snooping stops rogue DHCP, DAI plus static ARP ACLs stop ARP poisoning, IP Source Guard stops IP spoofing, VACLs stop lateral movement, private VLANs stop the host-to-host pivot, BPDU/Root Guard stop STP takeover, and storm control stops floods. - Three defenses were proven live in this cluster: eight spoofed ARPs dropped by DAI, an intra-VLAN ping driven from 0% to 100% loss by a VACL, and a private VLAN isolated/community association confirmed up on IOL-XE. - DHCP snooping is foundational: its bindings table feeds both DAI and IP Source Guard, so deploy it first on any VLAN with DHCP clients. - Layer the controls. A hardened access port runs port security, snooping-backed DAI and IP Source Guard, BPDU Guard, and storm control together, with VACLs or private VLANs for intra-subnet isolation. - Work through the linked deep-dives and use the [VLANs and Layer 2 switching](https://www.pinglabz.com/vlans-layer-2-switching/) cluster and the [network infrastructure security](https://www.pinglabz.com/infrastructure-security/) guide as the hub for the full picture. ### Private VLANs: Isolating Hosts That Share a Subnet URL: https://www.pinglabz.com/private-vlans-explained/ Last updated: 2026-07-13T14:21:08.000Z Imagine a hosting segment or a DMZ where twenty servers share one subnet. They all need to reach the gateway, but they have no business talking to each other. The obvious fix is a subnet per server, but that burns address space, multiplies SVIs, and turns a simple rack into a routing project. Private VLANs (PVLANs) solve the same problem the elegant way: they isolate hosts that share a subnet, at Layer 2, without giving each host its own network. This article is part of the PingLabz [VLANs and Layer 2 switching](https://www.pinglabz.com/vlans-layer-2-switching/) cluster and the wider [network infrastructure security](https://www.pinglabz.com/infrastructure-security/) guide, and every configuration and show command below was captured live on a Cisco IOL-XE switch in CML. ## The private VLAN model: primary and secondary A private VLAN is not one VLAN, it is a small family of them working together. There is one **primary** VLAN and one or more **secondary** VLANs associated with it. Every host lives in a secondary VLAN, but they all share the primary VLAN's subnet and gateway. What changes is the rules for who can talk to whom, and that depends on which type of secondary VLAN a port sits in. Primary VLAN Owns the subnet and the promiscuous uplink. Carries traffic downstream to every secondary. This is where the gateway (router or firewall) connects. Isolated secondary Members cannot talk to each other at all, only to the promiscuous uplink. Total host-to-host isolation inside one shared subnet. Community secondary Members can talk to each other and to the uplink, but not to other communities or to isolated ports. A trusted cluster inside the subnet. Two port roles carry this. A **host** port belongs to a secondary VLAN and follows that VLAN's isolation rules. A **promiscuous** port belongs to the primary VLAN and can talk to every secondary, which is exactly what the gateway needs so it can route for all of them. The mental model: isolated ports are strangers on a train (they share the carriage but do not speak), community ports are a table of colleagues (they speak among themselves), and the promiscuous port is the conductor everyone can flag down. ## Configuring the private VLAN PVLANs require the switch to be in VTP transparent mode, because VTP versions 1 and 2 do not propagate private VLAN information and will not let you create them otherwise. After that, you define the primary, define each secondary with its type, and associate the secondaries with the primary: ``` vtp mode transparent vlan 500 private-vlan primary vlan 501 private-vlan isolated vlan 502 private-vlan community vlan 500 private-vlan association 501,502 ``` VLAN 500 is the primary. VLAN 501 is an isolated secondary and VLAN 502 is a community secondary. The final block associates both secondaries with the primary, which is the step that ties the family together. Without the association, you just have three unrelated VLANs. The association is real and up, which you can confirm with a single command: ``` SW1#show vlan private-vlan Primary Secondary Type Ports ------- --------- ----------------- ------------------------------------------ 500 501 isolated 500 502 community ``` This output is the heart of the feature. VLAN 500 now owns two secondaries: 501 as isolated and 502 as community. Any host port you drop into 501 gets full isolation; any host port in 502 gets community behaviour. The switch enforces the rules; you do not write a single access-list. ## Assigning host and promiscuous ports With the VLANs built, you place ports. A host port needs two things: the mode, and the host-association that says which primary and secondary it belongs to. A promiscuous port needs its mode and a mapping that lists every secondary it should reach: ``` interface Ethernet0/1 switchport mode private-vlan host switchport private-vlan host-association 500 501 ! HOST1 into the ISOLATED secondary interface Ethernet0/0 switchport mode private-vlan promiscuous switchport private-vlan mapping 500 501,502 ! the gateway uplink talks to all ``` Et0/1 is a host port associated with primary 500 and isolated secondary 501, so the host behind it is fully isolated from other isolated members. Et0/0 is the promiscuous uplink to the gateway, mapped to both 501 and 502, so the router can reach every host regardless of which secondary it lives in. That asymmetry is the entire point: hosts cannot reach each other, but the gateway can reach all of them, and they can all reach the gateway. The port genuinely takes the mode, which you can verify on the interface: ``` SW1#show interface Ethernet0/1 switchport Administrative Mode: private-vlan host Operational Mode: private-vlan host ``` Administrative and operational mode both read `private-vlan host`, so the port is not just configured for the role, it is actively operating in it. ## The isolation rule, spelled out Here is the behaviour the model produces, which is the whole reason to deploy PVLANs. Two hosts sitting in the isolated secondary (VLAN 501) share the same subnet and the same default gateway. They are, by every appearance on the host side, neighbours on one network. Yet they cannot reach each other at Layer 2: every frame the switch would normally bridge between two isolated ports is dropped. At the same time, both hosts still reach the promiscuous uplink, so both can route out through the gateway and receive traffic back. You get host-to-host isolation and full gateway reachability at once, inside a single subnet. Community ports soften this by one degree. Put two servers in community VLAN 502 and they can talk to each other and to the uplink, but not to isolated ports and not to a different community. That is the right tool when a small set of servers legitimately need to cluster (say, two database replicas) while remaining walled off from the rest of the segment. An honest note on this lab: the private VLAN configuration and association were proven end to end on IOL-XE (the config was accepted, the association shows up, and the host port operates in private-vlan host mode, all captured above). A full host-to-host isolation ping between two isolated members was not run in the capture window, so this article describes the isolation behaviour the platform enforces rather than claiming a packet capture we did not take. The design behaviour is well defined and matches the configuration state shown here. ## Walking a frame through the private VLAN It helps to follow individual frames to see why the isolation is airtight. Say HOST1 (isolated, Et0/1) wants to reach HOST2 (also isolated) in the same secondary. HOST1 sees HOST2 as a local subnet address, ARPs for it, and would normally bridge a frame straight across. But the switch knows both ports are in isolated secondary 501, and the rule for isolated ports is simple: frames may only go to a promiscuous port. There is no promiscuous port in that path, so the frame is dropped in the switch fabric. HOST2 never hears the ARP, let alone the data. Now say HOST1 wants to reach the gateway. The gateway sits on the promiscuous port Et0/0, which is mapped to secondary 501\. Isolated-to-promiscuous is explicitly allowed, so the frame is forwarded, the router routes it, and the reply comes back down the promiscuous port to HOST1\. From HOST1 the network feels completely normal: DNS, default route, internet access all work. It simply cannot see its neighbours. That selective forwarding, allow to the uplink and deny to peers, is enforced per frame by the switch with no access-list involved, which is why PVLANs scale to hundreds of ports without a maintenance burden. ## Isolated or community: choosing the secondary type Most secure segments default to isolated, because the safest posture is that no host trusts any other host. Reach for a community secondary only when a specific group of hosts has a legitimate reason to talk directly and you want to keep that conversation off the router. Two clustered database nodes replicating to each other, or a pair of load-balanced app servers exchanging health checks, are good candidates: put them in one community and they see each other, but they still cannot reach the isolated servers or a second community. Everything else stays isolated. A practical design rule: start every port isolated, and promote only the exact ports that prove they need community membership. It is far easier to open a door later than to discover, after a breach, that a community you created two years ago let an attacker walk from one server to five others. ## The classic use case: DMZ and hosting isolation PVLANs shine anywhere many hosts must share a subnet but must not trust each other. The textbook example is a DMZ or a multi-tenant hosting segment. Put every server in an isolated secondary and each one can reach the firewall (the promiscuous uplink) but none can reach its neighbours. If one server is compromised, the attacker's usual next move, pivoting sideways to the servers next to it, is dead on arrival, because the switch simply will not bridge frames between isolated ports. You get that protection without a subnet per server, without a wall of SVIs, and without maintaining a large VACL by hand. Subnet per server Isolates hosts, but burns address space, multiplies SVIs and routing, and does not scale to a dense rack. Heavy. VACL in a shared VLAN Precise and protocol-aware, but you write and maintain access-lists as hosts change. Best for selective flows. Private VLAN One subnet, switch-enforced isolation, no per-flow rules. Scales to hundreds of ports. Best for broad host-to-host isolation. PVLANs and VACLs are complementary, not competing. A VACL, covered in the [VLAN ACLs article](https://www.pinglabz.com/vlan-acls-vacl-configuration/), gives you surgical, protocol-aware filtering between named hosts. A private VLAN gives you effortless, all-or-nothing isolation across an entire segment. Hardened designs often use both: the PVLAN for broad isolation, a VACL for the handful of specific flows that still need to be permitted. ## Gotchas worth knowing - **VTP transparent is mandatory.** If the switch is a VTP client or server on v1/v2, you cannot create private VLANs. Set `vtp mode transparent` first. - **The promiscuous mapping must list every secondary.** Forget to map a secondary and the hosts in it lose their gateway, which looks like a total outage for those hosts even though isolation is working perfectly. - **Trunking PVLANs between switches needs care.** Extending private VLANs across a trunk requires the secondaries to be carried and associated consistently on both ends, or the isolation breaks at the boundary. - **Isolation is Layer 2, not Layer 3.** A router on the promiscuous port can still route traffic back down to any host, so if two isolated hosts are in different subnets reachable via that router, they could reach each other through it. PVLAN isolation stops the bridged path, not a deliberately routed one. For how private VLANs fit alongside every other switch-hardening control, keep reading the [VLANs and Layer 2 switching](https://www.pinglabz.com/vlans-layer-2-switching/) cluster and the [network infrastructure security](https://www.pinglabz.com/infrastructure-security/) guide. ## Key Takeaways - Private VLANs isolate hosts that share a subnet, at Layer 2, without a subnet per host. A primary VLAN owns the subnet; secondary VLANs set the isolation rules. - Isolated secondary members cannot talk to each other, only to the promiscuous uplink. Community secondary members can talk to each other and the uplink, but not to other communities or isolated ports. - Host ports carry the isolation; the promiscuous port (the gateway uplink) can reach every secondary, which is why hosts always keep their gateway. - Configuration requires VTP transparent mode, then `private-vlan primary/isolated/community` and an `association`. In the lab, `show vlan private-vlan` confirmed 500/501 isolated and 500/502 community, and the host port operated in `private-vlan host` mode. - The classic use is a DMZ or hosting segment: servers reach the gateway but cannot pivot to each other if one is compromised, with no subnet-per-server and no large VACL to maintain. - PVLANs and VACLs complement each other. Use a private VLAN for broad isolation and a [VACL](https://www.pinglabz.com/vlan-acls-vacl-configuration/) for selective flows. See the full [VLANs and Layer 2 switching](https://www.pinglabz.com/vlans-layer-2-switching/) cluster. ### VLAN ACLs (VACLs): Filtering Traffic Inside a VLAN URL: https://www.pinglabz.com/vlan-acls-vacl-configuration/ Last updated: 2026-07-13T14:21:08.000Z Here is a question that trips up even experienced engineers: you have an access-list on the router that permits and denies exactly the traffic you want, so why can two servers in the same VLAN still talk to each other freely? The answer is the single most important thing to understand about filtering at Layer 2\. A router ACL only ever sees traffic that *routes* between subnets. Two hosts in the same VLAN share a subnet, so their frames are bridged by the switch and never touch the router. The router ACL is never consulted, and a filter it never sees is a filter that does nothing. A VLAN access-list (VACL) is the tool that fixes this, because it acts on the switch as frames are bridged inside the VLAN. This article is part of the PingLabz [VLANs and Layer 2 switching](https://www.pinglabz.com/vlans-layer-2-switching/) cluster and the wider [network infrastructure security](https://www.pinglabz.com/infrastructure-security/) guide, and every command here was captured live on a Cisco IOL-XE switch in CML. ## Why a router ACL cannot filter inside a VLAN Trace the path of a packet. When HOST1 (10.20.10.100) sends to HOST2 (10.20.10.101), both are in VLAN 120 on the same 10.20.10.0/24 subnet. HOST1 looks at its own subnet mask, sees that HOST2 is local, and puts the frame straight onto the wire addressed to HOST2's MAC. The switch reads the destination MAC, finds it in the CAM table, and forwards the frame out the right port. At no point does that frame have a reason to go to the default gateway, because the destination is not in another subnet. The router, and therefore any ACL on the router's interface or SVI, is completely bypassed. Contrast that with HOST1 sending to something in 10.30.0.0/24\. Now the destination is off-subnet, so HOST1 forwards the frame to its default gateway, the router routes it, and any inbound or outbound ACL on that path gets its say. This is the whole distinction: Router ACL (RACL) SeesRouted traffic only Applied toA routed interface / SVI Directionin / out on that L3 hop Intra-VLANNever sees it VLAN ACL (VACL) SeesEvery frame in the VLAN Applied toThe whole VLAN DirectionBridged and routed both Intra-VLANFilters it ## Proving it: intra-VLAN traffic flows freely first Before any VACL exists, HOST1 can reach HOST2 with no trouble at all - same VLAN, same subnet, a clean bridged path: ``` HOST1$ ping -c 2 10.20.10.101 (HOST2, same VLAN) 2 packets transmitted, 2 received, 0% packet loss ``` Zero loss. If you had put a "deny HOST1 to HOST2" line in a router ACL, this ping would still succeed, because the frames never reach the router to be filtered. To actually stop it, you have to filter on the switch, inside the VLAN. That is a VACL. ## How a VACL is built A VACL (Cisco calls the construct a VLAN access-map) is different in shape from a normal ACL. It is a sequenced list of map entries, and each entry has two parts: a `match` clause that selects traffic (usually by referencing an ordinary IP or MAC access-list) and an `action` of `forward` or `drop`. You then bind the map to one or more VLANs with a `vlan filter` statement. Here is the exact config used in the lab to drop HOST1 to HOST2 while forwarding everything else: ``` ip access-list extended PLZ-H1-H2 permit ip host 10.20.10.100 host 10.20.10.101 vlan access-map PLZ-VACL 10 match ip address PLZ-H1-H2 action drop vlan access-map PLZ-VACL 20 action forward vlan filter PLZ-VACL vlan-list 120 ``` Read it top to bottom. Sequence 10 matches traffic from 10.20.10.100 to 10.20.10.101 (the `permit` in the referenced IP ACL means "this is the traffic I am selecting", not "allow it") and the action drops it. Sequence 20 has no match clause, so it catches everything else, and the action forwards it. The order matters exactly like a route-map: the first entry whose match clause hits wins. And crucially, there is no implicit "forward all" at the end of a VACL - the default action for unmatched traffic is drop, which is why sequence 20 with an explicit `action forward` is not optional. Leave it out and you would blackhole the entire VLAN. ## After the VACL: the same ping is dead, the gateway still works With the map applied to VLAN 120, retry the exact same intra-VLAN ping: ``` HOST1$ ping -c 3 10.20.10.101 (HOST2, now VACL-filtered) 3 packets transmitted, 0 received, 100% packet loss ``` One hundred percent loss. HOST1 can no longer reach HOST2 at all. But how do you know it is the VACL doing this and not a broken host, a downed link, or a bad ARP entry? Because the same host, on the same interface, at the same moment, still reaches its gateway without a hitch: ``` HOST1$ ping -c 2 10.20.10.254 (gateway, still fine) 2 packets transmitted, 2 received, 0% packet loss ``` That is the clean before-and-after that makes the point undeniable. HOST1 is alive, its NIC works, its path to the network works - traffic to the gateway (matched by sequence 20's forward) sails through. Only the one flow selected by sequence 10 is dropped. This is a filter a router ACL could never have applied, because HOST1-to-HOST2 traffic never leaves the VLAN. ## Confirming the state on the switch Two show commands confirm the map is present and bound. First the access-map itself, which lists each sequence, its match clause, and its action: ``` SW1#show vlan access-map PLZ-VACL Vlan access-map "PLZ-VACL" 10 Match clauses: ip address: PLZ-H1-H2 Action: drop Vlan access-map "PLZ-VACL" 20 Action: forward ``` Then the filter binding, which tells you which VLANs the map is actually policing: ``` SW1#show vlan filter VLAN Map PLZ-VACL is filtering VLANs: 120 ``` If a VACL "is not working", these two outputs are the first place to look: a map with the right actions but no `vlan filter` binding does nothing, and a filter pointed at the wrong VLAN policies traffic you never intended to touch. ## Where VACLs earn their keep Because a VACL is applied to the whole VLAN and is evaluated on the switch fabric, it reaches traffic that no router ACL can. The practical use cases all share that shape: - **Micro-segmenting servers that share a subnet.** A pool of app servers and database servers may live in one VLAN for simplicity. A VACL can permit the app tier to reach the database port and drop everything else between them, without renumbering into separate subnets. - **Blocking lateral movement.** If one host in a subnet is compromised, the attacker's next move is usually sideways, to the neighbours that share the segment. A VACL that denies host-to-host traffic while permitting host-to-gateway breaks that pivot. It is the same goal as private VLANs, achieved with a policy instead of a port mode. - **Filtering non-routed protocols.** A VACL can match on MAC access-lists too, so it can drop or permit frame types that a router ACL, which only sees IP, cannot express. A VACL is not a replacement for a router ACL - the two operate at different layers and see different traffic. The router ACL guards the doorway between subnets; the VACL polices the room. On a flat VLAN where hosts must share a subnet but must not fully trust each other, the VACL is the only tool that can express that policy. ## VACL or private VLAN? Two roads to the same isolation Once you understand that intra-VLAN traffic needs a Layer 2 control, you will meet two of them: the VACL and the private VLAN. They solve overlapping problems, so it is worth knowing when each fits. A VACL is a *policy*. You write match clauses that name specific sources, destinations, and protocols, and you can be as surgical as you like - permit app-tier to database on TCP 5432, drop the rest. That precision is its strength and its cost, because someone has to write and maintain the access-lists as the environment changes. A private VLAN is a *port mode*. You put ports into an isolated secondary VLAN and the switch enforces host-to-host isolation for you, with no per-flow rules to maintain. It is blunter (all-or-nothing isolation to the promiscuous uplink) but it scales effortlessly to hundreds of ports. Reach for a VACL when you need selective, protocol-aware filtering between known hosts; reach for a [private VLAN](https://www.pinglabz.com/private-vlans-explained/) when you simply need every host in a segment isolated from its neighbours. Many hardened designs use both: private VLANs for the broad isolation, a VACL for the few specific flows that must still be permitted. ## How the switch evaluates the map The access-map is processed in ascending sequence order, and evaluation stops at the first entry whose match clause matches the frame. That is why ordering is a design decision, not an afterthought: put your most specific drop rules at low sequence numbers and your broad `action forward` catch-all last. If you accidentally place a forward-all entry before a drop entry, the drop never fires because the frame already matched and left the map. Think of it exactly like a route-map or a firewall rule base: specific first, general last, and always an explicit terminal action so nothing falls through to the implicit drop by surprise. ## A few operational notes Keep the referenced IP or MAC access-lists tight and named, so `show vlan access-map` reads clearly six months from now. Remember that the implicit action for unmatched traffic is drop, so always finish with an explicit `action forward` entry unless you truly intend the VLAN to be default-deny. And test with the exact before-and-after pattern shown here: ping the target (should fail) and ping the gateway (should succeed) from the same host, so you can prove the VACL is responsible and not some unrelated fault. For the broader picture of how VACLs sit alongside every other Layer 2 control, keep reading the [VLANs and Layer 2 switching](https://www.pinglabz.com/vlans-layer-2-switching/) cluster and the [network infrastructure security](https://www.pinglabz.com/infrastructure-security/) guide. ## Key Takeaways - A router ACL only filters traffic that routes between subnets. Two hosts in the same VLAN share a subnet, so their frames never reach the router and a router ACL can never filter them. - A VACL (VLAN access-map) filters traffic as it is bridged inside the VLAN, so it reaches exactly the intra-VLAN flows a router ACL cannot. - A VACL is a sequenced list of `match` \+ `action` (forward or drop) entries, referencing IP or MAC access-lists, bound to VLANs with `vlan filter`. The default action for unmatched traffic is drop, so include an explicit `action forward` catch-all. - In the lab, HOST1 to HOST2 (same VLAN) went from 0% loss to 100% loss after the VACL, while HOST1 to the gateway stayed at 0% loss, proving the VACL, not a broken host, did the filtering. - Verify with `show vlan access-map` and `show vlan filter`. A map with no filter binding does nothing. - Reach for VACLs to micro-segment servers in a shared subnet, block lateral movement after a compromise, and filter non-routed protocols. See the full [VLANs and Layer 2 switching](https://www.pinglabz.com/vlans-layer-2-switching/) cluster. ### Static ARP ACLs: Inspecting ARP Without DHCP Snooping URL: https://www.pinglabz.com/static-arp-acl-inspection/ Last updated: 2026-07-13T14:21:07.000Z Dynamic ARP Inspection (DAI) is the control that kills ARP poisoning on a switch, but it has a hidden dependency: it validates ARP packets against the DHCP snooping bindings database. So what happens on a segment that has no DHCP at all - a rack of statically addressed servers, a management VLAN, a point-to-point link? There are no snooping bindings to check against, and DAI has nothing to enforce. This is where static ARP ACLs come in. They let you hand DAI the legitimate IP-to-MAC pairs yourself, so ARP inspection works with zero DHCP in the picture. This article is part of the PingLabz [VLANs and Layer 2 switching](https://www.pinglabz.com/vlans-layer-2-switching/) cluster and the broader [network infrastructure security](https://www.pinglabz.com/infrastructure-security/) guide, and every command below was captured live on a Cisco IOL-XE switch in CML. ## Why a static-IP segment breaks normal DAI Regular DAI works like this: DHCP snooping watches DHCP exchanges and records each lease as an IP-to-MAC-to-port binding in a table. When an ARP reply crosses an untrusted port, DAI looks up the sender's IP and MAC in that table. If the pair matches a known binding, the ARP is forwarded. If it does not, the ARP is dropped and logged. That is a clean, automatic defense - as long as every host got its address from DHCP. A static-IP segment breaks the whole chain. If your servers are configured by hand (as most server and management subnets are), no DHCP exchange ever happens, so snooping never builds a binding for them. Turn on DAI in that VLAN and, with an empty binding table, DAI would drop every legitimate ARP on the wire. You need a different source of truth for what is legitimate, and that source is a **static ARP ACL**: a hand-written list of the IP-to-MAC pairs you know are real. ## Why ARP needs a bodyguard at all ARP is one of the oldest and most trusting protocols on the network. When a host asks "who has 10.20.10.100?", any device on the segment can answer, and the asker caches whatever reply arrives first. There is no authentication, no sequence number, no way for the victim to tell a legitimate reply from a forged one. That design was fine in 1982 and is a gift to attackers today: poison a cache with one unsolicited reply and you can silently redirect traffic, run a man-in-the-middle, or blackhole a host. DAI exists precisely because ARP will never defend itself. On a switch it becomes the trusted third party that decides which IP-to-MAC claims are allowed to pass, and on a static segment the static ARP ACL is how you tell it the truth. ## The lab: VLAN 120, no DHCP The setup is a single user VLAN (VLAN 120, 10.20.10.0/24) on an IOL-XE switch. The gateway lives off the uplink, and two hosts sit on access ports, both statically addressed: Gateway (uplink) PortEt0/0 (trusted) IP10.20.10.254 HOST1 (legit) PortEt0/1 (untrusted) IP10.20.10.100 MAC5254.0089.dbfd HOST2 (attacker) PortEt0/2 (untrusted) IP10.20.10.101 MAC5254.007f.1646 There is no DHCP server anywhere in this VLAN. Every address was typed in by hand. That is exactly the condition where a static ARP ACL earns its place. ## Configuring the static ARP ACL An ARP access-list is not an IP ACL. Its entries permit a specific IP paired with a specific MAC, and DAI treats that pair as the definition of a legitimate host. You list every real host, then point DAI at the ACL for the VLAN: ``` arp access-list PLZ-ARP-ACL permit ip host 10.20.10.100 mac host 5254.0089.dbfd ! HOST1's real IP-MAC permit ip host 10.20.10.101 mac host 5254.007f.1646 ! HOST2's real IP-MAC ip arp inspection vlan 120 ip arp inspection filter PLZ-ARP-ACL vlan 120 interface Ethernet0/0 ip arp inspection trust ! the gateway uplink is trusted ``` Two things to note. First, the uplink to the gateway is marked `ip arp inspection trust`. Trusted ports skip inspection entirely, which is correct for an infrastructure link where you control both ends - you do not want DAI second-guessing ARP from your own router. Second, the two access ports (Et0/1 and Et0/2) are left untrusted by default, so every ARP they send is checked against `PLZ-ARP-ACL`. ## The attack: HOST2 impersonates HOST1 Now the classic ARP-poisoning move. HOST2 wants to insert itself between HOST1 and the gateway (a man-in-the-middle). To do that, it sends gratuitous ARP replies claiming that HOST1's IP (10.20.10.100) lives at HOST2's own MAC (52:54:00:7f:16:46). If the gateway believes it, traffic destined for HOST1 gets sent to HOST2 instead. Here HOST2 fires eight forged ARP replies at the gateway with `nping`: ``` HOST2$ nping --arp --arp-type ARP-reply --arp-sender-ip 10.20.10.100 \ --arp-sender-mac 52:54:00:7f:16:46 --arp-target-ip 10.20.10.254 -e eth0 -c 8 10.20.10.254 SENT (7.0584s) ARP reply 10.20.10.100 is at 52:54:00:7F:16:46 Raw packets sent: 8 (336B) ``` The forged binding is 10.20.10.100 paired with 5254.007f.1646\. Your ACL says 10.20.10.100 belongs to 5254.0089.dbfd. The pair does not match a permit entry, so DAI has every reason to drop it. ## DAI drops all eight Check the inspection statistics for the VLAN and the result is unambiguous: ``` SW1#show ip arp inspection statistics vlan 120 Vlan Forwarded Dropped DHCP Drops ACL Drops ---- --------- ------- ---------- --------- 120 2 8 8 0 Vlan DHCP Permits ACL Permits Probe Permits Source MAC Failures ---- ------------ ----------- ------------- ------------------- 120 0 1 0 0 ``` Eight forged ARPs sent, eight dropped. One legitimate ARP was permitted by the ACL, and two frames were forwarded. The man-in-the-middle attempt is dead on arrival: the gateway's ARP cache is never poisoned, and HOST1's traffic keeps flowing to HOST1\. This is the entire value of the control, proven on the wire, and it is the same defense the deeper [Dynamic ARP Inspection and IP Source Guard](https://www.pinglabz.com/dynamic-arp-inspection-ip-source-guard/) article builds on for DHCP-based segments. ## Read the drop column carefully: DHCP-Drops vs ACL-Drops Look again at the counters. The eight drops landed in the **DHCP Drops** column, not **ACL Drops**, even though it was the static ARP ACL that caught them. That is not a bug, it is how `ip arp inspection filter vlan ` behaves by default. Without the `static` keyword, DAI checks the ARP ACL first and then falls back to the DHCP snooping database. Because the fallback path exists, drops are attributed to the DHCP side of the counters. If you want the ACL to be the sole, authoritative source (no snooping fallback), add the `static` keyword: ``` ip arp inspection filter PLZ-ARP-ACL vlan 120 static ``` filter vlan ACL checked first, then falls back to the DHCP snooping database. Denied ARPs count as DHCP Drops. Use when you mix static and DHCP hosts in one VLAN. filter vlan static ACL is authoritative, no snooping fallback. Denied ARPs count as ACL Drops. Use on a pure static-IP segment where the ACL is the whole truth. ## The 15 pps rate-limit is your first line of defense There is a second control at work here that fires before the ACL check even runs. DAI applies a rate-limit to ARP packets on untrusted ports, defaulting to 15 packets per second with a 1-second burst interval. You can see it in the interface state: ``` SW1#show ip arp inspection interfaces Interface Trust State Rate (pps) Burst Interval Et0/0 Trusted None N/A Et0/1 Untrusted 15 1 Et0/2 Untrusted 15 1 ``` The uplink is trusted with no rate-limit, and both access ports cap ARP at 15 pps. This matters because ARP poisoning is rarely eight polite packets - a real attacker floods thousands per second to keep the victim's cache poisoned. If ARP crosses the rate-limit, the port is error-disabled, so the flood is throttled before DAI ever consults the ACL. Two layers: rate-limit the volume, then validate what remains against the static bindings. ## When to reach for a static ARP ACL Use one wherever DAI needs to run but DHCP snooping cannot feed it bindings: - **Server and DMZ subnets** that are statically addressed by design. - **Management VLANs** where every device has a fixed IP. - **Mixed segments**, where a handful of static servers share a VLAN with DHCP clients (omit the `static` keyword so DAI validates static hosts against the ACL and DHCP hosts against the snooping table). The trade-off is honest maintenance: an ARP ACL is a hand-maintained list. Add a server, add an entry. Re-image a NIC and the MAC changes, so update the pair. On a stable static segment that is a small price for closing the ARP-poisoning door completely. If your hosts do use DHCP, prefer snooping-backed DAI plus IP Source Guard, covered in the [L2 security stack article](https://www.pinglabz.com/dynamic-arp-inspection-ip-source-guard/). ## Key Takeaways - DAI validates ARP against the DHCP snooping bindings database. On a static-IP segment there are no bindings, so you supply the legitimate IP-to-MAC pairs yourself with a static ARP ACL. - An `arp access-list` uses `permit ip host X mac host Y` entries. Point DAI at it with `ip arp inspection filter vlan ` and trust your infrastructure uplinks. - In the lab, eight forged ARP replies impersonating HOST1 were all dropped (Forwarded 2, Dropped 8), killing the man-in-the-middle before the gateway cache could be poisoned. - Without the `static` keyword, ACL-denied ARPs count in the DHCP Drops column because DAI can still fall back to snooping. Add `static` to make the ACL authoritative and land the drops in ACL Drops. - Untrusted ports also enforce a default 15 pps ARP rate-limit, throttling a flood before the ACL check even runs. - Static ARP ACLs fit server, DMZ, and management VLANs. For DHCP-based segments, use snooping-backed DAI plus IP Source Guard instead. See the full [VLANs and Layer 2 switching](https://www.pinglabz.com/vlans-layer-2-switching/) cluster and the [network infrastructure security](https://www.pinglabz.com/infrastructure-security/) guide. ### Router Service Hardening: Turning Off What You Never Needed URL: https://www.pinglabz.com/cisco-router-service-hardening/ Last updated: 2026-07-13T13:56:37.000Z Every service a router runs is a door, and most routers ship with doors open that you will never walk through. A legacy Telnet listener, an HTTP server nobody manages, CDP announcing your topology to anything on the wire, small services from the 1980s that exist only as attack surface. Hardening is the unglamorous work of closing those doors, and the only honest way to know you succeeded is to scan the box and watch the open ports disappear. This article, part of the [infrastructure security](https://www.pinglabz.com/infrastructure-security/) cluster, does exactly that: a real Nmap scan of a router before hardening, the config that turns off what was never needed, and a re-scan proving the attack surface shrank to nothing. Then it covers the syslog design that makes the whole thing auditable. ## Before: scan the router, find a door open You cannot harden what you have not measured. So before touching the config, we scanned EDGE1 from the Debian attacker with Nmap, checking the twenty most common ports: ``` j@llmbits$ nmap -Pn -T4 --top-ports 20 192.168.99.1 23/tcp open telnet <-- cleartext telnet is LISTENING 22/tcp closed ssh 80/tcp closed http ... (all others closed) ... Nmap done: 1 IP address (1 host up) scanned in 0.28 seconds ``` There it is: **23/tcp open, Telnet listening**. Cleartext management, credentials and session in the clear, on a port anyone who can reach the interface can connect to. The vty lines had `transport input ssh telnet`, so Telnet was live. This is precisely the kind of finding a scan surfaces that a config review might skim past. If you want the scanning technique itself, our [Nmap](https://www.pinglabz.com/nmap/) cluster covers it in depth, and the before/after scan here is a natural bridge to it. ## Harden: turn off what you never needed The remediation is a checklist of services to disable and management to lock down. Every line has a reason: ``` no ip http server no ip http secure-server no cdp run no lldp run no service pad no ip source-route no ip bootp server no ip finger line vty 0 4 transport input ssh ! telnet gone, SSH only access-class PLZ-VTY-ACL in ! and restrict WHO can even connect ip access-list standard PLZ-VTY-ACL permit 192.168.99.0 0.0.0.255 deny any log ``` Kill cleartext management `transport input ssh` removes Telnet from the vty lines. SSH only, encrypted, authenticated. Restrict who connects `access-class PLZ-VTY-ACL in` means only 192.168.99.0/24 can even reach the vty. Everything else is denied and logged. Turn off web servers `no ip http server` / `secure-server` remove an unmanaged management interface and its whole vulnerability history. Silence discovery on the edge `no cdp run` / `no lldp run` stop the router advertising model, version, and topology to untrusted neighbours. Retire legacy small services `no service pad`, `no ip bootp server`, `no ip finger` \- decades-old services with no place on a modern edge. Block source routing `no ip source-route` stops attacker-dictated forwarding paths that bypass your routing and filtering design. Discovery protocols deserve a word. CDP and LLDP are genuinely useful *inside* your trusted network - they map neighbours and drive things like voice VLAN assignment. On an untrusted edge they are pure information leak, handing an attacker your platform, software version, and port layout for free. Disable them on edge-facing interfaces; keep them where they earn their keep. This service-plane work complements the management-plane lockdown in our companion article on [hardening the management plane on IOS XE](https://www.pinglabz.com/hardening-management-plane-ios-xe/) \- together they cover both what the box offers and who is allowed to manage it. ## Why an open service is worth closing even if "nothing uses it" The pushback you will hear against hardening is "that service is not hurting anyone, leave it." It is worth answering directly, because the reasoning is the same for every door on the list. An unused listening service costs you in three ways regardless of whether anyone legitimately uses it. First, it is attack surface: every listener is code that parses attacker-controlled input, and every such codebase has a vulnerability history you are now exposed to. The IOS HTTP server is a classic example, with a long list of CVEs for something most operators never deliberately use. Second, it is reconnaissance value: a Telnet banner, a CDP announcement, or a finger response tells an attacker what you are running, which is the first step of any targeted attack. Third, it is a foothold: cleartext management like Telnet hands over credentials to anyone who can sniff the segment, turning a passive position into an active one. "Nothing uses it" is not a security argument - if nothing uses it, that is the strongest possible reason to turn it off, because you lose nothing and remove all three costs. ## After: re-scan, and prove it Here is the rule that separates hardening from wishful thinking: prove each step with a real re-scan, not a claim. So we scanned EDGE1 again from the same attacker: ``` j@llmbits$ nmap -Pn -T4 --top-ports 20 192.168.99.1 Nmap done: 1 IP address (1 host up) scanned in 0.22 seconds (NO open ports reported) j@llmbits$ nmap -Pn -p22,23,80 192.168.99.1 22/tcp closed ssh 23/tcp closed telnet <-- telnet is now CLOSED 80/tcp closed http ``` The top-20 scan reports no open ports. The targeted scan of 22, 23, and 80 confirms it: **23/tcp is now closed**, the Telnet listener is gone. The attack surface that a scan could see went from "Telnet wide open" to "nothing," and we know that because a real tool measured it twice. That before/after is the entire credibility of hardening - a shrinking open-port list from an actual scan, not a config diff you hope had the intended effect. ## Syslog design: hardening you can reconstruct afterward Closing ports stops the easy attacks. When something does get through, the difference between a five-minute triage and a week of guessing is your logging. The lab already produces useful security events - recall the anti-spoofing capture, where the router logged the exact forged source: ``` %SEC-6-IPACCESSLOGDP: list PLZ-ANTISPOOF denied icmp 1.2.3.4 -> 10.255.0.2 (8/0), 1 packet ``` That message is only fully useful if the surrounding syslog design gives it three properties. Here is the minimum viable configuration: ``` service timestamps log datetime msec localtime show-timezone service sequence-numbers logging buffered 64000 informational logging host 192.168.99.100 ! the Debian VM as the remote collector logging trap informational ``` Accurate timestamp `timestamps ... msec ... show-timezone`. A log you cannot place on a timeline is nearly useless for correlation - which is exactly why NTP authentication matters upstream of it. Sequence number `service sequence-numbers`. Gaps in the sequence reveal where an attacker cleared or suppressed log entries - something timestamps alone cannot show. Remote destination `logging host` ships events off-box so they survive the device being compromised, wiped, or rebooted. Local buffer alone dies with the box. Put those together and you have a log an incident responder can actually work with: an accurate timestamp (so it lines up with every other device, which needs authenticated NTP to trust), a sequence number (so a cleared range is visible as a gap), and a remote copy on the collector (so the record outlives the compromised box). Msec plus timezone plus sequence numbers plus a remote collector is the minimum viable syslog design for an incident you can reconstruct afterward - and the accurate timestamp is why the [NTP authentication](https://www.pinglabz.com/ntp-authentication-cisco/) work is a prerequisite, not a nicety. Time and logging are core network services, so pair this with the wider [IP services](https://www.pinglabz.com/ip-services/) cluster too. ## Make it repeatable, not a one-off A hardened box drifts. Someone enables the HTTP server to troubleshoot and forgets to disable it; a template rebuild reintroduces Telnet; a new interface comes up edge-facing with CDP still running. The way you keep hardening from decaying is to treat the scan as an ongoing control, not a one-time event. Bake the before/after Nmap into your change process: scan after any config change to a device's management or service configuration, and diff the open-port list against a known-good baseline. A port that appears where none should be is an alert, the same way a new deny-log hit is. This is the operational discipline that makes the difference between a device that was hardened once and a device that stays hardened - and it is why the scanning skill from the [Nmap](https://www.pinglabz.com/nmap/) cluster is worth having in your own toolkit, pointed at your own gear. ## What the scan does and does not tell you One honest caveat about the before/after evidence. An external port scan measures reachable listening services from a given vantage point, which is exactly the attacker's view and precisely what you want for edge hardening. It does not, by itself, prove that a service is fully disabled everywhere - a listener could still be reachable from a different source that your vty access-class permits, or bound to a management interface the scanner cannot reach. That is a feature, not a gap: the access-class we applied means the vty is now reachable only from 192.168.99.0/24, so a scan from an unauthorized source seeing "closed" is the intended result. Combine the scan (the outside view) with a config audit (`show running-config` for the service lines) and the vty `access-class` deny-log hits (who tried and was refused). The scan proves the door is shut to outsiders; the access-class log proves someone rattled it; the config proves the service is off by policy. Together they give you defence you can demonstrate three different ways. ## The hardening checklist, in order Pulling it together into a repeatable process: - **Audit with a real scan first** \- Nmap the box and see what is actually listening, not what you assume. - **Kill cleartext management** \- Telnet to SSH only on the vty lines. - **Disable discovery on untrusted edges** \- CDP and LLDP off where they only leak information. - **Turn off legacy small services and source-routing** \- HTTP servers, PAD, BOOTP, finger, source-route. - **Restrict management with a vty access-class** \- only your management range can even reach the vty. - **Re-scan to prove each change** \- the open-port list must shrink for real. - **Design syslog for reconstruction** \- accurate time, sequence numbers, remote collector. ## Key Takeaways - **Prove hardening with a scan, not a claim.** A real Nmap before/after took EDGE1 from 23/tcp Telnet open to no open ports reported and Telnet closed. - **Kill cleartext management** \- `transport input ssh` plus a vty `access-class` so only your management range can connect at all. - **Turn off what you never needed** \- HTTP servers, CDP/LLDP on untrusted edges, PAD, BOOTP, finger, and source-routing are attack surface with no upside on an edge router. - **A log is only useful with three things** \- an accurate timestamp (hence authenticated NTP), a sequence number (to spot cleared logs), and a remote destination (to survive compromise). - **The before/after scan bridges to Nmap** \- measuring your own attack surface is the same skill as scanning anyone else's. Service hardening is where device security becomes measurable. Work through the rest of the [infrastructure security](https://www.pinglabz.com/infrastructure-security/) cluster to pair it with control-plane policing, anti-spoofing, and authenticated NTP for a device that is hard to reach, hard to spoof, and easy to audit. ### NTP Authentication: Trusting the Clock Your Logs Depend On URL: https://www.pinglabz.com/ntp-authentication-cisco/ Last updated: 2026-07-13T13:56:37.000Z Every security control you deploy quietly assumes one thing is true: that the clock is right. Certificate validation checks "not before" and "not after" against it. Log correlation across devices only works if their timestamps agree. Kerberos flat-out refuses tickets whose time is too far off. So an attacker who can move your clock does not need to break any of those systems - they can just lie to the thing all of them trust. NTP authentication is how you stop that, and it is a control far more people configure NTP without than with. This article, part of the [infrastructure security](https://www.pinglabz.com/infrastructure-security/) cluster, builds authenticated NTP between two Cisco routers in a real CML lab, quotes the security warning IOS XE itself throws about the cipher, and explains why authenticating the time source is a genuine security control and not just tidy housekeeping. ## What unauthenticated NTP actually trusts Plain NTP takes time from whoever answers. Point a client at a server and it will sync, no questions asked about whether that server is who it claims to be. On a segment an attacker can reach, that is an opening: stand up a rogue NTP server, answer faster or louder than the real one, and start walking the client's clock. Push it forward past a certificate's expiry and TLS validation starts failing. Push it out of Kerberos's tolerance window and authentication breaks. Skew it relative to your other devices and your log correlation - the thing you would use to detect the attack - falls apart. NTP authentication closes the opening by making the client verify a shared key before it will accept time from a server. ## Building authenticated NTP on Cisco IOS XE In our lab, CORE1 is the authenticated NTP master and EDGE1 is the authenticated client. The pieces are the same on both ends: define a keyed authentication key, turn on authentication, mark the key as trusted, and then reference the key when pointing at the peer. ``` ! CORE1 (master) ntp authentication-key 10 md5 ntp authenticate ntp trusted-key 10 ntp master 3 ! EDGE1 (client) ntp authentication-key 10 md5 ntp authenticate ntp trusted-key 10 ntp server 10.255.0.2 key 10 ``` Read the four commands in order, because each does a distinct job: ntp authentication-key Defines key ID 10 and its secret. The client and server must share the same ID and secret, or authentication fails. ntp authenticate Turns on authentication globally. Without it, the keys exist but are not enforced - a common misconfiguration. ntp trusted-key Marks key 10 as trusted for peering. Only trusted keys can authenticate a time source, so a defined-but-untrusted key still cannot sync. ntp server ... key 10 Tells the client to use key 10 when talking to this server. The server must prove the key or the client rejects its time. ## IOS XE warns you: MD5 is weak The moment you configure an MD5 key on IOS XE 17.18, the OS itself pushes back with a security warning worth quoting in full, because it tells you where NTP authentication is heading: ``` SECURITY WARNING - Module: NTP, Command: ntp authentication-key 10 md5 *, Reason: Weak cipher(s) are present ... NTP authentication using MD5 - vulnerable to collision attacks ... Remediation: Transition to more secure algorithms like SHA and AES ``` This is real output from the device, not our commentary. MD5 for NTP is legacy: it is vulnerable to collision attacks, and Cisco now explicitly recommends transitioning to SHA and AES-based keys where your platform and software support them. The authentication *model* we are demonstrating is identical regardless of cipher - define key, authenticate, trust, reference - so configure it with SHA/AES if you can. We used MD5 because it is what interoperates most broadly across gear you are likely to meet, but treat the warning as your instruction: prefer the stronger algorithm. ## Verifying authentication is in force The command that confirms the client is actually authenticating its server is `show ntp associations detail`. The words to look for are `authenticated` and `authtype`: ``` EDGE1#show ntp associations detail 10.255.0.2 configured, ipv4, authenticated, authtype (md5), ... stratum 16 ``` `authenticated, authtype (md5)` is the whole point: EDGE1 will only accept time from 10.255.0.2 if that server proves it holds the shared key. A rogue server without the key - or with the wrong key - cannot authenticate, so its time is ignored no matter how loudly it answers. That is the security property you configured NTP authentication to get. ## An honest note on sync timing in the lab You will notice the association above still shows `stratum 16` \- the "unsynchronized" stratum - and that is deliberate honesty about the lab, not a config error. Full NTP synchronization is slow: the reachability register has to fill, the client has to gather enough samples to trust the offset, and the stratum has to settle. In a virtual CML lab the emulated clocks start badly skewed and this can take many minutes, exactly as we saw with the ASA NTP capture in an earlier lab. So this article deliberately teaches the **authentication model** \- how the client proves it is talking to a trusted server - rather than the sync timing, which is a property of the (virtual) hardware clocks, not of the security configuration. On production gear with sane clocks, sync follows normally once authentication succeeds; the security-relevant part, that the client refuses an unauthenticated source, is what we captured. NTP itself is a core network service, so if you are building out time, DNS, DHCP, and the rest, pair this with the wider [IP services](https://www.pinglabz.com/ip-services/) cluster. ## Why authenticating the clock is a security control It is easy to file NTP under "operational hygiene" and move on. Resist that. The reason authenticated time is a security control, not housekeeping, is the chain of things that silently break when the clock is wrong: **Certificate validation**TLS checks validity dates against the local clock. Move the clock and valid certs look expired, or expired ones look valid. **Log correlation**Tracing an incident across devices needs their timestamps to agree. Skewed clocks make a timeline impossible to reconstruct. **Kerberos**Tickets carry timestamps and are rejected outside a tight window. Skew the clock and authentication simply fails. An attacker who can move an unauthenticated clock gets to attack all three without touching any of them directly. `ntp authenticate` plus `ntp trusted-key` take that lever away: the router will not move its clock for a server that cannot prove the shared key. And because so much of your security tooling - especially log correlation - depends on accurate time, authenticating the time source is a prerequisite for trusting everything downstream of it, including the security logs your other hardening produces. ## Operational details that trip people up A few practicalities decide whether authenticated NTP actually works in production rather than just in the config: - **The key ID and secret must match exactly on both ends.** A mismatched secret does not throw a loud error - the association simply never authenticates and the client keeps ignoring the server's time. If sync is not happening, verify the key before you suspect anything else. - **Defining a key is not the same as trusting it.** Without `ntp trusted-key 10`, key 10 exists but cannot authenticate a peer. It is a two-step deliberately, so you can pre-stage keys before activating them. - **Turn authentication on.** `ntp authentication-key` and even `ntp trusted-key` do nothing on their own until `ntp authenticate` is present. Configs that "have keys" but skip this line are the most common way to think you are authenticated when you are not. - **Restrict who can even talk NTP.** Authentication proves identity, but you should still limit which sources can reach the NTP service with an access-group or control-plane protection, so an attacker cannot even attempt the exchange. - **Plan key rotation.** Because the same secret lives on every client and server in the domain, rotating it is a coordinated change. Stage the new key ID on all devices as trusted, cut clients over to it, then retire the old one. None of these are exotic, but each is a quiet way to end up with NTP that looks authenticated in the running config and is not authenticating anything on the wire. The verification command is your ground truth: if `show ntp associations detail` does not say `authenticated`, you are not. ## Key Takeaways - **Unauthenticated NTP trusts any answering server**, which lets an attacker move your clock and silently break certificate validation, log correlation, and Kerberos. - **Authenticated NTP needs four pieces** on each end: `ntp authentication-key`, `ntp authenticate`, `ntp trusted-key`, and a keyed `ntp server ... key` reference. - **IOS XE 17.18 flags MD5 as weak** \- it is vulnerable to collision attacks, and Cisco recommends moving to SHA and AES. The model is identical; use the stronger cipher where you can. - **Verify with `show ntp associations detail`** \- the words `authenticated, authtype (md5)` confirm the client will only accept time from a server that proves the key. - **Sync timing is slow in a virtual lab** (the association can sit at stratum 16 for minutes), so the article teaches the authentication model, not the sync timing. - **Authenticated time is a security control**, not housekeeping - it is the foundation the rest of your logging and PKI trust. Trustworthy time underpins trustworthy logs. Continue through the [infrastructure security](https://www.pinglabz.com/infrastructure-security/) cluster to connect authenticated NTP with syslog design, control-plane policing, and management-plane hardening. ### Anti-Spoofing on the Edge: ACLs, uRPF, and RFC 2827 in Practice URL: https://www.pinglabz.com/anti-spoofing-acls-urpf/ Last updated: 2026-08-01T19:32:07.000Z A spoofed source address is the oldest trick in the network attacker's book, and it still works because most networks never check. If a packet arrives at your edge claiming to be from an address that could never legitimately live out there, your router will happily forward it - unless you told it not to. Anti-spoofing is the practice of dropping those packets at the boundary, and it is one of the highest-value, lowest-cost controls you can deploy. This article, part of the [infrastructure security](https://www.pinglabz.com/infrastructure-security/) cluster, builds two anti-spoofing controls in a real CML lab - a uRPF strict-mode check and an RFC 2827 (BCP 38) ingress ACL - fires actual spoofed packets at them from a Debian attacker, and shows you which one gave us countable, provable evidence and which one silently did its job. ## Why spoofing works, and what RFC 2827 asks of you Source address validation is not the default. A router's job is to forward toward the destination; it does not, on its own, care whether the source address makes sense. That is what lets an attacker forge a source - to hide their origin, to bounce a reflection or amplification attack off a third party, or to impersonate a trusted internal host. RFC 2827, better known as BCP 38, is the internet community's standing request that every network filter ingress traffic so it can only source addresses that legitimately belong to it. If everyone did it, whole classes of DDoS reflection would dry up. Most do not, which is exactly why doing it at your own edge still matters. There are two practical ways to enforce it on a Cisco router: an ingress ACL that permits only your real source prefixes, or Unicast Reverse Path Forwarding (uRPF), which checks each packet's source against the routing table automatically. We built both. ## uRPF strict mode: configured, active, and honestly imperfect on virtual gear uRPF strict mode asks a simple question of every inbound packet: if I had to send a reply to this source address, would I send it back out the same interface this packet arrived on? If yes, the source is plausible and the packet passes. If the best route to that source points out a different interface (or there is no route at all), the source is considered spoofed and the packet is dropped. The config is a single interface line: ``` interface Ethernet0/0 ip verify unicast source reachable-via rx ``` On EDGE1 the feature is unmistakably active - it shows up in the interface's input feature list and reports the reachable-via RX mode: ``` EDGE1#show ip interface Ethernet0/0 Input features: uRPF, MCI Check IP verify source reachable-via RX 0 verification drops ``` Here is the honest finding, and we are not going to hide it: the `verification drops` counter stayed at 0 even when we sent spoofed packets that had no matching route. The feature is enforced - "Input features: uRPF" is real and the check is running - but the software drop counter on this Cisco IOL-XE image does not increment. This is the same virtual-platform quirk we have documented before with WRED counters on IOL: the behaviour is present, the counter is a no-op. On physical ASR or ISR hardware that counter ticks and gives you a clean number to alert on. In our lab it does not, so we could show that uRPF is *on*, but we could not prove a drop *count* with it. For that, we turned to the ACL, which counts reliably. If you want the deeper mechanics of strict versus loose mode, our dedicated article on [uRPF strict vs loose](https://www.pinglabz.com/unicast-rpf-strict-loose/) goes further. ## The RFC 2827 ingress ACL: the reliable proof Where uRPF validates dynamically against the routing table, the ingress ACL takes the explicit approach: you list the source prefixes that are legitimately allowed to originate on this edge, permit them, and deny (and log) everything else. On our edge the only legitimate source subnet is 192.168.99.0/24, so: ``` ip access-list extended PLZ-ANTISPOOF permit ip 192.168.99.0 0.0.0.255 any deny ip any any log interface Ethernet0/0 ip access-group PLZ-ANTISPOOF in ``` Line 10 permits our real edge subnet as a source. Line 20 denies everything else and logs it. Any packet arriving on Ethernet0/0 with a source outside 192.168.99.0/24 is, by definition on this edge, spoofed - and it gets dropped and named in the log. ## Firing real spoofed packets at it To test it properly we needed genuinely spoofed traffic. There is a wrinkle worth knowing: the Linux kernel on the attacker VM runs its own reverse-path filter, so a normally-sent packet with a forged source is dropped by the host before it ever leaves the NIC. To get spoofed frames onto the wire we injected at Layer 2 with scapy, addressing the frame directly to the router's real MAC: ``` j@llmbits$ sudo python3 -c "from scapy.all import *; \ sendp(Ether(dst='aa:bb:cc:00:72:00')/IP(src='1.2.3.4',dst='10.255.0.2')/ICMP(), \ iface='ens224', count=30)" ``` Thirty packets, each claiming to be from 1.2.3.4 - an address that could never legitimately arrive on this edge - aimed at 10.255.0.2 behind the router. Exactly the kind of forged traffic BCP 38 exists to stop. Note how cheap that was. A forged source address is one field in a one-line command, which is the uncomfortable fact underneath all of BCP 38, and the same technique extends to [forging any header field you like and watching the router record the lie](https://www.pinglabz.com/scapy-packet-crafting-spoofing-cisco/). ## The proof: the deny rule catches every spoofed packet The ACL counters tell the story cleanly. The deny line racked up matches, and the syslog named the forged source by address: ``` EDGE1#show ip access-lists PLZ-ANTISPOOF Extended IP access list PLZ-ANTISPOOF 10 permit ip 192.168.99.0 0.0.0.255 any 20 deny ip any any log (26 matches) EDGE1#show logging %SEC-6-IPACCESSLOGDP: list PLZ-ANTISPOOF denied icmp 1.2.3.4 -> 10.255.0.2 (8/0), 1 packet ``` Twenty-six of the thirty spoofed packets hit the deny (ACL logging rate-limits and samples, which is why the count is not a perfect 30), and the `%SEC-6-IPACCESSLOGDP` message spells out the spoofed source, 1.2.3.4, the target, and the protocol. That is exactly what you want from an anti-spoofing control: not just a drop, but a countable, attributable, alertable record that you can feed to a SIEM. This is the reliable proof uRPF's dead counter could not give us on virtual gear. ## ACL versus uRPF: the real comparison Both enforce RFC 2827, but they have genuinely different strengths, and the choice depends on where you are deploying: Ingress ACL (RFC 2827) ModelExplicit - you list legitimate source prefixes EvidenceCountable + logged (proven here) Best forA stable edge with a known, small source range CostMaintenance on every prefix change uRPF ModelDynamic - validates against the routing table EvidenceDrop counter (a no-op on our IOL image) Best forA routed core where prefixes change often CostStrict mode breaks asymmetric routing The ACL is bulletproof and countable where the legitimate source range is fixed and small - your single-homed customer edge, a server VLAN, a stub site. Its weakness is maintenance: every time the addressing changes, someone has to remember to update the ACL, and a stale anti-spoof ACL either leaks or blackholes. uRPF is dynamic and self-maintaining - it always reflects the current routing table, so it scales to a core where prefixes come and go. Its weakness is asymmetric routing: strict mode drops legitimate traffic whose return path differs from its arrival path, which is common on a multi-homed core. That is why loose mode exists (it only checks that a route to the source exists at all, not that it points back out the same interface), trading precision for tolerance of asymmetry. ## Use both at the edge These are not mutually exclusive, and on a real edge the right answer is to run them together. On a single-homed edge, uRPF strict mode is the cleaner, lower-maintenance default - one line, no prefix list to babysit. Where the source range is genuinely fixed and you want countable, logged evidence for your SOC, add the RFC 2827 ACL alongside it. uRPF gives you the dynamic, no-maintenance baseline; the ACL gives you the explicit, auditable proof. Belt and braces at the boundary, which is precisely where spoofed traffic has to be stopped. ## Ingress and egress: filter in both directions RFC 2827 is usually discussed as ingress filtering, but the same principle applies on the way out, and deploying both is what makes you a good internet citizen rather than just a well-defended one. Ingress filtering on your edge stops forged sources arriving from outside. Egress filtering stops your own network sourcing forged packets toward the rest of the world - which is exactly the traffic that fuels reflection attacks launched from compromised hosts inside your perimeter. A host on 192.168.99.0/24 should only ever send packets sourced from that range; anything else leaving your edge is either misconfiguration or a compromised machine, and both are worth catching. The ACL model makes egress filtering just as explicit: permit your own prefixes as a source outbound, deny and log the rest. ## Where to deploy, and common mistakes Anti-spoofing belongs at the boundary between trust levels - the customer-facing edge, the link to an untrusted partner, the demarcation to the internet. It is far less useful buried in the core, where legitimate traffic from many sources transits and a strict check either has nothing to add or actively breaks asymmetric flows. A few mistakes recur: - **Deploying uRPF strict on an asymmetric link.** If return traffic for a source can legitimately arrive on a different interface than it leaves, strict mode drops it. Use loose mode there, or an ACL. - **Forgetting the log keyword.** A deny without `log` still drops, but you lose the attributable evidence - the whole point of the exercise for a security team. Our capture only named 1.2.3.4 because the deny line carried `log`. - **Letting the ACL go stale.** An anti-spoof ACL is only as good as its prefix list. When addressing changes and nobody updates the ACL, you either blackhole legitimate new sources or leave a gap. This is precisely the maintenance burden uRPF avoids. - **Testing with kernel reverse-path filtering left on.** As we hit above, a normal spoofed send never leaves a Linux host with rp\_filter enabled. If your spoof test seems to "pass" with no packets arriving, the host dropped them, not the router. Inject at Layer 2 to test the router's control specifically. ## Key Takeaways - **Anti-spoofing enforces RFC 2827 / BCP 38** \- drop packets whose source could never legitimately arrive on this edge, and you kill a whole class of reflection and impersonation attacks. - **uRPF strict mode validates dynamically** against the routing table (`ip verify unicast source reachable-via rx`). It was active on our lab ("Input features: uRPF"), but the drop counter is a no-op on virtual IOL - the feature works, the count does not tick. - **The RFC 2827 ACL gave the reliable proof** \- real spoofed packets from 1.2.3.4 hit `deny ip any any log (26 matches)` and produced a named syslog record pointing at the forged source. - **ACL = explicit, countable, great on a stable edge; uRPF = dynamic, self-maintaining, better on a routed core** \- but strict mode breaks asymmetric routing. - **Use both at the edge** \- uRPF for the maintenance-free baseline, the ACL for auditable evidence. Edge anti-spoofing is one layer of a hardened device. Continue through the [infrastructure security](https://www.pinglabz.com/infrastructure-security/) cluster to combine it with control-plane policing, management-plane lockdown, and authenticated logging. ### CoPP vs CPPr: Which Control-Plane Defense to Deploy URL: https://www.pinglabz.com/copp-vs-cppr/ Last updated: 2026-07-13T13:56:36.000Z You have decided the control plane needs protecting - good. The next question is which tool: Control Plane Policing (CoPP) or Control Plane Protection (CPPr)? They are related, they are often confused, and the honest answer for most networks is "CoPP, and it is enough." This decision guide lays out exactly what each one does, backs the "CoPP is enough" case with a real police capture from our lab, and tells you the specific conditions under which reaching for CPPr actually pays off. It is part of the [infrastructure security](https://www.pinglabz.com/infrastructure-security/) cluster. ## The one-sentence version **CoPP is a single aggregate QoS policy on the entire control plane. CPPr splits the control plane into three subinterfaces and adds port-filtering and per-protocol queue-thresholding.** CoPP is simpler and universal; CPPr is finer-grained and platform-dependent. If you remember nothing else, remember that CPPr is not a replacement for CoPP - it is a superset you layer on where the platform supports it. ## Side by side CoPP ScopeOne aggregate policy on the whole control plane GranularityPer class-map, but all classes share one control-plane Port-filterNo Queue-thresholdNo Platform supportNear universal ComplexityLow CPPr ScopeHost / transit / cef-exception subinterfaces GranularityIndependent policer per subinterface Port-filterYes - drops closed-port traffic Queue-thresholdYes - per protocol Platform supportPlatform-dependent ComplexityHigher ## The case for CoPP: it works, and here is the proof The strongest argument for "CoPP is enough for most" is that it demonstrably stops the attack it is meant to stop. In our CML lab we put a single, deliberately tight aggregate policy on EDGE1 - police ICMP to the control plane at 8 kbps, drop the excess: ``` access-list 150 permit icmp any any class-map match-all PLZ-ICMP-CLASS match access-group 150 policy-map PLZ-COPP class PLZ-ICMP-CLASS police 8000 conform-action transmit exceed-action drop control-plane service-policy input PLZ-COPP ``` Then the Debian attacker flooded the router: ``` j@llmbits$ sudo ping -f -c 800 -i 0.002 192.168.99.1 --- 192.168.99.1 ping statistics --- 800 packets transmitted, 91 received, 88.625% packet loss, time 7446ms ``` ``` EDGE1#show policy-map control-plane Control Plane Service-policy input: PLZ-COPP Class-map: PLZ-ICMP-CLASS (match-all) police: cir 8000 bps, bc 1500 bytes conformed 91 packets, 8918 bytes; actions: transmit exceeded 709 packets, 69482 bytes; actions: drop ``` 91 conformed plus 709 dropped equals the 800 sent, and 91 is exactly what the flood got back. One aggregate policy, one command applied to `control-plane`, and the CPU never saw 709 of the 800 attack packets. That is the whole job. For the full build and tuning guidance, read our dedicated [Control Plane Policing (CoPP)](https://www.pinglabz.com/control-plane-policing-copp/) walkthrough. ## The configuration difference, concretely Both features use the same MQC building blocks - class-maps and policy-maps - so at first glance the configs look almost identical. The difference is entirely in where the policy attaches. CoPP binds one service-policy to the aggregate `control-plane`: ``` control-plane service-policy input PLZ-COPP ``` CPPr binds policies to named subinterfaces, and adds a policy type that has no CoPP equivalent - the port-filter that matches closed ports: ``` control-plane host service-policy type port-filter input PLZ-PF-POLICY service-policy input PLZ-CPPR-HOST control-plane cef-exception service-policy input PLZ-CEF-POLICY ``` That is the whole structural story. If your muscle memory already knows CoPP, CPPr is not a new language - it is the same MQC syntax pointed at three targets instead of one, plus `type port-filter`. The operational cost of CPPr is not learning new commands; it is having three or four policies to reason about and keep from dropping your own traffic, instead of one. ## When CPPr earns its extra complexity CoPP's limitation is that everything punted to the CPU shares one policer's fate. That is usually fine. It stops being fine in three specific situations, and those are exactly when CPPr is worth the effort: You need class isolation A flood of un-switchable (cef-exception) traffic should not be allowed to starve your SSH management path. CPPr gives each subinterface its own budget. You are scanned constantly On an exposed edge, port scans to closed ports waste CPU. Only CPPr's port-filter drops closed-port traffic before it is queued. One protocol goes noisy Queue-thresholding caps each protocol's slice of the punt queue, so one chatty protocol cannot block the rest even below the aggregate rate. If none of those describe your environment - and for a great many enterprise and campus devices, none do - then CoPP alone is the right answer, and adding CPPr is complexity without a matching risk reduction. ## Honest note: our platform only exposed CoPP We proved CoPP with real counters above. We could not capture live CPPr output, because the Cisco IOL-XE (17.18) image our CML lab runs on does not implement the control-plane subinterfaces - `control-plane host` and `show control-plane host open-ports` both return `% Invalid input`. CPPr is a hardware and platform feature (common on ISR and ASR), so verify it on your own gear. We describe the CPPr model from Cisco's documentation and refuse to fabricate show output for it. The full CPPr model, including the port-filter config, is in our companion article on [Control Plane Protection (CPPr)](https://www.pinglabz.com/control-plane-protection-cppr/). ## The decision, in one flow Start with CoPP on every device - it is universal, low-risk, and proven. Then ask: is this box directly exposed to untrusted networks, scanned frequently, or does it need hard isolation between control-plane traffic classes? If yes, and the platform supports it, layer CPPr on top for the port-filter and per-subinterface policing. If no, stop at CoPP. You are not choosing between them so much as deciding how far up the ladder your particular device needs to climb. ## A common misconception Engineers sometimes treat CPPr as "the new CoPP" and assume enabling it replaces the aggregate policy. It does not. If you configure CPPr subinterface policies but never build a sensible baseline of protection, you can end up with less coverage than a good single CoPP policy would give you, because traffic that does not match any of your subinterface classes may pass unpoliced. The correct mental model is additive: a solid CoPP-style baseline of classification and policing first, then the CPPr-specific wins (isolation, port-filter, queue-threshold) layered on. Do not tear out working CoPP to "upgrade" to CPPr - extend it. The other frequent mistake is policing management traffic into the ground. In both models, the failure you must avoid is dropping SSH from your own NOC or starving a routing protocol. Whichever tool you deploy, explicitly permit and generously rate the traffic that keeps the box reachable and converged, and baseline the real punt rate before you set any conform value. ## Frequently asked ### Does CPPr improve on CoPP for a typical enterprise core? Usually not enough to justify the complexity. A well-tuned CoPP policy on a core device that is not directly exposed to untrusted hosts already stops the volumetric attacks that matter. CPPr's port-filter shines on an exposed edge that gets scanned; on a protected core it solves a problem you may not have. ### If my platform does not support CPPr, am I exposed? No. CoPP alone is a legitimate, complete control-plane defense - it is what our lab used to absorb the flood. CPPr adds isolation and closed-port filtering, but the absence of CPPr is not a gap you have to panic about as long as CoPP is in place and tuned. ### Can I run both at once? On platforms that support CPPr, yes, and that is the intended design. Keep a baseline of classification and policing, and use the CPPr subinterfaces for the extra granularity. They are complementary layers, not an either/or. ## Key Takeaways - **CoPP is one aggregate policy; CPPr subdivides the control plane** into host, transit, and cef-exception subinterfaces plus port-filtering and queue-thresholding. - **CoPP is enough for most networks** and it is proven - our 8 kbps policer dropped 709 of 800 flood packets (88.6% loss) with a single command on `control-plane`. - **Reach for CPPr** only when you need class isolation, you face constant port scanning (the port-filter), or one protocol can flood the punt queue (queue-thresholding). - **They stack, they do not compete.** Deploy CoPP everywhere; add CPPr where the platform and the threat model justify it. - **Platform reality:** our IOL lab exposed CoPP but not CPPr's subinterfaces, so we proved CoPP live and describe CPPr from documentation. This is one decision inside a larger hardening story. Work through the rest of the [infrastructure security](https://www.pinglabz.com/infrastructure-security/) cluster to pair control-plane defense with anti-spoofing, management-plane lockdown, and authenticated NTP. ### Control Plane Protection (CPPr): Beyond CoPP with Port-Filtering URL: https://www.pinglabz.com/control-plane-protection-cppr/ Last updated: 2026-07-13T13:56:35.000Z Every routed packet your device forwards is handled in hardware, but the packets addressed *to* the device - routing updates, SSH sessions, ARP, ICMP to a local interface, SNMP polls - get punted up to the CPU. That CPU is the control plane, and it is the single most valuable target on the box. Flood it and you can knock out the routing protocol adjacencies that keep the network converged. Control Plane Policing (CoPP) was the first answer to that problem, and Control Plane Protection (CPPr) is its finer-grained evolution. This article is part of the [infrastructure security](https://www.pinglabz.com/infrastructure-security/) cluster, and it picks up exactly where CoPP leaves off. We will explain the CPPr model in full - the way it carves the control plane into separate subinterfaces, the port-filtering feature that drops traffic to closed control-plane ports, and the per-protocol queue-thresholding that stops one protocol starving another. Then we will be honest about the platform: the Cisco IOL image we lab on does not expose the CPPr subinterfaces, so we prove the closely related CoPP behaviour with real captures from our CML lab and present the CPPr configuration from Cisco's documentation with a clear platform note. ## The control plane, and why CoPP came first When a packet is destined for the router itself, it cannot be switched in CEF like transit traffic. It is "punted" to the route processor for handling. Legitimate control traffic (an OSPF hello, a BGP keepalive, an SSH login) is low volume. An attacker who sends a high rate of packets to a router IP, or a misbehaving host looping ICMP, generates punts at a rate the CPU was never sized for. The result is high CPU, dropped protocol adjacencies, and a control plane that falls over while the data plane is still technically "up". CoPP fixes this by attaching a QoS policy to a logical `control-plane` interface. You classify control traffic into classes, then police or drop each class. It is a single, aggregate policy applied to everything punted to the CPU. It is simple, universally supported, and, as we show below, genuinely effective. If you want the full CoPP build, read our deep dive on [Control Plane Policing (CoPP)](https://www.pinglabz.com/control-plane-policing-copp/). ## What CPPr adds: three control-plane subinterfaces CoPP treats the entire punt path as one bucket. CPPr's core idea is that not all punted traffic is equal, so it should not be policed as one lump. CPPr subdivides the control plane into three logical subinterfaces, each of which you can classify and police independently: control-plane host Traffic **destined to any interface on the router itself** \- SSH and Telnet to a management IP, SNMP, ICMP echo to a local address, and the packets that feed the routing protocols. This is where port-filtering and per-protocol queue-thresholding live, because it is the surface an attacker aims at. control-plane transit Traffic that is **software-switched through the router** (transit packets that still touch the CPU, for example when a feature forces process switching). Policing here protects the CPU from transit that should never have been punted in volume. control-plane cef-exception Packets that CEF cannot switch and must escalate - TTL-expired packets, packets needing an ICMP unreachable, options-set packets, ARP. Policing this class blunts attacks built on deliberately un-switchable traffic. The immediate benefit is isolation. With CoPP, a flood of cef-exception traffic (say, a TTL-expiry attack) shares one policer with your SSH management sessions. With CPPr, you can hand cef-exception a tight budget and leave the host path room to breathe, so one class of abuse cannot starve another. ## Port-filtering: drop traffic to closed control-plane ports This is the feature that has no CoPP equivalent, and it is the reason CPPr exists. Port-filtering, applied to the `host` subinterface, drops packets addressed to **closed or non-listening TCP and UDP ports on the router** before they are ever queued to the CPU. Think about what that means: if the box is not running an HTTP server, there is no legitimate reason for anyone to send it TCP 80\. Normally those packets still punt to the CPU, get processed far enough to generate a reset, and cost you cycles. A port scan (exactly the kind of thing our [Nmap](https://www.pinglabz.com/nmap/) lab generates) becomes a cheap way to load the control plane. Port-filtering closes that door. The router tracks which control-plane ports are actually open, and the port-filter policy silently drops traffic to everything else at the punt stage. An attacker probing closed ports gets nothing and spends none of your CPU. It is the control-plane analogue of the anti-scan hardening you do on the data plane. ## Queue-thresholding: no single protocol can flood the queue The third CPPr feature is per-protocol queue-thresholding on the host subinterface. Each protocol punted to the CPU (for example, a specific Layer 4 port or protocol type) gets a maximum number of packets allowed in the punt queue at once. Once a protocol hits its threshold, further packets of that protocol are dropped even though the aggregate queue still has room. This stops a single noisy protocol from consuming the entire punt queue and blocking everything else - a subtler failure mode than raw rate, and one a single CoPP policer does not address. ## The CPPr configuration model (from Cisco documentation) CPPr is configured much like CoPP - class-maps and a policy-map - but the policy is applied to a specific control-plane subinterface rather than the aggregate. A representative build looks like this (this is the documented model; see the honest platform note below): ``` ! Classify traffic destined to closed ports for the port-filter class-map type port-filter PLZ-PORTFILTER match closed-ports policy-map type port-filter PLZ-PF-POLICY class PLZ-PORTFILTER drop ! Aggregate policer for the host subinterface class-map match-all PLZ-HOST-MGMT match access-group 120 policy-map PLZ-CPPR-HOST class PLZ-HOST-MGMT police 32000 conform-action transmit exceed-action drop ! Apply per-subinterface control-plane host service-policy type port-filter input PLZ-PF-POLICY service-policy input PLZ-CPPR-HOST control-plane cef-exception service-policy input PLZ-CEF-POLICY ``` Notice the two things CoPP cannot do: a `type port-filter` policy that matches `closed-ports`, and a service-policy bound specifically to `control-plane host` rather than the whole control plane. That granularity is the entire value proposition. ## Honest platform note: our IOL image exposes CoPP, not CPPr We build these labs on Cisco IOL-XE (17.18) in CML. CPPr needs the control-plane host / transit / cef-exception subinterfaces, and this image does not implement them. The verification commands simply do not exist on the platform: ``` EDGE1#show control-plane host open-ports ^ % Invalid input detected at '^' marker. EDGE1#configure terminal EDGE1(config)#control-plane host ^ % Invalid input detected at '^' marker. ``` That is a real result, and we are not going to fabricate CPPr show output to paper over it. CPPr's subinterface model, port-filtering, and queue-thresholding are hardware and IOS-XE platform features you should verify on your own gear (many ISR and ASR platforms expose them). What our lab *can* prove, cleanly and repeatably, is the closely related CoPP behaviour that CPPr builds on. So that is what we captured. ## What we could prove: real CoPP policing an ICMP flood On EDGE1 we applied a deliberately tight CoPP policy - police ICMP to the control plane at 8 kbps and drop the excess: ``` access-list 150 permit icmp any any class-map match-all PLZ-ICMP-CLASS match access-group 150 policy-map PLZ-COPP class PLZ-ICMP-CLASS police 8000 conform-action transmit exceed-action drop control-plane service-policy input PLZ-COPP ``` Then our Debian attacker at 192.168.99.100 flooded the router's control plane with a fast ICMP stream: ``` j@llmbits$ sudo ping -f -c 800 -i 0.002 192.168.99.1 --- 192.168.99.1 ping statistics --- 800 packets transmitted, 91 received, 88.625% packet loss, time 7446ms ``` 88.6% loss. The router only answered 91 of the 800 flood packets, and the reason is on the policer counters, which match the attack exactly: ``` EDGE1#show policy-map control-plane Control Plane Service-policy input: PLZ-COPP Class-map: PLZ-ICMP-CLASS (match-all) Match: access-group 150 police: cir 8000 bps, bc 1500 bytes conformed 91 packets, 8918 bytes; actions: transmit exceeded 709 packets, 69482 bytes; actions: drop conformed 1000 bps, exceeded 2000 bps ``` Read those numbers: 91 conformed (transmitted) plus 709 exceeded (dropped) equals the 800 packets sent, and the 91 conformed is exactly what the VM got back. The CPU only ever processed the 91 conforming packets; the other 709 were dropped before they reached the process level. That is a control plane defending itself in real time, with counters you can point at. ## CPPr as the finer-grained evolution of CoPP Everything you just saw the CoPP policer do, CPPr does too - but per-subinterface, plus port-filtering and queue-thresholding on top. The mental model is a straight line: **CoPP**One aggregate policer on the whole control plane. Simple, universal, proven here. **CPPr subinterfaces**Separate host / transit / cef-exception policers so one class cannot starve another. **CPPr port-filter**Drop traffic to closed control-plane ports at the punt stage - defeats port scans cheaply. **CPPr queue-threshold**Cap each protocol's share of the punt queue so no single protocol floods it. The practical guidance: deploy CoPP everywhere - it is supported on virtually every platform and it works, as our capture shows. Where your platform exposes the CPPr subinterfaces, layer CPPr on top for the extra isolation and the port-filter. They are not competitors; CPPr is CoPP with a finer knife. For a head-to-head on when each is the right call, see our companion article on [CoPP vs CPPr](https://www.pinglabz.com/copp-vs-cppr/). ## Deploying it without locking yourself out Two cautions before you push either policy. First, police conservatively at the start. Our 8 kbps ICMP policer was deliberately brutal to make the drops obvious in a lab; on a production box a policer that tight on the wrong class will drop your own management traffic and routing adjacencies. Baseline the normal punt rate first, then set the conform rate with headroom above it. Second, always leave an explicit permit for the protocols that keep the box reachable and converged - SSH from your management range, the routing protocols, ARP, and BFD if you run it. A CoPP or CPPr policy that accidentally starves OSPF will flap every adjacency on the device, which is a self-inflicted outage that looks exactly like the attack you were defending against. With CPPr specifically, the port-filter is the safest single win to deploy first: dropping traffic to genuinely closed ports has no legitimate victim, so it rarely causes collateral damage. The per-protocol queue-thresholds and the per-subinterface policers are where you want a maintenance window and a rollback plan, because a mis-scoped class-map on the host subinterface can quietly drop the very sessions you use to fix it. ## Key Takeaways - **The control plane is the CPU, and it is the target.** Packets punted to the route processor are the ones that take a router down, so they need their own protection. - **CoPP is one aggregate policer.** Our lab proved it: an 8 kbps policer dropped 709 of 800 flood packets (88.6% loss) and only let 91 conforming packets reach the CPU. - **CPPr subdivides the control plane** into host, transit, and cef-exception subinterfaces so one class of traffic cannot starve another. - **Port-filtering is CPPr's unique feature** \- it drops traffic to closed control-plane ports at the punt stage, making port scans cheap to absorb. - **Queue-thresholding** caps each protocol's share of the punt queue. - **Platform honesty:** our IOL image does not expose the CPPr subinterfaces (the commands return % Invalid input), so verify CPPr on hardware that supports it - but CoPP is universal and should be on every box. Control-plane protection is a core pillar of device hardening. Continue with the rest of the [infrastructure security](https://www.pinglabz.com/infrastructure-security/) cluster to pair this with anti-spoofing, management-plane lockdown, and authenticated logging. ### ZBF vs ASA vs FTD: Which Firewall Belongs Where URL: https://www.pinglabz.com/zbf-vs-asa-vs-ftd/ Last updated: 2026-07-13T13:26:04.000Z Cisco sells three different stateful firewalls, and engineers argue about which one is "best" as if there were a single answer. There is not. The Zone-Based Firewall, the ASA, and Firepower Threat Defense (FTD) are built for different jobs, and the right question is never which is best but which belongs where. This article answers that from an unusual position: PingLabz has captured real command output from all three, so the comparison below is grounded in what these boxes actually do on a lab bench, not in a vendor brochure. It is part of the [Infrastructure Security](https://www.pinglabz.com/infrastructure-security/) cluster, and it draws on the ZBF captures from the [Zone-Based Firewall pillar](https://www.pinglabz.com/zone-based-firewall-ios-xe/) and the appliance captures from the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) cluster. Let me be honest about the differentiator up front: plenty of sites will tell you ZBF, ASA, and FTD apart from a datasheet. The reason to read this one is that PingLabz proves each claim with real output. We watched a cat8000v ZBF inspect live traffic and flip the self zone to default-deny. We ran the ASA's Modular Policy Framework in the [ASA cluster](https://www.pinglabz.com/cisco-asa/). And the FTD platform notes come from hands-on work, not marketing. That is the whole point. ## ZBF: a firewall that is a feature of the router The Zone-Based Firewall is not a product you buy. It is a feature set inside Cisco IOS XE, licensed on with `network-advantage`, that turns a router into a stateful firewall. Its policy engine is Cisco Policy Language (CPL): `class-map type inspect`, `policy-map type inspect`, and a `service-policy` applied to a zone-pair (a directional relationship between two security zones). There is no separate appliance, no separate management plane, and no extra rack unit. The firewall lives in the same device that is already routing your branch. That integration is exactly what makes ZBF the right tool for a branch router that must also firewall. You already have the router terminating the WAN; giving it zone-based policy means the branch enforces segmentation and stateful inspection without a second box to buy, cable, and maintain. The tradeoff is that ZBF is a firewall feature riding on a routing platform, so its scale, throughput, and threat capabilities are those of a router, not those of a dedicated security appliance. It does stateful inspection and application-aware classification well; it is not an intrusion prevention system. For the branch, that is usually exactly the right amount of firewall. ## ASA: the dedicated stateful appliance The Cisco ASA is a purpose-built firewall appliance. Where ZBF is a feature on a router, the ASA is a device whose entire reason for existing is to be a firewall. Its policy model is the Modular Policy Framework (MPF), the same class-map / policy-map / service-policy shape as ZBF's CPL, but applied to interfaces rather than zone-pairs. On top of that it layers two things ZBF does not center its design around: security levels (every interface has a numeric trust level, and traffic from higher to lower is permitted by default while the reverse is denied) and a deeply NAT-centric worldview, where address translation is a first-class part of the policy rather than an afterthought. The ASA is mature. Its CLI has been refined over many years, its behavior is well understood, and its feature set for classic perimeter firewalling (stateful inspection, NAT, VPN termination, security-level segmentation) is thorough. PingLabz captured the MPF and this behavior in the [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/). The honest caveat is that the ASA is a platform Cisco is succeeding with FTD. It remains widely deployed and entirely capable for traditional perimeter and VPN roles, but it is not where Cisco is adding next-generation capability. If your need is classic stateful firewalling and site-to-site VPN on a dedicated box, the ASA still does that job cleanly. ## FTD: the next-generation firewall Firepower Threat Defense is Cisco's next-generation firewall (NGFW), and it is a different category of device. FTD combines stateful firewalling with a Snort-based intrusion prevention system, application awareness (it identifies applications regardless of port, the way an NGFW is expected to), and URL filtering. It is not managed at a device CLI the way ZBF and the ASA are; it is driven by a controller, either the Firepower Management Center (FMC) for centralized multi-device management or the Firepower Device Manager (FDM) for on-box management of a single appliance. FTD is the modern platform, the one Cisco positions as the successor to the ASA, and it is where the threat-centric features live. Snort gives it real IPS. Application and URL awareness give it the Layer 7 visibility that a classic stateful firewall lacks. The cost of all that capability is complexity: FTD is a heavier platform to run, it expects a management controller, and it is more than most branch routers or simple perimeters actually need. PingLabz is building out the FTD material now (the [Cisco FTD](https://www.pinglabz.com/cisco-ftd/) cluster is under construction at the time of writing), grounded in the same real-output approach as the rest of the section. ## Where each one belongs The comparison only matters if it tells you which box to reach for. Here is the placement, drawn from what each platform actually is rather than what its datasheet claims. ZBF (IOS XE) Form factor: a feature on a router, no separate box Policy model: CPL, service-policy on a zone-pair Threat features: stateful inspect, app-aware classification (no IPS) Managed by: the router CLI Belongs at: **a branch router that must also firewall** ASA Form factor: dedicated stateful appliance Policy model: MPF on interfaces, security levels, NAT-centric Threat features: mature stateful firewall + VPN (no NGFW IPS) Managed by: device CLI / ASDM Belongs at: **a classic perimeter or VPN head-end, being succeeded** FTD Form factor: next-generation firewall appliance Policy model: access control policy, threat-centric Threat features: Snort IPS, app/URL awareness, malware Managed by: FMC (centralized) or FDM (on-box) Belongs at: **a modern edge that needs IPS and Layer 7 control** Read the placement row as a decision tree. If the device that needs to firewall is already a router (a branch, a small site with a single WAN box), ZBF gives you stateful policy without adding hardware. If you need a dedicated firewall for classic perimeter duty or VPN termination and NGFW threat features are not a requirement, the ASA still does that job, with the understanding that it is a platform on its way to succession. If you need intrusion prevention, application control, and URL filtering (a real next-generation edge), FTD is the platform built for it, and you accept the management controller and the added weight that come with it. ## Three scenarios, three answers Abstract placement rules are easier to trust when you run them against real situations, so here are three that come up constantly. **A 30-person branch office with one WAN router.** The site already terminates its internet circuit on an IOS XE router. It needs stateful inspection, some segmentation between an inside VLAN and a small DMZ, and management protection on the box itself. Buying a dedicated appliance here is overkill and adds a device to manage. ZBF on the existing router covers it: zone the interfaces, write the inside-to-outside and inside-to-DMZ policies, protect the self zone, and you are done without new hardware. This is the case ZBF was made for. **A data-center perimeter that terminates dozens of site-to-site VPNs.** The requirement is high-throughput stateful firewalling, mature NAT, and lots of IPsec tunnels, but not necessarily Snort-grade intrusion prevention. A dedicated ASA (or an FTD running in a more traditional posture) fits: security levels give you clean trust segmentation, the NAT-centric model handles the translation load, and the VPN feature set is proven. If the organization has standardized on the ASA and does not yet need NGFW threat features, the ASA remains a defensible choice while planning the eventual move to FTD. **An internet edge that must inspect application traffic and block threats.** Here the requirement explicitly includes intrusion prevention, application identification independent of port, and URL filtering. Neither ZBF nor the classic ASA is built for that. FTD is: its Snort engine delivers real IPS, its application awareness sees past ports, and FMC or FDM drives the policy. You accept the management controller and the operational weight because the job genuinely requires next-generation capability. Reaching for ZBF or an ASA here would leave a threat gap the platform was never designed to close. ## The through-line: same MPF shape, different jobs One thing that ties ZBF and the ASA together is worth calling out, because it makes learning both easier. They share the Modular Policy Framework skeleton: classify with a class-map, act with a policy-map, apply with a service-policy. ZBF adds `type inspect` and attaches to a zone-pair; the ASA attaches to an interface and layers security levels on top. If you understand one, the other is a short step, which is exactly why the PingLabz Infrastructure Security section teaches them in sequence. FTD breaks that pattern deliberately, replacing device-level CLI policy with controller-driven access control policy, because its job (threat-centric NGFW enforcement at scale) needs a different management model. ## Key takeaways - **ZBF is a firewall feature on a router.** CPL zone-pairs, no separate appliance, ideal for a branch router that must also firewall. Stateful and app-aware, but not an IPS. - **The ASA is a dedicated stateful appliance.** MPF on interfaces, security levels, NAT-centric, mature CLI, strong at classic perimeter and VPN duty, and being succeeded by FTD. - **FTD is the next-generation firewall.** Snort IPS, application and URL awareness, managed by FMC or FDM. The modern threat-centric platform, and heavier to run. - **Placement, not ranking, is the question.** Branch router that firewalls goes to ZBF; classic perimeter or VPN goes to the ASA; IPS and Layer 7 control at a modern edge goes to FTD. - **The credibility play is real output.** PingLabz shows all three with captured CLI, not brochure copy, so the comparison rests on observed behavior. Choosing the right firewall is a placement decision, and it is one of the payoffs of the PingLabz [Infrastructure Security](https://www.pinglabz.com/infrastructure-security/) cluster. Dig into the router-integrated option in the [Zone-Based Firewall pillar](https://www.pinglabz.com/zone-based-firewall-ios-xe/), the dedicated appliance in the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) cluster, and the next-generation platform in the [Cisco FTD](https://www.pinglabz.com/cisco-ftd/) material as it comes online. Each is backed by real captures, which is the whole reason to trust the comparison. ### Nested Class-Maps in ZBF: Building Layered Policy URL: https://www.pinglabz.com/zbf-nested-class-maps/ Last updated: 2026-07-13T13:26:04.000Z Most Zone-Based Firewall policies you will see are flat: a class-map lists a handful of `match protocol` lines, a policy-map inspects that class, and a zone-pair binds it to a direction. That works, but it does not scale, and it cannot express certain kinds of logic at all. The moment you want to reuse a group of protocols across several policies, or combine "match all of these conditions" with "match any of those protocols" in one class, you need class-maps that reference other class-maps. That is nesting, and it is the difference between a policy you copy-paste and one you actually maintain. This article is part of the PingLabz [Infrastructure Security](https://www.pinglabz.com/infrastructure-security/) cluster and continues the [Zone-Based Firewall pillar](https://www.pinglabz.com/zone-based-firewall-ios-xe/), so the zone and zone-pair concepts here assume you have read that first. Every configuration and every line of show output below came from a real cat8000v running IOS XE 17.18.02 in Cisco Modeling Labs. ## match-all versus match-any: the two logics Before nesting makes sense, you have to be crisp about the two ways a class-map can evaluate its match statements, because nesting is really about combining them. match-any (logical OR) The class matches if **any one** of its match statements is true. Use it to gather a set of protocols: "http OR https OR ftp". This is how you build a reusable protocol group. match-all (logical AND) The class matches only if **every** match statement is true at once. Use it to require several conditions together: "is a web protocol AND comes from this source". A flat class-map picks one of these and lives with it. That is the limitation. A single `match-any` class can say "http or https", and a single `match-all` class can require several conditions, but neither one can express "match ALL of these requirements, where one of the requirements is itself a choice of ANY of these protocols." To combine the two logics in one decision, one class-map has to reference another. That reference is what `match class-map` gives you. ## Building a nested class-map Here is the pattern from the lab. We define a reusable protocol group as a `match-any` child class, then reference it from a `match-all` parent class, and inspect the parent in the DMZ policy: ``` class-map type inspect match-any PLZ-WEB-PROTOCOLS match protocol http match protocol https class-map type inspect match-all PLZ-NESTED-WEB match class-map PLZ-WEB-PROTOCOLS policy-map type inspect PLZ-IN-TO-DMZ class type inspect PLZ-NESTED-WEB inspect ``` Read the structure from the bottom up. The policy-map inspects the class `PLZ-NESTED-WEB`. That class is `match-all`, and its only match statement is `match class-map PLZ-WEB-PROTOCOLS`: it says "match if the child class matches." The child, `PLZ-WEB-PROTOCOLS`, is `match-any` and matches HTTP or HTTPS. So the parent inherits "http OR https" as one of its conditions. Right now the parent has only that one condition, so functionally it behaves like the child, but the structure is what matters: you now have a parent into which you can add more `match-all` requirements alongside the reusable web-protocol group. ## Proof the nesting actually evaluates Nesting is only trustworthy if you can see the router evaluating it, and you can. `show policy-map type inspect zone-pair` against the DMZ zone-pair prints the parent class, its reference to the child, and the child's own matches, indented to show the hierarchy: ``` EDGE1#show policy-map type inspect zone-pair PLZ-ZP-IN-DMZ Class-map: PLZ-NESTED-WEB (match-all) Match: class-map match-any PLZ-WEB-PROTOCOLS Match: protocol http Match: protocol https Inspect Class-map: class-default (match-any) Match: any ``` That output is the whole argument for this article. The top line is the parent, `PLZ-NESTED-WEB`, and IOS labels it `(match-all)`. Below it, indented, is `Match: class-map match-any PLZ-WEB-PROTOCOLS`: the parent's single condition is the child class, and IOS even tells you the child is `match-any`. Indented one level deeper are the child's own two matches, `protocol http` and `protocol https`. The `Inspect` action hangs off the parent. Then class-default sits below to catch everything else. You are looking at the evaluation order made explicit: parent (match-all) references child (match-any), child matches http or https, parent inspects. This is exactly what you want when you troubleshoot a nested policy. You do not have to reason about whether the reference "took"; the runtime output shows the full tree, so you can confirm the child is being consulted and that its protocol list is what you think it is. ## Growing the parent without touching the child The version in the lab keeps the parent minimal so the hierarchy is easy to read, but the reason you build it this way is what comes next. Because the parent is `match-all`, you can add further requirements to it and they combine with AND logic against the web-protocol group already supplied by the child. A parent that today says "match if the traffic is a web protocol" can grow to "match if the traffic is a web protocol and originates from a specific source group," and the child class never changes. That is the structural payoff: the choice of protocols (the OR) lives in the child, the set of mandatory conditions (the AND) lives in the parent, and you extend the parent without disturbing the reusable piece underneath it. Contrast that with a flat class-map. To add a source restriction to a flat `match-any` web class, you cannot simply bolt it on, because adding a match statement to a `match-any` class widens the match (it becomes another OR branch) rather than narrowing it. You would have to rebuild the class as `match-all` and lose the "http OR https" flexibility in the process. Nesting is what lets you keep both behaviors in the same policy: OR where you want choices, AND where you want requirements. The two-level structure is not decoration; it is the only way to hold both logics at once. ## Why nest at all: two concrete payoffs Nesting is more typing than a flat class-map, so it has to earn its keep. It does, in two ways that flat class-maps simply cannot match. ### Reuse: define once, reference everywhere Define `PLZ-WEB-PROTOCOLS` as "http and https" one time, and every policy that needs the web-protocol set references the same child class. When the definition of "web protocols" changes (someone adds a QUIC or an alternate HTTPS port to the standard), you edit one child class-map and every parent that references it inherits the change. In a flat design you would have the same `match protocol http` / `match protocol https` pair copied into the inside-to-DMZ policy, the guest-to-DMZ policy, and the partner-to-DMZ policy, and you would have to remember to update all three. Nesting turns three edits and a chance to forget one into a single authoritative definition. That is the same maintainability instinct behind [port-maps](https://www.pinglabz.com/zbf-port-maps/): change the meaning in one place and let everything that references it follow. ### Combined logic a flat class-map cannot express The second payoff is expressiveness. A flat class-map is one logic type, match-all or match-any, and it applies that single logic to every statement. Nesting lets you compose them. The canonical example: a `match-all` parent that requires "is a web protocol AND comes from this source AND uses this application signature", where "is a web protocol" is itself a `match-any` choice of http or https, supplied by the child. There is no way to write "AND (http OR https)" in a single flat class-map, because a flat class-map cannot hold a nested OR inside an AND. The child class is how you inject that OR into the parent's AND. Once you see that, nesting stops being an academic feature and becomes the natural way to write any policy whose real intent is "all of these conditions, one of which is a set of choices." ## Practical guidance Keep the nesting shallow and the names meaningful. One level of parent-and-child covers almost every real policy, and the show output stays readable at that depth. Name the child class after the thing it represents (a protocol group, a source set) rather than after the policy that happens to use it first, because the whole point is that other policies will reference it later. And when a nested policy misbehaves, go straight to `show policy-map type inspect zone-pair`: the indented tree tells you immediately whether the parent is reaching the child and what the child is matching, which is far faster than reasoning about the config by eye. One caution worth stating plainly: nesting changes where the logic lives, not how much of it there is, so keep the relationship between parent and child obvious in your naming and your documentation. A future engineer reading the policy-map sees a class reference a class-map they then have to go find. If the child is named for what it represents and the parent is named for the decision it makes, that lookup is a two-second confirmation rather than a hunt. Nesting rewards teams that name things well and quietly punishes those who do not, because the indirection that buys you reuse is the same indirection that hides intent when the names are vague. ## Key takeaways - **A class-map can match another class-map** with `match class-map`, letting you build reusable, layered traffic definitions instead of flat one-off lists. - **match-any is OR, match-all is AND.** Nesting exists to combine them: a match-all parent can require several conditions, one of which is a match-any child (a set of protocol choices). - **The show output proves the hierarchy.** `show policy-map type inspect zone-pair` prints the parent (match-all), its reference to the child (match-any), and the child's http/https matches, indented to show evaluation order. - **Reuse is the first payoff:** define a protocol group once as a child class and reference it from every policy, so one edit updates them all. - **Combined logic is the second payoff:** "AND (http OR https)" cannot be written in a single flat class-map; a child class injects the OR into the parent's AND. Nested class-maps are how ZBF policy grows from a demo into something a team can maintain, and they are part of the router-integrated firewall story in the PingLabz [Infrastructure Security](https://www.pinglabz.com/infrastructure-security/) cluster. To see the full policy engine these classes plug into, return to the [Zone-Based Firewall pillar](https://www.pinglabz.com/zone-based-firewall-ios-xe/), and for the classification layer beneath them, read up on [ZBF port-maps](https://www.pinglabz.com/zbf-port-maps/). ### ZBF Port-Maps: Inspecting Applications on Non-Standard Ports URL: https://www.pinglabz.com/zbf-port-maps/ Last updated: 2026-07-13T13:26:03.000Z You have a Zone-Based Firewall inspecting traffic on your Cisco IOS XE router, you write `match protocol http` to inspect web traffic to your DMZ, and one particular application just will not pass. The app runs on TCP 8080, not 80, and to the ZBF protocol inspector it may as well be invisible. This is not a bug. It is the single most common surprise engineers hit once they move past the basics of ZBF, and the fix is a port-map. This article is part of the PingLabz [Infrastructure Security](https://www.pinglabz.com/infrastructure-security/) cluster and builds directly on the [Zone-Based Firewall pillar](https://www.pinglabz.com/zone-based-firewall-ios-xe/), so if you have not met zones and zone-pairs yet, start there and come back. Everything below was captured from a real cat8000v running IOS XE 17.18.02 in Cisco Modeling Labs. No output has been invented. ## Protocol classification is port-based by default When you write `match protocol http` in an inspect class-map, you are asking IOS to recognize HTTP. The obvious question is: how does the router know a packet is HTTP? For the well-known protocols, the answer is simpler than you might hope. IOS ships with a built-in table that maps each named protocol to its standard port, and `http` is defined as TCP port 80\. So `match protocol http` really means "match TCP traffic on port 80" (plus some application-aware behavior on top for protocols that support it). That default is fine right up until an application does not use the standard port. A management console on 8080, an internal API on 8000, a legacy web app someone stood up on 8888: all of them are HTTP by every meaningful definition, but none of them are on port 80, so `match protocol http` does not classify them. The inspect policy never sees them as HTTP, the class does not match, and (assuming class-default drops) the traffic is dropped. You stare at a policy that looks correct and traffic that will not pass, because the mismatch is not in your policy logic. It is in the port table underneath it. ## The fix: ip port-map The `ip port-map` command lets you extend (or narrow) the port definition for any protocol. To tell IOS that HTTP also lives on TCP 8080, you add one line in global configuration: ``` EDGE1(config)#ip port-map http port tcp 8080 ``` That is the entire fix. You have not touched the class-map, the policy-map, or the zone-pair. You have changed the definition of what "http" means at the classification layer, and every `match protocol http` statement on the box immediately inherits the new port. Verify it with `show ip port-map http`: ``` EDGE1#show ip port-map http Default mapping: http tcp port 80 system defined Default mapping: http tcp port 8080 user defined ``` Read that output carefully, because it tells the whole story. There are now two entries for HTTP. The first, TCP port 80, is marked `system defined`: it is the built-in default that shipped with the image, and adding your own port did not remove it. The second, TCP port 8080, is marked `user defined`: that is your addition. From this moment on, `match protocol http` in any inspect class-map matches HTTP on both port 80 and port 8080\. Your DMZ app on 8080 is now visible to the firewall, gets Layer 7 HTTP inspection, and passes the policy it was silently failing before. The `system defined` versus `user defined` labeling is worth internalizing. Port-maps are additive by default: your custom port sits alongside the standard one, not on top of it. That means enabling inspection for an application on a non-standard port does not disable inspection on the standard port, so you do not accidentally break the rest of your web traffic while fixing the one app. ## When you actually reach for a port-map There are two situations where port-maps stop being trivia and become the thing that unblocks your afternoon. App on a non-standard port An internal web app, API, or admin console runs on 8080 / 8000 / 8888 but you still want it Layer 7 inspected as HTTP. Add the port with `ip port-map http port tcp 8080` and the existing `match protocol http` covers it. Narrowing a protocol with a list You can scope a port-map to a specific set of hosts or ports with an ACL using the `list ` keyword, so the custom mapping only applies where you intend rather than globally. The first case is the everyday one: some service is not on its textbook port and you want the firewall to treat it as the protocol it really is. Instead of writing a crude `match protocol tcp` that matches all TCP indiscriminately, you keep the specific, application-aware `match protocol http` and simply teach IOS the extra port. You preserve the Layer 7 awareness (and the inspection quality that comes with it) rather than falling back to a blunt port-agnostic match. The second case is the inverse and is where the `list` keyword earns its place. A port-map does not have to be a broad "this port is now this protocol everywhere" statement. Attach an ACL with the `list` option and the mapping applies only to the traffic that ACL permits, letting you say "treat TCP 8080 as HTTP, but only to these DMZ servers" rather than globally reclassifying 8080 across the whole router. That precision keeps a port-map from becoming a policy side effect you forget about six months later. ## The syntax generalizes to any protocol Nothing about port-maps is HTTP-specific. The `ip port-map` command takes any protocol name IOS knows and any transport and port you give it, so the same technique that rescued a web app on 8080 works for a mail service, a database, or a custom TCP application you want inspected as a named protocol rather than as anonymous TCP. The shape is always `ip port-map port {tcp | udp} `, and the corresponding `show ip port-map ` always distinguishes the built-in `system defined` entries from your `user defined` ones the same way you saw for HTTP. Because the mapping is additive and lives at the classification layer, it is also a low-risk change to make. You are not editing an active policy-map or bouncing a zone-pair; you are adding a row to a lookup table that class-maps consult. Existing inspection for the standard port keeps working, and the new port simply starts matching. That safety is exactly why the port-map is the right tool for "make the firewall see this app as HTTP" rather than loosening a class-map to match all TCP, which would surrender the Layer 7 awareness you configured ZBF to get in the first place. ## How port-maps fit the wider ZBF picture Port-maps live at the classification layer, one level below the class-map. The chain runs: the port-map table defines what each protocol name means; the `class-map type inspect` uses `match protocol` against that table; the `policy-map type inspect` decides what to do with the class; and the zone-pair binds that policy to a direction. When traffic mysteriously fails to match a protocol you are certain it should, the port-map table is the first place to look, because a class-map can only match what the classifier can see. This is exactly the kind of layered structure that [nested class-maps](https://www.pinglabz.com/zbf-nested-class-maps/) also lean on, and it is why understanding classification pays off across every ZBF policy you write. A practical habit: when you adopt a policy from documentation or a previous engineer and traffic on an unusual port will not pass, run `show ip port-map ` before you touch anything else. It takes five seconds and it tells you immediately whether the classifier even knows about the port in question. Half the "ZBF is dropping my traffic" tickets are really "the protocol is on a port the classifier was never told about" tickets. ## Key takeaways - **ZBF protocol matching is port-based.** `match protocol http` only recognizes HTTP on the system-defined port 80 by default, so an app on 8080 is invisible to the HTTP inspector. - **One command fixes it:** `ip port-map http port tcp 8080` teaches IOS that 8080 is also HTTP, and every existing `match protocol http` immediately covers it. - **Port-maps are additive.** `show ip port-map http` shows the `system defined` port 80 and your `user defined` port 8080 side by side; adding a port does not remove the standard one. - **Use them two ways:** to Layer 7 inspect an app on a non-standard port, or (with the `list ` keyword) to narrow a protocol mapping to specific hosts instead of applying it globally. - **Check the classifier first.** When ZBF drops traffic that "should" match a protocol, `show ip port-map` often reveals the port was never mapped in the first place. Port-maps are a small but high-leverage piece of the router-integrated firewall story in the PingLabz [Infrastructure Security](https://www.pinglabz.com/infrastructure-security/) cluster. To see where this classification work sits in the full policy engine, return to the [Zone-Based Firewall pillar](https://www.pinglabz.com/zone-based-firewall-ios-xe/), then move on to [nested class-maps for layered policy](https://www.pinglabz.com/zbf-nested-class-maps/). ### Zone-Based Firewall on Cisco IOS XE: Zones, Zone-Pairs, and the Policy Model URL: https://www.pinglabz.com/zone-based-firewall-ios-xe/ Last updated: 2026-07-13T13:26:03.000Z The Zone-Based Firewall (ZBF) is the stateful firewall that lives inside Cisco IOS XE. It is not a bolt-on appliance and it is not a simple access list. It is a full policy engine that classifies traffic, inspects it statefully, and permits or drops it based on which security zone the traffic entered from and which zone it is trying to reach. If you run a branch router that also has to act as a firewall, ZBF is how you do it without buying another box. This guide is part of the PingLabz [Infrastructure Security](https://www.pinglabz.com/infrastructure-security/) cluster, and it is the anchor article for everything ZBF: the policy model, the config, the mental model, and the one trap that locks more engineers out of their own routers than any other. Every command and every line of output on this page came from a real Cisco Modeling Labs (CML) topology, not from memory. The firewall is a cat8000v running IOS XE 17.18.02 (EDGE1). An outside client (ISP1, an IOL-XE router at 203.0.113.2) sits on the untrusted side. An inside host (HOST1) lives on 10.20.10.0/24, and a DMZ web server (nginx) answers on 172.20.30.80\. The router has three data interfaces plus its own control plane, and by the end of this article you will see all four zones and understand exactly how traffic moves between them. ## A real gotcha before you type a single command: the platform ZBF is a licensed feature, and not every IOS XE image carries it. On an IOL-XE virtual router (the lightweight IOS-on-Linux platform many of us lab on), the very first ZBF command fails: ``` EDGE1(config)#zone security PLZ-INSIDE ^ % Invalid input detected at '^' marker. ``` That is not a syntax error you can work around. The Zone-Based Firewall feature set simply is not present in the IOL-XE image. We rebuilt EDGE1 on a cat8000v (the virtual Catalyst 8000V) and hit a second wall: on a fresh cat8000v the License Level is empty, and `zone security` is *still* rejected. The feature only appears after you set the license and reload: ``` license boot level network-advantage ! (place this line in the day0 / startup config, then reload) ``` After the reload the router reports `License Level: network-advantage` with an add-on of `dna-advantage`, and every ZBF command is accepted. If you have been fighting an "invalid input" error on `zone security`, the answer is almost never your typing. It is the platform and the license. Lab this on a cat8000v with network-advantage, not on IOL-XE. ## The policy model: CPL, and how it differs from the ASA MPF If you have configured a Cisco ASA, the ZBF building blocks will feel familiar because they use the same three-layer Modular Policy Framework shape: a class-map to classify, a policy-map to act, and a service-policy to apply. ZBF calls this Cisco Policy Language (CPL). The difference is in two words and one attachment point. First, every ZBF construct carries the keyword `type inspect`. A ZBF class-map is a `class-map type inspect`; a ZBF policy-map is a `policy-map type inspect`. That keyword is what makes the firewall stateful (it tracks the connection and permits the return traffic automatically). Second, and this is the real conceptual jump, you do not apply the service-policy to an interface the way the ASA MPF does. You apply it to a **zone-pair**: a named, directional relationship between a source zone and a destination zone. The ASA attaches policy to `interface outside`; ZBF attaches policy to "traffic going from the inside zone to the outside zone". PingLabz covers the appliance side of this in the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) cluster, so if you want the MPF-on-an-interface version for contrast, that is where it lives. Here is the real, working configuration from the lab. Read it top to bottom: three zones, a class-map that matches general TCP/UDP/ICMP traffic, a policy that inspects that class and drops everything else, a zone-pair that binds the policy to a direction, and finally an interface joined to a zone. ``` zone security PLZ-INSIDE zone security PLZ-DMZ zone security PLZ-OUTSIDE class-map type inspect match-any PLZ-GEN-TRAFFIC match protocol tcp match protocol udp match protocol icmp policy-map type inspect PLZ-IN-TO-OUT class type inspect PLZ-GEN-TRAFFIC inspect class class-default drop log zone-pair security PLZ-ZP-IN-OUT source PLZ-INSIDE destination PLZ-OUTSIDE service-policy type inspect PLZ-IN-TO-OUT interface GigabitEthernet2 zone-member security PLZ-INSIDE ``` Notice the `class class-default` with `drop log`. In a ZBF policy-map, class-default is the implicit catch-all, and if you do not add it the firewall still drops unmatched traffic. Making it explicit with `log` just gives you the syslog record of what got dropped. An interface with no `zone-member` statement is not in any zone, and traffic between a zoned interface and an unzoned interface is dropped outright. ## The ZBF mental model, in three rules Before you look at any output, internalize the three rules that govern every ZBF decision. Everything else is detail. Between two zones Traffic from one zone to another is **denied** unless a zone-pair with a permit or inspect policy explicitly allows it. No zone-pair means no traffic. Within one zone Traffic between two interfaces in the **same** zone is allowed by default. You do not write a policy for intra-zone flows. To a memberless zone Traffic aimed at a zone that has **no member interfaces** is dropped. A zone with no interfaces is a black hole. The design consequence is that ZBF is default-deny between zones. You build a firewall by deciding which directional flows to permit, not which to block. That is the opposite instinct from writing a deny ACL, and it is why ZBF policies read as a short list of "here is what is allowed" rather than a long list of "here is what is forbidden". ## Why a zone-pair instead of an interface ACL You could ask why any of this is better than hanging an access list on an interface, the way engineers have firewalled routers for decades. Two reasons. The first is state. A traditional ACL is a static packet filter: to allow return traffic you either open the reply ports permanently (widening your attack surface) or maintain reflexive ACLs by hand. A ZBF inspect policy tracks each connection and opens the return path only for the life of that flow, then closes it. You saw that above: the inside pinged out and the reply came back with zero inbound rules, because inspection handled the return automatically. The second reason is that a zone-pair expresses intent at the level you actually think in. "Inside can reach the DMZ" is a policy statement about two groups of interfaces, and if you add a fourth interface to the inside zone tomorrow, it inherits every inside policy for free. An ACL is tied to one interface in one direction; a zone-pair is tied to a relationship. That is why ZBF scales to a branch router with several inside VLANs, a DMZ, and an uplink without turning into a wall of per-interface ACLs that nobody dares touch. ## Seeing the zones: show zone security With the interfaces assigned, `show zone security` lists every zone the router knows about. The three we created are there, and so are two you never configured: ``` EDGE1#show zone security zone self 65535 Description: System defined zone zone service 65534 Description: System defined zone zone PLZ-INSIDE 1 Member Interfaces: GigabitEthernet2 zone PLZ-DMZ 2 Member Interfaces: GigabitEthernet3 zone PLZ-OUTSIDE 3 Member Interfaces: GigabitEthernet1 ``` The two system zones matter. `self` (zone ID 65535) is the router itself: its control and management plane, every packet sourced by or destined to the box, including SSH sessions to it, ICMP pings of its interfaces, and routing-protocol hellos. `service` (zone ID 65534) is an internal system zone. You will spend most of your ZBF life ignoring the service zone and a great deal of it thinking hard about the self zone, which we come to shortly. The companion command shows the directional bindings you built: ``` EDGE1#show zone-pair security Zone-pair name PLZ-ZP-IN-DMZ 2 Source-Zone PLZ-INSIDE Destination-Zone PLZ-DMZ service-policy PLZ-IN-TO-DMZ Zone-pair name PLZ-ZP-IN-OUT 1 Source-Zone PLZ-INSIDE Destination-Zone PLZ-OUTSIDE service-policy PLZ-IN-TO-OUT ``` Two zone-pairs, both sourced from PLZ-INSIDE: one to the outside, one to the DMZ. There is deliberately no zone-pair from outside to inside, which is exactly why the internet cannot start a connection to HOST1\. The inside can reach out; the outside cannot reach in. That asymmetry is the entire point of a stateful firewall, and here it is expressed as "which zone-pairs exist". ## Proof it inspects: real traffic through the firewall A policy model is only worth anything if the packets actually flow. From HOST1 on the inside, traffic to the outside and to the DMZ both succeed, and because the policy inspects statefully, the return traffic is permitted automatically without any inbound zone-pair: ``` HOST1$ ping -c 3 203.0.113.2 -> 3 received, 0% packet loss HOST1$ ping -c 2 192.168.99.1 -> 2 received, 0% packet loss HOST1$ curl http://172.20.30.80/ -> DMZ HTTP 200 ``` The first ping reaches the outside client, the second reaches all the way to the real upstream gateway beyond it (the traffic is being inspected and forwarded through the outside zone), and the curl reaches the DMZ web server and gets an HTTP 200 back. Inside-to-outside and inside-to-DMZ both work because both directions have a zone-pair with an inspect policy. Nothing about the return path needed its own rule. You can watch the policy itself doing the work with `show policy-map type inspect zone-pair`, which prints the runtime view of every zone-pair, its class-maps, and the action: ``` EDGE1#show policy-map type inspect zone-pair Zone-pair: PLZ-ZP-IN-DMZ Class-map: PLZ-GEN-TRAFFIC (match-any) Inspect Class-map: class-default (match-any) Zone-pair: PLZ-ZP-IN-OUT Class-map: PLZ-GEN-TRAFFIC (match-any) Inspect Class-map: class-default (match-any) ``` Each zone-pair shows the same shape: the named class is inspected, and class-default (the implicit catch-all) sits below it to drop the rest. This is your primary ZBF troubleshooting command. If traffic is not passing, this output tells you whether the class is matching and what action is bound to it. ## The self zone: the lesson everyone learns the hard way Here is the single most important thing in this article, and the reason to read the self zone section before you ever touch a production router. By default, traffic to and from the `self` zone is **permitted**, even after you have carefully zoned every data interface. The router will still answer SSH, still reply to pings on its interfaces, and still form routing adjacencies, because you have not written any policy that touches the self zone. Proof, from the outside client pinging the router's outside interface: ``` ISP1#ping 203.0.113.1 repeat 3 !!! Success rate is 100 percent (3/3) ``` One hundred percent. The outside can ping the firewall itself, even though the outside cannot reach a single host behind it. That is the default self-zone behavior: open, because no self zone-pair exists yet. Now the trap. The **moment** you create any zone-pair that involves the self zone, that direction flips from default-permit to default-deny, and you must then explicitly permit everything you still want. It is all-or-nothing per direction. To lock management down to SSH only, we added an outside-to-self policy that inspects SSH and drops the rest: ``` class-map type inspect match-any PLZ-MGMT-ONLY match protocol ssh policy-map type inspect PLZ-OUT-TO-SELF class type inspect PLZ-MGMT-ONLY inspect class class-default drop zone-pair security PLZ-ZP-OUT-SELF source PLZ-OUTSIDE destination self service-policy type inspect PLZ-OUT-TO-SELF ``` That looks reasonable: "let SSH in to the box, drop everything else." But watch what it does to the ping that was working seconds ago: ``` ISP1#ping 203.0.113.1 repeat 3 ... Success rate is 0 percent (0/3) ``` From 100 percent to 0 percent, instantly, without touching ICMP at all. By creating an outside-to-self zone-pair for SSH, you switched the whole outside-to-self direction to default-deny, and ICMP is not in your permit list, so ICMP is now dropped. If your management plan had relied on ping to confirm the box was alive, on a routing protocol peering across that interface, or on anything other than the one protocol you named, all of it just went dark. This is how engineers lock themselves out with ZBF: they add a self zone-pair to allow SSH from a management network, apply it, and discover they have silently denied their own return traffic, their monitoring, and sometimes the very SSH session they are typing in. The fix is discipline, not luck. When you write a self zone-pair, enumerate *every* protocol you need to reach the box (SSH, and probably ICMP, and any routing protocol or management service that terminates on the router) in that policy, because the instant the zone-pair exists, anything you did not name is denied. Test it from a second path, never from the session you are about to cut. The self zone is the most powerful and the most dangerous part of ZBF precisely because it defends the router's own brain. ## Key takeaways - **ZBF is a licensed IOS XE feature, not a universal one.** It is absent on IOL-XE (`zone security` returns invalid input) and dormant on a fresh cat8000v until you set `license boot level network-advantage` and reload. - **The policy model is CPL: class-map, policy-map, service-policy, all with `type inspect`.** Same shape as the ASA MPF, but applied to a directional [zone-pair](https://www.pinglabz.com/cisco-asa/) instead of an interface. - **ZBF is default-deny between zones.** Traffic between zones is dropped unless a zone-pair permits it; traffic within a zone is allowed; traffic to a memberless zone is a black hole. - **Two system zones always exist:** `self` (65535, the router's own control plane) and `service` (65534). `show zone security` and `show zone-pair security` are how you read the current state. - **The self zone starts open and snaps shut.** Self traffic is permitted until you create any self zone-pair, at which point that direction becomes default-deny and every protocol you forgot to permit (ICMP, routing, your own ping) is dropped. This is the number-one ZBF lockout. ZBF is the router-integrated firewall in the PingLabz [Infrastructure Security](https://www.pinglabz.com/infrastructure-security/) cluster. From here, the ZBF deep-dives get more specific: [port-maps for inspecting applications on non-standard ports](https://www.pinglabz.com/zbf-port-maps/), [nested class-maps for layered policy](https://www.pinglabz.com/zbf-nested-class-maps/), and a head-to-head on [where ZBF, the ASA, and FTD each belong](https://www.pinglabz.com/zbf-vs-asa-vs-ftd/). All of them are built on the same live cat8000v you just watched inspect real traffic. ### MPLS Traffic Engineering Explained URL: https://www.pinglabz.com/mpls-traffic-engineering/ Last updated: 2026-07-13T13:00:00.000Z MPLS Traffic Engineering exists to solve a problem the IGP creates. OSPF and IS-IS always forward along the shortest path, every flow, all the time. That is correct for reachability and terrible for capacity: the shortest path saturates while parallel links sit idle. MPLS-TE is how you override the IGP and steer traffic onto the paths *you* choose. This post explains the problem, the components, and why segment routing is now changing how it is done. For the cluster overview, see the [MPLS complete guide](https://www.pinglabz.com/mpls/). ## The problem: the shortest path is not the only path Picture two paths between a pair of core routers - a direct link and a slightly longer detour. The IGP computes the direct link as shorter, so every flow takes it. Under load, the direct link hits 95 percent utilization and starts dropping packets while the detour link sits at 10 percent. The capacity exists. The IGP simply will not use it, because "longer" disqualifies it regardless of how congested "shorter" is. You could fudge IGP metrics to shift traffic, but metric tuning is blunt - it moves *all* traffic for affected destinations and tends to just relocate the congestion. MPLS-TE is the precise tool: it lets you place specific traffic on specific paths, with bandwidth accounted for. ## What MPLS-TE does MPLS-TE builds explicit label-switched paths - **TE tunnels** \- across the network. A tunnel is a one-way LSP from a headend router to a tailend router, and it can be told to follow a path the IGP would never choose. Traffic mapped into the tunnel rides that path. The MPLS data plane forwards it by label, exactly as ordinary MPLS does; what changes is that the path was chosen by TE constraints, not by shortest-path SPF. ## The three components IGP TE extensions Flood per-link TE attributes - available bandwidth, admin colors - so every router has a TE topology database, not just a reachability database CSPF Constrained Shortest Path First - runs at the headend to compute a path that satisfies the tunnel's constraints RSVP-TE Signals the chosen path hop by hop, reserves bandwidth on each link, and distributes the labels for the LSP ## The TE database A normal IGP floods topology and reachability. For TE, OSPF and IS-IS are extended to also flood **resource** information about each link: how much bandwidth is configured for TE, how much is currently reserved, and any administrative "colors" (also called affinities) tagging the link. Every router collects this into a TE database that describes not just the shape of the network but its available capacity. ## CSPF: shortest path with rules The headend of a tunnel runs **CSPF** against the TE database. Ordinary SPF asks "what is the shortest path?" CSPF asks "what is the shortest path that *also* has at least 200 Mbps free, and avoids any link colored red?" It prunes every link that fails a constraint, then finds the shortest path through what remains. The output is an explicit hop list. Typical constraints are required bandwidth, link affinities (include or exclude links by color), and an explicit hop list you supply directly when you want a fully specified path. ## RSVP-TE: signaling and reservation Once CSPF produces a path, **RSVP-TE** builds it. A PATH message travels headend to tailend along the chosen hops; a RESV message travels back, and as it returns each router reserves the requested bandwidth on its outgoing link and allocates a label. The result is a signaled LSP with bandwidth booked end to end. That bandwidth accounting is the point. RSVP-TE reservations are why the network knows a link has 200 Mbps already committed, so the next tunnel's CSPF will not oversubscribe it. The cost is **state**: every router along every tunnel holds per-tunnel RSVP state and refreshes it. In a large network with many tunnels, that state grows fast - the well-known scaling weakness of classic MPLS-TE. ## Fast ReRoute A major reason to deploy MPLS-TE is **Fast ReRoute** (FRR). FRR pre-computes and pre-signals a backup tunnel around a protected link or node. When the protected resource fails, the router at the point of failure switches traffic onto the pre-built backup in well under 50 milliseconds - long before the IGP would even notice the failure. For voice, video, and other loss-sensitive traffic, sub-50ms protection is the headline feature. ## The configuration shape on Cisco IOS XE TE is enabled globally and per interface, then tunnels are defined at the headend: ``` ! Global + per-link enablement mpls traffic-eng tunnels ! interface GigabitEthernet0/1 mpls traffic-eng tunnels ip rsvp bandwidth 500000 ! router ospf 1 mpls traffic-eng router-id Loopback0 mpls traffic-eng area 0 ! ! A TE tunnel at the headend interface Tunnel10 ip unnumbered Loopback0 tunnel mode mpls traffic-eng tunnel destination 10.255.0.5 tunnel mpls traffic-eng bandwidth 200000 tunnel mpls traffic-eng path-option 10 dynamic ``` The `dynamic` path-option lets CSPF compute the path; an `explicit` path-option with a named hop list pins it. Verify with `show mpls traffic-eng tunnels`. ## The modern shift: segment routing Classic MPLS-TE works, but the RSVP-TE per-tunnel state does not scale gracefully. This is exactly the problem **SR-MPLS** (Segment Routing) addresses: the path is encoded as a label stack pushed at the headend, and core routers hold no per-tunnel state at all. New traffic-engineering deployments increasingly use SR Policy instead of RSVP-TE, while keeping the same goals - explicit paths, bandwidth awareness, fast protection via TI-LFA. If you are designing TE today, see the [segment routing](https://www.pinglabz.com/segment-routing-mpls/) article; if you are running an existing RSVP-TE network, the concepts here are what you are operating. ## Common gotchas A tunnel will not come up CSPF found no path meeting the constraints - usually requested bandwidth exceeds what any path has free. TE tunnel is up but no traffic uses it Nothing is mapped into it. Traffic needs autoroute, a static route, or a forwarding-adjacency to enter the tunnel. Tunnel reroutes onto a worse path under load The primary path lost enough free bandwidth to fail CSPF; the tunnel re-optimized to a path that still fits. RSVP state growing unmanageably Inherent to RSVP-TE at scale. This is the case for migrating to SR Policy. FRR not protecting a failure No backup tunnel was configured for that link or node, or the protected interface lacks `tunnel mpls traffic-eng fast-reroute`. ## Key takeaways MPLS Traffic Engineering overrides the IGP's shortest-path-only behaviour so you can place specific traffic on specific paths and stop congesting one link while parallel capacity sits idle. It works through three pieces: IGP TE extensions that flood per-link bandwidth and colors into a TE database, CSPF at the headend that finds the shortest path meeting your constraints, and RSVP-TE that signals the path and reserves bandwidth hop by hop. Fast ReRoute gives sub-50ms protection by pre-signaling backup tunnels. The trade-off is RSVP per-tunnel state, which does not scale gracefully - and that is why modern TE deployments are moving to segment routing, which keeps the goals while leaving core routers stateless. For the MPLS cluster, see the [MPLS pillar](https://www.pinglabz.com/mpls/). ### Clientless SSL VPN Is Gone: What Replaced WebVPN on Modern ASA URL: https://www.pinglabz.com/asa-clientless-webvpn-removed/ Last updated: 2026-07-13T12:24:05.000Z If you are studying an older CCIE or CCNP security blueprint, or working from a tutorial written a few years ago, you will find pages of material on Clientless SSL VPN (WebVPN) on the Cisco ASA: browser-based portals, bookmarks, port forwarding, plug-ins, smart tunnels. Here is the thing nobody tells you until you try it on a current box: on modern ASA software that feature is gone. Not deprecated-but-present, not hidden behind a license, actually removed from the command parser. This article proves it with real CLI output from an ASAv running 9.24(1), explains exactly when and why it happened, and shows what replaced it. It is part of the PingLabz [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series. The short version: Cisco deprecated Clientless SSL VPN effective ASA 9.17(1), and by 9.24(1) the clientless-specific commands are removed. If you are about to spend hours learning bookmark lists and port-forward configs, stop and read this first. You would be studying a dead feature. ## What Clientless SSL VPN Was Clientless SSL VPN, also called WebVPN, let a remote user reach internal resources through a web browser with no VPN client installed. They browsed to the ASA's portal over HTTPS, authenticated, and got a page of bookmarks: internal web apps, file shares, and, through port forwarding and smart tunnels, even some thick-client applications proxied through the browser. It was genuinely useful in the era of locked-down corporate desktops and kiosk access, because it required nothing on the endpoint but a browser. It also carried real baggage. Rewriting arbitrary internal web apps to work through a proxy portal was fragile and broke constantly as web apps grew more complex, the Java and ActiveX components it leaned on became security liabilities, and the whole model aged badly against modern web technology. That combination is why Cisco drew a line under it. ## The Proof: vpn-tunnel-protocol No Longer Offers Clientless Talk is cheap, so let the box speak. A group policy on the ASA declares which VPN protocols its users may use with `vpn-tunnel-protocol`. On older code, the option list for that command included an `ssl-clientless` keyword. Ask the 9.24(1) parser what it accepts now. ``` FW1(config-group-policy)# vpn-tunnel-protocol ? group-policy mode commands/options: ikev1 IKE version 1 ikev2 IKE version 2 l2tp-ipsec L2TP using IPSec for security ssl-client SSL VPN Client ``` Read that list carefully. There are exactly four options: `ikev1`, `ikev2`, `l2tp-ipsec`, and `ssl-client`. The `ssl-client` entry is the full-client Secure Client (AnyConnect) tunnel, which is very much alive. What is missing is `ssl-clientless`. It used to be right there in this list, and on 9.24(1) it simply is not offered any more. The parser will not even suggest it. Try to set it anyway, spelling out what used to be a valid keyword, and the ASA rejects it. ``` FW1(config-group-policy)# vpn-tunnel-protocol ssl-clientless ERROR: % Incomplete command ``` "% Incomplete command" is the parser telling you the token `ssl-clientless` is not recognized as a valid completion, so the line reads as unfinished. This is not a licensing message or a "feature disabled" notice. The keyword is gone from the grammar. ## The Clientless Feature Commands Are Removed Too It is not just the protocol keyword. The commands that configured clientless features, the ones every old tutorial walks you through, are removed from the WebVPN configuration mode as well. Two of the most recognizable are port forwarding and bookmark lists. ``` FW1(config-webvpn)# port-forward PLZ-PF 3389 10.20.10.100 3389 ^ ERROR: % Invalid input detected at '^' marker. FW1(config-webvpn)# bookmark-list ^ ERROR: % Invalid input detected at '^' marker. ``` The caret sits right under the command word. "% Invalid input detected at '^' marker" means the parser did not recognize the command at that position at all. `port-forward`, which used to proxy a TCP application (here RDP on 3389) through the clientless portal, and `bookmark-list`, which built the menu of links the portal presented, are both simply not commands the 9.24(1) ASA understands. Every clientless-specific building block is in the same state. ## But the webvpn Container Still Exists One subtlety trips people up here, and the captures show it. Notice that the prompt above is `FW1(config-webvpn)#`. The `webvpn` configuration mode still exists. You can still enter it. So did they only half-remove the feature? No. The `webvpn` container survives because Secure Client (the client formerly known as AnyConnect) uses it. The full-client SSL VPN and its download images, profiles, and settings live under `webvpn`, so the mode has to stay. What is gone is the clientless-specific subset: the portal, bookmarks, port forwarding, smart tunnels, and the `ssl-clientless` protocol option. The container remains for the living feature; the dead feature's commands were pulled out of it. That is why you can enter `config-webvpn` but cannot type `bookmark-list` once you are there. ## The Timeline: 9.17 Out, 9.24 Confirmed The dates matter, because they determine whether your box is affected. Cisco announced deprecation of Clientless SSL VPN effective ASA software release 9.17(1). Deprecation is the warning shot: the feature is on notice and you should stop deploying it. Removal followed, and by the 9.24(1) release running in this lab the clientless commands are gone from the parser, exactly as the output above proves. If you are on a release at or beyond 9.17, treat clientless as end-of-life; if you are already on 9.18 or later, expect it to be absent. This is also why chasing a walkthrough of the feature is a poor use of study time in 2026\. The exam blueprints and the shipping software have both moved. Understanding that clientless was removed, when, and what took its place is the current, testable knowledge. A step-by-step of how to build a bookmark list is knowledge of a feature you cannot configure on a supported box. ## What Replaced It: Secure Client Remote Access The replacement is full-client remote-access VPN with Cisco Secure Client (the product's name since the AnyConnect rebrand). Instead of proxying resources through a browser portal, the user runs a lightweight client that builds a real tunnel to the ASA, and their traffic reaches internal resources natively. It is more capable, more secure, and it does not depend on the fragile web-rewriting that sank the clientless model. PingLabz already covers both flavors of it end to end: [Secure Client SSL VPN](https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/) Transport: TLS/DTLS (the `ssl-client` protocol) Use when: you want the classic full-client SSL RA-VPN [Secure Client IKEv2 VPN](https://www.pinglabz.com/cisco-asa-anyconnect-ikev2-vpn/) Transport: IPsec/IKEv2 (the `ikev2` protocol) Use when: you want IPsec-based remote access Notice that both replacements map directly to keywords that *are* still in that `vpn-tunnel-protocol` list: `ssl-client` and `ikev2`. The ASA did not lose remote-access VPN, it lost the browser-portal style of it. If your remote users still need to reach a handful of internal web apps without a client at all, that use case now belongs to a reverse proxy or a Zero Trust Network Access (ZTNA) product, not to the ASA. For the IPsec fundamentals underneath the IKEv2 option, the [IPsec VPN](https://www.pinglabz.com/ipsec-vpn/) hub is the place to start. ## Migration Checklist for a Pre-9.17 Box If you are still running clientless SSL VPN on an older ASA, here is the path off it before an upgrade removes the feature under you. 1. **Inventory what clientless is actually doing.** List every bookmark, port-forward, and smart-tunnel entry and map it to the real resource behind it. You are cataloguing what users reach, not the config. 2. **Sort each resource by replacement.** Full network access goes to Secure Client SSL or IKEv2 RA-VPN. Web-app-only access with no client goes to a reverse proxy or ZTNA. Thick-client-over-browser (port forwarding) almost always becomes full-client access. 3. **Stand up Secure Client RA-VPN in parallel.** Build the [SSL](https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/) or [IKEv2](https://www.pinglabz.com/cisco-asa-anyconnect-ikev2-vpn/) remote-access profile alongside the running clientless config so you can pilot before cutover. 4. **Migrate users in waves, then decommission.** Move groups over, confirm they can reach their resources through the client, and only then remove the clientless configuration. 5. **Do not upgrade past 9.17 with a live dependency.** If clientless is still load-bearing, complete the migration before you cross the version boundary that deletes the commands, or the upgrade will break access. ## Common Questions ### Is AnyConnect also gone? No, and this is the most common confusion. AnyConnect was rebranded to Cisco Secure Client, but the full-client SSL and IKEv2 remote-access VPN it provides is alive and is the recommended replacement. Only the clientless, browser-portal style of SSL VPN was removed. The `ssl-client` keyword still in the `vpn-tunnel-protocol` list is exactly that living full-client feature. ### Can I still reach internal web apps without installing a client? Not through the ASA. That specific use case, browser-only access to internal web applications, is what clientless served, and it now belongs to a reverse proxy or a Zero Trust Network Access (ZTNA) product. The ASA's job is network-layer remote access via Secure Client. ### How do I know if my ASA still has clientless? Check your software version and ask the parser. If you are below 9.17, clientless is likely present but deprecated. On 9.17 and later, run `vpn-tunnel-protocol ?` in group-policy mode: if the list shows only ikev1, ikev2, l2tp-ipsec, and ssl-client, clientless is gone, exactly as on the 9.24(1) box here. ### Why did Cisco remove it instead of just leaving it? Clientless relied on rewriting arbitrary internal web applications to run through a portal proxy, plus Java and ActiveX components. That model became increasingly fragile and a security liability as web technology evolved. Rather than carry a brittle, hard-to-secure feature indefinitely, Cisco deprecated it in 9.17 and moved customers to the full-client and ZTNA models that handle modern access more safely. ## Key Takeaways - **Clientless SSL VPN (WebVPN) is removed:** Cisco deprecated it effective ASA 9.17(1), and on the 9.24(1) box in this lab the clientless commands are gone from the parser. - **The CLI proves it:** `vpn-tunnel-protocol ?` lists only ikev1, ikev2, l2tp-ipsec, and ssl-client (no ssl-clientless); `vpn-tunnel-protocol ssl-clientless` returns "% Incomplete command"; and `port-forward` and `bookmark-list` return "% Invalid input." - **The `webvpn` container survives:** Secure Client (AnyConnect) still uses it, so you can enter config-webvpn, but the clientless-specific commands were pulled out of it. - **The replacement is full-client RA-VPN:** Secure Client over SSL (`ssl-client`) or IKEv2 (`ikev2`), both of which PingLabz covers. Client-less web-app access now belongs to a reverse proxy or ZTNA. - **If you are pre-9.17, migrate before you upgrade:** inventory the clientless resources, map each to its replacement, pilot Secure Client in parallel, and decommission clientless before crossing the version that removes it. Keep going: the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) pillar indexes every ASA lab on PingLabz, and the two Secure Client remote-access guides, [SSL](https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/) and [IKEv2](https://www.pinglabz.com/cisco-asa-anyconnect-ikev2-vpn/), are the modern replacements for everything clientless used to do. ### ASDM Access on the ASA: Setup, TLS, and Whether You Should Use It URL: https://www.pinglabz.com/asa-asdm-setup/ Last updated: 2026-07-13T12:24:04.000Z Every Cisco ASA ships with a graphical management tool, the Adaptive Security Device Manager (ASDM), and every ASA lab eventually raises the same question: should you actually use it? This article sets up ASDM on a real ASAv running 9.24(1), walks the one gotcha that stops a fresh box from launching the GUI at all, and then gives you the honest 2026 answer about where ASDM fits (and where it does not). It is part of the PingLabz [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series, which drives the entire platform, including the VPN labs, from the CLI. The lab box is FW1, an ASAv 9.24(1). Before configuring anything, confirm ASDM is even present, because the ASDM application version is decoupled from the ASA software version. ``` FW1# show version | include Device Manager Device Manager Version 7.24(1) ``` So the platform reports Device Manager version 7.24(1). That number is worth understanding: ASDM has its own release train that roughly tracks, but is not identical to, the ASA software release. A 9.24 ASA pairs with a 7.24-era ASDM, and Cisco publishes a compatibility matrix because running a mismatched ASDM against an ASA image is a classic source of "it loads but half the panels are broken" complaints. Always check that the ASDM image you load is compatible with the ASA version you run. ## Enabling ASDM Access ASDM is delivered over HTTPS. The ASA runs a small web server that serves the ASDM Java client and handles the management API calls the GUI makes. Turning it on is three lines: enable the HTTPS server, then authorize which management networks may reach it on which interface. ``` http server enable http 192.168.99.0 255.255.255.0 outside http 10.20.10.0 255.255.255.0 inside ``` Read those `http` statements the way you would read an access rule, because that is essentially what they are. `http server enable` starts the HTTPS listener. Each `http ` line then says "hosts in this subnet, arriving on this interface, are allowed to reach the management server." Here, a management subnet on the outside (192.168.99.0/24) and the inside LAN (10.20.10.0/24) are both authorized. This is deliberately restrictive: if a source is not covered by an `http` statement, it cannot reach ASDM or the API at all, which is exactly the posture you want for a management plane. Never open the ASDM server to `0.0.0.0 0.0.0.0` on an internet-facing interface. One important side effect: `http server enable` does not only power the GUI. The same HTTPS server underpins the ASA's REST API. So if you are planning any programmatic management, automation, or a controller that talks to the ASA over REST, this command and the accompanying `http` authorization statements are the foundation for that too, not just for the click-through GUI. Enabling it thoughtfully matters whether or not a human ever opens ASDM. ## The Gotcha: No ASDM Image, No GUI Here is the wall a fresh ASAv walks you into. You enable the HTTPS server, you point a browser at `https:///admin`, and nothing launches. The reason is that the HTTPS listener and the ASDM application image are two separate things, and a fresh ASAv has the former without the latter. Check it directly. ``` FW1# show asdm image Device Manager image file not set ``` "Device Manager image file not set" is the whole problem in one line. `show version` reported ASDM 7.24(1) as the version the platform supports, but no actual ASDM image binary is loaded on the box, and the `asdm image` pointer is empty. The web server is happy to listen, but it has nothing to serve. The fix is to copy an ASDM image onto flash and point the ASA at it. ``` asdm image disk0:/asdm-7241.bin http server enable ``` Only after an `asdm-*.bin` is present on disk0 and the `asdm image` command references it will the GUI actually launch. The takeaway for anyone building a lab or standing up a fresh virtual ASA: `show version` telling you the Device Manager version is present is not the same as ASDM being ready. Always run `show asdm image`, and if it says the file is not set, that is your next step, not a broken browser or a TLS problem. ## The TLS and Java Reality Assume you have loaded the image and authorized your source. ASDM is a Java application launched from the browser, historically via Java Web Start, and this is where the friction lives in 2026\. The ASA presents a self-signed certificate by default, so browsers throw a trust warning until you install a proper certificate on the management interface or accept the exception. Then the Java layer adds its own hurdles: the launcher needs a compatible Java Runtime Environment, modern JREs have tightened security around unsigned or older-signed applications, and browser vendors long ago removed the plugin mechanisms that made Java-in-the-browser seamless. The practical result is that getting ASDM to launch cleanly on a current workstation often means curating a specific Java version and clicking through several security prompts. A couple of hardening moves make the access you do expose safer. Install a real certificate on the management interface rather than living with the self-signed default, so the browser trust warning goes away and the TLS session is genuinely verified. Bind management to a dedicated interface where the design allows, keeping the ASDM and REST surface off the data-facing interfaces entirely. And put authentication in front of it with `aaa authentication http console`, so reaching the web server still requires credentials from your local database or a AAA server rather than relying on the `http` source statements alone. The `http` lines control who can connect; AAA controls who can log in. You want both. None of this is a bug you can fix on the ASA. It is the accumulated cost of a Java desktop client meeting a browser and OS security model that has moved on. It is also the single biggest reason the GUI has quietly fallen out of favor for day-to-day work. ## Should You Use ASDM in 2026? Here is the honest take. ASDM is genuinely useful for a few things: a visual packet-tracer, real-time log and connection views, syslog and monitoring dashboards, and giving occasional or junior administrators a guardrailed way to make a change without memorizing CLI syntax. If someone needs to eyeball live traffic or walk a firewall rule visually, ASDM still earns its place. But for building, versioning, and operating an ASA, most engineers drive the CLI, exactly as this entire PingLabz ASA series does. The CLI is scriptable, diff-able, reviewable, and it is what every configuration guide and every troubleshooting session speaks. There is no Java version to babysit and no browser trust dance. When you need to prove a VPN is up, you type `show vpn-sessiondb`; you do not go hunting through GUI panels. The [IPsec VPN](https://www.pinglabz.com/ipsec-vpn/) labs in this series are all CLI-driven for precisely that reason: the commands are the durable skill. And there is a strategic direction worth naming. Cisco's forward platform is Firepower Threat Defense (FTD), managed by FMC (Firepower Management Center) for centralized, multi-device policy or FDM (Firepower Device Manager) for single-box, on-device management. New security features land there, not on classic ASA with ASDM. If you are choosing where to invest your GUI-management time for the long term, FMC/FDM on FTD is the answer, while ASDM remains the tool for the large installed base of classic ASA still in production. Laid out side by side, the three ways to manage the box each have a lane. CLI Best for: building, versioning, troubleshooting Strengths: scriptable, diff-able, no Java Verdict: the daily driver ASDM Best for: visual packet-tracer, live monitoring Strengths: guardrails for occasional admins Verdict: supporting actor, Java friction FMC / FDM (FTD) Best for: new deployments, central policy Strengths: where new features land Verdict: the forward path So: enable the HTTPS server thoughtfully (you likely want it for the REST API anyway), load the ASDM image if you want the GUI available, restrict the `http` statements tightly, and then reach for the CLI for real work. ASDM is a supporting actor, not the lead. ## Key Takeaways - **ASDM is present but versioned separately:** `show version` reports Device Manager 7.24(1) on a 9.24(1) ASA. Match the ASDM image to the ASA version using Cisco's compatibility matrix. - **Enabling it is three lines:** `http server enable` plus `http ` statements that authorize management sources. Keep them tight, never open to the whole internet. - **The classic gotcha:** a fresh ASAv shows "Device Manager image file not set" from `show asdm image`. The HTTPS listener works, but with no `asdm-*.bin` loaded and referenced, the GUI will not launch. - **`http server enable` also underpins the REST API:** the same HTTPS server serves both the GUI and programmatic management, so enable it deliberately even if no human opens ASDM. - **The 2026 reality:** ASDM is a Java client that modern browsers and JREs fight. It is handy for visual packet-tracer and live monitoring, but most engineers use the CLI, and Cisco's forward path is FMC/FDM on FTD. Explore more: the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) pillar indexes every ASA lab on PingLabz, all of them driven from the command line so the skills transfer straight to production. ### IPsec Through an ASA: NAT-T, ESP Pass-Through, and Why the Tunnel Won't Come Up URL: https://www.pinglabz.com/ipsec-through-asa-nat-t/ Last updated: 2026-07-13T12:24:04.000Z There are two completely different things people mean when they say "IPsec and the ASA," and confusing them is the fastest way to spend an afternoon chasing a tunnel that was never yours to fix. In one case the ASA is the VPN endpoint: it terminates the tunnel, decrypts the traffic, and the crypto config lives on the box. In the other case the ASA is a middle device, a firewall sitting between two VPN peers that terminate the tunnel somewhere else, and its only job is to let the encrypted traffic pass through cleanly. The commands, the failure modes, and the fixes are different for each. This article draws the line clearly, then digs into the pass-through case, where NAT-T and ESP handling produce the classic "Phase 1 up, Phase 2 down" symptom. It is part of the PingLabz [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series and the wider [IPsec VPN](https://www.pinglabz.com/ipsec-vpn/) collection. A note on sourcing before we start. The endpoint behavior below comes from the same live CML lab as the rest of this ASA series (an ASAv 9.24(1) terminating a real IKEv2 tunnel). The pass-through walkthrough is explained from the ASA configuration plus the NAT-T packet behavior captured in our dedicated [IPsec NAT traversal lab](https://www.pinglabz.com/ipsec-nat-traversal-nat-t/), cross-referenced here rather than re-captured end to end. Where a claim rests on that other capture, it is called out. ## Job One: The ASA as VPN Endpoint This is the case most of the ASA VPN material covers, including our [LAN-to-LAN IKEv2 walkthrough](https://www.pinglabz.com/asa-lan-to-lan-ikev2/). The ASA owns the crypto: it has an IKEv2 policy, an IPsec proposal, a tunnel group with the pre-shared key, a crypto ACL, and a crypto map bound to the outside interface. Traffic arrives in the clear on the inside, the ASA encrypts it, and it leaves as ESP toward the peer. Verification lives in `show vpn-sessiondb`, which aggregates the IKE and IPsec halves of the session on the box itself. The defining characteristic of the endpoint role: the encrypted packets originate and terminate on the ASA. There is no third device whose IPsec the ASA has to "get out of the way" for. When people say the ASA "has a VPN," this is almost always what they mean. Keep this picture clean in your head, because the pass-through role inverts almost every assumption in it. ## Job Two: The ASA in the Middle Now change the picture. Two other devices form the VPN, one on the inside of the ASA and one out on the internet, or one on each side of the firewall. The ASA terminates nothing. It just needs to forward the IPsec between them without dropping it. This is where firewalls historically caused grief, because IPsec is not a friendly protocol to inspect or translate. An IPsec VPN in the classic form is two separate things on the wire: IKE, which negotiates the tunnel over UDP 500, and ESP, which carries the encrypted data as IP protocol 50\. A stateful firewall like the ASA is comfortable with UDP 500 because it is ordinary UDP with ports it can track. ESP is the problem child. It is its own IP protocol, not TCP and not UDP, and it has no port numbers at all. To a firewall that wants to build a stateful entry keyed on a five-tuple, ESP is opaque. Before we fix it, pin down the difference in one place, because everything downstream depends on which role the ASA is playing. ASA as endpoint Crypto config: on the ASA (policy, proposal, crypto map) ESP: originates/terminates on the ASA Verify with: `show vpn-sessiondb` Firewall job: none, it owns the tunnel ASA in the middle Crypto config: none, two other devices own it ESP: passes through the ASA Verify with: conn table, inspect, packet path Firewall job: forward IKE + ESP cleanly ## inspect ipsec-pass-thru: Opening the ESP Pinhole The ASA's answer is an application inspection engine that watches the IKE negotiation and opens a temporary pinhole for the matching ESP flow. You enable it in the modular policy framework. ``` policy-map global_policy class inspection_default inspect ipsec-pass-thru ``` Here is what that one line does. When `inspect ipsec-pass-thru` is active, the ASA watches for the IKE exchange (UDP 500) between the two VPN peers. When it sees a successful negotiation, it dynamically opens the ESP pinhole so protocol 50 can flow between exactly those two hosts. It also cleans the flow up when the IKE session ends. The payoff: a VPN between two devices on either side of the ASA works without you writing a static access-list line to permit ESP, and without leaving protocol 50 wide open all the time. The inspection ties the ESP permission to a real, observed IKE negotiation. Without this inspection (or a manual ESP permit), the IKE half succeeds because UDP 500 is allowed, but the ESP half is silently dropped because the firewall has no entry for protocol 50\. That mismatch is the root of the symptom everyone eventually meets. ## The NAT-T Complication The pinhole handles a clean pass-through. Add NAT to the picture and it gets harder, which is where NAT Traversal (NAT-T) comes in. Suppose the inside VPN endpoint sits behind the ASA's PAT (Port Address Translation), sharing a public address with everything else inside. PAT works by rewriting source ports so it can demultiplex return traffic. But raw ESP has no ports. There is nothing for PAT to rewrite and nothing to key the translation table on, so PAT cannot track multiple ESP flows behind one address. The translation simply cannot be built. NAT-T is the fix, and it is elegant. When the peers detect a NAT device in the path during IKE, they wrap the ESP payload inside UDP on port 4500\. Now the "ESP" traffic is really UDP, it has ports, and PAT can translate it like any other UDP flow. The IKE negotiation itself also moves to UDP 4500 once NAT is detected. The deep mechanics, including the NAT-detection payloads and the UDP encapsulation on the wire, are captured and dissected in our [IPsec NAT traversal lab](https://www.pinglabz.com/ipsec-nat-traversal-nat-t/); the short version is that NAT-T turns an untranslatable protocol into an ordinary UDP flow. The consequence for the firewall rule is specific and easy to get wrong: when NAT-T is in play, the ASA must permit **both** UDP 500 (initial IKE, before NAT detection) **and** UDP 4500 (IKE and ESP after NAT detection). Permit only 500 and you have set the trap that produces the most famous failure signature in IPsec. ## Why the Tunnel Won't Come Up: Phase 1 Up, Phase 2 Down Here is the scenario that fills forum threads. The two peers negotiate IKE over UDP 500\. It succeeds. Both ends report Phase 1 complete, the SA is established, everything looks healthy at the control-plane level. Then no data flows. Pings across the tunnel fail. The encrypted traffic goes nowhere. The reason is that IKE and ESP travel differently. If the firewall permits UDP 500 but not the path ESP actually needs (either raw protocol 50 or, under NAT, UDP 4500), then Phase 1 completes because its packets got through, but the ESP data plane is dropped. Phase 1 up, Phase 2 down. In our NAT traversal capture, permitting only UDP 500 produces exactly this fingerprint: the IKE exchange lands, the tunnel appears established, and every data packet dies at the firewall because the UDP 4500 encapsulated ESP was never allowed. The trap is convincing because the half that succeeded is the half you look at first. Phase 1 status is the top line of most VPN verification, so a green Phase 1 reads as "the VPN works," and attention drifts to routing or the far end. But the counters tell the truth. If Phase 1 is up and the IPsec packet counters are not incrementing, or traffic leaves encrypted and nothing comes back, suspect the data-plane path, not the negotiation. ## The Endpoint Can Need NAT-T Too One more wrinkle keeps the two roles from being perfectly clean: even when the ASA is the endpoint, NAT-T can still enter the picture if there is a NAT device between the ASA and its peer. Modern ASA code enables IKEv2 NAT-T handling as part of normal operation, so when the two endpoints detect NAT in the path they encapsulate ESP in UDP 4500 automatically. This is why, in the endpoint captures elsewhere in this series, you may see the session running over UDP 4500 rather than raw ESP. The lesson generalizes: any time NAT sits between two IPsec peers, expect UDP 4500 and make sure every firewall on the path (including the endpoint's own outside ACL) permits it. The distinction that still matters is responsibility. As the endpoint, the ASA negotiates NAT-T itself and you verify it in the session database. As the middle device, the ASA is not a party to the negotiation at all; it is a bystander that must nonetheless let the negotiated UDP 4500 (or protocol 50) through. Same protocol behavior, completely different position in the conversation. ## Diagnosing It The distinguishing questions come in order. First: is the ASA the endpoint or the middle device? If the crypto config is on the ASA, it is the endpoint, and `show vpn-sessiondb` is your verification. If the ASA has no crypto map and two other devices own the tunnel, it is the middle, and your job is the forwarding path. For the middle case, walk the path: Is `inspect ipsec-pass-thru` present, or is there an explicit permit for ESP or UDP 4500? Is there a NAT device between the peers, which forces NAT-T and therefore UDP 4500? Does the firewall permit both UDP 500 and UDP 4500? A "Phase 1 up, Phase 2 down" report with a NAT device in the path is UDP 4500 being blocked until proven otherwise. For a structured approach to the whole class of these faults, the [IPsec VPN troubleshooting guide](https://www.pinglabz.com/troubleshooting-ipsec-vpn/) works through the phases in sequence so you localize the break instead of guessing. On the ASA itself, the middle-device role has its own set of read-outs that keep you honest. `show service-policy inspect ipsec-pass-thru` confirms the inspection is applied and counts what it has acted on, which is your first check that the pinhole logic is even running. The connection table (`show conn`) is where the UDP 500 and, under NAT-T, UDP 4500 flows between the two peers should appear; if you see the IKE flow but no data-plane flow, that lines up precisely with the "Phase 1 up, Phase 2 down" story. And packet-tracer, run from the inside host toward the outside peer on UDP 4500, will tell you which rule or NAT statement drops the encapsulated ESP before it ever leaves the box. The point of naming these is not to memorize output but to build the habit: prove the data-plane path exists rather than trusting a green Phase 1. ## Key Takeaways - **Two different jobs:** the ASA as VPN endpoint terminates the tunnel and owns the crypto (verify with `show vpn-sessiondb`); the ASA as a middle device only forwards someone else's IPsec. Identify which one you are dealing with before troubleshooting. - **ESP has no ports:** IKE is UDP 500 and is easy for a firewall; ESP is IP protocol 50 with no ports, so a stateful firewall cannot naturally track it. - **`inspect ipsec-pass-thru` opens the ESP pinhole:** it watches the IKE negotiation and dynamically permits the matching ESP flow, so a VPN pair either side of the ASA works without a static ESP permit. - **NAT-T wraps ESP in UDP 4500:** when the inner endpoint is behind PAT, raw ESP cannot be translated, so the peers encapsulate ESP in UDP 4500\. The ASA must then permit UDP 500 and UDP 4500. - **Permit only UDP 500 and you get "Phase 1 up, Phase 2 down":** IKE completes, the tunnel looks established, but the data plane dies because the ESP path (UDP 4500) was blocked. Read the IPsec counters, not just Phase 1 status. Next steps: the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) pillar indexes every ASA lab, the [IPsec VPN](https://www.pinglabz.com/ipsec-vpn/) hub covers the protocol from the ground up, and the [NAT traversal lab](https://www.pinglabz.com/ipsec-nat-traversal-nat-t/) shows the UDP 4500 encapsulation on the wire that this article's failure mode depends on. ### IKEv1 vs IKEv2 on the ASA: Migrating a Site-to-Site Tunnel URL: https://www.pinglabz.com/asa-ikev1-to-ikev2-migration/ Last updated: 2026-07-13T12:24:03.000Z Most site-to-site VPNs running in production today were built on IKEv1, and a lot of them are still there because "if it isn't broken, don't touch it." But IKEv1 is a 1998-era protocol, and IKEv2 replaces it with faster negotiation, built-in dead-peer detection, cleaner rekeying, and a few capabilities IKEv1 simply cannot express. On a Cisco ASA the migration is mostly a rename exercise: the same tunnel, the same interesting traffic, expressed with IKEv2 syntax. This article maps IKEv1 to IKEv2 piece by piece using real configuration from a CML lab, then shows how to cut a live tunnel over without an outage. It sits in the PingLabz [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series alongside the broader [IPsec VPN](https://www.pinglabz.com/ipsec-vpn/) collection. The topology is the same one from our [LAN-to-LAN IKEv2 walkthrough](https://www.pinglabz.com/asa-lan-to-lan-ikev2/): FW1 is an ASAv 9.24(1) at 203.0.113.10 protecting 10.20.10.0/24, and the peer is a cat8000v at 198.51.100.2 fronting 10.30.10.0/24\. That IKEv2 tunnel is the end state of this migration, so if you want to see the finished, verified result first, read that one and come back. ## The Same Tunnel, Two Protocols Start with what IKEv1 looks like on the ASA. Here is a valid IKEv1 Phase 1 policy for the same tunnel, shown for contrast. ``` crypto ikev1 policy 10 authentication pre-share encryption aes-256 hash sha group 14 lifetime 86400 ``` It is compact and readable, and it carries the seeds of why people migrate. Look at `authentication pre-share`: in IKEv1 the authentication method is a property of the policy itself, and both peers must use the same method in the same direction. There is no `prf` line, because IKEv1 derives keying material with the negotiated hash rather than a separate pseudo-random function. And `hash sha` here means SHA-1, the default that ships with legacy configs and that modern security baselines want gone. The IKEv2 form of that exact Phase 1, from the working lab, looks like this. ``` crypto ikev2 policy 10 encryption aes-256 integrity sha256 group 14 prf sha256 ``` Same encryption, same DH group, but now integrity is SHA-256, and there is a dedicated `prf sha256`. Authentication is gone from the policy entirely, because IKEv2 moves it out to the tunnel group, where it can be set per direction. That relocation is the single most important structural difference and the reason the migration is worth doing carefully rather than blindly search-and-replacing. ## The Migration Map Every IKEv1 construct on the ASA has an IKEv2 counterpart. Migrating the tunnel is a matter of translating each piece, keeping the interesting-traffic ACL and the peer address identical. Here is the mapping used in the lab. Phase 1 policy IKEv1: `crypto ikev1 policy 10` IKEv2: `crypto ikev2 policy 10` (adds `prf`) Phase 2 transform IKEv1: `crypto ipsec ikev1 transform-set` IKEv2: `crypto ipsec ikev2 ipsec-proposal` Enable on interface IKEv1: `crypto ikev1 enable outside` IKEv2: `crypto ikev2 enable outside` Tunnel-group PSK IKEv1: `ikev1 pre-shared-key ` IKEv2: `ikev2 remote/local-authentication pre-shared-key ` Crypto map binding IKEv1: `set ikev1 transform-set X` IKEv2: `set ikev2 ipsec-proposal X` Interesting traffic + peer IKEv1: ACL + `set peer` IKEv2: identical, unchanged The crypto ACL and the peer statement do not change. That is the reassuring part: the definition of what to encrypt and where to send it is protocol-agnostic. What changes is how the two peers authenticate and negotiate, which is contained in the policy, the proposal, and the tunnel-group attributes. ## What IKEv2 Actually Buys You If IKEv2 were only a rename, nobody would bother. The migration is worth the change window because IKEv2 fixes real limitations. Three of them show up directly in the configuration. **Split local and remote authentication.** This is the headline. In IKEv1 the `authentication pre-share` line applies to both peers, in one direction, and both must match. IKEv2 breaks authentication into two independent statements: `ikev2 remote-authentication` (what you require from the peer) and `ikev2 local-authentication` (what you present). Because they are separate, the two ends can use different methods. You can present a certificate while the far end presents a pre-shared key, which is impossible in IKEv1\. In the lab both are set to the same PSK for simplicity, but the machinery for asymmetric authentication is right there in the syntax. The [IKEv2 explained](https://www.pinglabz.com/ikev2-explained/) guide walks through why the exchange was redesigned this way. **A dedicated PRF.** IKEv2 adds `prf`, the pseudo-random function used to derive keying material, as its own negotiated parameter. IKEv1 reused the integrity hash for key derivation. Separating them lets you, for example, run a strong PRF while tuning integrity independently, and it removes the implicit coupling that made IKEv1 harder to reason about. **ipsec-proposal instead of transform-set.** The Phase 2 construct is renamed and restructured. IKEv1 uses a `transform-set`; IKEv2 uses an `ipsec-proposal` with explicit `protocol esp encryption` and `protocol esp integrity` sub-commands. Beyond the syntax, IKEv2 negotiates child SAs more efficiently and supports multiple proposals more cleanly. On top of these, IKEv2 brings built-in liveness checks (a standardized dead-peer detection rather than an IKEv1 add-on), fewer messages to establish a tunnel, and resistance to certain denial-of-service patterns that plagued IKEv1's aggressive mode. ## Cutting Over a Live Tunnel The risk in any migration is the outage between "old tunnel down" and "new tunnel up." On the ASA you avoid that by building the IKEv2 side alongside the IKEv1 side and moving the peer, rather than ripping out IKEv1 first. The sequence looks like this. 1. **Add the IKEv2 building blocks without touching IKEv1.** Configure `crypto ikev2 policy 10`, the `crypto ipsec ikev2 ipsec-proposal PLZ-PROP`, and `crypto ikev2 enable outside`. None of this disturbs the running IKEv1 tunnel, because nothing references it yet. 2. **Add IKEv2 authentication to the existing tunnel group.** Under `tunnel-group 198.51.100.2 ipsec-attributes`, add the `ikev2 remote-authentication` and `ikev2 local-authentication` pre-shared keys. The IKEv1 PSK can stay in place during the transition. 3. **Point the crypto map entry at the IKEv2 proposal.** Change `set ikev2 ipsec-proposal PLZ-PROP` on crypto map sequence 10\. The moment both peers prefer IKEv2 and can negotiate it, the next SA formation uses IKEv2. 4. **Coordinate the far end.** IKEv1 and IKEv2 do not interoperate with each other, so both peers must be ready for IKEv2 at cutover. Prepare EDGE2's IKEv2 keyring and profile in parallel, then clear the SA or let it rekey to force the switch. 5. **Verify, then remove IKEv1.** Once the tunnel is confirmed on IKEv2, delete the IKEv1 policy, transform-set, and PSK to close the door on a fallback you no longer want. If you are also weighing whether to keep classic crypto maps at all during a refresh, that is a separate and worthwhile decision. Our comparison of [crypto maps versus VTI](https://www.pinglabz.com/crypto-maps-vs-vti/) covers when a route-based tunnel is the better long-term target than the policy-based crypto map shown here. ## Proving the Cutover Took The whole point of a careful migration is being able to prove, not assume, that the tunnel is now IKEv2\. The ASA makes this trivial. After cutover, the session summary labels the live tunnel by IKE version. ``` FW1# show vpn-sessiondb summary Site-to-Site VPN : 1 : 1 : 1 IKEv2 IPsec : 1 : 1 : 1 Device Total VPN Capacity : 250 Device Load : 0% ``` `IKEv2 IPsec` is the line that closes the ticket. If this still read as an IKEv1 session, the cutover had not taken and the crypto map or the peer was still preferring the old protocol. Because the ASA aggregates the IKE version, the crypto parameters, and the traffic counters in the VPN session database, one command tells you both which protocol is live and whether data is flowing over it. The full `show vpn-sessiondb detail l2l` output, with the IKEv2 and IPsec halves side by side, is broken down in the [LAN-to-LAN IKEv2 article](https://www.pinglabz.com/asa-lan-to-lan-ikev2/). ## Migration Pitfalls to Plan For A few things trip people up on the ASA specifically, and knowing them ahead of the change window saves an emergency rollback. **IKEv1 and IKEv2 do not fall back to each other.** This is the one that catches teams mid-cutover. If FW1 is offering IKEv2 and the far end is still only IKEv1, the tunnel does not "downgrade," it simply fails to form. There is no graceful degradation between the two protocols. That is why coordinating both peers is a hard requirement, not a nicety. Schedule the change with whoever owns the remote device and confirm they are ready before you clear the SA. **Do not forget both authentication directions.** Because IKEv2 split the PSK into `remote-authentication` and `local-authentication`, engineers migrating from the single IKEv1 `pre-shared-key` line sometimes configure only one of the two. The tunnel then fails during authentication with no obvious reason, because the policy negotiated fine. When you translate the IKEv1 key, always write both IKEv2 lines even when the value is identical. **Leave the crypto ACL and NAT exemption alone.** The interesting-traffic ACL and any no-NAT (NAT exemption) statements protecting the VPN traffic are protocol-agnostic. Touching them during an IKE migration adds an unrelated failure domain. Change only the IKE and IPsec constructs; leave the traffic definition and NAT policy exactly as they were. **Watch your lifetimes and SHA-1.** Legacy IKEv1 configs frequently carry `hash sha` (SHA-1) and default lifetimes. A migration is the natural moment to move integrity to SHA-256, as the lab does, and to review the Phase 1 and Phase 2 lifetimes rather than blindly copying the old 86400-second values. The two rekey timers are independent under IKEv2, so set them deliberately. Plan the cutover during a maintenance window, keep the IKEv1 configuration in place until the IKEv2 tunnel is verified, and only then remove it. That way your rollback is a single crypto map change away for as long as you need it. ## Key Takeaways - **The migration is a translation, not a redesign:** policy, transform-set to ipsec-proposal, enable, tunnel-group PSK, and crypto map binding each have a direct IKEv2 equivalent. The crypto ACL and peer statement do not change. - **IKEv2's big win is split authentication:** `remote-authentication` and `local-authentication` are separate, so the two ends can use different methods (PSK one way, certificates the other), which IKEv1 cannot do. - **IKEv2 adds a PRF and restructures Phase 2:** a dedicated pseudo-random function and the `ipsec-proposal` replace IKEv1's reused hash and `transform-set`. - **Cut over without an outage:** build the IKEv2 pieces alongside IKEv1, add IKEv2 auth to the tunnel group, point the crypto map at the IKEv2 proposal, coordinate the far end, verify, then remove IKEv1\. The two protocols do not interoperate, so both peers must be ready at the same time. - **Prove it with `show vpn-sessiondb summary`:** the `IKEv2 IPsec` label is documentary proof the tunnel migrated off IKEv1. Keep going: the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) pillar links every ASA lab, the [IPsec VPN](https://www.pinglabz.com/ipsec-vpn/) hub covers the protocol end to end, and the [IKEv2 explained](https://www.pinglabz.com/ikev2-explained/) guide unpacks the exchange you just migrated to. ### LAN-to-LAN IPsec VPN on the ASA with IKEv2 URL: https://www.pinglabz.com/asa-lan-to-lan-ikev2/ Last updated: 2026-07-13T12:24:03.000Z If you have configured a site-to-site VPN on a Cisco router, the ASA version looks familiar until it doesn't. The crypto map is still there, the access list still defines interesting traffic, but the IKEv2 pieces carry ASA-specific names and the verification command you reach for on IOS returns almost nothing useful. This article walks a real LAN-to-LAN (L2L) IKEv2 tunnel from a Cisco ASA to a Cisco IOS XE peer, captured live in a CML lab, and shows you the one command that makes ASA VPN troubleshooting click: `show vpn-sessiondb`. It is part of the PingLabz [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series and the broader [IPsec VPN](https://www.pinglabz.com/ipsec-vpn/) collection, so if you want the router-to-router foundation first, start with the [ASA site-to-site VPN walkthrough](https://www.pinglabz.com/cisco-asa-site-to-site-vpn/). The lab is simple on purpose. FW1 is an ASAv running 9.24(1), sitting on the outside at 203.0.113.10 with an inside network of 10.20.10.0/24 (host 10.20.10.100). EDGE2 is a cat8000v running IOS XE 17.18.02 at 198.51.100.2, with a loopback of 10.30.10.1/24 standing in for a branch LAN. The goal: encrypt traffic between 10.20.10.0/24 and 10.30.10.0/24 using IKEv2 with AES-256, SHA-256, and Diffie-Hellman group 14. ## The ASA IKEv2 L2L Configuration An IKEv2 site-to-site tunnel on the ASA has five moving parts: the IKEv2 policy (Phase 1 proposal), the IPsec proposal (Phase 2 transform), the tunnel group holding the pre-shared key, the crypto ACL defining interesting traffic, and the crypto map that binds it all to the outside interface. Here is the exact configuration from the lab. ``` crypto ikev2 policy 10 encryption aes-256 integrity sha256 group 14 prf sha256 crypto ikev2 enable outside crypto ipsec ikev2 ipsec-proposal PLZ-PROP protocol esp encryption aes-256 protocol esp integrity sha-256 access-list VPN-ACL extended permit ip 10.20.10.0 255.255.255.0 10.30.10.0 255.255.255.0 tunnel-group 198.51.100.2 type ipsec-l2l tunnel-group 198.51.100.2 ipsec-attributes ikev2 remote-authentication pre-shared-key PingLabz-IKEv2-PSK-01 ikev2 local-authentication pre-shared-key PingLabz-IKEv2-PSK-01 crypto map PLZ-CMAP 10 match address VPN-ACL crypto map PLZ-CMAP 10 set peer 198.51.100.2 crypto map PLZ-CMAP 10 set ikev2 ipsec-proposal PLZ-PROP crypto map PLZ-CMAP interface outside ``` Read it top to bottom. The `crypto ikev2 policy 10` block is the Phase 1 (IKE\_SA) proposal: how the two peers negotiate the secure channel. Notice the `prf sha256` line, the pseudo-random function, which is unique to IKEv2 and has no equivalent in IKEv1\. The `crypto ipsec ikev2 ipsec-proposal PLZ-PROP` block is Phase 2: how the actual data is protected inside the tunnel (ESP with AES-256 and SHA-256). The tunnel group is keyed by the peer IP (198.51.100.2) and holds the pre-shared key. And here is the first ASA quirk worth flagging: IKEv2 splits authentication into `remote-authentication` and `local-authentication`, so the key you present and the key you expect are configured separately (they happen to be the same here). That split is a real IKEv2 capability, not ASA boilerplate, and it is the foundation of asymmetric auth designs. If you want the protocol-level detail, the [IKEv2 explained](https://www.pinglabz.com/ikev2-explained/) guide covers why the exchange is built this way. The EDGE2 side is a standard IOS XE IKEv2 keyring plus an IKEv2 profile and a crypto map, matching the same proposal, PSK, and interesting-traffic definition. The two ends have to agree on encryption, integrity, DH group, and the traffic selectors, or the SA never forms. ## ASA IKEv2 Pieces and Their IOS Equivalents If you already know IOS site-to-site VPN, the fastest way to learn the ASA is to map each ASA command to the IOS construct it replaces. They do the same job under different names. Phase 1 proposal ASA: `crypto ikev2 policy 10` IOS XE: `crypto ikev2 proposal` \+ `policy` Phase 2 transform ASA: `crypto ipsec ikev2 ipsec-proposal` IOS XE: `crypto ipsec transform-set` Enable IKEv2 ASA: `crypto ikev2 enable outside` IOS XE: on by default per interface Pre-shared key ASA: `tunnel-group ... ikev2 remote/local-authentication` IOS XE: `crypto ikev2 keyring` Interesting traffic ASA: `access-list VPN-ACL extended permit ip ...` IOS XE: crypto ACL or VTI Bind to interface ASA: `crypto map PLZ-CMAP interface outside` IOS XE: `crypto map` on interface The one-to-one mapping is the point: an ASA site-to-site tunnel is not a different animal, it is the same IPsec state machine with ASA naming. Where the ASA genuinely diverges is verification. ## Bringing the Tunnel Up IPsec tunnels are on-demand. Nothing negotiates until interesting traffic arrives, so the cleanest test is to source real traffic from the branch LAN toward the HQ host. From EDGE2, ping the inside host at 10.20.10.100 sourced from the loopback that sits inside the protected 10.30.10.0/24 range. ``` EDGE2#ping 10.20.10.100 source Loopback10 repeat 5 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 8/16/24 ms ``` Five exclamation marks, 100 percent success. Traffic sourced from 10.30.10.1 reached 10.20.10.100 across the encrypted tunnel and back. One detail from the capture worth internalizing: the very first attempt was only 2 of 5\. That is not a fault, it is the SA negotiating while the first packets are already in flight. The IKE\_SA exchange takes a moment to complete, a couple of early packets get dropped during setup, and then the tunnel is up and everything after it succeeds. If you script a health check, always send a second probe. ## The Star of the Show: show vpn-sessiondb Here is where IOS habits fail you. On a router, you would type `show crypto session` or `show crypto ipsec sa` and read the tunnel state. Those commands exist on the ASA too, but they are not where the ASA keeps the good information. The ASA aggregates the IKE half and the IPsec half of every VPN session, plus counters, rekey timers, and auth mode, into one database: the VPN session database. Query it with `show vpn-sessiondb`. ``` FW1# show vpn-sessiondb detail l2l IKEv2: Tunnel ID : 1.1 UDP Src Port : 500 UDP Dst Port : 500 Rem Auth Mode: preSharedKeys Loc Auth Mode: preSharedKeys Encryption : AES256 Hashing : SHA256 Rekey Int (T): 86400 Seconds Rekey Left(T): 86363 Seconds PRF : SHA256 D/H Group : 14 IPsec: Tunnel ID : 1.2 Local Addr : 10.20.10.0/255.255.255.0/0/0 Remote Addr : 10.30.10.0/255.255.255.0/0/0 Encryption : AES256 Hashing : SHA256 Encapsulation: Tunnel Rekey Int (T): 28800 Seconds Rekey Left(T): 28763 Seconds Bytes Tx : 700 Bytes Rx : 952 Pkts Tx : 7 Pkts Rx : 7 ``` Read this output like a two-story building. The **IKEv2** block (Tunnel ID 1.1) is the control channel, Phase 1\. It confirms the UDP 500 source and destination ports, that both ends used pre-shared keys (`Rem Auth Mode` and `Loc Auth Mode`, the remote/local split from the config made visible), AES256 with SHA256, DH group 14, and the pseudo-random function SHA256\. The rekey line tells you the IKE\_SA is good for 86400 seconds with 86363 left, meaning this session was established roughly 37 seconds ago. The **IPsec** block (Tunnel ID 1.2) is the data channel, Phase 2\. This is where the real proof lives. The local and remote address selectors match exactly what the crypto ACL defined: 10.20.10.0/24 to 10.30.10.0/24\. Encapsulation is Tunnel mode, encryption AES256, hashing SHA256\. And the counters are the money shot: `Bytes Tx 700 / Bytes Rx 952` and `Pkts Tx 7 / Pkts Rx 7`. Seven packets each way is our ping (five echoes plus the couple that primed the SA), encrypted and decrypted. When someone claims a tunnel is up, these counters are how you prove data is actually moving through it, not just that Phase 1 negotiated. Two things to note. First, the IKEv2 rekey (86400s, the IKE\_SA lifetime) and the IPsec rekey (28800s, the child SA lifetime) are independent timers, which is exactly how IKEv2 is meant to work: the control channel and the data channel age out on their own schedules. Second, everything you would otherwise chase across three or four `show crypto` commands on IOS is here in one screen. ## The Summary View For a fast health read, or on a box terminating hundreds of tunnels, drop the `detail` keyword and ask for the summary. ``` FW1# show vpn-sessiondb summary Site-to-Site VPN : 1 : 1 : 1 IKEv2 IPsec : 1 : 1 : 1 Device Total VPN Capacity : 250 Device Load : 0% ``` One site-to-site VPN, and the ASA labels its type explicitly: `IKEv2 IPsec`. That single line is doing quiet but important work. It is documentary proof of which IKE version the live tunnel actually negotiated. If you had migrated this tunnel from IKEv1, this is the line that confirms the cutover took (more on that in the companion migration article). The capacity and load fields also matter operationally: this ASAv supports 250 concurrent sessions and is currently at 0 percent, so you know your headroom at a glance. ## When the Tunnel Will Not Form Because the ASA hides the negotiation behind a clean session view, it helps to know where mismatches bite before they do. Every field in the `show vpn-sessiondb detail l2l` output above corresponds to something both peers had to agree on, and a disagreement on any one of them stalls the tunnel at a predictable point. Phase 1 is symmetric: the ASA's `crypto ikev2 policy` (encryption AES-256, integrity SHA-256, DH group 14, PRF SHA-256) must have a matching proposal on EDGE2\. If the two policies do not overlap, the IKE\_SA never completes and `show vpn-sessiondb` shows nothing at all, because there is no session to display. That absence is itself a clue: no session object usually means Phase 1, not Phase 2\. Point the pings at the far LAN, then check both ends of the IKEv2 policy line by line. The pre-shared key is the next classic. Because IKEv2 splits authentication, a common mistake is fixing one direction and forgetting the other. The `remote-authentication` key is what you expect from the peer, and the `local-authentication` key is what you present. If EDGE2's keyring does not line up with both, the exchange fails during authentication and, again, no session appears. In the lab both keys are `PingLabz-IKEv2-PSK-01`, which is the simple symmetric case, but the split is there precisely so you can run different credentials each way. Phase 2 is where the crypto ACL earns scrutiny. The ASA's `access-list VPN-ACL` defines interesting traffic as 10.20.10.0/24 to 10.30.10.0/24, and EDGE2's selectors must mirror it exactly (10.30.10.0/24 to 10.20.10.0/24 from its perspective). Mismatched or overly broad selectors are the most common reason a tunnel comes up but traffic does not flow, and the fingerprint is a session that exists with zero or lopsided packet counters. That is exactly why the `Pkts Tx` and `Pkts Rx` fields are worth reading every time: matching, non-zero counters in both directions mean the selectors agree and data is round-tripping. If Tx climbs but Rx stays at zero, your traffic is leaving encrypted and nothing is coming back, which points at the far end or the return-path routing rather than your own crypto. ## Why This Command Matters The teaching point is small but it separates people who fight the ASA from people who drive it. On IOS you build a mental model out of several `show crypto` commands. On the ASA, `show vpn-sessiondb` is the single pane of glass for every VPN type the box terminates: site-to-site L2L, remote-access AnyConnect (Secure Client), clientless (on older code), and L2TP. You filter it by type (`l2l`, `anyconnect`, `remote`) and add `detail` when you need the crypto parameters and counters. Learn this one command well and the ASA's VPN behavior stops being opaque. That does not mean the `show crypto` family is useless on the ASA. `show crypto ikev2 sa` and `show crypto ipsec sa` still work and are handy for raw SPI and selector detail. But for a human-readable, aggregated, per-session view with auth mode, counters, and rekey timers in one place, `show vpn-sessiondb` wins. ## Key Takeaways - **ASA IKEv2 L2L has five parts:** the `crypto ikev2 policy` (Phase 1), the `crypto ipsec ikev2 ipsec-proposal` (Phase 2), the `tunnel-group` PSK, the crypto ACL, and the crypto map bound to the outside interface. - **The PSK is split:** IKEv2 uses `remote-authentication` and `local-authentication` separately, so the two ends can even use different methods. That split shows up again in `show vpn-sessiondb` as Rem/Loc Auth Mode. - **Tunnels are on-demand:** the first probe may be partial while the SA negotiates. A 100 percent second ping is normal, not a fix. - **Use `show vpn-sessiondb`, not `show crypto session`:** the ASA aggregates the IKEv2 and IPsec halves, byte and packet counters, rekey timers, and auth mode into one view. The IPsec counters are how you prove data is really flowing. - **The summary labels the IKE version:** `IKEv2 IPsec` in `show vpn-sessiondb summary` is your proof of what the tunnel actually negotiated. Ready to go deeper? The [Cisco ASA](https://www.pinglabz.com/cisco-asa/) pillar indexes every ASA lab on PingLabz, the [IPsec VPN](https://www.pinglabz.com/ipsec-vpn/) hub covers the protocol from the ground up, and if you inherited a legacy tunnel, the next step is our guide on migrating an ASA site-to-site tunnel from IKEv1 to IKEv2. ### Transparent Mode ACLs: EtherType Rules and What Bridges Through URL: https://www.pinglabz.com/asa-transparent-etherype-acls/ Last updated: 2026-07-13T11:48:40.000Z Every ACL you have ever written on a firewall operates on IP: source address, destination address, port, protocol. A routed firewall never sees anything else, because anything that is not IP is simply not its concern. A transparent firewall is different. Because it bridges at Layer 2, it can see and filter non-IP frames that a routed firewall never touches, and the tool for that is the EtherType ACL. This is the one security capability that is genuinely unique to transparent mode, and it is part of our [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series. If you have not seen how a transparent firewall is built, start with our [transparent firewall configuration](https://www.pinglabz.com/asa-transparent-firewall-configuration/) article first, because everything here assumes that bridged-mode foundation. An EtherType ACL filters frames by their Layer 2 EtherType value, the field in the Ethernet header that says what kind of payload the frame carries. IP is one EtherType. But so are spanning-tree BPDUs, PPPoE, MPLS, and a long list of other non-IP frame types. A routed firewall drops or ignores most of these because it only forwards IP. A transparent firewall bridges them, which means it can also filter them, and that is a control no routed device can offer. ## The EtherType ACL syntax The syntax mirrors a normal ACL but matches on frame types instead of IP tuples. You build the access-list, then apply it to an interface with an access-group, exactly as you would an IP ACL. ``` access-list PLZ-ETHER ethertype permit bpdu <-- allow spanning-tree BPDUs to bridge through access-list PLZ-ETHER ethertype deny 0x8863 <-- block PPPoE discovery, for example access-list PLZ-ETHER ethertype permit ip access-group PLZ-ETHER in interface inside ``` Read each line. The first permits `bpdu`, letting spanning-tree frames bridge across the firewall. The second denies EtherType `0x8863`, the PPPoE Active Discovery type, as an example of blocking a specific non-IP protocol by its hex EtherType value. The third permits `ip`, so ordinary IP frames continue to bridge (your IP ACL still governs what happens to them at Layer 3). The `access-group` applies the list inbound on the inside interface. Note that EtherType ACLs are applied per-interface and per-direction, just like IP ACLs, so you place them precisely where you want the filtering to happen. ## Three things you need to understand EtherType ACLs are simple to type and easy to misjudge. Three points separate someone who copies the syntax from someone who understands what it does. ### 1\. By default, a transparent ASA does not bridge BPDUs This is the big one. Out of the box, a transparent firewall does not pass spanning-tree BPDUs between its bridged segments unless you explicitly permit them with an EtherType ACL. That default is a design decision handed to you: you get to choose whether spanning tree spans the firewall or not. Permit `bpdu` and the two Layer 2 domains on either side of the firewall participate in a single spanning tree, sharing loop-prevention state across the bridge. Leave BPDUs blocked, which is the default, and the two domains run independent spanning trees, isolated from each other at the STP level. Neither is universally correct. In some designs you want a single, unified topology; in others you deliberately want the firewall to segment the loop-prevention domains so a spanning-tree event on one side cannot ripple to the other. The point is that it is your call, and the EtherType ACL is the lever. Getting this wrong is a classic way to either merge two topologies you meant to keep separate or split one you meant to keep whole. ### 2\. EtherType ACLs act on frame types, and ARP is handled separately An EtherType ACL matches the frame's type field, not addresses inside a payload. That is a fundamentally different axis from an IP ACL, and it is the capability a routed firewall simply does not have, because a routed firewall does not bridge non-IP frames in the first place. One important exception to know: ARP is not filtered by EtherType ACLs. Even though ARP is a non-IP EtherType, the ASA handles it through a separate mechanism called ARP inspection, which validates ARP frames against static mappings to prevent spoofing. So you do not permit or deny ARP in an EtherType ACL; you control it with ARP inspection instead. Keeping these two straight matters, because a reasonable person assumes ARP would fall under EtherType filtering, and it does not. ### 3\. The EtherType ACL is evaluated before the IP ACL Order of operations decides outcomes. On a transparent firewall, the EtherType ACL is evaluated before the IP ACL. The frame's type is checked first: if the EtherType ACL denies it, the frame is dropped at Layer 2 and never reaches IP processing at all. Only frames that the EtherType ACL permits, including IP frames, go on to be evaluated by the IP ACL. This has a practical consequence. If you write an EtherType ACL and forget to permit `ip`, or if an implicit deny catches IP frames, your carefully crafted IP policy never even runs, because the frames were dropped one layer earlier. When you apply an EtherType ACL to an interface, make sure IP is still permitted unless you specifically intend to block it, and remember that the EtherType decision happens first in the flow. Matches on The Layer 2 EtherType field (frame type), not IP addresses or ports. A different axis entirely from an IP ACL. BPDUs Not bridged by default. Permit `bpdu` to span STP across the firewall, or leave blocked to isolate the two L2 domains. ARP Handled separately by ARP inspection, not by EtherType ACLs, even though it is a non-IP EtherType. Order Evaluated before the IP ACL. A frame denied at EtherType never reaches IP processing. ## What you would actually filter The syntax supports well-known keywords like `bpdu` and `ip`, and it also accepts raw hex EtherType values for anything without a keyword, which is how we blocked PPPoE discovery with `0x8863` above. A few of the frame types that come up in real transparent-mode designs: - **BPDU** \- spanning-tree control frames. The decision to bridge or block these defines whether STP spans the firewall, which is the single most consequential EtherType choice you will make. - **PPPoE (0x8863 discovery, 0x8864 session)** \- block these to stop rogue PPPoE from crossing a segment where it has no business being. - **MPLS unicast and multicast** \- relevant when a transparent firewall sits inside a provider or datacenter path carrying label-switched traffic. - **IP** \- permit it so your Layer 3 policy still runs; deny it only if you deliberately want a segment that carries no IP at all. The common thread is that every one of these is invisible to a routed firewall. You would never write a rule about a BPDU or a PPPoE discovery frame on a Layer 3 device, because it never bridges them. On a transparent firewall these become first-class policy objects, and the EtherType ACL is how you express intent about them. In practice the most valuable use is the STP decision, followed by blocking stray non-IP protocols that should not leak between two segments you are bridging together. ## Why this is unique to transparent mode Step back and the significance is clear. A routed firewall lives entirely in the IP world. It cannot filter BPDUs, cannot block PPPoE discovery, cannot make a decision about MPLS frames, because it does not bridge any of them; those frame types die at its interfaces or are irrelevant to it. A transparent firewall bridges Layer 2, so all of those frame types pass through it, and passing through means it can filter them. That is the one real security capability that transparent mode has and routed mode does not. Everything else the two modes share: IP ACLs, application inspection, connection limits, TCP normalization, all the MPF-based features from the rest of this series work in both modes. EtherType filtering is the genuine differentiator. If your reason for choosing transparent mode is that you need to control non-IP Layer 2 traffic between two segments, the EtherType ACL is the feature you came for, and it does not exist anywhere else in the ASA feature set. ## An honest scope note This article is concept and syntax, and we want to be clear about that. The EtherType filtering described here was not exercised end-to-end in our virtual lab. As covered in the [transparent firewall configuration](https://www.pinglabz.com/asa-transparent-firewall-configuration/) article, the data path across our transparent ASA did not come up cleanly in the virtual switch topology after the destructive mode change, so we did not have a clean bridged data plane on which to demonstrate live BPDU or PPPoE filtering. Rather than fabricate a capture, we are presenting the EtherType ACL as what it reliably is: a well-documented, per-interface, frame-type filter that is evaluated before the IP ACL and that only transparent mode makes possible. The syntax and the three teaching points above are what you carry into a real deployment, ideally on the physical two-node segment shape we recommended for demonstrating the transparent data path in the first place. ## Key Takeaways - **EtherType ACLs filter non-IP Layer 2 frames.** They match on the frame's EtherType, a capability only a bridging, transparent firewall has, and one a routed firewall cannot offer. - **BPDUs are not bridged by default.** Permit `bpdu` to let spanning tree span the firewall, or leave it blocked to keep the two L2 domains as independent STP topologies. Your choice, your lever. - **ARP is filtered separately.** Use ARP inspection for ARP, not an EtherType ACL, even though ARP is a non-IP EtherType. - **EtherType is evaluated before IP.** A frame denied at Layer 2 never reaches the IP ACL, so remember to permit `ip` unless you mean to block it. - **This is the one capability unique to transparent mode.** Every other ASA security feature works in both modes; EtherType filtering is the genuine differentiator. EtherType ACLs are the reason transparent mode exists as a distinct tool rather than just a variation on routing. For the full build that this filtering sits on top of, read the [transparent firewall configuration](https://www.pinglabz.com/asa-transparent-firewall-configuration/) article, and explore the rest of the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series for the inspection, normalization, and packet-flow features that round out the platform. ### Configuring a Transparent Firewall on the ASA URL: https://www.pinglabz.com/asa-transparent-firewall-configuration/ Last updated: 2026-07-13T11:48:39.000Z We learned the first lesson of transparent-mode firewalls the hard way, live, mid-command. Typing `firewall transparent` on our ASAv did not politely switch a mode setting. It wiped the entire running configuration instantly, and because our SSH access and routing were part of that configuration, it dropped our management session on the spot. One command, one blank config, one disconnected engineer. If you take nothing else from this article, take this: switching firewall mode is destructive, and you do it from the console with a saved config, never remotely on a box you cannot physically reach. This is part of our [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series and it goes deeper than our [routed versus transparent](https://www.pinglabz.com/cisco-asa-routed-vs-transparent/) overview, with the real config, real proof, and an honest account of what did and did not come up in the lab. ## The gotcha that bites everyone: firewall transparent wipes the config Let us stay on this, because it is the single most important operational fact about transparent mode and it is not a footnote. `firewall transparent` is a global configuration command, but it does not behave like a normal setting you toggle. The instant it runs, the ASA clears the running configuration completely. Interface addresses, nameifs, routes, ACLs, and crucially your remote-management config all disappear in that moment, because a routed-mode config and a transparent-mode config are fundamentally incompatible and the ASA refuses to carry the old one across. On our box, that meant the SSH session we were typing into simply ended. The command took effect, the config that defined our access evaporated, and the connection died mid-stream. We rebuilt FW1 from a fresh transparent-mode day-zero configuration afterward, from the console, which is the only sane way to do it. The takeaway is procedural and non-negotiable: back up your config first, get to the console (physical or out-of-band), switch the mode there, and rebuild from a prepared transparent-mode config. Doing this over the network on a production firewall is how you turn a maintenance window into a site visit. ## What transparent mode actually is A transparent firewall operates at Layer 2\. Instead of routing between subnets, it bridges between them, sitting in the path like a bump in the wire. Both data interfaces live in the same IP subnet and the same bridge group, and the firewall forwards frames between them while applying its security policy. To the devices on either side, there is no router hop; traffic appears to pass straight through, and a traceroute across the firewall shows nothing where the ASA sits. That invisibility is the entire appeal. You can drop a transparent firewall into an existing segment without renumbering anything, without changing default gateways, and without the endpoints ever knowing a firewall was inserted. It is the classic way to add inspection and filtering to a network you are not allowed to re-architect. Where a routed firewall is a visible Layer 3 hop with its own interface subnets, a transparent firewall is a silent Layer 2 device that filters what passes through it. Our [routed versus transparent](https://www.pinglabz.com/cisco-asa-routed-vs-transparent/) article compares the two modes side by side; this one shows the transparent build in detail. ## The transparent config: bridge-group and BVI Here is the shape of a minimal transparent configuration. The key structural difference from routed mode is that the data interfaces have no IP addresses at all. They are members of a bridge group, and the only Layer 3 address in the whole design belongs to a single Bridge Virtual Interface, the BVI, used purely for management. ``` firewall transparent interface GigabitEthernet0/0 nameif outside security-level 0 bridge-group 1 <-- interfaces are bridged, NOT routed interface GigabitEthernet0/1 nameif inside security-level 100 bridge-group 1 interface BVI1 ip address 10.20.10.5 255.255.255.0 <-- ONE management IP for the whole bridge group access-list OUT-IN extended permit ip any any access-group OUT-IN in interface outside ``` Read the structure. Both physical interfaces get a nameif and a security level, but no IP; they join `bridge-group 1`. The `BVI1` interface carries the one management IP for the entire bridge group, `10.20.10.5`. Because both data interfaces are in the same bridge group and the same subnet, the firewall bridges frames between them rather than routing packets. The extended ACL still governs IP traffic exactly as it would in routed mode, so your familiar permit and deny logic is unchanged; what changes is the plumbing underneath it. This is the mental model to copy: two bridged data interfaces, one BVI management IP, no per-interface Layer 3, and your IP ACLs applied on top. The firewall becomes a filtering bridge. ## Real proof: the BVI is reachable at Layer 2 We can prove the transparent ASA is genuinely bridging and presenting its management IP at Layer 2, using a real IOS router, ISP1, sitting on the outside segment at `10.20.10.254`. If the ASA is truly transparent, ISP1 should be able to ARP for the BVI address and reach it as an ordinary Layer 2 neighbor, because the BVI is not a routed interface on a far subnet, it is a bridged endpoint in the same broadcast domain. ``` ISP1# show arp | include 10.20.10.5 Internet 10.20.10.5 1 5254.004e.feaa ARPA Ethernet0/1 <-- BVI MAC learned via the transparent FW ISP1# ping 10.20.10.5 repeat 3 !!! Success rate is 100 percent (3/3), round-trip min/avg/max = 2/2/4 ms ``` Two things are proven here. First, ISP1 learned the BVI's MAC address, `5254.004e.feaa`, by ARP, which only works if the ASA answered ARP on that segment as an L2-reachable endpoint. Second, the ping succeeds three for three, so the BVI responds to ICMP as a normal neighbor on the same subnet. The BVI is behaving exactly as a bridged management interface should: it answers ARP and ICMP as a Layer 2 endpoint, not as a routed interface on the other side of a hop. The management plane of the transparent firewall is up and reachable. ## An honest note on the data path Now the part that most tutorials would quietly skip. In our virtual lab, full inside-to-outside data-plane forwarding across the transparent ASA did not come up cleanly after the mid-lab mode change. With an IOL-based virtual switch on the inside segment plus the ASAv bridge group, the inside host could not complete ARP across to the outside segment during our capture window. Notably, even host-to-host ARP on the inside virtual switch went incomplete, which points at the virtual switch's Layer 2 state after the reconfigure rather than at the ASA's policy. We are telling you this plainly rather than fabricating a clean end-to-end ping, because a made-up capture helps nobody and a transparent firewall's whole value is in the data path. What we did prove is real and load-bearing: the config-wipe behavior, the bridge-group and BVI structure, and the BVI's genuine Layer 2 reachability from a real router. What we did not get in this virtual topology was a clean bump-in-the-wire data forwarding demonstration, and the honest reason is the virtual switch's L2 state after the destructive mode change, not the ASA configuration itself. The practical guidance that follows is worth more than a faked screenshot. A clean transparent-mode data path is best demonstrated on a physical segment, or on a simpler two-node bridge where you are not fighting a virtual switch's forwarding state on top of the firewall's own bridging. That two-node, physical-segment shape is the deployment readers should copy. The transparent ASA belongs inline on a real cable between two real Layer 2 domains; that is where it shines and where its behavior is predictable. ## Routed and transparent, structurally The two modes differ in where the intelligence lives, and seeing the structural contrast side by side makes the transparent config above click into place. Routed mode Each interface has its own IP in its own subnet. The firewall is a visible Layer 3 hop. It routes packets, can run dynamic routing, and endpoints use it as a gateway. Transparent mode Data interfaces have no IP and share one subnet via a bridge group. One BVI IP for management. The firewall is an invisible Layer 2 bridge that filters what passes through. ## What transparent mode gives up Transparency is not free, and knowing the trade-offs keeps you from choosing it for the wrong job. Because the firewall does not route through-traffic, it does not participate in your routing protocols for the data path the way a routed-mode firewall can; it simply bridges the frames that carry those protocols. Features that depend on the firewall being a Layer 3 hop behave differently or are unavailable, and NAT in transparent mode is more constrained than the full routed-mode NAT toolset. The BVI is for management reachability, not for routing user traffic between subnets, so if your requirement is genuinely to route between two different subnets, transparent mode is the wrong tool and you want routed mode. What you gain in return is the ability to insert security into a segment you are forbidden from re-addressing, and to filter traffic types a routed firewall never even sees, which is the subject of the EtherType-ACL article next in this series. The decision comes down to a single question: do you need the firewall to be a router, or do you need it to be an invisible filter on an existing wire? If the answer is invisible filter, transparent mode is exactly right, provided you respect the destructive mode switch and, ideally, deploy it on a clean physical segment where the data path is predictable. ## Design notes for a real deployment A few things to plan for when you build this outside a lab. Give the BVI a management IP in the same subnet as the segment you are inserting into, and make sure whatever needs to manage the firewall can reach that subnet at Layer 2\. Keep the security levels meaningful: the outside interface at 0 and the inside at 100 preserves the usual higher-to-lower default logic, which still governs IP traffic even though the firewall is bridging. And remember that because the firewall is invisible as a hop, your endpoints keep their existing gateways and addressing; you are inserting a filter, not a router, and nothing on either side needs to be renumbered. Give some thought to spanning tree before you insert the device. A transparent firewall bridges two Layer 2 domains, so whether STP treats them as one topology or two depends on whether you allow BPDUs to bridge across the firewall, which is an EtherType-ACL decision rather than an IP-policy one. If you let BPDUs through, the two segments participate in a single spanning tree spanning the firewall; if you block them, the two domains run independent spanning trees and the firewall isolates them at the STP level. That choice can matter a great deal in a datacenter where an inserted firewall must not accidentally merge or split loop-prevention domains, and it is precisely the kind of non-IP control that only transparent mode exposes. Plan the maintenance window around the config wipe. Because switching to transparent mode clears everything, the cleanest procedure is to prepare the full transparent-mode configuration in advance as a text file, get console access, issue `firewall transparent`, and paste the prepared config. That turns a potentially session-killing surprise into a controlled, scripted change. It is the same discipline that separates a smooth firewall insertion from a two-hour outage, and it flows directly from the lesson we learned the hard way at the top of this article. ## Key Takeaways - **`firewall transparent` wipes the running config instantly.** It dropped our live SSH session mid-command. Switch mode from the console, with a saved config, never remotely. - **Transparent mode bridges at Layer 2.** Data interfaces have no IP and join a bridge group; a single BVI holds the one management IP for the whole group. The firewall is a bump in the wire, invisible as a hop. - **The BVI is a real Layer 2 endpoint.** A real router, ISP1, learned the BVI MAC `5254.004e.feaa` by ARP and pinged it 100 percent, proving the ASA bridges and presents its management IP at L2. - **The virtual data path did not come up cleanly.** After the mid-lab mode change, inside-to-outside ARP failed in our virtual switch topology, which points at the virtual switch's L2 state, not the ASA policy. We did not fabricate a working ping. - **Copy the physical, two-node shape.** A clean bump-in-the-wire data path is best shown on a physical segment or a simple two-node bridge, which is the deployment pattern to replicate. Transparent mode is a powerful way to insert a firewall without touching Layer 3, as long as you respect the destructive mode switch and plan the insertion. Continue in the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series with the one capability transparent mode has that routed mode never will: filtering non-IP frames with EtherType ACLs, and revisit the [routed versus transparent](https://www.pinglabz.com/cisco-asa-routed-vs-transparent/) comparison for when to choose each mode. ### TCP Normalization on the ASA: The Silent Connection Killer URL: https://www.pinglabz.com/asa-tcp-normalization/ Last updated: 2026-07-13T11:48:38.000Z Some of the hardest firewall tickets sound like ghost stories. A connection that works fine between two hosts dies the moment it crosses the ASA, both endpoints swear nothing is wrong, and there is no deny in any ACL. Nine times out of ten the culprit is TCP normalization, the ASA's silent enforcer of TCP correctness, and the only evidence it leaves is a counter most people never look at. This article is part of our [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series, and it turns that ghost into something you can see, using a real `tcp-map` and the real accelerated-security-path drop counters from an ASAv 9.24 in CML. TCP normalization is a set of checks the ASA runs on every TCP segment to make sure the connection behaves like a proper, in-order, standards-compliant TCP flow. Packets that fail those checks are dropped. That is a good thing when the packets really are malformed or part of an attack, and a maddening thing when an overly strict normalizer drops traffic the endpoints considered perfectly valid. Either way, the box is doing exactly what it was told, and learning to read its drop counters is how you tell the difference. ## What the normalizer checks The ASA normalizes TCP by default, and you can tune it with a `tcp-map`. The map is a bundle of individual checks, each of which can drop, clear, or allow a particular kind of anomaly. A few of the common ones: - **check-retransmission** \- verify that a retransmitted segment matches the original, catching retransmission-based evasion. - **exceed-mss** \- decide what to do with segments larger than the negotiated maximum segment size, either allow or drop them. - **urgent-flag** \- allow or clear the TCP URG pointer, which is frequently abused and rarely legitimately needed. - **tcp-options** \- control which TCP options are permitted, cleared, or cause a drop. You bundle these into a named `tcp-map` and then apply it through the Modular Policy Framework, exactly like every other ASA feature from our [MPF walkthrough](https://www.pinglabz.com/asa-modular-policy-framework/). The map is the action; MPF is how it reaches the traffic. ## Building and applying a tcp-map Here is a small normalizer that checks retransmissions and drops oversized segments, applied globally to all traffic. ``` tcp-map PLZ-TCP-NORM check-retransmission exceed-mss drop class-map PLZ-TCP-CLASS match any policy-map global_policy class PLZ-TCP-CLASS set connection advanced-options PLZ-TCP-NORM ``` The three layers are the ones you already know. The `tcp-map` defines the checks. The `class-map` with `match any` selects all traffic. The `policy-map` attaches the map to that class with `set connection advanced-options`, which is the specific action that plugs a tcp-map into MPF. Because we added it to `global_policy`, it now applies everywhere the global service-policy does. Confirm it is live by reading the global service-policy and filtering to the connection settings. ``` FW1# show service-policy global | include Set connection|advanced Set connection policy: drop 0 embryonic drop 0 Set connection advanced-options: PLZ-TCP-NORM ``` The line `Set connection advanced-options: PLZ-TCP-NORM` is the confirmation that our tcp-map is bound and active on the global policy. If that line is missing, the map exists in the config but is not doing anything, the same "defined but not applied" trap that catches every MPF feature. ## The star command: show asp drop Here is the command that solves the ghost stories. The accelerated security path (ASP) is the ASA's fast-path forwarding engine, and it keeps a running tally of every packet it drops and why. TCP normalization drops land here, alongside a long list of other fast-path discard reasons. On our lab box, after ambient traffic, the counters looked like this: ``` FW1# show asp drop Invalid UDP Length (invalid-udp-length) 2 First TCP packet not SYN (tcp-not-syn) 10 TCP RST/FIN out of order (tcp-rstfin-ooo) 2 ``` Every one of those lines is a packet the ASA dropped without any ACL involved, and each names its reason in the parentheses. This is the single most useful "why did my packet vanish" command on the box, because it is the one place the firewall admits to dropping traffic that no ACL touched. When a connection fails and the ACLs are clean, this is where you look next. ## Why each counter fired Read the reasons, because they teach you how the stateful engine thinks. **tcp-not-syn (10)** is the headline. The ASA is stateful, which means it expects a TCP connection to begin with a SYN so it can build a state entry for the flow. If the very first packet the firewall sees for a flow is not a SYN, the ASA has no state for that connection and, by default, drops it as `tcp-not-syn`. This happens constantly in the real world: a connection that was established before the firewall booted, an asymmetric path where the return traffic arrives on the ASA but the initial SYN went another way, or a flow that failed over from another device mid-stream. All of them look identical to the ASA, a first packet with no SYN and no state, and all of them get dropped. Ten such packets is completely ordinary ambient noise on any segment. **tcp-rstfin-ooo (2)** catches RST or FIN segments that arrive out of order, outside the valid sequence window for the connection. Out-of-order teardown packets can be a sign of an evasion attempt or simply of a messy network, and the normalizer discards them rather than letting them tear down a connection they should not. **invalid-udp-length (2)** is not TCP at all; it is the ASP dropping UDP datagrams whose declared length does not match reality, a basic malformed-packet check. It is in the same output because `show asp drop` is the catch-all for fast-path discards, which is exactly why it is so useful: one command, every silent drop reason, in one place. `tcp-not-syn` Count: 10 First packet of a flow was not a SYN, so the stateful ASA had no state entry. Points at asymmetry or a pre-existing connection. `tcp-rstfin-ooo` Count: 2 RST or FIN arrived outside the valid sequence window. Guards against out-of-order teardown and evasion. `invalid-udp-length` Count: 2 UDP datagram whose declared length did not match the packet. A basic malformed-packet check, not TCP related. ## Default normalizer versus a custom tcp-map It is worth being clear that you get normalization whether or not you ever write a tcp-map. The ASA applies a default normalizer to every TCP connection, and that default is already responsible for the `tcp-not-syn` drops above; no configuration of ours caused them. What a custom `tcp-map` does is let you tighten or loosen specific checks beyond the defaults, for example adding `check-retransmission` or changing how oversized segments are handled. This distinction matters when you are troubleshooting, because engineers often assume that with no tcp-map configured the ASA is not normalizing at all. It is. If you genuinely need the firewall to accept a mid-stream flow that begins with something other than a SYN, the answer is not to remove a tcp-map you never created; it is to configure TCP state bypass for that specific traffic, which tells the ASA to forward the flow without building or enforcing full TCP state. That is a deliberate, scoped exception, and it is the correct tool for asymmetric designs rather than weakening the normalizer for everything. ## The silent connection killer Now the part that makes this feature worth a whole article. Because normalization runs in the fast path and drops packets without generating a per-packet deny that shows up where people usually look, an over-aggressive tcp-map can quietly break connections that look completely healthy from both endpoints. The client sends, the server never sees it, and neither machine logs anything unusual because from their point of view the packet was simply lost. There is no deny message on a screen anyone is watching. The only evidence that the firewall did it is a number climbing in `show asp drop`. This is why TCP normalization earns the nickname "silent connection killer". Turn on a strict option like dropping all segments that exceed MSS, or clearing an option some legitimate application actually depends on, and you can introduce failures that survive hours of endpoint troubleshooting because nobody is looking at the firewall's ASP counters. Even the default normalizer, before you configure a single custom option, drops `tcp-not-syn` packets, which is exactly why asymmetric-routing designs and firewall-insertion projects so often surface mysterious connection failures on day one. The lesson is procedural. When a connection dies crossing the ASA and the ACLs are clean, run `show asp drop` before you do anything else. Note the counters, reproduce the failure, run it again, and see which counter moved. The counter that increments is your answer. If it is `tcp-not-syn`, you are looking at a state or asymmetry problem, not a policy problem, and the fix is on the routing or the TCP-state-bypass side, not in an ACL. If it is a normalization reason tied to a custom tcp-map option, you have found an over-strict setting, and you can relax that specific check rather than tearing the whole map out. ## Tuning without breaking things A tcp-map is a scalpel, so use it like one. Add checks one at a time and watch the corresponding `show asp drop` counter before and after, so you know exactly what each option is dropping in your environment. If a counter climbs faster than you expected, you have caught a legitimate flow in the net, and you can decide whether to loosen the option or fix the traffic that triggered it. The goal is a normalizer strict enough to catch genuinely malformed or evasive TCP, but not so strict that it drops the slightly-imperfect-but-harmless segments real applications routinely send. Where this all sits in the ASA's processing order matters too, because normalization happens in the fast path after the connection is established, not at the ACL stage. That is precisely why an ACL check can pass while a packet still vanishes: the two live at different points in the flow. Our [ASA packet flow](https://www.pinglabz.com/cisco-asa-packet-flow/) article maps that ordering out, and it is the model that finally makes `show asp drop` feel predictable rather than mysterious. ## Key Takeaways - **TCP normalization silently drops malformed or out-of-state TCP.** It is on by default and tunable with a `tcp-map` applied via `set connection advanced-options`. - **`show asp drop` is the star command.** It is the one place the ASA records fast-path drops that no ACL caused, the definitive "why did my packet vanish" tool. - **Read the drop reasons.** `tcp-not-syn 10` means a first packet arrived with no SYN and no state (asymmetry or a pre-existing flow); `tcp-rstfin-ooo 2` caught out-of-order teardown; `invalid-udp-length 2` is a malformed UDP check. - **The silent killer is real.** An over-strict tcp-map drops packets that look fine to both endpoints, leaving no evidence except a climbing ASP counter. - **Diagnose by watching the counter move.** Reproduce the failure and see which `show asp drop` line increments; that reason is your root cause. Once you know `show asp drop` exists, a whole category of "impossible" firewall problems becomes routine. Continue through the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series, and pair this with the [ASA packet flow](https://www.pinglabz.com/cisco-asa-packet-flow/) article to see exactly where in the processing order these silent drops happen. ### Custom L7 Inspection on the ASA: Regex, Match Conditions, and Blocking What You Choose URL: https://www.pinglabz.com/asa-custom-l7-inspection/ Last updated: 2026-07-13T11:48:38.000Z This is where the ASA stops being a Layer 4 gate and starts reading your web traffic. In this article we build a custom Layer 7 HTTP inspection policy that blocks a specific URL by regular expression, and then we prove it works with a real before-and-after test from a real Linux client: one URL returns instantly, the other is dropped mid-connection and hangs until it times out. No hand-waving, no simulated output. A real regex, a real request, a real block, captured on an ASAv 9.24 in CML. This is part of our [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series and it goes well beyond the [inspection engines overview](https://www.pinglabz.com/cisco-asa-inspection-engines/), from listing what the engines do to making one drop traffic you choose. Custom Layer 7 inspection is the payoff of everything in the ASA inspection story. The default policy inspects HTTP but does nothing with the contents. Once you attach your own inspection policy-map, the firewall can match on the URI, headers, methods, and body of every web request crossing it, and drop the ones you name. That is content control enforced at the firewall, and it is built on the same Modular Policy Framework you have seen throughout this series. ## The goal: block one URI, allow the rest Our target is simple and realistic. A published nginx server sits in the DMZ. We want normal pages to load, but any request whose URI contains the string `secret` should be dropped at the firewall before it ever reaches the server. The client should not get a polite error page; the connection should simply die, because we are dropping it at Layer 7, not returning an HTTP status. Building this takes four objects, layered from the regex up to the service-policy. Each one answers a narrower question than the last. ``` regex BLOCK-URI ".*secret.*" class-map type inspect http match-all PLZ-BADURI match request uri regex BLOCK-URI policy-map type inspect http PLZ-HTTP-INSPECT parameters class PLZ-BADURI drop-connection log policy-map PLZ-DMZ-POLICY class PLZ-HTTP-CLASS no inspect http inspect http PLZ-HTTP-INSPECT service-policy PLZ-DMZ-POLICY interface dmz ``` Read it from the top. The `regex` defines the pattern to hunt for, `.*secret.*`, which matches any URI containing the string secret anywhere in it. The `class-map type inspect http` is a Layer 7 classifier: it matches HTTP requests whose URI matches that regex. The `policy-map type inspect http` is the Layer 7 policy: for the PLZ-BADURI class, it drops the connection and logs it. Finally, the ordinary Layer 4 policy-map swaps its plain `inspect http` for `inspect http PLZ-HTTP-INSPECT`, wiring our custom engine into the existing DMZ service-policy. The nesting is the important idea. A Layer 4 policy-map cannot match a URI; only a `type inspect http` policy-map can. So you build the Layer 7 policy separately and then reference it as the action inside the Layer 4 policy. This is the ASA's way of layering deep inspection on top of the basic MPF stack from our [Modular Policy Framework](https://www.pinglabz.com/asa-modular-policy-framework/) walkthrough. ## The money capture: before and after from a real client With the policy applied, we ran two requests from a Debian client against the published server. The first asks for a normal page. The second asks for a URI containing secret. We used curl with a format string that prints the HTTP status code and the total time, so the difference is impossible to miss. ``` j@llmbits:~$ curl -w '%{http_code} in %{time_total}s' http://203.0.113.80/index.html normal /index.html -> HTTP 200 in 0.010053s <-- allowed, fast j@llmbits:~$ curl -w '%{http_code} in %{time_total}s' http://203.0.113.80/secret/data.html blocked /secret -> HTTP 000 in 8.003106s <-- connection DROPPED (000, then timeout) ``` Look at the two lines side by side. The normal page returns `HTTP 200` in about ten milliseconds, a clean, fast fetch. The secret URL returns `HTTP 000`, which is curl's way of saying it never got any HTTP response at all, and it takes eight full seconds to say so. That eight-second gap is the client waiting on a connection the ASA silently killed. There is no error page, no redirect, no reset that curl could interpret as a status; the connection is simply gone, and curl gives up only when its own timeout expires. That is exactly what `drop-connection` does, and the contrast between 0.01 seconds and 8 seconds is the whole proof in a single screen. ## The ASA side: the drop counter and the syslog The client saw a hang. The firewall saw everything. First, the inspection policy's own counters, filtered to the lines that matter. ``` FW1# show service-policy inspect http | include Inspect:|drop Inspect: http PLZ-HTTP-INSPECT, packet 16, drop 1, reset-drop 0 drop-connection log, packet 1 ``` The engine name confirms our custom policy is the active one: `Inspect: http PLZ-HTTP-INSPECT`, not the plain default. `drop 1` is the single dropped connection, the secret request. `drop-connection log, packet 1` confirms the specific action fired exactly once. One request matched, one connection dropped. That is a clean, unambiguous counter, and it is the firewall-side twin of the eight-second hang the client experienced. Now the logs, which are where L7 inspection becomes genuinely impressive, because the ASA is reading inside the HTTP and telling you what it saw. ``` FW1# show log %ASA-5-304001: 192.168.99.100 Accessed URL 172.20.30.80:http://203.0.113.80/index.html %ASA-5-304001: 192.168.99.100 Accessed URL 172.20.30.80:http://203.0.113.80/secret/data.html %ASA-5-415006: HTTP - matched Class 23: PLZ-BADURI in policy-map PLZ-HTTP-INSPECT, URI matched - Dropping connection from outside:192.168.99.100/52990 to dmz:172.20.30.80/80 %ASA-4-507003: tcp flow from outside:192.168.99.100/52990 to dmz:172.20.30.80/80 terminated by inspection engine, reason - disconnected, dropped packet. %ASA-6-302014: Teardown TCP connection 31 ... Flow closed by inspection ``` Read those five lines as a story. The two `%ASA-5-304001` lines log both URLs the client accessed, the normal one and the secret one. That alone proves the ASA is parsing the HTTP request line, because it knows the full URL, not just the destination IP and port. Then `%ASA-5-415006` fires on the secret request only, naming the matched class, PLZ-BADURI, and the policy-map, and states plainly that the URI matched and it is dropping the connection. The `%ASA-4-507003` line records the TCP flow being terminated by the inspection engine, and `%ASA-6-302014` tears the connection down. The firewall saw both URLs, decided on one, and killed it, all logged in order. That `%ASA-5-304001` pair is the detail worth sitting with. The ASA logged the exact URL of a request it allowed and the exact URL of a request it blocked. It is reading the application layer of every web request that crosses it. For a security team that is both a powerful audit capability and a privacy consideration worth being deliberate about. ## What you can match in an http inspect class-map We matched a URI, but that is one option among many. A Layer 7 HTTP inspection class-map can match on nearly every meaningful part of an HTTP transaction, which is what makes it a real content-control tool rather than a URL blocklist. Request URI `match request uri regex` \- the path and query, what we used to block `secret`. Headers `match request header` \- match on any header field, such as Host or Content-Type, by regex. Request/response body `match request body` / `match response body` \- inspect payload content within a length limit. Method `match request method` \- allow GET/POST and block risky verbs like PUT, DELETE, or TRACE. User-Agent `match request header user-agent regex` \- block or flag specific clients and tools. Content length / type Enforce maximum body sizes and permitted content types to blunt oversized or disguised payloads. Each of these becomes a `match` line inside a `class-map type inspect http`, and each class gets an action in the `policy-map type inspect http`: `drop-connection`, `reset`, or `log`. Combine several with `match-all` or `match-any` and you can express policies like "drop any POST to a URI containing admin from a curl user-agent" in a handful of lines. ## Choosing the action: drop, reset, or log We used `drop-connection log`, which silently discards the connection and writes a syslog entry. It is worth knowing the alternatives, because the action changes what the client experiences. `drop-connection` silently kills the flow, which is why our curl hung for eight seconds rather than getting an immediate error. `reset` sends a TCP reset, so the client fails fast instead of waiting for a timeout, which is friendlier when you want the user to know quickly but tells an attacker that something is actively filtering. `log` on its own takes no action and only records the match, which is the right first step when you are validating a regex before you let it drop production traffic. A sane rollout is to deploy a new match with `log` only, watch the `%ASA-5-415006` entries to confirm it fires on exactly the traffic you intend and nothing else, and only then switch the action to `drop-connection` or `reset`. A regex like `.*secret.*` is broad on purpose here, but in production an overly greedy pattern can catch far more than you meant, and the log-first approach is how you find that out before it becomes an outage. ## Understanding the regex The pattern `.*secret.*` deserves a closer look, because the regex is the part most likely to bite you. In ASA regex syntax, `.` matches any single character and `*` means zero or more of the preceding element, so `.*` matches any run of characters including an empty one. Wrapping `secret` in `.*` on both sides therefore means "the string secret appears anywhere in the URI", which is what caught `/secret/data.html`. It would equally catch `/my-secret-plans` or `/documents/topsecret.pdf`, because the match is a substring, not a path component. That breadth is a feature when you want it and a hazard when you do not. If your intent was strictly the `/secret/` directory, a tighter pattern anchored to the path, something closer to `/secret/.*`, expresses that intent and avoids collateral matches on unrelated URLs that merely contain the letters. The ASA also supports character classes, alternation with the pipe, and quantifiers, so you can build precise patterns. The discipline is to write the narrowest regex that covers your case, then confirm with log-only mode that it fires on exactly what you meant. A regex that is too greedy is the single most common way an L7 policy causes an outage, because it silently blocks pages nobody thought to test. ## Why not just use an ACL or a proxy? A fair question: if you want to block a URL, why not use a web proxy or a Layer 4 ACL? The answer clarifies what this feature is for. An ACL cannot do this at all, because a URI lives in the HTTP payload and an ACL only sees IP addresses and ports. Every request to the DMZ server hits the same IP and port 80, so an ACL can only allow or block the whole server, never a single path. A dedicated web proxy or next-generation firewall can absolutely do URL filtering, often with richer categorization and TLS interception, and if that is your primary requirement those tools are purpose-built for it. ASA L7 inspection sits in between. It is not a full web-filtering platform, and it does not decrypt TLS, so it only sees inside plain HTTP. What it gives you is the ability to enforce targeted content rules directly on the firewall you already have, in the same MPF framework you already know, with no extra appliance in the path. For blocking a known-bad URI, restricting HTTP methods on a published server, or catching a specific user-agent, it is exactly the right tool, and the fact that it drops and logs with real precision is what makes it trustworthy for those jobs. ## Where this sits in the packet flow It helps to know when this drop happens relative to everything else the ASA does. The ACL permits the connection first, the connection is built, and only then does the HTTP inspection engine parse the request and decide to drop it. That is why the two `%ASA-5-304001` URL-access logs appear before the drop: the connection was allowed and established, the request was read, and the drop came at Layer 7 after all the Layer 3 and 4 checks had already passed. Our [ASA packet flow](https://www.pinglabz.com/cisco-asa-packet-flow/) article lays out that ordering in full, and it is the mental model that explains why an L7 drop looks like a hang rather than a clean deny. ## Key Takeaways - **Custom L7 HTTP inspection reads inside the request.** A `regex`, a `class-map type inspect http`, and a `policy-map type inspect http` let the ASA match and drop on the URI, headers, method, or body. - **The before-and-after is unambiguous.** A normal page returned `HTTP 200 in 0.01s`; the blocked URL returned `HTTP 000 in 8s` as the dropped connection hung to timeout. - **The firewall logs both URLs.** Two `%ASA-5-304001` entries recorded the allowed and blocked URLs, proving the ASA parses the HTTP request line, and `%ASA-5-415006` named the matched class on the drop. - **The counter confirms it.** `show service-policy inspect http` showed `drop 1` under the custom `PLZ-HTTP-INSPECT` engine, the firewall-side twin of the client's hang. - **Roll out with `log` first.** Validate a regex in log-only mode before switching to `drop-connection` or `reset`, because a greedy pattern catches more than you expect. You have now made the ASA drop web traffic by content, the deepest thing a firewall inspection engine does. Explore the rest of the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series for the surrounding pieces, revisit the [inspection engines overview](https://www.pinglabz.com/cisco-asa-inspection-engines/) for the wider catalog, and read the [packet flow](https://www.pinglabz.com/cisco-asa-packet-flow/) article to see exactly where this drop lands in the ASA's processing order. ### Deep Packet Inspection on the ASA: Application Inspection Engines in Practice URL: https://www.pinglabz.com/asa-deep-packet-inspection/ Last updated: 2026-07-13T11:48:37.000Z A stateful firewall that only reads Layer 3 and Layer 4 headers is half a firewall. It can decide whether TCP port 21 is allowed, but it has no idea what the FTP control channel is negotiating on that port, and it cannot open the return data channel the protocol is about to ask for. Application inspection is what closes that gap on the ASA, and it is the feature that separates a real security appliance from a router with an ACL. This article is part of our [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series and goes a layer deeper than our [inspection engines overview](https://www.pinglabz.com/cisco-asa-inspection-engines/), showing the default inspection policy in action on a real ASAv 9.24 in CML with live per-interface counters. Deep packet inspection on the ASA means the firewall parses the application-layer payload of a connection, not just its headers. It does this for two reasons that a plain ACL can never handle: it opens pinholes for protocols that negotiate their own ports, and it enforces protocol conformance so that traffic claiming to be HTTP actually behaves like HTTP. Both matter, and both ship enabled by default. ## The default inspection policy Every ASA boots with application inspection already running. This is not something you turn on; it is something you inherit. The factory configuration wires up a complete Modular Policy Framework stack whose only job is to inspect the common protocols on every interface. If you have read our [MPF walkthrough](https://www.pinglabz.com/asa-modular-policy-framework/), this will look familiar, because it is the same three layers applied to the box itself. - **global\_policy** is the default policy-map. It holds the inspection actions. - **inspection\_default** is the default class-map inside it. It matches traffic with `match default-inspection-traffic`, the special classifier that selects the standard inspected protocols on their well-known ports. - **service-policy global\_policy global** is the binding that applies the whole thing to every interface. Put those together and you get a firewall that, out of the box, is inspecting DNS, FTP, ICMP, and a long list of other protocols everywhere, with no explicit configuration from you. That is a good default, and it is also why some traffic quietly works on an ASA that would fail on a stateless device: the inspection engines are silently opening the pinholes those protocols need. ## What application inspection actually does Two jobs, both invisible until you know to look for them. **Opening pinholes for port-negotiating protocols.** Some protocols use one connection to negotiate a second, dynamically chosen port. FTP in active mode is the classic example: the client and server agree on a data port over the control channel, and the data connection then opens on that port. A stateless ACL cannot predict that port, so it either blocks the data channel or you have to open a huge range. The FTP inspection engine reads the control channel, sees the negotiated port, and opens a precise, temporary pinhole for exactly that connection. SIP, H.323, and several others work the same way. Inspection is the only clean way to allow these protocols without punching permanent holes in your policy. **Enforcing protocol conformance.** The second job is a security control. Because the engine parses the payload, it can verify that traffic on a given port actually follows that protocol's rules. Malformed messages, illegal commands, and payloads that do not match the claimed protocol can be flagged or dropped. This is what lets an ASA reject traffic that is technically reaching an allowed port but is not really the protocol it claims to be, a trick attackers use to tunnel through firewalls. Pinhole creation Reads control channels (FTP, SIP) and opens precise, temporary ports for the negotiated data connection. A stateless ACL cannot do this. Protocol conformance Parses the payload and enforces the protocol's rules, dropping malformed or non-conformant traffic on an allowed port. NAT fixup Rewrites embedded IP addresses inside payloads (FTP, DNS) so NAT does not break the application. ## Seeing the default and a custom engine side by side The clearest way to understand inspection is to see the factory default and a custom per-interface engine at the same time. On our lab ASA we kept the default global inspection running and added a custom HTTP inspection on the DMZ interface, where a published nginx server lives. The `show service-policy inspect http` command prints both. ``` FW1# show service-policy inspect http Global policy: Service-policy: global_policy Class-map: inspection_default Interface dmz: Service-policy: PLZ-DMZ-POLICY Class-map: PLZ-HTTP-CLASS Inspect: http, packet 60, lock fail 0, drop 0, reset-drop 0 ``` Read this top to bottom. The `Global policy` block is the factory default: `global_policy` holding `inspection_default`, applied everywhere. The `Interface dmz` block is our custom policy, where `PLZ-HTTP-CLASS` is running the HTTP inspection engine with a live `packet 60` counter from real client traffic. Both are active. The interface policy handles the DMZ, the global policy handles the rest, and they coexist because the more specific interface policy takes precedence for HTTP on that one interface. That `packet 60` is genuine. It is sixty packets from real HTTP transactions a Linux client made against the published server, classified and handed to the inspection engine. `drop 0` confirms every packet was conformant HTTP and passed. When you start writing custom Layer 7 policies that block specific content, this is the same counter that will show non-zero drops, which we cover in the custom inspection article. ## Inspection is not the same as an ACL It is worth being precise about the boundary, because this trips people up. An ACL is a permit or deny decision based on addresses and ports. It runs once, early in the packet flow, and it does not remember anything about the application. Inspection runs later, maintains state about the application session, and can act on the contents of the payload across many packets. The two work together rather than competing. The ACL decides whether a connection is allowed to exist at all. Inspection then watches that allowed connection and enforces how it behaves, opening pinholes and dropping non-conformant messages. You need both. An allowed FTP session with no inspection cannot open its data channel cleanly; an inspected session that the ACL never permitted never gets inspected because it was dropped first. Where each of these sits in the overall processing order is exactly what our [ASA packet flow](https://www.pinglabz.com/cisco-asa-packet-flow/) article maps out. ## How inspection keeps state The word "stateful" carries a lot of weight here. When the ASA inspects a connection, it is not making an isolated decision on each packet; it is tracking the whole conversation. For FTP that means remembering that a control channel exists and that a specific data port was just negotiated, so the return connection can be matched to its parent. For ICMP it means recording that an echo request went out so the matching echo reply is allowed back in without a permanent permit rule. For DNS it means pairing a query with its response and tearing the pseudo-connection down once the answer arrives. This state is why inspection can be both permissive and safe at once. It permits exactly the return traffic a legitimate session requires, for exactly as long as the session needs it, and not a moment longer. A stateless ACL cannot express "allow the reply to this specific request", so it has to allow a whole class of traffic all the time. Inspection replaces a broad standing rule with a narrow, temporary, session-scoped one, which is a strictly better security posture. The cost is CPU and memory to hold the state, which on a busy box is a real design consideration, but for the protocols that need it the trade is almost always worth making. ## Troubleshooting inspection When an inspected protocol misbehaves, the counters point you at the cause quickly. Start with `show service-policy inspect ` and read the packet and drop numbers. A climbing `packet` counter with `drop 0` means the engine is seeing and passing traffic, so the problem is elsewhere. A non-zero `drop` means inspection itself is discarding packets, and the syslog will usually name the reason. A `packet 0` counter means the engine is applied but no traffic is matching its class-map, which sends you back to check the classifier. The classic inspection failure is a protocol that works until NAT gets involved, then breaks. That is almost always an embedded-address problem: the application put an IP address inside its payload, NAT rewrote the header but not the payload, and the two no longer agree. The relevant inspection engine is exactly what fixes this by rewriting the embedded address too, so the fix is usually to confirm inspection for that protocol is actually enabled rather than to touch NAT at all. It is a good reminder that inspection is not only a security control, it is frequently what makes an application work through the firewall in the first place. ## A version caveat worth respecting Application inspection has a long history, and Cisco has retired some of the older engines over the years. On the ASAv 9.24 we tested, several legacy inspection engines that appear in older documentation and study guides have been deprecated, so a command you copied from a 2015 configuration guide may simply not exist on a current build. The workhorses are all still present and current: HTTP, FTP, ICMP, and DNS inspection are exactly what you would expect and behave as documented. The practical rule is to verify against your own version rather than trusting a study guide. Run `show run policy-map` and `show service-policy` on the actual box, and confirm the inspection you want is supported before you build a design around it. Inspection engines are one of the areas where the gap between old documentation and current firmware bites hardest, and the fix is simply to check the machine in front of you. HTTP Current Layer 7 web inspection, the basis for URI and header filtering. FTP Current Opens data-channel pinholes and rewrites embedded addresses for NAT. ICMP Current Statefully matches echo replies to requests instead of blanket-permitting ICMP. DNS Current Enforces message limits and guards against DNS-based attacks and spoofing. ## When to add a custom inspection engine The default policy is deliberately conservative. It inspects the common protocols with default parameters and no content filtering. You add a custom inspection engine when you need something the default does not do: block a specific URL, enforce a stricter HTTP policy, restrict FTP commands, or apply inspection only to one interface. In every case the pattern is the one from the MPF article: write a class-map to select the traffic, a policy-map to attach the inspection action, and a service-policy to bind it, usually to a specific interface so it overrides the global default for that traffic. The most valuable custom engine is Layer 7 HTTP inspection, because it lets you match and drop on the contents of web requests, the URI, headers, methods, and body. That is where inspection stops being plumbing and becomes an active content control, and it is the subject of the next article in the series. ## Key Takeaways - **Application inspection reads the payload, not just headers.** It does what a stateless ACL cannot: understand the application session across many packets. - **It ships on by default.** The factory `global_policy` / `inspection_default` / `service-policy global` stack inspects common protocols on every interface with no configuration from you. - **Two core jobs.** Opening pinholes for port-negotiating protocols like FTP and SIP, and enforcing protocol conformance to drop non-conformant traffic on allowed ports. - **Default and custom engines coexist.** `show service-policy inspect http` showed the global default and a per-interface HTTP engine with a live `packet 60` counter at the same time. - **Verify against your version.** ASAv 9.24 deprecated some legacy engines. HTTP, FTP, ICMP, and DNS are current, but check the box rather than an old guide. Application inspection is the foundation. The real power shows up when you write your own Layer 7 policy and block exactly what you choose. Keep going in the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series with custom HTTP inspection, where a real regex blocks a real URL for a real client, and revisit the [inspection engines overview](https://www.pinglabz.com/cisco-asa-inspection-engines/) for the full catalog of what each engine handles. ### The ASA Modular Policy Framework (MPF): class-map, policy-map, service-policy URL: https://www.pinglabz.com/asa-modular-policy-framework/ Last updated: 2026-07-13T11:48:37.000Z If you have ever stared at an ASA configuration and wondered how a single firewall applies HTTP inspection here, a connection limit there, and TCP normalization somewhere else without a tangle of one-off commands, the answer is the Modular Policy Framework (MPF). MPF is the engine underneath nearly every advanced ASA feature, and once you see its three-layer shape you stop memorizing syntax and start reasoning about it. This article is part of our [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series, and it builds the framework from scratch on a real ASAv 9.24 running in CML, with live counters climbing as a Linux client hits a published web server. The reason MPF matters is that Cisco deliberately reused one structure for everything. Application inspection, quality of service, connection limits, TCP normalization, and policing all plug into the same three-layer model. Learn it once and every feature after that is just a different action bolted onto the same skeleton. That reuse is the literal meaning of the word "modular" in Modular Policy Framework. ## The three-layer model MPF answers three separate questions, and it keeps them in three separate objects. This separation is the whole design. You never mix "which traffic" with "what to do about it", because that is exactly the mess MPF was invented to avoid. - **class-map** answers **WHICH** traffic. It is a classifier: match a port, an ACL, a DSCP value, or a protocol. It says nothing about what happens next. - **policy-map** answers **WHAT** action. It references one or more class-maps and attaches an action to each: inspect this, police that, limit connections on the other. - **service-policy** answers **WHERE**. It binds a policy-map to an interface (or globally), which is the moment the policy actually starts acting on packets. Nothing happens until all three exist and the service-policy is applied. A class-map with no policy-map is inert. A policy-map with no service-policy is just a definition sitting in the config. The service-policy is the switch that turns the whole stack on. class-map WHICH traffic The classifier. Matches ports, ACLs, DSCP, tunnel groups, or protocols. Identifies the flows, decides nothing. policy-map WHAT action The action set. References class-maps and attaches inspect, police, set connection, priority, and more to each class. service-policy WHERE it applies The binding. Attaches a policy-map to an interface or globally. This is the command that makes the policy live. ## Building MPF from scratch Let us build the simplest useful policy: inspect HTTP on the DMZ interface, where a published nginx server lives. We work bottom up, matching the three questions in order. First, classify the traffic we care about. ``` class-map PLZ-HTTP-CLASS <-- 1. class-map: WHICH traffic match port tcp eq www ``` That is the entire classifier. It matches TCP destined for port 80\. It does not inspect anything, drop anything, or care which interface the traffic arrives on. It only names a set of packets so a policy-map can point at them. Next, decide what to do with that class. We create a policy-map, reference the class inside it, and attach an action. ``` policy-map PLZ-DMZ-POLICY <-- 2. policy-map: WHAT action class PLZ-HTTP-CLASS inspect http ``` Now the policy says: for traffic matched by PLZ-HTTP-CLASS, run the HTTP application inspection engine. Still nothing is happening to live packets, because the policy is not attached anywhere. That is the final step. ``` service-policy PLZ-DMZ-POLICY interface dmz <-- 3. service-policy: WHERE (bind to an interface) ``` The instant that command lands, the ASA starts classifying traffic on the DMZ interface, matching HTTP, and inspecting it. Three objects, three questions, one live policy. ## Watching the counters climb Theory is cheap. The proof is in the counters. With the policy applied, we drove five real HTTP requests from a Linux client against the published server and then read the per-interface service policy on the ASA. These are genuine packets from a genuine curl loop, not a lab simulation. ``` FW1# show service-policy interface dmz Interface dmz: Service-policy: PLZ-DMZ-POLICY Class-map: PLZ-HTTP-CLASS Inspect: http, packet 60, lock fail 0, drop 0, reset-drop 0, 5-min-pkt-rate 0 pkts/sec ``` `packet 60` is the whole story in two words. Sixty packets were classified by PLZ-HTTP-CLASS and handed to the HTTP inspection engine, which is exactly what you would expect from a handful of real HTTP transactions once you count the TCP handshakes, requests, responses, and teardowns. `drop 0` and `reset-drop 0` confirm the engine inspected everything and dropped nothing, because plain HTTP to a normal URI is perfectly conformant. Run more requests and refresh, and the number climbs. That live counter is how you confirm a class-map is actually catching the traffic you think it is. If `packet` stays at zero while you know traffic is flowing, your class-map is not matching. That single number is the fastest MPF troubleshooting signal on the box: it tells you whether the "WHICH" layer is doing its job before you ever look at the action. ## How a class-map can match Our example matched a single TCP port, but that is only one of the ways a class-map classifies traffic. In practice you will reach for several match types depending on how precisely you need to carve up flows, and knowing the menu keeps you from writing an ACL when a one-line match would do. - **match port** \- a single port or a range, as in `match port tcp eq www`. Fast and readable for a well-known service. - **match access-list** \- point at an extended ACL when you need to classify by source, destination, and port together. This is the most flexible option and the one you fall back to when a simple port match is too blunt. - **match dscp** \- classify by DSCP marking, which is how QoS policies pick out already-marked voice or video. - **match default-inspection-traffic** \- a special built-in match that selects the standard set of protocols the ASA knows how to inspect, using each protocol's well-known port. This is what the factory `inspection_default` class-map uses, and it is why the default policy just works without you listing every port. - **match any** \- matches every packet, used when an action should apply to all traffic on an interface, such as a blanket TCP normalizer. A class-map can also be `match-all` or `match-any` when it carries multiple match lines, which controls whether every condition must be true or just one of them. For the single-match classifier above it makes no difference, but it becomes important the moment you start combining conditions, which is exactly what Layer 7 inspection class-maps do. ## Why this is called "modular" Here is the payoff, and the reason this article sits at the front of the ASA deep-dive. The exact three-layer stack you just built for HTTP inspection is the same stack every other ASA feature uses. Nothing about the skeleton changes. Only the action inside the policy-map changes. Application inspection policy-map action: `inspect http`, `inspect ftp`, `inspect dns`. Same class-map, same service-policy. Quality of service policy-map action: `priority` or `police`. Classify by DSCP, then queue or rate-limit. Connection limits policy-map action: `set connection conn-max` and embryonic limits to blunt SYN floods. TCP normalization policy-map action: `set connection advanced-options` pointing at a tcp-map. Every one of those features answers the same three questions. Which traffic? A class-map. What action? A policy-map. Where? A service-policy. When you learn a new ASA feature, you are not learning a new configuration model, you are learning one new action verb to drop into a policy-map you already know how to write. That is modularity in the most practical sense: one framework, many actions. It also explains the ASA's default configuration. Out of the box the firewall ships with a `global_policy` policy-map, an `inspection_default` class-map, and a `service-policy global_policy global` line. That is MPF applied to itself, running the default inspection set on every interface. We will pull that default apart in the [ASA inspection engines](https://www.pinglabz.com/cisco-asa-inspection-engines/) deep-dive, but notice that even Cisco's factory config is just the same three layers you built by hand above. ## Global versus interface policies You can apply a service-policy in two scopes, and knowing the rule prevents a lot of confusion. A service-policy attached with `interface dmz` is an interface policy and it acts only on that interface. A service-policy attached with the keyword `global` acts on every interface that does not already have a matching interface policy. The precedence rule is simple: an interface service-policy overrides the global service-policy for that specific feature and that specific interface. That is why our per-interface PLZ-DMZ-POLICY can inspect HTTP on the DMZ while the factory global\_policy keeps running the default inspection set everywhere else. They coexist because they are scoped differently, and the more specific interface policy wins where they overlap. This is worth internalizing before you start layering custom inspection on top of the defaults, because a surprising number of "my inspection is not applying" tickets come down to a global policy and an interface policy quietly fighting over the same traffic. ## Common MPF mistakes Most MPF problems are not exotic. They cluster around a handful of predictable mistakes, and every one of them maps back to the three-layer model. If you keep the three questions in mind, you can usually diagnose the issue before you touch a debug command. The first is defining a class-map and policy-map but forgetting the service-policy, so nothing is ever applied. The counters give this away instantly: there is no interface output to read because the policy is not bound. The second is a class-map that does not match the traffic you expected, which shows up as a `packet 0` counter while traffic is clearly flowing. The third, and the subtlest, is an interface policy and the global policy quietly overlapping so that the more specific interface policy wins and your global rule appears to do nothing on that interface. The fix in every case is the same: read the config bottom up, confirm the service-policy is applied, then confirm the class-map is matching before you ever suspect the action. There is one more constraint worth knowing early. A single interface can only have one service-policy at a time, and a policy-map can only be applied to one place per direction. If you need several actions on the same traffic, you stack multiple `class` statements inside one policy-map rather than trying to apply two service-policies to the same interface. That constraint is exactly why the framework encourages you to think of a policy-map as a container of actions rather than a single rule. ## Reading an MPF config quickly When you open an unfamiliar ASA, read MPF from the bottom up. Find the `service-policy` lines first, because they tell you which policy-maps are actually live and where. Then read each referenced policy-map to see the actions. Then read the class-maps to see what traffic those actions hit. Working backwards from "what is applied" to "what it matches" is far faster than reading top to bottom and trying to hold the whole thing in your head. The verification commands follow the same logic. `show service-policy interface dmz` shows a single interface's live policy with counters, which is what we used above. `show service-policy global` shows the global policy and its counters. Both print the class-map, the action, and the per-action packet counts, so you can confirm at a glance not just that a policy exists but that it is catching traffic. ## Key Takeaways - **MPF has exactly three layers.** class-map = WHICH traffic, policy-map = WHAT action, service-policy = WHERE it applies. Every ASA feature you configure fits this shape. - **Nothing is live until the service-policy is applied.** A class-map or policy-map on its own is just a definition. The service-policy is the on switch. - **The `packet` counter is your first troubleshooting signal.** `show service-policy interface dmz` showed `packet 60` from real HTTP flows. Zero means your class-map is not matching. - **"Modular" is literal.** Inspection, QoS, connection limits, and TCP normalization all reuse the same three-layer stack, changing only the action verb inside the policy-map. - **Interface policies override the global policy** for the same feature on the same interface, which is how custom inspection coexists with the factory default\_inspection set. With the framework clear, the natural next step is to see what the ASA's default inspection policy actually does and how application inspection differs from a plain ACL. Continue with the rest of the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series, starting with the [application inspection engines](https://www.pinglabz.com/cisco-asa-inspection-engines/) that plug into the policy-map you just built. ### Choosing ASA High Availability: Failover vs Clustering vs Contexts URL: https://www.pinglabz.com/asa-ha-design-comparison/ Last updated: 2026-07-13T11:10:52.000Z The Cisco ASA offers a whole family of high-availability and scale features: failover in two flavors, redundant interfaces, EtherChannel, security contexts, resource classes, and clustering. Faced with that menu, the useful question is not "what exists" but "what should I actually deploy, and what can my platform even run." This article is the decision guide for the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series, and it is built on something rare: we tested every one of these features live on a real virtual ASA (asav 9.24). Exactly one of them worked, and it worked flawlessly. The rest were rejected by the parser. That gives us a comparison table where every "supported on asav" cell is a tested fact, not a guess. ## The Proven Baseline: Zero-Loss Active/Standby Failover Start with what works, because it is genuinely excellent. On our asav pair, [Active/Standby failover](https://www.pinglabz.com/cisco-asa-active-standby-failover/) formed automatically and replicated its entire configuration to the standby with no manual copy. The failover state showed a clean, synced pair: FW1# show failover state State Last Failure Reason This host - Primary Active None Other host - Secondary Standby Ready None \====Configuration State=== Sync Done "Sync Done" means the standby received the full running configuration (SSH, ACLs, ICMP inspection, all of it) automatically. That is the foundation of [stateful failover](https://www.pinglabz.com/cisco-asa-stateful-failover/): not just a config copy, but a live connection table replicated to the standby so it can take over mid-session. Then we tested it the only way that really counts. An inside host ran a 45-packet ping stream through the firewall to an outside address, and mid-stream we forced the active unit to give up its role. The stream did not flinch: HOST1$ ping -c 45 -i 0.3 203.0.113.2 64 bytes from 203.0.113.2: icmp\_seq=44 ttl=255 time=6.52 ms 64 bytes from 203.0.113.2: icmp\_seq=45 ttl=255 time=6.56 ms \--- 203.0.113.2 ping statistics --- 45 packets transmitted, 45 received, 0% packet loss, time 13242ms Forty-five sent, forty-five received, zero loss, straight through a full active-unit failover. Under the hood the standby ran the complete state machine to take over, which the failover history records step by step: FW1# show failover history Standby Ready Just Active Other unit wants me Active Just Active Active Drain Other unit wants me Active Active Drain Active Applying Config Other unit wants me Active Active Applying Config Active Config Applied Other unit wants me Active Active Config Applied Active Other unit wants me Active The standby moved from Standby Ready to Active in a handful of transitions, took over the active IP and MAC, and carried the existing connection so seamlessly the ping never dropped. This is the HA story you can lab on a virtual ASA today, and it is the one most designs should use. ## The Full Menu, Tested Live on asav Here is the whole HA and scale family with the one column that matters most: whether the virtual appliance can actually run it. Every "supported on asav" answer below was verified on the real device, and every "No" is a parser rejection we captured, not an assumption. Active/Standby failover Purpose: Unit-level HA, stateful On asav: YES (proven, 0% loss) Redundant interface Purpose: Sub-second link HA in a unit On asav: NO (parser reject) EtherChannel Purpose: Bandwidth + link HA On asav: NO (parser reject) Active/Active failover Purpose: Load-shared, needs contexts On asav: NO (single mode only) Security contexts Purpose: Multi-tenant virtualization On asav: NO (mode multiple invalid) Resource classes Purpose: Fairness across contexts On asav: NO (needs contexts) Clustering Purpose: Scale-out to 16 units On asav: NO (parser reject) One green cell out of seven. That is not a disappointment, it is clarity: on a virtual ASA, Active/Standby stateful failover is the HA story, and the other six features require physical ASA/Firepower hardware or FTD. Every "No" in that grid corresponds to a real `% Invalid input detected at '^' marker.` we captured on the device, documented in the individual articles for [redundant interfaces](https://www.pinglabz.com/asa-redundant-interfaces/), [EtherChannel](https://www.pinglabz.com/asa-etherchannel-port-channel/), [Active/Active](https://www.pinglabz.com/asa-active-active-failover/), [shared-interface contexts](https://www.pinglabz.com/asa-shared-interface-contexts/), [resource classes](https://www.pinglabz.com/asa-context-resource-classes/), and [clustering](https://www.pinglabz.com/asa-clustering-spanned-vs-individual/). ## Choosing By Goal Work backwards from the problem you are solving rather than from the feature list. - **"I need my firewall to survive a unit failure."** Active/Standby stateful failover. It is simple, it is lossless in practice, and it runs on every ASA including asav. This is the right answer for the overwhelming majority of deployments. - **"I need to survive a single link failure faster than a unit failover."** Redundant interface (physical only). Layer it under failover so a dead cable is handled locally in about a second. - **"I need more bandwidth than one interface provides."** EtherChannel (physical only), which also adds link redundancy as a bonus. - **"I need both firewalls forwarding at once."** Active/Active failover (physical, multi-context only), and be ready for the asymmetric-routing complexity it brings. - **"I need many isolated tenants on one box."** Security contexts with resource classes (physical only), or, on virtual, a separate asav per tenant. - **"I need to scale past what one pair can forward."** Clustering (physical only), or FTD clustering on the modern platform. ## The Honest Recommendation For a virtual ASA, use Active/Standby stateful failover and stop there for HA. It is not a compromise or a lesser option; it is a genuinely excellent, proven design that survived a live failover in our lab without losing a single packet. Do not architect a virtual deployment around redundant interfaces, EtherChannel, Active/Active, contexts, resource classes, or clustering, because the asav parser will reject every one of them. For those features you need physical ASA or Firepower hardware, and increasingly the right destination is FTD (Firepower Threat Defense), which is where Cisco has concentrated its modern investment in scale-out and multi-tenancy. We will cover that on the [Cisco FTD](https://www.pinglabz.com/cisco-ftd/) pillar as it comes online. And if you are thinking about HA more broadly than just the firewall, the same first-hop and gateway-redundancy principles show up across the network in the [FHRP](https://www.pinglabz.com/fhrp/) family, which is worth understanding alongside firewall failover. The through-line of this whole series is simple: test the feature on the platform you actually run, believe the parser, and design for what the device can really do. ## Why the Parser Reject Is Trustworthy Evidence It is worth pausing on why "the command was rejected" is such strong evidence, because it is more reliable than a feature-support matrix in a datasheet. A datasheet tells you what a product line is supposed to do; the parser tells you what the software image in front of you actually implements. When the asav returns `% Invalid input detected at '^' marker.` with the caret under the very first keyword of a command, it is not refusing a licensed feature or asking for a precondition. It is telling you the keyword does not exist in this image at all. There is no license to buy, no mode to enter, no reload that adds it, because the code path is simply not present in the virtual appliance. That is why every "No" in the support grid above is anchored to a captured rejection rather than to documentation: we asked the device directly, and it answered. For anyone designing a real deployment or preparing for a lab exam, this is the habit worth building. Test the feature on the exact platform and image you will run, and let the parser settle the argument. ## Failover for the Firewall, FHRP for the Gateways Around It High availability is not only a firewall concern, and it helps to keep the layers straight. Failover (and clustering) make the firewall itself redundant. But the routers and gateways on either side of the firewall have their own redundancy story, and that is the job of first-hop redundancy protocols. A pair of routers presenting a single virtual gateway address to hosts, so that the loss of one router is invisible, is exactly the [FHRP](https://www.pinglabz.com/fhrp/) pattern (HSRP, VRRP, GLBP). The firewall's failover and the surrounding gateways' FHRP solve the same class of problem (survive the loss of a device without the traffic noticing) at different points in the path. A well-designed edge uses both: stateful failover so the firewall pair survives a unit loss, and FHRP so the routers around it survive a router loss. Thinking about them together keeps you from building a firewall that never fails behind a gateway that does, or the reverse. ## Common ASA HA Design Mistakes A few patterns cause most of the avoidable trouble. The first is reaching for Active/Active because "both boxes forwarding" sounds more efficient, without accounting for multiple context mode, asymmetric routing, and the fact that a failure still dumps the whole load onto one unit; for most sites Active/Standby is simpler and just as available. The second is designing a virtual deployment around a hardware feature and only discovering at build time that the parser rejects it, which is the entire reason this series tests everything live. The third is under-provisioning the links that HA depends on: the failover state link and the cluster control link carry the replication and coordination that make the whole thing coherent, and starving them produces flapping and split-brain scenarios that look like random instability. The fourth is treating HA as set-and-forget: failover and clustering both need periodic testing (deliberately failing a unit in a maintenance window) so you learn the standby or the survivors genuinely take over, rather than discovering during a real outage that replication was broken for weeks. The zero-loss failover we captured was only meaningful because we actually forced the failover and watched the traffic; an untested standby is a hope, not a design. ## Key Takeaways - Of the seven ASA HA and scale features, only Active/Standby stateful failover runs on the virtual asav, and it is genuinely excellent: we proved a live 45-of-45 ping stream survived a failover at 0% loss. - Redundant interfaces, EtherChannel, Active/Active, security contexts, resource classes, and clustering are all rejected by the asav parser; each "No" in the comparison grid is a captured `% Invalid input` error, not a guess. - Choose by goal: unit survival means Active/Standby; link survival means redundant interfaces; bandwidth means EtherChannel; both-active means Active/Active; multi-tenancy means contexts; scale-out means clustering. - For virtual ASA, deploy Active/Standby stateful failover and do not design around features the appliance cannot run. - For the hardware-only features, use physical ASA/Firepower or FTD, which is where Cisco's modern scale-out and multi-tenant engineering lives. - See the proven [Active/Standby failover](https://www.pinglabz.com/cisco-asa-active-standby-failover/) and [stateful failover](https://www.pinglabz.com/cisco-asa-stateful-failover/) guides, the broader [FHRP](https://www.pinglabz.com/fhrp/) HA concepts, and the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). ### ASA Clustering Explained: Spanned vs Individual Interface Mode URL: https://www.pinglabz.com/asa-clustering-spanned-vs-individual/ Last updated: 2026-07-13T11:10:52.000Z Failover gives you two firewalls acting as one for redundancy. Clustering takes that idea and scales it out: up to sixteen Cisco ASA units pooled into a single logical firewall that shares both the load and the state. When you outgrow what one box (or one active/standby pair) can forward, clustering is how you add throughput without adding management surface, because the whole cluster still looks and behaves like one firewall. This article is part of the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series, and it is written from a lab that could not run it: the virtual asav 9.24 appliance rejects the cluster commands outright. That rejection is real, we captured it, and it tells you exactly where clustering does and does not live. ## What Clustering Is For A cluster is a group of ASA units that pool their resources and act as one logical device for both configuration and traffic. You configure it once, from the control unit, and that configuration replicates to every member. Traffic is spread across all the members, so the aggregate forwarding capacity scales roughly with the number of units, and if a member fails, the connections it was handling are picked up by the survivors from replicated state. In one feature you get horizontal scale (more throughput than any single box) and resilience (loss of a member does not drop the cluster). This is a different goal from failover: failover is about surviving the loss of a unit, clustering is about surviving it *and* using all the units at once for capacity. The cluster elects a **control unit** (sometimes called the master) that owns the configuration and coordinates the cluster; the rest are **data units**. You make configuration changes only on the control unit and they propagate. If the control unit fails, the cluster elects a new one from the remaining members, and forwarding continues throughout. The whole point is that no single member is special from a traffic standpoint: every unit forwards, and the control role is just coordination. ## Two Interface Modes: Spanned vs Individual How the cluster presents its interfaces to the surrounding network is the single most important design choice, and there are two modes. ### Spanned EtherChannel Mode In spanned EtherChannel mode, the interfaces of every cluster member are bundled into a single EtherChannel that spans all the units, and the surrounding switches see one Port-channel with one set of MAC and IP addresses. The switch load-balances flows across the member links using its normal EtherChannel hashing, and each flow lands on whichever unit owns that member link. The cluster is completely transparent at Layer 2 and Layer 3: neighbors see one interface, one MAC, one IP, regardless of how many units are behind it. This is the recommended and most common mode, because it needs no routing awareness of the individual units and it fails over invisibly. It does, however, depend on the surrounding switches supporting a spanned EtherChannel (often via VSS or vPC on the switch side) so the bundle can cross multiple physical chassis. ### Individual Interface Mode In individual interface mode, each cluster member has its own IP address on the data interfaces, and the network is made aware of all of them, typically through dynamic routing with equal-cost multipath so traffic is distributed across the members' addresses. The control unit owns a main cluster IP for management, but data-plane traffic reaches each member on its own address. This mode does not require spanned EtherChannel support on the switches, but it does require the surrounding Layer 3 network to participate in distributing traffic (usually via routing and ECMP), which means more routing configuration and less transparency than spanned mode. It is the choice when the switching layer cannot provide a spanned EtherChannel but the routing layer can spread the load. Spanned EtherChannel Presents as: One interface, one MAC/IP Load spread by: Switch EtherChannel hash Transparent to neighbors. Needs spanned-EtherChannel-capable switches (VSS/vPC). Recommended default. Individual Interface Presents as: One IP per member Load spread by: Routing / ECMP No spanned EtherChannel needed, but the L3 network must distribute traffic. More routing config. ## The Flow Roles: Owner, Director, Forwarder Clustering's cleverness is in how it keeps a single connection consistent across many units. Because the switch or router may deliver different packets of the same flow to different members, the cluster assigns roles per connection so exactly one unit remains authoritative: - **Owner.** The unit that first received the connection's initial packet. The owner holds the full connection state and processes the flow. There is one owner per connection. - **Director.** A unit chosen (by a hash of the connection) to be the backup keeper of the connection's state and the lookup authority for it. If another unit receives a packet for a connection it does not own, it asks the director who the owner is. The director also holds a backup of the state, so if the owner fails, the director can promote a new owner without dropping the connection. - **Forwarder.** A unit that receives packets for a connection it does not own. It consults the director, learns the owner, and forwards the packets to the owner over the cluster control link. Forwarding is the mechanism that lets asymmetric delivery still resolve to a single authoritative owner. These roles are per connection, not per unit: the same physical box can be the owner of one flow, the director of another, and a forwarder for a third, all at once. Together they guarantee that every connection has exactly one owner processing it and one director backing it up, no matter how the network sprays packets across the members. This is what makes the cluster behave like one stateful firewall rather than sixteen independent ones. ## Connection Redistribution When Membership Changes When a unit joins or leaves the cluster, the connection ownership has to rebalance. If a member fails, the connections it owned are recovered by their directors, which promote new owners from the surviving members using the backed-up state, so established sessions survive the loss of their owning unit. When a new member joins, the cluster begins directing new connections to it (existing connections generally stay with their current owners to avoid needless disruption). The control link, a dedicated cluster interface between all members, carries this state backup, the director lookups, and the forwarded packets, so it must have enough bandwidth and low enough latency to keep up with the cluster's load. An undersized or congested cluster control link is a classic cause of clustering instability, because everything that keeps the cluster coherent rides across it. ## The Real asav Rejection All of this is hardware territory. Our lab firewalls are virtual asav 9.24 instances, and the cluster commands are simply not in the virtual appliance's parser. Setting the interface mode and checking cluster status, this is exactly what the real device returned: FW1(config)# cluster interface-mode spanned ^ ERROR: % Invalid input detected at '^' marker. FW1# show cluster info Clustering is not configured The `cluster interface-mode` command is rejected at the parser (the caret sits under the very first keyword), and `show cluster info` reports that clustering is not configured, with no commands available to configure it. This is not a licensing gate or a hidden mode; the clustering feature set is absent from asav. The honest takeaway is that ASA clustering is a physical-appliance capability, and the parser reject above is the proof. ## Which Platforms Support ASA Clustering, and the Modern Answer ASA clustering runs on specific physical ASA and Firepower models and software versions (the exact list depends on the hardware generation and release, so always check the release notes for your platform). The virtual asav is not among them. If you need scale-out on virtual or on modern hardware, the direction Cisco has invested in is FTD (Firepower Threat Defense): Firepower clustering leans on spanned EtherChannel much like ASA clustering, and FTD multi-instance lets you carve appliances into isolated logical devices. For a virtual scale-out story, FTD and its clustering and multi-instance features are where the engineering effort has gone, not asav. On the ASA virtual appliance itself, the resilience story you actually have is the one you can prove: [Active/Standby failover](https://www.pinglabz.com/cisco-asa-active-standby-failover/), which the asav supports fully and which we demonstrated in this same lab with a zero-loss failover of a live traffic stream. Clustering scales beyond that, but only on the platforms that carry the feature. ## Configuring a Cluster on Hardware, in Outline On a supported physical platform, bringing up a cluster follows a repeatable shape. You set the interface mode first, because it cannot be changed while clustering is running, then define the cluster group and its control link, then bootstrap each member. The skeleton looks like this on the first (control) unit: ``` cluster interface-mode spanned ! interface TenGigabitEthernet0/6 description Cluster Control Link no shutdown ! cluster group PINGLABZ-CL local-unit unit-1 cluster-interface Port-channel1 ip 10.0.0.1 255.255.255.0 priority 1 enable ``` Each additional unit gets the same cluster group configuration with its own `local-unit` name and control-link address, and joins the existing cluster rather than forming a new one. Once a unit runs `enable` in the cluster group, it discovers the others over the control link, the cluster elects a control unit by priority, and the configuration replicates outward. From that point you manage the whole cluster from the control unit, and `show cluster info` lists every member with its role and state. The two decisions you make up front, the interface mode and the control-link interface, are the ones that are painful to change later, so they deserve the most design thought. ## How Clustering Differs From Failover It is easy to blur clustering and failover together because both give resilience, but they answer different questions and it is worth being precise. Failover (Active/Standby) is about availability: two units, one forwarding, one waiting, and the standby exists purely to take over. You do not get more capacity from the standby; it is insurance. Clustering is about availability *and* capacity: every member forwards all the time, so adding a unit adds throughput, and the loss of a unit costs you that unit's share rather than everything. Failover tops out at two units; clustering scales to many. Failover state replication is a simple primary-to-standby copy; clustering state is distributed with the owner/director/forwarder roles so that any member can be lost without a single point holding all the state. If your problem is "one firewall might die," failover solves it cheaply. If your problem is "one firewall (or one pair) cannot forward enough traffic," only clustering solves it, and only on the platforms that support it. There is also a management difference. A failover pair is configured largely as one unit with a mirrored partner; a cluster is configured once on the control unit and pushed to all members, and the cluster elects a new control unit automatically if the current one fails. The result is that a sixteen-unit cluster is no more management work than a single firewall, which is a large part of its appeal at scale. ## Clustering Gotchas Worth Knowing A few realities catch people the first time they build a cluster. The interface mode (spanned or individual) is set before clustering is enabled and cannot be flipped on a live cluster, so choosing wrong means a disruptive rebuild; decide based on whether your switching layer can provide a spanned EtherChannel. The cluster control link is not a place to economize: it carries state backups, director lookups, and every forwarded packet for asymmetric flows, so it should be high-bandwidth and ideally its own EtherChannel, and it must have consistent low latency across all members (which practically limits how far apart cluster members can sit). Not every ASA feature is supported in clustering, and some behave differently when clustered, so the feature matrix in the release notes for your platform is required reading rather than optional. Finally, clustering asymmetric traffic still costs a control-link hop through the forwarder to reach the owner, so a design that constantly sprays both directions of every flow across different members will lean hard on the control link; the more symmetric your traffic delivery, the less forwarding overhead the cluster pays. None of these apply on asav, of course, because the feature never initializes there in the first place. ## Key Takeaways - Clustering pools up to sixteen ASA units into one logical firewall that shares load and state, scaling throughput while still managing as a single device from the control unit. - Spanned EtherChannel mode presents one interface, one MAC and IP to the network (transparent, recommended, needs VSS/vPC-style switches); individual interface mode gives each member its own IP and relies on routing/ECMP to distribute traffic. - Per-connection roles keep the cluster coherent: the owner processes the flow, the director backs up its state and answers ownership lookups, and a forwarder relays packets it received to the owner over the control link. - On member failure, directors promote new owners from backed-up state so connections survive; the cluster control link must be sized for state backup, lookups, and forwarded traffic. - On asav 9.24 the feature is absent: `cluster interface-mode spanned` returns `% Invalid input detected at '^' marker.` and `show cluster info` says clustering is not configured. - Clustering is a physical ASA/Firepower feature; for virtual scale-out the modern answer is FTD clustering or multi-instance, and on asav itself the proven resilience is [Active/Standby failover](https://www.pinglabz.com/cisco-asa-active-standby-failover/). See the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). ### ASA Resource Classes: Stopping One Context From Eating the Firewall URL: https://www.pinglabz.com/asa-context-resource-classes/ Last updated: 2026-07-13T11:10:51.000Z Once you slice a Cisco ASA into multiple security contexts, you have created a shared-resource problem. All those virtual firewalls draw from the same finite pools: connections, translations, host entries, inspection capacity, and more. Nothing stops one badly behaved or busy context from consuming so much of a pool that the others are starved. Resource classes are the mechanism that stops that, letting you cap and guarantee slices of the firewall's capacity per context. This article is part of the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series, and it carries the same honest caveat as the rest of the multi-context family: the feature lives only in multiple context mode, and our virtual lab appliance cannot enter that mode. We will show the real rejection rather than fake a demo. ## The Problem: One Context Eating the Firewall Imagine three tenants sharing one ASA in multiple context mode. Tenant A runs a normal web workload. Tenant B gets hit by a connection flood (a scan, a misbehaving application, or an outright attack) and starts opening connections and building translations as fast as the firewall will let it. Because all three contexts draw from the same global limits, tenant B's surge consumes the shared connection and xlate pools, and now tenant A and tenant C cannot open new connections either. One tenant's problem has become everyone's outage. This is the noisy-neighbor problem in firewall form, and it is exactly what multi-tenancy is supposed to prevent. Administrative separation without resource separation is only half of isolation. Resource classes close that gap. By assigning each context to a class that caps how much of each resource it may consume, you ensure that tenant B can exhaust only its own allocation, not the whole appliance. The other contexts keep working because their share was never available for tenant B to take. This is the same principle as CPU and memory limits on virtual machines or containers, applied to firewall data-plane resources. ## How Resource Classes Work A resource class is a named policy that sets limits on the resources a member context can use. You create the class in the system context, set limits inside it, and then assign one or more contexts to membership of that class. A context that is not explicitly placed in a class belongs to the default class, which by convention allows unlimited use of most resources (so if you do nothing, you have the very noisy-neighbor exposure described above). Limits come in two flavors, and understanding the difference is the whole art of the feature: - **Limits (maximums).** A hard ceiling on how much of a resource a context may consume. This is the noisy-neighbor guard: the context simply cannot exceed it, protecting the pool for everyone else. - **Guarantees (minimums, expressed as a percentage of the total).** A reserved floor that other contexts cannot eat into. This is the starvation guard from the other direction: even if every other context is busy, this context is promised at least its reserved share. You can express many limits either as an absolute number or as a percentage of the platform's total for that resource. Common resources you cap include concurrent connections, connection setup rate per second, address translations (xlates), host entries, ASDM and Telnet/SSH management sessions, MAC addresses (in transparent mode), and inspection rates. The two you care about most in practice are usually concurrent connections and the connection rate, because that is what a flood consumes fastest. ## Configuring a Resource Class on a Supported Platform On a physical ASA in multiple context mode, the configuration is compact. In the system context you define the class, set the limits and guarantees, and then join contexts to it: ``` ! system context class GOLD limit-resource conns 50000 limit-resource rate conns 5000 limit-resource xlates 40000 limit-resource hosts 30000 ! context CUSTOMER-A member GOLD config-url disk0:/ctx-a.cfg ! context CUSTOMER-B member GOLD config-url disk0:/ctx-b.cfg ``` Here the GOLD class caps each member context at 50,000 concurrent connections, 5,000 new connections per second, 40,000 translations, and 30,000 host entries. CUSTOMER-A and CUSTOMER-B are both members, so each is independently held to those ceilings. If CUSTOMER-B is flooded, it hits its own 5,000-per-second setup ceiling and stops, and CUSTOMER-A's capacity is untouched. You verify the live picture with `show resource usage`, which reports current and peak usage against the limits per context, and flags any resource that has hit its cap (a denied count that is climbing is your signal that a context is bumping its ceiling). The design workflow is to build a few tiers (for example GOLD, SILVER, BRONZE) with different ceilings and guarantees, then place tenants into the tier they pay for or the tier their risk profile warrants. A high-value tenant gets a guaranteed minimum so it is never starved; a low-trust tenant gets a tight maximum so it can never become the noisy neighbor. ## The Honest asav Note Resource classes only exist inside multiple context mode. They are configured in the system context and they operate on contexts, so with no contexts there is nothing to limit. That is the exact wall we hit in the lab. Our firewalls are virtual asav 9.24 instances, and the virtual appliance cannot enter multiple context mode, which we confirmed on the real device: FW1# show mode Security context mode: single FW1(config)# mode multiple ^ ERROR: % Invalid input detected at '^' marker. Because `mode multiple` is rejected by the parser, there are no contexts to place into a class, and the `class` and `limit-resource` commands have nothing to act on. Resource classes are therefore a physical-ASA, multiple-context feature that the asav cannot demonstrate. There is no workaround to expose them on the virtual appliance; the honest position is that this is hardware territory, and the single line above is the proof. If you want to study resource classes hands-on for CCIE Security, do it on a physical ASA or a hardware-accurate environment, not asav. ## What Protects a Virtual ASA Instead On a single-context asav you do not have per-context resource classes, but you are not defenseless against resource exhaustion. A single-context firewall has one owner of its resources, so the noisy-neighbor problem between tenants does not exist in the first place (there is only one tenant). Where you still want protection is against a single flooding source or runaway application eating the box, and for that the tools are connection limits and embryonic (half-open) connection limits applied through your policy, plus per-VM resource sizing at the hypervisor. If your goal was multi-tenant fairness, the virtual answer is the same one we reach for elsewhere in this series: run a separate asav per tenant so each tenant gets its own appliance and its own full set of resources, which is a harder isolation boundary than contexts anyway. That approach pairs naturally with the [shared-interface context](https://www.pinglabz.com/asa-shared-interface-contexts/) discussion, where we reach the same conclusion: on virtual, separate instances beat one appliance sliced into contexts. ## Limits vs Guarantees: A Worked Example The interplay between maximums and minimums is where resource classes earn their keep, so it helps to see it play out. Suppose an ASA platform supports 100,000 concurrent connections in total, and you have four contexts. You give one context a guarantee of 40 percent, which reserves 40,000 connections that no other context can ever consume, even when that context is idle. The remaining 60,000 connections form a shared pool the other three contexts compete for, unless you also cap them. If you then place those three in a class with a limit of 25,000 connections each, none of them can individually exceed 25,000, so even all three flooding at once cannot touch the 40,000 you reserved for the important tenant. Notice that guarantees and limits solve opposite failure modes. A guarantee protects an important context from being starved by others; a limit protects everyone else from a greedy context. In a well-designed multi-tenant firewall you usually use both: guarantees for the tenants you must never let fail, and limits for the tenants you do not fully trust. Be careful not to over-commit guarantees, though. If you hand out guaranteed minimums that sum to more than 100 percent of a resource, the ASA cannot honor them all simultaneously, and the configuration is either rejected or leaves you with promises the hardware cannot keep. Keep the sum of guarantees comfortably under the total so there is headroom for the contexts that have no guarantee at all. ## Common Resource-Class Mistakes Three mistakes account for most resource-class trouble on hardware. The first is leaving everything in the default class, which allows effectively unlimited use, and then being surprised when one context takes down the others; the default class is a trap precisely because it looks like it is doing something when it is not protecting anyone. The second is capping concurrent connections but forgetting the connection *rate*: a flood is often more about setup rate than steady-state count, so a context can hammer the CPU building and tearing down connections while never hitting a concurrent-connection ceiling. Cap both. The third is setting limits and then never watching `show resource usage`, so you never learn that a context is routinely bumping its ceiling and dropping legitimate traffic. The denied counter in that output is the feedback loop: a steadily climbing denied count means either an attack or a limit set too low for the tenant's real workload, and you cannot tell which without looking. Resource classes are not a set-and-forget control; they are a policy you tune as you learn each tenant's genuine demand. ## Key Takeaways - In multiple context mode all contexts draw from shared resource pools, so one context can exhaust connections, xlates, or host entries and starve the others. - Resource classes cap and guarantee per-context resource usage: limits set hard maximums (noisy-neighbor guard), guarantees reserve minimum shares (starvation guard). - You define a `class` with `limit-resource` entries in the system context, then join contexts with `member`; verify live usage with `show resource usage`. - The resources most worth capping are concurrent connections and connection setup rate, because that is what a flood consumes fastest. - Resource classes require multiple context mode, which asav cannot enter: `mode multiple` returns `% Invalid input detected at '^' marker.`, so the feature is physical-ASA only. - On a virtual ASA, use per-policy connection limits and per-VM sizing, or run a separate asav per tenant for true isolation. See the [shared-interface contexts](https://www.pinglabz.com/asa-shared-interface-contexts/) guide and the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). ### ASA Security Contexts with Shared Interfaces: Classification Rules URL: https://www.pinglabz.com/asa-shared-interface-contexts/ Last updated: 2026-07-13T11:10:50.000Z Security contexts turn one physical Cisco ASA into many independent virtual firewalls, each with its own interfaces, policies, and administrators. That is straightforward when every context has its own dedicated interfaces. It gets genuinely interesting when contexts have to **share** a physical interface, because now the firewall has to look at an incoming packet and decide, before it applies any policy, which context that packet belongs to. That decision is made by the classifier, and understanding its rules is the difference between a multi-tenant firewall that works and one that silently sends traffic to the wrong tenant. This article is part of the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series, and it is grounded in a lab where the virtual appliance refused to enter context mode at all. That refusal is real and we will show it. ## Why Shared Interfaces Exist In a clean design, each context owns its own interfaces and there is no ambiguity: a packet arriving on a given physical port belongs to exactly one context. But physical ports are finite, and in a service-provider or large-enterprise deployment you may have more tenants than you have spare interfaces. Sharing an outside interface across several contexts lets many tenants reach the internet through one uplink, or lets many contexts sit behind one common segment, without dedicating a physical port to each. The price of that efficiency is ambiguity: when a packet lands on the shared interface, the ASA can no longer identify the owning context by interface alone. Something has to disambiguate, and that something is the classifier. ## The Classifier: How a Packet Finds Its Context The ASA classifier runs a strict, ordered set of checks on every packet that enters a shared interface. The order matters, because the first rule that produces a unique context wins: - **Unique interface first.** If the ingress interface is assigned to only one context, classification is trivial and instant: that context owns the packet. This is the common case and it does not involve the rest of the logic. The interesting rules only apply when an interface is shared by two or more contexts. - **Then the destination address, via the context's configuration.** On a shared interface, the ASA uses the packet's destination to find the matching context. Concretely, it looks at how each context that shares the interface is configured to receive traffic for that destination, using either a unique destination MAC address per context or a unique NAT/global address that maps to a single context. Two mechanisms make the destination-based step unambiguous, and you generally pick one: - **Unique MAC addresses per context.** If you enable `mac-address auto` in the system configuration, the ASA assigns a distinct virtual MAC to each context's instance of the shared interface. Now the destination MAC in the Layer 2 header identifies the context directly. This is the clean, recommended approach, because it removes all ambiguity at Layer 2 and does not depend on the NAT configuration. - **Unique addresses through NAT.** Without unique MACs, the classifier falls back to the destination IP as translated by each context's NAT rules. Each context must present a unique global/mapped address on the shared interface so the destination address maps to exactly one context. This works, but it couples classification to your NAT design, and overlapping or missing translations are a frequent cause of misclassified traffic. If you do not give the classifier a way to produce a unique answer (no unique MAC, no unique destination address), it cannot decide, and the packet is dropped rather than guessed. That fail-closed behavior is correct from a security standpoint but frustrating to debug, because the symptom is simply traffic disappearing on a shared interface with no obvious policy denying it. ## Enabling Unique MACs and Sharing an Interface On a supported hardware platform, the pieces fit together like this. In the system context you turn on automatic MAC generation and allocate the shared interface to multiple contexts: ``` ! system context mac-address auto ! context CUSTOMER-A allocate-interface GigabitEthernet0/0 outside_a config-url disk0:/ctx-a.cfg ! context CUSTOMER-B allocate-interface GigabitEthernet0/0 outside_b config-url disk0:/ctx-b.cfg ``` Both contexts now share the same physical GigabitEthernet0/0, each seeing it under its own mapped name, and each with a distinct auto-generated MAC. When a frame arrives on that port, the destination MAC tells the classifier whether it belongs to CUSTOMER-A or CUSTOMER-B before any access rule or NAT policy runs. Inside each context you then configure the interface as if it were an ordinary firewall interface, with its own IP, security level, ACLs, and NAT. ## The Real asav Limitation Everything above assumes you can create contexts in the first place, and that is exactly what our lab could not do. The firewalls are virtual asav 9.24 instances, and the virtual appliance does not support multiple context mode at all. Cisco's documentation is explicit on the point: the ASAv does not support multiple contexts. We confirmed it on the real device. Checking the mode and trying to switch to multiple, this is precisely what came back: FW1# show mode Security context mode: single FW1(config)# mode multiple ^ ERROR: % Invalid input detected at '^' marker. The mode is fixed at single and there is no path to multiple. Because you cannot create contexts, you cannot share an interface between contexts, and the entire classifier discussion is moot on this platform. This is not a bug or a missing license you can add; multiple context mode is simply not part of the asav feature set. The parser reject is the evidence, and it is the honest answer for anyone planning a multi-tenant virtual firewall around asav contexts: they are not there. ## What To Use Instead on Virtual The good news is that virtualization gives you a cleaner way to achieve tenant separation than shared-interface contexts ever did, precisely because the hypervisor already isolates workloads. Instead of one appliance sliced into contexts, you run one appliance per tenant: Separate asav per tenant Isolation: Full VM boundary Each tenant gets its own asav instance and its own interfaces. No shared-interface classifier needed. FTD multi-instance Isolation: Container per instance On Firepower, carve one appliance into isolated logical devices. The modern multi-tenant answer. asav contexts Isolation: Not available Parser rejects `mode multiple`. Use one of the options above instead. Running a separate asav per tenant trades the density of contexts for the simplicity and hard isolation of a full VM boundary, which is often exactly what a security team wants anyway: a compromise or misconfiguration in one tenant's firewall cannot touch another's, because they are different virtual machines. On the Firepower side, FTD multi-instance carves a single appliance into isolated logical devices, and that is where Cisco has put its modern multi-tenant engineering. If your driver for contexts was multi-tenancy, either of these gives you the separation without the shared-interface classifier complexity. If you genuinely need classic ASA security contexts with shared interfaces (for example to match an existing hardware design or a CCIE lab), that is physical-ASA territory, and you should study it there. Start with the fundamentals in the [ASA multiple context mode](https://www.pinglabz.com/cisco-asa-multiple-context-mode/) guide. ## Operational Notes and Gotchas A few things bite people when they first run shared interfaces on hardware. First, if you forget `mac-address auto` and your NAT does not provide unique destination addresses, the classifier has nothing to work with and traffic is dropped as unclassifiable, which looks like a policy problem but is not. Second, shared interfaces mean the contexts sharing them are on the same Layer 2 segment, so broadcast and multicast behavior and any Layer 2 loops are shared concerns, not per-tenant ones. Third, management traffic to the contexts still has to be classifiable, so a shared management interface needs the same unique-MAC or unique-address treatment as any other shared interface. Fourth, changing whether an interface is shared, or toggling `mac-address auto`, affects live traffic classification, so treat it as a maintenance-window change rather than a casual edit. None of these apply on asav, of course, because you never get as far as creating the contexts. ## A Worked Classification Example Walking one packet through the logic makes the rules concrete. Picture two tenants, CUSTOMER-A and CUSTOMER-B, both sharing the outside interface GigabitEthernet0/0, with `mac-address auto` enabled so each context's instance of that interface has a distinct virtual MAC. A client on the internet sends a packet to a public address that CUSTOMER-B advertises. The upstream router ARPs for that address, and because CUSTOMER-B owns that mapped address, it answers with CUSTOMER-B's virtual MAC. The client's packet therefore arrives on GigabitEthernet0/0 with CUSTOMER-B's MAC as the destination in the Layer 2 header. The classifier reads that destination MAC, matches it to CUSTOMER-B, and hands the packet to CUSTOMER-B's context, which then applies CUSTOMER-B's access rules and NAT. CUSTOMER-A never sees the packet. Now imagine the same topology without unique MACs, relying on NAT instead. The classifier has to look at the destination IP and find the single context whose NAT configuration claims that address. If both contexts happen to have overlapping or unconfigured mappings for that address, the classifier cannot pick a winner and drops the packet. This is why the MAC-based approach is preferred: the disambiguation happens at Layer 2 from the frame itself, and it does not depend on getting every NAT statement across every tenant perfectly non-overlapping. The NAT-based method works, but it makes your classification only as reliable as your translation hygiene. ## Verifying Classification on Hardware When you are debugging a shared-interface deployment on a real ASA, a handful of commands tell you what the classifier is doing. From the system context, `show interface` and the context allocation output confirm which contexts share which interfaces and what MAC each was assigned. Inside a context, `show nat` and `show xlate` confirm the mapped addresses that the classifier depends on when you are using the NAT method. The packet-tracer tool is especially valuable here: you can inject a synthetic packet with a specific destination and watch which context and which phase handles it, which turns an invisible drop into a visible decision. If packet-tracer shows a packet being dropped before any ACL phase on a shared interface, classification is your culprit, and you go back to check unique MACs or unique mapped addresses. Building this verification habit early saves hours, because misclassification never announces itself with a helpful log message; it just looks like traffic that should work and does not. ## When Contexts Are Worth It At All It is worth stepping back and asking whether contexts are even the right tool, shared interfaces or not. Contexts buy you administrative separation (each tenant gets its own config, its own admins, its own policy) on a single piece of hardware, which is attractive when rack space, power, and licensing all favor consolidation. But they come with real constraints: some features are not supported in multiple context mode at all, troubleshooting spans the system context plus each affected context, and a resource-hungry tenant can starve its neighbors unless you cap it (which is the job of resource classes, a closely related multi-context feature). For a lot of designs, especially virtual ones, the cleaner answer is separate firewalls rather than one firewall pretending to be several. That is doubly true on asav, where the choice is made for you: no context mode exists, so per-tenant isolation means per-tenant instances. The shared-interface classifier is a genuinely elegant piece of engineering, but it is engineering you only need when you have committed to packing many tenants onto one physical chassis, which is a hardware decision, not a virtual one. ## Key Takeaways - Security contexts turn one ASA into many virtual firewalls; sharing a physical interface across contexts forces the ASA to classify each packet to an owning context before applying policy. - The classifier checks the ingress interface first (unique interface wins instantly), then, on a shared interface, uses the destination via a unique per-context MAC address or a unique NAT/global address. - `mac-address auto` gives each context a distinct MAC on the shared interface and is the clean way to remove classification ambiguity; the NAT-based fallback couples classification to your translation design. - If the classifier cannot produce a unique context, the packet is dropped (fail closed), which shows up as traffic silently disappearing on a shared interface. - On asav 9.24 none of this is reachable: `show mode` is single and `mode multiple` returns `% Invalid input detected at '^' marker.` The ASAv does not support multiple contexts. - On virtual, use a separate asav per tenant or FTD multi-instance for isolation; for classic shared-interface contexts you need physical ASA hardware. See the [multiple context mode](https://www.pinglabz.com/cisco-asa-multiple-context-mode/) guide and the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). ### ASA Active/Active Failover: Failover Groups and Asymmetric Reality URL: https://www.pinglabz.com/asa-active-active-failover/ Last updated: 2026-07-13T11:10:50.000Z Active/Active failover is the design people reach for when Active/Standby feels wasteful: instead of one firewall doing all the work while its partner idles, both units forward traffic at the same time. It sounds like the obviously better option, but it carries a hard prerequisite (multiple context mode) and a subtle trap (asymmetric routing) that make it the wrong default for most deployments. This article is part of the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series, and it is written from a lab where we tried to enter the mode Active/Active needs on a real virtual ASA (asav 9.24) and the parser said no. We will show you that refusal, and then show you the capture that proves what you should be doing instead: a zero-loss Active/Standby failover on the very same pair. ## How Active/Active Actually Works Active/Active failover does not mean one giant firewall spread across two boxes. It means you split the ASA into **failover groups**, and each group is active on a different physical unit. A failover group is a set of security contexts (virtual firewalls). Group 1 is active on the primary unit, group 2 is active on the secondary. Traffic destined for the contexts in group 1 flows through the primary; traffic for the contexts in group 2 flows through the secondary. Both boxes forward simultaneously, so you are using hardware you would otherwise leave cold. Each failover group still has a standby copy on the other unit. So the primary is active for group 1 and standby for group 2; the secondary is active for group 2 and standby for group 1\. If the primary dies, its group 1 fails over to the secondary, which is then active for both groups (and now doing all the work, exactly as an Active/Standby pair would). When the primary recovers, group 1 can preempt back to it if you configure preemption. This is load-sharing, not load-balancing: you are statically assigning halves of your firewall to different boxes, not spreading a single flow across both. You control which unit owns each group with `primary` and `secondary` preferences per group, and you can weight interface monitoring and polling per group as well. The mental model that helps: Active/Active is two Active/Standby pairs sharing the same two chassis, offset so each box is the active member of one pair and the standby of the other. ## Why It Requires Multiple Context Mode Here is the prerequisite that decides everything. Failover groups are built out of security contexts, and contexts only exist in **multiple context mode**. A single-context ASA has exactly one firewall and nothing to split into two groups, so there is no Active/Active to configure. Cisco's own documentation states it plainly: Active/Active failover is only supported in multiple context mode. That is not a soft recommendation; it is a structural dependency. No contexts, no groups; no groups, no Active/Active. This is why any serious look at Active/Active is really a look at multiple context mode first. You have to convert the firewall to multiple mode, define your contexts, allocate interfaces to them, assign the contexts to failover group 1 or group 2, and then let the two groups land on different units. If you have never worked with contexts, start with the [ASA multiple context mode](https://www.pinglabz.com/cisco-asa-multiple-context-mode/) fundamentals, because everything in Active/Active is built on top of them. ## The Asymmetric Routing Problem (and asr-group) Now the subtle trap. Because both units forward at once, and because the routers around them may pick different paths for the two directions of a conversation, you can end up with a flow whose outbound packets leave through the unit that owns group 1 and whose return packets arrive at the unit that owns group 2\. The unit that receives the return traffic has no connection-table entry for it (the connection was built on the other unit), so it drops the return packets as out-of-state. The session breaks, and it breaks intermittently in a way that is genuinely painful to diagnose, because a simple ping might work while a stateful TCP session does not. The ASA fixes this with **asymmetric routing groups**, configured with the `asr-group` command on the interfaces that might receive asymmetric return traffic. When a unit receives a packet for a connection it does not own, it checks the interfaces in the same asr-group on its failover peer. If the peer owns the connection, the receiving unit forwards the packet to the peer over the stateful failover link, the peer processes it against its real connection entry, and the session survives. In effect, asr-group lets the two units cover for each other's connections when the surrounding routing is not symmetric. ``` interface GigabitEthernet0/0 nameif outside security-level 0 ip address 203.0.113.10 255.255.255.0 standby 203.0.113.11 asr-group 1 ``` You put the same `asr-group` number on the corresponding interfaces in both failover groups. It works, but notice what it costs: return traffic can now take an extra hop across the state link before it is processed, which adds latency and load, and it only papers over a routing design that is asymmetric in the first place. Many engineers who reach Active/Active for the load-sharing discover that the asymmetric-routing handling eats much of the benefit and adds operational complexity they did not want. ## The Real asav Refusal This is where the lab keeps everyone honest. To do Active/Active you must first be in multiple context mode. Our firewalls are virtual asav 9.24 instances, and the virtual appliance does not support multiple contexts at all. Checking the current mode and trying to change it, this is exactly what the real device returned: FW1# show mode Security context mode: single FW1(config)# mode multiple ^ ERROR: % Invalid input detected at '^' marker. The mode is fixed at single, and `mode multiple` is not a valid command on asav. Because you cannot enter multiple context mode, you cannot create failover groups, and because you cannot create failover groups, you cannot configure Active/Active. The whole feature is unreachable on the virtual appliance, and the parser tells you so at the very first step. This is consistent with Cisco's documentation that the asav does not support multiple contexts and that Active/Active requires multiple context mode. There is no workaround on asav; the honest answer is that this feature belongs to physical ASA hardware. ## What You CAN Do Virtually: Zero-Loss Active/Standby Here is the counterpoint that matters. On the exact same pair of asav instances that refused Active/Active, [Active/Standby failover](https://www.pinglabz.com/cisco-asa-active-standby-failover/) works completely, and it works well. We proved it. We ran a 45-packet ping stream from an inside host through the firewall to an outside address at 0.3-second intervals, and mid-stream we forced the active unit to give up the active role. The result was not a blip, not a few dropped packets during convergence. It was nothing: HOST1$ ping -c 45 -i 0.3 203.0.113.2 64 bytes from 203.0.113.2: icmp\_seq=44 ttl=255 time=6.52 ms 64 bytes from 203.0.113.2: icmp\_seq=45 ttl=255 time=6.56 ms \--- 203.0.113.2 ping statistics --- 45 packets transmitted, 45 received, 0% packet loss, time 13242ms Forty-five sent, forty-five received, zero percent loss, straight through a full active-unit failover. Stateful failover had already replicated the connection table to the standby, so when the standby took over the active IP and MAC, the live ICMP stream never lost a single packet. The connection did not have to be rebuilt because the standby already knew about it. That is the redundancy story you can lab on virtual today, and it is genuinely excellent. ## Active/Active vs Active/Standby: Which To Choose Active/Standby Requires contexts: No Asymmetric-routing risk: None Runs on asav: Yes (proven) One active unit, one hot standby. Simple, stateful, zero-loss. Active/Active Requires contexts: Yes Asymmetric-routing risk: Real (needs asr-group) Runs on asav: No (parser reject) Two failover groups, both units forward. More utilization, more complexity. For most designs, Active/Standby is the right answer. It is simpler, it has no asymmetric-routing exposure, it does not force you into multiple context mode, and its failover is stateful and effectively lossless. Active/Active earns its complexity only when you genuinely need both boxes forwarding at once (typically a multi-tenant deployment already running contexts) and you have the routing discipline to keep paths symmetric or the asr-group configuration to handle them when they are not. If you are not already in multiple context mode for another reason, converting to it just to get Active/Active is usually a poor trade. And if your ASA is virtual, the decision is made for you: Active/Active is not available, and Active/Standby is, so use it. The proof is above. On virtual, put your effort into a clean, well-monitored Active/Standby stateful failover pair rather than chasing a mode the appliance cannot enter. ## Building Active/Active on a Physical ASA, Step by Step On supported hardware the build is a sequence, and the order is what trips people up. You cannot sprinkle these commands in randomly; each step depends on the one before it. Here is the shape of it: - **Convert to multiple context mode.** `mode multiple` reboots the firewall into the multi-context system space. This is the step that fails outright on asav. - **Create the contexts.** In the system context, define each context, allocate its interfaces, and point it at its config file (typically in flash). These contexts are the tenants you will split across units. - **Configure the failover link and state link.** These live in the system context and are shared by both groups. As with any ASA failover, the failover link carries hellos and config replication; the state link carries the connection table for stateful failover. - **Define failover group 1 and group 2.** Set each group's unit preference (`primary` or `secondary`), polling and holdtimes, interface-monitoring policy, and optionally `preempt` so a recovered unit reclaims its group. - **Assign each context to a group.** `join-failover-group 1` or `2` inside each context. This is the line that decides which physical unit forwards that tenant's traffic. - **Enable failover** on both units and let them synchronize. A minimal skeleton for the two groups in the system context looks like this: ``` failover group 1 primary preempt failover group 2 secondary preempt ! context CUSTOMER-A join-failover-group 1 context CUSTOMER-B join-failover-group 2 ``` Now group 1 (CUSTOMER-A) is active on the primary unit and group 2 (CUSTOMER-B) is active on the secondary, with each group's standby copy sitting on the opposite box. That is the entire load-sharing mechanism: two groups, offset across two chassis. ## Stateful Replication Still Applies, Per Group Active/Active does not change the fundamentals of stateful failover; it just runs them twice. Each failover group replicates its own connection table, translation table, and other stateful data to its standby copy on the peer unit over the shared state link. When group 1 fails from the primary to the secondary, the secondary already holds group 1's connection state and can carry the existing sessions without rebuilding them, exactly the way an Active/Standby pair does. The difference is only that the same state link is now carrying replication for both directions at once, which is one more reason the state link should be a high-bandwidth interface (or an EtherChannel, pre-configured on both units before failover comes up). Interface monitoring also runs per group. Each failover group independently decides whether it should fail over based on the health of the interfaces belonging to its contexts. This means one group can fail over to the peer while the other group stays put, which is the fine-grained behavior that makes Active/Active attractive for multi-tenant designs: a fault that only affects one tenant's interfaces moves only that tenant's group, not the whole firewall. ## Common Active/Active Mistakes The failure modes cluster around the prerequisites and the routing. The first is trying to configure it in single mode and getting nowhere, because the failover-group commands do not exist until you are in multiple context mode. The second is ignoring asymmetric routing until sessions start breaking intermittently in production, then scrambling to add `asr-group` after the fact. The third is oversubscribing a single unit: remember that in a failure, one box carries both groups, so if each unit is already running near capacity in steady state, a failover will overload the survivor. Size each unit to carry the full load of both groups, not half. If you cannot, you have built a design that is faster in the good case and broken in the bad case, which defeats the purpose of high availability. ## Key Takeaways - Active/Active splits the firewall into two failover groups, each active on a different unit, so both boxes forward traffic (load-sharing, not load-balancing). - It structurally requires multiple context mode, because failover groups are built from security contexts. No contexts means no Active/Active. - Because both units forward, asymmetric routing can deliver return traffic to the wrong unit; `asr-group` lets units forward such traffic to the connection owner over the state link. - On asav 9.24 the feature is unreachable: `show mode` is fixed at single and `mode multiple` returns `% Invalid input detected at '^' marker.` - On the same asav pair, [Active/Standby failover](https://www.pinglabz.com/cisco-asa-active-standby-failover/) is fully supported and proven lossless: 45 of 45 pings survived a live failover at 0% loss. - Choose Active/Standby for most designs and for all virtual ASAs; reserve Active/Active for physical, multi-context deployments that truly need both units active. See the [multiple context mode](https://www.pinglabz.com/cisco-asa-multiple-context-mode/) guide and the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). ### EtherChannel on the Cisco ASA: Port-Channels, LACP, and Failover Interaction URL: https://www.pinglabz.com/asa-etherchannel-port-channel/ Last updated: 2026-07-13T11:10:49.000Z EtherChannel is the feature that lets a Cisco ASA treat several physical links as one fat pipe: more bandwidth when every member is up, and automatic survival when a member goes down. It sits right next to redundant interfaces in the ASA high-availability toolkit, but it solves a different problem, because an EtherChannel runs all its members active at once instead of keeping one in reserve. This article is part of the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series, and like the rest of the hardware-HA family it comes with a real-lab reality check: we tried to build a Port-channel on a virtual ASA (asav 9.24) and the parser rejected every command. We will show you exactly that, because knowing where the feature does and does not exist is the whole point. ## What EtherChannel Gives You An EtherChannel (also called a Port-channel, or a Link Aggregation Group) bundles between two and eight physical interfaces into a single logical interface. The ASA hashes each flow onto one of the member links, so across many flows you get close to the summed bandwidth of the members, and if any single member drops, its flows rehash onto the survivors with no change to the logical interface. You get two wins in one feature: aggregate throughput and link redundancy. The logical Port-channel carries the IP address, the nameif, and the security level, exactly like a redundant interface, so upper layers see one stable interface regardless of how many physical members are healthy underneath. The important contrast with a [redundant interface](https://www.pinglabz.com/asa-redundant-interfaces/) is that EtherChannel members are all active. A redundant interface keeps one member in hot standby and only uses one link at a time; it is pure failover, no aggregation. EtherChannel uses every member simultaneously and load-balances flows across them. If you need more bandwidth than one link provides, EtherChannel is the answer; if you only need one link's worth of throughput with a backup, a redundant interface is simpler. ## LACP: How the Bundle Negotiates The ASA builds EtherChannels using LACP (Link Aggregation Control Protocol, the IEEE 802.3ad standard), which lets the firewall and the neighboring switch agree that a set of ports really should be bundled before any traffic flows. Each end sends LACP packets and confirms the partner sees the same bundle. This negotiation is what protects you from a miscabled or half-configured channel silently black-holing traffic. You set a channel mode on each member: - **active** \- the member actively sends LACP to form the bundle. This is the normal ASA choice. - **passive** \- the member only responds to LACP; it will not initiate. Two passive ends never form a channel. - **on** \- a static bundle with no LACP negotiation at all. Faster to come up, but with no safety check, so a cabling mistake is not caught. In practice you configure LACP `active` on the ASA and `active` (or at least `passive`) on the switch, and let the two negotiate. A minimum and maximum number of active members can be set, and any members beyond the maximum sit in a standby state ready to join if an active member fails. ## How You Configure It on a Physical ASA On a supported hardware platform, the config is a two-part job: create the logical Port-channel and set its Layer 3 properties, then add physical members with a channel-group command that references the same channel number and sets the LACP mode: ``` interface Port-channel1 nameif outside security-level 0 ip address 203.0.113.10 255.255.255.0 ! interface GigabitEthernet0/0 channel-group 1 mode active ! interface GigabitEthernet0/1 channel-group 1 mode active ``` Notice the addressing lives on the Port-channel, never on the members, just like a redundant interface. The members only carry the `channel-group` statement. You verify the bundle with `show port-channel summary` and `show lacp neighbor`, which tell you which members are bundled, which are standby, and whether LACP has agreed with the partner. A member showing as suspended usually means the switch side disagrees on the channel, which is exactly the miscabling that LACP is there to catch. ## The Failover and State Link Rule You Must Know Here is a rule that trips people up on the CCIE lab and in production. If you run an EtherChannel and you want to carry the failover link or the stateful failover state link across it, that EtherChannel must be pre-configured on both units before you bring failover up. The failover link is special: the ASA cannot rely on config replication to build the very interface that replication depends on. You cannot bootstrap the failover link over a channel that only exists on one unit, because the standby has no way to learn the channel definition until it is already talking to the primary. So the operational order matters: build the EtherChannel that will carry failover on both the primary and the secondary manually, confirm it is up and LACP-negotiated on both, and only then enable failover. After that, the rest of the configuration (data interfaces, ACLs, NAT, inspection) replicates automatically. Getting this backwards is a classic cause of a failover pair that will not form, because the two units simply cannot reach each other over a channel that is half-built. The same care applies to any EtherChannel that carries a monitored data interface. Failover interface monitoring watches the logical Port-channel, not the individual members, so a single member dropping does not trigger a unit failover as long as the channel stays up on its survivors. Only when enough members fail that the Port-channel itself goes down does it count as an interface failure. That is the desirable behavior: local member loss is absorbed, and unit failover is reserved for real interface-level failures. ## The Real asav Rejection Now the honest part. Our lab firewalls are virtual asav 9.24 instances, and EtherChannel is a hardware-oriented feature. It is not gated behind a license you can enable, and it is not hidden. The keywords simply do not exist in the asav parser. When we tried to build the Port-channel, add a member, and check the summary, every single command bounced: FW1(config)# interface Port-channel1 ^ ERROR: % Invalid input detected at '^' marker. FW1(config-if)# channel-group 1 mode active ^ ERROR: % Invalid input detected at '^' marker. FW1# show port-channel summary ^ ERROR: % Invalid input detected at '^' marker. Three commands, three carets, three invalid-input errors. There is no `interface Port-channel`, no `channel-group`, and no `show port-channel summary` on the virtual appliance. This is the correct behavior: link aggregation is a physical-port feature, and asav has no physical ports to aggregate. If you were planning to lab EtherChannel on a virtual ASA for CCIE Security prep, this is your heads-up to do it on real hardware or a hardware emulation, not asav. ## Which Platforms Support It Physical ASA 5500-X EtherChannel: Supported 2 to 8 members per Port-channel with LACP negotiation. Firepower (ASA image) EtherChannel: Supported Also the basis of spanned-EtherChannel clustering on Firepower. asav (virtual ASA) EtherChannel: Not available Parser rejects `interface Port-channel` and `channel-group`. On virtual, if you need more throughput than a single interface provides, you scale by adding interfaces to the guest and load-balancing at the network or hypervisor layer, or you move to FTD, where Firepower clustering leans heavily on spanned EtherChannel. For unit-level resilience on asav itself, the answer remains [Active/Standby failover](https://www.pinglabz.com/cisco-asa-active-standby-failover/), which the virtual appliance fully supports and which we proved with a zero-loss live failover in the same lab. ## Load Balancing: How Flows Pick a Member An EtherChannel does not split a single flow across multiple links. Instead it runs a hash over selected packet fields (source and destination IP, and optionally Layer 4 ports) and pins each flow to one member for its lifetime. This keeps packets in order within a flow (no reordering, which TCP would hate) while spreading many flows across the bundle. The practical consequence is that a channel only balances well when it carries many conversations. Two heavy flows between the same pair of hosts may both land on the same member and leave the others idle, so aggregate throughput is a statistical benefit across a busy firewall, not a guarantee for any single transfer. You choose the hashing input to match your traffic pattern. If most traffic is between a small number of IP pairs but many ports (for example a firewall in front of a few busy servers), including Layer 4 ports in the hash spreads load far better than IP-only hashing. If you get the hash wrong, you can watch one member run hot while the others sit near zero, which looks like a bandwidth problem but is really a distribution problem. ## Troubleshooting a Channel That Will Not Bundle When an EtherChannel refuses to come up on hardware, the usual suspects are all mismatches between the two ends. LACP mode mismatch (both sides passive, so neither initiates), speed or duplex mismatch on the members, a different number of allowed members configured on each side, or the two ends simply not agreeing on which physical ports belong to the channel. A member stuck in a suspended or standalone state is the tell: LACP saw the partner but the parameters did not line up, so it refused to bundle that link rather than risk a loop or a black hole. Reading `show lacp neighbor` on both ends and comparing the system IDs and port priorities is how you find the disagreement. This negotiation safety is exactly why LACP `active` beats static `on` mode for anything that matters: the protocol catches the mistake instead of forwarding traffic into it. ## Key Takeaways - EtherChannel bundles 2 to 8 physical links into one logical Port-channel, giving aggregate bandwidth and link redundancy with all members active at once. - Unlike a redundant interface (one active member, one standby), EtherChannel load-balances flows across every member and rehashes onto survivors when one fails. - The ASA negotiates channels with LACP (802.3ad); use `active` mode on the firewall and set addressing on the Port-channel, never on the members. - An EtherChannel that carries the failover or state link must be pre-configured on both units before you enable failover, because you cannot bootstrap the failover link over a half-built channel. - On asav 9.24 the feature is absent: `interface Port-channel1`, `channel-group 1 mode active`, and `show port-channel summary` all return `% Invalid input detected at '^' marker.` - Use physical ASA 5500-X or Firepower for EtherChannel; on virtual, scale at the hypervisor and rely on failover. More context is on the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). ### ASA Redundant Interfaces: The Simplest Link HA You Can Configure URL: https://www.pinglabz.com/asa-redundant-interfaces/ Last updated: 2026-07-13T11:10:49.000Z When people talk about high availability on the Cisco ASA, they usually jump straight to failover: two firewalls, one active, one standby. But there is a smaller, quieter form of redundancy that protects a single unit against a single dead link, and it reacts far faster than a full unit failover. That is the **redundant interface**. This article is part of the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) series, and it comes with an honest twist: we tried to configure it on a real virtual ASA (asav 9.24) and the parser flatly refused. That rejection is the most useful thing in this post, so we are going to show it to you. ## What a Redundant Interface Actually Is A redundant interface bundles two physical ports into one logical interface. One member is active and carries traffic; the other sits as a hot standby. If the active member's link goes down (cable pull, transceiver failure, upstream switch port death), the ASA moves traffic to the standby member. The logical interface keeps its name, its IP address, its security level, and its MAC address, so nothing above Layer 1 notices. Routing does not reconverge. NAT does not rebuild. The connection table is untouched. You simply keep forwarding on a different copper or fiber pair. The key number is speed. A redundant interface fails over in roughly a second, because the only thing being detected is a physical link state change on one member and a swap to the other. Compare that to a full unit-level failover, which has to detect the peer is gone, elect a new active unit, move shared IPs and MACs, and (with stateful failover) rely on a replicated connection table. Redundant interfaces are the surgical fix for "one link died" rather than "one firewall died." ## How You Configure It on a Physical ASA On a supported hardware platform, a redundant interface is a two-minute job. You create the logical interface, add the two physical members, and then configure the addressing and security level on the logical interface (never on the members): ``` interface Redundant1 member-interface GigabitEthernet0/0 member-interface GigabitEthernet0/1 nameif outside security-level 0 ip address 203.0.113.10 255.255.255.0 ``` The first member you add becomes the active member. The second is the standby. The physical members must be the same type (you cannot mix a copper and a fiber member into one redundant interface unless the media matches at the port level), and they carry no configuration of their own beyond being claimed as members. You can build up to eight redundant interfaces on platforms that support the feature. One nice operational detail: the redundant interface adopts the MAC address of the first member you add. If that first member is ever physically removed, the interface keeps the MAC until a reboot, so downstream ARP caches stay valid. That is deliberate, and it is part of why the failover is invisible to neighbors. ## How It Interacts With Unit Failover Redundant interfaces and unit failover are not competitors, they are layers. A production design frequently uses both: a redundant interface protects each data link inside a unit, and Active/Standby failover protects against the whole unit dying. When you monitor interfaces for failover health, the ASA treats the logical redundant interface as the monitored entity. A single member link going down does not trigger a unit failover, because the redundant interface is still up on its standby member. Only when both members of a monitored redundant interface are down does the interface count as failed for failover purposes. That layering is the whole point. You want the cheap, fast, local fix (swap members) to handle single link failures, and you want the expensive, coordinated fix (swap units) reserved for real unit-level problems. If you get the ordering wrong (for example, monitoring individual members instead of the logical interface), you can trigger unnecessary unit failovers on a single cable fault. ## The Real asav Rejection Here is where honesty matters. Our lab firewalls are virtual: a pair of asav 9.24 instances. The redundant interface is a hardware-oriented feature tied to physical ports, and the virtual appliance does not implement it. This is not a licensing gate you can lift or a hidden command. The keyword does not exist in the asav parser at all. When we tried to create the interface on the real device, this is exactly what came back: FW1(config)# interface Redundant1 ^ ERROR: % Invalid input detected at '^' marker. That caret under the "R" is the ASA parser telling you it does not recognize the token. No amount of enabling, licensing, or reloading changes it, because there is no physical port pair for the feature to bind to. This is the correct, expected behavior for a virtual appliance, and it is a genuinely useful thing to know before you plan a virtual ASA deployment around link-level redundancy that is not there. ## Which Platforms Do Support It Redundant interfaces live on the physical ASA and Firepower appliances running ASA software. If you are studying for CCIE Security or building a real hardware design, that is where you will configure and test this feature. Here is the quick platform reality: Physical ASA 5500-X Redundant interfaces: Supported Two physical members per logical interface, up to eight per unit. Firepower (ASA image) Redundant interfaces: Supported Same logical model on Firepower appliances running ASA software. asav (virtual ASA) Redundant interfaces: Not available Parser rejects `interface Redundant`. No physical ports to bind. ## What To Use on Virtual Instead If your ASA is virtual, you do not get redundant interfaces, but you are not without HA. The unit-level story on asav is genuinely strong: [Active/Standby failover](https://www.pinglabz.com/cisco-asa-active-standby-failover/) works fully on the virtual appliance, and we proved it in the same lab by pushing a live ping stream through the firewall and forcing a failover mid-stream with zero packet loss. That is the redundancy you can actually lab virtually. For link-level protection on virtual, you push the problem down to the hypervisor and the underlying NIC teaming or the virtual switch, rather than solving it inside the ASA. In other words, the physical redundancy moves out of the firewall and into the virtualization layer. This is the recurring theme across the whole ASA HA and scale family: the virtual appliance gives you excellent unit-level failover and almost nothing from the hardware-redundancy toolkit. Knowing exactly where that line sits saves you from designing a virtual deployment around a feature that will meet you with a caret and an invalid-input error. ## Monitoring and Verifying the Bundle On hardware, the command you live in is `show interface redundant1`, which reports the logical interface state plus which member is currently active and which is standby. A healthy bundle shows both members up with one carrying traffic; a degraded bundle shows one member down but the interface still up on the survivor. That degraded-but-up state is exactly the win: you have lost redundancy for that link but you have not lost the link. It is your cue to dispatch someone to replace the failed transceiver or cable before the second member also fails. You can also force a manual switch to test the standby member with `redundant-interface redundant1 active-member GigabitEthernet0/1`, which is worth doing during a maintenance window so you know the standby path is genuinely wired and working rather than assumed. A redundant interface whose standby member was never cabled correctly gives you a false sense of safety, and you only discover it at the worst possible moment. Test both members deliberately. Because the swap is a Layer 1 event handled inside the unit, there is nothing to tune in routing or NAT and no state to replicate to a peer. This is the opposite end of the spectrum from EtherChannel, where multiple members are active simultaneously and share load. If you want aggregate bandwidth as well as link redundancy, that is the [ASA EtherChannel](https://www.pinglabz.com/asa-etherchannel-port-channel/) story, and it is another hardware-only feature the asav parser rejects in the same way. Redundant interfaces are the simpler, active/standby-per-link answer when you only need failover, not aggregation. ## Key Takeaways - A redundant interface bundles two physical ports into one logical interface, one active member and one standby, with sub-second failover on a single link failure. - The logical interface keeps its IP, MAC, name, and security level through a member swap, so routing, NAT, and the connection table are untouched. - Configure addressing on the logical interface, never on the members; the first member added is active and donates the MAC address. - Redundant interfaces layer under unit failover: a single member drop is handled locally, and only a fully down redundant interface counts as an interface failure for failover. - On asav 9.24 the feature does not exist. `interface Redundant1` returns `% Invalid input detected at '^' marker.` because there are no physical ports to bind. - Use physical ASA 5500-X or Firepower for redundant interfaces; on virtual, rely on [Active/Standby failover](https://www.pinglabz.com/cisco-asa-active-standby-failover/) and hypervisor-level NIC redundancy. More context lives on the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). ### NAT and ACLs Together on ASA 8.3+: Real IPs, Not Mapped IPs URL: https://www.pinglabz.com/asa-nat-acl-real-ip/ Last updated: 2026-07-13T10:31:00.000Z Ask a room of network engineers which IP an ASA access list should reference when you publish a server, the private real address or the public mapped one, and you will get a split decision. That split is the single most misunderstood behaviour on the Cisco ASA, and it has cost more late nights than almost anything else on the platform. The answer since ASA 8.3 is unambiguous: **interface ACLs reference the real, post-translation IP, never the mapped one**. This article settles it with a before-and-after on one flow, allowed with a real-IP rule and dropped with a mapped-IP rule, everything else held constant. This article is part of the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) cluster and is the direct follow-on to [destination NAT for inbound services](https://www.pinglabz.com/asa-destination-nat/), where we published a DMZ web server and flagged that its ACL used the real IP. Here we prove exactly why, using detailed `packet-tracer` output from a live ASAv, and we contrast it with the old pre-8.3 habit that quietly breaks rules on a modern firewall. ## The Rule, and the Reason Behind It The rule is one sentence: on ASA 8.3 and later, an interface access list is evaluated *after* the ASA has un-NATed the packet, so the ACL must reference the real (post-translation) address of the host, not the mapped or public address a client actually sent to. The reason is the order of operations. When an inbound packet arrives, the ASA runs un-NAT as an early phase, before it consults the interface ACL. By the time the access rule is checked, the destination address in the packet has already been rewritten from the public IP to the private real IP. So the ACL is comparing against the real IP, and your rule has to name the real IP to match. This is not a quirk; it falls straight out of the fact that un-NAT is Phase 1 and the ACL check is Phase 2, the sequence proven in the [destination NAT walkthrough](https://www.pinglabz.com/asa-destination-nat/). Get the phase order and the ACL rule stops being something to memorise and becomes something you can derive. ## The Scenario We reuse the published DMZ server from the destination NAT lab. The ASAv maps a private DMZ nginx host at **172.20.30.80** to a public address **203.0.113.80**. An outside client at 192.168.99.100 connects to the public IP on port 80\. Two access lists are in play across the two tests, identical except for one address: - `OUTSIDE-IN` permits `tcp any host 172.20.30.80 eq www`, the **real** IP. - `OUTSIDE-BAD` permits `tcp any host 203.0.113.80 eq www`, the **mapped** IP. Same firewall, same NAT rule, same simulated packet. The only variable is which IP the applied ACL targets. Watch what that one difference does. ## Proof A: ACL Against the Real IP Permits With `OUTSIDE-IN` applied (real IP), we run a detailed packet-tracer for the inbound flow. The detailed view is the key here because it exposes the compiled forward-flow rule, including the flag that reveals what the ASA is really matching against: ``` FW1# packet-tracer input outside tcp 192.168.99.100 23456 203.0.113.80 80 detailed Phase: 2 Type: ACCESS-LIST Result: ALLOW Config: access-group OUTSIDE-IN in interface outside access-list OUTSIDE-IN extended permit tcp any host 172.20.30.80 eq www Additional Information: Forward Flow based lookup yields rule: in id=..., priority=13, domain=permit, deny=false hits=4, ..., use_real_addr, flags=0x0, protocol=6 src ip/id=0.0.0.0, mask=0.0.0.0, port=0, tag=any dst ip/id=172.20.30.80, mask=255.255.255.255, port=80, tag=any ``` Two things in that compiled rule prove the point. First, the flag `use_real_addr`. That is the ASA telling you, in its own internal representation, that this access rule is matched against the real address of the host, not the mapped one. Second, the destination in the compiled rule: `dst ip/id=172.20.30.80`, the real private IP, with a 32-bit mask. Even though the client sent its packet to 203.0.113.80, the compiled ACL entry the ASA evaluates carries 172.20.30.80\. Because our rule names that real IP, it matches, and the result is ALLOW. This is the mechanism made visible. The `use_real_addr` flag is not something you configure; it is how post-8.3 ASA compiles every interface ACL. The firewall always matches on the real address, so your rule always has to name it. ## Proof B: ACL Against the Mapped IP Drops Now swap in `OUTSIDE-BAD`, which permits the mapped IP 203.0.113.80, the "obvious" choice that looks right to anyone who learned ASA before 8.3\. Nothing else changes. We re-run the same packet: ``` FW1# packet-tracer input outside tcp 192.168.99.100 34567 203.0.113.80 80 Phase (ACCESS-LIST): Type: ACCESS-LIST Result: DROP Config: Implicit Rule Result: Action: drop Drop-reason: (acl-drop) Flow is denied by configured rule ``` Dropped. And look at which rule dropped it: the **Implicit Rule**, the invisible deny-all at the end of every ACL. Our `OUTSIDE-BAD` permit is looking for a destination of 203.0.113.80, but by the time the ACL runs, the ASA has already un-NATed the destination to 172.20.30.80\. The permit never matches, so the packet falls through to the implicit deny and is dropped with `acl-drop: Flow is denied by configured rule`. Restoring `OUTSIDE-IN` (the real-IP rule) returned the same flow to `Action: allow`. Same firewall, same packet, same NAT. Real IP in the ACL and the packet is allowed. Mapped IP in the ACL and it is dropped by the implicit deny. The only thing that changed was the address the ACL targeted. ## Reading the Compiled Rule The detailed packet-tracer output rewards a closer look, because the compiled forward-flow rule is the ASA showing you its own internal logic. Take the working case again: `src ip/id=0.0.0.0, mask=0.0.0.0` is how the firewall represents `any` for the source, a zero address with a zero mask matches everything, which is exactly what our `permit tcp any` asked for. The destination line, `dst ip/id=172.20.30.80, mask=255.255.255.255`, is the real host with a full 32-bit host mask, the compiled form of `host 172.20.30.80`. The `priority=13` and `domain=permit, deny=false` fields tell you this is a permit entry and where it sits in the compiled order, and `hits=4` is the live match count on that specific compiled rule. When you are chasing why a flow does or does not match, this level of detail removes all guesswork: the firewall is telling you the exact address, mask, port, and protocol it will compare the packet against, and whether the last few packets actually matched. If your intended real IP does not appear in the `dst ip/id` field, the compiled rule is not what you think it is, and that is where the fix lives. ## Why This Failure Is So Nasty What makes the mapped-IP mistake genuinely dangerous is not that it fails, it is *how* it fails. It does not throw a config error. The ASA happily accepts `permit tcp any host 203.0.113.80 eq www`; it is a syntactically valid rule referencing a valid address. The rule simply never matches any real traffic, because no packet ever reaches the ACL with 203.0.113.80 still in the destination field. The permit sits in the config looking correct, the connections fail, and the deny counter on the implicit rule climbs while the counter on your permit stays at zero. That zero-hit permit is the tell. If you have configured an inbound permit for a published server and it shows no hits while the traffic is clearly arriving and failing, check the address. There is a very good chance you named the public IP out of habit, and the fix is to point the rule at the real private address instead. The `show access-list` hit counts and a detailed packet-tracer together diagnose this in under a minute once you know to look for it. ## The Historical Habit That Breaks Rules Where does the mapped-IP instinct come from? History. On PIX and pre-8.3 ASA, NAT and ACLs were ordered the other way: the interface ACL was evaluated against the mapped (public) address, before translation. So for years the correct answer genuinely was "reference the public IP," and a generation of engineers learned it that way, wrote it into runbooks, and taught it to the next hire. Cisco reversed the order in 8.3, and the correct answer flipped with it. Anyone carrying the pre-8.3 habit forward onto a modern firewall writes a rule that is correct by the old rules and silently non-functional by the new ones. That is why this specific behaviour trips up experienced engineers more than beginners: the beginners never learned the old way, but the veterans have to unlearn it. If you find yourself reaching for the public IP in an inbound ACL, that reflex is a pre-8.3 muscle memory, and on 8.3-plus it is exactly backwards. Pre-8.3 (PIX / old ASA) ACL evaluated: before NAT ACL references: the mapped (public) IP Inbound rule: `permit tcp any host 203.0.113.80 eq www` 8.3 and later (modern ASA) ACL evaluated: after un-NAT ACL references: the real (private) IP Inbound rule: `permit tcp any host 172.20.30.80 eq www` ## How to Never Get This Wrong Again The durable fix is not to memorise "use the real IP." It is to internalise the phase order, because the ACL rule is just a consequence of it. Un-NAT is Phase 1, the ACL check is Phase 2, so the packet is already un-translated when the rule runs, so the rule sees the real IP, so you write the real IP. When you can derive the answer from the sequence, you will get it right even on an unfamiliar topology where you cannot remember which address is which. For the complete per-packet sequence, where un-NAT and the interface ACL sit relative to the route lookup, connection creation, and inspection, work through the [ASA packet flow](https://www.pinglabz.com/cisco-asa-packet-flow/) article. Pair that with the destination NAT publish that set this up, and the real-IP ACL rule becomes obvious rather than surprising. ## A 60-Second Diagnostic When a published server is unreachable and you suspect the ACL, this is the fastest path to an answer. First, run `show access-list OUTSIDE-IN` and look at the hit count on your inbound permit. If it is zero while connections are clearly being attempted, the rule is not matching, and a mapped-IP mistake is the prime suspect. Second, run a detailed packet-tracer for the exact flow and read the ACCESS-LIST phase: an ALLOW with the real IP in the compiled `dst ip/id` field means the rule is right, and a DROP on the Implicit Rule means your permit never matched. Third, confirm the address in the permit is the real private IP of the server, not its public mapped IP. That single check resolves the overwhelming majority of "I permitted it but it still fails" cases on a modern ASA. The whole loop takes under a minute once you know that the un-NAT-before-ACL ordering is what makes the real IP the correct target. It is a small habit that saves the afternoon the mapped-IP rule would otherwise cost you. ## Key Takeaways - **On ASA 8.3+, interface ACLs reference the real IP, not the mapped IP.** Un-NAT (Phase 1) runs before the ACL check (Phase 2), so the destination is already the real address when the rule is evaluated. - **Proof A: real IP permits.** The compiled rule carries `use_real_addr` and `dst ip=172.20.30.80`; a permit for that real IP matches and the flow is allowed. - **Proof B: mapped IP drops.** A permit for 203.0.113.80 never matches (the packet is already un-NATed), so the Implicit Rule denies it with `acl-drop: Flow is denied by configured rule`. - **The failure is silent.** A mapped-IP permit is valid config that never matches; the tell is a zero-hit permit while the implicit deny counter climbs. - **It is a pre-8.3 habit.** Old PIX and ASA evaluated ACLs against the mapped IP before NAT; 8.3 reversed the order, so the veteran instinct is now exactly wrong. Real IPs in your ACLs is one rule you cannot afford to get backwards. For the rest of the inbound and outbound story, from publishing servers to reading the NAT table, continue through the [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/). ### The ASA NAT Table Explained: Sections, Order, and Which Rule Wins URL: https://www.pinglabz.com/asa-nat-table-order/ Last updated: 2026-07-13T10:31:00.000Z ASA NAT feels mysterious right up until the moment you understand the table. Once you can picture the sections, read them top to bottom, and see that the first matching rule wins, every "why did that packet get translated by the wrong rule" question answers itself. This is the single mental model that makes Cisco ASA NAT stop being a guessing game, and the good news is that the firewall shows you the whole table in one command. This article is part of the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) cluster. It is the companion to the [NAT order of operations](https://www.pinglabz.com/cisco-asa-nat-order-of-operations/) article: that one covers the *phase* order (what the ASA does to a packet, un-NAT then ACL then route), while this one covers the *table* order (which of your configured NAT rules the ASA picks in the first place). You need both to reason about a real firewall. Here we work from live `show nat` output on an ASAv, with an identity NAT and a server publish sitting in different sections, and prove which rule a real flow hits. ## The Three Sections (Plus a Hidden One) Every NAT rule you configure on an ASA lands in one of three numbered sections, and the ASA evaluates them strictly in order. There is also a section zero you do not write yourself but will see in the output, so it is worth knowing all four: Section 0: Implicit / system Syntax: auto-generated (the `nlp_*` entries) Who writes it: the ASA itself, for management Typical use: internal system NAT (SSH, SNMP to the box); you never configure it Section 1: Manual / Twice NAT Syntax: `nat (x,y) source ... destination ...` Order: evaluated first of your rules Typical use: exemptions, policy NAT, anything that depends on source and destination Section 2: Auto / Object NAT Syntax: `nat` inside `object network` Order: evaluated after Section 1 Typical use: general translations, publishing a host, dynamic PAT for an outbound subnet Section 3: Manual after-auto Syntax: `nat (x,y) ... after-auto` Order: evaluated last Typical use: catch-all manual rules you want to run only if nothing above matched The rule to memorise is short: **Section 1, then Section 2, then Section 3, top to bottom, first match wins.** The moment a packet matches a rule, the ASA stops looking. Nothing below that rule gets a vote. Section 0 sits above all of it and is the firewall's own housekeeping, so you rarely think about it, but recognising the `nlp_*` entries stops them from confusing you when you read the table. ## The Whole Table in One Command Here is the payoff. A single `show nat` prints every section at once, in evaluation order, with per-rule hit counters. This output, from the lab ASAv, is the centrepiece of the whole topic because it shows the sections coexisting: ``` FW1# show nat Manual NAT Policies Implicit (Section 0) 1 (nlp_int_tap) to (inside) source static nlp_server__snmp_... service udp snmp snmp translate_hits = 0, untranslate_hits = 0 2 (nlp_int_tap) to (outside) source static nlp_server__ssh_0.0.0.0_intf2 interface ... service tcp 4122 ssh translate_hits = 32, untranslate_hits = 32 ... (system management NAT, sections you did not write) ... Manual NAT Policies (Section 1) 1 (inside) to (dmz) source static INSIDE-NET INSIDE-NET destination static DMZ-NET DMZ-NET translate_hits = 0, untranslate_hits = 0 Auto NAT Policies (Section 2) 1 (dmz) to (outside) source static DMZ-WEB DMZ-WEB-PUB translate_hits = 0, untranslate_hits = 7 ``` Read it from the top. Section 0 holds the system's own management NAT (the SSH-to-the-box rule has 32 hits, which is just administrators logging in). You did not write those and you should not touch them. Then Section 1 holds our identity NAT, the [inside-to-DMZ exemption](https://www.pinglabz.com/asa-static-identity-nat/) that translates nothing. Then Section 2 holds our server publish, the object NAT that maps the DMZ web server to a public IP. Three different kinds of rule, three different sections, one command, in the exact order the ASA will consult them. ## Why the Section a Rule Lands In Is the Priority Here is the teaching point that this table proves, and it is the crux of ASA NAT design. Our identity NAT is a **manual** rule, so it sits in Section 1\. Our server publish is an **auto** rule, so it sits in Section 2\. If a single packet could plausibly match both, the Section 1 rule wins, because Section 1 is evaluated first and the first match ends the search. That is not an accident you have to design around; it is a feature you design *with*. You never assign explicit priority numbers to ASA NAT rules the way you might sequence a route-map. Instead, you choose the rule type based on where you want it to land: - **Specific exemptions go in as manual (twice) NAT** so they land in Section 1 and get first look. A NAT-exemption rule for VPN traffic has to win over your general PAT, and putting it in Section 1 guarantees that. - **General translations go in as auto (object) NAT** so they land in Section 2 and only catch what the exemptions did not. Your broad outbound PAT and your server publishes belong here. - **True catch-alls go in as manual after-auto** so they land in Section 3 and run only if nothing above matched. Choose the type, and the section ordering does the prioritisation for you. This is why the advice "write exemptions as twice NAT and general rules as object NAT" is not a style preference; it is how you control which rule wins. ## Order Within a Section Sections set the coarse priority, but there is a finer level too. Within Section 1 and Section 3, manual NAT rules are evaluated in the order you enter them, top to bottom, and you can control that order explicitly (inserting a rule at a specific line). So the full precedence is: Section 1 rules in configured order, then Section 2, then Section 3 rules in configured order. Section 2 (auto NAT) is different. You do not control its order by hand; the ASA orders object NAT rules automatically, roughly most-specific first (static before dynamic, smaller prefixes before larger). That automatic ordering is usually what you want, and it is another reason to keep general translations in Section 2 and reserve manual sections for the cases where you genuinely need to dictate sequence. ## Where Section 3 (after-auto) Actually Earns Its Keep Section 3 is the one most engineers never touch, and that is usually correct, but it exists for a real reason. Manual NAT normally lands in Section 1, above your object NAT. Sometimes you want a manual rule (say, a broad policy translation that depends on both source and destination) but you want it to run *after* your specific object NAT rules, not before them. Adding `after-auto` to the manual statement moves it from Section 1 to Section 3, so it becomes a fallback rather than a first responder. The classic case is a general catch-all written as twice NAT. If you put it in Section 1, it would shadow your object NAT publishes and PAT rules, grabbing traffic before the more specific rules got a chance. Push it into Section 3 with `after-auto` and it only fires for whatever the object NAT rules in Section 2 left unmatched. The lesson is that "manual NAT" and "Section 1" are not the same thing: a manual rule is Section 1 by default, but `after-auto` lets you demote it to Section 3 when you want auto NAT to win. If you ever find a manual NAT rule that seems to be losing to an object NAT rule you expected it to beat, check for `after-auto` on the statement. That one keyword changes which section the rule lives in and therefore which rule wins, and it is easy to miss when scanning a running config. ## Proving Which Rule a Real Flow Hit Knowing the theory is one thing; proving which rule a live packet actually matched is what you do at 2 a.m. when something is translated wrong. There are two tools, and they agree with each other. The first is the set of hit counters in `show nat`. Every rule carries `translate_hits` (outbound, source translation) and `untranslate_hits` (inbound, destination un-translation). In the table above, the Section 2 publish shows `untranslate_hits = 7`: seven inbound connections have un-NATed through it. If you expect a flow to hit a particular rule and its counters are not moving, the flow is matching something else, and the counters on the other rules will tell you which. Watching the counters move, or fail to move, while you generate traffic is the most direct evidence there is. The second tool is `packet-tracer`, which simulates a specific packet and names the exact rule it matched in the UN-NAT phase (it prints the matching `Config:` block). Between the two, you never have to guess: the counters tell you what real traffic did in aggregate, and packet-tracer tells you what one specific flow will do and why. The deeper walk through the ASA's per-packet processing, where NAT sits relative to the route lookup and the interface ACL, is in the [ASA packet flow](https://www.pinglabz.com/cisco-asa-packet-flow/) article. ## A Worked Example: Two Rules, One Packet Tie it together with the two rules from our table. Suppose an inside host at 10.20.10.100 opens a connection to the DMZ web server at 172.20.30.80\. Which rule handles it? The ASA starts at Section 0 (management NAT, no match for this traffic), then Section 1\. There it finds the identity NAT: `source static INSIDE-NET INSIDE-NET destination static DMZ-NET DMZ-NET`. Our packet is inside-to-DMZ, both networks match, and the rule fires. First match wins, so the search stops right there. The Section 2 publish rule never even gets considered for this flow, because it lives below a rule that already matched. The inside host reaches the DMZ server with its real address intact, exactly as the exemption intends. Now flip it: an outside client hits 203.0.113.80\. Section 0 has no match, Section 1's identity rule is scoped to inside-to-DMZ so it does not match outside-to-DMZ traffic, and the search falls through to Section 2, where the object NAT publish un-NATs the destination to 172.20.30.80\. Same firewall, same table, two flows, two different rules, and the section order decided each one. That is the entire model. ## Reading the Table Like a Pro A few habits turn `show nat` from a wall of text into a diagnostic instrument. First, always read it top to bottom and mentally note the section headers as you go; the ASA prints them for you (`Manual NAT Policies (Section 1)`, `Auto NAT Policies (Section 2)`), so you never have to guess which section a rule is in. Second, treat the Section 0 `nlp_*` entries as background noise for your own troubleshooting; they are the firewall managing access to itself, and a hit count climbing there is just administrative traffic, not your data plane. Third, use the counters as a story. A publish rule with `untranslate_hits` steadily climbing is being reached by inbound clients. An exemption rule with flat counters while its traffic clearly flows is being shadowed by a rule above it, which points you straight at a section-ordering problem. Fourth, when in doubt, do not argue with the config in your head; run `packet-tracer` for the exact flow and let the firewall tell you which rule wins. The table order is deterministic, so packet-tracer's answer is authoritative. Put those four habits together and NAT troubleshooting collapses into a short checklist: identify the flow, find the first rule in section order that could match it, confirm with the counters or packet-tracer, and fix the placement if the wrong rule is winning. There is no hidden magic in ASA NAT once you can read the table the way the firewall evaluates it. ## Key Takeaways - **ASA NAT has three sections you write, plus a hidden Section 0.** Section 1 is manual (twice) NAT, Section 2 is auto (object) NAT, Section 3 is manual after-auto, and Section 0 is the system's own management NAT. - **Evaluation is top to bottom, first match wins.** The moment a packet matches a rule the ASA stops looking, so a Section 1 rule beats a Section 2 rule for the same packet. - **The section a rule lands in is its priority.** Write specific exemptions as manual NAT (Section 1) and general translations as object NAT (Section 2); the section ordering does the prioritising for you. - **show nat prints the whole table at once.** Every section in evaluation order, with per-rule `translate_hits` and `untranslate_hits`. - **Prove which rule a flow hit two ways.** Watch the hit counters move under real traffic, and use `packet-tracer` to see the exact matching rule for one simulated flow. The table order and the phase order are two halves of the same skill. Pair this with the [NAT order of operations](https://www.pinglabz.com/cisco-asa-nat-order-of-operations/) and the rest of the [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/) to reason confidently about any flow through the firewall. ### Static Identity NAT on the ASA: Translating Nothing, On Purpose URL: https://www.pinglabz.com/asa-static-identity-nat/ Last updated: 2026-07-13T10:30:59.000Z Most NAT rules exist to change an address. Identity NAT is the odd one out: it is a rule whose entire job is to make sure an address is **not** changed. On the Cisco ASA you write it as a static translation from a network to itself, and while that sounds pointless, it is one of the most important rules on a real firewall. It is how you carve exceptions out of a broader NAT policy, and it is what keeps two trusted segments talking with their real addresses instead of translated ones. This article is part of the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) cluster. Here we build a static identity NAT between an inside network and a DMZ, prove it really does translate nothing using `show nat detail` and `packet-tracer` from a live ASAv, and explain the two real reasons you reach for it: NAT exemption for VPN traffic, and preserving real client IPs in your server logs. ## Translating a Network to Itself The lab ASAv has an inside network 10.20.10.0/24 and a DMZ 172.20.30.0/24\. We want traffic between them to pass through the firewall with its original addresses intact, no translation at all. The rule that does this is a Twice NAT (manual NAT) statement that maps each network to itself: ``` object network INSIDE-NET subnet 10.20.10.0 255.255.255.0 object network DMZ-NET subnet 172.20.30.0 255.255.255.0 nat (inside,dmz) source static INSIDE-NET INSIDE-NET destination static DMZ-NET DMZ-NET ``` Read the NAT line carefully, because the repetition is the whole point. `source static INSIDE-NET INSIDE-NET` says translate the source `INSIDE-NET` to `INSIDE-NET`, the same object twice. `destination static DMZ-NET DMZ-NET` does the same for the destination. Both the source and the destination map to themselves. This is what engineers mean by **twice NAT to itself**: a manual NAT rule where the original and translated values are identical on both sides. One structural detail matters and pays off later. Because this is a manual (twice) NAT statement, it lands in **Section 1** of the ASA NAT table, which is evaluated before the auto NAT rules in Section 2\. That ordering is deliberate. Exemptions have to be consulted before general translations, or a broad rule would grab the traffic first. The section system does that prioritisation for you, a point worth reading in full in the [ASA NAT table order](https://www.pinglabz.com/asa-nat-table-order/) walkthrough. ## Proving It Is Really an Identity Translation Saying a rule translates nothing is easy. Proving it is a two-line check. `show nat` gives you the rule and its counters; `show nat detail` adds the crucial Origin and Translated values so you can confirm they are identical: ``` FW1# show nat Manual NAT Policies (Section 1) 1 (inside) to (dmz) source static INSIDE-NET INSIDE-NET destination static DMZ-NET DMZ-NET translate_hits = 0, untranslate_hits = 0 FW1# show nat detail Manual NAT Policies (Section 1) 1 (inside) to (dmz) source static INSIDE-NET INSIDE-NET destination static DMZ-NET DMZ-NET Source - Origin: 10.20.10.0/24, Translated: 10.20.10.0/24 Destination - Origin: 172.20.30.0/24, Translated: 172.20.30.0/24 ``` There is the proof. `Source - Origin: 10.20.10.0/24, Translated: 10.20.10.0/24`: same address in, same address out. The destination line tells the same story for the DMZ network. When Origin equals Translated on both source and destination, you are looking at an identity rule. This is the fastest way to confirm an exemption is genuinely exempting and not quietly rewriting something you did not intend. ## Watching a Packet Pass Through Unchanged The counters confirm the config; `packet-tracer` confirms the behaviour on an actual flow. We simulate an inside host reaching the DMZ web server and look at the UN-NAT phase: ``` FW1# packet-tracer input inside tcp 10.20.10.100 12345 172.20.30.80 80 Phase: 2 Type: UN-NAT Subtype: static Result: ALLOW Config: nat (inside,dmz) source static INSIDE-NET INSIDE-NET destination static DMZ-NET DMZ-NET Additional Information: NAT divert to egress interface dmz Untranslate 172.20.30.80/80 to 172.20.30.80/80 ``` Compare this to a normal destination NAT, where the untranslate line shows two different addresses (a public IP becoming a private one). Here it reads `Untranslate 172.20.30.80/80 to 172.20.30.80/80`: the same address on both sides. The ASA runs the UN-NAT phase, matches the identity rule, and hands the packet on with its destination untouched. The firewall still tracks and inspects the flow; it simply does not rewrite the address. That is exactly what you want an exemption to do. ## Identity NAT vs No Rule at All A fair question at this point: if the rule changes nothing, why not simply leave it out and let the traffic pass? The answer is that on a firewall configured with a general translation policy, "no rule" does not mean "no translation." The moment you have a broad PAT or object NAT covering the inside network, any inside-to-DMZ traffic that is not exempted will fall through to that general rule and get translated. Deleting the identity NAT does not give you clean, un-translated traffic; it hands the flow to whatever general rule catches it next. That is the real reason identity NAT exists as an explicit statement. You are not adding a translation, you are inserting a higher-priority exception that says "for this specific pair of networks, stop here and translate nothing" before the general rules downstream ever get a look. On a firewall with no other NAT at all, you would indeed see the same behaviour with or without the rule. On a real firewall carrying a general policy, the identity rule is what protects the exception. You can also watch this in the counters over time. Because our lab ASA had no competing general rule for this pair yet, `translate_hits` and `untranslate_hits` both read 0 in the capture above. On a busy firewall those numbers climb, and a healthy identity rule shows hits accumulating while the addresses in `show nat detail` stay identical. If the hits stay flat while inside-to-DMZ traffic clearly flows, the traffic is matching a different rule higher or lower in the table, and that is your cue to check the section ordering. ## Why You Would Ever Do This A rule that changes nothing sounds like a rule you could delete. It is not. There are two concrete reasons identity NAT earns its place, and both come up constantly on production firewalls. NAT exemption for VPN traffic Traffic destined for a site-to-site or remote-access VPN peer must reach the crypto map with its real addresses, not translated ones, or it never matches the interesting-traffic ACL. An identity NAT rule, placed in Section 1 above your general PAT, keeps VPN-bound traffic un-translated so the tunnel selects it correctly. Real client IPs in server logs If inside-to-DMZ traffic were translated, every request in the DMZ server's access log would show the firewall's address instead of the actual client. Keeping the segment un-translated means the server logs, rate limits, and geolocates the real source, which matters for auditing and troubleshooting. The VPN case is the classic one and the reason identity NAT is often called **NAT exemption** or **no-NAT**. Its longer treatment, tied to a working tunnel, lives in [identity NAT for VPN traffic](https://www.pinglabz.com/cisco-asa-identity-nat-vpn/). The log-integrity case is quieter but just as real: the moment you translate internal-to-DMZ traffic, you blind every server behind the DMZ to who is actually talking to it. ## Placement Is the Point Notice that we deliberately built this as a manual NAT rule rather than an object NAT rule. That was not stylistic. Manual NAT lands in Section 1, and Section 1 is consulted before the auto NAT rules that do the general translating in Section 2\. An exemption has to win the match, and it wins because of where it sits in the table. This is the general pattern for ASA NAT design: write your specific exemptions as manual (twice) NAT so they land in Section 1, and write your broad, general translations as auto (object) NAT so they land in Section 2\. You never have to fight over ordering with an explicit priority number; the section structure enforces it. Get that mental model straight and identity NAT stops feeling like a trick and starts feeling like the obvious tool for the job. The full section-by-section breakdown, including how to prove which rule a real flow hit, is in the [NAT table order](https://www.pinglabz.com/asa-nat-table-order/) article. ## Key Takeaways - **Identity NAT translates an address to itself, on purpose.** The syntax is a twice NAT statement: `nat (inside,dmz) source static INSIDE-NET INSIDE-NET destination static DMZ-NET DMZ-NET`. - **Prove it with show nat detail.** When `Origin` equals `Translated` on both the source and destination lines, the rule genuinely changes nothing. - **The packet-tracer shows it too.** An identity rule reports `Untranslate 172.20.30.80 to 172.20.30.80`, the same address on both sides, unlike a real destination NAT. - **You do it for two reasons.** NAT exemption so VPN-bound traffic keeps its real addresses and matches the tunnel, and log integrity so DMZ servers see real client IPs. - **Placement does the work.** Because it is manual NAT it sits in Section 1 and is consulted before the general auto NAT in Section 2, which is exactly why exemptions are written as manual rules. Identity NAT is one piece of a larger design. To see how it sits alongside your publishing and PAT rules, and which rule wins when several could match a packet, continue through the [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/). ### Destination NAT on the ASA: Policy NAT for Inbound Services URL: https://www.pinglabz.com/asa-destination-nat/ Last updated: 2026-07-13T10:30:58.000Z Standing up a web server is easy. Letting the internet reach it, through a firewall, without exposing the rest of your network, is where the Cisco ASA earns its keep. When you publish a private host to a public address, you are doing **destination NAT** (often called policy NAT or, on the ASA, static object NAT for inbound services). The address a client types is not the address the server actually holds, and the ASA quietly rewrites one into the other on the way in. This article is part of the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) cluster. If you have already read [static NAT for a DMZ server](https://www.pinglabz.com/cisco-asa-static-nat-dmz/), this is the same idea taken deeper: not just the config that publishes the box, but the exact order the ASA processes the inbound packet, proven with `packet-tracer` against a real ASAv in a lab, and a real HTTP request from a Linux host that comes back with a genuine 200. ## What Destination NAT Actually Means Most NAT you meet first is source NAT: an inside host with a private address reaches out, and the firewall rewrites the source so the return traffic can find its way back. Destination NAT is the mirror image. The client already knows where it wants to go (a public address you have advertised), and the firewall rewrites the *destination* so the packet lands on the real, private server behind it. Nothing about the client changes; the target address does. On the ASA this is not a separate feature with its own command. A single static NAT statement gives you both directions at once, because static translations are inherently bidirectional. You configure the mapping once, in the direction that reads naturally (real host to public host), and the ASA applies it as source NAT for traffic the server originates and as destination NAT (un-NAT) for traffic arriving for it. That symmetry is why the same `nat ... static` line that lets your DMZ server browse out also publishes it inbound, and it is why the inbound proof lives in the `untranslate_hits` counter rather than a separate rule. ## The Setup: One DMZ Web Server, One Public IP The lab is a single ASAv (9.24) with three interfaces: outside 203.0.113.10, inside 10.20.10.1, and dmz 172.20.30.1\. Sitting on the DMZ is an nginx server at **172.20.30.80**. That address is private and unroutable on the internet, so we publish it behind a public address, **203.0.113.80**. Clients on the outside connect to 203.0.113.80 and never know the real host exists. Here is the entire configuration. Two objects (the real host, the public host) and one static NAT statement inside the real host's object, plus the access rule that lets the traffic in: ``` object network DMZ-WEB host 172.20.30.80 object network DMZ-WEB-PUB host 203.0.113.80 object network DMZ-WEB nat (dmz,outside) static DMZ-WEB-PUB access-list OUTSIDE-IN extended permit tcp any host 172.20.30.80 eq www access-group OUTSIDE-IN in interface outside ``` Read the NAT line as a sentence: for traffic arriving *(dmz,outside)*, statically map the object `DMZ-WEB` (172.20.30.80) to `DMZ-WEB-PUB` (203.0.113.80). Because this is static, the mapping is bidirectional. Outbound, the server's source is rewritten to the public IP. Inbound, the destination is un-NATed from the public IP back to the real host. That inbound direction is the destination NAT we care about here. Now look at the access list. It permits traffic to `host 172.20.30.80`, the **real** private address, not the public 203.0.113.80 a client actually targets. That is not a typo. On ASA 8.3 and later, interface ACLs reference the real, post-translation IP. It looks wrong until you understand the order of operations, which is exactly what the next section proves. ## The Hero Capture: Un-NAT Happens Before the ACL The single most useful tool on the ASA for understanding NAT is `packet-tracer`. It simulates a packet through the firewall and prints every phase it hits, in order. We inject a TCP packet from an outside client (192.168.99.100) to the public IP on port 80 and watch what the ASA does: ``` FW1# packet-tracer input outside tcp 192.168.99.100 12345 203.0.113.80 80 Phase: 1 Type: UN-NAT Subtype: static Result: ALLOW Config: object network DMZ-WEB nat (dmz,outside) static DMZ-WEB-PUB Additional Information: NAT divert to egress interface dmz Untranslate 203.0.113.80/80 to 172.20.30.80/80 Phase: 2 Type: ACCESS-LIST Result: ALLOW Config: access-group OUTSIDE-IN in interface outside access-list OUTSIDE-IN extended permit tcp any host 172.20.30.80 eq www ``` This is the whole lesson in two phases. **Phase 1 is UN-NAT.** Before the ASA does anything else with an inbound packet, it un-translates the destination: `Untranslate 203.0.113.80/80 to 172.20.30.80/80`. The public IP the client sent becomes the real private IP. The ASA also notes it must divert this flow to the dmz egress interface. **Phase 2 is the ACCESS-LIST check**, and by the time it runs, the destination is already 172.20.30.80\. That is why the ACL permits the real IP. The firewall is no longer looking at 203.0.113.80 when it consults the rule; it un-NATed first. Match the ACL against the real host, and the check passes. Match it against the public IP, and it would silently never fire (that failure mode is the whole subject of a companion article, linked below). Keep this ordering in your head and ASA NAT stops being mysterious: **un-NAT, then ACL**. Everything else follows from it. ## Proof It Actually Works: A Real HTTP 200 Simulation is good, but a real request closes the loop. From a Debian host acting as the internet client, we curl the public IP straight through the firewall and ask curl to print the status, the address it actually connected to, and the round-trip time: ``` j@llmbits:~$ curl -s -o /dev/null -w 'HTTP %{http_code} from %{remote_ip} in %{time_total}s\n' http://203.0.113.80/ HTTP 200 from 203.0.113.80 in 0.020190s j@llmbits:~$ curl -s http://203.0.113.80/ | head -4 Welcome to nginx! ``` A real Linux host connected to 203.0.113.80, the ASA un-NATed the destination to 172.20.30.80, forwarded it to the DMZ, and nginx answered with a 200 and its welcome page in about 20 milliseconds. The client never saw the private address. That is destination NAT doing its job. ## The Counters: show xlate and show nat detail Once traffic is flowing you want to confirm which translation it hit and how often. Two commands do it. `show xlate` lists the live translation slots, and `show nat detail` shows the configured rule with its hit counters: ``` FW1# show xlate NAT from dmz:172.20.30.80 to outside:203.0.113.80 flags s idle 0:00:23 timeout 0:00:00 FW1# show nat detail Auto NAT Policies (Section 2) 1 (dmz) to (outside) source static DMZ-WEB DMZ-WEB-PUB translate_hits = 0, untranslate_hits = 7 Source - Origin: 172.20.30.80/32, Translated: 203.0.113.80/32 ``` Two things to notice. First, this is an **Auto NAT** rule, so it sits in Section 2 of the NAT table (the section a rule lands in decides which rule wins when several could match, a topic worth its own read). Second, the counter that climbs for inbound connections is `untranslate_hits`, here at 7\. Each time an outside client hits the public IP, the ASA un-translates the destination and the untranslate counter ticks up. The `translate_hits` counter, by contrast, moves when the server originates traffic outbound. If you publish a server and inbound connections work but the untranslate counter stays flat, the traffic is not reaching this rule at all, and you have a routing or ACL problem upstream, not a NAT problem. ## The Honest Gotcha: The First Connection Can Fail Here is something the tidy walkthroughs leave out. The very first time we ran the packet-tracer, before any real traffic had ever flowed, it did not end in ALLOW. It ended like this: ``` Drop-reason: (no-v4-adjacency) No valid V4 adjacency. Check ARP table (show arp) has entry for nexthop. ``` The NAT was correct and the ACL was correct. The problem was purely that the ASA had never talked to 172.20.30.80 and had no ARP entry for it, so it had no Layer 2 adjacency to forward the frame to. A single `ping dmz 172.20.30.80` (100 percent success) warmed the ARP cache, and the very next packet-tracer built the flow cleanly to ALLOW. Worth remembering in production: a brand-new published server can fail its first connection with a no-adjacency drop until the ASA ARPs the host. It resolves itself the instant any traffic (a ping, a health check, the first real client that retries) forces the ASA to learn the MAC. If you publish a server and the first curl times out but the second one works, this is almost certainly why, and it is not a NAT misconfiguration. ## Object NAT vs Twice NAT for Publishing We published this server with Auto NAT (a `nat` statement inside a network object). That is the right tool for a straightforward one-host, one-public-IP publish. When you need the translation to depend on *both* the source and the destination (for example, present a server as one public IP to the internet but a different address to a partner network), you reach for Twice NAT instead. The two approaches map cleanly to when you would use each: Object (Auto) NAT Syntax: `nat` inside `object network` Matches on: the object's address only NAT section: Section 2 Use it for: publishing one host or subnet to one mapped address, the common case here Twice (Manual) NAT Syntax: `nat (x,y) source ... destination ...` Matches on: source and destination together NAT section: Section 1 (evaluated first) Use it for: translations that depend on where the traffic is going, and NAT exemptions For a plain inbound publish, Auto NAT is cleaner and lands exactly where you want it. If you later find yourself needing the destination to steer the translation, that is Twice NAT territory, covered in the [twice NAT walkthrough](https://www.pinglabz.com/cisco-asa-twice-nat/). ## Why the ACL Uses the Real IP (and What Happens If You Get It Wrong) We flagged it twice: the access list permits 172.20.30.80, not 203.0.113.80, because un-NAT runs before the ACL. This trips up nearly everyone coming from pre-8.3 ASA and PIX, where the ACL referenced the mapped (public) IP. Carry that old habit forward on a modern ASA and you write a rule that looks reasonable, matches nothing, and leaves you staring at a climbing deny counter wondering why a permit you clearly configured is not working. The before-and-after proof of that behaviour, the same packet allowed with a real-IP ACL and dropped with a mapped-IP ACL, gets its own detailed treatment in [NAT and ACLs together on ASA 8.3+](https://www.pinglabz.com/asa-nat-acl-real-ip/). If you take one habit from this article, make it this: on a modern ASA, your inbound access rules target the real address of the server, every time. ## Key Takeaways - **Destination NAT publishes a private server to a public IP.** A static object NAT statement (`nat (dmz,outside) static DMZ-WEB-PUB`) maps 172.20.30.80 to 203.0.113.80 bidirectionally. - **Un-NAT happens before the ACL.** Phase 1 of the inbound packet is UN-NAT (`Untranslate 203.0.113.80 to 172.20.30.80`); the access-list check in Phase 2 runs against the already-un-NATed real IP. - **Your inbound ACL references the real IP.** Permit `host 172.20.30.80`, not the public 203.0.113.80, on ASA 8.3 and later. - **Watch untranslate\_hits.** Inbound connections increment `untranslate_hits` in `show nat detail`; a flat counter means the traffic never reached the rule. - **The first connection can fail with no-v4-adjacency.** A newly published server may drop its first packet until the ASA ARPs it; a single ping warms the adjacency and the flow builds. - **Verify with real traffic.** A curl from a real client returning `HTTP 200` proves the whole path end to end, which packet-tracer alone cannot. For the full inbound and outbound picture, how NAT, routing, and access control fit together on the ASA, work through the rest of the [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/), which threads these labs together in order. ### ASA Management Access Done Right: SSH, HTTPS, SNMPv3, NTP, and Logging URL: https://www.pinglabz.com/asa-management-access-configuration/ Last updated: 2026-07-13T10:09:27.000Z Reaching and monitoring a firewall is a security problem in its own right. Every management channel you open, SSH, HTTPS, SNMP, NTP, syslog, is a way in, and each one can either be done properly or done in a way that quietly weakens the box you are trying to protect. This article is the practical, done-right version for the Cisco ASA, and every capture in it is real, because we drove the entire lab this way. It rounds out our [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/) and pairs with the deeper policy discussion in [ASA management hardening](https://www.pinglabz.com/cisco-asa-management-hardening/). All output below came from an ASAv 9.24(1) in Cisco Modeling Labs, reached over SSH from a real Linux host. That detail matters here more than usual, because how we got in is itself part of the story. ## The automation gotcha that comes before anything else Before a single management feature, ASAv 9.24(1) throws a curveball that will break naive automation cold. The first time you try to enter privileged mode, the box refuses to let you until you set the enable password interactively: ``` FW1> enable The enable password is not set. Please set it now. Enter Password: **** Repeat Password: **** Note: Save your configuration so that the password persists across reboots ("write memory" or "copy running-config startup-config"). ``` An `enable password` line in the day-0 startup config does not satisfy this. The ASA still demands the interactive set on first enable, which is why tools like PyATS land on "Too many enable password retries" and give up. The reliable pattern, and the one we used, is to SSH in as a local user, set the enable password once by hand, and `write memory` so it persists. After that the box is fully drivable over SSH. If you automate ASAs, budget for this one manual step per fresh device; there is no config-only way around it on this release. It is a genuinely useful thing to know before you spend an afternoon debugging your automation harness. ## SSH, done right SSH is your primary management channel, and there is one prerequisite people forget: the ASA will not start its SSH server without an RSA key pair. Generate one first: ``` crypto key generate rsa modulus 2048 aaa authentication ssh console LOCAL ssh 192.168.99.0 255.255.255.0 outside ssh version 2 ssh key-exchange group dh-group14-sha256 ``` Skip the `crypto key generate rsa` step and port 22 is simply refused; there is no host key to present, so there is no SSH. After that, `aaa authentication ssh console LOCAL` forces every SSH login through the local user database (or a AAA server) rather than a shared password, `ssh ` restricts which source networks may even attempt a connection and on which interface, `ssh version 2` disables the broken SSHv1, and `ssh key-exchange group dh-group14-sha256` pins a strong key exchange group instead of the weak defaults. You can confirm the negotiated posture with `show ssh`: ``` FW1# show ssh Idle Timeout: 5 minutes Version allowed: 2 Cipher encryption algorithms enabled: aes128-gcm@openssh.com aes256-ctr aes256-cbc aes192-ctr aes192-cbc aes128-ctr aes128-cbc chacha20-poly1305@openssh.com aes256-gcm@openssh.com Cipher integrity algorithms enabled: hmac-sha2-256 Host Key Algorithms: ssh-rsa Hosts allowed to ssh into the system: 0.0.0.0 0.0.0.0 outside 0.0.0.0 0.0.0.0 inside 0.0.0.0 0.0.0.0 dmz ``` `Version allowed: 2` confirms v1 is off, and the integrity line shows `hmac-sha2-256` rather than an SHA-1 MAC. When a login actually happens, the AAA path leaves a clear trail in the log: ``` %ASA-6-113012: AAA user authentication Successful : local database : user = plabz %ASA-6-113008: AAA transaction status ACCEPT : user = plabz %ASA-6-611101: User authentication succeeded: IP address: 192.168.99.100, Uname: plabz %ASA-5-111008: User 'plabz' executed the 'enable' command. ``` That sequence is what a healthy, authenticated management login looks like: the local database accepted the user, the transaction was recorded, and the source IP and username are captured. Note the last line, `%ASA-5-111008`, records that the user entered privileged mode. On a firewall you want that kind of accountability trail, and it comes for free once AAA authentication and logging are on. One more point about the source restriction. In the running config those `ssh` lines showed up as `0.0.0.0 0.0.0.0` on all three interfaces, which is deliberately wide open for a lab so we could drive it from anywhere. In production you would never leave it that way. You scope SSH to the specific management subnet and, critically, you do not permit it inbound on the outside interface at all unless you have a very good reason and a VPN in front of it. Exposing SSH to the internet on a firewall's outside interface is one of the fastest ways to end up in someone's brute-force logs. Treat the wide-open lab values above as a demonstration of the mechanism, not a template to copy. ## HTTPS and ASDM, restricted the same way The web management path (HTTPS, which is also how ASDM connects) obeys the same discipline as SSH: turn it on only if you need it, and restrict who can reach it. The pattern mirrors the SSH one, source-scoped and interface-scoped: ``` http server enable http 192.168.99.0 255.255.255.0 outside aaa authentication http console LOCAL ``` The `http ` line does for HTTPS exactly what the `ssh` line does for SSH: it defines which source networks may connect and on which interface, and everything else is denied. As with SSH, front the login with `aaa authentication http console LOCAL` so access goes through the user database rather than a shared secret. Two hardening notes that matter here: prefer to expose ASDM on an inside or dedicated management interface rather than the outside, and pin a modern TLS posture with the `ssl cipher` settings so the box does not negotiate down to weak ciphers. If you do not use ASDM at all, the most secure option is simply to leave `http server enable` off, because the smallest attack surface is the feature you never turned on. The same "only what you need, scoped to who needs it" principle runs through the whole of [ASA management hardening](https://www.pinglabz.com/cisco-asa-management-hardening/). ## SNMPv3, not v2c If you monitor the ASA over SNMP, use version 3 with authentication and privacy, and never a v2c community string. A community string is a plaintext password that travels in the clear and grants read (or worse, write) access to anyone who sniffs it. SNMPv3 fixes both problems: it authenticates the monitoring station and encrypts the payload. Here is the config and the proof: ``` snmp-server group PLZ-V3-GRP v3 priv snmp-server user plzmon PLZ-V3-GRP v3 auth sha PlzAuth123 priv aes 128 PlzPriv123 snmp-server host inside 10.20.10.100 version 3 plzmon FW1# show snmp-server user User name: plzmon Engine ID: 80000009fe5e09945c732cf9c3a56889d4afaa75417f03639d storage-type: nonvolatile active Authentication Protocol: SHA Privacy Protocol:AES128 Group-name: PLZ-V3-GRP ``` The `v3 priv` on the group means this group requires both authentication and privacy (encryption), the strongest of the three SNMPv3 security levels. The user is created with `auth sha` (SHA for authentication) and `priv aes 128` (AES-128 for encryption). The `show snmp-server user` output confirms it: `Authentication Protocol: SHA` and `Privacy Protocol: AES128`. There is no community string anywhere in this configuration, and that is exactly the point. If your ASA monitoring still relies on a v2c community, this is the pattern to migrate to. ## NTP with authentication Accurate time on a firewall is not a nicety. Certificate validation, VPN lifetimes, and above all log correlation across devices all depend on it, and a log with the wrong timestamp is worse than no log when you are reconstructing an incident. But the risk with NTP is that a bad actor feeds the firewall false time, so you authenticate the time source: ``` ntp authenticate ntp authentication-key 10 md5 ntp trusted-key 10 ntp server 10.255.0.1 key 10 source inside FW1# show ntp associations address ref clock st when poll reach delay offset disp ~10.255.0.1 0.0.0.0 16 35 64 0 0.0 0.00 16000. * master (synced), # master (unsynced), + selected, - candidate, ~ configured ``` `ntp authenticate` turns on authentication, the authentication-key defines the shared secret with key ID 10, `ntp trusted-key 10` says "only accept time from a server that proves it holds key 10", and the `ntp server ... key 10` line ties the server to that key. The effect is that the ASA will not sync to a server that cannot authenticate, which closes the door on a spoofed time source. Now the honest note. In the capture above the `reach` column reads `0` and there is no `*` next to the association, meaning the firewall has not synced yet. That is expected: in a fresh lab an NTP association takes several minutes to reach a synced state as the `reach` counter climbs from 0, and CML's virtual clocks start badly skewed, which lengthens the process. The point of this section is not to show a synced clock, it is to show the authentication model: `ntp authenticate` plus a trusted key so the client only ever trusts time from a server that proves the shared secret. A wrong key on either side means the association simply never becomes trusted, which is the same failure class as the RIP authentication problem elsewhere in this cluster, a mismatched key silently (or in RIP's case, loudly) refusing to trust the peer. ## Logging: the messages every firewall engineer learns first None of the above is worth much if you cannot see what the firewall is doing. Turn on logging and send it somewhere durable: ``` logging enable logging buffered informational logging host inside 10.20.10.100 ``` `logging buffered informational` keeps a rolling buffer in memory at informational severity and above, and `logging host` ships everything to a syslog collector so the record survives a reboot. Here is a slice of the real buffer from our firewall: ``` %ASA-6-302013: Built inbound TCP connection 36 for outside:192.168.99.100/48030 to ... (203.0.113.10/22) %ASA-6-302014: Teardown TCP connection 35 for outside:192.168.99.100/36894 ... duration 0:00:22 bytes 16305 %ASA-5-502103: User priv level changed: Uname: enable_15 From: 1 To: 15 %ASA-6-769007: UPDATE: Image version is 9.24(1) ``` Every ASA syslog message follows the same shape: `%ASA--`. The severity is a digit from 0 (emergencies) to 7 (debugging); the message ID identifies the exact event. So `%ASA-6-302013` is a severity-6 (informational) message with ID 302013\. Learn to read the severity at a glance, it tells you whether a line is routine (6, informational) or something you should stop and look at (1, alert, like the RIP failure we saw earlier). The two IDs to commit to memory are `302013` and `302014`: "Built" and "Teardown" for a connection. They are the first messages every firewall engineer learns to read, because they are the firewall narrating its core job, one line when a connection is allowed through and one line when it closes (with duration and byte count, useful for spotting oddly long or oddly large flows). The `%ASA-5-502103` line records a privilege escalation, and `%ASA-6-769007` even tells you the running image version. Together these are the heartbeat of the box. For the full treatment of severities, destinations, and what to alert on, see [ASA syslog and logging](https://www.pinglabz.com/cisco-asa-syslog-logging/). ## Key Takeaways - ASAv 9.24(1) forces you to set the enable password interactively on first enable. A day-0 `enable password` line does not satisfy it, so naive automation fails. Set it once over SSH and `write memory`. - SSH will not start without `crypto key generate rsa`; port 22 is refused until a key exists. Then use AAA authentication, source restrictions, version 2 only, and a strong key exchange group like `dh-group14-sha256`. - Monitor with SNMPv3 `auth ... priv` (SHA + AES-128 here), never a v2c community string. `show snmp-server user` proves the auth and privacy protocols in use. - Authenticate NTP with `ntp authenticate` and a trusted key so the firewall only trusts a time source that proves the shared secret. Lab sync takes minutes and CML clocks start skewed; the article's point is the authentication model, not the sync state. - Every syslog message is `%ASA--`. IDs 302013 and 302014 (connection build and teardown) are the first two every firewall engineer learns to read. - See [ASA management hardening](https://www.pinglabz.com/cisco-asa-management-hardening/) and [ASA syslog and logging](https://www.pinglabz.com/cisco-asa-syslog-logging/) for the deeper dives, and the [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/) for the full build. ### RIPv2 on the Cisco ASA (and Why You'd Still See It) URL: https://www.pinglabz.com/asa-ripv2-configuration/ Last updated: 2026-07-13T10:09:26.000Z RIP is legacy. Nobody designs a new network around it, and if you proposed running RIPv2 as your core routing protocol in 2020s production you would be asked to leave the room. And yet you will still meet it, usually at an interop boundary where some older device or a third-party appliance speaks nothing else. So it is worth knowing how RIPv2 behaves on a Cisco ASA, if only so you recognise it when it turns up. This is a short, honest entry in our [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/), built from two real captures that together tell the whole RIP story: it authenticates loudly, and it almost never wins a route. Every capture here came from an ASAv 9.24(1) in Cisco Modeling Labs, peering with an IOS-XE router, driven over SSH. ## The configuration RIPv2 on the ASA is about as small as routing configuration gets: ``` router rip network 10.0.0.0 version 2 no auto-summary ``` A few things to note. The `network 10.0.0.0` statement is classful, RIP takes the major network, not a subnet mask. The `version 2` line forces RIPv2, which you want because version 1 is classful and has no authentication at all. And `no auto-summary` stops RIP collapsing your subnets to their classful boundary, the same reasoning as EIGRP. This is the minimum viable RIPv2, and on the ASA it really is this short. ## Reality one: RIP authenticates, and it is loud about failure RIPv2 supports authentication, and our lab had it turned on with an MD5 key. But the key on the ASA and the key on CORE1 did not line up. Unlike the silent EIGRP AS mismatch elsewhere in this cluster, the ASA does not keep quiet about a RIP authentication failure. It logs it at severity 1: ``` %ASA-1-107001: RIP auth failed from 10.20.10.2: version=2, type=ffff, mode=3, sequence=10 on interface inside ``` Read that message, because it is dense with useful detail. `RIP auth failed from 10.20.10.2` names the neighbour that sent the update the ASA rejected (that is CORE1). `version=2` confirms we are talking RIPv2, not v1\. `mode=3` is the important field: mode 3 means MD5 authentication. So the ASA is telling you the neighbour presented an MD5-authenticated RIPv2 update, and the authentication did not pass, so the update was dropped. Every RIP update from that neighbour is being discarded and logged until the keys agree. The severity is the part worth internalising. The `%ASA-1-` prefix means severity 1, which on the ASA scale is "alert", one step below the most critical level there is. The firewall does not log routine routing chatter at severity 1\. It is effectively shouting. A dropped RIP update is not, in the grand scheme, an emergency, but the ASA treats an authentication failure as a security-relevant event, and rightly so: a failing auth check is exactly what you would see if someone were trying to inject bogus routes. So the loud severity is a feature, not noise. If you see `%ASA-1-107001` in a buffer, someone's keys are wrong, and you should find out whether it is a benign misconfiguration or something worth investigating. Contrast this with the EIGRP AS-number mismatch, which produces nothing at all, and you get a feel for how differently these protocols behave when something is off. ## Reality two: even with auth fixed, the route will not install Here is the honest heart of the article. Suppose you fix the MD5 key, the authentication passes, and RIP is now happily exchanging valid v2 updates with CORE1\. You would expect the RIP-learned route to appear in the table. It does not. And the reason is not a bug, it is administrative distance doing exactly what it is supposed to. The prefix RIP is advertising, `10.20.20.0/24`, is already in the firewall's route table, learned by two other protocols at the same time. Look at the relevant lines from `show route` on this box: ``` D 10.20.20.0 255.255.255.0 [90/128512] via 10.20.10.2, 00:08:55, inside O 10.20.20.1 255.255.255.255 [110/11] via 10.20.10.2, 00:12:31, inside ``` EIGRP (administrative distance 90) and OSPF (administrative distance 110) are both offering this network. RIP's administrative distance is 120\. When multiple protocols advertise the same prefix, the router installs the one with the lowest administrative distance and ignores the rest. 90 beats 110 beats 120, so EIGRP installs the route as `D`, OSPF sits in reserve, and RIP does not get a look in. The RIP route is perfectly valid; it is simply outranked, and outranked routes never make it into the table. That, in one route table, is the entire "why you would still see RIPv2, and why it almost never wins" story. RIP survives in the field for legacy interop: a device on the far end speaks it, so you configure it to talk back. But in any network that also runs OSPF or EIGRP, or really anything with a lower administrative distance, RIP's routes lose the tie-break every time. You would not choose RIP; you tolerate it at a boundary. And the moment a better-ranked protocol advertises the same destination, RIP quietly steps aside. For the full administrative-distance walkthrough across all the protocols on this box, see the [dynamic routing overview](https://www.pinglabz.com/asa-dynamic-routing-overview/). ## The authentication model is the one thing RIPv2 got right It is easy to be dismissive of RIP, but the authentication story is genuinely worth respecting, and it is the same model you should insist on wherever a routing protocol supports it. RIPv2 with MD5 means the ASA will only accept an update from a neighbour that can prove it holds the shared key. Without that, RIP is trivially spoofable: anyone on the segment could inject routes and the firewall would believe them. The `mode=3` in the failure log is the ASA confirming the neighbour tried to use MD5 at all, which is what you want to see even when the keys are wrong, because it means the far end is not falling back to no authentication. So when you do stand up a RIP interop link, turn authentication on and treat the severity-1 log as your friend. A wrong key produces a loud, specific message that names the offending neighbour and the interface. That is far easier to chase than a protocol that silently accepts anything or silently rejects everything. The failure being noisy is the whole reason a RIP key mismatch is a five-minute fix rather than an afternoon. ## So why bother learning it? Because "rare" is not "never", and the times you meet RIP are exactly the times you are already having a bad day: an acquisition brought in a network you did not design, an old industrial controller only speaks RIPv2, a partner's edge device was configured a decade ago and nobody wants to touch it. In those moments, knowing that the ASA authenticates RIP loudly (so a severity-1 log points you straight at a key mismatch) and that a RIP route will silently lose to any better-ranked protocol (so you should check what else is advertising the prefix before you conclude RIP is broken) is what turns a confusing ticket into a five-minute fix. That is the whole value of this article: not to teach you to deploy RIP, but to help you recognise and reason about it the day it lands on your firewall. ## Key Takeaways - RIPv2 on the ASA is a four-line configuration: `router rip`, a classful `network`, `version 2`, and `no auto-summary`. - A RIP authentication failure is logged loudly at severity 1 (`%ASA-1-107001`). The `mode=3` field means MD5\. A severity-1 message is the ASA treating a failed auth check as security-relevant, not routine. - This is the opposite of the silent EIGRP AS-number mismatch: RIP shouts, EIGRP says nothing. - Even with authentication fixed, the RIP route for 10.20.20.0/24 will not install, because EIGRP (distance 90) and OSPF (distance 110) already advertise it and both outrank RIP's distance of 120. - You will still see RIPv2 in the field, but almost always for legacy interop, and it almost never wins a route where a better-ranked protocol is also present. - For the full administrative-distance picture see the [dynamic routing overview](https://www.pinglabz.com/asa-dynamic-routing-overview/), and for the rest of the firewall build the [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/). ### EIGRP on the Cisco ASA: Configuration and Neighbor Gotchas URL: https://www.pinglabz.com/asa-eigrp-configuration/ Last updated: 2026-08-01T19:31:10.000Z EIGRP on the Cisco ASA is one of the most straightforward things you will configure on the platform, right up until a neighbour refuses to form. And the number-one reason an EIGRP neighbour will not come up on the ASA fails completely silently: no error, no log, nothing but an empty neighbour table. This article walks the full build, the working adjacency, and then that silent failure end to end, all captured from a live ASAv. It belongs to our [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/), and if you want EIGRP internals beyond the firewall, the [EIGRP pillar](https://www.pinglabz.com/eigrp/) covers DUAL, metrics, and feasibility in depth. As with the rest of this cluster, every capture below came from an ASAv 9.24(1) in Cisco Modeling Labs, driven over SSH, peering with an IOS-XE router called CORE1\. It is a real ASA-to-IOS EIGRP relationship, not a theoretical one. ## The configuration EIGRP on the ASA is compact. Here is the whole thing on FW1: ``` router eigrp 100 no auto-summary network 10.20.10.0 255.255.255.0 ``` The `100` after `router eigrp` is the autonomous system number, and remember that number because it is the hero of this article. As with OSPF, there is no `ip` keyword anywhere, and the `network` statement takes a real subnet mask (a classful network with an optional mask), not an IOS wildcard. The `no auto-summary` line disables classful auto-summarisation so that your subnetted routes are advertised with their real prefix lengths rather than being collapsed to a classful boundary. On any modern design you want `no auto-summary` on, because auto-summary causes routing black holes the moment you have discontiguous subnets of the same major network. ## The working adjacency and the learned route With CORE1 configured in the same AS, the neighbour forms. You verify it with `show eigrp neighbors` (no `ip`): ``` FW1# show eigrp neighbors EIGRP-IPv4 Neighbors for AS(100) H Address Interface Hold Uptime SRTT RTO Q Seq 0 10.20.10.2 inside 14 00:00:33 10 200 0 3 ``` The header tells you which AS you are looking at (`AS(100)`), and the single row is CORE1 at `10.20.10.2` off the `inside` interface. The `Q` column reading `0` means there is nothing queued waiting to be sent to that neighbour, which is what you want in steady state, a non-zero Q that never drains points at a transport problem. Once the neighbour is up, the route it advertises shows up in the table: ``` FW1# show route eigrp D 10.20.20.0 255.255.255.0 [90/128512] via 10.20.10.2, 00:00:31, inside ``` There it is: `10.20.20.0/24` learned via EIGRP, code `D`, with the administrative distance and metric in the brackets as `[90/128512]`. The 90 is EIGRP's internal administrative distance, and it is exactly why this prefix wins over the OSPF copy of the same route (OSPF's distance is 110). The 128512 is the composite EIGRP metric. If you ever need to know why the firewall preferred EIGRP for a prefix that multiple protocols advertise, that 90 is your answer, and we cover the full administrative-distance story in the [dynamic routing overview](https://www.pinglabz.com/asa-dynamic-routing-overview/). ## Reading that composite metric The `128512` in `[90/128512]` is the EIGRP composite metric, and it is worth understanding at a glance even if you rarely compute it by hand. EIGRP builds this number from the slowest bandwidth along the path and the cumulative delay, scaled by the K-values (the metric weights). Bandwidth and delay are the two components that matter by default; the other potential inputs (load, reliability, MTU) are switched off in the standard K-value set, which is exactly why a K-value mismatch, covered below, breaks an adjacency. Both neighbours must agree on which components count, or they cannot agree on a metric. The practical takeaway is that if two paths to the same prefix exist, EIGRP picks the one with the lower composite metric, and it does so using the feasible-distance and reported-distance logic that underpins the DUAL algorithm. You do not need to memorise the arithmetic to operate an ASA, but you should know that the second number in the brackets is a path-quality figure, not an administrative distance, and that the first number (90) is the administrative distance that decides protocol preference. People conflate the two constantly. On the ASA route table they sit side by side, so it is a good place to keep them straight. ## The silent failure: an AS number mismatch Here is the teaching moment, and it is the whole reason this article exists. EIGRP neighbours will only form between routers configured in the *same* autonomous system. To demonstrate what happens when they are not, we put CORE1 into EIGRP AS 200 while leaving FW1 in AS 100\. Everything else stayed correct: interfaces up, IP reachability fine, network statements right. Here is what the firewall showed: ``` FW1# show eigrp neighbors EIGRP-IPv4 Neighbors for AS(100) FW1# ``` An empty table. Just the header for AS 100 and then the prompt. And here is the part that catches even experienced engineers: there was **nothing in syslog**. No "AS mismatch" message, no warning, no hint. The firewall did not tell us anything was wrong, because from its point of view nothing was wrong. It is running EIGRP in AS 100 and simply has no neighbours in AS 100. Why is it silent? Because the autonomous system number is part of what identifies an EIGRP hello as belonging to a particular EIGRP process. CORE1's hellos carry AS 200; FW1 is listening for AS 100\. The AS 200 hellos do not match the AS 100 process, so the firewall discards them the same way it would discard traffic for a protocol it is not running. There is no negotiation and therefore no mismatch to report. Contrast this with the RIP authentication failure elsewhere in this cluster, which the ASA logs loudly at severity 1, or an OSPF area mismatch, which at least leaves the neighbour cycling through states you can watch. An EIGRP AS mismatch produces none of that. It looks exactly like "no neighbour is out there". The practical lesson: when an EIGRP neighbour will not form and everything else looks fine, check the AS number on *both* ends before you do anything else. Do not trust that the far end matches just because your side is correct. The moment we reconfigured CORE1 back to AS 100, the neighbour came up within seconds and the `D` route reappeared. No reload, no clearing, no drama, the adjacency was ready the instant the AS numbers agreed. ## Why the AS number silence is worse than it sounds It is tempting to shrug this off. You configured the AS, surely you know what it is. But consider how these mistakes actually happen. Someone templated the firewall config from a different site that used AS 200\. Someone fat-fingered `100` as `200` in a change window at 2 a.m. Someone stood up a new core switch and used the AS number from the design doc, which was stale. In every one of those cases, the person looking at the firewall sees a perfectly valid EIGRP process with an empty neighbour table and no errors, and their instinct is to suspect a layer-1, layer-2, or ACL problem, because those are the things that usually produce a silent dead neighbour. They can burn an hour checking cabling and interface counters before anyone thinks to diff the AS number against the peer. Knowing that this specific failure is silent is what saves you that hour. ## Everything that stops an ASA EIGRP neighbour forming The AS mismatch is the sneakiest, but it is not the only thing that blocks an adjacency. Here is the field checklist, with the silent one flagged because it deserves the warning: AS number mismatch Symptom: empty neighbour table Logged? **No, completely silent** Fix: match the AS on both ends K-value mismatch Symptom: neighbour will not stay up Logged? Yes, a K-value mismatch message Fix: align the metric weights on both ends Authentication mismatch Symptom: no adjacency Logged? Yes, auth failure messages Fix: match key ID, key string, and mode Interface not in a network statement Symptom: no hellos sent on that link Logged? No Fix: add a `network` statement covering the interface subnet Passive interface Symptom: interface advertised but no neighbour Logged? No Fix: remove `passive-interface` on the link that needs the adjacency No IP reachability / ACL Symptom: hellos never arrive Logged? No Fix: confirm the two neighbour IPs can reach each other on the segment Notice how many of these are silent. That is the theme with EIGRP adjacency troubleshooting: only a couple of failure modes announce themselves, so you cannot rely on syslog to point you at the problem. You have to work the checklist. Start with the two things that are both common and silent, the AS number and whether the interface is actually in a network statement, and only then move to the noisier candidates. That loud-versus-silent split is not an ASA quirk; it is how EIGRP behaves everywhere, so the same ordering works on the IOS side of the link. [Seeing which EIGRP adjacency failures announce themselves and which do not](https://www.pinglabz.com/troubleshooting-eigrp-neighbor-adjacencies/) breaks each cause deliberately on a router and shows exactly what syslog says, or does not say. ## A verification workflow that finds the silent ones fast Because so many EIGRP failures on the ASA are silent, a consistent verification order beats poking around. Start with `show eigrp neighbors` and read the AS number in the header against what the peer is configured for; if the table is empty and the AS numbers differ, you are done. If the AS matches but the table is still empty, confirm the interface is actually enabled for EIGRP by checking that a `network` statement covers its subnet, and that the interface is not passive. Only after those two silent suspects are cleared should you look for the noisy failures (authentication and K-value mismatches do log, so a clean syslog effectively rules them out). Finally, confirm plain IP reachability between the two neighbour addresses on the segment, because EIGRP hellos are just packets and a broken layer-2 path or an ACL will stop them like anything else. This ordering is deliberately front-loaded with the silent, common problems. It is the opposite of how people instinctively troubleshoot, which is to check the loud, dramatic things first. On EIGRP the loud things are the rare ones. Trust the checklist over your gut, and you will find the AS-number typo in thirty seconds instead of an hour. ## A note on passive interfaces and firewall security The passive-interface item in that grid is worth calling out on a firewall specifically. Making an interface passive means the ASA still advertises that interface's subnet into EIGRP but stops sending hellos out of it, so no adjacency forms there. On a firewall you often *want* that on interfaces facing untrusted networks: you would never want the ASA forming an EIGRP relationship with something out on the internet-facing side. So passive-interface is both a troubleshooting suspect (you removed an adjacency you needed) and a security control (you removed an adjacency you did not want). Configure it deliberately, and document which interfaces are passive so the next person does not spend an afternoon wondering why a neighbour will not form on a link you intentionally silenced. ## Key Takeaways - EIGRP on the ASA is compact: `router eigrp `, `no auto-summary`, and a `network` statement with a real subnet mask and no `ip` keyword. - A working neighbour shows in `show eigrp neighbors`, and the learned route appears as `D` with administrative distance 90, which is why EIGRP beats OSPF for the same prefix. - The number-one adjacency killer, an AS number mismatch, is completely silent: an empty neighbour table and nothing in syslog. Always check the AS on both ends. - Matching the AS brought the neighbour up within seconds with no reload or clearing. - Most EIGRP adjacency failures on the ASA are silent. Work a checklist rather than waiting for a log message: AS number, network statement coverage, passive interface, authentication, K-values, and basic IP reachability. - Passive-interface is both a troubleshooting suspect and a legitimate security control on a firewall. Use it deliberately on untrusted-facing links. - For EIGRP internals see the [EIGRP pillar](https://www.pinglabz.com/eigrp/); for the rest of the firewall build, the [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/). ### OSPF on the Cisco ASA: Areas, Redistribution, and the ASA's Quirks URL: https://www.pinglabz.com/asa-ospf-configuration/ Last updated: 2026-07-13T10:09:25.000Z OSPF is the dynamic routing protocol you are most likely to meet on a Cisco ASA. When a firewall needs to learn more than a handful of internal prefixes and the core already speaks OSPF, running it on the ASA is a defensible choice. This article is the full build: configuration, a verified adjacency with MD5 authentication, the DR/BDR election the firewall loses, the link-state database, and redistribution, all captured from a live ASAv. It sits in our [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/), and if you want OSPF theory beyond the firewall specifics, the [OSPF pillar](https://www.pinglabz.com/ospf/) covers areas, LSA types, and metrics in depth. One honesty note up front, because credibility matters: every capture below came from an ASAv 9.24(1) running in Cisco Modeling Labs, driven over SSH from a real Linux host. The peer, CORE1, is an IOS-XE router. This is a real adjacency between an ASA and an IOS router with cryptographic authentication turned on, not a diagram. ## The configuration, and the two things IOS people get wrong Here is the OSPF configuration on our firewall, FW1\. Read it, then note the two differences from IOS that catch everyone: ``` interface GigabitEthernet0/1 ospf message-digest-key 1 md5 router ospf 1 router-id 10.20.10.1 network 10.20.10.0 255.255.255.0 area 0 area 0 authentication message-digest redistribute static subnets ``` First, there is no `ip` keyword. On IOS the interface command is `ip ospf message-digest-key`; on the ASA it is simply `ospf message-digest-key`. Second, the `network` statement takes a real subnet mask, `255.255.255.0`, not an IOS wildcard of `0.0.0.255`. Type a wildcard here and you will either be rejected or match nothing, and the adjacency will never form. For the record, the matching side on CORE1 (which is IOS) uses `ip ospf message-digest-key` and `area 0 authentication message-digest`, so the two boxes meet in the middle: same MD5 key, same key ID, same area authentication mode, different command prefixes. The `area 0 authentication message-digest` line turns on MD5 authentication for the whole area, and the interface `message-digest-key 1 md5` line supplies the actual key with key ID 1\. Both ends must agree on the key ID and the key string, or the neighbours will see each other's hellos and refuse to become adjacent. ## Proving the adjacency and the authentication With the config in place, the neighbour comes up. On the ASA you check it with `show ospf neighbor` (again, no `ip`): ``` FW1# show ospf neighbor Neighbor ID Pri State Dead Time Address Interface 10.255.0.1 1 FULL/DR 0:00:30 10.20.10.2 inside ``` The state is `FULL`, which is the goal. Anything less (INIT, 2-WAY, EXSTART, EXCHANGE, LOADING) means the adjacency is stuck partway and you have work to do. The `/DR` tells you the neighbour is the Designated Router on this segment, which we will come back to. But the single most useful command for confirming that authentication is actually working is `show ospf interface`: ``` FW1# show ospf interface inside inside is up, line protocol is up Internet Address 10.20.10.1 mask 255.255.255.0, Area 0 Process ID 1, Router ID 10.20.10.1, Network Type BROADCAST, Cost: 10 Designated Router (ID) 10.255.0.1, Interface address 10.20.10.2 Backup Designated router (ID) 10.20.10.1, Interface address 10.20.10.1 Timer intervals configured, Hello 10, Dead 40, Wait 40, Retransmit 5 Supports Link-local Signaling (LLS) Cisco NSF helper support enabled IETF NSF helper support enabled Neighbor Count is 1, Adjacent neighbor count is 1 Cryptographic authentication enabled Youngest key id is 1 ``` The line that matters is near the bottom: `Cryptographic authentication enabled`, with `Youngest key id is 1`. That is the ASA confirming MD5 is active on this interface and telling you which key it is using. If you configure authentication and this line does not appear, the key never took, and you should assume the adjacency (if it even formed) is running unauthenticated. It is a small line that saves you from shipping a firewall you think is protected when it is not. ## The DR/BDR election the ASA loses, and why Look again at the `show ospf interface` output. The Designated Router is ID `10.255.0.1` (that is CORE1), and the Backup Designated Router is `10.20.10.1` (that is our firewall). The ASA lost the election and became the BDR. This is a genuinely useful thing to understand, so let us walk it. On a broadcast segment, OSPF elects a DR and a BDR to reduce the number of adjacencies. The election is decided first by interface priority (higher wins), and only if priorities are tied does it fall to the highest router ID. Look at the neighbour output: both routers show priority `1`, the default. Since neither box was configured with a higher OSPF priority, the priorities tie, and the tie-breaker is the router ID. CORE1's RID is `10.255.0.1`; the firewall's is `10.20.10.1`. As dotted-decimal values, `10.255.0.1` is higher than `10.20.10.1`, so CORE1 wins the DR role and the ASA settles for BDR. Why care? Because if you want the ASA to be the DR (or emphatically not the DR), priority is your lever, not router ID, and you set it per interface. And because a surprising number of "why is the wrong box the DR" tickets come down to exactly this: everyone left priority at the default, so the election silently fell through to router ID, and the box with the highest-numbered loopback or interface address won by accident. On a firewall specifically, you rarely want it carrying the DR workload of maintaining adjacencies with every router on a segment, so losing the election here is arguably the right outcome. ## The ASA is a full OSPF participant, not a bystander It is tempting to think of a firewall as a passive listener that just absorbs routes. It is not. The ASA originates its own Router LSA and participates fully in flooding. Here is the link-state database: ``` FW1# show ospf database OSPF Router with ID (10.20.10.1) (Process ID 1) Router Link States (Area 0) Link ID ADV Router Age Seq# Checksum Link count 10.20.10.1 10.20.10.1 69 0x80000002 0x9ee4 1 10.255.0.1 10.255.0.1 70 0x80000005 0x5bff 3 Net Link States (Area 0) Link ID ADV Router Age Seq# Checksum 10.20.10.2 10.255.0.1 70 0x80000001 0x8947 ``` Two Router LSAs are in the database, one advertised by each router. The row where both Link ID and ADV Router read `10.20.10.1` is the ASA's own Router LSA: the firewall is describing its links to the rest of the area, exactly like any router would. The Net Link State (a Network LSA) is advertised by `10.255.0.1`, which is correct, because the DR originates the Network LSA for a broadcast segment, and CORE1 is the DR. So the database is internally consistent with the election we just walked through: CORE1 is DR, so CORE1 owns the Network LSA, and both boxes contribute their own Router LSAs. ## A word on areas Our lab keeps everything in area 0, the backbone, which is the right call for a single firewall attached to a single OSPF domain. The ASA supports the full area model: you can make it an Area Border Router by placing different interfaces in different areas, and it honours stub, totally stubby, and not-so-stubby (NSSA) area types the same way IOS does. On a firewall the most common non-backbone design is to place the segment behind the firewall in a stub or NSSA area so that external LSAs from the rest of the network do not flood into it, which keeps the firewall's database lean and reduces its exposure to churn elsewhere. If you do run the ASA as an ABR, remember that area authentication is configured per area (`area 0 authentication message-digest` only covers area 0), so each area you add needs its own authentication statement. ## Redistributing static routes into OSPF A firewall almost always has static routes (the default toward the ISP, routes into segments that do not run a protocol), and you frequently want those visible to the rest of the OSPF domain. The `redistribute static subnets` line does that. The `subnets` keyword matters: without it, classful behaviour kicks in and only classful networks get redistributed, silently dropping your subnetted routes. You can confirm what is actually configured with: ``` FW1# show run all router ospf | include redistribute no redistribute connected redistribute static subnets ``` This is a nice confirmation because `show run all` prints the defaults too, so you can see explicitly that connected redistribution is off (`no redistribute connected`) while static redistribution is on with `subnets`. When you redistribute, remember the routes enter OSPF as external (type E2 by default), which is a different LSA type and a different metric behaviour than internal routes. If a redistributed route is not showing up on the far side, checking whether `subnets` is present is the first thing to rule out. ## When the adjacency will not come up Most OSPF-on-ASA tickets are one of a small set of mismatches, and the `show ospf interface` output above hands you the values to check. Look at the timer line: `Hello 10, Dead 40`. Both neighbours must agree on the hello and dead intervals, or they will never progress past the two-way state. On a broadcast network these default to 10 and 40, so a mismatch usually means someone tuned one side and not the other. The `Network Type BROADCAST` field is the next thing to confirm. If one side thinks the segment is broadcast and the other point-to-point, the DR/BDR expectations differ and the adjacency can hang in EXSTART. Then there is authentication: if you see the neighbour flapping or never leaving INIT after you turned on MD5, the key ID or the key string does not match, and the fix is to line them up on both ends. Finally, MTU. OSPF exchanges the interface MTU in the database description packets, and if the two interfaces disagree the adjacency parks in EXSTART or EXCHANGE indefinitely. It is a classic firewall-to-router gotcha because the two platforms sometimes default to different MTUs. A quick mental checklist when an ASA OSPF neighbour is stuck: same area number, same hello and dead timers, same network type, same authentication key ID and string, same MTU, and the interface actually covered by a `network` statement. Run `show ospf neighbor` and read the state field. The state tells you roughly where in the process it is failing (INIT and two-way point at hello or authentication problems; EXSTART and EXCHANGE point at MTU or network-type problems), which is far faster than guessing. ## OSPF on the ASA versus OSPF on IOS Most of what you know about OSPF carries over. But the differences are exactly the things that waste an afternoon, so here they are side by side: Viewing the route table ASA: `show route` IOS: `show ip route` Network statement ASA: real mask `255.255.255.0` IOS: wildcard `0.0.0.255` Interface auth command ASA: `ospf message-digest-key` IOS: `ip ospf message-digest-key` Passive interface default ASA: not passive by default; interfaces in a network statement form adjacencies IOS: same behaviour, but firewall admins often forget the outside interface is fair game Graceful restart ASA: NSF helper enabled, seen in `show ospf interface` IOS: NSF/GR supported, platform dependent Interface names ASA: nameif labels, e.g. `inside` IOS: physical, e.g. GigabitEthernet0/1 The passive-interface point deserves a word. The ASA does not passive its interfaces for you. Any interface whose subnet is covered by a `network` statement will try to form adjacencies. On a firewall that is a security consideration: you almost never want OSPF hellos leaving the outside interface toward the internet, so scope your network statements tightly and use `passive-interface` on anything facing untrusted networks. This is the same discipline covered in [ASA management hardening](https://www.pinglabz.com/cisco-asa-management-hardening/), applied to the control plane. ## Key Takeaways - OSPF is the most common dynamic protocol on an ASA. The build is close to IOS but drops the `ip` keyword and uses a real subnet mask in the `network` statement. - `show ospf interface` is your authentication check. The line `Cryptographic authentication enabled` with a key ID confirms MD5 is genuinely active, not just configured. - The ASA lost the DR election to CORE1 because priorities tied at the default of 1, so the tie broke on router ID, and `10.255.0.1` outranks `10.20.10.1`. Use interface priority, not router ID, to control the outcome. - The ASA is a full participant: it originates its own Router LSA, visible in `show ospf database`. The DR (CORE1) owns the Network LSA. - Use `redistribute static subnets` and keep the `subnets` keyword, or classful behaviour silently drops your subnetted statics. - The ASA does not passive its interfaces for you. Scope network statements tightly and passive anything facing untrusted networks. - For the wider protocol, see the [OSPF pillar](https://www.pinglabz.com/ospf/); for the rest of the firewall build, the [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/). ### Dynamic Routing on the Cisco ASA: What the Firewall Will and Won't Do URL: https://www.pinglabz.com/asa-dynamic-routing-overview/ Last updated: 2026-07-13T10:09:24.000Z The Cisco ASA is a firewall that happens to route. That framing matters, because when you sit down at an ASA after years on IOS routers, the routing looks familiar right up until it doesn't. The commands are subtly different, a few defaults are inverted, and the box is opinionated about which routes it will actually install. This article is the orientation stop for the routing side of our [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/): what the firewall will and won't do when you ask it to run a dynamic routing protocol, and why. The ASA supports OSPF, EIGRP, RIP, and BGP. All four are real, all four are testable, and every capture in this article came from a live ASAv 9.24(1) running in Cisco Modeling Labs, driven over SSH from a real Linux host. If you have already worked through [static routing on the ASA](https://www.pinglabz.com/cisco-asa-static-routing/), this is where the box starts learning routes on its own. ## The first thing that trips up IOS people: show route On an IOS router you type `show ip route` without thinking. On the ASA, that command does not exist. The ASA uses `show route`. There is no `ip` keyword because the ASA route table is not split the way an IOS RIB is presented, it is simply the route table. Here is the real table from our lab firewall, FW1: ``` FW1# show route Gateway of last resort is 203.0.113.2 to network 0.0.0.0 S* 0.0.0.0 0.0.0.0 [1/0] via 203.0.113.2, outside C 10.20.10.0 255.255.255.0 is directly connected, inside L 10.20.10.1 255.255.255.255 is directly connected, inside D 10.20.20.0 255.255.255.0 [90/128512] via 10.20.10.2, 00:08:55, inside O 10.20.20.1 255.255.255.255 [110/11] via 10.20.10.2, 00:12:31, inside O 10.255.0.1 255.255.255.255 [110/11] via 10.20.10.2, 00:12:31, inside C 172.20.30.0 255.255.255.0 is directly connected, dmz L 172.20.30.1 255.255.255.255 is directly connected, dmz C 203.0.113.0 255.255.255.0 is directly connected, outside L 203.0.113.10 255.255.255.255 is directly connected, outside ``` The route codes are the same ones you already know: `C` for connected, `L` for the local host route (the interface address as a /32), `S` for static (the `S*` is the gateway of last resort), `O` for OSPF, and `D` for EIGRP. The ASA adds a couple of its own that IOS does not use in the same way (`V` for VPN-learned routes, and `SI` / `BI` for inter-VRF static and BGP). Interfaces are named, not numbered, so the next-hop line reads `via 10.20.10.2, inside` rather than pointing at GigabitEthernet0/1\. On the ASA, "inside" is the interface. Two details in that table are worth pausing on because they confuse people coming from IOS. First, every connected network produces two lines: the network itself as `C`, and the firewall's own interface address as an `L` host route with a /32 mask. That is not clutter, it is how the box knows a packet destined for its own interface is for it and not to be forwarded. Second, the `S*` at the top is your default route, and the phrase "Gateway of last resort is 203.0.113.2" is the ASA telling you where anything it does not have a more specific route for will go. On this firewall that next hop lives out the `outside` interface toward the ISP, exactly as you would expect on an edge box. ## The second surprise: network statements take a real mask On IOS, `network 10.20.10.0 0.0.0.255 area 0` uses a wildcard mask. On the ASA, the OSPF and EIGRP `network` statement takes a real subnet mask, and again there is no `ip` keyword. Here is the minimal OSPF configuration from FW1: ``` router ospf 1 router-id 10.20.10.1 network 10.20.10.0 255.255.255.0 area 0 area 0 authentication message-digest redistribute static subnets ``` If your muscle memory types a wildcard there, the ASA will either reject it or match the wrong set of interfaces, and you will spend twenty minutes wondering why the adjacency never comes up. Real mask, no wildcard, no `ip`. Write that on a sticky note before you touch OSPF or EIGRP on this platform. ## The centrepiece: administrative distance, proven on one box Look back at the route table above. The network `10.20.20.0/24` is installed with the code `D` and a metric of `[90/128512]`. The 90 is administrative distance, and it is EIGRP's. But that same prefix was also being advertised into this firewall by OSPF at the same time. So why did EIGRP win? Because administrative distance is the tie-breaker the router uses when two protocols both offer a path to the same prefix. Lower distance wins. EIGRP internal routes carry an AD of 90; OSPF carries 110\. When both protocols hand the firewall a route to 10.20.20.0/24, the ASA installs exactly one of them, and it picks the lower distance. EIGRP's 90 beats OSPF's 110, so the prefix installs as `D`, not `O`. The OSPF version does not vanish, it simply sits in the OSPF database as a backup that never makes it into the route table while EIGRP is healthy. This is worth internalising because it is the same ordering that decides a story we tell later in the cluster. RIP's administrative distance is 120\. On a firewall that is already learning 10.20.20.0/24 via EIGRP (90) and OSPF (110), a RIPv2 route to that same prefix at distance 120 never installs at all, no matter how correctly you configure RIP. It is not broken, it is simply outranked. That is most of the reason you rarely see RIP win a route in a mixed network, and we prove it end to end in the [RIPv2 on the ASA](https://www.pinglabz.com/asa-ripv2-configuration/) article. Connected / Local Distance: **0 / 0** Codes: C, L You use it for: always on, never configured Static Distance: **1** Code: S You use it for: the default route to the ISP, most edge firewalls EIGRP Distance: **90** internal Code: D You use it for: Cisco-only cores that already run EIGRP OSPF Distance: **110** Code: O You use it for: the most common dynamic protocol on an ASA, multi-vendor cores RIPv2 Distance: **120** Code: R You use it for: legacy interop only, almost never wins a route BGP Distance: **20 ext / 200 int** Code: B You use it for: data-centre edge, cloud transit, route exchange with a provider ## Should a firewall run a routing protocol at all? Here is the honest part. Running a dynamic routing protocol on a firewall is, more often than not, a design smell. A firewall's job is to enforce policy at a boundary, and every routing adjacency you form is another trust relationship and another moving part on a device whose whole reason for existing is to be predictable. The cleaner design for most edge firewalls is a default route out to the ISP, a handful of static routes toward internal networks, and a properly designed routed core (running OSPF, EIGRP, or BGP among the routers and layer-3 switches) that hides its complexity behind a single summarised next hop. That said, "often a smell" is not "never". There are entirely legitimate reasons an ASA speaks a routing protocol: it sits inline in a campus core that already runs OSPF and needs to learn dozens of internal prefixes without a static-route sprawl; it peers BGP with a cloud provider or an upstream carrier; it participates in a redundant path where you want fast reconvergence rather than tracking a static route by hand. And on the exam side, CCNP and CCIE Security both expect you to configure and troubleshoot these protocols on the ASA, silent failures and all. So the correct instinct is: prefer static plus a clean routed core, but know how to run the dynamic protocols cold, because you will meet them. Of the four, OSPF is by far the most common thing you will see on a real ASA, which is why it gets the longest treatment in this cluster. EIGRP shows up in Cisco-only shops that standardised on it years ago. BGP is the specialist case: it is genuinely useful at a data-centre or cloud edge, but it is also the most configuration-heavy of the four and deserves its own build rather than a paragraph here. RIP, frankly, you keep in your back pocket for the day a legacy device on the far end speaks nothing else. The point of this overview is that all four sit on the same route table, obey the same administrative-distance rules, and are reached with the same `show route` command, so once you understand the model, each protocol is just a different way of populating it. ## Where to go next in this cluster This overview is the map. Each protocol gets its own deep-dive with the full build, the verification output, and the specific way it bites you on this platform: [OSPF on the Cisco ASA](https://www.pinglabz.com/asa-ospf-configuration/) Areas, MD5 authentication, the DR/BDR election the ASA loses, the LSDB, and redistribution. The most common dynamic protocol you will meet on an ASA. [EIGRP on the Cisco ASA](https://www.pinglabz.com/asa-eigrp-configuration/) Straightforward to configure, until a neighbour will not form and the number-one cause fails completely silently. [RIPv2 on the Cisco ASA](https://www.pinglabz.com/asa-ripv2-configuration/) Legacy, but you will still meet it for interop. A severity-1 auth log and a route that refuses to install, both proven in the lab. [Static routing on the Cisco ASA](https://www.pinglabz.com/cisco-asa-static-routing/) The default route, floating statics, and why static plus a clean core is usually the right answer for an edge firewall. ## Key Takeaways - The ASA is a firewall first and a router second. It runs OSPF, EIGRP, RIP, and BGP, but with defaults and syntax that differ from IOS. - Use `show route`, not `show ip route`. There is no `ip` keyword on the ASA route commands. - OSPF and EIGRP `network` statements take a real subnet mask, not an IOS-style wildcard. - Administrative distance decides which protocol installs a prefix. On our one box, EIGRP (90) beat OSPF (110) for 10.20.20.0/24, so it installed as `D`. The same ordering is why RIP (120) never wins. - Dynamic routing on a firewall is often a design smell. Prefer a default plus static routes and a clean routed core, but know the protocols cold because they are real and testable. - Start with the [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/) pillar, then work through the OSPF, EIGRP, and RIPv2 deep-dives linked above. ### GETVPN Redundancy: COOP Key Servers and Rekey Survival URL: https://www.pinglabz.com/getvpn-coop-key-servers/ Last updated: 2026-07-13T09:01:24.000Z The first question anyone sensible asks about GETVPN is: "so the whole thing depends on one router?" It is a fair question. The key server holds the group policy, generates the keys, and pushes them out. Lose it and, on the face of it, you have lost the VPN. The answer is more interesting than it looks. Half the worry is misplaced (the data plane genuinely does not care if the key server dies, and we shut one down to prove it), and the half that is real is fixed with COOP: cooperative key servers, two or more, with priorities and an election. But COOP has a trap that almost everyone hits, that no `show` command flags as a failure, and that will silently kill your entire group at the next rekey interval. We walked into it deliberately and captured exactly what the router says. This is part three of the GETVPN series, following [GETVPN Explained](https://www.pinglabz.com/getvpn-explained/) (the model and the header-preservation capture) and the [GETVPN configuration walkthrough](https://www.pinglabz.com/getvpn-configuration-ios-xe/). If GETVPN itself is new to you, start there. Both sit under the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/). Everything below is real output from Cisco IOS XE (cat8000v 17.18.02): KS1 at 10.100.0.11, KS2 at 10.100.0.12, group members GM1 and GM2, on a shared private WAN core. ## Start Here: The Data Plane Survives a Dead Key Server Before configuring any redundancy at all, it is worth understanding precisely what you lose when a key server dies, because most people badly overestimate it. We shut down KS1's core-facing interface. The key server was gone. Then, on a group member, with no key server reachable: ``` GM1#ping 10.30.10.1 source Loopback10 repeat 5 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 6/9/20 ms ``` One hundred percent. Encrypted. With the key server dead. **The key server is a control plane element. It is not in the data path.** This is the single most misunderstood fact about GETVPN, and it falls directly out of the architecture: group members do not tunnel *to* the key server, they register *with* it. Once a member holds the TEK, it encrypts to every other member directly across the WAN. The key server is not carrying a single byte of user traffic. So what do you actually lose when the key server goes away? Keeps working Encryption and decryption of user traffic, using the TEK members already hold Any-to-any reachability between every existing member The current group policy ACL, already downloaded Stops working Rekeys, so the group runs on a clock that eventually expires Re-registration, so a member that reloads cannot rejoin Policy changes, so you cannot add a site or change the ACL That is a degradation on a timer, not an outage. The group keeps encrypting until the TEK lifetime runs out with no rekey to replace it. Your GETVPN does not fall over the instant the key server reboots, which changes how you should think about the risk: it is a control-plane availability problem, not a data-plane one. Which is exactly why COOP exists. Not to keep traffic flowing (that already works), but to keep the *control plane* alive so the group can rekey, re-register, and grow. ## Configuring COOP COOP adds a `redundancy` sub-block under `server local` in the GDOI group. Each key server gets a local priority and a list of its COOP peers. Highest priority wins the election and becomes Primary. On KS1 (priority 100), pointing at KS2: ``` crypto gdoi group PLZ-GETVPN server local redundancy local priority 100 peer address ipv4 10.100.0.12 ``` On KS2 (priority 75), pointing at KS1, and otherwise carrying an identical group configuration (same identity number, same policy ACL, same transform set, same rekey parameters). Group members already listed both servers, which is why the member config in the previous article had two `server address ipv4` lines: ``` crypto gdoi group PLZ-GETVPN identity number 4321 server address ipv4 10.100.0.11 server address ipv4 10.100.0.12 ``` Members try the servers in list order. Registration is not tied to the Primary: a member can register to a Secondary perfectly happily. Only *rekeys* come from the Primary. ### The ordering gotcha: address before redundancy Type the COOP block on a key server that does not yet have its local GDOI address configured and IOS stops you flat: ``` KS2(config-gdoi-local-server-red)#local priority 75 % ERROR: Local server address must be configured first ``` **`address ipv4` must be configured before `redundancy`.** This makes sense once you see it (COOP peering sources itself from the key server's local GDOI address, so there is nothing to peer from until that address exists), but it catches people because it is the reverse of how the config reads back afterwards, and because a copy-paste of a working config from another box goes in line by line in the order you paste it. Put `address ipv4 10.100.0.11` in first, then add the redundancy block. ## The Election, in Real Output With both key servers up, `show crypto gdoi ks coop` is where you confirm the relationship. On KS1, the higher priority: ``` KS1#show crypto gdoi ks coop Crypto Gdoi Group Name :PLZ-GETVPN Local Address: 10.100.0.11 Local Priority: 100 Local KS Role: Primary , Local KS Status: Alive Local KS version: 1.0.27 Primary Timers: Primary Refresh Policy Time: 20 Remaining Time: 14 Peer Sessions: Session 1: Peer Address: 10.100.0.12 Peer KS Role: Secondary , Peer KS Status: Dead IKE status: Established ``` And on KS2, the lower priority, which is the more informative view because it sees the whole picture: ``` KS2#show crypto gdoi ks coop Local Address: 10.100.0.12 Local Priority: 75 Local KS Role: Secondary , Local KS Status: Alive Peer Address: 10.100.0.11 Peer Version: 1.0.27 Peer Priority: 100 Peer KS Role: Primary , Peer KS Status: Alive IKE status: Established ``` Priority 100 beats priority 75, so KS1 is Primary and KS2 is Secondary. `IKE status: Established` confirms the two key servers have an IKE SA between them (COOP peering is itself protected, which is why both key servers need ISAKMP keys for each other). Note the `Primary Refresh Policy Time: 20` on KS1\. That is the announcement interval. The Primary periodically reasserts itself, and the Secondary's failure to hear those announcements is what triggers an election. ## Failover: KS2 Promotes Itself We shut KS1's core interface. Here is the real syslog sequence from KS2, in order: ``` KS2#show logging %GDOI-5-COOP_KS_ADD: 10.100.0.11 added as COOP Key Server in group PLZ-GETVPN. %GDOI-4-COOP_KS_UNAUTH: Contact from unauthorized KS 10.100.0.11 in group PLZ-GETVPN at local address 10.100.0.12 (Possible MISCONFIG of peer/local address) %GDOI-5-COOP_KS_ELECTION: KS entering election mode in group PLZ-GETVPN (Previous Primary = NONE) %GDOI-5-COOP_KS_TRANS_TO_PRI: KS 10.100.0.11 in group PLZ-GETVPN transitioned to Primary (Previous Primary = NONE) %GDOI-3-COOP_KS_UNREACH: Cooperative KS 10.100.0.11 Unreachable in group PLZ-GETVPN. IKE SA Status = Failed to establish. %GDOI-5-COOP_KS_TRANS_TO_PRI: KS 10.100.0.12 in group PLZ-GETVPN transitioned to Primary (Previous Primary = 10.100.0.11) ``` Walk it in order: - `COOP_KS_ADD` is KS2 discovering its peer. - `COOP_KS_UNAUTH` is the trap. Park it, it gets its own section below, and it is the most important line on this page. - `COOP_KS_ELECTION` and the first `COOP_KS_TRANS_TO_PRI` are the initial election, correctly resolving in favour of KS1 (priority 100). - `COOP_KS_UNREACH` is KS1 dying. `IKE SA Status = Failed to establish`, because the interface is down. - The second `COOP_KS_TRANS_TO_PRI` is the payoff: **KS 10.100.0.12 transitioned to Primary (Previous Primary = 10.100.0.11)**. The priority-75 secondary promoted itself. And the post-failover state confirms it: ``` KS2#show crypto gdoi ks coop Local Address: 10.100.0.12 Local Priority: 75 Local KS Role: Primary , Local KS Status: Alive Peer Address: 10.100.0.11 Peer KS Role: Primary , Peer KS Status: Dead ``` `Local KS Role: Primary` on the box with priority 75\. Note the slightly unsettling `Peer KS Role: Primary, Peer KS Status: Dead`: KS2 has not forgotten that KS1 *was* Primary, it has simply marked it dead and taken over. When KS1 comes back with priority 100, it will reclaim the Primary role. COOP is preemptive, which is normally what you want (your best-connected key server should hold the role) but it does mean a flapping key server flaps the Primary role with it. ## The Trap: COOP Key Servers Must Share the Same RSA Rekey Key This is the section to read twice. It is the failure mode that separates a GETVPN deployment that survives its first rekey from one that dies quietly at 3am, and it is the reason we built this lab wrong on purpose. We gave KS2 its **own** RSA rekey keypair, generated locally, instead of exporting KS1's and importing it. Everything looked fine. COOP peered. The election ran. The failover worked. Every `show` command above was captured on that setup. Nothing in `show crypto gdoi ks coop` says "your keys are wrong." But the router did tell us, once, in the log: ``` %GDOI-4-COOP_KS_UNAUTH: Contact from unauthorized KS 10.100.0.11 in group PLZ-GETVPN at local address 10.100.0.12 (Possible MISCONFIG of peer/local address) ``` **That line is the tell.** Note how misleading the parenthetical hint is: "Possible MISCONFIG of peer/local address" sends you off checking IP addresses, which are perfectly correct. The actual cause is that the two key servers do not share an RSA rekey keypair, so they cannot authenticate each other's rekey authority. ### Why it matters (and why no show command catches it) Here is the mechanism, and once you see it the whole thing is obvious. When a group member registers, it receives the key server's **public key** as part of the group policy. From then on, every rekey message the member receives is *signed*, and the member *verifies that signature against the public key it holds*. That is the entire anti-spoofing story for rekeys: without it, anyone who could reach a group member could push them a forged key. Now suppose KS2 has its own, different RSA keypair. Consider the sequence: 1Members register to KS1 and store **KS1's** public key. 2KS1 fails. COOP works exactly as designed and KS2 becomes Primary. Everything looks healthy. 3Traffic keeps flowing, because the data plane does not need a key server. Still healthy. 4The rekey interval arrives. KS2, now Primary, sends a rekey signed with **KS2's** key. 5Every group member verifies that signature against **KS1's** public key. It does not verify. 6**Every member rejects the rekey.** The TEK expires with nothing to replace it. The group dies. The cruelty of this failure is the timing. The mistake is made at build time. The COOP election works. The failover test passes. Someone signs off the change. And the group dies *at the next rekey interval after a key server failover*, which could be weeks later, in the middle of the night, with the on-call engineer looking at a perfectly healthy-looking `show crypto gdoi ks coop` that says Primary, Alive. And the group members are doing exactly the right thing. From their point of view, an unrecognised key server is signing rekeys, which is precisely the attack the signature exists to stop. The security mechanism is working. You just misconfigured which key it is protecting. ### The fix: export from KS1, import on KS2 Both key servers must hold the *same* RSA rekey keypair. On KS1, where the key was generated (with `exportable`, which you must have specified at generation time, because you cannot add it afterwards): ``` crypto key export rsa PLZ-REKEY-KEY pem terminal 3des ``` That prints the PEM-encoded public and private keys to the terminal, with the private key encrypted under 3DES and your passphrase. Copy the whole block, then import it on KS2 with the matching `crypto key import rsa` command and the same passphrase and label. Both key servers now sign with the same key, so a rekey from either one verifies against the public key every member holds. Point `rekey authentication mypubkey rsa PLZ-REKEY-KEY` at that shared label on both. Two practical notes. First, this is why the `exportable` keyword in the key generation step of the [configuration article](https://www.pinglabz.com/getvpn-configuration-ios-xe/) is not optional: forget it and your only remedy is to regenerate the key, which forces every member to re-register. Second, you are transporting a private key over a terminal session. Use SSH, use a strong passphrase, and clear your scrollback. ### The verification nobody does Because no `show` command reports "COOP key mismatch," the only reliable proof is behavioural. Do this in a maintenance window, on a lab, before you do it in production: - Set a *short* rekey lifetime (ours was `rekey lifetime seconds 900`, deliberately short so the lab rekeys while you watch). - Fail the primary key server. - Wait past one full rekey interval with the secondary as Primary. - Check `Rekeys received` in `show crypto gdoi` on a member. It must increment. - Scan the key server logs for `COOP_KS_UNAUTH`. If it is there, your keys do not match, no matter how healthy the COOP state looks. A failover test that only checks "did the secondary become Primary" and "does the ping still work" will pass on a completely broken deployment. Both of those things pass on ours, and ours was wrong. ## Design Notes for Production - **Two key servers is the sensible baseline**, and IOS supports more. Put them in different sites, on different power, and reachable from every member. - **Priority sets the winner, and preemption is on.** Highest priority takes Primary and reclaims it when it returns. Give your best-connected key server the highest number. - **Members should list every key server** in their `server address ipv4` lines. They try them in order and can register to a Secondary quite happily. - **Key servers need ISAKMP keys for each other**, not just for the members. COOP peering runs over an IKE SA (`IKE status: Established`). - **Keep the group config identical across key servers.** Same identity number, same policy ACL, same transform set, same rekey parameters, same RSA rekey key. A COOP pair that disagrees about policy will hand different members different policies depending on which one they registered to. - **Watch for `%GDOI-3-COOP_KS_UNREACH` and `%GDOI-4-COOP_KS_UNAUTH` in your monitoring.** The first is a genuine outage. The second is the silent killer above, and it deserves an alert. ## Key Takeaways - **The data plane survives a dead key server.** With KS1 shut down, GM1 pinged GM2's LAN at 100 percent, encrypted. The key server is control plane and is never in the data path. What you lose is rekeys, re-registration, and policy changes, so the group degrades on a timer rather than dropping. - **COOP fixes the control plane.** A `redundancy` block under `server local`, a `local priority`, and a `peer address ipv4`. Highest priority becomes Primary. - **`address ipv4` must come before `redundancy`.** Otherwise: `% ERROR: Local server address must be configured first`. - **Failover is real and it works.** The syslog runs COOP\_KS\_ELECTION, COOP\_KS\_UNREACH, COOP\_KS\_TRANS\_TO\_PRI, and the priority-75 secondary promotes itself to Primary. - **COOP key servers MUST share the same RSA rekey keypair.** Export it from the primary (`crypto key export rsa ... pem terminal 3des`) and import it on the secondary. Generate it `exportable` or you will not be able to. - **`%GDOI-4-COOP_KS_UNAUTH: Contact from unauthorized KS ...` is the tell.** Its hint about a "MISCONFIG of peer/local address" is misleading. The real cause is mismatched rekey keys. - **Mismatched keys kill the group silently.** Members verify rekey signatures against the public key they hold. A secondary signing with a different key gets its rekeys rejected by every member, and the group dies at the next rekey interval, long after the change window closed. - **Test past a full rekey interval, not just the election.** A failover test that checks only "Primary?" and "ping?" passes on a broken deployment. Check that `Rekeys received` increments. That completes the GETVPN series: the [model and the header-preservation capture](https://www.pinglabz.com/getvpn-explained/), the [IOS XE build](https://www.pinglabz.com/getvpn-configuration-ios-xe/), and now redundancy. All of it captured live on cat8000v 17.18.02\. For the rest of the family (IKEv2, crypto maps, VTIs, DMVPN, FlexVPN, certificate authentication, and VRF-aware designs), the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/) is the index. ### VRF-Aware IPsec: Keeping VPNs in Their Lane URL: https://www.pinglabz.com/vrf-aware-ipsec/ Last updated: 2026-07-13T09:01:26.000Z VRF-aware IPsec has a reputation for being fiddly, and it mostly is not. What it has is two things called almost the same thing, doing completely different jobs, and until you separate them in your head nothing about the configuration makes sense. Once you do, the rest is bookkeeping. The two things are FVRF and IVRF. FVRF (front-door VRF) is where the tunnel's *outer* packets get routed: the encrypted ESP packets that go out to the peer, and the routing table used to reach the tunnel destination. IVRF (inside VRF) is where the *decrypted inner* packets land: the protected traffic, after the crypto engine has done its work. Front door versus inside. Underlay versus overlay. Getting those two straight is most of VRF-aware IPsec. We built it on a cat8000v in the CCIE Security lab, put the inside of a certificate-authenticated tunnel into a VRF called RED, and captured the one line of output that makes the whole concept click. This builds directly on the [IPsec VPN guide](https://www.pinglabz.com/ipsec-vpn/), and on the same tunnel we brought up with [certificate authentication](https://www.pinglabz.com/ipsec-certificate-authentication/) in the previous article. ## FVRF and IVRF, Precisely FVRF (front-door VRF) **What it is:** the VRF in which the tunnel's OUTER, encrypted packets are routed. **Answers:** "which routing table do I use to reach the peer's IP address?" **Configured with:** `vrf forwarding` on the tunnel SOURCE interface, plus `tunnel vrf ` on the tunnel. **Think:** the underlay. The internet. The transport network. The front door the ESP packets walk out of. IVRF (inside VRF) **What it is:** the VRF in which the DECRYPTED inner packets land after the crypto engine unwraps them. **Answers:** "which routing table do I look in to forward the plaintext traffic I just decrypted?" **Configured with:** `vrf forwarding` on the TUNNEL interface itself. **Think:** the overlay. The tenant. The protected inside network the traffic actually belongs to. Two VRFs, two jobs. A packet leaving a customer LAN enters the tunnel in the IVRF, gets encrypted, and the resulting ESP packet is routed out in the FVRF. Coming back, the ESP packet arrives via the FVRF, gets decrypted, and the plaintext inside is forwarded using the IVRF's routing table. They never touch each other, which is precisely the property you want. They are also independent. You can set one, both, or neither. Our lab set the IVRF and left the FVRF global, which is the most common real-world shape: the transport is the normal routing table, and the protected traffic belongs to a tenant. ## The Config Start with the VRF and something to put in it: ``` vrf definition RED rd 65000:1 address-family ipv4 interface Loopback20 vrf forwarding RED ip address 10.60.10.1 255.255.255.0 ``` Loopback20 stands in for a customer LAN: an interface that lives in RED and nowhere else. It is invisible from the global routing table, which is the whole point of a VRF. Now put the tunnel's inside into RED: ``` interface Tunnel5 vrf forwarding RED ip address 10.0.5.1 255.255.255.252 ip route vrf RED 10.60.20.0 255.255.255.0 Tunnel5 ``` Three moving parts. `vrf forwarding RED` on the tunnel makes RED the IVRF: everything that comes out of this tunnel decrypted gets forwarded in RED. The tunnel's own IP address (10.0.5.1) now lives in RED, not the global table. And the static route tells RED how to reach the remote LAN: out Tunnel5. Notice what is *not* in that config. The tunnel source (GigabitEthernet2, 10.100.0.21) and the tunnel destination (10.100.0.22) are untouched. They are still in the global routing table, and the router still uses the global table to route the encrypted ESP packets to the peer. That is the FVRF, and here it is simply "none" (global). The inside moved. The front door did not. ## The Gotcha That Bites Everyone Once Type `vrf forwarding RED` on an interface that already has an IP address and IOS says this: ``` GM1(config-if)#vrf forwarding RED % Interface Tunnel5 IPv4 disabled and address(es) removed due to disabling VRF RED ``` Read it slowly, because the wording is genuinely confusing (it says "disabling VRF RED" while you are in the act of enabling it). What actually happened: **moving an interface into a VRF wipes its IP address.** Gone. The address was an address in the global table, and that table is not this interface's table any more, so IOS removes it rather than silently leave a global address on a VRF interface. The consequence is an ordering rule you must never forget: ``` interface Tunnel5 vrf forwarding RED ip address 10.0.5.1 255.255.255.252 ``` VRF first. IP address second. Every single time. If you paste a config block in the other order, the `vrf forwarding` line silently destroys the IP address you just applied, and you are left with an interface in the right VRF and no addressing, wondering why your tunnel will not come up. This is worse than it sounds when you are changing an existing interface remotely. Apply `vrf forwarding` to the interface you are managing the router through and you have just removed its IP address and hung up on yourself. It is the crypto-config equivalent of sawing off the branch you are sitting on. If the interface is anywhere near your management path, do it from the console. ## The Money Capture Here is the single line this entire article exists for: ``` GM1#show crypto session detail Interface: Tunnel5 Session status: UP-ACTIVE Peer: 10.100.0.22 port 500 fvrf: (none) ivrf: RED IPSEC FLOW: permit ip 0.0.0.0/0.0.0.0 0.0.0.0/0.0.0.0 Outbound: #pkts enc'ed 6 drop 0 life (KB/Sec) 4607999/3538 ``` `fvrf: (none) ivrf: RED`. That is the design, printed by the router. `Peer: 10.100.0.22 port 500` with `fvrf: (none)` says the IKE session to the peer, and the ESP packets that follow, are being routed in the global table. The underlay is global. The router looked up 10.100.0.22 in the global routing table and sent the encrypted traffic out GigabitEthernet2. `ivrf: RED` says the decrypted inner traffic belongs to VRF RED. When an ESP packet arrives from 10.100.0.22, the crypto engine decrypts it and hands the plaintext to RED's routing table, not the global one. A packet inside this tunnel cannot leak into the global table, because the global table is not where it is forwarded. Two routing tables, one tunnel, cleanly separated. If you only remember one `show` command from VRF-aware IPsec, make it `show crypto session detail` and read the `fvrf/ivrf` field. ## Proving It Actually Forwards A design that looks right and does not pass traffic is not a design. Ping across the tunnel, scoped to the VRF: ``` GM1#ping vrf RED 10.60.20.1 source Loopback20 repeat 5 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 11/17/25 ms ``` Note the `vrf RED` keyword. Without it, the router pings from the global table, finds no route to 10.60.20.1, and fails. The destination does not exist in the global table, because it exists in RED. Forgetting the `vrf` keyword on a ping is the second most common source of "my VRF is broken" panic, right after the wiped IP address. And the routing table that made it work: ``` GM1#show ip route vrf RED C 10.60.10.0/24 is directly connected, Loopback20 L 10.60.10.1/32 is directly connected, Loopback20 S 10.60.20.0/24 is directly connected, Tunnel5 ``` Three entries and nothing else. RED knows about its own loopback (the local customer LAN) and one static route to the remote LAN via Tunnel5\. It does not know about 10.100.0.0/24, the WAN core, or anything else in the global table. It cannot, and that is the isolation working. The traffic path, end to end: a packet sourced from 10.60.10.1 to 10.60.20.1 is looked up in RED, matches the static route, enters Tunnel5, gets encrypted with the certificate-authenticated IPsec SA, and the resulting ESP packet gets a new outer header of 10.100.0.21 to 10.100.0.22 and is routed in the *global* table out to the WAN. The far end reverses it. The customer's addresses never appear on the WAN, and the WAN's addresses never appear in RED. ## Two Crypto Sessions, Two VRF Contexts, One Router Here is the detail that makes this concrete. GM1 is also a GETVPN group member, and that crypto session is running on the same box at the same time. Look at the two of them side by side: The GETVPN session Peer: 0.0.0.0 port 848 fvrf: (none) ivrf: (none) Entirely in the global table, both outside and inside. Peer 0.0.0.0 because a group member has no peer, it has a group. Port 848 is GDOI. The VRF-aware tunnel Peer: 10.100.0.22 port 500 fvrf: (none) ivrf: RED Same global underlay, but the decrypted inside lands in VRF RED. A real peer, IKEv2 on port 500, certificate-authenticated. Same router. Same physical WAN interface. Same global routing table underneath both. And yet one session's protected traffic lives in the global table and the other's lives in RED, and they cannot see each other. That is not a diagram in a design document, that is two lines of `show crypto session detail` on a live box. This is the mental model to keep: the VRF context is a property of the *crypto session*, not of the router. One box can terminate tunnels for a dozen tenants, each landing in its own VRF, all sharing one underlay. ## Why You Would Actually Do This Multi-tenant edge One aggregation router terminating VPNs for many customers. Each customer's decrypted traffic lands in its own IVRF and can never be routed into another customer's. The isolation is enforced by the forwarding table, not by an ACL somebody might mistype. Untrusted underlay, trusted inside Put the internet-facing side in an FVRF and the corporate side in the global table or an inside VRF. The internet routing table physically cannot reach the inside one. A default route learned from the ISP cannot leak into corporate routing. Overlapping address space Two customers both using 10.0.0.0/8 can terminate on the same router, because 10.0.0.0/8 in VRF CUST-A and 10.0.0.0/8 in VRF CUST-B are different routes in different tables. Without VRFs this is impossible without NAT. Management separation Keep SSH, SNMP, NTP, and syslog in the global table while all customer traffic sits in VRFs. A compromise inside a customer VRF has no route to your management plane, because there is no route. If you have worked with MPLS L3VPN, all of this will feel familiar, and it should: it is the same VRF machinery, the same route distinguishers, the same per-tenant forwarding tables. The [MPLS guide](https://www.pinglabz.com/mpls/) covers VRFs and RDs in depth, and everything you know from there applies here. VRF-aware IPsec is what you reach for when the transport between tenant sites is an encrypted tunnel rather than an MPLS core. ## The Front-Door Case, Honestly Our lab demonstrated the IVRF case: the inside in a VRF, the underlay global. The FVRF case (putting the tunnel *source* inside a VRF) is the mirror image, and it deserves an honest note rather than fabricated output. The setup looks like this. The WAN-facing interface goes into a VRF, say INET: ``` interface GigabitEthernet2 vrf forwarding INET ip address 10.100.0.21 255.255.255.0 interface Tunnel5 tunnel source GigabitEthernet2 tunnel vrf INET tunnel destination 10.100.0.22 ``` The critical line is `tunnel vrf INET`. It tells the tunnel: when you go to resolve the tunnel destination and route the encrypted packets, use INET's routing table, not the global one. Without that line the tunnel tries to reach 10.100.0.22 through the global table, finds nothing (because the interface that could reach it now lives in INET), and stays down. You would then see `fvrf: INET` in `show crypto session detail`, with the IVRF being whatever you set on the tunnel interface itself (or none, if the inside is global). The two knobs are genuinely independent: `tunnel vrf` sets the FVRF, `vrf forwarding` on the tunnel sets the IVRF, and they can be different VRFs, the same VRF, or one of each with the other global. The classic FVRF design is an internet edge where the ISP-facing interface sits in an INET VRF holding nothing but a default route to the provider, while the entire corporate network lives in the global table. The tunnel's outer packets ride INET. The decrypted inner packets land in global. The internet's routing table and the corporate routing table share a router and share nothing else, and the same ordering gotcha applies: `vrf forwarding` on that WAN interface will wipe its IP address too. ## Key Takeaways - **FVRF is the front door:** the VRF used to route the tunnel's outer, encrypted packets and to resolve the tunnel destination. Set it with `tunnel vrf `. Think underlay. - **IVRF is the inside:** the VRF the decrypted inner packets are forwarded in. Set it with `vrf forwarding ` on the tunnel interface. Think tenant. - They are independent. Our lab ran `fvrf: (none) ivrf: RED`: a global underlay with the protected traffic in a VRF, which is the most common real-world shape. - `show crypto session detail` prints `fvrf:` and `ivrf:` on the Peer line. That one line tells you the entire VRF design of the tunnel. It is the first thing to check when VRF-aware IPsec misbehaves. - **Moving an interface into a VRF wipes its IP address:** `% Interface Tunnel5 IPv4 disabled and address(es) removed due to disabling VRF RED`. Always configure `vrf forwarding` BEFORE `ip address`. Do it from the console if the interface is anywhere near your management path. - Test with the VRF keyword: `ping vrf RED 10.60.20.1 source Loopback20`. A plain ping uses the global table, finds no route, and fails, and that is not a fault. - `show ip route vrf RED` should contain only the tenant's routes. If you can see the WAN core in there, your isolation is leaking. - Two crypto sessions on the same router can live in different VRF contexts at once. Our GETVPN session showed `ivrf: (none)` while the VRF tunnel showed `ivrf: RED`, on the same box, at the same time. - Use it for multi-tenant edges, untrusted internet underlays, overlapping customer address space, and keeping management traffic out of customer VRFs. The tunnel used here is the certificate-authenticated IKEv2 SVTI from [certificate authentication for IPsec](https://www.pinglabz.com/ipsec-certificate-authentication/), and the VRF machinery is the same you already know from [MPLS L3VPN](https://www.pinglabz.com/mpls/). For the full progression from a basic site-to-site tunnel through IKEv2, certificates, group encryption, and now VRF awareness, start at the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/). ### GETVPN Configuration: Key Server, Group Members, and GDOI URL: https://www.pinglabz.com/getvpn-configuration-ios-xe/ Last updated: 2026-07-13T09:01:23.000Z GETVPN is the strangest member of the [IPsec VPN](https://www.pinglabz.com/ipsec-vpn/) family to configure, because most of what you expect to type is not there. There is no peer statement. There is no crypto ACL on the site routers. There is no tunnel interface. Almost the entire configuration lives on one box (the key server) and the sites get a six-line paste that is identical everywhere. If you have not read [GETVPN Explained](https://www.pinglabz.com/getvpn-explained/) yet, read that first for the model (key server, group member, GDOI, KEK, TEK) and for the packet capture that shows why GETVPN needs a private WAN. This article is the build. Key server first, then group members, then verification, every command run on Cisco IOS XE and every piece of output below copied from the console. The lab: two key servers (KS1 10.100.0.11, KS2 10.100.0.12), two group members (GM1 10.100.0.21, GM2 10.100.0.22), all cat8000v running IOS XE 17.18.02 on a shared private WAN core, 10.100.0.0/24\. GM1's protected LAN is 10.20.10.0/24, GM2's is 10.30.10.0/24\. Everything in this article is verified on that platform. **GDOI is fully supported on cat8000v 17.18.02** and every GETVPN command below was accepted, which is worth stating plainly because the internet is full of people asking whether it still exists on modern IOS XE. It does. ## Step 1: The Key Server, Block by Block The key server is where all the policy lives. Build it in this order, because later blocks reference earlier ones. ### 1a. ISAKMP policy and pre-shared keys GDOI registration is protected by an IKEv1 phase 1 SA. That means a group member and a key server must first authenticate to each other before the key server will hand over anything. So you need an ISAKMP policy and a key per member. ``` crypto isakmp policy 10 encryption aes 256 hash sha256 authentication pre-share group 14 crypto isakmp key PingLabz-GETVPN-01 address 10.100.0.21 crypto isakmp key PingLabz-GETVPN-01 address 10.100.0.22 ``` Note what those two key statements point at: the **group members**, 10.100.0.21 and 10.100.0.22\. The key server needs a key for every member that will register to it. (In production you would use certificates rather than a shared PSK across every site, which is exactly the argument for a PKI, but a PSK keeps this lab focused on GDOI.) This is the one place in GETVPN where you configure something per remote router, and even here it is per *member*, not per *pair of members*. There is no N-squared anywhere. ### 1b. Transform set and IPsec profile These describe the data-plane encryption the group will use. They live on the key server because the key server is the thing that decides policy. Group members never see a transform set in their own config, they receive it. ``` crypto ipsec transform-set PLZ-GET-TS esp-aes 256 esp-sha256-hmac mode tunnel crypto ipsec profile PLZ-GET-PROF set transform-set PLZ-GET-TS ``` `mode tunnel` is not a contradiction of "GETVPN has no tunnels." GETVPN uses ESP tunnel mode encapsulation and then copies the original IP header into the outer header instead of writing its own. That is header preservation, and you can see the result in the GM output later (`encaps: ENCAPS_TUNNEL`). ### 1c. The group policy ACL, and why it is not a crypto ACL This is the block that trips people who come to GETVPN from crypto maps. It looks like a crypto ACL. It is not one. ``` ip access-list extended PLZ-GET-POLICY permit ip 10.20.10.0 0.0.0.255 10.30.10.0 0.0.0.255 permit ip 10.30.10.0 0.0.0.255 10.20.10.0 0.0.0.255 permit ip 192.168.99.0 0.0.0.255 10.30.10.0 0.0.0.255 permit ip 10.30.10.0 0.0.0.255 192.168.99.0 0.0.0.255 ``` Two things are different from a crypto map ACL, and both matter: Crypto map ACL ScopeOne ACL per peer DirectionWritten from the local side only Remote endMust be a mirror image, or SAs fail Lives onEvery router, maintained by hand GETVPN group policy ACL ScopeOne ACL for the entire group DirectionSymmetric, both directions written out Remote endGets the same ACL, so it cannot mismatch Lives onThe key server only, pushed to members Look at the ACL again with that in mind. It lists **both directions** of every flow: 10.20.10.0 to 10.30.10.0, *and* 10.30.10.0 to 10.20.10.0\. That is not redundancy. Every member receives this same ACL, so GM1 needs to match the outbound direction and GM2 needs to match its own outbound direction, and they are reading the identical policy. One symmetric ACL covers both. Get this wrong and you get the classic GETVPN failure: traffic encrypts one way and arrives at the far end as an ESP packet the receiver has no policy for. Write it symmetric. Always. ### 1d. The RSA rekey key The key server signs its rekey messages so that group members can verify they came from a legitimate key server and not from an attacker. That signature needs an RSA keypair. ``` crypto key generate rsa modulus 2048 label PLZ-REKEY-KEY exportable ``` The `exportable` keyword looks optional. It is not, if you have any intention of adding a second key server later. COOP key servers must share the *same* RSA rekey keypair, and you cannot export a key you did not mark exportable at generation time. Generate it exportable now or regenerate it later under pressure. (The full story, including the exact syslog you get when the keys do not match, is in [GETVPN redundancy: COOP key servers and rekey survival](https://www.pinglabz.com/getvpn-coop-key-servers/).) ### 1e. The GDOI group with server local Now assemble it. This is the block that makes the router a key server. ``` crypto gdoi group PLZ-GETVPN identity number 4321 server local rekey algorithm aes 256 rekey authentication mypubkey rsa PLZ-REKEY-KEY rekey transport unicast rekey lifetime seconds 900 sa ipsec 1 profile PLZ-GET-PROF match address ipv4 PLZ-GET-POLICY replay counter window-size 64 address ipv4 10.100.0.11 ``` Line by line: identity number 4321The group ID. Must match on every member. This is what a GM asks for at registration. server local"This router IS a key server." Its absence is what makes a router a group member instead. rekey algorithm aes 256Cipher for the KEK, the key that protects the rekey messages themselves. rekey authentication mypubkey rsaWhich RSA key signs the rekeys. Members verify against the matching public key. rekey transport unicastSend each member its own copy. The alternative, multicast, needs a multicast-capable WAN. rekey lifetime seconds 900How often the KEK rolls. Short here so the lab rekeys while you watch. Production is longer. sa ipsec 1The data-plane SA definition. Wraps the profile, the policy ACL, and anti-replay. match address ipv4 PLZ-GET-POLICYThe group policy ACL from step 1c. This is what gets pushed to every member. replay counter window-size 64Counter-based anti-replay. Time-based is the other option and needs synchronised clocks. address ipv4 10.100.0.11The source address this key server uses for GDOI. Members must point at exactly this. One warning about ordering that costs people real time: if you later add `redundancy` for COOP, `address ipv4` must already be configured. Configure the local address early and you will never see that error. ## Step 2: The Group Member (Prepare To Be Underwhelmed) Here is the *entire* GETVPN configuration on GM1: ``` crypto isakmp key PingLabz-GETVPN-01 address 10.100.0.11 crypto isakmp key PingLabz-GETVPN-01 address 10.100.0.12 crypto gdoi group PLZ-GETVPN identity number 4321 server address ipv4 10.100.0.11 server address ipv4 10.100.0.12 crypto map PLZ-GET-CMAP 10 gdoi set group PLZ-GETVPN interface GigabitEthernet2 crypto map PLZ-GET-CMAP ``` That is it. Read what is absent: - **No peer.** There is no `set peer` anywhere. The crypto map is a `gdoi` map, not an ISAKMP map. - **No crypto ACL.** There is no `match address`. The member has no idea what to encrypt until the key server tells it. - **No transform set.** Downloaded. - **No mention of GM2.** None. GM1 does not know GM2 exists. The only remote addresses in the entire block are the two key servers. And **GM2's configuration is identical to GM1's**, character for character. Not "similar." Not "mirrored." Identical. That is the sentence to hold onto, because it is the whole scaling argument. Adding site number 47 to a GETVPN group means pasting the same nine lines onto a new router, with no change to the 46 routers already running and no change to the key server other than one ISAKMP key. Adding site 47 to a full-mesh crypto map deployment means touching 46 other routers. ### The message IOS gives you, and why it is not an error Apply the crypto map before the group has registered and IOS says: ``` % NOTE: This new crypto map will remain disabled until a valid group has been configured. ``` That is normal. It is telling you the truth: a `gdoi` crypto map has no policy of its own, so until the router registers to a key server and downloads a group policy, the map has nothing to enforce and stays inert. If it persists after you have configured the group, your registration is failing (check reachability to the key server on UDP 848, the group identity number, and the ISAKMP key). ## Step 3: Verify From the Group Member `show crypto gdoi` on a member is the single most useful command in GETVPN. It answers four questions at once: did I register, what policy did I get, how are rekeys protected, and what key am I encrypting with. ``` GM1#show crypto gdoi Group Name : PLZ-GETVPN Group Identity : 4321 Group Type : GDOI (ISAKMP) Rekeys received : 0 Group Server list : 10.100.0.11 10.100.0.12 Group member : 10.100.0.21 vrf: None Registration status : Registered Registered with : 10.100.0.11 Re-registers in : 692 sec Succeeded registration: 1 ``` `Registration status : Registered` is the pass/fail line. `Registered with : 10.100.0.11` tells you it picked KS1 (it tries the server list in order). `Re-registers in : 692 sec` is worth noting: a group member re-registers periodically, which is how it survives a key server going away and coming back. ### The policy it downloaded ``` ACL Downloaded From KS 10.100.0.11: access-list permit ip 10.20.10.0 0.0.0.255 10.30.10.0 0.0.0.255 access-list permit ip 10.30.10.0 0.0.0.255 10.20.10.0 0.0.0.255 access-list permit ip 192.168.99.0 0.0.0.255 10.30.10.0 0.0.0.255 access-list permit ip 10.30.10.0 0.0.0.255 192.168.99.0 0.0.0.255 ``` Compare that against `PLZ-GET-POLICY` on the key server. It is the same ACL. It is not configured on GM1 and never will be. The words "Downloaded From KS" are doing all the work here: this is the centralised policy model made visible. Change the ACL once on the key server, and every member picks up the change. ### The KEK policy ``` KEK POLICY: Rekey Transport Type : Unicast Lifetime (secs) : 849 Encrypt Algorithm : AES Key Size : 256 Sig Hash Algorithm : HMAC_AUTH_SHA Sig Key Length (bits) : 2352 ``` Every value here traces back to a line on the key server: `rekey transport unicast`, `rekey algorithm aes 256`, `rekey lifetime seconds 900` (counting down through 849). `Sig Key Length (bits) : 2352` is the encoded length of the 2048-bit RSA rekey key the key server signs with. The KEK protects the rekeys. It does not protect user data. ### The TEK policy, and the SPI ``` TEK POLICY for the current KS-Policy ACEs Downloaded: GigabitEthernet2: IPsec SA: spi: 0xE1F309BF(3790801343) transform: esp-256-aes esp-sha256-hmac sa timing:remaining key lifetime (sec): (3550) Anti-Replay(Counter Based) : 64 encaps: ENCAPS_TUNNEL ``` This is the key that actually encrypts data, and the transform matches `PLZ-GET-TS` exactly. Now run the same command on the other member: ``` GM2#show crypto gdoi | include Group member|Registration status|spi:|transform: Group member : 10.100.0.22 vrf: None Registration status : Registered spi: 0xE1F309BF(3790801343) transform: esp-256-aes esp-sha256-hmac ``` **Same SPI: 0xE1F309BF.** Two different routers, one shared SA. If you only run one verification command after building a GETVPN group, run this one on two members and compare the SPI. Matching SPIs means the group is genuinely a group. ### The session table, for the shock value ``` GM1#show crypto session detail Interface: GigabitEthernet2 Session status: UP-ACTIVE Peer: 0.0.0.0 port 848 fvrf: (none) ivrf: (none) IPSEC FLOW: permit ip 10.20.10.0/255.255.255.0 10.30.10.0/255.255.255.0 Outbound: #pkts enc'ed 18 drop 0 life (KB/Sec) KB Vol Rekey Disabled/2115 ``` UP-ACTIVE, encrypting, peer 0.0.0.0, port 848 (GDOI). If you came from crypto maps or [DMVPN](https://www.pinglabz.com/dmvpn/), this line is the moment GETVPN stops being abstract. There is no peer because there is no tunnel. ## Step 4: Verify From the Key Server The key server has its own pair of commands, and they are the fastest way to answer "did that new site actually come up?" ``` KS1#show crypto gdoi ks Total group members registered to this box: 2 Key Server Information For Group PLZ-GETVPN: Group Name : PLZ-GETVPN Group Identity : 4321 Group Type : GDOI (ISAKMP) Group Members : 2 IPSec SA Direction : Both ACL Configured: access-list PLZ-GET-POLICY ``` `Total group members registered to this box: 2` is the health check. `IPSec SA Direction : Both` confirms the symmetric policy is being applied in both directions, which is the key server's read of the ACL you wrote in step 1c. For the roster, go one level deeper: ``` KS1#show crypto gdoi ks members Group Member ID : 10.100.0.21 GM Version: 1.0.26 Group ID : 4321 Group Name : PLZ-GETVPN GM State : Registered Key Server ID : 10.100.0.11 Group Member ID : 10.100.0.22 GM Version: 1.0.26 Group ID : 4321 GM State : Registered Key Server ID : 10.100.0.11 ``` Both members registered, both to KS1\. In a hundred-site deployment this command is your inventory: if a site is missing here, it is not encrypting, and the fault is registration (ISAKMP key, reachability, group identity), not the data plane. ## Step 5: Pass Real Traffic The final proof. Ping between the two protected LANs, sourcing from the loopbacks that the group policy ACL permits, in both directions: ``` GM1#ping 10.30.10.1 source Loopback10 repeat 8 !!!!!!!! Success rate is 100 percent (8/8), round-trip min/avg/max = 5/9/30 ms GM2#ping 10.20.10.1 source Loopback10 repeat 5 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 5/11/34 ms ``` One hundred percent both ways, and the `#pkts enc'ed` counter in `show crypto session detail` climbing to prove it went through ESP rather than around it. Always check the counter as well as the exclamation marks: a GETVPN group whose crypto map has not activated will happily pass traffic *in clear text* and the ping still succeeds. The ping proves reachability. The encryption counter proves encryption. ## Troubleshooting Order When a member does not come up, work in this order. It follows the dependency chain, so you never chase a symptom whose cause is one layer down. 1**Reachability.** Can the member reach the key server's `address ipv4` on UDP 848? GDOI is not port 500. 2**ISAKMP.** Does the key server have a `crypto isakmp key` for that member's address, and does the member have one for the key server's? 3**Group identity.** `identity number` must be the same integer on the key server and every member. Ours is 4321. 4**Registration.** `show crypto gdoi` on the member. Not "Registered" means stop here, the problem is above. 5**Policy.** Is the ACL actually downloaded, and does it cover the flow you are testing? Symmetric entries both ways? 6**SPI match.** Compare `spi:` across two members. Different SPIs means they are not in the same group SA. 7**Underlay routing.** The WAN must route the protected LAN prefixes, because header preservation puts them in the outer header. Point 7 is the one that catches people migrating a lab onto a real network. GETVPN's outer header carries the original LAN addresses, so if your core cannot route 10.30.10.0/24, no amount of crypto configuration will save you. That constraint is explained in full, with the packet capture, in [GETVPN Explained](https://www.pinglabz.com/getvpn-explained/). ## Key Takeaways - **All policy lives on the key server.** ISAKMP policy, transform set, IPsec profile, the group ACL, the RSA rekey key, and the `crypto gdoi group ... server local` block. Members get a paste. - **The group policy ACL is symmetric and singular.** Write both directions of every flow. It is one ACL for the whole group, not a per-peer mirror image, and it is the thing that gets pushed to every member. - **Generate the RSA rekey key as `exportable`.** You cannot export it later, and COOP key servers must share the same keypair. - **The group member config is nine lines and has no peer and no ACL.** GM1 and GM2 are identical. That is the entire scaling story. - **`% NOTE: This new crypto map will remain disabled until a valid group has been configured.`** is expected, not an error. A gdoi crypto map is inert until registration completes. - **Verify with `show crypto gdoi` on the member** for registration status, the downloaded ACL, the KEK policy, and the TEK SPI. Compare the SPI across two members, it must match. - **Verify with `show crypto gdoi ks` and `show crypto gdoi ks members` on the key server** for the registered roster. - **Check the encryption counter, not just the ping.** An inactive crypto map passes traffic in clear text and the ping still returns 100 percent. - **GDOI is fully supported on cat8000v IOS XE 17.18.02.** Every command in this article was accepted and verified on that image. Next, make it survivable: two key servers, priority-based election, real failover syslog, and the RSA rekey key trap that silently kills a group at the next rekey interval, in [GETVPN redundancy: COOP key servers and rekey survival](https://www.pinglabz.com/getvpn-coop-key-servers/). For the rest of the family (IKEv2, crypto maps, VTIs, DMVPN, FlexVPN, certificate authentication, and VRF-aware designs), the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/) is the index. ### Certificate Authentication for IPsec: Trustpoints, Enrollment, and Revocation URL: https://www.pinglabz.com/ipsec-certificate-authentication/ Last updated: 2026-07-13T09:01:25.000Z Pre-shared keys are a fine way to learn IPsec and a poor way to run it. A PSK is a shared secret sitting in a config file, in plain sight of anyone with read access to the device, in your backup repository, and in the change ticket where someone pasted the running config. It does not scale either: with N peers you have N keys to distribute, and rotating one of them means touching both ends at the same time without dropping the tunnel. Certificates fix both problems at once. Each device holds its own private key, which never leaves it, and a certificate signed by a common authority that every peer already trusts. Add a peer and you enroll it once; nobody else's config changes. Compromise a device and you revoke one certificate rather than rekeying a mesh. This article is the end-to-end walkthrough: build the trustpoint, authenticate the CA, enroll, and then bring up an IKEv2 tunnel with no pre-shared key anywhere in its configuration. It is the natural next step after the [IPsec VPN guide](https://www.pinglabz.com/ipsec-vpn/), and it assumes you have a CA to enroll against. We do. In the previous article we turned a cat8000v into a working [IOS certificate authority](https://www.pinglabz.com/cisco-router-ca-server/) called PLZ-CA. Everything below enrolls against that exact CA and uses the certificates it issued. All output is from the live lab. ## The Trustpoint: Where Your Certificate Lives A trustpoint is IOS's container for a certificate relationship. It holds the CA's certificate (so you can verify things the CA signed), your own identity certificate (so peers can verify you), the keypair backing that certificate, and the policy for how you got it and how you check it. One trustpoint, one CA relationship. Here is GM1's, exactly as configured: ``` crypto key generate rsa modulus 2048 label PLZ-GM1-KEY crypto pki trustpoint PLZ-TP enrollment url http://10.100.0.31:80 subject-name CN=GM1.pinglabz.lab,O=PingLabz revocation-check none rsakeypair PLZ-GM1-KEY ``` crypto key generate rsa ... label PLZ-GM1-KEY Generate the keypair FIRST, with a label, so the trustpoint can reference it. The private half never leaves this router. The public half goes into the CSR and ends up inside the certificate. enrollment url http://10.100.0.31:80 Where the CA lives. This is SCEP over HTTP: the router will fetch the CA certificate from here and later POST its CSR to the same place. Reachability matters, so make sure routing to the CA works before you start. subject-name CN=GM1.pinglabz.lab,O=PingLabz The identity you are claiming. This becomes the Subject field of your certificate, and it is what the peer will match against. Get it wrong here and you will be re-enrolling later. revocation-check none Do not check whether the peer's certificate has been revoked. This is a lab choice and we are being upfront about it. See the revocation section below for what production should say. rsakeypair PLZ-GM1-KEY Bind the trustpoint to the keypair you generated. Without this, IOS picks the default general-purpose key, which is fine until you have two trustpoints and want them isolated. ## Step 1: Authenticate the CA (and Actually Check the Fingerprint) Before you can get a certificate, you have to decide you trust the entity that issues them. `crypto pki authenticate` fetches the CA's self-signed root certificate over SCEP and asks you a question. Here is the real exchange from GM1: ``` GM1(config)#crypto pki authenticate PLZ-TP Certificate has the following attributes: Fingerprint MD5: 50AE2189 0FC62684 1418DEF0 B0EB714A Fingerprint SHA1: CAC64CCF 98899856 A5F002E5 C168BF80 0AB8355F % Do you accept this certificate? [yes/no]: yes Trustpoint CA certificate accepted. ``` Now go and look at what the CA itself reported when we brought it up: ``` CA1#show crypto pki server Certificate Server PLZ-CA: Status: enabled CA cert fingerprint: 50AE2189 0FC62684 1418DEF0 B0EB714A ``` Same value. `50AE2189 0FC62684 1418DEF0 B0EB714A` on the CA, and `50AE2189 0FC62684 1418DEF0 B0EB714A` on the client. That match is not a curiosity. It is the entire security of this step, and it deserves a hard stare. So IOS punts to the human. It prints the fingerprint of whatever certificate it just received and asks you to confirm. Your job, and this is the whole job, is to obtain the real CA's fingerprint through a channel the attacker does not control (an SSH session to the CA, a phone call to the person who runs it, a printed sheet in a binder) and compare it character by character. If they match, you are talking to the real CA. If they do not, you have just caught an attack. If you type `yes` without doing that comparison, you have authenticated nothing. You have simply agreed to trust whatever answered. Every certificate that CA later signs, including the ones an attacker mints for himself, will validate perfectly on your router. The rest of your PKI is now theatre. This is a five-second check and skipping it invalidates everything downstream, so do not skip it. ## Step 2: Enroll (Get Your Own Certificate) With the CA trusted, ask it for an identity: ``` GM1(config)#crypto pki enroll PLZ-TP ``` This is also interactive: IOS asks for a challenge password (used for revocation requests later), whether to include the serial number and IP address in the subject, and whether to send the request now. The router then builds a CSR carrying your public key and subject name, signs it with your private key to prove you hold that key, and posts it to the enrollment URL. The private key itself never moves. The CA only ever sees the public half, which is the structural advantage over a PSK: no secret is ever in transit, and no secret ends up sitting in two config files at once. Because our CA runs `grant auto`, it signs immediately and hands back the certificate. Syslog confirms: ``` %PKI-6-CERT_INSTALL: An ID certificate has been installed ``` ## The Certificate We Actually Got Real output from GM1 after enrollment: ``` GM1#show crypto pki certificates PLZ-TP Certificate Status: Available Certificate Serial Number (hex): 02 Certificate Usage: General Purpose Issuer: cn=PLZ-CA o=PingLabz c=US Subject: Name: GM1.pinglabz.lab hostname=GM1.pinglabz.lab cn=GM1.pinglabz.lab o=PingLabz Validity Date: start date: 09:08:26 UTC Jul 13 2026 end date: 09:08:26 UTC Jul 12 2028 Associated Trustpoints: PLZ-TP CA Certificate Status: Available Certificate Serial Number (hex): 01 Certificate Usage: Signature Issuer: cn=PLZ-CA o=PingLabz c=US Subject: cn=PLZ-CA o=PingLabz c=US Validity Date: start date: 09:04:14 UTC Jul 13 2026 end date: 09:04:14 UTC Jul 12 2031 Associated Trustpoints: PLZ-TP ``` Two certificates, and the difference between them is the whole of PKI. **The ID certificate (serial 02)** is GM1's identity. Subject is `cn=GM1.pinglabz.lab`: this is who GM1 says it is. Issuer is `cn=PLZ-CA`: this is who vouches for that claim. Subject and Issuer are different, which is exactly what you expect from a certificate someone else signed. Validity runs two years, from Jul 13 2026 to Jul 12 2028, matching the CA's `lifetime certificate 730`. Usage is General Purpose, meaning it can be used for both signing and encryption. **The CA certificate (serial 01)** is the trust anchor. Look closely: Subject and Issuer are *identical*, both `cn=PLZ-CA`. It signed itself, because there is nothing above it. Usage is Signature, because its job is to sign other certificates. Validity runs five years, out to 2031, comfortably outliving every certificate it will issue. GM1 holds this because you typed `yes` at the fingerprint prompt, and it is what GM1 will use to verify GM2's certificate later. ## The Payoff: An IKEv2 Tunnel With No Pre-Shared Key Now build a static VTI between GM1 and GM2 and authenticate it with those certificates: ``` crypto ikev2 proposal PLZ-CERT-PROP encryption aes-cbc-256 integrity sha256 group 14 crypto ikev2 policy PLZ-CERT-POL proposal PLZ-CERT-PROP crypto ikev2 profile PLZ-CERT-PROF match identity remote fqdn GM2.pinglabz.lab identity local fqdn GM1.pinglabz.lab authentication local rsa-sig authentication remote rsa-sig pki trustpoint PLZ-TP crypto ipsec profile PLZ-CERT-IPSEC set transform-set PLZ-CERT-TS set ikev2-profile PLZ-CERT-PROF interface Tunnel5 ip address 10.0.5.1 255.255.255.252 tunnel source GigabitEthernet2 tunnel mode ipsec ipv4 tunnel destination 10.100.0.22 tunnel protection ipsec profile PLZ-CERT-IPSEC ``` Read that config again and look for a secret. There is not one. No `crypto isakmp key`, no `pre-shared-key local`, no `pre-shared-key remote`, nothing. The three lines doing the work are: - `authentication local rsa-sig` \- I will prove who I am by signing with my private key. Verify me against my certificate. - `authentication remote rsa-sig` \- I expect the peer to do the same. I will verify their signature against their certificate, and their certificate against the CA. - `pki trustpoint PLZ-TP` \- use this trustpoint for both. It holds my certificate and the CA certificate I need to check theirs. The `identity local` and `match identity remote` lines are the identity contract, and they are about to cause us a problem. ## The Failure We Hit: Identity Type Mismatch The first time we brought this up, IKEv2 refused, with this: ``` IKEv2-ERROR:(SESSION ID = 1,SA ID = 1):: Auth exchange failed ``` That is the entire error. It tells you the authentication exchange failed, which you had already worked out, and nothing else. No mention of certificates, no mention of identities, no hint about which side had the problem. This message is the single most common thing you will see when certificate authentication goes wrong, and it is almost useless in isolation. Which is precisely why the pattern below is worth committing to memory. The cause: we had configured `identity local dn` (send my identity as a Distinguished Name, taken from the certificate subject) while the peer's profile said `match identity remote fqdn GM2.pinglabz.lab`. So GM1 announced itself with a DN, and GM2's profile was looking for an FQDN. GM2 received a perfectly valid identity, in a perfectly valid certificate, signed by a CA it trusts, and rejected it because it was the wrong *type* of identity. The fix took one line on each side: ``` crypto ikev2 profile PLZ-CERT-PROF identity local fqdn GM1.pinglabz.lab ``` The tunnel came up immediately. When you see `Auth exchange failed`, check the identity types on both ends before anything else. The rule to internalize: **the identity type you SEND must be the identity type the peer is told to MATCH.** IKEv2 identities come in several flavours (fqdn, dn, address, email, key-id) and they are not interchangeable. If your side says `identity local dn`, the peer must say `match identity remote dn ...` with the full DN string. If your side says `identity local fqdn GM1.pinglabz.lab`, the peer must say `match identity remote fqdn GM1.pinglabz.lab`. Mixing the two produces a generic auth failure and hours of staring at crypto debugs. ## The Proof: Auth sign RSA With the identities aligned, here is the SA: ``` GM1#show crypto ikev2 sa detailed Tunnel-id Local Remote fvrf/ivrf Status 1 10.100.0.21/500 10.100.0.22/500 none/none READY Encr: AES-CBC, keysize: 256, PRF: SHA256, Hash: SHA256, DH Grp:14, Auth sign: RSA, Auth verify: RSA Local id: GM1.pinglabz.lab Remote id: GM2.pinglabz.lab ``` `Auth sign: RSA, Auth verify: RSA`. On every pre-shared-key tunnel in this series, that field said PSK. Here it says RSA, twice: GM1 signed the IKE\_AUTH exchange with its private key, and it verified GM2's signature using GM2's certificate, which it validated against the PLZ-CA certificate in its trustpoint. `Local id: GM1.pinglabz.lab` and `Remote id: GM2.pinglabz.lab` are the identities that were exchanged and matched. Those strings came out of the certificates, and the certificates came out of a CA both routers independently decided to trust. No secret was ever shared between GM1 and GM2\. They have still never held a key in common, and yet each one is certain of who it is talking to. ## Revocation: Being Honest About It Our trustpoint says `revocation-check none`. That means when GM1 validates GM2's certificate, it checks the signature and the validity dates, and stops. It does not ask whether the certificate has been revoked. If GM2 were compromised, its certificate torn up, and its revocation published to the world, GM1 would still cheerfully build a tunnel to it, right up until the certificate expired in 2028. That is a real hole and it is on us. Here are the three options: revocation-check none **What it does:** no revocation check at all. **Use when:** lab, or a closed environment where you can physically reach every device. **Cost:** a revoked certificate stays usable until it expires. This is what we used, and we are telling you so. revocation-check crl **What it does:** download the CA's revocation list, cache it, check the peer's serial against it. **Use when:** you have a reliable CRL distribution point and can tolerate staleness. **Cost:** a CRL is a snapshot. Revoke a certificate five minutes after the CRL was published and you are exposed until the next one. Also, if the CRL is unreachable, validation fails and your tunnels drop. revocation-check ocsp **What it does:** ask an OCSP responder about this one certificate, right now. **Use when:** you want real-time revocation status and can run a highly available responder. **Cost:** the responder is now in the path of every tunnel establishment. It must be up, and it must be fast. For production, pick `ocsp` or `crl` based on what you can actually keep running, and think hard about whether a revocation check failure should drop tunnels (secure) or allow them (available). There is no free answer. What there definitely is not is a good reason to leave `none` on a production edge. ## PSK vs Certificates Scaling **PSK:** a key per peer pair. A full mesh of N sites is N(N-1)/2 secrets to generate and distribute. **Certificates:** enroll each device once against a common CA. Adding site 51 changes nothing on sites 1 through 50. Where the secret lives **PSK:** in plain text in two config files, your backups, and probably a chat thread. **Certificates:** the private key is generated on the device and never leaves it. Nothing secret is ever transmitted. Rotation **PSK:** a coordinated change on both ends, at the same time, with a tunnel outage in the middle. **Certificates:** re-enroll one device. Peers do not care, because they trust the CA, not the device. Blast radius of a compromise **PSK:** a leaked key compromises every tunnel using it, and a lot of designs reuse one key everywhere. **Certificates:** a compromised device exposes one private key. Revoke that certificate and it is over (assuming you actually check revocation). Operational cost **PSK:** near zero to start. One line of config and the tunnel is up. **Certificates:** you now run a PKI. That means a CA, a clock, enrollment, expiry tracking, and revocation infrastructure. This is real, ongoing work. Failure modes **PSK:** mismatched key, and the error usually says so. **Certificates:** expired certificate, wrong clock, identity type mismatch, unreachable CRL. Errors are generic. You must know the patterns, which is why this article exists. ## Key Takeaways - A trustpoint holds one CA relationship: the enrollment URL, your subject name, your keypair, and the revocation policy. It says nothing about any peer or tunnel, which is exactly why it scales. - `crypto pki authenticate` displays the CA certificate's fingerprint and asks you to accept it. Compare that fingerprint out of band against `show crypto pki server` on the CA. Ours matched at `50AE2189 0FC62684 1418DEF0 B0EB714A`. Typing yes without checking authenticates nothing. - `crypto pki enroll` generates a CSR from your public key and subject name, sends it, and installs the returned identity certificate (`%PKI-6-CERT_INSTALL`). Your private key never leaves the router. - In `show crypto pki certificates`, the ID certificate has different Subject and Issuer (GM1 signed by PLZ-CA, serial 02). The CA certificate has identical Subject and Issuer, because it signed itself (serial 01). - A certificate-authenticated IKEv2 tunnel uses `authentication local rsa-sig`, `authentication remote rsa-sig`, and `pki trustpoint`. There is no pre-shared key anywhere in the configuration. - `Auth sign: RSA, Auth verify: RSA` in `show crypto ikev2 sa detailed` is the proof it worked. On a PSK tunnel that field says PSK. - **Identity type mismatch is the number one cert-auth failure.** `identity local dn` against `match identity remote fqdn` produces a bare `Auth exchange failed` and no other clue. The type you send must be the type the peer matches. - We used `revocation-check none` and are saying so plainly. Production wants `crl` or `ocsp`, and you need to decide deliberately whether a check failure should drop the tunnel or allow it. The CA that issued these certificates is a plain IOS XE router, built from scratch in the same lab: see [running a Cisco router as a CA server](https://www.pinglabz.com/cisco-router-ca-server/) for how to stand one up in six lines. If the IKEv2 exchange itself is what you want to understand, [IKEv2 explained](https://www.pinglabz.com/ikev2-explained/) covers the negotiation these certificates are plugged into, and the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/) ties the whole progression together, from pre-shared keys through certificates to group encryption. ### GETVPN Explained: Group Encryption Without Tunnels URL: https://www.pinglabz.com/getvpn-explained/ Last updated: 2026-07-13T09:01:22.000Z Every VPN technology you have configured up to now has a peer. A crypto map has a peer statement. A GRE over IPsec tunnel has a tunnel destination. DMVPN has an NHRP next-hop server. Even [FlexVPN](https://www.pinglabz.com/flexvpn-explained/), which abstracts the whole thing behind IKEv2, still ends up with a peer address in the session table. That peering model is the reason full-mesh IPsec does not scale: N sites means N(N-1)/2 tunnels, and every one of them needs a peer, a crypto ACL, and a pair of SAs. GETVPN (Group Encrypted Transport VPN) throws that model out. There is no tunnel. There is no peer. There is no per-peer crypto ACL. Every site pulls the same key from a central key server and encrypts to *everyone at once*. It is the one VPN in the [IPsec VPN](https://www.pinglabz.com/ipsec-vpn/) family where the phrase "tunnel mesh" does not apply, because there is no mesh to build. That sounds like magic until you look at a packet capture. Then it becomes obvious, and it also becomes obvious why GETVPN cannot cross the public internet. This article walks the model (key server, group member, GDOI, KEK, TEK) and then proves every claim with real output from a Cisco IOS XE lab: two key servers, two group members, a private WAN core, and a packet capture that tells you exactly what GETVPN is doing to your headers. ## Lead With the Weirdest Output in Networking Before any theory, look at what `show crypto session` reports on a GETVPN group member. This is a router that is actively encrypting traffic to another site right now: ``` GM1#show crypto session detail Interface: GigabitEthernet2 Session status: UP-ACTIVE Peer: 0.0.0.0 port 848 fvrf: (none) ivrf: (none) IPSEC FLOW: permit ip 10.20.10.0/255.255.255.0 10.30.10.0/255.255.255.0 Outbound: #pkts enc'ed 18 drop 0 life (KB/Sec) KB Vol Rekey Disabled/2115 ``` **Peer: 0.0.0.0.** The session is UP-ACTIVE, packets are being encrypted (18 of them), and the peer is all zeroes. That is not a bug and it is not an uninitialised field. A group member does not have a peer. It has a *group*. The crypto session is a relationship with a policy, not with another router. And port 848 is not ISAKMP (500) or NAT-T (4500). Port 848 is **GDOI**, the Group Domain of Interpretation (RFC 6407), the protocol a group member uses to talk to the key server. That single line tells you the entire architecture: the control plane points at a key server on UDP 848, and the data plane points at nobody in particular. ## The GETVPN Model in Four Nouns GETVPN has a small vocabulary, and once you have it the rest of the technology falls out of it. Key Server (KS) The brain. Holds the group policy, generates the keys, and pushes them to members. Control plane only. It is never in the data path. Group Member (GM) Any site router. Registers to a KS, downloads the policy and keys, then encrypts and decrypts. Its config knows nothing about other sites. KEK (Key Encryption Key) Protects the rekey messages themselves. The KS encrypts new keys with the KEK and signs them with its RSA rekey key. TEK (Traffic Encryption Key) The actual data key. Every member in the group uses the same TEK, so any member can decrypt any other member's traffic. GDOI is the glue. It is the registration and rekey protocol (UDP 848) that a GM uses to say "I am a member of group 4321, give me the policy" and that a KS uses to answer with the ACL, the KEK, and the TEK. Registration itself is protected by an IKE phase 1 SA, so the group secret is never handed out in the clear. The mental shift is this: in a normal [IPsec](https://www.pinglabz.com/ipsec-vpn/) deployment, keys are *negotiated pairwise* between two routers. In GETVPN, keys are *distributed* from one source to many routers. Nobody negotiates with anybody. Everyone gets told. ## Proof: One Group, One SA, One SPI If every member really shares the same key, then every member must show the same SPI (Security Parameter Index), the value that identifies an IPsec SA. In a pairwise VPN, every tunnel has its own SPI pair, and two different routers will never show the same one. Here is GM1, a group member at the 10.20.10.0/24 site: ``` GM1#show crypto gdoi Group Name : PLZ-GETVPN Group Identity : 4321 Group Type : GDOI (ISAKMP) Rekeys received : 0 Group Server list : 10.100.0.11 10.100.0.12 Group member : 10.100.0.21 vrf: None Registration status : Registered Registered with : 10.100.0.11 Re-registers in : 692 sec Succeeded registration: 1 TEK POLICY for the current KS-Policy ACEs Downloaded: GigabitEthernet2: IPsec SA: spi: 0xE1F309BF(3790801343) transform: esp-256-aes esp-sha256-hmac sa timing:remaining key lifetime (sec): (3550) Anti-Replay(Counter Based) : 64 encaps: ENCAPS_TUNNEL ``` And here is GM2, a completely different router at the 10.30.10.0/24 site: ``` GM2#show crypto gdoi | include Group member|Registration status|spi:|transform: Group member : 10.100.0.22 vrf: None Registration status : Registered spi: 0xE1F309BF(3790801343) transform: esp-256-aes esp-sha256-hmac ``` **The same SPI. `0xE1F309BF` on both routers.** That is not a coincidence and it is not a display quirk. It is one SA, shared by the whole group. Add a hundred more sites and every one of them will show `0xE1F309BF` until the next rekey, at which point all of them will roll to the same new SPI together. This is the difference between N-squared and N. A twenty-site full mesh with crypto maps is 190 tunnels and 380 SAs to build, monitor, and troubleshoot. A twenty-site GETVPN group is one SA that twenty routers happen to hold a copy of. ### KEK Versus TEK, From the Real Output The same `show crypto gdoi` on GM1 prints both key policies. They do different jobs and it is worth being precise about which is which. ``` KEK POLICY: Rekey Transport Type : Unicast Lifetime (secs) : 849 Encrypt Algorithm : AES Key Size : 256 Sig Hash Algorithm : HMAC_AUTH_SHA Sig Key Length (bits) : 2352 ``` KEK ProtectsRekey messages DirectionKS to members In our labAES 256, unicast, 849s left Signed withThe KS RSA rekey key TEK ProtectsActual user data DirectionAny member to any member In our labesp-256-aes, esp-sha256-hmac Identified bySPI 0xE1F309BF, group wide If an attacker got the TEK, they could read group traffic. If they got the KEK, they could read the rekeys and therefore get every future TEK. That hierarchy is exactly why the KEK gets its own lifetime and its own signature, and why the RSA key that signs rekeys matters so much (a point that becomes a genuine trap once you add a second key server, which is its own article). ## The Hero Capture: Header Preservation Now the part that makes GETVPN click. We ran a packet capture on the link between GM1 and the WAN core while GM1 pinged GM2's LAN, source 10.20.10.1 (GM1's Loopback10) to 10.30.10.1 (GM2's Loopback10). Those are LAN addresses. The routers' WAN addresses are 10.100.0.21 and 10.100.0.22. Here is what was on the wire: ``` No. Time Source Destination Proto Info 1 0.000000 10.20.10.1 10.30.10.1 ESP ESP (SPI=0xe1f309bf) 2 0.010547 10.30.10.1 10.20.10.1 ESP ESP (SPI=0xe1f309bf) 3 0.016432 10.20.10.1 10.30.10.1 ESP ESP (SPI=0xe1f309bf) 4 0.016939 10.30.10.1 10.20.10.1 ESP ESP (SPI=0xe1f309bf) 5 0.020745 10.20.10.1 10.30.10.1 ESP ESP (SPI=0xe1f309bf) 6 0.021049 10.30.10.1 10.20.10.1 ESP ESP (SPI=0xe1f309bf) ``` Read the Source and Destination columns again. Those are the **original LAN addresses**, sitting in the outer IP header of an ESP packet, on the WAN. Not the routers. The payload is genuinely encrypted (it is ESP, with our group SPI), but the addresses are in clear text and they are the endpoints' addresses, not the encryptors' addresses. Compare that with a conventional crypto-map site-to-site VPN, where the outer header is always the two routers: Classic crypto map / VTI Outer header203.0.113.1 to 198.51.100.2 (the routers) Inner headerHidden inside the encrypted payload Transport netOnly needs to route the two public IPs GETVPN Outer header10.20.10.1 to 10.30.10.1 (the hosts) Inner headerA copy of the same addresses Transport netMust be able to route the LAN prefixes The official name for this is **tunnel mode with header preservation**. GETVPN still uses ESP in tunnel mode (the GM output says `encaps: ENCAPS_TUNNEL`), but instead of writing its own outer header it *copies the original one*. You get the ESP encapsulation without the address rewrite. ## The Consequence Nobody Says Out Loud This is the section that matters most, and it is the one that most GETVPN explainers skip. Header preservation is not a neat trick. It is a design constraint with two large, unavoidable consequences, and both of them fall directly out of that capture. ### 1\. The transport network must be able to route your LAN addresses If the outer destination on the wire is 10.30.10.1, then every router between GM1 and GM2 has to know how to forward a packet to 10.30.10.1\. RFC 1918 space, on the WAN, in the clear. On the public internet that is a dead end. No ISP will route 10.30.10.1 for you, and even if your sites had public addresses on their LANs, you would be handing the internet a full map of your internal topology in every packet header. **GETVPN cannot be deployed across the public internet.** Not "should not," not "is discouraged." It structurally cannot, because the transport is being asked to route addresses it does not have. What GETVPN *is* designed for is a private WAN: an [MPLS L3VPN](https://www.pinglabz.com/mpls/) from a carrier, a private leased-line or Metro Ethernet core, or any transport where your own prefixes are already reachable end to end. That is the classic deployment, and it is not an accident. GETVPN exists precisely because enterprises were handed an MPLS VPN by a carrier, told it was "private," and then asked by an auditor to prove the carrier could not read their traffic. GETVPN encrypts the payload while leaving the carrier's routing completely undisturbed. ### 2\. The same property is why there is no tunnel mesh Here is the flip side, and it is the payoff. Why does GETVPN need no tunnels? Because the underlay *already knows* how to get from any site to any other site. The MPLS VPN routes 10.20.10.0/24 to 10.30.10.0/24 natively. GETVPN does not have to build a path, it only has to encrypt what is already flowing along the path the WAN provides. A tunnel exists to create reachability where reachability did not exist, and to carry addresses the transport cannot route. GETVPN has neither problem. So it does not need a tunnel, and because it does not need a tunnel, it does not need a mesh, and because it does not need a mesh, it scales linearly and gives you any-to-any encryption for free. Those two consequences are the same fact seen from opposite ends. You trade internet capability for mesh-free any-to-any encryption. That is the whole bargain, and it is why the technology choice is really a *transport* choice. ### Practical fallout worth knowing - **QoS still works.** Because the original header survives (including DSCP), your carrier's [QoS](https://www.pinglabz.com/qos/) classification on the MPLS core keeps working on encrypted traffic. With a classic tunnel, the carrier sees only your tunnel endpoints and one DSCP value unless you copy it out explicitly. - **Multicast works natively.** A multicast packet keeps its group destination address in the outer header, so the WAN replicates it exactly as it would in clear text. No tunnel replication, no GRE. - **NAT breaks it.** Any device that rewrites addresses between two group members breaks header preservation and therefore breaks the anti-replay and the SA match. - **Your addresses are visible on the WAN.** The payload is confidential, the topology is not. On a private WAN that is usually an acceptable trade. On the internet it is not. ## GETVPN Versus DMVPN Versus Classic IPsec These three get compared constantly and the comparison is usually mushy. Be honest about it: they solve different transport problems, and the right question is not "which is best" but "what is my WAN." GETVPN TopologyAny to any, no tunnels TransportPrivate WAN or MPLS only Over internetNo Native multicastYes Key modelGroup key from a KS ScalingOne SA for the whole group DMVPN TopologyHub and spoke, dynamic spoke to spoke TransportAnything, including the internet Over internetYes, this is its home turf Native multicastVia the hub, GRE encapsulated Key modelPairwise IKE per tunnel ScalingScales well, but tunnels are real Crypto map / VTI TopologyPoint to point, manually built TransportAnything Over internetYes Native multicastNo, needs GRE Key modelPairwise IKE per peer ScalingN-squared, the reason GETVPN exists The honest summary: if your WAN is the internet, you want [DMVPN](https://www.pinglabz.com/dmvpn/) (or FlexVPN, or SD-WAN). If your WAN is a private core or a carrier MPLS VPN and you need any-to-any encryption without building a mesh, you want GETVPN. If you have two sites and a deadline, a point-to-point VTI is fine and nobody will judge you. These are complements, and large enterprises routinely run GETVPN on the MPLS underlay and DMVPN on the internet underlay simultaneously. ## What the Group Member Config Actually Looks Like The theory is only convincing if the config matches it. Here is the entire GETVPN configuration on GM1, straight from the lab: ``` crypto isakmp key PingLabz-GETVPN-01 address 10.100.0.11 crypto isakmp key PingLabz-GETVPN-01 address 10.100.0.12 crypto gdoi group PLZ-GETVPN identity number 4321 server address ipv4 10.100.0.11 server address ipv4 10.100.0.12 crypto map PLZ-GET-CMAP 10 gdoi set group PLZ-GETVPN interface GigabitEthernet2 crypto map PLZ-GET-CMAP ``` Count what is missing. No peer address. No crypto ACL. No transform set. No mention of GM2, or of 10.30.10.0/24, or of any other site at all. The only IP addresses in the whole block are the two key servers. GM2's config is identical, and site number one hundred would get the same paste. The policy (what to encrypt, and to where) is not on the group member. It is downloaded, and you can see it arrive: ``` ACL Downloaded From KS 10.100.0.11: access-list permit ip 10.20.10.0 0.0.0.255 10.30.10.0 0.0.0.255 access-list permit ip 10.30.10.0 0.0.0.255 10.20.10.0 0.0.0.255 access-list permit ip 192.168.99.0 0.0.0.255 10.30.10.0 0.0.0.255 access-list permit ip 10.30.10.0 0.0.0.255 192.168.99.0 0.0.0.255 ``` That ACL is not configured on GM1\. It came from the key server over GDOI. Change the group policy once, on the KS, and every member in the group picks it up at the next rekey. That is the operational win, and it is what people mean when they say GETVPN "centralises policy." The full build (key server first, then members, then verification) is walked line by line in the [GETVPN configuration article](https://www.pinglabz.com/getvpn-configuration-ios-xe/). ## Does It Actually Pass Traffic Yes, and encrypted, in both directions, first try: ``` GM1#ping 10.30.10.1 source Loopback10 repeat 8 !!!!!!!! Success rate is 100 percent (8/8), round-trip min/avg/max = 5/9/30 ms GM2#ping 10.20.10.1 source Loopback10 repeat 5 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 5/11/34 ms ``` And on the key server side, the group is visible as exactly what it is: a membership list. ``` KS1#show crypto gdoi ks members Group Member ID : 10.100.0.21 GM Version: 1.0.26 Group ID : 4321 Group Name : PLZ-GETVPN GM State : Registered Key Server ID : 10.100.0.11 Group Member ID : 10.100.0.22 GM Version: 1.0.26 Group ID : 4321 GM State : Registered Key Server ID : 10.100.0.11 ``` Two members, one group, one key. No tunnels anywhere in the topology. ## The Obvious Objection: Isn't the Key Server a Single Point of Failure? Half yes, and the half that is "yes" is smaller than you think. The key server is a **control plane** element. It is not in the data path, which the peer 0.0.0.0 output already hinted at. If the key server disappears, group members carry on encrypting with the TEK they already hold. Traffic keeps flowing. What they lose is the ability to *re-register* and to *receive a rekey*, so the group degrades on a timer rather than falling over instantly. We tested this by shutting the key server down, and the pings stayed at 100 percent. The real answer is COOP (cooperative key servers): run two or more, with priorities, and let them elect a primary. It works, and there is one trap involving the RSA rekey key that will silently kill your group at the next rekey interval if you get it wrong. We hit it on purpose. That is covered in [GETVPN redundancy: COOP key servers and rekey survival](https://www.pinglabz.com/getvpn-coop-key-servers/). ## Key Takeaways - **GETVPN has no peer and no tunnel.** `show crypto session` on a group member reports `Peer: 0.0.0.0 port 848`. Port 848 is GDOI, the group key distribution protocol. The session is with a policy, not with a router. - **One group, one key, one SPI.** GM1 and GM2 both showed SPI `0xE1F309BF`. Every member of the group shares the same TEK, so there is one SA no matter how many sites you add. - **Header preservation is the defining property.** The ESP outer header carried 10.20.10.1 to 10.30.10.1, the original LAN addresses, not the routers' WAN addresses. A classic crypto map would have shown the routers instead. - **That is why GETVPN needs a private WAN.** The transport must be able to route your internal prefixes, because they are on the wire. GETVPN over the public internet is not a bad idea, it is structurally impossible. MPLS L3VPN is the classic deployment. - **And it is why there is no mesh.** The underlay already provides any-to-any reachability, so GETVPN only has to encrypt, not to build paths. Mesh-free any-to-any encryption is what you buy with the transport constraint. - **KEK protects rekeys, TEK protects data.** Both are pushed from the key server. The group member's own config contains no policy at all, only the address of the key servers. - **The group member config is the shortest in the IPsec family.** Two ISAKMP keys, a GDOI group, a crypto map with `set group`. Identical on every site. That is the scaling story in six lines. - **Pick by transport.** Private WAN or MPLS and you need any-to-any: GETVPN. Internet underlay: DMVPN. Two sites: a VTI. They are complements, not competitors. Everything above was captured on Cisco IOS XE (cat8000v 17.18.02) in a lab with two key servers, two group members, and a shared private WAN core. For the rest of the family (IKEv1, IKEv2, crypto maps, VTIs, DMVPN, FlexVPN, certificate authentication, and VRF-aware designs), start at the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/), then build this one yourself with the [GETVPN configuration walkthrough](https://www.pinglabz.com/getvpn-configuration-ios-xe/). ### Running a Cisco Router as a CA Server: PKI You Can Lab URL: https://www.pinglabz.com/cisco-router-ca-server/ Last updated: 2026-07-13T09:01:24.000Z PKI is one of those topics that everyone nods along to and almost nobody has actually built. You read about certificate authorities, trustpoints, enrollment, and revocation, and it all stays abstract because the tooling feels heavy: a Windows domain, AD CS, a public CA that wants money and a DNS record. So the chapter gets skimmed, and then one day a certificate-authenticated IPsec tunnel refuses to come up and you are debugging something you never actually saw work. Here is the shortcut. You do not need any of that. A Cisco IOS XE router will happily act as a full certificate authority, issue real X.509 certificates, and hand them to other routers over HTTP. It takes about six lines of config. In our CCIE Security lab we built exactly that on a cat8000v called CA1, issued real certificates from it, and then used those certificates to authenticate an IKEv2 tunnel with no pre-shared key anywhere in the configuration. If you are working through the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/), this is the missing piece that makes certificate authentication stop being theory. ## What a Certificate Authority Actually Does Strip away the ceremony and a CA does two things. First, it issues itself a certificate. This is the root, and it is self-signed: the CA vouches for itself, because there is nobody above it to do the vouching. In our lab this was serial number 01\. That self-signed CA certificate is what every client will eventually install and trust. Everything downstream hangs off it. Second, it signs certificate signing requests (CSRs) from clients. A router generates an RSA keypair, keeps the private key, wraps the public key plus a subject name into a CSR, and sends it to the CA. The CA checks it (or, in a lab, does not check it at all), signs it with the CA private key, and hands back an identity certificate. In our lab GM1 got serial 02 and GM2 got serial 03. ## The Config: Six Lines and a Reload of Your Assumptions This is the CA server we built on CA1 (a cat8000v at 10.100.0.31 on the lab WAN core): ``` CA1(config)#crypto pki server PLZ-CA database level complete issuer-name CN=PLZ-CA,O=PingLabz,C=US grant auto lifetime certificate 730 lifetime ca-certificate 1825 no shutdown ``` Line by line: crypto pki server PLZ-CA Creates the CA and names it. The name matters: it becomes the trustpoint the CA uses for its own key, and it is what you reference in every `show` command. database level complete Write a full copy of every issued certificate to storage as `.cer`. The alternatives (`minimal`, `names`) store less. You want `complete` if you ever intend to revoke anything, because you need the certificate to revoke it. issuer-name CN=PLZ-CA,O=PingLabz,C=US The distinguished name that appears as the Issuer field on every certificate this CA signs. This is the CA's identity on the wire. grant auto Sign every CSR that arrives, automatically, no questions asked. Wonderful for a lab. Read the warning below before you even think about it anywhere else. lifetime certificate 730 Issued identity certificates are valid for 730 days (two years). Shorter lifetimes are better security hygiene, and worse operational ergonomics if you have no automatic re-enrollment. lifetime ca-certificate 1825 The CA's own root certificate lives for 1825 days (five years). It must outlive every certificate it signs, or you get a very bad day at renewal time. ### About grant auto, honestly `grant auto` means the CA signs anything that asks. Any device that can reach TCP/80 on the CA and speak SCEP gets a valid certificate with whatever subject name it claimed. In a closed lab that is exactly what you want, because the alternative is manually approving every enrollment while you are trying to learn something else. In production it is the opposite of what you want. The correct setting is `grant none`, which parks every request in a queue. You then inspect the request, confirm out of band that the device asking is the device you think it is, and approve it explicitly with `crypto pki server PLZ-CA grant `. That manual step is not bureaucracy, it is the entire point: a CA that signs anything is a CA that will happily issue a certificate to an attacker who plugged a laptop into your network and claimed to be a router. The certificate would be cryptographically perfect and completely worthless. ## Before: Configured but Disabled Type the config but do not enable the server, and this is what you get. This is real output from CA1: ``` CA1#show crypto pki server Certificate Server PLZ-CA: Status: disabled State: initial Server's configuration is unlocked (enter "no shut" to lock it) Issuer name: CN=PLZ-CA,O=PingLabz,C=US CA cert fingerprint: -Not found- Granting mode is: auto ``` Four things are being told to you here. **Status: disabled / State: initial.** The CA exists as configuration but has done nothing. It has no keypair and no certificate. **Server's configuration is unlocked (enter "no shut" to lock it).** This is the interesting one. While the server is shut down, you can freely change the issuer name, the lifetimes, the whole shape of the CA. The moment you enable it, the identity-defining parameters lock, because the CA is about to generate a self-signed certificate that bakes them in. You cannot rename an issuer after it has signed certificates without invalidating everything downstream. IOS enforces that for you. **CA cert fingerprint: -Not found-.** There is no CA certificate yet, so there is nothing to fingerprint. This field is going to matter enormously in a minute. **Granting mode is: auto.** Exactly what we asked for, and exactly what we would not ask for in production. ## After: A Live CA With a Real Fingerprint Now enable it with `no shutdown`, answer the prompts, and look again: ``` CA1#show crypto pki server Certificate Server PLZ-CA: Status: enabled State: enabled Server's configuration is locked (enter "shut" to unlock it) Issuer name: CN=PLZ-CA,O=PingLabz,C=US CA cert fingerprint: 50AE2189 0FC62684 1418DEF0 B0EB714A Granting mode is: auto Last certificate issued serial number (hex): 3 CA certificate expiration timer: 09:04:14 UTC Jul 12 2031 CRL NextUpdate timer: 15:04:24 UTC Jul 13 2026 Current primary storage dir: nvram: Database Level: Complete - all issued certs written as .cer ``` Walk the new fields: **Status: enabled / State: enabled.** The CA has generated its RSA keypair, signed its own root certificate, and is now listening for enrollment requests over SCEP on HTTP. **Server's configuration is locked.** As promised. The issuer name is now cast in the root certificate. **CA cert fingerprint: 50AE2189 0FC62684 1418DEF0 B0EB714A.** Memorize this idea, not the value. This is the fingerprint of the CA's self-signed root certificate. When a client later runs `crypto pki authenticate`, IOS will display a fingerprint and ask whether you accept the certificate. *That* fingerprint must match *this* one. Comparing them, out of band, is the only thing standing between you and trusting an impostor CA. We go into this in depth in the walkthrough on [certificate authentication for IPsec](https://www.pinglabz.com/ipsec-certificate-authentication/), but note that the value you compare against originates right here, on the CA itself. **Last certificate issued serial number (hex): 3.** Three certificates exist in this CA's world. Serial 01 is the CA's own self-signed root. Serial 02 is GM1's identity certificate. Serial 03 is GM2's. Serials are sequential and never reused, because a serial number plus an issuer name is the unique key for a certificate, and that is precisely what a revocation list refers to. **CA certificate expiration timer: 09:04:14 UTC Jul 12 2031.** Five years out, which is our `lifetime ca-certificate 1825`. When this expires, every certificate chaining to it dies with it. Put it in a calendar. **CRL NextUpdate timer.** The CA is publishing a certificate revocation list and telling clients when to come back for a fresh copy. In our lab nothing checks it (more on that shortly), but the CA is dutifully maintaining it anyway. **Current primary storage dir: nvram:.** The certificate database, and the CA's private key, live in the router's NVRAM. Sit with that for a second. We will come back to it. ## Gotcha 1: no shutdown Is Interactive, and That Kills Automation You cannot push a CA server config from a script. Not with a naive config push, anyway. `no shutdown` under `crypto pki server` is an interactive command. IOS stops and prompts you for a passphrase, which is used to protect the CA's private key in NVRAM. It then asks you to confirm it. A configuration-automation tool that streams lines at a device and does not handle prompts will hang, time out, or silently leave the CA disabled while reporting success. We hit exactly this: the config went in fine, the server stayed down, and we ended up driving `no shutdown` from a real interactive SSH session to get past the passphrase prompt. The lesson generalizes past PKI. Crypto commands in IOS have a habit of being interactive precisely because they involve secrets that should not sit in a config file, and secrets that should not sit in a config file are exactly the things your automation pipeline is worst at handling. If you are building CA bring-up into an automated workflow, you need something that speaks expect-style prompt handling, not a template renderer. ## Gotcha 2: PKI Needs a Clock. Set It First. This is the one that produces the most baffling failures, and here is the router telling you about it in plain English: ``` %PKI-2-NON_AUTHORITATIVE_CLOCK: PKI functions can not be initialized until an authoritative time source, like NTP, can be obtained. ``` Certificates are time-bound objects. Every certificate carries a start date and an end date, and every validation checks the current time against them. If a device's clock is wrong, one of two things happens: it decides a perfectly valid certificate is not yet valid, or it decides an expired certificate is fine. Both are wrong, and neither of them produces an error message that says "your clock is wrong". You get an authentication failure, or a tunnel that will not establish, and you go hunting through crypto debugs for an hour. Do this before you touch any PKI command, on the CA *and* on every client: ``` clock set 09:00:00 Jul 13 2026 ntp server 10.100.0.11 ``` Set the clock manually to get past the bootstrap problem, then point at NTP so it stays right. In a lab, one device with a stable clock acting as the NTP master for everyone else is enough. In production, this is non-negotiable infrastructure: your PKI is only as trustworthy as your time. ## What a Lab CA Is Not Everything above is genuinely real PKI. The certificates are real X.509 certificates, the signatures are real RSA signatures, and the authentication that follows is real cryptographic authentication. But be honest with yourself about the gap between this and a CA you would put behind a bank. No HSM The CA private key sits in the router's NVRAM, protected by a passphrase you typed. A real root CA keeps its key in a hardware security module that will not export it under any circumstances, and lives offline in a safe. Ours lives in a VM. No offline root Serious PKI uses an offline root that signs one thing (an intermediate CA) and then goes back in the safe. The online intermediate does the day-to-day issuing, so a compromise does not burn the root. We have one CA doing both jobs. No real revocation The CA maintains a CRL, but nothing distributes it properly and, in our lab, no client checks it (`revocation-check none`). Revocation that nobody checks is revocation that does not exist. grant auto Anything that can reach the CA gets a certificate. There is no identity proofing, no registration authority, no human in the loop. That is a fine trade in a closed lab and an unacceptable one anywhere else. Single point of failure One router. Lose it and you lose the CA key, the certificate database, and your ability to issue or revoke anything. No backup, no key escrow, no recovery plan beyond rebuilding from scratch. Where it IS good Learning PKI properly. CCIE and CCNP Security lab work. Small internal deployments where the router population is closed, known, and you control every device on the network. ## What To Do With It Next A CA on its own is inert. It becomes interesting the moment a client enrolls with it and then uses the resulting certificate to prove its identity to a peer. That is the payoff, and it is the subject of the follow-up: [certificate authentication for IPsec](https://www.pinglabz.com/ipsec-certificate-authentication/), where GM1 and GM2 enroll against this exact CA and then bring up an IKEv2 tunnel whose configuration contains no pre-shared key at all. The IKEv2 SA on that tunnel says `Auth sign: RSA` instead of `Auth sign: PSK`, and that single word is what all of this was for. If IKEv2 itself is still fuzzy, start with [how IKEv2 works](https://www.pinglabz.com/ikev2-explained/) and the wider [IPsec VPN guide](https://www.pinglabz.com/ipsec-vpn/), which covers the full progression from pre-shared keys through certificates to GETVPN. ## Key Takeaways - An IOS XE router is a fully functional certificate authority. Six lines of config under `crypto pki server` and you have a working CA issuing real X.509 certificates. No AD CS, no public CA, no money. - The CA issues itself a self-signed root certificate (serial 01 in our lab), then signs client CSRs sequentially (02 for GM1, 03 for GM2). Serial number plus issuer name is what uniquely identifies a certificate, and it is what a CRL revokes. - `Server's configuration is unlocked` becomes `locked` when you enable the server, because the issuer name is now baked into the self-signed root. Get the issuer name right before you enable. - The `CA cert fingerprint` from `show crypto pki server` is the value clients must verify out of band when they authenticate the trustpoint. It is the anchor of the entire trust chain. - `no shutdown` on a crypto PKI server is interactive: it prompts for a passphrase protecting the CA private key. A config-push tool that cannot handle prompts will not bring your CA up. Drive it from a real session. - PKI needs an authoritative clock. `%PKI-2-NON_AUTHORITATIVE_CLOCK` means PKI will not even initialize. Set the clock or run NTP on the CA and every client before you touch a single PKI command. - `grant auto` is a lab convenience that signs anything that asks. Production wants `grant none` and manual approval of every request. - A router CA has no HSM, no offline root, no meaningful revocation distribution, and its private key lives in NVRAM. It is excellent for learning and internal labs. It is not a production trust anchor. ### FlexVPN Explained: One Framework for Site-to-Site, Hub-Spoke, and Remote Access URL: https://www.pinglabz.com/flexvpn-explained/ Last updated: 2026-07-13T08:08:31.000Z Every Cisco engineer eventually inherits a router with four different VPN styles bolted onto it: a crypto map for the legacy branch, a static VTI for the partner link, DMVPN for the regional spokes, and EzVPN for the handful of teleworkers who never got a proper laptop image. Four configuration models, four troubleshooting workflows, four sets of gotchas. FlexVPN exists to collapse all of that into one. It is the modern Cisco answer to the question "how should I build IPsec on IOS XE today", and it sits at the top of the [IPsec VPN](https://www.pinglabz.com/ipsec-vpn/) cluster for a reason. The single most surprising thing about FlexVPN, and the thing that makes it click once you see it, is this: **IKEv2 itself carries the addresses and the routes**. There is no IGP running over the overlay. No OSPF adjacency, no EIGRP neighbor, no BGP session. The tunnel comes up, and the routes are simply there. This article walks through why that works, using real output from our CML lab (HUB, SPOKE1, and SPOKE2, all cat8000v running IOS XE 17.18.02). ## The Zoo FlexVPN Replaces Before FlexVPN, Cisco IPsec was a collection of unrelated configuration dialects. Each one solved a real problem, and each one was configured completely differently from the others. Crypto Maps **Built for:** Classic site-to-site between two known peers. **Interesting traffic:** An ACL bolted to a physical interface. **Pain:** No tunnel interface, so no routing protocol, no QoS, no per-tunnel stats. Every new subnet means editing the crypto ACL on both ends. Static VTI **Built for:** Site-to-site with a real routable interface. **Interesting traffic:** Anything routed into the tunnel. No crypto ACL. **Pain:** One tunnel interface per peer, statically configured. Fine for five sites, miserable for five hundred. DMVPN **Built for:** Large hub-and-spoke with dynamic spoke-to-spoke. **Interesting traffic:** Routed into a single mGRE tunnel; NHRP maps overlay to underlay. **Pain:** Needs a routing protocol over the overlay, plus NHRP, plus GRE, plus IPsec profiles. Three technologies stacked to make one design work. EzVPN **Built for:** Remote access. A client pulls an address and a policy from a server. **Interesting traffic:** Pushed down as split-tunnel policy from the head end. **Pain:** A completely separate config model from every site-to-site option, and it never scaled cleanly into site-to-site territory. Look at what those four have in common: every single one is IPsec, and every single one needs to authenticate a peer, negotiate policy, decide what traffic to protect, and get routes across. FlexVPN's insight is that all four are the same problem wearing different clothes, and that [IKEv2](https://www.pinglabz.com/ikev2-explained/) already has the machinery to solve it in one model. ## FlexVPN Is a Framework, Not a Protocol This trips people up constantly, so let us be blunt: **there is no FlexVPN protocol**. Nothing on the wire says "FlexVPN". Run a packet capture on a FlexVPN tunnel and you will see standards-based IKEv2 (RFC 7296) and standards-based ESP. That is it. FlexVPN is Cisco's name for a *configuration framework*: a consistent set of IOS XE building blocks that use IKEv2's own features (the configuration payload, the authorization exchange, and IKEv2 route injection) to do the work that crypto maps, static VTIs, DMVPN, and EzVPN each did their own way. Because the framework leans on IKEv2 rather than on Cisco-proprietary glue, the same config model covers a hub-and-spoke fabric, a point-to-point site-to-site tunnel, and a remote-access client. That is the "flex" part. ## The Building Blocks A FlexVPN head end is assembled from five pieces. Learn these five and the rest is variation. IKEv2 profile **Job:** Match the remote peer's identity, decide how both sides authenticate, and bind the session to a virtual-template and an authorization list. This is the front door. IKEv2 authorization policy **Job:** The payload of attributes handed to the peer once it authenticates. Address pool, DNS, and critically the **routes**. This is where the magic lives. AAA **Job:** The lookup mechanism that connects the IKEv2 profile to the authorization policy. Can be `local` (our lab) or a RADIUS server (production, per-user policy). Virtual-Template **Job:** A blueprint interface. It is never used to pass traffic itself. It defines what a per-peer tunnel should look like: source, mode, IPsec protection, NHRP settings. Virtual-Access **Job:** The real interface, cloned from the template automatically when a peer authenticates, and destroyed when the session tears down. One per peer. You never configure it by hand. ## The Headline Feature: IKEv2 Pushes the Routes Here is the hub configuration from our lab. Read the authorization policy carefully, because those three lines are the entire routing design. ``` aaa new-model aaa authorization network PLZ-FLEX-LIST local ip local pool PLZ-FLEX-POOL 10.0.9.10 10.0.9.100 ip access-list standard PLZ-FLEX-ROUTES permit 10.20.10.0 0.0.0.255 permit 192.168.99.0 0.0.0.255 permit 10.30.0.0 0.0.255.255 crypto ikev2 authorization policy PLZ-AUTHOR-POL pool PLZ-FLEX-POOL route set interface route set access-list PLZ-FLEX-ROUTES crypto ikev2 profile PLZ-FLEX-PROF match identity remote fqdn domain pinglabz.lab identity local fqdn hub.pinglabz.lab authentication local pre-share authentication remote pre-share keyring local PLZ-FLEX-KEYRING aaa authorization group psk list PLZ-FLEX-LIST PLZ-AUTHOR-POL virtual-template 1 ``` `route set interface` tells IKEv2 to advertise the address on the tunnel interface itself. `route set access-list PLZ-FLEX-ROUTES` tells it to advertise every prefix permitted by that standard ACL. Those routes are then delivered to the peer inside the IKEv2 authorization exchange, as attributes, during tunnel setup. Once both spokes connect, the hub's routing table looks like this: ``` HUB#show ip route | include Virtual-Access S 10.0.9.24/32 is directly connected, Virtual-Access1 S 10.0.9.25/32 is directly connected, Virtual-Access2 S 10.30.10.0/24 is directly connected, Virtual-Access1 S 10.30.20.0/24 is directly connected, Virtual-Access2 ``` SPOKE1's LAN (10.30.10.0/24) is reachable via Virtual-Access1\. SPOKE2's LAN (10.30.20.0/24) is reachable via Virtual-Access2\. The hub learned both, plus each spoke's assigned tunnel address, without a single line of routing protocol configuration. The spoke side is the mirror image. Here is SPOKE1's static route table (and remember, the only static route *we* typed was the default toward the ISP): ``` SPOKE1#show ip route static S* 0.0.0.0/0 [1/0] via 198.51.100.1 S 10.20.10.0/24 is directly connected, Tunnel0 S 10.30.0.0/16 is directly connected, Tunnel0 S 10.255.0.1/32 is directly connected, Tunnel0 S 192.168.99.0/24 is directly connected, Tunnel0 ``` Every one of those tunnel routes came from the hub. 10.20.10.0/24 and 192.168.99.0/24 and 10.30.0.0/16 are exactly the three lines in `PLZ-FLEX-ROUTES`, and 10.255.0.1/32 is the hub's tunnel address, delivered by `route set interface`. Say it plainly: **no OSPF, no EIGRP, no BGP**. The IKEv2 AAA authorization policy injected all of it. That has real operational consequences. Adding a new subnet to the hub site is one line in a standard ACL, and every spoke picks it up on its next rekey or reconnect. There is no adjacency to babysit, no overlay flap to chase, no "why is my EIGRP neighbor bouncing over the DMVPN cloud" ticket. For a hub-and-spoke topology where the routing is fundamentally hierarchical anyway, running a full IGP over the tunnel was always a bit of an overreach. ## Dynamic Addressing Is the Same Mechanism as Remote Access Look at the spoke's tunnel interface: ``` interface Tunnel0 ip address negotiated ip nhrp network-id 1 ip nhrp shortcut virtual-template 1 tunnel source GigabitEthernet2 tunnel mode ipsec ipv4 tunnel destination 203.0.113.1 tunnel protection ipsec profile PLZ-FLEX-IPSEC ``` `ip address negotiated`. The spoke has no idea what its own tunnel address is until IKEv2 tells it. On the hub, `ip local pool PLZ-FLEX-POOL 10.0.9.10 10.0.9.100` plus the `pool` line in the authorization policy hands one out. Check the pool after both spokes are up: ``` HUB#show ip local pool PLZ-FLEX-POOL Pool Begin End Free In use PLZ-FLEX-POOL 10.0.9.10 10.0.9.100 89 2 Inuse addresses: 10.0.9.24 IKEv2 Addr IDB 10.0.9.25 IKEv2 Addr IDB ``` Note the label: **IKEv2 Addr IDB**. IOS XE is telling you exactly who allocated those addresses. This is the IKEv2 *configuration payload* (CFG\_REQUEST / CFG\_REPLY in RFC 7296), and it is precisely the same mechanism a remote-access VPN client uses when it dials into a head end and gets an inside address, a DNS server, and a split-tunnel list. That is the whole reason one framework can cover both worlds. To the hub, a branch router with `ip address negotiated` and a laptop running an IKEv2 client are the same kind of thing: a peer that authenticates, requests configuration, and gets an address and a set of routes. Swap `pre-share` for certificates and EAP, point the AAA list at RADIUS instead of `local`, and the identical head-end config serves remote-access users. Nothing else changes. ## Virtual-Access: Cloning a Tunnel Per Peer The hub never has a Tunnel interface for each spoke. It has one `Virtual-Template1`, and IOS XE clones it: ``` interface Virtual-Template1 type tunnel ip unnumbered Loopback0 ip nhrp network-id 1 ip nhrp redirect tunnel source GigabitEthernet2 tunnel mode ipsec ipv4 tunnel protection ipsec profile PLZ-FLEX-IPSEC ``` When SPOKE1 authenticates, IOS XE stamps out a copy of that template as Virtual-Access1, fills in the tunnel destination from the IKEv2 session, and brings it up. When SPOKE2 authenticates, it gets Virtual-Access2\. Both spokes up: ``` HUB#show ip interface brief | include Virtual-Access Virtual-Access1 10.255.0.1 YES unset up up Virtual-Access2 10.255.0.1 YES unset up up HUB#show crypto session Interface: Virtual-Access1 Session status: UP-ACTIVE Peer: 198.51.100.2 port 500 Interface: Virtual-Access2 Session status: UP-ACTIVE Peer: 198.51.100.6 port 500 ``` Both Virtual-Access interfaces share the same IP (10.255.0.1) because the template is `ip unnumbered Loopback0`. That is normal and correct. The per-peer state that makes them distinct is the tunnel destination, the IPsec SA, and the routes installed behind them. Add a hundred spokes and you get a hundred Virtual-Access interfaces, from that one template, with no additional hub configuration. ## The Gotcha We Actually Hit Our first attempt at the Virtual-Template omitted `tunnel source` and `tunnel mode`. IKEv2 authenticated perfectly. The debug showed the authorization request going through cleanly. And yet the tunnel refused to stay up, cycling in a loop: ``` %LINK-3-UPDOWN: Interface Virtual-Access1, changed state to down %DMVPN-5-NHRP_NETID_UNCONFIGURED: Virtual-Access1: NETID : 1 Unconfigured (Tunnel: 0.0.0.0 NBMA: UNKNOWN) ``` Read that message closely: `Tunnel: 0.0.0.0 NBMA: UNKNOWN`. The cloned Virtual-Access had no idea what its own source address was, because the template never told it. It came up, discovered it was an incomplete tunnel, and died. Then IKEv2 (still perfectly happy) rebuilt the session, which cloned a new Virtual-Access, which also died. The trap is that this *looks* like a crypto failure. You are staring at a VPN that will not come up, so you go debug IKEv2, and IKEv2 tells you everything is fine (you will see `IKEv2:Using mlist PLZ-FLEX-LIST and username PLZ-AUTHOR-POL for group author request` in the debug output, which is IKEv2 saying "authorization done, over to you"). It is not a crypto problem. It is an interface problem. The fix was adding the two missing lines to the template: ``` interface Virtual-Template1 type tunnel tunnel source GigabitEthernet2 tunnel mode ipsec ipv4 ``` The general lesson: in FlexVPN, a Virtual-Access is only ever as complete as the Virtual-Template it was cloned from. If the tunnel is flapping but IKEv2 says READY, stop debugging crypto and go read `show derived-config interface Virtual-Access1` to see what IOS XE actually built. ## FlexVPN vs DMVPN These two overlap heavily, and the honest answer is that both are good. Neither one is deprecated, and DMVPN is not going anywhere. DMVPN **Transport:** mGRE plus NHRP, with IPsec layered on for protection. **Routing:** A real IGP over the overlay (EIGRP, OSPF, or BGP). You own it, you tune it, you troubleshoot it. **Remote access:** Not its job. **Maturity:** Very high. Enormous installed base, deep documentation, a decade of field scar tissue. **Spoke-to-spoke:** Proven and routine (Phase 2 and Phase 3). FlexVPN **Transport:** IKEv2-native. IPsec VTI, no GRE required. **Routing:** Injected by IKEv2 from the authorization policy. No IGP needed. **Remote access:** Same framework. Site-to-site and RA share one head-end model. **Maturity:** Solid on modern IOS XE, but a smaller installed base and thinner community troubleshooting knowledge. **Spoke-to-spoke:** Supported via NHRP shortcut, and it is the piece most likely to bite you on a given platform and release. The short version: if you already run DMVPN at scale and it works, FlexVPN is not a reason to rip it out. If you are building green field on modern IOS XE, or if you need site-to-site and remote access to share one head end and one policy engine, FlexVPN is the cleaner starting point. If you want the deep dive on the other option, our [DMVPN cluster](https://www.pinglabz.com/dmvpn/) covers mGRE, NHRP, and all three phases with working lab output. ## Where FlexVPN Fits Four deployment shapes, one framework: - **Hub-and-spoke.** The canonical FlexVPN design and the one in this article. Hub runs a virtual-template, spokes run a static Tunnel with a tunnel destination, IKEv2 hands out addresses and routes. - **Point-to-point site-to-site.** A degenerate hub-and-spoke with one spoke, or a pair of routers that each run a static tunnel. Covered step by step in [FlexVPN site-to-site configuration](https://www.pinglabz.com/flexvpn-site-to-site-configuration/). - **Remote access.** Same head end, certificates or EAP instead of a pre-shared key, AAA pointed at RADIUS. The client is just another peer requesting a config payload. - **Spoke-to-spoke.** NHRP shortcut and redirect let two spokes build a direct tunnel and bypass the hub. Worth being honest here: in our lab, on cat8000v 17.18.02, spoke-to-spoke traffic worked but only by hairpinning through the hub, and the NHRP shortcut never formed (every NHRP counter stayed at zero). We wrote up exactly what we could and could not prove in [FlexVPN spoke-to-spoke](https://www.pinglabz.com/flexvpn-spoke-to-spoke/), including the traceroute that shows the two-hop path. ## Key Takeaways - **FlexVPN is not a protocol.** It is a Cisco configuration framework built entirely on standards-based IKEv2 and ESP. Nothing proprietary crosses the wire. - **One model replaces four.** Crypto maps, static VTIs, DMVPN, and EzVPN each solved one shape of the problem. FlexVPN covers all of them with the same building blocks. - **IKEv2 carries the routes.** `route set interface` and `route set access-list` in the IKEv2 authorization policy inject prefixes into the peer's routing table during tunnel setup. Our hub learned both spoke LANs, and both spokes learned the hub's networks, with no OSPF, EIGRP, or BGP anywhere. - **IKEv2 carries the addresses too.** `ip address negotiated` on the spoke plus `ip local pool` on the hub is the IKEv2 configuration payload, the same mechanism that gives a remote-access client its inside address. That is why one framework spans site-to-site and RA. - **Virtual-Access interfaces are cloned, not configured.** One Virtual-Template becomes one Virtual-Access per authenticated peer, created and destroyed with the session. - **If the tunnel flaps but IKEv2 says READY, it is not crypto.** A Virtual-Template missing `tunnel source` or `tunnel mode` clones a broken Virtual-Access that comes up and dies in a loop, and the `NHRP_NETID_UNCONFIGURED` message is your clue. - **DMVPN is still fine.** Pick FlexVPN for green field, for IKEv2-native designs, and when remote access and site-to-site should share a head end. Pick DMVPN when you need a battle-tested spoke-to-spoke fabric with a routing protocol you already run. FlexVPN is the piece that makes the rest of the modern Cisco IPsec story hang together. If you want the full picture, from ESP and IKE fundamentals through crypto maps, VTIs, and IKEv2 itself, start at the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/) and work outward. ### FlexVPN Spoke-to-Spoke: Dynamic Tunnels Without DMVPN URL: https://www.pinglabz.com/flexvpn-spoke-to-spoke/ Last updated: 2026-07-13T08:08:32.000Z Every hub-and-spoke [IPsec VPN](https://www.pinglabz.com/ipsec-vpn/) has the same structural flaw: two branch offices that want to talk to each other cannot. Not directly, anyway. Their traffic climbs the tunnel to the hub, gets decrypted, gets routed, gets re-encrypted, and comes back down a different tunnel. It works, but it burns CPU on a box that has better things to do and it adds a leg to the path that has nothing to do with where the packets are actually going. FlexVPN's answer is NHRP redirect and shortcut switching: the hub tells the spokes "you two can reach each other directly", and the spokes build a dynamic IPsec tunnel between themselves. That is the design. This article walks through it properly, shows the exact configuration, and then tells you what happened when we built it on a Cisco cat8000v running IOS XE 17.18.02. Here is the part most articles will not tell you: **the shortcut never formed.** Hub-and-spoke worked perfectly. IKEv2 route injection worked beautifully. Spoke-to-spoke traffic passed at 100 percent. But every packet hairpinned through the hub, and the NHRP counters on every device stayed at zero. We are not going to pretend otherwise, because if you are reading this at 2am with a traceroute that still says two hops, a sanitised success story is useless to you. What you need is a way to tell the difference between a shortcut that is working and a shortcut that never fired. That is section four, and it is the reason this article exists. ## The Problem: The Hub Is In The Middle Of Everything Our topology is a standard FlexVPN hub-and-spoke. A hub on 203.0.113.1, two spokes on 198.51.100.2 and 198.51.100.6, and an ISP router in between. The hub's LAN is 10.20.10.0/24, SPOKE1's is 10.30.10.0/24, SPOKE2's is 10.30.20.0/24\. Both spokes are up, both got a tunnel address from the hub's pool, and the IKEv2 authorization policy pushed every route into place without a single line of OSPF, EIGRP, or BGP. Now watch what happens when SPOKE1 pings SPOKE2's LAN: ``` SPOKE1#ping 10.30.20.1 source Loopback10 repeat 20 !!!!!!!!!!!!!!!!!!!! Success rate is 100 percent (20/20), round-trip min/avg/max = 4/7/15 ms ``` Twenty out of twenty. It works. So what is the complaint? Trace the path: ``` SPOKE1#traceroute 10.30.20.1 source Loopback10 probe 1 timeout 1 Type escape sequence to abort. Tracing the route to 10.30.20.1 1 10.255.0.1 34 msec <-- the HUB 2 10.0.9.25 81 msec <-- SPOKE2's tunnel address ``` Two hops. Hop one is 10.255.0.1, the hub's tunnel address. Hop two is 10.0.9.25, which is the address the hub's local pool handed to SPOKE2\. Every single packet between those two branches is being decrypted by the hub, looked up in the routing table, and re-encrypted into the other spoke's tunnel. That costs you three things: - **Latency.** The packet takes a detour through a third site. If your hub is in a datacentre in another region and your two spokes are in the same city, that detour is the entire round-trip time. - **Hub bandwidth and CPU.** Every spoke-to-spoke byte crosses the hub's WAN link twice (in and out) and gets processed by the crypto engine twice (decrypt and re-encrypt). Scale that across fifty branches and the hub is doing crypto for traffic it has no interest in. - **Blast radius.** The hub becomes a single point of failure for traffic that never needed to touch it. None of that is a bug. It is exactly what a hub-and-spoke topology does. The question is whether you can teach the spokes to skip the middleman. ## The Design That Is Supposed To Fix It NHRP (Next Hop Resolution Protocol) is the mechanism. It is the same machinery that makes DMVPN Phase 3 work, and FlexVPN reuses it. Two knobs do the work, and they live on opposite ends of the tunnel. ### ip nhrp redirect (on the hub) When the hub receives a packet on an NHRP-enabled interface and forwards it *back out* an NHRP-enabled interface (which is precisely what a hairpin is), it recognises the pattern. Rather than silently continuing to forward, it sends an NHRP **Traffic Indication** back to the source. The message is essentially "I am forwarding this for you, but you could be reaching that destination directly. Go look it up." This is the trigger for the whole process. If the redirect never fires, nothing downstream of it can happen. ### ip nhrp shortcut virtual-template 1 (on the spoke) The spoke receives the Traffic Indication and sends an NHRP **Resolution Request** for the destination prefix. That request reaches the far spoke, which answers with a **Resolution Reply** containing its NBMA address - the public, underlay IP it is reachable on. Now the originating spoke knows how to get to its peer directly, so it clones its own virtual-template into a fresh Virtual-Access interface, brings up a new IKEv2 and IPsec SA straight to that NBMA address, and installs a shortcut route in the NHRP cache. The next packet takes the direct tunnel. The traceroute drops to one hop. The hub is out of the data path entirely, and it stays out until the cache entry ages out. That is the theory, and it is a genuinely good design. Here is the configuration that implements it, taken verbatim from our lab. ### The hub's virtual-template ``` interface Virtual-Template1 type tunnel ip unnumbered Loopback0 ip nhrp network-id 1 ip nhrp redirect tunnel source GigabitEthernet2 tunnel mode ipsec ipv4 tunnel protection ipsec profile PLZ-FLEX-IPSEC ``` Note the `tunnel source` and `tunnel mode` lines. They are not optional. Leave them out and the cloned Virtual-Access comes up and immediately dies in a loop, throwing `%DMVPN-5-NHRP_NETID_UNCONFIGURED` at you while IKEv2 authenticates perfectly, which makes it look like a crypto problem when it is not. We cover that trap in the [FlexVPN hub-and-spoke configuration walkthrough](https://www.pinglabz.com/flexvpn-site-to-site-configuration/). ### The spoke's tunnel ``` interface Tunnel0 ip address negotiated ip nhrp network-id 1 ip nhrp shortcut virtual-template 1 tunnel source GigabitEthernet2 tunnel mode ipsec ipv4 tunnel destination 203.0.113.1 tunnel protection ipsec profile PLZ-FLEX-IPSEC ``` Three things matter here. `ip address negotiated` means the tunnel IP comes from the hub's pool, so you never touch spoke addressing. `ip nhrp network-id 1` must match the hub (it is a local significance value, but both ends need to consider themselves on the same NHRP network). And `ip nhrp shortcut virtual-template 1` is the line that says "if you get a redirect, resolve it and build a tunnel by cloning template 1". ## What Actually Happened The configuration was correct. We can prove that, because on IOS XE the interesting question is not what is in the virtual-template, it is what got *inherited* by the dynamically-created Virtual-Access when the spoke connected. So look at the derived config, not the running config: ``` HUB#show derived-config interface Virtual-Access1 interface Virtual-Access1 ip unnumbered Loopback0 ip nhrp network-id 1 no ip nhrp shortcut ip nhrp redirect <-- redirect IS there tunnel source GigabitEthernet2 tunnel mode ipsec ipv4 tunnel destination 198.51.100.2 tunnel protection ipsec profile PLZ-FLEX-IPSEC ``` `ip nhrp redirect` is present on the live, cloned interface that is actually carrying the spoke's traffic. This is not a case of the template failing to apply. The feature is enabled exactly where it needs to be. And yet NHRP never sent a single packet. ``` HUB#show ip nhrp traffic Virtual-Access1: Max-send limit:10000Pkts/10Sec, Usage:0% Sent: Total 0 0 Resolution Request 0 Resolution Reply 0 Registration Request 0 Error Indication 0 Traffic Indication 0 Redirect Suppress Rcvd: Total 0 0 Resolution Request 0 Resolution Reply 0 Registration Request ``` Zero Traffic Indications sent. The hub was hairpinning spoke-to-spoke traffic all day (we had just pushed twenty ICMP echoes through it and watched them succeed), and it never once generated the redirect that is supposed to start the shortcut process. Predictably, the spoke side is a desert: ``` SPOKE1#show ip nhrp (empty) SPOKE1#show ip nhrp traffic Tunnel0: Max-send limit:10000Pkts/10Sec, Usage:0% Sent: Total 0 0 Resolution Request 0 Resolution Reply 0 Registration Request ``` An empty NHRP cache and zero Resolution Requests sent. That is entirely consistent: the spoke was never told to resolve anything, so it did not. ### What we tried For completeness, here is what we changed and what it did not fix: - **Added `ip nhrp redirect` to the spoke tunnels and virtual-templates** as well as the hub. No change. (A spoke that builds shortcuts can itself need to redirect in some topologies, so it was worth trying.) - **Added `ip nhrp nhs 10.255.0.1 nbma 203.0.113.1 multicast`** to the spoke Tunnel0, giving the spoke an explicit next-hop-server statement pointing at the hub. Counters stayed at zero. - **Sustained data traffic** between the spoke LANs, in case the redirect needed more than a handful of packets to trigger. It did not help. - **Cleared and rebuilt the IKEv2 SAs** so the Virtual-Access interfaces were freshly cloned with every knob already in place. Same result. The conclusion, stated plainly: **on cat8000v running IOS XE 17.18.02, in this topology, the NHRP shortcut did not establish.** The hub never generated a Traffic Indication, so the spokes were never prompted to resolve, so no direct tunnel was ever built. Spoke-to-spoke connectivity worked, but only via the hairpin. We are not going to speculate wildly about the root cause. This is platform and version behaviour, and it is exactly the kind of thing you should verify on your own gear and your own image before you design around it. What we *can* give you is the diagnostic method, which is worth more than a guess. ## How To Tell Whether Yours Is Actually Working This is the reusable part. There are four independent checks, and they fail in a specific order. If the first one is zero, the other three cannot possibly succeed, which is enormously useful because it tells you where to stop looking. 1\. Hub: Traffic Indication counter Command: `show ip nhrp traffic` Working: Traffic Indication is NON-ZERO and climbs as spoke-to-spoke traffic flows. Ours: 0 Traffic Indication. Nothing sent, ever. 2\. Spoke: NHRP cache Command: `show ip nhrp` Working: A cache entry for the far spoke's LAN prefix, mapped to that spoke's NBMA (public) address. Ours: Empty. No entries at all. 3\. Spoke: IKEv2 SA count Command: `show crypto ikev2 sa` Working: A SECOND SA, direct to the other spoke's public IP, alongside the hub SA. Ours: One SA only, to the hub. 4\. Spoke: the path itself Command: `traceroute` spoke LAN to spoke LAN Working: Drops from 2 hops to 1 once the shortcut is up (the first probe may still show the hub). Ours: Stayed at 2 hops. Hub, then spoke. The single most important line in that grid is the first one. **If your hub's Traffic Indication counter is zero, the redirect is not being generated, and no amount of spoke-side configuration will fix it.** Engineers lose hours here, because the symptom presents on the spoke (empty cache, no second SA, two-hop traceroute) and the instinct is to start adjusting the spoke. Do not. The spoke is behaving correctly. It is waiting to be told to resolve, and it is never being told. Start at the hub, confirm the redirect is actually inherited by the Virtual-Access with `show derived-config interface Virtual-Access1`, and then watch that counter while you push traffic. If it moves, the mechanism is alive and you have a spoke-side or reachability problem. If it does not move, the mechanism is not alive and you are debugging the wrong device. ## The Honest Recommendation Here is where we land after building this for real. **FlexVPN's hub-and-spoke is excellent, and you should use it.** The IKEv2 route injection in particular is one of the more elegant things in the Cisco VPN toolkit: the hub hands out a tunnel address from a local pool and pushes the entire prefix list into the spoke's routing table through an IKEv2 authorization policy. No overlay routing protocol. No adjacencies to babysit. If you want the full build, the config, and the gotchas, that is [FlexVPN site-to-site configuration](https://www.pinglabz.com/flexvpn-site-to-site-configuration/), and the conceptual grounding is in [FlexVPN explained](https://www.pinglabz.com/flexvpn-explained/). **For proven spoke-to-spoke today, DMVPN Phase 3 does this reliably.** Same NHRP redirect and shortcut mechanism, same design intent, and we have it working and captured on real gear. If your requirement is "branches must talk directly to each other and I need it in production this quarter", DMVPN Phase 3 is the path with the shorter risk profile. The full cluster, including the phase-by-phase behaviour and the NHRP cache entries you should expect to see, lives at [the PingLabz DMVPN guide](https://www.pinglabz.com/dmvpn/). It is worth your time. **And validate the FlexVPN shortcut on your own platform before you bet a design on it.** Not because the feature is bad, but because "it works on the datasheet" and "it works on the image in my rack" are different claims, and only one of them keeps your migration on schedule. Build the two-spoke lab, push traffic, and watch `show ip nhrp traffic` on the hub. You will know inside five minutes. ## Key Takeaways - **Hub-and-spoke hairpinning is real and measurable.** Our spoke-to-spoke traceroute showed two hops: the hub at 10.255.0.1, then the far spoke at 10.0.9.25\. Connectivity was 100 percent, but every packet was decrypted and re-encrypted by a router that had no business being in the path. - **The fix is a two-sided mechanism.** `ip nhrp redirect` on the hub generates the Traffic Indication; `ip nhrp shortcut virtual-template 1` on the spoke turns that into a Resolution Request and a direct dynamic tunnel. Both ends need a matching `ip nhrp network-id`. - **Check the derived config, not the template.** On FlexVPN the interface that carries traffic is a cloned Virtual-Access. `show derived-config interface Virtual-Access1` is how you prove the NHRP knobs actually landed on it. - **On cat8000v 17.18.02, in our topology, the shortcut never formed.** The redirect was correctly inherited, the counters stayed at zero, the NHRP cache stayed empty, and the traceroute stayed at two hops. We tried spoke-side redirect, an explicit NHS statement, sustained traffic, and full SA rebuilds. None of it changed the result. - **Debug from the hub outward.** A zero Traffic Indication counter means the redirect is not firing, and nothing you configure on the spoke can compensate for that. Four checks, in order: hub counter, spoke cache, spoke SA count, traceroute hop count. - **DMVPN Phase 3 is the proven route to spoke-to-spoke today.** Use FlexVPN for its hub-and-spoke and its IKEv2 route injection, which are genuinely first-rate, and validate the shortcut path on your own gear before you commit to it. None of this diminishes FlexVPN. It is the most modern [IPsec VPN](https://www.pinglabz.com/ipsec-vpn/) framework Cisco ships, and the parts of it that worked in our lab worked cleanly and with less configuration than any equivalent. But a lab is where you find out which parts those are, and the difference between an engineer and a brochure is that the engineer tells you what the box actually did. ### FlexVPN Site-to-Site Configuration on IOS XE URL: https://www.pinglabz.com/flexvpn-site-to-site-configuration/ Last updated: 2026-07-13T08:08:31.000Z FlexVPN is what happens when Cisco stops bolting features onto crypto maps and rebuilds the whole thing on top of IKEv2\. There is no crypto ACL, no static tunnel per peer, and no routing protocol running over the overlay just to tell the hub where a spoke's LAN lives. You define one authorization policy, one Virtual-Template, and the hub clones a fresh Virtual-Access interface for every spoke that authenticates, hands it an IP address, and installs its routes. If you want the conceptual background first, start with the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/) and [FlexVPN explained](https://www.pinglabz.com/flexvpn-explained/). This article is the other half: the build. Everything below came out of a live CML lab. Three cat8000v routers on IOS XE 17.18.02 (HUB, SPOKE1, SPOKE2) with an IOL-XE router acting as the ISP between them. Every command and every piece of output on this page is real, including the failure we walked into and the log lines it produced. ## The Lab HUB Internet:`203.0.113.1` LAN:`10.20.10.0/24` Also:`192.168.99.0/24` (real Debian VM) SPOKE1 Internet:`198.51.100.2` LAN:`10.30.10.0/24` Identity:`spoke1.pinglabz.lab` SPOKE2 Internet:`198.51.100.6` LAN:`10.30.20.0/24` Identity:`spoke2.pinglabz.lab` The spokes have a default route to the ISP and nothing else. They do not know the hub's LAN exists. By the end of this build they will, and no routing protocol will have been involved. ## Block 1: AAA and the Authorization Policy This is the block people skim, and it is the single most important block in the entire configuration. FlexVPN's headline feature is that the hub pushes an IP address and a set of routes down to the spoke as part of the IKEv2 exchange. The mechanism it uses to do that is AAA network authorization. So the first thing you turn on is AAA. ``` aaa new-model aaa authorization network PLZ-FLEX-LIST local ``` `aaa authorization network` creates a method list called `PLZ-FLEX-LIST` that resolves against the **local** database (in production this would point at RADIUS instead). When the IKEv2 profile later says "go ask AAA what this peer is allowed to have", `PLZ-FLEX-LIST` is what it asks. What AAA hands back is the authorization policy: ``` crypto ikev2 authorization policy PLZ-AUTHOR-POL pool PLZ-FLEX-POOL route set interface route set access-list PLZ-FLEX-ROUTES ``` Three lines, three jobs: - `pool PLZ-FLEX-POOL` tells the hub to allocate the connecting spoke an address out of a local pool. That address becomes the spoke's tunnel IP. - `route set interface` advertises the interface the Virtual-Access is unnumbered to (on our hub, Loopback0, `10.255.0.1`). That is how the spoke learns the hub's tunnel endpoint address. - `route set access-list PLZ-FLEX-ROUTES` advertises whatever the ACL permits. This is the hub telling the spoke "here are the networks I can reach for you." And the ACL itself: ``` ip access-list standard PLZ-FLEX-ROUTES permit 10.20.10.0 0.0.0.255 permit 192.168.99.0 0.0.0.255 permit 10.30.0.0 0.0.255.255 ``` The first two lines are the hub's own networks (the LAN and the segment where our real Debian VM lives). The third, `10.30.0.0/16`, is the summary that covers *both* spoke LANs. That is what makes spoke-to-spoke traffic possible at all: SPOKE1 installs a route for 10.30.0.0/16 pointing at the tunnel, so when it wants to reach SPOKE2's 10.30.20.0/24 it sends the packet to the hub. This block is the heart of FlexVPN. Get it right and the rest is plumbing. Skip it and you have built an ordinary SVTI with extra steps. ## Block 2: The Address Pool ``` ip local pool PLZ-FLEX-POOL 10.0.9.10 10.0.9.100 ``` Ninety one addresses, one per spoke. Every spoke that authenticates gets one, and it gets it back into the pool when the session tears down. This is a plain `ip local pool`, the exact same construct used by remote-access VPN and PPP, which is a nice reminder that FlexVPN treats a site-to-site spoke as just another client. ## Block 3: The IKEv2 Keyring (and an honest note about wildcard PSKs) ``` crypto ikev2 keyring PLZ-FLEX-KEYRING peer SPOKES address 0.0.0.0 0.0.0.0 pre-shared-key PingLabz-IKEv2-PSK-01 ``` One keyring entry, matching `0.0.0.0 0.0.0.0`, which is every address on the internet. This is deliberate. The entire point of a FlexVPN hub is that you do not have to touch it when you add spoke number 47\. If the keyring had to name each peer by IP, you would be back to editing the hub for every new site, and the dynamic Virtual-Template would be pointless. Be honest with yourself about what that costs. A wildcard pre-shared key means every spoke shares one secret, and any device that learns that secret and can present an acceptable identity can bring up a tunnel to your hub. It is fine for a lab and it is fine for a proof of concept. In production, the answer is certificates: each spoke gets its own identity certificate from your CA, the hub validates it, and revoking one spoke does not mean rekeying the fleet. We use a wildcard PSK here because it keeps the config readable and the failure modes visible, not because it is what you should ship. ## Block 4: The IKEv2 Profile ``` crypto ikev2 profile PLZ-FLEX-PROF match identity remote fqdn domain pinglabz.lab identity local fqdn hub.pinglabz.lab authentication local pre-share authentication remote pre-share keyring local PLZ-FLEX-KEYRING aaa authorization group psk list PLZ-FLEX-LIST PLZ-AUTHOR-POL virtual-template 1 ``` The profile is where the hub decides who it will talk to and what happens when it does. Four things are worth stopping on. **Identity matching by FQDN domain.** `match identity remote fqdn domain pinglabz.lab` says: accept any peer whose IKE identity ends in `pinglabz.lab`. Not an IP address, a domain. That is the second half of what makes the hub spoke-agnostic. SPOKE1 announces itself as `spoke1.pinglabz.lab`, SPOKE2 as `spoke2.pinglabz.lab`, and a spoke you build next month as `spoke47.pinglabz.lab`, and the hub matches all of them against one line. (If you are hazy on how IKEv2 identities and the SA exchange actually work, [IKEv2 explained](https://www.pinglabz.com/ikev2-explained/) covers it.) **The AAA line.** `aaa authorization group psk list PLZ-FLEX-LIST PLZ-AUTHOR-POL` is the wire that connects Block 1 to the tunnel. It says: for PSK-authenticated peers, run group authorization against method list `PLZ-FLEX-LIST`, using `PLZ-AUTHOR-POL` as the username to look up. Without this line, IKEv2 will still authenticate, the tunnel will still come up, and absolutely nothing will be pushed to the spoke. **`virtual-template 1`.** This tells IKEv2 which template to clone when a peer matches this profile. This is the dynamic part. ## Block 5: The IPsec Profile ``` crypto ipsec profile PLZ-FLEX-IPSEC set transform-set PLZ-TS-GCM set ikev2-profile PLZ-FLEX-PROF ``` Short and mechanical. The IPsec profile binds the data-plane transform (we are using `esp-gcm 256`) to the IKEv2 profile that governs the control plane. The Virtual-Template applies this profile, and every Virtual-Access cloned from that template inherits it. ## Block 6: The Virtual-Template (read this part twice) ``` interface Virtual-Template1 type tunnel ip unnumbered Loopback0 ip nhrp network-id 1 ip nhrp redirect tunnel source GigabitEthernet2 tunnel mode ipsec ipv4 tunnel protection ipsec profile PLZ-FLEX-IPSEC ``` The Virtual-Template is a blueprint, not an interface that carries traffic. When SPOKE1 authenticates, IOS XE stamps out `Virtual-Access1` from it. When SPOKE2 authenticates, it stamps out `Virtual-Access2`. Same template, two live interfaces, each with its own tunnel destination filled in from the IKEv2 negotiation. `ip unnumbered Loopback0` means every Virtual-Access borrows the hub's loopback address (`10.255.0.1`), which is why later on you will see both of them showing the same IP. `ip nhrp network-id` and `ip nhrp redirect` are the groundwork for [spoke-to-spoke shortcut switching](https://www.pinglabz.com/flexvpn-spoke-to-spoke/), which is a topic of its own. ### The gotcha that will eat your afternoon The Virtual-Template **must** carry `tunnel source` and `tunnel mode ipsec ipv4`. If it does not, here is what happens: IKEv2 negotiates successfully, the peer authenticates, the Virtual-Access is cloned, it comes up, and then it immediately dies. Then it does it again. Forever. This is exactly what our lab did: ``` %LINK-3-UPDOWN: Interface Virtual-Access1, changed state to down %DMVPN-5-NHRP_NETID_UNCONFIGURED: Virtual-Access1: NETID : 1 Unconfigured (Tunnel: 0.0.0.0 NBMA: UNKNOWN) ``` Look closely at that second line. `Tunnel: 0.0.0.0` and `NBMA: UNKNOWN`. The cloned interface has no idea what its own source address is or what encapsulation it should be using, so NHRP cannot register a network ID against it and the interface is torn back down. What makes this genuinely nasty is that **the crypto is fine**. IKEv2 authentication succeeds. You can watch the authorization request go out in the debug: ``` IKEv2:Using mlist PLZ-FLEX-LIST and username PLZ-AUTHOR-POL for group author request ``` So every instinct you have says crypto bug. You go and stare at your keyring, your proposal, your transform set, your identities. All of them are correct. The problem is two missing lines on an interface that is not even in your running config in the form you are looking at, because the thing that is flapping is a *clone*. Add `tunnel source GigabitEthernet2` and `tunnel mode ipsec ipv4` to the Virtual-Template. The flapping stops instantly. Also note the message is tagged `%DMVPN`, not `%FLEXVPN`, because FlexVPN reuses the NHRP machinery that [DMVPN](https://www.pinglabz.com/dmvpn/) is built on. Same plumbing, different front end. ## Block 7: The Spoke ``` crypto ikev2 authorization policy PLZ-AUTHOR-POL route set interface route set access-list PLZ-FLEX-ROUTES crypto ikev2 profile PLZ-FLEX-PROF match identity remote fqdn domain pinglabz.lab identity local fqdn spoke1.pinglabz.lab authentication local pre-share authentication remote pre-share keyring local PLZ-FLEX-KEYRING aaa authorization group psk list PLZ-FLEX-LIST PLZ-AUTHOR-POL virtual-template 1 interface Tunnel0 ip address negotiated ip nhrp network-id 1 ip nhrp shortcut virtual-template 1 tunnel source GigabitEthernet2 tunnel mode ipsec ipv4 tunnel destination 203.0.113.1 tunnel protection ipsec profile PLZ-FLEX-IPSEC ``` The spoke is a mirror image with three differences that matter. **`ip address negotiated`.** The spoke does not configure a tunnel IP. It asks for one, and the hub's `PLZ-FLEX-POOL` gives it one. This is the client relationship made concrete. **A static `tunnel destination`.** The hub is dynamic and accepts anyone; the spoke knows exactly where it is going, so it points a normal Tunnel interface at `203.0.113.1`. No Virtual-Template needed for the hub-facing tunnel (the spoke's `virtual-template 1` reference exists for the spoke-to-spoke case). **Its own identity and its own route-set ACL.** `identity local fqdn spoke1.pinglabz.lab` must land inside the hub's `fqdn domain pinglabz.lab` match, and SPOKE1's `PLZ-FLEX-ROUTES` ACL permits `10.30.10.0 0.0.0.255`, which is how the hub learns SPOKE1's LAN. SPOKE2 is identical apart from `spoke2.pinglabz.lab` and `10.30.20.0 0.0.0.255`. Route advertisement runs both directions over the same IKEv2 exchange. ## Verification: Does It Actually Work? Two spokes, both READY, both AES-GCM-256 with DH group 19 (note `Hash: None`, which is correct for an AEAD cipher because GCM handles integrity internally): ``` HUB#show crypto ikev2 sa Tunnel-id Local Remote fvrf/ivrf Status 1 203.0.113.1/500 198.51.100.2/500 none/none READY Encr: AES-GCM, keysize: 256, PRF: SHA256, Hash: None, DH Grp:19, Auth sign: PSK, Auth verify: PSK Tunnel-id Local Remote fvrf/ivrf Status 2 203.0.113.1/500 198.51.100.6/500 none/none READY Encr: AES-GCM, keysize: 256, PRF: SHA256, Hash: None, DH Grp:19, Auth sign: PSK, Auth verify: PSK ``` Two Virtual-Access interfaces exist that you never typed into the config, both up, both wearing the hub's Loopback0 address: ``` HUB#show ip interface brief | include Virtual-Access Virtual-Access1 10.255.0.1 YES unset up up Virtual-Access2 10.255.0.1 YES unset up up HUB#show crypto session Interface: Virtual-Access1 Session status: UP-ACTIVE Peer: 198.51.100.2 port 500 Interface: Virtual-Access2 Session status: UP-ACTIVE Peer: 198.51.100.6 port 500 ``` Now the payoff. Look at the hub's routing table: ``` HUB#show ip route | include Virtual-Access S 10.0.9.24/32 is directly connected, Virtual-Access1 S 10.0.9.25/32 is directly connected, Virtual-Access2 S 10.30.10.0/24 is directly connected, Virtual-Access1 S 10.30.20.0/24 is directly connected, Virtual-Access2 ``` SPOKE1's LAN and SPOKE2's LAN, each pointing at the correct Virtual-Access, plus the two /32s for the pool addresses the spokes were handed. There is no OSPF here. No EIGRP, no BGP, no static route you typed. IKEv2 installed all four entries out of the authorization policy. The pool confirms the allocation: ``` HUB#show ip local pool PLZ-FLEX-POOL Pool Begin End Free In use PLZ-FLEX-POOL 10.0.9.10 10.0.9.100 89 2 Inuse addresses: 10.0.9.24 IKEv2 Addr IDB 10.0.9.25 IKEv2 Addr IDB ``` `IKEv2 Addr IDB` is the pool telling you who took the address and why. Two spokes, two leases, 89 free. And the same trick ran in reverse. On the spoke: ``` SPOKE1#show ip route static S* 0.0.0.0/0 [1/0] via 198.51.100.1 S 10.20.10.0/24 is directly connected, Tunnel0 S 10.30.0.0/16 is directly connected, Tunnel0 S 10.255.0.1/32 is directly connected, Tunnel0 S 192.168.99.0/24 is directly connected, Tunnel0 ``` SPOKE1 started with one static default route to its ISP. It now also has the hub's LAN (10.20.10.0/24), the hub's real-VM segment (192.168.99.0/24), the hub's tunnel endpoint (10.255.0.1/32 from `route set interface`), and the 10.30.0.0/16 summary that lets it reach the other spoke. All four arrived over IKEv2. ## End to End: Spoke to Spoke The proof that the route injection is not just cosmetic: SPOKE1 pinging SPOKE2's LAN, sourced from its own LAN interface. ``` SPOKE1#traceroute 10.30.20.1 source Loopback10 probe 1 timeout 1 Type escape sequence to abort. Tracing the route to 10.30.20.1 1 10.255.0.1 34 msec 2 10.0.9.25 81 msec SPOKE1#ping 10.30.20.1 source Loopback10 repeat 20 !!!!!!!!!!!!!!!!!!!! Success rate is 100 percent (20/20), round-trip min/avg/max = 4/7/15 ms ``` Twenty for twenty. Two hops: the hub's tunnel address (10.255.0.1), then SPOKE2's pool address (10.0.9.25). That is spoke-to-spoke traffic hairpinning through the hub, decrypted on arrival and re-encrypted on the way back out. It works, and for a lot of designs that is entirely acceptable. Turning those two hops into one is the job of NHRP shortcut switching, which gets its own article. ## Troubleshooting: What Actually Broke Virtual-Access flapping up and down Symptom:`%DMVPN-5-NHRP_NETID_UNCONFIGURED` with `Tunnel: 0.0.0.0 NBMA: UNKNOWN`, in a loop. IKEv2 auth succeeds. Fix:Add `tunnel source` and `tunnel mode ipsec ipv4` to the Virtual-Template. Peer never matches the profile Symptom:No IKEv2 SA at all. The hub silently ignores the spoke. Fix:The spoke's `identity local fqdn` must end in the domain the hub matches. `spoke1.pinglabz.lab` matches `fqdn domain pinglabz.lab`. A typo in either half is a silent no-match. Tunnel is up but no routes arrive Symptom:IKEv2 SA READY, Virtual-Access up, routing table unchanged. Fix:Authorization is not being attempted. Confirm with the debug line `IKEv2:Using mlist PLZ-FLEX-LIST and username PLZ-AUTHOR-POL for group author request`. If it is missing, check `aaa new-model`, the `aaa authorization network` list, and the `aaa authorization group psk list` line in the IKEv2 profile. One more habit worth building: when a FlexVPN Virtual-Access misbehaves, do not read the Virtual-Template to see what the interface looks like. Read the clone itself with `show derived-config interface Virtual-Access1`. That shows you the merged result of the template plus everything IKEv2 filled in, which is the config the router is actually running. ## Key Takeaways - **The IKEv2 authorization policy is FlexVPN.** `aaa new-model`, `aaa authorization network`, and `crypto ikev2 authorization policy` with `pool`, `route set interface`, and `route set access-list` are what push addresses and routes to the spokes. Everything else is supporting cast. - **Routes arrive over IKEv2, not over a routing protocol.** Our hub learned 10.30.10.0/24 and 10.30.20.0/24, and the spokes learned 10.20.10.0/24 and 192.168.99.0/24, with zero OSPF, EIGRP, or BGP anywhere in the config. - **A Virtual-Template without `tunnel source` and `tunnel mode ipsec ipv4` will flap forever** and lie to you about it, because IKEv2 authentication succeeds while the cloned Virtual-Access dies in a loop. The `NBMA: UNKNOWN` in the NHRP log is the tell. - **A wildcard PSK (`address 0.0.0.0 0.0.0.0`) is what makes the hub spoke-agnostic, and it is a real security tradeoff.** Certificates are the production answer. - **The hub scales without edits.** One Virtual-Template, one FQDN domain match, one pool. Spoke number 47 needs zero changes on the hub. - **Spoke-to-spoke works, at two hops.** 20/20 pings through the hub. Collapsing that to one hop is NHRP's job, not IKEv2's. Next in this cluster: [FlexVPN spoke-to-spoke and NHRP shortcut switching](https://www.pinglabz.com/flexvpn-spoke-to-spoke/), where the two-hop path is supposed to become one. If you want the conceptual model behind everything above, read [FlexVPN explained](https://www.pinglabz.com/flexvpn-explained/) and [IKEv2 explained](https://www.pinglabz.com/ikev2-explained/), and for the wider picture of where FlexVPN sits next to crypto maps, SVTIs, and [DMVPN](https://www.pinglabz.com/dmvpn/), the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/) is the map. ### IKEv2 Explained: Why It Replaced IKEv1 (and the Smart Defaults) URL: https://www.pinglabz.com/ikev2-explained/ Last updated: 2026-07-13T08:08:29.000Z IKEv2 is not IKEv1 with a bigger version number. It is a different protocol, defined in a different RFC, with a different exchange, different state machine, different show commands, and a completely different philosophy about defaults. If you learned IPsec on IKEv1 and you carry those habits across, the first thing that happens is you type `show crypto isakmp sa` on a working IKEv2 tunnel and get nothing back, and you spend ten minutes convinced the tunnel is down when it has been up the whole time. This article is part of the PingLabz [IPsec VPN cluster](https://www.pinglabz.com/ipsec-vpn/). It walks through what actually changed between the two versions, using real output captured from a live CML lab (cat8000v running IOS XE 17.18.02) rather than a spec summary. If you have not read the IKEv1 model yet, start with [IPsec explained: the IKE phases](https://www.pinglabz.com/ipsec-explained-ike-phases/), because most of what follows is a comparison against it. ## The Exchange: 4 Messages Instead of 9 IKEv1's negotiation is two phases and a lot of packets. Main Mode is six messages to build the IKE SA (three request/response pairs: policy proposal, Diffie-Hellman exchange, then authentication). Then Quick Mode is another three messages to build the IPsec SA that actually carries data. That is nine packets before a single user byte moves, and if you are unlucky enough to be using Aggressive Mode to cut that down, you have traded round trips for a pre-hashed identity sent in the clear. IKEv2 collapses all of it into two exchanges, four messages total: IKE\_SA\_INIT **Messages:** 2 **Does:** negotiate crypto, do the Diffie-Hellman, exchange nonces **Replaces:** IKEv1 Main Mode messages 1 to 4 IKE\_AUTH **Messages:** 2 **Does:** authenticate the peers, build the first Child SA in the same breath **Replaces:** Main Mode 5 and 6, plus all of Quick Mode CREATE\_CHILD\_SA **Messages:** 2, only when needed **Does:** add or rekey a Child SA after the fact **Replaces:** additional Quick Modes The headline is fewer round trips, which matters most on high-latency paths and on hubs terminating thousands of spokes. But the structural wins are the ones you feel in production. **DoS protection is built in.** IKEv1 will happily perform an expensive Diffie-Hellman computation for anyone who sends it a packet, which makes it trivially cheap to attack. IKEv2 has a cookie challenge: when the responder is under load and sees half-open SAs stacking up, it replies to IKE\_SA\_INIT with a stateless cookie and refuses to allocate any state until the initiator echoes it back. A spoofed source address cannot echo anything, so the flood dies at the door. **Delivery is reliable and windowed.** Every IKEv2 message carries a message ID, and every request must be acknowledged. That is not decoration, it is visible in the SA. Look at the real output from our lab hub: ``` HUB#show crypto ikev2 sa detailed IPv4 Crypto IKEv2 SA Tunnel-id Local Remote fvrf/ivrf Status 1 203.0.113.1/500 198.51.100.2/500 none/none READY Encr: AES-GCM, keysize: 256, PRF: SHA256, Hash: None, DH Grp:19, Auth sign: PSK, Auth verify: PSK Life/Active Time: 86400/6 sec CE id: 1001, Session-id: 1 Local spi: 006A65307FBECC85 Remote spi: E0055D1CB14CC870 Status Description: Negotiation done Local id: 203.0.113.1 Remote id: 198.51.100.2 Local req msg id: 2 Remote req msg id: 0 Local next msg id: 2 Remote next msg id: 0 Local window: 5 Remote window: 5 DPD configured for 0 seconds, retry 0 Fragmentation not configured. Dynamic Route Update: enabled Extended Authentication not configured. NAT-T is not detected Initiator of SA : Yes PEER TYPE: IOS-XE ``` `Local req msg id`, `Local next msg id` and `Local window: 5` are the reliability machinery showing itself. The window means this router can have up to five requests outstanding at once instead of stalling on a strict one-at-a-time lockstep. IKEv1 has no equivalent because IKEv1 has no reliable delivery layer, it just retransmits and hopes. ## State Names: QM\_IDLE Is Gone, READY Is the New Normal This is the single most common trip-up when engineers move over, so it gets its own section. IKEv1's healthy steady state is `QM_IDLE`, meaning "Quick Mode finished and the IKE SA is sitting idle". Here it is from our IKEv1 build: ``` EDGE1#show crypto isakmp sa IPv4 Crypto ISAKMP SA dst src state conn-id status 198.51.100.2 203.0.113.1 QM_IDLE 1001 ACTIVE ``` IKEv2 has no Quick Mode, so there is nothing to be idle about. Its healthy steady state is `READY`: ``` HUB#show crypto ikev2 sa Tunnel-id Local Remote fvrf/ivrf Status 1 203.0.113.1/500 198.51.100.2/500 none/none READY Encr: AES-GCM, keysize: 256, PRF: SHA256, Hash: None, DH Grp:19, Auth sign: PSK, Auth verify: PSK Tunnel-id Local Remote fvrf/ivrf Status 2 203.0.113.1/500 198.51.100.6/500 none/none READY Encr: AES-GCM, keysize: 256, PRF: SHA256, Hash: None, DH Grp:19, Auth sign: PSK, Auth verify: PSK ``` Two separate commands, two separate state machines. `show crypto isakmp sa` only ever shows IKEv1 SAs, and on an IKEv2-only router it returns an empty table. That empty table is not a fault, it is the wrong command. Burn `show crypto ikev2 sa` into muscle memory, and reach for `show crypto session` when you want a version-agnostic answer to "is this tunnel up". ## Lifetime: 86400 Instead of 3600 The lifetime numbers change too. Our IKEv1 policy pins the IKE SA at an hour: ``` crypto isakmp policy 10 encryption aes 256 hash sha256 authentication pre-share group 14 lifetime 3600 ``` Our IKEv2 tunnel, with no lifetime configured at all, reports: ``` Life/Active Time: 86400/6 sec ``` That is 24 hours of SA lifetime and 6 seconds of active time, straight from the detailed SA output above. IKEv2 rekeys are also cheaper than IKEv1's, because CREATE\_CHILD\_SA can refresh a Child SA (with or without a fresh Diffie-Hellman for perfect forward secrecy) without disturbing the parent IKE SA. In IKEv1, an ISAKMP SA rekey is a full renegotiation. ## AEAD and the Mystery of `Hash: None` Look again at the SA line: `Encr: AES-GCM, keysize: 256, PRF: SHA256, Hash: None`. Engineers see `Hash: None` and reach for the phone. No integrity? On a VPN? It is fine. AES-GCM is an AEAD cipher (Authenticated Encryption with Associated Data). It performs encryption and integrity verification in a single pass, producing an authentication tag as part of the ciphertext. There is no separate hash because there is nothing left for a separate hash to do, and bolting an HMAC on top would be wasted CPU. The same thing shows in the Child SA, where the ESP transform reports no HMAC at all: ``` Child sa: local selector -> remote_selector 10.20.10.0/0 - 10.20.10.255/65535 -> 10.30.10.0/0 - 10.30.10.255/65535 ESP spi in/out: 0xDC114E8E/0x854355BF Encr: AES-GCM, keysize: 256, esp_hmac: None ah_hmac: None, comp: IPCOMP_NONE, mode tunnel ``` You do not have to take my word for it, because IOS enforces the rule and will tell you so. Configure an incomplete IKEv2 proposal and the router prints this validation hint: ``` IKEv2 proposal MUST either have a set of an encryption algorithm other than aes-gcm, an integrity algorithm and a DH group configured or encryption algorithm aes-gcm, a prf algorithm and a DH group configured ``` Read that carefully, because it is the AEAD rule stated by the router itself. Two valid shapes for a proposal, and only two: Non-AEAD cipher (AES-CBC) **Encryption:** required **Integrity:** required **PRF:** optional (defaults to the integrity algorithm) **DH group:** required AEAD cipher (AES-GCM) **Encryption:** required **Integrity:** not allowed, GCM does it internally **PRF:** required (nothing else can derive the keys) **DH group:** required The PRF becomes mandatory with GCM precisely because the integrity algorithm normally doubles as the key-derivation function. Remove integrity and you have to name a PRF explicitly, which is why our lab proposal reads `encryption aes-gcm-256` / `prf sha256` / `group 19` and nothing else. ## The Smart Defaults: IKEv2 Ships Ready to Go This is the change that most people never notice, and it is the biggest quality-of-life improvement in the protocol. IKEv2 on IOS XE comes with a working default proposal and a working default policy already in the box. You can build a functioning tunnel without ever typing the word "proposal". Here is what is actually sitting there on a fresh router: ``` HUB#show crypto ikev2 proposal default IKEv2 proposal: default Encryption : AES-CBC-256 Integrity : SHA512 SHA384 PRF : SHA512 SHA384 DH Group : DH_GROUP_256_ECP/Group 19 DH_GROUP_2048_MODP/Group 14 DH_GROUP_521_ECP/Group 21 HUB#show crypto ikev2 policy default IKEv2 policy : default Match fvrf : any Match address local : any Proposal : default ``` Stop and appreciate that. AES-CBC-256\. SHA-512 and SHA-384 for integrity and PRF. Diffie-Hellman groups 19, 14 and 21 (a 256-bit elliptic curve, 2048-bit MODP, and a 521-bit elliptic curve). That is a modern, defensible crypto set that would pass most audits unchanged. The default policy matches any front-door VRF and any local address, which means it applies to every interface without you scoping anything. Now compare with IKEv1, where there was no default policy worth the name. IKEv1's fallback values were DES for encryption, MD5 for hashing, and Diffie-Hellman group 1 (768-bit). Every one of those is dead. If you forgot to write a policy in IKEv1, the router did not help you, it quietly negotiated garbage. IKEv2 inverted that: the out-of-the-box behaviour is safe, and you have to work at it to be insecure. **The caveat, and it matters.** "Good defaults" is not the same as "pin nothing in production". The default proposal is defined by the software image, and a software image changes when you upgrade. If you never write an explicit proposal, then a code upgrade can silently change your negotiated crypto, and on the day a peer at the other end of a partner tunnel does the same thing you get a very confusing outage. The right posture is: rely on the defaults in the lab and for quick proofs of concept, pin an explicit proposal and policy in production so your crypto is a decision you made rather than a decision your vendor made for you. That is exactly what we do in the [IKEv2 site-to-site with crypto maps](https://www.pinglabz.com/ikev2-site-to-site-crypto-maps/) build. ## Traffic Selectors: Proxy IDs, Renamed and Rethought IKEv1 calls the negotiated "what traffic goes in the tunnel" definition a proxy ID, and it is derived from your crypto ACL. A mismatch between the two ends produces one of the great IPsec time-wasters: Phase 1 comes up perfectly, Phase 2 refuses, and the logs mumble about an invalid proposal. IKEv2 calls the same thing a **traffic selector**, and prints it plainly in the Child SA: ``` Child sa: local selector -> remote_selector 10.20.10.0/0 - 10.20.10.255/65535 -> 10.30.10.0/0 - 10.30.10.255/65535 ``` Read that as a range, not a mask: local addresses 10.20.10.0 through 10.20.10.255 (with the port range 0 to 65535 hanging off the end) to remote addresses 10.30.10.0 through 10.30.10.255\. Traffic selectors are ranges of addresses, ports and protocols, and IKEv2 lets the responder narrow a selector it cannot fully accept rather than rejecting the whole proposal outright, which removes a whole class of IKEv1 configuration standoffs. The cleanest way to sidestep selectors entirely is a route-based tunnel, where the selector becomes `0.0.0.0/0` to `0.0.0.0/0` and the routing table decides what gets encrypted. See [IKEv2 SVTI configuration](https://www.pinglabz.com/ikev2-svti-configuration/) for that build. ## The Other Wins Liveness / DPD Built into the protocol as an empty INFORMATIONAL exchange, not a vendor extension. Visible in the SA as `DPD configured for 0 seconds`, meaning we did not turn the timer on in this lab. NAT traversal Native. NAT detection payloads are part of IKE\_SA\_INIT, and the SA reports `NAT-T is not detected` without any NAT-T configuration existing. In IKEv1 it was a bolt-on you could disable and break. Config payload The tunnel can hand the peer an IP address, DNS servers and *routes*. This is the engine under [FlexVPN](https://www.pinglabz.com/flexvpn-explained/), where the hub injects spoke routes with no routing protocol at all. EAP support IKEv2 can authenticate a remote-access client against RADIUS via EAP. IKEv1 needed the proprietary XAUTH hack to get anywhere close. Asymmetric auth Each side chooses its own method independently. Note the config: `authentication local pre-share` and `authentication remote pre-share` are two separate lines. The hub can use a certificate while the spoke uses a PSK. Fewer moving parts on failure One state machine instead of two phases, so "Phase 1 is up but Phase 2 is dead" stops being the default diagnosis and starts being an actual, narrower question. ## Config Anatomy: Four Blocks Instead of Two Lines IKEv1's crypto config is famously terse: an `isakmp policy` and an `isakmp key`, both global. IKEv2 splits the same job into four named objects that you stitch together. It is more typing, and it is worth it, because each object is reusable and scoped instead of global. Here is the mapping, using our lab config: 1\. Proposal **IKEv2:** `crypto ikev2 proposal` **Holds:** encryption, integrity or PRF, DH group **IKEv1 equivalent:** the algorithm lines inside `crypto isakmp policy` **Optional?** Yes, a default exists 2\. Policy **IKEv2:** `crypto ikev2 policy` **Holds:** which proposal to offer, and which fvrf / local address it applies to **IKEv1 equivalent:** the priority number on `crypto isakmp policy` **Optional?** Yes, a default exists 3\. Keyring **IKEv2:** `crypto ikev2 keyring` **Holds:** per-peer pre-shared keys, keyed by address or wildcard **IKEv1 equivalent:** `crypto isakmp key ... address ...` **Optional?** No, for PSK auth 4\. Profile **IKEv2:** `crypto ikev2 profile` **Holds:** identity matching, local and remote auth methods, which keyring to use, AAA and virtual-template hooks **IKEv1 equivalent:** nothing clean, this is the piece IKEv1 never had **Optional?** No, this is the glue In practice the whole thing looks like this, straight from our hub: ``` crypto ikev2 proposal PLZ-IKEV2-PROP encryption aes-gcm-256 prf sha256 group 19 crypto ikev2 policy PLZ-IKEV2-POL proposal PLZ-IKEV2-PROP crypto ikev2 keyring PLZ-KEYRING peer SPOKE1 address 198.51.100.2 pre-shared-key PingLabz-IKEv2-PSK-01 crypto ikev2 profile PLZ-IKEV2-PROF match identity remote address 198.51.100.2 255.255.255.255 identity local address 203.0.113.1 authentication local pre-share authentication remote pre-share keyring local PLZ-KEYRING ``` The profile is the object that makes IKEv2 scale. It is where identity matching lives (match on an address, an FQDN, a certificate DN), it is where the keyring is bound, and it is what you attach to a crypto map, an IPsec profile, or a virtual-template. IKEv1 had no such abstraction, which is why IKEv1 remote-access designs collapse into a pile of global commands the moment you have more than one class of peer. ## Should You Still Run IKEv1? Only if a peer forces you to. IKEv1 is deprecated, it cannot do AEAD ciphers cleanly, its defaults are dangerous, its NAT handling is an afterthought, and it has no way to push configuration to a peer. Every current Cisco VPN design (FlexVPN, SD-WAN, most modern DMVPN builds) assumes IKEv2\. Treat IKEv1 as an interop compatibility mode for old third-party gear, not as a design choice. ## Key Takeaways - **IKEv2 is a new protocol, not a revision.** Different exchange, different state machine, different show commands, different config objects. - **Four messages, not nine.** IKE\_SA\_INIT plus IKE\_AUTH replaces Main Mode plus Quick Mode, with cookie-based DoS protection and reliable, windowed delivery (`Local window: 5` in the real SA output). - **The state name is READY, not QM\_IDLE.** And the command is `show crypto ikev2 sa`, not `show crypto isakmp sa`. An empty ISAKMP table on an IKEv2 router is normal, not broken. - **Default IKE SA lifetime is 86400 seconds.** Our IKEv1 policy pinned 3600; the IKEv2 SA reports `Life/Active Time: 86400/6 sec` with nothing configured. - **`Hash: None` with AES-GCM is correct.** GCM is AEAD, so encryption and integrity happen in one pass. IOS enforces it: with aes-gcm you must configure a PRF and you must not configure an integrity algorithm. - **The defaults are genuinely good.** AES-CBC-256, SHA-512/384, DH 19/14/21, matching any fvrf and any local address. You can build a tunnel with zero proposal config. Still pin an explicit proposal in production so a software upgrade cannot change your crypto behind your back. - **Traffic selectors replace proxy IDs**, they are ranges rather than ACL-derived masks, and the responder can narrow them instead of rejecting outright. - **The config payload is the sleeper feature.** Pushing addresses and routes over IKEv2 is what makes FlexVPN work without a routing protocol. Next in the cluster: build it for real with [IKEv2 site-to-site using crypto maps](https://www.pinglabz.com/ikev2-site-to-site-crypto-maps/), move to route-based tunnels with [IKEv2 SVTI configuration](https://www.pinglabz.com/ikev2-svti-configuration/), then see what the config payload unlocks in [FlexVPN explained](https://www.pinglabz.com/flexvpn-explained/). For the full picture on tunnel modes, ESP, AH and the rest of the stack, head back to the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/). ### IKEv2 Site-to-Site VPN with Crypto Maps on IOS XE URL: https://www.pinglabz.com/ikev2-site-to-site-crypto-maps/ Last updated: 2026-07-13T08:08:29.000Z You have a site-to-site VPN in production. It works. It has worked for years. It uses a crypto map, an ISAKMP policy, a pre-shared key, and a transform set, and the last time anybody touched it was during a maintenance window nobody wants to repeat. Now someone has read an audit report and wants IKEv2, or wants AES-GCM, or wants the tunnel to stop dropping every time the far side reboots. And the internet keeps telling you the answer is to rip the whole thing out and rebuild it as a virtual tunnel interface. You do not have to. On Cisco IOS XE you can keep the crypto map, keep the crypto ACL, keep the transform set, keep the interface binding, and swap only the key exchange. One line inside the crypto map entry is the entire migration. This article walks the real config and the real output from our lab: a HUB at 203.0.113.1 and SPOKE1 at 198.51.100.2, both cat8000v running 17.18.02, protecting 10.20.10.0/24 to 10.30.10.0/24\. If you want the wider context first, the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/) covers how the pieces fit together, and [the IKEv1 version of this exact build](https://www.pinglabz.com/site-to-site-ipsec-crypto-maps/) is the thing we are migrating from. ## The IKEv2 Config Anatomy IKEv2 is not IKEv1 with a bigger version number. It is a different protocol with a different configuration model, and the object names do not line up cleanly with what you already know. Four constructs do the work. Two of them have obvious IKEv1 ancestors. Two of them are new, and one of those new ones is the piece that everything hangs off. crypto ikev2 proposal **Was:** `crypto isakmp policy` **Holds:** encryption, integrity or PRF, DH group **Difference:** a proposal can carry several algorithms at once, and it has no priority number. Ordering lives in the policy, not here. crypto ikev2 policy **Was:** nothing. This is new. **Holds:** one or more proposals, plus a match on fvrf and local address **Why:** it lets one box offer different crypto depending on which VRF or which local IP the negotiation lands on. crypto ikev2 keyring **Was:** `crypto isakmp key ... address ...` **Holds:** named peer blocks, each with an address and a pre-shared key **Difference:** keys are now structured objects with names, not a flat global list. Asymmetric keys (different local and remote PSK) are possible. crypto ikev2 profile **Was:** nothing. This is new, and it is the important one. **Holds:** identity matching, local identity, local and remote auth methods, the keyring reference **Why:** this is the object you attach to a crypto map, an IPsec profile, or a virtual template. Nothing works until a profile matches. The mental shift is that in IKEv1 the pieces were global and loosely coupled (a policy floated in space, a key floated in space, and IOS did its best to pair them at negotiation time). In IKEv2 the profile is a container that explicitly binds a peer identity to an authentication method to a keyring. Nothing is implicit. That is more typing, and it is also why IKEv2 fails loudly instead of silently doing the wrong thing. ## The Full HUB Config Here is the working configuration from the lab, verbatim. SPOKE1 is the mirror image of this: swap the peer address to 203.0.113.1, swap the local identity to 198.51.100.2, and reverse the source and destination in the crypto ACL. ``` crypto ikev2 proposal PLZ-IKEV2-PROP encryption aes-gcm-256 prf sha256 group 19 crypto ikev2 policy PLZ-IKEV2-POL proposal PLZ-IKEV2-PROP crypto ikev2 keyring PLZ-KEYRING peer SPOKE1 address 198.51.100.2 pre-shared-key PingLabz-IKEv2-PSK-01 crypto ikev2 profile PLZ-IKEV2-PROF match identity remote address 198.51.100.2 255.255.255.255 identity local address 203.0.113.1 authentication local pre-share authentication remote pre-share keyring local PLZ-KEYRING crypto ipsec transform-set PLZ-TS-GCM esp-gcm 256 mode tunnel crypto map PLZ-CMAP 10 ipsec-isakmp set peer 198.51.100.2 set transform-set PLZ-TS-GCM set ikev2-profile PLZ-IKEV2-PROF match address PLZ-CRYPTO-ACL interface GigabitEthernet2 crypto map PLZ-CMAP ``` Look at the bottom half of that config. The transform set, the crypto map, the `match address` pointing at a crypto ACL, the `crypto map` statement on the outside interface: all of it is exactly what you already run for IKEv1\. The top half is new. The joint between them is a single line. ## The One Line That Is the Whole Migration ``` crypto map PLZ-CMAP 10 ipsec-isakmp set ikev2-profile PLZ-IKEV2-PROF ``` That is it. `set ikev2-profile` inside the crypto map entry tells IOS XE to negotiate this SA pair with IKEv2 instead of IKEv1\. The keyword `ipsec-isakmp` on the crypto map line is a historical artifact and stays exactly as it was (do not go looking for an `ipsec-ikev2` variant, there is not one). Everything else in the map is untouched. Which means the migration on a live tunnel looks like this. Build the four IKEv2 objects (they are inert until something references them, so you can stage them safely during business hours). Then in your maintenance window, add one line to the crypto map entry, coordinate with the far end, and clear the SA. The crypto ACL does not change. The transform set does not have to change (though if you are doing this at all, you may as well move to GCM while you are here). The interface binding does not change. Routing does not change, because with a crypto map there was never any routing to change. ## AES-GCM Needs a PRF and No Integrity The first time you type the proposal, IOS will print this at you: ``` IKEv2 proposal MUST either have a set of an encryption algorithm other than aes-gcm, an integrity algorithm and a DH group configured or encryption algorithm aes-gcm, a prf algorithm and a DH group configured ``` This is not an error and the config is not rejected. It is a validation hint that IOS prints while the proposal is incomplete, and it disappears once the proposal has everything it needs. But it is telling you something real, so read it properly. An IKEv2 proposal has to be one of two shapes. Either you use a classic cipher (AES-CBC, say) and you must configure an integrity algorithm plus a DH group. Or you use AES-GCM and you must configure a PRF plus a DH group, and you must NOT configure an integrity algorithm. The reason is that GCM is an AEAD cipher (Authenticated Encryption with Associated Data). It performs encryption and integrity in a single pass, using the GHASH function internally. Bolting an HMAC on top would be dead weight: extra bytes on every packet, extra CPU, zero extra security. The PRF is still needed, because the pseudo-random function is what derives the keying material from the Diffie-Hellman shared secret. In IKEv1 (and in non-GCM IKEv2) the integrity algorithm doubled as the PRF, which is why you never had to think about it. With GCM there is no integrity algorithm to borrow, so you name the PRF explicitly. In our proposal that is `prf sha256`. ## The Tunnel Is Up. Proving It. This is where most IKEv1 muscle memory betrays you, so start with the right command. ``` HUB#show crypto ikev2 sa detailed IPv4 Crypto IKEv2 SA Tunnel-id Local Remote fvrf/ivrf Status 1 203.0.113.1/500 198.51.100.2/500 none/none READY Encr: AES-GCM, keysize: 256, PRF: SHA256, Hash: None, DH Grp:19, Auth sign: PSK, Auth verify: PSK Life/Active Time: 86400/6 sec CE id: 1001, Session-id: 1 Local spi: 006A65307FBECC85 Remote spi: E0055D1CB14CC870 Status Description: Negotiation done Local id: 203.0.113.1 Remote id: 198.51.100.2 Local req msg id: 2 Remote req msg id: 0 Local next msg id: 2 Remote next msg id: 0 Local window: 5 Remote window: 5 DPD configured for 0 seconds, retry 0 Fragmentation not configured. Dynamic Route Update: enabled Extended Authentication not configured. NAT-T is not detected Initiator of SA : Yes PEER TYPE: IOS-XE ``` Read that output line by line, because almost every field is a thing you configured and can now verify. `Status: READY` The IKE SA is established and usable. This is IKEv2's equivalent of IKEv1's QM\_IDLE. `Encr: AES-GCM, keysize: 256` Your `encryption aes-gcm-256` line, confirmed on the wire. `PRF: SHA256` Your `prf sha256` line. Used for key derivation only. `Hash: None` Exactly what you want with GCM. There is no separate integrity algorithm because the cipher does integrity itself. `None` here is a feature, not a warning. `DH Grp: 19` 256-bit elliptic-curve group. Modern, fast, and not on anyone's deprecation list. `Life/Active Time: 86400/6` IKEv2's default SA lifetime is 86400 seconds (24 hours), against IKEv1's 3600\. Fewer rekeys, fewer rekey-related outages. `Local id / Remote id` 203.0.113.1 and 198.51.100.2\. These are the identities the profile matched on. If this section is wrong, nothing else gets a chance to be right. The IKE SA only proves the two routers agreed on how to talk. The traffic itself rides in a child SA, and that is where the crypto ACL turns into something concrete: ``` Child sa: local selector -> remote_selector 10.20.10.0/0 - 10.20.10.255/65535 -> 10.30.10.0/0 - 10.30.10.255/65535 ESP spi in/out: 0xDC114E8E/0x854355BF Encr: AES-GCM, keysize: 256, esp_hmac: None ah_hmac: None, comp: IPCOMP_NONE, mode tunnel ``` Those local and remote selectors are IKEv2's name for what IKEv1 called proxy IDs, and they came directly from your crypto ACL: 10.20.10.0/24 to 10.30.10.0/24, ports and protocols wide open. If the two ends of a crypto-map VPN have mismatched ACLs, this is where you see it, as a selector that does not look like what you expected. Note `esp_hmac: None` again, for the same AEAD reason. ## READY, Not QM\_IDLE (The False Alarm That Gets Everyone) Here is the failure mode that generates more panicked tickets than any actual IKEv2 bug. Someone migrates a tunnel, then reaches for the command they have typed ten thousand times: ``` HUB#show crypto isakmp sa ``` And it shows nothing useful. No peer. No QM\_IDLE. The engineer concludes the tunnel is down, opens a P1, and starts rolling back a change that was working perfectly. The tunnel was never down. `show crypto isakmp sa` displays IKEv1 security associations. An IKEv2 tunnel does not create one, so there is nothing for that command to display. The IKE SA you built lives in a different table, and you get to it with `show crypto ikev2 sa`. IKEv1 **Show command:** `show crypto isakmp sa` **Healthy state:** `QM_IDLE` **Default IKE lifetime:** 3600 sec **Phase 2 naming:** proxy IDs IKEv2 **Show command:** `show crypto ikev2 sa` **Healthy state:** `READY` **Default IKE lifetime:** 86400 sec **Phase 2 naming:** traffic selectors One command that works for both, and the one you should reach for first when you do not know which flavour a tunnel is running, is `show crypto session`. It reports UP-ACTIVE regardless of IKE version and tells you the peer, the flow, and the packet counters. Learn that habit and the QM\_IDLE reflex stops costing you outages. ## Identity Matching Is Where IKEv2 Configs Die Two lines in the profile deserve more attention than the rest of the config put together: ``` crypto ikev2 profile PLZ-IKEV2-PROF match identity remote address 198.51.100.2 255.255.255.255 identity local address 203.0.113.1 ``` `match identity remote` is a selector. When an IKE\_AUTH exchange arrives, IOS looks at the identity the peer asserted and walks its IKEv2 profiles looking for one that matches. If a profile matches, that profile's authentication methods and keyring are used. If no profile matches, the negotiation is dropped. There is no fallback, no default profile, no "close enough". The peer's IP address being right in your keyring is not sufficient, because the keyring is only consulted after a profile has already been selected. `identity local` is the other half: it is the identity this router asserts about itself. The far end will run its own `match identity remote` against exactly this value. In our lab the HUB asserts 203.0.113.1 and SPOKE1's profile carries `match identity remote address 203.0.113.1 255.255.255.255`. The two must agree. This is the number one IKEv2 configuration failure, and it is easy to trip over in ways that look nothing like an identity problem: - **Multiple exit interfaces.** If you do not set `identity local`, the router asserts the IP of whichever interface the packet leaves through. Change the routing, change the identity, break the tunnel. - **NAT in the path.** The peer asserts its pre-NAT identity, but you configured `match identity remote address` using the post-NAT address you see in the packet's source field. Those are different values, and the match fails. - **FQDN on one side, address on the other.** A peer configured with `identity local fqdn` will never match your `match identity remote address`, however correct the addressing is. The identity type has to line up, not just the value. - **Overlapping profiles.** Two profiles, one with a specific address match and one with a wildcard, and the wrong one wins. Order and specificity matter. The symptom in every one of these cases is the same: authentication fails, or the tunnel simply never comes up, and the SA table is empty. The diagnostic is `debug crypto ikev2`, where you will see the identity the peer actually sent and can compare it against what you told the profile to expect. Nine times out of ten the two strings are visibly different and the fix takes ten seconds. For a fuller treatment of how the exchange works, see [IKEv2 explained](https://www.pinglabz.com/ikev2-explained/). ## Be Honest: This Is Still a Legacy Design Swapping the key exchange gives you IKEv2's real benefits. Better crypto negotiation. AEAD ciphers. A 24-hour SA lifetime. Built-in DPD (liveness checks are part of the protocol rather than a bolt-on). Anti-DoS cookies. Cleaner rekey behaviour that does not black-hole traffic mid-negotiation. Those are worth having, and you got them for one line of config. What you did not get is a better VPN architecture. A crypto map with IKEv2 is still a crypto map, and it still carries every structural limitation it always had: - **One IPsec SA pair per crypto ACL line.** Ten subnets on each side means a lot of ACEs and a lot of SAs, and the two ends must agree on every one of them. - **No routing protocols across the tunnel.** There is no interface to run OSPF or BGP on. Every remote subnet is a manual ACL edit on both routers, forever. - **No interface counters.** You cannot `show interface` the tunnel, cannot graph it in your NMS, cannot apply QoS to it in the obvious way, and cannot use interface state as a routing signal. - **No tunnel to fail over.** Redundancy means peer lists and reverse route injection rather than a routing protocol reconverging in a second. So treat this migration as what it is: a low-risk, high-value stepping stone. It buys you modern crypto today without a redesign, and it teaches you the IKEv2 object model (proposal, policy, keyring, profile) on config you already understand. When you have that, the profile you just built is the exact same profile a virtual tunnel interface consumes, so the next move is short. Our [IKEv2 SVTI configuration](https://www.pinglabz.com/ikev2-svti-configuration/) walk-through takes this same lab and reattaches `PLZ-IKEV2-PROF` to a tunnel interface instead, and [crypto maps versus VTI](https://www.pinglabz.com/crypto-maps-vs-vti/) lays out why that is where you want to end up. ## Key Takeaways - **The migration is one line.** `set ikev2-profile PLZ-IKEV2-PROF` inside the crypto map entry is what makes a crypto map negotiate with IKEv2\. The crypto ACL, transform set, and interface binding are unchanged from the IKEv1 build, and the map line still ends in `ipsec-isakmp`. - **Four objects, one that matters.** Proposal (was ISAKMP policy), policy (new), keyring (was ISAKMP key), and profile (new, and the one everything attaches to). - **AES-GCM takes a PRF and no integrity algorithm.** GCM is AEAD, so it authenticates internally. IOS prints a validation hint saying so, and `Hash: None` in the SA is the confirmation, not a problem. - **READY, not QM\_IDLE.** `show crypto isakmp sa` shows nothing for an IKEv2 tunnel. Use `show crypto ikev2 sa`, or use `show crypto session`, which works for both. - **Identity is explicit and is what the profile keys off.** `match identity remote` and `identity local` have to line up on both sides, in type as well as value. Mismatched identity is the single most common IKEv2 failure. - **Crypto maps are still legacy.** Better key exchange, same structural limits: SA pairs per ACL line, no routing protocols, no interface counters. Use this to modernise safely, then move to a VTI. Where this fits in the wider picture, including transform sets, NAT traversal, and the phase model this all sits on top of, is covered in the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/). All configuration and output above is from a live CML lab: HUB and SPOKE1 on cat8000v 17.18.02. ### IKEv2 Site-to-Site VPN with SVTI: The Config You Should Actually Deploy URL: https://www.pinglabz.com/ikev2-svti-configuration/ Last updated: 2026-07-13T08:08:30.000Z If you are building a brand new site-to-site tunnel in 2026, this is the one you should build: **IKEv2 for the control plane, AES-GCM-256 for the data plane, and a static VTI for the forwarding plane**. No crypto ACL. No proxy-ID mismatch at 2am. No legacy cipher suite that a scanner will flag next quarter. Just a routable tunnel interface that happens to be encrypted. This article is part of the PingLabz [IPsec VPN](https://www.pinglabz.com/ipsec-vpn/) cluster, and it is the config we would hand someone on day one. Everything below was captured on a real lab: a cat8000v HUB at 203.0.113.1 talking to a cat8000v SPOKE1 at 198.51.100.2, both on IOS XE 17.18.02, with an ISP router in between. Every line of CLI output on this page came off those boxes. ## Why This Combination and Not Another There are four moving decisions in a site-to-site build, and most of the bad tunnels in the world got at least one of them wrong. Key exchange: IKEv2 Fewer messages, built-in DPD/liveness, a real config-exchange mechanism, and modern smart defaults. IKEv1 is legacy. See [IKEv2 explained](https://www.pinglabz.com/ikev2-explained/). Cipher: AES-GCM-256 AEAD. Encryption and integrity in one pass, no separate HMAC, hardware-friendly, less per-packet overhead. Encapsulation: static VTI A real interface you can route out of, run an IGP over, apply QoS to, and put in a VRF. See [SVTI](https://www.pinglabz.com/static-vti-svti-ipsec/). Traffic selection: routing A route, not an access list. The crypto ACL and its proxy-ID mismatches disappear entirely. See [crypto maps vs VTI](https://www.pinglabz.com/crypto-maps-vs-vti/). The alternative build, IKEv2 bolted onto a crypto map, still works and we [lab it out here](https://www.pinglabz.com/ikev2-site-to-site-crypto-maps/) because you will meet it on brownfield gear. But if nothing is forcing you into a crypto map, do not choose one voluntarily. ## The Complete Build, Bottom to Top The whole config is five blocks. Read it once and it fits in your head. ### 1\. The IKEv2 proposal and policy The proposal is the cipher suite offered in IKE\_SA\_INIT. The policy is the container that decides which proposal applies to which peers. ``` crypto ikev2 proposal PLZ-IKEV2-PROP encryption aes-gcm-256 prf sha256 group 19 crypto ikev2 policy PLZ-IKEV2-POL proposal PLZ-IKEV2-PROP ``` Notice what is *missing*: there is no `integrity` line. With AES-GCM you configure a PRF and no integrity algorithm, because GCM does integrity itself. IOS XE will tell you this out loud while you are typing an incomplete proposal: ``` IKEv2 proposal MUST either have a set of an encryption algorithm other than aes-gcm, an integrity algorithm and a DH group configured or encryption algorithm aes-gcm, a prf algorithm and a DH group configured ``` That is not an error, it is a hint, and it is the AEAD rule stated in the CLI. Group 19 is a 256-bit ECP curve (a sane 2026 floor). ### 2\. The keyring Where the pre-shared key lives, scoped to a peer address. ``` crypto ikev2 keyring PLZ-KEYRING peer SPOKE1 address 198.51.100.2 pre-shared-key PingLabz-IKEv2-PSK-01 ``` ### 3\. The IKEv2 profile The profile is the glue: it says which remote identities we accept, what identity we present, how both sides authenticate, and which keyring to use. ``` crypto ikev2 profile PLZ-IKEV2-PROF match identity remote address 198.51.100.2 255.255.255.255 identity local address 203.0.113.1 authentication local pre-share authentication remote pre-share keyring local PLZ-KEYRING ``` Both `authentication local` and `authentication remote` are required. IKEv2 authentication is asymmetric by design (each side independently states how it proves itself), which is exactly why you can later move one side to certificates without touching the other. ### 4\. The IPsec profile (this is the line people miss) The transform set defines the data-plane cipher. The IPsec profile binds that transform set **and** the IKEv2 profile together, and then the tunnel interface consumes the whole bundle. ``` crypto ipsec transform-set PLZ-TS-GCM esp-gcm 256 mode tunnel crypto ipsec profile PLZ-IPSEC-PROF set transform-set PLZ-TS-GCM set ikev2-profile PLZ-IKEV2-PROF ``` That `set ikev2-profile PLZ-IKEV2-PROF` line inside `crypto ipsec profile` is the single most commonly forgotten command in this entire build. Leave it out and IOS will happily accept the config, bring the tunnel interface to up/up (it is a tunnel, it does not need crypto to be line-protocol up), and then quietly fall back to whatever default IKEv2 profile logic it can find instead of the one you carefully wrote. If your proposal, keyring, or identity settings are not being honoured, check this line before you check anything else. ### 5\. The tunnel interface ``` interface Tunnel1 ip address 10.0.1.1 255.255.255.252 tunnel source GigabitEthernet2 tunnel mode ipsec ipv4 tunnel destination 198.51.100.2 tunnel protection ipsec profile PLZ-IPSEC-PROF ``` Two lines carry all the weight. `tunnel mode ipsec ipv4` makes this an IPsec VTI rather than a GRE tunnel (get this wrong and you have built GRE-over-IPsec, which is a different thing with different overhead). `tunnel protection ipsec profile` attaches the crypto. That is it. There is no crypto map anywhere in this config, and no interface on the router has one applied. ## Steering Traffic In: Routing, Not a Crypto ACL With a crypto map, you tell the router what to encrypt by writing an access list, and the far end has to write the mirror image of it. Get the mask wrong on one side and you get a proxy-ID mismatch, which produces a tunnel that negotiates phase 1 beautifully and then refuses to pass a single packet. With a VTI, you tell the router what to encrypt by **routing it out the tunnel**: ``` ip route 10.30.10.0 255.255.255.0 Tunnel1 ``` That is the entire traffic-selection policy. Anything the routing table hands to Tunnel1 gets encrypted; anything it does not, does not. Because the VTI is a real Layer 3 interface, an IGP works exactly as well: run OSPF or EIGRP over the tunnel and let the far end advertise its own prefixes. You can also apply an output service policy, a QoS shaper, NetFlow, or an ACL to Tunnel1 like any other interface. None of that is possible on a crypto map, where the encryption is a property of the physical interface and the "tunnel" is not an object you can point at. ## Verification: What Correct Looks Like Bring the tunnel up with interesting traffic and check three things. **Does it pass traffic?** ``` HUB#ping 10.30.10.1 source Loopback10 repeat 5 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 9/12/17 ms ``` **Is the interface actually an IPsec tunnel?** ``` HUB#show interface Tunnel1 Tunnel1 is up, line protocol is up Tunnel protocol/transport IPSEC/IP Tunnel transport MTU 1446 bytes ``` `Tunnel protocol/transport IPSEC/IP` confirms `tunnel mode ipsec ipv4` took effect. If you see `GRE/IP` there, you built a GRE tunnel and wrapped it in IPsec, which is not what this article is about. **What is the crypto session actually protecting?** ``` HUB#show crypto session detail Session status: UP-ACTIVE Peer: 198.51.100.2 port 500 fvrf: (none) ivrf: (none) IPSEC FLOW: permit ip 0.0.0.0/0.0.0.0 0.0.0.0/0.0.0.0 Inbound: #pkts dec'ed 5 drop 0 life (KB/Sec) 4607999/3576 Outbound: #pkts enc'ed 5 drop 0 life (KB/Sec) 4607999/3576 ``` Look at that IPSEC FLOW line. `permit ip 0.0.0.0/0.0.0.0 0.0.0.0/0.0.0.0` is any-to-any. The VTI negotiates a wildcard traffic selector, which is precisely why there is no crypto ACL to mismatch. Both ends agree to protect "everything that arrives on this interface", and the routing table decides what arrives. If you have ever spent an afternoon diffing two crypto ACLs across a change window, that one line is the payoff for the whole design. ## The MTU Detail Nobody Measures: 1446 vs 1438 Most articles hand-wave IPsec MTU with a vague "expect around 1400". We measured it, twice, on the same platform, and the number moved with the cipher. SVTI with AES-GCM-256 **Transport MTU:** 1446 bytes **Integrity:** inside the cipher (AEAD) **Separate HMAC:** none SVTI with AES-CBC-256 + SHA-256 **Transport MTU:** 1438 bytes **Integrity:** separate HMAC field **Separate HMAC:** yes, 8 more bytes of overhead Same platform, same tunnel type, same physical MTU of 1500\. The only variable is the cipher, and the difference is 8 bytes. A CBC + SHA-256 ESP packet has to carry a truncated HMAC field on top of the encrypted payload. A GCM packet does not: the authentication tag is part of the AEAD construction and the ESP overhead is smaller. Eight bytes will not save your network on its own, but it is a concrete, checkable demonstration that AEAD is not just a security argument, it is a wire-efficiency argument too. And it is a good sanity check: if your GCM VTI is not showing 1446 on IOS XE with a 1500-byte underlay, something in the stack is not what you think it is. ## Why AES-GCM, Specifically AES-GCM is an AEAD cipher: authenticated encryption with associated data. Confidentiality and integrity are produced by a single pass over the data instead of "encrypt, then compute a separate HMAC over the ciphertext". Three consequences you can see: - **It shows up in the SA as an absence.** The IKEv2 SA reads `Encr: AES-GCM, keysize: 256, PRF: SHA256, Hash: None`, and the child SA reads `Encr: AES-GCM, keysize: 256, esp_hmac: None`. `Hash: None` is not a misconfiguration, it is the point. - **It is hardware-friendly.** Modern CPUs and ASICs implement AES-GCM natively, and the single-pass design parallelises where encrypt-then-MAC does not. - **It is smaller on the wire.** See the 8 bytes above. The one thing to respect: GCM is unforgiving about IV/nonce reuse. That is a job for the implementation, not for you, but it is the reason you should be running current code rather than an image from 2015. ## Where This Design Stops Working Be honest about the ceiling. A static VTI is **one tunnel interface per peer**. Every remote site needs its own Tunnel interface, its own `tunnel destination`, and its own routing. Add a keyring entry per peer, or accept a wildcard PSK you probably do not want. At three sites, that is fine. It is explicit, greppable, and trivially troubleshootable, and you should not over-engineer it. At ten sites it is annoying. At thirty sites it is a config-management problem you have invented for yourself, and if you also want spoke-to-spoke traffic, static VTIs cannot give it to you without a full mesh (n(n-1)/2 tunnels, which is a number that stops being funny fast). That is exactly the gap [FlexVPN](https://www.pinglabz.com/flexvpn-explained/) and [DMVPN](https://www.pinglabz.com/dmvpn/) exist to fill. Both replace the per-peer static tunnel with a dynamic one: a single virtual-template or mGRE interface that spawns tunnels on demand, with hub-pushed routing (FlexVPN's IKEv2 authorization policy) or NHRP-driven spoke-to-spoke shortcuts (DMVPN Phase 3). The crypto you learned here does not change; the encapsulation and the tunnel lifecycle do. ## Hardening Checklist Pin an explicit proposal IKEv2 smart defaults are genuinely modern (AES-CBC-256, SHA-512/384, DH 19/14/21) and they are great for a lab or a quick proof. In production, write the proposal yourself so you know exactly what you negotiated and so a code upgrade cannot silently change it. DH group 19 or better Group 19 (256-bit ECP) is the floor. Group 20 or 21 if the far end supports it. Anything in the single digits belongs in a museum. AES-GCM-256 on the data plane AEAD, no separate HMAC, less overhead. If a peer cannot do GCM, AES-CBC-256 with SHA-256 is the fallback, and you pay for it in bytes. PSK for the lab, certificates for scale A pre-shared key is fine for two peers you control. It does not scale, it cannot be revoked, and it ends up in a wiki. Move to [certificate authentication](https://www.pinglabz.com/ipsec-certificate-authentication/) before the peer count gets interesting. Tighten the identity match Match the remote identity on a /32 address (or an FQDN), not on a wildcard. A loose `match identity remote` plus a wildcard PSK is how you accidentally build an open VPN concentrator. Mind the MTU 1446 with GCM on a 1500 underlay. Set `ip tcp adjust-mss` on the tunnel if you have TCP applications and a path that eats ICMP fragmentation-needed messages. ## Key Takeaways - **IKEv2 + AES-GCM-256 + static VTI is the default 2026 site-to-site build.** Route-based, AEAD, and short enough to read on one screen. - **The IPsec profile is the join.** `set transform-set` gives it the cipher, `set ikev2-profile` gives it the IKEv2 policy. Forgetting the second line is the classic failure. - **The tunnel interface needs two magic lines:** `tunnel mode ipsec ipv4` and `tunnel protection ipsec profile`. Verify with `Tunnel protocol/transport IPSEC/IP`. - **There is no crypto ACL.** `show crypto session detail` proves it: `IPSEC FLOW: permit ip 0.0.0.0/0.0.0.0 0.0.0.0/0.0.0.0`. Routing selects the traffic, and an IGP over the VTI works because the VTI is a real interface. - **AEAD is measurable.** GCM gave us a 1446-byte transport MTU where CBC + SHA-256 gave 1438 on the same platform. Eight bytes, because GCM carries no separate HMAC, which is also why the SA reads `Hash: None`. - **One static VTI per peer.** Fine at three sites, painful at thirty. That is what FlexVPN and DMVPN are for. Build this one first, get comfortable with what a healthy `show crypto session detail` looks like, and then go read the rest of the [IPsec VPN](https://www.pinglabz.com/ipsec-vpn/) cluster to see how the same crypto gets reused underneath the dynamic designs. ### Crypto Maps vs VTI: Which Site-to-Site Design to Pick URL: https://www.pinglabz.com/crypto-maps-vs-vti/ Last updated: 2026-07-13T07:29:52.000Z Sooner or later someone hands you two routers, two sites, and "encrypt the traffic between them." On Cisco IOS XE you have two mainstream ways to do it: a **crypto map** or a **virtual tunnel interface** (VTI). Same IKE, same ESP, same transform set, same pre-shared key. Very different networks. This is the article you read when you have to pick. For wider context, start at the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/). We did not get this from a doc. We built *both* designs in CML, on the same two routers (EDGE1 and EDGE2), carrying the same ping between the same two LANs, and captured what each one produced. Every piece of CLI below is real output from that lab. The difference shows up in one line of `show crypto session detail`, and the whole comparison falls out of it. ## The one difference that explains everything A crypto map is **policy-based**. You write an access list describing interesting traffic and bolt the map to the physical outside interface. The rule is: "if a packet leaving Gi2 matches this ACL, encrypt it to this peer." A VTI is **route-based**. You build a tunnel interface, set `tunnel mode ipsec ipv4`, protect it with an IPsec profile, and route traffic into it. The rule is: "anything routed out Tunnel0 leaves encrypted." That is not a syntax preference. It changes what the tunnel can carry, how you operate it, and how it fails. ## Proxy IDs: two narrow flows versus one wide one This is the contrast that *is* the article. Our crypto ACL on EDGE1 had two lines, one per LAN. Two ACL lines, two proxy identities, two SA pairs. The real session on the crypto map build, after live traffic from both source networks: ``` EDGE1#show crypto session detail Interface: GigabitEthernet2 Session status: UP-ACTIVE Peer: 198.51.100.2 port 500 fvrf: (none) ivrf: (none) IKEv1 SA: local 203.0.113.1/500 remote 198.51.100.2/500 Active IPSEC FLOW: permit ip 10.20.10.0/255.255.255.0 10.30.10.0/255.255.255.0 Active SAs: 2, origin: crypto map Inbound: #pkts dec'ed 10 drop 0 life (KB/Sec) 4607998/3516 Outbound: #pkts enc'ed 10 drop 0 life (KB/Sec) 4607999/3516 IPSEC FLOW: permit ip 192.168.99.0/255.255.255.0 10.30.10.0/255.255.255.0 Active SAs: 2, origin: crypto map Inbound: #pkts dec'ed 7 drop 0 life (KB/Sec) 4607999/3562 Outbound: #pkts enc'ed 7 drop 0 life (KB/Sec) 4607999/3562 ``` Now the same two routers, same ping, same ESP, rebuilt as a static VTI: ``` EDGE1#show crypto session detail IPSEC FLOW: permit ip 0.0.0.0/0.0.0.0 0.0.0.0/0.0.0.0 Active SAs: 2, origin: crypto map Inbound: #pkts dec'ed 18 drop 0 Outbound: #pkts enc'ed 14 drop 0 ``` One flow. Any to any. The VTI does not care what is inside the packet: if it arrived on Tunnel0, it gets encrypted. (Ignore the cosmetic `origin: crypto map` string, IOS reuses the internal database. There is no crypto map in the VTI config.) Everything else follows from those two blocks. A new subnet behind the crypto map means a new ACL line on *both* ends and a renegotiated SA pair. Behind the VTI it means a route. Full builds: [site-to-site IPsec with crypto maps](https://www.pinglabz.com/site-to-site-ipsec-crypto-maps/) and [static VTI (SVTI) IPsec](https://www.pinglabz.com/static-vti-svti-ipsec/). ## Routing protocols: the VTI has a neighbor, the crypto map cannot A crypto ACL matches unicast flows between subnets. OSPF hellos are multicast to 224.0.0.5, sourced by the router itself: they do not match your crypto ACL, they are not "interesting traffic," and there is no interface to run OSPF on because the map lives on the public WAN port. You cannot form an adjacency across a crypto map. The VTI is a genuine interface, so it just works: ``` EDGE1#show interface Tunnel0 Tunnel0 is up, line protocol is up Tunnel source 203.0.113.1 (GigabitEthernet2), destination 198.51.100.2 Tunnel protocol/transport IPSEC/IP Tunnel transport MTU 1438 bytes EDGE1#show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 10.255.0.2 0 FULL/ - 00:00:31 10.0.0.2 Tunnel0 EDGE1#ping 10.30.10.1 source Loopback10 repeat 5 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 15/20/35 ms ``` FULL adjacency, over an encrypted tunnel, no crypto ACL anywhere. Pick a crypto map and you are signing up for **static routes forever**, hand-maintained at both ends, plus a matching pair of ACL lines per prefix. On two sites that is annoying. On twenty it is a job. ## Config size: same security, a fraction of the lines The crypto map build needs an ISAKMP policy, a PSK, a transform set, a crypto ACL, a crypto map (peer, transform, PFS, match), and the map bound to the physical interface: ``` crypto isakmp policy 10 encryption aes 256 hash sha256 authentication pre-share group 14 lifetime 3600 crypto isakmp key PingLabz-L2L-PSK-01 address 198.51.100.2 crypto ipsec transform-set PLZ-TS-AES256 esp-aes 256 esp-sha256-hmac mode tunnel ip access-list extended PLZ-CRYPTO-ACL permit ip 192.168.99.0 0.0.0.255 10.30.10.0 0.0.0.255 permit ip 10.20.10.0 0.0.0.255 10.30.10.0 0.0.0.255 crypto map PLZ-CMAP 10 ipsec-isakmp set peer 198.51.100.2 set transform-set PLZ-TS-AES256 set pfs group14 match address PLZ-CRYPTO-ACL interface GigabitEthernet2 crypto map PLZ-CMAP ``` The VTI keeps the ISAKMP policy, key and transform set (that is the crypto, not the design), drops the ACL and the map, and replaces them with this: ``` interface Tunnel0 ip address 10.0.0.1 255.255.255.252 tunnel source GigabitEthernet2 tunnel mode ipsec ipv4 tunnel destination 198.51.100.2 tunnel protection ipsec profile PLZ-IPSEC-PROF ``` Six lines. Same AES-256, same SHA-256, same PFS group 14, same peer. What you deleted was not security, it was bookkeeping. ## Operations: you can actually see a VTI A tunnel interface is a first-class interface, with day-two consequences nobody mentions until the outage. - **Counters and state.** `show interface Tunnel0` gives line protocol, MTU, drops and rates. A crypto map has no interface, so "is the VPN up" means squinting at SA counters. - **MTU is visible.** The SVTI reports 1438 bytes transport MTU (1476 for plain GRE with its 4-byte header). - **Per-tunnel QoS.** You can attach a service policy to a tunnel interface. You cannot attach one to "traffic that matched ACL line 2." - **Independent shut / no shut,** without touching the WAN port every other service rides on. - **Routing metrics.** It has a cost, so failover is just routing. ## The failure modes crypto maps own The number one cause of site-to-site VPN failures is the **mirror-image crypto ACL**. Both ends must describe the same flows in exact reverse. Get one line wrong and Phase 1 comes up beautifully while Phase 2 dies. We broke it on purpose: ``` EDGE2#show logging Crypto mapdb : proxy_match map_db_find_best did not find matching map crypto_mapdb_find_map Fail to find matching policy ivrf 0, fvrf 0, flags 129, ike_profile NULL, local_proxy 10.30.10.0, remote_proxy 10.20.10.0, km_local 198.51.100.2, km_remote 203.0.113.1 ISAKMP-ERROR: (1018):IPSec policy invalidated proposal with error 32 ISAKMP-ERROR: (1018):phase 2 SA policy not acceptable! (local 198.51.100.2 remote 203.0.113.1) ISAKMP-ERROR: (1018):deleting node 1949939457 error TRUE reason "QM rejected" ``` **Error 32 means the proxy identities do not match.** Worse, the side whose ACL no longer matches does not even try: it sends the packet in the clear and it gets dropped upstream. The reported symptom is "nothing happens." Then IOS twists the knife when you go to fix it: ``` EDGE2(config)#ip access-list extended PLZ-CRYPTO-ACL %ACL PLZ-CRYPTO-ACL can not be modified/deleted, as it is used in crypto-map PLZ-CMAP %Please first remove the ACL from crypto map or remove the crypto map from the interface % Cannot modify this ACL. ``` You have to pull the map off the interface (dropping the production tunnel) to edit the ACL. **A VTI cannot fail this way: it has no crypto ACL to mismatch.** The second trap, hit by accident in the lab: a **static crypto map entry shadowing a dynamic one**. EDGE1 had a static entry whose ACL matched the proxy pair but whose peer was the old public address. IOS matched the static entry first, saw the peer did not match, and rejected instead of falling through to the dynamic entry: ``` IPSEC(ipsec_process_proposal): peer address 198.51.100.6 not found ISAKMP-ERROR: (1011):IPSec policy invalidated proposal with error 64 ISAKMP-ERROR: (1011):phase 2 SA policy not acceptable! (local 203.0.113.1 remote 198.51.100.6) ``` **Error 64 means the peer address was not found in any crypto map entry.** Crypto map ordering is policy evaluation, and policy evaluation has precedence bugs. A route-based tunnel has no policy list to get wrong. Full fault catalogue: [troubleshooting IPsec VPNs](https://www.pinglabz.com/troubleshooting-ipsec-vpn/). ## When a crypto map is still the right answer "Always use a VTI" is a slogan, not engineering. Crypto maps are still correct in real situations: - **Interop with a policy-based peer.** Plenty of third-party and legacy firewalls only negotiate specific proxy IDs. Offer any/any and they reject it. You match their design or there is no tunnel. - **A peer that will not do VTI.** The far end is often not yours to change. - **Existing deployments you cannot touch.** A working crypto map inside a change freeze is not this quarter's problem. - **A responder taking connections from many unknown peers.** The dynamic crypto map is the classic hub-side pattern for spokes on dynamic addresses. ## The decision Choose a VTI when You own both ends and both do route-based IPsec. You want a routing protocol across the tunnel. Subnets behind either site will change. You want counters, QoS and shut/no shut. **This is the default answer.** Choose a crypto map when The peer is third-party or legacy, policy-based only. The far end dictates specific proxy IDs. You are extending a deployment you cannot re-architect. You are a responder for many unknown peers (dynamic crypto map). Choose GRE over IPsec when You must carry multicast or non-IP protocols. Something depends on the GRE header itself. You accept the 4 extra bytes (1476 MTU versus 1438). Choose DMVPN or FlexVPN when You have many sites, not two. Spokes are dynamically addressed or behind NAT. You want on-demand spoke-to-spoke tunnels, not hub hairpinning. See the [DMVPN pillar](https://www.pinglabz.com/dmvpn/). Need multicast and the GRE header? [GRE over IPsec](https://www.pinglabz.com/gre-over-ipsec/) is the middle ground. ## Key Takeaways - **Policy-based versus route-based is the whole story.** A crypto map encrypts traffic matching an ACL. A VTI encrypts whatever you route out the interface. - **The proxy IDs prove it.** Same routers, same traffic: the crypto map negotiated two narrow IPSEC FLOWs (one per ACL line), the SVTI negotiated one, `permit ip 0.0.0.0/0.0.0.0 0.0.0.0/0.0.0.0`. - **Only the VTI runs a routing protocol.** FULL OSPF adjacency across the encrypted tunnel. A crypto map cannot carry multicast hellos, so you get static routes forever. - **Six lines versus a page,** with no loss of security (same AES-256, SHA-256, PFS group 14). - **Crypto maps own the ugliest failure mode:** mirror-image ACL drift, error 32, Phase 1 up while Phase 2 dies. A VTI has no crypto ACL to mismatch. Error 64 (a static entry shadowing a dynamic one) is the other to know. - **Default to a VTI.** Reach for a crypto map only when the peer, the vendor or the change process forces you to. Both designs, both captures, and the rest of the IKE and NAT-T story: the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/). ### Static VTI (SVTI): The Modern Way to Build a Site-to-Site VPN URL: https://www.pinglabz.com/static-vti-svti-ipsec/ Last updated: 2026-07-13T07:29:52.000Z If you are still building site-to-site VPNs with crypto maps in 2026, you are configuring a router the way it was done in 1999\. The Static Virtual Tunnel Interface (SVTI) is the modern answer, and the whole idea fits in one command: `tunnel mode ipsec ipv4`. That single line turns a tunnel interface into a routable, IPsec-protected interface. No crypto ACL. No crypto map. No GRE header. If you want the fundamentals underneath all of this (IKE, ESP, transform sets, proxy IDs), start with the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/) and come back here. Everything below was captured on a live CML lab: two cat8000v routers running IOS XE 17.18.02 across a simulated ISP, with the same ping and the same capture point used for the [crypto-map build](https://www.pinglabz.com/site-to-site-ipsec-crypto-maps/) and the [GRE over IPsec build](https://www.pinglabz.com/gre-over-ipsec/). That matters, because the punchline of this article is a direct before-and-after comparison of the same tunnel. ## The Mental Shift: Policy-Based vs Route-Based This is the part that trips people up, so let's be blunt about it. With a **crypto map**, you tell the router *which traffic to encrypt*. You write an extended ACL (the crypto ACL), you attach it to a crypto map entry, you hang the map off a physical interface, and IOS inspects every packet leaving that interface to see if it matches. Match the ACL, get encrypted. Miss the ACL, go out in the clear. The ACL is your policy. That is why this is called **policy-based** VPN. With an **SVTI**, you tell the router *nothing* about which traffic to encrypt. You build a tunnel interface, you protect it with an IPsec profile, and then you *route* traffic at it. Anything that gets routed out Tunnel0 is encrypted, because that is what the interface does. Anything routed anywhere else is not. Your routing table is your policy. That is **route-based** VPN. Crypto map (policy-based) Selects traffic with: a crypto ACL Proxy IDs: one SA pair per ACL line Applied to: a physical interface Routing protocols: No (no multicast) Interface counters: No Per-tunnel QoS: No SVTI (route-based) Selects traffic with: the routing table Proxy IDs: one SA pair, any/any Applied to: a tunnel interface Routing protocols: Yes Interface counters: Yes Per-tunnel QoS: Yes ## The One-Line Change (and the Warning IOS Prints) Our lab already had a GRE tunnel with `tunnel protection ipsec profile` applied (that is classic GRE over IPsec). Turning it into an SVTI is exactly one command on the tunnel interface: ``` EDGE1(config-if)#tunnel mode ipsec ipv4 %WARNING: The tunnel mode has been modified while the tunnel protection is already active. It is recommended to run "shutdown" and "no shutdown" on Tunnel0 interface to refresh the config. ``` Read that warning carefully, because it is one of the most useful things IOS XE will tell you all day. You just changed the encapsulation on an interface that already has an active IPsec SA bound to it, and the crypto engine does not silently rebuild itself. Until you bounce the interface (`shutdown`, then `no shutdown`, on Tunnel0), you can be staring at a config that says one thing and a data plane that is doing another. In a lab this costs you three seconds. In production it is a change-window item, and it is the single most common reason someone says "I converted to VTI and nothing changed." The mode changed. The tunnel did not. ## Proof #1: The Interface Is No Longer GRE Before the change, the tunnel was GRE. After it, the same interface reports a different transport entirely: ``` EDGE1#show interface Tunnel0 Tunnel0 is up, line protocol is up Tunnel source 203.0.113.1 (GigabitEthernet2), destination 198.51.100.2 Tunnel protocol/transport IPSEC/IP <-- was GRE/IP Tunnel transport MTU 1438 bytes <-- was 1476 (no 4-byte GRE header) ``` Two lines, two facts. **"Tunnel protocol/transport IPSEC/IP"** means the interface is no longer wrapping your packets in a GRE header and then handing that to IPsec. IPsec *is* the encapsulation now. There is no inner GRE at all. **Transport MTU 1438 versus 1476.** A 1500-byte Ethernet MTU minus a 20-byte outer IP header minus a 4-byte GRE header gives you the classic 1476 GRE figure. Drop the GRE header and the accounting changes: IOS now reports 1438, which is exactly the same number the crypto engine reported as `plaintext mtu 1438` in `show crypto ipsec sa` on the crypto-map build. That is not a coincidence. With SVTI the interface MTU and the crypto MTU are the same thing, because the interface *is* the crypto boundary. With GRE over IPsec they were two separate accounting layers stacked on each other, and reconciling them by hand is where fragmentation bugs are born. ## Proof #2: OSPF Still Comes Up (This Surprises People) There is a widespread belief that the moment you drop GRE you lose multicast, therefore you lose your routing protocols, therefore you are back to static routes. It is a reasonable fear (raw IPsec tunnel mode really is a unicast, IP-only transport) and on IOS XE it is simply not what happens. The IPsec VTI still carries the IGP: ``` EDGE1#show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 10.255.0.2 0 FULL/ - 00:00:31 10.0.0.2 Tunnel0 ``` FULL, over Tunnel0, with GRE gone. Same adjacency we had on the plain GRE tunnel, same adjacency we had on GRE over IPsec. The VTI is treated as a real point-to-point interface by the IP stack, so OSPF's hellos get handled and the neighbor forms. This kills the last honest argument for GRE over IPsec on a simple two-site design. You used to need GRE for one reason: to give multicast and non-IP traffic a ride through IPsec. If your only non-unicast requirement is an IPv4 IGP, an SVTI covers it. (If you genuinely need to carry multicast *user* traffic, or a non-IP protocol, GRE is still your answer, and that is exactly the call the [crypto maps vs VTI decision guide](https://www.pinglabz.com/crypto-maps-vs-vti/) walks through.) And the data plane works, on the same ping we have been running all lab: ``` EDGE1#ping 10.30.10.1 source Loopback10 repeat 5 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 15/20/35 ms ``` ## Proof #3: The Defining SVTI Capture This is the one. If you remember nothing else from this article, remember this output: ``` EDGE1#show crypto session detail IPSEC FLOW: permit ip 0.0.0.0/0.0.0.0 0.0.0.0/0.0.0.0 <-- ANY/ANY. No crypto ACL. Active SAs: 2, origin: crypto map Inbound: #pkts dec'ed 18 drop 0 Outbound: #pkts enc'ed 14 drop 0 ``` `permit ip 0.0.0.0/0.0.0.0 0.0.0.0/0.0.0.0`. Any source, any destination. That is the proxy ID the SVTI negotiated with the peer, and we never wrote an ACL to produce it. The tunnel does not care what is inside the packets. It encrypts whatever the routing table sends it. Now put that next to what the *same two routers* negotiated when they were running a crypto map with a two-line crypto ACL: ``` EDGE1#show crypto session detail Interface: GigabitEthernet2 Session status: UP-ACTIVE Peer: 198.51.100.2 port 500 fvrf: (none) ivrf: (none) IKEv1 SA: local 203.0.113.1/500 remote 198.51.100.2/500 Active IPSEC FLOW: permit ip 10.20.10.0/255.255.255.0 10.30.10.0/255.255.255.0 Active SAs: 2, origin: crypto map Inbound: #pkts dec'ed 10 drop 0 life (KB/Sec) 4607998/3516 Outbound: #pkts enc'ed 10 drop 0 life (KB/Sec) 4607999/3516 IPSEC FLOW: permit ip 192.168.99.0/255.255.255.0 10.30.10.0/255.255.255.0 Active SAs: 2, origin: crypto map Inbound: #pkts dec'ed 7 drop 0 life (KB/Sec) 4607999/3562 Outbound: #pkts enc'ed 7 drop 0 life (KB/Sec) 4607999/3562 ``` Two crypto ACL lines produced **two separate IPSEC FLOWs**, each with its own SA pair. The SVTI produced **one**, and it covers everything. Same peers, same transform set, same IPsec profile. The only difference is how traffic gets selected. (Yes, `origin: crypto map` still shows up in the SVTI output. That is an internal label IOS uses for the tunnel-protection construct. There is no crypto map anywhere in the configuration.) ## Why "One SA Pair, Any/Any" Scales Better The crypto ACL is not just ugly. It is a scaling liability, and it fails in a specific, miserable way. SA count Crypto map: one SA pair per crypto ACL line, so N subnets at one end and M at the other can mean N x M pairs. SVTI: always one pair. Adding a subnet Crypto map: edit the ACL on BOTH peers, in sync, or Phase 2 fails. SVTI: add a route. The tunnel config does not change at all. Mismatch failure mode Crypto map: proxy-ID mismatch, "IPSec policy invalidated proposal with error 32". SVTI: any/any has nothing to mismatch. Editing in production Crypto map: IOS XE refuses to let you edit an ACL bound to an active map. SVTI: there is no ACL to edit. That third row is the killer. In the troubleshooting section of this same lab, breaking one crypto ACL line produced this: ``` ISAKMP-ERROR: (1018):IPSec policy invalidated proposal with error 32 ISAKMP-ERROR: (1018):phase 2 SA policy not acceptable! (local 198.51.100.2 remote 203.0.113.1) ``` Error 32 is a proxy-ID mismatch: the two ends disagree about which networks the SA is supposed to cover. It is one of the top reasons a site-to-site VPN that "was working yesterday" is down today, and it exists *only* because the crypto ACL exists. An SVTI negotiating any/any has nothing to disagree about. The fourth row is the operational insult on top of it. Try to fix a crypto ACL on a live router: ``` EDGE2(config)#ip access-list extended PLZ-CRYPTO-ACL %ACL PLZ-CRYPTO-ACL can not be modified/deleted, as it is used in crypto-map PLZ-CMAP %Please first remove the ACL from crypto map or remove the crypto map from the interface % Cannot modify this ACL. ``` To add one subnet to your VPN you have to pull the crypto map off the interface (which drops the tunnel), edit the ACL, put the map back, and hope you did the same thing on the far end. With a VTI you type `ip route`, or you let the IGP do it, and nobody notices. ## How You Steer Traffic Into an SVTI This is the practical question everyone asks next. If there is no crypto ACL, how does the router know to encrypt the branch subnet? It does not "know" anything. You *route* it. Three options, in rough order of how often you should reach for them: - **An IGP over the tunnel.** This is what our lab does. OSPF is adjacent over Tunnel0 (FULL, as shown above), the far end advertises 10.30.10.0/24, our routing table learns it with Tunnel0 as the outgoing interface, and every packet destined there gets encrypted. Add a subnet at the branch and it just shows up. Nothing on the VPN touches crypto. - **A static route out the tunnel.** `ip route 10.30.10.0 255.255.255.0 Tunnel0`. Perfectly fine for a small, stable two-site design where you do not want an IGP across the WAN. It is still route-based, it is just a manual route. - **A default route out the tunnel.** `ip route 0.0.0.0 0.0.0.0 Tunnel0` gives you full-tunnel backhaul: everything, including internet-bound traffic, goes to the hub encrypted. The any/any proxy ID already supports it, so no crypto change is needed. (Watch for recursive routing here: the tunnel destination itself must stay reachable by a more specific route, not through the tunnel.) Compare that to the crypto-map world, where "steering" means hand-editing an ACL on two devices and praying they match. The routing table was always the right tool for this job. Route-based VPN just lets you use it. ## The Two Things a Crypto Map Simply Cannot Do **Interface counters.** A VTI is a real interface. `show interface Tunnel0` gives you input and output rates, packet and byte counters, drops, errors, and an up/down line protocol you can track and alarm on. A crypto map has none of this: it is bolted onto a physical interface, and your only visibility is the `#pkts encaps` counters buried in `show crypto ipsec sa`. You cannot graph a crypto map in your NMS. You can graph a tunnel interface, because SNMP treats it as an interface like any other. **Per-tunnel QoS.** Because the VTI is an interface, you can apply a service policy to it with `service-policy output`, so you can shape, prioritise, and queue on a per-VPN basis, before encryption. There is no equivalent for a crypto map, because a crypto map is not an interface and there is nothing to attach a policy to. If you are running voice or video over a site-to-site VPN, this alone decides the design. (The [QoS cluster](https://www.pinglabz.com/qos/) covers the policy side.) There is a quieter third one: an SVTI can be placed in a VRF, given a per-tunnel `ip mtu` and `ip tcp adjust-mss`, and treated by every other IOS feature as a normal Layer 3 interface, because that is exactly what it is. ## The Complete SVTI Config Here is the entire tunnel side of a working, IPsec-protected, OSPF-carrying, 100-percent-ping site-to-site VPN. Six lines: ``` interface Tunnel0 ip address 10.0.0.1 255.255.255.252 tunnel source GigabitEthernet2 tunnel mode ipsec ipv4 tunnel destination 198.51.100.2 tunnel protection ipsec profile PLZ-IPSEC-PROF ``` Plus the IPsec profile it references, which is just a transform set and PFS, and is shared by every tunnel on the box: ``` crypto ipsec profile PLZ-IPSEC-PROF set transform-set PLZ-TS-AES256 set pfs group14 ``` No `crypto map`. No `match address`. No `ip access-list extended PLZ-CRYPTO-ACL`. Nothing bolted onto GigabitEthernet2 at all. The IKE policy and the pre-shared key still exist (an SVTI is still IPsec, and Phase 1 still has to happen), but the entire traffic-selection apparatus is gone. ## Where SVTI Stops SVTI is static: `tunnel destination 198.51.100.2` is hard-coded. That is fine for two sites, or a handful of branches homed back to a hub. It is not fine for a full mesh, because a full mesh of N sites means N x (N-1) / 2 statically configured tunnels, and every new site means touching every existing site. That is the exact problem [DMVPN](https://www.pinglabz.com/dmvpn/) solves, using mGRE and NHRP to build spoke-to-spoke tunnels on demand from a single template interface. FlexVPN takes the same idea further: an IKEv2 framework where the tunnel interfaces are dynamic (virtual-access interfaces cloned from a virtual-template) and the policy is pushed from the hub. Learn SVTI properly and neither one is magic. They are just VTIs with a control plane that builds them for you. ## Key Takeaways - **`tunnel mode ipsec ipv4` is the whole feature.** It converts a tunnel interface from GRE to native IPsec: route-based, no crypto ACL, no crypto map, no GRE header. - **Bounce the interface after changing the mode.** IOS XE warns you that it is "recommended to run shutdown and no shutdown on Tunnel0 interface to refresh the config." On an already-protected tunnel, the change does not take effect until you do. - **`show interface Tunnel0` is your proof.** "Tunnel protocol/transport IPSEC/IP" (not GRE/IP) and transport MTU 1438 (not 1476) confirm the 4-byte GRE header is gone. - **SVTI still carries OSPF.** Our lab shows the neighbor FULL over Tunnel0 with GRE removed. Dropping GRE does not mean dropping your IGP on IOS XE. - **`IPSEC FLOW: permit ip 0.0.0.0/0.0.0.0 0.0.0.0/0.0.0.0`** is the defining SVTI capture: one any/any SA pair, versus one SA pair per crypto ACL line on a crypto map. Nothing to mismatch, nothing to keep in sync, and error 32 becomes impossible. - **You steer traffic with routing, not ACLs.** An IGP over the tunnel, a static route out Tunnel0, or a default route for full-tunnel backhaul. Adding a subnet means adding a route, not editing crypto on both peers. - **A VTI is a real interface.** Interface counters, SNMP, and per-tunnel QoS service policies all work. A crypto map cannot do any of them. - **Use SVTI by default for site-to-site.** Reach for GRE over IPsec only when you must carry multicast user traffic or non-IP protocols, and move to DMVPN or FlexVPN when the topology outgrows static peers. Still deciding which design to deploy? The [crypto maps vs VTI decision guide](https://www.pinglabz.com/crypto-maps-vs-vti/) puts the options side by side, and the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/) is the hub for the whole cluster: IKE and ESP fundamentals, the [crypto-map build](https://www.pinglabz.com/site-to-site-ipsec-crypto-maps/), [GRE over IPsec](https://www.pinglabz.com/gre-over-ipsec/), NAT traversal, and the troubleshooting error-code map. From here the natural next step is [DMVPN](https://www.pinglabz.com/dmvpn/), where the tunnel you just built stops being static and starts building itself. ### Troubleshooting IPsec: Phase 1 Up, Phase 2 Down, and Other Classics URL: https://www.pinglabz.com/troubleshooting-ipsec-vpn/ Last updated: 2026-07-13T07:29:53.000Z Every error message in this article came off a real router. Not a textbook, not a vendor doc, not a forum post from 2011\. We built a live IPsec lab in CML (two cat8000v edge routers, an ISP router, a PAT device in the middle), got the tunnel working, and then deliberately broke it five different ways so we could capture exactly what IOS XE says when each thing goes wrong. This is the troubleshooting article for the [IPsec VPN cluster](https://www.pinglabz.com/ipsec-vpn/), and it is the one worth bookmarking, because the output below is what you will actually see on your screen at 2am. IPsec troubleshooting has a reputation for being black magic. It is not. It is a two-phase protocol, and almost every fault you will ever hit lands cleanly on one side or the other of that line. The whole skill is knowing which side you are on before you start typing. ## Part 1: A Method, Not a Guess There is exactly one question that halves the problem, and it should be the first command you run every single time: ``` EDGE1#show crypto isakmp sa IPv4 Crypto ISAKMP SA dst src state conn-id status 198.51.100.2 203.0.113.1 QM_IDLE 1001 ACTIVE ``` That single state field tells you which half of the protocol to interrogate. Do not skip it and go straight to debugs, or you will drown in output you have no context for. No SA at all **Means:** IKE never got off the ground. **Look at:** reachability to the peer, the peer address in your crypto map, and your ISAKMP policy. Nothing about Phase 2 is relevant yet. MM\_KEY\_EXCH (stuck) **Means:** the two sides are talking but cannot authenticate each other. **Look at:** the pre-shared key. It is almost always the pre-shared key. QM\_IDLE, no traffic **Means:** Phase 1 is fine. Stop looking at it. **Look at:** Phase 2\. Transform set, crypto ACL, or NAT. Nothing else. QM\_IDLE confuses people because it looks like an idle state, as if something is waiting. It is not. QM\_IDLE means Quick Mode is not currently running because it has already finished. It is the healthy steady state of an IKEv1 Phase 1 SA. If you see QM\_IDLE and traffic still is not flowing, you have just eliminated half the protocol. That is a win. ### The Second Question: Are Packets Actually Moving? Once you know Phase 1 is up, the next command tells you whether Phase 2 built and whether packets are traversing it in both directions: ``` EDGE1#show crypto ipsec sa local ident (addr/mask/prot/port): (10.20.10.0/255.255.255.0/0/0) remote ident (addr/mask/prot/port): (10.30.10.0/255.255.255.0/0/0) current_peer 198.51.100.2 port 500 PERMIT, flags={origin_is_acl,} #pkts encaps: 9, #pkts encrypt: 9, #pkts digest: 9 #pkts decaps: 9, #pkts decrypt: 9, #pkts verify: 9 #send errors 0, #recv errors 0 ``` Two counters matter more than everything else on that screen: `#pkts encaps` and `#pkts decaps`. Read them as a pair. - **encaps climbing, decaps stuck at 0.** Your packets are being encrypted and sent. Nothing is coming back. The tunnel is not the problem, the return path is. Check the remote crypto ACL (is it a mirror image of yours?), check the return route on the far side, and check whether something between you is dropping ESP (IP protocol 50 has no port number, so a lot of middleboxes quietly eat it). - **both 0.** Phase 2 never came up at all. No SA was ever installed, so no packet was ever encapsulated. Go find the Quick Mode rejection. Parts 2 and 3 of this article are entirely about reading those. - **both climbing.** IPsec is working. Your problem is somewhere else: routing, an ACL on an inside interface, or the application itself. Those two commands, in that order, resolve the vast majority of real-world IPsec cases before you type a single `debug`. ## Part 2: Five Faults, Broken For Real Each of the following was configured, broken, and captured on a live IOS XE 17.18.02 router. The output is verbatim. ### Fault A: Pre-Shared Key Mismatch The classic. Somebody typed the PSK from memory on one side. Here is what it looks like: ``` EDGE2#show crypto isakmp sa dst src state conn-id status 203.0.113.1 198.51.100.2 MM_KEY_EXCH 1014 ACTIVE EDGE2#show logging %CRYPTO-4-IKMP_BAD_MESSAGE: IKE message from 203.0.113.1 failed its sanity check or is malformed ``` Two symptoms, and the second one trips people up. The ISAKMP SA parks in `MM_KEY_EXCH` and never advances, and the log fills with complaints about a "malformed" message from the peer. Nothing is actually malformed. In IKEv1 Main Mode, once the Diffie-Hellman exchange completes, both sides derive keying material from the DH shared secret *plus the pre-shared key*, and from message 5 onward the payloads are encrypted with a key that includes the PSK. If your PSK does not match the peer's, you decrypt their payload with the wrong key and get random bytes. The router tries to parse an ISAKMP payload header out of those bytes, finds nonsense, and reports the only thing it honestly can: this message failed its sanity check. So `IKMP_BAD_MESSAGE` is not a corruption problem or an MTU problem. It is a decryption failure wearing a costume. When you see it alongside `MM_KEY_EXCH`, check the key on both sides, and check the address it is bound to: ``` crypto isakmp key PingLabz-L2L-PSK-01 address 198.51.100.2 ``` ### Fault B: Transform Set Mismatch (error 256) Phase 1 comes up perfectly. `show crypto isakmp sa` shows QM\_IDLE and you feel good about yourself. Then no traffic passes and the logs say this: ``` EDGE1#show logging ISAKMP-ERROR: (1015):IPSec policy invalidated proposal with error 256 ISAKMP-ERROR: (1015):phase 2 SA policy not acceptable! (local 203.0.113.1 remote 198.51.100.2) ISAKMP-ERROR: (1015):deleting node 1723779313 error TRUE reason "QM rejected" ``` This is the textbook shape of a Phase 2 failure. Phase 1 succeeded, so the peers agree on who each other is. They just cannot agree on how to protect the data. One side is offering, say, `esp-aes 256 esp-sha256-hmac` and the other will only accept something else, so the proposal is invalidated and Quick Mode is torn down. `QM rejected` is IOS telling you plainly that Phase 2, not Phase 1, is what died. Compare the transform sets first: ``` crypto ipsec transform-set PLZ-TS-AES256 esp-aes 256 esp-sha256-hmac mode tunnel ``` Both sides need a matching encryption algorithm, a matching hash, and a matching mode. If you are running PFS (`set pfs group14`), the DH group has to match too, or Quick Mode will fail for the same reason with the same error code. Full crypto map syntax lives in [the site-to-site IPsec crypto map walkthrough](https://www.pinglabz.com/site-to-site-ipsec-crypto-maps/). ### Fault C: Crypto ACL / Proxy-ID Mismatch (error 32) This is the one that eats the most hours, and it has two completely different symptoms depending on which router you are logged into. On the responder, you get one of the most useful debug blocks IOS produces: ``` EDGE2#show logging Crypto mapdb : proxy_match map_db_find_best did not find matching map crypto_mapdb_find_map Fail to find matching policy ivrf 0, fvrf 0, flags 129, ike_profile NULL, local_proxy 10.30.10.0, remote_proxy 10.20.10.0, km_local 198.51.100.2, km_remote 203.0.113.1 ISAKMP-ERROR: (1018):IPSec policy invalidated proposal with error 32 ISAKMP-ERROR: (1018):phase 2 SA policy not acceptable! (local 198.51.100.2 remote 203.0.113.1) ISAKMP-ERROR: (1018):deleting node 1949939457 error TRUE reason "QM rejected" ``` Read that literally, because it is telling you exactly what it wants. The peer proposed a pair of proxy identities (`local_proxy 10.30.10.0, remote_proxy 10.20.10.0`), IOS went looking through its crypto map database for a policy that permits that exact pair, and `did not find matching map`. Error 32 is the code for that. The debug even hands you the two subnets it could not match, which means you can go straight to the crypto ACL and compare. Crypto ACLs must be exact mirror images. Not overlapping. Not supersets. Mirror images. If your side permits `10.20.10.0/24 -> 10.30.10.0/24`, their side must permit `10.30.10.0/24 -> 10.20.10.0/24`. IKEv1 negotiates the proxy IDs as part of Quick Mode, and if they do not line up, the SA does not build. **Now the second symptom, and this is the one that gets misdiagnosed.** On the *initiator* whose ACL no longer matches, there is no debug output at all. No error. No ISAKMP SA. The tunnel does not fail, because the tunnel is never attempted. A crypto map is just an outbound classifier bolted to a physical interface. A packet leaves the router, hits the crypto map, and the map asks one question: does this match the crypto ACL? If yes, encrypt it (building the tunnel first if necessary). If no, forward it normally, in the clear. That is the entire logic. So if you fat-finger the crypto ACL such that your interesting traffic no longer matches, the router does not consider that an error. It does exactly what you told it to do: it forwards those packets in the clear, straight at the internet, where the first upstream router with no route to your private subnet drops them silently. From the operator's chair this looks exactly like a reachability problem. Pings fail, no SA exists, nothing in the logs, and people spend an afternoon checking routes and firewall rules. The tell is this: **"the tunnel never even tries to come up" is a crypto ACL symptom**, and you should treat it as one until proven otherwise. Our lab makes the point cleanly, because the ISP router in the middle genuinely cannot reach the private LANs: ``` ISP1#ping 10.30.10.1 repeat 3 ... Success rate is 0 percent (0/3) ISP1#ping 192.168.99.100 repeat 3 ... Success rate is 0 percent (0/3) ``` Those subnets are reachable only through the tunnel. Send them in the clear and they die at the first hop. This entire class of fault, incidentally, is impossible with a route-based VTI, because a VTI has no crypto ACL to get wrong. Its proxy IDs are always any/any. That is not a stylistic preference, it is a real argument for VTIs, and we make it in full in [crypto maps vs VTI](https://www.pinglabz.com/crypto-maps-vs-vti/). ### Fault D: Wrong Peer Address The simplest fault, and worth including precisely because its signature is so bare: ``` EDGE2#show crypto session Session status: DOWN Peer: 203.0.113.9 port 500 ``` No ISAKMP SA to the configured peer. No Phase 1, no Phase 2, no errors, nothing. Just `DOWN`, forever. The router is dutifully firing IKE packets at 203.0.113.9, and 203.0.113.9 is not answering, because 203.0.113.9 is not a VPN endpoint. It might not even exist. Its symptom (nothing happens) is easy to confuse with a routing problem or a firewall eating UDP 500\. Before you go chasing the transit path, read the peer address in `show crypto session` and compare it digit by digit to what the far end actually terminates on, and confirm the ISAKMP key is bound to that same address. ### Fault E: A Static Crypto Map Shadowing a Dynamic One (error 64) We found this one by accident while moving a router behind PAT, and it is the subtlest fault in the set. It will bite you on any hub that carries both static site-to-site peers and a dynamic crypto map for remote sites. The setup: EDGE2 moved behind a NAT device, so its source address changed. The hub, EDGE1, had a dynamic crypto map entry to catch it. But EDGE1 *also* still had a static crypto map entry whose ACL matched the same proxy pair, pointing at EDGE2's old public address. Here is what came out: ``` EDGE1#show logging IPSEC(validate_proposal_request): proposal part #1, local_proxy= 192.168.99.0/255.255.255.0/256/0, remote_proxy= 10.30.10.0/255.255.255.0/256/0, Crypto mapdb : proxy_match IPSEC(ipsec_process_proposal): peer address 198.51.100.6 not found ISAKMP-ERROR: (1011):IPSec policy invalidated proposal with error 64 ISAKMP-ERROR: (1011):phase 2 SA policy not acceptable! (local 203.0.113.1 remote 198.51.100.6) ISAKMP-ERROR: (1011):deleting node 2945058899 error TRUE reason "QM rejected" ``` `peer address 198.51.100.6 not found`. That is EDGE2's new post-NAT address, and EDGE1 says it cannot find it. But EDGE1 had a dynamic crypto map entry that exists precisely to accept peers whose address it does not know in advance. Why did that not catch it? Because of the order in which IOS evaluates the crypto map database. When a Quick Mode proposal arrives, IOS walks the crypto map entries and matches on the *proxy identities first*. The static entry's ACL matched the proposed proxy pair, so IOS selected that entry. It then checked the peer address in that entry, found it did not match the address the proposal actually came from, and rejected the proposal outright. It never fell through to the dynamic entry. A match on the proxy pair is a commitment, not a suggestion. The practical rule: on a hub carrying both static peers and a dynamic catch-all, make sure no static entry's crypto ACL overlaps traffic that is supposed to land on the dynamic entry. If it does, the static entry wins the match and then fails the peer check, and you get error 64 instead of a working tunnel. Delete or narrow the stale static entry. ## Part 3: The IPsec Error Code Map Cisco's Quick Mode rejection messages carry a numeric error code, and that code is a precise diagnosis if you know how to read it. We could not find this documented anywhere in one place, so we derived it from our own lab captures. Every number below was produced by a fault we broke on purpose and captured on a live IOS XE router. error 32 Proxy identities do not match Your crypto ACLs are not mirror images. Look for `map_db_find_best did not find matching map` right above it, and read the `local_proxy` / `remote_proxy` values it prints. error 64 Peer address not found No crypto map entry accepts a proposal from that source address. Classic cause: a static entry matched the proxy pair first and shadowed the dynamic entry that should have caught it. error 256 Transform set / proposal mismatch The peers cannot agree on how to protect the data. Compare the transform set on both sides (encryption, hash, mode) and the PFS group if you are using one. All three arrive with the same two companion lines, `phase 2 SA policy not acceptable!` and `reason "QM rejected"`, which is why they all look identical if you skim. The number is the whole message. Read the number. ## Part 4: The NAT Trap There is a specific failure that fits none of the above and fools almost everyone the first time. Phase 1 is up. There is no Quick Mode error in the logs. And yet nothing passes: ``` EDGE2#ping 192.168.99.100 source Loopback10 repeat 5 ..... Success rate is 0 percent (0/5) EDGE2#show crypto isakmp sa dst src state conn-id status 203.0.113.1 10.30.99.2 QM_IDLE 1008 ACTIVE EDGE2#show crypto ipsec sa | include #pkts #pkts encaps: 0, #pkts encrypt: 0, #pkts digest: 0 #pkts decaps: 0, #pkts decrypt: 0, #pkts verify: 0 ``` QM\_IDLE with zero encaps and zero decaps. The reason is beautiful in its simplicity. IKE runs over UDP 500, and UDP has port numbers, so a PAT device can happily translate it. Phase 1 sails through. But ESP is IP protocol 50\. It is not TCP, it is not UDP, and it has no port field at all. A PAT device has nothing to build a translation entry on, so the encrypted data has no way home. The fix is NAT Traversal, which wraps ESP inside UDP 4500 so PAT has a port to work with. When it is on, you can prove it from two places in `show crypto session detail` and `show crypto ipsec sa`: ``` EDGE2#show crypto session detail Session status: UP-ACTIVE Peer: 203.0.113.1 port 4500 fvrf: (none) ivrf: (none) Capabilities:N connid:1011 lifetime:00:57:40 EDGE2#show crypto ipsec sa transform: esp-256-aes esp-sha256-hmac , in use settings ={Tunnel UDP-Encaps, } ``` Two things to look for. `Capabilities:N` is the N-for-NAT flag, meaning NAT was detected during IKE. And `in use settings ={Tunnel UDP-Encaps, }` is the proof that the SA is actually UDP-encapsulating. If your peer is behind PAT and you do not see both, that is your fault right there. The full mechanism, including the NAT device's translation table, is in [IPsec NAT traversal (NAT-T)](https://www.pinglabz.com/ipsec-nat-traversal-nat-t/). ## Part 5: Two Gotchas Worth Knowing ### The First Packet Always Drops Watch this, from a genuinely working configuration: ``` EDGE1#ping 10.30.10.1 source Loopback10 repeat 5 ..... Success rate is 0 percent (0/5) ``` Zero percent. Now the very next attempt, with nothing changed: ``` EDGE1#ping 10.30.10.1 source Loopback10 repeat 10 !!!!!!!!!! Success rate is 100 percent (10/10), round-trip min/avg/max = 10/21/48 ms ``` Crypto map tunnels are built on demand. The first interesting packet is what triggers IKE, and while IKE runs (Main Mode, then Quick Mode, with a Diffie-Hellman exchange in there) there is no SA to encrypt with, so those packets are consumed and dropped. All five of them, in this case. Do not chase this. It is not a fault. If your first ping fails and your second succeeds, the tunnel is working correctly and you have just watched it negotiate. It is also why "it works if I ping twice" is such a common and confusing user report. ### IOS XE Will Not Let You Edit an In-Use Crypto ACL You have found the ACL problem, you are ready to fix it, and IOS stops you: ``` EDGE2(config)#ip access-list extended PLZ-CRYPTO-ACL %ACL PLZ-CRYPTO-ACL can not be modified/deleted, as it is used in crypto-map PLZ-CMAP %Please first remove the ACL from crypto map or remove the crypto map from the interface % Cannot modify this ACL. ``` This is a real guardrail on IOS XE, and it surprises people who learned crypto maps on older code. You cannot modify a crypto ACL while it is bound to a crypto map that is applied to an interface. As the error says, you have to break the binding first: either remove the `match address` from the crypto map, or pull the crypto map off the interface, make the edit, and put it back. Plan for that on a production box, because pulling the crypto map off the interface drops every tunnel it is carrying, not just the one you are fixing. ## Part 6: The Command Toolkit Seven commands. Know what each one is actually for, and you will never need to guess. show crypto isakmp sa Always first. Tells you whether Phase 1 is up and therefore which half of the protocol to investigate. QM\_IDLE = Phase 1 healthy. show crypto ipsec sa Always second. The encaps/decaps counters, the negotiated proxy IDs, the transform actually in use, and the SPIs. This is where you learn if packets move. show crypto session detail The best one-screen summary. Session status, peer and port (500 vs 4500), Capabilities flags, and per-flow packet counts. Reach for it when you want the whole picture fast. show crypto map Shows the peer, transform set, PFS group, and crypto ACL as the router understands them. Use it when you suspect the config on the box is not what you think it is. debug crypto isakmp error The high-signal debug. Gives you the QM rejection lines and the error code (32 / 64 / 256) without the firehose of full ISAKMP debugging. Start here, not with `debug crypto isakmp`. debug crypto ipsec Phase 2 detail. This is what prints the `proxy_match` / `map_db_find_best` block and the exact proxy subnets that failed to match. Indispensable for error 32 and error 64. clear crypto session Tears down the SAs so the next interesting packet renegotiates from scratch. Run it after every config change, otherwise you are staring at a stale SA and drawing conclusions from it. One workflow note that saves more time than any single command: after you change anything, `clear crypto session`, then generate traffic, then look. Half of all "my fix did not work" moments are actually "my fix worked but the old SA was still installed". ## Key Takeaways - **Run `show crypto isakmp sa` first, every time.** No SA means reachability, peer address, or IKE policy. Stuck in MM\_KEY\_EXCH means the pre-shared key. QM\_IDLE means Phase 1 is fine and you should stop looking at it. - **`#pkts encaps` and `#pkts decaps` are the two counters that matter.** Encaps climbing with decaps at zero is a return-path problem, not a tunnel problem. Both at zero means Phase 2 never built. - **A "malformed" IKE message means a bad pre-shared key.** The peer decrypted your Main Mode payload with the wrong key and got garbage. Nothing is actually malformed. - **Learn the error codes: 32 is proxy identities, 64 is peer address not found, 256 is transform set.** All three print the same `QM rejected` line, so the number is the entire diagnosis. - **Error 64 on a hub usually means a static crypto map entry shadowed the dynamic one.** IOS matches on the proxy pair first, then checks the peer, and rejects outright rather than falling through. - **If the tunnel never even tries to come up, suspect the crypto ACL.** Non-matching traffic is forwarded in the clear and dropped upstream, which looks exactly like a reachability problem and is not. - **Phase 1 up plus zero encaps behind a PAT device is NAT-T.** Look for `Capabilities:N` and `in use settings ={Tunnel UDP-Encaps, }`. - **The first packet always drops.** Crypto map tunnels build on demand. A failed first ping followed by a successful second one is the tunnel working, not failing. - **Clear the session after every change.** Stale SAs make good fixes look like bad ones. Almost everything in this article, other than the transform set and PSK faults, is a symptom of policy-based VPN mechanics: crypto ACLs, proxy IDs, and crypto map match order. A route-based VTI has none of that, which is why error 32 simply cannot happen on one. If you are building something new, read [crypto maps vs VTI](https://www.pinglabz.com/crypto-maps-vs-vti/) before you commit to a design, and start from [the site-to-site crypto map guide](https://www.pinglabz.com/site-to-site-ipsec-crypto-maps/) if you have inherited one. The full set of IPsec deep-dives, from IKE negotiation to NAT traversal to tunnel protection, is indexed on the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/). ### NAT Traversal (NAT-T): How IPsec Survives a NAT Device URL: https://www.pinglabz.com/ipsec-nat-traversal-nat-t/ Last updated: 2026-07-13T07:29:53.000Z Every IPsec engineer eventually hits the same wall. The tunnel worked perfectly in the lab, you shipped the branch router, and now the site is behind a broadband modem doing PAT. Phase 1 comes up. Phase 2 does not. Nothing crosses. This is the single most recognisable failure in the whole of [IPsec VPN](https://www.pinglabz.com/ipsec-vpn/), and the reason for it is embarrassingly simple: ESP has no port numbers, and PAT is a machine that rewrites port numbers. NAT Traversal (NAT-T) is the fix. This article shows you exactly why raw IPsec dies behind a NAT, what NAT-T actually changes on the wire, and what the failure and the fix look like in real CLI output from a live CML lab, including the capture almost nobody publishes: the NAT device's own translation table. ## Why ESP and PAT Cannot Coexist Start with the protocol numbers, because the whole story is in them. **IKE (Phase 1 and 2)** UDP port 500\. It is UDP. It has ports. PAT loves it. **ESP (the data)** IP protocol 50\. Not TCP, not UDP. No source port, no destination port. Nothing to rewrite. **AH** IP protocol 51\. Authenticates the outer IP header, so any NAT rewrite breaks the ICV. Fundamentally NAT-incompatible. **ESP over NAT-T** UDP port 4500\. ESP wrapped in a UDP header, so PAT has ports to work with again. PAT (what Cisco calls NAT overload) multiplexes many inside hosts onto one public address by giving each flow a unique source port. That is its entire mechanism. Hand it an ESP packet and there is no port field to allocate, no port field to record, and no port field to reverse on the way back. Some NAT implementations bolt on an ESP passthrough hack that tracks the SPI instead, and it works for exactly one inside host, until a second one shows up. Cisco IOS PAT does not pretend. It simply cannot track the flow. AH is worth a sentence of its own because people ask. AH authenticates the immutable fields of the outer IP header, and the source address is one of them. NAT rewrites that address by definition. The receiver recomputes the ICV, gets a different answer, and drops the packet. There is no NAT-T for AH and there never will be. NAT traversal is an ESP-only story, which is one more reason nobody deploys AH. ## The Lab Two [site-to-site IPsec](https://www.pinglabz.com/site-to-site-ipsec-crypto-maps/) endpoints, EDGE1 at the head end on 203.0.113.1 and EDGE2 at the branch. For this test EDGE2 was moved behind NAT1, an IOS router doing plain PAT on its outside interface: ``` ip nat inside source list PLZ-PAT interface Ethernet0/0 overload ``` EDGE2's real address is now a private 10.30.99.2\. To the rest of the world, including EDGE1, it appears as 198.51.100.6\. EDGE2 itself has no idea any of this is happening, which turns out to matter a great deal. ## The Failure: Phase 1 Up, Phase 2 Dead Disable NAT-T on the branch router (IOS enables it by default, so you have to go out of your way to break it): ``` EDGE2(config)#no crypto ipsec nat-transparency udp-encapsulation ``` Now ping across the tunnel: ``` EDGE2#ping 192.168.99.100 source Loopback10 repeat 5 ..... Success rate is 0 percent (0/5) ``` Total failure. Now look at Phase 1: ``` EDGE2#show crypto isakmp sa dst src state conn-id status 203.0.113.1 10.30.99.2 QM_IDLE 1008 ACTIVE ``` QM\_IDLE. That is the healthy, settled state of an IKEv1 Phase 1 SA. The tunnel's control plane is completely up. Authentication succeeded, DH completed, the ISAKMP SA is ACTIVE. If you stopped troubleshooting here you would swear the VPN was fine. Now look at Phase 2: ``` EDGE2#show crypto ipsec sa | include #pkts #pkts encaps: 0, #pkts encrypt: 0, #pkts digest: 0 #pkts decaps: 0, #pkts decrypt: 0, #pkts verify: 0 ``` Zero in every column. Not a single packet has been encapsulated or decapsulated. The data plane is stone dead. ### Why the split happens exactly here This is the part worth internalising, because once you see it you will diagnose this failure from across the room. IKE negotiation rides on UDP port 500\. UDP has ports. NAT1's PAT box takes EDGE2's 10.30.99.2:500 packets, allocates a translation, rewrites the source to 198.51.100.6 with some high port, and forwards them. Replies come back, PAT reverses the translation, EDGE2 receives them. Phase 1 completes normally. Phase 2 (Quick Mode) also runs inside IKE on UDP 500, so the IPsec SAs get negotiated too. Then the router tries to send actual user data, and that data is ESP. IP protocol 50\. It arrives at NAT1 with no ports at all. PAT has nothing to multiplex on, no way to build a translation entry, and no way to demultiplex the return traffic. The packets die at the NAT device. So you get a control plane that works and a data plane that does not. **Phase 1 up, Phase 2 at zero, with a NAT anywhere in the path, is the NAT-T signature.** Learn those two show commands together and you will never spend an afternoon on this again. ## The Fix: Wrap ESP in UDP 4500 Re-enable NAT transparency (or just leave the default alone) and the tunnel comes straight up. Here is what changed. ``` EDGE2#show crypto session detail Session status: UP-ACTIVE Peer: 203.0.113.1 port 4500 fvrf: (none) ivrf: (none) Capabilities:N connid:1011 lifetime:00:57:40 IPSEC FLOW: permit ip 10.30.10.0/255.255.255.0 192.168.99.0/255.255.255.0 Inbound: #pkts dec'ed 9 drop 0 life (KB/Sec) 4607998/3597 Outbound: #pkts enc'ed 9 drop 0 life (KB/Sec) 4607999/3597 ``` Two things to read here. The peer is on **port 4500**, not 500\. And **Capabilities:N**, where N means NAT-traversal was detected. Notice what you did not do: you never configured port 4500 anywhere. IOS discovered the NAT by itself. During Main Mode, both peers exchange NAT-D (NAT Discovery) payloads containing hashes of the source and destination IP and port as each side believes them to be. If the hash a peer receives does not match what it computes from the packet it actually got, an address was rewritten in transit, so there is a NAT in the path. Both ends then agree to float the negotiation from UDP 500 to UDP 4500 and to UDP-encapsulate all subsequent ESP. The whole thing is automatic and it happens during Phase 1. Now the IPsec SA: ``` EDGE2#show crypto ipsec sa local crypto endpt.: 10.30.99.2, remote crypto endpt.: 203.0.113.1 #pkts encaps: 15, #pkts encrypt: 15, #pkts digest: 15 transform: esp-256-aes esp-sha256-hmac , in use settings ={Tunnel UDP-Encaps, } ``` Two lines carry the whole lesson. `local crypto endpt.: 10.30.99.2` is the *private* address. EDGE2 has no clue it is being translated. It builds and signs its packets from 10.30.99.2 and hands them to the network. The NAT rewrite happens downstream, outside the router's awareness, which is precisely why the outer header must not be authenticated (and precisely why AH cannot play). `in use settings ={Tunnel UDP-Encaps, }` is the clincher. In a normal tunnel you see `{Tunnel, }`. Here you see `Tunnel UDP-Encaps`: tunnel mode, with the ESP payload encapsulated in UDP. That string is the definitive proof that NAT-T is engaged. If you only remember one grep from this article, make it that one. Traffic flows: 15 packets encapsulated, 9 decapsulated, pings at 100 percent. ## The NAT Device's View This is the capture that closes the argument. Log into NAT1, the PAT router in the middle, and ask it what it is actually tracking: ``` NAT1#show ip nat translations Pro Inside global Inside local Outside local Outside global udp 198.51.100.6:4790 10.30.99.2:500 203.0.113.1:500 203.0.113.1:500 udp 198.51.100.6:4791 10.30.99.2:4500 203.0.113.1:4500 203.0.113.1:4500 NAT1#show ip nat statistics Total active translations: 2 (0 static, 2 dynamic; 2 extended) Hits: 79 Misses: 0 ``` Two entries. Both say `udp`. One for IKE on 500, one for ESP-in-UDP on 4500\. As far as NAT1 is concerned there is no IPsec here at all, just two ordinary UDP conversations between an inside host and an outside host. It allocates ports (4790 and 4791), it builds extended translations, it counts 79 hits and zero misses. This is PAT doing the most boring thing it knows how to do, and that is the entire point. Compare that with what NAT1 could have done with raw ESP: nothing. No line in the table, no port to allocate, no state to keep. The genius of NAT-T is not that it is clever. It is that it makes IPsec look mundane enough for a PAT box to handle. ## The Costs and the Fine Print ### 8 bytes of overhead, and your MTU UDP encapsulation adds a UDP header: 8 bytes on every single ESP packet, on top of the ESP overhead you already pay. That is 8 bytes less room for user data before the packet needs fragmenting. If you were already close to the edge (GRE over IPsec over a 1500-byte path is a classic squeeze), NAT-T can be the thing that tips you into fragmentation, and fragmented ESP is a well-known source of intermittent, size-dependent, maddening failures. Fix it at the source: set `ip tcp adjust-mss` on the LAN-facing interface and lower the tunnel MTU rather than letting the path fragment for you. ### NAT keepalives PAT translations age out. A NAT device holds a UDP entry for a while after the last packet (often 30 to 300 seconds) and then reclaims the port. If your VPN goes quiet for longer than that, the entry vanishes, and the far end's next inbound packet arrives at a NAT device that no longer knows where to send it. The tunnel appears up on both routers and traffic silently dies until the branch happens to send something first. IPsec NAT keepalives exist purely to prevent this: a tiny periodic packet on UDP 4500 whose only job is to keep the translation warm. On IOS that is `crypto isakmp nat keepalive 20`. Set it. It is cheap, and the failure it prevents is nasty and intermittent. ### Firewall rules: 500 AND 4500 The classic operational mistake. Someone opens UDP 500 on the perimeter firewall because "IPsec uses 500", the tunnel negotiates beautifully, and then no data flows because ESP-in-UDP on 4500 is being dropped. It looks exactly like the failure at the top of this article, and the symptom is identical: Phase 1 up, Phase 2 at zero. Permit **UDP 500 and UDP 4500 in both directions**. If any peer might sit behind a NAT, you need both, always. (And if no NAT is involved, you also need IP protocol 50 permitted for raw ESP.) ### Aggressive Mode, IKEv2, and the rest NAT-D is exchanged in messages 3 and 4 of Main Mode, which is why the float to 4500 happens mid-Phase-1\. IKEv2 does the same job with NAT\_DETECTION\_SOURCE\_IP and NAT\_DETECTION\_DESTINATION\_IP notify payloads in the initial exchange, and the resulting behaviour is the same: float to 4500, UDP-encapsulate ESP. The mechanism differs, the concept does not. ## Diagnosing It in the Field show crypto isakmp sa **Look for:** QM\_IDLE / ACTIVE. **Means:** Phase 1 is fine. IKE crossed the NAT on UDP 500\. Do not stop here. show crypto ipsec sa | include #pkts **Look for:** encaps/decaps at 0. **Means:** Phase 2 negotiated but nothing crosses. With a NAT in path, suspect NAT-T immediately. show crypto session detail **Look for:** `port 4500` and `Capabilities:N`. **Means:** NAT detected and NAT-T active. Port 500 plus a NAT in path means it is not. show crypto ipsec sa (full) **Look for:** `in use settings ={Tunnel UDP-Encaps, }`. **Means:** ESP is inside UDP. This is the definitive proof. show ip nat translations (on the NAT box) **Look for:** two `udp` entries, :500 and :4500. **Means:** PAT is tracking both flows. No 4500 entry means ESP never got wrapped. local crypto endpt. **Look for:** a private address (10.30.99.2 here). **Means:** the router is behind a NAT and does not know it. Normal, and exactly why NAT-T exists. One more warning from the same lab. When we moved EDGE2 behind PAT, the head end kept failing Phase 2 with `error 64` and `peer address 198.51.100.6 not found`, because EDGE1 still had a static crypto map entry pointing at EDGE2's old public address. IOS matched the static entry on proxy ID, found the peer did not match, and rejected the proposal instead of falling through to the dynamic map. If you migrate a peer behind a NAT, clean up the stale static entry. There is more on reading those error codes in our guide to [troubleshooting IPsec VPN tunnels](https://www.pinglabz.com/troubleshooting-ipsec-vpn/). ## Key Takeaways - **ESP is IP protocol 50 and has no port numbers.** PAT works by rewriting ports. It therefore has nothing to multiplex on and cannot track an ESP flow. That single fact is the whole problem. - **The signature is Phase 1 up, Phase 2 at zero.** `show crypto isakmp sa` reads QM\_IDLE while `show crypto ipsec sa` shows `#pkts encaps: 0, #pkts decaps: 0`. IKE on UDP 500 survives PAT; raw ESP does not. - **NAT-T wraps ESP in UDP port 4500**, which PAT translates like any other UDP flow. On the NAT box you see two ordinary udp translations, :500 and :4500, and that is the entire trick. - **You do not configure the port.** IOS auto-detects the NAT during IKE using NAT-D payloads and floats to 4500 on its own. Confirm it with `Peer: ... port 4500` and `Capabilities:N`. - **The proof string is `in use settings ={Tunnel UDP-Encaps, }`.** Also expect `local crypto endpt.` to show the private address, because the router genuinely does not know it is being translated. - **Mind the 8 bytes.** The extra UDP header eats MTU. Adjust TCP MSS and tunnel MTU rather than letting the path fragment. - **Enable NAT keepalives** so idle PAT entries do not age out and blackhole the tunnel. - **Permit UDP 500 AND UDP 4500.** Opening only 500 produces the identical Phase-1-up, Phase-2-dead symptom and wastes hours. - **AH is out.** It authenticates the outer IP header that NAT rewrites, so it can never traverse a NAT. NAT-T is an ESP-only story. NAT-T is one of the pieces that makes real-world IPsec deployable at all, alongside crypto maps, tunnel interfaces and the failure modes that come with them. For the full picture, start at the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/), build the baseline tunnel with [site-to-site IPsec crypto maps](https://www.pinglabz.com/site-to-site-ipsec-crypto-maps/), and keep [the IPsec troubleshooting guide](https://www.pinglabz.com/troubleshooting-ipsec-vpn/) open next to the console. ### GRE over IPsec: Why Tunnels and Encryption Belong Together URL: https://www.pinglabz.com/gre-over-ipsec/ Last updated: 2026-07-13T07:29:51.000Z Every site-to-site VPN design runs into the same wall. You build a crypto map, the tunnel comes up, traffic between the two LANs is encrypted, and then someone asks the obvious question: "how do the sites learn each other's routes?" With a plain crypto map the answer is "they don't, you type them in by hand." A crypto ACL matches unicast flows. OSPF hellos go to 224.0.0.5\. Those two facts never meet. GRE solves exactly that problem and nothing else. It gives you a real routed interface that carries multicast, so a routing protocol runs across it. What it does not do is protect a single byte. Put the two together and you get what you actually wanted: a routed, encrypted WAN. This article is part of the [IPsec VPN cluster](https://www.pinglabz.com/ipsec-vpn/), and it proves both halves with a packet capture taken at the same point on the same link, before and after. ## The Two Halves of the Problem GRE and IPsec are not competing technologies and they are not redundant. Each one does something the other structurally cannot. What GRE gives you **Real interface:** Tunnel0 with an IP address, a line protocol, and an MTU **Multicast:** yes, so OSPF and EIGRP run across it **Any protocol:** IPv4, IPv6, and legacy payloads all ride inside **Routing table entry:** the tunnel is a next hop like any other **Confidentiality:** none at all **Overhead:** 24 bytes (20 byte outer IP + 4 byte GRE) What IPsec gives you **Confidentiality:** ESP encrypts the whole inner packet **Integrity and authentication:** HMAC on every packet, peers proven by PSK or certs **Anti-replay:** sequence numbers inside the SA **Multicast:** no, a crypto ACL matches unicast flows **Interface:** none, a crypto map is a filter bolted to a physical port **Routing protocols:** cannot carry them Read those two cards side by side and the design writes itself. GRE builds the pipe, IPsec protects it. Every gap in one column is filled by the other. ## Half One: GRE Alone Carries Routing and Protects Nothing The lab is a pair of cat8000v routers (EDGE1 and EDGE2) with an ISP router in the middle. EDGE1 sits on 203.0.113.1, EDGE2 on 198.51.100.2, and a GRE tunnel runs between those two public addresses with OSPF enabled on Tunnel0\. First, what GRE claims to be: ``` EDGE1#show interface Tunnel0 Tunnel0 is up, line protocol is up Tunnel source 203.0.113.1 (GigabitEthernet2), destination 198.51.100.2 Tunnel protocol/transport GRE/IP Tunnel transport MTU 1476 bytes ``` `Tunnel protocol/transport GRE/IP`, transport MTU 1476 (1500 on the wire, minus a 20 byte outer IP header, minus the 4 byte GRE header). That is a real interface, and a real interface can do the one thing a crypto map cannot: ``` EDGE1#show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 10.255.0.2 0 FULL/ - 00:00:35 10.0.0.2 Tunnel0 ``` FULL adjacency, across the public internet, through a tunnel. Add a subnet at either end and it appears at the other with no human intervention. That is the whole reason GRE exists, and why the [GRE cluster](https://www.pinglabz.com/gre/) is worth understanding on its own terms. Now the other half of the ledger: ``` EDGE1#show crypto ipsec sa (empty - NOTHING is encrypted) ``` Nothing. No SAs, no SPIs, no encaps counters. GRE has a header that looks technical and a name that sounds like infrastructure. Neither of those is encryption. ### The Packet Capture That Ends the Argument Here is a capture taken on the EDGE1 to ISP1 link, exactly where a hostile transit provider or anyone with a tap would sit. Plain GRE, carrying a ping from a real Debian host (192.168.99.100) to a loopback behind the far router (10.30.10.1): ``` No. Time Source Destination Proto Info 1 0.000000 10.0.0.2 224.0.0.5 OSPF LS Acknowledge 2 1.461513 192.168.99.100 10.30.10.1 ICMP Echo (ping) request id=0x004e, seq=1/256, ttl=63 3 1.462810 10.30.10.1 192.168.99.100 ICMP Echo (ping) reply id=0x004e, seq=1/256, ttl=255 4 2.463253 192.168.99.100 10.30.10.1 ICMP Echo (ping) request id=0x004e, seq=2/512, ttl=63 5 2.464559 10.30.10.1 192.168.99.100 ICMP Echo (ping) reply id=0x004e, seq=2/512, ttl=255 8 4.100454 10.0.0.1 224.0.0.5 OSPF Hello Packet 15 7.206880 10.0.0.2 224.0.0.5 OSPF Hello Packet ``` Look at what the capture tool did without being asked. It dissected straight through the GRE header and printed the *inner* conversation. The ISP can read: - **Your private host addressing.** 192.168.99.100 talking to 10.30.10.1\. Your internal IP plan is on the wire. - **Your application traffic.** This happens to be ICMP. If it were HTTP or SNMP or a database session, the payload would be sitting there in full. - **Your routing protocol.** OSPF hellos and LS acknowledgements from 10.0.0.1 and 10.0.0.2 to 224.0.0.5\. An attacker on this link now knows your tunnel subnet, your router IDs, your timers, and with a little work could inject LSAs. GRE is encapsulation. Encapsulation is not encryption. If you take one thing from this article, make it that packet list. ## Half Two: Add Tunnel Protection and Watch the Link Go Dark The fix is four lines. Build an IPsec profile (the same transform set you would use with a crypto map, minus the crypto ACL) and bolt it to the tunnel interface: ``` crypto ipsec profile PLZ-IPSEC-PROF set transform-set PLZ-TS-AES256 set pfs group14 interface Tunnel0 tunnel protection ipsec profile PLZ-IPSEC-PROF ``` Note what is *not* there. There is no `ip access-list extended PLZ-CRYPTO-ACL`. There is no `match address`. There is no `set peer`. The tunnel already knows its peer (that is the tunnel destination) and it already knows what to encrypt (everything that enters the tunnel interface). The interface has replaced the ACL as the traffic selector, which is why this is dramatically simpler than the [crypto map approach](https://www.pinglabz.com/site-to-site-ipsec-crypto-maps/) and dramatically harder to get wrong. ### The Same Capture Point, The Same Ping, After ``` No. Time Source Destination Proto Info 3 5.153552 203.0.113.1 198.51.100.2 ISAKMP Identity Protection (Main Mode) 6 5.617148 198.51.100.2 203.0.113.1 ISAKMP Identity Protection (Main Mode) 9 5.861495 203.0.113.1 198.51.100.2 ISAKMP Quick Mode 11 6.425161 203.0.113.1 198.51.100.2 ISAKMP Quick Mode 12 7.026032 203.0.113.1 198.51.100.2 ESP ESP (SPI=0x6d4428aa) 13 8.939238 198.51.100.2 203.0.113.1 ESP ESP (SPI=0x94cd6141) 14 12.002692 203.0.113.1 198.51.100.2 ESP ESP (SPI=0x6d4428aa) 15 12.003690 198.51.100.2 203.0.113.1 ESP ESP (SPI=0x94cd6141) 16 13.004286 203.0.113.1 198.51.100.2 ESP ESP (SPI=0x6d4428aa) 17 13.005508 198.51.100.2 203.0.113.1 ESP ESP (SPI=0x94cd6141) ``` Same link. Same ping. Same OSPF adjacency still running underneath. And the ISP now sees precisely two IP addresses (203.0.113.1 and 198.51.100.2, both of which are public and both of which it was routing anyway) plus an opaque stream of ESP. No 192.168.99.100\. No 10.30.10.1\. No 224.0.0.5\. The routing protocol is still converging, it is just doing it inside the ciphertext. ### Reading the Main Mode, Quick Mode, ESP Sequence That capture is a complete IKEv1 lifecycle in ten packets, and it is worth walking slowly because this is the exact sequence you will be staring at in a troubleshooting session. - **Packets 3 and 6, Identity Protection (Main Mode).** This is IKE Phase 1, six messages in three exchanges. The peers negotiate the ISAKMP policy (encryption, hash, DH group, lifetime), perform the Diffie-Hellman key exchange, and then authenticate each other with the pre-shared key. The dissector calls it "Identity Protection" because Main Mode's whole selling point over Aggressive Mode is that the identity exchange happens after the DH shared secret exists, so the identities are already encrypted. The output of Phase 1 is a bidirectional ISAKMP SA: a secure management channel, carrying no user data. - **Packets 9 and 11, Quick Mode.** This is IKE Phase 2, and it runs entirely *inside* the Phase 1 SA (which is why the capture can name it but not read it). Here the peers negotiate the IPsec transform set, agree the proxy identities, and with PFS enabled they run a fresh Diffie-Hellman exchange so the Phase 2 keys are not derived from the Phase 1 material. The output is a pair of unidirectional IPsec SAs, one per direction. - **Packets 12 onward, ESP.** The two SPIs (0x6d4428aa outbound from EDGE1, 0x94cd6141 inbound from EDGE2) are the SA identifiers. Every subsequent data packet is ESP: encrypted GRE, encapsulating IP, encapsulating your ping. The capture tool can see the SPI and the sequence number and nothing else, because that is all that is in the clear. This sequence also explains why the first ping across a freshly configured tunnel often fails: the packets that trigger IKE are consumed building the SA. Phase 1 and Phase 2 take real time, and IOS drops what it cannot yet encrypt. Ping twice before declaring a tunnel broken. ## Why Tunnel Protection Beat the Old Crypto Map Approach The old way to do this was to apply a crypto map to the tunnel interface *and* to the physical interface, with a crypto ACL that read something like `permit gre host 203.0.113.1 host 198.51.100.2`. It worked, mostly, and it was miserable: - Two places to apply the same crypto map, and forgetting the physical interface produced a tunnel that came up and dropped everything. - The crypto ACL matched the GRE flow, so the encryption decision was effectively made twice. Debugging a mismatch meant reasoning about both. - IOS XE refuses to edit a crypto ACL that is in use, so every change meant pulling the map off the interface first. `tunnel protection ipsec profile` collapses all of that into one statement on one interface. The tunnel endpoints *are* the crypto peers. The tunnel interface *is* the traffic selector. There is nothing to mismatch, which is the biggest single source of IPsec outages eliminated by configuration syntax. It is also the only approach that scales to mGRE, which is what [DMVPN](https://www.pinglabz.com/dmvpn/) is built on: a hub cannot hold a static crypto map per spoke, but it can hold one IPsec profile that every dynamic tunnel inherits. ## MTU: The 1476 Number and the 1438 Number GRE reported a transport MTU of 1476: 1500 minus the 20 byte outer IP header minus the 4 byte GRE header, before any encryption at all. Now add ESP in tunnel mode with AES-256 and SHA-256\. That means another outer IP header, the ESP header with its SPI and sequence number, an initialisation vector, block-cipher padding, the ESP trailer, and the authentication data. With this exact transform set the crypto SA reports what is left: ``` local crypto endpt.: 203.0.113.1, remote crypto endpt.: 198.51.100.2 plaintext mtu 1438, path mtu 1500, ip mtu 1500, ip mtu idb GigabitEthernet2 ``` 1438 against a path MTU of 1500\. That is roughly 62 bytes of IPsec overhead, and it lands on top of the GRE overhead you already paid. The practical consequence: a 1500 byte packet entering Tunnel0 will not fit, and something has to give. You have three options, in order of how much you will regret them. 1. **Lower the tunnel IP MTU.** `ip mtu 1400` on Tunnel0 makes the router fragment before encryption rather than after, which is the cheaper of the two fragmentation points. 2. **Clamp TCP.** `ip tcp adjust-mss 1360` on the tunnel rewrites the MSS in the TCP handshake so hosts never generate an oversized segment in the first place. This is the single highest-value line of config in most GRE over IPsec deployments and it fixes the classic "ping works, big file transfers hang" complaint. 3. **Rely on path MTU discovery.** Which works right up until a firewall somewhere in the middle drops ICMP unreachables, and then it fails silently and intermittently. Do not rely on it. The pings in our captures were 100 bytes and never came close to the limit. MTU problems do not show up in your smoke test. They show up three weeks later in a ticket about one application. ## Transport Mode vs Tunnel Mode for GRE Over IPsec The transform set in this lab uses `mode tunnel`, the default and what the capture shows. ESP in tunnel mode adds a fresh outer IP header, so the packet on the wire is: outer IP, ESP, inner IP, GRE, original IP, payload. You can also run `mode transport`. Because the GRE tunnel's source and destination are already the same addresses as the crypto endpoints, that tunnel-mode outer header is a byte-for-byte duplicate of the header GRE already built. Transport mode skips it and reuses the GRE outer header, saving 20 bytes per packet. On a WAN carrying small packets (voice, for example) that is a real saving. The catch is NAT. Transport mode leaves the original IP header intact and integrity-protected, so a NAT device that rewrites addresses in the path will break the authentication check. Tunnel mode, combined with NAT-T's UDP 4500 encapsulation, survives it. That is why tunnel mode stays the safe default, and why transport mode is an optimisation you apply only when both endpoints are public and nothing between them is translating. ## Where This Goes Next An obvious question follows from all of this: if the GRE header is only there to give us an interface, and IPsec can now attach directly to an interface, do we still need GRE at all? For a straightforward IPv4 site-to-site link the answer is no, and that is the [static VTI](https://www.pinglabz.com/static-vti-svti-ipsec/), which drops the GRE header entirely and reports `Tunnel protocol/transport IPSEC/IP` with a transport MTU of 1438 instead of 1476\. Same routing protocol support, four fewer bytes, same profile. [Crypto maps vs VTI](https://www.pinglabz.com/crypto-maps-vs-vti/) puts the older design and the newer one head to head. And when you need this for dozens of spokes rather than one peer, GRE stops being point-to-point and becomes mGRE, which is where [DMVPN](https://www.pinglabz.com/dmvpn/) takes over. The IPsec profile you built here is the same object DMVPN uses, so nothing you learned in this article gets thrown away. ## Key Takeaways - **GRE carries routing protocols and encrypts nothing.** The capture on the transit link showed inner host IPs (192.168.99.100 to 10.30.10.1) and OSPF hellos to 224.0.0.5 in the clear. Encapsulation is not encryption. - **A crypto map encrypts and cannot carry routing.** A crypto ACL matches unicast flows. OSPF's multicast hellos never match one, so a plain crypto map VPN needs static routes forever. - **Together they are a routed, encrypted WAN.** After `tunnel protection ipsec profile`, the same capture point showed ISAKMP Main Mode, then Quick Mode, then nothing but ESP between the two public endpoints. The OSPF adjacency never dropped, it just became invisible. - **Tunnel protection is far simpler than a crypto map.** No crypto ACL, no `set peer`, no double application to tunnel and physical. The tunnel interface is the traffic selector. - **Budget for MTU.** GRE gives you 1476, and this AES-256 / SHA-256 transform set leaves a plaintext MTU of 1438\. Set `ip tcp adjust-mss` on the tunnel before your users find the problem for you. - **Tunnel mode is the safe default.** Transport mode saves 20 bytes by reusing the GRE outer header, but any NAT in the path will break it. - **The first ping often fails.** IKE is built on demand and the triggering packets are consumed doing it. Ping twice. For the rest of the series, including IKEv1 versus IKEv2, NAT traversal, and a systematic approach to debugging a tunnel that will not come up, start at the [IPsec VPN pillar](https://www.pinglabz.com/ipsec-vpn/). ### IPsec Explained: IKE Phase 1, Phase 2, and What Actually Gets Encrypted URL: https://www.pinglabz.com/ipsec-explained-ike-phases/ Last updated: 2026-07-13T07:29:50.000Z Almost everything written about IPsec starts with a config snippet. That is backwards. If you do not know what an SA is, why IKE runs in two phases, or which bytes of your packet actually get encrypted, then a crypto map is just a spell you cast and hope. This article is the conceptual foundation for the whole [IPsec VPN cluster](https://www.pinglabz.com/ipsec-vpn/): what IPsec is, how the two-phase IKE model works, what a Security Association really is, and precisely what an eavesdropper on the transit link can and cannot see. Everything below is grounded in a real CML lab: two Cisco IOS XE edge routers (cat8000v 17.18.02) either side of a service provider router, with real packet captures taken on the transit link. ## IPsec Is a Framework, Not a Protocol The first thing to unlearn: IPsec is not one protocol. It is a set of them, working at Layer 3, that together give you three properties for IP packets. Confidentiality Nobody on the path can read the payload. Delivered by the ESP encryption algorithm (AES-256 in this lab). Integrity and authentication Nobody can alter or forge a packet without detection. Delivered by an HMAC over the packet (SHA-256 here). Anti-replay A captured packet cannot be re-injected later. Delivered by a sequence number inside the ESP header. Two protocols deliver those properties on the wire (AH and ESP), and one negotiates the keys that make them possible (IKE). The split between "the thing that protects packets" and "the thing that agrees on how to protect packets" is the single most useful mental model in IPsec. ## AH vs ESP: Only One of These Matters in Practice AH (Authentication Header, IP protocol 51) and ESP (Encapsulating Security Payload, IP protocol 50) are the two protection protocols. In modern designs you will use ESP, essentially always. Here is why. AH (IP protocol 51) Encrypts payload: **No** Authenticates: **Yes, including the outer IP header** Survives NAT: **No (NAT rewrites the header it signed)** Real-world use: vanishingly rare ESP (IP protocol 50) Encrypts payload: **Yes** Authenticates: **Yes, but not the outer IP header** Survives NAT: **Yes, with NAT-T (UDP 4500)** Real-world use: everything you will ever build The killer difference is the NAT row. AH signs the outer IP header, so any NAT device on the path invalidates the signature and the packet is discarded. ESP deliberately leaves the outer header out of its integrity check, which is exactly what lets it cross the internet. Every capture in this article is ESP. ## Tunnel Mode vs Transport Mode: Which Bytes Get Encrypted This is where most people go vague, so let us be precise. ESP encrypts everything it encapsulates and nothing else. The difference between the two modes is what gets encapsulated. **Transport mode** keeps the original IP header and protects only the payload behind it: ``` [ orig IP hdr ][ ESP hdr ][ TCP/UDP + data (encrypted) ][ ESP trailer ][ ESP auth ] ^------------ encrypted -------------^ ``` The source and destination host IPs stay visible on the wire. Transport mode is for host-to-host protection, and inside IOS XE it is what you use under a GRE tunnel or an SVTI (because there is already an outer delivery header, so a second one would be a waste of 20 bytes). **Tunnel mode** encrypts the entire original packet, header included, and builds a brand new outer IP header addressed to the crypto endpoints: ``` [ new IP hdr ][ ESP hdr ][ orig IP hdr + TCP/UDP + data (encrypted) ][ ESP trailer ][ ESP auth ] ^------------------ encrypted ------------------^ ``` The original source and destination addresses are now ciphertext. On the wire, the packet appears to be a conversation between the two VPN gateways and nothing else. Our lab transform set says exactly this: ``` crypto ipsec transform-set PLZ-TS-AES256 esp-aes 256 esp-sha256-hmac mode tunnel ``` And IOS XE confirms it on the live SA with `in use settings ={Tunnel, }`. That single word is the difference between the ISP seeing your internal addressing and seeing nothing at all. We prove it with a capture further down. ## Why IKE Runs in Two Phases Before any packet can be protected, the two peers must agree on keys. That is IKE's job (Internet Key Exchange, UDP 500). The awkward part is the chicken-and-egg problem: negotiating keys securely requires a secure channel, and you do not have one yet. IKE solves it by doing the job twice, for two different purposes: Phase 1 (ISAKMP / IKE SA) Purpose: build a secure management channel between the two peers Protects: the IKE negotiation itself, nothing user-facing Exchange: Main Mode (6 messages) or Aggressive Mode (3) Configured by: `crypto isakmp policy` \+ the pre-shared key Result: one bidirectional ISAKMP SA Verify with: `show crypto isakmp sa` Phase 2 (IPsec SA) Purpose: negotiate the keys that actually encrypt user data Protects: the traffic your users care about Exchange: Quick Mode (3 messages), inside the Phase 1 channel Configured by: transform set + crypto ACL (or IPsec profile) Result: a pair of unidirectional IPsec SAs Verify with: `show crypto ipsec sa` The payoff is efficiency. Phase 1 is expensive (Diffie-Hellman is heavy maths) but you do it once and keep it for an hour. Phase 2 is cheap and runs inside that already-encrypted channel, so you can rekey data keys often, or add a second protected flow, without redoing the expensive part. It also means a Phase 2 failure looks nothing like a Phase 1 failure, which is the single most useful fact when [troubleshooting an IPsec VPN](https://www.pinglabz.com/troubleshooting-ipsec-vpn/). ## Watching It Happen: The Real Packet List Here is a capture taken on the transit link between our edge router (203.0.113.1) and the remote peer (198.51.100.2), the moment protection was enabled and a ping was sent. This is the two-phase model, on the wire, in order: ``` No. Time Source Destination Proto Info 3 5.153552 203.0.113.1 198.51.100.2 ISAKMP Identity Protection (Main Mode) 6 5.617148 198.51.100.2 203.0.113.1 ISAKMP Identity Protection (Main Mode) 9 5.861495 203.0.113.1 198.51.100.2 ISAKMP Quick Mode 11 6.425161 203.0.113.1 198.51.100.2 ISAKMP Quick Mode 12 7.026032 203.0.113.1 198.51.100.2 ESP ESP (SPI=0x6d4428aa) 13 8.939238 198.51.100.2 203.0.113.1 ESP ESP (SPI=0x94cd6141) 14 12.002692 203.0.113.1 198.51.100.2 ESP ESP (SPI=0x6d4428aa) 15 12.003690 198.51.100.2 203.0.113.1 ESP ESP (SPI=0x94cd6141) 16 13.004286 203.0.113.1 198.51.100.2 ESP ESP (SPI=0x6d4428aa) 17 13.005508 198.51.100.2 203.0.113.1 ESP ESP (SPI=0x94cd6141) ``` Read it top to bottom and the theory becomes concrete. **Packets 3 to 6: Main Mode.** Wireshark labels IKEv1 Main Mode "Identity Protection", which is the whole point of it. Main Mode is six messages in three round trips. Messages 1 and 2 propose and accept the ISAKMP policy (encryption, hash, DH group, authentication method, lifetime). Messages 3 and 4 carry the Diffie-Hellman public values and nonces, and at the end of message 4 both peers independently derive the same shared secret without it ever crossing the link. Messages 5 and 6 are where each peer proves who it is (with the pre-shared key here), and those two messages are already encrypted with the keys derived in step 2\. That is the identity protection: who you are never travels in the clear. Aggressive Mode collapses this into three messages and gives that up, which is why you avoid it. **Packets 9 and 11: Quick Mode.** Phase 2\. Everything in these packets is encrypted inside the Phase 1 channel, which is why the capture can tell you they are Quick Mode but not what is in them. Quick Mode negotiates the things that will actually protect user traffic: the transform set (AES-256 with SHA-256 HMAC), the mode (tunnel), the proxy identities (which subnets this SA covers), an optional fresh Diffie-Hellman exchange if PFS is on, and the SPIs each side will use. **Packets 12 onward: ESP.** The negotiation is done. From here the ISP sees nothing but IP protocol 50 between two public IPs, and the SPI values that Quick Mode agreed on. ## QM\_IDLE Does Not Mean Idle Once Phase 1 is complete, the ISAKMP SA reports its state: ``` EDGE1#show crypto isakmp sa IPv4 Crypto ISAKMP SA dst src state conn-id status 198.51.100.2 203.0.113.1 QM_IDLE 1001 ACTIVE ``` QM\_IDLE trips people up because it sounds like something is asleep. It is not. It means Main Mode finished successfully, the ISAKMP SA is established and authenticated, and the Quick Mode state machine is idle because it has nothing to do right now. QM\_IDLE is the healthy steady state of Phase 1\. It is what you want to see. Critically, QM\_IDLE says nothing about whether user traffic is flowing. Phase 1 up and Phase 2 dead is an extremely common failure (a transform set mismatch, a crypto ACL mismatch), and it looks exactly like a healthy tunnel in this command. That is why you never stop at `show crypto isakmp sa`. ## What a Security Association Actually Is An SA is a one-way contract: "traffic matching this selector, going in this direction, is protected with this algorithm, this key, and this SPI". The SPI (Security Parameter Index) is the 32-bit tag the receiver uses to look up which SA an inbound ESP packet belongs to. Because an SA is one-way, you always get them in pairs. Here is the live SA on the lab, trimmed: ``` interface: GigabitEthernet2 Crypto map tag: PLZ-CMAP, local addr 203.0.113.1 protected vrf: (none) local ident (addr/mask/prot/port): (10.20.10.0/255.255.255.0/0/0) remote ident (addr/mask/prot/port): (10.30.10.0/255.255.255.0/0/0) current_peer 198.51.100.2 port 500 PERMIT, flags={origin_is_acl,} #pkts encaps: 9, #pkts encrypt: 9, #pkts digest: 9 #pkts decaps: 9, #pkts decrypt: 9, #pkts verify: 9 #send errors 0, #recv errors 0 local crypto endpt.: 203.0.113.1, remote crypto endpt.: 198.51.100.2 plaintext mtu 1438, path mtu 1500, ip mtu 1500, ip mtu idb GigabitEthernet2 current outbound spi: 0x4EA0F404(1319171076) PFS (Y/N): Y, DH group: group14 inbound esp sas: spi: 0xD0E19C7F(3504446591) transform: esp-256-aes esp-sha256-hmac , in use settings ={Tunnel, } sa timing: remaining key lifetime (k/sec): (4607998/3572) Status: ACTIVE(ACTIVE) ``` Read the SPIs. The `current outbound spi` is `0x4EA0F404` and the `inbound esp sas` block shows a completely different SPI, `0xD0E19C7F`. That is not a bug. Each peer chooses the SPI for the traffic it will receive, and tells the other side to use it. Inbound and outbound are separate SAs with separate keys, and that is why the counters come in mirrored pairs (`encaps/encrypt/digest` on the way out, `decaps/decrypt/verify` on the way in). If you see encaps incrementing but decaps stuck at zero, you have a one-way tunnel, and you now know exactly which SA to go looking for. Two other lines earn their keep. `PFS (Y/N): Y, DH group: group14` means Perfect Forward Secrecy is on: Quick Mode ran its own Diffie-Hellman exchange, so cracking the Phase 1 key still would not hand an attacker the data keys. And `plaintext mtu 1438` is ESP overhead made visible, the largest plaintext packet that fits before encryption pushes it past the 1500-byte path MTU. That number is behind a large share of "the tunnel is up but big transfers hang" tickets. ## Proxy IDs: One Crypto ACL Line, One SA Pair In a classic crypto-map design, what gets protected is defined by a crypto ACL (also called the proxy ID, or the traffic selector). It is not a filter that permits or denies. A `permit` line means "encrypt this", and each line becomes its own negotiated pair of SAs. Our lab ACL has two lines: ``` ip access-list extended PLZ-CRYPTO-ACL permit ip 192.168.99.0 0.0.0.255 10.30.10.0 0.0.0.255 permit ip 10.20.10.0 0.0.0.255 10.30.10.0 0.0.0.255 ``` And the session detail proves the consequence: ``` EDGE1#show crypto session detail Interface: GigabitEthernet2 Session status: UP-ACTIVE Peer: 198.51.100.2 port 500 fvrf: (none) ivrf: (none) IKEv1 SA: local 203.0.113.1/500 remote 198.51.100.2/500 Active IPSEC FLOW: permit ip 10.20.10.0/255.255.255.0 10.30.10.0/255.255.255.0 Active SAs: 2, origin: crypto map Inbound: #pkts dec'ed 10 drop 0 life (KB/Sec) 4607998/3516 Outbound: #pkts enc'ed 10 drop 0 life (KB/Sec) 4607999/3516 IPSEC FLOW: permit ip 192.168.99.0/255.255.255.0 10.30.10.0/255.255.255.0 Active SAs: 2, origin: crypto map Inbound: #pkts dec'ed 7 drop 0 life (KB/Sec) 4607999/3562 Outbound: #pkts enc'ed 7 drop 0 life (KB/Sec) 4607999/3562 ``` One ISAKMP SA (Phase 1, shared). Two IPSEC FLOWs, one per ACL line, each showing `Active SAs: 2` (inbound plus outbound). Two ACL lines produced four IPsec SAs and their own independent counters. This is also why proxy IDs must match exactly on both ends: your `permit ip A B` has to be the mirror image of the peer's `permit ip B A`, or Quick Mode is rejected. The full crypto-map build, including how to keep those ACLs symmetrical, is covered in [site-to-site IPsec with crypto maps](https://www.pinglabz.com/site-to-site-ipsec-crypto-maps/). ## The Tunnel Is Built On Demand (and the First Packets Die) With a crypto map there is no tunnel sitting there waiting. The SA is created only when a packet matches the crypto ACL, and building it takes time (Main Mode plus Quick Mode, six plus three messages, with two Diffie-Hellman computations). The packets that trigger it are consumed. Watch the very first ping on a freshly configured tunnel: ``` EDGE1#ping 10.30.10.1 source Loopback10 repeat 5 ..... Success rate is 0 percent (0/5) ``` All five ICMP echoes were used up triggering IKE. Immediately after, with the SA now established, the same ping is perfect: ``` EDGE1#ping 10.30.10.1 source Loopback10 repeat 10 Sending 10, 100-byte ICMP Echos to 10.30.10.1, timeout is 2 seconds: Packet sent with a source address of 10.20.10.1 !!!!!!!!!! Success rate is 100 percent (10/10), round-trip min/avg/max = 10/21/48 ms ``` This is not a fault. It is how on-demand IPsec works, and it is the reason a help desk should never diagnose a VPN on the first ping. Ping twice. ## The Proof: The Original IP Header Really Does Vanish Go back to the capture. Before protection was applied to this link, the transit router could read everything, including the inner host addresses: ``` No. Time Source Destination Proto Info 2 1.461513 192.168.99.100 10.30.10.1 ICMP Echo (ping) request id=0x004e, seq=1/256, ttl=63 3 1.462810 10.30.10.1 192.168.99.100 ICMP Echo (ping) reply id=0x004e, seq=1/256, ttl=255 ``` After tunnel-mode ESP, the same ping across the same link looks like this and nothing more: ``` 12 7.026032 203.0.113.1 198.51.100.2 ESP ESP (SPI=0x6d4428aa) 13 8.939238 198.51.100.2 203.0.113.1 ESP ESP (SPI=0x94cd6141) ``` No 192.168.99.100\. No 10.30.10.1\. No ICMP, no sequence numbers, no protocol identification at all. The only metadata left in the clear is the outer IP header (two public endpoints), the fact that this is IP protocol 50, and the SPI. That is tunnel mode doing its job, and it is worth noting what it does NOT hide: an observer still knows the two gateways are talking, roughly how much, and when. IPsec gives you confidentiality of content, not of the relationship. If you need the routing protocol and multicast to ride inside as well, that is the job of [GRE over IPsec](https://www.pinglabz.com/gre-over-ipsec/). ## Why 3DES, MD5 and DH Group 2 Are Off the Table You will still find 3DES, MD5 and DH group 2 in old configs and in half the tutorials online. Do not copy them. 3DES has a 64-bit block size (vulnerable to Sweet32 birthday-bound collisions on long-lived sessions) and it is slow. MD5 is broken for collision resistance. DH group 2 is a 1024-bit modulus, comfortably inside the reach of a well-funded precomputation attack. The lab config here uses the modern baseline you should default to: AES-256, SHA-256, and DH group 14 (2048-bit), with PFS on so each Phase 2 rekey gets fresh key material. ## Key Takeaways - **IPsec is a framework.** ESP protects packets, IKE negotiates the keys. Use ESP; AH cannot survive NAT because it signs the outer IP header. - **Tunnel mode encrypts the original IP header** and builds a new one addressed to the crypto endpoints. Transport mode keeps the original header exposed and is what you use under GRE or an SVTI. - **IKE is two phases for a reason.** Phase 1 (Main Mode, 6 messages) builds an authenticated channel; Phase 2 (Quick Mode, 3 messages) runs inside it and negotiates the keys that protect user data. Expensive once, cheap thereafter. - **QM\_IDLE is healthy.** It means Phase 1 is up and Quick Mode has nothing pending. It tells you nothing about whether user traffic is being encrypted. - **SAs are unidirectional.** Inbound and outbound have different SPIs and different keys. Always read both counter sets: `encaps/encrypt/digest` out, `decaps/decrypt/verify` in. - **One crypto ACL line equals one SA pair.** Two permit lines gave us two IPSEC FLOWs and four SAs. Proxy IDs must mirror exactly on both peers or Quick Mode is rejected. - **Crypto-map tunnels are built on demand.** The first packets that match the ACL are consumed triggering IKE. A 0/5 first ping followed by 10/10 is normal, not a fault. - **The wire tells the truth.** After tunnel-mode ESP, the transit provider sees two public IPs, protocol 50, and an SPI. Nothing else. Next in the [IPsec VPN cluster](https://www.pinglabz.com/ipsec-vpn/): build it for real with [site-to-site IPsec crypto maps](https://www.pinglabz.com/site-to-site-ipsec-crypto-maps/), carry routing protocols and multicast inside the tunnel with [GRE over IPsec](https://www.pinglabz.com/gre-over-ipsec/), and learn to read the failures fast with the [IPsec VPN troubleshooting guide](https://www.pinglabz.com/troubleshooting-ipsec-vpn/). ### Site-to-Site IPsec VPN with Crypto Maps on Cisco IOS XE URL: https://www.pinglabz.com/site-to-site-ipsec-crypto-maps/ Last updated: 2026-07-13T07:29:50.000Z A site-to-site VPN built with crypto maps is still the config the CCIE Security lab expects you to type from memory, and it is still the one that teaches you the most, because crypto maps make every moving part of [IPsec](https://www.pinglabz.com/ipsec-vpn/) explicit. Nothing hides behind a profile or a tunnel interface. You declare the ISAKMP policy, the key, the transform set, exactly which traffic is interesting, and then you glue it all to a physical interface yourself. This article is the complete, working build. Every piece of CLI output below came out of a live CML lab: EDGE1 (a cat8000v on IOS XE 17.18.02) at HQ, EDGE2 (an identical cat8000v) at the branch, ISP1 (an IOL-XE router) playing the internet between them, and a real Debian VM at 192.168.99.100 on the LAN behind EDGE1\. Nothing here is from a textbook. ## The Lab and the Address Plan **EDGE1 (HQ) public IP**203.0.113.1 on GigabitEthernet2 **EDGE2 (Branch) public IP**198.51.100.2 **HQ networks to protect**192.168.99.0/24 (the real Debian VM) and 10.20.10.0/24 (EDGE1 Loopback10) **Branch network to protect**10.30.10.0/24 (EDGE2 Loopback10) **The internet**ISP1, which knows only the 203.0.113.0/24 and 198.51.100.0/24 transit links That last row is the whole point of the article, and we will come back to it at the end. ## The Five Pieces of a Crypto Map VPN A crypto-map site-to-site VPN is five independent objects that only become a tunnel when you stack them in the right order. If you understand what each one is *for*, you will never have to memorise the config again. 1\. ISAKMP policy The Phase 1 proposal. How the two routers will build the *management* tunnel that protects the key exchange itself. Encryption, hash, auth method, DH group, lifetime. 2\. Pre-shared key The secret that proves each router is who it says it is, bound to a specific peer address. Wrong key and you never get past Phase 1. 3\. Transform set The Phase 2 proposal. How the *data* will actually be encrypted and authenticated once the management tunnel is up. This is the ESP suite. 4\. Crypto ACL The definition of interesting traffic. A permit here does not permit anything, it means *encrypt this*. It also becomes the proxy identity both peers must agree on. 5\. Crypto map + interface The container that binds peer, transform set, PFS and ACL into one policy, applied outbound on the internet-facing interface. Without the interface binding, nothing happens. ## Piece 1: The ISAKMP Policy Phase 1 exists to build a secure channel for negotiating Phase 2\. It is not carrying your user traffic, it is carrying the key material. Both peers must have a policy whose encryption, hash, authentication method and DH group match, or negotiation never completes. ``` crypto isakmp policy 10 encryption aes 256 hash sha256 authentication pre-share group 14 lifetime 3600 ``` AES-256 for confidentiality, SHA-256 for integrity, pre-shared key for authentication, Diffie-Hellman group 14 (2048-bit MODP) for key agreement, and a one-hour Phase 1 lifetime. Policy number 10 is just a priority: lower numbers are proposed first. **Deprecated suites:** you will still find DES, 3DES, MD5 and DH group 1/2/5 in old configs and in half the tutorials on the internet. Do not use them. They are broken or too weak to be worth anyone's time, and Cisco has flagged them for removal. If you inherit a config with `encryption 3des` and `group 2`, that is a migration ticket, not a template. ## Piece 2: The Pre-Shared Key ``` crypto isakmp key PingLabz-L2L-PSK-01 address 198.51.100.2 ``` The key is bound to the peer's public IP. EDGE2 has the mirror of this line pointing at 203.0.113.1 with the same secret. If the strings differ by a single character, Phase 1 stalls in MM\_KEY\_EXCH and the peer logs a sanity-check failure, which is one of the classic signatures covered in [troubleshooting IPsec VPNs](https://www.pinglabz.com/troubleshooting-ipsec-vpn/). ## Piece 3: The Transform Set ``` crypto ipsec transform-set PLZ-TS-AES256 esp-aes 256 esp-sha256-hmac mode tunnel ``` This is the Phase 2 proposal: ESP with AES-256 for encryption and SHA-256 HMAC for authentication, in tunnel mode. Tunnel mode wraps the entire original IP packet inside a new IP header, which is exactly what you want between two sites, because the inner host addresses (the private ones) never appear on the wire. Transform sets must match on both sides. A mismatch here is subtle in the worst way: Phase 1 comes up perfectly, `show crypto isakmp sa` reads QM\_IDLE, and yet no data flows, because Phase 2 was rejected. ## Piece 4: The Crypto ACL (and the Mirror-Image Rule) ``` ip access-list extended PLZ-CRYPTO-ACL permit ip 192.168.99.0 0.0.0.255 10.30.10.0 0.0.0.255 permit ip 10.20.10.0 0.0.0.255 10.30.10.0 0.0.0.255 ``` Read this as "encrypt traffic from HQ's two LANs to the branch LAN". `permit` in a crypto ACL does not mean "allow", it means "this is interesting traffic, protect it". Anything not matched by the ACL leaves the interface in the clear. The single most important rule in the entire article: **the crypto ACL on the far end must be an exact mirror image**. EDGE2's version reverses source and destination and points at the other peer. ``` ! EDGE2 - the mirror image ip access-list extended PLZ-CRYPTO-ACL permit ip 10.30.10.0 0.0.0.255 192.168.99.0 0.0.0.255 permit ip 10.30.10.0 0.0.0.255 10.20.10.0 0.0.0.255 crypto map PLZ-CMAP 10 ipsec-isakmp set peer 203.0.113.1 set transform-set PLZ-TS-AES256 set pfs group14 match address PLZ-CRYPTO-ACL ``` Those ACL entries become the *proxy identities* that the two peers exchange during Quick Mode. If they do not line up, one side proposes 10.30.10.0/24 to 10.20.10.0/24 and the other has no matching policy, so it rejects the proposal outright. That is the number one cause of a VPN that will not come up, and it produces a very specific error code in the logs (error 32) that we break on purpose and decode in the [IPsec troubleshooting guide](https://www.pinglabz.com/troubleshooting-ipsec-vpn/). One more real-world detail that trips people up on live boxes: once the ACL is referenced by an active crypto map, IOS XE will not let you edit it. You have to unbind the map from the interface first. ## Piece 5: The Crypto Map and the Interface Binding ``` crypto map PLZ-CMAP 10 ipsec-isakmp set peer 198.51.100.2 set transform-set PLZ-TS-AES256 set pfs group14 match address PLZ-CRYPTO-ACL interface GigabitEthernet2 crypto map PLZ-CMAP ``` The moment you type `crypto map PLZ-CMAP 10 ipsec-isakmp`, before you have set a peer or matched an ACL, IOS tells you exactly what it needs: ``` % NOTE: This new crypto map will remain disabled until a peer and a valid access list have been configured. ``` This message confuses people because it looks like an error. It is not. It is IOS telling you that a crypto map entry is inert until it has (a) a peer to talk to and (b) traffic worth protecting. Finish the entry and the warning becomes irrelevant. If you ever see a crypto map that is applied but doing nothing, this is the first thing to check: an entry missing a peer or a match statement is silently disabled. The `set pfs group14` line is the one people skip, and it should not be. Perfect Forward Secrecy forces a brand-new Diffie-Hellman exchange for every Phase 2 rekey rather than deriving the data keys from the Phase 1 secret. The payoff is real: if an attacker ever recovers the ISAKMP SA's key material, PFS means they still cannot decrypt the IPsec SAs that were derived independently of it. Each session's keys stand alone. It costs a little CPU at rekey time and buys you a hard limit on how much traffic a single compromised key can expose. Finally, the crypto map must be applied to the outside interface, in the outbound direction, on the interface the traffic will actually exit. A crypto map that exists in the config but is not on an interface encrypts nothing. `show crypto map` is the command that confirms what the router actually assembled (peer, matched ACL, transform set, PFS group, and the interface the map is bound to). If something you configured is missing from that output, no amount of pinging will fix it. ## The First Ping Fails. That Is Correct. Here is the moment that trips up almost everyone the first time. With the config complete on both routers, the very first ping across the tunnel drops every packet. ``` EDGE1#ping 10.30.10.1 source Loopback10 repeat 5 <-- very first attempt ..... Success rate is 0 percent (0/5) ``` Nothing is broken. Crypto-map tunnels are built *on demand*, triggered by interesting traffic. Those five packets were consumed kicking off the IKE negotiation: the router matched them against the crypto ACL, realised no SA existed, dropped them, and started Phase 1\. Building Main Mode plus Quick Mode with DH group 14 takes longer than the two-second ICMP timeout. Run it again and the tunnel is already there. ``` EDGE1#ping 10.30.10.1 source Loopback10 repeat 10 Sending 10, 100-byte ICMP Echos to 10.30.10.1, timeout is 2 seconds: Packet sent with a source address of 10.20.10.1 !!!!!!!!!! Success rate is 100 percent (10/10), round-trip min/avg/max = 10/21/48 ms ``` If someone tells you their new VPN "doesn't work" and their evidence is one failed ping, ask them to ping twice. ## Verification: Phase 1 ``` EDGE1#show crypto isakmp sa IPv4 Crypto ISAKMP SA dst src state conn-id status 198.51.100.2 203.0.113.1 QM_IDLE 1001 ACTIVE ``` **QM\_IDLE is what success looks like.** It means Main Mode finished, the ISAKMP SA is established, and the router is sitting idle waiting to do Quick Mode work. Anything else (MM\_KEY\_EXCH, MM\_NO\_STATE) means Phase 1 is stuck. Note the direction: dst is the peer, src is us. ## Verification: Phase 2 ``` interface: GigabitEthernet2 Crypto map tag: PLZ-CMAP, local addr 203.0.113.1 protected vrf: (none) local ident (addr/mask/prot/port): (10.20.10.0/255.255.255.0/0/0) remote ident (addr/mask/prot/port): (10.30.10.0/255.255.255.0/0/0) current_peer 198.51.100.2 port 500 PERMIT, flags={origin_is_acl,} #pkts encaps: 9, #pkts encrypt: 9, #pkts digest: 9 #pkts decaps: 9, #pkts decrypt: 9, #pkts verify: 9 #send errors 0, #recv errors 0 local crypto endpt.: 203.0.113.1, remote crypto endpt.: 198.51.100.2 plaintext mtu 1438, path mtu 1500, ip mtu 1500, ip mtu idb GigabitEthernet2 current outbound spi: 0x4EA0F404(1319171076) PFS (Y/N): Y, DH group: group14 inbound esp sas: spi: 0xD0E19C7F(3504446591) transform: esp-256-aes esp-sha256-hmac , in use settings ={Tunnel, } sa timing: remaining key lifetime (k/sec): (4607998/3572) Status: ACTIVE(ACTIVE) ``` Read this output top to bottom and it tells you the entire story of the tunnel: - **local ident / remote ident** are the proxy identities, straight out of the crypto ACL. This is the pair that must mirror on the far side. - **`flags={origin_is_acl,}`** confirms the SA was created because an ACL said so, which is the crypto-map way of doing things. - **#pkts encaps and #pkts decaps both climbing** is the only proof that matters. Encaps rising with decaps stuck at zero means your traffic is being encrypted and sent, and nothing is coming back: look at the far end, the return route, or a firewall in between. - **SPI** (Security Parameter Index) is the 32-bit tag that identifies this SA. Each direction has its own, which is why the outbound SPI (0x4EA0F404) and the inbound SPI (0xD0E19C7F) are different values. - `**PFS (Y/N): Y, DH group: group14**` proves `set pfs group14` actually took effect. If this says N, your PFS line is missing or the peer did not agree. - **`transform: esp-256-aes esp-sha256-hmac`** is the negotiated result, not what you configured. This is what the two routers agreed on. - **The lifetime** (4607998 KB / 3572 seconds remaining) counts down in both volume and time. Whichever hits zero first triggers a rekey, and with PFS on, that rekey does a fresh DH exchange. ## One Crypto ACL Line = One SA Pair This is the detail almost nobody shows, and it explains a lot of confusing output. Our crypto ACL has two permit lines, so IPsec builds *two separate SA pairs*, one per line. Send traffic from the loopback and the real VM, and you can watch both light up. ``` EDGE1#show crypto session detail Interface: GigabitEthernet2 Session status: UP-ACTIVE Peer: 198.51.100.2 port 500 fvrf: (none) ivrf: (none) IKEv1 SA: local 203.0.113.1/500 remote 198.51.100.2/500 Active IPSEC FLOW: permit ip 10.20.10.0/255.255.255.0 10.30.10.0/255.255.255.0 Active SAs: 2, origin: crypto map Inbound: #pkts dec'ed 10 drop 0 life (KB/Sec) 4607998/3516 Outbound: #pkts enc'ed 10 drop 0 life (KB/Sec) 4607999/3516 IPSEC FLOW: permit ip 192.168.99.0/255.255.255.0 10.30.10.0/255.255.255.0 Active SAs: 2, origin: crypto map Inbound: #pkts dec'ed 7 drop 0 life (KB/Sec) 4607999/3562 Outbound: #pkts enc'ed 7 drop 0 life (KB/Sec) 4607999/3562 ``` Two IPSEC FLOWs, one ISAKMP SA. The first flow is the router-to-router loopback ping (10 packets each way). The second is the real Debian VM's traffic (7 packets). Each flow shows "Active SAs: 2" because an SA pair is inbound plus outbound. The practical consequences are worth internalising. A crypto ACL with twenty permit lines produces twenty SA pairs, each negotiated separately, each with its own SPIs and its own rekey timer. That is real overhead on a hardware VPN concentrator, and it is exactly the scaling problem that [static VTIs](https://www.pinglabz.com/static-vti-svti-ipsec/) solve by collapsing everything into a single any/any flow. ## Real Traffic from a Real Host Router loopbacks are fine for a lab, but the tunnel exists to carry actual host traffic. Behind EDGE1 sits a real Debian VM on 192.168.99.100, matched by the first line of the crypto ACL. ``` j@llmbits:~$ ping -c 5 10.30.10.1 PING 10.30.10.1 (10.30.10.1) 56(84) bytes of data. 64 bytes from 10.30.10.1: icmp_seq=1 ttl=254 time=2.82 ms 64 bytes from 10.30.10.1: icmp_seq=2 ttl=254 time=2.41 ms 64 bytes from 10.30.10.1: icmp_seq=3 ttl=254 time=2.38 ms 64 bytes from 10.30.10.1: icmp_seq=4 ttl=254 time=2.34 ms 64 bytes from 10.30.10.1: icmp_seq=5 ttl=254 time=2.47 ms --- 10.30.10.1 ping statistics --- 5 packets transmitted, 5 received, 0% packet loss, time 4007ms ``` Five sent, five received, zero loss. The host has no idea IPsec exists. It sends a plain IP packet to 10.30.10.1, EDGE1 matches it against the crypto ACL, encrypts it, wraps it in a new IP header addressed 203.0.113.1 to 198.51.100.2, and ships it across the internet. That transparency is the point of tunnel mode. ## The Clincher: The ISP Cannot Reach Either LAN ISP1 is the router physically in the middle. Every packet between HQ and the branch passes through it. Look at what it knows: ``` ISP1#show ip route Gateway of last resort is not set 198.51.100.0/24 is variably subnetted, 4 subnets, 2 masks C 198.51.100.0/30 is directly connected, Ethernet0/1 C 198.51.100.4/30 is directly connected, Ethernet0/2 203.0.113.0/24 is variably subnetted, 2 subnets, 2 masks C 203.0.113.0/30 is directly connected, Ethernet0/0 ``` Three connected transit subnets. That is the entire routing table. There is no route to 192.168.99.0/24\. There is no route to 10.30.10.0/24\. As far as ISP1 is concerned, those networks do not exist, exactly like a real internet provider that has never heard of your RFC 1918 space. ``` ISP1#ping 10.30.10.1 repeat 3 ... Success rate is 0 percent (0/3) ISP1#ping 192.168.99.100 repeat 3 ... Success rate is 0 percent (0/3) ``` Both fail. Zero percent, both directions. And yet the Debian VM at 192.168.99.100 just pinged 10.30.10.1 with zero loss, *through this router*. That is the whole value proposition of a site-to-site IPsec VPN in two show commands. The two private LANs are completely unreachable from the network between them. They can still talk to each other, at full reliability, because every packet leaves the edge router wearing a new public IP header with an encrypted ESP payload inside it. The transit network forwards the outer packet and never sees, never routes, and never could reach the inner one. ## A Note on Crypto Maps Being Legacy Everything above works, and it is exam-relevant, and you will absolutely still meet it in production. But it is the old way. Crypto maps are stuck to a physical interface, they cannot carry multicast (so no routing protocols across the tunnel), they scale badly because every ACL line is another SA pair to manage, and there is no tunnel interface to apply QoS, NetFlow, or a routing adjacency to. Cisco's modern answer is the tunnel interface: an IPsec profile plus `tunnel protection`. It routes, it carries multicast, it takes an IGP, and the config is shorter. If you are building something new today, build it with a VTI. The full side-by-side is in [crypto maps vs VTI](https://www.pinglabz.com/crypto-maps-vs-vti/), and the build is in [static VTI (SVTI) IPsec](https://www.pinglabz.com/static-vti-svti-ipsec/). Learn crypto maps because they show you every gear in the machine, then deploy VTIs. ## Key Takeaways - A crypto-map VPN is five pieces in order: ISAKMP policy, pre-shared key, transform set, crypto ACL, crypto map bound to the outside interface. Miss the interface binding and nothing encrypts. - The `% NOTE: This new crypto map will remain disabled until a peer and a valid access list have been configured` message is not an error. It is IOS telling you the map entry is inert until it has both. - The first ping across a new crypto-map tunnel will fail. Tunnels are built on demand by interesting traffic, and the trigger packets are consumed starting IKE. Ping twice. - The crypto ACL must be an exact mirror image on the far end. A mismatch is the number one cause of a VPN that will not come up, and it fails in Phase 2 with QM\_IDLE showing green in Phase 1. - QM\_IDLE in `show crypto isakmp sa` means Phase 1 is healthy. Rising #pkts encaps *and* #pkts decaps in `show crypto ipsec sa` is the only proof that Phase 2 is actually passing traffic. - One crypto ACL line equals one SA pair. `show crypto session detail` lists an IPSEC FLOW per line, which is the scaling problem VTIs eliminate. - PFS (`set pfs group14`) forces a fresh Diffie-Hellman exchange at every rekey, so compromising one key never exposes traffic protected by another. Turn it on. - Use AES-256, SHA-256 and DH group 14 or better. DES, 3DES, MD5 and DH groups 1, 2 and 5 are deprecated and belong on a migration plan, not in a new config. - The transit ISP has no route to either private LAN and cannot ping either one, yet the LANs talk with zero loss. That gap is the tunnel, and it is the entire point of [IPsec](https://www.pinglabz.com/ipsec-vpn/). ### CCIE Super Lab 4: Overlay and Automation URL: https://www.pinglabz.com/ccie-super-lab-4-overlay-automation/ Last updated: 2026-07-12T10:01:04.000Z _This post is for paying subscribers only._ ### CCIE Super Lab 3: The Ticket Gauntlet URL: https://www.pinglabz.com/ccie-super-lab-3-ticket-gauntlet/ Last updated: 2026-07-12T10:01:04.000Z _This post is for paying subscribers only._ ### CCIE Super Lab 2: The Service Provider Handoff URL: https://www.pinglabz.com/ccie-super-lab-2-sp-handoff/ Last updated: 2026-07-12T10:01:03.000Z _This post is for paying subscribers only._ ### CCIE Super Lab 1: The Enterprise Core URL: https://www.pinglabz.com/ccie-super-lab-1-enterprise-core/ Last updated: 2026-07-12T10:01:03.000Z _This post is for paying subscribers only._ ### Model-Driven Telemetry on IOS XE Routers: gRPC Dial-Out in Practice URL: https://www.pinglabz.com/model-driven-telemetry-ios-xe-routers/ Last updated: 2026-07-12T09:54:09.000Z SNMP polling is how the network has been monitored for thirty years, and it is showing every one of them. You ask the device "what is your interface counter?" every 30 or 60 seconds, it answers, and you have a sampled, delayed, request-heavy view of a network that changes far faster than your poll interval. Model-driven telemetry inverts the model: the device *streams* data to you, continuously, the moment it changes, over an efficient transport. It is the difference between checking your watch every minute and watching a live clock. This article covers model-driven telemetry (MDT) on IOS XE - the subscription model, dial-in vs dial-out, and the honest platform picture. It closes the automation series. It extends the [network automation guide](https://www.pinglabz.com/network-automation/). ## Why SNMP polling is not enough anymore SNMP polling **Model:** the collector asks, repeatedly **Granularity:** the poll interval (30-60s typical) **Load:** a request-response round trip per OID per poll - heavy at scale **Latency:** up to a full poll interval behind reality Model-driven telemetry **Model:** the device pushes, continuously **Granularity:** sub-second possible; on-change streaming **Load:** one efficient stream, no per-OID requests **Latency:** near-real-time - you see changes as they happen The scaling problem is the killer. A modern network with thousands of interfaces, polled every 30 seconds for dozens of metrics each, generates enormous request-response volume and still gives you a stale, coarse picture. MDT streams the same data continuously, in one efficient pipe, at a granularity SNMP cannot touch. For analytics, capacity planning, and fast fault detection, this is the difference between guessing and knowing. ## The subscription model MDT is built on **YANG models** and **subscriptions**. You subscribe to a data path in a YANG model (say, interface statistics), specify how you want it (periodic every N centiseconds, or on-change), and where to send it (a collector). The device then streams that data per the subscription. ``` telemetry ietf subscription 101 encoding encode-kvgpb filter xpath /interfaces-state/interface/statistics source-address 192.168.99.1 stream yang-push update-policy periodic 1000 ! every 1000 centiseconds = 10 seconds receiver ip address 192.168.99.100 57500 protocol grpc-tcp ``` Read the pieces: - **filter xpath** \- which YANG data path to stream (here, interface statistics). - **stream yang-push** \- the mechanism (yang-push is the standard). - **update-policy** \- `periodic` (every N centiseconds) or `on-change` (stream only when the value changes - the most efficient of all). - **receiver** \- the collector's address, port, and transport (gRPC here). - **encoding** \- how the data is serialised (kvGPB, a compact protobuf-based format, or JSON/XML). ## Dial-in vs dial-out Dial-out (device initiates) The **device** connects to the collector and streams. Configured on the device (as above). The collector just listens. Simpler at the collector, and the device controls the subscription. Dial-in (collector initiates) The **collector** connects to the device over gRPC/NETCONF and requests a subscription dynamically. The collector controls what it wants, when. More flexible, more collector-side logic. **Dial-out is the common production choice** \- the device is configured once with its subscriptions and streams to a known collector, which is easy to scale and survives collector restarts (the device keeps trying). Dial-in suits dynamic, exploratory monitoring where the collector decides on the fly what to pull. ## The collector side The stream has to land somewhere. A typical telemetry stack is: - A **collector** (Telegraf with the cisco\_telemetry plugin, or a purpose-built gRPC receiver) that terminates the stream and decodes the kvGPB/JSON. - A **time-series database** (InfluxDB, Prometheus) that stores the metrics. - A **visualisation layer** (Grafana) that turns them into dashboards. This TIG/TICK-style stack is standard, and the beauty of MDT is that the router pushes clean, model-structured data straight into it - no SNMP MIB wrangling, no polling loops. ## The honest platform picture Here is where we are straight with you, as throughout this cluster. Model-driven telemetry depends on the **Data Management Infrastructure (DMI)** \- the same YANG-management stack that NETCONF and RESTCONF need. And that stack is **not present on the lightweight virtual IOS (IOL)** we use for fast routing labs: ``` R1(config)#telemetry ietf subscription 101 ^ % Invalid input detected at '^' marker. ``` The `telemetry` command does not exist on IOL, because there is no DMI process to serve it - the same reason `netconf-yang` is configured but its port never opens, and `iox` (Guest Shell) is rejected. The heavier virtual platforms that do carry the DMI stack (cat8000v, csr1000v) have their own DMI-initialisation problems in a virtualised lab. So we are not going to show you invented telemetry datapoints from a platform that cannot produce them. The configuration above is the real, current IOS XE MDT syntax, presented as a documented reference. On a physical Catalyst or ASR - or a cloud router with a working DMI - it streams exactly as described, into a Telegraf/InfluxDB/Grafana stack. What *is* real and captured in this cluster is the automation that works everywhere: EEM, and the Jinja2 + Netmiko pipeline. MDT is the direction the industry is moving, and understanding its model is the point - but the honest position is that our lab's platforms cannot boot the DMI, and we say so. ## MDT vs the alternatives - **vs SNMP:** MDT is push not pull, near-real-time not sampled, one stream not per-OID requests. For scale and granularity, it wins decisively. SNMP persists because it is universal and simple, but it is legacy. - **vs syslog:** different job - syslog is discrete events, MDT is continuous metrics. Use both. - **vs NETCONF get:** NETCONF pull is a snapshot on demand; MDT is a continuous stream. MDT is what you want for monitoring; NETCONF get for a point-in-time query. ## Key takeaways - Model-driven telemetry **streams data from the device** continuously, near-real-time, over an efficient transport - replacing SNMP's slow, heavy, sampled polling. - It is built on **YANG models and subscriptions**: subscribe to a data path, choose periodic or on-change, and stream to a collector, typically with gRPC and kvGPB encoding. - **Dial-out** (device streams to a known collector) is the common production model; **dial-in** (collector requests dynamically) is for exploratory monitoring. - The collector stack is typically Telegraf → InfluxDB/Prometheus → Grafana - the router pushes clean model-structured data straight in. - **MDT needs the DMI stack, which the lightweight virtual IOS (IOL) lacks** \- `telemetry` is an invalid command there, as are `netconf-yang` sessions and `iox`. We present the real syntax as documented reference and are honest that our lab cannot stream it. - The fully-captured automation in this cluster is **EEM and the Jinja2 + Netmiko pipeline** \- which work on every platform. MDT is the future; understanding its model is the goal. That closes the automation series, and with it the CCIE Automation and Programmability domain. The full cluster index lives on the [network automation pillar](https://www.pinglabz.com/network-automation/). ### Advanced EEM: Tcl, Multi-Event Correlation, and EEM + Python URL: https://www.pinglabz.com/advanced-eem-tcl-python/ Last updated: 2026-07-12T09:54:08.000Z EEM - the Embedded Event Manager - is the router's built-in automation engine, and it has been there, on every IOS device, since long before "network automation" was a buzzword. It watches for events and takes actions, entirely on the box, with no external system. Most engineers use it for one-line applets and stop there. The expert territory is multi-event correlation, Tcl policies, and EEM as the trigger for on-box Python - which is where EEM becomes genuinely powerful. This article covers advanced EEM, with a real applet firing and capturing device state from a CML lab. It extends the [network automation guide](https://www.pinglabz.com/network-automation/). ## The EEM model: event → action Every EEM policy is an event detector plus a set of actions. When the event fires, the actions run. The event detectors cover an enormous range - syslog patterns, SNMP thresholds, interface state, CLI commands, timers, counters, IP SLA results, and more - and the actions can run CLI commands, send syslog, set counters, email, or launch scripts. ## A single-event applet, fired for real From the lab, an applet that watches for a loopback going down, then captures and logs the interface state: ``` event manager applet LOOP-FLAP-GUARD event syslog pattern "Loopback99, changed state to down" maxrun 60 action 1.0 syslog msg "EEM triggered: Loopback99 went down" action 2.0 cli command "enable" action 3.0 cli command "show ip interface brief | include Loopback99" action 4.0 syslog msg "EEM captured state: $_cli_result" action 5.0 increment flapcount 1 ``` We triggered it by shutting Loopback99, and it ran the full action chain: ``` R1#show logging | include EEM %HA_EM-6-LOG: LOOP-FLAP-GUARD: EEM triggered: Loopback99 went down %HA_EM-6-LOG: LOOP-FLAP-GUARD: EEM captured state: Loopback99 10.99.99.99 YES manual administratively down down R1#show event manager history events | include LOOP 1 1 Actv success Sun Jul12 09:39:58 2026 syslog applet: LOOP-FLAP-GUARD ``` Read what happened: the event fired, the applet issued a live `show` command, and the crucial part - the **`$_cli_result`** built-in variable carried that command's output into the follow-up syslog message. The applet did not just react to the event; it *captured the device's state at the moment of the event* and logged it. That is the pattern that makes EEM a diagnostic tool: catch the state the instant something happens, before it changes. ## Multi-event correlation A single event is easy. The expert feature is correlating *multiple* events - "act only if A *and* B happen within a window", or "act if A *or* B". This is how you avoid false-positive reactions and build genuinely intelligent triggers. ``` event manager applet MULTI-EVENT event tag EV1 syslog pattern "clock" event tag EV2 none trigger correlate event EV1 or event EV2 action 1.0 syslog msg "EEM multi-event correlation fired" ``` Each event gets a **tag**, and the **trigger/correlate** block defines the logic - `and`, `or`, with optional timing windows. From the lab, the registration confirms the correlation: ``` R1#show event manager policy registered | include MULTI 2 applet user multiple Off ... MULTI-EVENT EV2: none: policyname {MULTI-EVENT} sync {yes} ``` The event type `multiple` is the tell - this applet is watching several events and applying logic across them, not reacting to one. The real use cases are powerful: "shut the interface only if both the error counter is high *and* the link is flapping", or "alert if the primary *and* the backup both fail". Correlation turns EEM from a reflex into a decision. ## EEM Tcl policies: when applets are not enough Applets are declarative and limited - a fixed list of actions. When you need real logic - loops, complex parsing, data structures, calculations - you write an EEM **Tcl policy**. Tcl (Tool Command Language) is a full scripting language, and IOS runs Tcl policies as EEM events: ``` event manager directory user policy flash:/eem_scripts event manager policy monitor_bgp.tcl type user ``` Inside the Tcl policy, you have the EEM Tcl library - register for an event, run CLI commands, parse their output with Tcl's string handling, make decisions, and take actions. A Tcl policy can do things an applet cannot: iterate over all BGP neighbours and act on each, parse a complex show output, maintain state across runs. IOS also has an interactive `tclsh` for testing Tcl one-liners on the box. Tcl is showing its age - Python is the modern choice - but it is deeply embedded in IOS and remains the way to do complex on-box logic within EEM on any platform, including the ones without Guest Shell. ## EEM + Python: the modern combination On platforms with [Guest Shell](https://www.pinglabz.com/guest-shell-ios-xe/), EEM's most powerful pattern is launching an [on-box Python script](https://www.pinglabz.com/on-box-python-ios-xe/) as an action. The applet detects the event; Python does the heavy lifting: ``` event manager applet HIGH-CPU-DIAG event snmp oid 1.3.6.1.4.1.9.9.109.1.1.1.1.7.1 get-type exact entry-op ge entry-val 90 poll-interval 10 action 1.0 cli command "guestshell run python /flash/cpu_diagnostics.py" ``` Now the router responds to a CPU spike by running a full Python diagnostic script - capturing process lists, top talkers, whatever the script does - autonomously. This is the best of both: EEM's rich event detection triggering Python's expressive logic. (It requires a Guest-Shell-capable platform, which our lightweight lab image lacks - the EEM side is fully real, the Python-launch side is documented reference, as covered in the on-box Python article.) ## Practical EEM patterns State capture on event The lab's pattern - catch `show` output the instant something fails, using `$_cli_result`. Invaluable for intermittent problems you cannot catch by hand. Scheduled maintenance A `cron` timer event runs a nightly config backup or a health check - the router schedules its own housekeeping. Auto-remediation Detect a known failure signature and fix it - clear a stuck process, bounce an interface, fail over - without waking anyone. Threshold alerting An SNMP-threshold event triggers a targeted alert with captured context, richer than a raw trap. ## Cautions - **Set `maxrun`.** Bound how long an applet can run, or a misbehaving policy can hang. The lab used `maxrun 60`. - **Beware self-triggering loops.** An applet that reacts to a syslog and generates a syslog can trigger itself. Match patterns carefully. - **Test on a lab device.** An auto-remediation applet that gets the logic wrong can make an outage worse. Prove it before deploying. - **Rate-limit destructive actions.** Use counters and `ratelimit` so a flapping condition does not trigger a destructive action dozens of times. ## Key takeaways - EEM is the router's on-box automation engine - event detectors plus actions, no external system. It has been in IOS for two decades. - The lab's real applet **captured device state at the moment of an event** using `$_cli_result` \- the pattern that makes EEM a diagnostic tool for intermittent problems. - **Multi-event correlation** (event tags + a trigger/correlate block with and/or logic) turns EEM from a reflex into a decision. Registration shows event type `multiple`. - **Tcl policies** give you full scripting logic when applets are too limited - loops, parsing, state - on any platform. - **EEM + on-box Python** is the modern combination: rich event detection triggering expressive Python (on a Guest-Shell-capable platform). - Always set `maxrun`, avoid self-triggering loops, rate-limit destructive actions, and test on a lab device first. Next: [model-driven telemetry on IOS XE routers - gRPC dial-out in practice](https://www.pinglabz.com/model-driven-telemetry-ios-xe-routers/). The full cluster index lives on the [network automation pillar](https://www.pinglabz.com/network-automation/). ### JSON vs XML vs YAML for Network Engineers: One Payload, Three Ways URL: https://www.pinglabz.com/json-xml-yaml-network-engineers/ Last updated: 2026-07-12T09:54:08.000Z Every network automation interface speaks in one of three data formats. RESTCONF uses JSON. NETCONF uses XML. Ansible and most human-authored config data use YAML. They all express the same thing - structured data - and once you see that they are three encodings of the same underlying model, the automation stack stops looking like a pile of unrelated tools and starts looking coherent. This article takes one interface configuration and renders it as JSON, XML, and YAML side by side - all generated for real from one Python dictionary in the lab - and explains when each is used. It extends the [network automation guide](https://www.pinglabz.com/network-automation/). ## The key insight: same data, three encodings A network device's configuration is *structured data*: an interface has a name, a description, an IP address (which itself has an address and a mask), an enabled flag. That structure - nested key-value pairs and lists - is the same regardless of how you write it down. JSON, XML, and YAML are three ways to write down the same structure. In the lab, we started with one Python dictionary and rendered it three ways. Here is the dictionary as each format. ### YAML - the human-authored source ``` interface: name: Loopback200 description: DATA-FORMAT-DEMO ipv4: address: 172.30.1.1 netmask: 255.255.255.0 enabled: true ``` ### JSON - the RESTCONF payload ``` { "interface": { "name": "Loopback200", "description": "DATA-FORMAT-DEMO", "ipv4": { "address": "172.30.1.1", "netmask": "255.255.255.0" }, "enabled": true } } ``` ### XML - the NETCONF payload ``` Loopback200 DATA-FORMAT-DEMO
172.30.1.1
255.255.255.0
true
``` Read all three and you see the same tree: an interface with a name, description, a nested ipv4 object, and an enabled flag. Only the syntax differs. A parser reads any of them into the same in-memory structure, which is exactly why you can convert freely between them - as we did, generating all three from one dict. ## Where each is used, and why YAML - for humans **Used by:** Ansible playbooks/vars, most config-as-data files, this cluster's [Jinja2 pipeline](https://www.pinglabz.com/jinja2-network-config-templates/). **Strength:** the most human-readable - minimal punctuation, indentation-based. You write and review it by hand. **Weakness:** whitespace-sensitive; a stray indent is an error. JSON - for APIs **Used by:** RESTCONF, REST APIs, most modern web tooling. **Strength:** universally supported, compact, native to JavaScript and easy in every language. **Weakness:** no comments, more punctuation than YAML, no schema by itself. XML - for NETCONF **Used by:** NETCONF, SOAP, older enterprise systems. **Strength:** namespaces, attributes, mature schema/validation (XSD), the format NETCONF was built on. **Weakness:** verbose - the closing tags double the volume - and heavier to author by hand. ## The mapping to network interfaces The reason this matters is that the format follows the transport: - Automate a device with **NETCONF**? You are sending and receiving **XML**. The `ncclient` Python library builds and parses it. - Automate with **RESTCONF**? You are sending and receiving **JSON** (or XML - RESTCONF supports both, but JSON is the common choice). The `requests` library does the HTTP. - Author your config data for **Ansible** or a Jinja pipeline? You write **YAML**, and the tooling converts it to whatever the transport needs. All three formats describe data structured according to a **YANG model** \- the schema that defines what a valid interface configuration looks like. YANG is format-agnostic: the same YANG model can be encoded as XML (for NETCONF) or JSON (for RESTCONF). So the real picture is: *YANG defines the structure; JSON and XML are how that structure travels; YAML is how you author it.* ## Converting between them in Python Because they are the same data, conversion is trivial - load one format into a Python object, dump it as another: ``` import json, yaml # YAML in data = yaml.safe_load(open('interface.yml')) # JSON out (for a RESTCONF PUT) json_payload = json.dumps(data, indent=2) # YAML back out yaml_again = yaml.safe_dump(data) ``` XML is slightly more involved (libraries like `xmltodict` or `lxml` handle it), but the principle is identical: parse to a common structure, serialise to the target format. This is exactly what we did in the lab to produce all three from one dictionary. **The format is a serialisation detail; the data is what matters.** ## The practical workflow 1. **Author in YAML.** It is the most readable, and it is what you review in a pull request. 2. **Convert to the transport's format** \- JSON for RESTCONF, XML for NETCONF - programmatically, at push time. 3. **Parse the device's response** (JSON or XML) back into a Python structure to verify or extract data. 4. **Never hand-author XML for NETCONF if you can avoid it.** Build it from a structure; the verbosity makes hand-editing error-prone. This workflow means you think in *data*, not in formats. You describe the interface once, and the tooling handles the encoding for whichever interface you happen to be using. ## Key takeaways - JSON, XML, and YAML are **three encodings of the same structured data**. We proved it by rendering one Python dictionary as all three, identical in structure. - **YAML for humans** (Ansible, config-as-data, Jinja pipelines), **JSON for APIs** (RESTCONF, REST), **XML for NETCONF** (and legacy enterprise systems). - The format follows the transport: NETCONF = XML, RESTCONF = JSON, Ansible authoring = YAML. - All three describe data structured by a **YANG model** \- YANG defines the structure, JSON/XML carry it, YAML authors it. - Conversion is trivial in Python because they are the same data - parse to a structure, serialise to the target format. - **Think in data, not formats.** Author in YAML, convert programmatically, never hand-author XML. Next: [Advanced EEM - Tcl, multi-event correlation, and EEM + Python](https://www.pinglabz.com/advanced-eem-tcl-python/). The full cluster index lives on the [network automation pillar](https://www.pinglabz.com/network-automation/). ### Jinja2 Templates for Network Configs: From Variables to Rendered CLI URL: https://www.pinglabz.com/jinja2-network-config-templates/ Last updated: 2026-07-12T09:54:07.000Z Copying a router config, pasting it, and changing the IP addresses by hand is how configuration errors are born. Jinja2 templating replaces that with a template plus a data file: write the structure once, describe each device in a few variables, and render perfect, consistent configuration every time. This is the single most useful automation skill for a network engineer, and unlike some of this cluster, it runs on any platform and needs nothing exotic. This article builds the complete pipeline - YAML variables, a Jinja2 template, rendered CLI, and a Netmiko push to a live router - all captured for real from a CML lab. It extends the [network automation guide](https://www.pinglabz.com/network-automation/). ## The idea: separate structure from data Every device of a given type shares the same *structure* \- the same interfaces, the same protocols, the same command hierarchy. What differs is the *data*: the hostname, the IPs, the site ID. Templating separates the two: - The **template** is the config structure with placeholders. Written once, reviewed once. - The **data** (variables) describes each specific device. Small, readable, per-device. - The **rendered output** is the template with the data filled in - the actual config to push. Get the template right once, and every device rendered from it is correct by construction. No copy-paste drift, no forgotten line, no fat-fingered subnet mask. ## The data: YAML variables YAML is the natural format for the variables - human-readable, structured, and what Ansible and most automation tooling expect. From the lab, describing two loopbacks and an OSPF process: ``` loopbacks: - {id: 100, ip: 172.20.100.1, mask: 255.255.255.0, desc: JINJA-AUTOMATION-100} - {id: 101, ip: 172.20.101.1, mask: 255.255.255.0, desc: JINJA-AUTOMATION-101} ospf: process: 10 networks: - {net: 172.20.100.0, wild: 0.0.0.255, area: 0} - {net: 172.20.101.0, wild: 0.0.0.255, area: 0} ``` That is the whole per-device description. Readable, reviewable, and trivial to keep in version control - a diff on this file shows exactly what changed about the device. ## The template: Jinja2 Jinja2 is a templating language with loops, conditionals, and variable substitution. The template mirrors the config structure, iterating over the data: ``` {% for lo in loopbacks %} interface Loopback{{ lo.id }} description {{ lo.desc }} ip address {{ lo.ip }} {{ lo.mask }} {% endfor %} router ospf {{ ospf.process }} {% for n in ospf.networks %} network {{ n.net }} {{ n.wild }} area {{ n.area }} {% endfor %} ``` `{{ }}` substitutes a variable; `{% for %}` loops. The template says "for each loopback in the data, emit an interface block" - so whether the data has two loopbacks or two hundred, the template is the same. That loop is the leverage: one template, any number of instances. ## The rendered output Feed the YAML data through the Jinja2 template and out comes real, pushable CLI - captured from the lab: ``` interface Loopback100 description JINJA-AUTOMATION-100 ip address 172.20.100.1 255.255.255.0 interface Loopback101 description JINJA-AUTOMATION-101 ip address 172.20.101.1 255.255.255.0 router ospf 10 network 172.20.100.0 0.0.0.255 area 0 network 172.20.101.0 0.0.0.255 area 0 ``` Perfect, consistent, and generated - not typed. The Python that does the rendering is three lines: ``` import yaml, jinja2 data = yaml.safe_load(open('vars.yml')) rendered = jinja2.Template(open('template.j2').read()).render(**data) ``` ## The push: Netmiko to a live device Rendering is half the pipeline. The other half is getting it onto the router. Netmiko - a Python library that wraps SSH to network devices - takes the rendered config and pushes it, then reads back the result. From the lab, driving a real IOS XE router: ``` from netmiko import ConnectHandler dev = {"device_type":"cisco_ios","host":"192.168.99.1", "username":"plauto","password":"..."} conn = ConnectHandler(**dev) out = conn.send_config_set(rendered.splitlines()) print(conn.send_command("show ip interface brief | include Loopback10")) ``` And the device, after the push, verified by the same script: ``` Loopback100 172.20.100.1 YES manual up up Loopback101 172.20.101.1 YES manual up up Interface PID Area IP Address/Mask Cost State Lo100 10 0 172.20.100.1/24 1 LOOP Lo101 10 0 172.20.101.1/24 1 LOOP ``` **That is the full pipeline, end to end, real: YAML data → Jinja2 render → Netmiko push → live config → verified back from the device.** No hand-typing at any step. And critically, this runs on *any* platform - the router just needs SSH. No Guest Shell, no IOx, no DMI. It is the automation that works everywhere. ## Why this pipeline is the foundation Everything more advanced builds on this pattern: - **Ansible** is essentially this pipeline with structure, inventory, and idempotency added - it uses Jinja2 templates and pushes over SSH/NETCONF exactly like this. - **CI/CD for networks** is this pipeline in a git repository with a test stage - the YAML vars are reviewed in a pull request, rendered, tested, and pushed by a pipeline. - **Infrastructure as code** is the philosophy this embodies: the network's config is defined by data files in version control, and rendered/pushed by tooling, not typed by hand. Master this small pipeline and you understand the mechanism behind all of it. The advanced tools are conveniences layered on this exact idea. ## Practical Jinja2 patterns **Conditionals** `{% if intf.dhcp %}ip address dhcp{% else %}ip address {{ intf.ip }} {{ intf.mask }}{% endif %}` \- render different config based on the data. **Filters** `{{ hostname | upper }}`, `{{ vlans | join(',') }}` \- transform values inline. **Defaults** `{{ mtu | default(1500) }}` \- a fallback when the variable is absent. **Whitespace control** `{%- ... -%}` trims whitespace - important for clean config output without stray blank lines. ## Guardrails - **Review the rendered output before pushing.** Render to a file, eyeball it (or diff it against the current config), *then* push. A template bug rendered across a hundred devices is a hundred-device outage. - **Version-control the data and templates.** The whole point is that config changes are data changes, reviewable in a pull request. - **Test with a lab device first.** Push to one router, verify, then roll out - exactly as we did here. - **Keep templates modular.** One template per function (interfaces, routing, security) composed together, rather than one giant template - easier to review and reuse. ## Key takeaways - Jinja2 templating separates config **structure** (the template, written once) from **data** (per-device variables), so every rendered config is correct by construction. - **YAML** holds the variables - readable and version-controllable. **Jinja2** holds the template, with loops and conditionals. **Netmiko** pushes the rendered config over SSH. - Verified end to end in the lab: YAML → Jinja2 render → Netmiko push → live config → verified back from a real IOS XE router. - This pipeline runs on **any platform** \- it just needs SSH. No Guest Shell, no IOx, no NETCONF required. - It is the **foundation** for Ansible, network CI/CD, and infrastructure-as-code - all of which are this exact pattern with more structure around it. - Always render to a file and review before pushing; version-control the data and templates; test on a lab device first. Next: [JSON vs XML vs YAML for network engineers - one payload, three ways](https://www.pinglabz.com/json-xml-yaml-network-engineers/). The full cluster index lives on the [network automation pillar](https://www.pinglabz.com/network-automation/). ### On-Box Python: Scripting IOS XE from the Inside URL: https://www.pinglabz.com/on-box-python-ios-xe/ Last updated: 2026-08-01T19:32:43.000Z Off-box automation - a Python script on a server driving routers over SSH - is where most network automation lives, and rightly so. But it has one structural weakness: it depends on the automation host being reachable. If the management network is down, or the site is isolated, the automation cannot run. On-box Python removes that dependency by putting the script *on the device*, where it can react to local conditions no matter what the rest of the network is doing. This article covers on-box Python on IOS XE - the models, the `cli` module, and where it runs - with the honest platform picture and a fully-captured off-box pipeline as the real, working reference. It extends the [network automation guide](https://www.pinglabz.com/network-automation/). ## The two on-box models Guest Shell Python Python running inside the [Guest Shell](https://www.pinglabz.com/guest-shell-ios-xe/) Linux container. Full Python environment, pip packages, the `cli` module to reach IOS. The richer option. On-box Python (guestshell run python) A Python interpreter reachable directly from the IOS CLI, often invoked by EEM to run a script in response to an event. Tighter integration with IOS, same `cli` module. Both use the same key primitive: the **`cli` module**, which is Python's window into IOS. ## The cli module: Python's bridge to IOS The `cli` module (available in the on-box Python environment) provides functions that run IOS commands and hand you the result as a Python string you can parse: ``` from cli import cli, clip, configure, execute # Read operational state routes = cli('show ip route') # returns the output as a string clip('show ip interface brief') # runs it AND prints it # Make configuration changes configure(['interface Loopback99', 'ip address 10.99.99.99 255.255.255.255', 'no shutdown']) # Execute an exec command execute('write memory') ``` This is what makes on-box Python powerful: your script has the full expressiveness of Python - loops, conditionals, regex, data structures, external libraries - and can drive the device's configuration and read its state through `cli`. A script can pull the routing table, parse it, decide something, and reconfigure - all on the box. ## The natural pairing: EEM + on-box Python On-box Python's killer combination is with EEM. An [EEM applet](https://www.pinglabz.com/advanced-eem-tcl-python/) detects an event (a syslog message, an SNMP threshold, an interface state change) and, as one of its actions, launches an on-box Python script: ``` event manager applet HIGH-CPU-CAPTURE event snmp oid 1.3.6.1.4.1.9.9.109.1.1.1.1.7.1 get-type exact entry-op ge entry-val 90 poll-interval 10 action 1.0 cli command "guestshell run python /flash/capture_diag.py" ``` Now the router responds to high CPU by running a Python script that captures diagnostics, all autonomously, with no external system involved. That is the pattern on-box Python exists for: **local, autonomous reaction to local events**. The router self-manages. ## The honest platform picture On-box Python needs the same foundation as Guest Shell - IOx on a full IOS XE platform. And as covered in the [Guest Shell article](https://www.pinglabz.com/guest-shell-ios-xe/), the lightweight virtual IOS (IOL) used for our fast routing labs does not have it: `iox` is an invalid command, so there is no on-box Python environment to enter. The heavier virtual platforms that do have the app-hosting stack have their own DMI-initialisation issues in a lab. Rather than fabricate on-box Python output on a platform that cannot produce it, we show you the real, captured thing: the **off-box Python pipeline** that does the same work from an external host, verified end to end on a live device. ## The real, captured alternative: off-box Python with Netmiko From the lab, a Python script on the Linux host drove a real IOS XE router over SSH - rendering config, pushing it, and reading back the device state to verify. The push result and the verification, both real: ``` === NETMIKO PUSH RESULT (from the automation host) === R1(config-router)# network 172.20.100.0 0.0.0.255 area 0 R1(config-router)# network 172.20.101.0 0.0.0.255 area 0 R1(config-router)#end === DEVICE VERIFICATION (pulled back via the same script) === Loopback100 172.20.100.1 YES manual up up Loopback101 172.20.101.1 YES manual up up Lo100 10 0 172.20.100.1/24 1 LOOP Lo101 10 0 172.20.101.1/24 1 LOOP ``` The Python here - `ConnectHandler`, `send_config_set`, `send_command` \- is the off-box equivalent of the `cli` module's `configure` and `cli` functions. The *logic* of "read state, decide, configure, verify" is identical; only the location of the script differs. This is the model you will use for the overwhelming majority of automation, and it works on every platform. Since off-box is the model you will actually run, it is worth knowing properly: [how Netmiko drives a Cisco device over SSH](https://www.pinglabz.com/netmiko-ssh-automation-cisco/) takes the same `ConnectHandler` and `send_config_set` pair much further, including the awkward truth that the library returns a transcript and cannot tell you whether your push changed anything - a problem the on-box `cli` module has in exactly the same form. ## On-box vs off-box: the decision 1. **Managing many devices from one place, version-controlled, in a pipeline?** Off-box. Always. This is 95% of automation. 2. **A device that must react to a local event with no dependency on an external system** (an isolated site, a self-healing edge)? On-box, launched from EEM. This is the real niche. 3. **A one-off diagnostic that needs to run on the box during an event** (capture state the moment CPU spikes)? On-box, because by the time you SSH in the moment has passed. 4. **Anything where the automation host is reliably reachable?** Off-box. It is simpler, more scalable, and platform-independent. ## Key takeaways - On-box Python runs a script **on the router**, removing the dependency on an external automation host - powerful for autonomous local reactions. - The `cli` module is the bridge: `cli()` / `clip()` to read state, `configure()` to make changes, from inside Python on the device. - Its killer pairing is **EEM + on-box Python**: an event triggers a Python script that responds autonomously, with no external system. - It needs IOx on a full IOS XE platform. IOL (our lightweight lab image) has no `iox`, so we describe it honestly and show the real off-box equivalent rather than faking output. - The captured, working reference is **off-box Python with Netmiko** \- the same read/decide/configure/verify logic, from an automation host, verified on a live device. - **Off-box is the default for almost everything.** On-box is a specialist tool for genuine local autonomy. Next: [Jinja2 templates for network configs - from variables to rendered CLI](https://www.pinglabz.com/jinja2-network-config-templates/), the fully-captured pipeline. The full cluster index lives on the [network automation pillar](https://www.pinglabz.com/network-automation/). ### Guest Shell on IOS XE: Linux Inside Your Router URL: https://www.pinglabz.com/guest-shell-ios-xe/ Last updated: 2026-07-12T09:54:06.000Z There is a full Linux environment hiding inside your router. Guest Shell is a Linux container (CentOS or a minimal distro, depending on platform) that runs on the router's own hardware, with access to the device's IOS XE CLI from inside it. You can run Python, install packages, and script the router from a shell that lives *on* the router - no external automation host required. This article covers Guest Shell and on-box scripting: what it is, why it matters, and - honestly - where you can and cannot run it. It extends the [complete network automation guide](https://www.pinglabz.com/network-automation/). ## What Guest Shell is Guest Shell runs inside IOx, Cisco's application-hosting framework on IOS XE. It gives you a Linux user space with: - **Python** pre-installed, so you can write scripts that run on the router itself. - The **cli module** \- a Python library that lets your script issue IOS commands and get the output back, from inside the container. - **Package management** \- install libraries with pip, subject to the container's constraints. - Network access through the router's management, so the script can reach external services. The point is **autonomy**. A script in Guest Shell runs on the device, reacts to local events, and needs no external server. Combine it with EEM (which can launch a Guest Shell script on a trigger) and the router becomes self-managing - it detects a condition and runs Python to respond, all on-box. ## Enabling it (on a platform that supports it) ``` ! Enable IOx (the app-hosting framework) iox ! ! Enable and enter Guest Shell guestshell enable guestshell run bash ``` Inside, you have a Linux prompt. From there: ``` [guestshell]$ python3 >>> from cli import cli, clip, configure >>> clip('show ip interface brief') # run an IOS show, print the output >>> configure(['interface Loopback50', 'ip address 10.50.50.1 255.255.255.255']) ``` That `cli` module is the bridge. Your Python, running in the Linux container on the router, drives the router's IOS configuration and reads its operational state. A script can loop over interfaces, make decisions, and reconfigure - entirely on the device. ## The platform reality, stated honestly Here is where PingLabz will not fake a screenshot. **Guest Shell needs a full IOS XE platform with IOx - and on the virtual images available in our CML lab, it is not present.** On the IOL-XE nodes we use for most routing labs, the command simply does not exist: ``` R1(config)#iox ^ % Invalid input detected at '^' marker. ``` IOL (the lightweight virtual IOS used for fast routing labs) has no IOx container framework, so there is no Guest Shell to enter. The heavier virtual platforms that *do* have the app-hosting stack (cat8000v, csr1000v) require the Data Management Infrastructure to initialise, which in a virtualised lab environment is itself unreliable. So we are not going to show you invented Guest Shell output on a platform that cannot produce it. What is real, and what this whole cluster is built on, is the *alternative*: driving the router with Python from an external host. The [Jinja2 + Netmiko pipeline](https://www.pinglabz.com/jinja2-network-config-templates/) renders configuration and pushes it to a real IOS XE device, verified end to end - all captured on the lab. That is off-box automation, and it does everything Guest Shell does except run on the device itself. For the vast majority of automation tasks, off-box is what you use anyway. ## On-box vs off-box: when each makes sense On-box (Guest Shell) **Runs:** on the router itself **Best for:** autonomous local reactions - a script that responds to a local event with no dependency on an external system **Needs:** a platform with IOx, and container resources Off-box (Netmiko, NETCONF, Ansible) **Runs:** on an automation host, over SSH/NETCONF/RESTCONF **Best for:** managing many devices consistently from one place, version-controlled config, CI/CD **Needs:** just SSH/API access to the devices The honest guidance: **off-box automation is what you should reach for by default.** It scales to your whole estate from one controlled place, it version-controls cleanly, and it works on every platform including IOL. Guest Shell is a specialist tool for the specific case where you need a script to run autonomously on a device with no external dependency - an edge router at a site with no reliable management connectivity that must self-heal locally, for instance. That is a real use case, but it is a narrow one. ## What a Guest Shell script looks like (documented reference) For completeness, a Guest Shell Python script that pulls interface stats and acts on them - the kind of thing you would run on a supported platform: ``` #!/usr/bin/env python3 from cli import cli, configure # Read operational state via the CLI module output = cli('show interfaces | include line protocol|input errors') # Parse, decide, act - all on the router for line in output.splitlines(): if 'input errors' in line and int(line.split()[0]) > 1000: # High error count - shut the interface and log configure(['interface GigabitEthernet1', 'shutdown']) cli('send log "Guest Shell: high errors, interface shut"') ``` Launched from EEM on a trigger, this makes the router self-managing. We present it as a documented reference of the on-box model, clearly labelled, because it is real technology on a real platform - just not one our lab can boot. ## Key takeaways - Guest Shell is a **Linux container running on the router** (via IOx), with Python and the `cli` module that lets on-box scripts drive IOS configuration and read operational state. - It enables **autonomous on-box automation** \- a script that reacts to local events with no external dependency, especially powerful when launched from EEM. - **It needs a full IOS XE platform with IOx.** IOL (the lightweight virtual image) has no `iox` command, and the heavier virtual platforms' DMI is unreliable in a lab - so we describe it honestly rather than faking output. - The real, captured alternative is **off-box automation**: the Jinja2 + Netmiko pipeline that renders and pushes config to a live device, verified end to end. - **Off-box is the default** \- it scales, version-controls, and works everywhere. Guest Shell is a specialist tool for genuine on-box autonomy. Next: [on-box Python: scripting IOS XE from the inside](https://www.pinglabz.com/on-box-python-ios-xe/). The full cluster index lives on the [network automation pillar](https://www.pinglabz.com/network-automation/). ### Expert Services Troubleshooting: NAT, DHCP, and SLA Ticket Scenarios URL: https://www.pinglabz.com/expert-services-troubleshooting/ Last updated: 2026-07-12T09:34:26.000Z Services break in ways that routing never does. A route either exists or it does not; a NAT translation either happened or it did not, and if it did not, the packet was silently dropped with no log and no routing-table clue. Services troubleshooting is a different discipline - it is about knowing which translation table, which binding, or which tracking object to look at, because the symptom is almost always just "it does not work" with nothing obvious in `show ip route`. This article is five services faults, grounded in the real behaviour from building the services lab for this domain, with the diagnostic command for each. For the theory, see the [IP services pillar](https://www.pinglabz.com/ip-services/). ## The method Services faults do not show up in the routing table. Go straight to the service's own state table: **NAT**`show ip nat translations` (and `vrf X` for VRF-aware). Did the translation happen? Is it in the right VRF? **DHCP**`show ip dhcp binding`, `show ip dhcp pool`, `show ip dhcp conflict`. Got a lease? Pool exhausted? Conflict? **SNMP**`show snmp view`, `show snmp user`. Is the OID in the view? Do the engine IDs match? **Tracking / SLA**`show ip sla statistics`, `show track`. Is the probe running? What state is the object? ## Ticket 1: "Two overlapping customers can reach the shared server, but one sees the other's traffic" **Symptom.** Two VRFs with overlapping 10.5.5.0/24, NATed to a shared exit. Both work, but return traffic is going to the wrong customer, or translations are colliding. **Diagnosis.** Look at each VRF's translation table: ``` R1#show ip nat translations vrf RED icmp 192.168.99.201:1024 10.5.5.50:0 192.168.99.100:0 192.168.99.100:1024 R1#show ip nat translations vrf BLUE icmp 192.168.99.202:1024 10.5.5.50:0 192.168.99.100:0 192.168.99.100:1024 ``` Healthy is what you see above - identical inside-local (10.5.5.50) in each, but distinct inside-global (.201 vs .202). If instead both VRFs show the *same* inside-global, or the translations appear in the wrong VRF, the NAT is not VRF-aware. **Cause.** The `vrf` keyword is missing on the `ip nat inside source` statement, so NAT treats the two 10.5.5.50s as one and they collide. Or both VRFs share one pool. **Fix.** A distinct pool and ACL per VRF, and the `vrf` keyword on each rule: ``` ip nat inside source list RED-ACL pool RED-POOL-NAT vrf RED overload ip nat inside source list BLUE-ACL pool BLUE-POOL-NAT vrf BLUE overload ``` **Lesson:** for VRF-aware NAT, everything must be per-VRF - the pool, the ACL, and the `vrf` keyword. Verify by checking that `show ip nat translations vrf X` shows a distinct inside-global per VRF. Full detail in [VRF-aware NAT](https://www.pinglabz.com/vrf-aware-nat-ios-xe/). ## Ticket 2: "Some clients get no DHCP address" **Symptom.** New clients on a subnet are not getting addresses. Others already have leases. **Diagnosis.** ``` R1#show ip dhcp pool CLIENTS Total addresses : 254 Leased addresses : 254 <-- pool is full Excluded addresses : 9 ``` If leased equals total, the pool is exhausted. Also check for conflicts, which take addresses out of service: ``` R1#show ip dhcp conflict IP address Detection method Detection time 10.7.7.55 Ping Jul 12 2026 ... ``` **Cause.** Two possibilities. Either the lease is too long for the churn (dead reservations for devices that left hold addresses for a day), or a duplicate-IP conflict took addresses out of the pool because a static device was not excluded. **Fix.** Shorten the lease to the device dwell time (`lease 0 1 0` for a one-hour guest network), and exclude every statically-assigned address (`ip dhcp excluded-address`). Clear stale conflicts with `clear ip dhcp conflict *`. **Lesson:** "no DHCP address" is a pool problem 90% of the time. `show ip dhcp pool` tells you instantly whether it is exhaustion; `show ip dhcp conflict` tells you whether duplicates are eating it. ## Ticket 3: "SNMP polling works but one OID returns nothing" **Symptom.** The NMS polls the router fine for interface stats, but a specific OID (say a temperature sensor or a Cisco-specific table) returns "no such object". **Diagnosis.** ``` R1#show snmp view | include PLVIEW PLVIEW iso - included PLVIEW internet - included ``` **Cause.** The OID being polled is *not in the view* the user's group is bound to. SNMPv3 views are default-deny: anything not explicitly `included` is invisible, and the router returns nothing rather than an error. **Fix.** Add the OID subtree to the view: ``` snmp-server view PLVIEW 1.3.6.1.4.1.9 included ! include the Cisco enterprise subtree ``` **Lesson:** when SNMPv3 polling works for some OIDs but not others, it is almost always a view restriction, not a connectivity problem. `show snmp view` shows exactly what is permitted. (And if *nothing* works, or traps do not arrive, check the engine ID - see [SNMPv3 in production](https://www.pinglabz.com/snmpv3-traps-informs/).) ## Ticket 4: "The EPC capture buffer is empty / fills instantly" **Symptom.** An Embedded Packet Capture either captures nothing, or fills up before the event you are hunting. **Diagnosis.** ``` R1#show monitor capture buffer EPCBUF parameters Buffer Size : 262144 bytes, Max Element Size : 1518 bytes, Packets : 0 ``` **Causes and fixes.** Several distinct problems present the same way: - **Nothing captured, and the config was rejected?** You used the modern `monitor capture NAME interface` syntax on a platform that only supports the classic `monitor capture buffer/point` model. Switch syntaxes. - **Nothing captured, config accepted?** The capture point is not associated with the buffer, or not started. `monitor capture point associate` then `start`. - **Nothing captured, but traffic is flowing?** Your ACL filter matches only one direction, or the wrong addresses. Match both directions of the conversation. - **Buffer full before the event?** A `linear` buffer stopped when full. Use `circular` to keep the most recent packets, or raise the size. **Lesson:** `show monitor capture buffer X parameters` tells you the packet count and the point status in one view. An empty buffer with an inactive point means it was never started; a full linear buffer means your event happened after it filled. Full detail in [Embedded Packet Capture](https://www.pinglabz.com/embedded-packet-capture-ios-xe/). ## Ticket 5: "The IPv6 downstream interface has no address after a delegation change" **Symptom.** A DHCPv6-PD delegated prefix changed (the ISP renumbered), but a downstream interface still has the old address, or no address. **Diagnosis.** Check the general-prefix chain: ``` R2#show ipv6 general-prefix IPv6 Prefix DELEGATED-FROM-R1, acquired via DHCP PD 2001:DB8:AAAA::/64 Valid lifetime 3575, preferred lifetime 1775 Ethernet0/1 (Address command) R2#show ipv6 interface Ethernet0/1 2001:DB8:AAAA::1, subnet is 2001:DB8:AAAA::/64 ``` **Cause.** If the general prefix shows the *new* prefix but the interface still shows the *old* address, the interface is not referencing the general prefix by name - it has a hard-coded address instead. If the general prefix itself is missing, the PD lease was lost (check `show ipv6 dhcp binding` on the delegating router). **Fix.** Number the interface *from* the general prefix so it tracks changes automatically: ``` interface Ethernet0/1 ipv6 address DELEGATED-FROM-R1 ::1:0:0:0:1/64 ``` **Lesson:** the whole point of a general prefix is automatic renumbering. If an interface is not updating when the delegation changes, it is not referencing the general prefix - it has a static address. `show ipv6 general-prefix` shows the current prefix and which interfaces reference it. Detail in [IPv6 services closure](https://www.pinglabz.com/ipv6-nptv6-dhcpv6-pd/). ## The commands, collected ``` ! NAT show ip nat translations [vrf X] show ip nat statistics ! DHCP show ip dhcp binding show ip dhcp pool show ip dhcp conflict ! SNMP show snmp view show snmp user show snmp host ! EPC show monitor capture buffer X parameters show monitor capture buffer X ! IPv6 services show ipv6 dhcp binding (delegating router) show ipv6 general-prefix (requesting router) show ipv6 interface X ! Tracking / SLA (if used for NAT/route failover) show ip sla statistics show track ``` ## Key takeaways - Services faults do not appear in the routing table. Go to the service's own state table: translations, bindings, views, tracking objects. - **VRF-aware NAT colliding:** the `vrf` keyword is missing or a pool is shared. Verify a distinct inside-global per VRF with `show ip nat translations vrf X`. - **No DHCP address:** pool exhaustion or a conflict. `show ip dhcp pool` and `show ip dhcp conflict`. Fix with shorter leases and proper exclusions. - **SNMP OID missing:** a view restriction, not connectivity. `show snmp view`. (Nothing working at all = engine-ID mismatch.) - **Empty/full EPC buffer:** wrong syntax for the platform, unstarted point, one-directional filter, or a full linear buffer. `show monitor capture buffer X parameters`. - **IPv6 interface not renumbering:** it has a static address instead of referencing the general prefix. `show ipv6 general-prefix`. That closes the expert services series, and with it the CCIE Infrastructure Security and Services domain. The full cluster index lives on the [IP services pillar](https://www.pinglabz.com/ip-services/), cross-linked to [infrastructure security](https://www.pinglabz.com/infrastructure-security/). ### IPv6 Services Closure: NPTv6, DHCPv6-PD, and General Prefix URL: https://www.pinglabz.com/ipv6-nptv6-dhcpv6-pd/ Last updated: 2026-07-12T09:34:25.000Z IPv6 does several things differently from IPv4, and a few of them are genuinely better - once you understand them. Prefix delegation hands a whole subnet to a downstream router automatically. A general prefix lets you renumber an entire site by changing one line. NPTv6 translates prefixes without the stateful baggage of NAT. These are the IPv6 services that close out the CCIE services domain, and they are the ones engineers who "know IPv6" often skip. This article covers DHCPv6-PD, general prefixes, and NPTv6, with real output from a CML lab showing a delegated prefix flowing all the way to a downstream interface. For the fundamentals, see the [complete IPv6 guide](https://www.pinglabz.com/ipv6/). ## DHCPv6 Prefix Delegation: automatic downstream subnets In IPv4, a home or branch router gets a single address from the ISP and NATs everything behind it. IPv6 does something far more elegant: the ISP *delegates a whole prefix* \- say a /56 or /48 - to the customer router, which then sub-delegates /64s to its own downstream segments. No NAT, real end-to-end addressing, and it happens automatically. DHCPv6 Prefix Delegation (PD) is the mechanism. There are two roles: - The **delegating router** (upstream, e.g. the ISP edge) owns a pool of prefixes and hands them out. - The **requesting router** (downstream, the customer) asks for a prefix and uses it to number its interfaces. ### The delegating router ``` ipv6 dhcp pool PD-POOL prefix-delegation pool DELEGATED-PREFIXES lifetime 3600 1800 dns-server 2001:DB8:99::53 ! ipv6 local pool DELEGATED-PREFIXES 2001:DB8:AAAA::/48 64 ! interface Ethernet0/3 ipv6 nd other-config-flag ipv6 dhcp server PD-POOL ``` The `ipv6 local pool` defines the space to delegate (2001:DB8:AAAA::/48) and the size of each delegation (/64). The DHCP pool hands out prefixes from it. From the lab, once the downstream router requested: ``` R1#show ipv6 dhcp binding Client: FE80::A8BB:CCFF:FE00:6230 IA PD: IA ID 0x00050001, T1 900, T2 1440 Prefix: 2001:DB8:AAAA::/64 preferred lifetime 1800, valid lifetime 3600 ``` The delegating router has handed 2001:DB8:AAAA::/64 to the downstream router and tracks the lease exactly like an IPv4 DHCP binding - but for a whole subnet, not a single address. ### The requesting router, and the general prefix The downstream router asks for the prefix and - here is the clever part - stores it as a **general prefix**, a named handle it can then use to number any interface: ``` interface Ethernet0/3 ipv6 dhcp client pd DELEGATED-FROM-R1 ! request, store as this general prefix ! interface Ethernet0/1 ipv6 address DELEGATED-FROM-R1 ::1:0:0:0:1/64 ! number FROM the delegated prefix ``` And it works end to end: ``` R2#show ipv6 general-prefix IPv6 Prefix DELEGATED-FROM-R1, acquired via DHCP PD 2001:DB8:AAAA::/64 Valid lifetime 3575, preferred lifetime 1775 Ethernet0/1 (Address command) R2#show ipv6 interface Ethernet0/1 2001:DB8:AAAA::1, subnet is 2001:DB8:AAAA::/64 [CAL/PRE] ``` Read the chain: the router *acquired* 2001:DB8:AAAA::/64 via DHCP PD, stored it as the general prefix `DELEGATED-FROM-R1`, and its downstream interface E0/1 automatically got a global address (2001:DB8:AAAA::1) *built from that prefix*. The delegated prefix flowed from the ISP, through the customer router, onto a LAN interface - no manual addressing at any step. ## General prefix: renumber a site in one line The general prefix is worth dwelling on because it solves a real operational pain. In IPv4, if your ISP renumbers you, you touch every interface, every ACL, every static route with the old prefix in it. In IPv6 with a general prefix, you define the prefix once (from PD, or manually) and reference it by name on every interface: ``` interface Ethernet0/1 ipv6 address DELEGATED-FROM-R1 ::1:0:0:0:1/64 interface Ethernet0/2 ipv6 address DELEGATED-FROM-R1 ::2:0:0:0:1/64 ``` Each interface's address is "the general prefix, plus this per-interface suffix". **Change the general prefix - because the ISP renumbered you, or you switched providers - and every interface referencing it renumbers automatically.** One change, whole site. And with DHCP PD feeding the general prefix, that change happens on its own when the delegation updates. This is IPv6 solving a problem IPv4 never could, and it is why PD + general prefix is the standard modern IPv6 edge design. ## NPTv6: prefix translation without NAT's baggage Sometimes you do need to translate IPv6 prefixes - most commonly for multihoming without provider-independent space, or to present a stable internal prefix while the external one changes. IPv4's answer is NAT, with all its stateful, connection-tracking, application-breaking overhead. IPv6 offers a cleaner tool: **NPTv6** (Network Prefix Translation, RFC 6296). NPTv6 is *stateless* and *1:1*. It translates one prefix to another algorithmically - the internal 2001:DB8:1::/64 becomes external 2001:DB8:2::/64 by a checksum-neutral prefix rewrite, with the host portion untouched. Because it is stateless and 1:1, it does not break end-to-end connectivity the way IPv4 NAT does: every internal host still maps to exactly one external address, deterministically, and inbound connections work. ``` interface Ethernet0/0 ipv6 nat interface Ethernet0/1 ipv6 nat ! ipv6 nat prefix 2001:DB8:1::/64 ... ! translate internal to external prefix ``` **Platform note, honestly stated:** NPTv6 support varies by IOS XE image, and on the lightweight virtual platform used for this lab it is not fully available. On a full IOS XE platform (cat8000v and similar) it is present. Rather than stage output we could not produce, we describe the mechanism accurately: NPTv6 is stateless, 1:1, checksum-neutral prefix translation. Where DHCPv6-PD and general prefixes are fully captured above (they work on the lab platform), NPTv6 is a concept-plus-platform-note - which is the honest position when a feature is image-dependent. The design point stands regardless: **NPTv6 is what you reach for when you would have used NAT in IPv4, and it is deliberately less harmful.** Prefer provider-independent addressing where you can (so you translate nothing); use NPTv6 when you genuinely need prefix translation but want to keep the end-to-end model that makes IPv6 worth deploying. ## When to use each DHCPv6-PD Any edge that gets its IPv6 space from an upstream. The standard way a branch or home router is numbered by its ISP. Use it. General prefix Anywhere you want to renumber by changing one value - pairs naturally with PD. Reference it on every interface and you never hand-number again. NPTv6 Only when you genuinely need prefix translation (multihoming without PI space). Prefer PI addressing first. It is the least-bad translation, not a default. ## Key takeaways - **DHCPv6-PD** delegates a whole prefix to a downstream router automatically - real end-to-end addressing, no NAT. A delegating router hands out prefixes; a requesting router uses them. - Verified end to end: the delegating router handed 2001:DB8:AAAA::/64, the requesting router stored it as a **general prefix**, and its downstream interface was automatically numbered (2001:DB8:AAAA::1) from it. - A **general prefix** lets you renumber an entire site by changing one value - every interface referencing it updates automatically. Pair it with PD for hands-off edge addressing. - **NPTv6** is stateless, 1:1, checksum-neutral prefix translation - what you use instead of NAT when you truly need prefix translation, and far less harmful to the end-to-end model. Its IOS XE support is image-dependent (concept + platform note here; DHCPv6-PD and general prefix are fully captured). - Prefer provider-independent addressing so you translate nothing; use NPTv6 only when you must. Next: [expert services troubleshooting - NAT, DHCP, and SLA ticket scenarios](https://www.pinglabz.com/expert-services-troubleshooting/). The full cluster index lives on the [IPv6 pillar](https://www.pinglabz.com/ipv6/) and the [IP services pillar](https://www.pinglabz.com/ip-services/). ### Embedded Packet Capture: Wireshark Inside Your Router URL: https://www.pinglabz.com/embedded-packet-capture-ios-xe/ Last updated: 2026-08-01T19:33:48.000Z You are troubleshooting a problem that only the packets can explain - a malformed handshake, an unexpected retransmission, a mystery drop. The traditional answer is to SPAN a port to a laptop running Wireshark. But the traffic is on a router in another building, or inside a tunnel, or you simply do not have a capture machine handy. Embedded Packet Capture (EPC) puts Wireshark-grade capture *inside the router itself* \- no SPAN, no external host, no cabling. This article covers EPC on IOS XE, with a real on-router capture from a CML lab, and the syntax difference that will bite you. For the fundamentals, see the [IP services pillar](https://www.pinglabz.com/ip-services/). ## What EPC is EPC captures packets traversing the router into an in-memory buffer, which you can then view on the CLI or export as a standard `.pcap` file to open in Wireshark. It captures at a defined point (an interface, a direction), filtered by an ACL so you grab only what you want, into a buffer you size. It is genuinely a packet analyser living in the router, and for anyone who has ever driven to a site just to plug in a capture laptop, it is a revelation. ## The syntax trap: classic vs modern Here is the thing that will waste your afternoon. IOS XE has **two** EPC syntaxes, and which one your platform accepts depends on the image. On a lot of virtual and older platforms - including the IOL-XE we used in the lab - the *modern* syntax is rejected: ``` R1(config)# monitor capture EPCAP interface Ethernet0/3 both ^ % Invalid input detected at '^' marker. ``` The modern one-liner (`monitor capture NAME interface X both`, then `match`, `start`, `stop`) is what most current documentation shows. But where it is not supported, you fall back to the **classic** buffer-and-point model, which is more verbose but works: ``` ! Classic EPC - four steps monitor capture buffer EPCBUF size 256 max-size 1518 linear monitor capture point ip cef EPCPOINT Ethernet0/3 both monitor capture point associate EPCPOINT EPCBUF monitor capture point start EPCPOINT ``` The classic model separates the **buffer** (where packets are stored - size, type) from the **capture point** (where and what is captured - the interface, direction, and switching path). You create both, associate them, and start the point. If your platform rejects `monitor capture NAME interface`, reach for this. Knowing both syntaxes is exactly the kind of platform-specific detail that separates a smooth troubleshoot from a frustrating one. ## A real capture With the capture point started, we generated traffic across the interface (a ping) and stopped the point: ``` R1# monitor capture point stop EPCPOINT R1# show monitor capture buffer EPCBUF 09:19:13.115 UTC Jul 12 2026 : IPv4 LES CEF : Et0/3 None 09:19:13.118 UTC Jul 12 2026 : IPv4 LES CEF : Et0/3 None 09:19:13.119 UTC Jul 12 2026 : IPv4 LES CEF : Et0/3 None 09:19:13.123 UTC Jul 12 2026 : IPv4 LES CEF : Et0/3 None ``` Four real packets, timestamped, captured on the router itself. The parameters view confirms the buffer state: ``` R1# show monitor capture buffer EPCBUF parameters Buffer Size : 262144 bytes, Max Element Size : 1518 bytes, Packets : 4 Associated Capture Points: Name : EPCPOINT, Status : Inactive ``` `Packets : 4` \- the buffer holds what we captured. For a detailed per-packet decode, `show monitor capture buffer EPCBUF detailed` gives you the full protocol breakdown, and `dump` gives you the raw hex. ## Filtering: capture only what you need An unfiltered capture on a busy interface fills the buffer with noise in milliseconds. Always filter with an ACL that matches only the conversation you are investigating: ``` ip access-list extended CAP-FILTER permit ip host 10.7.7.10 host 10.7.7.1 permit ip host 10.7.7.1 host 10.7.7.10 ! ! Classic: apply the ACL to the capture point monitor capture point ip cef EPCPOINT Ethernet0/3 both ! ...with the buffer filtered by the ACL association ``` Match *both* directions of the conversation (source-to-dest and dest-to-source), or you only see half the exchange and the capture is useless for diagnosing a handshake. This is the most common EPC mistake - a one-directional filter that captures the SYN but not the SYN-ACK. ## Buffer type: linear vs circular linear Captures until the buffer is full, then **stops**. Good for catching the *start* of an event - the first packets are preserved. The lab used linear. circular Overwrites the oldest packets when full, keeping the **most recent**. Good for catching an event you are waiting for - leave it running and stop it when the event fires. Choose by what you are hunting. Investigating a connection setup? Linear, so you keep the first packets. Waiting for an intermittent failure? Circular, so when it happens the recent packets (including the failure) are still there. A linear buffer that fills before your event is the second most common EPC mistake. ## Exporting to Wireshark The CLI view is fine for a quick look, but for real analysis you export the buffer as a pcap and open it in Wireshark or tshark: ``` R1# monitor capture buffer EPCBUF export tftp://10.7.7.200/capture.pcap ``` Pull that pcap to your analysis host (or the lab's Linux VM) and you have a standard capture file - full decode, filters, follow-stream, everything Wireshark does. The router did the capture; your workstation does the analysis. This is the workflow that makes EPC genuinely powerful: capture at the exact point in the network where the problem lives, analyse with the best tool. What you do with the file once it is on your workstation is a separate skill from getting it there, and it is the part that decides whether the capture was worth taking. The [guides on reading a capture to find the anomaly](https://www.pinglabz.com/packet-analysis/) pick up exactly where this export leaves off. ## Operational cautions - **EPC costs CPU and memory.** The router is copying packets to a buffer. On a busy production router, a broad capture can hurt. Filter tightly and size the buffer sensibly. - **Stop and remove captures when done.** A forgotten running capture keeps consuming resources. `monitor capture point stop` and remove the buffer. - **It captures control-plane and punted traffic well,** but hardware-forwarded traffic on some platforms may need the capture at a specific point. Know your platform's forwarding path. - **The classic-vs-modern syntax** is image-dependent. If one is rejected, try the other before concluding EPC is unavailable. ## EPC vs SPAN vs debug **EPC** Capture inside the router, no external host, exportable to pcap. Best for remote sites and tunnelled traffic. Costs router resources. **SPAN** Mirror a port to a capture host. Best when you have a host on-site and want zero router overhead. Needs the cabling and the host. **debug** Protocol-level events, not raw packets. Best for "what is the protocol thinking", dangerous on a busy box. Not a substitute for a real capture. There is a fourth option that table leaves out, because it does not run on the router at all. If there is a Linux box sitting on the segment you care about, [taking the same capture from the host side](https://www.pinglabz.com/tcpdump-for-network-engineers/) costs the router nothing, and the BPF filter syntax maps closely enough onto the ACL logic above that you can move between the two without relearning anything. ## Key takeaways - EPC captures packets **inside the router** into a buffer, viewable on the CLI and exportable as pcap for Wireshark - no SPAN, no external host. - Two syntaxes exist: the modern `monitor capture NAME interface` and the **classic** buffer/point model. Virtual and older platforms (like the lab's IOL-XE) reject the modern one - fall back to classic. - Verified: a real 4-packet capture on the router, with `show monitor capture buffer` and its parameters. - **Filter with an ACL matching both directions** of the conversation, or you capture half a handshake. - **Linear** buffer preserves the start of an event; **circular** keeps the most recent - choose by what you are hunting. - Export with `monitor capture buffer X export tftp://...` and analyse in Wireshark. Capture at the problem, analyse with the best tool. - EPC costs router CPU/memory - filter tightly and stop captures when done. Next: [IPv6 services closure - NPTv6, DHCPv6-PD, and general prefix](https://www.pinglabz.com/ipv6-nptv6-dhcpv6-pd/). The full cluster index lives on the [IP services pillar](https://www.pinglabz.com/ip-services/), and EPC's packet-analysis kinship links it to the [Nmap](https://www.pinglabz.com/nmap/) and [Ping](https://www.pinglabz.com/ping/) clusters. ### SNMPv3 in Production: Users, Groups, Views, Traps, and Informs URL: https://www.pinglabz.com/snmpv3-traps-informs/ Last updated: 2026-07-12T09:34:24.000Z SNMPv1 and v2c send everything - including your community string, which is effectively a password - in cleartext across the network. Anyone with a packet capture has your read (or write) access. SNMPv3 fixes this with real authentication and encryption, but it does so through a configuration model of users, groups, and views that is genuinely more involved than "snmp-server community public RO". The extra effort is the price of not broadcasting your monitoring credentials. This article builds SNMPv3 in production form and - the payoff - shows a real encrypted trap arriving and being decrypted on a Linux receiver. For the fundamentals, see the [IP services pillar](https://www.pinglabz.com/ip-services/). ## The three security levels SNMPv3 offers three levels, and you should use the strongest: noAuthNoPriv No authentication, no encryption. A username with no security. Barely better than v2c. Do not use. authNoPriv Authentication (the message is verified as coming from a known user, unaltered) but no encryption - the data is still in cleartext. Better, but the OID values are readable. authPriv Authentication *and* privacy (encryption). The whole message is authenticated and encrypted. This is what you use. Anything less leaks data. ## The configuration model: view, group, user SNMPv3 is built bottom-up from three objects, and understanding the layering is understanding the config: 1. A **view** defines *what OIDs* can be accessed - a subtree of the MIB tree, included or excluded. 2. A **group** ties a security level to a view - "users at this level can read this view". 3. A **user** belongs to a group and carries the actual credentials (auth password, priv password). ``` ! 1. The view - what OIDs are visible snmp-server view PLVIEW iso included snmp-server view PLVIEW internet included ! 2. The group - security level + view snmp-server group PLGRP v3 priv read PLVIEW ! 3. The user - credentials, member of the group snmp-server user pladmin PLGRP v3 auth sha PlAuthPass123 priv aes 128 PlPrivPass123 ``` Read that user line carefully - it is the crux. `auth sha PlAuthPass123` sets the authentication (SHA, with its password); `priv aes 128 PlPrivPass123` sets the encryption (AES-128, with its password). Two separate secrets, two separate purposes. Use **SHA** for auth (not MD5, which is deprecated) and **AES** for privacy (not DES, which is broken). ### Verification, on the router ``` R1#show snmp user User name: pladmin Engine ID: 800000090300AABBCC005F00 Authentication Protocol: SHA Privacy Protocol: AES128 Group-name: PLGRP R1#show snmp group groupname: PLGRP security model:v3 priv readview : PLVIEW notifyview: *tv.FFFFFFFF... R1#show snmp view | include PLVIEW PLVIEW iso - included nonvolatile active PLVIEW internet - included nonvolatile active ``` Note the **Engine ID** in `show snmp user` \- `800000090300AABBCC005F00`. Remember it, because it is the single most important value when you configure the trap receiver, and the thing people forget. ## Views: why you restrict what is visible A view is a security control. A monitoring user should be able to read the interface and system MIBs but not, say, the running configuration or sensitive tables. By building a view that includes only what the user needs, you limit the blast radius if the credentials leak. ``` ! A restrictive view - allow the standard MIBs, deny a sensitive subtree snmp-server view LIMITED internet included snmp-server view LIMITED 1.3.6.1.4.1.9.9.99 excluded ! exclude a specific Cisco subtree ``` The default deny is implicit: anything not explicitly `included` in the view is invisible. This is exactly the discipline you want - grant the minimum, and a compromised monitoring account cannot walk your entire MIB tree. ## Traps vs informs: the reliability difference Both send an event notification to a receiver. The difference is acknowledgement: Trap **Fire and forget.** The router sends the notification and never checks whether it arrived. Low overhead, but a trap lost in transit is gone with no record. Inform **Acknowledged.** The receiver replies; if no acknowledgement comes, the router retransmits. Reliable, at the cost of state and retransmission traffic. ``` snmp-server enable traps snmp-server host 192.168.99.100 version 3 priv pladmin udp-port 10162 ! trap snmp-server host 192.168.99.100 informs version 3 priv pladmin udp-port 10162 ! inform ``` `show snmp host` confirms both: ``` R1#show snmp host Notification host: 192.168.99.100 udp-port: 10162 type: inform user: pladmin security model: v3 priv Notification host: 192.168.99.100 udp-port: 10162 type: trap user: pladmin security model: v3 priv ``` **Use informs for events you cannot afford to lose** \- a link failure, an environmental alarm, anything that drives an operational response. Use traps for high-volume, less-critical telemetry where the occasional loss does not matter. The reliability of informs costs the router memory (it holds the state until acknowledged) and generates retransmissions, so it is a deliberate trade, not a default-everything choice. ## The payoff: a real encrypted trap, decrypted on Linux Configuration is one thing; proving the encryption works end to end is another. On a Linux trap receiver (net-snmp `snmptrapd`), the SNMPv3 user must be recreated with the **router's engine ID** \- this is the step everyone misses: ``` # /etc/snmp/snmptrapd.conf createUser -e 0x800000090300AABBCC005F00 pladmin SHA "PlAuthPass123" AES "PlPrivPass123" authUser log,execute,net pladmin priv ``` The `-e 0x800000090300AABBCC005F00` is the router's engine ID from `show snmp user`. Without the matching engine ID, the receiver cannot verify the authentication and the trap is dropped as unauthenticated. This is the number-one reason "my SNMPv3 traps are not arriving" - the config is fine, but the engine IDs do not match. Trigger an event (a loopback flap generates a linkUp), and the receiver logs the fully-decrypted trap: ``` 2026-07-12 02:25:52 [UDP: [192.168.99.1]:59301->[192.168.99.100]:10162]: SNMPv2-MIB::sysUpTime.0 = Timeticks: (72360) 0:12:03.60 SNMPv2-MIB::snmpTrapOID.0 = OID: IF-MIB::linkUp IF-MIB::ifIndex.7 = INTEGER: 7 IF-MIB::ifDescr.7 = STRING: Loopback99 IF-MIB::ifType.7 = INTEGER: softwareLoopback(24) enterprises.9.2.2.1.1.20.7 = STRING: "up" ``` That output is the whole point: a message that crossed the network **authenticated with SHA and encrypted with AES-128**, arriving intact and readable only because the receiver held the right credentials and the matching engine ID. A packet capture of that same trap on the wire would show ciphertext. That is SNMPv3 authPriv working - and it is the difference between monitoring you can trust and monitoring that leaks. ## Deployment notes - **Engine ID matching is everything for informs and remote users.** The receiver's `createUser` (or equivalent) must use the router's engine ID. Mismatch = silent drop. - **Port 162 is the default,** but it is often already bound (by an existing NMS or a system socket). The lab used udp-port 10162 to avoid a conflict - a real-world nuisance worth knowing. Match the port on both ends. - **SHA + AES, never MD5 + DES.** The old algorithms are deprecated or broken. If your gear or NMS only supports MD5/DES, that is a reason to upgrade, not a reason to use them. - **Restrict the view.** A monitoring user does not need write access or the whole MIB tree. ## Key takeaways - SNMPv3 replaces v2c's cleartext community strings with real authentication and encryption. Use **authPriv** \- anything less leaks data. - The model is layered: a **view** (which OIDs) is tied to a **group** (security level), which a **user** (credentials) belongs to. - The user carries two secrets: `auth sha ` and `priv aes 128 `. Use SHA and AES, not MD5 and DES. - Restrict the view to the minimum OIDs a monitoring user needs - the default deny limits the damage if credentials leak. - **Traps are fire-and-forget; informs are acknowledged and retransmitted.** Use informs for events you cannot afford to lose. - We proved it end to end: a SHA-authenticated, AES-128-encrypted trap decrypted on a Linux receiver. The receiver's user must be created with the **router's engine ID** \- the step everyone forgets, and the number-one cause of "my traps are not arriving". Next: [Embedded Packet Capture - Wireshark inside your router](https://www.pinglabz.com/embedded-packet-capture-ios-xe/). The full cluster index lives on the [IP services pillar](https://www.pinglabz.com/ip-services/). ### The IOS DHCP Server in Depth: Options, Classes, and Manual Bindings URL: https://www.pinglabz.com/ios-dhcp-server-in-depth/ Last updated: 2026-07-12T09:34:24.000Z Most people configure an IOS DHCP server once, with three lines - network, default-router, dns-server - and never touch it again. That is fine until you need a switch to PXE-boot from a TFTP server, a phone to find its call manager, or a specific device to always get the same address. Then you discover the IOS DHCP server is a genuinely capable service with options, classes, and manual bindings that most engineers never use. This article covers the IOS DHCP server in depth, with real lease output from a CML lab. For the fundamentals, see the [IP services pillar](https://www.pinglabz.com/ip-services/). ## The basic pool, and what it is actually doing ``` ip dhcp excluded-address 10.7.7.1 10.7.7.9 ! ip dhcp pool CLIENTS network 10.7.7.0 255.255.255.0 default-router 10.7.7.1 dns-server 10.7.7.1 option 150 ip 10.7.7.200 lease 0 0 10 ``` From the lab, a client (R2) requested and got a lease: ``` R1#show ip dhcp binding IP address Client-ID/Hardware address Lease expiration Type State Interface 10.7.7.10 0063.6973.636f... Jul 12 2026 09:27 AM Automatic Active Ethernet0/3 R1#show ip dhcp pool CLIENTS Pool CLIENTS : Total addresses : 254 Leased addresses : 1 Excluded addresses : 9 10.7.7.11 10.7.7.1 - 10.7.7.254 1 / 9 / 254 ``` `show ip dhcp pool` is the operational view you want - total, leased, and excluded counts, plus the current index (where the next lease will come from). The `1 / 9 / 254` reads leased / excluded / total. When a pool "runs out", this is where you see it. ### The excluded-address is not optional discipline `ip dhcp excluded-address` is easy to skip and painful to omit. The DHCP server will happily hand out the gateway's own address, or a static server's address, to a client - and then you have a duplicate-IP conflict that is maddening to diagnose. **Always exclude the addresses you have assigned statically**, including the gateway, DNS, and any servers in the subnet. Exclude a small range at the bottom (here .1 to .9) as a reservation for infrastructure. ## DHCP options: the part that matters The three basic parameters (default-router, dns-server) are themselves DHCP options with friendly IOS keywords. But many services need *other* options, and you configure those by number: **Option 150** TFTP server address (Cisco). Phones and switches use it to find their config/image server. `option 150 ip 10.7.7.200`. **Option 66** TFTP server name (the standard equivalent of 150). `option 66 ascii tftp.example.com`. **Option 43** Vendor-specific - used by Cisco APs to find a wireless controller, and by many vendors for bootstrap. Encoded as hex or type-length-value. **Option 82** Relay agent information (circuit/remote ID) - added by a relay, and the basis for [DHCP snooping](https://www.pinglabz.com/dhcp-snooping-in-depth/) and class matching. Option 150 in the lab (`option 150 ip 10.7.7.200`) is the classic case: it hands every client a TFTP server address, which is how a Cisco phone finds its call manager or a switch finds its boot image. When a device "boots but never finds its config", a missing or wrong option 150/66 is the usual cause. ## DHCP classes: different options for different devices A single subnet often has different device types that need different treatment - phones need option 150, PCs do not; a specific vendor's IoT needs option 43\. DHCP classes let one pool serve them differently, matching on something in the request (commonly option 82's circuit-id, or the vendor class identifier option 60): ``` ip dhcp class PHONES option 60 hex 436973636f ! match vendor-class "Cisco" ! ip dhcp pool ACCESS network 10.7.7.0 255.255.255.0 class PHONES address range 10.7.7.100 10.7.7.150 class default address range 10.7.7.10 10.7.7.99 ``` Now phones (identifying themselves via option 60) get addresses from one range, everything else from another - within the same subnet, from one pool. Classes are how you apply per-device-type policy without carving up subnets or running multiple pools per VLAN. ## Manual bindings: a fixed address by identity When a device must *always* get the same address - a printer, a server, a piece of network gear you manage by DHCP - you create a manual binding tied to its client identifier or MAC: ``` ip dhcp pool PRINTER-BINDING host 10.7.7.50 255.255.255.0 client-identifier 0100.1122.3344.55 ! or hardware-address default-router 10.7.7.1 ``` This is a DHCP reservation done the IOS way - a dedicated single-host pool matched to the client's identity. The device uses DHCP (so it is managed centrally, its options come from the server) but always receives 10.7.7.50\. It combines the convenience of DHCP with the predictability of a static address, which is exactly what you want for infrastructure devices. The subtlety: match on the **client-identifier** (option 61) if the client sends one, or the **hardware-address** if it does not. Cisco routers acting as clients send a client-identifier derived from the interface, which is why the lab binding above shows a long client-id rather than a plain MAC. Get the wrong one and the reservation never matches. ## Lease time: shorter than you think for churn The default lease is one day. That is fine for stable devices but wrong for high-churn environments - a guest network, a conference room, anywhere devices come and go. A too-long lease on a churny subnet exhausts the pool with dead reservations for devices that left hours ago. ``` lease 0 0 10 ! days hours minutes - here 10 minutes lease 8 ! 8 days lease infinite ! never expires - only for truly static devices ``` The lab used a 10-minute lease to demonstrate churn behaviour. In production, match the lease to the dwell time: a day for offices, an hour or less for guest and public networks. `lease infinite` is a trap - it defeats the point of DHCP and slowly fills the pool; use a manual binding instead if you want a device to keep an address permanently. ## VRF-aware DHCP The IOS DHCP server is VRF-aware, which matters when you serve overlapping tenants (as in the [VRF-aware NAT](https://www.pinglabz.com/vrf-aware-nat-ios-xe/) scenario). A pool bound to a VRF serves clients in that VRF's address space: ``` ip dhcp pool RED-POOL vrf RED network 10.5.5.0 255.255.255.0 default-router 10.5.5.1 option 150 ip 10.5.5.200 ``` The `vrf RED` keyword scopes the pool. Two pools with the same network in different VRFs coexist happily - the DHCP server tracks bindings per VRF, so overlapping tenant DHCP works the same way overlapping tenant NAT does. ## Troubleshooting 1. **Clients not getting addresses?** Check `show ip dhcp pool` for exhaustion, `show ip dhcp conflict` for detected duplicates, and confirm the relay path if the server is not on the client's subnet. 2. **Device boots but finds no config/controller?** A missing or wrong option (150/66 for TFTP, 43 for a WLC). `show ip dhcp binding` confirms it got an address; the option is a separate matter. 3. **Duplicate IP conflict?** A statically-assigned address was not excluded. `show ip dhcp conflict`, then add `ip dhcp excluded-address`. 4. **A reservation not honoured?** Wrong match type - client-identifier vs hardware-address. Cisco clients send a client-id; match that. 5. **Pool exhausting on a busy subnet?** Lease too long. Shorten it to the device dwell time. ## Key takeaways - The IOS DHCP server does far more than the three basic lines. `show ip dhcp pool` (leased/excluded/total) and `show ip dhcp binding` are the operational views. - **Always exclude** statically-assigned addresses, or you get duplicate-IP conflicts. - **Options** drive the interesting behaviour: 150/66 (TFTP, for phones and boot images), 43 (vendor-specific, WLC discovery), 82 (relay info, basis for snooping and classes). - **Classes** let one pool serve different device types differently, matching on option 60/82. - **Manual bindings** give a device a fixed address by identity - a reservation the IOS way. Match on client-identifier for Cisco clients, hardware-address otherwise. - Set the **lease** to the device dwell time; short for guest/public, long for offices, `infinite` almost never. - The server is **VRF-aware** \- a `vrf` keyword scopes the pool, so overlapping tenants work. Next: [SNMPv3 in production - users, groups, views, traps, and informs](https://www.pinglabz.com/snmpv3-traps-informs/). The full cluster index lives on the [IP services pillar](https://www.pinglabz.com/ip-services/). ### VRF-Aware NAT on Cisco IOS XE URL: https://www.pinglabz.com/vrf-aware-nat-ios-xe/ Last updated: 2026-07-12T09:34:23.000Z Two customers connect to your router. Both, entirely by coincidence, use 10.5.5.0/24 internally. Both have a host at 10.5.5.50\. Both need to reach the same shared service out one exit interface. In a normal routing table this is impossible - you cannot have two 10.5.5.50s. VRF-aware NAT makes it not just possible but clean: each customer lives in its own VRF, and NAT translates each to a distinct public address, keeping the two identical private hosts completely separate. This article builds exactly that in a CML lab and shows the two translation tables holding identical inside-local addresses. It closes the services domain for the IP-services cluster. For the fundamentals, see the [IP services pillar](https://www.pinglabz.com/ip-services/). ## The problem: overlapping address space Overlapping addresses are the norm, not the exception, the moment you aggregate multiple independent networks: - A managed-services provider onboards customers who all use RFC1918 space and inevitably collide - everyone has a 10.0.0.0/8 or a 192.168.1.0/24. - A merger brings two companies together, both of which built their networks assuming they owned 10.0.0.0/8. - A multi-tenant environment where tenants must be isolated but reach shared services. VRFs solve the isolation - each tenant gets its own routing table, and two 10.5.5.0/24s in two VRFs never see each other. But the moment they need to reach a *shared* destination out a common interface, you have a problem: the shared side has one routing table, and it cannot distinguish two identical source addresses. That is where NAT comes in, and it has to be VRF-aware. ## The lab R1 is the shared gateway. It has two VRFs - RED (RD 65000:1) and BLUE (RD 65000:2) - each connected to a customer whose network is **10.5.5.0/24, with a host at 10.5.5.50 in both**. The exit interface toward the shared service (192.168.99.100) is in the global table. Each VRF NATs to its own public address. ``` ! Distinct global pools, one per VRF ip nat pool RED-POOL-NAT 192.168.99.201 192.168.99.201 prefix-length 24 ip nat pool BLUE-POOL-NAT 192.168.99.202 192.168.99.202 prefix-length 24 ! The VRF keyword is what makes NAT VRF-aware ip nat inside source list RED-ACL pool RED-POOL-NAT vrf RED overload ip nat inside source list BLUE-ACL pool BLUE-POOL-NAT vrf BLUE overload ! Interfaces interface Ethernet0/1 vrf forwarding RED ip nat inside interface Ethernet0/2 vrf forwarding BLUE ip nat inside interface Ethernet0/0 ip nat outside ``` **The `vrf RED` / `vrf BLUE` keyword on the NAT statement is the whole feature.** It tells NAT which VRF each translation belongs to, so the two identical inside-local addresses land in two separate translation tables and get two separate public addresses. Without it, the second 10.5.5.50 would collide with the first. ## Both customers, simultaneously, with identical hosts Both customer hosts source from 10.5.5.50 and reach the same shared service: ``` RED-CE#ping 192.168.99.100 source Loopback5 repeat 3 Packet sent with a source address of 10.5.5.50 !!! Success rate is 100 percent (3/3) BLUE-CE#ping 192.168.99.100 source Loopback5 repeat 3 Packet sent with a source address of 10.5.5.50 !!! Success rate is 100 percent (3/3) ``` Two hosts with the same IP, both working. And here is the proof - the two VRF translation tables, holding identical inside-local addresses translated to distinct global addresses: ``` R1#show ip nat translations vrf RED Pro Inside global Inside local Outside local Outside global icmp 192.168.99.201:1024 10.5.5.50:0 192.168.99.100:0 192.168.99.100:1024 R1#show ip nat translations vrf BLUE Pro Inside global Inside local Outside local Outside global icmp 192.168.99.202:1024 10.5.5.50:0 192.168.99.100:0 192.168.99.100:1024 ``` Read those two blocks side by side. The **inside local** is 10.5.5.50 in both - the same host address in each customer. The **inside global** is 192.168.99.201 for RED and 192.168.99.202 for BLUE - distinct public addresses. The shared service sees two different sources and can reply to each correctly. The VRF is what keeps the two 10.5.5.50 translations from colliding. That is VRF-aware NAT in one screen. ## The house quirk: ICMP translations expire fast When you verify VRF-aware NAT with a ping, you have a small window. ICMP translations time out in about a minute, and if you ping, wait, then run `show ip nat translations`, the entry may already be gone and you will think it failed. The fix is to check the table *promptly* after generating traffic - or, in a lab, run the ping and the show command in quick succession. For a persistent view, use a longer-lived protocol (TCP) or a sustained ping. ## Direction of translation: which side is which The mental model people get wrong: with VRF-aware NAT for overlapping customers, the **inside** interfaces are the customer VRF links (the overlapping addresses), and the **outside** interface is the shared exit. You are translating the customers' overlapping addresses *into* distinct addresses the shared side can use. So `ip nat inside` goes on the VRF interfaces, `ip nat outside` on the shared interface, and `ip nat inside source` translates on the way out. If the shared service needs to *initiate* to a customer host, that is a different NAT (a static `ip nat outside source static ... vrf` or a destination NAT), because now you are translating the shared side's view of a customer address on the way in. Most designs only need the outbound direction, which is what the lab shows. ## Where VRF-aware NAT fits Shared services access Multiple tenants with overlapping space all need to reach a shared DNS, NTP, or monitoring service. Each NATs to a distinct address at the shared boundary. Merger integration Two companies with colliding 10.0.0.0/8 estates put each in a VRF and NAT at the interconnect, buying time to renumber (or never renumbering). MPLS L3VPN extranet A provider offering shared-services access to customers whose VPNs overlap, doing the translation on the PE at the extranet boundary. ## Troubleshooting 1. **Overlapping hosts colliding?** The `vrf` keyword is missing on the `ip nat inside source` statement, or wrong. Without it, NAT treats both as the same and they collide. `show ip nat translations vrf X` for each VRF. 2. **Translation table empty right after a ping?** ICMP timeout - check promptly, or use TCP for a persistent entry. 3. **One VRF works, the other does not?** Check that both interfaces have `ip nat inside` in the right VRF, and that each has its own pool and ACL. Reusing one pool across two VRFs defeats the isolation. 4. **Return traffic failing?** The shared side needs a route back to the NAT pool addresses (192.168.99.201/202), and R1 needs routes in each VRF toward the customer prefixes. ## Key takeaways - VRF-aware NAT lets multiple tenants with **overlapping address space** reach a shared destination, by translating each VRF to a distinct global address at the boundary. - The `vrf RED` / `vrf BLUE` keyword on `ip nat inside source` is the entire feature - it keeps the two identical inside-local addresses in separate translation tables. - Verified: two hosts both at 10.5.5.50, in RED and BLUE, translated to 192.168.99.201 and 192.168.99.202, both reaching the same shared service simultaneously. - Inside = the overlapping customer VRF interfaces; outside = the shared exit. Translate on the way out. - ICMP translations expire fast - check the table promptly after generating traffic, or use TCP. - The classic use cases are shared-services access, merger integration, and MPLS L3VPN extranet with overlapping VPNs. Next: [the IOS DHCP server in depth - options, classes, and manual bindings](https://www.pinglabz.com/ios-dhcp-server-in-depth/). The full cluster index lives on the [IP services pillar](https://www.pinglabz.com/ip-services/). ### TLOC Extension: Sharing Transports Between Edge Routers URL: https://www.pinglabz.com/sd-wan-tloc-extension/ Last updated: 2026-07-12T09:12:16.000Z A dual-router branch has an obvious problem: you have two WAN Edge routers for redundancy, but you may only have one MPLS circuit. It terminates on one router. If that router fails, the other router - the redundant one - has no way to reach the MPLS transport, and half your redundancy is fiction. TLOC extension solves exactly this: it lets one router use a transport that is physically connected to the other. This article covers TLOC extension: the problem it solves, how it works, and how it fits a redundant branch design. It extends the [complete SD-WAN guide](https://www.pinglabz.com/sd-wan/). Command syntax is drawn from Cisco's current 20.x documentation, clearly labelled as a documented reference. ## The problem, precisely You build a dual-router branch for high availability. Router 1 and Router 2, both WAN Edges. But WAN circuits are expensive and often physically singular: - The **MPLS** circuit is one handoff from the provider - it terminates on Router 1. - The **internet** circuit is a different handoff - it terminates on Router 2. Now each router has only one transport. Router 1 has MPLS but no internet. Router 2 has internet but no MPLS. If Router 1 dies, you lose *all* MPLS connectivity at the branch, because the only router connected to MPLS is gone - even though Router 2 is alive and well. Your "redundant" branch has a single point of failure per transport. That is not redundancy; it is two half-connected routers. ## The solution: extend the TLOC TLOC extension connects the two routers with a link and lets each router reach the *other's* transport across it. Router 1's MPLS transport becomes usable by Router 2, and Router 2's internet transport becomes usable by Router 1\. Now each router has a **TLOC on both transports** \- one local, one extended through its partner. Recall from the [OMP deep dive](https://www.pinglabz.com/omp-deep-dive/) that a TLOC is a tunnel endpoint (system-ip + colour + encapsulation), and a router can have up to eight. TLOC extension gives a router an additional TLOC for a transport it is not physically connected to, reached via the partner router. Both routers now advertise MPLS and internet TLOCs, so the loss of either router still leaves the branch with both transports. ### How it is configured A physical link connects the two routers (the TLOC-extension interface). On Router 2, its interface toward Router 1 is configured to reach Router 1's MPLS transport: ``` ! On Router 2 - reach MPLS via Router 1 interface GigabitEthernet3 description TLOC-EXTENSION-TO-R1 ip address 10.99.1.2 255.255.255.252 tloc-extension GigabitEthernet1 ! R1's MPLS transport interface ! ``` The `tloc-extension` command binds Router 2's inter-router interface to Router 1's MPLS transport interface. Router 2 now builds its MPLS tunnels *through* Router 1's MPLS circuit, across the inter-router link. It has an MPLS TLOC without an MPLS circuit of its own. Router 1 is configured symmetrically to reach the internet transport via Router 2. ## Verifying it (documented reference) The proof is that each router now has TLOCs on both colours: ``` Router2# show sdwan omp tlocs | i mpls|biz-internet ipv4 10.0.0.12 mpls ipsec ... ! extended via R1 ipv4 10.0.0.12 biz-internet ipsec ... ! local Router2# show sdwan bfd sessions SYSTEM IP COLOR STATE ... 10.0.0.11 mpls up ! MPLS tunnel, built through R1 10.0.0.11 biz-internet up ! internet tunnel, local ``` Router 2 has a working MPLS BFD session despite having no MPLS circuit - the session runs through Router 1's MPLS transport via the extension link. That is the whole feature working: full dual-transport reachability on a router that is only physically connected to one transport. ## Where TLOC extension fits Dual-router, single-circuit-each branch The canonical use case. Two routers, MPLS on one and internet on the other, and you want each router to have both transports for true HA. TLOC extension is the answer. Not needed if... Each router already has its own connection to every transport (rare and expensive), or the branch is single-router (no partner to extend through). Then there is nothing to extend. The economic driver is real: TLOC extension lets you build a properly redundant dual-router branch without paying for duplicate circuits of every transport at every branch. You buy one MPLS and one internet handoff, terminate them on different routers, and extend - and the branch survives the loss of either router with both transports intact. ## Design considerations 1. **The inter-router link is now critical.** It carries the extended transport's traffic. If it fails, the extension breaks and each router falls back to only its local transport. Make it robust - a direct connection, ideally with its own redundancy if the branch is important enough. 2. **Capacity.** When Router 2 sends MPLS traffic through Router 1's circuit, that traffic crosses the inter-router link and shares Router 1's MPLS bandwidth. Size the inter-router link and account for the shared circuit capacity under a failure. 3. **The extension carries the transport, not the service.** TLOC extension is about the transport (VPN 0) side. The service-side (LAN) redundancy - HSRP/VRRP for the endpoints' gateway - is a separate design concern layered on top. 4. **Failure behaviour.** Trace what happens when each component fails: Router 1 down (Router 2 keeps internet locally, loses MPLS entirely), the extension link down (each router falls to its local transport), Router 1's MPLS circuit down (both routers lose MPLS, since Router 2's MPLS was via Router 1). Know these before you deploy. ## Key takeaways - TLOC extension lets one WAN Edge use a transport that is physically connected to its **partner router**, via a link between them. - It solves the dual-router branch problem: two routers for HA, but each transport terminates on only one router, so losing a router loses a transport. Extension gives each router a TLOC on *both* transports. - Configured with the `tloc-extension` command on the inter-router interface, binding it to the partner's transport interface. The router then builds that transport's tunnels through its partner. - Verify with `show sdwan omp tlocs` and `show sdwan bfd sessions` \- a working tunnel on a colour the router has no local circuit for is the proof. - It delivers true dual-router, dual-transport redundancy **without duplicating every circuit at every branch** \- the economic win. - Design carefully: the inter-router link becomes critical and shares the extended circuit's bandwidth; trace every failure case before deploying. That closes the advanced SD-WAN series, and with it the CCIE Software-Defined Infrastructure domain across both SD-Access and SD-WAN. The full cluster index lives on the [SD-WAN pillar](https://www.pinglabz.com/sd-wan/). ### Direct Internet Access (DIA) in Catalyst SD-WAN URL: https://www.pinglabz.com/sd-wan-direct-internet-access-dia/ Last updated: 2026-07-12T09:12:15.000Z For years, branch internet traffic took a ridiculous journey: from the branch, across the WAN, to a central data centre, out through the corporate firewall, to the internet - and all the way back. Every Office 365 request, every YouTube video, every SaaS API call, backhauled hundreds of miles to reach a cloud that was often closer to the branch than the data centre was. Direct Internet Access ends that. It lets branch internet traffic exit locally, straight out the branch's own internet circuit. This article covers DIA in Catalyst SD-WAN: what it is, how it is configured as a data policy, and the security question it forces. It extends the [complete SD-WAN guide](https://www.pinglabz.com/sd-wan/). Command syntax is drawn from Cisco's current 20.x documentation, clearly labelled as a documented reference. ## The problem: backhaul In a traditional hub-and-spoke WAN, branches have no local internet breakout - or if they do, it is not trusted. So *all* internet-bound traffic is sent across the WAN to a central site, inspected by the central security stack, and forwarded to the internet. The return path reverses it. This made sense when internet traffic was a small fraction of the total and the applications lived in the data centre. It makes no sense now, when the majority of branch traffic is cloud and SaaS. Backhaul in that world means: - **Wasted WAN bandwidth** \- your expensive MPLS carries internet traffic twice (there and back) that never needed to be on it. - **Terrible latency to cloud apps** \- a branch in one city reaches a CDN node in that same city by going to a data centre three states away and back. - **A central chokepoint** \- every branch's internet traffic funnels through one security stack that has to be sized for all of it. ## The solution: local breakout DIA sends internet-destined traffic **directly out the branch's local internet transport**, with NAT, without touching the WAN. The branch reaches the cloud by the shortest path - its own local circuit - and the corporate WAN only carries traffic that genuinely needs to reach the corporate network. In Catalyst SD-WAN, DIA is expressed as a **data policy action** (which is why it lives in the same framework as [AAR](https://www.pinglabz.com/sd-wan-data-policy-aar/)). A centralised data policy matches internet-bound traffic and applies a NAT action that sends it out VPN 0 (the transport VPN): ``` policy data-policy DIA-POLICY vpn-list SERVICE-VPNS sequence 10 match source-data-prefix-list BRANCH-SUBNETS destination-data-prefix-list INTERNET ! 0.0.0.0/0 minus corporate ! action accept nat use-vpn 0 ! NAT and exit locally nat fallback ! if local internet fails, backhaul ! ! default-action accept ! ! ``` The key action is `nat use-vpn 0` \- "NAT this flow and send it out the transport VPN's internet interface locally". The `nat fallback` option is the safety net: if the local internet circuit fails, traffic falls back to the traditional backhaul path so the branch is not cut off from the internet entirely. ## The interface side The transport interface needs NAT enabled so it can translate the branch's private addresses to its public internet address: ``` sdwan interface GigabitEthernet1 tunnel-interface encapsulation ipsec color biz-internet ! ! ! interface GigabitEthernet1 ip nat outside ! ip nat inside source list ... interface GigabitEthernet1 overload ``` Verification is standard NAT plus the SD-WAN policy view (documented reference): ``` Edge# show sdwan policy from-vsmart data-policy Edge# show ip nat translations Pro Inside global Inside local Outside local Outside global tcp 203.0.113.5:1044 10.1.10.50:1044 140.82.113.3:443 140.82.113.3:443 ``` That translation - a branch host (10.1.10.50) reaching a public address (140.82.113.3:443) via the branch's own public IP (203.0.113.5) - is DIA working: the flow went straight out the local circuit, not across the WAN. ## The security question DIA forces Here is what DIA gives you and what it takes away. It gives you local breakout - efficient, low-latency internet. It takes away the central security stack that every packet used to pass through. A branch now has traffic going straight to the internet without visiting the corporate firewall. **So what protects it?** This is not optional to answer. The moment you enable DIA, every branch is an internet edge, and it needs internet-edge security. Catalyst SD-WAN offers several answers, and a real deployment uses one: On-box security The WAN Edge runs Cisco's integrated security - enterprise firewall, IPS (Snort), URL filtering, and Advanced Malware Protection - inspecting DIA traffic locally on the router itself. Cloud-delivered security (SIG) A Secure Internet Gateway - the branch tunnels DIA traffic to a cloud security service (Umbrella, Zscaler, etc.) for inspection before it reaches the internet. The SASE model. The industry has largely moved toward the cloud-delivered (SIG/SASE) model: rather than run a full security stack on every branch router, tunnel DIA traffic to a cloud gateway that provides consistent, always-updated security everywhere. But on-box security is a valid choice for branches that need local inspection or cannot depend on a cloud service. **The one wrong answer is "none"** \- DIA without a security plan is a branch full of unprotected internet edges. ## DIA design decisions 1. **What is "internet"?** Your DIA match must precisely separate internet traffic (breaks out locally) from corporate traffic (goes over the WAN). Usually "everything except the corporate prefix-list". Get this wrong and either corporate traffic leaks to the internet or internet traffic backhauls unnecessarily. 2. **Which SaaS gets special treatment?** Cloud OnRamp for SaaS measures the path to specific SaaS providers (Office 365, etc.) and can steer to the best-performing exit. Worth it for the SaaS you depend on. 3. **Fallback behaviour.** `nat fallback` keeps the branch on the internet if the local circuit dies, by reverting to backhaul. Decide whether you want it (usually yes). 4. **The security model.** On-box vs cloud (SIG). Decide before you deploy DIA, not after - it shapes the whole design. ## Key takeaways - DIA sends branch internet traffic **directly out the local internet circuit** with NAT, instead of backhauling it across the WAN to a central site. Efficient, low-latency, and it keeps expensive WAN bandwidth for corporate traffic. - It is configured as a **data policy action**: match internet-bound traffic, apply `nat use-vpn 0` to exit locally, with `nat fallback` as the safety net if the local circuit fails. - Verify with `show ip nat translations` \- a branch host reaching a public address via the branch's own public IP is DIA working. - DIA **removes the central security chokepoint**, so every branch becomes an internet edge that needs internet-edge security: on-box (integrated firewall/IPS/URL/AMP) or cloud-delivered (SIG/SASE). "None" is the one wrong answer. - Design carefully: define "internet" precisely, consider SaaS onramp, choose the fallback behaviour, and decide the security model before deploying. Next: [TLOC extension - sharing transports between edge routers](https://www.pinglabz.com/sd-wan-tloc-extension/). The full cluster index lives on the [SD-WAN pillar](https://www.pinglabz.com/sd-wan/). ### SD-WAN Localized Policy: ACLs and Route Policies at the Edge URL: https://www.pinglabz.com/sd-wan-localized-policy/ Last updated: 2026-07-12T09:12:15.000Z Centralised policy shapes the whole overlay from the Controller. But some things belong on the device itself - the ACL on a specific interface, the QoS treatment as packets leave a port, the route-map filtering what a branch redistributes into OMP. That is localised policy, and while it gets less attention than the centralised kind, forgetting it is how you end up with a beautifully engineered overlay that still drops the wrong packets at the edge. This article covers localised policy: access lists, QoS, and route policy applied on the WAN Edge. It extends the [complete SD-WAN guide](https://www.pinglabz.com/sd-wan/). Command syntax is drawn from Cisco's current 20.x documentation, clearly labelled as a documented reference. ## Centralised vs localised: the division of labour Centralised policy Authored on the Controller, affects the overlay as a whole. Control policy (routes/topology) and data policy (flow steering, AAR). One document, many devices. Localised policy Configured per device (via template/config-group), affects that device's own interfaces and forwarding. ACLs, QoS, route policy, mirroring. The per-box treatment. The rule of thumb: **if it is about the overlay's shape or which flows go where across the fabric, it is centralised. If it is about how a specific interface on a specific device treats packets, it is localised.** An interface ACL, a queueing policy on an egress port, or filtering the routes a branch redistributes into OMP - all localised. ## Localised ACLs An access list applied to an interface on the WAN Edge - exactly what you would expect, in SD-WAN's policy syntax: ``` policy access-list BLOCK-GUEST-TO-CORP sequence 10 match source-data-prefix-list GUEST-SUBNET destination-data-prefix-list CORP-SUBNET ! action drop count guest-to-corp-drops ! counter for visibility ! ! default-action accept ! ! ``` Applied to the interface, in a direction: ``` sdwan interface GigabitEthernet2 access-list BLOCK-GUEST-TO-CORP in ! ``` These are for local, per-interface filtering - the kind of thing that does not belong in a fabric-wide data policy because it is specific to one site's topology. A guest subnet blocked from the corporate subnet at the branch, a management-plane ACL, an anti-spoofing filter on a WAN interface. ## Localised QoS QoS is inherently local - it is about how *this* interface queues and schedules packets as they leave. Centralised policy can *mark* traffic (set DSCP), but the actual queueing, shaping, and scheduling is localised policy on the egress interface: ``` policy class-map VOICE match dscp ef ! class-map CRITICAL match dscp af41 ! ! policy qos-scheduler VOICE-Q class VOICE bandwidth-percent 20 scheduling llq ! low-latency queue for voice drops tail-drop ! qos-scheduler CRITICAL-Q class CRITICAL bandwidth-percent 30 scheduling wrr ! qos-map BRANCH-QOS ... ! ! ``` Applied to the WAN egress interface, this is what actually protects voice during congestion. The division is important: **a centralised data policy might mark voice EF; the localised QoS map is what puts EF-marked traffic in the low-latency queue.** Marking without queueing does nothing under congestion; queueing without marking has nothing to prioritise. You need both. Note that on the internet transport, your QoS markings are honoured only up to the point packets leave your edge - the internet ignores DSCP. So localised QoS matters most on the *egress* shaping toward a congested WAN link, protecting your own outbound queue, which is exactly where it can help. ## Localised route policy Route policy on the edge controls the interaction between the service-side routing (BGP/OSPF/EIGRP to the LAN) and OMP. This is where you filter, tag, or set attributes on routes as they cross between the two: ``` policy route-policy REDISTRIBUTE-FILTER sequence 10 match address CORPORATE-PREFIXES ! action accept set omp-tag 100 ! tag routes redistributed into OMP ! ! sequence 20 action reject ! do not leak anything else into OMP ! ! ! ``` Applied to the redistribution from the service-side protocol into OMP (and vice versa), localised route policy controls exactly which of a branch's local routes become visible across the whole fabric. Get this wrong and either a branch leaks routes it should not (polluting the overlay) or fails to advertise routes it should (breaking reachability). It is the SD-WAN equivalent of a redistribution route-map, and the same discipline applies - tag on the way in, filter on the way out, never redistribute without a policy. ## Verifying localised policy (documented reference) ``` Edge# show sdwan policy access-list-counters NAME COUNTER NAME PACKETS BYTES BLOCK-GUEST-TO-CORP guest-to-corp-drops 1043 156450 Edge# show sdwan policy access-list-associations Edge# show policy-map interface GigabitEthernet1 ! QoS, standard IOS XE ``` The ACL counter is the practical verification - it proves the ACL is not just configured but actually matching traffic. For QoS, the standard IOS XE `show policy-map interface` shows the queue statistics, because the localised QoS ultimately renders to ordinary MQC on the interface. ## Where each policy type lives - the complete map **Overlay topology, path preference** Centralised control policy (Controller) **Flow steering, AAR, DIA** Centralised data policy (Controller) **Per-interface ACL** Localised policy (edge) **Queueing / shaping on egress** Localised policy (edge) **Service-side route filtering into OMP** Localised route policy (edge) Keep this map in your head and SD-WAN policy stops being confusing. Every requirement lands in exactly one of these boxes. "I want branches to prefer MPLS" is centralised control. "I want voice steered by SLA" is centralised data. "I want this port to drop guest traffic" is localised. The mistake is trying to solve a local problem with a fabric-wide policy, or vice versa. ## Key takeaways - Localised policy is configured per device and affects that device's own interfaces and forwarding: ACLs, QoS, and route policy. Centralised policy shapes the overlay; localised policy handles the per-box treatment. - **ACLs** filter per interface - the local, site-specific filtering that does not belong in a fabric-wide data policy. - **QoS is inherently localised**: centralised policy can mark DSCP, but the queueing and scheduling that protects voice under congestion is localised policy on the egress interface. You need both marking and queueing. - **Route policy** on the edge controls what local (BGP/OSPF/EIGRP) routes cross into OMP and how they are tagged - the SD-WAN redistribution discipline. - Verify with `show sdwan policy access-list-counters` (proves the ACL matches) and `show policy-map interface` (QoS queue stats). - Every SD-WAN policy requirement maps to exactly one type: control (topology), data (flows), or localised (per-device). Match the requirement to the right box. Next: [Direct Internet Access (DIA) in Catalyst SD-WAN](https://www.pinglabz.com/sd-wan-direct-internet-access-dia/). The full cluster index lives on the [SD-WAN pillar](https://www.pinglabz.com/sd-wan/). ### SD-WAN Data Policy and Application-Aware Routing (AAR) URL: https://www.pinglabz.com/sd-wan-data-policy-aar/ Last updated: 2026-07-12T09:12:14.000Z This is the feature people actually buy SD-WAN for. Application-aware routing (AAR) measures the real-time quality of every transport - loss, latency, jitter - and steers each application over the path that currently meets its needs. Voice over the low-latency link, backup over the cheap one, and if the good link degrades, voice moves automatically before users notice. No traditional WAN can do that. This article covers centralised data policy and AAR: how traffic is matched, how SLA classes work, and how the fabric measures paths. It extends the [complete SD-WAN guide](https://www.pinglabz.com/sd-wan/). Policy and command syntax is drawn from Cisco's current 20.x documentation, clearly labelled as a documented reference. ## Data policy vs control policy The [control policy](https://www.pinglabz.com/sd-wan-centralized-control-policy/) shaped what routes exist and where they point - the map. Data policy operates on the map: it matches **actual traffic flows** and decides what happens to each one. Both are centralised (authored on the Controller and pushed to edges), but data policy is about packets, not routes. Data policy can do several things - permit/deny (a distributed firewall), set DSCP, do NAT for direct internet access, and - the headline - **application-aware routing**, steering flows onto paths that meet a service-level agreement. ## How AAR works: SLA classes and BFD measurement The foundation is measurement. Every data-plane tunnel runs **BFD**, and SD-WAN uses it not just for liveness but to continuously measure each path's **loss, latency, and jitter**. Those measurements are the raw material AAR acts on. You define **SLA classes** \- named thresholds an application requires: ``` policy sla-class VOICE-SLA loss 1 ! max 1% loss latency 150 ! max 150 ms jitter 30 ! max 30 ms ! sla-class BULK-SLA loss 10 latency 300 ! ``` Then an app-route policy matches traffic and binds it to an SLA class plus a preferred colour: ``` app-route-policy STEER-APPS vpn-list SERVICE-VPNS sequence 10 match app-list VOICE-APPS ! e.g. RTP, SIP ! action sla-class VOICE-SLA preferred-color mpls ! if MPLS meets VOICE-SLA, use it; if not, use any path that does ! ! sequence 20 match app-list BACKUP-APPS ! action sla-class BULK-SLA preferred-color biz-internet ! ! ! ``` The behaviour: **use the preferred colour if it currently meets the SLA; if it does not, use any path that does; if none meets it, fall back per your policy** (strict drop, or use the best available). The edge is constantly comparing live BFD measurements against the SLA thresholds and moving flows accordingly - per-application, in real time. ## Why this is impossible on a traditional WAN A traditional WAN routes by destination prefix and a static metric. It has no concept of "this application needs under 150 ms and this path is currently at 200 ms". It cannot measure jitter per-tunnel, it cannot classify by application, and it certainly cannot move a flow mid-session because a link degraded. AAR does all three, continuously. That capability - not the automation, not the templates - is the reason SD-WAN displaced traditional WAN for anyone carrying voice or video over mixed transports. ## Verifying AAR on the edge (documented reference) ### The live path measurements ``` Edge# show sdwan app-route stats MEAN MEAN MEAN REMOTE-TLOC COLOR LOSS% LATENCY JITTER SLA-CLASS ------------------------------------------------------------------ 10.0.0.12 mpls 0.0 12 2 VOICE-SLA, BULK-SLA 10.0.0.12 biz-internet 2.1 45 18 BULK-SLA ``` This is AAR's decision-making laid bare. The MPLS path to the remote edge has 0% loss, 12 ms latency, 2 ms jitter - it meets both SLA classes, so voice can use it. The internet path is at 2.1% loss and 45 ms - it fails VOICE-SLA (over the 1% loss threshold) but meets BULK-SLA, so backup traffic uses it but voice does not. Watch this table during a brownout and you see flows move as a path drops below its SLA. ### What the edge is enforcing ``` Edge# show sdwan policy from-vsmart app-route-policy from-vsmart app-route-policy STEER-APPS ... Edge# show sdwan app-route sla-class SLA-CLASS INDEX LOSS LATENCY JITTER VOICE-SLA 1 1 150 30 BULK-SLA 2 10 300 - ``` ## Direct Internet Access as a data policy action A common data-policy action is steering internet-bound traffic straight out the local internet transport rather than backhauling it to a data centre - Direct Internet Access. In a data policy this is a match on the destination (or an application) with a `nat use-vpn 0` action, sending that traffic out the local internet circuit with NAT. DIA is important enough to have its own article in this cluster; the point here is that it is *expressed as a data policy*, in the same framework as AAR. ## Application identification AAR is only as good as its ability to recognise applications. SD-WAN uses several methods: - **NBAR2** \- deep packet inspection that recognises thousands of applications by signature. The workhorse. - **Custom applications** \- define your own by IP/port/domain for in-house apps NBAR2 does not know. - **SaaS / cloud onramp** \- specific optimisations for Office 365, Salesforce, and other SaaS, measuring the path to the actual SaaS endpoint. The quality of application identification is what separates a working AAR deployment from one that misclassifies traffic and steers it wrong. Getting the app-lists right (and validating them with `show sdwan app-fwd` statistics) is where the real deployment effort goes. ## Design guidance 1. **Start with a small set of SLA classes.** Voice, video, business-critical, bulk. You do not need dozens - a handful of tiers covers most requirements and stays maintainable. 2. **Set preferred colours deliberately.** Voice preferring MPLS with internet as the SLA-meeting fallback is the canonical design. Bulk preferring internet keeps expensive MPLS free for what needs it. 3. **Decide the fallback behaviour.** When no path meets the SLA, do you drop (strict, for traffic that is useless if degraded) or use the best available (for traffic that is better late than never)? This is a per-class choice with real consequences. 4. **Tune BFD carefully.** The measurement interval and multiplier determine how fast AAR reacts and how much control-plane load the measurement generates. Too aggressive and you get flapping; too slow and users notice degradation before AAR does. 5. **Validate application identification** before trusting the steering. A misidentified application is steered by the wrong policy. ## Key takeaways - AAR steers each application onto the path that currently meets its SLA, measured live via BFD (loss, latency, jitter). This is the headline reason SD-WAN replaced traditional WAN. - Data policy (centralised, on the Controller) matches actual flows and acts on them - permit/deny, DSCP, NAT/DIA, and application-aware routing. Control policy shapes routes; data policy steers packets. - You define **SLA classes** (loss/latency/jitter thresholds), then bind applications to an SLA class and a **preferred colour**: use the preferred path if it meets the SLA, else any path that does. - `show sdwan app-route stats` shows the live per-path measurements AAR decides on - the single most illuminating AAR command. - Application identification (NBAR2, custom apps, SaaS onramp) is what AAR acts on - get it right or steering is wrong. - Design with a few SLA classes, deliberate preferred colours, an explicit fallback per class, and carefully tuned BFD. Next: [SD-WAN localized policy - ACLs and route policies at the edge](https://www.pinglabz.com/sd-wan-localized-policy/). The full cluster index lives on the [SD-WAN pillar](https://www.pinglabz.com/sd-wan/). ### SD-WAN Centralized Control Policy: Shaping the Overlay URL: https://www.pinglabz.com/sd-wan-centralized-control-policy/ Last updated: 2026-07-12T09:12:14.000Z Here is the thing that makes SD-WAN fundamentally different from a traditional WAN: you do not configure routing policy on the routers. You configure it once, centrally, on the Controller, and it shapes what every WAN Edge sees. A centralised control policy is a single document that decides the topology, the path preferences, and the reachability of your entire overlay - and no edge device holds a copy. This article covers centralized control policy: what it is, how it shapes the fabric, and the classic use cases. It extends the [complete SD-WAN guide](https://www.pinglabz.com/sd-wan/). Command and policy syntax is drawn from Cisco's current 20.x documentation, clearly labelled as a documented reference. ## The mental model: policy at the route reflector Recall from the [OMP deep dive](https://www.pinglabz.com/omp-deep-dive/) that the Controller (vSmart) is a route reflector - every edge peers with it, and it redistributes OMP routes and TLOCs between them. A centralised control policy is a filter applied *at that reflector*, on the OMP updates flowing through it. Because it sits at the reflector, one policy controls what every edge learns. Change the policy once on the Controller, and the entire overlay's topology and path preferences change. This is the direct analogue of applying a route-map at a BGP route reflector - except here it is the standard operating model, not an advanced trick. **The crucial consequence: control policy shapes the control plane, not the data plane.** It decides which routes and TLOCs an edge *learns*, and with what preferences. It does not touch individual packets - that is data policy's job (covered next). Control policy draws the map; data policy directs the traffic on it. ## Anatomy of a control policy A centralised control policy has a familiar structure - a list of sequences, each with a match and an action, applied in a direction to a list of sites: ``` policy control-policy PREFER-MPLS sequence 10 match tloc color mpls ! action accept set preference 200 ! make MPLS TLOCs more preferred ! ! ! sequence 20 match route prefix-list CORPORATE-PREFIXES ! action accept set preference 100 ! ! ! default-action accept ! ! apply-policy site-list BRANCH-SITES control-policy PREFER-MPLS out ! 'out' = toward the edges ! ! ``` Read that structure and it is BGP policy in different clothing: match on TLOC (colour) or route (prefix), take an action (accept/reject), set an attribute (preference), and apply it in a direction to a set of sites. The `out` direction means "as OMP updates leave the Controller toward these edges" - which is how you control what those edges learn. ## The classic use cases ### 1\. Path preference (prefer one transport) The most common control policy: make branches prefer MPLS over internet (or vice versa) by setting a higher preference on one colour's TLOCs, as in the example above. Every branch in the site-list now prefers the chosen transport, set in one place. To flip the whole estate to prefer internet, you change one number on the Controller. ### 2\. Topology control: hub-and-spoke vs full mesh By default SD-WAN is a full mesh - every edge can build a direct tunnel to every other edge. Often you do not want that; you want spokes to reach each other only via a hub (for a firewall, for scale, for simplicity). Control policy builds the topology by controlling which TLOCs each site learns: ``` control-policy HUB-AND-SPOKE sequence 10 match route prefix-list SPOKE-PREFIXES ! action accept set tloc-list HUB-TLOCS ! rewrite the next-hop TLOC to the hub ! ! ! ``` By rewriting the TLOC of spoke prefixes to point at the hub, spokes reach each other *through* the hub - a hub-and-spoke overlay built entirely by policy, with no per-device configuration. This is the SD-WAN version of building topology, and it is one policy document. ### 3\. Reachability / segmentation Control policy can also decide which sites even *learn about* which prefixes - reject the OMP routes for a VPN toward sites that should not reach it, and those sites simply never see the prefix. This is segmentation at the control-plane level, complementary to VPN-based isolation. ### 4\. Extranet and shared services Leak specific prefixes between VPNs (a shared services VPN reachable from several customer VPNs) by accepting and re-originating those routes across VPN boundaries in policy - the SD-WAN analogue of VRF route-leaking, done centrally. ## Direction matters, and it is from the Controller's view The single most common mistake with control policy is the direction. It is applied **from the Controller's perspective**: out Applied to OMP updates *leaving the Controller toward the edges*. This controls what the edges **learn**. The common case - path preference, topology, reachability. in Applied to OMP updates *arriving at the Controller from the edges*. This controls what the Controller **accepts and re-advertises**. Less common - used to filter or tag at ingress. Get the direction backwards and your policy either does nothing or does the opposite of what you intended. Always reason from the Controller: "as these updates leave me toward those sites" (out) versus "as these updates arrive at me from those sites" (in). ## Verifying the effect on the edge The policy lives on the Controller, but its effect is visible on the edge - the routes and preferences an edge learned are what the policy shaped: ``` Edge# show sdwan omp routes vpn 1 10.20.0.0/24 detail ... PREFERENCE 200 ... ! the preference the control policy set ... TLOC 10.0.0.11 mpls ipsec ... Edge# show sdwan policy from-vsmart ! the policy the Controller pushed from-vsmart control-policy PREFER-MPLS ... ``` `show sdwan policy from-vsmart` is the key troubleshooting command: it shows the edge what centralised policy the Controller has applied to it. If a branch is not behaving as your policy intends, this tells you whether the policy actually reached it. ## Control policy vs the other policy types - **Centralised control policy** (this article): shapes OMP - routes, TLOCs, topology, preferences. Control plane. Applied at the Controller. - **Centralised data policy** (next article): matches actual traffic flows and steers them - the packet-level decisions, including application-aware routing. Also at the Controller. - **Localised policy**: ACLs, QoS, route policy applied on the edge itself. The per-device stuff. The division is clean and worth memorising: **control policy = which routes exist and where they point; data policy = which packets go where; localised policy = per-device forwarding treatment.** Most topology and path-preference work is control policy; most application steering is data policy. ## Key takeaways - A centralised control policy is applied **at the Controller** and shapes what every WAN Edge learns via OMP - one document controls the whole overlay's topology and path preferences. - It shapes the **control plane** (which routes and TLOCs exist, with what preference), not individual packets. It draws the map; data policy directs the traffic. - Classic use cases: **path preference** (prefer a transport), **topology** (build hub-and-spoke by rewriting TLOCs), **reachability/segmentation** (reject routes toward sites), and **extranet** (leak between VPNs). - Direction is from the **Controller's perspective**: `out` controls what edges learn (the common case), `in` controls what the Controller accepts. Getting it backwards is the number-one mistake. - `show sdwan policy from-vsmart` on an edge shows what centralised policy the Controller actually applied to it - the first troubleshooting command. - Remember the split: control policy (routes/topology) vs data policy (packet steering) vs localised policy (per-device forwarding). Next: [SD-WAN data policy and application-aware routing (AAR)](https://www.pinglabz.com/sd-wan-data-policy-aar/). The full cluster index lives on the [SD-WAN pillar](https://www.pinglabz.com/sd-wan/). ### OMP Deep Dive: Routes, TLOCs, Service Routes, and Path Selection URL: https://www.pinglabz.com/omp-deep-dive/ Last updated: 2026-07-12T09:12:13.000Z OMP is to Catalyst SD-WAN what BGP is to the internet: the one protocol that carries everything. Routes, next-hop locations, service advertisements, encryption keys, policy - all of it rides OMP, over a secure connection to the controllers. If you understand OMP, you understand how an SD-WAN fabric actually forwards a packet. If you do not, the whole thing stays magic. This article is the deep dive: OMP routes, TLOCs, service routes, and path selection. It extends the [complete SD-WAN guide](https://www.pinglabz.com/sd-wan/). As with the rest of this series, the command output is drawn from Cisco's current 20.x documentation and reflects real WAN Edge CLI - a documented reference, clearly labelled, because a full controller stack cannot be driven through a device CLI session. ## What OMP is The Overlay Management Protocol runs between each WAN Edge and the **Controller** (formerly vSmart), over a secure DTLS/TLS connection. The Controller is a route reflector for the entire overlay - edges do not peer with each other, they all peer with the Controllers, which redistribute the information. This is deliberately BGP-like, and the analogy holds all the way down. OMP carries three kinds of advertisement, and understanding the three is understanding OMP: OMP routes The prefixes reachable in each service VPN, and which TLOC (which edge, over which transport) to reach them through. The "what and where to reach it". TLOC routes Transport Locators - the identity of each tunnel endpoint: system-ip + colour + encapsulation. The "how to build the tunnel to that edge". The BGP next-hop analogue. Service routes Advertisements of services (firewall, IPS, load-balancer) available at a site, so traffic can be steered through them. The "what's available here". ## The TLOC: the concept that makes SD-WAN make sense A TLOC (Transport Locator) is the single most important idea in the data plane. It uniquely identifies a tunnel endpoint by three things: ``` system-ip : colour : encapsulation e.g. 10.0.0.11 : mpls : ipsec e.g. 10.0.0.11 : biz-internet : ipsec ``` One WAN Edge with two transports (MPLS and internet) has **two TLOCs** \- same system-ip, different colour. A remote edge learns both, and can build a tunnel over either transport to reach it. The **colour** is how SD-WAN reasons about transports abstractly: "mpls", "biz-internet", "public-internet", "lte" - the colour, not the underlay address, is what policy and path selection work with. An OMP route says "prefix 10.20.0.0/24 is reachable via TLOC 10.0.0.11:mpls:ipsec". Resolve the TLOC to its actual tunnel endpoint, build (or reuse) the IPsec tunnel over the MPLS transport, and forward. **That two-step - OMP route points at a TLOC, TLOC resolves to a tunnel - is the entire SD-WAN forwarding model.** It is exactly BGP's "prefix points at a next-hop, next-hop resolves to an interface", lifted into an overlay. ## Verifying OMP on the edge (documented reference) ### OMP routes ``` Edge# show sdwan omp routes PATH ATTRIBUTE VPN PREFIX FROM PEER ID LABEL STATUS TLOC IP COLOR ENCAP --------------------------------------------------------------------------------- 1 10.20.0.0/24 10.0.0.4 1 1002 C,I,R 10.0.0.12 mpls ipsec 1 10.20.0.0/24 10.0.0.4 2 1002 C,I,R 10.0.0.12 biz-internet ipsec ``` The same prefix, reachable via the same remote edge (10.0.0.12) over two colours. `C,I,R` \= chosen, installed, resolved - the healthy state. Two paths means the edge can load-share or fail over between MPLS and internet to reach that prefix. ### TLOCs ``` Edge# show sdwan omp tlocs ADDRESS FAMILY TLOC IP COLOR ENCAP FROM PEER STATUS ------------------------------------------------------------------ ipv4 10.0.0.11 mpls ipsec 0.0.0.0 C,Red,R ipv4 10.0.0.11 biz-internet ipsec 0.0.0.0 C,Red,R ipv4 10.0.0.12 mpls ipsec 10.0.0.4 C,I,R ipv4 10.0.0.12 biz-internet ipsec 10.0.0.4 C,I,R ``` The local edge's own two TLOCs (from peer 0.0.0.0 = self) and the remote edge's two, learned from the Controller (10.0.0.4). This is the map of every tunnel endpoint in the fabric. ### BFD - the tunnels themselves ``` Edge# show sdwan bfd sessions SYSTEM IP SITE ID STATE SOURCE-TLOC REMOTE-TLOC DST-IP PROTO UPTIME ------------------------------------------------------------------------------------ 10.0.0.12 200 up mpls mpls 198.51.100.2 ipsec 1:20:15 10.0.0.12 200 up biz-internet biz-internet 203.0.113.2 ipsec 1:20:14 ``` BFD runs inside every data-plane tunnel and is how SD-WAN measures each path's health - loss, latency, jitter - in real time. `show sdwan bfd sessions` is the data-plane counterpart to `show sdwan omp routes`: OMP tells you what the control plane knows; BFD tells you whether the tunnels actually work. Application-aware routing (covered later) is built directly on these BFD measurements. ## OMP path selection When multiple OMP routes exist for the same prefix, OMP selects among them with a best-path algorithm that is - once again - deliberately BGP-like: 1. **Valid and reachable** routes only (the TLOC must resolve and BFD must be up). 2. **Higher OMP route preference** wins (set by policy - the local-pref analogue). 3. **Lower origin metric.** 4. **OMP route origin type** and other tie-breaks. 5. **Lower TLOC preference,** then TLOC IP, as final tie-breaks. The number of equal-cost paths installed is governed by `send-path-limit` (how many the Controller advertises) and `ecmp-limit` (how many the edge installs). By default an edge advertises each route-TLOC tuple, and a device can have up to eight TLOCs. Understanding this is what lets you reason about "why is my traffic taking MPLS when I wanted internet" - it is an OMP best-path question, answered by policy setting route/TLOC preference. ## OMP vs BGP: the analogy in full **OMP route**\= a BGP prefix (NLRI) **TLOC**\= the BGP next-hop (but richer: colour + encap) **Controller (vSmart)**\= a BGP route reflector **OMP route preference**\= local preference **Centralised control policy**\= route-maps applied at the reflector If you know BGP, you already understand 80% of OMP - which is why the CCIE blueprint expects BGP mastery before SD-WAN. The remaining 20% is the TLOC (a next-hop that carries transport identity) and the fact that policy is applied centrally at the Controller rather than per-router. That is genuinely why the OMP deep dive is cross-linked from the [BGP pillar](https://www.pinglabz.com/bgp/): readers who know BGP ask "what is the SD-WAN version of this", and OMP is the answer. ## Troubleshooting OMP 1. **A prefix is missing on an edge?** `show sdwan omp routes` \- is it there but not chosen (a policy filtering it at the Controller), or absent entirely (the originating edge is not advertising it, or its control connection is down)? 2. **Route present but not forwarding?** The TLOC is not resolved, or the BFD session is down. Check `show sdwan omp tlocs` and `show sdwan bfd sessions` \- the tunnel to that TLOC must be up. 3. **Traffic on the wrong transport?** An OMP best-path outcome. Check the route/TLOC preferences a centralised control policy is setting. 4. **Nothing at all?** `show sdwan control connections` first. No control connection to the Controller means no OMP, means nothing. ## Key takeaways - OMP is SD-WAN's BGP - one protocol carrying routes, TLOCs, service routes, keys and policy, between each edge and the Controllers (which act as route reflectors). - A **TLOC** identifies a tunnel endpoint by system-ip + colour + encapsulation. An edge with two transports has two TLOCs. It is the BGP next-hop, enriched with transport identity. - Forwarding is two steps: an OMP route points at a TLOC; the TLOC resolves to an IPsec tunnel over a transport. That is the whole data-plane model. - `show sdwan omp routes` / `tlocs` / `bfd sessions` are the three commands that tell you what the control plane knows and whether the tunnels work. - OMP path selection is BGP-like: valid/reachable, then OMP route preference (local-pref), then metrics and TLOC tie-breaks. - Know BGP and you know most of OMP. The new parts are the TLOC and centralised (Controller-applied) policy. Next: [SD-WAN centralized control policy - shaping the overlay](https://www.pinglabz.com/sd-wan-centralized-control-policy/). The full cluster index lives on the [SD-WAN pillar](https://www.pinglabz.com/sd-wan/), cross-linked to [BGP](https://www.pinglabz.com/bgp/). ### Catalyst SD-WAN Templates: Device, Feature, and CLI Templates Compared URL: https://www.pinglabz.com/sd-wan-templates-device-feature-cli/ Last updated: 2026-07-12T09:12:13.000Z The first thing that trips people up in Catalyst SD-WAN is not OMP or TLOCs or policy - it is provisioning. How does configuration actually get onto a WAN Edge? There are three answers, they arrived in three different eras of the product, and knowing which one you are dealing with matters more than almost anything else when you inherit an SD-WAN deployment. This article covers CLI templates, feature templates, and configuration groups - what each is, why they exist, and which one to use in 2026\. It extends the [complete SD-WAN guide](https://www.pinglabz.com/sd-wan/). **A note on the output below.** A full Catalyst SD-WAN control-plane deployment (Manager, Validator, Controller, plus certificate onboarding) is a large, GUI-and-API-driven build that cannot be faithfully reproduced through a device CLI session. The command syntax and output shown here is drawn from Cisco's current 20.x documentation and reflects real WAN Edge CLI - it is presented as a documented reference, clearly labelled, not as a home-lab capture. Where a command is standard IOS XE that runs on the edge, it is the genuine article. ## The three provisioning models, in order of arrival CLI templates The original. You paste a full device configuration into Manager as a template, with variables for the per-device bits (hostname, IPs, site-id). Total control, total responsibility - you own every line. Used for corner cases the structured templates cannot express. Feature templates The structured middle era. Instead of raw config, you fill in forms - a template for VPN 0, one for the tunnel interface, one for OMP, and so on - then bundle them into a device template. Manager renders the CLI for you. Reusable, less error-prone, but a lot of small templates to manage. Configuration groups The current model (from 20.8, matured through 20.11+). A higher-level, intent-based grouping: you describe a *profile* (system, transport, service, WAN) and apply it to a group of devices, with far less repetition than feature templates. This is where Cisco is investing. ## CLI templates: full control, full responsibility A CLI template is exactly what it sounds like - the device's running configuration, with variables: ``` system host-name {{hostname}} system-ip {{system_ip}} site-id {{site_id}} organization-name "PingLabz-SDWAN" vbond {{vbond_ip}} ! sdwan interface GigabitEthernet1 tunnel-interface encapsulation ipsec color {{tloc_color}} exit exit ``` You get complete control - every knob, every corner case, anything the product supports. The price is that you also own every mistake, and there is no structure to reuse. CLI templates are the escape hatch: when a feature template cannot express what you need, you drop to CLI. In a well-run modern deployment they are the exception, not the rule. ## Feature templates: structured but proliferating Feature templates decompose the device config into functional blocks, each configured through a form in Manager: - A **system** feature template (system-ip, site-id, org-name, vBond). - A **VPN 0 (transport)** template plus a tunnel-interface template per transport. - A **VPN 512 (management)** template. - A **service VPN** template per customer VPN, plus interface templates. - An **OMP** template, a **BGP/OSPF** template if you run one to the service side, and so on. You bundle these into a **device template** and attach it to devices, supplying the per-device variables (via a CSV for bulk). Manager renders and pushes the CLI. Feature templates were a big step up from raw CLI - reusable, validated, less error-prone. Their weakness is proliferation: a real deployment ends up with dozens of feature templates and a matrix of device templates, and changing something common (an NTP server, a syslog host) can mean touching many templates. That pain is exactly what configuration groups set out to solve. ## Configuration groups: the current model Configuration groups (20.8 onward, and the model to learn now) raise the abstraction level. Instead of many small feature templates, you describe a handful of **profiles**: System profile Identity and global settings: system-ip, site-id, org-name, AAA, logging, NTP. Transport profile VPN 0, the tunnel interfaces, colours, and the underlay routing to reach the controllers. Service profile The service-side VPNs, LAN interfaces, and the routing to the local network. CLI profile (add-on) For the corner cases - drop to CLI for anything the structured profiles do not yet express, layered on top. The win is **reuse with less repetition**: common settings live in a profile once, feature parameterisation is cleaner, and applying a group to many devices is more intent-driven. Configuration groups also integrate with Manager's newer topology and monitoring workflows. For any new deployment on a recent version, this is where to start. ## Verifying what got pushed (on the edge) Whichever model produced it, the result is CLI on the WAN Edge, and you verify it with the `show sdwan` family - these are real IOS XE Catalyst SD-WAN commands (documented reference): ``` Edge# show sdwan running-config system system system-ip 10.0.0.11 site-id 100 organization-name "PingLabz-SDWAN" vbond 10.0.0.3 Edge# show sdwan control connections PEER PEER PEER SITE DOMAIN PEER PROT STATE TYPE PROT SYSTEM-IP ID ID PRIVATE IP vsmart dtls 10.0.0.4 1 1 10.0.0.4 dtls up vbond dtls 10.0.0.3 0 0 10.0.0.3 dtls up vmanage dtls 10.0.0.2 1 0 10.0.0.2 dtls up ``` `show sdwan control connections` is the single most important command on a WAN Edge - it tells you whether the device has joined the fabric (control connections to the Manager, Validator, and Controller all `up`). If these are not up, nothing else matters, and no template will help until they are. ## Choosing a model 1. **New deployment, recent version:** configuration groups. It is the current direction, cleaner to maintain, and integrates with the newer Manager workflows. 2. **Existing feature-template deployment:** keep it if it works, but plan the eventual move to configuration groups. Do not mix models for the same devices. 3. **A device that needs something no structured model expresses:** a CLI template, or a CLI add-on profile on top of a configuration group. The escape hatch, used sparingly. 4. **Bulk-provisioning many identical branches:** any structured model plus a device-variables CSV. This is where templating pays for itself. ## Key takeaways - Catalyst SD-WAN has three provisioning models: **CLI templates** (raw config, full control), **feature templates** (structured forms, reusable but proliferating), and **configuration groups** (intent-based profiles, the current model from 20.8+). - For any new deployment on a recent version, start with **configuration groups**. Feature templates are the previous generation; CLI templates are the escape hatch. - All three ultimately render CLI onto the WAN Edge, verified with the `show sdwan` command family. - `show sdwan control connections` is the first command to run on any edge - it confirms the device has joined the fabric (Manager, Validator, Controller all up). - Do not mix provisioning models for the same devices; migrate deliberately when you move from feature templates to configuration groups. Next: [OMP deep dive - routes, TLOCs, service routes, and path selection](https://www.pinglabz.com/omp-deep-dive/). The full cluster index lives on the [SD-WAN pillar](https://www.pinglabz.com/sd-wan/). ### SD-Access Design Questions: Where Fabrics Fit and Where They Don't URL: https://www.pinglabz.com/sd-access-design-considerations/ Last updated: 2026-07-12T09:03:25.000Z Not every network should be a fabric. SD-Access is a genuinely powerful architecture, but it is also a big commitment - a controller, an identity service, compatible hardware, and a operational model your team has to learn. The most valuable thing a senior engineer can do with SD-Access is know when *not* to deploy it. This article is the honest design conversation the vendor slide decks skip. It closes the SD-Access series. For the mechanics, see the earlier articles; this one is about judgement. For the wider context, see the [network virtualization pillar](https://www.pinglabz.com/network-virtualization/). ## What SD-Access is genuinely good at Give the architecture its due - where it fits, it fits well: - **Large campuses with a strong segmentation requirement.** If you need to separate guests, IoT, corporate, and contractors *consistently across thousands of ports* and enforce it by identity, SD-Access does this better than any pile of VLANs and ACLs. This is its home turf. - **Environments where users and devices move.** The anycast gateway and LISP roaming mean policy follows people between buildings without re-configuration. For a hospital, a university, a large corporate campus, this is real value. - **Organisations that already run ISE and want to extend identity-based policy into the network fabric itself.** If you have made the TrustSec investment, SD-Access is the natural extension. - **Greenfield builds at scale** where you can design the underlay, choose compatible hardware, and adopt the operational model from day one - rather than retrofitting. ## Where a fabric does not fit And now the part the design guides underplay: 1 **Small sites.** A branch with two switches and forty users does not need LISP, VXLAN, and a controller. The operational overhead dwarfs the benefit. Fabric-in-a-box exists, but for a genuinely small site a well-configured traditional access layer is simpler, cheaper, and easier to troubleshoot. 2 **No segmentation requirement.** If your security model does not actually need identity-based micro-segmentation, the single biggest reason to adopt SD-Access is absent. Much of the complexity buys you segmentation you are not going to use. 3 **Incompatible or mixed-vintage hardware.** SD-Access needs fabric-capable switches. A campus full of older or mixed gear means a forklift, and the cost of the hardware refresh often dominates the business case. 4 **A team without the skills or appetite.** SD-Access changes how you operate - you drive the network through Catalyst Center, troubleshoot LISP and VXLAN, and depend on the controller. A team that is not ready for that will fight the fabric, and a fabric fought is worse than a traditional network run well. 5 **Simple, stable networks that already work.** "If it ain't broke" is a legitimate engineering position. A traditional campus that meets its requirements, that the team understands, and that is stable does not automatically benefit from being rebuilt as a fabric. ## The honest cost of a fabric Every SD-Access deployment carries costs that are easy to underweight in the planning phase and impossible to ignore once you are live: - **The controller is a dependency.** Catalyst Center provisions and operates the fabric. It needs to be sized, made redundant, backed up, patched, and understood. It is another critical system, and when it has a problem, your ability to change the network has a problem. - **Troubleshooting is different, and initially harder.** "Why can't this host reach that host" now involves LISP map-caches, VXLAN encapsulation, SGT policy, and the controller's view - not just a routing table and a MAC table. The skills transfer, but there is a learning curve, and it lands during outages. - **You are more locked in.** A fabric is a Cisco architecture end to end. That is fine if it is a deliberate choice, but it is a choice, and it narrows your options. - **The underlay still has to be right.** All the fabric magic sits on an IGP with correct MTU and fast convergence. Get the boring part wrong and the clever part fails in confusing ways. ## The alternatives SD-Access competes with SD-Access is not the only way to get segmentation and identity-based policy: Traditional + TrustSec Run a conventional campus but layer TrustSec/SGTs on top for identity-based segmentation, without the full fabric. You get much of the security benefit with far less architectural change. VXLAN/EVPN fabric A standards-based fabric (BGP EVPN) gives you the overlay benefits with more vendor flexibility and a control plane many teams already know from the data centre - at the cost of the turnkey campus automation. Well-run traditional A conventional design with good VLAN hygiene, 802.1X, and ACLs is entirely adequate for a great many networks, and it is what most of your team already knows. The middle option is the one most under-considered: **traditional access plus TrustSec** gives you identity-based segmentation - the main reason to want SD-Access - without the controller, the LISP, or the hardware refresh. For an organisation whose real requirement is "segment by identity" rather than "automate the campus", it is often the better fit. And it is built on the real TrustSec technology PingLabz has captured standalone. ## The decision framework Ask these five questions, in order: 1. **Do you have a real, funded segmentation requirement?** If no, most of the reason for SD-Access is gone - stop here and consider a traditional design. 2. **Is it a large campus with mobility?** If yes, SD-Access's strengths align. If it is a handful of small sites, they do not. 3. **Is the hardware compatible, or is a refresh funded?** If not, the business case has to justify a forklift. 4. **Does the team have the skills, or a plan and appetite to build them?** A fabric run by a team that resents it is a liability. 5. **Have you considered traditional + TrustSec as the middle path?** If your real need is segmentation, this may deliver it with far less change. If you answer yes to the first four and have genuinely weighed the fifth, SD-Access is a strong choice and its complexity is buying you something real. If you are reaching for it because it is new, or because a vendor slide was compelling, pause. The best network is the one that meets the requirement with the least complexity your team can operate well - and sometimes that is a fabric, and sometimes it very much is not. ## Key takeaways - SD-Access fits **large campuses with a real segmentation requirement, user mobility, an existing ISE investment, and compatible hardware**. There it is genuinely strong. - It is a poor fit for **small sites, networks with no segmentation need, incompatible hardware, teams without the skills, and simple stable networks that already work**. - The honest costs: the controller is a critical dependency, troubleshooting is different and initially harder, you are more locked in, and the underlay still has to be right. - The most under-considered alternative is **traditional access + TrustSec** \- identity-based segmentation without the full fabric, on real technology. - Decide with a framework: real segmentation need → scale and mobility → hardware → team readiness → whether traditional+TrustSec is the better middle path. - The best network meets the requirement with the least complexity the team can operate well. Sometimes that is a fabric; often it is not - and knowing the difference is the expert skill. That closes the SD-Access series - an honest concept-and-components treatment of the 12.5% of the CCIE EI blueprint that cannot be labbed without Catalyst Center, grounded throughout in the real LISP, VXLAN, and TrustSec technology PingLabz has captured standalone. The full cluster index lives on the [network virtualization pillar](https://www.pinglabz.com/network-virtualization/). ### SD-Access Segmentation: Macro (VN) vs Micro (SGT) URL: https://www.pinglabz.com/sd-access-macro-micro-segmentation/ Last updated: 2026-07-12T09:03:24.000Z Segmentation is the entire reason most organisations buy SD-Access. Not the automation, not the roaming - the ability to say "guests cannot reach the finance systems, IoT devices cannot reach each other, and contractors only touch the two applications they need" and have it enforced consistently everywhere, without maintaining a thousand ACLs. SD-Access does this in two layers, and understanding the difference between them is understanding SD-Access segmentation. This article covers macro-segmentation (virtual networks) and micro-segmentation (SGTs). It is a concept piece - the policy is authored in Cisco Catalyst Center and ISE - but the underlying TrustSec/SGT technology is real and PingLabz has captured it standalone. For the wider context, see the [network virtualization pillar](https://www.pinglabz.com/network-virtualization/). ## Two layers, two questions Macro-segmentation (VN) **Question it answers:** which broad network do you belong to? **Mechanism:** the virtual network (VN) - a separate VRF and routing table, mapped to a VXLAN VNI. Total isolation. Traffic in one VN cannot reach another VN except through a controlled leak point (the fusion router). **Analogy:** different buildings. You are in the Employees building or the Guests building, and the buildings do not connect except through a guarded lobby. Micro-segmentation (SGT) **Question it answers:** within your building, which rooms can you enter? **Mechanism:** the Scalable Group Tag (SGT) - a tag on every packet, with policy (SGACLs) deciding which groups can talk to which. Fine-grained, within a VN. **Analogy:** rooms within a building. Everyone is in the same building (VN), but your badge (SGT) only opens certain doors. The mental model that makes this click: **a virtual network is a building; an SGT is a badge that controls which rooms you can enter within it.** You use both - macro to separate broad trust domains completely, micro to control who talks to whom inside a domain. ## Macro: virtual networks A virtual network is a complete routing separation. Employees, Guests, IoT, and Building-Management might each be their own VN. Each is a VRF on the fabric, mapped to its own VXLAN VNI, with its own routing table. **Traffic cannot cross between VNs at all** \- not filtered, not restricted, but genuinely separate, the same way two customers' VRFs in an [MPLS L3VPN](https://www.pinglabz.com/mpls-l3vpn/) are separate. When two VNs *do* need to communicate (say, IoT devices need to reach a shared DNS server), the traffic goes out to the **fusion router** at the border, which leaks the specific routes between the VRFs. This is deliberate: cross-VN communication requires an explicit, auditable route leak, not a rule someone forgot about. VRF route-leaking is standard, real technology PingLabz has captured. The design guidance: **use few, broad VNs.** A VN is heavyweight - it is a whole VRF, a whole routing table, a whole segment to manage. You do not want dozens. Employees, Guests, IoT, and maybe Building-Management is a typical set. Finer distinctions belong to SGTs, not VNs. ## Micro: scalable group tags Within a VN, you rarely want everyone talking to everyone. Employees and Employees-Contractors might share the Employees VN but need different access. IoT cameras and IoT sensors might share the IoT VN but should not talk to each other at all. That is micro-segmentation, and it uses SGTs. An SGT is a 16-bit tag assigned to an endpoint at onboarding (by ISE, based on identity) and carried in every packet. Policy is expressed as a matrix: **source SGT × destination SGT → permit or deny**, enforced by SGACLs. The beauty is that it is *topology-independent* \- the policy is "Contractors cannot reach Finance-Servers", not "10.1.5.0/24 cannot reach 10.2.9.0/24". Renumber, move, roam - the policy holds because it follows the tag, not the address. ### The SGT is real and PingLabz has captured it SGTs and TrustSec are not SD-Access-specific - they predate it and run on standalone IOS XE. From the real TrustSec/802.1X captures on the site, an endpoint carrying its assigned SGT: ``` Switch# show cts role-based sgt-map all Active IPv4-SGT Bindings Information IP Address SGT Source ============================================ 10.1.10.50 10 LOCAL (Employees) 10.1.10.60 20 LOCAL (Contractors) 10.1.20.30 30 LOCAL (IoT-Cameras) Switch# show cts role-based permissions IPv4 Role-based permissions from group 20:Contractors to group 100:Finance-Servers: Deny IP-00 ``` That last block is micro-segmentation in action: a policy denying Contractors (SGT 20) from reaching Finance-Servers (SGT 100), enforced regardless of IP address or location. This is real TrustSec, captured on IOS XE, and it is exactly the enforcement mechanism SD-Access uses - the difference is only that Catalyst Center and ISE author and distribute the policy centrally instead of you configuring the SGACL by hand. ## How the SGT travels in the fabric Inside the fabric, the SGT rides in the **VXLAN header** (SD-Access uses VXLAN-GPO, which has a field for the group tag). So the tag is carried with the packet across the entire fabric, and the destination edge node enforces the SGACL at egress. The endpoint's group membership follows it everywhere, and enforcement happens as close to the destination as possible. The one place the SGT can be lost is at an **IP-transit border handoff** \- plain IP transit does not carry the SGT past the border unless you extend TrustSec to the next device (via SXP, which propagates IP-to-SGT bindings, or inline tagging on a capable link). This is the segmentation consequence of the border-handoff choice covered in the [border handoff article](https://www.pinglabz.com/sd-access-border-handoff/): SDA transit preserves the SGT end to end, IP transit does not by default. ## Using both together A worked example makes the two-layer model concrete: - **Employees VN** (macro): contains Employees (SGT 10) and Contractors (SGT 20). Micro-policy: Contractors can reach the two applications they need and nothing else; Employees have broader access. Both are in the same building; different badges. - **IoT VN** (macro): completely isolated from Employees at the VN level - an IoT device cannot even *route* to an employee subnet. Within IoT, micro-policy stops cameras (SGT 30) talking to sensors (SGT 31), limiting lateral movement if one is compromised. - **Guests VN** (macro): isolated from everything internal; can only reach the internet via the border. No SGT policy needed beyond "guests reach the door and nothing else". **The rule of thumb: macro-segment by broad trust domain, micro-segment within it.** Use a handful of VNs for the fundamental separations, and SGTs for the fine-grained, identity-driven policy inside each. This keeps the VN count low (manageable) and pushes the complexity into SGT policy (topology-independent and centrally managed). ## What needs the controller, and what does not - **Needs Catalyst Center + ISE:** defining VNs, authoring the SGT policy matrix, distributing SGACLs, mapping identities to SGTs, the centralised policy administration. - **Real and standalone:** TrustSec/SGT enforcement itself (SGACLs, `show cts role-based`, SGT-in-packet), captured on IOS XE. VRF separation (the macro mechanism) is standard VRF technology, also captured. The fusion-router route leak between VNs is plain VRF route-leaking. Every enforcement mechanism SD-Access segmentation uses is real technology PingLabz has captured. What Catalyst Center adds is the central authoring and distribution of the policy - turning "configure this SGACL on every switch" into "draw the matrix once". We describe the orchestration honestly and point to the real TrustSec captures for the mechanism. ## Key takeaways - SD-Access segments in two layers: **macro** (virtual networks - separate VRFs, total isolation) and **micro** (SGTs - fine-grained policy within a VN). - The model: a VN is a building; an SGT is a badge controlling which rooms you enter within it. Use both. - VNs are heavyweight (a whole VRF each) - use **few, broad** ones. Cross-VN traffic requires an explicit route leak at the fusion router. - SGTs are topology-independent - the policy is "Contractors cannot reach Finance", not an IP-based ACL. It follows the endpoint everywhere. - The SGT rides in the VXLAN header across the fabric. It is preserved end-to-end over SDA transit but **dropped at an IP-transit border** unless you extend TrustSec. - TrustSec/SGT enforcement and VRF separation are real, standalone-captured technology; Catalyst Center + ISE provide the central policy authoring. We are clear which is which. Next: [SD-Access design questions - where fabrics fit and where they don't](https://www.pinglabz.com/sd-access-design-considerations/). The full cluster index lives on the [network virtualization pillar](https://www.pinglabz.com/network-virtualization/), cross-linked to [802.1X and TrustSec](https://www.pinglabz.com/802-1x/). ### SD-Access Border Handoff: IP Transit, SDA Transit, and L2 Handoff URL: https://www.pinglabz.com/sd-access-border-handoff/ Last updated: 2026-07-12T09:03:24.000Z A fabric that cannot talk to anything outside itself is a very expensive island. The border node is how an SD-Access fabric connects to the rest of the world - the data centre, the WAN, the internet, another fabric - and getting the handoff right is where fabric design meets traditional networking. It is also where the three most-confused SD-Access terms live: IP transit, SDA transit, and Layer 2 handoff. This article untangles them. It is a concept-and-architecture piece - the border provisioning needs Cisco Catalyst Center - but the technologies underneath (VRF-lite, BGP, LISP) are standard and real. For the wider context, see the [network virtualization pillar](https://www.pinglabz.com/network-virtualization/). ## The border's job Inside the fabric, everything is VXLAN-encapsulated, LISP-mapped, and SGT-tagged. Outside the fabric, none of that exists - it is ordinary IP routing. The border node is the translator. It **de-encapsulates** fabric traffic leaving the fabric and **encapsulates** external traffic entering it, and it decides how the fabric's virtual networks map to whatever is on the other side. The way it does that translation is the "transit" type, and there are three patterns. ## 1\. IP transit The simplest and most common. The border hands off to a traditional routed network using **VRF-lite** \- each fabric virtual network is mapped to a VRF on the border, and each VRF connects to the outside via its own routed sub-interface running a routing protocol (usually BGP). IP transit at a glance **How it hands off:** VRF-lite - one VRF and one BGP session per virtual network, over a routed link to a fusion router or the next-hop device. **What is preserved:** the VRF separation (macro-segmentation). Each VN stays isolated on the other side. **What is lost:** the SGT. IP transit does *not* carry the security group tag past the border unless you specifically extend TrustSec (SXP or inline tagging) to the next device. **Use when:** connecting the fabric to a traditional network, a data centre, or the internet - the everyday case. The border config is recognisable to anyone who has done [MPLS L3VPN](https://www.pinglabz.com/mpls-l3vpn/) or plain VRF-lite - a VRF per VN, a routed sub-interface per VRF, a BGP session per VRF: ``` vrf definition VN-EMPLOYEES rd 1:10 address-family ipv4 ! interface GigabitEthernet1/0/48.10 encapsulation dot1Q 10 vrf forwarding VN-EMPLOYEES ip address 172.16.10.1 255.255.255.252 ! router bgp 65001 address-family ipv4 vrf VN-EMPLOYEES neighbor 172.16.10.2 remote-as 65100 neighbor 172.16.10.2 activate ``` That is standard VRF-aware BGP - real, labbable, and exactly the material in the MPLS and route-control clusters. The **fusion router** on the other side is the device that (optionally) leaks routes between VRFs where the design needs shared services, and it is plain VRF route-leaking that PingLabz has captured before. ## 2\. SDA transit When you have *multiple fabrics* and want to connect them while keeping the fabric semantics intact end to end, IP transit is not enough - it drops the SGT and terminates the fabric at each border. SDA transit solves this by **extending the fabric across the transit itself**. SDA transit at a glance **How it hands off:** it does not hand off to a traditional network - it carries VXLAN (with the VNI and SGT intact) across the transit between fabric sites, coordinated by a transit control plane node. **What is preserved:** everything - the virtual network *and* the SGT travel end to end between fabrics. **Use when:** a multi-site SD-Access deployment (SD-Access for Distributed Campus) where policy must be consistent across sites. The key difference: with IP transit, the fabric *ends* at each border and traditional routing takes over. With SDA transit, the fabric *continues* across the transit - it is one policy domain spanning multiple sites. SDA transit is heavier (it needs a dedicated transit control plane and a capable underlay between sites) but it preserves segmentation and SGTs everywhere, which IP transit cannot. ## 3\. Layer 2 handoff Sometimes you need to extend a Layer 2 domain out of the fabric - to a legacy device, a data-centre subnet, or a piece of equipment that has to be in the same broadcast domain as fabric endpoints. Layer 2 handoff maps a fabric virtual network's Layer 2 segment to a traditional VLAN outside the fabric. Layer 2 handoff at a glance **How it hands off:** the border maps a fabric L2 VNI to an external 802.1Q VLAN, bridging the fabric segment to a traditional switched network. **What is preserved:** the Layer 2 adjacency - endpoints on both sides are in the same broadcast domain. **Use when:** migrating into the fabric (keeping old and new in one subnet during transition), or connecting equipment that genuinely requires L2 adjacency with fabric endpoints. **Caution:** you are extending a broadcast domain, so all the usual L2 risks apply - keep it minimal. Layer 2 handoff is mostly a *migration* tool. During a brownfield migration to SD-Access, you often need old (traditional VLAN) and new (fabric) hosts to coexist in the same subnet while you move them across. L2 handoff bridges the two. Once migration is complete, you generally remove it - a permanent stretched Layer 2 domain is something to minimise, not embrace. ## Choosing the transit **Fabric to data centre / WAN / internet** **IP transit.** VRF-lite handoff. Extend TrustSec (SXP/inline) if you need SGTs beyond the border. **Fabric to fabric, multi-site, consistent policy** **SDA transit.** Keeps VN + SGT intact across sites. **Migration or L2-adjacent legacy device** **Layer 2 handoff.** Temporary bridge; minimise and remove after migration. ## Border types: internal, external, anywhere You will also hear borders described by *what they know about external routes*: - **Internal border** \- connects to known networks (the data centre, specific internal subnets). It imports those specific routes into the fabric. - **External border** \- the default exit, connecting to the unknown (the internet). Endpoints use it as the gateway of last resort; it does not need to know every external route. - **Anywhere border** \- both roles on one device, connecting to both known internal networks and the unknown outside. Most fabrics have an external border for internet-bound traffic and one or more internal borders for the data centre - or an anywhere border doing both on a capable device. Deploy them redundantly, as with every fabric role. ## What needs the controller, and what does not - **Needs Catalyst Center:** assigning the border role, provisioning the VN-to-VRF mappings, configuring SDA transit, the fabric-side of the handoff. - **Real and standalone:** the external side of an IP transit handoff is **VRF-lite + VRF-aware BGP + route leaking on a fusion router** \- all standard, all captured in the MPLS and route-control clusters. The routing that carries fabric VNs into the wider network is ordinary VRF routing. The border's external handoff is, at its heart, the same VRF-lite and VPNv4-adjacent routing PingLabz has labbed extensively. The fabric side needs the controller; the traditional-network side is real networking you can build and verify. ## Key takeaways - The border node translates between the fabric's VXLAN/LISP/SGT world and the ordinary routed world outside. The "transit" type is how it does that translation. - **IP transit:** VRF-lite handoff, one VRF + BGP session per virtual network. Preserves VRF separation, drops the SGT (unless you extend TrustSec). The everyday case. - **SDA transit:** extends the fabric (VXLAN, VNI, and SGT) across the transit between fabric sites. For multi-site with consistent end-to-end policy. - **Layer 2 handoff:** bridges a fabric L2 segment to an external VLAN. Mostly a migration tool - minimise and remove. - Borders are also classed as internal (known routes), external (default exit), or anywhere (both). Deploy redundantly. - The external side of an IP-transit handoff is standard VRF-lite + BGP + fusion-router route leaking - real, labbable technology. The fabric side needs Catalyst Center. We are clear about the boundary. Next: [SD-Access segmentation - macro (VN) vs micro (SGT)](https://www.pinglabz.com/sd-access-macro-micro-segmentation/). The full cluster index lives on the [network virtualization pillar](https://www.pinglabz.com/network-virtualization/), cross-linked to [MPLS](https://www.pinglabz.com/mpls/). ### SD-Access Host Onboarding: From Port to Policy URL: https://www.pinglabz.com/sd-access-host-onboarding/ Last updated: 2026-07-12T09:03:23.000Z An endpoint plugs into a port. Somehow, seconds later, it is in the right virtual network, carrying the right security group tag, with the right policy following it wherever it goes in the fabric. That "somehow" is host onboarding, and it is where SD-Access stops being an architecture diagram and starts being something a user experiences. This article walks the journey from physical port to enforced policy. It is a concept-and-components piece - the orchestration needs Cisco Catalyst Center - but the underlying mechanisms (802.1X, SGT assignment, LISP registration) are real technology PingLabz has captured standalone. For the wider context, see the [network virtualization pillar](https://www.pinglabz.com/network-virtualization/). ## The journey, step by step 1 **Connect.** A device plugs into an edge node port. The port is a fabric access port, not an ordinary switchport - it is waiting to authenticate whatever connects. 2 **Authenticate.** The edge node challenges the device via 802.1X (or MAB for devices that cannot do 802.1X). The credentials go to ISE, the identity service. 3 **Authorize.** ISE decides who the device is and returns a policy: which virtual network it belongs to, and which SGT it carries. This is the moment identity becomes network placement. 4 **Assign.** The edge node places the endpoint into the assigned virtual network (a fabric VLAN mapped to a VNI) and stamps it with the SGT ISE returned. 5 **Register.** The edge node registers the endpoint's address with the LISP control plane node: "this EID is behind me." Now the fabric can find it. 6 **Forward.** The endpoint gets its gateway (the anycast gateway on the edge node), and traffic flows - VXLAN-encapsulated, SGT-tagged, policy-enforced. What makes this powerful: **the same user gets the same policy on any port, in any building, because the policy is tied to identity, not to a switchport or a VLAN.** Move to a different desk, plug into a different edge node, and you land in the same virtual network with the same SGT. That is the whole promise of SD-Access, delivered at onboarding time. ## The anycast gateway: why the endpoint never notices moving Every edge node in the fabric presents the *same* gateway IP and MAC for a given virtual network. This is the **anycast gateway**, and it is quietly one of the most important pieces of the design. Because every edge node answers to the same default gateway, an endpoint that moves from one edge node to another does not need to re-ARP, does not notice a gateway change, and does not drop its sessions. Its gateway is "here" no matter which edge node "here" is. Combined with LISP re-registration (the new edge node tells the control plane "this endpoint is now behind me"), roaming becomes seamless - the fabric updates its map, the endpoint stays blissfully unaware. ## 802.1X: the real, labbable part The authentication step is standard 802.1X, and it is entirely real technology that PingLabz has configured and captured on IOS XE (the whole [802.1X cluster](https://www.pinglabz.com/802-1x/) is built on real captures). The edge node's port config for fabric onboarding is 802.1X with a few fabric-specific additions: ``` interface GigabitEthernet1/0/10 description FABRIC-EDGE-PORT switchport mode access access-session port-control auto access-session host-mode multi-auth dot1x pae authenticator mab service-policy type control subscriber FABRIC-POLICY ``` And the verification is the same 802.1X you would run on any Catalyst: ``` Edge# show access-session interface Gi1/0/10 details Interface: GigabitEthernet1/0/10 MAC Address: 0011.2233.4455 IPv4 Address: 10.1.10.50 Status: Authorized Domain: DATA Oper host mode: multi-auth Authorized By: Authentication Server SGT: 0010-0 (Employees) ``` That `SGT: 0010` line is the key moment - the authorization returned a security group tag, and now every packet from this endpoint carries it. From here, policy follows the endpoint everywhere in the fabric. ## Closed, low-impact, and open modes How aggressively you enforce authentication is a deployment choice, and it maps to three modes you will hear about: Closed mode No access at all until authentication succeeds. The most secure, the most disruptive to deploy. A device that cannot authenticate gets nothing. Low-impact mode A pre-auth ACL allows limited access (DHCP, DNS, PXE) before authentication completes, then full policy applies after. The usual production choice. Open / monitor mode Authenticate and log, but do not enforce. For rolling out 802.1X without breaking anything while you find the devices that cannot authenticate. Deploy here first. **The standard rollout path is open → low-impact → closed.** Start in monitor mode to discover every device that will fail authentication (there are always more than you think - printers, badge readers, old IoT), fix or profile them, then tighten. Going straight to closed mode is how you take down a floor of users on day one. ## What needs the controller, and what does not - **Needs Catalyst Center + ISE:** defining the authentication policy, mapping identities to virtual networks and SGTs, provisioning the fabric edge ports as onboarding ports, the overall orchestration. - **Real and standalone:** 802.1X and MAB (fully captured in the 802.1X cluster), SGT assignment via RADIUS, LISP registration of the endpoint (the map-cache and database, captured standalone). Every mechanism the onboarding flow uses is real technology. The orchestration - the "plug in and it just works" experience - is what Catalyst Center and ISE provide. The individual mechanisms are all things you can build and verify by hand, and PingLabz has. We describe the orchestrated flow honestly as architecture, and point to the real captures for the pieces. ## Key takeaways - Host onboarding is a six-step journey: connect → authenticate (802.1X/MAB) → authorize (ISE returns VN + SGT) → assign (into the virtual network, tagged) → register (with the LISP control plane) → forward (VXLAN, SGT, policy). - The policy follows **identity, not the port**. The same user gets the same virtual network and SGT on any port in any building. - The **anycast gateway** \- the same gateway IP/MAC on every edge node - is why an endpoint can roam without noticing or dropping sessions. - The authentication step is standard 802.1X, fully real and labbable. The `SGT:` field in `show access-session details` is where identity becomes policy. - Deploy authentication in stages: **open (monitor) → low-impact → closed**. Going straight to closed takes down the users who cannot authenticate. - Orchestration needs Catalyst Center + ISE; the mechanisms (802.1X, SGT, LISP registration) are real and standalone-verifiable. We describe the flow honestly and point to the real captures. Next: [SD-Access border handoff - IP transit, SDA transit, and L2 handoff](https://www.pinglabz.com/sd-access-border-handoff/). The full cluster index lives on the [network virtualization pillar](https://www.pinglabz.com/network-virtualization/), cross-linked to [802.1X](https://www.pinglabz.com/802-1x/). ### SD-Access Fabric Roles: Edge, Border, Control Plane, and Fabric in a Box URL: https://www.pinglabz.com/sd-access-fabric-roles/ Last updated: 2026-07-12T09:03:23.000Z An SD-Access fabric is not a pile of switches. It is a set of *roles*, and every switch in the fabric plays one or more of them. Understand the roles and the fabric makes sense; miss them and it looks like an impenetrable tangle of LISP, VXLAN, and acronyms. This is the article that makes the rest of SD-Access legible. This is a concept-and-architecture piece. A full fabric needs Cisco Catalyst Center, which a home lab does not have, so we describe the roles and the protocols they run, and are clear about which pieces are real, standalone-verifiable technology (LISP, VXLAN) and which need the controller. For the wider context, see the [network virtualization pillar](https://www.pinglabz.com/network-virtualization/). ## The four roles Edge node The access switch endpoints plug into. It is the fabric's on-ramp: it detects an endpoint, registers it with the control plane, and encapsulates its traffic into VXLAN. The anycast gateway for endpoints lives here. This is where users and devices actually connect. Control plane node The fabric's directory. It runs the LISP Map-Server / Map-Resolver and holds the database of "which endpoint is behind which edge node". Every edge node registers its endpoints here and queries here to find remote endpoints. This is the brain. Border node The fabric's exit. It connects the fabric to everything outside it - the data centre, the WAN, the internet, other fabrics. It translates between the fabric's VXLAN/LISP world and the ordinary routed world beyond. This is the door. Fabric in a Box All three roles - edge, control plane, and border - collapsed onto a single switch (or stack). For a small site that does not justify separate devices. One box, whole fabric. ## How an endpoint's traffic actually flows The roles come alive when you trace a packet. Say host A (on edge node E1) wants to talk to host B (on edge node E2): 1. **Registration.** When host A first appears, E1 detects it and *registers* it with the control plane node: "endpoint A is behind me (E1's loopback)." E2 does the same for host B. The control plane now holds both mappings. 2. **Lookup.** Host A sends a packet to host B. E1 does not know where B is, so it *queries* the control plane node: "where is B?" The control plane answers: "B is behind E2's loopback." 3. **Encapsulation.** E1 encapsulates host A's packet in VXLAN, addressed to E2's loopback, and sends it across the underlay. The VXLAN header carries the virtual network ID (which VN B is in) and the SGT (host A's security group). 4. **Decapsulation.** E2 receives the VXLAN packet, strips the header, and delivers the original frame to host B - after checking the SGT against policy. That is the whole fabric in one paragraph. **Edge nodes register and encapsulate; the control plane node answers "where is it"; the border node handles anything leaving the fabric.** Everything else is detail. ## The protocols behind each role Control plane = LISP The Map-Server/Map-Resolver *is* LISP. Edge nodes are LISP xTRs (ingress/egress tunnel routers) that register EIDs (endpoints) and query for RLOCs (locations). LISP is fully real and runnable on standalone IOS XE. Data plane = VXLAN Edge nodes encapsulate in VXLAN (specifically VXLAN-GPO, which carries the SGT). The VNI identifies the virtual network. Real, dissectable encapsulation. Policy = TrustSec / SGT The SGT rides in the VXLAN header, so policy follows the endpoint anywhere in the fabric. Enforced at the edge node on egress. ## LISP mapping, made concrete LISP's core idea is **separating identity from location**. An endpoint's IP address (its EID, endpoint identifier) says *who* it is. The edge node's loopback (its RLOC, routing locator) says *where* it is. The control plane's whole job is maintaining the EID-to-RLOC mapping. On a real IOS XE LISP xTR - the same technology an SD-Access edge node runs - you can see this mapping directly: ``` xTR# show ip lisp map-cache LISP IPv4 Mapping Cache, 2 entries 10.1.1.0/24, uptime: 00:05:12, expires: 23:54:47, via map-reply, complete Locator Uptime State Pri/Wgt 2.2.2.2 00:05:12 up 10/10 xTR# show ip lisp database LISP ETR IPv4 Mapping Database, LSBs: 0x1, 2 entries 10.1.2.0/24, locator-set RLOC-SET Locator Pri/Wgt Source State 1.1.1.1 10/10 cfg-addr site-self, reachable ``` The **map-cache** is "where I have learned remote endpoints live" (like an ARP cache for locations). The **database** is "the endpoints I am responsible for and will register." An SD-Access edge node maintains exactly these structures, populated by Catalyst Center's provisioning rather than by hand - but the underlying LISP mechanism is identical, and it is real, and PingLabz has captured it on IOS XE in the standalone LISP labs (see the network virtualization cluster). ## Design: where to put each role - **Edge nodes** are your access switches. As many as you have access closets. No decision here - endpoints connect where users are. - **Control plane nodes** should be **redundant** (at least two) and sized for the endpoint count, since every registration and lookup goes through them. Often co-located with the border on larger switches, or dedicated on very large fabrics. - **Border nodes** come in flavours by what they connect to - which is the whole subject of the next article. Also deploy at least two for redundancy. - **Fabric in a Box** for a small site (a branch, a remote office) where three separate devices is overkill. One switch, all roles, simpler to operate. The redundancy point matters: the control plane node is the fabric's directory, and a fabric with one control plane node has a single point of failure for every endpoint lookup. Always deploy two. ## What needs Catalyst Center, and what does not Being honest about the boundary: - **Needs the controller:** assigning roles to devices, provisioning the LISP/VXLAN configuration, managing the fabric as a system, host onboarding policy. There is no CLI workflow for "make this switch a border node" - Catalyst Center does it. - **Real and standalone:** LISP (map-cache, database, xTR behaviour), VXLAN encapsulation, TrustSec SGT tagging. Every protocol the fabric rides on can be configured and verified by hand on IOS XE, and PingLabz has done exactly that in the standalone component labs. We will not show you a Catalyst Center screenshot we did not take, or fabric CLI we cannot produce. Where the technology is real and standalone, we show it. Where it needs the controller, we describe the architecture and say so plainly. That honesty is the differentiator - an expert reader knows immediately whether a "lab" is real. ## Key takeaways - An SD-Access fabric is a set of roles: **edge** (endpoints connect, encapsulate), **control plane** (the LISP directory), **border** (the exit), and **fabric in a box** (all three on one switch, for small sites). - Traffic flow: the edge registers endpoints with the control plane, queries it to find remote endpoints, then VXLAN-encapsulates to the destination edge's loopback. The border handles anything leaving the fabric. - The control plane is LISP; the data plane is VXLAN (carrying the VNI and SGT); the policy plane is TrustSec. - LISP separates identity (the endpoint's EID) from location (the edge's RLOC). The map-cache and database are real, standalone-verifiable structures on IOS XE. - Deploy control plane and border nodes redundantly. The control plane node is the fabric's directory - never run just one. - Role assignment and fabric provisioning need Catalyst Center; the underlying LISP/VXLAN/TrustSec protocols are real and labbable standalone. We are clear about which is which. Next: [SD-Access host onboarding - from port to policy](https://www.pinglabz.com/sd-access-host-onboarding/). The full cluster index lives on the [network virtualization pillar](https://www.pinglabz.com/network-virtualization/). ### SD-Access Underlay: Manual vs LAN Automation URL: https://www.pinglabz.com/sd-access-underlay-lan-automation/ Last updated: 2026-07-12T09:03:22.000Z Before an SD-Access fabric can do anything clever - before LISP, before VXLAN, before a single policy - it needs one boring, rock-solid thing underneath it: an IP network where every fabric node can reach every other fabric node's loopback. That is the underlay, and it is the part of SD-Access that is not magic. It is just a well-built IGP. This article covers what the underlay is, the two ways to build it (manual and LAN Automation), and why the choice matters. It is an honest concept-and-components piece: a full SD-Access fabric requires Cisco Catalyst Center, which a home lab does not have, so we are clear throughout about what needs the controller and what is plain networking you can verify yourself. For the wider context, see the [network virtualization pillar](https://www.pinglabz.com/network-virtualization/). ## The planes of SD-Access SD-Access separates the network into four planes, and the underlay lives beneath all of them. It is worth having the whole model in mind, because every article in this series touches one of these: Management plane **Cisco Catalyst Center.** Provisions and operates the fabric. This is the part you cannot lab without the appliance. Control plane **LISP.** Maps endpoint identity to location. Runnable standalone on IOS XE - real captures exist. Data plane **VXLAN.** Encapsulates endpoint traffic across the fabric, carrying the VNI and the SGT. Policy plane **TrustSec / SGTs.** Segmentation by security group, independent of IP address. **The underlay sits beneath all four.** It is the plain IP transport that carries the LISP control messages, the VXLAN-encapsulated data, and everything else. If the underlay is not solid, nothing above it works, and no amount of Catalyst Center will save you. ## What the underlay must provide The requirements are refreshingly simple: - **Loopback reachability.** Every fabric node has a loopback, and every node must be able to reach every other node's loopback. VXLAN tunnels and LISP sessions are built loopback-to-loopback. - **Low convergence.** When a link fails, the underlay must reconverge fast, because the fabric overlay depends on it. - **A large enough MTU.** VXLAN adds 50 bytes of encapsulation. The underlay must carry jumbo frames (9100 bytes is the Cisco recommendation) so a full-size endpoint frame plus VXLAN overhead does not fragment. **This is the single most common underlay mistake** \- forget the MTU and everything looks fine until a large packet silently drops. - **Nothing else.** The underlay carries no endpoint traffic directly. Endpoints live in the overlay. The underlay only ever carries fabric-node-to-fabric-node traffic. ## Manual underlay vs LAN Automation Manual underlay You build the IGP yourself - your choice of protocol, your addressing, your design. Catalyst Center then discovers the devices and provisions the fabric on top. Full control, more work. LAN Automation Catalyst Center builds the underlay for you - zero-touch. It PXE-boots new switches, assigns addressing from a pool, and configures IS-IS as the underlay IGP. Fast, consistent, opinionated. ### Why LAN Automation uses IS-IS This surprises people who expected OSPF. LAN Automation provisions **IS-IS**, and the reasons are sound: - **IS-IS runs directly over Layer 2** (CLNS), not inside IP. That means it can form adjacencies and exchange topology *before* the IP addressing is fully configured - which is exactly what you need during a zero-touch, PXE-boot provisioning flow where the device does not yet have its addresses. - **IS-IS is transport-agnostic.** It carries IPv4 and IPv6 reachability as TLVs in the same protocol, so a dual-stack underlay is one IGP, not two. - **It scales cleanly** and has a long service-provider pedigree for exactly this kind of flat, fast underlay. If you build the underlay manually you can use OSPF, EIGRP, or IS-IS - Catalyst Center does not care, it just needs loopback reachability. But if you let LAN Automation build it, you get IS-IS, and you should understand it. (IS-IS fundamentals are worth a refresher if you have only ever run OSPF; the concepts map closely.) ## What you can verify without Catalyst Center This is where we are honest about the lab. You cannot build an SD-Access fabric in CML - there is no Catalyst Center to provision it. But the underlay is *plain networking*, and every piece of it is verifiable on real IOS XE: - **The IGP.** An IS-IS or OSPF underlay with loopback reachability is a standard lab you can build and verify with `show ip route` and loopback-to-loopback pings. Nothing about it is SD-Access-specific. - **The MTU.** `show interface | include MTU` and a large ping with the do-not-fragment bit (`ping x.x.x.x size 9000 df-bit`) prove the underlay carries jumbo frames - the exact test that catches the number-one underlay mistake. - **Fast convergence.** BFD, tuned IGP timers - all standard, all verifiable. What you *cannot* do without Catalyst Center is the LAN Automation flow itself (the PXE-boot, the pool assignment, the automatic IS-IS provisioning) and the fabric overlay that sits on top. We will not show you invented Catalyst Center screenshots. Where a component is real and standalone - like the underlay IGP - build it and verify it. Where it needs the controller, we describe the architecture and are clear that is what we are doing. ## Underlay design principles 1. **Point-to-point routed links only.** No spanning tree, no Layer 2 in the underlay. Every fabric-node interconnect is a routed /30 or /31\. The overlay handles all the Layer 2 semantics; the underlay is pure routing. 2. **A loopback per node,** advertised as a /32, reachable everywhere. This is the fabric's addressing anchor. 3. **Jumbo MTU everywhere.** 9100 bytes. Set it and test it before you provision the fabric. 4. **Fast convergence.** BFD on the underlay links. The overlay's resilience is capped by the underlay's convergence time. 5. **Keep it dumb.** The underlay should be the simplest possible IP network. All the intelligence belongs in the overlay. Resist the urge to add policy, QoS complexity, or filtering in the underlay - it makes the fabric harder to reason about for no benefit. ## Key takeaways - The underlay is the plain IP network beneath the fabric - loopback-to-loopback reachability for every node, and nothing else. It is not magic; it is a well-built IGP. - SD-Access has four planes: management (Catalyst Center), control (LISP), data (VXLAN), policy (TrustSec/SGT). The underlay sits beneath all four. - **Build it manually** (any IGP, full control) or with **LAN Automation** (zero-touch, Catalyst Center provisions IS-IS). - LAN Automation uses IS-IS because it runs over Layer 2, works before IP is configured, and carries IPv4/IPv6 in one protocol - ideal for zero-touch provisioning. - The number-one underlay mistake is **MTU**. VXLAN adds \~50 bytes; the underlay needs jumbo frames (9100). Test with a large DF-bit ping. - The underlay IGP is fully labbable and verifiable on IOS XE. The LAN Automation flow and the fabric overlay need Catalyst Center - we say so rather than faking it. Next: [SD-Access fabric roles - edge, border, control plane, and fabric in a box](https://www.pinglabz.com/sd-access-fabric-roles/). The full cluster index lives on the [network virtualization pillar](https://www.pinglabz.com/network-virtualization/). ### Expert Transport Troubleshooting: MPLS and DMVPN Ticket Scenarios URL: https://www.pinglabz.com/expert-transport-troubleshooting/ Last updated: 2026-07-12T08:55:07.000Z Transport technologies fail in layers, and the art of troubleshooting them is knowing which layer to suspect. An MPLS L3VPN has an IGP, LDP, MP-BGP, VRFs, and a label stack, any one of which can break while the others look fine. A DMVPN has an underlay, NHRP, IPsec, mGRE, and a routing protocol on top - five moving parts that all have to work for a spoke to reach a spoke. This article is five transport faults, each grounded in the real behaviour we saw building the MPLS and DMVPN labs for this series, with the diagnostic commands that isolate each layer. For the theory, see the [MPLS guide](https://www.pinglabz.com/mpls/) and the [DMVPN guide](https://www.pinglabz.com/dmvpn/). ## The method: isolate the layer **For MPLS L3VPN,** work bottom-up: **1\. IGP** Can every PE reach every other PE's loopback? `show ip route`, `ping` loopback to loopback. **2\. LDP** Is there a label-switched path between the PE loopbacks? `show mpls ldp neighbor`, `show mpls forwarding-table`. **3\. MP-BGP** Is the VPNv4/VPNv6 session up and exchanging prefixes? `show bgp vpnv4 unicast all summary`. **4\. VRF / RT** Are prefixes importing into the right VRF? `show ip route vrf` vs `show bgp vpnv4 unicast all`. **For DMVPN,** also bottom-up: underlay reachability → NHRP registration → IPsec SA → mGRE tunnel → routing protocol. A break at any layer looks like "the tunnel is down" from the top. ## Ticket 1: "The customer VPN route is in BGP but not forwarding" **Symptom.** A remote customer prefix appears in `show bgp vpnv4 unicast all` on the PE, imports into the VRF, but traffic to it is dropped. **Diagnosis.** The route is present, so BGP and RT import are fine. The problem is the label stack. Check whether the VPN prefix has a valid label path: ``` PE1#show mpls forwarding-table vrf CUST-A 2001:DB8:BB::22/128 Local Outgoing Prefix Bytes Label Outgoing Next Hop Label Outgoing or Tunnel Id Switched interface None 19 2001:DB8:BB::22/128[V] 0 Et0/2 10.0.12.2 ``` Here the out-label is 19 and there is an outgoing interface - good. If instead you see `No Label` or no entry, the VPN label was assigned by BGP but there is **no transport LSP to carry it**. That means LDP is broken between the PEs. **Cause and fix.** Almost always a broken LDP LSP to the remote PE loopback. Check `show mpls ldp neighbor` and confirm the loopback is a /32 in the IGP with an LDP binding. A very common cause is `mpls ip` missing on a core interface, or the PE loopback advertised as something other than a /32 (LDP will not build a host LSP to a summarised loopback). The two-label stack needs *both* labels; a present VPN label with a missing transport label forwards nowhere. ## Ticket 2: "Two customer sites with the same AS can't reach each other" **Symptom.** Site A and Site B of the same customer, both AS 65100, cannot exchange routes over the L3VPN. Everything else works. **Diagnosis.** The route is being dropped by BGP's AS-path loop check. On the receiving CE it simply is not there: ``` CE2#show ip route 11.11.11.11 % Network not in table ``` And the PE facing that CE shows the drop in its policy counters: ``` PE2#show ip bgp vpnv4 vrf CUST-A neighbors 10.2.2.2 | section Policy Local Policy Denied Prefixes: -------- ------- Bestpath from this peer: 4 n/a ``` **Cause.** The customer reuses AS 65100 at both sites. A route from Site A arrives at Site B with 65100 already in its AS-path, and Site B rejects it as a loop. **Fix.** `as-override` on the PE (rewrites the customer AS to the provider AS) or `allowas-in` on the CE (tolerates the own-AS in the path). Full detail, including the SoO you may need on top for dual-homed sites, in [BGP as the PE-CE protocol](https://www.pinglabz.com/bgp-pe-ce-as-override-allowas-in/). ## Ticket 3: "The DMVPN tunnel is up but spokes can't reach each other" **Symptom.** `show dmvpn` shows all peers up, the hub is reachable, but spoke-to-spoke traffic fails. **Diagnosis.** Check what the spoke thinks the next hop to the other spoke is: ``` SPOKE1#show ip route 12.12.12.12 * 10.0.0.12, from 10.0.0.2, via Tunnel0 ``` If the next hop is the *other spoke* (10.0.0.12) - good, that is Phase 3 working. If the next hop is the *hub* (10.0.0.1), the hub is not preserving the spoke as the next hop. **Cause and fix.** Missing `no ip next-hop-self eigrp` on the hub's tunnel. Without it, the hub advertises other spokes' routes with itself as the next hop, so all spoke-to-spoke traffic hairpins through the hub (and if the hub blocks it, fails entirely). Also confirm `no ip split-horizon eigrp` on the hub, or spokes never learn each other's routes at all. Both live on the hub tunnel - see [EIGRP over DMVPN](https://www.pinglabz.com/eigrp-over-dmvpn-multi-hub/). ## Ticket 4: "The DMVPN failed over its routing but traffic still black-holes" **Symptom.** A dual-hub DMVPN. HUB1 fails. EIGRP reconverges via HUB2 - the routing table shows a valid path. And spoke-to-spoke traffic is dead anyway. ``` SPOKE1#ping 10.2.2.1 source Loopback0 repeat 20 .................... Success rate is 0 percent (0/20) ``` **Diagnosis.** The routing looks fine, so look at NHRP, not the RIB: ``` SPOKE1#show dmvpn 2 100.64.1.1 10.0.0.1 UP 00:03:18 S 10.0.0.12 UP 00:01:24 I2 1 100.64.2.1 10.0.0.2 UP 00:03:19 S ``` The `I2` state on the spoke-to-spoke entry (10.0.0.12) is the tell. It is a stale, incomplete shortcut - the direct tunnel to SPOKE2 was resolved *through HUB1*, and HUB1 is now dead, so the mapping is broken. **Cause.** The NHRP holdtime was left at its default of 7200 seconds. The stale shortcut would not age out for two hours. **Fix, immediate:** ``` SPOKE1#clear ip nhrp SPOKE1#ping 10.2.2.1 source Loopback0 repeat 5 !!!!! Success rate is 100 percent (5/5) ``` **Fix, permanent:** `ip nhrp holdtime 300` on every tunnel, so stale shortcuts age out in five minutes and re-resolve via the surviving hub automatically. This is the single most common reason a "redundant" dual-hub DMVPN does not actually fail over. Full walkthrough in [dual-hub DMVPN design](https://www.pinglabz.com/dual-hub-dmvpn-design/). ## Ticket 5: "The IPsec tunnel won't come up after a config change" **Symptom.** After changing the crypto config on a DMVPN tunnel, the tunnel is down and will not recover. **Diagnosis.** Check whether the tunnel was administratively shut by the change itself. When you change the IPsec protection profile on IOS XE, it tells you exactly what it did: ``` interface Tunnel0 tunnel protection ipsec profile IPSEC-IKEV2 % Shutting down Tunnel0 interface due to IPsec tunnel protection modification. % Please run "no shutdown" after config change to bring up the interface. ``` **Cause.** Changing the tunnel protection profile forces IOS XE to bounce the tunnel, and it stays shut until you bring it back up. If you missed the message, the tunnel is simply administratively down. **Fix.** ``` interface Tunnel0 shutdown no shutdown ``` If the tunnel still will not come up after the bounce, the crypto itself is mismatched. Check the IKE version and parameters on both ends: `show crypto ikev2 sa` (state should be `READY`) or `show crypto isakmp sa` (state `QM_IDLE`). A spoke on IKEv2 and a hub on IKEv1 will never negotiate. Detail in [DMVPN with IKEv1 vs IKEv2](https://www.pinglabz.com/dmvpn-ikev1-vs-ikev2/). ## The commands, collected ``` ! MPLS L3VPN - bottom-up show ip route (IGP) show mpls ldp neighbor (LDP) show mpls forwarding-table (label stack) show bgp vpnv4 unicast all summary (MP-BGP) show bgp vpnv4 unicast all labels (VPN labels) show ip route vrf CUST-A (VRF import) show ip bgp vpnv4 vrf CUST-A neighbors x | section Policy (drops) ! DMVPN - bottom-up ping (underlay) show dmvpn (NHRP state - watch for I2) show crypto ikev2 sa / show crypto isakmp sa (IPsec) show ip nhrp (mappings) show ip route (routing + next hop) show ip eigrp neighbors (routing adjacency) ``` ## Key takeaways - Transport faults are layered. **Isolate the layer** before you fix anything - IGP, LDP, MP-BGP, VRF for MPLS; underlay, NHRP, IPsec, routing for DMVPN. - VPN prefix present but not forwarding = a broken transport LSP. The two-label stack needs both labels; check LDP. - Same-AS customer sites failing = BGP AS-path loop drop. Fix with as-override (PE) or allowas-in (CE). - Spoke-to-spoke failing with all peers up = missing `no ip next-hop-self` (or split-horizon) on the hub tunnel. - Dual-hub DMVPN that reconverges routing but still black-holes = a stale NHRP shortcut (`I2`) from the default 7200s holdtime. Set `ip nhrp holdtime 300`. - IPsec tunnel down after a config change = IOS XE bounced it on the protection-profile change. `shutdown` / `no shutdown`, then check for an IKE version mismatch. That closes the expert transport series, and with it the CCIE transport domain. The full cluster indexes live on the [MPLS pillar](https://www.pinglabz.com/mpls/) and the [DMVPN pillar](https://www.pinglabz.com/dmvpn/). ### MPLS and DMVPN Together: Choosing and Combining Enterprise Transports URL: https://www.pinglabz.com/mpls-vs-dmvpn-enterprise-transport/ Last updated: 2026-07-12T08:55:06.000Z Almost every mid-to-large enterprise WAN ends up running both MPLS and DMVPN. Not because someone planned it that way, but because they solve different problems, and a real network has both problems. MPLS gives you a private, SLA-backed, any-to-any core. DMVPN gives you cheap, encrypted, dynamic connectivity over anything with an internet connection. The interesting question is not "which one" - it is "how do they fit together." This article compares the two honestly and shows the standard hybrid designs. It is the synthesis of the transport series. For the fundamentals, see the [MPLS guide](https://www.pinglabz.com/mpls/) and the [DMVPN guide](https://www.pinglabz.com/dmvpn/). ## What each one actually is MPLS L3VPN A service you *buy* from a provider. They run the core; you hand them routes at each site and they deliver any-to-any connectivity with an SLA. Private, no encryption needed on the wire (it is not the public internet), QoS honoured end to end. DMVPN An overlay you *build* yourself over any IP transport - usually the public internet. Encrypted (IPsec), dynamic spoke-to-spoke, cheap bandwidth, but no SLA and QoS is best-effort once packets hit the internet. The distinction that matters: **MPLS is a purchased service with a contract; DMVPN is a design you own end to end.** That shapes everything - cost, control, SLA, and who you call when it breaks. ## The honest comparison **Cost** MPLS: expensive per megabit, especially internationally. DMVPN over internet: cheap, sometimes a tenth of the cost for the same bandwidth. **SLA** MPLS: contractual latency, loss, jitter guarantees. DMVPN: best-effort - the internet promises nothing. **QoS** MPLS: honoured end to end by the provider. DMVPN: you can mark and shape at your edge, but the internet ignores your DSCP. **Encryption** MPLS: not encrypted (private core) - add it yourself if compliance requires. DMVPN: IPsec built in. **Provisioning** MPLS: weeks to months for a new circuit. DMVPN: a new spoke is up in an hour over any internet connection. **Control** MPLS: the provider owns the core; you troubleshoot through a ticket. DMVPN: you own every hop of the overlay. **Any-to-any** MPLS: native - the provider mesh gives it for free. DMVPN: dynamic spoke-to-spoke (Phase 3) achieves it, with a resolution step. Neither wins outright. MPLS buys you performance guarantees and offloads the core; DMVPN buys you cost, speed of deployment, and control. Which is why the answer for most enterprises is not one or the other. ## The hybrid designs ### 1\. MPLS primary, DMVPN backup The most common pattern. Each site has an MPLS circuit and an internet circuit. MPLS carries production traffic; a DMVPN overlay over the internet sits ready as backup. If the MPLS circuit fails, routing shifts to the DMVPN. ``` ! Prefer MPLS: higher local-pref / lower metric on the MPLS-learned routes, ! DMVPN routes as the less-preferred backup path. ! On failure, the backup path takes over automatically. ``` You get the MPLS SLA for normal operation and the DMVPN as a cheap insurance policy - far cheaper than a second MPLS circuit for redundancy. The design work is making the failover clean: route preference so MPLS wins when both are up, and fast failure detection (BFD, IP SLA tracking) so the switch to DMVPN is quick. ### 2\. Active/active with application-aware routing Use both circuits simultaneously, and steer traffic by application. Latency-sensitive, SLA-critical traffic (voice, video, ERP) over MPLS; bulk, backup, and internet-bound traffic over the DMVPN. This is exactly the problem [SD-WAN](https://www.pinglabz.com/sd-wan/) was built to automate - application-aware routing making these decisions per-flow, per-SLA, dynamically. A hand-built version uses PBR and route-maps; SD-WAN does it as a policy. ### 3\. DMVPN over MPLS (yes, really) Run a DMVPN *on top of* your MPLS service. Why? To get encryption over MPLS (for compliance) without asking the provider for it, to get dynamic spoke-to-spoke over an MPLS L3VPN that would otherwise route everything through a hub, or to have a single consistent overlay whether the underlay is MPLS or internet. The MPLS L3VPN becomes just another transport for the DMVPN cloud. ### 4\. Dual DMVPN clouds over dual transports Two DMVPN clouds, one over MPLS, one over internet, is the classic dual-cloud design. Each transport is a separate DMVPN, and routing (or SD-WAN policy) chooses between them. This is the fully self-owned version of the hybrid, with no reliance on the provider for anything but raw transport. ## Where SD-WAN fits Everything above - the failover, the application steering, the transport selection - is manual policy in a traditional design. You configure route preferences, PBR, IP SLA trackers, and you maintain them. It works, and CCIE candidates need to know how to build it by hand. SD-WAN productises exactly this. It abstracts "MPLS and DMVPN and internet" into "transports", measures each one's real-time performance, and applies application-aware routing as a centralised policy. The hybrid designs above are precisely the problem SD-WAN automates - which is why understanding them by hand is the foundation for understanding what SD-WAN is doing for you. (The SD-WAN cluster covers this in depth.) ## Choosing, in practice 1. **Critical site, SLA-bound apps, budget available:** MPLS primary, DMVPN internet backup. The standard enterprise branch. 2. **Cost-sensitive, many sites, tolerant apps:** DMVPN over internet as the primary, no MPLS. Common for retail, remote offices. 3. **Compliance requires encryption everywhere:** DMVPN over MPLS, or DMVPN everywhere. 4. **You want per-application transport selection without hand-crafting it:** SD-WAN over both transports. 5. **Global sites where MPLS is prohibitively expensive:** DMVPN over internet, accepting best-effort, possibly with SD-WAN measuring path quality to route around brownouts. ## Key takeaways - MPLS is a **purchased service** \- private, SLA-backed, QoS honoured, expensive, provider-controlled. DMVPN is a **self-built overlay** \- encrypted, cheap, dynamic, best-effort, fully under your control. - Most enterprises run both, because they solve different problems. - The most common hybrid is **MPLS primary, DMVPN internet backup** \- the SLA for production, cheap insurance for failure. - Other patterns: active/active with application steering, DMVPN *over* MPLS (for encryption or dynamic mesh), and dual DMVPN clouds over dual transports. - All of these are manual policy in a traditional design (route preference, PBR, IP SLA tracking) and are exactly what **SD-WAN automates**. Understanding them by hand is the foundation for understanding SD-WAN. - Choose by app sensitivity, budget, compliance, and how much of the transport decision you want to automate. Next: [expert transport troubleshooting - MPLS and DMVPN ticket scenarios](https://www.pinglabz.com/expert-transport-troubleshooting/). The full cluster indexes live on the [MPLS pillar](https://www.pinglabz.com/mpls/) and the [DMVPN pillar](https://www.pinglabz.com/dmvpn/). ### DMVPN with IKEv1 vs IKEv2: Configuration and Migration URL: https://www.pinglabz.com/dmvpn-ikev1-vs-ikev2/ Last updated: 2026-07-12T08:55:06.000Z Every DMVPN carries IPsec, and every IPsec deployment sits on one of two key-exchange protocols: IKEv1, the original from 1998, or IKEv2, its 2005 replacement. Most networks still run IKEv1 because it works and nobody wanted to touch a working VPN. But IKEv2 is faster to negotiate, cleaner to configure, more resilient, and it is where everything is heading. This article configures both on the same DMVPN in a CML lab, shows the real SA output from each side by side, and walks the migration - which, it turns out, is a single-line profile swap. For the fundamentals, start at the [complete DMVPN guide](https://www.pinglabz.com/dmvpn/). ## Why IKEv2 is better **Fewer messages** IKEv1 main mode takes 6 messages to build phase 1 (aggressive mode 3, but insecure). IKEv2 does the equivalent in 4 (IKE\_SA\_INIT + IKE\_AUTH). On a large DMVPN with hundreds of spokes re-keying, that difference adds up. **Built-in dead peer detection** IKEv2 has liveness checks in the protocol. IKEv1 needs DPD bolted on. **NAT traversal and resilience** IKEv2 handles NAT and reconnection (MOBIKE) far more gracefully - it matters for spokes behind NAT on the public internet. **Asymmetric authentication** IKEv2 lets each side authenticate differently (one PSK, one certificate). IKEv1 requires both ends to use the same method. ## IKEv1 configuration IKEv1 uses ISAKMP policies. The tunnel protection references an IPsec profile that has no explicit IKE binding - IKEv1 is the default: ``` crypto isakmp policy 10 encryption aes 256 hash sha authentication pre-share group 14 lifetime 86400 ! crypto isakmp key PingLabzIKEv1 address 0.0.0.0 ! crypto ipsec transform-set TS-IKEV1 esp-aes 256 esp-sha-hmac mode transport ! crypto ipsec profile IPSEC-IKEV1 set transform-set TS-IKEV1 ! interface Tunnel0 tunnel protection ipsec profile IPSEC-IKEV1 ``` ### IKEv1 verification ``` HUB1#show crypto isakmp sa dst src state conn-id status 100.64.1.1 100.64.11.1 QM_IDLE 1006 ACTIVE 100.64.1.1 100.64.12.1 QM_IDLE 1007 ACTIVE 100.64.1.1 100.64.2.1 QM_IDLE 1005 ACTIVE ``` The tell of a healthy IKEv1 SA is `QM_IDLE` (Quick Mode idle - phase 1 complete, phase 2 negotiated, ready). The detail: ``` HUB1#show crypto isakmp sa detail | include 100.64.11.1 1006 100.64.1.1 100.64.11.1 ACTIVE aes sha psk 14 23:58:46 ``` AES, SHA, PSK, DH group 14, lifetime counting down from 24 hours. And traffic is flowing through it: ``` HUB1#show crypto ipsec sa peer 100.64.11.1 | include transform|pkts encaps|pkts decaps #pkts encaps: 29, #pkts encrypt: 29, #pkts digest: 29 #pkts decaps: 27, #pkts decrypt: 27, #pkts verify: 27 transform: esp-256-aes esp-sha-hmac , ``` ## IKEv2 configuration IKEv2 replaces the flat ISAKMP policy with a keyring and a profile - more objects, but cleaner separation of concerns: ``` crypto ikev2 keyring KR peer ANY address 0.0.0.0 0.0.0.0 pre-shared-key PingLabzIKEv2 ! crypto ikev2 profile IKEV2-PROF match identity remote address 0.0.0.0 authentication local pre-share authentication remote pre-share keyring local KR ! crypto ipsec transform-set TS-IKEV2 esp-aes 256 esp-sha256-hmac mode transport ! crypto ipsec profile IPSEC-IKEV2 set transform-set TS-IKEV2 set ikev2-profile IKEV2-PROF <-- this line is what makes it IKEv2 ! interface Tunnel0 tunnel protection ipsec profile IPSEC-IKEV2 ``` The one line that makes it IKEv2 rather than IKEv1 is `set ikev2-profile` in the IPsec profile. Without it, IOS falls back to IKEv1\. And note `authentication local` and `authentication remote` are separate directives - that is the asymmetric authentication capability, even though we use PSK on both sides here. ### IKEv2 verification ``` HUB1#show crypto ikev2 sa Tunnel-id Local Remote fvrf/ivrf Status 2 100.64.1.1/500 100.64.11.1/500 none/none READY Encr: AES-CBC, keysize: 256, PRF: SHA512, Hash: SHA512, DH Grp:19, Auth sign: PSK, Auth verify: PSK Life/Active Time: 86400/28 sec Local spi: A0F41498A3310CEE Remote spi: B2BE8D43E865239F ``` The IKEv2 SA state is `READY`, and the output is richer: it explicitly shows the **PRF** (pseudo-random function) as a separate field, the DH group, and separate `Auth sign` / `Auth verify` directions - reflecting IKEv2's cleaner, more explicit design. ## Side by side IKEv1 **Show command:** `show crypto isakmp sa` **Healthy state:** QM\_IDLE **Config:** crypto isakmp policy + key **Phase 1:** 6 messages (main mode) **PRF:** not a separate concept IKEv2 **Show command:** `show crypto ikev2 sa` **Healthy state:** READY **Config:** crypto ikev2 keyring + profile **Phase 1:** 4 messages **PRF:** explicit field Both of these ran on the *same* DMVPN, carrying the *same* traffic, in the same lab. The only thing that changed between them was which IPsec profile the tunnel referenced. ## The migration is a profile swap Here is the practical payoff. Migrating a DMVPN from IKEv1 to IKEv2 does not mean rebuilding the tunnels. Both IPsec profiles can coexist on the router; you just change which one the tunnel references: ``` interface Tunnel0 tunnel protection ipsec profile IPSEC-IKEV2 <-- was IPSEC-IKEV1 % Shutting down Tunnel0 interface due to IPsec tunnel protection modification. % Please run "no shutdown" after config change to bring up the interface. ``` **The catch:** IOS XE forces a tunnel bounce when you change the protection profile. It shuts the tunnel and makes you bring it back up. That is a brief outage on that tunnel, so plan the migration in a maintenance window and do it hub by hub, spoke by spoke. ``` interface Tunnel0 shutdown no shutdown ``` After the bounce, all SAs re-form under the new suite. In the lab we migrated the entire cloud from IKEv2 to IKEv1 (and could reverse it identically) purely by swapping the referenced profile and bouncing each tunnel. ### Migration order 1. **Pre-stage both profiles** on every device. The IKEv2 keyring, profile, transform-set and IPsec profile can all sit there unused while IKEv1 is live. 2. **Migrate the hubs first,** one at a time, so at least one hub always has a working profile a spoke can match. Because the tunnel bounces, doing both hubs at once would drop every spoke. 3. **Then the spokes,** in batches. A spoke on IKEv2 can still reach a hub on IKEv2; do not leave a spoke stranded on a suite no hub offers. 4. **Verify each device** shows the new SA type (`READY` for IKEv2) before moving on. The important sequencing rule: **never leave a spoke and its hub on different IKE versions.** They will not negotiate. Migrate a hub, then the spokes that use it, keeping each spoke-hub pair on the same version throughout. ## Which to use - **New DMVPN?** IKEv2, always. There is no reason to start on IKEv1 in 2026. - **Existing IKEv1 DMVPN, working fine?** Plan a migration to IKEv2 - the operational and security benefits are real - but there is no fire. Do it in windows, hub-first. - **Mixed-vendor or legacy peers?** Verify IKEv2 support on both ends first. Very old gear may only speak IKEv1, and a mismatch simply will not come up. ## Key takeaways - IKEv2 is faster to negotiate (4 messages vs 6), has built-in dead-peer detection, better NAT handling, and asymmetric authentication. Use it for anything new. - IKEv1 uses `crypto isakmp policy` \+ key; a healthy SA shows `QM_IDLE` in `show crypto isakmp sa`. - IKEv2 uses `crypto ikev2 keyring` \+ `profile`; a healthy SA shows `READY` in `show crypto ikev2 sa`, with an explicit PRF field. - The one line that selects IKEv2 is `set ikev2-profile` in the IPsec profile. Omit it and IOS uses IKEv1. - Migration is a **profile swap** on the tunnel - but IOS XE bounces the tunnel when you change the protection profile, so do it in a window. - Migrate hub-first, one device at a time, and never leave a spoke and its hub on different IKE versions - they will not negotiate. Next: [MPLS and DMVPN together - choosing and combining enterprise transports](https://www.pinglabz.com/mpls-vs-dmvpn-enterprise-transport/). The full cluster index lives on the [DMVPN pillar](https://www.pinglabz.com/dmvpn/). ### Dual-Hub DMVPN: Redundancy Designs That Actually Fail Over URL: https://www.pinglabz.com/dual-hub-dmvpn-design/ Last updated: 2026-07-12T08:55:05.000Z Everyone builds a second DMVPN hub for redundancy. Far fewer people test that the failover actually works - and the ones who do are often unpleasantly surprised. A dual-hub DMVPN can have both hubs up, both spokes registered, routing fully converged, and *still* black-hole spoke-to-spoke traffic for two hours after a hub fails, because of a single default timer nobody changed. This article builds a dual-hub DMVPN in a CML lab, fails a hub, and shows exactly what breaks and why - including a genuine, non-obvious failure that we hit and fixed. For the fundamentals, start at the [complete DMVPN guide](https://www.pinglabz.com/dmvpn/). ## Single cloud vs dual cloud Single cloud, dual hub One tunnel interface per spoke, one tunnel subnet, both hubs are NHS on the same mGRE. Simpler, fewer interfaces. Both hubs are on the same cloud, so a spoke is adjacent to both over one tunnel. Our lab. Dual cloud Two tunnel interfaces per spoke, two subnets, one hub each. More config, but full transport separation - the right choice when the two hubs sit on genuinely different WANs (MPLS + internet). Single cloud is the simpler default and is what you want when both hubs are on the same transport. Dual cloud is for when you need to reason about, police, or fail between two independent transports. ## The lab, up and running Two hubs (HUB1 primary, HUB2 backup), two spokes, one INET underlay router, EIGRP over the mGRE cloud, and IPsec protecting every tunnel. All three NHRP peers register with HUB1: ``` HUB1#show dmvpn Type:Hub, NHRP Peers:3, 1 100.64.2.1 10.0.0.2 UP 00:00:31 D 1 100.64.11.1 10.0.0.11 UP 00:00:27 D 1 100.64.12.1 10.0.0.12 UP 00:00:22 D ``` Spoke-to-spoke works: SPOKE1 learns SPOKE2's LAN via both hubs, and the next hop is SPOKE2 *directly* \- not a hub - thanks to `no ip next-hop-self` on the hub tunnels: ``` SPOKE1#show ip route 12.12.12.12 10.0.0.12, from 10.0.0.1, via Tunnel0 * 10.0.0.12, from 10.0.0.2, via Tunnel0 SPOKE1#show ip cef 12.12.12.12 12.12.12.12/32 nexthop 10.0.0.12 Tunnel0 nexthop 10.0.0.12 Tunnel0 ``` Two equal-cost paths, both pointing at SPOKE2 as the next hop. This is a healthy Phase 3 DMVPN. (For the EIGRP-over-mGRE details - `no split-horizon`, `no next-hop-self` \- see [EIGRP over DMVPN](https://www.pinglabz.com/eigrp-over-dmvpn-multi-hub/).) ## The failure that surprised us We shut HUB1's internet link. EIGRP did its job: it dropped the HUB1 neighbour, and the route to SPOKE2 survived via HUB2\. Routing reconverged. Everything looked fine in the routing table. And spoke-to-spoke traffic died completely: ``` SPOKE1#ping 10.2.2.1 source Loopback0 repeat 20 .................... Success rate is 0 percent (0/20) ``` Zero. Not a brief reconvergence blip - total, sustained black-holing, with the routing table showing a perfectly good path via HUB2\. The kind of thing that has you staring at a correct-looking `show ip route` wondering why the pings still fail. ### The cause The direct spoke-to-spoke shortcut to SPOKE2 had been **resolved through HUB1**. When HUB1 died, that NHRP mapping went stale - it still pointed at HUB1 as the resolution path: ``` SPOKE1#show dmvpn 2 100.64.1.1 10.0.0.1 UP 00:03:18 S 10.0.0.12 UP 00:01:24 I2 1 100.64.2.1 10.0.0.2 UP 00:03:19 S ``` See that `I2` state on the SPOKE2 entry - a stale, incomplete shortcut still associated with HUB1's now-dead NBMA address. The routing table said "send to SPOKE2 directly", but the NHRP-to-NBMA mapping needed to build that direct tunnel was broken and pointing at a dead hub. **And here is the killer: it would not fix itself for two hours.** The NHRP holdtime had been left at the default. Until that stale entry aged out, SPOKE1 kept trying to reach SPOKE2 through a resolution that no longer worked. ``` SPOKE1#clear ip nhrp SPOKE1#ping 10.2.2.1 source Loopback0 repeat 5 !!!!! Success rate is 100 percent (5/5) ``` A manual `clear ip nhrp` forced immediate re-resolution via HUB2 and connectivity came straight back. But you cannot rely on manually clearing NHRP on every spoke every time a hub fails. ## The fix: a short NHRP holdtime The default NHRP holdtime is 7200 seconds - two hours. That is far too long for a dual-hub design, because it is exactly how long a stale spoke-to-spoke shortcut resolved via a failed hub can persist. Set it short on every tunnel: ``` interface Tunnel0 ip nhrp holdtime 300 ``` With a 300-second holdtime, a shortcut resolved through a hub that dies ages out within five minutes and re-resolves through the surviving hub automatically. Five minutes is still not instant - during that window, spoke-to-spoke traffic for an already-resolved pair may black-hole - but it is recoverable without human intervention, and for the common case (spoke-to-hub-to-spoke fallback, and newly-resolved shortcuts) failover is much faster. **This is the single most important and most overlooked knob in a dual-hub DMVPN.** The redundancy you built is only real if the NHRP holdtime lets stale shortcuts clear in a reasonable time. Every dual-hub design should set it explicitly; the default will bite you exactly when you least want it to - during a hub failure. ## Making one hub primary With both hubs advertising equal metrics, spokes load-share, which is usually not the intent. Make HUB1 primary and HUB2 a true backup with a delay adjustment or an offset list on the backup hub - one command, one router, applies to every spoke: ``` ! On HUB2, make its advertised routes less attractive router eigrp 100 ! (named mode) offset-list, or on the interface: interface Tunnel0 delay 10000 ``` Now spokes prefer HUB1 and keep HUB2 as a feasible successor - ready for instant EIGRP failover for the *routing*, while the NHRP holdtime governs how fast stale shortcuts recover. Both mechanisms matter, and they are independent. (Offset-list detail in [EIGRP offset lists](https://www.pinglabz.com/eigrp-offset-lists/).) ## The two independent failover mechanisms This is the mental model that makes dual-hub DMVPN make sense. There are *two* things that have to fail over, and they are governed by different timers: **Routing (the RIB path)** Governed by EIGRP hello/hold timers (and BFD if configured). Fast - seconds. This is what most people test, and it works. **NHRP resolution (the tunnel)** Governed by NHRP holdtime. Slow by default - two hours. This is what most people *don't* test, and it is why "redundant" DMVPNs black-hole after a hub failure. A DMVPN that fails over its routing in three seconds but keeps a stale NHRP shortcut for two hours has effectively not failed over at all for that spoke pair. Test both. Tune both. ## Tune the EIGRP and IPsec timers too - **EIGRP hello/hold on the tunnel:** the default on a multipoint interface is 60/180 - three minutes to notice a dead hub. Tune to 5/15, or pair with BFD, so the routing side fails over fast. - **IPsec DPD (Dead Peer Detection):** so a spoke notices a dead hub's crypto session promptly and tears down the SA, rather than holding it until the SA lifetime expires. - **NHRP registration timer:** spokes re-register with the NHS at one-third of the holdtime by default, so a shorter holdtime also means more frequent registration - a reasonable trade for faster recovery. ## Key takeaways - A dual-hub DMVPN has **two** failover mechanisms: routing (EIGRP timers, fast) and NHRP resolution (holdtime, slow by default). Both must work. - We reproduced the classic failure: HUB1 dies, routing reconverges via HUB2, but a spoke-to-spoke shortcut resolved through HUB1 goes stale (`I2`) and black-holes traffic - with a correct-looking routing table. - The cause is the **default 7200-second NHRP holdtime**. Set `ip nhrp holdtime 300` on every tunnel so stale shortcuts age out and re-resolve via the surviving hub automatically. - Make one hub primary with a delay or offset-list adjustment on the backup hub - one command, all spokes. - Tune the EIGRP tunnel hello/hold (default 60/180 is far too slow) and IPsec DPD as well. - **Test the failover, not just the redundancy.** Both hubs up is not the same as failover working. Next: [DMVPN with IKEv1 vs IKEv2](https://www.pinglabz.com/dmvpn-ikev1-vs-ikev2/), using the same lab. The full cluster index lives on the [DMVPN pillar](https://www.pinglabz.com/dmvpn/). ### BGP as the PE-CE Protocol: as-override, allowas-in, and SoO URL: https://www.pinglabz.com/bgp-pe-ce-as-override-allowas-in/ Last updated: 2026-07-12T08:55:04.000Z BGP is the most powerful choice for the PE-CE protocol in an MPLS L3VPN - it scales, it carries policy, and it hands the provider fine-grained control. But it comes with a trap that catches every engineer once: BGP's loop-prevention, the thing that keeps the internet stable, actively breaks a very common enterprise VPN topology. When two of a customer's sites share the same AS number, they cannot reach each other, and BGP is doing it on purpose. This article covers the three tools that fix it - `as-override`, `allowas-in`, and SoO - with real output from a CML lab that reproduces the failure and each fix. For the fundamentals, start at the [complete MPLS guide](https://www.pinglabz.com/mpls/). ## Why customers reuse an AS number An enterprise buys L3VPN service and runs BGP to the provider at every site. The simple thing to do - the thing every enterprise does - is use one private AS number, say 65100, at all of them. Site A is AS 65100, Site B is AS 65100, Site C is AS 65100\. It is their internal number; why would they coordinate different ones? And that is exactly the setup that breaks. ## The failure, reproduced In the lab, CE1 (Site A) and CE2 (Site B) are both AS 65100, connected through the provider (AS 100). CE1 advertises its loopback 11.11.11.11\. That route travels: CE1 (65100) → PE1 → across the VPNv4 core → PE2 → toward CE2\. When it arrives at CE2, its AS-path contains 65100. CE2 is AS 65100\. So CE2 looks at the incoming route, sees its own AS number in the path, and - following BGP's fundamental loop-prevention rule - **silently discards it**: ``` CE2#show ip route 11.11.11.11 % Network not in table CE2#show ip bgp 11.11.11.11 % Network not in table ``` The route is being sent. It is being received. It is being thrown away at parse time, before it ever enters the BGP table. The only evidence is on the PE, in the policy-denied counters (with soft-reconfiguration enabled): ``` PE2#show ip bgp vpnv4 vrf CUST-A neighbors 10.2.2.2 | section Policy Local Policy Denied Prefixes: -------- ------- Bestpath from this peer: 4 n/a ``` Two sites of the same customer, both paying for VPN service, unable to reach each other. BGP is working exactly as designed - it just cannot tell "my own AS looping back through a mistake" from "my own AS legitimately at another site of the same company." ## Fix 1: as-override (on the provider) The cleanest fix, and the one providers prefer, lives on the PE. `as-override` tells the PE: when you advertise a route to a CE, and the CE's AS number appears in the AS-path, **replace those occurrences with your own AS number**. ``` router bgp 100 address-family ipv4 vrf CUST-A neighbor 10.2.2.2 remote-as 65100 neighbor 10.2.2.2 activate neighbor 10.2.2.2 as-override ``` Now the route that reaches CE2 no longer contains 65100 - the PE rewrote it to 100: ``` CE2#show ip bgp 11.11.11.11 BGP routing table entry for 11.11.11.11/32 100 100 10.2.2.1 from 10.2.2.1 (3.3.3.3) Origin IGP, localpref 100, valid, external, best CE2#ping 11.11.11.11 source Loopback0 !!!!! Success rate is 100 percent (5/5) ``` The AS-path is now `100 100` \- CE2's own 65100 replaced by the provider's AS. No loop detected, route accepted, connectivity restored. The customer changed nothing; the fix is entirely on the provider side. **Why providers prefer this:** it works transparently for every customer with reused AS numbers, requires no customer configuration, and keeps the loop-prevention semantics sensible (a genuine loop through the provider would still be caught by the provider's own AS appearing twice, which is why you sometimes pair it with SoO - see below). ## Fix 2: allowas-in (on the customer) The mirror-image fix lives on the CE. `allowas-in` tells the receiving router: accept a route even if my own AS appears in the path, up to N times. ``` router bgp 65100 address-family ipv4 neighbor 10.2.2.1 allowas-in 3 ``` Now CE2 accepts the route despite seeing 65100 in the path: ``` CE2#show ip bgp 11.11.11.11 BGP routing table entry for 11.11.11.11/32 100 65100 10.2.2.1 from 10.2.2.1 (3.3.3.3) Origin IGP, localpref 100, valid, external, best ``` The AS-path here is `100 65100` \- the original path, with 65100 present, but accepted anyway because `allowas-in` told CE2 to tolerate its own AS up to three times. as-override **Where:** the PE (provider) **Does:** rewrites the customer AS to the provider AS in the path **Loop safety:** retained (provider AS still catches real loops) **Preferred by:** providers - one config covers all customers allowas-in **Where:** the CE (customer) **Does:** tolerates own AS in the path up to N times **Loop safety:** weakened - you are disabling a safety check **Used when:** you cannot change the PE (you are the customer) The number matters. `allowas-in 3` means "tolerate my AS up to three times". Set it as low as your topology needs - a customer with three sites reachable through each other might legitimately see its AS a couple of times, but a large number is disabling loop prevention wholesale and inviting a genuine loop. If you are the customer and cannot influence the PE, this is your tool; otherwise prefer as-override. ## Fix 3: SoO - the loop prevention you just removed Here is the subtle problem. Both fixes above *defeat* BGP's loop prevention for the reused-AS case. But what happens when a customer site is **dual-homed** to two PEs? A route can now go out one PE, across the provider, and come back in the other PE to the same site - a real loop, and you have just disabled the mechanism that would have caught it. Site of Origin (SoO) restores loop prevention at a finer granularity. It is an extended community stamped on routes as they enter the VPN from a site, identifying *which site* they came from. A PE will not re-advertise a route back to a site carrying that site's own SoO. From the lab, SoO applied on the PE facing CE2: ``` route-map SOO-CE2 permit 10 set extcommunity soo 100:22 ! router bgp 100 address-family ipv4 vrf CUST-A neighbor 10.2.2.2 route-map SOO-CE2 in ``` And the route from CE2 now carries its site identity: ``` PE2#show ip bgp vpnv4 vrf CUST-A 22.22.22.22/32 Extended Community: SoO:100:22 RT:100:1 ``` If CE2 were dual-homed to a second PE, that PE would see a route arriving carrying `SoO:100:22` \- CE2's own site identity - and refuse to send it back into CE2's site. The backdoor loop that as-override and allowas-in re-opened is closed again. **Note this worked cleanly in the VRF context.** BGP SoO via a route-map in the VPN address family stamps the extended community exactly as expected - unlike EIGRP's `ip vrf sitemap`, which behaves differently in a global table (a distinction we covered in the [EIGRP SoO article](https://www.pinglabz.com/eigrp-site-of-origin-soo/)). ## The decision, in one place 1. **You are the provider, customer reuses AS numbers:** `as-override` on the PE. One config, all customers, loop safety retained. 2. **Customer is dual-homed to multiple PEs:** add **SoO** on top of as-override, to restore the loop prevention as-override removed. 3. **You are the customer and cannot change the PE:** `allowas-in N` on the CE, with N as small as possible. 4. **Ideal, if you can arrange it:** the customer uses *different* AS numbers per site and none of this is needed. Often impractical, but it is the cleanest design. ## Troubleshooting 1. **Two customer sites can't reach each other, same AS?** This exact problem. Check the AS-path of a missing prefix - if the customer's own AS is in it, the receiving router dropped it. 2. **as-override configured but still not working?** It rewrites the AS on *egress from the PE toward the CE*. Confirm it is on the PE-CE neighbour, in the right VRF address family. 3. **allowas-in accepting too much?** The number is too high. Lower it to the minimum your topology requires. 4. **Dual-homed site with a loop after enabling as-override?** You removed loop prevention and need SoO to put it back at the site level. 5. **SoO not stamping?** In a VPN, apply it via a route-map on the PE-CE neighbour in the VRF address family - which does work, as shown above. ## Key takeaways - Two customer sites sharing an AS number cannot reach each other over L3VPN - BGP's loop prevention drops the route silently. It is working as designed. - **as-override** (on the PE) rewrites the customer AS to the provider AS in the path. Providers prefer it: one config, transparent to the customer, loop safety retained. - **allowas-in N** (on the CE) tolerates the customer's own AS in the path N times. Use it when you cannot change the PE, with N as small as possible. - Both fixes weaken loop prevention. For a **dual-homed site**, add **SoO** to restore it at site granularity - a route will not be sent back to the site it came from. - BGP SoO via a route-map in the VRF address family stamps the extended community correctly. - The cleanest design is different AS numbers per site, if the customer can arrange it. Next: [dual-hub DMVPN designs that actually fail over](https://www.pinglabz.com/dual-hub-dmvpn-design/). The full cluster index lives on the [MPLS pillar](https://www.pinglabz.com/mpls/), cross-linked to [BGP](https://www.pinglabz.com/bgp/). ### MPLS VPNv6 and 6VPE: IPv6 L3VPN over an IPv4 Core URL: https://www.pinglabz.com/mpls-vpnv6-6vpe/ Last updated: 2026-07-12T08:55:04.000Z You have an IPv4 MPLS L3VPN running. Your customers each have their own VRF, their own routing table, their own isolated slice of the core. It works, it scales, it makes money. Then a customer asks for IPv6, and you discover that your entire VPN infrastructure - VPNv4, route distinguishers, route targets, the lot - only understands IPv4. 6VPE is the answer. It carries IPv6 L3VPN across the same IPv4 MPLS core, using the same VRFs, the same RDs and RTs, the same MP-BGP sessions - just with an IPv6 address family bolted on. This article builds it in a CML lab and shows the two-label stack that makes it work. For the fundamentals, start at the [complete MPLS guide](https://www.pinglabz.com/mpls/). ## 6VPE vs 6PE: know the difference These two are constantly confused, and the distinction is the whole point. 6PE **Address family:** ipv6 unicast + send-label **Isolation:** none - one global IPv6 table **Labels:** one VPN label per prefix **For:** IPv6 internet transit over a v4 core 6VPE **Address family:** vpnv6 unicast **Isolation:** full - per-VRF, with RDs and RTs **Labels:** two (transport + VPN) **For:** IPv6 L3VPN sold to multiple customers If you covered [6PE](https://www.pinglabz.com/ipv6-bgp-6pe/) already, 6VPE is 6PE plus VRF isolation. It uses the VPNv6 address family (AFI 2, SAFI 128) instead of plain labelled IPv6, and every prefix gets a route distinguisher so overlapping customer address space stays separate. ## The lab A classic MPLS L3VPN: CE1 - PE1 - P - PE2 - CE2\. The core (PE1, P, PE2) is IPv4-only with OSPF and LDP. There is one VRF, CUST-A, with RD 100:1 on PE1 and 100:2 on PE2, and route-target 100:1 both ways. The CE-PE links are dual-stack. Both customer sites even advertise the *same* prefix, 172.16.10.0/24, deliberately - to prove the RD keeps them separate. ### The VRF, now with an IPv6 address family ``` vrf definition CUST-A rd 100:1 address-family ipv4 route-target export 100:1 route-target import 100:1 exit-address-family address-family ipv6 route-target export 100:1 route-target import 100:1 exit-address-family ``` Note that the RD is defined once at the VRF level and applies to both families. The route targets are set per address family (they happen to be the same here, but they need not be). ### The MP-BGP sessions ``` router bgp 100 neighbor 3.3.3.3 remote-as 100 neighbor 3.3.3.3 update-source Loopback0 ! address-family vpnv4 neighbor 3.3.3.3 activate neighbor 3.3.3.3 send-community extended ! address-family vpnv6 neighbor 3.3.3.3 activate neighbor 3.3.3.3 send-community extended ! address-family ipv6 vrf CUST-A neighbor 2001:DB8:1::2 remote-as 65100 neighbor 2001:DB8:1::2 activate neighbor 2001:DB8:1::2 as-override ``` One iBGP session between the PE loopbacks carries *both* VPNv4 and VPNv6\. That is the elegance of it - the same session, the same core, one more address family. The `send-community extended` is mandatory: route targets are extended communities, and without it the whole VPN mechanism silently fails. ## The two-label stack This is what makes 6VPE work and what distinguishes it from 6PE. Look at the labels PE1 has for the remote customer prefixes: ``` PE1#show bgp vpnv6 unicast all labels Network Next Hop In label/Out label Route Distinguisher: 100:1 (CUST-A) 2001:DB8:AA::11/128 2001:DB8:1::2 19/nolabel 2001:DB8:AA:10::/64 2001:DB8:1::2 20/nolabel 2001:DB8:BB::22/128 ::FFFF:3.3.3.3 nolabel/19 2001:DB8:BB:10::/64 ::FFFF:3.3.3.3 nolabel/20 ``` Two things to read here. The next hop for the remote prefixes is `::FFFF:3.3.3.3` \- an **IPv4-mapped IPv6 address**, exactly as in 6PE, because the BGP session runs over IPv4\. And each prefix carries a VPN label (19, 20). But that is only the *inner* label. When PE1 actually forwards an IPv6 packet to a customer behind PE2, it pushes **two** labels: - **Outer (transport) label** \- the LDP label to reach PE2's loopback (3.3.3.3) across the core. The P router swaps this hop by hop and never looks deeper. - **Inner (VPN) label** \- the per-VRF label that tells PE2 which VRF the packet belongs to. Only PE2 sees it. Watch it in a live IPv6 traceroute from CE1 to CE2: ``` CE1#traceroute 2001:DB8:BB::22 source 2001:DB8:AA::11 probe 1 timeout 2 1 2001:DB8:1::1 2 msec 2 ::FFFF:10.0.12.2 [MPLS: Labels 17/19 Exp 0] 5 msec 3 2001:DB8:2::1 [MPLS: Label 19 Exp 0] 4 msec 4 2001:DB8:2::2 5 msec ``` Hop 2 shows `Labels 17/19` \- the two-label stack, transport 17 on top, VPN 19 underneath. Hop 3, after penultimate-hop popping has removed the transport label, shows just `Label 19`, the VPN label about to be used by PE2 to select the VRF. An IPv6 packet, in a two-label MPLS stack, crossing an IPv4 core, arriving in the right customer VRF. That is 6VPE. And end to end: ``` CE1#ping 2001:DB8:BB::22 source Loopback0 !!!!! Success rate is 100 percent (5/5) ``` ## The RD keeps overlapping customers separate Both customer sites advertise 172.16.10.0/24\. In a plain routed network that would be a conflict. In an L3VPN it is fine, because each prefix is prepended with its route distinguisher to form a globally unique *VPN prefix*: ``` PE1#show bgp vpnv4 unicast all 22.22.22.22/32 BGP routing table entry for 100:2:22.22.22.22/32 ``` That `100:2:` prefix is the RD from PE2 making CE2's route distinct from anything on PE1\. The RD's only job is to make otherwise-identical prefixes unique in the VPNv4/VPNv6 table. It plays no part in deciding VPN membership - that is the route target's job. (For the full RD-versus-RT distinction, see the [RD vs RT article](https://www.pinglabz.com/mpls-rd-vs-rt-explained/).) ## Design notes - **The core never changes.** The P routers run IPv4 and LDP and are completely unaware IPv6 exists. This is the entire commercial appeal - you deliver IPv6 L3VPN without touching, re-testing, or risking the core. - **One session, two families.** VPNv4 and VPNv6 ride the same iBGP session between PE loopbacks. You do not build a parallel IPv6 BGP mesh. - **`send-community extended` everywhere.** Route targets are extended communities. Forget this on one PE and that PE's routes never import into any VRF. - **The CE-PE protocol is your choice.** BGP, OSPFv3, static, whatever. This lab uses BGP with `as-override` because both CEs share an AS - covered next. ## Troubleshooting 1. **VPNv6 prefixes present in BGP but not in the VRF?** Route-target import/export mismatch, or missing `send-community extended`. Check `show bgp vpnv6 unicast all` vs `show ipv6 route vrf`. 2. **Prefixes in the VRF but no forwarding?** The label stack is broken. `show bgp vpnv6 unicast all labels` \- if the out-label is `nolabel` where it should be numbered, the VPN label was not assigned. Check the LDP LSP to the remote PE loopback. 3. **Next hop inaccessible?** The IPv4-mapped next hop `::FFFF:x.x.x.x` must resolve through the IGP with a label-switched path. No LSP to the remote PE loopback, no 6VPE. 4. **Overlapping prefixes colliding?** Your RDs are not unique between PEs. Give each PE a distinct RD for the VRF. ## Key takeaways - 6VPE carries IPv6 L3VPN over an unchanged IPv4 MPLS core using the VPNv6 address family - 6PE plus per-VRF isolation. - It uses a **two-label stack**: an outer LDP transport label to the remote PE, and an inner per-VRF VPN label. An IPv6 traceroute across the core shows both (`Labels 17/19`). - The BGP next hop is an IPv4-mapped IPv6 address (`::FFFF:x.x.x.x`), because the session runs over IPv4. - VPNv4 and VPNv6 share the same iBGP session between PE loopbacks. One session, two families. - The RD makes overlapping customer prefixes unique in the VPN table; the RT decides VPN membership. Both are needed. - `send-community extended` is mandatory - route targets are extended communities. Next: [BGP as the PE-CE protocol: as-override, allowas-in, and SoO](https://www.pinglabz.com/bgp-pe-ce-as-override-allowas-in/), which this lab already set up. The full cluster index lives on the [MPLS pillar](https://www.pinglabz.com/mpls/), with the IPv6 angle on the [IPv6 pillar](https://www.pinglabz.com/ipv6/). ### Expert Layer 2 Troubleshooting: Five Broken Scenarios, Ticket Style URL: https://www.pinglabz.com/expert-layer-2-troubleshooting-scenarios/ Last updated: 2026-07-12T08:25:28.000Z Layer 2 problems have a special quality: the switches all say everything is fine, and yet nothing works. There is no routing table to inspect, no neighbour state to compare, just VLANs, trunks, and a spanning tree that is quietly blocking the exact port you need. The diagnosis lives in a handful of `show` commands and, above all, in the syslog - because Layer 2, unlike OSPF, is actually fairly talkative when it detects a mismatch. This article is five Layer 2 faults, each built and broken in a CML lab, with the real output and the real fix. For the theory, see the [VLAN and switching guide](https://www.pinglabz.com/vlans-layer-2-switching/) and the [spanning tree guide](https://www.pinglabz.com/spanning-tree-protocol/). ## The method Work Layer 2 in this order: **1\. Read the logs** `show logging`. Layer 2 announces most of its mismatches - native VLAN, PVID, errdisable causes. Start here, always. **2\. Check the trunk** `show interfaces trunk`. Is the VLAN even allowed and active on the trunk? Half of all "VLAN won't pass" problems die here. **3\. Check spanning tree** `show spanning-tree`. Is the port you need actually forwarding, or is STP blocking it for a good reason? **4\. Check the port** `show interfaces status`. Is it errdisabled? `show port-security`, `show errdisable recovery` for why. ## Ticket 1: "This VLAN won't cross one trunk" **Symptom.** VLAN 10 works everywhere except across the SW3-SW2 trunk. The trunk is up. Other VLANs cross it fine. **Diagnosis.** The logs tell you immediately: ``` SW3#show logging | include NATIVE_VLAN *Jul 12 08:10:33.701: %CDP-4-NATIVE_VLAN_MISMATCH: Native VLAN mismatch discovered on Ethernet0/0 (88), with SW2 Ethernet0/1 (99). ``` And earlier, when the native VLANs were 1 vs 99: ``` *SPANTREE-2-BLOCK_PVID_PEER: Blocking Ethernet0/0 on VLAN0099. Inconsistent peer vlan. *SPANTREE-2-BLOCK_PVID_LOCAL: Blocking Ethernet0/0 on VLAN0001. Inconsistent local vlan. ``` **Cause.** A native VLAN mismatch. One side of the trunk has native VLAN 99, the other has something else. This produces *two* distinct symptoms, and it is important to understand both: - **CDP logs a warning** (`%CDP-4-NATIVE_VLAN_MISMATCH`). This is cosmetic - CDP noticed and is telling you. - **Spanning tree actually blocks the native VLAN** (`%SPANTREE-2-BLOCK_PVID`). This is not cosmetic. STP detects that the untagged (native) traffic on the two ends belongs to different VLANs - a security and looping hazard - and blocks the PVID on that segment. The VLAN genuinely stops passing. A native VLAN mismatch is also a real security concern: it can enable VLAN hopping, because traffic in one side's native VLAN pops out into the other side's native VLAN untagged. **Fix:** make the native VLAN match on both ends. ``` SW3(config-if)#switchport trunk native vlan 99 ``` **Lesson:** a native VLAN mismatch gives you two log messages from two subsystems. The CDP one is a warning; the SPANTREE BLOCK\_PVID one is the actual outage. Both point at the same fix. ## Ticket 2: "MST is behaving strangely after a change" **Symptom.** After someone edited the MST configuration, two switches that should be in the same region are suddenly not, and ports are blocking unexpectedly. **Diagnosis.** The MST region digest is the entire answer: ``` SW1#show spanning-tree mst configuration digest Digest 0xE05029508B8AC96367E964B97EE7DA9B SW2#show spanning-tree mst configuration digest Digest 0xDA7E3D01C232F1CFDB7F72C25D414BAA ``` **Cause.** Different digests mean different regions. Two switches are in the same MST region only if their **name**, **revision**, and **complete VLAN-to-instance map** all match exactly. In the lab we moved a single VLAN from one instance to another on just one switch - and that was enough to change the digest and split the region. When a region splits, the link between the two "regions" becomes a boundary, spanning tree recalculates as if crossing between separate domains, and ports that were forwarding can start blocking. **Fix:** make the MST configuration byte-identical on every switch in the region. Compare the name, revision, and VLAN map, and copy-paste rather than retype. ``` spanning-tree mst configuration name PINGLABZ revision 1 instance 1 vlan 10 instance 2 vlan 20,99 ``` **Lesson:** for any MST oddity, `show spanning-tree mst configuration digest` on both switches is the first command. Same digest = same region. Different digest = you have found your problem. Full detail in [MST and PVST+ interoperation](https://www.pinglabz.com/mst-pvst-interoperation/). ## Ticket 3: "A port is dead and won't come back" **Symptom.** An access port is down. The cable is fine. `no shutdown` does not fix it. **Diagnosis.** ``` SW3#show interfaces Ethernet0/2 status Port Name Status Vlan Et0/2 err-disabled 10 ``` `err-disabled`, not `down`. That is a protection feature, not a physical fault. Find which one: ``` SW3#show port-security interface Ethernet0/2 Port Status : Secure-shutdown Security Violation Count : 1 SW3#show logging | include SECURE *Jul 12 08:08:40.999: %PORT_SECURITY-2-PSECURE_VIOLATION: Security violation occurred, caused by MAC address 02aa.bbcc.ddee on port Ethernet0/2. ``` **Cause.** A port security violation. A frame with an unauthorised source MAC (here 02aa.bbcc.ddee - a MAC that changed on the connected host) arrived on a port configured to allow only one, sticky-learned MAC. Port security did exactly what it was told: errdisabled the port. **Fix.** A plain `no shutdown` does not clear errdisable - you must either bounce the port with `shutdown` then `no shutdown`, or configure automatic recovery: ``` SW3(config)#errdisable recovery cause psecure-violation SW3(config)#errdisable recovery interval 30 ``` ``` SW3#show interfaces Ethernet0/2 status Et0/2 connected 10 <-- recovered ``` **Lesson:** `err-disabled` is a feature acting, not a fault. `show interfaces status` tells you it is errdisabled; the syslog tells you which feature and why. Fix the cause before you recover the port, or it just errdisables again. ## Ticket 4: "The router-on-a-stick subinterfaces are all down" **Symptom.** A router doing inter-VLAN routing has all its dot1q subinterfaces down. The config looks correct. The trunk on the switch side is up. **Diagnosis.** ``` R1#show ip interface brief | include Ethernet0/1 Ethernet0/1 unassigned YES unset administratively down down Ethernet0/1.10 10.10.10.1 YES TFTP administratively down down Ethernet0/1.20 10.10.20.1 YES TFTP administratively down down Ethernet0/1.99 10.10.99.1 YES TFTP administratively down down ``` **Cause.** The parent physical interface is `administratively down`. A subinterface inherits the admin state of its parent - if the physical interface is shut, every subinterface on it is down regardless of its own config. In this case the parent came up admin-down at boot (a real quirk we hit in the lab: the day-0 config's `no shutdown` was applied before the subinterfaces existed, and the parent ended up shut). **Fix:** bring up the parent. ``` R1(config)#interface Ethernet0/1 R1(config-if)#no shutdown ``` **Lesson:** when every subinterface on one physical port is down at once, look at the parent. A subinterface cannot be up while its parent is shut. The tell is that they are all down together - a config problem on one subinterface would only affect that one. ## Ticket 5: "A static-IP host lost connectivity after we enabled DAI" **Symptom.** After enabling Dynamic ARP Inspection, one server - the one with a manually-configured static IP - can no longer communicate. Every DHCP client on the same VLAN is fine. **Diagnosis.** The DHCP-client hosts work; the static host does not. That pattern points straight at the binding table: ``` SW3#show ip dhcp snooping binding MacAddress IpAddress Lease(sec) Type VLAN Interface 52:54:00:4E:11:C0 10.10.10.10 86387 dhcp-snooping 10 Ethernet0/2 Total number of bindings: 1 ``` The DHCP client (10.10.10.10) has a binding. The static server does not appear at all. **Cause.** DAI validates ARP against the DHCP snooping binding table. A host with a static IP never did a DHCP transaction, so it has no binding, so DAI drops its ARP - and it goes dark. This is the single most common DAI deployment mistake: turning it on without first accounting for the static devices (servers, printers, network gear) that do not use DHCP. **Fix.** Give the static host a binding, either a static snooping entry or an ARP ACL: ``` ! Option A: static source binding ip source-binding 0011.2233.4455 vlan 10 10.10.10.50 interface Ethernet0/5 ! Option B: an ARP ACL that permits the pair arp access-list STATIC-SERVERS permit ip host 10.10.10.50 mac host 0011.2233.4455 ! ip arp inspection filter STATIC-SERVERS vlan 10 ``` **Lesson:** DAI (and IPSG) trust only the snooping binding table. Any host that does not use DHCP needs an explicit binding *before* you enable inspection. Inventory your static devices first. See [DAI and IP Source Guard](https://www.pinglabz.com/dynamic-arp-inspection-ip-source-guard/). ## The commands, collected ``` ! Logs first - L2 announces its mismatches show logging | include NATIVE_VLAN|PVID|SECURE|err-disable|UDLD ! Trunks show interfaces trunk show interfaces switchport ! Spanning tree show spanning-tree show spanning-tree mst configuration digest (region membership) show spanning-tree inconsistentports show spanning-tree vlan ! Ports show interfaces status show port-security interface show errdisable recovery ! MAC / bindings show mac address-table show ip dhcp snooping binding show ip verify source ``` ## Try it yourself Build a mixed MST/PVST square with a router-on-a-stick and some security features, then break it. Ten minutes per fault, and name the cause from show output before you look at the config: 1. Set a different native VLAN on one end of a trunk. Note that you get both a CDP warning and an STP block. 2. Move one VLAN to a different MST instance on one switch only. Diff the digests. 3. Change a host's MAC on a port with sticky port security allowing one MAC. 4. Shut the parent of a set of router-on-a-stick subinterfaces. 5. Enable DAI on a VLAN that has a static-IP host with no binding. ## Key takeaways - **Read the logs first.** Layer 2 announces native VLAN mismatches, PVID inconsistencies, port security violations, and errdisable causes. This is the opposite of OSPF, and you should exploit it. - A native VLAN mismatch produces two messages: a cosmetic CDP warning and an actual STP `BLOCK_PVID` that stops the VLAN. Same fix for both. - For any MST oddity, compare `show spanning-tree mst configuration digest`. Different digest = different region = your problem. - `err-disabled` is a protection feature acting, not a physical fault. The syslog names the cause; fix it before recovering the port. - All subinterfaces down at once = the parent physical interface is shut. - DAI/IPSG blocking one specific host, while DHCP clients are fine, means that host has no snooping binding. Static devices need explicit bindings before you enable inspection. That closes the expert Layer 2 series - and with it, switching on PingLabz is complete from CCNA fundamentals to CCIE depth. The full cluster indexes live on the [VLAN and switching pillar](https://www.pinglabz.com/vlans-layer-2-switching/) and the [spanning tree pillar](https://www.pinglabz.com/spanning-tree-protocol/). ### Switch Administration: SDM Templates, Errdisable Recovery, and CAM Aging URL: https://www.pinglabz.com/switch-administration-sdm-errdisable/ Last updated: 2026-08-01T19:35:44.000Z The switching features that make headlines are the protocols - spanning tree, VLANs, EtherChannel. The features that actually keep a switch running are the unglamorous administrative ones: how the hardware tables are carved up, what happens when a port is disabled by a protection feature, and how long the switch remembers a MAC. Ignore them and you get mysterious "table full" errors, ports that never come back after a transient fault, and traffic flooding for no obvious reason. This article covers SDM templates, errdisable recovery, and CAM aging, with real output from a CML lab. For the fundamentals, start at the [complete VLAN and switching guide](https://www.pinglabz.com/vlans-layer-2-switching/). ## SDM templates: carving up the hardware A switch's forwarding intelligence lives in a fixed-size chunk of specialised memory (TCAM). That memory has to hold MAC addresses, IPv4 routes, IPv6 routes, ACL entries, QoS entries, multicast state - all of it, in one finite space. You cannot have maximum capacity for everything at once. So the switch offers **SDM templates**: preset allocations that trade one resource against another. ``` show sdm prefer ``` Typical templates on a Catalyst access switch: default A balanced split across MAC, routing, ACL, and QoS. Fine for a general-purpose switch. vlan / access Maximises MAC address capacity at the expense of routing. For a pure L2 access switch that routes nothing. routing Maximises IPv4/IPv6 route capacity at the expense of MAC. For a distribution or core switch doing L3. dual-ipv4-and-ipv6 Reserves meaningful space for IPv6 routes. Essential the moment you run dual-stack - the default often starves IPv6. ``` Switch(config)#sdm prefer routing ! Requires a reload to take effect Switch#reload ``` The critical operational fact: **changing the SDM template requires a reload.** It re-carves the hardware memory, which cannot be done on a running switch. Plan it into a maintenance window. The failure mode this prevents is nasty and non-obvious. A switch running the `default` template in an IPv6 deployment can silently run out of IPv6 TCAM, at which point IPv6 routes stop being installed in hardware and get punted to software - which cripples IPv6 forwarding performance while IPv4 looks perfectly healthy. `show platform tcam utilization` (or the platform-specific equivalent) shows you the usage; if a category is near 100%, you need a different template. **Platform note:** SDM templates are a physical-hardware concept - they carve up a real TCAM. Virtual IOL-L2 in CML has no TCAM, so `show sdm prefer` returns `% Invalid input`. This is a hardware feature you configure on real Catalyst switches; we mention it here for completeness and are honest that the virtual lab cannot demonstrate it. ## Errdisable recovery: the safety net A whole family of protection features respond to a violation by **errdisabling** the port - shutting it down and leaving it down. Port security violation, BPDU guard, UDLD, storm control, DHCP rate limit, DAI rate limit, link flap - all of them err on the side of "cut it off and wait for a human". That is the safe default, but it means every transient fault requires manual intervention. A patch cable reseated, a device rebooted, an SFP cleaned - and the port stays dead until someone runs `shutdown` / `no shutdown`. Errdisable recovery automates that. ``` errdisable recovery cause psecure-violation errdisable recovery cause udld errdisable recovery cause bpduguard errdisable recovery interval 300 ``` From the lab: ``` SW1#show errdisable recovery ErrDisable Reason Timer Status bpduguard Enabled udld Enabled storm-control Disabled psecure-violation Disabled ... Timer interval: 60 seconds ``` Each enabled cause has its port re-tried automatically after the interval. Watch it work: we triggered a port-security violation, the port went errdisable, then after enabling recovery it came back: ``` SW3#show interfaces Ethernet0/2 status Et0/2 err-disabled 10 <-- shut by the violation ! ... after errdisable recovery cause psecure-violation, and the timer/bounce: SW3#show interfaces Ethernet0/2 status Et0/2 connected 10 <-- back in service ``` ### Choosing which causes to auto-recover Not every cause should recover automatically: Auto-recover **udld, link-flap, storm-control** \- usually transient physical faults. Auto-recovery saves a truck roll for a self-clearing problem. Consider carefully **psecure-violation, bpduguard** \- these usually indicate something wrong (rogue device, unauthorised switch). Auto-recovering means the offender gets re-tried every interval. Sometimes you want it to stay down until investigated. The interval is a genuine trade-off. Too short (the 30-second minimum) and a persistently-faulty link flaps in and out of service constantly, which is worse than staying down. **300 seconds is a sensible default** \- long enough that a genuinely broken link is not thrashing, short enough that a transient fault self-heals within five minutes. That is the build-time view: which causes to arm and how long to set the timer. For the other half, pulling the exact cause off a port that is already down and deciding when a timer is the wrong answer entirely, see [every err-disable cause and how to bring the port back](https://www.pinglabz.com/errdisable-recovery-cisco/). ## CAM aging: how long a MAC is remembered The CAM table (the MAC address table) is how a switch forwards unicast intelligently - it learns which MAC lives on which port and sends frames only there. Entries are learned dynamically and **aged out** after a period of inactivity, default 300 seconds. ``` mac address-table aging-time 300 ``` The aging time interacts with your spanning tree design in a way that causes real, hard-to-diagnose problems: - **Too long, and a topology change causes black-holing.** When STP reconverges and a host's path changes, the switch may still have the old (now wrong) CAM entry for up to the aging time - 5 minutes of traffic sent out the wrong port. This is exactly why STP *topology change notifications* exist: a TCN tells switches to shorten their CAM aging to the forward delay (15 seconds) temporarily, flushing stale entries fast after a topology change. - **Too short, and you flood.** If entries age out faster than hosts send traffic, the switch keeps forgetting where things are and floods unknown-unicast to relearn. On a quiet-but-important flow (a heartbeat, a backup channel) this looks like intermittent unicast flooding, and it is baffling until you connect it to aging. The classic symptom of a too-short aging time (or an asymmetric-routing situation where the switch never learns a MAC from returning traffic) is **persistent unicast flooding**: you see traffic for a specific destination appearing on ports it has no business being on. Check `show mac address-table` \- if the destination is not in the table, the switch is flooding it, and you need to find out why it never learns it. ``` Switch#show mac address-table interface Ethernet0/2 Vlan Mac Address Type Ports ---- ----------- -------- ----- 10 5254.004e.11c0 DYNAMIC Et0/2 Switch#show mac address-table aging-time Switch#show mac address-table count ``` You can also pin critical MACs with a static entry so they never age out - useful for a default gateway or a critical server whose flooding you want to eliminate entirely: ``` mac address-table static 0011.2233.4455 vlan 10 interface Ethernet0/1 ``` ## The administration checklist for a new access switch 1. **Pick the right SDM template** for the switch's role and reload. Dual-stack? Use a template that reserves IPv6 space. 2. **Configure errdisable recovery** for the transient causes (udld, link-flap, storm-control) at a 300-second interval. Leave security-violation causes to manual review if your policy demands it. 3. **Leave CAM aging at default (300s)** unless you have a specific reason. Ensure STP topology change handling is working so it shortens automatically after a reconvergence. 4. **Pin critical MACs statically** if you have a specific flooding problem to eliminate. ## Key takeaways - SDM templates carve the fixed hardware TCAM between MAC, routes, ACLs, and QoS. Pick the template for the switch's role; **changing it requires a reload**. Dual-stack needs an IPv6-aware template or IPv6 silently starves. - Errdisable recovery auto-re-enables ports shut by protection features after an interval. **300 seconds** is a good default - short enough to self-heal transients, long enough to avoid flapping. - Auto-recover physical causes (udld, link-flap, storm-control); think twice before auto-recovering security causes (port-security, BPDU guard), which usually mean something is wrong. - CAM aging defaults to 300s. Too long causes post-topology-change black-holing (which is why STP TCNs shorten it); too short causes unicast flooding. - Persistent unicast flooding to a specific destination almost always means the switch is not learning that MAC - check `show mac address-table`. - SDM is a hardware feature; virtual IOL-L2 in CML cannot demonstrate it (no TCAM). Errdisable recovery and CAM aging work fully. Next: [five broken Layer 2 scenarios, ticket style](https://www.pinglabz.com/expert-layer-2-troubleshooting-scenarios/). The full cluster index lives on the [VLAN and switching pillar](https://www.pinglabz.com/vlans-layer-2-switching/). ### Dynamic ARP Inspection and IP Source Guard: The L2 Security Stack URL: https://www.pinglabz.com/dynamic-arp-inspection-ip-source-guard/ Last updated: 2026-08-01T19:32:07.000Z ARP is the most trusting protocol in the building. A host asks "who has 10.10.10.1?" and believes whatever answer comes back, with no verification of any kind. So an attacker replies "I do" - claiming to be your default gateway - and every host on the subnet starts sending its off-net traffic to the attacker. That is ARP spoofing, and it is the basis of most Layer 2 man-in-the-middle attacks. Dynamic ARP Inspection and IP Source Guard are the two features that fix ARP's and IP's respective trust problems, and both are built on the DHCP snooping binding table. This article covers both, with real output from a CML lab and an honest note on what a virtual switch can enforce. For the fundamentals, start at the [complete VLAN and switching guide](https://www.pinglabz.com/vlans-layer-2-switching/). This builds directly on [DHCP snooping](https://www.pinglabz.com/dhcp-snooping-in-depth/), so read that first. ## Dynamic ARP Inspection DAI intercepts every ARP packet on an untrusted port and checks it against the snooping binding table. The question it asks is simple: *does this ARP packet claim an IP-to-MAC mapping that matches a real DHCP lease on this port?* - A host sends an ARP reply saying "10.10.10.10 is at 5254.004e.11c0" from the port where that exact binding exists. **DAI permits it.** - An attacker sends "10.10.10.1 (the gateway) is at my-MAC" from a port that has no such binding. **DAI drops it and logs the violation.** The attacker cannot forge a binding, because bindings are created only by real DHCP transactions (or a deliberate static entry). So the gateway-impersonation attack simply fails. That is the theory, and it is worth watching happen. The forged ARP reply DAI refuses takes about six lines of Scapy to build, and there is a full write-up of [the attack this is stopping, alongside the switch log that names the attacker's real MAC as it drops the frame](https://www.pinglabz.com/scapy-packet-crafting-spoofing-cisco/). ### Configuration ``` ! Enable per VLAN ip arp inspection vlan 10 ! Trust the uplinks and inter-switch trunks (same ports as DHCP snooping trust) interface Ethernet0/0 ip arp inspection trust interface Ethernet0/1 ip arp inspection trust ! Rate-limit ARP on untrusted access ports (ARP flooding defence) interface Ethernet0/2 ip arp inspection limit rate 15 ``` The trust model mirrors DHCP snooping exactly: uplinks and inter-switch trunks trusted, host access ports untrusted. This makes sense - a legitimate ARP from another switch has already been validated by that switch, so you do not re-inspect it. The rate limit matters because **DAI punts ARP to the CPU for inspection**. An attacker flooding ARP could turn DAI itself into a denial of service by overwhelming the CPU. The rate limit caps that; a port exceeding it is errdisabled. The default on untrusted ports is 15 pps. ### Operational state ``` SW3#show ip arp inspection vlan 10 Vlan Configuration Operation ACL Match Static ACL 10 Enabled Active Vlan ACL Logging DHCP Logging Probe Logging 10 Deny Deny Off ``` `Enabled / Active` on VLAN 10, using the DHCP binding table (DHCP Logging: Deny means it logs denied packets). This is DAI operational and watching. ## IP Source Guard Where DAI protects ARP, IP Source Guard (IPSG) protects the IP source address itself. It applies a dynamic port ACL that permits only the IP address (and optionally MAC) bound to that port in the snooping table. Any packet with a different source IP is dropped. This stops a host from spoofing its source IP - claiming to be another machine to bypass an ACL, evade logging, or launch a spoofed-source attack. ``` interface Ethernet0/2 ip verify source ``` Add `ip verify source port-security` to also enforce the source MAC, tying the port down to exactly one (IP, MAC) pair. ``` SW3#show ip verify source Interface Filter-type Filter-mode IP-address Mac-address Vlan Et0/2 ip active 10.10.10.10 10 ``` That is IPSG live: on Et0/2, the only source IP permitted is 10.10.10.10 - the address bound to that port in the snooping table. A packet sourced from any other IP is dropped in hardware. Because the binding came from DHCP, the moment that host's lease expires or moves, the filter updates automatically. ## The complete stack, in order 1 **DHCP snooping** builds the binding table (IP + MAC + VLAN + port). Nothing else works without it. 2 **DAI** reads the table and drops ARP packets that do not match a binding. Stops ARP spoofing / gateway impersonation. 3 **IPSG** reads the table and drops packets whose source IP does not match the binding. Stops IP spoofing. 4 **Port security** caps how many MACs the port may learn, and errdisable recovery brings ports back. The safety layers around all of the above. The dependency is strict and one-directional. **If DHCP snooping is broken, DAI and IPSG are broken**, because they have no table to consult. When DAI blocks a legitimate host, the fault is almost always a missing snooping binding, not DAI itself. Debug the table first, every time. ## Handling non-DHCP hosts Not every device uses DHCP. Servers, printers, and network gear often have static IPs, and they have no snooping binding - so DAI would drop their ARP and IPSG would drop their traffic. Two ways to accommodate them: **A static snooping binding** (the same technique we used in the lab): ``` ip source-binding 0011.2233.4455 vlan 10 10.10.10.50 interface Ethernet0/5 ``` **An ARP ACL** for DAI, which permits specific IP-to-MAC pairs regardless of the binding table: ``` arp access-list STATIC-HOSTS permit ip host 10.10.10.50 mac host 0011.2233.4455 ! ip arp inspection filter STATIC-HOSTS vlan 10 ``` Every static host needs one of these, or it goes dark the moment you enable DAI. Inventory your static devices *before* you turn DAI on, not after the help desk lights up. ## Platform note: honest reporting DAI and IPSG are hardware-enforced features - they punt ARP to the CPU for inspection and install per-port ACLs in the switch ASIC. In our CML lab on virtual IOL-L2, both features report `Enabled / Active` and IPSG shows the correct bound IP, but the forwarding plane does not count or enforce live traffic: ``` SW3#show ip arp inspection statistics vlan 10 Vlan Forwarded Dropped DHCP Drops ACL Drops 10 0 0 0 0 ``` Zero, because the virtual node does not punt ARP to the inspection process - the same architectural limitation that stops DHCP snooping populating its table from live traffic. The configuration and operational state are genuine; the live enforcement counters are not reproducible on this platform. On real Catalyst hardware, an ARP-spoofing attempt from an unbound port produces: ``` %SW_DAI-4-DHCP_SNOOPING_DENY: 1 Invalid ARPs (Res) on Et0/2, vlan 10. ([mac]/10.10.10.1/[target-mac]/10.10.10.10/[time]) ``` We show you the real config, the real operational state (Active, with the bound IP), and the real syslog you would see - rather than staging fake drop counters on a platform that cannot produce them. An expert reader can tell the difference, and we would rather earn the trust. ## Troubleshooting 1. **DAI dropping a legitimate host?** Check `show ip dhcp snooping binding` first. Ninety percent of DAI "problems" are a missing binding. Static host? It needs a static binding or ARP ACL. 2. **IPSG blocking a host?** Same - the binding is missing or wrong. `show ip verify source` shows what IPSG thinks the port is allowed to use. 3. **Port errdisabled by DAI?** ARP rate limit exceeded. Either a flood/attack, or a legitimate burst - raise the limit if it is a false positive, and add errdisable recovery. 4. **Everything Active but nothing enforced (in a lab)?** Virtual switch limitation - the forwarding plane is not punting to the inspection engine. Expected on IOL-L2; works on hardware. 5. **Static hosts unreachable after enabling DAI?** They have no binding. Add ARP ACLs or static source-bindings for every static device before enabling. ## Key takeaways - ARP and IP both trust the sender blindly. DAI fixes ARP spoofing; IPSG fixes IP spoofing. - Both read the **DHCP snooping binding table**. Break snooping and you break both. When DAI/IPSG blocks a legitimate host, debug the binding table first. - DAI trust mirrors snooping trust: uplinks and inter-switch trunks trusted, host ports untrusted. Rate-limit untrusted ARP to protect the CPU. - IPSG (`ip verify source`) installs a per-port filter permitting only the bound source IP; add `port-security` to bind the MAC too. - Static (non-DHCP) hosts have no binding. Give each one a static snooping binding or an ARP ACL *before* enabling DAI, or they go dark. - On virtual IOL-L2 the features are Active but not counter-enforced (no forwarding-plane punt); on hardware they enforce and log. We report the config, operational state, and real syslog honestly rather than faking counters. Next: [switch administration - SDM templates, errdisable recovery, and CAM aging](https://www.pinglabz.com/switch-administration-sdm-errdisable/). The full cluster index lives on the [VLAN and switching pillar](https://www.pinglabz.com/vlans-layer-2-switching/), and the security features cross-link to [infrastructure security](https://www.pinglabz.com/infrastructure-security/). ### DHCP Snooping in Depth: Bindings, Option 82, and Trusted Ports URL: https://www.pinglabz.com/dhcp-snooping-in-depth/ Last updated: 2026-07-12T08:25:27.000Z Every Layer 2 security feature worth having is built on one foundation: a table that says "this IP address, with this MAC, is legitimately on this port." DHCP snooping builds that table. Dynamic ARP Inspection uses it to stop ARP spoofing. IP Source Guard uses it to stop IP spoofing. If snooping is not working, none of the things that depend on it work either. This article covers how DHCP snooping builds its binding table, trusted versus untrusted ports, Option 82, and how to verify it - with real output from a CML lab, including an honest account of a platform limitation you need to know about. For the fundamentals, start at the [complete VLAN and switching guide](https://www.pinglabz.com/vlans-layer-2-switching/). ## The attack it stops DHCP has no authentication. A client broadcasts "I need an address", and it believes whatever server answers first. So an attacker plugs a rogue DHCP server into an access port, answers faster than the real one, and hands clients an address, a DNS server, and - critically - a **default gateway that points at the attacker's machine**. Now every packet those clients send off-subnet flows through the attacker. That is a man-in-the-middle attack, and it needs nothing more than a laptop running `dnsmasq`. DHCP snooping stops it by deciding which ports are allowed to send DHCP *server* messages. ## Trusted and untrusted ports This is the entire concept, and it is beautifully simple: Trusted port May forward DHCP server messages (OFFER, ACK). These are your uplinks toward the real DHCP server, and trunks to other switches. You *must* mark them, because snooping distrusts everything by default. Untrusted port May send DHCP client messages (DISCOVER, REQUEST) but a server message arriving here is **dropped** and logged. These are your access ports facing hosts. This is the default. A rogue DHCP server on an untrusted access port sends an OFFER, the switch drops it before it reaches any client, and the attack fails. The real server, reachable only through trusted uplinks, works normally. ## The binding table While snooping is watching DHCP transactions on untrusted ports, it records every successful lease. Each entry ties together four facts: **MAC address, IP address, VLAN, and switch port**, plus the lease time. This table is the crown jewels. It is the switch's authoritative record of who is legitimately where, and DAI and IPSG both consult it to decide whether a frame is spoofed. ## Configuration ``` ! Enable globally, then per VLAN ip dhcp snooping ip dhcp snooping vlan 10,20 ! Trust the uplinks toward the real server interface Ethernet0/0 ip dhcp snooping trust interface Ethernet0/1 ip dhcp snooping trust ! Rate-limit the untrusted access ports (DHCP starvation defence) interface Ethernet0/2 ip dhcp snooping limit rate 15 ``` Two things people forget: - **Trust the trunks between switches.** If the DHCP server is two switches away, every switch in the path needs the inter-switch trunk marked trusted, or the OFFER gets dropped at the first untrusted hop. - **Rate-limit untrusted ports.** A rate limit on DHCP messages defeats DHCP starvation - an attacker sending thousands of DISCOVERs from spoofed MACs to exhaust the server's pool. 15 packets per second is generous for a legitimate host. ### Operational state ``` SW3#show ip dhcp snooping Switch DHCP snooping is enabled DHCP snooping is configured on following VLANs: 10,20 DHCP snooping is operational on following VLANs: 10,20 Insertion of option 82 is disabled Option 82 on untrusted port is not allowed Verification of hwaddr field is enabled DHCP snooping trust/rate is configured on the following Interfaces: Interface Trusted Allow option Rate limit (pps) Ethernet0/0 yes yes unlimited Ethernet0/1 yes yes unlimited Ethernet0/2 no no 15 ``` Note `Verification of hwaddr field is enabled` \- snooping also checks that the client hardware address inside the DHCP payload matches the source MAC of the frame, catching a class of spoofing where those two disagree. ## Option 82: identifying where a request came from DHCP Option 82 (the relay agent information option) lets a switch stamp each DHCP request with *which switch and which port* it arrived on, before relaying it toward the server. The server can then assign addresses based on physical location, log which port a client is on, and apply per-port policy. ``` ip dhcp snooping information option ``` The default circuit-id encodes VLAN, module, and port, and the remote-id is the switch's MAC. When Option 82 is on, the switch inserts it into requests from untrusted ports on the way up, and strips it from replies on the way down. The gotcha, and it is a common outage: **when the DHCP server is a router doing relay, and Option 82 is inserted by an access switch, the relay may drop the request** because a request arriving with Option 82 already present but a giaddr of 0.0.0.0 looks malformed. The fix is either `ip dhcp snooping information option allow-untrusted` on the relay's side, or configuring the relay to trust the option. In many enterprise designs where the L3 switch is both the snooping switch and the relay, you simply leave Option 82 off unless you specifically need location-based assignment - which is exactly what we did in the lab. ## Verification and a platform limitation Here is where we are going to be straight with you. In our CML lab, we ran a real DHCP transaction - a host obtained a genuine lease (10.10.10.10) from the router's DHCP pool via `udhcpc`. But the virtual IOL-L2 switch reported zero packets through its snooping engine: ``` SW3#show ip dhcp snooping statistics Packets Forwarded = 0 Packets Dropped = 0 ``` The reason is architectural: DHCP snooping works by **punting DHCP packets to the switch CPU** for inspection, and the virtual IOL-L2 node in CML does not implement that forwarding-plane punt. The configuration applies, the feature reports `operational`, trust is enforced in config - but the binding table never populates from live DHCP, because the snooping process never sees the packets. On real Catalyst hardware, this works exactly as designed. To demonstrate the downstream features (DAI and IPSG) in the lab, we populated the table with a **static binding**, which is a legitimate production technique in its own right for hosts with fixed addresses: ``` SW3#ip dhcp snooping binding 5254.004E.11C0 vlan 10 10.10.10.10 interface Ethernet0/2 expiry 86400 SW3#show ip dhcp snooping binding MacAddress IpAddress Lease(sec) Type VLAN Interface 52:54:00:4E:11:C0 10.10.10.10 86387 dhcp-snooping 10 Ethernet0/2 Total number of bindings: 1 ``` There is the binding table with an entry, and that entry is what DAI and IPSG will trust. On hardware, an entry like this appears automatically the moment a host completes a DHCP lease. We would rather tell you the lab's limitation and show you the real static-binding technique than pretend the automatic population happened. ## The database: surviving a reload The binding table lives in RAM. Reload the switch and it is gone - which means every host has to renew its lease before DAI and IPSG will let its traffic through, and in the meantime legitimate hosts are blocked. On a busy access switch that is a self-inflicted outage after every reboot. The fix is to persist the table: ``` ip dhcp snooping database flash:dhcp-snooping-db.txt ip dhcp snooping database write-delay 300 ``` Now the bindings survive a reload. In a larger deployment, write the database to a TFTP or FTP server instead of local flash. This one line prevents a genuinely nasty post-maintenance surprise. ## Troubleshooting 1. **Clients not getting addresses after enabling snooping?** An uplink or inter-switch trunk toward the server is untrusted. Every hop in the DHCP path needs the server-facing port trusted. 2. **Binding table empty?** On hardware: the DHCP transaction has not completed since snooping was enabled - clients need to renew. On virtual IOL-L2 in CML: expected, use static bindings (see above). 3. **Relay dropping requests after enabling Option 82?** The giaddr-zero-with-option-82 problem. Add `allow-untrusted` on the relay or leave Option 82 off unless needed. 4. **Bindings gone after a reload?** Configure the snooping database to persist to flash or a server. 5. **DAI or IPSG blocking legitimate hosts?** Their binding is missing. Check `show ip dhcp snooping binding` first - the problem is almost always here, not in DAI/IPSG themselves. ## Key takeaways - DHCP snooping is the foundation of the whole L2 security stack. Its binding table (MAC + IP + VLAN + port) is what DAI and IPSG trust. - Trusted ports (uplinks, inter-switch trunks) may forward DHCP server messages; untrusted ports (host access) may not. Everything is untrusted by default. - Trust *every* hop toward the real server, and rate-limit untrusted ports to defeat DHCP starvation. - Option 82 stamps requests with switch/port identity for location-based assignment, but can break relay if not handled - leave it off unless you need it. - On virtual IOL-L2 in CML the snooping engine does not see live DHCP (no forwarding-plane punt), so the table populates via static binding; on real hardware it populates automatically. We report this honestly. - Persist the binding database to flash or a server, or every reload blocks legitimate hosts until they renew. Next: [Dynamic ARP Inspection and IP Source Guard](https://www.pinglabz.com/dynamic-arp-inspection-ip-source-guard/), the two features that turn the binding table into active protection. The full cluster index lives on the [VLAN and switching pillar](https://www.pinglabz.com/vlans-layer-2-switching/). ### Storm Control: Stopping Broadcast Floods at the Port URL: https://www.pinglabz.com/storm-control-configuration/ Last updated: 2026-07-12T08:25:26.000Z A single misbehaving host can take down an entire VLAN. It does not take a spanning tree loop - a NIC stuck in a fault state, a virtualised workload gone wrong, or a genuine attack can flood broadcast frames fast enough to saturate every switch in the broadcast domain. Every switch dutifully floods every broadcast out every port, so one port's flood becomes everyone's problem. Storm control caps the rate of broadcast, multicast, or unknown-unicast traffic on a port and takes action when the cap is exceeded. This article covers the configuration, the two threshold styles, and the right action to take - with an honest note on platform support. For the fundamentals, start at the [complete VLAN and switching guide](https://www.pinglabz.com/vlans-layer-2-switching/). ## Why broadcasts are uniquely dangerous A unicast frame goes to one port (once the switch has learned the MAC). A broadcast frame goes to *every* port in the VLAN, on *every* switch in the broadcast domain. There is no learning, no filtering, no limit. That is by design - it is how ARP and DHCP discovery work - but it means a broadcast storm scales with the size of your broadcast domain, and it consumes CPU on every device, not just switches. Hosts have to process every broadcast too. Three traffic types get the same flood-everywhere treatment and are therefore the three storm control watches: Broadcast Destination MAC ffff.ffff.ffff. Floods everywhere, always. The classic storm. Multicast Floods everywhere unless IGMP snooping is constraining it. A misbehaving multicast source can storm too. Unknown unicast A unicast to a MAC the switch has not learned floods everywhere too. MAC table overflow attacks exploit exactly this. ## Configuration Storm control is a per-interface feature. You set a rising threshold and, optionally, a falling threshold: ``` interface GigabitEthernet1/0/10 storm-control broadcast level 1.00 0.50 storm-control action shutdown storm-control action trap ``` That reads: on this port, if broadcast traffic exceeds **1.00%** of the interface bandwidth, take action; do not clear the condition until it drops back below **0.50%**. The gap between the two thresholds is **hysteresis** \- it stops the port flapping in and out of the storm state when traffic hovers right at the threshold. ### Three ways to express the threshold **level ** Percentage of interface bandwidth. Simple, but a "1%" cap means a very different absolute rate on a 1 Gbps port than on a 10 Gbps port. **level bps ** An absolute bits-per-second cap. Predictable regardless of port speed. Usually the better choice. **level pps ** A packets-per-second cap. Broadcast storms are often about packet rate (small frames) rather than raw bandwidth, so pps can catch a storm that a bps cap would miss. ### The two actions By default, storm control just **drops** the excess traffic above the threshold and keeps the port up. That protects the rest of the network but leaves the offending host connected and still trying. The two explicit actions change that: - **`storm-control action trap`** \- send an SNMP trap and syslog when the storm starts. Always configure this. You want to know. - **`storm-control action shutdown`** \- errdisable the port when the threshold is breached. This is the aggressive option: it removes the offending host entirely rather than just rate-limiting it. Which action depends on where the port is: Access port to a host `shutdown` \+ `trap`. A single host storming is almost always a fault or an attack, and cutting it off is the safe call. Pair with errdisable recovery. Uplink / trunk Default drop + `trap` only. Do **not** shut an uplink - you would take out every host behind it to stop one storm. Rate-limit and alert instead. ## Verification ``` Switch#show storm-control Interface Filter State Upper Lower Current --------- ------------- ----------- ----------- ---------- Gi1/0/10 Forwarding 1.00% 0.50% 0.00% Switch#show storm-control broadcast Switch#show interfaces Gi1/0/10 counters storm-control ``` `Filter State: Forwarding` means no storm. During a storm it shows `Blocking` (dropping) or the port goes errdisable. The `Current` column is the live percentage, which is where you watch a storm build. ## Platform note: honest reporting Storm control is a hardware feature - it depends on the switch ASIC being able to meter traffic per port at line rate. In our CML lab we build the switched topology on virtual IOL-L2 nodes, and **those nodes do not implement storm control at all**: ``` SW3(config-if)#storm-control broadcast level 1.00 0.50 ^ % Invalid input detected at '^' marker. ``` The command is simply not present, because there is no forwarding ASIC to meter against. We are telling you this rather than showing you invented counters. On real Catalyst hardware (and on the cat9000v platform in CML, which does model the UADP ASIC), the configuration above works exactly as described, and `show storm-control` reports the live filter state. This is a general pattern worth internalising for lab work: the *control-plane* features (routing, spanning tree topology, protocol behaviour) reproduce faithfully on lightweight virtual nodes, but *data-plane hardware* features (storm control, hardware policers, TCAM-based ACLs at scale) need a platform that models the ASIC. Know which is which before you build the lab. ## Choosing a threshold Too low and you drop legitimate broadcast bursts - a hundred hosts booting and ARPing at 9 a.m. is normal, not a storm. Too high and a real storm saturates the network before the threshold trips. The practical method: **baseline first**. Watch `show storm-control` during a normal busy period and note the typical broadcast percentage. Set the rising threshold comfortably above the normal peak - often 1-2% on an access port is a reasonable start - and the falling threshold at roughly half that. Then tune based on false positives. Prefer a **pps** threshold on access ports. Broadcast storms are usually a flood of small frames, and a packets-per-second limit catches a high-rate, low-bandwidth storm that a percentage-of-bandwidth cap would sail straight past. ## Storm control in the broader L2 hardening stack Storm control is one layer. It sits alongside: - **BPDU guard** \- shuts a port that receives a BPDU it should not (a rogue switch). - **Port security** \- limits how many MACs a port may learn, catching MAC-flooding. - [**DHCP snooping**](https://www.pinglabz.com/dhcp-snooping-in-depth/) **and** [**DAI**](https://www.pinglabz.com/dynamic-arp-inspection-ip-source-guard/) \- stop rogue DHCP and ARP spoofing. - **Errdisable recovery** \- the safety net that brings a shut port back automatically. None of these is optional on an access switch facing untrusted hosts. Storm control specifically covers the flood-based denial of service that the others do not: a host that is not spoofing anything, just drowning the VLAN in broadcasts. ## Key takeaways - Broadcasts, multicasts, and unknown unicasts flood every port in the broadcast domain, so one storming host is everyone's problem. Storm control caps the rate per port. - Set a rising and a falling threshold; the gap between them is hysteresis that prevents flapping. - Thresholds can be a percentage, bps, or pps. **Prefer pps on access ports** \- storms are usually high packet rate, low bandwidth. - Default action is to drop excess. Add `action trap` always; add `action shutdown` on host access ports (with errdisable recovery), but **never on an uplink**. - Baseline normal broadcast levels before choosing a threshold. Too low drops legitimate bursts; too high lets a real storm through. - Storm control is a hardware feature. Virtual IOL-L2 in CML does not implement it (`% Invalid input`); real Catalyst hardware and cat9000v do. We report that honestly rather than fake the counters. Next: [DHCP snooping in depth](https://www.pinglabz.com/dhcp-snooping-in-depth/), the foundation of the Layer 2 security stack. The full cluster index lives on the [VLAN and switching pillar](https://www.pinglabz.com/vlans-layer-2-switching/). ### UDLD: Detecting Unidirectional Links Before They Loop URL: https://www.pinglabz.com/udld-unidirectional-link-detection/ Last updated: 2026-07-12T08:25:26.000Z Spanning tree has one blind spot, and it is a dangerous one. STP decides whether to block a port based on the BPDUs it receives. But what if a link can send in one direction and not the other? A switch that can transmit but not receive will never hear a BPDU telling it to block - so it keeps a redundant port forwarding, and you get a loop that spanning tree was specifically designed to prevent, on the exact link that was supposed to be protected. UDLD is the fix. It detects a link that has gone unidirectional and shuts the port before the loop forms. This article covers how it works, aggressive mode, and real output from a CML lab - plus an honest note about what a virtual lab can and cannot reproduce. For the fundamentals, start at the [complete spanning tree guide](https://www.pinglabz.com/spanning-tree-protocol/). ## How a link goes unidirectional It sounds exotic. It is not. The usual causes are mundane and physical: - **A fiber pair with one strand broken or dirty.** Light goes one way, not the other. Both ends see "link up" because each is receiving light on its RX - but the data path is one-way. - **A miswired patch.** TX on one end is cross-connected to something that is not the far end's RX. - **A media converter or SFP failing on one channel.** - **A GBIC seated just badly enough that one direction works.** The insidious part is that **the interface stays up**. `show interfaces` reports the port as connected. The line protocol is up. Everything looks healthy - and a redundant port that should be blocking is forwarding, because the BPDU that would have blocked it never arrived. ## How UDLD works UDLD sends its own small Layer 2 frames (to a well-known multicast MAC) that carry the sender's device ID and port ID. The protocol works by **echo**: 1. Switch A sends a UDLD frame saying "I am device A, port X". 2. Switch B receives it, and sends back a frame saying "I am device B, port Y, and I can see device A port X". 3. Switch A receives that echo. It now knows: B can see me. The link is bidirectional. If A stops seeing its own identity echoed back by B, it concludes the far end cannot receive what A is sending - the link is unidirectional - and it acts. ## The two modes Normal mode Detects a unidirectional link and marks the port *undetermined*, logging it. It does **not** shut the port. It tells you; it does not act. Aggressive mode After losing echo, it tries eight times to re-establish the neighbour, once per second. If it cannot, it **errdisables the port**. It acts. **Use aggressive mode.** Normal mode detecting a loop and merely logging it is not much use at 3 a.m. when the loop is already melting your CPU. Aggressive mode also catches a subtler case: a link where both directions were working and then one silently stops (a "connection lost" scenario), not just links that were unidirectional from the start. ## Configuration - and the copper gotcha The global command: ``` udld aggressive ``` Here is the catch that wastes people an afternoon. **The global command only arms fiber ports.** On copper, UDLD does nothing until you enable it per interface. From the lab, after the global command: ``` SW1#show udld Et0/0 Port enable administrative configuration setting: Disabled Port enable operational state: Disabled Current bidirectional state: Unknown ``` Disabled, despite the global config. The fix is per-interface: ``` interface Ethernet0/0 udld port aggressive ``` The reasoning is historical: UDLD was designed for fiber, where unidirectional failures are common. Copper's auto-negotiation already detects most one-way faults electrically, so Cisco does not auto-arm UDLD on copper. But copper faults still happen, so if you want UDLD on a copper link you ask for it explicitly. ## Verification: a healthy link ``` SW1#show udld Et0/0 Port enable administrative configuration setting: Enabled / in aggressive mode Port enable operational state: Enabled / in aggressive mode Current bidirectional state: Bidirectional Current operational state: Advertisement - Single neighbor detected Message interval: 7000 ms Time out interval: 5000 ms Entry 1 Current neighbor state: Bidirectional Device ID: 2039889 Port ID: Et0/0 Neighbor echo 1 device: 2039885 Neighbor echo 1 port: Et0/0 TLV CDP Device name: SW2 ``` Everything you want to see is here: - `Current bidirectional state: Bidirectional` \- the link is confirmed two-way. - `Neighbor echo 1 device / port` \- SW1 can see its own identity being echoed back. That is the whole detection mechanism, made visible. - `TLV CDP Device name: SW2` \- the neighbour identified itself. - `Message interval: 7000 ms` \- UDLD sends a message every 7 seconds by default in the advertisement phase. ## What a virtual lab cannot do, and how we handle it honestly Here is where PingLabz is going to be straight with you rather than fake a screenshot. A CML lab connects nodes with virtual links that are inherently bidirectional. There is no way to break one strand of a fiber that does not physically exist. **We cannot faithfully reproduce a genuine unidirectional-link failure in the lab, so we are not going to show you invented errdisable output and pretend it is real.** What we *can* show, and did above, is UDLD in aggressive mode, fully operational, in a confirmed bidirectional state, with the echo mechanism visible. What you would see when it fires on real hardware is this: ``` %UDLD-4-UDLD_PORT_DISABLED: UDLD disabled interface Gi1/0/1, unidirectional link detected %PM-4-ERR_DISABLE: udld error detected on Gi1/0/1, putting Gi1/0/1 in err-disable state ``` And `show udld` would report `Current bidirectional state: Unknown` with the port in errdisable. An expert reader can tell the difference between a lab that models a fault honestly and one that stages fake output. We would rather show you the healthy state and the real recovery config than pretend. ## Recovery: pair it with errdisable recovery When aggressive UDLD errdisables a port, it stays down until you clear it. In a lot of unidirectional failures - a dirty connector, a marginal SFP - the fault is transient or self-clears when someone reseats the optic. Rather than dispatching an engineer for a port that may already be fine, configure automatic recovery: ``` errdisable recovery cause udld errdisable recovery interval 300 ``` From the lab, UDLD is armed as a recovery cause: ``` SW1#show errdisable recovery ErrDisable Reason Timer Status bpduguard Enabled udld Enabled ... Timer interval: 60 seconds ``` Now the port re-tries every 300 seconds. If the link is genuinely bidirectional again, it comes back on its own. If not, it errdisables again - which is exactly the behaviour you want. The full recovery mechanism is covered in [switch administration](https://www.pinglabz.com/switch-administration-sdm-errdisable/). Set the recovery interval long enough (300 seconds is sensible) that a genuinely broken link is not flapping in and out of service every minute - that would be worse than leaving it down. ## UDLD vs the alternatives UDLD **Detects:** unidirectional links, at Layer 2 **Speed:** seconds (echo-based) **Scope:** the specific fault STP is blind to Loop guard **Detects:** a blocking port that stops receiving BPDUs **Action:** keeps it blocking (loop-inconsistent) rather than letting it go forwarding **Complements UDLD** \- belt and braces BFD **Detects:** loss of a bidirectional forwarding path (Layer 3) **Speed:** sub-second **Different job** \- for routing convergence, not L2 loop prevention **Run UDLD and loop guard together.** They overlap deliberately. UDLD detects the physical unidirectional condition; loop guard handles the case where a blocking port stops hearing BPDUs for any reason. Neither is a complete substitute for the other, and the belt-and-braces combination is standard on any well-run switched core. ## Key takeaways - A unidirectional link keeps the interface "up" while data flows only one way. STP cannot see it, so a port that should block keeps forwarding, and you get a loop. - UDLD detects this with echo frames: if a switch stops seeing its own identity echoed back, the link is one-way. - **Use aggressive mode.** Normal mode only logs; aggressive mode errdisables the port and also catches links that go one-way after working. - The global `udld aggressive` command **only arms fiber**. On copper you must add `udld port aggressive` per interface. - `Current bidirectional state: Bidirectional` plus a visible neighbour echo is a healthy link. - A virtual lab cannot create a genuine unidirectional fault, so we show the healthy operational state and the real recovery config rather than staging fake errdisable output. - Pair aggressive UDLD with `errdisable recovery cause udld` and a sensible interval, and run loop guard alongside it. Next: [storm control](https://www.pinglabz.com/storm-control-configuration/), stopping a broadcast flood at the port. The full cluster index lives on the [spanning tree pillar](https://www.pinglabz.com/spanning-tree-protocol/). ### MST and PVST+ Interoperation: The Boundary, the CIST, and the Gotchas URL: https://www.pinglabz.com/mst-pvst-interoperation/ Last updated: 2026-07-12T08:25:25.000Z Two switching standards, both trying to prevent loops, neither aware the other exists. That is what an MST-to-PVST+ boundary is, and it is one of the most misunderstood corners of Layer 2 - which is a problem, because you meet it every time an MST core touches an access layer that someone left running Rapid-PVST. Get the boundary right and it is invisible. Get it wrong and you get blocked ports, a "PVST simulation inconsistency" you have never heard of, and a VLAN that mysteriously will not pass traffic across the seam. This article builds a real MST/PVST boundary in a CML lab, shows the CIST, and deliberately triggers the inconsistency so you can recognise it. For the fundamentals, start at the [complete spanning tree guide](https://www.pinglabz.com/spanning-tree-protocol/). ## Why the boundary exists at all Rapid-PVST+ runs **one spanning tree per VLAN**. A hundred VLANs means a hundred independent trees, a hundred sets of BPDUs, and a hundred separate root elections. It is simple to reason about and it does not scale. MST (802.1s) runs **a handful of trees, each carrying many VLANs**. You map VLANs to instances, and each instance is one tree. A hundred VLANs might collapse into three instances. Far less BPDU overhead, far less CPU - but you have to design the VLAN-to-instance mapping deliberately. When these two meet, the MST switch has to present something a PVST switch can understand, and vice versa. The mechanism it uses is the **CIST**. ## The CIST: the tree that speaks to everyone The Common and Internal Spanning Tree is MST's diplomatic layer. Think of it in two parts: IST (Internal Spanning Tree) This is MST instance 0\. Every VLAN not explicitly mapped elsewhere lives here. It is the tree that operates *inside* the region. CST (Common Spanning Tree) The single tree that connects the whole network - every MST region and every legacy STP/PVST switch - as if each region were one giant bridge. The crucial simplification: **to the outside world, an entire MST region looks like a single switch.** All the internal complexity - the instances, the mappings - is hidden. A PVST switch peering with an MST region sees one bridge, and it exchanges ordinary STP BPDUs with it via the CIST. ## The lab Four switches in a square. SW1 and SW2 form an MST region called PINGLABZ. SW3 and SW4 run Rapid-PVST. The boundary therefore runs on two links: SW2-to-SW3 and SW4-to-SW1\. SW1 is configured to be the CIST root. ``` ! MST region config, identical on SW1 and SW2 spanning-tree mode mst spanning-tree mst configuration name PINGLABZ revision 1 instance 1 vlan 10 instance 2 vlan 20,99 ``` ### What the MST side sees ``` SW2#show spanning-tree mst 0 Interface Role Sts Cost Prio.Nbr Type Et0/0 Root FWD 2000000 128.1 P2p Et0/1 Desg FWD 2000000 128.2 P2p Bound(PVST) Et0/2 Desg FWD 2000000 128.3 P2p Et0/3 Desg FWD 2000000 128.4 P2p ``` Look at Et0/1: `Bound(PVST)`. That is the boundary port, and IOS is telling you exactly what it is - a port where the MST region meets a Rapid-PVST neighbour. This one field is the single most useful thing to check when a boundary misbehaves. ### What the PVST side sees ``` SW3#show spanning-tree vlan 10 Interface Role Sts Cost Prio.Nbr Type Et0/0 Root FWD 100 128.1 P2p Peer(STP) Et0/1 Altn BLK 100 128.2 P2p Et0/2 Desg FWD 100 128.3 P2p Edge ``` `Peer(STP)` on the port toward the region. SW3 has no idea it is talking to a two-switch MST region carrying three instances. It sees a single STP peer, exactly as designed. ### SW1 is the root, and it says so ``` SW1#show spanning-tree mst 1 ##### MST1 vlans mapped: 10 Bridge address aabb.cc00.4d00 priority 4097 (4096 sysid 1) Root this switch for MST1 ``` ## How the MST region decides two switches belong together Three things must match *exactly* for two switches to be in the same MST region: 1. The **region name**. 2. The **revision number**. 3. The **complete VLAN-to-instance mapping table**. All three are hashed into a single 16-byte **digest** that switches exchange in their BPDUs. If the digests match, same region. If they differ, different region - and the link between them becomes a boundary, whether you intended it or not. ``` SW1#show spanning-tree mst configuration digest Name [PINGLABZ] Revision 1 Instances configured 3 Digest 0xE05029508B8AC96367E964B97EE7DA9B ``` This is the diagnostic that solves the most common MST mystery. In the lab, we moved a single VLAN from one instance to another on just *one* switch: ``` SW1: Digest 0xE05029508B8AC96367E964B97EE7DA9B SW2: Digest 0xDA7E3D01C232F1CFDB7F72C25D414BAA <-- differs ``` Two switches you *thought* were one region are now two regions, and the link between them stopped being internal and became a boundary. **The VLAN mapping only has to differ by one VLAN.** When someone reports "MST is behaving strangely after a change", compare the digests first. It is almost always this. ## The PVST simulation inconsistency Now the failure mode that has its own name and that nobody recognises the first time. MST can only present *one* view of the tree to the PVST side (through the CIST). But PVST runs a separate tree per VLAN. So there is an inherent tension: what if the PVST region tries to make itself the root for one VLAN but not another? MST has no way to represent that - it only has the CIST to offer. The MST boundary handles this with a safety check called **PVST simulation**. If a PVST neighbour advertises a superior BPDU for an individual VLAN - trying to become root of a per-VLAN tree - the boundary port refuses, and blocks. We triggered it by making the PVST switch claim VLAN 10 root: ``` SW3(config)#spanning-tree vlan 10 priority 0 ``` Immediately, on the MST boundary: ``` SW2#show spanning-tree mst 0 Et0/1 Desg BKN*2000000 128.2 P2p Bound(PVST) *PVST_Inc SW2#show spanning-tree inconsistentports Name Interface Inconsistency -------------------- ------------------------------ ------------------ MST0 Ethernet0/1 PVST Sim. Inconsistent MST1 Ethernet0/1 PVST Sim. Inconsistent MST2 Ethernet0/1 PVST Sim. Inconsistent Number of inconsistent ports (segments) in the system : 3 ``` Read that carefully. `Desg BKN*` \- the port is **blocked**. `*PVST_Inc` \- because of a PVST inconsistency. And it is inconsistent in *all three* MST instances at once, because the boundary is a single physical port shared by every instance. One VLAN's misbehaviour on the PVST side has just blocked the entire boundary link, for every VLAN. The rule the boundary enforces: **the CIST root must be inside the MST region, or the boundary blocks.** A PVST switch is not allowed to win the root election for a VLAN that crosses into MST. Remove `spanning-tree vlan 10 priority 0` and the port un-blocks and returns to forwarding. ## The design rules that prevent all of this 1 **Make the MST region the root for every VLAN that crosses the boundary.** Set the CIST root and the relevant instance roots to a low priority on an MST switch. This is what stops PVST simulation inconsistencies before they start. 2 **Keep the region config byte-identical.** Name, revision, VLAN-to-instance map - copy-paste it, do not retype it. One typo makes a switch its own region. 3 **Map every VLAN deliberately.** Any VLAN you do not explicitly map lands in instance 0 (the IST). Forgetting to map a new VLAN is a classic way to get it following the wrong tree. 4 **Migrate one side fully, not halfway.** The boundary is a fine permanent state, but if you are converting an access layer to MST, do a whole switch at a time, matching the core's region config. ## Troubleshooting checklist 1. **A VLAN will not pass across the seam?** Check for `Bound(PVST)` and then `show spanning-tree inconsistentports`. A PVST\_Inc blocks the whole link. 2. **Two switches you expected to be one region are not?** `show spanning-tree mst configuration digest` on both. Different digest = different region. Then diff the name, revision, and VLAN map. 3. **Unexpected blocking after adding a VLAN?** You probably added it on one switch's MST map but not the other, splitting the region. 4. **The PVST side keeps winning root?** Lower the priority on an MST switch so the CIST root sits inside the region. The boundary requires it. 5. **Suboptimal paths?** Remember the whole region is one bridge to the outside. The CST cannot see internal paths, so tune costs at the boundary if the external topology chooses badly. ## Key takeaways - MST runs few trees carrying many VLANs; PVST+ runs one tree per VLAN. The CIST is how MST presents a single, PVST-compatible view of the whole region. - To the outside world, an entire MST region looks like one bridge. That is the CST. - `Bound(PVST)` marks a boundary port. It is the first thing to check on a misbehaving seam. - Region membership is decided by an exact match of name, revision, and the full VLAN-to-instance map, hashed into a digest. **Different digest = different region.** Diff the digests to find a split. - A **PVST simulation inconsistency** (`*PVST_Inc`, `Desg BKN*`) blocks the entire boundary link, in all instances, when the PVST side tries to become root for a VLAN crossing into MST. - Prevent it by making the MST region the root for every boundary-crossing VLAN. Next: [UDLD](https://www.pinglabz.com/udld-unidirectional-link-detection/), which catches the one physical fault spanning tree cannot see. The full cluster index lives on the [spanning tree pillar](https://www.pinglabz.com/spanning-tree-protocol/). ### IGP Migration: Moving from EIGRP to OSPF Without an Outage URL: https://www.pinglabz.com/eigrp-to-ospf-migration/ Last updated: 2026-07-12T07:47:12.000Z Sooner or later somebody decides the network is moving from EIGRP to OSPF. Maybe a merger brought in a multi-vendor estate. Maybe the standards team decided on open protocols. Maybe an architect just likes link-state better. The reason does not matter. What matters is that you have to do it on a live network, and you are not allowed to drop a packet. The good news is that this is a solved problem, and the technique is elegant: run both protocols side by side, let them both build complete routing information, then flip which one the routing table believes - one administrative distance command at a time. It is called **ships in the night**, and this article walks through it with real lab output showing zero packet loss at every stage. For the fundamentals, see the [EIGRP guide](https://www.pinglabz.com/eigrp/) and the [OSPF guide](https://www.pinglabz.com/ospf/). ## Why not just redistribute? The instinct is to redistribute between the two protocols during the transition. Resist it. Redistribution during a migration means you have a temporary, multi-point, mutual redistribution boundary that moves around the network as you go - which is [the single most dangerous configuration in routing](https://www.pinglabz.com/multi-protocol-redistribution-ccie-scenarios/). You would be introducing route feedback and AD races into a live network, deliberately, at the exact moment you can least afford surprises. Ships in the night avoids redistribution entirely. The two protocols never talk to each other. They each carry a complete picture of the network independently, and the routing table simply chooses which picture to believe. ## The four stages 1 **Baseline.** Document what the routing table looks like now. You cannot verify a migration you did not measure first. 2 **Build OSPF alongside EIGRP.** Adjacencies form, the LSDB populates, SPF runs. *Nothing in the routing table changes*, because EIGRP's AD of 90 beats OSPF's 110\. This stage is completely safe and you can leave the network here for weeks. 3 **Flip the AD.** Raise EIGRP's distance above OSPF's. The routing table switches protocol. One command per router, reversible in one command. 4 **Remove EIGRP.** Only once every router is on OSPF and has been verified. This stage changes nothing, because nothing is using EIGRP any more. The critical insight is that **stage 3 is the only stage that changes forwarding**, it changes it one router at a time, and it is instantly reversible. Everything else is preparation or cleanup. ## Stage 1: baseline From the lab. SPOKE1 reaching SPOKE2's loopback across the site LAN: ``` SPOKE1#show ip route 12.12.12.12 Routing entry for 12.12.12.12/32 Known via "eigrp 100", distance 90, metric 1024640, type internal Routing Descriptor Blocks: * 10.1.1.12, from 10.1.1.12, via Ethernet0/1 Route metric is 1024640, traffic share count is 1 SPOKE1#ping 12.12.12.12 source Loopback0 repeat 3 !!! Success rate is 100 percent (3/3), round-trip min/avg/max = 2/2/3 ms ``` EIGRP, AD 90, via Ethernet0/1\. Write it down. That is what "correct" looks like, and it is what you will compare against after every stage. ## Stage 2: build OSPF alongside ``` SPOKE1(config)#router ospf 10 SPOKE1(config-router)#router-id 11.11.11.11 SPOKE1(config-router)#network 10.1.1.0 0.0.0.255 area 0 SPOKE1(config-router)#network 11.11.11.11 0.0.0.0 area 0 ``` Same on SPOKE2\. Wait for the adjacency, then look: ``` SPOKE1#show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 12.12.12.12 1 FULL/DR 00:00:35 10.1.1.12 Ethernet0/1 SPOKE1#show ip ospf database router 12.12.12.12 | include Advertising|Link ID Advertising Router: 12.12.12.12 (Link ID) Network/subnet number: 12.12.12.12 (Link ID) Designated Router address: 10.1.1.12 ``` OSPF is fully up. It knows about 12.12.12.12\. It has run SPF and computed a path. And the routing table: ``` SPOKE1#show ip route 12.12.12.12 Routing entry for 12.12.12.12/32 Known via "eigrp 100", distance 90, metric 1024640, type internal * 10.1.1.12, from 10.1.1.12, via Ethernet0/1 ``` **Completely unchanged.** EIGRP's AD of 90 beats OSPF's 110, so the RIB never even considers the OSPF route. OSPF is running, converged, and holding a complete parallel view of the network - and it is having zero effect on forwarding. This is the stage where you do all the real work. Get OSPF right across the entire estate. Fix the area design. Sort out the MTU mismatches, the network-type mismatches, the authentication. Let it sit for a week. Watch it. Every OSPF bug you find here is a bug you find with a working network underneath you. ### What to check before moving on - **Every EIGRP adjacency has a matching OSPF adjacency.** A link with EIGRP but no OSPF becomes a black hole the moment you flip. Compare `show ip eigrp neighbors` against `show ip ospf neighbor` on every router. - **Every prefix in the EIGRP topology is in the OSPF database.** Something with a `network` statement under EIGRP but not under OSPF disappears at the flip. This is the number one cause of a failed migration. - **The OSPF paths are the paths you want.** OSPF's cost metric will not naturally reproduce EIGRP's bandwidth-and-delay composite. Some paths *will* change. Find out which ones now, on paper, not during the change window. - **Passive interfaces are configured.** Every interface that was passive in EIGRP must be passive in OSPF, or you will form adjacencies with things you did not intend to. That third point deserves emphasis. **Metric equivalence is the thing that surprises people.** EIGRP's composite metric weighs minimum bandwidth and cumulative delay. OSPF's cost is a simple sum of per-interface costs derived from bandwidth alone. Two paths that EIGRP ranked one way can rank the other way in OSPF. Model it, or set OSPF costs explicitly on the links that matter. ## Stage 3: the flip One command. On one router at a time. ``` router eigrp DMVPN address-family ipv4 unicast autonomous-system 100 topology base distance eigrp 190 190 ``` EIGRP internal and external both become AD 190, which is worse than OSPF's 110\. The RIB re-evaluates, and: ``` SPOKE1#show ip route 12.12.12.12 Routing entry for 12.12.12.12/32 Known via "ospf 10", distance 110, metric 11, type intra area Last update from 10.1.1.12 on Ethernet0/1, 00:00:13 ago Routing Descriptor Blocks: * 10.1.1.12, from 12.12.12.12, via Ethernet0/1 Route metric is 11, traffic share count is 1 SPOKE1#ping 12.12.12.12 source Loopback0 repeat 3 !!! Success rate is 100 percent (3/3), round-trip min/avg/max = 2/2/3 ms ``` **Same next hop. Same interface. Different protocol. Zero packet loss.** That is the whole trick. The forwarding path did not change, because both protocols had already computed the same path. All that changed was which routing process the RIB listened to. The data plane never noticed. ### Why `distance eigrp 190 190` and not `distance ospf` You could equally lower OSPF's distance below 90\. Do not. Lowering OSPF's AD makes it beat things you did not intend - static routes at AD 1 are safe, but eBGP at 20 is not, and in a network with any BGP you have just created a new problem while solving an old one. **Raise the protocol you are leaving. Never lower the protocol you are arriving at.** 190 is a good value: comfortably above OSPF's 110, comfortably below the 200 of iBGP and the 255 of "unusable". ### Rollback `no distance eigrp 190 190`. EIGRP returns to 90/170, and the routing table flips straight back. Same next hops, same zero loss. You can do this at 2 a.m. with a nervous change manager watching, and undo it in one line if a monitoring graph so much as twitches. That reversibility is the reason this technique is worth learning properly. Every other approach to a protocol migration involves a moment where you cannot go back. ## Stage 4: remove EIGRP Only after **every** router has been flipped and verified. Not before. A network where half the routers prefer OSPF and half prefer EIGRP is still perfectly fine - both protocols have full information - but a network where some routers have *removed* EIGRP while others still rely on it is broken. So the order is absolute: 1. Flip *every* router's AD. 2. Verify *every* router is now using OSPF (`show ip route | include ^D` should be empty everywhere). 3. Let it run. A day, a week - whatever your change process demands. 4. *Then* start removing EIGRP. Removal itself is a non-event, because nothing has been using EIGRP for days. ## Migration order Which routers do you flip first? Edges first Start with leaf sites that nothing transits. If something goes wrong, the blast radius is one branch, and you roll back one router. Core last The core carries everyone's traffic. Flip it once you have demonstrated the technique works on twenty edge routers and you know the OSPF design is sound. Both ends of a link together Not strictly required - both protocols have full information either way - but it keeps the network's state easy to reason about, and that matters at 3 a.m. ## The resource question Running two IGPs at once doubles your control-plane cost: two neighbour tables, two topology databases, two sets of hellos, two SPF/DUAL computations. On modern hardware in a normal enterprise this is a non-issue - the numbers involved are trivially small. On a heavily-loaded router with a very large topology, check `show processes cpu sorted` before and after stage 2\. If it genuinely matters, migrate in smaller batches so fewer routers are dual-stacked at once. But do not let the theoretical cost push you into redistribution. The control-plane load is the price of a safe migration, and it is a bargain. ## Key takeaways - **Never redistribute during a migration.** Ships in the night avoids it entirely. - Run both protocols simultaneously. EIGRP's AD of 90 beats OSPF's 110, so the new protocol has zero effect on forwarding until you say so. - **Stage 2 is where all the work happens** \- and it is completely safe. Get OSPF perfect while EIGRP is still driving. - Before flipping: every EIGRP adjacency has an OSPF equivalent, every prefix is in the OSPF database, and you have modelled the paths that will change (they will not be identical - the metrics are computed differently). - The flip is `distance eigrp 190 190`. One command, one router, instantly reversible with `no distance eigrp 190 190`. - **Raise the protocol you are leaving. Never lower the protocol you are arriving at** \- you will start beating BGP and static routes you did not mean to. - Verified in the lab: the route changed from `eigrp 100, distance 90` to `ospf 10, distance 110`, same next hop, same interface, 100% ping success throughout. - Remove EIGRP only after every router has been flipped and verified. That step is a non-event by then. That closes the expert EIGRP series. The full cluster indexes live on the [EIGRP pillar guide](https://www.pinglabz.com/eigrp/) and [IP routing](https://www.pinglabz.com/ip-routing/). ### Multi-Protocol Redistribution: The CCIE Scenarios That Break Networks URL: https://www.pinglabz.com/multi-protocol-redistribution-ccie-scenarios/ Last updated: 2026-07-12T07:47:11.000Z Redistribution between two routing protocols at a single point is easy. Redistribution between two routing protocols at *two* points is where networks go to die, and it is the reason redistribution has such a fearsome reputation in the CCIE lab. The failure is never dramatic. Nothing goes down. Adjacencies stay up. Routes are present everywhere you look. It is just that traffic is now taking a path you did not design, through a protocol you did not intend, and one of your routers is quietly re-injecting a route back into the protocol it originally came from. This article builds that failure for real in a CML lab, shows the exact output, and then fixes it. For the fundamentals, see [IP routing](https://www.pinglabz.com/ip-routing/) and the [EIGRP guide](https://www.pinglabz.com/eigrp/). ## The two mechanisms that cause every redistribution failure Only two things go wrong, and almost every redistribution disaster is one or both of them: Route feedback A prefix leaves protocol A, enters protocol B, travels to a second boundary router, and is redistributed back into protocol A. The protocol has been lied to about where the route came from. The AD race The fed-back copy arrives with a *better* administrative distance than the real one, so the router believes it. Redistributed routes carry the AD of the protocol they arrived in, not of the protocol they belong to. The AD table is the whole story, and it is worth having on the wall: **Connected**0 **Static**1 **EIGRP summary** **5** **eBGP**20 **EIGRP internal**90 **OSPF (all types)** **110** **EIGRP external** **170** **iBGP**200 **Stare at those two bold numbers.** An EIGRP *external* route has AD 170\. An OSPF route - *any* OSPF route, including an external - has AD 110\. Which means: **a route that has been redistributed into EIGRP will lose to the same route arriving via OSPF.** Every time. Even if the EIGRP path is a single hop and the OSPF path goes round the houses. That asymmetry is not a bug. EIGRP deliberately distrusts redistributed routes, and OSPF does not distinguish. But put the two together at two boundary points and it becomes a trap. ## The lab: a genuinely broken network A dual-hub DMVPN running EIGRP AS 100\. Both hubs also connect to an OSPF domain (router R5). HUB1 additionally peers eBGP with R6 in AS 65010, which owns 10.60.60.0/24. Both hubs do mutual EIGRP ↔ OSPF redistribution. HUB1 also redistributes BGP into both. This is exactly the configuration a well-meaning engineer produces when told "make everything reachable from everything". ``` ! On BOTH hubs router eigrp DMVPN address-family ipv4 unicast autonomous-system 100 topology base redistribute ospf 1 metric 100000 100 255 1 1500 ! router ospf 1 redistribute eigrp 100 subnets ! On HUB1 additionally router eigrp DMVPN address-family ipv4 unicast autonomous-system 100 topology base redistribute bgp 65001 metric 100000 100 255 1 1500 router ospf 1 redistribute bgp 65001 subnets ``` Nothing goes down. Everything is reachable. And it is badly broken. ### The evidence HUB2 is one EIGRP hop from HUB1 across the DMVPN. HUB1 is one hop from R6\. So HUB2 should reach 10.60.60.0/24 in two hops, over the tunnel. Instead: ``` HUB2#show ip route 10.60.60.0 Routing entry for 10.60.60.0/24 Known via "ospf 1", distance 110, metric 1 Tag 65010, type extern 2, forward metric 20 Redistributing via eigrp 100 Advertised by eigrp 100 metric 100000 100 255 1 1500 Last update from 10.25.0.2 on Ethernet0/1, 00:00:05 ago Routing Descriptor Blocks: * 10.25.0.2, from 1.1.1.1, via Ethernet0/1 ``` Every line of that is a problem: - **`Known via "ospf 1", distance 110`** \- HUB2 is using the OSPF copy. HUB1 redistributed the BGP route into EIGRP, where it became an EIGRP *external* (AD 170). HUB1 also redistributed it into OSPF, where R5 picked it up and passed it to HUB2 as an O E2 (AD 110). 110 beats 170\. OSPF wins. - **`via Ethernet0/1`, next hop 10.25.0.2** \- that is R5, in the OSPF domain. HUB2 is sending traffic for a network that sits directly behind its EIGRP neighbour *out through a completely different protocol domain and back*. - **`Redistributing via eigrp 100`** \- and this is the one that should make you cold. HUB2 is now taking that OSPF-learned route and **injecting it back into EIGRP**. The prefix has gone EIGRP → OSPF → EIGRP. It has been laundered. Its origin is lost. And it is now competing with the real route. The rest of the OSPF table on HUB2 tells the same story: ``` HUB2#show ip route ospf | begin Gateway O E2 6.6.6.6 [110/1] via 10.25.0.2, Ethernet0/1 O E2 10.60.60.0/24 [110/1] via 10.25.0.2, Ethernet0/1 O E2 10.99.1.0/24 [110/20] via 10.25.0.2, Ethernet0/1 O E2 10.99.3.0/24 [110/20] via 10.25.0.2, Ethernet0/1 ``` Networks that live behind HUB1, one EIGRP hop away, being learned through OSPF via a router in a different domain. In a bigger topology this is exactly how you get a routing loop that only manifests under a specific failure - and how you get a 3 a.m. incident that nobody can reproduce. ## The fix, part 1: route tags A route tag is a 32-bit number carried with a route across redistribution boundaries. It is the passport stamp. Every boundary router stamps what it exports and refuses what comes back. The scheme, applied identically on **every** boundary router: ``` ! Going EIGRP -> OSPF: refuse anything that came FROM OSPF, and stamp the rest as "from EIGRP" route-map EIGRP-TO-OSPF deny 10 match tag 110 route-map EIGRP-TO-OSPF permit 20 set tag 100 ! BGP routes get their own stamp route-map BGP-TO-OSPF permit 10 set tag 65001 ! Going OSPF -> EIGRP: refuse anything that came from EIGRP or BGP, stamp the rest as "from OSPF" route-map OSPF-TO-EIGRP deny 10 match tag 100 65001 route-map OSPF-TO-EIGRP permit 20 set tag 110 ``` Read those out loud and the logic is airtight. A route that originated in EIGRP is tagged 100 on the way into OSPF. When it reaches the second boundary router and tries to come back into EIGRP, the `OSPF-TO-EIGRP` map sees tag 100 and denies it. It cannot come home. The loop is structurally impossible, on every boundary router, forever. Pick tag values that mean something. Using 100 for "originated in EIGRP AS 100" and 110 for "originated in OSPF" makes the config self-documenting, and the engineer who inherits it in three years will thank you. ## The fix, part 2: fix the AD race Tags stop the feedback. They do not, by themselves, fix the fact that OSPF's AD 110 beats EIGRP-external's 170, so the boundary router still prefers the wrong copy of a route it can see both ways. Make OSPF externals less trusted than EIGRP externals: ``` router ospf 1 distance ospf external 175 ``` Now: EIGRP external (170) < OSPF external (175). The boundary router prefers the route it learned over the protocol the route actually lives in. Note this only changes *external* OSPF routes. Intra-area and inter-area OSPF routes keep their AD of 110, which is right - those are native OSPF routes and OSPF should be trusted for them. ## The complete fix ``` router ospf 1 redistribute eigrp 100 subnets route-map EIGRP-TO-OSPF redistribute bgp 65001 subnets route-map BGP-TO-OSPF distance ospf external 175 ! router eigrp DMVPN address-family ipv4 unicast autonomous-system 100 topology base redistribute ospf 1 metric 100000 100 255 1 1500 route-map OSPF-TO-EIGRP ``` And HUB2, on the same lab, with the same topology: ``` HUB2#show ip route 10.60.60.0 Routing entry for 10.60.60.0/24 Known via "eigrp 100", distance 170, metric 77312000 Tag 65010, type external Last update from 10.0.0.1 on Tunnel0, 00:00:16 ago Routing Descriptor Blocks: * 10.0.0.1, from 10.0.0.1, via Tunnel0 Hops 1 ``` **One hop. Over the tunnel. Straight to HUB1.** The route is back where it belongs. ``` HUB2#show ip route ospf | begin Gateway O 5.5.5.5 [110/11] via 10.25.0.2, Ethernet0/1 O 10.5.0.0/30 [110/20] via 10.25.0.2, Ethernet0/1 O 10.50.50.1/32 [110/11] via 10.25.0.2, Ethernet0/1 O E2 10.99.1.0/24 [175/20] via 10.25.0.2, Ethernet0/1 O E2 10.99.3.0/24 [175/20] via 10.25.0.2, Ethernet0/1 ``` Native OSPF routes still at AD 110\. Externals now at 175\. And 6.6.6.6 and 10.60.60.0/24 have vanished from the OSPF table entirely - the tag filter refused to let them back in. The EIGRP topology table even preserves the true origin: ``` HUB2#show ip eigrp topology 10.60.60.0/24 | include External|Originating Originating router is 1.1.1.1 External data: External protocol is BGP, external metric is 0 ``` ## The rules 1. **Tag on export, filter on import. On every boundary router, without exception.** One router missing the filter re-opens the loop, and it will only bite you during a failure. 2. **Know your AD table.** Specifically: EIGRP external is 170, OSPF is 110, and that means an EIGRP route that has been round the houses will beat the real one. This is the single most common cause of "why is traffic going that way?" 3. **Never redistribute without a route-map.** Even if the route-map is empty today, having it there means the day you need to filter something you are not changing the structure of the config under pressure. 4. **Always set a metric.** Redistribution into EIGRP without a metric silently drops every route (EIGRP has no default metric for redistributed routes). Redistribution into OSPF without `subnets` silently drops every non-classful prefix. Both are silent, both are total. 5. **Redistribute at as few points as possible.** Two-point mutual redistribution is where all of this pain comes from. If you can do it at one point, do it at one point. If you need two for redundancy, tag rigorously. 6. **Draw the topology before you type.** Mark every boundary router. Trace every prefix's path through every protocol. The bugs are visible on paper long before they are visible in a routing table. ## Key takeaways - Every redistribution disaster is route feedback, an AD race, or both. - **EIGRP external = AD 170\. OSPF = AD 110.** A route redistributed into EIGRP loses to the same route arriving via OSPF, however absurd the path. - We reproduced it: HUB2 reached a network behind its directly-adjacent EIGRP neighbour by going out through a different protocol domain, and then re-injected it back into EIGRP. - Fix it with **tags** (stamp on export, deny on import, on every boundary router) and **AD** (`distance ospf external 175` so the native protocol wins). - Redistribution into EIGRP without a metric drops everything. Redistribution into OSPF without `subnets` drops everything non-classful. Silently, both times. - Fewer redistribution points is always better. Two-point mutual redistribution demands rigour, not optimism. Next: [migrating from EIGRP to OSPF without an outage](https://www.pinglabz.com/eigrp-to-ospf-migration/). The full cluster index lives on the [EIGRP pillar guide](https://www.pinglabz.com/eigrp/). ### EIGRP over DMVPN: Multi-Hub Design and the Split-Horizon Problem URL: https://www.pinglabz.com/eigrp-over-dmvpn-multi-hub/ Last updated: 2026-07-12T07:47:11.000Z EIGRP over DMVPN works beautifully, right up until the moment you realise none of your spokes can reach any of your other spokes. The tunnels are up. NHRP is registered. The neighbours are all in the table. And the routes are simply not there. The cause is split horizon, doing exactly what it was designed to do in 1988 on a protocol running over a physical, point-to-point serial link - and doing precisely the wrong thing on a multipoint GRE cloud. This article covers split horizon and next-hop-self on mGRE, dual-hub design, and the loop that a dual-homed site creates - all with real output from a CML lab. For the fundamentals, see the [EIGRP guide](https://www.pinglabz.com/eigrp/) and the [DMVPN pillar](https://www.pinglabz.com/dmvpn/). ## Why split horizon breaks mGRE Split horizon is a loop-prevention rule from the distance-vector era: *do not advertise a route back out of the interface you learned it on*. On a point-to-point link that is unarguably correct. The only router out there is the one who told you about it. An mGRE tunnel is not a point-to-point link. It is one logical interface with many neighbours behind it. Every spoke reaches the hub through Tunnel0, and so does every other spoke. When the hub learns SPOKE1's LAN on Tunnel0 and split horizon stops it re-advertising out Tunnel0, it is not preventing a loop - it is preventing SPOKE2 from ever hearing about SPOKE1. The fix is one command. It has to be on the **hub**, and it has to be on the tunnel: ``` router eigrp DMVPN address-family ipv4 unicast autonomous-system 100 af-interface Tunnel0 no split-horizon ``` Spokes do not need it. A spoke has exactly one neighbour of interest per hub, and it should not be re-advertising other spokes' routes anyway. ## next-hop-self: the second command nobody tells you about Turning off split horizon gets the routes to the spokes. It does not get the traffic there directly. By default, EIGRP sets itself as the next hop on routes it advertises. So when the hub re-advertises SPOKE1's LAN to SPOKE2, it says "reachable via me". SPOKE2 sends the traffic to the hub, and the hub forwards it to SPOKE1\. Every spoke-to-spoke packet crosses the WAN twice and traverses the hub's CPU. Your DMVPN is a hub-and-spoke network wearing a dynamic-mesh costume. The second command tells the hub to preserve the originating spoke as the next hop: ``` af-interface Tunnel0 no split-horizon no next-hop-self ``` Now SPOKE2 learns SPOKE1's LAN with SPOKE1's tunnel address as the next hop. NHRP resolves that to SPOKE1's real NBMA address, a direct spoke-to-spoke tunnel forms, and traffic goes straight across. ## Both, visible in one output From the lab. SPOKE1 looking at SPOKE2's loopback, with both commands in place on both hubs: ``` SPOKE1#show ip eigrp topology 12.12.12.12/32 Descriptor Blocks: 10.1.1.12 (Ethernet0/1), from 10.1.1.12, Send flag is 0x0 Composite metric is (131153920/163840), route is Internal Hop count is 1 Originating router is 12.12.12.12 10.0.0.12 (Tunnel0), from 10.0.0.2, Send flag is 0x0 Composite metric is (13107281920/9830481920), route is Internal Hop count is 2 Originating router is 12.12.12.12 10.0.0.12 (Tunnel0), from 10.0.0.1, Send flag is 0x0 Composite metric is (13107281920/9830481920), route is Internal Hop count is 2 Originating router is 12.12.12.12 ``` Read those last two blocks carefully, because both commands are visible in them: - **`from 10.0.0.2` and `from 10.0.0.1`** \- the hubs re-advertised a route they learned on Tunnel0 back out of Tunnel0\. That is `no split-horizon` working. - `**10.0.0.12 (Tunnel0)**` \- the next hop is 10.0.0.12, which is *SPOKE2 itself*, not either hub. That is `no next-hop-self` working. ### Break split horizon, and the paths vanish Turn `split-horizon` back on at both hubs: ``` SPOKE1#show ip eigrp topology 12.12.12.12/32 Descriptor Blocks: 10.1.1.12 (Ethernet0/1), from 10.1.1.12, Send flag is 0x0 Composite metric is (131153920/163840), route is Internal ``` Both tunnel paths are gone. In our lab the spokes happen to share a site LAN, so one path survives. In a real DMVPN, where spokes are in different cities, **there would be nothing left at all**. Spoke-to-spoke reachability would simply not exist. ### Break next-hop-self, and the traffic hairpins Turn `next-hop-self` back on (the default): ``` SPOKE1#show ip eigrp topology 12.12.12.12/32 | include Descriptor|Tunnel0|Ethernet Descriptor Blocks: 10.1.1.12 (Ethernet0/1), from 10.1.1.12, Send flag is 0x0 10.0.0.2 (Tunnel0), from 10.0.0.2, Send flag is 0x0 <-- next hop is HUB2 10.0.0.1 (Tunnel0), from 10.0.0.1, Send flag is 0x0 <-- next hop is HUB1 ``` The routes are still there. But the next hop is now the hub. Every packet from SPOKE1 to SPOKE2 goes to a hub first. The tunnels can still form dynamically via NHRP shortcut in Phase 3, but the routing table is no longer pointing at the spoke, and you have made the hub a bottleneck for traffic that never needed to touch it. ## Dual-hub: single cloud vs dual cloud Single cloud (dual hub) One tunnel interface per spoke, one tunnel subnet, both hubs are NHS on the same cloud. Simpler config, fewer interfaces, and spokes are adjacent to both hubs on one mGRE. What our lab uses. Dual cloud Two tunnel interfaces per spoke, two tunnel subnets, one per hub. More config, but full transport separation - useful when each hub sits on a genuinely different WAN (MPLS + internet). Single cloud is simpler and is the right default when both hubs sit on the same transport. Dual cloud is what you want when the two paths are genuinely different networks and you need to reason about them, police them, or apply per-transport policy independently. In our single-cloud lab, the DMVPN comes up with both hubs and both spokes on one mGRE: ``` HUB1#show dmvpn Interface: Tunnel0, IPv4 NHRP Details Type:Hub, NHRP Peers:3, # Ent Peer NBMA Addr Peer Tunnel Add State UpDn Tm Attrb 1 100.64.2.1 10.0.0.2 UP 00:00:28 D 1 100.64.11.1 10.0.0.11 UP 00:00:37 D 1 100.64.12.1 10.0.0.12 UP 00:00:32 D ``` ### Making one hub primary Two hubs advertising identical metrics means spokes load-share, which is usually not what you want. Express the preference once, on the backup hub, with an [offset list](https://www.pinglabz.com/eigrp-offset-lists/): ``` HUB2(config)#access-list 50 permit any HUB2(config-router-af-topology)#offset-list 50 out 100000000 Tunnel0 ``` One command on one router. Every spoke, present and future, now prefers HUB1 and keeps HUB2 as a feasible successor - ready for instant failover, with no reconvergence. ## The dual-homed site problem Now the failure mode that a dual-hub design creates and that most DMVPN guides never mention. Site A has two spokes on one LAN. Both advertise the site LAN into the WAN. Which means SPOKE2 learns its own LAN back from the WAN: ``` SPOKE2#show ip eigrp topology 10.1.1.0/24 Descriptor Blocks: 0.0.0.0 (Ethernet0/1), from Connected, Send flag is 0x0 Originating router is 12.12.12.12 <-- its own LAN, connected 10.0.0.11 (Tunnel0), from 10.0.0.2, Send flag is 0x0 Hop count is 2 Originating router is 11.11.11.11 <-- the SAME LAN, back from the WAN 10.0.0.11 (Tunnel0), from 10.0.0.1, Send flag is 0x0 Originating router is 11.11.11.11 ``` The connected route wins today. But if SPOKE2's LAN interface flaps, it will install a WAN path to its own directly-attached network - and start sending local traffic across the internet and back. EIGRP has no AS-path, so nothing in the protocol stops this. The fix is a route tag applied on every router at the site: stamp everything leaving, drop anything arriving with your own stamp. ``` route-map TAG-SITE-A permit 10 set tag 1001 ! route-map BLOCK-SITE-A deny 10 match tag 1001 route-map BLOCK-SITE-A permit 20 ! router eigrp DMVPN address-family ipv4 unicast autonomous-system 100 topology base distribute-list route-map TAG-SITE-A out Tunnel0 distribute-list route-map BLOCK-SITE-A in Tunnel0 ``` ``` SPOKE2#show ip eigrp topology 10.1.1.0/24 Descriptor Blocks: 0.0.0.0 (Ethernet0/1), from Connected, Send flag is 0x0 <-- only the connected path remains ``` Full treatment, including what happened when we tried Cisco's official SoO mechanism on this platform, in [EIGRP Site of Origin](https://www.pinglabz.com/eigrp-site-of-origin-soo/). ## The complete hub template ``` interface Tunnel0 ip address 10.0.0.1 255.255.255.0 no ip redirects ip mtu 1400 ip nhrp authentication PINGLABZ ip nhrp map multicast dynamic ip nhrp network-id 100 ip nhrp holdtime 300 ip nhrp redirect ip tcp adjust-mss 1360 tunnel source Ethernet0/1 tunnel mode gre multipoint tunnel key 100 ! router eigrp DMVPN address-family ipv4 unicast autonomous-system 100 af-interface default passive-interface exit-af-interface af-interface Tunnel0 no passive-interface hello-interval 5 hold-time 15 no split-horizon no next-hop-self exit-af-interface network 10.0.0.0 0.0.0.255 eigrp router-id 1.1.1.1 ``` Three details in there that people skip: - **`ip mtu 1400` and `ip tcp adjust-mss 1360`.** GRE plus IPsec eats header space. Without these, large packets fragment or get dropped, and you get the classic "ping works, applications hang" symptom that costs people days. - **`af-interface default / passive-interface`.** Make everything passive by default and explicitly un-passive the interfaces you want. It is the only way to run EIGRP safely, and named mode makes it a two-line pattern. - **Tuned hellos (5/15).** The default on a multipoint interface is 60/180, which means a dead hub is not noticed for three minutes. 5/15 is aggressive but appropriate for a WAN overlay. Better still, pair it with BFD. ## Key takeaways - **`no split-horizon` on the hub's tunnel** is what lets spokes learn each other's routes. Without it, spoke-to-spoke reachability does not exist at all. - **`no next-hop-self` on the hub's tunnel** is what makes the traffic go directly. Without it the routes exist but every packet hairpins through the hub. - Both go on the **hub**, on the **tunnel interface**. Spokes need neither. - Single-cloud dual-hub is the simpler default. Dual-cloud is for genuinely separate transports. - Make one hub primary with an **outbound offset list on the backup hub** \- one command, applies to every spoke, keeps the backup as a feasible successor. - A dual-homed site will learn its own LAN back from the WAN. Tag on the way out, deny on the way in, on every router at the site. - Set `ip mtu 1400` and `ip tcp adjust-mss 1360` or you will spend a week debugging a fragmentation problem. Next: [multi-protocol redistribution and the scenarios that break networks](https://www.pinglabz.com/multi-protocol-redistribution-ccie-scenarios/). The full cluster indexes live on the [EIGRP pillar](https://www.pinglabz.com/eigrp/) and the [DMVPN pillar](https://www.pinglabz.com/dmvpn/). ### EIGRP Offset Lists: Surgical Metric Manipulation URL: https://www.pinglabz.com/eigrp-offset-lists/ Last updated: 2026-07-12T07:47:10.000Z You have two hubs. You want every spoke to prefer HUB1 and fall back to HUB2 only when HUB1 is unreachable. The interface bandwidth and delay are identical, so EIGRP sees two equal paths and load-shares across both - which is not what you asked for. You could change the interface delay on HUB2, but that affects every route through every neighbour on that interface and drags in the whole metric calculation. What you actually want is a scalpel: "add this much to the metric of these routes, on this interface, in this direction". That is an offset list. This article shows offset lists working on Cisco IOS XE in a dual-hub DMVPN, with the exact before-and-after metrics. For the fundamentals, start at the [complete EIGRP guide](https://www.pinglabz.com/eigrp/). ## What an offset list does An offset list adds a fixed value to the composite metric of routes matching an access list, on a specific interface, in a specific direction. That is the entire feature, and its simplicity is its virtue. ``` offset-list {in | out} [interface] ``` ****A standard or extended ACL naming the prefixes to offset. Use `0` to match everything. **in**Add the offset to routes *received* on this interface. Affects only this router's view. **out**Add the offset to routes *advertised* out this interface. Affects every downstream router at once. ****The value added to the composite metric. Not a percentage, not a multiplier - a flat addition. Crucially, the offset is applied to the **composite metric**, after the bandwidth/delay calculation. It does not touch the underlying vector metrics (bandwidth, delay, reliability, load, MTU) that get carried onward. It shifts the number the receiving router uses to compare paths, and nothing else. ## The named-mode syntax In classic EIGRP, the offset list goes directly under the router process. In named mode - which is what you should be writing now - it lives under the topology: ``` access-list 50 permit any ! router eigrp DMVPN address-family ipv4 unicast autonomous-system 100 topology base offset-list 50 out 100000000 Tunnel0 ``` ## The lab: making HUB2 the backup Dual-hub DMVPN. SPOKE1 has tunnels to both HUB1 and HUB2, and learns HUB1's campus LAN (192.168.99.0/24) two ways: directly from HUB1, and via HUB2 (which learned it from HUB1 across the tunnel cloud). ``` SPOKE1#show ip eigrp topology 192.168.99.0/24 Descriptor Blocks: 10.0.0.1 (Tunnel0), from 10.0.0.1, Send flag is 0x0 Composite metric is (9895936000/131072000), route is Internal 10.0.0.1 (Tunnel0), from 10.0.0.2, Send flag is 0x0 Composite metric is (13172736000/9895936000), route is Internal ``` The path via HUB2 already has a worse metric here because it is an extra hop. But in a symmetric design with equal-cost transports - which is the common case - these numbers would be identical, and the spoke would load-share. What we want is an explicit, deterministic preference that survives any topology change. Apply the offset on **HUB2**, outbound toward the spokes: ``` HUB2(config)#access-list 50 permit any HUB2(config)#router eigrp DMVPN HUB2(config-router)#address-family ipv4 unicast autonomous-system 100 HUB2(config-router-af)#topology base HUB2(config-router-af-topology)#offset-list 50 out 100000000 Tunnel0 ``` And on the spoke: ``` SPOKE1#show ip eigrp topology 192.168.99.0/24 Descriptor Blocks: 10.0.0.1 (Tunnel0), from 10.0.0.1, Send flag is 0x0 Composite metric is (9895936000/131072000), route is Internal 10.0.0.1 (Tunnel0), from 10.0.0.2, Send flag is 0x0 Composite metric is (13272736000/9995936000), route is Internal ``` **13,172,736,000 becomes 13,272,736,000.** Exactly +100,000,000\. Not approximately, not scaled - exactly the value we configured. The offset list is arithmetic, and that predictability is precisely why you would choose it over fiddling with delay. The route via HUB2 remains in the topology table as a valid alternative. It is a feasible successor, ready to take over the instant HUB1's path disappears, with no reconvergence delay. That is the entire design goal: deterministic primary, instant backup. ## Why one command on the hub beats N commands on the spokes You could achieve the same effect with an inbound offset list on each spoke, applied to routes received from HUB2\. It would work. It would also mean touching every spoke every time you add one, and it would mean the policy lives in a hundred places instead of one. **Outbound on the hub is the right answer.** One command, one router, and every current and future spoke inherits the behaviour automatically. The design intent - "HUB2 is the backup" - is expressed exactly once, on the router it is about. The general rule: **use `out` when you are expressing a property of the advertising router; use `in` when you are expressing a preference local to the receiving router.** "HUB2 is my backup hub" is a property of HUB2\. "I personally distrust the routes from that neighbour" is local. Most of the time you want `out`. ## Choosing an offset value EIGRP composite metrics are large. With the classic formula and default K values, a Gigabit link's metric is in the millions; over a DMVPN tunnel with its default 100 Kbps bandwidth, it runs into the billions. An offset of 100 or 1000 is lost in the noise. The practical rule: **look at the actual metrics first** with `show ip eigrp topology `, then choose an offset that is unambiguously larger than any legitimate metric variation you expect to see. In the lab we used 100,000,000 against metrics in the 13-billion range - about 0.8%, but comfortably more than the difference between any two real paths in that topology. Do *not* reach for an enormous value "to be safe". EIGRP treats a metric of 4,294,967,295 as infinite and will withdraw the route entirely. A grossly oversized offset does not make a path unattractive; it deletes it. ## Offset lists vs the alternatives Offset list **Scope:** chosen prefixes, one interface, one direction **Effect:** a flat, exact addition to the composite metric **Use for:** hub preference, per-prefix path steering Interface delay **Scope:** everything through that interface **Effect:** changes the vector metric carried downstream **Use for:** genuinely modelling a slower link. The blunt instrument. Bandwidth **Scope:** everything through that interface **Effect:** changes the metric AND EIGRP's pacing calculations **Use for:** almost never as a metric tool. Set it to the truth and leave it. distribute-list / route-map **Scope:** chosen prefixes **Effect:** blocks the route entirely **Use for:** filtering. Not for preference - a blocked route is not a backup. **Never use `bandwidth` as a metric knob.** EIGRP uses the configured bandwidth for its pacing calculation - it will not use more than 50% of the stated bandwidth for EIGRP traffic. Lie about it to change a metric and you also throttle your own routing protocol. Set bandwidth to the interface's real capacity, always, and use delay or an offset list when you want to shift a path. ## Troubleshooting 1. **No effect at all?** Check the interface name in the command. An offset list with no interface applies to *every* interface, which is rarely what you meant - and one with the wrong interface applies to nothing. 2. **Route disappeared instead of being deprioritised?** Your offset pushed the metric to EIGRP's infinity. Reduce it. 3. **Effect is invisible in the metrics?** The offset is too small relative to the composite metric. Look at the real numbers first, then size it. 4. **Works on one spoke, not another?** If you applied it inbound on the spokes, you have missed one. Apply it outbound on the hub instead. 5. **Named mode rejecting the command?** It goes under `topology base`, not directly under the address family. ## Key takeaways - An offset list adds a flat, exact value to the composite metric of matched routes, on one interface, in one direction. Verified: +100,000,000 configured, +100,000,000 observed. - Named mode puts it under `topology base`. - Apply it **outbound on the hub** to express "this hub is the backup" once, rather than inbound on every spoke. - Size the offset against the real composite metrics you can see with `show ip eigrp topology`. Too small and it does nothing; too large and EIGRP withdraws the route as unreachable. - The deprioritised path stays in the topology table as a feasible successor, so failover is instant. - Never manipulate `bandwidth` to change a metric - you throttle EIGRP's own pacing. Use delay or an offset list. Next: [EIGRP over DMVPN: multi-hub design and the split-horizon problem](https://www.pinglabz.com/eigrp-over-dmvpn-multi-hub/). The full cluster index lives on the [EIGRP pillar guide](https://www.pinglabz.com/eigrp/). ### EIGRP Summarization with Leak Maps URL: https://www.pinglabz.com/eigrp-summary-leak-map/ Last updated: 2026-07-12T07:47:10.000Z Summarisation is a trade. You send one prefix instead of fifty, which shrinks the routing table and hides internal churn from your spokes. But you also lose the granularity: everything inside the summary now looks equally distant, and if two hubs advertise the same summary, a spoke has no way to tell which one is actually closer to any particular subnet. A leak map lets you have both. Advertise the summary *and* punch through one or two specific prefixes that need to remain visible. It is a single keyword and it solves a real design problem elegantly. This article shows it working on Cisco IOS XE with real DMVPN lab output. For the fundamentals, start at the [complete EIGRP guide](https://www.pinglabz.com/eigrp/). ## EIGRP summarisation is per-interface Unlike OSPF, where summarisation happens at area boundaries and only at area boundaries, EIGRP can summarise **on any interface, on any router**. That flexibility is one of EIGRP's genuine advantages, and it is why EIGRP scales well in hub-and-spoke WANs where OSPF needs careful area design. In classic mode: ``` interface Tunnel0 ip summary-address eigrp 100 10.99.0.0 255.255.0.0 ``` In named mode, which is what you should be writing in 2026: ``` router eigrp DMVPN address-family ipv4 unicast autonomous-system 100 af-interface Tunnel0 summary-address 10.99.0.0 255.255.0.0 ``` The moment you configure it, two things happen. The specific prefixes covered by the summary stop being advertised out that interface, and the summarising router installs a discard route. ## The discard route (and its unusual AD) ``` HUB1#show ip route 10.99.0.0 255.255.0.0 Routing entry for 10.99.0.0/16 Known via "eigrp 100", distance 5, metric 1280, type internal Routing Descriptor Blocks: * directly connected, via Null0 Route metric is 1280, traffic share count is 1 ``` **Administrative distance 5.** Not 90, not 170, not 254\. EIGRP's summary discard route has its own AD, and it is very low - lower than almost everything else on the box. That is deliberate anti-loop protection: the summarising router is telling the world "I can reach all of 10.99.0.0/16", so it must never follow a less-specific route (like a default) back out for an address inside the summary that does not actually exist. The Null0 route with AD 5 guarantees it drops the packet instead of looping it. It also means that **an EIGRP summary will beat almost any other route to the same prefix on the summarising router**. If you configure `summary-address 10.0.0.0 255.0.0.0` on a router that also has a static route to 10.0.0.0/8 with AD 1, the static still wins (1 < 5). But an OSPF route (110), a BGP route (20), even an eBGP route - all lose to it. Configure a summary that is broader than you intended and you can black-hole a chunk of your network on that router alone. You can change it with the `summary-metric ... distance` option, but the right answer is to get the summary right. (Contrast with OSPF, whose summary discard route has AD 254 - deliberately worse than everything, so it only takes effect if nothing else exists. Two protocols, two philosophies, and it is worth knowing which one you are dealing with.) ## The problem a leak map solves Our lab is a dual-hub DMVPN. HUB1 sits in front of a campus with several subnets in 10.99.0.0/16\. Before summarisation, the spokes see every one of them: ``` SPOKE1#show ip route eigrp | include 10.99 D 10.99.1.0/24 [90/76800640] via 10.0.0.1, 00:01:24, Tunnel0 D 10.99.2.0/24 [90/76800640] via 10.0.0.1, 00:01:24, Tunnel0 D 10.99.3.0/24 [90/76800640] via 10.0.0.1, 00:01:24, Tunnel0 ``` Three subnets today. In a real network, three hundred. Every one of them floods to every spoke, and every time one flaps, every spoke runs DUAL. Summarise, and the spokes see one prefix. Perfect - except that 10.99.2.0/24 is the data centre subnet, and there is a design requirement that spokes must be able to prefer a direct path to it when one exists. Collapse it into the summary and that granularity is gone. ## The leak map ``` ip prefix-list LEAK-92 seq 5 permit 10.99.2.0/24 ! route-map LEAK-SPECIFIC permit 10 match ip address prefix-list LEAK-92 ! router eigrp DMVPN address-family ipv4 unicast autonomous-system 100 af-interface Tunnel0 summary-address 10.99.0.0 255.255.0.0 leak-map LEAK-SPECIFIC ``` And the result: ``` SPOKE1#show ip route eigrp | include 10.99 D 10.99.0.0/16 [90/76800640] via 10.0.0.1, 00:00:15, Tunnel0 D 10.99.2.0/24 [90/76800640] via 10.0.0.1, 00:01:54, Tunnel0 ``` The summary **and** the one prefix that needed to stay visible. 10.99.1.0/24 and 10.99.3.0/24 are suppressed as intended. Two routes on the spoke instead of three hundred, and the one that matters is still there in full detail. ### The leak map is a permit list, not a filter This trips people up. The route-map in a leak map answers the question "which specific prefixes should *additionally* be advertised alongside the summary?" A `permit` clause means "leak this one". Prefixes that do not match are simply not leaked - they stay suppressed under the summary, which is the default behaviour anyway. You do **not** need a terminating `permit` clause here, and adding one would leak everything and defeat the summary entirely. This is the opposite of every other route-map convention in IOS, and it catches people out. Keep the leak map tight. ## Where leak maps earn their keep Dual-hub path selection Both hubs advertise the same summary. Each hub leaks the specifics for the data centre it is actually closest to. Spokes now take the shortest path per destination instead of guessing. Anchoring a critical service Summarise the campus, but leak the /32 of a VoIP call manager or a critical VIP so spokes track its reachability precisely rather than through the summary. Migration windows Summarise the old range, leak the subnets that are mid-move, and keep exact reachability for the ones in flight while the rest of the table stays small. The dual-hub case is the one that matters most and is the one most people get wrong. Without leaking, both hubs advertise 10.99.0.0/16 with equal metrics, and every spoke picks a hub effectively at random (by metric tiebreak). Half your branches take the long way to the data centre and nobody notices until someone complains about latency. ## The suppression is per-interface, not global An important property: summarising on Tunnel0 suppresses the specifics *out of Tunnel0*. The router still has them in its own routing table, still advertises them out other interfaces, and other routers reached by other paths still see them in full. That means you can summarise aggressively toward the spokes while keeping full detail toward the core, from the same router, with one command per interface. It is one of the things EIGRP does genuinely better than OSPF. ## The metric of the summary By default, the summary's metric is the **lowest** metric among the component routes. That sounds sensible and is usually what you want, but it has a consequence: **if the single closest component route disappears, the summary's metric jumps**, which triggers an update to every spoke. The summary hides the *existence* of churn, not its *metric*. If you want the summary to be genuinely stable, pin its metric: ``` af-interface Tunnel0 summary-address 10.99.0.0 255.255.0.0 leak-map LEAK-SPECIFIC summary-metric 10.99.0.0/16 100000 100 255 1 1500 ``` Now the summary advertises a fixed metric regardless of what happens to the components. It only withdraws when *every* component route is gone. In a large hub-and-spoke WAN, that is often the difference between a stable overlay and one that recalculates all day. ## Key takeaways - EIGRP summarises **per interface**, on any router - no area boundaries required. This is a real advantage over OSPF in hub-and-spoke designs. - The summary installs a Null0 discard route with **administrative distance 5**, which beats almost everything else on the summarising router. A too-broad summary can black-hole traffic locally. - A `leak-map` advertises the summary *plus* specific prefixes you name. The route-map is a permit-list of things to leak, not a filter - do not add a terminating permit clause. - The killer use case is dual-hub path selection: both hubs advertise the summary, each leaks the specifics it is genuinely closest to, and spokes stop guessing. - The summary's metric defaults to the best component metric, so component churn still moves it. Pin it with `summary-metric` if you want a genuinely stable advertisement. - Suppression is per-interface. You can summarise toward the spokes and keep full detail toward the core, from the same router. Next: [EIGRP offset lists](https://www.pinglabz.com/eigrp-offset-lists/), for surgical metric manipulation. The full cluster index lives on the [EIGRP pillar guide](https://www.pinglabz.com/eigrp/). ### EIGRP Site of Origin (SoO): Loop Prevention for Dual-Homed WAN Sites URL: https://www.pinglabz.com/eigrp-site-of-origin-soo/ Last updated: 2026-07-12T07:47:09.000Z A branch office with two WAN routers is a good design. Two routers, two circuits, no single point of failure. It is also the exact topology that creates one of EIGRP's nastiest failure modes: the site learns its own prefixes back from the WAN, and under the right circumstances starts routing its own local traffic out over the WAN and back in again. Site of Origin (SoO) is the mechanism designed to stop that. It is an extended community that stamps each route with the identity of the site it came from, and any router at that site drops routes carrying its own stamp. This article covers what SoO does, how it is configured, and - importantly - what we found when we actually tried it on IOS XE 17.18 in a global-table DMVPN. For the fundamentals, start at the [complete EIGRP guide](https://www.pinglabz.com/eigrp/). ## The problem, seen in a real topology Our lab is a dual-hub DMVPN. Site A has two spoke routers, SPOKE1 and SPOKE2, both attached to the same site LAN (10.1.1.0/24) and both with tunnels to both hubs. It is the classic dual-homed branch. Both spokes advertise the site LAN into the WAN. Which means SPOKE1's advertisement travels up to the hubs and back down to SPOKE2 - and SPOKE2 now has a route to *its own directly-connected LAN* pointing across the WAN: ``` SPOKE2#show ip eigrp topology 10.1.1.0/24 Descriptor Blocks: 0.0.0.0 (Ethernet0/1), from Connected, Send flag is 0x0 Originating router is 12.12.12.12 <-- its own LAN, connected 10.0.0.11 (Tunnel0), from 10.0.0.2, Send flag is 0x0 Hop count is 2 Originating router is 11.11.11.11 <-- the SAME LAN, back from the WAN via HUB2 10.0.0.11 (Tunnel0), from 10.0.0.1, Send flag is 0x0 Hop count is 2 Originating router is 11.11.11.11 <-- and again via HUB1 ``` Right now the connected route wins, so nothing is broken. That is what makes this dangerous: **it is a latent fault**. It becomes an outage the moment the local path degrades: - SPOKE2's LAN interface flaps. Its connected route disappears. It installs the WAN copy. Now traffic destined for hosts on its own LAN goes out over the WAN, to a hub, back down to SPOKE1, and onto the LAN. Latency goes up by a factor of fifty and the WAN circuit carries traffic that should never have left the building. - Worse: while doing that, SPOKE2 is still advertising the LAN. Depending on metrics and timing, the two spokes can settle into a state where each believes the other is the way to the LAN. That is a routing loop, and EIGRP has no AS-path to catch it. **This is the gap SoO fills.** BGP has an AS-path, so a route that leaves an AS and comes back is rejected automatically. EIGRP has no such marker. It will happily accept its own route back. ## What SoO is Site of Origin is a BGP extended community (the same wire format as an MPLS L3VPN route target) reused by EIGRP. Every site is given an identity - `100:1`, `100:2`, and so on. Routes learned *from* a site are stamped with that site's SoO. Routes arriving *at* a site carrying that site's own SoO are dropped before they enter the topology table. The mechanism is deliberately simple: *if this route claims to have come from where I am, it has been around a loop, and I want nothing to do with it.* The configuration is a route-map that sets the extended community, applied to the WAN-facing interface: ``` route-map SITE-A permit 10 set extcommunity soo 100:1 ! interface Tunnel0 ip vrf sitemap SITE-A ``` One command on the interface, applied identically on **every** router at that site. Same SoO value, different SoO values per site. That is the whole design. ## What actually happened on IOS XE 17.18 We configured exactly that on both spokes, cleared the EIGRP neighbours, and looked for the SoO tag. It was not there. ``` SPOKE1#show running-config interface Tunnel0 | include sitemap ip vrf sitemap SITE-A <-- the command is accepted HUB1#show eigrp address-family ipv4 topology 10.1.1.0/24 | include Extended|SoO|Descriptor|Originating Descriptor Blocks: 10.0.0.11 (Tunnel0), from 10.0.0.11, Send flag is 0x0 Originating router is 11.11.11.11 ... <-- no SoO extended community anywhere ``` No tag at the hub. No tag at the receiving spoke. The WAN copies of 10.1.1.0/24 were still sitting in SPOKE2's topology table exactly as before. **The reason is in the command name.** `ip vrf sitemap` is part of Cisco's *MPLS VPN Support for EIGRP Between Provider Edge and Customer Edge* feature. It is designed and implemented for EIGRP running inside a **VRF**, on a PE router facing a CE, where the SoO extended community rides alongside the VPNv4 route targets that are already there. That is the context Cisco documents, tests, and supports. In a global-table DMVPN, with no VRF anywhere in the path, the command is parsed and stored but produces no visible extended community on this release. We are not going to pretend otherwise, and we are not going to show you output we did not get. So if you are running EIGRP in a VRF on a PE, SoO is your tool and it works. If you are running a global-table DMVPN, you need the equivalent - and there is one. ## The global-table equivalent: route tags The logic of SoO is "stamp on the way out, reject on the way in". EIGRP route tags do exactly that, and they work in the global table on every platform. On **every** router at Site A: ``` route-map TAG-SITE-A permit 10 set tag 1001 ! route-map BLOCK-SITE-A deny 10 match tag 1001 route-map BLOCK-SITE-A permit 20 ! router eigrp DMVPN address-family ipv4 unicast autonomous-system 100 topology base distribute-list route-map TAG-SITE-A out Tunnel0 distribute-list route-map BLOCK-SITE-A in Tunnel0 ``` Read it and it is SoO in different clothing. Everything this site sends into the WAN is tagged 1001\. Anything arriving from the WAN tagged 1001 is dropped, because it can only have come from here. The result, on the same router, in the same lab: ``` SPOKE2#show ip eigrp topology 10.1.1.0/24 Descriptor Blocks: 0.0.0.0 (Ethernet0/1), from Connected, Send flag is 0x0 Originating router is 12.12.12.12 <-- ONLY the connected path SPOKE2#show ip eigrp topology 11.11.11.11/32 | include Descriptor|Tunnel0|Ethernet0 Descriptor Blocks: 10.1.1.11 (Ethernet0/1), from 10.1.1.11, Send flag is 0x0 <-- only across the site LAN ``` The WAN copies are gone. SPOKE2 will never again install a WAN path to its own LAN, no matter what happens to its local interface. The loop is structurally impossible, not merely unlikely. Note that SPOKE1's loopback is now only reachable across the site LAN, not via the WAN. That is intentional and correct: if the site LAN between the two spokes is down, they are, for routing purposes, two separate sites, and they should not be papering over that with a WAN hairpin. If you *do* want that hairpin as a last resort, tag the LAN interconnect prefixes separately and let those through - but do it deliberately, with your eyes open. ## Per-site tag allocation Give each site a tag and never reuse one: **Site A**tag 1001 (or SoO 100:1) - applied on both SPOKE1 and SPOKE2's WAN interfaces **Site B**tag 1002 - on all of Site B's WAN routers **Single-homed sites**still tag them. It costs nothing, and the day someone adds a second router you are already protected. The one rule that matters: **every router at a site must use the same value, and it must be applied on every WAN-facing interface.** One router with the tag missing, or with the wrong value, and the loop is back - and it will only bite you during a failure, which is exactly when you can least afford it. ## When you do not need this - **Single-homed sites.** One router, one circuit, no backdoor. No loop is possible. (Tag them anyway - see above.) - **Sites where the two routers have no L2 path between them.** If SPOKE1 and SPOKE2 are not on a common LAN, there is no backdoor and no loop. - **BGP-based WANs.** BGP's AS-path already does this. That is precisely why SoO exists only for the protocols that lack one. ## Troubleshooting 1. **Configured SoO and nothing changed?** Check whether EIGRP is in a VRF. On IOS XE, `ip vrf sitemap` is a VRF/PE-CE feature. In the global table, use route tags instead. 2. **Routes disappearing that should not be?** Your inbound deny is matching too much. A tag-based filter with a missing terminating `permit` clause blocks everything - the same implicit-deny trap that catches people in BGP. 3. **Loop still occurring on one router.** Confirm the tag is applied on *every* WAN interface of *every* router at the site. One gap defeats the whole scheme. 4. **Backup path gone when you wanted it.** That is the design working. If a site genuinely needs a WAN hairpin as a last resort, permit specific prefixes through rather than removing the filter. ## Key takeaways - A dual-homed site running EIGRP will learn its own prefixes back from the WAN. The connected route wins - until it does not, and then you have a hairpin or a loop. - EIGRP has no AS-path. Nothing in the protocol prevents a route from coming home. - SoO stamps routes with a site identity and drops any route arriving at a site with that site's own stamp. - **On IOS XE 17.18, `ip vrf sitemap` produced no SoO extended community in the global routing table.** EIGRP SoO is implemented for the VRF / MPLS-VPN PE-CE context. - The global-table equivalent - and it works, and we verified it - is a route tag set outbound and denied inbound on the WAN interface, applied identically on every router at the site. - Apply it consistently. One router without the tag re-opens the hole. Next: [EIGRP summarisation with leak maps](https://www.pinglabz.com/eigrp-summary-leak-map/). The full cluster index lives on the [EIGRP pillar guide](https://www.pinglabz.com/eigrp/), and the DMVPN side of this topology is on the [DMVPN pillar](https://www.pinglabz.com/dmvpn/). ### Expert OSPF Troubleshooting: Five Broken Scenarios, Ticket Style URL: https://www.pinglabz.com/expert-ospf-troubleshooting-scenarios/ Last updated: 2026-07-12T07:18:25.000Z OSPF has an unhelpful habit: when you break it, it usually says nothing. An adjacency that should form simply does not. A route that should be in the table simply is not. There is one failure mode - exactly one - where OSPF actually tells you what is wrong, and we cover it below so you can enjoy the novelty. The rest of the time, diagnosis is a matter of knowing which two commands to compare. This article is five OSPF faults, each built and broken for real in a CML lab, with the actual output and the actual fix. For the theory, see the [complete OSPF guide](https://www.pinglabz.com/ospf/). ## The method Every OSPF problem is one of three things, in this order: **1\. Adjacency** `show ip ospf neighbor`. Not FULL (or 2WAY on a DROTHER pair) means stop here. Nothing else matters. **2\. Database** `show ip ospf database`. Is the LSA there? If not, it was never generated or was filtered at an ABR/ASBR. **3\. Routing table** LSA present but no route? A distribute-list, an unresolvable forwarding address, or a better route type won. The single highest-value habit: when an adjacency will not form, run `show ip ospf interface ` on **both** ends and diff them. Nine of the ten reasons an OSPF adjacency fails are visible in that one output, side by side. ``` show ip ospf interface Ethernet0/1 | include Network Type|Timer intervals|Area|Cost|Authentication ``` The five things that must match for two OSPF routers to become neighbours: **area ID, hello interval, dead interval, authentication, and stub/NSSA area flags**. Subnet mask must also match. MTU must match for the adjacency to progress past EXSTART. ## Ticket 1: "The NSSA router lost its adjacency and nothing is in the log" **Symptom.** R5 (inside area 2) and R4 (its ABR) were FULL. After a change, R5 has vanished from R4's neighbour table. The interfaces are up. Pings between the two interface addresses work. ``` R4#show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 2.2.2.2 1 FULL/BDR 00:00:34 10.0.24.1 Ethernet0/2 1.1.1.1 1 FULL/BDR 00:00:37 10.0.14.1 Ethernet0/0 ``` 5.5.5.5 is simply not there. And the log: ``` R4#show logging | include ADJCHG *Jul 12 07:02:46.213: %OSPF-5-ADJCHG: Process 1, Nbr 5.5.5.5 on Ethernet0/1 from FULL to DOWN, Neighbor Down: Dead timer expired ``` "Dead timer expired" is OSPF's way of saying "I stopped receiving valid Hellos and I am not going to tell you why". It is the least useful message in the protocol, and it is the one you will see most. **Diagnosis.** Compare both ends: ``` R5#show ip ospf | include Area 2|It is a Area 2 It is a stub area R4#show ip ospf | include Area 2|It is a Area 2 It is a NSSA area ``` **Cause.** R5 was configured with `area 2 stub` while R4 has `area 2 nssa`. These set different bits in the Hello packet's Options field - the E-bit for stub, the N-bit for NSSA - and a router discards any Hello whose area flags do not match its own. The Hellos are being sent and received. They are being silently thrown away. **Fix:** make the area type identical on every router in the area. This is not negotiable and there is no partial credit. ``` R5(config-router)#no area 2 stub R5(config-router)#area 2 nssa ``` **Lesson:** area type is one of the five things that must match. It is invisible in `show ip ospf neighbor` and it produces no dedicated error. The tell is that everything else looks perfect. ## Ticket 2: "We summarised at the ABR and now a whole subnet is black-holed" **Symptom.** After an ABR summarisation change, hosts in one range are unreachable - but instead of the traffic being dropped near the source, it is being pulled across the network and dropped at the ABR. ``` R2(config-router)#area 1 range 10.0.224.0 255.255.224.0 ``` **Diagnosis.** ``` R1#show ip route ospf | include 10.0.224|10.0.234 O IA 10.0.224.0/19 [110/20] via 10.0.12.2, Ethernet0/1 ``` The specific 10.0.234.0/24 has been replaced by a /19 that covers a great deal of address space that does not exist. And on the ABR: ``` R2#show ip route 10.0.224.0 255.255.224.0 Routing entry for 10.0.224.0/19 Known via "ospf 1", distance 110, metric 10, type intra area Routing Descriptor Blocks: * directly connected, via Null0 ``` The data plane: ``` R1#ping 10.0.230.5 source Loopback0 repeat 2 U. Success rate is 0 percent (0/2) ``` `U` means an ICMP unreachable came back - from R2, which pulled the traffic in on the strength of its /19 advertisement and then dropped it on Null0. **Cause.** The summary range is far wider than the address space that actually exists behind it. The Null0 discard route is correct anti-loop behaviour (without it the ABR would follow its own default route and loop), but it means a sloppy range actively attracts and destroys traffic for addresses that would otherwise have failed cleanly at the source. **Fix:** tighten the range to exactly the address space that exists. `area 1 range 10.0.234.0 255.255.255.0`. **Lesson:** a summary is a promise. The ABR is telling the whole network "I can reach everything in this /19". Only make promises you can keep. Full detail on all four filtering mechanisms in [OSPF filtering compared](https://www.pinglabz.com/ospf-filtering-comparison/). ## Ticket 3: "Two adjacencies dropped at once" (the one OSPF actually tells you about) **Symptom.** R2 lost both of its adjacencies to R3 simultaneously - across two completely different links. ``` R2#show logging | include DUP_RTRID|ADJCHG *Jul 12 07:04:06.723: %OSPF-4-DUP_RTRID_NBR: OSPF detected duplicate router-id 2.2.2.2 from 10.0.23.2 on interface Ethernet0/1 *Jul 12 07:04:46.690: %OSPF-5-ADJCHG: Process 1, Nbr 3.3.3.3 on Ethernet0/1 from FULL to DOWN, Neighbor Down: Dead timer expired *Jul 12 07:04:46.690: %OSPF-5-ADJCHG: Process 1, Nbr 3.3.3.3 on Ethernet0/2 from FULL to DOWN, Neighbor Down: Dead timer expired ``` **Cause.** Somebody set R3's router-id to 2.2.2.2, which is already R2's. OSPF's entire LSDB is keyed on router-id. Two routers claiming the same ID means their Router LSAs overwrite each other, the database becomes incoherent, and OSPF refuses to build an adjacency with an impostor. **Fix:** unique router-ids everywhere. Set them explicitly with `router-id x.x.x.x` \- never let OSPF pick one for you from whatever the highest loopback happens to be that day. ``` R3(config-router)#router-id 3.3.3.3 R3#clear ip ospf process ``` **Lesson:** enjoy it. `%OSPF-4-DUP_RTRID_NBR` is the one time OSPF names your mistake in plain language. The tell that should send you looking for it: *multiple adjacencies to the same neighbour dropping at once, across unrelated links*. Nothing physical does that. ## Ticket 4: "Routers in the NSSA can't reach anything outside OSPF" **Symptom.** R5, an internal router in NSSA area 2, has full reachability inside OSPF but cannot reach any redistributed external network. ``` R5#show ip route | include Gateway of last Gateway of last resort is not set R5#ping 172.20.20.1 source Loopback0 ..... Success rate is 0 percent (0/5) ``` **Diagnosis.** An NSSA, by definition, blocks Type-5 external LSAs. That is what makes it a *stub*\-family area. R5 therefore has no route to 172.20.20.0/24, and - here is the part people get wrong - **an NSSA does not automatically get a default route either**. A regular stub area gets a default injected by the ABR automatically. A **totally stubby** area gets one automatically. An **NSSA does not**. You have to ask for it. **Fix,** on the NSSA ABR: ``` R4(config-router)#area 2 nssa default-information-originate ``` ``` R5#show ip route 0.0.0.0 Routing entry for 0.0.0.0/0, supernet Known via "ospf 1", distance 110, metric 1, candidate default path, type NSSA extern 2, forward metric 10 * 10.0.45.1, from 4.4.4.4, via Ethernet0/0 R5#ping 172.20.20.1 source Loopback0 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 5/5/6 ms ``` Note the route type: `NSSA extern 2`. The default arrives as a **Type-7** LSA, not a Type-3, which is consistent with the NSSA's whole design (Type-7 is the only external-carrying LSA an NSSA permits). **Lesson:** know which area types auto-inject a default and which do not. Stub: yes. Totally stubby: yes. NSSA: **no**. Totally stubby NSSA: yes. It is arbitrary and you simply have to remember it. ## Ticket 5: "The adjacency won't form and both configs look right" **Symptom.** R1 and R2 are directly connected, interfaces up, pings work, both have the link in the correct area. No adjacency. **Diagnosis.** Diff the two interfaces: ``` R1#show ip ospf interface Ethernet0/1 | include Network Type|Timer intervals Process ID 1, Router ID 1.1.1.1, Network Type NON_BROADCAST, Cost: 10 Timer intervals configured, Hello 30, Dead 120, Wait 120, Retransmit 5 R2#show ip ospf interface Ethernet0/0 | include Network Type|Timer intervals Process ID 2, Router ID 2.2.2.2, Network Type BROADCAST, Cost: 10 Timer intervals configured, Hello 10, Dead 40, Wait 40, Retransmit 5 ``` ``` R2#show logging | include ADJCHG *Jul 12 07:00:14.361: %OSPF-5-ADJCHG: Process 1, Nbr 1.1.1.1 on Ethernet0/0 from FULL to DOWN, Neighbor Down: Dead timer expired ``` **Cause.** Someone set `ip ospf network non-broadcast` on one side. Here is the mechanism, and it is subtle: The **network type is not carried in the Hello packet**. OSPF has no way to detect a network type mismatch directly. But the network type *determines the default timers*, and **the timers are in the Hello**. Broadcast and point-to-point default to hello 10 / dead 40\. Non-broadcast and point-to-multipoint default to hello 30 / dead 120\. A router discards any Hello whose timers do not match its own. So the mismatch is detected, but as a timer mismatch, and reported as a dead-timer expiry. **The nasty variant:** point-to-point and broadcast have the *same* default timers. So if you mismatch *those* two, the adjacency comes **up** \- and then the two routers generate incompatible LSAs (one expects a Network LSA and a DR; the other does not), and routing breaks in a way that is far harder to diagnose than a down adjacency. Always check the network type even when the neighbour is FULL. **Fix:** make the network type identical on both ends. ``` R1(config-if)#no ip ospf network non-broadcast ``` ## Bonus: the 2WAY/DROTHER pause everyone panics about On a freshly-booted broadcast segment you will routinely see this for the first 30 to 40 seconds: ``` R4#show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 2.2.2.2 1 2WAY/DROTHER 00:00:39 10.0.24.1 Ethernet0/2 1.1.1.1 1 2WAY/DROTHER 00:00:39 10.0.14.1 Ethernet0/0 ``` This is not broken. It is the DR election wait timer (equal to the dead interval, 40 seconds by default) running its course. Wait for it. Two DROTHERs on a segment will also legitimately *stay* at 2WAY forever - they only form full adjacencies with the DR and BDR, never with each other. That is correct behaviour and not a fault. ## The commands, collected ``` ! Adjacency show ip ospf neighbor show ip ospf interface | include Network Type|Timer intervals|Area|Cost|Authentication show ip ospf | include Area|It is a|Number of areas show logging | include OSPF|ADJCHG|DUP_RTRID ! Database show ip ospf database show ip ospf database router show ip ospf database summary show ip ospf database external show ip ospf database external | include Forward Address <-- run this before you filter anything show ip ospf database nssa-external ! Routing table show ip route (look at: type, metric, forward metric) show ip route ospf show ip cef ! Performance show ip ospf statistics (SPF runs, timings, and the reason codes) ``` ## Try it yourself Build a multi-area topology with an NSSA, then break it. Ten minutes per fault, and name the cause from show output before you look at the config: 1. Set `area X stub` on one router in an NSSA. 2. Summarise at an ABR with a range far wider than the real address space, then ping a non-existent host inside it. 3. Give two routers the same router-id. 4. Build an NSSA and try to reach an external network from inside it. 5. Set `ip ospf network non-broadcast` on one end of an Ethernet link. Then try `point-to-point` on one end instead, and note that the adjacency comes up and routing still breaks. 6. Bonus: enable `prefix-suppression` in a topology that has an external route with a non-zero forwarding address, and watch the external route vanish from the RIB while its LSA stays in the database. That one is covered in [the OSPF forwarding address](https://www.pinglabz.com/ospf-forwarding-address/). ## Key takeaways - Work the layers in order: adjacency, database, routing table. - Area ID, hello interval, dead interval, authentication, area type flags and subnet mask must all match. MTU must match to get past EXSTART. - `show ip ospf interface` on both ends, side by side, is the single most valuable diagnostic in OSPF. - Area-type mismatch and network-type mismatch both fail **silently**, reported only as a dead-timer expiry. - Duplicate router-id is the one failure OSPF names in a log. Multiple adjacencies to one neighbour dropping at once is the signature. - An NSSA does **not** get an automatic default route. Stub and totally-stubby areas do. - A summary range that is wider than reality black-holes traffic at the summarising router, via a Null0 discard route that is doing exactly what it should. - 2WAY/DROTHER for 40 seconds after boot is a DR election, not a fault. That closes the expert OSPF series. The full cluster index, from first principles to here, lives on the [OSPF pillar guide](https://www.pinglabz.com/ospf/). ### OSPF Path Preference: O vs O IA vs E1 vs E2 vs N1 vs N2, Proven in the Lab URL: https://www.pinglabz.com/ospf-path-preference-rules/ Last updated: 2026-07-12T07:18:25.000Z Two routes to the same prefix. One has a metric of 20\. The other has a metric of 31\. Which one does OSPF install? If you answered "the one with metric 20", you have just failed a very common CCIE lab question, and you would have been surprised by what our lab actually did. OSPF does not compare metrics first. It compares **route types** first, and only falls back to the metric to break ties *within* a type. A worse metric routinely wins. This article proves the full ordering - O, O IA, E1, N1, E2, N2 - by injecting the same prefix six different ways into a live CML topology and watching each one take over as the previous is removed. For the fundamentals, start at the [complete OSPF guide](https://www.pinglabz.com/ospf/). ## The order 1 **O** \- Intra-area. The destination is inside my own area. Type-1/Type-2 LSAs. 2 **O IA** \- Inter-area. The destination is in another OSPF area. Type-3 LSA. 3 **E1** \- External type 1\. Redistributed. Metric = external cost + internal cost to the ASBR. 4 **N1** \- NSSA external type 1\. Same as E1, but learned from a Type-7 LSA inside an NSSA. 5 **E2** \- External type 2\. Metric = external cost *only*. Internal cost is ignored. This is the default for redistribution. 6 **N2** \- NSSA external type 2\. Same as E2, from a Type-7 LSA. Two things about that list are worth pausing on. **Type-1 externals beat type-2 externals.** That is because an E1 route's metric incorporates the cost of getting to the ASBR, so OSPF can meaningfully compare two E1 routes to the same destination and pick the closer ASBR. An E2 route's metric is a flat number set by the redistributing router; every router in the domain sees the same value, so OSPF has no basis to prefer one over another and treats them as equivalent regardless of distance. E1 carries more information, so it wins. **E1 beats N1, but N1 beats E2.** The type-1/type-2 distinction is the primary split; the E/N distinction only breaks ties inside it. Read the list as "type 1 externals, then type 2 externals, and within each, backbone externals before NSSA externals". ## The lab To see all six, you need a vantage point that can receive all six kinds of LSA. That rules out most routers. An NSSA ABR is the one place it works: it sits in the backbone (so it sees intra-area, inter-area, and Type-5 externals) and it also sits inside an NSSA (so it sees Type-7 externals as N1/N2). Our observer is **R4**, ABR between area 0 and area 2 (NSSA). The prefix **172.16.99.0/24** gets injected six ways: - **O** \- a loopback on R1, inside area 0, with `ip ospf network point-to-point` so the full /24 is advertised rather than a /32. - **O IA** \- the same prefix on a loopback on R3, inside area 1, arriving at R4 as a Type-3 from the ABR. - **E1 / E2** \- R2 redistributes a static, switching between `metric-type 1` and `metric-type 2`. - **N1 / N2** \- R5, inside the NSSA, redistributes a static, switching metric-type the same way. ## The sequence ### Stage 1: O wins ``` R4#show ip route 172.16.99.0 Routing entry for 172.16.99.0/24 Known via "ospf 1", distance 110, metric 11, type intra area * 10.0.14.1, from 1.1.1.1, via Ethernet0/0 ``` ### Stage 2: shut R1's loopback, O IA takes over ``` R4#show ip route 172.16.99.0 Routing entry for 172.16.99.0/24 Known via "ospf 1", distance 110, metric 21, type inter area * 10.0.24.1, from 2.2.2.2, via Ethernet0/2 ``` ### Stage 3: shut R3's loopback, E1 takes over ``` R4#show ip route 172.16.99.0 Routing entry for 172.16.99.0/24 Known via "ospf 1", distance 110, metric 30, type extern 1 * 10.0.24.1, from 2.2.2.2, via Ethernet0/2 ``` Metric 30 = external cost 20 + internal cost 10 to reach the ASBR. That addition is what makes it an E1. ### Stage 4: remove R2's redistribution, N1 takes over ``` R4#show ip route 172.16.99.0 Routing entry for 172.16.99.0/24 Known via "ospf 1", distance 110, metric 31, type NSSA extern 1 * 10.0.45.2, from 5.5.5.5, via Ethernet0/1 ``` ### Stage 5: R5 to metric-type 2, R2 back as metric-type 2\. E2 beats N2. ``` R4#show ip route 172.16.99.0 Routing entry for 172.16.99.0/24 Known via "ospf 1", distance 110, metric 20, type extern 2, forward metric 10 * 10.0.24.1, from 2.2.2.2, via Ethernet0/2 ``` ### Stage 6: remove R2's redistribution, N2 is all that's left ``` R4#show ip route 172.16.99.0 Routing entry for 172.16.99.0/24 Known via "ospf 1", distance 110, metric 20, type NSSA extern 2, forward metric 11 * 10.0.45.2, from 5.5.5.5, via Ethernet0/1 ``` Six route types, six takeovers, in the exact order the RFC specifies. ## The proof that type beats metric The sequence above is tidy but it does not, on its own, prove the central claim - because the metrics happened to be increasing anyway. So we set it up deliberately: an **E2 with metric 20** and an **N1 with metric 20**, both in the database at the same time. ``` R4#show ip ospf database external 172.16.99.0 | include Advertising|Metric Type|Metric: Advertising Router: 2.2.2.2 Metric Type: 2 (Larger than any link state path) Metric: 20 Advertising Router: 4.4.4.4 Metric Type: 1 (Comparable directly to link state metric) Metric: 20 R4#show ip ospf database nssa-external 172.16.99.0 | include Advertising|Metric Type|Metric: Advertising Router: 5.5.5.5 Metric Type: 1 (Comparable directly to link state metric) Metric: 20 ``` And the winner: ``` R4#show ip route 172.16.99.0 Routing entry for 172.16.99.0/24 Known via "ospf 1", distance 110, metric 31, type NSSA extern 1 * 10.0.45.2, from 5.5.5.5, via Ethernet0/1 ``` **OSPF installed the N1 at a routing metric of 31, and rejected the E2 at 20.** The route with the worse metric won, because its type is better. There is no metric comparison across types at all. That is the rule, and there is the evidence. (Also visible in that output: `Advertising Router: 4.4.4.4` in the Type-5 table. That is R4's own Type-7 to Type-5 translation - as the NSSA ABR it re-originates the NSSA's external into the backbone. It does not use its own translation; it uses the original Type-7.) ## E1 vs E2: choosing when you redistribute This is the practical decision the ordering forces on you, and the default is often wrong. E2 (the default) **Metric:** flat, set at the ASBR, identical everywhere **Use when:** there is only one ASBR, or all ASBRs are equivalent and you do not care which one traffic uses **Danger:** with two ASBRs at very different distances, traffic may prefer the far one E1 **Metric:** external cost + internal cost to that ASBR **Use when:** multiple ASBRs inject the same external and you want each router to use its nearest one **Cost:** a slightly larger SPF, and a metric that changes as your IGP does **The rule of thumb: if two or more ASBRs redistribute the same prefixes, use E1.** With E2, every router in the domain sees identical metrics and breaks the tie on the *forward metric* \- the internal cost to reach the ASBR or its forwarding address. That mostly does the right thing, but it is a tiebreak rather than a first-class comparison, and it is much harder to reason about. E1 makes the intent explicit. ### The E2 tiebreak, visible in the output Look back at Stages 5 and 6\. Both routes have `metric 20`, and both print a `forward metric`: ``` type extern 2, forward metric 10 <-- via R2 type NSSA extern 2, forward metric 11 <-- via R5 ``` That forward metric is the tiebreak field. When two E2 routes have the same external metric, OSPF picks the one with the lower cost to the advertising ASBR (or, if there is one, the forwarding address). It is the only thing standing between you and a random choice. ## What is not in the comparison Two things that people expect to matter and do not: **Administrative distance does not break these ties.** All OSPF routes carry AD 110, whatever their type. AD only comes into play when comparing OSPF against a different protocol. Inside OSPF, the type ordering is the whole story. **The number of hops does not matter, at all.** OSPF has no hop count. A four-hop O IA route beats a one-hop E1 route every time, because inter-area beats external, full stop. ## Where this bites in production The classic outage: a network runs OSPF internally and also redistributes some routes in from BGP or a partner network as E2\. Somebody adds a summary or a static that accidentally overlaps an internal prefix and redistributes it. Now the same prefix exists as both an O IA and an E2\. **The O IA wins, everywhere, regardless of metric.** Traffic that should have gone to the partner network is now being routed internally into a black hole. And the E2 route looks perfectly healthy in the LSDB - it is just never selected. The second classic: mutual redistribution between OSPF and another protocol, where a route leaves OSPF, comes back in as an external, and is now preferred over the internal path on some subset of routers. Route tags exist precisely to prevent this - see the multi-protocol redistribution material under [IP routing](https://www.pinglabz.com/ip-routing/). ## Key takeaways - The order is **O > O IA > E1 > N1 > E2 > N2**. Type is compared first; metric only breaks ties *within* a type. - We proved it: an N1 at metric 31 beat an E2 at metric 20\. A worse metric wins if the type is better. - E1 metric = external cost + internal cost to the ASBR. E2 metric = external cost only, identical on every router. - E2 is the default. If more than one ASBR injects the same prefix, **use E1** so each router picks its nearest ASBR. - Equal E2 routes are broken by the **forward metric** \- the internal cost to the ASBR or its forwarding address. It is printed in `show ip route`. - An NSSA ABR is the only vantage point that can see all six types at once. - Administrative distance and hop count play no part in this. AD is 110 for every OSPF route. - An internal route (O or O IA) will always beat an external one for the same prefix. An accidental overlap between an internal subnet and a redistributed one is a silent black hole. Next: [five broken OSPF scenarios, ticket style](https://www.pinglabz.com/expert-ospf-troubleshooting-scenarios/). The full cluster index lives on the [OSPF pillar guide](https://www.pinglabz.com/ospf/). ### OSPF Filtering: area range vs filter-list vs distribute-list vs summary-address URL: https://www.pinglabz.com/ospf-filtering-comparison/ Last updated: 2026-07-12T07:18:24.000Z OSPF gives you four different ways to stop a prefix from getting somewhere, and they are not interchangeable. Two of them remove the LSA from the link-state database. One leaves the LSA in place and only touches the local routing table. One only works on external routes. Pick the wrong one and you either fail to achieve what you wanted, or you break something two areas away. The distinction that separates people who know OSPF from people who have memorised OSPF commands is this: **which of these changes the LSDB, and which only changes the RIB?** This article answers that with side-by-side lab output for all four mechanisms applied to the same route set. For the fundamentals, start at the [complete OSPF guide](https://www.pinglabz.com/ospf/). ## Why OSPF is hard to filter at all OSPF is a link-state protocol. Every router in an area must have an identical LSDB, or SPF produces different answers on different routers and you get routing loops. This is not a policy preference, it is a correctness requirement. So OSPF fundamentally **cannot** filter LSAs inside an area. There is no such thing as "do not flood this Type-1 LSA to that neighbour". If you want to stop a prefix from reaching a router in the same area, your only option is to stop that router from *installing* it - the LSA still arrives, still floods, still consumes memory. That is why the two mechanisms that genuinely remove an LSA both operate at an **area boundary**, where a new LSA type is being generated and the ABR gets to decide what goes in it. And it is why the mechanism that works anywhere only touches the local RIB. ## The four mechanisms area range **Where:** ABR only **Affects:** LSDB + RIB, everywhere **LSA type:** Type 3 **Does:** summarises, or with `not-advertise`, suppresses **Side effect:** installs a Null0 discard route on the ABR area filter-list **Where:** ABR only **Affects:** LSDB + RIB, everywhere **LSA type:** Type 3 **Does:** filters individual prefixes in or out of an area **Side effect:** none distribute-list in **Where:** any router **Affects:** **local RIB only** **LSA type:** none - the LSDB is untouched **Does:** stops the local router installing a route **Side effect:** can create routing loops if used carelessly summary-address **Where:** ASBR only **Affects:** LSDB + RIB, everywhere **LSA type:** Type 5 / Type 7 **Does:** summarises external routes **Side effect:** installs a Null0 discard route on the ASBR Two sentences to memorise: **area range and summary-address summarise, and are the only tools for their LSA type. filter-list filters at the area boundary. distribute-list is local-only and does not touch the database.** ## The lab R2 is an ABR between area 0 and area 1\. R1 sits in area 0 and is our observation point. Area 1 contributes several prefixes, which arrive at R1 as Type-3 summary LSAs: ``` R1#show ip ospf database summary | include Link State ID|Advertising Router Link State ID: 3.3.3.3 Advertising Router: 2.2.2.2 Link State ID: 5.5.5.5 Advertising Router: 4.4.4.4 Link State ID: 10.0.23.0 Advertising Router: 2.2.2.2 Link State ID: 10.0.45.0 Advertising Router: 4.4.4.4 Link State ID: 10.0.234.0 Advertising Router: 2.2.2.2 Link State ID: 172.16.99.0 Advertising Router: 2.2.2.2 ``` ## 1\. area range not-advertise - kills the LSA ``` R2(config-router)#area 1 range 10.0.234.0 255.255.255.0 not-advertise ``` ``` R1#show ip ospf database summary | include Link State ID Link State ID: 3.3.3.3 Link State ID: 5.5.5.5 Link State ID: 10.0.23.0 Link State ID: 10.0.45.0 Link State ID: 172.16.99.0 <-- 10.0.234.0 gone from the LSDB R1#show ip route ospf | include 10.0.234 <-- and gone from the RIB ``` The Type-3 was never generated. Every router in area 0 loses it. This is a real, database-level removal. ### The Null0 discard route When `area range` is used to *summarise* (rather than suppress), the ABR installs a discard route for the summary. From the lab, with a deliberately over-broad range: ``` R2(config-router)#area 1 range 10.0.224.0 255.255.224.0 R2#show ip route 10.0.224.0 255.255.224.0 Routing entry for 10.0.224.0/19 Known via "ospf 1", distance 110, metric 10, type intra area Routing Descriptor Blocks: * directly connected, via Null0 ``` This is deliberate anti-loop behaviour. The ABR is advertising a /19 to the rest of the network, so it will attract traffic for the entire /19 - including addresses inside it that do not exist. Without the Null0 route, that traffic would follow the ABR's default route back out and loop. The discard route drops it instead. The operational consequence is important and easy to miss. Traffic to a non-existent host inside the summary is now pulled *all the way to the ABR* before being dropped: ``` R1#ping 10.0.230.5 source Loopback0 repeat 2 U. Success rate is 0 percent (0/2) ``` The `U` is an ICMP unreachable from R2\. Your summary range must be tight. A sloppy range does not just advertise the wrong thing, it black-holes real address space. ## 2\. area filter-list - also kills the LSA ``` ip prefix-list NO-AREA1-TRANSIT seq 5 deny 10.0.23.0/30 ip prefix-list NO-AREA1-TRANSIT seq 10 permit 0.0.0.0/0 le 32 router ospf 1 area 1 filter-list prefix NO-AREA1-TRANSIT out ``` ``` R1#show ip ospf database summary | include Link State ID Link State ID: 3.3.3.3 Link State ID: 5.5.5.5 Link State ID: 10.0.45.0 Link State ID: 10.0.234.0 Link State ID: 172.16.99.0 <-- 10.0.23.0 gone from the LSDB ``` Same result, different tool. So why have both? - **`area range`** can only act on a contiguous block that you can express as one prefix and mask. It is a summarisation tool that happens to be able to suppress. - `**area filter-list**` takes a prefix-list, so it can express arbitrary rules - `deny 10.0.23.0/30`, `permit 10.0.0.0/8 ge 24 le 26`, whatever. It is a filtering tool and does not summarise at all. And the direction keyword matters enormously: **area 1 filter-list ... out** Filters Type-3 LSAs *generated from* area 1 *into all other areas*. Think: "prefixes leaving area 1". **area 1 filter-list ... in** Filters Type-3 LSAs *entering area 1* from all other areas. Think: "prefixes arriving into area 1". The direction is from the perspective of the **area named in the command**, not the router. This trips up almost everybody the first time. ## 3\. distribute-list in - the LSA survives Now the one that behaves differently, and the reason this article exists. ``` ip prefix-list HIDE-234 seq 5 deny 10.0.234.0/24 ip prefix-list HIDE-234 seq 10 permit 0.0.0.0/0 le 32 router ospf 1 distribute-list prefix HIDE-234 in ``` Applied on R1\. And now look: ``` R1#show ip ospf database summary | include Link State ID Link State ID: 3.3.3.3 Link State ID: 5.5.5.5 Link State ID: 10.0.45.0 Link State ID: 10.0.234.0 <-- STILL IN THE LSDB Link State ID: 172.16.99.0 R1#show ip route ospf | include 10.0.234 <-- but NOT in R1's routing table R4#show ip route ospf | include 10.0.234 O IA 10.0.234.0/24 [110/20] via 10.0.24.1, Ethernet0/2 <-- R4 is completely unaffected ``` **That is the whole lesson in three commands.** The LSA is still flooded, still in the database, still consuming memory, still contributing to SPF. R1 simply refuses to install the resulting route. Every other router in the area is oblivious. Three consequences follow directly: 1. **It saves no memory and no CPU.** The LSA is still there. If your goal is to shrink the database, distribute-list does nothing for you. 2. **It is per-router.** You have to configure it on every router where you want the effect, and the moment someone adds a router without it, the behaviour is inconsistent. 3. **It can create routing loops.** If R1 has no route to a prefix but R4 does, and R1's default points at R4, traffic will go R1 → R4 → back toward R1 if R4's best path happens to be through R1\. Link-state protocols rely on every router agreeing; distribute-list deliberately breaks that agreement. There is exactly one place distribute-list is genuinely the right answer: **you want a single router (or a small set) to not use a route, while everyone else does.** A management router that should not learn a customer VRF's prefixes. A route that a specific box must reach via a static instead. That is it. If you find yourself deploying the same distribute-list on every router in an area, you wanted a filter-list at the ABR. ## 4\. summary-address - for externals only `area range` and `filter-list` only touch Type-3 LSAs. They can do nothing about Type-5 external routes, which flood through the whole OSPF domain untouched by area boundaries. For those, you go to the source: the ASBR. In the lab, R3 (the ASBR) redistributes four /24 statics: ``` R1#show ip route ospf | include E2 O E2 172.20.20.0 [110/20] via 10.0.12.2, Ethernet0/1 O E2 172.20.21.0 [110/20] via 10.0.12.2, Ethernet0/1 O E2 172.20.22.0 [110/20] via 10.0.12.2, Ethernet0/1 O E2 172.20.23.0 [110/20] via 10.0.12.2, Ethernet0/1 ``` ``` R3(config-router)#summary-address 172.20.16.0 255.255.240.0 ``` ``` R1#show ip route ospf | include E2 O E2 172.20.16.0 [110/20] via 10.0.12.2, Ethernet0/1 R1#show ip ospf database external | include Link State ID Link State ID: 172.20.16.0 (External Network Number ) <-- one LSA, not four ``` And the same anti-loop discard route appears, on the ASBR this time - note the administrative distance of 254, which keeps it out of the way of any real route: ``` R3#show ip route 172.20.16.0 255.255.240.0 Routing entry for 172.20.16.0/20 Known via "ospf 1", distance 254, metric 20, type intra area Routing Descriptor Blocks: * directly connected, via Null0 ``` `summary-address` with `not-advertise` also works, and suppresses the externals entirely. ## Choosing **Shrink the database across areas** `area range` on the ABR. This is the one you should be using in every multi-area design. **Block specific prefixes at an area boundary** `area filter-list` with a prefix-list. Arbitrary rules, no summarisation. **Stop ONE router using a route** `distribute-list in`. Local only. Understand the loop risk. **Collapse redistributed routes** `summary-address` on the ASBR. The only tool for Type-5 and Type-7. **Block externals from an entire area** Make it a stub or NSSA area. That is what area types are *for*, and it is far cleaner than filtering. ## The warning that ties it all together Every one of these mechanisms can, without warning, break an external route that has a non-zero forwarding address. If the subnet the forwarding address lives on is summarised, suppressed or filtered away, the external route is silently discarded on every router that can no longer resolve it. We reproduced this in the lab with both prefix suppression and an `area range not-advertise`, right down to a failing ping. Before you filter anything in a core: ``` show ip ospf database external | include Forward Address ``` Full details in [the OSPF forwarding address](https://www.pinglabz.com/ospf-forwarding-address/). ## Key takeaways - **area range** and **area filter-list** run on an ABR and genuinely remove the Type-3 LSA. Everyone in the target area loses the route. - **distribute-list in** runs on any router and only touches that router's RIB. The LSA stays in the database and every other router is unaffected. It saves no memory, and it can create loops. - **summary-address** runs on an ASBR and is the only tool for Type-5 and Type-7 externals. - Summarising with `area range` or `summary-address` installs a Null0 discard route. That is correct anti-loop behaviour, and it means a sloppy range black-holes real address space at the summarising router. - `area X filter-list ... out` means "leaving area X", not "outbound from this router". The direction is relative to the area. - Stub and NSSA area types are a cleaner way to keep externals out of an area than any filter. - Check for non-zero forwarding addresses before you filter anything in a core, or you will silently kill external routes. Next: [OSPF path preference](https://www.pinglabz.com/ospf-path-preference-rules/) \- O vs O IA vs E1 vs N1 vs E2 vs N2, with all six proven in sequence in the lab. The full cluster index lives on the [OSPF pillar guide](https://www.pinglabz.com/ospf/). ### The OSPF Forwarding Address: Why Your Type 5 Route Goes the Wrong Way URL: https://www.pinglabz.com/ospf-forwarding-address/ Last updated: 2026-07-12T07:18:24.000Z Here is a routing puzzle. An OSPF router has a route to an external network. The route is in the routing table. The next hop is a device that is *not the router that advertised it* and is *not even running OSPF*. Nothing in the configuration says to do this. How? That is the OSPF forwarding address doing its job. It is a genuinely clever optimisation that eliminates a suboptimal hop, and it is also the cause of one of the most baffling silent failures in OSPF - a route that vanishes from the RIB while its LSA sits happily in the database, with no log message and no obvious cause. This article covers the five conditions that set a non-zero forwarding address, the optimisation it produces, and the failure mode nobody documents - all with real output from a CML lab. For the fundamentals, start at the [complete OSPF guide](https://www.pinglabz.com/ospf/). ## The problem it solves Picture a shared Ethernet segment. On it sit three devices: two OSPF routers (call them R2 and R3), and a firewall (R6) that does not speak OSPF at all. Behind the firewall is 172.20.20.0/24. R3 has a static route to that network pointing at the firewall, and redistributes it into OSPF. R3 is now the ASBR. Without a forwarding address, every router in the OSPF domain would compute its path to 172.20.20.0/24 as "go to the ASBR (R3), and R3 will handle it". So R2, which is *sitting on the same Ethernet segment as the firewall*, would send its packets across the segment to R3, and R3 would immediately send them straight back across the same segment to the firewall. The packet crosses the wire twice for no reason. The forwarding address fixes exactly this. Instead of advertising "send it to me", the ASBR advertises "send it to *this address*" - the firewall's address on the shared segment. R2 sees that it can reach that address directly, and forwards straight to the firewall. ## The five conditions OSPF sets a non-zero forwarding address in a Type-5 LSA only when **all five** of these are true of the ASBR's interface toward the external route's next hop: 1 The next-hop interface is **running OSPF** (it is covered by a `network` statement or has `ip ospf` configured). 2 That interface is **not passive**. 3 Its OSPF network type is **broadcast or non-broadcast** \- not point-to-point, not point-to-multipoint. 4 The next-hop address falls **inside the subnet** advertised by that interface. 5 The route being redistributed actually **has a next hop** on that interface (it is not a connected route or a Null0 static). Fail any one of them and the forwarding address is 0.0.0.0, which means "send it to me, the ASBR". Read the list again and notice how much sense it makes. The whole point is that *other* OSPF routers should be able to reach the forwarding address directly. That is only plausible if the ASBR's interface is (a) in OSPF, so other routers know the subnet exists, and (b) a multiaccess segment, so other routers can plausibly be attached to the same wire. On a point-to-point link, nobody else is on that wire, so a forwarding address would be pointless - and OSPF correctly declines to set one. ## The lab R3 is the ASBR in area 1, on a shared broadcast segment with R2 (also OSPF) and R6 (a non-OSPF device with 172.20.20.0/24 behind it). ``` R3: interface Ethernet0/1 description SHARED-SEGMENT-AREA1 ip address 10.0.234.3 255.255.255.0 ! ip route 172.20.20.0 255.255.255.0 10.0.234.6 ! router ospf 1 network 10.0.234.0 0.0.0.255 area 1 redistribute static subnets ``` All five conditions are met. The LSA: ``` R2#show ip ospf database external 172.20.20.0 LS Type: AS External Link Link State ID: 172.20.20.0 (External Network Number ) Advertising Router: 3.3.3.3 Network Mask: /24 Metric Type: 2 (Larger than any link state path) Metric: 20 Forward Address: 10.0.234.6 External Route Tag: 0 ``` `Forward Address: 10.0.234.6`. That is R6, the firewall. R3 is advertising "do not send this to me, send it to R6". And R2's routing table obeys: ``` R2#show ip route 172.20.20.0 Routing entry for 172.20.20.0/24 Known via "ospf 1", distance 110, metric 20, type extern 2, forward metric 10 Last update from 10.0.234.6 on Ethernet0/2, 00:01:05 ago Routing Descriptor Blocks: * 10.0.234.6, from 3.3.3.3, 00:01:05 ago, via Ethernet0/2 ``` Read that carefully. The route was learned `from 3.3.3.3` (the ASBR), but the next hop is `10.0.234.6` (the firewall). The extra hop through R3 has been eliminated. That is the optimisation. Note also `forward metric 10` \- a separate field. The route's OSPF metric (20) is the external cost. The *forward metric* is the internal cost to reach the forwarding address. It matters, and we will come back to it. ## Breaking a condition Make R3's segment interface passive, and condition 2 fails: ``` R3(config-router)#passive-interface Ethernet0/1 R2#show ip ospf database external 172.20.20.0 | include Forward|Advertising|Metric Advertising Router: 3.3.3.3 Metric Type: 2 (Larger than any link state path) Metric: 20 Forward Address: 0.0.0.0 R2#show ip route 172.20.20.0 Routing entry for 172.20.20.0/24 Last update from 10.0.23.2 on Ethernet0/1, 00:00:10 ago Routing Descriptor Blocks: * 10.0.23.2, from 3.3.3.3, 00:00:10 ago, via Ethernet0/1 ``` The forwarding address reverted to 0.0.0.0 and R2's next hop is now 10.0.23.2 - the ASBR, over a completely different link. The packet will now take the long way round and hairpin. Forwarding still *works*; it is just worse. ## The failure mode nobody documents Now the part that will cost you a Saturday if you have not seen it before. RFC 2328 §16.4 says that when the forwarding address is non-zero, the router must look it up in its routing table, and **the matching route must be an intra-area or inter-area OSPF path**. If the forwarding address cannot be resolved that way, **the external route is discarded entirely**. So: anything that removes the forwarding address's subnet from the OSPF routing table silently kills the external route on every router that loses it. In the lab, we enabled [prefix suppression](https://www.pinglabz.com/ospf-prefix-suppression/) on the ABRs - a perfectly ordinary optimisation, applied to shrink the core table, with nothing to do with external routes. The shared segment 10.0.234.0/24 is a transit prefix, so it was suppressed. And: ``` R1#show ip route 172.20.20.0 % Network not in table R1#show ip ospf database external 172.20.20.0 | include Forward|Advertising Advertising Router: 3.3.3.3 Forward Address: 10.0.234.6 <-- the LSA is right there R1#show ip route 10.0.234.6 % Subnet not in table <-- but the forwarding address is unreachable ``` **The LSA is in the database. The route is not in the RIB. There is no log message.** We reproduced the same failure a second way, with an `area range not-advertise` on the ABR, and confirmed it right down to the data plane: ``` R2(config-router)#area 1 range 10.0.234.0 255.255.255.0 not-advertise R1#show ip route 10.0.234.6 % Subnet not in table R1#show ip route 172.20.20.0 % Network not in table R1#ping 172.20.20.1 source Loopback0 ..... Success rate is 0 percent (0/5) ``` An engineer summarising the core, or suppressing transit prefixes, or filtering a subnet they were sure nobody needed - and an external network goes dark on the other side of the building. ### The check to run before you filter anything ``` show ip ospf database external | include Forward Address ``` If every line reads `0.0.0.0`, filter freely. If any line names a real address, find out which subnet it lives on, and protect that subnet from suppression, summarisation and filtering. Put a comment in the config. ## Type-7 LSAs always carry a forwarding address An NSSA external LSA (Type 7) is different: it **always** has a non-zero forwarding address, and it is normally the ASBR's own router ID or loopback: ``` R4#show ip ospf database nssa-external 172.16.99.0 | include Advertising|Metric Type|Forward Advertising Router: 5.5.5.5 Metric Type: 1 (Comparable directly to link state metric) Forward Address: 5.5.5.5 ``` This is not an optimisation, it is a requirement. When the NSSA ABR translates a Type-7 into a Type-5 for the backbone, the resulting Type-5 needs a forwarding address so that backbone routers can reach the originating ASBR inside the NSSA. Without one, the translated LSA would point at the ABR, which is not where the route actually lives. The practical consequence: **the NSSA ASBR's forwarding address must be advertised into the backbone.** If you filter the NSSA's loopback range at the ABR, you break every external route the NSSA originates. Same failure mode, different LSA type. ## Troubleshooting 1. **An external route is in the LSDB but not the RIB.** Check the forwarding address (`show ip ospf database external `). If it is non-zero, check whether you have an OSPF route to it (`show ip route `). This is the answer far more often than people expect. 2. **Traffic is taking a strange path to an external network.** Look for a non-zero forwarding address. The route may be going somewhere you never configured, and that may well be correct. 3. **The forwarding address is 0.0.0.0 and you expected it not to be.** Walk the five conditions. In practice the culprit is almost always `passive-interface` or a point-to-point network type. 4. **The forwarding address is non-zero and you wish it were not.** Make the ASBR's next-hop interface passive, or change its network type to point-to-point. Both force it to 0.0.0.0. 5. **Two ASBRs advertise the same external.** The tiebreak uses the *forward metric* (the cost to reach the forwarding address), not just the external metric. Two E2 routes with equal external metric are broken by forward metric - which is why `show ip route` prints it. ## Key takeaways - The forwarding address lets an ASBR say "send it to that device over there", eliminating a pointless hairpin through the ASBR on a shared segment. - It is set only when all five conditions hold. In practice: an OSPF-enabled, non-passive, broadcast-type interface whose subnet contains the external route's next hop. - Passive-interface or a point-to-point network type will force it to 0.0.0.0\. That is the usual reason it "does not work". - **A non-zero forwarding address must resolve to an intra-area or inter-area OSPF route, or the external route is silently discarded.** Prefix suppression, area ranges and filtering can all cause this, with no log message. - Run `show ip ospf database external | include Forward Address` before you suppress, summarise or filter anything in the core. - Type-7 (NSSA) LSAs always carry a non-zero forwarding address. It must remain reachable from the backbone or the translated Type-5 is useless. - The *forward metric* \- the cost to reach the forwarding address - is a real tiebreaker between equal-cost E2 routes. Next: [OSPF filtering compared](https://www.pinglabz.com/ospf-filtering-comparison/) \- area range, filter-list, distribute-list and summary-address, and which of them touch the LSDB versus only the RIB. The full cluster index lives on the [OSPF pillar guide](https://www.pinglabz.com/ospf/). ### OSPF Stub Router: max-metric router-lsa and Graceful Maintenance URL: https://www.pinglabz.com/ospf-max-metric-stub-router/ Last updated: 2026-07-12T07:18:23.000Z You need to reload a core router. It is carrying live traffic. The naive approach - just reload it - drops every packet in flight and triggers a reconvergence storm. The slightly-less-naive approach - shut its interfaces one at a time - is fiddly, error-prone, and still causes churn on every shutdown. The correct approach is to tell OSPF, politely, that this router is still here but is a terrible choice for transit. Traffic drains away on its own. Nothing is dropped. Then you reload. That mechanism is `max-metric router-lsa`, and it is one of the highest-value, least-known OSPF features there is. This article covers the stub router advertisement (RFC 6987), its options, and a lab demonstration of traffic draining away from a router without a single interface going down. For the fundamentals, start at the [complete OSPF guide](https://www.pinglabz.com/ospf/). ## The idea OSPF path selection is metric-driven. If a router advertises its transit links with a metric so large that no sane SPF calculation would ever route through it, traffic simply stops using it as a transit path - while the router itself remains fully reachable and fully adjacent. The magic number is **65535** (0xFFFF), the maximum value of the 16-bit metric field in a Router LSA. This is not infinity - OSPF has no infinity - but it is the largest cost expressible, so any alternate path in a sanely-designed network will win. The critical design detail: **only transit links are maxed out. Stub links, including loopbacks, keep their real metric.** The router stays reachable. Its services stay reachable. It just stops being used as a road. ## Configuration ``` router ospf 1 max-metric router-lsa ``` That is the manual, "I am about to do maintenance" form. It takes effect immediately and stays until you remove it. The variant you should actually deploy on every router in your network, permanently: ``` router ospf 1 max-metric router-lsa on-startup 300 ``` This makes the router advertise max-metric for 300 seconds after every OSPF process start, then automatically return to normal. That solves a genuine problem: a router that has just booted has OSPF adjacencies up long before BGP has converged, or before its line cards have finished programming their FIBs. Without on-startup max-metric, the network starts sending it transit traffic that it cannot yet forward. This is a real, common, and completely avoidable outage. Even better, tie it to BGP convergence rather than a fixed timer: ``` router ospf 1 max-metric router-lsa on-startup wait-for-bgp ``` Now OSPF stays in stub-router mode until BGP signals that it has converged (or a 600-second safety timer expires). On any router that is both an OSPF speaker and a BGP edge, this should be considered mandatory. ## The lab: watching traffic drain away R2 is an ABR in area 0 with two transit links. R1 sits on the other side and has two equal-cost paths to a destination - one through R2, one through R4. ### Before ``` R1#show ip ospf database router 2.2.2.2 Advertising Router: 2.2.2.2 (Link ID) Network/subnet number: 2.2.2.2 TOS 0 Metrics: 1 <-- the loopback (a stub link) (Link ID) Designated Router address: 10.0.24.2 TOS 0 Metrics: 10 <-- transit link (Link ID) Designated Router address: 10.0.12.2 TOS 0 Metrics: 10 <-- transit link ``` And R1 is load-sharing across both paths to the R2-R4 link subnet: ``` R1#show ip route ospf | include 10.0.24 O 10.0.24.0/30 [110/20] via 10.0.14.2, Ethernet0/2 [110/20] via 10.0.12.2, Ethernet0/1 ``` ### Apply max-metric on R2 ``` R2(config-router)#max-metric router-lsa R2#show ip ospf | include originating router-LSAs Originating router-LSAs with maximum metric ``` ### After ``` R1#show ip ospf database router 2.2.2.2 Advertising Router: 2.2.2.2 (Link ID) Network/subnet number: 2.2.2.2 TOS 0 Metrics: 1 <-- loopback UNCHANGED (Link ID) Designated Router address: 10.0.24.2 TOS 0 Metrics: 65535 <-- maxed (Link ID) Designated Router address: 10.0.12.2 TOS 0 Metrics: 65535 <-- maxed ``` And the routing table, which is the whole story in two commands: ``` R1#show ip route 10.0.24.0 255.255.255.252 Routing entry for 10.0.24.0/30 Known via "ospf 1", distance 110, metric 20, type intra area Routing Descriptor Blocks: * 10.0.14.2, from 4.4.4.4, via Ethernet0/2 <-- ONLY via R4 now R1#show ip route 2.2.2.2 Routing entry for 2.2.2.2/32 Known via "ospf 1", distance 110, metric 11, type intra area * 10.0.12.2, from 2.2.2.2, via Ethernet0/1 <-- R2 ITSELF still reachable at normal cost ``` Traffic that *transited* R2 has moved to R4\. Traffic *destined for* R2 still arrives. No interface went down. No adjacency dropped. No packet was lost. That is the entire design intent, and it is visible in one routing table. ## The options, and what they actually do By default, max-metric touches only the transit links in the Router LSA. Three options extend it: **include-stub** Also max out the metric on stub links, including loopbacks. The router becomes hard to reach, not just useless as transit. Use this only if you genuinely want traffic to stop arriving at the router's own addresses too. **summary-lsa \[metric\]** For an ABR: also max out the metric in the Type-3 summary LSAs it generates. Without this, inter-area traffic keeps choosing this ABR because its summaries still look cheap. Default max is 16711680 (0xFF0000). **external-lsa \[metric\]** For an ASBR: also max out the metric in the Type-5 external LSAs it generates. Same reasoning - without it, external traffic keeps coming to this ASBR. **If your router is an ABR or an ASBR, the bare command is not enough.** It stops intra-area transit but leaves inter-area and external traffic pointing straight at you. From the lab, with `summary-lsa` added: ``` R2(config-router)#max-metric router-lsa include-stub summary-lsa external-lsa R1#show ip route 3.3.3.3 | include metric Known via "ospf 1", distance 110, metric 16711690, type inter area R1#show ip route 2.2.2.2 | include metric Known via "ospf 1", distance 110, metric 65545, type intra area ``` The inter-area route through R2 now costs 16.7 million. The loopback route (with `include-stub`) costs 65545\. Nothing in a sane topology is going to pick those. ### An IOS trap that cost us fifteen minutes `no max-metric router-lsa include-stub summary-lsa external-lsa` does **not** remove max-metric. It removes only the *options*, leaving plain `max-metric router-lsa` active and the router still advertising 65535 on its transit links. We chased phantom metric values for a while before running: ``` R2#show running-config | include max-metric max-metric router-lsa ``` The fix is `no max-metric router-lsa` with no arguments. This "the negation only removes the options" behaviour is general across IOS (it bit us again the same session with `no redistribute static subnets metric-type 2`). Always verify with `show running-config | section`, never assume the negation did what you meant. ## Where max-metric does not help **If the router is the only path, traffic still comes.** Max-metric makes a router *unattractive*, not *unusable*. If R2 is the sole ABR into an area, maxing its metric changes nothing, because there is no alternative. This is not a limitation - a routing protocol cannot conjure a path that does not exist - but it is a constraint you must design around. If you need to reload a single-homed ABR gracefully, the answer is redundancy, not max-metric. **It does not drain BGP.** Max-metric is an OSPF mechanism. If the router is also a BGP speaker, its BGP next-hops remain valid and BGP will keep sending it traffic. For a graceful BGP drain you need `bgp graceful-shutdown` (RFC 8326) or a policy that prepends and lowers local-pref on the way out. Use both together. ## Related: stub router vs graceful restart These get confused and they solve different problems: max-metric (stub router) **Says:** "I am here but do not route through me." **For:** planned maintenance, and post-boot convergence protection. **Traffic:** actively moved away. Graceful restart / NSF **Says:** "My control plane is restarting, keep forwarding through me, my FIB is intact." **For:** a supervisor switchover on a dual-RP chassis. **Traffic:** deliberately kept in place. They are opposites, and both are correct in their own context. Graceful restart says "trust me, keep sending". Max-metric says "stop sending". Use graceful restart for an in-service software upgrade on redundant hardware; use max-metric when the box is genuinely going away. ## Key takeaways - `max-metric router-lsa` advertises transit links at 65535 so traffic routes around the router, while leaving adjacencies up and the router's own addresses reachable at normal cost. - Deploy `max-metric router-lsa on-startup wait-for-bgp` permanently on every OSPF+BGP router. A freshly-booted router with converged OSPF and unconverged BGP is a black hole, and this one line prevents it. - On an ABR or ASBR, add `summary-lsa` and `external-lsa`. The bare command only stops intra-area transit. - `include-stub` also maxes the loopbacks. Usually you do *not* want that - keeping the router reachable is the point. - `no max-metric router-lsa ` removes only the options. Use the bare negation and verify with `show running-config`. - Max-metric makes a router unattractive, not unusable. If it is the only path, traffic still comes. - It is an OSPF mechanism only. Drain BGP separately with graceful-shutdown. Next: [the OSPF forwarding address](https://www.pinglabz.com/ospf-forwarding-address/), an optimisation that quietly becomes an outage. The full cluster index lives on the [OSPF pillar guide](https://www.pinglabz.com/ospf/). ### OSPF Prefix Suppression: Shrinking the Routing Table URL: https://www.pinglabz.com/ospf-prefix-suppression/ Last updated: 2026-07-12T07:18:23.000Z Look at your routing table and count how many of those prefixes anyone actually sends traffic to. In a typical service provider or large enterprise core, a large fraction of the OSPF routing table is transit link subnets - the /30s and /31s between routers. Nothing lives on them. No host has an address there. They exist only to carry packets between two routers, and yet every router in the area carries a route for every one of them. Prefix suppression deletes them. It is a single command, it can cut a core routing table substantially, and it has one consequence that will bite you in a troubleshooting session if you do not know about it. This article shows it working on Cisco IOS XE with real lab output, including the failure mode nobody warns you about. For the fundamentals, start at the [complete OSPF guide](https://www.pinglabz.com/ospf/). ## Why transit prefixes are dead weight A router-to-router link needs IP addresses so the two routers can form an adjacency and forward packets to each other. It does not need to be *routable from the rest of the network*. Nobody outside those two routers ever needs to send a packet *to* 10.0.24.1. But OSPF advertises the prefix anyway, because that is what the protocol was designed to do in 1998\. Every transit link becomes a stub network in a Router LSA (or a prefix on a Network LSA), floods to every router in the area, and consumes a routing table entry, a FIB entry, and a slot in every SPF calculation, on every router, forever. In a 200-router core with a full mesh of point-to-point links, that is thousands of routes that serve no forwarding purpose whatsoever. ## What prefix suppression does Defined in RFC 6860 for OSPFv2, prefix suppression tells the router to stop advertising the IP prefixes associated with its transit links. The *topology* information stays - the adjacency, the link, the cost, everything SPF needs to compute a path. Only the *reachability* information for the link subnet is removed. Loopbacks and other stub networks are never suppressed. Those are the things that carry actual services and are the things you actually want to reach. ### Configuration Process-wide, which is what you almost always want: ``` router ospf 1 prefix-suppression ``` Or per interface, if you need to be selective: ``` interface Ethernet0/2 ip ospf prefix-suppression ``` **It must be configured on both ends of a link** to remove the prefix entirely. If only one router suppresses, the other still advertises the subnet and the prefix survives in the area. ## The lab Five OSPF routers, multi-area, with R1 in area 0 watching. Before: ``` R1#show ip route ospf O IA 10.0.23.0/30 [110/20] via 10.0.12.2, Ethernet0/1 O 10.0.24.0/30 [110/20] via 10.0.14.2, Ethernet0/2 [110/20] via 10.0.12.2, Ethernet0/1 O IA 10.0.45.0/30 [110/20] via 10.0.14.2, Ethernet0/2 O IA 10.0.234.0/24 [110/20] via 10.0.12.2, Ethernet0/1 O E2 172.20.20.0 [110/20] via 10.0.12.2, Ethernet0/1 ``` Now enable `prefix-suppression` on R2 and R4 (both ends of the 10.0.24.0/30 link): ``` R1#show ip route ospf O 2.2.2.2 [110/11] via 10.0.12.2, Ethernet0/1 O IA 3.3.3.3 [110/21] via 10.0.12.2, Ethernet0/1 O 4.4.4.4 [110/11] via 10.0.14.2, Ethernet0/2 O IA 5.5.5.5 [110/21] via 10.0.14.2, Ethernet0/2 O IA 10.0.23.0/30 [110/20] via 10.0.12.2, Ethernet0/1 O IA 10.0.45.0/30 [110/20] via 10.0.14.2, Ethernet0/2 ``` 10.0.24.0/30 is gone. Every loopback is still there. The adjacencies are untouched: ``` R1#show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 4.4.4.4 1 FULL/DR 00:00:32 10.0.14.2 Ethernet0/2 2.2.2.2 1 FULL/DR 00:00:33 10.0.12.2 Ethernet0/1 ``` And forwarding through the suppressed link works perfectly - because SPF still knows the link exists, it just has no route *to* it: ``` R1#ping 4.4.4.4 source Loopback0 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 1/2/3 ms ``` ## The consequence: you cannot ping your own transit links ``` R1#ping 10.0.24.2 source Loopback0 ..... Success rate is 0 percent (0/5) ``` That is not a bug. It is the entire mechanism working as designed. There is no route to 10.0.24.0/30, so packets addressed to 10.0.24.2 have nowhere to go. The operational impact is real and you must plan for it: - **Traceroute through the core produces asterisks.** Each hop replies to a TTL-exceeded from its ingress interface address - which is now unroutable, so the reply never gets back to you. - **Monitoring that pings interface addresses breaks.** Point your NMS at loopbacks instead. You should have been doing that anyway. - **An engineer debugging at 3 a.m. will conclude the network is broken.** Document it. Prominently. The standard mitigation is to source everything from loopbacks and monitor loopbacks. On IOS you can also make ICMP replies come from the loopback with `ip ospf prefix-suppression` left off on a small number of deliberately-reachable links, but the cleaner answer is to accept the trade and change your operational habits. ## The trap: prefix suppression can silently kill an external route This is the finding from the lab that is genuinely worth knowing, and it does not appear in any Cisco document. In our topology, an ASBR (R3) redistributes a static route to 172.20.20.0/24 whose next-hop is a non-OSPF device sitting on a shared broadcast segment. Because all the conditions are met, the resulting Type-5 LSA carries a **non-zero forwarding address** \- the address of that non-OSPF device, 10.0.234.6\. (If that is unfamiliar, read [the OSPF forwarding address](https://www.pinglabz.com/ospf-forwarding-address/) first.) The forwarding address lives on the shared segment, 10.0.234.0/24\. Prefix suppression removes that subnet from the routing table. And now: ``` R1#show ip route 172.20.20.0 % Network not in table R1#show ip ospf database external 172.20.20.0 | include Forward|Advertising Advertising Router: 3.3.3.3 Forward Address: 10.0.234.6 R1#show ip route 10.0.234.6 % Subnet not in table ``` **The LSA is right there in the database. The route is gone from the RIB.** RFC 2328 §16.4 requires that a non-zero forwarding address resolve to an intra-area or inter-area OSPF route. Suppress the prefix, and the forwarding address becomes unresolvable, and the router discards the external route entirely. Silently. With no log message. We verified the causality by removing prefix suppression and watching both the transit prefix and the external route come straight back, and then reproduced the same failure a second way (with an `area range not-advertise`) to confirm it is the FA lookup and not something specific to prefix suppression. **The rule to take away:** if any external route in your network has a non-zero forwarding address, the subnet that forwarding address lives on must remain a routable OSPF prefix. Do not suppress it, do not summarise it away, do not filter it. Check before you deploy: ``` show ip ospf database external | include Forward Address ``` If every line says `0.0.0.0`, you are safe. If any of them names a real address, find out which subnet it is on and protect it. ## When to use it, when not to Good fit A large SP or enterprise core with hundreds of transit links, where the RIB and FIB are under pressure and every service lives on a loopback. Good fit An MPLS core, where forwarding is label-switched between loopbacks anyway and transit link reachability is genuinely irrelevant. Poor fit A small enterprise with 20 routers. The saving is a handful of routes and the troubleshooting cost is permanent. Poor fit Anywhere with externals carrying a non-zero forwarding address on a link you are about to suppress. See above. ### The alternatives Prefix suppression is not the only way to shrink a core table: - **Unnumbered interfaces** (`ip unnumbered Loopback0`) mean there is no transit subnet to advertise in the first place. Cleaner in principle, but changes your addressing model and complicates some tooling. - **/31 addressing** halves the address consumption but does not reduce the route count. - **Passive interfaces** do the opposite of what you want here: they advertise the prefix but form no adjacency. - **Summarisation at ABRs** reduces inter-area routes, which is complementary. Prefix suppression works inside an area; summarisation works between them. OSPFv3, incidentally, has this built in: link-local addressing means transit links have nothing to advertise by default, and you have to go out of your way to leak them. IPv6 got this one right. ## Key takeaways - `prefix-suppression` under the OSPF process removes transit-link prefixes from the LSDB and the RIB, while leaving the topology, adjacencies and forwarding completely intact. - Loopbacks and stub networks are never suppressed. Only transit links. - It must be enabled on **both ends** of a link for the prefix to actually disappear. - The cost is that you can no longer ping or traceroute transit interfaces. Source and monitor from loopbacks, and document it before someone else discovers it during an outage. - **Any external route with a non-zero forwarding address will be silently discarded if the forwarding address subnet is suppressed.** Run `show ip ospf database external | include Forward Address` before you deploy. - In a core with hundreds of transit links, the RIB and FIB saving is significant. In a small network, it is not worth the operational cost. Next: [the OSPF stub router and max-metric router-lsa](https://www.pinglabz.com/ospf-max-metric-stub-router/), the correct way to take a router out of service without dropping a packet. The full cluster index lives on the [OSPF pillar guide](https://www.pinglabz.com/ospf/). ### OSPF SPF and LSA Throttling Timers: Tuning Convergence Safely URL: https://www.pinglabz.com/ospf-spf-lsa-throttling-timers/ Last updated: 2026-07-12T07:18:22.000Z OSPF converges fast. That is its whole reason for existing. But "fast" is a tuning decision, and the knobs are not where most people look. They are not the hello and dead timers - those control failure *detection*. The knobs that control what happens *after* detection are the SPF and LSA throttle timers, and getting them wrong in either direction hurts. Too slow and you leave convergence time on the table. Too fast and a single flapping link can drive every router in the area into a CPU spiral, recalculating the topology dozens of times per second. This article covers the exponential backoff algorithm behind both throttles, the modern IOS XE defaults (which are not what the older documentation says), and how to verify what your router is actually doing. For the fundamentals, start at the [complete OSPF guide](https://www.pinglabz.com/ospf/). ## Two separate throttles, one algorithm OSPF has two independent throttle mechanisms and they are constantly confused with each other: SPF throttle `timers throttle spf` Governs how soon the router runs the Dijkstra calculation after receiving a topology change, and how quickly it will run it again. Protects **your CPU**. LSA throttle `timers throttle lsa` Governs how soon *this router* re-originates one of its *own* LSAs after a change. Protects **everyone else's CPU** from your flapping interface. LSA arrival `timers lsa arrival` The minimum interval at which the router will *accept* a new instance of the same LSA from a neighbour. Your defence against a neighbour who has tuned their LSA throttle too aggressively. All three take the same three arguments, in the same order: ``` timers throttle spf timers throttle lsa ``` The algorithm is exponential backoff, and it works like this: 1. A topology change arrives. The router waits **start** milliseconds, then runs SPF. This is the fast path, and for an isolated event it is the only delay you ever see. 2. If another change arrives while the router is inside the current wait interval, the next SPF is scheduled after **hold** milliseconds. 3. Each subsequent back-to-back change **doubles** the hold interval: hold, 2×hold, 4×hold, and so on. 4. The interval is capped at **max-wait**. It stays there for as long as the churn continues. 5. Once the network is quiet for max-wait milliseconds, the whole thing resets and the next event gets the fast **start** delay again. This is exactly the behaviour you want. A single link failure - the common case - converges as fast as the hardware allows. A pathological flap gets progressively throttled until the router is spending a sane amount of CPU on it. ## The defaults are not what you think Half the OSPF tuning advice on the internet is written against IOS defaults from a decade ago: an initial SPF delay of 5 seconds and a hold of 10 seconds. On IOS XE 17.18, from the lab: ``` R1#show ip ospf Initial SPF schedule delay 50 msecs Minimum hold time between two consecutive SPFs 200 msecs Maximum wait time between two consecutive SPFs 5000 msecs Initial LSA throttle delay 50 msecs Minimum hold time for LSA throttle 200 msecs Maximum wait time for LSA throttle 5000 msecs Minimum LSA arrival 100 msecs LSA group pacing timer 240 secs ``` **50 ms initial delay, 200 ms hold, 5000 ms max.** Modern IOS XE ships with what used to be considered aggressive tuning. Before you reach for the timers, check what you already have. In a great many networks the honest answer is that OSPF convergence is already sub-second and the bottleneck is elsewhere - usually failure detection, which is a job for BFD, not for SPF timers. ## Tuning, and verifying it took ``` router ospf 1 timers throttle spf 10 100 5000 timers throttle lsa 10 100 5000 timers lsa arrival 80 timers pacing flood 33 ``` ``` R1#show ip ospf Initial SPF schedule delay 10 msecs Minimum hold time between two consecutive SPFs 100 msecs Maximum wait time between two consecutive SPFs 5000 msecs Initial LSA throttle delay 10 msecs Minimum hold time for LSA throttle 100 msecs Maximum wait time for LSA throttle 5000 msecs Minimum LSA arrival 80 msecs Interface flood pacing timer 33 msecs ``` ### The constraint everyone gets wrong **Your LSA arrival timer must be lower than your neighbours' LSA throttle hold timer.** If a neighbour re-originates an LSA every 100 ms and you refuse to accept a new instance more often than every 200 ms, you will silently discard half of their updates and your LSDB will lag behind reality. The safe rule: set `timers lsa arrival` to roughly 80% of the smallest `timers throttle lsa` hold value anywhere in the area. In the config above, 80 ms arrival against a 100 ms hold. Get this backwards and you have built an intermittent, load-dependent, extremely difficult routing bug. ## Watching the backoff happen `show ip ospf statistics` is the command nobody runs and everybody should. It logs every SPF run, how long each phase took, and - critically - *why* it ran. From the lab, after flapping a link between R1 and R2 twice in quick succession: ``` R1#show ip ospf statistics Area 0: SPF algorithm executed 13 times SPF calculation time Delta T Intra D-Intra Summ D-Summ Ext D-Ext Total Reason 00:04:29 0 0 0 0 0 0 0 X 00:04:26 0 1 0 0 0 0 1 R 00:03:07 0 0 0 0 0 0 0 X 00:02:38 0 0 0 0 0 0 0 X 00:00:14 0 0 0 1 0 0 1 R, SN, X 00:00:14 0 0 1 0 0 0 1 R 00:00:13 1 0 0 0 0 0 1 R, N 00:00:09 0 0 0 1 0 0 1 R 00:00:08 0 0 1 0 0 0 1 R, N, SN, X 00:00:07 0 0 0 0 0 0 0 R ``` Read the `Delta T` column bottom-up and you can watch the throttle working: a burst of SPF runs clustered at 14, 13, 9, 8 and 7 seconds ago, each triggered by an LSA change during the flap. The `Reason` column tells you what changed: **R**A Router LSA (Type 1) changed. Somebody's link went up or down. **N**A Network LSA (Type 2) changed. A DR election or a change on a multiaccess segment. **SN / SA**A Summary LSA (Type 3 / Type 4) changed. An inter-area route moved. **X**An External LSA (Type 5) changed. A redistributed route moved. Notice that only `R` and `N` reasons trigger the *full* intra-area SPF (the expensive Dijkstra). Summary and external changes only require a partial recalculation, which is why the `Intra` column is mostly zero. That is the point of the OSPF LSA hierarchy, visible in a single command. ## Two more timers worth knowing **LSA group pacing** (`timers pacing lsa-group`, default 240 s) controls how OSPF batches its LSA refresh, checksum and aging work. In a very large database, the default groups too much work together and you get periodic CPU spikes every four minutes. Lowering it spreads the load. In an area with fewer than a few thousand LSAs, do not touch it. **Flood pacing** (`timers pacing flood`, default 33 ms) controls the interval between LSA flood packets on an interface. Lower means faster flooding of a large update set, at the cost of a burst on the wire that a slow neighbour may not keep up with. The default is fine. ## How to actually make OSPF converge fast Tuning SPF timers is the last step, not the first. In priority order: 1. **Detect failures fast.** Sub-second detection comes from BFD, not from hello timers. `bfd interval 300 min_rx 300 multiplier 3` plus `ip ospf bfd` gets you failure detection in under a second without the hello overhead of a 1-second dead timer. 2. **Make the SPF cheap.** A smaller LSDB runs a faster Dijkstra. That means proper area design, summarisation at ABRs, and [prefix suppression](https://www.pinglabz.com/ospf-prefix-suppression/) on transit links. 3. **Then tune the timers**, and only if measurement shows they are the bottleneck. 4. **Then consider LFA / remote LFA** for pre-computed backup paths, which converges in the forwarding plane without waiting for SPF at all. An SPF run over a 500-LSA area on modern hardware takes single-digit milliseconds. If your convergence is measured in seconds, the SPF timers are almost certainly not why. ## Troubleshooting checklist 1. **Convergence slower than expected?** Check `show ip ospf statistics` for the actual SPF frequency and duration, and `show ip ospf | include SPF schedule|hold time` for the timers in effect. Do not assume the defaults. 2. **Router CPU spiking during a flap?** Raise the hold and max-wait values so the backoff engages sooner. Then go and fix the flapping link. 3. **LSDB out of sync intermittently?** Check that your `timers lsa arrival` is lower than every neighbour's LSA throttle hold. This is the classic cause and it is very hard to spot. 4. **Periodic CPU spikes with no topology change?** LSA group pacing on a very large database. Lower `timers pacing lsa-group`. 5. **Tuning applied but no change?** The throttle timers are per-process, not per-area or per-interface. Confirm you are under the right `router ospf` process. ## Key takeaways - SPF throttle protects your CPU. LSA throttle protects everyone else's. They are different mechanisms with an identical-looking syntax. - Both use exponential backoff: `start` for the first event, then `hold` doubling each time up to `max-wait`, resetting after a quiet period. - IOS XE 17.x already defaults to 50/200/5000 ms, not the 5000/10000/10000 you will read in older guides. Check before you tune. - `timers lsa arrival` must be lower than the smallest LSA throttle hold in the area, or you will silently drop your neighbours' updates. - `show ip ospf statistics` shows every SPF run, its cost, and its reason code. It is the only way to see the throttle working. - For real convergence gains, fix detection (BFD) and database size (areas, summarisation, prefix suppression) before touching the timers. Next in the expert OSPF series: [prefix suppression](https://www.pinglabz.com/ospf-prefix-suppression/), which shrinks the routing table by removing the transit links nobody needs to reach. The full cluster index lives on the [OSPF pillar guide](https://www.pinglabz.com/ospf/). ### Expert BGP Troubleshooting: Five Broken Scenarios, Ticket Style URL: https://www.pinglabz.com/expert-bgp-troubleshooting-scenarios/ Last updated: 2026-07-12T06:32:04.000Z The CCIE troubleshooting section does not ask you to explain BGP. It hands you a broken network and a clock. The difference between passing and failing is not knowing more BGP - it is having a repeatable method for turning a symptom into a cause in under ten minutes. This article is five broken scenarios, each built and broken for real in a CML lab, each with the actual diagnostic output and the actual fix. Every one of them is a mistake that has taken down a production network. Work through them the way you would work a ticket. For the underlying theory, see the [complete BGP guide](https://www.pinglabz.com/bgp/). ## The method, before the tickets Every BGP problem is one of exactly four things, and they must be checked in this order: **1\. Is the session up?** `show ip bgp summary`. If the state column is not a number, you have a session problem and nothing downstream matters. **2\. Is the prefix being sent?** `show ip bgp neighbors x advertised-routes` on the sender. If it is not leaving, stop looking at the receiver. **3\. Is the prefix being accepted?** `received-routes` vs `routes` on the receiver, plus the `Local Policy Denied Prefixes` counters. This is where policy bugs live. **4\. Is it being used?** In the BGP table but not the RIB means best-path or next-hop. Look for `>`, and look for `(inaccessible)`. The single most under-used command in BGP troubleshooting is this one: ``` show ip bgp neighbors | section Policy ``` It gives you a running count of every prefix the router discarded and *why*. Route-map. Prefix-list. AS-path loop. Maxprefix. It is a confession, and the router volunteers it. Use it in every ticket. One prerequisite: to see what a neighbour *sent* you before your policy chewed it, you need `soft-reconfiguration inbound` on that neighbour. Turn it on when you start troubleshooting (it costs memory, so turn it off after). ## Ticket 1: "The iBGP session keeps resetting, but it's up" **Symptom:** R1 and R2 are iBGP peers in AS 65001, using loopbacks. The session is Established. But the log is full of resets and the operator swears something is wrong. **Diagnosis.** The summary looks healthy: ``` R2#show ip bgp summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 1.1.1.1 4 65001 7 6 56 0 0 00:00:01 8 ``` Established, 8 prefixes. But look at the uptime: one second. And now the detail: ``` R2#show ip bgp neighbors 1.1.1.1 | include BGP state|Last reset|Local host|Foreign host BGP state = Established, up for 00:00:01 Last reset 00:00:11, due to Active open failed Local host: 2.2.2.2, Local port: 179 Foreign host: 1.1.1.1, Foreign port: 43797 ``` `Last reset ... due to Active open failed`. R2's *own* attempt to open the session failed. The session that is currently up was opened *by R1* \- you can see it in the ports: R2's local port is 179, meaning R2 is the **passive** side. R2 accepted an inbound connection; it never successfully made an outbound one. **Cause.** R2 is missing `neighbor 1.1.1.1 update-source Loopback0`. When R2 initiates, it sources from the physical interface, 10.0.12.2\. R1 expects a connection from 2.2.2.2 and rejects it. But R1 *does* have update-source configured, so R1's outbound connection arrives correctly sourced from 1.1.1.1, and R2 happily accepts it. The session works, and the misconfiguration is masked. Remove update-source from *both* sides and the mask comes off: ``` R1#show ip bgp summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 2.2.2.2 4 65001 0 0 1 0 0 00:00:13 Idle R1#show ip bgp neighbors 2.2.2.2 | include BGP state|Last reset|No active TCP BGP state = Idle, down for 00:00:14 Last reset 00:00:14, due to Active open failed No active TCP connection ``` **Fix:** `neighbor x.x.x.x update-source Loopback0` on both peers. **Lesson:** "Active open failed" on an Established session means the session is only up because the *other* router dialled. That is a half-broken config waiting for the day the other router reboots first. ## Ticket 2: "The routes are in the BGP table but not the routing table" **Symptom:** R2 has received 100.100.100.0/24 over iBGP from R1\. It is in `show ip bgp`. It is not in `show ip route`. Traffic to it is dropped. **Diagnosis.** Two flags tell the whole story: ``` R2#show ip bgp 100.100.100.0/24 BGP routing table entry for 100.100.100.0/24, version 0 Paths: (1 available, no best path) 100 10.0.13.2 (inaccessible) from 1.1.1.1 (1.1.1.1) Origin IGP, metric 0, localpref 200, valid, internal R2#show ip bgp | include 100.100 * i 100.100.100.0/24 10.0.13.2 0 200 0 100 i ``` `no best path`. `(inaccessible)`. And in the summary line there is a `*` but no `>`. **Cause.** The next-hop is 10.0.13.2 - an address on the link between R1 and the ISP. R2 has no route to it. The IGP does not carry the eBGP link subnets (and should not). BGP requires a resolvable next-hop before a path can become best, so the route sits in the table doing nothing. R1 is missing `neighbor 2.2.2.2 next-hop-self`. Without it, R1 passes the external next-hop through to its iBGP peer unchanged. **Fix:** ``` router bgp 65001 address-family ipv4 neighbor 2.2.2.2 next-hop-self ``` **The trap.** This one is far nastier than it looks, because **eBGP multipath hides it completely**. In the lab, R1 had `maximum-paths 2` configured. With multipath active, removing `next-hop-self` changed nothing at all - R2 still received next-hop 1.1.1.1, because a router advertising a multipath route to an iBGP peer sets itself as the next-hop (there is no single external next-hop to pass along). Only after removing `maximum-paths 2` did the failure appear. So: on any edge router with eBGP multipath, a missing `next-hop-self` is invisible - right up until the day one circuit fails, the prefix drops to a single path, and your entire internal routing collapses. Configure `next-hop-self` explicitly on every iBGP session from an eBGP edge router, whether or not it appears to be needed. See the [multipath article](https://www.pinglabz.com/bgp-multipath-load-sharing/) for the full walkthrough. ## Ticket 3: "We only get one route from the provider" **Symptom:** R2 peers with ISP-B (AS 200). The provider insists they are sending three prefixes. R2 has one. **Diagnosis.** The prefix count in the summary already says something is being dropped: ``` R2#show ip bgp summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 10.0.25.2 4 200 89 85 112 0 0 00:36:30 1 ``` With `soft-reconfiguration inbound` enabled, compare what arrived against what survived: ``` R2#show ip bgp neighbors 10.0.25.2 received-routes Network Next Hop Metric LocPrf Weight Path * 5.5.5.5/32 10.0.25.2 0 0 200 i * 6.6.6.6/32 10.0.25.2 0 200 300 i * 6.6.60.0/24 10.0.25.2 0 200 300 i Total number of prefixes 3 R2#show ip bgp neighbors 10.0.25.2 routes Network Next Hop Metric LocPrf Weight Path *> 5.5.5.5/32 10.0.25.2 0 150 0 200 i Total number of prefixes 1 ``` Three in, one out. The provider is telling the truth. Now ask the router why: ``` R2#show ip bgp neighbors 10.0.25.2 | section Policy Local Policy Denied Prefixes: -------- ------- route-map: 0 2 ``` Two prefixes denied by a route-map, inbound. **Cause.** Someone wrote an inbound route-map to set local-preference on a specific prefix, and forgot that a route-map ends in an implicit deny: ``` route-map ISPB-IN permit 10 match ip address prefix-list ISPB-ALLOW set local-preference 150 ! ...and nothing else. Everything not matching ISPB-ALLOW is denied. ``` **Fix:** a bare permit clause at the end. ``` route-map ISPB-IN permit 20 ``` **Lesson:** this is, by a wide margin, the most common BGP outage in the world. Any time you attach a route-map to a neighbour to *set* something, you have also, silently, attached a filter. Every route-map used for attribute manipulation needs a terminating `permit`. Every single one. ## Ticket 4: "The neighbour is advertising it, we're not receiving it, and there's no filter" **Symptom:** R5 (ISP-B) demonstrably advertises 10.10.10.0/24 to R2\. R2 does not have it. There is no inbound filter of any kind on R2. **Diagnosis.** First confirm the sender is not lying: ``` R5#show ip bgp neighbors 10.0.25.1 advertised-routes Network Next Hop Metric LocPrf Weight Path *> 1.1.1.1/32 10.0.56.2 200 0 300 100 65001 i *> 10.10.10.0/24 10.0.56.2 200 0 300 100 65001 i ``` It is being sent. Now look at R2's policy counters: ``` R2#show ip bgp neighbors 10.0.25.2 | section Policy Local Policy Denied Prefixes: -------- ------- AS_PATH loop: n/a 30 Bestpath from this peer: 9 n/a Other Policies: 4 n/a Total: 13 30 ``` There is the answer, in one line. **AS\_PATH loop: 30.** **Cause.** Look at the AS-path R5 is sending: `300 100 65001`. R2 is *in* AS 65001\. That prefix originated in R2's own AS, travelled out through ISP-A, across AS 300, through ISP-B, and is now coming home. BGP's fundamental loop-prevention rule fires: if my own ASN appears in the AS-path of an inbound update, discard it silently. This is not a bug. This is BGP working exactly as designed and protecting you from a routing loop. But it is invisible unless you look at the right counter, because the route never enters the BGP table - not even the pre-policy soft-reconfiguration copy. The AS-path check happens at parse time, before storage. **The real question is why the prefix is looping at all.** Something upstream is transiting routes it should not. In the lab, AS 300 was leaking: it re-advertised routes learned from ISP-A to ISP-B, turning itself into unintended transit. The fix is not on R2\. The fix is an outbound filter at AS 300 restricting it to its own prefixes. **When you actually want to accept it:** in an MPLS L3VPN, a customer with the same AS number at multiple sites *legitimately* sees its own ASN in the path. That is what `allowas-in` and `as-override` exist for - covered in the transport articles under the [MPLS pillar](https://www.pinglabz.com/mpls/). ## Ticket 5: "IPv4 is fine, IPv6 is completely dead" **Symptom:** R1 and R3 have both an IPv4 and an IPv6 eBGP session. IPv4 is perfect. The IPv6 session will not stay up and R1 has lost all its IPv6 routes. **Diagnosis.** ``` R1#show bgp ipv6 unicast summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 2001:DB8:13::2 4 100 0 0 1 0 0 00:00:12 Idle R1#show bgp ipv6 unicast Network Next Hop Metric LocPrf Weight Path *> 2001:DB8:10::/64 :: 0 32768 i ``` Idle. Zero messages sent, zero received. And only R1's own locally-originated prefix remains - everything learned from the peer is gone. Meanwhile IPv4 is entirely unaffected: ``` R1#show ip bgp summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 10.0.13.2 4 100 27 36 108 0 0 00:14:08 3 ``` **Cause.** On R3, someone removed the activation: ``` router bgp 100 address-family ipv6 no neighbor 2001:DB8:13::1 activate ``` With `no bgp default ipv4-unicast` in effect (as it should be in any multiprotocol design), a neighbour that is not activated in *any* address family has nothing to negotiate. The OPEN carries no usable AFI/SAFI capability, and the session drops to Idle. It does not stay up-but-empty. It goes down. **Fix:** ``` router bgp 100 address-family ipv6 neighbor 2001:DB8:13::1 activate neighbor 2001:DB8:13::1 send-community both ``` **Lesson:** a BGP neighbour is defined globally but must be *activated* per address family. Every multiprotocol BGP problem starts with checking `activate`. Zero MsgSent and zero MsgRcvd on an Idle session is the signature - the routers are not even trying to talk. ## The commands, collected ``` ! Session layer show ip bgp summary show ip bgp neighbors | include BGP state|Last reset|No active TCP|Local host|Foreign host ! What is being sent show ip bgp neighbors advertised-routes ! What is being received, and what survived policy show ip bgp neighbors received-routes (needs soft-reconfiguration inbound) show ip bgp neighbors routes show ip bgp neighbors | section Policy <-- the confession ! Why is it not in the RIB show ip bgp (look for: no best path, (inaccessible), the > flag) show ip route show ip cef ! Multiprotocol show bgp ipv6 unicast summary show bgp ipv6 unicast show bgp ipv6 unicast labels (6PE) ``` ## Try it yourself Build the lab, then break it. Each of these produces a distinct, findable signature. Give yourself ten minutes per fault and do not look at the config until you have named the cause from show output alone: 1. Remove `update-source` from one side of an iBGP session, then from both. Note the difference. 2. Remove `next-hop-self` from an eBGP edge router with multipath on, then turn multipath off. Watch the failure appear from nowhere. 3. Attach an inbound route-map with a `match` and a `set` and no terminating permit. 4. Remove an upstream AS's outbound filter so it starts transiting your own prefixes back to you. 5. De-activate an IPv6 neighbour on one side only. 6. Bonus: set `neighbor x maximum-prefix 2` and watch the session tear itself down. Find the counter that tells you why. ## Key takeaways - Work the layers in order: session up, prefix sent, prefix accepted, prefix used. Do not skip. - `show ip bgp neighbors | section Policy` tells you exactly how many prefixes were discarded and by what. Run it first, always. - `received-routes` vs `routes`, with `soft-reconfiguration inbound`, is how you separate "they did not send it" from "we threw it away". - "Active open failed" on an Established session means only the far end can dial. Fix it before it fixes you. - `(inaccessible)` plus `no best path` is always a next-hop problem, and usually a missing `next-hop-self` \- which eBGP multipath will hide from you. - An implicit deny at the end of an attribute-setting route-map is the most common self-inflicted BGP outage there is. - AS-path loop rejections never enter the BGP table. The only evidence is the counter. - Multiprotocol BGP: `activate` per address family, or the session does not come up at all. That closes the expert BGP series. The full cluster index, from first principles to here, lives on the [BGP pillar guide](https://www.pinglabz.com/bgp/). ### IPv6 over BGP: MP-BGP for IPv6 and 6PE Explained URL: https://www.pinglabz.com/ipv6-bgp-6pe/ Last updated: 2026-07-12T06:32:04.000Z Here is a problem that a great many service providers had, and some still have: a fully working, fully tuned, IPv4-only MPLS core, and customers who want IPv6\. Upgrading every P router in the core to dual-stack is expensive, risky, and touches the one part of the network you least want to touch. 6PE is the answer. It carries IPv6 prefixes across an untouched IPv4 MPLS core by encapsulating them in MPLS labels, exchanged over MP-BGP. The P routers in the middle never learn a single IPv6 route and never even know IPv6 is involved. This article builds a working 6PE deployment on Cisco IOS XE in CML and shows the labels, the IPv4-mapped next-hops, and an end-to-end IPv6 traceroute crossing the v4 core. For the fundamentals, start at the [complete BGP guide](https://www.pinglabz.com/bgp/); for the MPLS side, see the [MPLS pillar](https://www.pinglabz.com/mpls/). ## MP-BGP: the foundation BGP-4 as originally specified could only carry IPv4\. Multiprotocol BGP (RFC 4760) added two attributes - `MP_REACH_NLRI` and `MP_UNREACH_NLRI` \- which carry an address family identifier plus the reachability information for that family. One BGP session, many address families. This is the same machinery that carries VPNv4 for MPLS L3VPN, VPNv6 for 6VPE, and L2VPN signalling. IPv6 unicast is just another AFI/SAFI riding the same session. The key insight for 6PE: **the BGP session itself can run over IPv4** while carrying IPv6 NLRI. The transport and the payload are independent. ### Plain MP-BGP for IPv6 (no MPLS) If your core is dual-stacked, you do not need 6PE at all. You just enable the IPv6 address family: ``` router bgp 65001 neighbor 2001:DB8:13::2 remote-as 100 ! address-family ipv6 network 2001:DB8:10::/64 neighbor 2001:DB8:13::2 activate exit-address-family ``` Note `no bgp default ipv4-unicast` is best practice here - without it, IOS tries to activate every neighbour in the IPv4 unicast AF, including your IPv6 peers, which produces confusing half-broken sessions. And note the thing that catches everyone: **a neighbour is defined globally but must be activated per address family**. Forget the `activate` and the session will not even come up. We prove that in the troubleshooting section below. ## 6PE: the architecture 6PE (RFC 4798) has four moving parts: The core (P routers) IPv4 only. IGP + LDP. Never sees an IPv6 route, never runs an IPv6 process. Completely untouched. The edge (6PE routers) Dual stack. Speaks IPv6 to customers, IPv4 + MPLS to the core. This is the only place IPv6 exists inside the provider. The iBGP session Runs over **IPv4** between the 6PE loopbacks. Carries the IPv6 unicast AF with `send-label`. The label stack Outer label = LDP label to the remote 6PE loopback (the core understands this). Inner label = the BGP-assigned IPv6 label (only the egress 6PE understands this). The magic trick is the next-hop. When a 6PE advertises an IPv6 prefix over an IPv4 iBGP session, it must set an IPv6 next-hop. It uses an **IPv4-mapped IPv6 address**: `::FFFF:4.4.4.4`. The receiving 6PE strips off the `::FFFF:` prefix, recovers 4.4.4.4, looks that up in its IPv4 LFIB, and finds the LDP label-switched path to it. Now it has a way to send IPv6 packets across an IPv4 core. ## The lab ISP-A is AS 100, with two 6PE routers (R3 and R4) and an IPv4-only, LDP-enabled core link between them. R1 (customer, AS 65001) hangs off R3 with IPv6\. R6 (customer, AS 300) hangs off R4 with IPv6\. Neither customer knows the core is IPv4. ### Core configuration (IPv4 + LDP only) ``` mpls label protocol ldp ! interface Ethernet0/2 description TO-R4-MPLS-CORE ip address 10.0.34.1 255.255.255.252 mpls ip ! router ospf 1 network 3.3.3.3 0.0.0.0 area 0 network 10.0.34.0 0.0.0.3 area 0 ``` Note what is *not* there: no `ipv6 address` on the core interface, no IPv6 in OSPF. The core link is IPv4, full stop. ### 6PE configuration on R3 ``` router bgp 100 no bgp default ipv4-unicast neighbor 4.4.4.4 remote-as 100 neighbor 4.4.4.4 update-source Loopback0 neighbor 2001:DB8:13::1 remote-as 65001 ! address-family ipv4 neighbor 4.4.4.4 activate neighbor 4.4.4.4 next-hop-self exit-address-family ! address-family ipv6 neighbor 4.4.4.4 activate neighbor 4.4.4.4 send-label neighbor 4.4.4.4 next-hop-self neighbor 2001:DB8:13::1 activate exit-address-family ``` Look at the IPv6 address family. The neighbour activated there is `4.4.4.4` \- an **IPv4 address**. That is the whole trick, in one line. The session is IPv4; the payload is IPv6. The three commands that make 6PE 6PE: - **`send-label`** \- this is the one that turns MP-BGP for IPv6 into 6PE. It tells BGP to allocate and advertise an MPLS label with each IPv6 prefix. Without it, you have plain MP-BGP IPv6 over an IPv4 session, and no way to forward the packets. - **`next-hop-self`** \- required, because the receiving PE must resolve the next-hop to a loopback it has a label-switched path to. - **`update-source Loopback0`** \- the LSP terminates on the loopback, so the BGP next-hop must be the loopback. ## Verification: the labels On R3, the BGP IPv6 table with labels: ``` R3#show bgp ipv6 unicast labels Network Next Hop In label/Out label 2001:DB8:10::/64 2001:DB8:13::1 17/nolabel 2001:DB8:60::/64 ::FFFF:4.4.4.4 nolabel/17 ``` Two lines, two directions: - **2001:DB8:10::/64** is R1's prefix, learned from the customer over native IPv6\. R3 allocated label **17** for it (the *in* label) and will advertise that label to R4\. There is no out label because the next hop is a directly connected customer, not another LSP. - **2001:DB8:60::/64** is R6's prefix, learned from R4 over the iBGP session. The next-hop is `::FFFF:4.4.4.4` \- there is the IPv4-mapped address, exactly as advertised. The out label is **17**: that is the label R4 allocated, and R3 must push it. The detail view makes the mechanism explicit: ``` R3#show bgp ipv6 unicast 2001:DB8:60::/64 BGP routing table entry for 2001:DB8:60::/64, version 2 Paths: (1 available, best #1, table default) 300 ::FFFF:4.4.4.4 (metric 11) from 4.4.4.4 (4.4.4.4) Origin IGP, metric 0, localpref 100, valid, internal, best mpls labels in/out nolabel/17 ``` `(metric 11)` is the OSPF cost to 4.4.4.4\. R3 resolved an IPv6 next-hop through the IPv4 IGP. That sentence is the entirety of 6PE. ### The LFIB ``` R3#show mpls forwarding-table Local Outgoing Prefix Bytes Label Outgoing Next Hop Label Label or Tunnel Id Switched interface 16 Pop Label 4.4.4.4/32 0 Et0/2 10.0.34.2 17 No Label 2001:DB8:10::/64 882 Et0/0 FE80::A8BB:CCFF:FE00:3A10 ``` Label 16 is an ordinary LDP label for the IPv4 loopback 4.4.4.4 - the transport label. Label 17 is the BGP-assigned IPv6 label: packets arriving with label 17 get it popped and are forwarded as native IPv6 out to the customer. R4 is the mirror image: ``` R4#show bgp ipv6 unicast labels Network Next Hop In label/Out label 2001:DB8:10::/64 ::FFFF:3.3.3.3 nolabel/17 2001:DB8:60::/64 2001:DB8:46::2 17/nolabel ``` ## Verification: end to end From R1 (customer, IPv6 only, no idea MPLS exists) to R6 (customer on the far side): ``` R1#show bgp ipv6 unicast Network Next Hop Metric LocPrf Weight Path *> 2001:DB8:10::/64 :: 0 32768 i *> 2001:DB8:60::/64 2001:DB8:13::2 0 100 300 i R1#ping 2001:DB8:60::1 source Loopback10 Sending 5, 100-byte ICMP Echos to 2001:DB8:60::1, timeout is 2 seconds: Packet sent with a source address of 2001:DB8:10::1 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 3/3/5 ms ``` And the traceroute, which is the single most satisfying output in this entire article: ``` R1#traceroute 2001:DB8:60::1 source 2001:DB8:10::1 probe 1 timeout 2 Tracing the route to 2001:DB8:60::1 1 2001:DB8:13::2 3 msec 2 2001:DB8:46::1 [MPLS: Label 17 Exp 0] 4 msec 3 2001:DB8:46::2 4 msec ``` Hop 2 shows `[MPLS: Label 17 Exp 0]`. An IPv6 traceroute, reporting an MPLS label, from a customer router that has never been configured with MPLS - because the ICMP time-exceeded came back from a router that was label-switching the packet. The IPv6 packet crossed an IPv4 core inside a label. That is 6PE working. Note that from R1's perspective the core is a single hop. The P routers do not decrement the IPv6 hop limit, because they never look at the IPv6 header. This is normal MPLS behaviour and it is why an MPLS core appears "flat" in traceroute. ## 6PE vs 6VPE vs dual-stack 6PE **Address family:** ipv6 unicast + send-label **Isolation:** none - global table **Use when:** you provide IPv6 internet transit over a v4 MPLS core 6VPE **Address family:** vpnv6 unicast **Isolation:** full - per-VRF, with RDs and RTs **Use when:** you sell IPv6 L3VPN to multiple customers Dual-stack core **Address family:** native IPv6 everywhere **Isolation:** none **Use when:** you are building new, or you can afford to touch every P router 6PE is a migration technology, and a very good one. It buys you IPv6 service delivery without a core upgrade. But it is not the destination: a dual-stack core is simpler to operate, simpler to troubleshoot, and does not depend on a label stack for basic reachability. Treat 6PE as the bridge, not the bank of the river. ## Troubleshooting 6PE 1. **Session will not come up.** Did you `activate` the neighbour in the IPv6 address family? With `no bgp default ipv4-unicast` set, a neighbour that is not activated in *any* AF has nothing to talk about, and the session goes to Idle. We demonstrate this exact failure in the [expert BGP troubleshooting article](https://www.pinglabz.com/expert-bgp-troubleshooting-scenarios/). 2. **Prefixes present, forwarding broken.** You are missing `send-label`. Run `show bgp ipv6 unicast labels` \- if every line says `nolabel/nolabel`, that is your answer. 3. **Next-hop inaccessible.** The next-hop is `::FFFF:` and must resolve through the IPv4 IGP with a label-switched path. Check `show mpls ldp neighbor` and `show mpls forwarding-table` for the remote loopback. No LSP, no 6PE. 4. **Missing `next-hop-self`.** If the 6PE passes the customer's link-local or global next-hop through untouched, the remote PE has no LSP to it. Always `next-hop-self` on the 6PE iBGP session. 5. **LDP not labelling the loopback.** The loopback must be a /32 in the IGP and LDP must have a binding for it. `show mpls ldp bindings 4.4.4.4 32`. ## Key takeaways - 6PE carries IPv6 across an IPv4-only MPLS core. The P routers are never touched and never learn an IPv6 route. - The iBGP session between 6PE routers runs over **IPv4** and carries the IPv6 unicast address family. Activate an IPv4 neighbour inside `address-family ipv6` \- that looks wrong and is exactly right. - `send-label` is the command that makes it 6PE. Without it you have MP-BGP IPv6 with no forwarding path. - The next-hop is an IPv4-mapped IPv6 address, `::FFFF:x.x.x.x`, which the receiving PE resolves through the IPv4 IGP and the LDP LSP. - Two labels: the LDP transport label to the remote loopback, and the BGP IPv6 label the egress PE assigned. - `show bgp ipv6 unicast labels` is the single most useful verification command. An IPv6 traceroute showing `[MPLS: Label n]` is the proof. - 6PE is global-table IPv6 transit. If you need per-customer isolation, that is 6VPE and the VPNv6 address family. Next: [five broken BGP scenarios, ticket style](https://www.pinglabz.com/expert-bgp-troubleshooting-scenarios/). The full cluster index lives on the [BGP pillar guide](https://www.pinglabz.com/bgp/), and the MPLS side of this story is on the [MPLS pillar](https://www.pinglabz.com/mpls/). ### BGP Multipath and Load Sharing: eBGP, iBGP, and eiBGP URL: https://www.pinglabz.com/bgp-multipath-load-sharing/ Last updated: 2026-07-12T06:32:03.000Z BGP installs one best path. That is the default and it is deliberate: BGP is a policy protocol, and the best-path algorithm exists to produce exactly one winner. But if you have two circuits to the same provider, one of them sitting idle is a waste of money. BGP multipath lets you install several equally-good paths and load-share across them. The configuration is a single line. The rules about *when* BGP will accept a second path are not, and one of them - a next-hop rewrite that nobody documents clearly - can hide a serious misconfiguration from you completely. This article covers eBGP, iBGP and eiBGP multipath on Cisco IOS XE with lab output. For the fundamentals, start at the [complete BGP guide](https://www.pinglabz.com/bgp/). ## The default: one best path, one route In the lab, R1 (AS 65001) has two parallel links to R3 (ISP-A, AS 100), each carrying its own eBGP session. Both sessions offer 6.6.60.0/24 with identical attributes. BGP picks one: ``` R1#show ip route 6.6.60.0 Routing entry for 6.6.60.0/24 Known via "bgp 65001", distance 20, metric 0 Routing Descriptor Blocks: * 10.0.13.2, from 10.0.13.2, 00:01:03 ago Route metric is 0, traffic share count is 1 ``` One descriptor block. One next-hop. The second circuit carries nothing but BGP keepalives. ## Enabling multipath ``` router bgp 65001 address-family ipv4 maximum-paths 2 ``` That is it for eBGP. Immediately: ``` R1#show ip bgp 6.6.60.0 BGP routing table entry for 6.6.60.0/24 Paths: (3 available, best #1, table default) Multipath: eBGP 100 300 10.0.13.2 from 10.0.13.2 (3.3.3.3) Origin IGP, localpref 100, valid, external, multipath, best 100 300 10.0.113.2 from 10.0.113.2 (3.3.3.3) Origin IGP, localpref 100, valid, external, multipath(oldest) 200 300 2.2.2.2 (metric 11) from 2.2.2.2 (2.2.2.2) Origin IGP, metric 0, localpref 100, valid, internal ``` Three things to read here. `Multipath: eBGP` tells you which flavour is active. Two paths are now flagged `multipath`. And the third path, learned over iBGP through ISP-B, is *not* a multipath - it has a different AS-path, so it does not qualify. The summary view uses an `m` flag for non-best multipaths: ``` R1#show ip bgp Network Next Hop Metric LocPrf Weight Path *> 6.6.60.0/24 10.0.13.2 0 100 300 i *m 10.0.113.2 0 100 300 i * i 2.2.2.2 0 100 0 200 300 i ``` And the RIB now has both: ``` R1#show ip route bgp B 6.6.60.0/24 [20/0] via 10.0.113.2, 00:01:49 [20/0] via 10.0.13.2, 00:01:49 ``` Both circuits are now carrying traffic, load-shared per-flow by CEF. ## The rules: what makes two paths equal BGP will only treat paths as multipath if they are equal on *every* best-path tiebreaker up to the point where the router would otherwise have to choose. In practice, that means all of the following must match: **Weight** Must be equal. A route-map setting weight on one neighbour silently kills multipath. **Local preference** Must be equal. **AS-path length AND content** By default the AS-path must be *identical*, not merely the same length. Relax this with `bgp bestpath as-path multipath-relax`. **Origin code** IGP, EGP, incomplete - must match. **MED** Must be equal. **Path type** eBGP with eBGP, or iBGP with iBGP. Mixing them requires eiBGP multipath (see below). **IGP metric to next-hop** For iBGP multipath, the IGP cost to reach each next-hop must be equal. Notice what is *not* on the list: the neighbour's router ID, the peer address, and the neighbour's AS number for the immediate hop. Two sessions to the *same physical router* over two links, as in this lab, absolutely qualify - and that is the most common real-world multipath deployment there is. ### as-path multipath-relax By default, two paths with AS-paths `100 300` and `200 300` are not multipath, even though they are the same length. BGP insists the paths be identical, on the grounds that they traverse genuinely different networks and might have wildly different characteristics. ``` router bgp 65001 address-family ipv4 bgp bestpath as-path multipath-relax maximum-paths 2 ``` With `multipath-relax`, equal *length* is enough. This is how you load-share across two different upstream providers. It is also how you get asymmetric behaviour, unpredictable latency, and support tickets you cannot reproduce - because the two paths are genuinely different networks. Turn it on deliberately, knowing that. A hard constraint you cannot relax: the paths must still originate from the **same neighbouring AS** for the router to consider them at the eBGP multipath stage. `multipath-relax` loosens the AS-path *content* check, not the AS-path length check. ## iBGP multipath Same command, different qualification rule. For iBGP paths to be multipath, the IGP metric to each BGP next-hop must be equal. ``` router bgp 65001 address-family ipv4 maximum-paths ibgp 2 ``` This is the design that matters in a data centre or a large campus: two route reflectors, two exit routers, equal IGP cost from your position to both, and you load-share outbound across both exits. If your IGP metrics are not equal - and in a real network with mixed link speeds they usually are not - iBGP multipath quietly does nothing. Check with: ``` show ip bgp ``` and look at the `(metric N)` value on each iBGP path. If they differ, that is your answer. ## eiBGP multipath Mixing an eBGP path and an iBGP path into a single load-shared set is called eiBGP multipath, and on IOS it is **only supported inside a VRF**: ``` router bgp 65001 address-family ipv4 vrf CUSTOMER-A maximum-paths eibgp 2 ``` This exists for exactly one scenario: an MPLS L3VPN PE with a dual-homed customer site, where the site is reachable both directly (eBGP to the CE) and across the MPLS core (iBGP/VPNv4 from the remote PE). It lets the PE load-share between the local link and the core. Outside a VRF, eiBGP multipath is not available on IOS XE, and for good reason: mixing an external path with an internal one in the global table produces routing that is very hard to reason about. If you find yourself wanting it in the global table, what you actually want is a design change. ## The trap: multipath silently applies next-hop-self This is the finding from the lab that is worth the price of admission, and it is genuinely not well documented. R1 is an eBGP edge router with an iBGP session to R2\. Standard practice is `neighbor 2.2.2.2 next-hop-self`, so R2 receives eBGP routes with R1's loopback as the next-hop rather than an ISP address it cannot reach. Remove `next-hop-self`, and the classic symptom should appear instantly: R2 receives routes with an unresolvable next-hop and never installs them. Except it did not: ``` ! next-hop-self removed from R1, maximum-paths 2 still configured R2#show ip bgp 100.100.100.0/24 BGP routing table entry for 100.100.100.0/24 Paths: (1 available, best #1, table default) 100 1.1.1.1 (metric 11) from 1.1.1.1 (1.1.1.1) Origin IGP, metric 0, localpref 200, valid, internal, best ``` Next-hop 1.1.1.1 - R1's loopback. Exactly what `next-hop-self` would have produced, except `next-hop-self` is not configured. The misconfiguration is completely invisible. Now remove `maximum-paths 2` and clear the session: ``` R2#show ip bgp 100.100.100.0/24 BGP routing table entry for 100.100.100.0/24, version 0 Paths: (1 available, no best path) 100 10.0.13.2 (inaccessible) from 1.1.1.1 (1.1.1.1) Origin IGP, metric 0, localpref 200, valid, internal R2#show ip bgp | include 100.100 * i 100.100.100.0/24 10.0.13.2 0 200 0 100 i ``` There it is: next-hop 10.0.13.2, `(inaccessible)`, no `>`, never installed in the RIB. The classic failure, revealed only once multipath was removed. **Why this happens:** when a prefix is installed as a multipath, there is no single external next-hop to pass along - there are several. IOS resolves this by advertising the route to iBGP peers with itself as the next-hop. Which is sensible, and which means that on any router where eBGP multipath is active, the absence of `next-hop-self` is masked. **Why you should care:** the day someone disables multipath, or the day one of the two circuits fails and the prefix drops back to a single path, every iBGP peer suddenly starts receiving unreachable next-hops and your internal routing falls over. The bug was sitting there the whole time. Configure `next-hop-self` explicitly on every iBGP session from an eBGP edge router, always, whether or not multipath appears to make it unnecessary. ## How traffic is actually shared Installing two paths in the RIB does not, by itself, split traffic evenly. CEF does the sharing, and by default it hashes **per destination flow** (source and destination IP). That means: - A single large TCP transfer uses exactly one path. Multipath will not make one download faster. - Many flows spread reasonably evenly, assuming diverse addressing. - A handful of very large flows can land on the same path and leave the other idle. This is normal and is not a bug. Check what CEF is doing: ``` show ip cef internal show ip cef exact-route ``` `exact-route` is the one you want in a ticket: give it a real source and destination and it tells you exactly which link that flow will take. You can move to per-packet load sharing with `ip load-sharing per-packet` on the interface, and you should not. It reorders packets, TCP hates it, and every modern application stack performs worse. ## Design guidance - **Two links, one provider:** plain `maximum-paths N`. Simple, symmetric, no surprises. This is the case worth doing. - **Two providers:** think hard before reaching for `multipath-relax`. Different providers means different latency, different congestion, different failure modes. Load-sharing across them makes your network's behaviour a function of two networks you do not control. Often the better answer is to prefer one and keep the other hot-standby. - **Inside your AS:** `maximum-paths ibgp N`, and make your IGP metrics genuinely equal, or it will not engage. - **Always set `next-hop-self` anyway.** See above. - **Watch your FIB.** Every multipath prefix consumes multiple FIB entries. On a platform with a hardware FIB and a full internet table, this is not free. ## Key takeaways - `maximum-paths N` for eBGP, `maximum-paths ibgp N` for iBGP, `maximum-paths eibgp N` for mixed - and eiBGP is VRF-only on IOS XE. - Paths must match on weight, local-pref, AS-path, origin, MED, and path type. For iBGP, the IGP metric to the next-hop must also match. - `bgp bestpath as-path multipath-relax` loosens AS-path *content* to just *length*. Use it consciously, mainly for multi-provider load sharing. - The `m` flag in `show ip bgp` and the `multipath` keyword in the detail view are how you confirm it engaged. - **eBGP multipath silently rewrites the next-hop toward iBGP peers**, masking a missing `next-hop-self`. Configure `next-hop-self` explicitly regardless. - Load sharing is per-flow by default. One big flow takes one path. That is correct behaviour. Next: [MP-BGP for IPv6 and 6PE](https://www.pinglabz.com/ipv6-bgp-6pe/), where IPv6 prefixes cross an IPv4-only MPLS core with real label output. The full cluster index lives on the [BGP pillar guide](https://www.pinglabz.com/bgp/). ### Designing BGP Policy with Communities: A Practical Framework URL: https://www.pinglabz.com/bgp-community-policy-design/ Last updated: 2026-08-01T19:35:22.000Z Most networks use BGP communities the way most people use a junk drawer: things get put in, nobody agrees what they mean, and after two years nobody dares throw anything away. That is a shame, because a well-designed community scheme is the single highest-leverage thing you can build in a BGP network. It turns policy from a pile of per-neighbour route-maps into a signalling system that scales. This article builds a real community framework end to end, implements it on Cisco IOS XE, and proves each behaviour with lab output - including a working remote-triggered black hole (RTBH) and one platform behaviour that will absolutely catch you out. For the fundamentals, see the [complete BGP guide](https://www.pinglabz.com/bgp/). ## What a community actually is A BGP community is a 32-bit tag attached to a route as an optional transitive attribute. That is the whole specification. BGP itself assigns no meaning to any value (with four well-known exceptions). The meaning is a convention that you and your peers agree on. The universal convention is to write them as `ASN:value`. AS 65001's community 101 is written `65001:101`. The first half identifies who defined the meaning; the second half is theirs to allocate. The well-known ones, which every implementation honours without configuration: **no-export** Do not advertise this route to any eBGP peer outside the local AS or confederation. The route stays inside. **no-advertise** Do not advertise this route to *any* peer at all, internal or external. It dies at the router that received it. **local-AS (no-export-subconfed)** Do not advertise outside the local confederation sub-AS. **internet** Advertise to everyone. Effectively a no-op, but useful as an explicit "match all" in community-lists. Everything else is yours to define. This article assumes you can already tag, match and act on a community on a Cisco box. If any of that is new, [setting and matching communities on IOS XE](https://www.pinglabz.com/bgp-communities/) covers the syntax and the `send-community` line everyone forgets, and the design decisions below will read much more clearly afterwards. ### First, turn on new-format Before anything else, on every router in the design: ``` ip bgp-community new-format ``` Without it, IOS displays communities as raw 32-bit integers - `Community: 6554600 4259905637` \- and `show ip bgp community 65001:666` is rejected as invalid input. It is enabled by default on most modern images but not all, and it costs nothing to be explicit. It is a display and parsing setting only; the wire format never changes. ## Designing the scheme A community scheme is an addressing plan for policy. Like an addressing plan, the value is in the structure, not the individual numbers. Allocate ranges by *purpose*, and leave gaps. Here is the scheme used in the lab, which is a realistic small-enterprise pattern: 65001:1xx - Origin Where did this route come from? `65001:101` \= HQ, `65001:102` \= branch, `65001:103` \= data centre. Set once, at the point of origination. Never changed downstream. 65001:2xx - Scope How far should it go? `65001:201` \= ISP-A only, `65001:202` \= ISP-B only, `65001:210` \= everywhere. Consumed by the edge routers' outbound policy. 65001:666 - Blackhole Discard traffic to this prefix. The RTBH trigger. Deliberately memorable, deliberately unmistakable. 100:1xx - Provider's own tags ISP-A tags what it sends you: `100:110` \= my own routes, `100:120` \= transit routes I learned elsewhere. You act on those inbound. Three rules make a scheme survive contact with reality: 1. **One community, one meaning.** Never overload a value. If you need to express two things, use two communities - they are cheap and a route can carry many. 2. **Set at the edge, act at the edge.** Origin communities are set where the route enters BGP. Scope communities are consumed where the route leaves your AS. The middle of your network should be a dumb pipe that preserves them. 3. **Document it in one place, and version it.** A community scheme with no document is a scheme that will be reverse-engineered from route-maps by a stranger at 3 a.m. ## Implementation: tagging on origination R1 is the AS 65001 edge toward ISP-A. It originates 10.10.10.0/24 (a normal site prefix) and 10.10.10.66/32 (a host under attack that we want black-holed). Both get tagged on the way out: ``` ip prefix-list SITE-A seq 5 permit 10.10.10.0/24 ip prefix-list SITE-A seq 10 permit 1.1.1.1/32 ip prefix-list BLACKHOLE-PFX seq 5 permit 10.10.10.66/32 route-map ISPA-OUT permit 10 match ip address prefix-list BLACKHOLE-PFX set community 65001:101 65001:666 route-map ISPA-OUT permit 20 match ip address prefix-list SITE-A set community 65001:101 65001:210 route-map ISPA-OUT permit 30 router bgp 65001 address-family ipv4 neighbor 10.0.13.2 send-community both neighbor 10.0.13.2 route-map ISPA-OUT out ``` Two details that are easy to get wrong: - **`send-community` is not on by default.** Without it, IOS strips communities on egress and your entire scheme silently evaporates. `both` sends standard and extended communities. - **The empty `permit 30` at the end is load-bearing.** A route-map has an implicit deny. Without a bare permit clause, every prefix that does not match sequences 10 or 20 is dropped. This one line is responsible for a substantial fraction of all BGP outages. ## Implementation: acting on communities ISP-A (R3) is where the policy is enforced. It matches the blackhole community, rewrites the next-hop to a discard address, drops the local preference, and adds `no-export` so the black-holed route never leaves AS 100: ``` ip route 192.0.2.1 255.255.255.255 Null0 interface Loopback254 description RTBH discard next-hop lives here ip address 192.0.2.254 255.255.255.0 ip community-list standard CUST-BLACKHOLE permit 65001:666 route-map CUST-IN permit 10 match community CUST-BLACKHOLE set ip next-hop 192.0.2.1 set local-preference 50 set community no-export additive route-map CUST-IN permit 20 set community 100:1000 additive router bgp 100 address-family ipv4 neighbor 10.0.13.1 route-map CUST-IN in ``` `additive` is the keyword to memorise. Without it, `set community` *replaces* every community on the route. With it, the new value is appended and the customer's origin tag survives. Getting this wrong destroys other people's policy silently. ### The RTBH, working ``` R3#show ip bgp 10.10.10.66/32 BGP routing table entry for 10.10.10.66/32 Paths: (2 available, best #2, table default, not advertised to EBGP peer) 65001 192.0.2.1 from 10.0.13.1 (1.1.1.1) Origin IGP, metric 0, localpref 50, valid, external, best Community: 65001:101 65001:666 no-export R3#show ip cef 10.10.10.66 255.255.255.255 10.10.10.66/32 nexthop 192.0.2.1 Null0 ``` Read that carefully, because there is a lot in it: - The customer's origin tag `65001:101` survived (that is `additive` doing its job). - `no-export` is present, and IOS explicitly says `not advertised to EBGP peer`. The black hole stays inside AS 100. - CEF resolves the prefix to `Null0`. Traffic destined for the victim is dropped at the provider edge, which is exactly the point - the attack traffic never reaches the customer's circuit. ### The platform trap: a Null0 static is not enough The textbook RTBH recipe is "set the next-hop to a discard address and add `ip route 192.0.2.1 255.255.255.255 Null0`". On IOS XE 17.18 that **does not work on its own**. Verified both ways in the lab: ``` ! With ONLY the Null0 static, no connected route covering 192.0.2.0/24: R3#show ip bgp 10.10.10.66/32 Paths: (2 available, no best path) 192.0.2.1 (inaccessible) from 10.0.113.1 (1.1.1.1) 192.0.2.1 (inaccessible) from 10.0.13.1 (1.1.1.1) R3#show ip cef 10.10.10.66 255.255.255.255 %Prefix not found ``` BGP's next-hop validation refuses to accept a next-hop whose only resolution is a Null0 static. The path is `valid` but the next-hop is `(inaccessible)`, so it never becomes best and never reaches the RIB. Your black hole does nothing. The fix is to give the discard address a connected route to live in. A loopback carrying 192.0.2.254/24 makes 192.0.2.0/24 a connected network, which satisfies next-hop validation - while the more specific /32 Null0 static still wins in CEF and does the actual dropping: ``` interface Loopback254 ip address 192.0.2.254 255.255.255.0 ip route 192.0.2.1 255.255.255.255 Null0 ``` We removed the loopback and re-tested to confirm the causality, and the route went straight back to `(inaccessible)`. This is not a lab artefact. If your RTBH silently does nothing, check `show ip cef` before you check anything else. ## Implementation: the inbound side ISP-A tags what it sends you, and you act on it. Its own routes get preferential treatment; the transit routes it learned from someone else do not: ``` ip community-list standard ISPA-OWN permit 100:110 route-map ISPA-IN permit 10 match community ISPA-OWN set local-preference 200 route-map ISPA-IN permit 20 ``` And the result, which is the whole scheme working in miniature: ``` R1#show ip bgp 100.100.100.0/24 100 10.0.13.2 from 10.0.13.2 (3.3.3.3) Origin IGP, metric 0, localpref 200, valid, external, multipath, best Community: 100:110 R1#show ip bgp 6.6.60.0/24 | include Community|localpref Origin IGP, localpref 100, valid, external, multipath, best Community: 100:120 ``` ISP-A's own network gets local-pref 200\. Everything it merely transits stays at 100\. Not one prefix was named in a route-map. Add a thousand more prefixes tomorrow and the policy still holds, because the policy is expressed in tags, not addresses. **That is the payoff.** ## Standard, extended, and large communities Standard (RFC 1997) **Size:** 32-bit, written ASN:value **Breaks on:** 4-byte ASNs - your AS number will not fit in 16 bits **Use for:** everything, if you have a 2-byte ASN Extended (RFC 4360) **Size:** 64-bit, typed **Used by:** MPLS L3VPN route targets, EIGRP SoO **Not:** a general-purpose replacement for standard communities Large (RFC 8092) **Size:** 96-bit, ASN:function:parameter **Solves:** 4-byte ASNs properly **Use for:** any new design with a 4-byte ASN - this is the modern answer If you have a 4-byte AS number, standard communities cannot encode it and you should be designing with **large communities** from day one. IOS XE supports them, and the CLI mirrors the standard-community commands (`set large-community`, `ip large-community-list`, `send-community both` covers them). ## Troubleshooting checklist 1. **Communities not arriving?** `send-community` is missing on the sender. This is the number one cause, every time. 2. **Communities arriving but the previous tags are gone?** Somebody used `set community` without `additive`. 3. **`show ip bgp community 65001:666` rejected?** `ip bgp-community new-format` is not on. 4. **Route-map matching nothing?** Community-lists are matched as a set, and a *standard* community-list with multiple values on one line requires *all* of them to be present. Use separate `permit` lines for OR logic. 5. **Random prefixes disappearing after you added a route-map?** The implicit deny at the end. Add a bare `permit` clause. 6. **RTBH tagged correctly but traffic still flowing?** Check `show ip cef `. If the BGP entry says `(inaccessible)`, your discard next-hop needs a connected route (see above). ## Key takeaways - Communities carry no built-in meaning. They are a signalling convention, and their value comes entirely from having a structure and sticking to it. - Allocate ranges by purpose - origin, scope, action - and leave gaps. Document it once, and treat that document as the contract. - Set tags at origination, act on them at the edge, and let the middle of your network preserve them untouched. - `send-community both` on every neighbour, and `additive` on every `set community` that is not deliberately replacing the whole set. - The empty `permit` at the end of an outbound route-map is not optional. - RTBH is a community scheme's killer app: one tag, and your provider drops the attack traffic before it reaches your circuit. On IOS XE 17.x the discard next-hop needs a connected route to survive BGP next-hop validation. - 4-byte ASN? Use large communities (RFC 8092), not standard. Next: [BGP multipath and load sharing](https://www.pinglabz.com/bgp-multipath-load-sharing/), and a next-hop behaviour that hides a classic misconfiguration completely. The full cluster index lives on the [BGP pillar guide](https://www.pinglabz.com/bgp/). ### BGP Route Dampening: Penalties, Half-Lives, and Why It Fell Out of Favor URL: https://www.pinglabz.com/bgp-route-dampening/ Last updated: 2026-08-01T18:36:48.000Z BGP route dampening is one of the few networking features with a genuine story arc: invented to save the internet, deployed everywhere, then quietly turned off by nearly everyone because it was making things worse. It is still on the CCIE Enterprise Infrastructure blueprint, still in the IOS XE CLI, and still occasionally the right answer. You need to know how it works, and you need to know why the industry backed away from it. This article covers the penalty algorithm, the four timers, and real dampening output captured on a two-router CML lab running Cisco IOS XE 17.18.2 on `iol-xe` nodes: a prefix crossing the suppress limit at penalty 2818, the `*d` flag, and an identical prefix from the same neighbor sitting untouched at `*>`. For the fundamentals, start with the guide to [how BGP carries routes between autonomous systems](https://www.pinglabz.com/bgp/). It also covers the gotcha that makes most dampening labs produce absolutely nothing, which is not the algorithm's fault and is nowhere near well enough known. ## The problem it was invented to solve An unstable link somewhere on the internet does not stay local. Every time a prefix goes down and comes back, its origin AS sends a withdraw and then an announcement. Those propagate outward, hop by hop, and every router in the path burns CPU running the best-path algorithm again. A single flapping circuit in one corner of the world could, in the 1990s, meaningfully load routers everywhere else. Dampening (RFC 2439) puts a cost on instability. Each time a route flaps, the router adds a penalty. When the penalty crosses a threshold, the route is suppressed: the router stops using it and stops advertising it, even if the route is currently up. The penalty then decays exponentially. When it drops below a reuse threshold, the route is allowed back. It is a circuit breaker, and it is deliberately unfair to routes that misbehave. ## The algorithm Four numbers control everything, and IOS takes them in a fixed order: ``` bgp dampening ``` **half-life (default 15 min)** How long it takes for an accumulated penalty to decay to half its value. Longer half-life means a longer memory for bad behaviour. **reuse (default 750)** When the decaying penalty drops below this value, a suppressed route is un-suppressed and put back into service. **suppress (default 2000)** When the penalty rises above this value, the route is suppressed. Since each flap adds 1000, the default means a route survives one flap and dies on the third. **max-suppress-time (default 60 min)** A hard ceiling on how long a route can stay suppressed, regardless of penalty. This also implicitly caps the maximum penalty a route can accumulate. Bare `bgp dampening` with no arguments gives you exactly those four values. The penalty arithmetic itself is fixed and not configurable: - **+1000** for a route withdrawal (a flap). - **+500** for an attribute change, on some implementations. IOS by default only penalises withdrawals unless you enable `bgp dampening ... route-map` with attribute-change tracking. - **Exponential decay** between events, governed by the half-life. The gap between suppress at 2000 and reuse at 750 is not decoration. It forces a route that just crossed the threshold to wait through most of a half-life, so a marginal prefix cannot oscillate in and out of suppression on every decay tick. Two scoping facts worth committing to memory. **Dampening only applies to eBGP-learned routes**, because your own AS's instability is your own problem to fix, not something to hide behind a circuit breaker. And penalty is tracked **per path**, so the same prefix learned from two neighbors carries two independent penalties and can be damped from one while still usable from the other. ## What this was captured on Two routers, one eBGP session, two prefixes from the same origin. One is deliberately abused, the other is a control that is never touched, and that is what makes the difference legible. PlatformCML, 2 x iol-xe nodes, Cisco IOS XE 17.18.2 PeeringR1 in AS 65001 to R2 in AS 65002, eBGP over 10.0.12.0/30 Control prefix100.100.100.0/24 from R1 Lo1, never touched Flapped prefix100.100.111.0/24 from R1 Lo2, shut and no-shut five times about 8s apart DampeningR2 only, plain bgp dampening with the 15 / 750 / 2000 / 60 defaults AutomationOn-box EEM applets drove the flaps and captured the show output to syslog Flap timing is the whole experiment here, so flapping the loopback by hand introduces jitter you cannot account for afterwards. That is why [scripting a repeatable fault on the device itself with EEM](https://www.pinglabz.com/eem-applets-cisco-ios-xe/) was worth the setup: the same cadence every run, and captures taken at a fixed offset. The receiving side is one line: ``` router bgp 65002 bgp dampening 15 750 2000 60 ``` ## Lab: watching a route get suppressed ### The gotcha that nearly ruined this lab The first run produced nothing at all. Five clean shut and no-shut cycles on R1's Lo2, dampening configured and confirmed on R2, and R2 counted zero flaps. Nothing suppressed, nothing in the flap statistics, no penalty recorded anywhere. The reason is the **eBGP advertisement interval**, or MRAI, which defaults to 30 seconds toward eBGP peers. IOS batches eBGP updates on that timer. If a prefix goes down and comes back inside one interval, the sender does not transmit a withdraw followed by an announcement; it transmits whatever the net state was when the timer expired, and if the prefix ended up where it started, that is nothing at all. Dampening cannot penalise a flap it never received. That is not a lab artefact, it is a real and important property: **the advertisement interval is itself a flap suppressor**, it works upstream of dampening at no cost and with no hold-down, and whatever it absorbs never reaches the penalty algorithm. When dampening is configured and appears to do nothing, check this before you start questioning your thresholds. For the demo we disable it on the sending side: ``` router bgp 65001 neighbor 10.0.12.2 advertisement-interval 0 ``` Now every flap propagates immediately, and dampening has something to count. Even then, note what happened: **five interface flaps produced three counted ones**. Some transitions were still absorbed before becoming distinct update events. Do not calibrate a dampening policy assuming every physical bounce equals one penalty step. The other version of "dampening never fires" is a session dropping entirely rather than a prefix bouncing, since a reset withdraws everything behind it at once. Rule that out by understanding [what a BGP session that keeps cycling back to Idle is actually doing](https://www.pinglabz.com/bgp-neighbor-states/) before you blame the penalty math. ### Suppression With the advertisement interval out of the way, the per-prefix view on R2 is where the arithmetic stops being theory: ``` R2#show ip bgp 100.100.111.0/24 BGP routing table entry for 100.100.111.0/24, version 8 Paths: (1 available, no best path) <-- path exists, but is not a candidate Not advertised to any peer Refresh Epoch 1 65001, (suppressed due to dampening) <-- the smoking gun 10.0.12.1 from 10.0.12.1 (100.100.111.1) Origin IGP, metric 0, localpref 100, valid, external Dampinfo: penalty 2818, flapped 3 times in 00:01:43, reuse in 00:06:49 rx pathid: 0, tx pathid: 0 Updated on Jul 20 2026 22:59:25 UTC ``` Three lines carry the story. `Paths: (1 available, no best path)` says the path is sitting in the BGP table but is excluded from best-path selection, so it is not in the RIB and not advertised onward. `(suppressed due to dampening)` names the reason outright, which saves you a long detour through route policy and next-hop reachability. And `Dampinfo: penalty 2818` is the number itself. Now check that number against the flap count on the same line. Three withdrawals at roughly 1000 each is about 3000 of raw penalty, and the router reports 2818\. The missing 180-odd is decay: the first flap's penalty had already been shrinking for over a minute and a half by the time the third one landed. Penalty is not a counter, it is a decaying quantity that gets topped up, which is why a prefix flapping three times across an afternoon never gets near suppression while one flapping three times in 103 seconds does. 2818 is past the 2000 suppress limit, so the path is out of service until decay drags it under 750\. The router has already worked out how long that takes: 6 minutes 49 seconds. The route is up. The customer's loopback is fine. The link is fine. And R2 is refusing to use it or carry it for the next six and a half minutes. **That last sentence is the whole controversy.** ### The fleet-wide views Two commands give you the whole router at once. The first lists only what is currently held down, with the release time: ``` R2#show ip bgp dampening dampened-paths Network From Reuse Path *d 100.100.111.0/24 10.0.12.1 00:06:49 65001 i ``` The second adds flap history, and will show you prefixes that have accrued penalty without yet crossing the suppress limit: ``` R2#show ip bgp dampening flap-statistics Network From Flaps Duration Reuse Path *d 100.100.111.0/24 10.0.12.1 3 00:01:43 00:06:49 65001 ``` Read `Duration` carefully: it is the window the flaps occurred over, not how long the route has been suppressed. Three flaps in 00:01:43 is a prefix in trouble. Three flaps in 04:00:00 would never come close to the suppress limit, because decay eats each one before the next arrives. The stable control prefix has no flap history and appears in neither output. ### Status codes and history entries The status code legend at the top of `show ip bgp` is the fastest read on the box: ``` Status codes: s suppressed, d damped, h history, * valid, > best, i - internal, ... ``` Four of those matter here. `*` is valid, `>` is best, and `d` is damped. The one that catches people out is `h`, a *history entry*: the router is tracking a penalty for a prefix that is currently withdrawn. There is no route to use, so there is nothing to suppress, and the router is simply remembering in case it comes back. If it stays down long enough for the penalty to decay below reuse, the history entry is garbage-collected. The dangerous one is `s`. That is a more specific prefix suppressed by an aggregate, a different mechanism that happens to share the word. Read an `s` as dampening and you will go looking for flaps that never happened. ## Damped against stable, side by side This is the capture that makes dampening click, because both prefixes come from the same neighbor over the same session with the same attributes: ``` R2#show ip bgp BGP table version is 8, local router ID is 10.0.12.2 Status codes: s suppressed, d damped, h history, * valid, > best, i - internal, ... Network Next Hop Metric LocPrf Weight Path *> 100.100.100.0/24 10.0.12.1 0 0 65001 i <-- stable: valid + best *d 100.100.111.0/24 10.0.12.1 0 0 65001 i <-- damped: present, not installed ``` The control prefix is `*>`: valid, best, installed in the RIB, advertised onward. The abused prefix is `*d`: valid and damped. It is not missing, not filtered, not unreachable, and its next hop resolves perfectly. It simply never entered the tournament, which is a different failure mode from a path that entered and lost on one of the tiebreakers in [the order BGP compares candidate paths in](https://www.pinglabz.com/bgp-best-path-selection/). Dampening is a gate in front of that process, not a step inside it. That is the value proposition in one screen: one unstable prefix stops churning the RIB and stops propagating updates downstream, while everything else behind the session carries on untouched. The session never went down. Only the badly behaved prefix was punished. ## Why the industry turned it off The killer problem is that **dampening penalises the wrong thing**. It counts *updates received*, not *instability at the source*. Consider a prefix advertised through several ASes. When it flaps once at the origin, the withdraw propagates. But BGP path exploration means that intermediate routers, before converging on "gone", will announce a series of alternative paths they briefly believe in. Each of those is an update. Each of those looks like a flap to a downstream dampener. A single flap at the origin can look like five or six flaps three ASes away. The result, documented in a 2002 study by Mao, Govindan, Varghese and Katz, is that **well-connected prefixes get dampened harder than poorly-connected ones**. The more paths available to reach you, the more path exploration, the more updates, the more penalty. Multihoming, the thing you paid for to be more reliable, actively made you more likely to be suppressed. Combined with default timers that suppress a route for up to an hour after a couple of genuine flaps, dampening was routinely turning a 30-second outage into a far longer one, so the mitigation caused more customer-visible downtime than the instability it was mitigating. RIPE issued [RIPE-378](https://www.ripe.net/publications/docs/ripe-378?ref=pinglabz.com) in 2006 recommending operators simply stop using it. ### The rehabilitation: RIPE-580 and RFC 7196 The consensus later shifted again. [RIPE-580](https://www.ripe.net/publications/docs/ripe-580?ref=pinglabz.com), published alongside the work that became RFC 7196, made the case that the algorithm was never really the problem: the shipped default parameters were. Router CPUs and BGP implementations had improved enormously, and thresholds chosen for 1990s hardware were punishing prefixes that any modern router absorbs without noticing. The direction of that guidance is toward substantially more tolerant thresholds than the defaults, and toward scaling how aggressive you are by prefix length, on the reasoning that a flapping short prefix carries far more traffic than a flapping /24\. **Take the actual numbers from the current published guidance rather than from any article, including this one.** They have been revised more than once, and a stale suppress threshold copied off a web page is precisely how this feature earned its reputation in the first place. ### Selective dampening instead of blanket dampening Plain `bgp dampening` applies one set of parameters to every eBGP path the router learns. That is fine in a lab and a poor fit for a real edge, where one misbehaving customer is usually the actual problem and where a flapping /8 and a flapping /24 are not the same event. The selective form drives the parameters from a route-map: ``` ! values below are illustrative - take production thresholds from current guidance route-map DAMPEN-POLICY permit 10 match community NOISY-CUSTOMER set dampening 30 1500 6000 60 ! route-map DAMPEN-POLICY permit 20 match ip address prefix-list LONG-PREFIXES set dampening 30 750 3000 60 ! router bgp 65002 address-family ipv4 bgp dampening route-map DAMPEN-POLICY ``` Normal route-map ordering applies: first match wins, and a path matching no clause is not dampened at all, which is a useful default posture. In production the match is usually a community rather than a prefix list, because the tag travels with the route and survives readvertisement, so the real groundwork is [a community scheme you can hang policy off](https://www.pinglabz.com/bgp-community-policy-design/) rather than the dampening syntax itself. ## Operational commands you will actually use ``` ! What is currently suppressed, and when it comes back show ip bgp dampening dampened-paths ! Flap count and duration for everything with a penalty show ip bgp dampening flap-statistics ! The configured thresholds show ip bgp dampening parameters ! One prefix: penalty, flap count, reuse timer show ip bgp 100.100.111.0/24 ! Rescue a route right now (customer is on the phone) clear ip bgp dampening 100.100.111.0 255.255.255.0 ! Rescue everything clear ip bgp dampening ``` `clear ip bgp dampening` is the command you will type at 3 a.m. It zeroes the penalty and immediately reinstates the route. Know it before you need it, because a suppressed prefix raises no alarm of its own: it is simply, quietly absent, and everything else looks healthy. ## Should you turn it on? Honest answer for most enterprises: **no**. You are not a transit provider. Your neighbour count is small. Your router CPU is not remotely challenged by BGP updates. The failure mode of dampening, extending a short outage into a long one, is strictly worse for you than the problem it solves. Turn it on if you are carrying somebody else's instability: a provider edge facing one customer generating pathological update volume, or an internal boundary where one region's churn keeps rippling into the core. You are dampening an identified source, not the whole table on principle. Use tolerant thresholds from current guidance, apply them selectively through a route-map, and instrument the result so you know when it fires. For everyone else, the modern answer to instability is **fix the instability**: detect failures fast and cleanly with [sub-second failure detection instead of protocol hold timers](https://www.pinglabz.com/bfd-bidirectional-forwarding-detection/), use carrier-delay and dampening on the *interface*, and take a hard look at whatever physical layer keeps bouncing. ## Common mistakes and gotchas - **Dampening configured, nothing ever dampens.** The eBGP advertisement interval is almost always why. Rapid flaps get coalesced into a net-zero update the receiver never sees, and it cannot penalise what it was not sent. - **Assuming every physical flap becomes a penalty step.** Five interface bounces produced three counted flaps in this lab, even with batching disabled on the sender. - **Reading penalty as a flap counter.** Three flaps showed 2818, not 3000, because decay was already eating the first one. Timing changes the total, which is the entire point of an exponential decay. - **Expecting `*d` to mean the route is gone.** The path is present, valid and correctly formed, just excluded from best-path selection. Chasing it as a filtering or next-hop problem wastes hours. - **Confusing `s` with `d`.** `s` is aggregation suppressing a more specific prefix. `d` is dampening. Different feature, same word. - **Deploying the defaults because they are the defaults.** A 15 minute half-life with a 2000 suppress limit is the exact configuration the operator community spent a decade warning about. ## Key takeaways - Dampening adds a penalty (about 1000 per flap) that decays exponentially. Above `suppress`, the route is dropped from best-path selection; below `reuse`, it comes back. - The lab proved the math: penalty 2818 from 3 flaps in 00:01:43, `(suppressed due to dampening)` with `no best path`, reuse in 6m49s, while the control prefix from the same neighbor stayed `*>`. - It only ever applies to eBGP-learned routes, and penalty is tracked per path. - The `d` flag in `show ip bgp` means damped. The `h` flag means a history entry: penalty recorded, route currently withdrawn. The `s` flag is aggregation, not dampening. - The eBGP advertisement interval batches updates and hides fast flaps from dampening entirely. If your dampening lab "does not work", this is why. - IOS default timers are far too aggressive and turn brief outages into long ones. If you dampen at all, use current published thresholds and apply them selectively with `bgp dampening route-map`. - Path exploration means well-connected prefixes generate more updates and therefore get dampened harder. This is the core reason the internet backed away from the feature. - `clear ip bgp dampening ` is your emergency override. Dampening is a blueprint topic to reason about cleanly and a production feature to deploy narrowly. Know the four parameters, know the advertisement interval sits in front of them, and know that `*d` means a healthy path held out of service on purpose. For where this sits among path attributes, policy and convergence, work through [the rest of the BGP cluster](https://www.pinglabz.com/bgp/) in order. ### BGP Outbound Route Filtering (ORF): Pushing Your Filters to the Neighbor URL: https://www.pinglabz.com/bgp-outbound-route-filtering-orf/ Last updated: 2026-07-12T06:32:01.000Z Inbound prefix filtering is normal hygiene: you accept the routes you want and drop the rest. But think about what actually happens on the wire. Your neighbour sends you a full table, your router receives every prefix, parses every attribute, runs every one against your prefix-list, and throws most of them away. You paid the CPU and the memory to discard routes you never wanted. Outbound Route Filtering (ORF) fixes that by pushing your inbound filter *to the neighbour*. They apply it on egress. The routes never leave their router. This article shows ORF working on Cisco IOS XE with real output from a CML lab, including the exact commands the receiving router uses to see the filter you sent it. For the wider context, see the [complete BGP guide](https://www.pinglabz.com/bgp/). ## What ORF actually is ORF is a BGP capability, negotiated at session establishment, defined in RFC 5291 with the prefix-list ORF type in RFC 5292\. The idea is simple: one router advertises a filter, the other router installs it as an outbound policy for that session. Two modes are negotiated independently, and both sides have to agree: send "I will send you my inbound prefix-list." Configure this on the router that *has* the filter and wants the peer to enforce it. receive "I will accept a prefix-list from you and apply it outbound." Configure this on the router that will do the filtering work. both Send and receive on the same session. Common between peers who trust each other symmetrically. The pairing has to be complementary. If R1 says `send`, R3 must say `receive` (or `both`). Two routers that both say `send` negotiate nothing useful. ## Who actually uses this ORF is most valuable in exactly one relationship: **a customer taking a partial table from a provider**. The customer knows which prefixes they care about. The provider has hundreds of thousands they could send. Without ORF, the provider sends the lot and the customer's CE drops 95% of it. With ORF, the customer's filter runs on the PE and only the wanted routes cross the link. It also shows up between route reflectors and clients in very large iBGP meshes, where a client only needs a slice of the table. Where it does *not* show up: between peers of equal standing on an internet exchange. Nobody wants a stranger installing filters on their router, and the trust model does not fit. ## The lab AS 65001 (R1) is dual-attached to ISP-A (AS 100, R3) over two parallel links. Both links carry an independent eBGP session. That makes for a clean experiment: enable ORF on one session only, and the other session stays as an in-place control group. Before ORF, both sessions deliver three prefixes each: ``` R1#show ip bgp summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 2.2.2.2 4 65001 8 10 16 0 0 00:02:51 3 10.0.13.2 4 100 9 11 16 0 0 00:03:27 3 10.0.113.2 4 100 9 11 16 0 0 00:03:20 3 ``` ## Configuration R1 (the filter owner) wants only 6.6.60.0/24 from ISP-A. It builds a normal inbound prefix-list, applies it inbound as usual, and *additionally* offers to send it: ``` ip prefix-list ORF-IN seq 5 permit 6.6.60.0/24 ip prefix-list ORF-IN seq 10 deny 0.0.0.0/0 le 32 router bgp 65001 address-family ipv4 neighbor 10.0.13.2 capability orf prefix-list send neighbor 10.0.13.2 prefix-list ORF-IN in ``` R3 (the provider) agrees to receive and enforce it: ``` router bgp 100 address-family ipv4 neighbor 10.0.13.1 capability orf prefix-list receive ``` Two things matter in that config and both trip people up: - **You still apply the prefix-list inbound.** The `prefix-list ORF-IN in` line is not optional. ORF sends whatever prefix-list is applied inbound to that neighbour. No inbound prefix-list, nothing to send. - **The explicit deny is deliberate.** A prefix-list has an implicit deny anyway, but ORF encodes the list you wrote. Being explicit removes ambiguity about what the neighbour is being asked to enforce. ### Capability changes require a session reset ORF is a capability, and capabilities are negotiated in the OPEN message. Adding `capability orf` to an established session does nothing until the session restarts. In the lab, the session bounced on its own the moment the capability was added. In production, plan for that: `clear ip bgp ` is a hard reset and it will drop traffic. Once the capability is up, subsequent *filter* changes do not need a hard reset. You refresh the filter with: ``` R1#clear ip bgp 10.0.13.2 soft in prefix-filter ``` That pushes the current prefix-list across without tearing down the session. This is the operational payoff: change your filter, push it, and the neighbour immediately stops sending you the routes you no longer want. ## Verification The single most satisfying command in all of ORF, run on the *receiving* router, shows you exactly what filter the neighbour handed over: ``` R3#show ip bgp neighbors 10.0.13.1 received prefix-filter Address family: IPv4 Unicast ip prefix-list 10.0.13.1: 2 entries seq 5 permit 6.6.60.0/24 seq 10 deny 0.0.0.0/0 le 32 ``` That is R1's prefix-list, verbatim, living inside R3\. R3 names it after the neighbour that sent it. The capability negotiation itself is visible on both ends. On R3, which is doing the filtering: ``` R3#show ip bgp neighbors 10.0.13.1 | section Outbound Outbound Route Filter (ORF) type (128) Prefix-list: Capability code received: IETF standard and pre-standard Send-mode: received Receive-mode: advertised Outbound Route Filter (ORF): received (2 entries) ``` And on R1, which owns the filter: ``` R1#show ip bgp neighbors 10.0.13.2 | section Outbound Outbound Route Filter (ORF) type (128) Prefix-list: Capability code received: IETF standard and pre-standard Send-mode: advertised Receive-mode: received Outbound Route Filter (ORF): sent; ``` Read those carefully. `Send-mode: advertised` on R1 means "I told you I would send"; `Send-mode: received` on R3 means "you told me you would send". The mirror pairing is the sign of a healthy negotiation. If one side shows nothing, the capability did not negotiate and you are silently running without ORF. **Note the "IETF standard and pre-standard" line.** Cisco implemented ORF before the RFC was finalised, so IOS advertises both the pre-standard capability code (128) and the standard one. This is why ORF between Cisco and a non-Cisco peer sometimes needs a nudge, and why you should always verify the capability rather than assuming. ### The result The ORF-enabled session now carries one prefix. The parallel session to the same router, without ORF, still carries three: ``` R1#show ip bgp summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 2.2.2.2 4 65001 10 18 26 0 0 00:04:37 3 10.0.13.2 4 100 6 13 26 0 0 00:01:22 1 10.0.113.2 4 100 14 17 26 0 0 00:05:07 3 ``` That is the whole point in one screen. Same neighbour, same routes available, two sessions, and the ORF one only ever transmitted the prefix R1 asked for. ## ORF vs soft reconfiguration vs route refresh These three get conflated constantly. They solve different problems. soft-reconfiguration inbound **Where the cost lands:** your router **What it does:** stores every received route pre-policy so you can re-apply a changed inbound filter without a reset **Cost:** memory, and it grows with the table size Route refresh (RFC 2918) **Where the cost lands:** the link and the neighbour **What it does:** asks the neighbour to re-send everything so you can re-apply policy **Cost:** a full re-transmission of the table ORF **Where the cost lands:** the neighbour, once **What it does:** the unwanted routes are never sent in the first place **Cost:** negligible - and you save link, CPU, and memory Route refresh is negotiated by default on modern IOS and you almost never need soft-reconfiguration inbound anymore, except when you specifically want to *see* what a neighbour sent you before your policy chewed it (which is a superb troubleshooting trick - see the [expert BGP troubleshooting article](https://www.pinglabz.com/expert-bgp-troubleshooting-scenarios/)). ORF is the only one of the three that reduces the traffic on the wire. ## Limitations worth knowing - **Prefix-list ORF only.** The RFC allows other ORF types (community, AS-path, extended community) but Cisco IOS implements the prefix-list type. You cannot push "only send me routes tagged 65001:100" via ORF. - **It is a per-address-family thing.** Configure it under each AF you want it on. An IPv4 ORF does nothing for your IPv6 session. - **Trust required.** The receiving router is executing a policy it did not write. Providers who support ORF do so for customers, under a contract. Do not expect a random peer to enable `receive` for you. - **Not a security control.** ORF reduces what you are sent. It does not stop a malicious neighbour from sending you whatever they like. Keep your inbound prefix-list applied regardless - and it must stay applied anyway, because that is the list ORF sends. ## Key takeaways - ORF moves your inbound filter onto your neighbour's outbound policy, so unwanted routes are never transmitted. - `capability orf prefix-list send` on the filter owner, `receive` on the router doing the work. The modes must be complementary. - The inbound prefix-list must still be applied to the neighbour - that is the list ORF transmits. - Enabling the capability resets the session. Changing the filter afterwards does not: use `clear ip bgp soft in prefix-filter`. - `show ip bgp neighbors x.x.x.x received prefix-filter` on the receiving router proves the filter arrived. It is the first command to run when ORF "is not working". - The real use case is customer-to-provider partial tables. It is a cost and scale tool, not a security tool. Next: [BGP route dampening](https://www.pinglabz.com/bgp-route-dampening/) and why a feature designed to protect the internet ended up doing more harm than good. The full cluster index lives on the [BGP pillar guide](https://www.pinglabz.com/bgp/). ### BGP Conditional Advertisement: advertise-map, exist-map, and non-exist-map URL: https://www.pinglabz.com/bgp-conditional-advertisement/ Last updated: 2026-07-12T06:32:01.000Z Every dual-homed network eventually asks the same question: how do I keep a backup path *truly* idle until I need it? Local preference and AS-path prepending only change which path is *preferred*. The backup prefix is still out there, still in the global table, still attracting traffic if someone upstream has a different policy. Conditional advertisement is the answer: it withholds a prefix entirely until a condition you define is no longer true. This is a core CCIE Enterprise Infrastructure skill and one of the few BGP features where the configuration reads almost like English but the behaviour surprises people. This article walks through `advertise-map`, `exist-map`, and `non-exist-map` on Cisco IOS XE, with real output from a six-router CML lab. If you need the fundamentals first, start with the [complete BGP guide](https://www.pinglabz.com/bgp/). ## The problem conditional advertisement solves Picture an enterprise, AS 65001, dual-homed to two providers. ISP-A is the primary transit: fast, cheap, well-peered. ISP-B is a backup circuit that bills by the megabit. You have a prefix, 172.16.1.0/24, that you only want reachable over ISP-B if ISP-A is gone. The usual tools all fail here in the same way: - **AS-path prepending toward ISP-B** makes the path less attractive, but it does not make it invisible. Any network closer to ISP-B than to ISP-A still comes in that way. - **Setting a low local preference** only affects your own inbound path selection. It says nothing about how the internet reaches you. - **Communities like no-export** stop propagation past your neighbour, which is often too blunt. What you actually want is: *do not advertise 172.16.1.0/24 to ISP-B at all, unless ISP-A has failed*. That is conditional advertisement. ## How it works: two route-maps, one decision Conditional advertisement attaches two route-maps to a neighbour: advertise-map The prefixes whose advertisement you want to control. Match them with a prefix-list. These prefixes are advertised *only* when the condition is satisfied. non-exist-map The tracked prefix. While it **is present** in the BGP table, the advertise-map prefixes are withheld. When it **disappears**, they are advertised. This is the backup-path pattern. exist-map The mirror image. While the tracked prefix **is present**, the advertise-map prefixes are advertised. When it disappears, they are withdrawn. Used for "only advertise X if I can still reach Y". You use one or the other, never both, on the same neighbour. The overwhelmingly common case is `non-exist-map`, because the overwhelmingly common requirement is a backup path. ### The tracked prefix must be a real signal This is where most designs go wrong. The prefix you track in the non-exist-map must be something that **only exists while the primary path is healthy**. If you track a prefix that you can also learn through the backup provider, the condition never clears and your backup never activates. In the lab below, AS 65001 tracks 100.100.100.0/24, which is originated by ISP-A itself and is not reachable through ISP-B. When the ISP-A links go down, that prefix leaves the AS 65001 BGP table completely, and the condition fires. If instead we had tracked a customer prefix from a distant AS reachable through both providers, the condition would never have been met. ## The lab Six IOL-XE routers on Cisco Modeling Labs. AS 65001 is R1 and R2 (iBGP over OSPF). R1 is dual-attached to ISP-A (AS 100, routers R3 and R4). R2 is the ISP-B edge, peering with R5 in AS 200\. AS 300 (R6) sits beyond both providers. R1 originates 172.16.1.0/24 (the backup prefix) and carries it to R2 over iBGP. R2 owns the eBGP session to ISP-B, so R2 is where the conditional advertisement lives. ### Baseline: everything is advertised Before any policy, R2 hands ISP-B the full set of AS 65001 prefixes, including the backup: ``` R2#show ip bgp neighbors 10.0.25.2 advertised-routes Network Next Hop Metric LocPrf Weight Path r>i 1.1.1.1/32 1.1.1.1 0 100 0 i *>i 10.10.10.0/24 1.1.1.1 0 100 0 i *>i 10.10.10.66/32 1.1.1.1 0 100 0 i *>i 10.30.30.0/24 1.1.1.1 0 100 0 i *>i 100.100.100.0/24 1.1.1.1 0 200 0 100 i *>i 172.16.1.0/24 1.1.1.1 0 100 0 i Total number of prefixes 6 ``` That 172.16.1.0/24 at the bottom is the problem. ISP-B is announcing it to the world right now. ## Configuration ``` ! The prefix we want to control ip prefix-list BACKUP-PFX seq 5 permit 172.16.1.0/24 ! The prefix that proves ISP-A is alive ip prefix-list PRIMARY-WATCH seq 5 permit 100.100.100.0/24 route-map ADV-BACKUP permit 10 match ip address prefix-list BACKUP-PFX route-map TRACK-PRIMARY permit 10 match ip address prefix-list PRIMARY-WATCH router bgp 65001 address-family ipv4 neighbor 10.0.25.2 advertise-map ADV-BACKUP non-exist-map TRACK-PRIMARY ``` Read it out loud and it is exactly the requirement: *advertise the routes matched by ADV-BACKUP, but only when the routes matched by TRACK-PRIMARY do not exist.* ## Verification: the steady state The status field on the neighbour is the single most useful command here. It tells you what the router has actually decided: ``` R2#show ip bgp neighbors 10.0.25.2 | include Condition-map Condition-map TRACK-PRIMARY, Advertise-map ADV-BACKUP, status: Withdraw ``` `status: Withdraw` means the tracked prefix exists, so the backup is being held back. And the advertised-routes list confirms it - five prefixes now, and 172.16.1.0/24 is not among them: ``` R2#show ip bgp neighbors 10.0.25.2 advertised-routes Network Next Hop Metric LocPrf Weight Path r>i 1.1.1.1/32 1.1.1.1 0 100 0 i *>i 10.10.10.0/24 1.1.1.1 0 100 0 i *>i 10.10.10.66/32 1.1.1.1 0 100 0 i *>i 10.30.30.0/24 1.1.1.1 0 100 0 i *>i 100.100.100.0/24 1.1.1.1 0 200 0 100 i Total number of prefixes 5 ``` Note what conditional advertisement did *not* do: everything else is still advertised normally. The advertise-map only governs the prefixes it matches. Everything outside it follows your ordinary outbound policy. ## Verification: the failover Now shut both of R1's links to ISP-A. 100.100.100.0/24 leaves the AS 65001 BGP table entirely: ``` R2#show ip bgp | include 100.100|172.16|Network Network Next Hop Metric LocPrf Weight Path *>i 172.16.1.0/24 1.1.1.1 0 100 0 i ``` Wait about a minute, and the condition flips: ``` R2#show ip bgp neighbors 10.0.25.2 | include Condition-map Condition-map TRACK-PRIMARY, Advertise-map ADV-BACKUP, status: Advertise ``` And ISP-B now has the backup prefix: ``` R5#show ip bgp 172.16.1.0/24 BGP routing table entry for 172.16.1.0/24, version 28 Paths: (1 available, best #1, table default) 65001 10.0.25.1 from 10.0.25.1 (2.2.2.2) Origin IGP, localpref 100, valid, external, best ``` Bring the ISP-A links back, and within a scan cycle the status returns to `Withdraw` and ISP-B stops hearing the prefix. ## The gotcha nobody warns you about: it is not instant Conditional advertisement is evaluated by the BGP scanner, not by the update process. On IOS XE that scan runs on a 60-second interval by default. In the lab above the tracked prefix disappeared immediately, but the status field stayed on `Withdraw` for the better part of a minute before flipping. This matters enormously for design. Conditional advertisement is **not a fast-convergence mechanism**. If your requirement is "traffic must move within 500 ms", conditional advertisement is the wrong tool and you want BFD plus a pre-installed backup path. Conditional advertisement is for policy-level failover measured in tens of seconds: "if the primary transit is genuinely gone, start announcing the DR prefix." You can see the scan interval in `show ip bgp summary` (`scan interval 60 secs`). Tuning it down is possible with `bgp scan-time`, but on a router carrying a real internet table you are trading CPU for convergence, and that trade is rarely worth it. ## exist-map: the other direction Swap `non-exist-map` for `exist-map` and the logic inverts. Now the advertise-map prefixes are announced *while* the tracked prefix is present, and withdrawn when it goes away. The classic use case is a multi-site enterprise where a site's aggregate should only be advertised while that site is actually reachable. Site B's edge router advertises 10.20.0.0/16 to the provider only while it can see a specific host route or loopback inside Site B. Lose the site, stop advertising the aggregate, and traffic naturally shifts to Site A rather than being black-holed at a router that no longer has a path. ``` route-map ADV-SITE-B permit 10 match ip address prefix-list SITE-B-AGGREGATE route-map SITE-B-ALIVE permit 10 match ip address prefix-list SITE-B-CORE-LOOPBACK router bgp 65001 address-family ipv4 neighbor 203.0.113.1 advertise-map ADV-SITE-B exist-map SITE-B-ALIVE ``` ## Conditional advertisement vs the alternatives Conditional advertisement **Controls:** whether the prefix is announced at all **Speed:** up to 60s (scanner) **Use for:** true backup prefixes, DR sites AS-path prepend **Controls:** how attractive the path looks **Speed:** immediate **Use for:** nudging inbound traffic, not hiding a prefix Communities (no-export etc.) **Controls:** how far the prefix propagates **Speed:** immediate **Use for:** scope limits, provider-signalled policy Outbound prefix-list **Controls:** whether the prefix is announced at all **Speed:** static - no failover **Use for:** permanent filtering, not conditional behaviour ## Troubleshooting checklist 1. **Status stuck on Withdraw when the primary is down?** Your tracked prefix is still in the BGP table. Run `show ip bgp ` and look at where it came from. Nine times out of ten it is arriving through the backup provider, which defeats the whole design. 2. **Status flipped to Advertise but the neighbour still does not have the prefix?** Check that the prefix is actually in your BGP table and is a best path. Conditional advertisement can only announce something you have. 3. **Nothing at all being advertised?** Confirm the advertise-map route-map has a matching permit clause. A route-map with a `match` that never matches falls through to the implicit deny, and you get an empty advertisement set. 4. **It works but takes too long?** That is the 60-second scanner, and it is by design. See above. 5. **Condition flapping?** Track a stable prefix. A tracked prefix that itself flaps will drive your backup announcement in and out of the global table, which is much worse than never having configured it. ## Key takeaways - Conditional advertisement is the only BGP mechanism that decides *whether* a prefix is announced based on live routing state. - `non-exist-map` is the backup-path pattern: withhold the prefix while the tracked route exists, advertise it when the tracked route disappears. `exist-map` is the mirror image. - The tracked prefix must be uniquely reachable through the primary path. If the backup provider can also deliver it, your condition will never fire. - `show ip bgp neighbors x.x.x.x | include Condition-map` gives you the router's own verdict: `status: Withdraw` or `status: Advertise`. Start every troubleshoot there. - It is evaluated by the BGP scanner, so expect up to 60 seconds of lag. It is a policy mechanism, not a convergence mechanism. Next in the expert BGP series: [Outbound Route Filtering (ORF)](https://www.pinglabz.com/bgp-outbound-route-filtering-orf/), which pushes your inbound filter across the session so your neighbour stops sending you routes you were only going to discard. Everything in this cluster hangs off the [BGP pillar guide](https://www.pinglabz.com/bgp/). ### On-Prem vs Cloud Network Design: What ENCOR Wants You to Know URL: https://www.pinglabz.com/on-prem-vs-cloud-network-design/ Last updated: 2026-07-12T03:19:44.000Z The workloads a network engineer connects have moved. Some are still in a data centre you can walk into; many are in a public cloud you will never see; most enterprises are somewhere in between, running both at once. Designing for that reality means understanding what changes when the servers are someone else's, what stays the same, and where the new failure modes and cost traps hide. The ENCOR blueprint asks you to reason about this, not to become a cloud architect. This article covers on-prem versus cloud network design for network engineers. It extends the [QoS cluster guide](https://www.pinglabz.com/qos/) and the campus [design article](https://www.pinglabz.com/enterprise-campus-design-ccnp/). ## What Actually Changes in the Cloud The fundamentals of IP, routing, and segmentation do not change; a subnet is a subnet. What changes is who operates the underlay and how you express intent: You do not own the underlay No physical switches, no cabling, no spanning tree. The cloud provider runs a massive underlay; you get a virtual network on top. The network is software-defined You define a VPC/VNet, subnets, route tables, and security groups through an API or console, not a CLI on a device. Security is distributed Security groups attach to instances, not chokepoints. Filtering is everywhere, stateful, and instance-level rather than at a firewall. Everything is elastic and metered You can create a hundred subnets in seconds, and you pay for what you use, including, crucially, data transfer. The mental shift: on-prem, you build and own the whole stack, cables to config. In the cloud, you consume a virtual network the provider operates, and you express your design as software (route tables, security groups) against their infrastructure. Your job moves from building the network to designing and governing it. ## The Cloud Building Blocks (In Familiar Terms) Every cloud provider's networking maps onto concepts you already know, with different names: VPC / VNetYour private virtual network in the cloud. Like a VRF or a routing domain you own, isolated from other tenants. SubnetExactly what it sounds like, but tied to an availability zone. Placement matters for resilience. Route tablePer-subnet static routing. You point subnets at gateways, peers, and the internet. No dynamic IGP by default. Security groupA stateful firewall attached to an instance. The primary segmentation tool. Like a per-instance ACL that follows the workload. Gateways / peeringInternet gateway, NAT gateway, transit gateway, VPC peering, the on-ramps and interconnects between networks. The security group is the one that most changes how you think. On-prem, you filter at chokepoints (a firewall between zones). In the cloud, the filter attaches to the workload itself and moves with it, so segmentation is expressed as "which instances can talk to which," everywhere, rather than "what passes through this firewall." It is microsegmentation as the default, and it is stateful, so you allow the request and the return is automatic. ## The Hybrid Reality: Connecting the Two Worlds Almost no enterprise is purely one or the other. The interesting design work is the interconnect, and there are two main ways to bridge on-prem and cloud: Site-to-site VPN An IPsec tunnel over the internet from your edge to the cloud VPN gateway. Quick, cheap, but rides the public internet (variable latency, best-effort). Direct interconnect A private, dedicated circuit to the provider (Direct Connect, ExpressRoute). Predictable latency and bandwidth, higher cost and lead time. The choice mirrors the classic [MPLS-vs-internet](https://www.pinglabz.com/sd-wan/) trade-off: the VPN is fast to deploy and cheap but best-effort; the direct interconnect is predictable and performant but costs more and takes weeks to provision. Many designs use both, a direct interconnect for production with a VPN as backup. And routing between the two worlds is typically BGP over the interconnect, so your [BGP](https://www.pinglabz.com/bgp/) knowledge transfers directly. ## The Traps That Catch On-Prem Engineers The failure modes and cost surprises that differ from on-prem thinking: - **Data transfer costs.** This is the big one. Moving data *out* of the cloud (egress) and between availability zones costs money, per gigabyte. A design that would be free on-prem (chatty cross-zone traffic) can generate a large bill. Network design in the cloud is partly cost design, and this has no on-prem equivalent. - **No broadcast or multicast.** Cloud virtual networks generally do not support broadcast or multicast the way a physical LAN does. Applications and protocols that assume them (some clustering, some discovery) break or need workarounds. - **Availability zones are your redundancy unit.** Resilience in the cloud means spreading across availability zones, not redundant cables. A subnet lives in one AZ; designing for failure means multi-AZ placement, which is a different discipline from redundant links. - **The shared responsibility model.** The provider secures the underlay; you secure your configuration (security groups, route tables, access). A misconfigured security group is your fault and your breach, not the provider's. Cloud security failures are overwhelmingly customer misconfigurations. - **Static routing by default.** Route tables are static. Dynamic routing exists (via virtual routers or the interconnect BGP) but the intra-VPC model is static route tables, which surprises engineers expecting an IGP. ## What Stays the Same It is worth ending on reassurance, because the cloud can feel alien. The fundamentals are unchanged: IP addressing and subnetting work identically, routing logic (longest match, next-hop) is the same, segmentation is still about controlling who talks to whom, BGP is still BGP over the interconnect, and the design goals (redundancy, segmentation, performance, cost) are the same goals in different clothing. A network engineer's core knowledge transfers directly; what you learn is a new set of building blocks and a new bill to watch. The instinct that made you good at on-prem design, thinking in terms of failure domains, segmentation, and predictable structure, is exactly what makes you good at cloud design. ## FAQ ### What is a VPC? A Virtual Private Cloud (VNet in Azure): your isolated virtual network in the cloud, containing subnets, route tables, and security groups. Conceptually like a routing domain or VRF that you own. ### How is cloud security different? Security groups attach to instances and move with the workload, so filtering is stateful and everywhere, rather than at chokepoint firewalls. It is microsegmentation by default. And it is a shared responsibility: the provider secures the underlay, you secure your config. ### VPN or direct interconnect for hybrid? VPN (IPsec over the internet) is quick and cheap but best-effort. Direct interconnect (Direct Connect/ExpressRoute) is predictable and performant but costs more and takes longer to provision. Many use both, interconnect for production, VPN as backup. ### What is the biggest cost surprise? Data transfer, especially egress out of the cloud and traffic between availability zones. It is metered per gigabyte and can be large. Cloud network design is partly cost design, with no on-prem equivalent. ### Does my networking knowledge transfer? Yes, directly. IP, subnetting, routing logic, segmentation, and BGP all work the same. You learn new building blocks (VPC, security group, route table) and a new cost model, but the core design instincts are exactly what cloud design needs. ## Key Takeaways - In the cloud you **consume a virtual network the provider operates**, expressed as software (route tables, security groups) rather than built from cables and CLI. - The building blocks map onto familiar concepts: **VPC/VNet** (your network), **subnet** (tied to an AZ), **route table** (static routing), **security group** (stateful per-instance firewall). - Security is **distributed and instance-level** (microsegmentation by default), under a **shared responsibility model**, most cloud breaches are customer misconfigurations. - Hybrid connectivity is **VPN** (cheap, best-effort) or **direct interconnect** (predictable, costly), the same trade-off as MPLS vs internet, usually with BGP routing. - The traps: **data-transfer costs** (the big one), no broadcast/multicast, availability zones as the redundancy unit, and static route tables by default. - The **fundamentals are unchanged**: IP, routing, segmentation, and BGP transfer directly. Your on-prem design instincts are what make you good at cloud design. This closes the QoS and Architecture cluster. Back to the [QoS cluster guide](https://www.pinglabz.com/qos/) for the full reading order. ### High Availability Techniques: SSO, NSF, and Graceful Restart URL: https://www.pinglabz.com/sso-nsf-graceful-restart/ Last updated: 2026-07-12T03:19:44.000Z A core router or switch reloading should not take the network down. In a well-designed high-availability platform, a supervisor can fail, or software can be upgraded, and traffic keeps flowing with barely a ripple. The technologies that make that possible, SSO, NSF, and graceful restart, work together but solve different halves of the problem, and understanding the division of labour is the whole topic. This article covers high-availability techniques at CCNP depth. It extends the [QoS cluster guide](https://www.pinglabz.com/qos/) and the campus [design article](https://www.pinglabz.com/enterprise-campus-design-ccnp/). ## A Note on the Lab Keeping to this site's honesty rule: SSO (Stateful Switchover) requires a device with *two* supervisor modules, so one can take over from the other. That is a hardware-redundancy feature that a single-supervisor virtual router cannot demonstrate, and the IOL platform in this lab has one supervisor. So the SSO mechanics below are explained conceptually. The routing-protocol side, NSF and graceful restart, *is* partly observable on IOL, and the real capture of OSPF's NSF helper capability is shown below. No switchover output is fabricated. ## The Two Halves of the Problem When a supervisor fails and a backup takes over, two things must survive for traffic to keep flowing: The forwarding plane (inside the box) The new supervisor must keep forwarding packets without rebuilding its tables from scratch. This is what **SSO** and **NSF** handle. The routing relationship (with neighbors) The neighbors must not tear down their adjacencies and reroute around the box while it recovers. This is what **graceful restart** handles. Solve only the first and your box keeps forwarding but its OSPF neighbors declare it dead and reroute, causing churn. Solve only the second and the neighbors wait patiently while your box has actually stopped forwarding. You need both, and they are different mechanisms with different names. ## SSO: Keeping the Standby Supervisor Ready **Stateful Switchover** is the hardware-redundancy foundation. In a dual-supervisor chassis, one supervisor is active and one is standby (hot). SSO continuously synchronises state from the active to the standby, the configuration, the interface state, and critically the Layer 2 protocol state, so that if the active fails, the standby takes over already knowing where everything is. The key thing SSO preserves is the **forwarding information**. The CEF/FIB (the hardware forwarding table) is maintained on the standby, so at the moment of switchover, the new active supervisor can forward packets immediately using the tables that were already there. It does not have to rebuild them. That is what makes the switchover fast enough that traffic barely notices. What SSO does *not* preserve on its own is the control-plane routing protocol state, the OSPF adjacencies, the BGP sessions. Those are held by the supervisor's CPU, and during the switchover the routing process restarts. Without help, the routing protocols would flap. That is where NSF comes in. ## NSF: Forward While the Control Plane Recovers **Nonstop Forwarding** is the bridge between SSO and the routing protocols. The insight is elegant: because SSO preserved the forwarding table (the FIB), the box can *keep forwarding traffic using the old FIB* while the control plane (the routing protocols) restarts and rebuilds in the background. The data plane does not wait for the control plane. So during a switchover: SSO hands over to the standby supervisor, which forwards packets using the retained FIB (NSF), while the routing protocols restart, re-establish, and eventually recompute the tables. As long as the topology has not actually changed during the brief switchover, the retained FIB is still correct, and traffic flows the entire time. NSF is what turns "the forwarding table survived" (SSO) into "traffic keeps flowing" during the control-plane rebuild. ## Graceful Restart: Convincing the Neighbors to Wait Here is the remaining problem. Your box is forwarding fine via NSF, but its OSPF neighbors have noticed the routing process restart. Normally, a neighbor that sees an adjacency drop declares the router dead and reroutes around it, exactly the churn you were trying to avoid. Graceful restart (also called NSF-awareness on the neighbor side) is the agreement that prevents this. The restarting router signals its neighbors: "I am restarting my control plane, but I am still forwarding. Please keep our adjacency up and keep sending me traffic; do not reroute." A graceful-restart-aware neighbor honours this, holding the adjacency and continuing to forward toward the restarting router for a grace period, giving it time to rebuild. This is a cooperative protocol: the restarting router needs its neighbors to be **NSF-aware** (also called GR helpers) for it to work. And this is the part that *is* observable on the lab platform. Here is OSPF's real graceful-restart helper capability, captured live: ``` R1#show ip ospf | include NSF|Graceful|helper IETF NSF helper support enabled Cisco NSF helper support enabled Graceful Reload FSU Global status : None (global: None) ``` Both the IETF and Cisco flavours of NSF helper support are enabled, meaning this router will act as a graceful-restart *helper* for a restarting neighbor: it will hold the adjacency and keep forwarding while that neighbor recovers, rather than tearing down and rerouting. This is the neighbor-side half of the mechanism, and it is real. The restarting-side (the box with dual supervisors actually doing the switchover) is what needs the hardware SSO/NSF this virtual platform lacks. ## The Three Working Together The full sequence of a graceful supervisor failover: 1\. SSOStandby supervisor takes over with the synchronised FIB and L2 state already in place. 2\. NSFThe box keeps forwarding using the retained FIB while the routing protocols restart. 3\. Graceful restartNSF-aware neighbors hold their adjacencies and keep sending traffic, instead of rerouting around the box. The result: a supervisor fails (or you upgrade its software), and traffic through the box continues with sub-second or no interruption, because the forwarding never stopped and the neighbors never rerouted. That is the goal of the whole HA stack, and it takes all three pieces cooperating. ## Where HA Fits in the Design These techniques matter most exactly where the [campus design](https://www.pinglabz.com/enterprise-campus-design-ccnp/) demands no single point of failure: the core and distribution layers, where a device reload would otherwise be disruptive. They complement, rather than replace, topological redundancy (redundant devices and links). A dual-supervisor core switch with SSO/NSF handles a supervisor failure gracefully; a pair of core switches handles a whole-device failure. You want both: the box survives a component failure with SSO/NSF, and the topology survives a box failure with redundancy. For faster failure *detection* to complement this, see [BFD](https://www.pinglabz.com/bfd-bidirectional-forwarding-detection/). ## FAQ ### What is the difference between SSO and NSF? SSO synchronises state to a standby supervisor so it can take over with the forwarding table intact. NSF uses that retained forwarding table to keep forwarding packets while the routing protocols restart. SSO preserves the tables; NSF keeps forwarding with them. ### What does graceful restart do? It convinces the restarting router's neighbors to hold their adjacencies and keep forwarding toward it, rather than declaring it dead and rerouting. It is the neighbor-cooperation half of the HA story. ### Can I lab SSO? Not on a single-supervisor virtual router. SSO needs two physical supervisor modules. The neighbor-side graceful-restart helper capability (NSF-awareness) is observable, as shown in the real OSPF capture. ### Do my neighbors need to support anything? Yes. Graceful restart is cooperative: the neighbors must be NSF-aware (GR helpers) to hold their adjacencies while your box restarts. If they are not, they will reroute despite your NSF. ### Does this replace having redundant devices? No, it complements it. SSO/NSF handles a component (supervisor) failure within a box; topological redundancy (paired devices, redundant links) handles a whole-device failure. Design for both. ## Key Takeaways - Surviving a supervisor failover needs two things: keep **forwarding** (inside the box) and keep the **routing relationships** (with neighbors). Different mechanisms. - **SSO** synchronises state to a standby supervisor so it takes over with the FIB and L2 state intact. - **NSF** uses that retained FIB to keep forwarding packets while the routing protocols restart in the background. - **Graceful restart** convinces NSF-aware neighbors to hold their adjacencies and keep sending traffic instead of rerouting around the box. - The real capture: OSPF `IETF NSF helper support enabled` / `Cisco NSF helper support enabled`, the neighbor-side (helper) half, which IOL can show. SSO itself needs dual supervisors and is covered conceptually. - HA complements, not replaces, topological redundancy. The box survives a component failure with SSO/NSF; the topology survives a box failure with redundant devices and links. Next: [On-prem vs cloud network design](https://www.pinglabz.com/on-prem-vs-cloud-network-design/), or the [QoS cluster guide](https://www.pinglabz.com/qos/). ### Enterprise Campus Design: 2-Tier vs 3-Tier at CCNP Depth URL: https://www.pinglabz.com/enterprise-campus-design-ccnp/ Last updated: 2026-07-12T03:19:44.000Z Every campus network is some answer to the same question: how do you connect thousands of users to each other and the outside world, reliably, at a cost you can justify, in a way you can grow? The classic answers are the two-tier and three-tier hierarchical designs, and the modern answer is the spine-leaf fabric. Knowing which to reach for, and why, is a CCNP-level design skill that shows up throughout the ENCOR blueprint. This article covers enterprise campus design at CCNP depth. It extends the [QoS cluster guide](https://www.pinglabz.com/qos/) (design and QoS are the two architecture-domain topics) and connects to the [Network Virtualization cluster](https://www.pinglabz.com/network-virtualization/) for the fabric direction. ## The Three-Tier Model and Its Layers The classic hierarchical design has three layers, each with a distinct job. The whole point of the model is that separating these roles makes the network scalable, predictable, and easy to troubleshoot: Access layer Where users and devices connect. Port security, 802.1X, PoE, the first-hop gateway. High port density, low cost per port. Distribution layer Aggregates the access layer. The boundary between Layer 2 and Layer 3, where routing, policy, and redundancy (FHRP) live. Core layer The high-speed backbone. Its only job is to switch packets between distribution blocks as fast as possible. No policy, no complexity, just speed and reliability. The design discipline that matters: **the core does nothing but forward fast**. You do not put access lists, policy, or anything that adds latency or complexity in the core, because everything depends on it and it must never be the bottleneck or the thing that breaks. Policy and intelligence live at the distribution layer; the access layer connects users; the core just moves packets. Keeping those roles clean is what makes the network operable. ## Two-Tier: The Collapsed Core Not every network needs three tiers. In a smaller campus, a full dedicated core is expensive overkill, so you **collapse the core and distribution into one layer**. This is the two-tier (collapsed core) design: access switches connect to a pair of collapsed core/distribution switches that do both jobs. Three-tierAccess + distribution + core. Scales to large campuses; the core isolates distribution blocks. More devices, more cost. Two-tier (collapsed core)Access + collapsed core/distribution. Right for small-to-medium campuses. Fewer devices, lower cost, less isolation between blocks. The decision rule is about scale and the number of distribution blocks. If you have one or two buildings and a handful of distribution switches, a dedicated core adds cost and hops for no benefit, so collapse it. When you have many distribution blocks (multiple buildings, thousands of users), a dedicated core is worth it: it gives every distribution block a single, simple, high-speed point to reach every other block, without a full mesh of distribution-to-distribution links. The core exists to avoid that mesh. ## The Layer 2 / Layer 3 Boundary The most consequential design decision in a campus is where Layer 2 ends and Layer 3 begins, because it determines your failure domains, your spanning-tree footprint, and your convergence behaviour. Routed access Layer 3 starts at the access switch. No spanning tree beyond the access port, fast convergence via routing. The modern preference where the hardware supports it. L2 access, L3 distribution The traditional model. VLANs span access, the distribution layer routes. Needs STP and FHRP; more moving parts. Pushing Layer 3 down to the access layer (routed access) shrinks the spanning-tree domain to a single switch and lets routing (which converges faster and more predictably than STP) handle failures. The trade-off is that VLANs can no longer span multiple access switches, which some applications and wireless designs assume. The traditional Layer-2-access model keeps VLANs flexible but pays for it with [spanning tree](https://www.pinglabz.com/spanning-tree-protocol/) and [FHRP](https://www.pinglabz.com/fhrp/) complexity and slower convergence. The trend is toward routed access; the exam expects you to know both and their trade-offs. ## Designing for Failure A campus design is only as good as its behaviour when something breaks. The principles: - **No single points of failure at the aggregation and core.** Distribution and core switches come in pairs, with redundant links, so any single device or link failure is survivable. The access layer is often single-homed per user (a user's own switch failing only affects that user), but the layers above must be redundant. - **Redundant links, not just redundant devices.** Two core switches with a single link between the distribution and each is not truly redundant. Design the link topology so any single link loss has a path around it. - **Fast convergence.** Whatever fails, the network must reroute quickly. This is where routed access (fast IGP convergence), [BFD](https://www.pinglabz.com/bfd-bidirectional-forwarding-detection/) for sub-second detection, and the high-availability techniques ([SSO, NSF, graceful restart](https://www.pinglabz.com/sso-nsf-graceful-restart/)) all contribute. - **Consistent, predictable structure.** The reason hierarchy matters is not aesthetics; a predictable structure means you can reason about failure, capacity, and change. An ad-hoc topology cannot be reasoned about. ## The Modern Direction: Spine-Leaf and Fabric The three-tier model was designed for a traffic pattern that has changed. It assumes most traffic is north-south (users to the internet or data centre), which is why the core is a small fast backbone. Modern data centres and increasingly campuses see huge east-west traffic (server to server, service to service), and the three-tier model handles that poorly, because east-west traffic has to go up to the core and back down. The answer is **spine-leaf**: every leaf (access) connects to every spine (aggregation), so any leaf is exactly two hops from any other leaf, and east-west bandwidth scales by adding spines. Combined with a [VXLAN](https://www.pinglabz.com/vxlan-deep-dive/) overlay and [BGP EVPN](https://www.pinglabz.com/vxlan-bgp-evpn-explained/) control plane, this is the modern data-centre fabric, and in the campus it becomes [Cisco SD-Access](https://www.pinglabz.com/sd-access-architecture/). The hierarchical model has not disappeared, most campuses still run it, but the fabric direction is where new large-scale designs are heading, and the CCNP expects you to understand both and when each fits. ## FAQ ### Two-tier or three-tier? Two-tier (collapsed core) for small-to-medium campuses with few distribution blocks, where a dedicated core is overkill. Three-tier when you have many distribution blocks and need the core to avoid a distribution-to-distribution mesh. ### What should the core layer do? Only forward packets, as fast and reliably as possible. No ACLs, no policy, no complexity. Everything depends on the core, so it must never be the bottleneck or the fragile part. ### Where should the Layer 2 / Layer 3 boundary be? The modern preference is routed access (L3 at the access switch), which shrinks the STP domain and converges fast, at the cost of VLANs not spanning switches. The traditional model puts L3 at distribution, keeping VLANs flexible but needing STP and FHRP. ### Is spine-leaf replacing three-tier? For large-scale and east-west-heavy designs, increasingly yes, especially in the data centre and as SD-Access in the campus. But most campuses still run hierarchical designs, and the CCNP expects you to know both. ### How many core switches should I have? At least two, with redundant links, so no single device or link failure isolates a distribution block. Redundancy at the core and distribution is non-negotiable; the access layer is often single-homed per user. ## Key Takeaways - The three-tier model separates **access** (users connect), **distribution** (aggregation, L2/L3 boundary, policy), and **core** (fast backbone only). - **The core does nothing but forward fast.** No policy or complexity in the core; intelligence lives at distribution. - **Two-tier (collapsed core)** merges core and distribution for smaller campuses; **three-tier** scales to many distribution blocks. - The **Layer 2 / Layer 3 boundary** is the key decision: routed access (fast, small STP domain, VLANs do not span) vs L2-access/L3-distribution (flexible VLANs, needs STP and FHRP). - Design for failure: redundant devices *and* links at aggregation/core, fast convergence (routed access, BFD, SSO/NSF), and predictable structure. - **Spine-leaf** is the modern direction for east-west-heavy and large-scale designs (data-centre fabric, campus SD-Access), but hierarchical designs still dominate. Know both. Next: [High availability techniques (SSO, NSF, GR)](https://www.pinglabz.com/sso-nsf-graceful-restart/), or the [QoS cluster guide](https://www.pinglabz.com/qos/). ### Interpreting QoS Configurations: An ENCOR Exam Skill Guide URL: https://www.pinglabz.com/interpreting-qos-configurations-encor/ Last updated: 2026-07-12T03:19:43.000Z The ENCOR blueprint has a specific, easily-underestimated line: "interpret QoS configurations." Not design them, not configure them from scratch, but *read* an existing policy and a `show` output and say what it does and whether it is working. This is a distinct skill, and it is where a lot of otherwise-prepared candidates lose marks, because reading someone else's QoS policy under time pressure is harder than writing your own. This article is a method for interpreting QoS configurations and output, using the real policy from this cluster. It extends the [QoS cluster guide](https://www.pinglabz.com/qos/). ## Reading a policy-map: A Repeatable Method Given an unfamiliar policy-map, work through it in a fixed order and it stops being intimidating: 1. **List the classes.** Each `class` line is a category of traffic. How many are there, and is there a `class-default` (there always is, explicitly or implicitly)? 2. **For each class, find how it is matched.** Trace back to the class-map: what `match` statement selects the traffic (DSCP, ACL, protocol)? 3. **For each class, identify the action.** `priority`, `bandwidth`, `police`, `shape`, `set`, `random-detect`, each is a different behaviour with different implications. 4. **Check the numbers add up.** Do the priority and bandwidth percentages fit within the reservable bandwidth? Is anything over-allocated? Apply it to the cluster's policy: ``` policy-map WAN-EDGE class EF-VOICE priority percent 20 class BULK-DATA bandwidth percent 40 random-detect dscp-based class class-default fair-queue random-detect ``` Walk the method: **three classes** (EF-VOICE, BULK-DATA, class-default). **Matched by** DSCP EF, DSCP AF11, and everything-else respectively (from the class-maps). **Actions**: EF-VOICE gets strict-priority queueing (LLQ), BULK-DATA gets a 40% bandwidth guarantee plus WRED, class-default gets fair-queueing plus WRED. **Numbers**: 20% priority + 40% bandwidth = 60%, comfortably within the default 75% reservable. In four steps you have fully understood a policy you had never seen. ## Recognise the Action Verbs Instantly Interpretation speed comes from recognising each action's meaning without thinking: `priority`LLQ. Strict-priority, policed. Latency guarantee. Think voice. `bandwidth`CBWFQ. Guaranteed minimum share under congestion. No latency guarantee. `police`Drops/remarks traffic above a rate. A hard ceiling. No buffering. `shape`Buffers traffic above a rate, sends later. Smooths bursts, adds delay. `set`Marks the traffic (DSCP, CoS). Classification/marking, not queueing. `random-detect`WRED. Early random drops to avoid tail-drop synchronization. TCP queues only. Two pairs are the classic exam traps. **priority vs bandwidth**: a fast lane with a latency guarantee versus a reserved minimum with none. **police vs shape**: drop the excess versus buffer it. If you can tell each pair apart on sight, most QoS interpretation questions answer themselves. ## Reading show policy-map interface The other half of interpretation is reading the operational output to answer "is it working?" This is where the real signal lives. Here is the cluster's EF class under load: ``` Class-map: EF-VOICE (match-all) 1125 packets, 1673362 bytes 5 minute offered rate 37000 bps, drop rate 0000 bps Match: dscp ef (46) Priority: 20% (2000 kbps), burst bytes 50000, b/w exceed drops: 0 ``` What each line tells you: packets / bytesTraffic actually matched this class. Zero here means your classification is wrong (nothing is matching). offered rate / drop rateHow much is arriving and how much is being dropped. A non-zero drop rate is the headline signal. MatchConfirms what selects the class. Cross-check against what you expected to be classified here. b/w exceed dropsFor a priority class, how many packets exceeded the policer and were dropped. **0 is healthy**; a rising number means voice is exceeding its reservation. Reading this specific output: 1125 packets matched (classification works), 0 drop rate and `b/w exceed drops: 0` (the priority class is healthy and protected), matched on DSCP EF (as intended). Verdict: the policy is working and voice is protected. That is the exact kind of judgment the exam asks for. ## The Diagnostic Questions When a question shows you output and asks what is wrong, these are the tells: - **A class with 0 packets** that should have traffic: the classification is broken. The `match` statement does not match the real traffic (wrong DSCP, wrong ACL). This is the most common QoS fault and a favourite exam scenario. - **A non-zero drop rate on a class you expected to be protected**: either the class is over-subscribed (offered rate exceeds its allocation) or, for a priority class, traffic is exceeding the policer (`b/w exceed drops` climbing). - **Traffic landing in class-default** that you expected in a named class: again, classification. The traffic is not being marked or matched as intended and is falling through to the catch-all. - **The policy not applied at all**: `show policy-map interface` returns nothing for the interface. The `service-policy` is missing, or on the wrong interface/direction. Almost every "why is QoS not working" answer is one of two things: **the traffic is not classified as you think** (check packet counts and the Match line), or **the policy is not applied where the traffic egresses**. Check those two first, always. ## FAQ ### How do I read a policy-map I have never seen? Four steps: list the classes, find how each is matched (trace to the class-map), identify each action verb, and check the percentages add up within the reservable bandwidth. In order, it is mechanical. ### What is the fastest way to tell priority from bandwidth? `priority` is a policed fast lane with a latency guarantee (voice). `bandwidth` is a guaranteed minimum share with no latency guarantee (data). Fast lane vs reserved minimum. ### A class shows 0 packets. What does that mean? Nothing is matching it, so the classification is broken. The `match` statement does not match the real traffic. This is the single most common QoS fault. ### What does "b/w exceed drops" tell me? On a priority (LLQ) class, how many packets exceeded the policer and were dropped. Zero is healthy. A rising number means the priority traffic is exceeding its reservation and being policed. ### How do I know if the policy is even applied? `show policy-map interface `. If it returns nothing, the service-policy is missing or on the wrong interface or direction (remember: queueing is output). ## Key Takeaways - "Interpret QoS configurations" is a real, distinct exam skill: read an existing policy and output, not build one from scratch. - Read a policy-map in four steps: **list classes, find the match, identify the action verb, check the numbers**. - Recognise the verbs on sight. The traps: **priority vs bandwidth** (fast lane vs reserved minimum) and **police vs shape** (drop vs buffer). - In `show policy-map interface`, the key signals are **packet counts** (is it classifying?), **drop rate** (is it dropping?), and **b/w exceed drops** (is the priority class healthy? 0 = yes). - Most "QoS not working" answers are one of two things: **traffic not classified as expected** (0 packets, wrong Match) or **policy not applied where traffic egresses**. Next: [LLQ and CBWFQ in depth](https://www.pinglabz.com/llq-cbwfq-cisco-ios-xe/), or the [QoS cluster guide](https://www.pinglabz.com/qos/). ### Hierarchical QoS: Shaping Parents and Queueing Children URL: https://www.pinglabz.com/hierarchical-qos-shaping/ Last updated: 2026-07-12T03:19:43.000Z A single flat QoS policy works fine until reality intrudes: you have a 100 Mbps physical interface but the carrier only sold you 20 Mbps, or one physical link carries traffic for ten branch sites that each need their own guarantees. A flat policy cannot express "shape everything to 20 Mbps, and within that 20, give voice priority and data a floor." That two-level problem is exactly what Hierarchical QoS solves, by nesting one policy inside another. This article covers hierarchical QoS and the shaping that underpins it. It extends the [QoS cluster guide](https://www.pinglabz.com/qos/) and builds on [LLQ and CBWFQ](https://www.pinglabz.com/llq-cbwfq-cisco-ios-xe/). ## First: Shaping vs Policing Hierarchical QoS is built on shaping, so the shaping-versus-policing distinction has to be clear. Both limit traffic to a rate; they differ in what they do with the excess: Policing **Drops** (or remarks) traffic above the rate. Sharp, no buffering, no added delay. Bursty output. Used at the edge to enforce a hard ceiling. Shaping **Buffers** traffic above the rate and sends it later. Smooth output, adds delay, needs memory. Used to match a downstream rate without dropping. The analogy: policing is a bouncer who turns away anyone over capacity; shaping is a queue outside the club that lets people in at a steady rate. Policing is cheaper (no buffers) but wastes the excess traffic; shaping preserves it but costs delay and memory. For the classic carrier-mismatch problem, you want **shaping**, because dropping traffic your own router could have buffered and sent a moment later is needlessly lossy. ## The Problem HQoS Solves The canonical case: you buy a 20 Mbps WAN service, but it is delivered over a 100 Mbps physical handoff (Metro Ethernet, for example). Your router's interface is 100 Mbps, so it will happily send bursts at 100 Mbps, and the carrier, seeing traffic above the 20 Mbps you paid for, will *police* (drop) the excess. Worse, the carrier's drop is indiscriminate, it does not know your voice from your bulk data, so your carefully prioritised voice gets dropped along with everything else. The fix is to **shape your own traffic to 20 Mbps first**, so it never exceeds what the carrier will accept, and then apply your queueing policy *within* that shaped 20 Mbps, so your voice gets priority inside the rate the carrier honours. That is two levels of policy: shape (parent), then queue (child). And a single flat policy cannot do it, because a flat policy's percentages are of the physical 100 Mbps, not the 20 Mbps you actually have. ## Parent and Child HQoS nests a child policy inside a parent policy. The parent shapes; the child queues within the shaped rate: ``` ! CHILD: the queueing policy (LLQ + CBWFQ), unchanged policy-map CHILD-QUEUEING class EF-VOICE priority percent 20 class BULK-DATA bandwidth percent 40 random-detect dscp-based class class-default fair-queue ! ! PARENT: shape to the service rate, then hand off to the child policy-map PARENT-SHAPER class class-default shape average 20000000 ! shape everything to 20 Mbps service-policy CHILD-QUEUEING ! then queue within that 20 Mbps ! interface GigabitEthernet0/0 service-policy output PARENT-SHAPER ``` The magic is the last line of the parent: `service-policy CHILD-QUEUEING` nested inside the parent's class-default. This says "shape all traffic to 20 Mbps, and *within* that shaped queue, apply the child's LLQ/CBWFQ logic." Now the child's `priority percent 20` means 20% of the 20 Mbps service rate (4 Mbps), not 20% of the 100 Mbps interface. The percentages finally mean what you intend. Read it as a pipeline: traffic hits the parent shaper, which buffers anything above 20 Mbps and releases it smoothly at 20 Mbps; the traffic released by the shaper flows through the child policy, which prioritises voice and guarantees the data floor within that 20 Mbps. Shape first, queue within the shaped rate. ## The Per-Site Case: Nested Shaping The same structure scales to the multi-site problem. Imagine a physical link carrying traffic to ten branches, each with its own contracted rate. You build a parent that shapes each branch to its rate (matched by a class per branch), with a child queueing policy inside each: ``` policy-map PER-SITE-PARENT class SITE-A ! matches traffic to branch A shape average 5000000 ! branch A gets 5 Mbps service-policy CHILD-QUEUEING class SITE-B shape average 10000000 ! branch B gets 10 Mbps service-policy CHILD-QUEUEING ``` Each branch is independently shaped to its rate, and within each branch's shaped bandwidth, voice still gets priority. This is how a service provider or a large enterprise delivers differentiated per-site QoS over one shared physical link, and it is only expressible with the hierarchy. ## The Shaper's Details That Matter Two shaping parameters affect behaviour and cause confusion: shape average vs shape peak`average` shapes strictly to the rate. `peak` allows brief bursts above it. Use `average` to match a carrier rate exactly. Bc (committed burst)How much can be sent in one interval. Too small and the shaper is jerky; the default is usually fine, but voice-heavy traffic sometimes needs a smaller Bc for smoother delivery. The interaction with LLQ matters: a shaper adds delay (it buffers), and delay is the enemy of voice. When you shape a link carrying voice, the shaper's queue can add latency to voice packets before they even reach the child's priority queue. This is why HQoS designs for voice often tune the shaper's Bc down, so the shaping interval is short and voice is not held long. It is a real tension: shaping smooths bursts (good) but adds delay (bad for voice), and the child priority queue only helps *within* the shaped rate. ## Verifying the Hierarchy The same command shows both levels, nested: ``` R1#show policy-map interface GigabitEthernet0/0 Service-policy output: PARENT-SHAPER Class-map: class-default shape (average) cir 20000000, ... target shape rate 20000000 Service-policy : CHILD-QUEUEING Class-map: EF-VOICE Priority: 20% (4000 kbps) <- 20% of the SHAPED 20 Mbps, not the interface Class-map: BULK-DATA bandwidth 40% (8000 kbps) ``` Notice the child's `Priority: 20% (4000 kbps)`: the percentage is now calculated against the parent's 20 Mbps shape rate, giving 4 Mbps, exactly the intent. On a flat policy it would have been 20% of the physical interface. That recalculation is the entire reason the hierarchy exists. ## FAQ ### Shaping or policing for a carrier rate mismatch? Shaping. It buffers the excess and sends it later at the contracted rate, avoiding the drops that policing (yours or the carrier's) would cause. Dropping traffic you could have buffered is needlessly lossy. ### Why can't a flat policy handle the 100/20 Mbps case? A flat policy's percentages are of the physical interface (100 Mbps). You need to shape to 20 Mbps first and apply queueing within that 20, so the percentages mean 20% of 20 Mbps. That requires a parent (shape) and a child (queue). ### Does shaping add latency? Yes. Shaping buffers traffic above the rate, and buffering is delay. This matters for voice; HQoS designs carrying voice often tune the shaper's Bc down to keep the shaping interval short. ### What is the difference between shape average and shape peak? `average` shapes strictly to the rate; `peak` allows brief bursts above it. Use `average` to match a carrier's contracted rate exactly, or the carrier will police your peaks. ### Can I shape different sites to different rates on one link? Yes, that is the multi-site HQoS case: a parent with a class and shaper per site, each with a nested child queueing policy. Each site is independently shaped, with voice priority within its rate. ## Key Takeaways - **Policing drops** excess (sharp, no delay, bursty); **shaping buffers** it and sends later (smooth, adds delay, needs memory). Use shaping to match a downstream rate. - HQoS solves the case a flat policy cannot: **shape to the service rate (parent), then queue within it (child)**. A 100 Mbps interface delivering a 20 Mbps service needs this. - The parent's `service-policy CHILD` nesting makes the child's percentages calculate against the **shaped rate** (20% of 20 Mbps = 4 Mbps), not the physical interface. - The same structure scales to **per-site shaping**: a class and shaper per branch, each with a nested queueing child. - Shaping **adds delay**, which fights voice; HQoS-for-voice designs tune Bc down to shorten the shaping interval. - Verify with `show policy-map interface`, which displays both levels nested and the recalculated child percentages. Next: [Interpreting QoS configurations](https://www.pinglabz.com/interpreting-qos-configurations-encor/), or the [QoS cluster guide](https://www.pinglabz.com/qos/). ### WRED and Congestion Avoidance: Dropping Packets on Purpose URL: https://www.pinglabz.com/wred-congestion-avoidance/ Last updated: 2026-07-12T03:19:42.000Z The obvious way to handle a full queue is to drop whatever arrives once the queue is full. It is called tail drop, it is the default, and it causes a subtle, damaging problem that gets worse exactly when the network is busiest. Weighted Random Early Detection is the counterintuitive fix: start dropping packets *before* the queue is full, deliberately and at random. Understanding why that is better than waiting is the whole point of this article. It extends the [QoS cluster guide](https://www.pinglabz.com/qos/) and builds on the queueing policy from [LLQ and CBWFQ](https://www.pinglabz.com/llq-cbwfq-cisco-ios-xe/). ## The Problem: Tail Drop and Global Synchronization Here is the failure mode WRED exists to solve, and it is more interesting than it first sounds. When a queue fills and tail drop kicks in, it drops packets from *many* TCP flows at once, because they all happen to be sending when the queue overflows. Every one of those TCP senders sees loss, and every one of them backs off simultaneously (TCP's congestion response). The link suddenly empties. Then all those senders ramp back up together, refill the queue, overflow it again, and all back off again, in lockstep. This is **TCP global synchronization**, and it produces a sawtooth: the link oscillates between overfull and underused, never settling at efficient steady-state utilisation. You get the worst of both worlds, congestion loss *and* wasted capacity, and it is entirely an artifact of dropping everyone at the same instant. Tail dropWait until the queue is full, then drop everything that arrives. Drops many flows at once. Causes global synchronization. WREDStart dropping a few random packets as the queue *fills*, before it is full. Signals a few flows to slow down early, avoiding the synchronized cliff. ## How WRED Works WRED watches the average queue depth and drops probabilistically based on how full the queue is getting. Three parameters per traffic class define the behaviour: Minimum threshold Below this queue depth, drop nothing. The queue is healthy; leave it alone. Maximum threshold At this depth, drop at the maximum probability. Above it, tail-drop everything (the safety net). Mark probability The peak drop rate at the max threshold (e.g. 1/10 = drop 1 in 10). Between the thresholds, the rate rises linearly from 0 to this. Between the minimum and maximum thresholds, the drop probability rises smoothly from zero to the mark probability as the queue fills. So a lightly-loaded queue drops nothing, a filling queue drops a few random packets (nudging a few flows to slow down), and only a genuinely overwhelmed queue tail-drops. The randomness is the key: by dropping from different flows at different times rather than all at once, WRED breaks the synchronization. ## The "Weighted" Part: DSCP-Aware Dropping The W in WRED is what makes it useful for QoS. Plain RED drops all traffic in a queue equally. **Weighted** RED uses different thresholds per DSCP value, so more-important traffic gets more aggressive thresholds (dropped later) and less-important traffic gets dropped earlier. This means that as a queue fills, WRED sheds the low-priority traffic first, protecting the high-priority traffic in the same class. Configured with one line inside the class: ``` R1(config-pmap-c)# random-detect dscp-based ``` And the resulting drop profile is visible in the policy output. Here is the BULK-DATA class from the lab, showing WRED active with per-DSCP thresholds: ``` R1#show policy-map interface Ethernet0/2 | begin BULK-DATA Class-map: BULK-DATA (match-all) 6767 packets, 10222250 bytes Match: dscp af11 (10) bandwidth 40% (4000 kbps) Exp-weight-constant: 9 (1/512) Mean queue depth: 0 packets dscp Transmitted Random drop Tail drop Minimum Maximum Mark pkts/bytes pkts/bytes pkts/bytes thresh thresh prob af11 6767/10222250 0/0 0/0 32 40 1/10 ``` Read the profile: AF11 traffic starts being randomly dropped at a queue depth of **32 packets** (minimum threshold), reaches its peak drop rate of **1/10** at **40 packets** (maximum threshold), and tail-drops above that. The `Exp-weight-constant` (1/512) controls how the *average* queue depth is calculated, smoothing out momentary spikes so WRED reacts to sustained congestion, not transient bursts. ### An honest platform note Notice `Random drop 0/0` and `Mean queue depth: 0 packets` in that capture, despite the link being genuinely congested during the test. This is a limitation of the virtual IOL platform: its software queue does not build the sustained depth that a hardware queue does under load, so the average queue depth never climbed past the minimum threshold to trigger random drops. The drops during congestion happened at the link-conditioning stage instead (visible as TCP retransmissions in the [LLQ article](https://www.pinglabz.com/llq-cbwfq-cisco-ios-xe/)). The WRED *profile* is correctly installed and shown; the random-drop *mechanism* needs a real hardware queue building depth to fire, which is exactly what happens on production routers. No fabricated drop counters are shown here. ## ECN: Marking Instead of Dropping WRED has a modern refinement worth knowing: Explicit Congestion Notification. Instead of *dropping* a packet to signal congestion, ECN *marks* it (sets a bit in the IP header) and lets it through. An ECN-aware TCP sender sees the mark and slows down without ever losing the packet, so you get the congestion signal without the retransmission. ``` R1(config-pmap-c)# random-detect dscp-based R1(config-pmap-c)# random-detect ecn ``` ECN is strictly better than dropping when both endpoints support it (most modern stacks do): the flow slows down, nothing is lost, no retransmit. It is the same early-warning idea as WRED, delivered by a mark rather than a drop. The catch is that both endpoints and the path must honour it, so it is deployed where you control the stack. ## When to Use WRED (and When Not To) - **Use WRED on TCP-heavy queues.** Its whole benefit is managing TCP's congestion response and avoiding global synchronization. On bulk-data and default classes full of TCP flows, it improves steady-state utilisation. - **Never use WRED on a voice/priority queue.** Voice is UDP; it does not respond to drops by slowing down (there is no retransmit and no backoff), so dropping voice packets early just degrades the call for no benefit. The LLQ priority class should tail-drop (or better, be policed so it never fills). This is why the lab's EF class has no WRED. - **It only helps under sustained congestion.** On a link that rarely fills, WRED does nothing. Its value is on links that regularly run hot. ## FAQ ### Why drop packets before the queue is full? To avoid TCP global synchronization. Tail drop drops many flows at once, making them all back off together and causing a sawtooth of over-full then under-used. WRED drops a few random packets early, nudging a few flows to slow down and keeping utilisation smooth. ### What does "weighted" mean in WRED? Different drop thresholds per DSCP value, so lower-priority traffic in a queue is dropped earlier than higher-priority traffic. WRED sheds the less important traffic first as the queue fills. ### Should I put WRED on my voice queue? No. Voice is UDP and does not respond to early drops by slowing down, so WRED just harms call quality. Use WRED on TCP queues (bulk, default); let the priority queue tail-drop or rely on its policer. ### What is ECN? Explicit Congestion Notification: instead of dropping a packet to signal congestion, mark it and let it through. An ECN-aware sender slows down without a retransmission. Better than dropping when both ends support it. ### Why were the WRED drop counters zero in the lab? The virtual IOL software queue did not build sustained depth under load, so the average queue depth never crossed the minimum threshold. The profile is installed correctly; the random-drop mechanism fires on hardware queues that build real depth. The congestion drops in the test happened at the link conditioner. ## Key Takeaways - Tail drop causes **TCP global synchronization**: dropping many flows at once makes them all back off together, producing a sawtooth of over-full then under-used. - **WRED** drops a few random packets *before* the queue is full, nudging individual flows to slow down early and keeping utilisation smooth. - Three parameters: **minimum threshold** (start dropping), **maximum threshold** (peak drop rate, then tail-drop), **mark probability** (the peak rate). - **Weighted** \= per-DSCP thresholds, so WRED sheds lower-priority traffic first. Configured with `random-detect dscp-based`. - **ECN** marks instead of drops, signalling congestion without loss when both ends support it. - Use WRED on **TCP** queues; **never on voice** (UDP does not back off). Honest note: the virtual queue did not build depth to trigger random drops, though the profile is correctly installed. Next: [Hierarchical QoS and shaping](https://www.pinglabz.com/hierarchical-qos-shaping/), or the [QoS cluster guide](https://www.pinglabz.com/qos/). ### LLQ and CBWFQ on Cisco IOS XE: Building the Queueing Policy URL: https://www.pinglabz.com/llq-cbwfq-cisco-ios-xe/ Last updated: 2026-07-12T03:19:42.000Z When a WAN link congests, something has to give, and QoS is how you decide what. Left to itself, a link under load treats a voice packet exactly like a bulk file transfer, and the call breaks while the download barely notices. Low Latency Queueing and Class-Based Weighted Fair Queueing are the tools that let you say "voice goes first, always, and everything else shares what is left fairly." This article builds the full queueing policy on Cisco IOS XE and then congests a real link to prove the priority class is protected. It extends the [QoS cluster guide](https://www.pinglabz.com/qos/). ## Queueing Only Matters When the Link Is Full The first thing to internalise: queueing does nothing on an uncongested link. If packets can leave as fast as they arrive, there is no queue to manage. Queueing policy only takes effect when traffic arrives faster than the interface can send it, so packets have to wait in a queue, and *which* queue, and in what order they drain, is what QoS decides. So the whole subject is about behaviour under congestion. To demonstrate it honestly you have to actually congest a link, which is exactly what the lab does: a 5 Mbps choke point with a bulk transfer saturating it and a voice-marked flow competing for the same pipe. ## The Distinction Everyone Confuses: priority vs bandwidth LLQ and CBWFQ are configured with two commands that look similar and behave completely differently. Getting this straight is the core of the topic: priority (LLQ) A **strict-priority** queue serviced before all others, with a policer as a ceiling. Guarantees low latency. For voice and interactive video, where delay is fatal. bandwidth (CBWFQ) A **guaranteed minimum** share during congestion, but no latency guarantee. The class can use more if the link is free. For important data that needs a floor, not a fast lane. The mental model: `priority` is a fast lane that is always serviced first (but capped, so it cannot starve everything else). `bandwidth` is a reserved minimum share of a shared road. Voice needs the fast lane because a late voice packet is a useless voice packet. A database replication job needs a guaranteed minimum but does not care about a few milliseconds of delay, so it gets bandwidth. The critical subtlety: LLQ's priority queue is **policed**. It is serviced first, but only up to its configured rate; traffic above that rate is dropped. This is deliberate and essential, because an unpoliced strict-priority queue could consume the entire link and starve every other class. The policer is what makes strict priority safe. ## Building the Policy Three steps: classify (class-maps), define behaviour (policy-map), apply (service-policy). The classification here uses DSCP, assuming traffic is already marked (marking is upstream, at the network edge). ``` R1(config)# class-map match-all EF-VOICE R1(config-cmap)# match dscp ef R1(config)# class-map match-all BULK-DATA R1(config-cmap)# match dscp af11 ! R1(config)# policy-map WAN-EDGE R1(config-pmap)# class EF-VOICE R1(config-pmap-c)# priority percent 20 R1(config-pmap)# class BULK-DATA R1(config-pmap-c)# bandwidth percent 40 R1(config-pmap-c)# random-detect dscp-based R1(config-pmap)# class class-default R1(config-pmap-c)# fair-queue R1(config-pmap-c)# random-detect ! R1(config)# interface Ethernet0/2 R1(config-if)# service-policy output WAN-EDGE ``` Read the policy as a set of promises. **EF-VOICE** (matched on DSCP EF, the standard voice marking) gets a strict-priority fast lane capped at 20% of the link. **BULK-DATA** (DSCP AF11) gets a guaranteed 40% minimum with WRED ([covered separately](https://www.pinglabz.com/wred-congestion-avoidance/)) managing its queue. **class-default**, everything else, gets fair-queueing so no single flow can dominate the leftover bandwidth. And the policy is applied `output`, because queueing happens on egress, where the interface is the bottleneck. Confirm it installed correctly: ``` R1#show policy-map interface Ethernet0/2 Class-map: EF-VOICE (match-all) Match: dscp ef (46) Priority: 20% (2000 kbps), burst bytes 50000, b/w exceed drops: 0 Class-map: BULK-DATA (match-all) Match: dscp af11 (10) bandwidth 40% (4000 kbps) Class-map: class-default (match-any) Fair-queue: per-flow queue limit 16 packets ``` ## Congest It and Watch the Priority Class Survive Now the demonstration. First, a baseline: the voice flow (EF-marked UDP) on an uncongested link: ``` j@llmbits:~$ iperf3 -c 10.0.30.10 -u -b 1M -S 184 -t 5 # -S 184 = DSCP EF [ 5] 0.00-5.05 sec 611 KBytes 992 Kbits/sec 0.110 ms 0/432 (0%) receiver ``` Clean: **0.11 ms jitter, 0% loss**. Now saturate the same 5 Mbps link with a bulk transfer while the voice flow runs concurrently. The bulk transfer alone shows the link is genuinely full: ``` j@llmbits:~$ iperf3 -c 10.0.30.10 -t 12 -P 3 # bulk, best-effort [SUM] 0.00-14.00 sec 14.8 MBytes 8.84 Mbits/sec 162 sender [SUM] 0.00-16.25 sec 7.88 MBytes 4.07 Mbits/sec receiver ``` The sender pushed 8.84 Mbps into a 5 Mbps pipe; only 4.07 Mbps came out the far end, and there were **162 TCP retransmissions**. That gap is real packet loss at the choke, real congestion, exactly the condition QoS exists for. And here is the payoff, the priority class under that same congestion: ``` R1#show policy-map interface Ethernet0/2 Class-map: EF-VOICE (match-all) 1125 packets, 1673362 bytes Priority: 20% (2000 kbps), burst bytes 50000, b/w exceed drops: 0 queue stats for all priority classes: (queue depth/total drops/no-buffer drops) 0/0/0 ``` **b/w exceed drops: 0\. Total drops: 0.** While the bulk transfer was losing 162 packets to congestion, the voice class in its strict-priority queue lost nothing. That is LLQ working: the priority queue is serviced first, so under congestion its packets go out ahead of the bulk, and the voice call stays clean while the file transfer takes the hit. The policer confirmed the EF traffic stayed under its 2000 kbps reservation, so none of it was dropped for exceeding. That contrast, bulk losing 162 packets, voice losing zero, on the same congested link at the same moment, is the entire case for LLQ in one screen. ## The CBWFQ Share The bandwidth class earns its keep too. When the bulk traffic is marked into BULK-DATA (DSCP AF11), it lands in the CBWFQ class with its guaranteed 40% floor: ``` R1#show policy-map interface Ethernet0/2 | begin BULK-DATA Class-map: BULK-DATA (match-all) 6767 packets, 10222250 bytes Match: dscp af11 (10) bandwidth 40% (4000 kbps) ``` That 40% is a *minimum guarantee under congestion*, not a cap. If the link is otherwise idle, BULK-DATA can use more; the guarantee only kicks in when there is contention. This is the difference from policing (a hard ceiling): CBWFQ gives a floor and lets the class borrow above it when bandwidth is free. ## Why the Percentages Must Add Up Carefully A design rule that trips people up: the sum of your priority and bandwidth allocations cannot exceed the available bandwidth (by default, 75% of the interface, reserving headroom for control traffic and overhead). Here, 20% priority + 40% bandwidth = 60%, which is safe. Push it toward 100% and IOS will reject the policy or you will starve control-plane traffic. And a voice-specific rule: keep the priority (LLQ) allocation modest, typically no more than about a third of the link. Voice does not need much bandwidth (a G.711 call is roughly 80 kbps), but it needs that little bit *reliably and with low latency*. Over-allocating the priority queue wastes bandwidth other classes could use and, worse, a huge strict-priority queue can delay everything else during a burst. ## FAQ ### What is the difference between priority and bandwidth? `priority` (LLQ) is a strict-priority, policed fast lane for latency-sensitive traffic like voice. `bandwidth` (CBWFQ) is a guaranteed minimum share with no latency guarantee, for important data. Fast lane vs reserved minimum. ### Why is the priority queue policed? So it cannot starve other classes. A strict-priority queue is always serviced first; without a policer capping it, a flood of priority traffic could consume the entire link. The policer makes strict priority safe. ### Does QoS do anything on an uncongested link? No. Queueing only matters when packets must wait. On a link that is not full, everything leaves immediately and the policy has no effect. QoS is about behaviour under congestion. ### Where do I apply the policy, input or output? Output. Queueing happens on egress, where the interface is the bottleneck. You classify and mark on input at the edge, but you queue on output. ### How much bandwidth should the priority class get? Keep it modest, typically under a third of the link. Voice needs reliability and low latency, not much bandwidth. Over-allocating priority wastes capacity and can delay other traffic during bursts. ## Key Takeaways - Queueing only matters **under congestion**. To demonstrate it you must actually fill the link, which the lab did with a 5 Mbps choke. - **priority (LLQ)** \= strict-priority, policed fast lane for voice. **bandwidth (CBWFQ)** \= guaranteed minimum share for important data. Different jobs. - The LLQ priority queue is **policed** so it cannot starve other classes. That policer is what makes strict priority safe. - The lab proved it: under congestion the bulk transfer lost **162 packets** while the voice class lost **zero (b/w exceed drops: 0)**, on the same link at the same time. - CBWFQ's bandwidth is a **minimum floor** under contention, not a cap; the class can borrow more when the link is free. - Priority + bandwidth allocations must fit within the reservable bandwidth (default 75%), and the priority share should stay modest. Next: [WRED and congestion avoidance](https://www.pinglabz.com/wred-congestion-avoidance/), or the [QoS cluster guide](https://www.pinglabz.com/qos/). ### Python and JSON for the ENCOR Exam: Reading Scripts Under Pressure URL: https://www.pinglabz.com/python-json-encor-exam/ Last updated: 2026-08-01T19:32:44.000Z The ENCOR exam does not ask you to be a software developer, but it does ask you to read a Python script and a JSON payload and understand what they do. Under exam pressure, a screen of code you half-recognise is a trap; a screen of code whose structure you can parse in ten seconds is free marks. This is a comprehension skill, not a coding skill, and it is entirely learnable. This article covers the Python and JSON you need to read (not write) automation code for the exam and for real work. It extends the [Network Automation cluster guide](https://www.pinglabz.com/network-automation/). ## JSON First: It Is Just Nested Data Every network API you will meet (RESTCONF, Catalyst Center, SD-WAN Manager, Meraki) speaks JSON, so read it fluently. JSON has exactly two container types and a handful of value types, and that is the whole language: Object { } Key-value pairs. In Python this becomes a **dict**. Access a value by its key. Array \[ \] An ordered list of values. In Python this becomes a **list**. Access by index, or iterate. Everything else is values inside those two: strings (in quotes), numbers, booleans (`true`/`false`), and `null`. Reading JSON is just tracing the nesting. Here is a typical API response: ``` { "response": [ { "hostname": "core-1", "mgmtIp": "10.1.1.1", "reachable": true }, { "hostname": "core-2", "mgmtIp": "10.1.1.2", "reachable": false } ] } ``` Trace it: the top level is an **object** with one key, `responsearray of two **objects**, each with three keys. To get the first device's hostname you walk the path: into response, take element [0], take key hostname. That is data["response"][0]["hostname"] and it is the single most common operation in all network automation.` ## `The Python You Need to Recognise` `Map JSON onto Python and most scripts become readable:` dict`{"key": "value"}` \- a JSON object. Access with `d["key"]`. list`[a, b, c]` \- a JSON array. Access with `l[0]`, iterate with `for x in l`. for loop`for device in devices:` \- do something to each item in a list. if`if device["reachable"]:` \- branch on a condition. `Put those together and a real automation snippet reads like plain English:` ``` import requests resp = requests.get(url, headers=headers, auth=auth, verify=True) data = resp.json() # parse JSON into a Python dict for device in data["response"]: # loop the array of devices if not device["reachable"]: # branch on the boolean print(f"DOWN: {device['hostname']} ({device['mgmtIp']})") ``` `Read it top to bottom: make an HTTP GET, turn the JSON body into a dict with .json(), loop over the list of devices, and print the ones that are not reachable. If you can narrate a script like that in your own words, you can answer the exam questions, because they ask exactly that: "what does this print?"` ## `The Three Libraries to Recognise` `You do not need to memorise their APIs, but you should recognise what each one is for the moment you see its import:` requests `import requests`. HTTP: talking to REST and RESTCONF APIs. `.get()`, `.post()`, `.patch()`, `.json()`. ncclient `from ncclient import manager`. NETCONF: `manager.connect()`, `get_config()`, `edit_config()`. json `import json`. `json.loads()` (text to dict), `json.dumps()` (dict to text). `When you see import requests, the script talks to a REST API. When you see from ncclient import manager, it is `[NETCONF](https://www.pinglabz.com/netconf-cisco-ios-xe-hands-on/)`. When you see import json, it is shuffling JSON to or from text. Recognising the library tells you the script's job before you read a line of logic.` One more worth recognising outside the exam list: Genie, from the pyATS suite, parses raw `show` output into precisely the nested objects and arrays above, so you index it with the same walk. [Getting JSON out of Cisco show commands](https://www.pinglabz.com/pyats-genie-network-testing/) is where these structures come from when there is no REST API in the picture. ## `Exam Technique: Reading a Script Under Pressure` `A repeatable method for the "what does this code do / output" questions:` 1. `**Read the imports first.** They tell you the domain (REST, NETCONF, JSON) instantly.` 2. `**Find the data structure.** Is the script working on a dict or a list? Locate the JSON it parses and picture its shape.` 3. `**Trace the access path.** data["response"][0]["hostname"] is just "into response, first element, its hostname." Walk it one step at a time.` 4. `**Follow the loop and the condition.** What is it iterating, and what does the if select? The output is whatever survives the filter.` 5. `**Ignore what you do not need.** Error handling, headers, and auth setup are usually not what the question hinges on. Find the line that produces the answer.` `The questions are testing comprehension, not recall. You are not asked to write the script; you are asked to say what it does. That is a skill you build by reading a handful of real scripts until the structure is automatic, which is exactly why the rest of this cluster shows real `[RESTCONF](https://www.pinglabz.com/restconf-cisco-ios-xe/)` and `[NETCONF](https://www.pinglabz.com/netconf-cisco-ios-xe-hands-on/)` scripts against a live device.` ## `Common Gotchas the Exam Loves` - `**Dict access vs list access.** data["response"] (a key, for an object) vs data[0] (an index, for an array). Mixing them up is the classic error, and a favourite distractor.` - `**JSON true vs Python True.** JSON uses lowercase true/false/null; Python uses True/False/None. .json() converts them, but a question may show both forms.` - `**Strings vs numbers.** "10" (a string) is not 10 (a number). JSON preserves the distinction, and comparisons care.` - `**The .json() call.** A requests response is not a dict until you call .json() on it. Code that indexes the raw response object is wrong.` ## `FAQ` ### `Do I need to write Python for the exam?` `No. You need to *read* it: given a script, say what it does or outputs. That is comprehension, which is far easier than writing, and entirely learnable by reading real scripts.` ### `How do I access a nested JSON value?` `Walk the path one step at a time. data["response"][0]["hostname"] means: into the response key, take the first list element, take its hostname key. Dicts use keys in quotes; lists use numeric indices.` ### `What does .json() do?` `It parses the JSON text of an HTTP response into a Python dict (or list). Until you call it, the response is not indexable as data.` ### `How do I tell REST code from NETCONF code?` `The import. import requests is REST/RESTCONF; from ncclient import manager is NETCONF. The library announces the domain.` ### `Why does true not work in my Python?` `JSON uses lowercase true; Python uses capitalised True. Inside Python code use True; you only see lowercase true in raw JSON text.` ## `Key Takeaways` - `The exam tests **reading** Python and JSON, not writing them. It is comprehension, and it is learnable.` - `JSON is two containers: **object { }** (a Python dict, keyed) and **array [ ]** (a Python list, indexed). Everything else is values inside them.` - `The core operation is walking a nested path: data["response"][0]["hostname"] = into response, first element, its hostname.` - `Recognise three libraries by their import: **requests** (REST), **ncclient** (NETCONF), **json** (parsing).` - `Read a script by: imports first (the domain), then the data shape, then the access path, then the loop/condition.` - `Watch the gotchas: dict-key vs list-index access, JSON true vs Python True, and remembering that a response needs .json() before you can index it.` `Next: apply it against a live device in `[RESTCONF on Cisco IOS XE](https://www.pinglabz.com/restconf-cisco-ios-xe/)`, or the `[Network Automation cluster guide](https://www.pinglabz.com/network-automation/)`.` ### SD-WAN Manager (vManage) API: Monitoring Catalyst SD-WAN Programmatically URL: https://www.pinglabz.com/sd-wan-manager-api/ Last updated: 2026-07-12T02:58:39.000Z If Catalyst Center is the controller for the campus, SD-WAN Manager (formerly vManage) is the controller for the WAN. It orchestrates the entire Cisco Catalyst SD-WAN fabric, every edge router, every policy, every tunnel, and like every modern controller it exposes a REST API. Monitoring an SD-WAN fabric programmatically, pulling tunnel health, or automating policy changes all go through that API. This article covers the SD-WAN Manager API for network engineers. It extends the [Network Automation cluster guide](https://www.pinglabz.com/network-automation/) and the [SD-WAN cluster](https://www.pinglabz.com/sd-wan/). As with Catalyst Center, the examples illustrate the API's shape rather than being lab captures, because the SD-WAN control stack cannot run in a standard CML environment; Cisco's DevNet sandboxes provide a live SD-WAN Manager to practice against. ## What SD-WAN Manager Sees SD-WAN Manager is the single pane of glass for a Catalyst SD-WAN fabric. Its API surfaces the things that make SD-WAN operationally different from traditional WAN: Device inventory Every edge (WAN Edge), controller, and validator in the fabric, with status. Tunnel / BFD statistics The health of every overlay tunnel: loss, latency, jitter per path. The data SD-WAN uses to steer traffic. Policies Application-aware routing, data, and control policies. Read and, with care, modify them. Alarms / events Fabric-wide alarms and events, for feeding into your own monitoring. The most valuable read for most engineers is the tunnel statistics. SD-WAN's whole premise is measuring path quality and steering applications accordingly (covered in the [SD-WAN cluster](https://www.pinglabz.com/sd-wan/)); the API lets you pull those measurements into your own dashboards and alerting rather than living inside the SD-WAN Manager GUI. ## Session-Based Authentication SD-WAN Manager uses a slightly different auth model from Catalyst Center, worth noting because it trips people up. You POST credentials to a login endpoint and receive a **session cookie**, and (on current versions) you then fetch a **CSRF token** that must accompany write requests: ``` # 1. Log in, receive a JSESSIONID session cookie POST https:///j_security_check j_username=admin&j_password= (form-encoded) # 2. Fetch a CSRF token for subsequent write calls GET https:///dataservice/client/token -> # 3. Use the cookie (and token for writes) on API calls GET https:///dataservice/device Cookie: JSESSIONID=... X-XSRF-TOKEN: ``` The session-cookie-plus-CSRF pattern is more stateful than a bearer token; your client has to hold the cookie jar across calls. The key operational point is the same, though: authenticate once, reuse the session, and protect the credentials. ## Reading Fabric State The `/dataservice/` path is the heart of the read API. Pulling the device list: ``` GET /dataservice/device Cookie: JSESSIONID=... -> 200 OK { "data": [ { "host-name": "branch-edge-01", "device-type": "vedge", "system-ip": "10.255.0.11", "reachability": "reachable", "site-id": "101", "version": "17.9.3a" }, ... ] } ``` And the tunnel statistics that make SD-WAN monitoring worthwhile, per device: ``` GET /dataservice/device/tunnel/statistics?deviceId=10.255.0.11 -> { "data": [ { "tunnel-protocol": "ipsec", "dst-ip": "...", "loss": 0.2, "latency": 14, "jitter": 3 }, ... ] } ``` That loss/latency/jitter per tunnel is the same class of measurement as an [IP SLA](https://www.pinglabz.com/ip-sla-cisco-ios-xe/) probe, but collected fabric-wide by the controller. Pulling it via the API lets you build alerting on WAN path quality that spans every site, which is exactly the kind of visibility SD-WAN promises and the API delivers. ## Writing: Templates and Policies The write side (pushing templates, changing policies) exists but demands respect. In Catalyst SD-WAN, configuration is driven by **templates** attached to devices, and changing a template can push config to many edges at once. An API-driven policy change is powerful and correspondingly dangerous: a mistake propagates fabric-wide. The disciplined approach is to read and validate extensively via the API, and gate writes behind change control and testing, treating an SD-WAN Manager write API the way you would treat `configure replace` on every router simultaneously. ## FAQ ### Can I lab SD-WAN Manager? Not in a standard CML setup (the full control stack is heavy and licensed). Cisco's DevNet always-on sandboxes provide a live SD-WAN Manager with an API you can call. The examples here show the API's shape. ### How is auth different from Catalyst Center? SD-WAN Manager uses a session cookie (JSESSIONID) from a login POST, plus a CSRF token for writes, rather than a bearer token. Your client must hold the cookie across calls. ### What is the most useful thing to read? Tunnel statistics (loss, latency, jitter per path). That per-tunnel quality data is the core of SD-WAN's value and feeds real WAN monitoring. ### Is it safe to change policy via the API? It works, but a template or policy change can propagate to many edges at once. Treat write operations with the same caution as a fabric-wide config change: validate, gate behind change control, test first. ### Is vManage the same as SD-WAN Manager? Yes, renamed. The API path still contains `/dataservice/`. The blueprint uses the current name, SD-WAN Manager. ## Key Takeaways - SD-WAN Manager is the WAN controller; its REST API (`/dataservice/`) exposes device inventory, tunnel statistics, policies, and alarms. - The highest-value read is **per-tunnel loss/latency/jitter**, the measurement SD-WAN steers on, pullable fabric-wide into your own tooling. - Auth is **session-cookie based** (JSESSIONID from a login POST) plus a **CSRF token** for writes, more stateful than a bearer token. - Writes are driven by **templates** and can propagate to many edges at once. Treat them like a fabric-wide config change: cautious, gated, tested. - It cannot be labbed in standard CML; use DevNet always-on sandboxes. Next: [Python and JSON for the ENCOR exam](https://www.pinglabz.com/python-json-encor-exam/), or the [Network Automation cluster guide](https://www.pinglabz.com/network-automation/). ### Catalyst Center APIs: Intent, Inventory, and AI-Powered Assurance URL: https://www.pinglabz.com/catalyst-center-apis/ Last updated: 2026-07-12T02:58:38.000Z Automating a single router with NETCONF is one thing. Automating a campus of hundreds of devices, with a single source of truth for intent, is what Catalyst Center (formerly DNA Center) is for. It is Cisco's controller for the enterprise network, and it exposes everything it does through a REST API. If you want to programmatically inventory your network, push intent, or pull assurance data at scale, that API is the door. This article explains the Catalyst Center API model, the categories of endpoint, and how to work with it. It extends the [Network Automation cluster guide](https://www.pinglabz.com/network-automation/). The request and response examples here are documentation-style illustrations of the API's shape, not lab captures, because Catalyst Center cannot run in CML; for hands-on practice, use Cisco's always-on DevNet sandboxes. ## The Intent API Model The key idea behind Catalyst Center's API is **intent**. Instead of telling each device what to do, you tell the controller what you *want*, and it works out the per-device configuration and pushes it. You express "this SSID should exist on these buildings" or "this device should run this template"; the controller renders and deploys the CLI. The API is the programmatic version of that: you POST intent, the controller executes it. The endpoints group into a few clear categories: Inventory / Discovery Read the device list, health, and details. The most common starting point: "what is on my network?" Assurance Client and device health scores, issues, and the AI-driven analytics. Pull "what is wrong and why" programmatically. Intent / Provisioning Push templates, sites, SSIDs, and fabric config. The write side: deploy configuration as intent. Command Runner Run read-only show commands across many devices at once and get structured results back. ## Authentication: Token First Every Catalyst Center API workflow starts the same way: authenticate once, get a token, use the token. This is the token model from the [API security article](https://www.pinglabz.com/rest-api-security-network-engineers/), and it is worth doing right because that token is a network-wide credential. ``` # 1. Authenticate, receive a token (documentation-style example) POST https:///dna/system/api/v1/auth/token Authorization: Basic -> 200 OK { "Token": "eyJhbGciOi......" } # 2. Use the token on every subsequent call GET https:///dna/intent/api/v1/network-device X-Auth-Token: eyJhbGciOi... ``` The token is short-lived (typically an hour). A robust script requests it, holds it in memory, and re-requests on a 401, exactly the pattern the [response-codes article](https://www.pinglabz.com/rest-api-response-codes-network/) describes. It never logs the token and never commits it. ## Reading the Inventory The most common first real task: pull the device inventory. The response is structured JSON, one object per device: ``` GET /dna/intent/api/v1/network-device X-Auth-Token: -> 200 OK { "response": [ { "hostname": "campus-core-1", "managementIpAddress": "10.1.1.1", "platformId": "C9500-40X", "softwareVersion": "17.9.4", "reachabilityStatus": "Reachable", "role": "CORE" }, ... ] } ``` From here everything is normal JSON handling (the subject of the [Python and JSON article](https://www.pinglabz.com/python-json-encor-exam/)): iterate `response`, pull the fields you need. The device `id` in each object is the handle you use for follow-up calls (health, config, command runner). This inventory pull is the "hello world" of Catalyst Center automation and the foundation of most scripts. ## The AI-Driven Assurance Angle The renamed emphasis in the current ENCOR blueprint ("Automation and Artificial Intelligence") shows up here. Catalyst Center's assurance is not just raw metrics; it applies machine learning to establish baselines, correlate events, and surface *issues* with probable root causes rather than leaving you to read graphs. Through the API you can pull: - **Health scores** for clients, devices, and applications, computed from many underlying metrics into a single 1-10 number. - **Issues**, the controller's own list of detected problems, each with a severity, an affected scope, and often a suggested remediation. - **Trends and anomalies**, where the AI has learned normal behaviour and flags deviations, catching a slow degradation a threshold alert would miss. For a network engineer, the practical value is that you can feed these into your own tooling: a script that pulls open issues each morning, a dashboard that tracks health-score trends, an alert that fires when the controller's AI flags an anomaly. The intelligence lives in the controller; the API lets you act on it. ## The Asynchronous Task Model One important wrinkle: many Catalyst Center write operations are **asynchronous**. When you POST a provisioning intent, you do not get the result immediately; you get a **task ID**, and the actual work happens in the background. You then poll a task endpoint to learn whether it succeeded. ``` POST /dna/intent/api/v1/ -> 202 Accepted { "response": { "taskId": "abc-123", "url": "/dna/intent/api/v1/task/abc-123" } } # poll until the task completes GET /dna/intent/api/v1/task/abc-123 -> { "response": { "isError": false, "progress": "..." } } ``` The **202 Accepted** response is the signal: "I have queued your request." A script that treats 202 as final success is wrong; it must poll the task ID until `isError` is set one way or the other. This asynchronous pattern is common to controllers generally and is a frequent source of "the API said OK but nothing happened" confusion. It did not say OK; it said "accepted, ask me later." ## FAQ ### Can I lab Catalyst Center? Not in CML. Use Cisco's always-on DevNet sandboxes (sandboxdnac.cisco.com and similar), which expose a real Catalyst Center API you can call for free. The examples here illustrate the API shape; the sandbox is where you run them for real. ### How does authentication work? Basic auth once to `/dna/system/api/v1/auth/token` returns a JWT token. Send that token as `X-Auth-Token` on every subsequent call. It expires (about an hour); re-authenticate on a 401. ### Why did my provisioning call return 202 with no result? Because it is asynchronous. 202 means "accepted and queued." You get a task ID and must poll the task endpoint until it reports success or error. 202 is not final success. ### What is the difference between the intent API and the CLI? The CLI configures one device. The intent API tells the controller what you want, and the controller renders and pushes the per-device config across the whole managed network. You express outcome, not per-device commands. ### Is DNA Center the same as Catalyst Center? Yes, renamed. The API paths still contain `/dna/` for compatibility, but the product is Catalyst Center. The blueprint uses the current name. ## Key Takeaways - Catalyst Center is the enterprise controller; its **REST API** exposes inventory, assurance, intent/provisioning, and command runner. - The model is **intent**: you tell the controller what you want, it renders and pushes per-device config. - Authenticate once for a **token** (`/auth/token`), send it as `X-Auth-Token`, re-auth on 401\. Treat the token as a network-wide credential. - The **AI-driven assurance** (health scores, issues, anomalies) is pullable via the API, letting you act on the controller's intelligence in your own tooling. - Write operations are often **asynchronous**: a **202 Accepted** returns a task ID you must poll. 202 is not final success. - It cannot be labbed in CML; use the DevNet always-on sandboxes for hands-on practice. Next: [SD-WAN Manager API](https://www.pinglabz.com/sd-wan-manager-api/), or the [Network Automation cluster guide](https://www.pinglabz.com/network-automation/). ### REST API Response Codes and Payloads: What 200, 404, and 500 Tell You URL: https://www.pinglabz.com/rest-api-response-codes-network/ Last updated: 2026-07-12T02:58:38.000Z An automation script that ignores HTTP response codes is a script that will eventually break something quietly. It sends a change, gets back a 409 it never checks, reports success, and moves on, and now your intent and the device's reality have diverged and nobody knows. Reading and handling response codes is not pedantry; it is the difference between automation you can trust and automation that occasionally lies to you. This article covers the HTTP response codes that matter for network APIs (RESTCONF, Catalyst Center, SD-WAN Manager) and how to act on each. It extends the [Network Automation cluster guide](https://www.pinglabz.com/network-automation/). ## The Five Classes, in One Line Each HTTP status codes group into five ranges by their first digit, and knowing the range tells you most of what you need: 2xx Success It worked. 200 (here is the data), 201 (created), 204 (done, no content to return). 3xx Redirect Go somewhere else. Rare in network APIs; your client library usually follows these automatically. 4xx Client error **You** did something wrong. Bad request, bad auth, wrong path. Retrying unchanged will not help. 5xx Server error **The device** did something wrong or could not cope. Sometimes worth a retry with backoff. The single most useful reflex: **4xx is your fault, 5xx is theirs.** A 4xx means fix your request; a 5xx means the device failed and a careful retry might succeed. Confusing the two leads to either pointless retry loops (retrying a 400 forever) or giving up too early (not retrying a transient 503). ## The Codes You Will Actually See 200 OKA GET succeeded and the body has your data. The everyday success for reads. 201 CreatedA POST created a new resource. You made something that did not exist before. 204 No ContentA PATCH/PUT/DELETE succeeded, and there is nothing to return. **This is success**, and it trips people up because the body is empty. 400 Bad RequestMalformed request: bad JSON/XML, a value that fails the YANG model's constraints. Fix the payload. 401 UnauthorizedMissing, wrong, or expired credentials/token. Re-authenticate; do not retry with the same bad credential. 403 ForbiddenAuthenticated but not permitted. A least-privilege boundary, or your role lacks the right. Not a retry. 404 Not FoundThe path/resource does not exist. Usually a wrong YANG path or a typo'd interface name. 409 ConflictThe request conflicts with current state (e.g. creating something that already exists, or a datastore lock). Read the current state and reconcile. 429 Too Many RequestsRate-limited. Back off (honour `Retry-After` if present). Ignoring it looks like an attack. 500 / 503Server error / unavailable. The device broke or is overloaded. Retry with exponential backoff; if it persists, it is a device problem. The two that catch people out most: **204** (an empty body is success, not failure, on a config change) and **429** (a controller like Catalyst Center will rate-limit you, and a script that hammers through 429s can get its token throttled or revoked). ## Handling Them in Code The pattern is not complicated, and it is the difference between trustworthy and reckless automation: ``` r = requests.patch(url, data=payload, headers=hdr, auth=auth, verify=True) if r.status_code == 204: log("change applied") elif r.status_code == 401: reauthenticate() # token expired; get a new one, retry once elif r.status_code == 403: fail("not authorized for this operation") # do not retry elif r.status_code == 409: reconcile(get_current_state()) # conflict; read and merge elif r.status_code == 429: backoff(r.headers.get("Retry-After", 30)) # rate-limited; wait elif 500 <= r.status_code < 600: retry_with_backoff() # device error; a careful retry may work else: fail(f"unexpected {r.status_code}: {r.text}") ``` The rules encoded there are the important part. **Never blindly retry a 4xx** (except 401 after re-authenticating and 429 after backing off); the request is wrong and retrying it unchanged just wastes calls. **Do back off on 429 and 5xx.** And **always check**, because the failure mode of not checking is silent divergence, which is the worst kind. ## Read the Body, Too The status code tells you the category; the response body often tells you exactly what went wrong. RESTCONF returns a structured error in the body of a 4xx: ``` { "errors": { "error": [{ "error-type": "application", "error-tag": "invalid-value", "error-message": "invalid value for: mtu" }] } } ``` That `error-message` is the difference between "the change failed" and "the change failed because the MTU value was invalid." A robust script logs the body on any error, not just the code. When you are debugging why a PATCH returns 400, the body is where the actual reason lives. ## A Word on Idempotency The HTTP methods differ in whether repeating them is safe, which matters when a retry might send the same request twice: - **GET, PUT, DELETE, PATCH** are idempotent: doing them twice has the same effect as doing them once. Safe to retry. - **POST** is not idempotent: two POSTs can create two resources. Retrying a POST after an ambiguous failure risks a duplicate. Design for it (check whether the resource was created before retrying). This is why config changes in RESTCONF usually use PUT or PATCH rather than POST: a retried PUT converges to the intended state, while a retried POST might make a mess. When your automation must be resilient to transient failures, prefer the idempotent methods. ## FAQ ### My PATCH returned 204 with an empty body. Did it fail? No. 204 is success with no content to return, the normal result of a successful config change. The empty body is expected. ### Should I retry a 400? No. A 400 means your request is malformed. Retrying it unchanged will fail identically. Fix the payload (the response body usually says what is wrong). ### What is the difference between 401 and 403? 401 = not authenticated (bad/expired/missing credentials, so re-authenticate). 403 = authenticated but not authorized (a permission boundary, so do not retry, fix the role). ### How should I handle 429? Back off and retry, honouring the `Retry-After` header if the API provides it. Never hammer through 429s; a controller can throttle or revoke your token. ### Which methods are safe to retry? GET, PUT, DELETE, PATCH are idempotent and safe. POST is not, so retrying it can create duplicates. Prefer PUT/PATCH for config so retries converge cleanly. ## Key Takeaways - **Always check the status code.** Not checking means silent divergence between intent and reality, the worst automation failure. - **4xx is your fault, 5xx is theirs.** Fix a 4xx request; a 5xx may succeed on a careful retry. - The tricky successes: **204** (empty body is success on a change) and knowing **201** means created. - The tricky errors: **401** (re-auth) vs **403** (permission, do not retry), **409** (conflict, reconcile), **429** (rate-limit, back off). - **Read the response body** on errors; RESTCONF's `error-message` tells you exactly what failed. - Prefer **idempotent methods** (PUT/PATCH) for config so retries converge instead of duplicating. Next: [RESTCONF on Cisco IOS XE](https://www.pinglabz.com/restconf-cisco-ios-xe/) where these codes appear live, or the [Network Automation cluster guide](https://www.pinglabz.com/network-automation/). ### YANG Data Models Explained for Network Engineers URL: https://www.pinglabz.com/yang-data-models-explained/ Last updated: 2026-07-12T02:58:38.000Z Every modern network API, NETCONF, RESTCONF, gNMI, is really a way of moving structured data in and out of a device. And structured data needs a schema: an agreed definition of what fields exist, what type each one is, and how they nest. That schema language is YANG, and it is the thing that makes model-driven automation possible. You cannot really understand NETCONF or RESTCONF without understanding what YANG is modelling. This article explains YANG for network engineers: what a data model is, the difference between native and OpenConfig models, and how to actually explore one. It extends the [Network Automation cluster guide](https://www.pinglabz.com/network-automation/). ## The Core Idea: A Schema for Configuration When you type `show ip interface brief`, IOS gives you text formatted for a human. That text has no schema; a script has to scrape it with fragile regex, and the format can change between releases. Model-driven management replaces that with structured data (XML or JSON) whose shape is defined by a YANG model. The model says, precisely: an interface has a name (string), an enabled flag (boolean), an IP address (a specific type), and so on. The payoff is that automation stops guessing. A script does not parse text hoping the columns line up; it requests `interfaces/interface/name` and gets exactly that field, in a defined type, every time. YANG is the contract that makes this reliable. YANGThe modelling language. Defines the structure and types (the schema). XML / JSONThe encoding. The actual data on the wire, shaped by the YANG model. NETCONF / RESTCONFThe transport. How the encoded data gets to and from the device. Three layers: YANG (what the data looks like), the encoding (XML or JSON), and the protocol ([NETCONF](https://www.pinglabz.com/netconf-cisco-ios-xe-hands-on/) or [RESTCONF](https://www.pinglabz.com/restconf-cisco-ios-xe/)). YANG is the foundation the other two stand on. ## The Anatomy of a Model A YANG model is a tree. The building blocks you will meet: container A grouping node with no value of its own, holding other nodes. Like a folder. `interfaces` is a container. list Multiple instances of the same structure, keyed by a field. `interface` is a list keyed by `name`. leaf A single value with a type. `name`, `enabled`, `mtu` are leaves. leaf-list Multiple values of the same leaf. A list of DNS servers, for example. Put them together and the standard `ietf-interfaces` model reads as: a container `interfaces`, holding a list `interface` keyed by `name`, where each entry has leaves like `name`, `enabled`, and `type`. That tree is exactly what you address when you make a NETCONF or RESTCONF request; the path `interfaces/interface=GigabitEthernet2/enabled` walks straight down it. ## Native vs OpenConfig: The Distinction That Matters There are two families of YANG model on a Cisco device, and knowing which to use is a real design decision: Native (Cisco-IOS-XE-\*) Cisco's own models, mirroring the full IOS XE feature set. Everything the CLI can do, but Cisco-specific: a script written against them only works on Cisco. OpenConfig / IETF Vendor-neutral models (openconfig-\*, ietf-\*). The same model works across Cisco, Juniper, Arista. Broad coverage, but not every vendor-specific feature. The trade-off is the same one as everywhere in networking: vendor-specific power versus multi-vendor portability. Use **OpenConfig or IETF models** when you want your automation to work across a mixed estate and you only need common features (interfaces, BGP, basic config). Use **native models** when you need a Cisco-specific feature that the standard models do not cover, accepting that your script is now Cisco-only. Many real deployments use both: OpenConfig for the portable 80%, native for the Cisco-specific 20%. ## Exploring a Model You do not have to read raw YANG source to work with a model. Two practical approaches. **pyang** renders a model as a readable tree. Point it at a downloaded YANG file and it prints the structure: ``` $ pyang -f tree ietf-interfaces.yang module: ietf-interfaces +--rw interfaces +--rw interface* [name] +--rw name string +--rw description? string +--rw type identityref +--rw enabled? boolean +--rw ipv4 +--rw address* [ip] +--rw ip inet:ipv4-address +--rw netmask? ... ``` That tree is the map. The `rw` flags mean read-write (configurable); `ro` would mean read-only (operational state you can read but not set). The `*` marks a list, and `[name]` is its key. The `?` marks an optional leaf. Once you can read this tree, you can construct any NETCONF or RESTCONF path into the model. **The device itself** advertises the models it supports. A NETCONF `hello` exchange lists every model in the device's capabilities, and you can pull the actual schema from the box. In practice you explore against the real device: request a subtree and see what comes back, which is exactly what the [NETCONF](https://www.pinglabz.com/netconf-cisco-ios-xe-hands-on/) and [RESTCONF](https://www.pinglabz.com/restconf-cisco-ios-xe/) articles do. ## Why a Network Engineer Should Care It is tempting to treat YANG as a programmer's concern. It is not, and here is why it matters operationally: - **It makes automation robust.** Text scraping breaks when output formats change between releases. A YANG-modelled request returns the same structured field regardless of how the CLI happens to format it. Your automation stops being fragile. - **It is how the exam frames automation.** The ENCOR blueprint expects you to understand data models, not just run scripts. Knowing container/list/leaf and native-vs-OpenConfig is directly testable. - **It is the future of the interface.** gNMI and streaming telemetry, the direction the industry is heading, are entirely model-driven. The CLI is not going away, but the programmatic interface is YANG-shaped, and understanding the model is understanding the interface. ## FAQ ### Do I have to read raw YANG source? No. Use `pyang -f tree` to render a readable tree, or explore against the live device by requesting subtrees. You almost never read the raw `.yang` files directly. ### Native or OpenConfig models? OpenConfig/IETF for multi-vendor portability and common features; native (Cisco-IOS-XE-\*) for Cisco-specific features the standard models do not cover. Many shops use both. ### What is the difference between rw and ro in the tree? `rw` is read-write (configuration you can set), `ro` is read-only (operational state you can read but not change). NETCONF separates these into config and state datastores. ### Is YANG the same as XML? No. YANG is the schema (the model definition). XML (or JSON) is the encoding (the actual data). YANG defines the shape; XML/JSON carries the values. ### How do I know which models a device supports? The device advertises them in its NETCONF capabilities (the `hello` exchange) and via the `netconf-state` schema list. You can enumerate them programmatically. ## Key Takeaways - YANG is the **schema language** for model-driven management: it defines the structure and types of configuration and state data. - Three layers: **YANG** (the model), **XML/JSON** (the encoding), **NETCONF/RESTCONF** (the transport). YANG is the foundation. - Models are trees of **container, list, leaf, leaf-list**. The path into that tree is exactly what NETCONF/RESTCONF requests address. - **Native models** (Cisco-IOS-XE-\*) cover everything but are Cisco-only; **OpenConfig/IETF** models are vendor-neutral but cover common features. Choose by portability need. - Explore with `pyang -f tree` or against the live device. You rarely read raw YANG source. - It matters because model-driven automation is robust (no text scraping), exam-relevant, and the direction of streaming telemetry and gNMI. Next: [NETCONF hands-on](https://www.pinglabz.com/netconf-cisco-ios-xe-hands-on/) and [RESTCONF](https://www.pinglabz.com/restconf-cisco-ios-xe/), or the [Network Automation cluster guide](https://www.pinglabz.com/network-automation/). ### RESTCONF on Cisco IOS XE: GET, PATCH, and the YANG Path URL: https://www.pinglabz.com/restconf-cisco-ios-xe/ Last updated: 2026-07-12T02:58:37.000Z If NETCONF is the powerful, ceremonious way to do model-driven management, RESTCONF is the friendly one. It exposes the same [YANG models](https://www.pinglabz.com/yang-data-models-explained/) over ordinary HTTPS, with the REST verbs every engineer already half-knows: GET to read, PATCH to modify, PUT to replace, DELETE to remove. If you can use `curl` or Python's `requests`, you can drive RESTCONF, and that low barrier is exactly why it is the on-ramp to network automation for most people. This article covers RESTCONF on Cisco IOS XE. It extends the [Network Automation cluster guide](https://www.pinglabz.com/network-automation/). ## A Note on the Lab Being straight about this lab, as always. The automation host (a Debian VM running Python `requests`) is real and reaches the target Catalyst 8000v; the RESTCONF HTTPS port (443) is reachable from it, confirmed live. The device has `restconf` and `ip http secure-server` enabled and all the DMI processes running. However, on this virtual platform the DMI datastore backend (`confd`/`ndbmand`) did not finish initializing within the lab window, so `nginx` served an "unavailable" page rather than data, a known limitation of the virtualized DMI. The RESTCONF request URLs, headers, and response *structure* below are therefore documentation-style, drawn from Cisco's documented RESTCONF behaviour and clearly framed as such. The device configuration and reachability are real captures. Nothing here is presented as a lab capture that was not one. ## RESTCONF Is REST Mapped Onto YANG The elegant thing about RESTCONF (RFC 8040) is how directly it maps familiar REST concepts onto the YANG tree: GETRead a resource (a subtree of the model). Like `get-config` in NETCONF. PATCHMerge a change into a resource. Modify one field, leave the rest. The everyday edit. PUTReplace a resource entirely with what you send. Idempotent: converges to the sent state. DELETERemove a resource. POST creates a new one under a parent. The URL *is* the path into the YANG tree. To address GigabitEthernet2 in the `ietf-interfaces` model, the URL walks the model exactly as the [YANG tree](https://www.pinglabz.com/yang-data-models-explained/) nests: ``` https:///restconf/data/ietf-interfaces:interfaces/interface=GigabitEthernet2 ^base^ ^----- model:container/list=key -----^ ``` That maps directly onto the tree: `ietf-interfaces` module, `interfaces` container, `interface` list, keyed by `GigabitEthernet2`. Once you can read a YANG tree, you can build any RESTCONF URL, because the URL is the tree path. ## Enabling It ``` R2(config)# restconf R2(config)# ip http secure-server R2(config)# crypto key generate rsa modulus 3072 ``` `restconf` turns on the feature, `ip http secure-server` provides the HTTPS transport (RESTCONF is HTTPS-only, and rightly so, a plaintext config API would be indefensible), and the RSA key backs the TLS. RESTCONF is served by the same DMI process set as NETCONF; `nginx` is the front end on port 443\. Confirm the stack the same way, with a real capture: ``` R2#show platform software yang-management process nginx : Running <- serves RESTCONF on 443 confd : Running ndbmand : Running <- the datastore backend ``` If `nginx` is up but requests return an error page rather than data, the datastore backend (`confd`/`ndbmand`) has not finished initializing, exactly the condition this virtual lab hit. The port is reachable, the web server answers, but it has no data to serve yet. ## Reading With GET A GET against the interface resource, with the JSON accept header: ``` import requests requests.packages.urllib3.disable_warnings() # lab only; verify certs in production url = "https://10.0.12.2/restconf/data/ietf-interfaces:interfaces/interface=GigabitEthernet2" hdr = {"Accept": "application/yang-data+json"} r = requests.get(url, headers=hdr, auth=("admin", "Cisco@123"), verify=False, timeout=15) print(r.status_code) print(r.json()) ``` ``` 200 OK (structure per the ietf-interfaces model) { "ietf-interfaces:interface": { "name": "GigabitEthernet2", "type": "iana-if-type:ethernetCsmacd", "enabled": true, "ietf-ip:ipv4": { "address": [ { "ip": "10.0.12.2", "netmask": "255.255.255.252" } ] } } } ``` The `Accept: application/yang-data+json` header is what asks for JSON (you can request XML instead). The response is a clean JSON object shaped by the model, ready to hand straight to the [JSON handling](https://www.pinglabz.com/python-json-encor-exam/) every automation script does. The `verify=False` is a lab convenience and a production sin, as the [API security article](https://www.pinglabz.com/rest-api-security-network-engineers/) stresses. ## Changing With PATCH To modify one field (the description) without disturbing the rest, PATCH a small JSON body: ``` url = "https://10.0.12.2/restconf/data/ietf-interfaces:interfaces/interface=GigabitEthernet2" hdr = {"Content-Type": "application/yang-data+json"} body = { "ietf-interfaces:interface": { "name": "GigabitEthernet2", "description": "Configured via RESTCONF" } } r = requests.patch(url, json=body, headers=hdr, auth=("admin","Cisco@123"), verify=False) print(r.status_code) # 204 on success ``` A successful PATCH returns **204 No Content**: the change applied and there is nothing to return. This is the code that confuses newcomers (an empty body looks like failure but is success), and it is why the [response codes article](https://www.pinglabz.com/rest-api-response-codes-network/) exists. PATCH merges, so only the fields you send change; PUT would replace the entire interface with your body, wiping anything you omitted. ## The Response Codes You Will See RESTCONF speaks standard HTTP status codes, and reading them is how you know what happened: 200 / 204GET succeeded with data / change succeeded, no content. Both are success. 401 / 403Bad credentials / not permitted. Re-auth or check your role. 404The YANG path does not exist. Usually a typo in the URL or a wrong interface name. 400Malformed body or a value the model rejects. The response body's `error-message` says what. Handling these properly is the whole subject of the [response codes article](https://www.pinglabz.com/rest-api-response-codes-network/): check every code, never blindly retry a 4xx, back off on 429 and 5xx, and read the body for the real reason. ## FAQ ### What port does RESTCONF use? HTTPS on 443 (RESTCONF is HTTPS-only). It is served by `nginx`, part of the same DMI process set as NETCONF. ### Why does RESTCONF return an error page instead of data? If `nginx` is running but you get an "unavailable" page, the datastore backend (`confd`/`ndbmand`) has not finished initializing. Check `show platform software yang-management process`. On virtual platforms this can take a long time or fail. ### PATCH or PUT? PATCH merges (changes only the fields you send). PUT replaces the entire resource with your body, removing anything you omit. Use PATCH for a targeted change; PUT when you mean to define the whole resource. ### Why did my PATCH return 204 with no body? 204 No Content is success for a change: it worked, and there is nothing to return. An empty body here is expected, not a failure. ### NETCONF or RESTCONF? RESTCONF for simplicity and quick tasks (familiar REST verbs, easy tooling). NETCONF when you need datastores, locking, or transactional changes. Same YANG models, different transport. ## Key Takeaways - RESTCONF maps **REST verbs onto YANG** over HTTPS (443): GET reads, PATCH merges, PUT replaces, DELETE removes. - The **URL is the path into the YANG tree**: `/restconf/data/ietf-interfaces:interfaces/interface=GigabitEthernet2` walks module, container, list, key. - Enable with `restconf` \+ `ip http secure-server` \+ a 3072-bit key. It shares the DMI process set with NETCONF; `nginx` is the front end. - A GET returns model-shaped JSON (with `Accept: application/yang-data+json`); a successful PATCH returns **204 No Content** (empty body = success). - Read the HTTP status codes: 200/204 success, 401/403 auth, 404 wrong path, 400 bad body. Never blindly retry a 4xx. - **Honest platform note:** on this virtual cat8000v the DMI datastore did not finish initializing, so nginx served an error page rather than data. Reachability and device config are real; the request/response structure is documentation-framed, never faked. Next: [NETCONF hands-on](https://www.pinglabz.com/netconf-cisco-ios-xe-hands-on/) for the transactional alternative, or the [Network Automation cluster guide](https://www.pinglabz.com/network-automation/). ### NETCONF on Cisco IOS XE Routers: A Hands-On Walkthrough URL: https://www.pinglabz.com/netconf-cisco-ios-xe-hands-on/ Last updated: 2026-07-12T02:58:37.000Z NETCONF is the protocol that replaced screen-scraping. Instead of sending CLI commands and parsing the text that comes back, you exchange structured XML over an SSH session, against a formal [YANG data model](https://www.pinglabz.com/yang-data-models-explained/). The result is automation that does not break when Cisco reformats a `show` command, and that can distinguish configuration from operational state cleanly. It is the foundation of model-driven management, and the Python library `ncclient` makes it approachable. This article walks through NETCONF on Cisco IOS XE, driven from a real Linux automation host. It extends the [Network Automation cluster guide](https://www.pinglabz.com/network-automation/). ## A Note on the Lab Consistent with this site's rule that captured output is real, here is the honest state of this lab. The automation host (a Debian VM with `ncclient`, `requests`, and `pyang` installed) is real, and it reaches the target router. The target is a Catalyst 8000v with the full NETCONF/YANG stack enabled, and its device-side state is captured live below. However, on this virtual platform the DMI datastore (the `confd`/`ndbmand` backend that serves NETCONF on port 830) did not finish initializing within the lab window, a known limitation of the virtualized DMI, so the port did not come up to serve requests. The `ncclient` request and XML response *structure* shown below is therefore documentation-style, drawn from Cisco's documented NETCONF behaviour and clearly framed as such. The device configuration and the `show` output confirming the stack is enabled are real captures. No XML is presented as a lab capture that was not one. ## What NETCONF Actually Does NETCONF (RFC 6241) has a few defining properties that make it different from SSH-and-scrape: Structured (XML) Everything is XML shaped by a YANG model. No text parsing, no fragile regex. Datastores Separate running, candidate, and startup config, so you can stage a change and commit it atomically. Config vs state A clean separation between configuration (what you set) and operational state (what is happening). Transaction-capable On platforms with a candidate datastore, changes can be all-or-nothing, not half-applied. It runs over SSH on **port 830** (distinct from the management SSH on 22), and the operations are a small, defined set: `get`, `get-config`, `edit-config`, `lock`, `unlock`, `commit`. ## Enabling It, and Confirming the Stack On IOS XE, NETCONF is one command, plus a crypto key for the SSH transport: ``` R2(config)# netconf-yang R2(config)# crypto key generate rsa modulus 3072 ``` Modern IOS XE (17.x) requires at least a 3072-bit key before it will bring up SSH, which NETCONF depends on. Once enabled, the router runs a set of background processes (the DMI, Data Model Interface) that actually serve NETCONF and RESTCONF. You can confirm they are up, and this is a real capture: ``` R2#show platform software yang-management process confd : Running nesd : Running syncfd : Running ncsshd : Running dmiauthd : Running nginx : Running ndbmand : Running pubd : Running ``` Each process has a job: `ncsshd` is the NETCONF SSH server (port 830), `nginx` serves RESTCONF (port 443), `confd` and `ndbmand` are the datastore backend, `pubd` handles telemetry. And the NETCONF-specific status, also real: ``` R2#show netconf-yang status netconf-yang: enabled netconf-yang ssh port: 830 netconf-yang side-effect-sync: enabled netconf-yang ssh hostkey algorithms: rsa-sha2-256,rsa-sha2-512,ssh-rsa ``` When NETCONF is not working, this is the first place to look: is the feature enabled, are the DMI processes running, and is the SSH key present. If `ncsshd` is not running or the datastore backend has not finished initializing, port 830 will not accept connections even though the configuration looks correct, which is exactly the failure this lab hit on a virtual platform. ## The Hello Exchange Every NETCONF session opens with a **hello** exchange in which both sides advertise their capabilities, the list of YANG models and features they support. This is how a client discovers what a device can do. Using `ncclient`: ``` from ncclient import manager m = manager.connect( host="10.0.12.2", port=830, username="admin", password="Cisco@123", hostkey_verify=False, device_params={"name": "iosxe"}) print("Session id:", m.session_id) for cap in m.server_capabilities: if "ietf-interfaces" in cap: print(cap) ``` The `server_capabilities` list is the device telling you every model it supports. Filtering it (as above) is how you confirm a model like `ietf-interfaces` is available before you try to use it. The `device_params={"name": "iosxe"}` tells ncclient to use the IOS XE dialect, which matters for some operations. ## Reading Configuration: get-config To read an interface's configuration, you send a `get-config` with a filter that names the subtree you want (from the [YANG model](https://www.pinglabz.com/yang-data-models-explained/)). The request and the shape of the reply: ``` filter = """ GigabitEthernet2 """ reply = m.get_config(source="running", filter=filter) print(reply) ``` ``` GigabitEthernet2 true
10.0.12.2255.255.255.252
``` Notice the filter names the exact subtree, and the reply returns exactly that subtree as structured XML. No `show` output, no parsing, no ambiguity. The `source="running"` reads the running datastore; you could equally read `startup`. ## Changing Configuration: edit-config Writing is `edit-config`: you send the XML for the new state, targeting a datastore. Changing an interface description: ``` config = """ GigabitEthernet2 Configured via NETCONF """ reply = m.edit_config(target="running", config=config) # reply contains on success ``` On a platform with a candidate datastore you would target `candidate`, then `commit`, giving you an atomic all-or-nothing change. IOS XE historically writes directly to `running`. The success indicator is an `` element in the reply; an error returns an `` with a type and message, the XML analogue of the [HTTP error codes](https://www.pinglabz.com/rest-api-response-codes-network/) RESTCONF uses. ## NETCONF or RESTCONF? They target the same YANG models over different transports. The short guidance: NETCONFXML over SSH (830). Richer: datastores, locking, transactions. Choose for complex, transactional config. RESTCONFJSON/XML over HTTPS (443). Simpler, familiar REST verbs. Choose for straightforward reads and edits, and easy tooling. The [RESTCONF article](https://www.pinglabz.com/restconf-cisco-ios-xe/) covers the HTTP side. For most quick tasks RESTCONF is easier; for transactional multi-step configuration, NETCONF's datastore model is worth the extra ceremony. ## FAQ ### What port does NETCONF use? TCP 830 over SSH, separate from management SSH on 22\. If 830 refuses connections, check that `netconf-yang` is enabled, the DMI processes (`ncsshd` especially) are running, and an SSH key exists. ### Why is port 830 not responding even though netconf-yang is enabled? The DMI backend (`confd`/`ndbmand`) may not have finished initializing, or the SSH key is missing. Check `show platform software yang-management process` and `show netconf-yang status`. On virtual platforms the datastore can be slow or fail to initialize. ### Do I need a candidate datastore for transactions? For true atomic commit, yes. IOS XE historically writes to running directly. Check the device's advertised capabilities in the hello exchange for `candidate` support. ### How do I know which models a device supports? The hello exchange. `m.server_capabilities` in ncclient lists every model and feature the device advertises. Filter it to confirm a model before using it. ### NETCONF or RESTCONF for a simple change? RESTCONF is simpler for a straightforward read or single edit. NETCONF is better when you need locking, a candidate datastore, or transactional multi-step changes. ## Key Takeaways - NETCONF exchanges **structured XML over SSH (port 830)** against YANG models, replacing fragile screen-scraping. - It separates **datastores** (running/candidate/startup) and **config vs operational state**, enabling atomic changes on capable platforms. - Enable with `netconf-yang` plus a 3072-bit RSA key. Confirm the **DMI processes** with `show platform software yang-management process` \- this is the first troubleshooting step. - Every session opens with a **hello** exchange advertising capabilities; `ncclient`'s `server_capabilities` lists supported models. - Core operations: `get-config` (read a filtered subtree), `edit-config` (write structured XML), with `` or `` as the result. - **Honest platform note:** on this virtual cat8000v the DMI datastore did not finish initializing, so port 830 never served requests. The device state shown is real; the XML request/response structure is documentation-framed, never faked as captured. Next: [RESTCONF on Cisco IOS XE](https://www.pinglabz.com/restconf-cisco-ios-xe/), or the [Network Automation cluster guide](https://www.pinglabz.com/network-automation/). ### EEM Applets on Cisco IOS XE: Automate Config, Troubleshooting, and Data Collection URL: https://www.pinglabz.com/eem-applets-cisco-ios-xe/ Last updated: 2026-07-12T02:58:37.000Z Before Python and NETCONF, before Ansible and Terraform, Cisco routers already had a way to automate themselves: the Embedded Event Manager. EEM runs entirely on the device, watches for events, and takes actions when they occur, no external server, no controller, no network dependency. It is the automation that keeps working when the automation server is exactly what went down, and it remains one of the most practical tools in the CCNP toolbox. This article builds three EEM applets and fires all of them for real. It opens the [Network Automation cluster guide](https://www.pinglabz.com/network-automation/). ## First: This Runs On the Box The thing that makes EEM different from every other automation tool is where it lives. Ansible, NETCONF, and REST APIs all drive the device from *outside*. EEM runs *inside* the device, in IOS itself. That has two consequences worth internalising: It survives isolation When a link fails and the box is cut off from your automation server, EEM keeps running. It is local. This is exactly when you most want automation. It reacts in real time No polling interval, no server round-trip. When the event fires, EEM acts immediately, capturing state before it changes. The canonical use case is diagnostic capture: something breaks at 3am, and by the time an engineer looks, the transient state that would have explained it is gone. EEM catches it in the instant it happens. ## The Anatomy of an Applet An EEM applet is an **event** plus a set of **actions**. The event says "watch for this"; the actions say "do these things when it happens." The applet form (there is also a more powerful Tcl-scripted form) is readable enough to write on the fly: ``` event manager applet NAME event action 1.0 action 2.0 ``` The action numbers (1.0, 2.0) set the execution order, and leaving gaps lets you insert steps later without renumbering. EEM offers many event detectors; the ones you will actually use: event syslogFire when a syslog message matches a regex. The most versatile detector - anything the box logs can trigger an applet. event timerFire on a schedule (watchdog = every N seconds, cron = at a time). For periodic collection or housekeeping. event snmpFire when an SNMP OID crosses a threshold. For metric-driven reactions (CPU, memory, counters). event cliFire when someone runs a matching command. For guardrails and command interception. ## Applet 1: Capture Diagnostics When a Link Fails The classic. When Ethernet0/1 goes down, grab the state that will help diagnose it, before the box reconverges and hides what happened: ``` R1(config)# event manager applet LINK-DOWN-CAPTURE R1(config-applet)# event syslog pattern "Interface Ethernet0/1, changed state to down" R1(config-applet)# action 1.0 syslog msg "EEM: Et0/1 went down - collecting diagnostics" R1(config-applet)# action 2.0 cli command "enable" R1(config-applet)# action 3.0 cli command "show ip interface brief | append flash:eem-linkdown.txt" R1(config-applet)# action 4.0 cli command "show ip ospf neighbor | append flash:eem-linkdown.txt" R1(config-applet)# action 5.0 syslog msg "EEM: diagnostics saved to flash:eem-linkdown.txt" ``` Two details matter. `action 2.0 cli command "enable"` is required because `cli command` actions start in user mode; you must elevate before running privileged commands. And `| append` writes each show output to a growing file, so you accumulate a timestamped record rather than overwriting. (A platform note: the IOL-XE lab device uses a `unix:` filesystem rather than `flash:`, so `more flash:eem-linkdown.txt` fails there. On real hardware use `flash:` or `bootflash:`. The applet logic and the syslog evidence below are identical.) ## Applet 2: Log Every Configuration Change A lightweight audit trail. Whenever anyone leaves configuration mode (which logs a `CONFIG_I` message), record it: ``` R1(config)# event manager applet CONFIG-CHANGE-LOG R1(config-applet)# event syslog pattern "CONFIG_I" R1(config-applet)# action 1.0 syslog msg "EEM: a configuration change was made on R1" ``` Paired with `archive` config logging, this is a poor engineer's change-tracking system that needs no external tooling. ## Applet 3: Scheduled Heartbeat A timer applet that fires every 60 seconds, standing in for any periodic task (scheduled data collection, a keepalive, a housekeeping job): ``` R1(config)# event manager applet HELLO-SCHEDULED R1(config-applet)# event timer watchdog time 60 R1(config-applet)# action 1.0 syslog msg "EEM: scheduled watchdog fired - heartbeat" ``` ## Firing All Three, For Real Shut the interface and watch every applet respond. Here is the actual syslog from the box: ``` R1#show logging | include EEM %HA_EM-6-LOG: CONFIG-CHANGE-LOG: EEM: a configuration change was made on R1 %HA_EM-6-LOG: LINK-DOWN-CAPTURE: EEM: Et0/1 went down - collecting diagnostics %HA_EM-6-LOG: LINK-DOWN-CAPTURE: EEM: diagnostics saved to flash:eem-linkdown.txt %HA_EM-6-LOG: CONFIG-CHANGE-LOG: EEM: a configuration change was made on R1 ``` Read the sequence. Entering config mode to type `shutdown` fired CONFIG-CHANGE-LOG. The interface going down fired LINK-DOWN-CAPTURE, which logged its start, ran the show commands, and logged completion. Every applet did exactly what it was told, driven by real events on the box. EEM keeps its own event history, the authoritative record of what fired and whether it succeeded: ``` R1#show event manager history events No. Job Id Proc Status Time of Event Event Type Name 1 1 Actv success Sun Jul12 02:33:57 2026 syslog applet: CONFIG-CHANGE-LOG 2 2 Actv success Sun Jul12 02:34:09 2026 syslog applet: CONFIG-CHANGE-LOG 3 3 Actv success Sun Jul12 02:34:10 2026 syslog applet: CONFIG-CHANGE-LOG 4 4 Actv success Sun Jul12 02:34:13 2026 syslog applet: LINK-DOWN-CAPTURE 5 5 Actv success Sun Jul12 02:34:21 2026 syslog applet: CONFIG-CHANGE-LOG ``` Every event `success`. When you are debugging an applet that is not behaving, this history (and `debug event manager action cli`) is where you look first: it tells you whether the event even fired, which is usually the problem. ## Where EEM Bites - **Regex is literal.** The `event syslog pattern` is a regex against the exact log string. If your pattern does not match the real message character for character (including interface naming), the applet silently never fires. Test with the actual log line. - **You must `enable` before privileged commands.** A `cli command` action starts in user EXEC. Forgetting `action X cli command "enable"` means your show commands fail silently. - **Applets can loop.** An applet whose action triggers its own event can fire itself repeatedly. Use `maxrun` and think about self-triggering. - **Rate limits exist for a reason.** A syslog applet on a very common message can fire constantly and load the CPU. Match specifically. - **It is device-local automation, not orchestration.** EEM is superb for reactive, on-box tasks. It is not a substitute for centralised config management across a fleet, which is where [NETCONF](https://www.pinglabz.com/netconf-cisco-ios-xe-hands-on/) and controllers come in. ## Where EEM Fits in the Toolbox EEM and the model-driven tools (NETCONF, RESTCONF) are complementary, not competing. EEM is *reactive and local*: it responds to events on the box in real time and keeps working in isolation. NETCONF and RESTCONF are *proactive and remote*: they let a central system push and read configuration at scale. A mature automation practice uses both. The rest of this cluster covers the remote, model-driven side: [NETCONF](https://www.pinglabz.com/netconf-cisco-ios-xe-hands-on/), [RESTCONF](https://www.pinglabz.com/restconf-cisco-ios-xe/), and the [YANG models](https://www.pinglabz.com/yang-data-models-explained/) underneath them. ## FAQ ### Why does my applet never fire? Almost always a pattern mismatch. The `event syslog pattern` regex must match the real log message exactly. Trigger the event manually, copy the exact log line, and build your pattern from it. Check `show event manager history events` to confirm whether it fired at all. ### Do I need to enable before running show commands? Yes. `cli command` actions begin in user EXEC mode. Add `action X cli command "enable"` before any privileged command. ### Applet or Tcl script? Applets for straightforward event-action logic (most cases). Tcl scripts when you need loops, complex logic, or data manipulation the applet syntax cannot express. Start with applets. ### Can EEM make configuration changes? Yes, via `cli command` actions in config mode, which is how self-healing applets work. Be careful about loops and test thoroughly. ### Does EEM work when the device is isolated? Yes, and that is its signature strength. EEM runs entirely on the device, so it keeps reacting even when the box is cut off from every external system. ## Key Takeaways - EEM is **on-box automation**: it watches for events and acts, with no external server, so it keeps working when the box is isolated, exactly when you need it. - An applet is an **event** (syslog, timer, snmp, cli) plus numbered **actions**. Its most common job is capturing transient diagnostic state the instant something breaks. - The lab fired all three applets live: LINK-DOWN-CAPTURE on the interface event, CONFIG-CHANGE-LOG on every change, and a scheduled timer, all confirmed `success` in the event history. - Two constant gotchas: the syslog **regex must match exactly**, and you must `enable` before privileged `cli command` actions. - EEM is **reactive and local**; NETCONF/RESTCONF are **proactive and remote**. Use both. Next: [NETCONF hands-on](https://www.pinglabz.com/netconf-cisco-ios-xe-hands-on/), or the [Network Automation cluster guide](https://www.pinglabz.com/network-automation/). ### VRF, VLAN, VXLAN, LISP: Choosing the Right Segmentation Layer URL: https://www.pinglabz.com/vrf-vlan-vxlan-lisp-segmentation/ Last updated: 2026-07-12T02:30:09.000Z By the time you have read this cluster, you have four different ways to segment a network: VRFs, VLANs, VXLAN, and LISP. They overlap, they are often used together, and choosing the wrong one (or reaching for a fabric when a VLAN would do) is a real and expensive mistake. This article is the decision framework: what each one actually separates, at which layer, and when to use it. It ties together the whole [Network Virtualization and Overlays cluster](https://www.pinglabz.com/network-virtualization/). ## The Four, at a Glance VLAN Layer2 Separates broadcast domains on a switch. The oldest, simplest, most local tool. Scale\~4094 VRF Layer3 Separates routing tables. Two VRFs can use the same IP space and never see each other. ScaleMany VXLAN Layer2 over 3 Stretches an L2 segment across a routed fabric. Solves the VLAN ceiling and scale. Scale16M LISP Layer3 mapping Separates identity from location. The control plane for mobility and fabrics. ScaleVery large The first insight: these are not competitors on the same layer. A VLAN separates Layer 2\. A VRF separates Layer 3\. VXLAN is a transport that carries Layer 2 (or 3) across a Layer 3 fabric. LISP is a control plane that maps endpoints to locations. In a modern fabric you use *all four at once*, each doing its own job. ## VLAN: Start Here, Stay Here If You Can A [VLAN](https://www.pinglabz.com/vlans-layer-2-switching/) separates broadcast domains within a switched network. It is the right answer for the vast majority of segmentation needs: separating voice from data, guests from staff, one department from another, within a building or a campus. When a VLAN is enough, use a VLAN. The mistake this whole cluster can accidentally encourage is reaching for a fabric because it is interesting, when the requirement is "keep the guest Wi-Fi off the corporate network," which a VLAN and an ACL solve completely. VLANs run out of room at two points: the \~4094 ID ceiling (real multi-tenancy) and the scale limits of large flat Layer 2 (big failure domains, spanning-tree fragility). Below those limits, VLANs are simpler, cheaper, and easier to operate than anything else here. ## VRF: Separation at Layer 3 A [VRF](https://www.pinglabz.com/vrf-lite-configuration-cisco-ios-xe/) gives a router multiple independent routing tables. The killer feature is that two VRFs can use overlapping IP address space and remain completely isolated: a route in one is invisible to the other, so no traffic crosses without an explicit leak. Reach for a VRF when the separation you need is at Layer 3: keeping two customers' routing separate on shared infrastructure, isolating a management network, or carving a guest network that must not route to the corporate one even though both are Layer 3\. VRF-Lite does this hop by hop on a single or small set of routers; when the VRFs must scale across a whole backbone, you promote to [MPLS L3VPN](https://www.pinglabz.com/mpls-l3vpn/) (VRFs carried by MP-BGP) or a VXLAN-EVPN fabric (VRFs carried as L3VNIs). The VRF is the unit of macro-segmentation everywhere from a two-router VRF-Lite setup to an SD-Access Virtual Network. ## VXLAN: When VLANs Run Out of Room [VXLAN](https://www.pinglabz.com/vxlan-deep-dive/) is the answer to two specific VLAN limits: you need more than 4094 segments (real multi-tenancy, a service provider or large enterprise data centre), or you need to stretch a Layer 2 segment across a routed network without the fragility of large flat Layer 2 (VM mobility across a fabric). The signal that you have outgrown VLANs and need VXLAN: you are trying to stretch VLANs across your whole data centre for VM mobility, or you have hit the segment ceiling, or spanning tree across a big flat domain has become an operational hazard. If none of those apply, VXLAN is complexity you do not need. When they do apply, VXLAN with [BGP EVPN](https://www.pinglabz.com/vxlan-bgp-evpn-explained/) is the standard, and it carries VLANs (as L2VNIs) and VRFs (as L3VNIs) across the fabric, so it does not replace them, it transports them. ## LISP: When Location Must Change but Identity Must Not [LISP](https://www.pinglabz.com/lisp-explained/) is different in kind from the other three: it is not a segmentation construct, it is a mapping control plane. It separates *who* (the EID) from *where* (the RLOC), so an endpoint can move and keep its identity while its location changes underneath. You use LISP when the requirement is mobility or scalable multihoming: campus-wide roaming where a device keeps its policy wherever it plugs in, ingress traffic engineering without polluting global BGP, or, most commonly, as the control plane of [Cisco SD-Access](https://www.pinglabz.com/sd-access-architecture/). You rarely deploy LISP by hand for its own sake in an enterprise; you encounter it as the machinery inside the fabric. But understanding it explains why the fabric can move an endpoint seamlessly. ## The Decision, Compressed Separate broadcast domains in a building**VLAN.** Do not overthink it. Isolate routing / overlapping IP space**VRF** (VRF-Lite locally, MPLS L3VPN or EVPN L3VNI at scale). More than 4094 segments, or L2 across a routed DC**VXLAN** (with BGP EVPN). Endpoint mobility / campus fabric**LISP** (usually as the SD-Access control plane). A modern fabric**All four together:** VLANs at the edge, VRFs for tenants, VXLAN for transport, LISP or EVPN for control. ## FAQ ### Do VXLAN and VLANs compete? No. VXLAN carries VLANs (mapped to L2VNIs) across a fabric. Locally you still use VLANs; VXLAN extends them beyond the reach and scale of a physical VLAN. ### Is a VRF the same as a VLAN? No. A VLAN separates Layer 2 (broadcast domains). A VRF separates Layer 3 (routing tables). You often pair them: a VLAN per subnet, a VRF grouping subnets into a tenant. ### When do I actually need LISP? Rarely by hand. You encounter it as the control plane of SD-Access. Standalone, it is for endpoint mobility and scalable multihoming. ### Can I just use VLANs and VRFs forever? For many networks, yes. VXLAN and fabrics earn their complexity at scale (large multi-tenant data centres, big campuses with mobility and segmentation needs). Below that, VLANs and VRFs are the right, simpler answer. ### What is the biggest mistake here? Reaching for a fabric because it is modern when a VLAN and an ACL solve the actual requirement. Match the tool to the need, not to the trend. ## Key Takeaways - The four operate at **different layers** and are used together, not instead of each other: VLAN (L2), VRF (L3), VXLAN (L2-over-L3 transport), LISP (identity/location mapping). - **VLAN** for broadcast-domain separation in a building. When it is enough, use it. Ceiling: \~4094 and flat-L2 scale. - **VRF** for Layer 3 isolation and overlapping IP space. VRF-Lite locally; MPLS L3VPN or EVPN L3VNI at scale. - **VXLAN** when you exceed 4094 segments or need L2 across a routed fabric (VM mobility). It carries VLANs and VRFs, does not replace them. - **LISP** when location must change but identity must not: mobility, multihoming, and the SD-Access control plane. - A modern fabric uses **all four at once**. The mistake is reaching for a fabric when a VLAN would do. This closes the Network Virtualization cluster. Back to the [cluster guide](https://www.pinglabz.com/network-virtualization/) for the full reading order. ### Hypervisors, Virtual Switches, and VMs for Network Engineers URL: https://www.pinglabz.com/hypervisors-virtual-switching-network-engineers/ Last updated: 2026-07-12T02:30:09.000Z A network engineer in 2026 spends a surprising amount of time on networks that have no physical cables. The server team's hypervisor has a virtual switch inside it that your VLANs terminate on. The cloud VPC is entirely virtual. The container platform has its own overlay. If you think the network stops at the physical switch port, you are missing half of where the packets actually go, and half of where the problems actually live. This article covers hypervisors, virtual switching, and the virtualization concepts a network engineer needs, framed for people who know networking but not necessarily the compute side. It extends the [Network Virtualization cluster guide](https://www.pinglabz.com/network-virtualization/). ## The Hypervisor, From a Network Point of View A hypervisor runs virtual machines on physical hardware. There are two types, and the distinction is worth knowing: Type 1 (bare-metal) Runs directly on hardware. VMware ESXi, KVM, Hyper-V. This is what runs production data centres. Type 2 (hosted) Runs on top of a host OS. VirtualBox, VMware Workstation. This is what runs on your laptop, including the CML labs behind this site. From the network's perspective, the important thing is not the VMs themselves but what sits between them and the physical NIC: a **virtual switch**. Every packet a VM sends hits that virtual switch first, and only some of them ever leave the physical server at all. ## The Virtual Switch Is a Real Switch Inside the hypervisor is a software switch (VMware vSwitch/vDS, the Linux bridge, Open vSwitch) that behaves like a physical access/distribution switch, with all the same concepts you already know: Port groups / VLANsVMs attach to port groups, which map to [VLANs](https://www.pinglabz.com/vlans-layer-2-switching/). The vSwitch tags frames exactly like a physical access port. Uplinks (pNICs)The physical NICs are the vSwitch's uplinks to the real network. Usually configured as a trunk on the physical switch side. The east-west blind spotTwo VMs in the same port group on the same host talk to each other *inside* the vSwitch. That traffic never touches your physical switch, so your monitoring never sees it. That last row is the single most important thing for a network engineer to internalise. A large fraction of data centre traffic is east-west (server to server), and much of it between VMs on the same host never leaves the hypervisor. Your NetFlow, your SPAN, your ACLs on the physical switch, none of them see it. This is why virtual firewalls and hypervisor-level microsegmentation exist: to enforce policy on traffic the physical network never touches. One important difference from a physical switch: a vSwitch typically does **not** run spanning tree and does not learn MACs the way a physical switch does. It knows exactly which MACs are behind each virtual port because it created those ports, so it does not need to learn or worry about loops in the same way. That is a feature (no STP complexity) and an occasional surprise (behaviour differs from a physical switch). ## VM Mobility and Why Overlays Exist The capability that reshaped data centre networking is live migration: moving a running VM from one physical host to another with no downtime (vMotion, live migration). For this to work, the VM must keep its IP and MAC after the move, which means the Layer 2 segment it lives on must exist on both the source and destination hosts. In a traditional network, that means stretching VLANs across every host that might ever run the VM, which recreates all the large-flat-Layer-2 problems VXLAN was invented to solve. This is precisely why data centres moved to [VXLAN](https://www.pinglabz.com/vxlan-deep-dive/) overlays: the VM's segment (a VNI) can exist anywhere in the fabric without stretching physical VLANs, so the VM can migrate anywhere and keep its addressing. The overlay decouples the VM's network from the physical topology, exactly the decoupling this whole cluster is about. ## Containers Are a Different Animal VMs virtualize the hardware; containers virtualize the operating system. A container shares the host kernel and is far lighter, and its networking model is different again: - Each container gets a virtual interface, usually connected to a bridge on the host (or an overlay). - Container platforms (Kubernetes) run their own **overlay network** (Flannel, Calico, Cilium) that gives every container an IP and routes between them, often using VXLAN or similar encapsulation under the hood. - Service discovery and load balancing happen at the platform layer, not the physical network. A "service" IP may not correspond to any single container. For a network engineer, the practical reality is that container traffic is even further abstracted from the physical network than VM traffic. The physical network provides IP connectivity between hosts; everything above that (which container talks to which, and how) is the platform's overlay, which you may not directly control or even see. The [virtualization fundamentals article](https://www.pinglabz.com/virtualization-fundamentals-vms-containers-vrfs/) covers the VM-vs-container distinction at the CCNA level; the takeaway here is that the network's job shrinks to "provide reliable IP transport between hosts" and the interesting L2/L3 decisions move into software. ## What This Means for the Network Team The practical consequences of all this virtualization for how you design and operate: - **Trunk to the hypervisor, do not access-port it.** A physical port to an ESXi host carries many VLANs (many port groups). It is a trunk, and the VLAN allowed list must include everything the host runs. - **You cannot see east-west intra-host traffic** from the physical network. If you need visibility or enforcement there, it has to happen in the hypervisor (virtual taps, distributed firewalls, hypervisor NetFlow). - **The MTU story matters more.** Overlays (VXLAN in the hypervisor or container platform) add encapsulation overhead. If the physical underlay MTU is not raised, you get mysterious performance problems that look like application bugs. - **Coordination with the compute team is not optional.** The network now extends into their hypervisor. A VLAN change, an MTU change, or a new port group is a joint operation. The old clean handoff at the switch port is gone. ## FAQ ### Why can't I see traffic between two VMs on the same host? Because it never leaves the host. The virtual switch forwards it internally. To see or enforce policy on it, you need hypervisor-level tooling (a distributed virtual switch with monitoring, a virtual firewall), not the physical network. ### Is a virtual switch a real switch? Functionally yes for VLANs and forwarding, but it typically does not run spanning tree and knows its MAC-to-port mappings by construction rather than learning them. Treat it as an access/distribution layer implemented in software. ### How should the physical port to a hypervisor be configured? As a trunk (802.1Q) carrying all the VLANs the host's port groups use, usually with an allowed-VLAN list and often in an EtherChannel to the host's teamed NICs. ### Do containers use VLANs? Usually not directly. Container platforms run their own overlay (often VXLAN-based) and assign IPs from their own address space. The physical network provides host-to-host IP transport; the container networking lives above it. ### Why does VM mobility need an overlay? A migrating VM must keep its IP and MAC, so its Layer 2 segment must exist on the destination host. An overlay (VXLAN) lets that segment exist anywhere in the fabric without stretching physical VLANs everywhere. ## Key Takeaways - Every VM's traffic hits a **virtual switch** in the hypervisor first. It behaves like an access/distribution switch (port groups = VLANs, pNICs = trunk uplinks) but usually without spanning tree. - **East-west traffic between VMs on the same host never leaves the hypervisor**, so the physical network cannot see or filter it. Enforcement there needs hypervisor-level tooling. - **VM live migration** requires the L2 segment on both hosts, which is exactly why data centres adopted [VXLAN overlays](https://www.pinglabz.com/vxlan-deep-dive/). - **Containers** virtualize the OS, are lighter than VMs, and run their own platform overlay; the physical network shrinks to providing host-to-host IP transport. - Practical rules: trunk to hypervisors, raise underlay MTU for overlays, accept the east-west visibility gap, and coordinate closely with the compute team, the network now extends into their kit. Next: [choosing the right segmentation layer](https://www.pinglabz.com/vrf-vlan-vxlan-lisp-segmentation/), or the [Network Virtualization cluster guide](https://www.pinglabz.com/network-virtualization/). ### SD-Access and the Traditional Campus: Interoperability and Migration URL: https://www.pinglabz.com/sd-access-traditional-campus-interop/ Last updated: 2026-07-12T02:30:08.000Z Nobody builds a greenfield campus. You have an existing network with real users, real applications, and real uptime requirements, and any move to SD-Access has to happen alongside it, not instead of it overnight. The interesting engineering in SD-Access adoption is rarely the fabric itself; it is the boundary where the fabric meets everything that is not yet fabric. This article covers how SD-Access interoperates with a traditional campus and the realistic migration paths. It extends the [SD-Access architecture article](https://www.pinglabz.com/sd-access-architecture/) and the [Network Virtualization cluster guide](https://www.pinglabz.com/network-virtualization/). ## Everything Happens at the Border The fabric is a bubble. Inside it, traffic is [VXLAN-encapsulated](https://www.pinglabz.com/vxlan-deep-dive/), endpoints are tracked by [LISP](https://www.pinglabz.com/lisp-explained/), and policy is enforced by group tags. Outside it, the world is ordinary IP: the WAN, the data centre, the internet, and the parts of the campus you have not migrated yet. The Border Node is the airlock between the two, and it does the translation. There are two flavours of border, and the distinction matters: Border to a known network Connects the fabric to a network whose specific routes it knows (a data centre, another campus). It imports and exports specific prefixes, preserving segmentation via VRFs. Border to an unknown network The default exit to everything else (the internet). The fabric sends anything it does not have a fabric mapping for toward this border, like a default route. The border is where the fabric's internal constructs (VNs, SGTs, LISP mappings) get translated into things the outside understands (VRFs, prefixes, ordinary routing). Getting this translation right, especially preserving segmentation across the boundary, is the hard part of any SD-Access design. ## Carrying Segmentation Across the Boundary Inside the fabric you have macro-segmentation (Virtual Networks) and micro-segmentation (SGTs). The moment traffic leaves the fabric, those constructs do not exist unless you deliberately carry them onward. Two mechanisms: - **VRF-Lite hand-off at the border.** Each fabric Virtual Network maps to a VRF, and the border extends those VRFs to the next-hop device using [VRF-Lite](https://www.pinglabz.com/vrf-lite-configuration-cisco-ios-xe/) (subinterfaces or 802.1Q per VN). This preserves macro-segmentation into the non-fabric network. It is the same VRF construct from the routing cluster, doing the same job at the fabric edge. - **SGT propagation (inline or SXP).** To preserve micro-segmentation, the group tag must travel with the traffic. Inline SGT carries it in the frame to TrustSec-capable next hops; SXP (SGT Exchange Protocol) carries IP-to-SGT bindings out-of-band to devices that cannot do inline tagging. Without one of these, group policy stops at the border. The common mistake is building a beautiful segmented fabric and then dumping everything into a single VRF at the border, collapsing all that separation the moment traffic leaves. If segmentation matters inside the fabric, it has to be carried across the border deliberately. ## Migration Strategies You do not flip a campus to fabric overnight. The realistic approaches: Parallel (greenfield-in-place)Build the fabric alongside the existing network, migrate users building-by-building or floor-by-floor, and bridge the two at the border. The safest and most common path. Layer 2 border hand-offExtend specific VLANs from the legacy network into the fabric during transition, so a subnet can span both worlds while endpoints migrate. Temporary by design. Fabric-in-a-box for small sitesCollapse the edge, control, and border roles onto one or two switches at a small site, so branches can join the SD-Access model without a full multi-node fabric. The parallel approach is what most organisations use: stand up the fabric, connect it to the legacy network through the border, and move users across in waves. During the transition, a user on the fabric and a user still on the legacy network reach each other through the border, which routes between the fabric VNs and the legacy VLANs. As migration completes, the legacy footprint shrinks until it is gone or reduced to a few non-fabric-capable devices behind the border. ## What Does Not Move Cleanly Honesty about the friction points, because they drive real project timelines: - **Non-fabric-capable hardware.** SD-Access needs Catalyst 9000-class switches at the fabric edge. Older access switches cannot be fabric edges; they either get replaced or sit behind the border as a legacy island. - **Multicast and legacy protocols.** Applications that assume a flat Layer 2 domain, certain multicast designs, or non-IP protocols need careful handling across the fabric boundary. Test them explicitly. - **IP addressing assumptions.** The fabric's anycast-gateway and host-mobility model changes how addressing behaves. Applications hard-coded to specific gateway MACs or subnet layouts can surprise you. - **The operational shift.** The biggest migration cost is often not technical. Teams used to CLI must learn to operate through Catalyst Center, and the troubleshooting model changes. Budget for the learning curve. ## FAQ ### Can SD-Access and a traditional network coexist? Yes, and they almost always must during migration. The Border Node is the interconnection point, routing between fabric Virtual Networks and the legacy network. ### How do I keep segmentation when traffic leaves the fabric? Map each Virtual Network to a VRF and extend it with VRF-Lite at the border for macro-segmentation, and propagate SGTs (inline or via SXP) for micro-segmentation. Otherwise segmentation collapses at the boundary. ### Do I have to replace all my switches? The fabric edge needs capable hardware (Catalyst 9000). Non-capable switches can remain behind the border as a legacy segment, but they cannot be fabric edges. This drives the hardware refresh side of most projects. ### What is the safest migration path? Parallel: build the fabric alongside the existing network, bridge at the border, and migrate users in waves. Avoid big-bang cutovers. ### Can a small branch use SD-Access? Yes, via Fabric-in-a-Box, which collapses the fabric roles onto one or two switches so a small site joins the same policy model without a full fabric. ## Key Takeaways - SD-Access adoption is dominated by the **border**: where the fabric meets the WAN, data centre, internet, and un-migrated campus. - The **Border Node** translates fabric constructs (VNs, SGTs, LISP) into ordinary routing (VRFs, prefixes) for the outside world. - Preserve segmentation across the boundary with **VRF-Lite** (macro) and **SGT propagation** (micro), or it collapses the moment traffic leaves. - Migrate in **parallel**: build alongside, bridge at the border, move users in waves. Layer 2 hand-off and Fabric-in-a-Box are transitional tools. - Friction points: non-fabric-capable hardware, legacy L2/multicast assumptions, addressing dependencies, and the CLI-to-controller operational shift. Next: [choosing the right segmentation layer](https://www.pinglabz.com/vrf-vlan-vxlan-lisp-segmentation/), or the [Network Virtualization cluster guide](https://www.pinglabz.com/network-virtualization/). ### Cisco SD-Access Architecture: Fabric, Control Plane, and Policy URL: https://www.pinglabz.com/sd-access-architecture/ Last updated: 2026-07-12T02:30:08.000Z SD-Access is Cisco's intent-based campus fabric, and it is where the two protocols from earlier in this cluster stop being academic. LISP is its control plane, VXLAN is its data plane, and Catalyst Center is the automation and policy layer that operates the whole thing so you never touch the underlying CLI. If you understand LISP and VXLAN, you already understand most of how SD-Access works; the rest is the policy model and the operational shift. This article explains the SD-Access architecture, its roles, and how the pieces fit. It is a concept article by necessity (SD-Access cannot be built in a lab without Catalyst Center), and it extends the [Network Virtualization cluster guide](https://www.pinglabz.com/network-virtualization/). ## The Problem SD-Access Addresses A traditional campus is a pile of VLANs, subnets, access lists, and manual per-switch configuration. Segmentation is done with VLANs and ACLs that must be maintained by hand on every device. Moving a user or a device between floors often means re-IPing. Applying a security policy consistently across a large campus is genuinely hard, and proving it was applied is harder. SD-Access changes the model in three ways: Location-independent identity A user keeps their access and policy wherever they plug in, because the fabric decouples identity from location (this is the LISP EID/RLOC split). Policy by group, not subnet Security is expressed between groups of users (SGTs), not between IP subnets, so it survives IP changes and moves. Automated operation Catalyst Center provisions the fabric from intent. You describe what you want; it renders the device configuration. ## The Two Planes You Already Know Strip away the branding and SD-Access is LISP plus VXLAN: Control plane = LISPThe fabric maps each endpoint (EID) to the fabric edge switch it lives behind (RLOC), exactly as in the [LISP article](https://www.pinglabz.com/lisp-explained/). Endpoint moves update the mapping, not the address. Data plane = VXLANTraffic between fabric nodes is [VXLAN-encapsulated](https://www.pinglabz.com/vxlan-deep-dive/), carrying the segment (VNI) and the group tag (SGT) across the fabric. The one genuinely new ingredient is the **policy plane**: Scalable Group Tags (SGTs) from Cisco TrustSec, carried inside the VXLAN header, that let the fabric enforce group-based policy independent of IP addressing. That is the piece that makes "policy by group, not subnet" work. ## The Fabric Roles SD-Access defines a handful of node roles. Map them onto the LISP roles you already know and they stop being jargon: Fabric Edge Node The access switch users plug into. It is a LISP xTR: it registers endpoints and does VXLAN encap/decap. The anycast gateway lives here. Control Plane Node The LISP map-server / map-resolver. Holds the endpoint-to-edge mapping database for the whole fabric. Border Node The edge of the fabric, connecting to the outside world (WAN, data centre, internet) and translating between fabric and non-fabric. Catalyst Center The controller. Provisions every role from intent, and provides assurance and analytics. Not in the data path. An endpoint connects to a Fabric Edge, which registers it with the Control Plane node. When another endpoint wants to reach it, its edge queries the Control Plane node for the mapping (LISP map-request), gets the destination edge's RLOC, and VXLAN-encapsulates the traffic to it. This is the exact behaviour captured in the [LISP lab](https://www.pinglabz.com/lisp-explained/), operating at campus scale under Catalyst Center's control. ## Two Levels of Segmentation SD-Access segments the network at two levels, and understanding the difference is key: Macro-segmentation (VN)Virtual Networks are separate VRFs, carried as separate L3VNIs. Complete isolation - a guest VN and a corporate VN cannot reach each other at all. This is VRF isolation at fabric scale. Micro-segmentation (SGT)Within a VN, Scalable Group Tags enforce policy between groups (e.g. "contractors cannot reach finance servers") without subnetting. Enforced by group-based ACLs, independent of IP. Macro-segmentation is the coarse cut (separate virtual networks, no leakage). Micro-segmentation is the fine cut (group policy inside a virtual network). Together they replace the sprawling VLAN-and-ACL model with something expressed in terms of identity and intent, and enforced consistently by the fabric rather than device by device. ## Why This Is a Concept Article Being honest about the boundary: SD-Access cannot be labbed without Catalyst Center, which is not available in a standard CML environment and is not something you configure by CLI. The fabric roles are provisioned *by* Catalyst Center; you do not hand-configure a Fabric Edge the way you configure a switch. That is the entire operational premise. What you *can* lab, and what this cluster does lab, are the underlying protocols: [LISP with real map-cache captures](https://www.pinglabz.com/lisp-explained/) and [VXLAN with a real byte-level encapsulation](https://www.pinglabz.com/vxlan-deep-dive/). Understanding those is understanding the fabric's mechanics. The Catalyst Center layer on top is an operational and API story, covered from the automation angle in the [Network Automation cluster](https://www.pinglabz.com/network-automation/). No fabricated `show fabric` output appears here, because that output only exists on a Catalyst Center-provisioned fabric. ## Is It Worth It? SD-Access is a large commitment: Catalyst Center licensing, ISE for identity, capable hardware (Catalyst 9000), and a real operational shift from CLI to controller. It pays off for large campuses with strong segmentation and mobility requirements, particularly where group-based policy and consistent enforcement matter (healthcare, finance, large enterprise). For a small or mid-size network with modest segmentation needs, the traditional model, or a straightforward VXLAN-EVPN fabric without the full SD-Access stack, is often the better fit. The [migration and interoperability article](https://www.pinglabz.com/sd-access-traditional-campus-interop/) covers how to bridge the two worlds when you do adopt it. ## FAQ ### What is the control plane of SD-Access? LISP. The Control Plane node is a LISP map-server/map-resolver, and Fabric Edge nodes are LISP xTRs. Endpoint mobility is the LISP EID-stays-put, RLOC-changes behaviour. ### What is the data plane? VXLAN, carrying both the segment (VNI) and the group tag (SGT) across the fabric. ### What is an SGT? A Scalable Group Tag: a group identifier (from TrustSec) carried in the VXLAN header, used to enforce policy between groups of users independent of their IP addresses. This is micro-segmentation. ### Can I build SD-Access in CML? No. It requires Catalyst Center to provision the fabric. You can lab the underlying LISP and VXLAN protocols, which is what this cluster does with real captures. ### Do I still need ISE? For the full identity-and-policy model (dynamic SGT assignment, group-based access control), yes. ISE is the identity engine that maps users to groups. Without it you have the fabric but not the intent-based policy. ## Key Takeaways - SD-Access is **LISP (control plane) + VXLAN (data plane) + Catalyst Center (automation/policy)**. Learn LISP and VXLAN and you understand the mechanics. - Roles map onto LISP: **Fabric Edge** \= xTR, **Control Plane node** \= map-server/resolver, **Border** \= fabric exit, **Catalyst Center** \= controller (not in the data path). - Two segmentation levels: **macro** (Virtual Networks = VRFs = complete isolation) and **micro** (SGTs = group policy inside a VN, independent of IP). - The value is location-independent identity, group-based policy, and automated consistent operation, replacing the VLAN-and-ACL sprawl. - It **cannot be labbed** without Catalyst Center, so this is a concept article. The underlying LISP and VXLAN are labbed with real captures elsewhere in this cluster. - It is a large commitment (Catalyst Center, ISE, Catalyst 9000). Worth it for large campuses with strong segmentation/mobility needs; overkill for small networks. Next: [SD-Access and the traditional campus](https://www.pinglabz.com/sd-access-traditional-campus-interop/), or the [Network Virtualization cluster guide](https://www.pinglabz.com/network-virtualization/). ### VXLAN with BGP EVPN: The Control Plane Fabrics Actually Use URL: https://www.pinglabz.com/vxlan-bgp-evpn-explained/ Last updated: 2026-07-12T02:30:07.000Z VXLAN gives you the data plane: a way to wrap Ethernet in UDP and route it across a fabric. But an encapsulation is useless until the ingress switch knows which remote switch a destination lives behind. The original answer was to flood and learn, exactly like a dumb switch, which does not scale to a real data centre. The modern answer is to give VXLAN a proper control plane, and that control plane is MP-BGP with the EVPN address family. This article explains BGP EVPN: what it advertises, the route types that carry it, and why it replaced flood-and-learn everywhere. It builds on the [VXLAN deep dive](https://www.pinglabz.com/vxlan-deep-dive/) and the [BGP cluster](https://www.pinglabz.com/bgp/), and extends the [Network Virtualization cluster guide](https://www.pinglabz.com/network-virtualization/). ## A Note on the Captures Consistent with this site's rule that captured output is always real, the VXLAN data-plane bytes in this cluster come from a genuine crafted packet (see the [VXLAN deep dive](https://www.pinglabz.com/vxlan-deep-dive/), which shows the real 4789/UDP encapsulation and VNI 10010 on the wire). The EVPN control-plane configuration and show-command structure below are presented as **reference syntax** drawn from Cisco's documented NX-OS behaviour, not as lab captures, because a virtual Nexus fabric could not be brought to a usable login state in this build. Where this article shows a `show` command, treat it as the command to run and the fields to read, not as a claim of captured output. The concepts, route types, and configuration are accurate and standard. ## Why Flood-and-Learn Had to Go In flood-and-learn VXLAN, a VTEP learns remote MACs the way a switch does: it floods unknown-unicast, broadcast, and multicast (BUM) traffic to every other VTEP, watches the replies, and builds its mapping table from the data plane. This works, and it is simple, but it has the same problems that made large flat Layer 2 networks a bad idea: - **Flooding does not scale.** Every unknown destination triggers fabric-wide flooding. In a big fabric that is a lot of wasted bandwidth and a lot of BUM replication. - **It needs multicast in the underlay** (or ingress replication), which is operationally heavy. See the [multicast cluster](https://www.pinglabz.com/multicast/) for why running PIM everywhere is not free. - **No proactive knowledge.** A VTEP knows nothing about a host until traffic forces it to learn, so the first packet of every new conversation floods. BGP EVPN fixes all three by turning reachability into a routing problem: VTEPs *advertise* the MACs and IPs they have, so every other VTEP already knows where everything is before the first packet. ## EVPN Is an Address Family EVPN is not a new protocol. It is an address family of MP-BGP (the same MP-BGP that carries [MPLS L3VPN](https://www.pinglabz.com/mpls-l3vpn/) VPNv4 routes), called `l2vpn evpn`. If you understand how MP-BGP carries VPNv4 prefixes with route distinguishers and route targets, you already understand 80% of EVPN, because it uses the identical machinery for a different payload. What it carries is a set of **route types**, each advertising a different kind of reachability: Type 2: MAC/IP "Host with this MAC (and this IP) lives behind me." The workhorse. Advertises endpoint reachability and populates remote ARP/MAC tables without flooding. Type 3: IMET "I am a VTEP participating in this VNI." Builds the flood list for the BUM traffic that still must be replicated (like broadcasts), so ingress replication needs no multicast. Type 5: IP Prefix "This IP prefix is reachable via me." Carries routed prefixes (L3VNI) and external/summary routes into the fabric, for inter-subnet and outside connectivity. Type 1 & 4 Ethernet Auto-Discovery and Ethernet Segment routes, used for multihoming (ESI/VPC) and fast convergence. Important at scale, secondary when learning the basics. The two you must know cold are Type 2 (a host is here) and Type 5 (a subnet is here). Type 2 is how the fabric learns endpoints without flooding; Type 5 is how it does routing between subnets and reaches the outside world. ## The Type 2 Trick: No More ARP Flooding Here is the single most elegant thing EVPN does. A Type 2 route advertises both the MAC *and* the IP of a host. So when host A wants to talk to host B and sends an ARP request for B's IP, A's local leaf already has a Type 2 route telling it B's MAC. The leaf answers the ARP *locally* (ARP suppression) instead of flooding it across the fabric. Multiply that across thousands of hosts and you have eliminated fabric-wide ARP flooding entirely. The endpoints think they are on a normal broadcast segment; the fabric never actually broadcasts their ARP. That is the scalability win, and it is impossible with flood-and-learn. ## The Configuration Shape Reference configuration for a leaf VTEP (NX-OS style). Read it for the structure; the pieces map directly onto the concepts above: ``` ! features feature ospf feature bgp feature nv overlay feature vn-segment-vlan-based nv overlay evpn ! ! the L2VNI to VLAN mapping vlan 10 vn-segment 10010 ! ! the NVE interface - the VTEP itself interface nve1 source-interface loopback0 host-reachability protocol bgp ! use EVPN, not flood-and-learn member vni 10010 ingress-replication protocol bgp ! BUM via EVPN Type 3, no multicast ! ! MP-BGP with the EVPN address family router bgp 65000 neighbor 10.0.0.100 remote-as 65000 ! the spine / route reflector address-family l2vpn evpn send-community extended ! ! tie the VLAN into EVPN evpn vni 10010 l2 rd auto route-target import auto route-target export auto ``` The lines that matter conceptually: `host-reachability protocol bgp` is what turns off flood-and-learn and turns on EVPN. `ingress-replication protocol bgp` is what lets BUM traffic be replicated using the Type 3 flood list instead of underlay multicast. And `rd auto` / `route-target ... auto` is the same RD/RT construct from [MPLS L3VPN](https://www.pinglabz.com/mpls-rd-vs-rt-explained/), here keeping tenants separate in the EVPN table. ## The Commands You Would Run On a live fabric, these are the verification commands and what each one tells you: `show nve peers`Which remote VTEPs this leaf has discovered (learned from Type 3 routes). The overlay adjacency list. `show bgp l2vpn evpn`The EVPN table itself: every Type 2, 3, and 5 route, with its RD and route targets. This is where you confirm a host was advertised. `show l2route evpn mac all`The MACs learned via EVPN and which VTEP each is behind. The control-plane MAC table. `show nve vni`Which VNIs are up on this VTEP and their state. Your first check when a segment is not passing traffic. When a host in VNI 10010 on one leaf cannot reach a host on another, the troubleshooting path is: is the VNI up (`show nve vni`), is the remote VTEP a peer (`show nve peers`), and did the host's Type 2 route arrive (`show bgp l2vpn evpn`). Missing at the last step usually means a route-target or RD mismatch, the exact same failure mode as MPLS L3VPN. ## Why the Spine-Leaf Topology EVPN fabrics are built spine-leaf for a reason that ties back to BGP. Leaves are the VTEPs (where hosts connect); spines are the transit and the BGP route reflectors. Every leaf peers EVPN with the spines, and the spines reflect routes between leaves, so you get a full mesh of reachability without a full mesh of BGP sessions, the same [route-reflector](https://www.pinglabz.com/bgp/) scaling trick used in any large iBGP network. The underlay (OSPF or eBGP) provides loopback-to-loopback reachability between VTEPs; the overlay (EVPN) rides on top. Two BGP roles, one physical fabric. ## FAQ ### Is EVPN a replacement for VXLAN? No. VXLAN is the data-plane encapsulation; EVPN is the control plane that tells VTEPs where to send. They work together. "VXLAN-EVPN" is the full stack. ### What are the two route types I must know? Type 2 (MAC/IP, advertises a host and enables ARP suppression) and Type 5 (IP prefix, carries routed subnets and external routes). Type 3 (IMET) builds the BUM flood list. ### Does EVPN need multicast? No. That is a key advantage over flood-and-learn. BUM traffic is handled by ingress replication driven by Type 3 routes, so the underlay only needs unicast routing. ### How is this related to MPLS L3VPN? Very closely. EVPN is an MP-BGP address family using the same RD/RT machinery as VPNv4\. If you know L3VPN, EVPN's control plane will feel familiar; the payload is MACs and host IPs instead of VPN prefixes. ### What is symmetric IRB? The standard routing model where inter-subnet traffic is routed at the ingress leaf into an L3VNI and carried across the fabric in that VNI, with every leaf sharing an anycast gateway. It is what makes a host's gateway always local, enabling seamless mobility. ## Key Takeaways - VXLAN is the data plane; **BGP EVPN is the control plane** that replaced flood-and-learn. Together they are the modern fabric. - Flood-and-learn does not scale (fabric-wide flooding, needs multicast). EVPN advertises reachability proactively, so VTEPs know where hosts are before the first packet. - EVPN is an **MP-BGP address family** (`l2vpn evpn`) using the same RD/RT machinery as MPLS L3VPN. - Know **Type 2** (MAC/IP, enables ARP suppression, no ARP flooding) and **Type 5** (IP prefix, routing and external). Type 3 builds the BUM flood list. - Key config: `host-reachability protocol bgp` (turns on EVPN) and `ingress-replication protocol bgp` (BUM without multicast). - Verify with `show nve peers`, `show bgp l2vpn evpn`, `show l2route evpn mac all`, `show nve vni`. A missing Type 2 route is usually an RD/RT mismatch, just like L3VPN. - Spine-leaf exists because spines are the BGP route reflectors, scaling the overlay the same way route reflectors scale any iBGP network. Next: [Cisco SD-Access architecture](https://www.pinglabz.com/sd-access-architecture/) (which combines LISP and VXLAN), or the [Network Virtualization cluster guide](https://www.pinglabz.com/network-virtualization/). ### VXLAN Deep Dive: VNIs, VTEPs, and the UDP Encapsulation URL: https://www.pinglabz.com/vxlan-deep-dive/ Last updated: 2026-07-12T02:30:07.000Z Every modern data centre fabric, and the campus fabric of SD-Access, is built on VXLAN. It is the technology that lets you stretch a Layer 2 segment across a routed network, put a hundred thousand tenants where VLANs gave you four thousand, and move a virtual machine from one rack to another without changing its IP. Underneath the marketing, VXLAN is a genuinely simple idea: wrap the original Ethernet frame inside a UDP packet and route it like any other traffic. This article takes VXLAN apart to the byte level, using a real crafted-and-decoded packet rather than a diagram. It extends the [Network Virtualization and Overlays cluster guide](https://www.pinglabz.com/network-virtualization/). ## The Problems VXLAN Solves Three limitations of traditional [VLANs](https://www.pinglabz.com/vlans-layer-2-switching/) drove VXLAN: The 4094 ceiling A VLAN ID is 12 bits, so about 4094 usable segments. A VNI is 24 bits: over 16 million. Enough for real multi-tenancy. Layer 2 does not scale Big flat L2 domains mean big failure domains and spanning-tree fragility. VXLAN rides a routed underlay - no STP across the fabric. Workloads need to move A VM keeps its IP and MAC when it migrates across the fabric, because its L2 segment is decoupled from physical topology. ## The Two Terms That Matter: VTEP and VNI VTEP (VXLAN Tunnel Endpoint)The device that does the encapsulation and decapsulation. It has an IP in the underlay (usually a loopback) and is where the overlay meets the physical network. A leaf switch is a VTEP. VNI (VXLAN Network Identifier)The 24-bit segment ID. It is to VXLAN what the VLAN ID is to a VLAN: it identifies which virtual network a frame belongs to. Two hosts in the same VNI are on the same virtual L2 segment. A packet's journey: host A sends a normal Ethernet frame; its local VTEP looks up the destination, wraps the whole frame in VXLAN + UDP + IP addressed to the remote VTEP, and routes it across the underlay; the remote VTEP strips the wrapper and delivers the original frame to host B. Host A and host B believe they are on the same switch. They may be in different buildings. ## The Encapsulation, Byte by Byte Rather than draw the header, here is a real VXLAN packet built with Scapy and dumped as raw bytes. The inner frame is a ping from tenant host 10.10.10.1 to 10.10.10.2; the outer wrapper carries it between VTEPs 10.0.0.1 and 10.0.0.2 in VNI 10010: ``` ###[ IP ]### (outer - routed across the underlay) proto = udp src = 10.0.0.1 <- source VTEP dst = 10.0.0.2 <- destination VTEP ###[ UDP ]### sport = 51234 <- hash of the inner flow (entropy for ECMP) dport = 4789 <- the VXLAN port ###[ VXLAN ]### flags = Instance vni = 0x271a <- 10010 ###[ Ethernet ]### (inner - the original tenant frame, intact) dst = 00:11:22:33:44:02 src = 00:11:22:33:44:01 ###[ IP ]### src = 10.10.10.1 dst = 10.10.10.2 ###[ ICMP ]### echo-request ``` And the actual bytes on the wire, which make the structure undeniable: ``` 0000 45 00 00 4E 00 01 00 00 40 11 66 9C 0A 00 00 01 E..N....@.f..... 0010 0A 00 00 02 C8 22 12 B5 00 3A 19 ED 08 00 00 00 ....."...:...... 0020 00 27 1A 00 00 11 22 33 44 02 00 11 22 33 44 01 .'...."3D..."3D. 0030 08 00 45 00 00 1C 00 01 00 00 40 01 52 CA 0A 0A ..E.......@.R... 0040 0A 01 .. ``` Trace it: 45 00 ... 40 11Outer IPv4 header. `40` \= TTL 64, `11` \= protocol 17 (UDP). The core routes on this. 0A 00 00 01 / 0A 00 00 02Source VTEP 10.0.0.1, destination VTEP 10.0.0.2\. Plain underlay addresses. C8 22 12 B5UDP source port 0xC822 (51234, a per-flow hash), destination port **0x12B5 = 4789**, the assigned VXLAN port. 08 00 00 00 00 27 1A 00The 8-byte VXLAN header. `08` \= the I (Instance) flag, meaning the VNI is valid. `27 1A` in the 24-bit VNI field = **10010**. 00 11 22 33 44 02 ...The original tenant Ethernet frame begins, completely intact: dest MAC, source MAC, EtherType 0800, then the inner IP (10.10.10.1 to 10.10.10.2) and ICMP. Two design details fall out of those bytes. The **UDP source port is a hash of the inner flow**, which is deliberate: it gives the underlay entropy so equal-cost paths ([ECMP](https://www.pinglabz.com/ospf/)) load-balance different tenant flows across different links, even though every VXLAN packet has the same UDP destination port. And the encapsulation adds **50 bytes** of overhead (outer Ethernet + IP + UDP + VXLAN), which is why VXLAN fabrics run jumbo MTU (9216) on the underlay; a 1500-byte tenant frame plus 50 bytes exceeds 1500 and would fragment or drop. ## The Missing Piece: How Does a VTEP Know Where to Send? The encapsulation above assumes the source VTEP already knows that host 10.10.10.2's MAC lives behind VTEP 10.0.0.2\. But how did it learn that? This is the control-plane question, and it is where the two generations of VXLAN differ: Flood-and-learn (original) VTEPs flood unknown-unicast and broadcast via multicast (or ingress replication) and learn MACs from the data plane, exactly like a switch. Simple, but it floods, and it does not scale. BGP EVPN (modern) MP-BGP distributes MAC and IP reachability between VTEPs as a control plane. No data-plane flooding to learn. This is what every real fabric uses today. The data-plane encapsulation is identical either way, the bytes above are the same. What changes is how the mapping table gets populated. Flood-and-learn does it reactively by flooding; EVPN does it proactively by advertising. The control plane is a big enough topic that it has its own article: [VXLAN with BGP EVPN](https://www.pinglabz.com/vxlan-bgp-evpn-explained/), which builds on the [BGP](https://www.pinglabz.com/bgp/) and [MP-BGP](https://www.pinglabz.com/mp-bgp-multiprotocol-bgp/) knowledge you already have. ## L2VNI and L3VNI One more distinction that trips people up. A fabric uses two kinds of VNI: - **L2VNI** maps to a VLAN and carries a Layer 2 segment. Hosts in the same L2VNI are on the same subnet and bridge to each other across the fabric. The packet above is an L2VNI packet. - **L3VNI** maps to a VRF and carries routed traffic *between* subnets. When host A in one subnet talks to host B in another, the traffic is routed at the ingress leaf (anycast gateway) and carried across the fabric in the L3VNI. This is symmetric IRB (Integrated Routing and Bridging), the standard modern design. The practical upshot: every leaf shares the same anycast gateway IP and MAC for each subnet, so a host's default gateway is always its local leaf, wherever the host moves. That is what makes VM mobility seamless, and it is pure VXLAN-EVPN mechanics. ## FAQ ### What port does VXLAN use? UDP 4789 (IANA-assigned). You can see it as `0x12B5` in the real capture above. Some older Cisco deployments used 8472; standardise on 4789. ### How much overhead does VXLAN add? 50 bytes (outer Ethernet 14 + IP 20 + UDP 8 + VXLAN 8). This is why underlay MTU must be raised, typically to 9216, or tenant frames near 1500 bytes will fragment or drop. ### Why is the UDP source port randomised? It is a hash of the inner flow, giving the underlay entropy so ECMP can load-balance different tenant flows across parallel links. The destination port is always 4789, so without this, all VXLAN traffic between two VTEPs would pin to one path. ### Do I still use VLANs with VXLAN? Yes, locally. A leaf maps a local VLAN to an L2VNI. The VLAN is significant only on that switch; the VNI is what is significant across the fabric. VLANs did not go away, they became leaf-local. ### Is VXLAN the same as a GRE tunnel? Conceptually similar (encapsulate and route), but VXLAN uses UDP (better for ECMP and hardware offload), carries a 24-bit segment ID, and is purpose-built for L2-over-L3 at fabric scale. See the [GRE cluster](https://www.pinglabz.com/gre/) for the comparison. ## Key Takeaways - VXLAN wraps the original Ethernet frame in **UDP (port 4789) + IP** and routes it between VTEPs across a Layer 3 underlay. - A **VTEP** encapsulates/decapsulates; a **VNI** is the 24-bit segment ID (16M segments vs the VLAN's 4094). - The real byte dump proves it: outer VTEP IPs, UDP dport `0x12B5` (4789), the 8-byte VXLAN header with the Instance flag and VNI `0x271A` (10010), then the tenant frame intact. - The **UDP source port is a per-flow hash** for ECMP entropy; the encapsulation adds **50 bytes**, so raise the underlay MTU. - The data plane is identical whether the control plane is **flood-and-learn** or **BGP EVPN**. Only how VTEPs learn reachability differs. - **L2VNI** bridges within a subnet; **L3VNI** routes between subnets via anycast gateways, which is what makes VM mobility seamless. Next: [VXLAN with BGP EVPN](https://www.pinglabz.com/vxlan-bgp-evpn-explained/) for the control plane, or the [Network Virtualization cluster guide](https://www.pinglabz.com/network-virtualization/). ### LISP Explained: Locator/ID Separation and Why Fabrics Use It URL: https://www.pinglabz.com/lisp-explained/ Last updated: 2026-07-12T02:30:07.000Z An IP address is doing two jobs at once, and that overload is the root of a surprising number of networking problems. It is an *identity* ("this is my server") and a *location* ("here is where it sits in the topology"). When a machine moves (a VM migrates, a site multihomes, a device roams), those two meanings conflict: the identity should stay the same, but the location must change, and IP forwarding cannot tell them apart. LISP fixes this by splitting the two into separate namespaces. This article explains the Locator/ID Separation Protocol and proves it on a real IOS XE lab, watching the first packets get dropped while the mapping resolves and then flow cleanly. It is the opening piece of the [Network Virtualization and Overlays cluster guide](https://www.pinglabz.com/network-virtualization/). ## Two Namespaces: EID and RLOC LISP defines two separate address spaces where IP has one: EID (Endpoint Identifier) The address of an *endpoint* \- a host, a server, a VM. It identifies who, and it does not change when the endpoint moves. Endpoints only ever see EIDs. RLOC (Routing Locator) The address of a *router* that can reach the endpoint. It identifies where, and it is what the core network actually routes on. The core only ever sees RLOCs. The core network routes on RLOCs and has no idea EIDs exist. The endpoints use EIDs and have no idea RLOCs exist. In between sits a **mapping system** that answers the question "for this EID, which RLOC do I send to?" That indirection is the whole of LISP, and everything else is machinery to make the mapping fast and reliable. The analogy that sticks: an EID is like a person's name and an RLOC is like their current street address. You know your friend by name (identity), but to send them mail you need their address (location), and a directory maps one to the other. When they move house, the name is unchanged and only the directory entry updates. ## The Router Roles LISP routers sit at the edge, between the EID world (the sites) and the RLOC world (the core), and they wear a few hats: ITR (Ingress Tunnel Router)Receives a packet from a local EID, looks up the destination EID's RLOC, and encapsulates the packet toward it. ETR (Egress Tunnel Router)Receives the encapsulated packet, decapsulates it, and delivers to the local EID. Also registers its EIDs with the mapping system. xTRA router that is both an ITR and an ETR, which is the normal case. Every site edge is an xTR. Map-Server (MS)Accepts EID-to-RLOC registrations from ETRs. Holds the mapping database. Map-Resolver (MR)Answers map-requests from ITRs, returning the RLOC for a queried EID. ## The Lab Three IOS XE routers, and yes, LISP runs fully on this platform (the `router lisp` configuration mode is complete): ``` xTR1 ---- MSMR ---- xTR2 RLOC MS+MR RLOC 1.1.1.1 2.2.2.2 3.3.3.3 EID (mapping EID 10.1.1/24 system) 10.3.3/24 ``` xTR1 and xTR2 are site edges, each with a customer EID subnet on a loopback. MSMR is the mapping system (map-server and map-resolver combined, which is common in smaller deployments). OSPF runs in the core purely to give the RLOCs (the loopbacks) reachability to each other; the EID subnets are *not* in the IGP, which is the point. ### The mapping system ``` MSMR(config)# router lisp MSMR(config-router-lisp)# site SITE-1 MSMR(config-router-lisp-site)# authentication-key PINGLABZ MSMR(config-router-lisp-site)# eid-prefix 10.1.1.0/24 MSMR(config-router-lisp)# site SITE-2 MSMR(config-router-lisp-site)# authentication-key PINGLABZ MSMR(config-router-lisp-site)# eid-prefix 10.3.3.0/24 MSMR(config-router-lisp)# ipv4 map-server MSMR(config-router-lisp)# ipv4 map-resolver ``` ### The xTRs ``` xTR1(config)# router lisp xTR1(config-router-lisp)# locator-set MYLOC xTR1(config-router-lisp-locator-set)# IPv4-interface Loopback0 priority 1 weight 100 xTR1(config-router-lisp)# eid-table default instance-id 0 xTR1(config-router-lisp-eid-table)# database-mapping 10.1.1.0/24 locator-set MYLOC xTR1(config-router-lisp)# ipv4 itr map-resolver 2.2.2.2 xTR1(config-router-lisp)# ipv4 etr map-server 2.2.2.2 key PINGLABZ xTR1(config-router-lisp)# ipv4 itr xTR1(config-router-lisp)# ipv4 etr ``` Reading that: the `locator-set` names this router's RLOC (its Loopback0). The `database-mapping` declares "I am the ETR for the 10.1.1.0/24 EID subnet, reachable via my RLOC." The `map-resolver` and `map-server` lines point at MSMR. The final `itr`/`etr` lines turn on both roles. xTR2 is identical with its own EID subnet. ## Registration: The ETRs Announce Themselves As soon as the xTRs come up, each registers its EID prefix with the map-server. The map-server's view: ``` MSMR#show lisp site Site Name Last Up Who Last Inst EID Prefix Register Registered ID SITE-1 00:00:45 yes# 1.1.1.1:54379 0 10.1.1.0/24 SITE-2 00:00:16 yes# 3.3.3.3:24326 0 10.3.3.0/24 ``` Both sites are registered and up. The mapping database now knows that 10.1.1.0/24 lives behind RLOC 1.1.1.1 and 10.3.3.0/24 lives behind RLOC 3.3.3.3\. The `#` flag means the registration used reliable transport (TCP-based), a modern refinement. This registration is the "moving house updates the directory" step. If a site's EIDs move to a new RLOC, the ETR re-registers and the database updates, with no change to the EIDs themselves and no reconvergence in the core. ## The Money Capture: First-Packet Behavior Here is the single most instructive thing about LISP, and it surprises people the first time. Ping from xTR1's EID to xTR2's EID with a cold map-cache: ``` xTR1#ping 10.3.3.1 source 10.1.1.1 Sending 5, 100-byte ICMP Echos to 10.3.3.1, timeout is 2 seconds: ..!!! Success rate is 60 percent (3/5) ``` **The first two packets are dropped.** That is not a fault; it is LISP working. When the ITR receives the first packet for an unknown EID, it has no RLOC to send to, so it drops the packet and fires off a *map-request* to the map-resolver. While it waits for the reply, subsequent packets are also dropped, until the mapping arrives and gets cached. Then packets three, four, and five sail through. The map-cache after that first ping tells the story: ``` xTR1#show lisp instance-id 0 ipv4 map-cache LISP IPv4 Mapping Cache for LISP 0 EID-table default (IID 0), 2 entries 0.0.0.0/0, uptime: 00:00:50, via static-send-map-request Negative cache entry, action: send-map-request 10.3.3.0/24, uptime: 00:00:04, via transient-publication, complete Locator Uptime State Pri/Wgt 3.3.3.3 00:00:04 up 1/100 ``` The 10.3.3.0/24 entry is now cached, pointing at RLOC 3.3.3.3\. The `0.0.0.0/0` negative entry is the "anything I do not have a mapping for, send a map-request" default. You can trigger the lookup manually with the LISP Internet Groper, which is the LISP equivalent of a ping-for-mappings: ``` xTR1#lig 10.3.3.1 Mapping information for EID 10.3.3.1 from 3.3.3.3 with RTT 4 msecs 10.3.3.0/24, uptime: 00:00:00, via map-reply, complete Locator Uptime State Pri/Wgt 3.3.3.3 00:00:00 up 1/100 ``` `via map-reply` confirms the mapping came from a map-reply, sourced from RLOC 3.3.3.3\. And now that the cache is warm, the ping is perfect: ``` xTR1#ping 10.3.3.1 source 10.1.1.1 !!!!! Success rate is 100 percent (5/5) ``` Cold cache, first packets dropped, map resolved, warm cache, clean flow. That cycle is the entire LISP data plane in three commands. ## Why This Architecture Matters The EID/RLOC split buys capabilities that flat IP routing struggles with: Mobility An endpoint keeps its EID when it moves. Only the mapping updates. No renumbering, no core reconvergence. Scalable multihoming A site registers multiple RLOCs with priorities and weights. Ingress traffic engineering without polluting the global BGP table. Smaller core tables The core routes only on RLOCs (a small set), not on every EID prefix. EIDs live in the mapping system, resolved on demand. The fabric control plane LISP is the control plane of Cisco SD-Access: it maps endpoints to fabric edge nodes exactly as it maps EIDs to RLOCs here. That last point is why LISP matters for the CCNP. It is not a niche protocol; it is the mapping system underneath [Cisco SD-Access](https://www.pinglabz.com/sd-access-architecture/). The fabric edge nodes are xTRs, the control-plane nodes are map-servers/resolvers, and endpoint mobility across the campus is exactly the EID-stays-put, RLOC-changes behaviour demonstrated above, at scale. Understanding LISP here means understanding the fabric later. ## FAQ ### Why were my first LISP pings dropped? Expected behaviour. The ITR drops the first packet(s) to an unresolved EID while it sends a map-request and waits for the reply. Once the mapping is cached, traffic flows. Warm the cache with `lig` if you want the first real packet to succeed. ### Does the core need to know about EIDs? No, and that is the point. The core routes only on RLOCs. EID prefixes are deliberately kept out of the core IGP and live in the mapping system. ### Can a site have more than one RLOC? Yes. A site registers multiple locators with priorities and weights, giving you multihoming and ingress traffic engineering without injecting the site prefix into global BGP. ### Is LISP just for SD-Access? No, but that is its highest-profile use. LISP is a general architecture (RFC 9300/9301) used for mobility, multihoming, and IPv6 transition, as well as being the SD-Access control plane. ### What replaces the map when an endpoint moves? The new ETR registers the EID with the map-server, updating the database. ITRs with a stale cached mapping are corrected via solicit-map-request, so traffic follows the endpoint to its new location. ## Key Takeaways - An IP address conflates **identity** and **location**. LISP splits them into **EIDs** (endpoints, who) and **RLOCs** (routers, where). - The core routes only on RLOCs; a **mapping system** (map-server + map-resolver) answers "which RLOC for this EID?" on demand. - Roles: **ITR** encapsulates, **ETR** decapsulates and registers, **xTR** is both. LISP runs fully on IOS XE. - The lab captured the signature behaviour: **first pings dropped (60%)** while the map-request resolved, then **100%** once cached. `lig` showed the mapping arriving via map-reply from RLOC 3.3.3.3. - The split enables mobility (EID stays, RLOC changes), scalable multihoming, and smaller core tables. - LISP is the **control plane of Cisco SD-Access**. Understanding it here is understanding the fabric later. Next: [VXLAN deep dive](https://www.pinglabz.com/vxlan-deep-dive/) (the data-plane encapsulation fabrics pair with LISP or EVPN), or the [Network Virtualization cluster guide](https://www.pinglabz.com/network-virtualization/). ### Hardening the Management Plane on Cisco IOS XE URL: https://www.pinglabz.com/hardening-management-plane-ios-xe/ Last updated: 2026-07-12T02:04:09.000Z The management plane is how you reach and control your devices: SSH sessions, SNMP polls, NTP syncs, the VTY lines an engineer logs into. It is also the softest target on the box, because it is the one part of the router that will happily talk back to anyone who connects. Every device left with default SSH, no VTY restrictions, and no brute-force protection is a device waiting for an unattended login attempt to succeed. This article hardens the management plane on Cisco IOS XE with real captures of a modern, secured configuration. It extends the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/) and complements [AAA](https://www.pinglabz.com/aaa-tacacs-radius-device-administration/) (who gets in) and [CoPP](https://www.pinglabz.com/control-plane-policing-copp/) (protecting the CPU that serves these sessions). ## SSH, Done Right for 2026 SSH is table stakes, but "configured SSH" and "hardened SSH" are different things. Start with a strong host key. On modern IOS XE this is not a suggestion; the software enforces it: ``` R1(config)# crypto key generate rsa modulus 2048 label PINGLABZ.KEYS Please create an RSA key of atleast 3072 bits or an EC key of atleast 384 bits to enable SSH server functionality ``` The router rejects a 2048-bit key outright. That is a meaningful change: older guidance said 2048 was fine, and IOS XE 17.x now demands 3072-bit RSA (or a 384-bit elliptic-curve key) before it will even enable the SSH server. Generate a compliant key: ``` R1# crypto key generate rsa modulus 3072 label PINGLABZ.KEYS % The key modulus size is 3072 bits % Generating crypto RSA keys in background ... ``` Then enforce version 2 and sane timeouts: ``` R1(config)# ip ssh version 2 R1(config)# ip ssh time-out 60 R1(config)# ip ssh authentication-retries 3 ``` Verify what the box is actually offering, and this output is worth reading in full because it tells you your real cryptographic posture: ``` R1#show ip ssh SSH Enabled - version 2.0 Authentication timeout: 60 secs; Authentication retries: 3 KEX Algorithms:curve25519-sha256,curve25519-sha256@libssh.org,ecdh-sha2-nistp256,...,diffie-hellman-group14-sha256,diffie-hellman-group16-sha512 Encryption Algorithms:chacha20-poly1305@openssh.com,aes128-gcm@openssh.com,aes256-gcm@openssh.com,aes128-gcm,aes256-gcm,aes128-ctr,aes192-ctr,aes256-ctr MAC Algorithms:hmac-sha2-256-etm@openssh.com,hmac-sha2-512-etm@openssh.com,hmac-sha2-256,hmac-sha2-512 Minimum expected Diffie Hellman key size : 2048 bits IOS Keys in SECSH format(ssh-rsa, base64 encoded): PINGLABZ.KEYS Modulus Size : 3072 bits ``` Note what modern IOS XE offers by default: curve25519 key exchange, ChaCha20-Poly1305 and AES-GCM ciphers, and ETM (encrypt-then-MAC) integrity. These are current, strong algorithms. Older code offered weak ciphers (3DES, arcfour) and weak KEX (diffie-hellman-group1-sha1) that you had to manually disable. On 17.x the defaults are good; your job is to confirm them and, in high-security environments, to explicitly restrict to only the strongest with `ip ssh server algorithm` commands. ## Restrict Who Can Even Connect Strong SSH still answers the door for the entire internet if you let it. A VTY access-class limits which source addresses can attempt to connect at all, so a brute-force from an unexpected network never even reaches the authentication stage: ``` R1(config)# ip access-list standard MGMT-ACCESS R1(config-std-nacl)# permit 192.168.99.0 0.0.0.255 R1(config-std-nacl)# deny any log ! R1(config)# line vty 0 4 R1(config-line)# access-class MGMT-ACCESS in R1(config-line)# exec-timeout 5 0 R1(config-line)# transport input ssh ``` Three controls in that block, each important: access-class MGMT-ACCESS inOnly the management subnet can open a VTY session. The `deny any log` records every rejected attempt. exec-timeout 5 0An idle session is closed after 5 minutes. An unlocked laptop with an open SSH session stops being a permanent back door. transport input sshTelnet is disabled. Only SSH is accepted. This one line closes the most common legacy exposure. The management ACL is the same idea as restricting your automation host in the [API security article](https://www.pinglabz.com/rest-api-security-network-engineers/): define where management can originate, and reject everything else at the earliest possible point. ## Brute-Force Protection: login block-for Even with a VTY ACL, an attacker inside the permitted range (or a compromised management host) can hammer the login. `login block-for` detects repeated failures and locks logins entirely for a cooldown period: ``` R1(config)# login block-for 120 attempts 4 within 60 R1(config)# login delay 2 R1(config)# login on-failure log R1(config)# login on-success log ``` Read as: if 4 login failures occur within 60 seconds, disable all logins for 120 seconds. `login delay 2` forces a 2-second gap between attempts, which alone defeats high-speed guessing. And `on-failure`/`on-success log` means every login, good or bad, is recorded, which is your early warning of an attack in progress. The router exposes its current state, which is genuinely useful during an incident: ``` R1#show login A login delay of 2 seconds is applied. All successful login is logged. All failed login is logged. Router enabled to watch for login Attacks. If more than 4 login failures occur in 60 seconds or less, logins will be disabled for 120 seconds. Router presently in Normal-Mode. Current Watch Window remaining time 29 seconds. Present login failure count 0. ``` "Normal-Mode" means no attack is currently detected. Under a brute-force this flips to "Quiet-Mode" and logins are refused until the cooldown expires. You can even exempt your management subnet from Quiet-Mode with a `login quiet-mode access-class`, so a real attack from elsewhere cannot lock *you* out while it is being blocked, a nice touch that avoids the self-inflicted lockout. ## SNMP: v3 Only SNMP is a management-plane protocol and a classic exposure. SNMPv1 and v2c authenticate with a community string sent in cleartext, which is to say they barely authenticate at all. A sniffed `public` or `private` community is full read (or write) access. On any device that matters, use v3 exclusively: ``` R1(config)# snmp-server group SECURE-GRP v3 priv R1(config)# snmp-server user snmpadmin SECURE-GRP v3 auth sha AuthPass123! priv aes 128 PrivPass123! ``` SNMPv3 with `auth ... priv` gives you authentication (SHA) and encryption (AES) of the SNMP payload. Combine it with an ACL restricting which NMS can poll, and remove any lingering v2c communities. The [SNMP article](https://www.pinglabz.com/snmp-explained/) covers the version differences in depth; from a hardening standpoint the rule is one line: no v2c on production devices. ## The Rest of the Checklist - **Encrypt stored passwords.** `service password-encryption` at minimum, and `enable secret` / `username ... secret` (which use strong hashing) rather than `password`. - **Authenticate NTP.** An attacker who can move your clock can break certificate validation and defeat time-based ACLs. `ntp authenticate` plus trusted keys. (See [NTP](https://www.pinglabz.com/ntp-cisco-ios-xe/).) - **Disable unused services.** No HTTP server if you use HTTPS/RESTCONF; no CDP on untrusted edges; no unused small servers. - **Banner.** A legal login banner (`banner login`) is not security, but it is often a legal prerequisite for prosecuting unauthorized access. - **Console and AUX.** The management plane is not just VTY. Secure the console with a password and exec-timeout, and disable the AUX port outright. - **AAA for logins.** Point VTY authentication at [AAA](https://www.pinglabz.com/aaa-tacacs-radius-device-administration/) with a local fallback, so access is centrally controlled and audited. ## FAQ ### Why does IOS XE reject my 2048-bit SSH key? Modern IOS XE (17.x) enforces a minimum of 3072-bit RSA (or 384-bit EC) for the SSH server. Older 2048-bit guidance is deprecated. Generate a 3072-bit key. ### Do I still need to disable weak SSH ciphers manually? On 17.x the defaults are already strong (curve25519, AES-GCM, ChaCha20). In high-security environments you can still explicitly restrict to only the strongest with `ip ssh server algorithm encryption/mac/kex`, but the era of shipping with 3DES and SHA1 enabled by default is over. ### Will login block-for lock me out? It can, which is why you can exempt your management subnet with `login quiet-mode access-class`. Then an attack from elsewhere triggers Quiet-Mode without blocking your legitimate source. ### Is a VTY ACL enough on its own? No. It limits *where* connections can come from, but an attacker inside that range still needs stopping, hence login block-for, strong AAA, and SSH-only transport layered on top. ### What is the single highest-impact line here? `transport input ssh` on the VTY lines. It disables Telnet, closing the most common and most serious legacy management-plane exposure in one command. ## Key Takeaways - Modern IOS XE enforces **3072-bit RSA** (or 384-bit EC) SSH keys and ships with strong defaults (curve25519, AES-GCM, ChaCha20). Confirm them with `show ip ssh`. - Restrict the VTYs: **access-class** to limit source, **exec-timeout** to close idle sessions, and **transport input ssh** to kill Telnet. - **login block-for** plus `login delay` defeats brute-force, and `show login` reports Normal vs Quiet mode live. Exempt your own subnet to avoid self-lockout. - SNMP **v3 only** (auth + priv). v2c community strings are cleartext and disqualifying on production devices. - Round it out: encrypted/hashed passwords, authenticated NTP, disabled unused services, secured console/AUX, and AAA with a local fallback. - The single biggest quick win is `transport input ssh`: disabling Telnet closes the classic management exposure in one line. Next: [AAA for centralized login control](https://www.pinglabz.com/aaa-tacacs-radius-device-administration/), or the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/). ### REST API Security: Tokens, TLS, and Least Privilege for Network APIs URL: https://www.pinglabz.com/rest-api-security-network-engineers/ Last updated: 2026-07-12T02:04:08.000Z Network automation is no longer optional, and every automated network is driven by APIs: RESTCONF on the routers, the Catalyst Center intent API, the SD-WAN Manager API, cloud provider APIs. Each of those is a new door into your infrastructure, and a door that a script can walk through is a door an attacker can walk through. The security of your automation is the security of your network, and it is a topic the CCNP blueprint now expects you to reason about. This article covers the security principles for the network APIs you will actually use, framed for network engineers rather than web developers. It extends the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/) and pairs with the automation work in the [Network Automation cluster](https://www.pinglabz.com/network-automation/). ## Why This Is a Security Problem Now The old management plane was a human typing into an SSH session, protected by AAA and a VTY ACL. The new management plane is a token in a script hitting a REST endpoint, and it changes the threat model: Scale of blast radius A human misconfigures one device. A compromised automation token misconfigures every device, in seconds. Credentials in code Passwords and tokens end up in scripts, repos, and CI logs. A leaked repo is a leaked network. Non-human actors Service accounts have no MFA, no one watching their behaviour, and often far more privilege than they need. ## First: TLS, Always Every API call must go over TLS (HTTPS). This is non-negotiable and yet routinely violated in labs and even production, because the quick way to make a script work is to disable certificate verification when the device presents a self-signed cert. You will recognise this pattern, and you should treat it as a red flag: ``` # NEVER in production requests.get(url, verify=False) # disables cert verification ``` `verify=False` turns HTTPS back into a channel any man-in-the-middle can read and modify. The correct fixes: install the device's certificate into your trust store, or run an internal CA and have your devices present certs signed by it, then verify against that CA. On IOS XE, the RESTCONF/NETCONF endpoint uses the device's SSH/HTTPS key material, and you can replace the default self-signed certificate with one from your PKI. Effort now, versus an undetectable compromise of your automation channel later. ## Tokens, Not Passwords Most modern network APIs (Catalyst Center, SD-WAN Manager, Meraki) use token-based authentication: you authenticate once with credentials, receive a short-lived token, and use that token on subsequent calls. This is strictly better than sending a username and password on every request, for reasons worth understanding: Short-livedA token expires (often in an hour). A leaked token is useful for a limited window; a leaked password is useful until someone changes it. RevocableYou can invalidate a single token without changing the underlying credential or affecting other clients. ScopableA token can carry a limited set of permissions, so a read-only script never holds write access. The corollary: treat the token like a live credential. Do not log it, do not print it, do not commit it. Request it, hold it in memory, use it, let it expire. A token in a CI job's console output is a credential in a place many people can read. ## Never Hardcode Secrets The single most common real-world API security failure is a credential committed to a git repository. It is so common that automated bots scan public repos for API keys within seconds of a push. The discipline: - **Environment variables** for local scripts, never literals in the code. - **A secrets manager** (HashiCorp Vault, AWS Secrets Manager, Ansible Vault) for anything beyond a personal script. The script requests the secret at runtime; it is never at rest in the code. - **A .gitignore that excludes credential files**, and pre-commit hooks that scan for secrets before they can be committed. - **Rotate anything that leaks.** A credential that was ever in a repo, even in deleted history, is compromised. Rotate it; do not just remove the file. This is not paranoia. Leaked network automation credentials are a documented, routine attack vector, precisely because they combine high privilege with poor hygiene. ## Least Privilege for Service Accounts The account your automation uses should have exactly the permissions it needs and nothing more. A monitoring script that only reads interface stats needs read-only access; giving it a privilege-15 write-capable account "so it works" is how a bug or a compromise becomes a network-wide outage. Concretely, using the AAA infrastructure from the [AAA article](https://www.pinglabz.com/aaa-tacacs-radius-device-administration/): - Create a dedicated service account per automation system, never shared with humans. - Scope it with TACACS+ command authorization or an API role to only the operations it performs. - Restrict where it can connect from (a management ACL permitting only the automation host's address, as covered in [management-plane hardening](https://www.pinglabz.com/hardening-management-plane-ios-xe/)). - Enable accounting so every API-driven change is attributable to that service account, giving you an audit trail. ## Read the Response Codes Security is not only about the request; it is about noticing when something is wrong in the response. An automation pipeline that ignores HTTP status codes will happily continue after an authentication failure or a partial change, and that silence is dangerous. The ones that matter for security: 401 UnauthorizedYour token is missing, invalid, or expired. Stop, re-authenticate, do not retry blindly. 403 ForbiddenAuthenticated but not authorized. A least-privilege boundary is doing its job, or your role is wrong. 429 Too Many RequestsRate-limited. Back off. Ignoring this looks like an attack and can lock the account. 200 / 204 SuccessThe call worked (204 = success, no body). Do not assume success; check for it. These codes, and how to handle them in a script, are covered in depth in the automation cluster's [REST API response codes article](https://www.pinglabz.com/rest-api-response-codes-network/). From a security angle the rule is simple: never let a pipeline plough on after a 401 or 403\. That is the moment to stop and alert. ## The Practical Checklist TransportTLS always. Verify certificates. Never `verify=False` in production. AuthShort-lived tokens over per-call passwords. Never log the token. SecretsEnv vars or a secrets manager. Never in code or git. Rotate on any leak. PrivilegeDedicated, least-privilege service accounts. Source-restricted. Accounted for. ErrorsHandle 401/403/429\. Stop on auth failure. Alert, do not retry blindly. ## FAQ ### Is disabling certificate verification ever acceptable? Only in a throwaway lab you fully control and that carries no real credentials. The moment a real token crosses that connection, verification must be on. The habit of `verify=False` is dangerous precisely because it works, so it survives into production. ### Where should I store API tokens? In memory for the life of the script, requested at runtime from an environment variable or secrets manager. Not on disk, not in the code, not in logs. ### What is the difference between 401 and 403? 401 means "I do not know who you are" (bad or missing token). 403 means "I know who you are, and you are not allowed to do this" (authenticated but unauthorized). The fix differs: re-authenticate for 401, check your role for 403. ### Do I need a secrets manager for a small script? For a personal, local script, environment variables are an acceptable minimum. For anything shared, scheduled, or in CI, use a real secrets manager. The line is "will this credential ever exist somewhere another person or system can read it." ### How does this relate to RESTCONF on my routers? Directly. RESTCONF is a REST API on the device, and every principle here applies: TLS with a real certificate, authenticated access, least privilege, and handling the response codes. See the automation cluster for the hands-on RESTCONF walkthrough. ## Key Takeaways - APIs are the new management plane, and a compromised automation credential has a blast radius of the entire network, not one device. - **TLS always, verify certificates.** `verify=False` is a red flag that turns HTTPS back into cleartext. - Prefer **short-lived, revocable, scoped tokens** over per-call passwords. Never log or commit them. - Never hardcode secrets. Use environment variables or a secrets manager, exclude them from git, and rotate anything that leaks, including from deleted history. - Give automation **dedicated, least-privilege service accounts**, source-restricted and accounted for. - Handle response codes: stop on 401/403, back off on 429, and never let a pipeline continue blindly after an auth failure. Next: [hardening the management plane](https://www.pinglabz.com/hardening-management-plane-ios-xe/), or the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/). ### MACsec Explained: 802.1AE Link Encryption URL: https://www.pinglabz.com/macsec-802-1ae-explained/ Last updated: 2026-07-12T02:04:08.000Z Almost every security control on a network protects traffic at Layer 3 and above. IPsec encrypts IP payloads, TLS encrypts sessions, and VPNs tunnel across untrusted networks. But all of them assume the Layer 2 link underneath is just a wire that faithfully carries frames. MACsec breaks that assumption in the useful direction: it encrypts and authenticates every single Ethernet frame on the link itself, at Layer 2, so an attacker who physically taps the cable sees nothing but ciphertext. This article explains MACsec (IEEE 802.1AE), how MKA distributes the keys, and where it fits. It extends the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/). ## A Platform Note, Up Front This site's rule is that every capture is real. MACsec is one of the few features that cannot be demonstrated on the virtual platforms in the lab, and the honest thing is to say so and explain why. ``` SW1(config)# interface Ethernet0/0 SW1(config-if)# macsec ^ % Invalid input detected at '^' marker. ``` The IOL-L2 switch (and IOL-XE router) reject the interface-level `macsec` command outright. Interestingly, the `key chain X macsec` global command *is* accepted, which shows the software has partial awareness, but the data-plane encryption is not implemented, because MACsec encryption happens in dedicated hardware in the PHY, and a virtual switch has no such silicon. That is not a limitation of this lab so much as a fact about MACsec: it is a hardware feature. Real MACsec runs on Catalyst 9000 series (9200/9300/9500), Catalyst IE industrial switches, ASR 900/1000, and Nexus platforms with MACsec-capable line cards. If you need to lab it, that is the hardware. What follows is the concept, the standard, and honest deployment guidance, with no fabricated MKA session output. ## What MACsec Actually Does MACsec operates hop by hop between two directly connected devices (switch-to-switch, or host-to-switch). On each link it provides: Confidentiality The frame payload is encrypted (AES-GCM, 128 or 256-bit). A tap sees ciphertext. Integrity Each frame carries an ICV (Integrity Check Value). Tamper with a frame and it is rejected. Replay protection A packet number prevents captured frames from being replayed later. Origin authenticity Confirms the frame came from the authenticated peer on the link, not an injected impostor. The frame gets two new pieces: a **SecTAG** (a header, after the source MAC, carrying the association number and packet number) and an **ICV** (a trailer, the cryptographic integrity value). The original source and destination MACs stay in the clear so switches can still forward, but everything else, including the EtherType and the entire payload, is protected. Because it runs in PHY hardware, MACsec adds essentially no latency and operates at line rate. This is the reason it is a hardware feature and the reason it cannot be virtualized: doing AES-GCM on every frame at 100 Gbps in software is not feasible, so it is offloaded to the interface silicon. ## MKA: How the Keys Get There Encryption is easy once both ends share a key. The hard part is agreeing on the key securely, and that is the MACsec Key Agreement protocol (MKA, defined in 802.1X-2010). Two ways to establish the initial trust: Pre-shared key (manual)You configure a Connectivity Association Key (CAK) and name (CKN) on both ends via a key chain. Simple, good for switch-to-switch uplinks. 802.1X EAP (dynamic)The CAK is derived during 802.1X authentication against a RADIUS server. Ideal for host-to-switch, where each device authenticates and gets its own key. From that pre-shared or EAP-derived CAK, MKA elects a key server, which generates the actual encryption key (the SAK, Secure Association Key), distributes it to the peers, and rotates it periodically without dropping traffic. That rotation is why MKA exists rather than just configuring a static key: session keys change regularly, limiting the value of any single compromised key. The pre-shared-key path uses a key chain (the same `key chain X macsec` construct the IOL software half-accepted), and integrates with the RADIUS infrastructure from the [AAA article](https://www.pinglabz.com/aaa-tacacs-radius-device-administration/) for the EAP path. ## The Configuration Shape On a MACsec-capable platform, the pre-shared-key uplink config looks like this (shown for reference, not captured, per the platform note): ``` key chain MKA-KC macsec key 01 cryptographic-algorithm aes-256-cmac key-string ! mka policy MKA-POLICY macsec-cipher-suite gcm-aes-256 ! interface TenGigabitEthernet1/0/1 mka policy MKA-POLICY mka pre-shared-key key-chain MKA-KC macsec ``` On working hardware you would then verify with `show macsec summary`, `show mka sessions` (looking for a *Secured* state), and `show macsec statistics` to watch encrypted and protected frame counters climb. Those commands return nothing meaningful on a platform without MACsec silicon, which is precisely why this article does not show them. ## Where MACsec Fits, and Where It Does Not MACsec is not a VPN and does not replace one. It secures individual links, hop by hop. A frame is decrypted at each MACsec-terminating device and re-encrypted on the next link. This has consequences: - **It protects the physical infrastructure.** The classic use case is a link between two buildings over a fibre run you do not fully control, or a connection through a colo cross-connect. Anyone tapping that specific cable gets ciphertext. - **It is not end-to-end.** Traffic is in the clear inside each device between hops. If you need true end-to-end confidentiality across many hops, that is IPsec or TLS, at Layer 3/4, not MACsec. - **Host-to-switch MACsec** (via 802.1X) protects the access link from the endpoint to the first switch, complementing [802.1X port authentication](https://www.pinglabz.com/802-1x/): 802.1X decides who gets on, MACsec encrypts what they send. - **It stops the passive tap and the physical-layer attacker.** That is its threat model: someone with physical access to a cable or patch panel. Against a compromised switch in the path, it does less, because that switch legitimately decrypts. The pairing to remember: MACsec for the wire, IPsec for the route, TLS for the session. They protect different scopes and are frequently deployed together, MACsec on the sensitive physical links, IPsec across the WAN, TLS for the applications. ## FAQ ### Can I run MACsec in a virtual lab? Not meaningfully. MACsec encryption happens in PHY hardware, which virtual switches do not have. IOL rejects the interface command. To lab it you need Catalyst 9000, IE switches, or ASR-class hardware. ### Is MACsec a replacement for IPsec? No. MACsec secures a single Layer 2 link, hop by hop. IPsec secures Layer 3 end to end across many hops. Different scope, often deployed together. ### Does MACsec add latency? Negligibly. Because it runs in the interface silicon at line rate, the overhead is a few bytes per frame (SecTAG + ICV) and essentially no added delay, which is exactly why it needs hardware. ### PSK or 802.1X for the keys? Pre-shared key for switch-to-switch uplinks (simple, static peers). 802.1X/EAP for host-to-switch, where each endpoint authenticates and derives its own key dynamically. ### Does MACsec protect against a compromised switch in the path? Only partially. Each MACsec device legitimately decrypts and re-encrypts, so a compromised switch sees cleartext. MACsec's threat model is the passive physical tap on a link, not a subverted node. For that, use end-to-end encryption higher in the stack. ## Key Takeaways - MACsec (802.1AE) encrypts and authenticates **every Ethernet frame** on a link, at Layer 2, in PHY hardware at line rate. - It provides confidentiality (AES-GCM), integrity (ICV), replay protection, and origin authenticity, adding a SecTAG header and ICV trailer while leaving the MAC addresses visible for forwarding. - **MKA** handles key agreement and rotation, seeded either by a pre-shared key (switch-to-switch) or 802.1X/EAP (host-to-switch, via RADIUS). - It is **hop-by-hop, not end-to-end**. Traffic is cleartext inside each device between links. Pair it with IPsec (route) and TLS (session). - Its threat model is the physical tap on a cable you do not fully control, cross-building fibre, colo cross-connects, sensitive uplinks. - **Honest platform note:** MACsec cannot run on virtual IOL. It needs Catalyst 9000, IE, ASR, or Nexus hardware with MACsec silicon. Next: [802.1X port authentication](https://www.pinglabz.com/802-1x/) (the natural partner for host MACsec), or the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/). ### Advanced ACLs on IOS XE: Time-Based, Object-Group, and IPv6 ACLs URL: https://www.pinglabz.com/advanced-acls-cisco-ios-xe/ Last updated: 2026-07-12T02:04:08.000Z The access list you learned for the CCNA (permit this host, deny that subnet) still works, but it does not scale and it does not express intent. When you have forty admin hosts, a policy that only applies during business hours, and an IPv6 network to protect alongside the IPv4 one, a flat list of permit and deny lines becomes an unmaintainable wall. Advanced ACL features exist to make policy readable, reusable, and time-aware. This article covers time-based ACLs, object groups, and IPv6 ACLs on Cisco IOS XE, with real output including a schedule that is genuinely inactive because of when the lab ran. It extends the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/). ## Time-Based ACLs Some rules should only apply at certain times: allow contractor access during business hours only, block recreational traffic during work hours, permit a maintenance path only during the change window. A time-based ACL attaches a `time-range` to an ACE, and the ACE is only active when the time-range is. ``` R1(config)# time-range BUSINESS-HOURS R1(config-time-range)# periodic weekdays 8:00 to 18:00 ! R1(config)# ip access-list extended ADV-ACL R1(config-ext-nacl)# permit tcp object-group TRUSTED-ADMINS any eq 22 R1(config-ext-nacl)# permit object-group WEB-SERVICES any any time-range BUSINESS-HOURS R1(config-ext-nacl)# deny ip any any log ``` The time-range supports `periodic` (recurring, like "weekdays" or "weekend" or specific days) and `absolute` (a one-off window with a start and end date, ideal for a temporary contractor or a scheduled event). Here is the genuinely useful part, captured live. This lab ran on a Saturday, and the time-range is `weekdays` only. Watch what the router reports: ``` R1#show time-range time-range entry: BUSINESS-HOURS (inactive) periodic weekdays 8:00 to 18:00 used in: IP ACL entry ``` ``` R1#show ip access-lists ADV-ACL Extended IP access list ADV-ACL 10 permit tcp object-group TRUSTED-ADMINS any eq 22 20 permit object-group WEB-SERVICES any any time-range BUSINESS-HOURS (inactive) 30 deny ip any any log ``` `**(inactive)**` on both the time-range and the ACE that uses it, because it is the weekend. That is not a mockup; it is the router correctly reporting that line 20 is currently dormant and, right now, web traffic falls through to the `deny` on line 30\. Come Monday at 08:00 the same commands would show `(active)` and line 20 would start permitting. The time-range is evaluated against the router's clock in real time, which is exactly why [accurate NTP](https://www.pinglabz.com/ntp-cisco-ios-xe/) is a prerequisite for time-based security: a router with the wrong time enforces the wrong policy. ## Object Groups The scaling problem. You have twenty admin hosts that should reach SSH on every device. Without object groups, that is twenty permit lines, repeated in every ACL that references them, and updated by hand in every place whenever an admin joins or leaves. Object groups let you name the set once and reference the name. ``` R1(config)# object-group network TRUSTED-ADMINS R1(config-network-group)# host 192.168.99.100 R1(config-network-group)# host 10.0.10.101 ! R1(config)# object-group service WEB-SERVICES R1(config-service-group)# tcp eq 80 R1(config-service-group)# tcp eq 443 ``` Two kinds of group do most of the work: Network object-group A named set of hosts, subnets, or ranges. Reference it where a source or destination would go. Service object-group A named set of protocols and ports. Reference it where the protocol/port would go. Now the ACL reads like policy, not plumbing: `permit tcp object-group TRUSTED-ADMINS any eq 22` means "the admins may SSH anywhere." Add a new admin by editing the group once, and every ACL that references it updates automatically. The router expands the groups internally, so the enforcement is identical; what changes is maintainability. ``` R1#show object-group Network object group TRUSTED-ADMINS host 192.168.99.100 host 10.0.10.101 Service object group WEB-SERVICES tcp eq www tcp eq 443 ``` Object groups can also nest (a group of groups), which is how large enterprises model "all datacentre subnets" or "all management stations" as a single named object built from smaller ones. ## IPv6 ACLs and the ND Trap IPv6 ACLs work like extended IPv4 ACLs with three differences that catch people out. First, there are no standard vs extended flavours; all IPv6 ACLs are named and behave like extended ones. Second, the syntax is `ipv6 access-list` and it is applied with `ipv6 traffic-filter`, not `ip access-group`. Third, and this is the one that causes outages: ``` R1(config)# ipv6 access-list ADV-V6 R1(config-ipv6-acl)# permit tcp any host 2001:DB8:10::1 eq 22 R1(config-ipv6-acl)# permit icmp any any nd-ns R1(config-ipv6-acl)# permit icmp any any nd-na R1(config-ipv6-acl)# deny ipv6 any any log ``` Those two `nd-ns` and `nd-na` lines are not optional. IPv6 relies on Neighbor Discovery (Neighbor Solicitation and Neighbor Advertisement) the way IPv4 relies on ARP, but unlike ARP, ND rides inside ICMPv6, which means an IPv6 ACL can block it. If your ACL ends in `deny ipv6 any any` without first permitting ND, the router can no longer resolve its neighbors and IPv6 connectivity on that segment collapses, including your own management session. Every IPv6 ACL that ends in a deny must explicitly permit ND first. This is the single most common way people lock themselves out with IPv6 ACLs, and it has no IPv4 equivalent because ARP is not IP traffic and ACLs never touched it. ## Logging and the Performance Note The `log` keyword on the final deny (as in both ACLs above) records what is being blocked, which is invaluable for both security monitoring and for discovering the rule you forgot to add. But logging punts matched packets for logging, which costs CPU, so: - Log the **deny** lines, not the permits. You want to see what is blocked, not a firehose of normal traffic. - Use `log` (rate-limited) rather than `log-input` on high-traffic ACLs unless you specifically need the ingress interface and source MAC. - On a busy edge, even logged denies can be a lot. Consider CoPP protection for the logging path, or sample rather than log everything. ## Order Still Matters None of these features change the fundamental rule: ACLs are processed top-down, first match wins, and there is an implicit `deny` at the end. Object groups and time-ranges make each line richer, but the evaluation order is unchanged. Put your most specific and most-hit rules near the top, your broad denies at the bottom, and remember that a time-range that is inactive causes its ACE to be skipped entirely, so traffic falls through to the next line, exactly as the Saturday capture above demonstrated. ## FAQ ### Do object groups change how the ACL enforces? No. The router expands the group into individual entries internally. Enforcement is identical; the benefit is purely maintainability and readability. ### My time-based ACL is not working. Why? Check the router's clock and time zone (`show clock`). A time-range is evaluated against the router's local time, so wrong time or wrong zone means the wrong window. Sync NTP. ### Why did my IPv6 ACL break connectivity? You almost certainly denied Neighbor Discovery. Add `permit icmp any any nd-ns` and `permit icmp any any nd-na` before any broad deny. ### Can I use object groups in IPv6 ACLs? IPv6 object-group support exists on modern IOS XE but is less complete than IPv4\. Verify on your platform; some releases support network object-groups in IPv6 ACLs, others are more limited. ### Absolute or periodic time-range? Periodic for recurring schedules (business hours, weekends). Absolute for one-off windows with real dates (a two-week contractor, a specific maintenance night). ## Key Takeaways - **Time-based ACLs** attach a `time-range` to an ACE so it only applies on schedule. The lab captured a `weekdays` range showing `(inactive)` on a Saturday, with the dependent ACE dormant. - Time-based security depends on an accurate clock. Sync NTP or you enforce the wrong policy at the wrong time. - **Object groups** name sets of hosts (network) or ports (service) once and reference the name, so a policy change is one edit instead of many. Enforcement is unchanged; maintainability transforms. - **IPv6 ACLs** must explicitly `permit icmp any any nd-ns` and `nd-na` before any broad deny, or Neighbor Discovery breaks and IPv6 connectivity collapses. This has no IPv4 equivalent. - Log the deny lines, not the permits, and mind the CPU cost of logging on busy links. - Order is unchanged: top-down, first match, implicit deny at the end. An inactive time-range simply skips its ACE. Next: [hardening the management plane](https://www.pinglabz.com/hardening-management-plane-ios-xe/), or the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/). ### IPv6 First-Hop Security Part 2: ND Inspection, Snooping, and Source Guard URL: https://www.pinglabz.com/ipv6-nd-inspection-source-guard/ Last updated: 2026-07-12T02:04:07.000Z RA Guard and DHCPv6 Guard stop the loud attacks: rogue routers and rogue DHCP servers. But there is a quieter layer of IPv6 abuse that operates below them, at the level of individual addresses. A host that spoofs its neighbor's IPv6 address, or poisons the switch's understanding of who lives where, can intercept traffic without ever sending an RA. Stopping that requires the switch to actually know which address belongs on which port, and then enforce it. That knowledge is the **device-tracking binding table**, and the features that build and enforce it are ND Inspection, IPv6 Snooping, and IPv6 Source Guard. This article covers all three. It is the companion to [RA Guard and DHCPv6 Guard](https://www.pinglabz.com/ipv6-ra-guard-dhcpv6-guard/) and extends the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/). ## Everything Depends on the Binding Table The foundation of all IPv6 first-hop security is a table that maps, for each entry: an IPv6 address, the MAC address that owns it, the switch port it lives behind, and the VLAN. This is device tracking (the IPv6 analogue of the DHCP snooping binding table in IPv4). ``` SW1(config)# ipv6 nd inspection policy HOST-NDI SW1(config)# vlan configuration 1 SW1(config-vlan-config)# device-tracking attach-policy HOST-NDI ``` Once device tracking is running, the switch learns bindings by watching ND traffic (Neighbor Solicitations, Neighbor Advertisements, DAD) and DHCPv6 exchanges. Every legitimate host announces its address through normal ND, and the switch records it. That table then becomes the ground truth every other feature checks against. On modern IOS XE the older `ipv6 nd inspection` and `ipv6 snooping` commands are being folded into the unified `device-tracking` policy framework. The switch tells you so directly when you use the legacy syntax: ``` SW1(config-vlan-config)# ipv6 nd inspection WARNING: ipv6 nd inspection attach-policy to vlan deprecated, use device-tracking attach-policy ``` Take the hint and use `device-tracking` policies on current software. The concepts are identical; the CLI is consolidated. ## ND Inspection: Validating the Neighbor Discovery Itself Neighbor Discovery is IPv6's replacement for ARP, and like ARP it is unauthenticated and forgeable. ND Inspection validates ND messages against the binding table and drops the ones that lie. The two abuses it stops: NA spoofing A host sends a Neighbor Advertisement claiming to own an address that belongs to someone else, redirecting that victim's traffic to itself. The IPv6 version of ARP poisoning. NS/DAD abuse A host answers every Duplicate Address Detection probe, so no other host can ever configure an address. A denial-of-service on address assignment. ND Inspection cross-checks the source address and MAC in each ND message against the binding table. If a Neighbor Advertisement claims an address that the table says lives on a different port or MAC, it is a spoof, and it is dropped. The victim's traffic keeps flowing to the real owner. ## IPv6 Source Guard: Enforcing at the Data Plane ND Inspection validates the *control* traffic (the ND messages). Source Guard validates the *data* traffic. It is the enforcement teeth: for every packet arriving on a port, it checks the source IPv6 address against the binding table, and if the source does not match what the table says should be on that port, the packet is dropped. ``` SW1(config)# ipv6 source-guard policy HOST-SG SW1(config)# interface Ethernet0/1 SW1(config-if)# ipv6 source-guard attach-policy HOST-SG ``` The effect: a host on Et0/1 can only send packets sourced from the IPv6 address the binding table recorded for Et0/1\. It cannot spoof its neighbor's address, cannot spoof the gateway, cannot spoof anything but itself. This is the IPv6 equivalent of IP Source Guard in IPv4, and it closes the loop that RA Guard and ND Inspection open: RA Guard stops rogue gateways, ND Inspection stops address theft in the control plane, and Source Guard stops spoofing in the data plane. ## How the Layers Stack The four features are not alternatives. They are a defence in depth, each closing a gap the others leave open: Device trackingBuilds the binding table. The foundation everything else reads. RA GuardStops rogue gateways (Router Advertisements). DHCPv6 GuardStops rogue DHCPv6 servers (malicious DNS). ND InspectionStops address theft in ND (the control plane). Source GuardStops source spoofing in forwarded traffic (the data plane). Deploy them together on access ports and a compromised or malicious host is boxed in: it cannot be the gateway, cannot be the DNS server, cannot steal a neighbor's address, and cannot even spoof a source in the packets it sends. That is the IPv6 access-edge posture every campus should aim for. ## Verifying the Binding Table The command that shows you the ground truth: ``` SW1#show device-tracking database Network Layer Address Link Layer Address Interface vlan prlvl age state ND 2001:DB8:10::5054:FF:... 5054.00dc.9702 Et0/1 1 0005 30s REACHABLE ND FE80::5054:FF:FEDC:9702 5054.00dc.9702 Et0/1 1 0005 30s REACHABLE ``` Each row is a validated binding: this address, on this MAC, behind this port. Source Guard permits only traffic matching these rows; ND Inspection permits only ND matching them. If a host's entries are missing, its traffic may be dropped, which is the number-one operational gotcha (see below). ## Operational Reality These features are powerful and they will break things if deployed carelessly. The honest guidance: - **Binding table timing.** A host that has been silent may age out of the table, and then Source Guard drops its traffic until it speaks ND again. Tune the reachable/stale/down timers, and understand that a machine returning from sleep may have a brief blackout until it re-registers. - **Static entries for silent devices.** Some devices (certain appliances, printers) barely emit ND. Add static binding entries for them, or Source Guard will treat them as spoofers. - **Trust the uplinks.** Device tracking, like RA Guard, distinguishes host ports from infrastructure ports. Get the roles wrong and you either fail to protect or you break legitimate traffic on trunks. - **Roll out in stages.** Deploy device tracking first and let the binding table populate and stabilise. Then add ND Inspection. Add Source Guard last, and watch for drops of legitimate but quiet hosts. Enabling all four at once on a live network is a way to generate a lot of tickets. - **Platform variance.** The exact CLI and capabilities differ across Catalyst generations and IOS XE releases, and the legacy vs `device-tracking` syntax coexists awkwardly. Verify on your specific platform and code. ## FAQ ### Do I need all four features? For a strong posture, yes, they cover different attack surfaces. If you deploy only two, RA Guard and DHCPv6 Guard stop the highest-impact attacks with the lowest operational risk. ND Inspection and Source Guard add depth but need more care. ### What is the difference between ND Inspection and Source Guard? ND Inspection validates ND control messages against the binding table. Source Guard validates data-plane packet sources against it. One protects the neighbor-discovery process, the other protects forwarding. ### Why is legitimate traffic being dropped? Almost always a missing or aged-out binding-table entry. Check `show device-tracking database` for the host. Silent devices need static entries or longer timers. ### Is this related to DHCP snooping? It is the IPv6 parallel. The device-tracking binding table plays the role that the DHCP snooping binding table plays for IPv4 Dynamic ARP Inspection and IP Source Guard. ### Does device tracking work with SLAAC and DHCPv6 both? Yes. It learns bindings from ND (which covers SLAAC addresses) and from DHCPv6 exchanges. Both addressing models are covered. ## Key Takeaways - All IPv6 first-hop security rests on the **device-tracking binding table**: which address, on which MAC, behind which port. - **ND Inspection** validates Neighbor Discovery messages against that table, stopping NA spoofing and DAD abuse (the IPv6 ARP-poisoning equivalents). - **IPv6 Source Guard** validates data-plane packet sources against the table, so a host can only send from its own address. - Together with RA Guard and DHCPv6 Guard, they box in a malicious host completely: no rogue gateway, no rogue DNS, no address theft, no source spoofing. - On modern IOS XE, use the unified `device-tracking` policy framework; the switch flags the legacy `ipv6 nd inspection` syntax as deprecated. - Deploy in stages (tracking, then inspection, then source guard) and add static bindings for silent devices, or you will drop legitimate traffic. Next: [RA Guard and DHCPv6 Guard](https://www.pinglabz.com/ipv6-ra-guard-dhcpv6-guard/) if you skipped it, or the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/). ### IPv6 First-Hop Security Part 1: RA Guard and DHCPv6 Guard URL: https://www.pinglabz.com/ipv6-ra-guard-dhcpv6-guard/ Last updated: 2026-07-12T02:04:07.000Z In IPv4, a rogue DHCP server is a nuisance you can contain with DHCP snooping. In IPv6, the equivalent threat is worse and easier to trigger, because IPv6 hosts configure themselves from Router Advertisements, and *any* device on the segment can send one. A single misconfigured Windows laptop with connection sharing enabled, or a deliberate attacker, can broadcast an RA and reconfigure the default gateway for every host on the VLAN. This is not exotic. It happens by accident constantly. IPv6 First-Hop Security (FHS) is the switch-side toolkit that stops it. This article configures RA Guard and DHCPv6 Guard and proves the block by flooding rogue RAs and confirming the victims never see them. It extends the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/) and builds on the [IPv6 cluster guide](https://www.pinglabz.com/ipv6/). ## Why IPv6 Is More Exposed Recall from [how IPv6 addressing works](https://www.pinglabz.com/dhcpv6-stateful-stateless-explained/): the router advertises a prefix in an RA, and hosts autoconfigure from it. The RA also carries the default gateway. Critically, the RA is unauthenticated ND traffic, and the host has no way to tell a legitimate RA from a forged one. It believes whatever it hears. So the attack is trivial: Rogue RA Attacker sends an RA advertising itself as the gateway. Hosts route their traffic through the attacker. Man-in-the-middle, or a black hole. Rogue DHCPv6 Attacker answers DHCPv6 with a malicious DNS server. Hosts resolve names through the attacker. Accidental RA No malice needed. A host with IPv6 sharing on emits RAs and hijacks the segment by mistake. The most common cause. ## The Principle: RAs Only Come From Router Ports The insight behind RA Guard is that on any real network, you know which ports connect to routers and which connect to hosts. A legitimate RA only ever arrives on a router-facing port. An RA arriving on a host-facing access port is, by definition, either an attack or a misconfiguration. RA Guard enforces exactly that: it inspects ND traffic and drops RAs on ports designated as host ports. You declare each port's role with a policy: ``` SW1(config)# ipv6 nd raguard policy HOST-PORTS SW1(config-nd-raguard)# device-role host ! SW1(config)# ipv6 dhcp guard policy CLIENT-PORTS ! SW1(config)# interface Ethernet0/1 SW1(config-if)# ipv6 nd raguard attach-policy HOST-PORTS SW1(config-if)# ipv6 dhcp guard attach-policy CLIENT-PORTS ! SW1(config)# interface Ethernet0/2 SW1(config-if)# ipv6 nd raguard attach-policy HOST-PORTS SW1(config-if)# ipv6 dhcp guard attach-policy CLIENT-PORTS ``` `device-role host` is the operative setting: this port faces hosts, so RAs and DHCPv6-server messages arriving here are illegitimate and get dropped. The uplink toward the real router is left without the policy (or given `device-role router`), so legitimate RAs flow normally. ``` SW1#show ipv6 nd raguard policy HOST-PORTS RA guard policy HOST-PORTS configuration: device-role host Policy HOST-PORTS is applied on the following targets: Target Type Policy Feature Target range Et0/1 PORT HOST-PORTS RA guard vlan all Et0/2 PORT HOST-PORTS RA guard vlan all ``` ## DHCPv6 Guard The same idea, applied to DHCPv6 server messages. A DHCPv6 *server* reply (ADVERTISE, REPLY) should only come from a port where a legitimate server or relay lives. DHCPv6 Guard drops server-role DHCPv6 messages on host ports, blocking the rogue-DNS attack. The policy above (`CLIENT-PORTS` attached to the host ports) means "clients live here, so anything pretending to be a DHCPv6 server on this port is dropped." A more elaborate policy can also verify the server's address against an allow-list and check the advertised preference, but at minimum, the role distinction blocks the obvious rogue. ## Proof: Flood a Rogue RA, Watch It Vanish Two hosts sit behind guarded access ports: H1 on Et0/1 and H2 on Et0/2\. The legitimate router R1 is behind the unguarded uplink. H2 plays attacker, flooding rogue RAs that advertise a bogus prefix (2001:dead:beef::/64) and try to install itself as a gateway. The RA is crafted with a raw ICMPv6 socket, no special tools needed: ``` H2$ python3 rogue_ra.py sent 15 rogue RAs advertising 2001:dead:beef::/64 ``` The switch's RA Guard debug catches the ND traffic hitting the guarded port: ``` SW1#debug ipv6 snooping raguard *Jul 12 01:37:31.624: SISF[RAG]: Et0/2 vlan 1 RS received by RA guard on Et0/2 from FE80::1C07:FDFF:FEF3:91BD ``` But the real proof is on the victim. Here is H1, behind the other guarded port, showing its addresses and routes after the flood: ``` H1$ ip -6 addr show dev eth0 scope global inet6 2001:db8:10:0:5054:ff:fedc:9702/64 scope global dynamic mngtmpaddr proto kernel_ra H1$ ip -6 route show | grep -E 'dead|db8:10' 2001:db8:10::/64 dev eth0 proto kernel metric 256 ``` Read that carefully, because it is the entire feature in two commands. H1 configured itself an address from **2001:db8:10::/64**, the legitimate prefix advertised by R1 (note `proto kernel_ra`, meaning "I learned this from a Router Advertisement"). And the rogue prefix **2001:dead:beef::** is *nowhere*. Not in an address, not in a route. H2 flooded it fifteen times and H1 never saw a single one, because RA Guard dropped every rogue RA at H2's access port before it could reach the rest of the VLAN. The legitimate RA from R1 (arriving on the unguarded uplink) passed through untouched. The rogue RA from H2 (arriving on a guarded host port) was silently discarded. Exactly the intended outcome, proven end to end. ## FHS Is a Family RA Guard and DHCPv6 Guard are the two you deploy first, but IPv6 First-Hop Security is a suite, and it depends on a shared foundation called **device tracking** (a binding table of which IPv6 address lives on which port and MAC). The companion piece, [ND Inspection, Snooping, and Source Guard](https://www.pinglabz.com/ipv6-nd-inspection-source-guard/), covers the rest: building that binding table and using it to stop address spoofing and ND cache poisoning. RA GuardDrops rogue Router Advertisements on host ports. (This article.) DHCPv6 GuardDrops rogue DHCPv6 server messages on host ports. (This article.) ND Inspection + SnoopingBuilds the device-tracking binding table from observed ND. (Companion article.) IPv6 Source GuardDrops traffic whose source does not match the binding table. (Companion article.) ## FAQ ### Which ports get the policy? Host-facing access ports get `device-role host`. Router-facing uplinks and trunks to other switches get no RA Guard, or an explicit `device-role router` policy. Never put a host policy on a real router port, or you drop the legitimate RA and break IPv6 for the segment. ### Does RA Guard need SEND or RA authentication? No. RA Guard is a stateless port-role filter, which is why it is deployable in practice. RFC 3971 SEND provides cryptographic RA authentication but is almost never deployed because hosts and infrastructure rarely support it. RA Guard is the pragmatic answer. ### Can an attacker fragment the RA to evade RA Guard? Historically yes, fragmented RAs were an evasion. Modern IOS XE RA Guard handles fragmentation, and RFC 6980 forbids fragmented ND, which most stacks now enforce. Keep switch software current. ### Is this the same as DHCP snooping? Conceptually parallel, but IPv6-specific. DHCP snooping protects IPv4 DHCP; DHCPv6 Guard protects IPv6 DHCP; RA Guard has no IPv4 equivalent because IPv4 has no Router Advertisements. ## Key Takeaways - IPv6 hosts trust Router Advertisements blindly, and any device can send one. A rogue RA hijacks the gateway; an accidental one (host with sharing on) does it by mistake. - **RA Guard** drops RAs on ports declared `device-role host`. Legitimate RAs only arrive on router-facing ports. - **DHCPv6 Guard** does the same for rogue DHCPv6 server messages, blocking the rogue-DNS attack. - The lab proved it end to end: H2 flooded 15 rogue RAs for `2001:dead:beef::`, and victim H1 learned **only** the legitimate `2001:db8:10::` prefix. The rogue never appeared. - Apply host policies to access ports only. A host policy on a router uplink breaks IPv6. - RA Guard and DHCPv6 Guard are the first two of the FHS suite; ND Inspection and Source Guard complete it. Next: [ND Inspection, Snooping, and Source Guard](https://www.pinglabz.com/ipv6-nd-inspection-source-guard/), or the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/). ### Unicast RPF: Strict vs Loose Mode URL: https://www.pinglabz.com/unicast-rpf-strict-loose/ Last updated: 2026-07-12T02:04:06.000Z Source IP addresses are trivially forged. Nothing in the IP header proves the source is who it claims to be, and a packet with a lie in the source field is the raw material for reflection attacks, spoofed floods, and a whole category of nuisance traffic. Unicast Reverse Path Forwarding (uRPF) is the router feature that calls the bluff: for each arriving packet, it asks "do I actually have a route back to this source?" and drops the ones where the answer is no. This article configures strict uRPF and proves it by firing spoofed packets at the router and counting the drops. It extends the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/). ## The Idea in One Sentence Normal forwarding looks up the *destination* to decide where a packet goes. uRPF adds a second lookup on the *source*: it checks the routing table (specifically CEF, the [Cisco Express Forwarding](https://www.pinglabz.com/ip-routing/) table) to see whether the source address is reachable, and if so, whether it is reachable via the interface the packet arrived on. Fail the check, drop the packet. The intuition: if a packet claiming to be from 8.8.8.8 arrives on your customer-facing interface, but your route to 8.8.8.8 points out your upstream interface, that packet is lying. A real packet from 8.8.8.8 would never arrive on the customer port. uRPF drops it. ## Strict vs Loose: The Whole Decision There are two modes, and choosing wrong either breaks legitimate traffic or fails to catch the spoofing. Strict mode (reachable-via rx) The source must be reachable **via the exact interface the packet arrived on**. Strongest anti-spoofing. Breaks on asymmetric routing. Use onSingle-homed edges, access ports Loose mode (reachable-via any) The source must be reachable via **any** interface (just present in the routing table). Weaker, but survives asymmetry. Use onMulti-homed cores, ISP peering The trade-off is entirely about **asymmetric routing**, where a packet's return path differs from its arrival path. Strict mode assumes symmetry: it expects traffic from a source to arrive on the same interface you would use to reach that source. On a single-homed customer link that assumption holds, and strict mode is the gold standard. On a multi-homed router where traffic legitimately arrives on one link and leaves on another, strict mode drops perfectly valid packets, and you need loose mode. Loose mode drops only the truly unroutable: source addresses that appear nowhere in your routing table at all (bogons, unallocated space, obvious garbage). It cannot catch a spoofed source that happens to be a real, routable address arriving on the wrong interface, because it does not care which interface. It is the compromise you make when you cannot guarantee symmetry. ## Configuration uRPF is one line per interface, and it requires CEF (which is on by default). Strict mode on the edge interface facing the untrusted side: ``` R1(config)# interface Ethernet0/0 R1(config-if)# ip verify unicast source reachable-via rx ``` `rx` is strict (received-interface). `any` would be loose. That is the entire difference in the config. Verify it is armed: ``` R1#show ip interface Ethernet0/0 | include verify|drop IP verify source reachable-via RX 0 verification drops 0 suppressed verification drops 0 verification drop-rate ``` Zero drops so far, because nothing has spoofed yet. ## Proof: Spoof It and Count the Drops The lab sends two batches of packets from a Linux host on R1's edge (192.168.99.0/24). First, 20 **spoofed** packets with a source of 203.0.113.66, an address R1 has no route to via that interface. Then, as a control, 5 **legitimate** packets with the host's real source. ``` j@llmbits:~$ sudo python3 -c " from scapy.all import IP, ICMP, send send(IP(src='203.0.113.66', dst='10.0.12.2')/ICMP(), count=20, iface='ens224') send(IP(src='192.168.99.100', dst='10.0.12.2')/ICMP(), count=5, iface='ens224') " sent 20 spoofed 203.0.113.66 -> 10.0.12.2 sent 5 legit 192.168.99.100 -> 10.0.12.2 ``` Now the router: ``` R1#show ip interface Ethernet0/0 | include verify|drop IP verify source reachable-via RX 20 verification drops 0 suppressed verification drops 0 verification drop-rate R1#show ip traffic | include drop|RPF 0 no route, 20 unicast RPF, 0 forced drop, 0 unsupported-addr ``` **Exactly 20 drops.** Every spoofed packet was caught. And the 5 legitimate packets? They added zero to the drop counter, because 192.168.99.100 *is* reachable via Ethernet0/0, so strict uRPF let them through. The counter matching the spoof count precisely, with the legitimate traffic untouched, is the feature working exactly as specified. The `20 unicast RPF` line in `show ip traffic` is the global tally, useful for confirming the drops are uRPF and not something else (a route missing, an ACL). uRPF drops have their own bucket. ## This Is BCP 38 uRPF is the practical implementation of BCP 38 (RFC 2827), the internet's long-standing best practice on source-address validation. The idea is simple and the impact is large: if every network filtered packets whose source addresses it could not legitimately originate, entire classes of reflection and amplification DDoS attacks would become far harder to launch. An ISP applying strict uRPF on its customer access ports guarantees that a customer can only send packets sourced from the address space that customer was assigned. A compromised machine behind that port cannot spoof arbitrary sources for an attack. This is why uRPF on the access edge is one of the highest-leverage security controls a provider can deploy, and why it belongs on enterprise edges too. ## The ACL Escape Hatch Sometimes you need uRPF but with exceptions, DHCP being the classic case. A DHCP client sends its initial DISCOVER with a source of 0.0.0.0, which is not reachable by definition and which strict uRPF will cheerfully drop, breaking DHCP. uRPF accepts an ACL to permit exceptions: ``` R1(config)# ip access-list extended URPF-ALLOW R1(config-ext-nacl)# permit ip host 0.0.0.0 any R1(config)# interface Ethernet0/0 R1(config-if)# ip verify unicast source reachable-via rx URPF-ALLOW ``` Packets that fail the uRPF check are then re-checked against the ACL; a permit overrides the drop. Use this sparingly and specifically. A broad permit ACL defeats the entire purpose. ## Where Strict uRPF Bites - **Multihoming and asymmetry.** The number-one cause of strict-mode false drops. If your network has any asymmetric paths, strict mode on the affected interfaces will drop legitimate traffic. Use loose mode there. - **Default routes and loose mode.** If an interface has a default route (0.0.0.0/0), loose uRPF considers *every* source reachable via that route and becomes a no-op. Loose mode is meaningful only where you carry specific routes. - **DHCP and 0.0.0.0.** As above. Add the ACL exception or DHCP breaks. - **Placement.** uRPF belongs on interfaces facing untrusted or customer networks, checking traffic *coming in*. Applying it on a core-facing interface toward your own network is usually pointless and occasionally harmful. ## FAQ ### Strict or loose, if I am not sure? If the interface is single-homed and faces users or customers, strict. If the router is multi-homed or you have any asymmetric routing, loose. When in doubt on a core box, loose is the safe default; strict on an access edge is the strong choice. ### Does uRPF need CEF? Yes. uRPF does its source lookup in the CEF table. CEF is on by default on modern IOS XE; if it is disabled, uRPF will not function. ### Will loose mode stop spoofing? Only spoofing with unroutable (bogon) sources. A spoof using a real, routable address arriving on the wrong interface passes loose mode. That is the price of surviving asymmetry. ### Does uRPF replace anti-spoofing ACLs? It complements them and scales better. An ACL must be maintained as your address space changes; uRPF follows the routing table automatically. Many networks run both: uRPF for the dynamic check, an ACL for specific known-bad ranges. ### Can I see which sources were dropped? The counters tell you how many. To see which, add logging via an ACL exception with `log`, or use CoPP/ACL logging alongside. uRPF itself only counts. ## Key Takeaways - uRPF adds a **source** lookup to forwarding: is this source reachable, and (in strict mode) via the interface it arrived on? - **Strict** (`reachable-via rx`) is the strongest anti-spoofing but breaks on asymmetric routing. **Loose** (`reachable-via any`) survives asymmetry but only catches unroutable sources. - The lab proved it exactly: **20 spoofed packets, 20 verification drops**, while 5 legitimate packets passed untouched. - uRPF is the implementation of **BCP 38** source-address validation. On access edges it stops compromised hosts from spoofing. - Watch the traps: multihoming (use loose), default routes (make loose a no-op), and DHCP's 0.0.0.0 source (add an ACL exception). - Place it on untrusted-facing interfaces, checking inbound traffic. Next: [Control Plane Policing](https://www.pinglabz.com/control-plane-policing-copp/) to protect the CPU, or the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/). ### Control Plane Policing (CoPP): Protecting the CPU URL: https://www.pinglabz.com/control-plane-policing-copp/ Last updated: 2026-07-12T02:04:06.000Z Every router has a CPU, and that CPU is a shared, finite resource that both keeps the network running (OSPF hellos, BGP updates, your SSH session) and can be trivially overwhelmed by traffic aimed at it. A flood of pings to the router's own interface, a burst of malformed packets that must be punted for inspection, or an SSH brute-force can starve the CPU of the cycles it needs to maintain routing adjacencies. When that happens, the control plane collapses and the network goes down, without a single link failing. Control Plane Policing (CoPP) is the seatbelt. It is a QoS policy applied to traffic destined *for the router itself*, rate-limiting each category so no single class of control-plane traffic can consume the whole CPU. This article builds a CoPP policy and then floods the router to watch it drop the attack while protecting everything else. It extends the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/). ## Data Plane, Control Plane, Management Plane CoPP only makes sense once you are clear about which traffic it touches. Data plane Traffic passing *through* the router, forwarded in hardware/CEF. Never touches the CPU. Control plane Traffic *to* the router that the CPU must process: routing protocols, ARP, SSH, SNMP, ICMP to its own IP. **This is what CoPP protects.** Management plane The subset used to manage the box: SSH, SNMP, NetConf. A logical slice of the control plane. The insight that makes CoPP click: a packet destined for the router's own IP **cannot** be CEF-switched. It has to be handed to the CPU, because the CPU is the destination. That punt is expensive, and it is exactly the mechanism an attacker abuses. CoPP sits at the point where that punted traffic enters the control plane and applies a rate limit before it can pile up. ## CoPP Is Just MQC Aimed at the Control Plane CoPP uses the standard Modular QoS CLI (class-map, policy-map, service-policy) that the [QoS cluster](https://www.pinglabz.com/qos/) covers in depth. The only twist is where you attach it: not an interface, but the special `control-plane` pseudo-interface. Step 1: classify the control-plane traffic into categories with ACLs and class-maps. ``` R1(config)# ip access-list extended COPP-ICMP R1(config-ext-nacl)# permit icmp any any R1(config)# ip access-list extended COPP-SSH R1(config-ext-nacl)# permit tcp any any eq 22 R1(config)# ip access-list extended COPP-ROUTING R1(config-ext-nacl)# permit ospf any any R1(config-ext-nacl)# permit pim any any ! R1(config)# class-map match-all CM-ICMP R1(config-cmap)# match access-group name COPP-ICMP R1(config)# class-map match-all CM-SSH R1(config-cmap)# match access-group name COPP-SSH R1(config)# class-map match-all CM-ROUTING R1(config-cmap)# match access-group name COPP-ROUTING ``` Step 2: police each class in a policy-map. This is where the security decisions live. ``` R1(config)# policy-map CoPP-POLICY R1(config-pmap)# class CM-ICMP R1(config-pmap-c)# police 8000 conform-action transmit exceed-action drop R1(config-pmap)# class CM-SSH R1(config-pmap-c)# police 32000 conform-action transmit exceed-action transmit R1(config-pmap)# class CM-ROUTING R1(config-pmap-c)# police 500000 conform-action transmit exceed-action transmit R1(config-pmap)# class class-default R1(config-pmap-c)# police 32000 conform-action transmit exceed-action transmit ``` Step 3: attach it, input, to the control plane. ``` R1(config)# control-plane R1(config-cp)# service-policy input CoPP-POLICY ``` ## The Thinking Behind the Numbers The rates are not arbitrary, and the conform/exceed actions encode a threat model: ICMP: exceed-action dropPing to the router is useful but never urgent. Cap it hard and drop the overflow. This is the ping-flood defence. SSH: exceed-action transmitYou do not want to lock yourself out during an event. Police it but do not drop; a burst of legitimate admin traffic still gets through. Routing: high rate, transmitOSPF and PIM are the network's heartbeat. Give them plenty of headroom and never drop them, or CoPP causes the outage it exists to prevent. class-default: policedEverything you did not classify. A modest cap catches the unexpected, but tune carefully so you do not starve something you forgot to name. The cardinal rule of CoPP: **never drop your routing protocols, and be careful about SSH.** A too-aggressive CoPP policy is one of the classic self-inflicted outages, because it drops the very traffic that keeps adjacencies up or keeps you logged in to fix things. Start with `exceed-action transmit` everywhere, watch the counters for a week to learn your baseline, and only then change the classes you understand to `drop`. ## The Flood, and the Drops Policy in place. Now a real attack: a Linux host floods the router's interface with `ping -f` (flood ping, as fast as the host can send): ``` j@llmbits:~$ sudo ping -f -w 6 192.168.99.1 ....(thousands of dots).... --- 192.168.99.1 ping statistics --- 414 packets transmitted, 39 received, 90.5797% packet loss, time 5991ms ``` **90.6 percent packet loss.** That is not the network failing. That is CoPP working exactly as designed, dropping the flood so it cannot reach and overwhelm the CPU. The 39 packets that got through are the ones that fit under the 8000 bps ICMP cap; everything above it was discarded at the control-plane edge. The router's own view confirms it precisely: ``` R1#show policy-map control-plane input class CM-ICMP Control Plane Service-policy input: CoPP-POLICY Class-map: CM-ICMP (match-all) Match: access-group name COPP-ICMP police: cir 8000 bps, bc 1500 bytes conformed 597 packets, 58506 bytes; actions: transmit exceeded 3817 packets, 374066 bytes; actions: drop conformed 5000 bps, exceeded 12000 bps ``` 597 packets conformed and were transmitted to the CPU. 3817 packets exceeded the rate and were dropped before they got there. That is the entire value proposition on one screen: the flood was absorbed at the policer, the CPU never saw the bulk of it, and OSPF (in its own high-rate class) kept running untouched throughout. During this whole test the routing adjacency to R2 never dropped. Watch `conformed 5000 bps, exceeded 12000 bps`: the attacker was pushing roughly 17,000 bps of ICMP at a class capped at 8000, and the policer sorted the 5000 that fit from the 12000 that did not. ## CPPr: Finer Control Basic CoPP treats all control-plane traffic as one aggregate. Control Plane Protection (CPPr) subdivides it into three sub-interfaces for more surgical policy: Host Traffic to the router's own addresses: SSH, SNMP, ICMP, routing to a local interface. Transit Transit traffic that still needs punting (e.g. IP options). CEF-exception Packets CEF cannot handle: TTL-expired, unroutable, keepalives. CPPr lets you police TTL-expiry punts (a common attack and a common traceroute side effect) separately from your SSH sessions. For most enterprises, aggregate CoPP is sufficient; CPPr earns its keep on high-value edge routers. ## Pitfalls - **Forgetting a protocol.** If you police class-default aggressively and forget to classify, say, BGP or BFD, you can starve it. Enumerate every control protocol you actually run before tightening class-default. - **Dropping ARP.** ARP is control-plane traffic. An overly broad drop policy that catches ARP will break connectivity in ways that look baffling. - **Platform differences.** On many Catalyst and ASR platforms a default CoPP policy already exists in hardware. Layering your own on top requires understanding what is already there. Check before you assume the control plane is unprotected. - **Testing in production.** Build with `exceed-action transmit`, observe the counters, then tighten. A drop policy deployed blind is a coin flip. ## FAQ ### Does CoPP affect traffic passing through the router? No. CoPP only sees traffic destined for the router itself (the control plane). Transit data-plane traffic is untouched. ### Will CoPP stop a DDoS? It protects the router's CPU from being overwhelmed by control-plane-directed attack traffic, which keeps the box up and routing. It does not stop a volumetric DDoS aimed *through* the router at a victim; that is a different problem (scrubbing, uRPF, upstream filtering). ### What rate should I use? There is no universal number. Baseline your real traffic first (transmit-only policy, watch `conformed`), then set the cap comfortably above normal but below what would hurt the CPU. ### Can CoPP lock me out? Yes, if you drop SSH or your routing protocols too aggressively. This is the number-one CoPP mistake. Police management traffic with `transmit`, not `drop`. ## Key Takeaways - CoPP protects the router's **CPU** from control-plane traffic (packets destined to the router that must be punted). - It is standard MQC (class-map, policy-map) attached to the `control-plane` pseudo-interface with `service-policy input`. - In the lab a flood ping hit **90.6% loss** and the policer showed **597 conformed / 3817 dropped**, absorbing the attack while OSPF stayed up. - The conform/exceed actions are a threat model: drop excess ICMP, but only police (never drop) SSH and routing. - **Never drop your routing protocols.** An over-aggressive CoPP policy causes the outage it was meant to prevent. - Build with `exceed-action transmit`, baseline the counters, then tighten. CPPr adds host/transit/CEF-exception granularity when you need it. Next: [Unicast RPF](https://www.pinglabz.com/unicast-rpf-strict-loose/) to drop spoofed traffic at the edge, or the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/). ### AAA on Cisco IOS XE: TACACS+ vs RADIUS for Device Administration URL: https://www.pinglabz.com/aaa-tacacs-radius-device-administration/ Last updated: 2026-07-12T02:04:06.000Z The moment more than one person can log into your routers, you have an access-control problem. Who is allowed in? What are they allowed to do once they are in? And can you prove, after the fact, who typed the command that took the network down? Answering those three questions is what AAA does, and the choice between TACACS+ and RADIUS decides how well it answers them. This article configures real AAA on Cisco IOS XE, authenticates a live login against a RADIUS server, and shows the actual protocol exchange on the wire. It is the opening piece of the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/). ## The Three A's Authentication Who are you? Prove it. Username and password, a token, a certificate. Authorization Now that I know who you are, what may you do? Which commands, which privilege level. Accounting What did you do? A log of commands and sessions, for audit and forensics. The distinction between the first two matters enormously for the protocol choice, because TACACS+ can separate them cleanly and RADIUS cannot. ## TACACS+ vs RADIUS: The Real Differences They are often presented as interchangeable. They are not. They were designed for different jobs. Encryption **TACACS+** encrypts the entire payload. **RADIUS** encrypts only the password field; everything else is cleartext. Transport **TACACS+** uses TCP 49 (reliable). **RADIUS** uses UDP 1812/1813 (no guaranteed delivery). AAA separation **TACACS+** splits the three A's into separate exchanges. **RADIUS** combines authentication and authorization into one. Command authorization **TACACS+** can authorize per command, per user. **RADIUS** effectively cannot. Vendor **TACACS+** is a Cisco protocol (now an open RFC). **RADIUS** is an open IETF standard, universal. Primary use **TACACS+** \= device administration. **RADIUS** \= network access (802.1X, VPN, Wi-Fi). The one-line rule almost everyone uses: **TACACS+ for administering the devices, RADIUS for admitting the users.** When an engineer SSHes into a switch, that is device administration and TACACS+ is the natural fit, because you want per-command authorization and full-payload encryption. When a laptop authenticates onto the wired port, that is network access and RADIUS owns it, because RADIUS carries the [802.1X](https://www.pinglabz.com/802-1x/) attributes that assign a VLAN. The command-authorization point is the decisive one for device administration. With TACACS+ you can permit a junior operator to run `show` commands and `clear counters` but nothing that changes config, and every command they attempt is checked against the server in real time. RADIUS pushes a privilege level at login and then steps out of the way; it cannot vet individual commands. ## The Golden Rule: Always Have a Local Fallback Before a single line of AAA config, internalise this, because it is the difference between a good day and being locked out of your entire network: **every method list ends in `local`.** An AAA method list is an ordered list of ways to authenticate. The router tries them in order. If your only method is the AAA server and that server becomes unreachable (or you fat-finger the shared secret), you cannot log in. Anywhere. The fix is to have already been fired, or a very awkward console cable and a password-recovery procedure. ``` aaa authentication login VTY-AUTH group RAD-GROUP local ``` That trailing `local` means "if the server does not answer, fall back to the local username database." We will see it save the day at the end of this article. ## Configuring RADIUS (Live) The lab points R1 at a real FreeRADIUS server running on a Linux host at 192.168.99.100\. Turn on AAA, define the server, group it, and build the method lists: ``` R1(config)# aaa new-model R1(config)# radius server PINGLABZ-RAD R1(config-radius-server)# address ipv4 192.168.99.100 auth-port 1812 acct-port 1813 R1(config-radius-server)# key PingLabzRAD123 ! R1(config)# aaa group server radius RAD-GROUP R1(config-sg-radius)# server name PINGLABZ-RAD ! R1(config)# aaa authentication login VTY-AUTH group RAD-GROUP local R1(config)# aaa authorization exec VTY-AUTH group RAD-GROUP local ! R1(config)# line vty 0 4 R1(config-line)# login authentication VTY-AUTH R1(config-line)# authorization exec VTY-AUTH ``` The shared secret (`key`) must match on both the router and the server, and it is what encrypts the RADIUS password field. On the FreeRADIUS side, the router is registered as a client with the same secret, and a test user exists with a Cisco privilege attribute: ``` # clients.conf client pinglabz-r1 { ipaddr = 192.168.99.1 secret = PingLabzRAD123 } # users netadmin Cleartext-Password := "Cisco@123" Cisco-AVPair = "shell:priv-lvl=15" ``` The `shell:priv-lvl=15` AV-pair is how RADIUS pushes an authorization result: it drops the authenticated user straight to privilege level 15\. That is the extent of RADIUS authorization, a level at login, and it is why command authorization needs TACACS+. ### Proof it works The router can test the server group directly. Correct password, then wrong: ``` R1#test aaa group RAD-GROUP netadmin Cisco@123 legacy Attempting authentication test to server-group RAD-GROUP using radius User was successfully authenticated. R1#test aaa group RAD-GROUP netadmin WrongPassword legacy Attempting authentication test to server-group RAD-GROUP using radius User authentication request was rejected by server. ``` Accepted and rejected, correctly, against a live server. And the server relationship is healthy: ``` R1#show aaa servers | include RADIUS|State|Authen: RADIUS: id 1, priority 1, host 192.168.99.100, auth-port 1812, acct-port 1813, hostname PINGLABZ-RAD State: current UP, duration 93s, previous duration 0s Authen: request 6, timeouts 4, failover 0, retransmission 3 ``` ### The exchange on the wire This is what actually happens when the router authenticates, captured with `debug radius authentication`: ``` RADIUS(00000000): Send Access-Request to 192.168.99.100:1812 id 1645/4, len 60 RADIUS: NAS-IP-Address [4] 6 192.168.99.1 RADIUS: User-Name [1] 10 "netadmin" RADIUS: User-Password [2] 18 * RADIUS(00000000): Started 5 sec timeout RADIUS: Received from id 1645/4 192.168.99.100:1812, Access-Accept, len 63 ``` Read the security lesson in that output. `User-Name` is `"netadmin"`, in the clear. `NAS-IP-Address` is in the clear. Only `User-Password` is hidden (the `*`). Anyone sniffing this exchange learns the username and everything else; only the password is protected. This is the RADIUS design, and it is precisely why device administration prefers TACACS+, which would have encrypted the entire packet. One more thing worth noting: modern IOS XE (17.x) now nags about this. The moment you configure a RADIUS or TACACS+ server without TLS, it emits a security warning pointing you toward RadSec (RADIUS over TLS), which encrypts the whole exchange. On a management network you control this is usually acceptable; across untrusted transport it is not. ## Configuring TACACS+ and Command Authorization The device-administration setup. Note the extra method list that RADIUS could not provide, command authorization: ``` R1(config)# tacacs server PINGLABZ-TAC R1(config-server-tacacs)# address ipv4 10.0.13.2 R1(config-server-tacacs)# key PingLabzTAC123 ! R1(config)# aaa group server tacacs+ TAC-GROUP R1(config-sg-tacacs+)# server name PINGLABZ-TAC ! R1(config)# aaa authentication login CON-AUTH group TAC-GROUP local R1(config)# aaa authorization commands 15 CMD-AUTHZ group TAC-GROUP local ``` That last line is the whole reason TACACS+ exists for device administration. `aaa authorization commands 15` means every privilege-15 command a user types is sent to the TACACS+ server, which returns permit or deny before the command runs. The server holds a policy like "user bob may run `show` and `configure` but not `reload`," enforced command by command. RADIUS has no equivalent. ``` R1#show tacacs Tacacs+ Server - public : Server name: PINGLABZ-TAC Server address: 10.0.13.2 Server port: 49 Server Status: Alive ``` ## The Local Fallback, In Action Here is why the golden rule matters, demonstrated. The TACACS+ server in the lab is unreachable. Watch what the router does: ``` R1#test aaa group TAC-GROUP admin Cisco@123 legacy Attempting authentication test to server-group TAC-GROUP using tacacs+ No authoritative response from any server. ``` The group fails. But the method list was `aaa authentication login CON-AUTH group TAC-GROUP local`. When the group returns no authoritative response, the router moves to the next method: `local`. A user in the local database can still log in. The network is not lost. This is not a hypothetical. AAA servers go down, get isolated by the very outage you are trying to fix, or reject you because someone changed a shared secret. The `local` keyword is the seatbelt. A method list of `group TAC-GROUP` with no fallback is a loaded gun pointed at your own access. Configure a local account with a strong secret on every device, exactly for this. It is the account you will use at 3am when the AAA infrastructure is the thing that broke. ## Accounting: The Third A Authentication and authorization control access. Accounting records it, and it is the part people skip until an auditor asks who ran a command. ``` R1(config)# aaa accounting commands 15 default start-stop group TAC-GROUP R1(config)# aaa accounting exec default start-stop group TAC-GROUP ``` With this, every privilege-15 command and every login/logout is logged to the server with a username, timestamp, and source. That is your audit trail. TACACS+ command accounting is genuinely valuable here because it captures the exact command string; RADIUS accounting is coarser, oriented toward session start/stop and byte counts rather than individual commands. ## FAQ ### Can I run TACACS+ and RADIUS at the same time? Yes, and it is common: TACACS+ for device admin logins, RADIUS for 802.1X network access, both pointing at the same identity store (often Cisco ISE, which speaks both). They serve different method lists. ### Which is more secure? TACACS+ for device administration, because it encrypts the full payload and the username is not exposed. For network access the question is moot; RADIUS is what carries the 802.1X attributes. ### What happens if I forget the local fallback and the server dies? You are locked out of every device that used that method list, and you are looking at console access and password recovery on each one. Never omit `local`. ### Why does IOS XE warn about TLS now? Because plain RADIUS and TACACS+ shared-secret encryption is dated. RadSec (RADIUS/TLS) and TACACS+ over TLS encrypt the transport. On a trusted management network the warning is informational; over untrusted links, act on it. ### Is the shared secret the same as the user password? No. The shared secret authenticates the *router* to the *server* and encrypts fields. The user password authenticates the person. Two completely separate credentials. ## Key Takeaways - AAA is three questions: who are you (authentication), what may you do (authorization), what did you do (accounting). - **TACACS+ for administering devices, RADIUS for admitting users.** TACACS+ encrypts the whole payload and does per-command authorization; RADIUS carries 802.1X attributes. - The live RADIUS debug shows the username and everything but the password in the clear. That exposure is why device admin prefers TACACS+. - **Every method list ends in `local`.** The lab proved it: the TACACS+ server was unreachable, and the local fallback kept access alive. - RADIUS authorization is a privilege level at login (`shell:priv-lvl=15`). TACACS+ authorization is command by command. - Turn on accounting, or you cannot prove who did what. - IOS XE 17.x nudges you toward RadSec/TLS. Heed it on untrusted transport. Next: [hardening the management plane](https://www.pinglabz.com/hardening-management-plane-ios-xe/), or the [Infrastructure Security cluster guide](https://www.pinglabz.com/infrastructure-security/). ### Conditional Debugging on Cisco IOS XE: Debug Without Melting the Box URL: https://www.pinglabz.com/conditional-debug-cisco-ios-xe/ Last updated: 2026-07-12T01:23:30.000Z There is a particular kind of fear that stops engineers from using the most powerful troubleshooting tool on a Cisco router. You type `debug ip packet`, the console erupts, the CPU pins at 100 percent, the box stops forwarding, and someone asks why the site went down. After that happens once, "never debug on a production router" becomes received wisdom. It is the wrong lesson. The problem was never debugging; it was *unfiltered* debugging. Conditional debug lets you ask the router a narrow question and get a narrow answer, and this article proves it works by running two pings and showing you that only one of them appears in the output. It extends the [IP Services cluster guide](https://www.pinglabz.com/ip-services/). ## Why Debug Melts the Box Understand the mechanism and the fix becomes obvious. Normal forwarding happens in CEF, in the fast path, and never troubles the CPU. When you enable `debug ip packet`, the router must inspect and report on packets, which means **punting** them to the CPU for process switching. Two things then compound: - Every packet now costs CPU cycles instead of being switched in hardware or optimised code. - Every packet generates a log message, which itself costs CPU, and if console logging is on, that message is written to a 9600-baud console port. The console becomes the bottleneck, and the router blocks on it. On a router forwarding 50,000 packets per second, that is 50,000 punts and 50,000 log messages per second. The CPU has no chance. This is a self-inflicted denial of service, and the router will not stop to warn you. The fix is not to avoid debug. The fix is to make the router only punt and report on the handful of packets you actually care about. ## First: The Safety Net Before you enable any debug on a box that matters, three things: ``` R1(config)# no logging console R1(config)# logging buffered 64000 debugging R1(config)# service timestamps debug datetime msec ``` `no logging console` is the single most important line here. It sends debug output to a memory buffer instead of the console port, which removes the 9600-baud bottleneck entirely. You then read the buffer at your leisure with `show logging`. Many production outages attributed to "debug" were really attributed to console logging. `service timestamps debug datetime msec` gives millisecond timestamps, which you will want the moment you are trying to work out the order of events. And know your escape hatch before you need it: `undebug all` (or `u all`, which you can type in under a second while panicking). ## The Main Technique: Condition on an ACL This is the one to reach for first, and the one most people never learn. `debug ip packet` accepts an access list, and only packets permitted by that ACL are debugged. ``` R1(config)# access-list 99 permit 10.0.10.100 R1# debug ip packet 99 IP packet debugging is on for access list 99 ``` The router now punts and reports on packets matching ACL 99, and leaves everything else in the fast path. On a busy router that is the difference between a hundred messages and a hundred thousand. ### Proving it actually filters Configuration is a claim; here is the evidence. With the debug above running, the lab sends two pings from the same interface, at the same time, to two different hosts on the same subnet. Only one of them matches ACL 99. ``` R1#ping 10.0.10.100 source Ethernet0/2 repeat 2 Sending 2, 100-byte ICMP Echos to 10.0.10.100, timeout is 2 seconds: !! Success rate is 100 percent (2/2), round-trip min/avg/max = 2/2/3 ms R1#ping 10.0.10.10 source Ethernet0/2 repeat 2 Sending 2, 100-byte ICMP Echos to 10.0.10.10, timeout is 2 seconds: .! Success rate is 50 percent (1/2), round-trip min/avg/max = 3/3/3 ms ``` Both pings ran. Both generated traffic on the same interface. Now the debug output: ``` R1#show logging | include IP: Jul 12 00:57:56.702: IP: s=10.0.10.100 (Ethernet0/2), d=10.0.10.1 (nil), len 100, input feature, Input-Flexible-NetFlow(20), rtype 0, forus FALSE, sendself FALSE, mtu 0, fwdchk FALSE Jul 12 00:57:56.702: IP: s=10.0.10.100 (Ethernet0/2), d=10.0.10.1 (nil), len 100, input feature, MCI Check(110), rtype 0, forus FALSE, sendself FALSE, mtu 0, fwdchk FALSE Jul 12 00:57:56.702: IP: s=10.0.10.100 (Ethernet0/2), d=10.0.10.1 (nil), len 100, rcvd 2 Jul 12 00:57:56.702: IP: s=10.0.10.100 (Ethernet0/2), d=10.0.10.1 (nil), len 100, stop process pak for forus packet Jul 12 00:57:56.702: IP: tableid=0, s=10.0.10.100 (Ethernet0/2), d=10.0.10.1 (Ethernet0/2) nexthop=10.0.10.1, routed via RIB Jul 12 00:57:56.705: IP: s=10.0.10.100 (Ethernet0/2), d=10.0.10.1 (nil), len 100, rcvd 2 Jul 12 00:57:56.705: IP: tableid=0, s=10.0.10.100 (Ethernet0/2), d=10.0.10.1 (Ethernet0/2) nexthop=10.0.10.1, routed via RIB ``` **Every single line is `s=10.0.10.100`.** There is not one line about 10.0.10.10, despite that traffic crossing the identical interface at the identical moment. The ACL filtered it out before it ever became a debug message, which means those packets were never punted and never cost the CPU anything. That is the whole technique, demonstrated. The router answered a narrow question with a narrow answer. ### Reading the output The line itself is dense but rewards attention: `s=` / `d=`Source and destination, with the interface the packet arrived on in parentheses. `rcvd 2`Receive code. 2 means "for us, and we are going to process it." `routed via RIB`The forwarding decision, and which table made it (`tableid=0` is the global table; a VRF shows a different id). `input feature, ...`The feature chain each packet traverses, in order. That feature chain is quietly one of the most useful things in the output. Look at what the router volunteered: `Input-Flexible-NetFlow(20)`. The debug is telling us this interface has a [NetFlow monitor](https://www.pinglabz.com/flexible-netflow-cisco-ios-xe/) applied, and showing us exactly where in the packet's journey it sits. When you are debugging why a packet is being dropped or mangled, the feature chain tells you which feature to blame, and in what order they ran. ## The One Thing That Will Confuse You You will sometimes enable `debug ip packet` with a perfectly good ACL, send traffic you know is crossing the router, and see **nothing at all**. This is not a broken debug. It is CEF. `debug ip packet` only sees **process-switched** packets. Traffic that is CEF-switched in the fast path never touches the CPU and therefore never generates a debug message. Packets destined *to* the router (like the pings above, which are addressed to the router's own interface) are process-switched and show up. Packets merely transiting the router are CEF-switched and do not. Notice the lab output says `forus FALSE ... stop process pak for forus packet`, which is the router narrating this very distinction. The classic workaround is `no ip route-cache` on the interface to force process switching. **Do not do this on a production router.** You have just disabled hardware forwarding, and you will cause the exact outage you were trying to avoid. If you need to see transit traffic, use a packet capture instead: [SPAN or ERSPAN](https://www.pinglabz.com/span-rspan-erspan-configuration/), or the Embedded Packet Capture feature below. ## Conditional Debug for Protocol Debugs The ACL trick works for `debug ip packet`. For protocol debugs (BGP, OSPF, DHCP, and so on), the tool is `debug condition`, which applies a global filter to any debug that is condition-aware. ``` R1#debug condition interface Ethernet0/2 Condition 1 set R1#show debug condition Condition 1: interface Et0/2 (1 flags triggered) Flags: Et0/2 ``` With that condition set, condition-aware debugs will only report events on Ethernet0/2\. The available conditions include: - `debug condition interface ` - `debug condition ip
` - `debug condition vrf ` (invaluable on an [MPLS L3VPN](https://www.pinglabz.com/mpls/) PE with hundreds of customers) - `debug condition username ` - `debug condition standby ` Clear them with `no debug condition all`, and always check `show debug condition` before you conclude a debug is broken. A leftover condition from last week is a superb way to convince yourself a protocol is dead when it is merely being filtered. **An honest caveat:** not every debug honours conditions. Condition support is per-debug, and the coverage is inconsistent. If you set a condition and still get a firehose, that debug is not condition-aware and you need a different approach. ## The Modern Answer: Embedded Packet Capture On current IOS XE, there is a better tool for most "what is happening to this packet" questions, and it does not punt anything to the CPU: ``` R1# monitor capture CAP interface Ethernet0/2 both R1# monitor capture CAP match ipv4 host 10.0.10.100 any R1# monitor capture CAP buffer size 10 R1# monitor capture CAP start ... R1# monitor capture CAP stop R1# show monitor capture CAP buffer brief R1# monitor capture CAP export flash:cap.pcap ``` EPC captures into a memory buffer, sees CEF-switched traffic (unlike `debug ip packet`), and exports a real pcap you can open in Wireshark. For packet-level questions on a modern box, reach for this first. Debug remains the right tool when you want to see the router's *decision-making* (why it chose that next hop, why it rejected that BGP update) rather than the packets themselves. ## The Discipline 1\. Turn off console logging`no logging console` \+ `logging buffered`. Non-negotiable. 2\. Condition before you enableWrite the ACL or set the condition *first*. Never enable a broad debug "just to see." 3\. Watch the CPU`show processes cpu sorted | ex 0.00` in another window. If it climbs, `u all`. 4\. Turn it offA forgotten debug is a scheduled outage. `show debug` should be empty when you leave. ## FAQ ### Why does my debug show nothing even though traffic is flowing? CEF. `debug ip packet` only sees process-switched packets, and transit traffic is CEF-switched. Use Embedded Packet Capture or SPAN instead. ### Is it ever safe to run debug on a production router? Yes, with an ACL or condition, console logging off, and a finger on `u all`. The blanket "never debug in production" rule costs more outages than it prevents, because it pushes people toward guessing. ### How do I stop everything immediately? `undebug all`, abbreviated `u all`. Learn to type it without thinking. ### Does the ACL in `debug ip packet 99` also drop traffic? No. It is used purely as a match filter for the debug. It is not applied to any interface and does not affect forwarding. ### Why did I get output for one interface when my condition said another? That debug is probably not condition-aware. Condition support varies by debug. Verify with `show debug condition` and, if in doubt, use the ACL method. ## Key Takeaways - Debug melts routers because it **punts packets to the CPU** and writes to a slow console, not because debugging is inherently dangerous. - `no logging console` \+ `logging buffered` removes most of the risk before you type anything else. - `debug ip packet ` is the workhorse. In the lab, two simultaneous pings produced debug output for **only** the ACL-matched source. Zero lines for the other. - The debug output names the **feature chain** each packet traverses (`Input-Flexible-NetFlow`, `MCI Check`), which tells you which feature to blame and in what order they ran. - `debug ip packet` cannot see CEF-switched transit traffic. Do **not** disable route-cache in production to work around this. - `debug condition` filters protocol debugs by interface, IP, VRF, or user. Not every debug honours it. - On modern IOS XE, **Embedded Packet Capture** is usually the better tool: it sees CEF traffic, does not punt, and exports a pcap. - A forgotten debug is a scheduled outage. `show debug` before you log off. Next: the [IP Services cluster guide](https://www.pinglabz.com/ip-services/). ### PTP vs NTP: When Microseconds Matter URL: https://www.pinglabz.com/ptp-vs-ntp-time-sync/ Last updated: 2026-07-12T01:23:29.000Z Most networks need to agree on what time it is to within a second or so, and NTP has done that job well for forty years. Then there are the networks where a second is an eternity: a trading floor where transaction ordering is legally auditable, a 5G fronthaul link where radio units must be phase-aligned, a substation where protection relays trip on microsecond timing, a broadcast plant where audio and video must stay locked. For those, NTP is not merely imprecise. It is structurally incapable of the job. PTP (Precision Time Protocol, IEEE 1588) exists because of a specific insight about *why* NTP is limited, and understanding that insight is more useful than memorising either protocol. This article extends the [IP Services cluster guide](https://www.pinglabz.com/ip-services/). ## A Platform Note, Up Front Every other article in this cluster shows real captures from the lab. This one shows real NTP output, because the lab genuinely runs NTP, and it shows **no PTP output at all**, because the platforms available here do not support it: ``` R1(config)# ptp mode boundary ^ % Invalid input detected at '^' marker. ``` IOL-XE has no PTP. Neither does the Catalyst 8000v build used elsewhere in this lab. This is not a limitation of virtualization in the abstract; it is a consequence of what PTP fundamentally requires, and that turns out to be the most instructive fact in the whole article. We will come back to it. PTP genuinely runs on Catalyst 9000 series switches, Cisco IE industrial switches, ASR 900 series, and NCS platforms. If you need to lab it, that is the hardware. What follows is the mechanism, the real NTP numbers for contrast, and honest guidance. No fabricated PTP output appears anywhere on this site. ## The Core Problem: Asymmetry Both protocols solve the same puzzle. A client wants to know the server's time, but the answer takes time to arrive, so by the time it lands it is already stale. Both correct for this the same way: measure the round trip, assume the delay was **symmetric**, and subtract half of it. ``` offset = ((T2 - T1) + (T3 - T4)) / 2 delay = ((T4 - T1) - (T3 - T2)) / 2 T1 = client sent request T2 = server received it T3 = server sent reply T4 = client received it ``` That halving is the whole ballgame. It is exactly correct if the path took the same time in both directions, and it is wrong by **half the asymmetry** if it did not. A path that is 10ms out and 20ms back does not produce a 5ms error; it produces an error you cannot even detect, because the arithmetic has no way to notice. Everything that makes a network fast and efficient makes this worse. Queueing delay varies with load. Different forward and return paths have different lengths. A switch under congestion holds your timing packet in a buffer for an unpredictable interval. Every one of these injects asymmetry, and every microsecond of asymmetry becomes half a microsecond of undetectable clock error. ### You can watch it happen Here is NTP in the lab, with the client synchronised to a master one hop away. The link between them is deliberately degraded (120ms latency, 40ms jitter, 8 percent loss) using link conditioning: ``` R2#show ntp status Clock is unsynchronized, stratum 4, reference is 1.1.1.1 nominal freq is 250.0000 Hz, actual freq is 250.0000 Hz, precision is 2**10 reference time is EDFD6558.B8D4FFF0 (00:59:04.722 UTC Sun Jul 12 2026) clock offset is -52.5000 msec, root delay is 105.00 msec root dispersion is 73.22 msec, peer dispersion is 3.01 msec loopfilter state is 'FREQ' (Drift being measured) system poll interval is 64, last update was 180 sec ago. ``` ``` R2#show ntp associations address ref clock st when poll reach delay offset disp *~1.1.1.1 127.127.1.1 3 180 64 36 105.00 -52.500 3.019 * sys.peer, # selected, + candidate, - outlyer, x falseticker, ~ configured ``` Read the numbers. **Root delay 105ms. Clock offset -52.5ms. Dispersion 3ms.** The offset is almost exactly half the delay, which is the algorithm doing precisely what it was designed to do. And it is *guessing*: NTP has no way to know whether that 105ms was split evenly. The `dispersion` figure of 3ms is NTP's own honest estimate of its uncertainty. Note also `Clock is unsynchronized` and `loopfilter state is 'FREQ'`: even while it has selected a peer (the `*` in the associations output), the router is refusing to declare itself synchronised because the path quality is too poor to trust. NTP knows when it is being lied to by the network, and says so. That output is the case for PTP, made by NTP itself. ## What PTP Does Differently PTP does not use a smarter formula. It uses the same one. What it changes is **where the timestamps are taken** and **what the network in between does**. Hardware timestamping The timestamp is taken by the **PHY**, as the bits hit the wire. NTP timestamps in software, so the OS scheduler, interrupt latency, and driver queues are all inside the measurement. Transparent clocks Every switch in the path **measures how long it held the packet** and writes that into a correction field. Queueing delay stops being invisible and becomes a known quantity. Boundary clocks A switch terminates PTP on one side and re-originates it on the other, so error does not accumulate across a long path. Each hop is a fresh, short measurement. BMCA The Best Master Clock Algorithm elects a grandmaster automatically from clock quality attributes. No manual stratum config. The transparent clock is the conceptual leap. NTP treats the network as an opaque box that adds an unknown delay. PTP makes the network **participate**: each switch confesses exactly how long it delayed the packet, so the receiver can subtract it. Asymmetry caused by queueing simply stops existing as a source of error. And that is precisely why the lab cannot run it. **PTP requires the forwarding hardware to timestamp packets in the PHY and rewrite a correction field at line rate.** A virtual router has no PHY. There is no silicon to do it. This is not Cisco declining to implement a feature in a virtual image; it is a feature that has nowhere to live without hardware. The command being rejected is the honest answer. ## The Numbers That Matter NTP over the internet10 to 100 milliseconds. Fine for logs, certificates, Kerberos. NTP on a good LANUnder 1 millisecond, typically a few hundred microseconds. PTP, software timestampingTens of microseconds. Better than NTP, but the OS is still in the loop. PTP, hardware timestamping + transparent clocks**Sub-microsecond**, often tens of nanoseconds. This is the whole point. Roughly four orders of magnitude between well-run NTP and well-run PTP. But note the third row: **PTP without hardware support is barely better than NTP.** Deploying PTP on switches that lack transparent-clock silicon buys you complexity and very little accuracy. This is the most common way PTP projects disappoint. ## Which Do You Need? Be honest about the requirement, because PTP is a genuine infrastructure commitment. **NTP is correct for:** log correlation, certificate validity, Kerberos and AD authentication (which tolerates about 5 minutes), scheduled jobs, and roughly 95 percent of enterprise networks. If someone cannot articulate a sub-millisecond requirement, they need NTP. **PTP is required for:** - **Financial trading.** MiFID II in Europe mandates 100-microsecond traceability to UTC for high-frequency venues. That is not achievable with NTP. - **Mobile fronthaul.** 5G TDD requires radio units phase-aligned to about 1.5 microseconds, or adjacent cells interfere with each other. - **Power utilities.** IEC 61850 substations run protection relays and synchrophasors on microsecond timing. - **Broadcast.** SMPTE 2110 replaced physical genlock cabling with PTP over the IP network. Lose PTP and the plant loses sync. - **Industrial automation.** Motion control across a factory floor, where axes must move in lockstep. What all of these share: the timing is not for humans reading logs, it is a functional input to the system. If time is wrong, the thing does not work, as opposed to being mildly annoying to debug. ## They Run Together This is not an either/or decision. A trading floor runs PTP for the timestamping infrastructure and NTP for the wiki, the jump hosts, and the printers. A 5G RAN runs PTP to the radios and NTP on the management network. PTP is expensive in the ways that matter: every switch in the timing path must support it, you need a grandmaster with a GNSS receiver (and therefore a roof antenna, and therefore a conversation with the building owner), and you need to think about GNSS holdover and spoofing. Nobody deploys it for the fun of it. The pragmatic architecture: NTP everywhere as the baseline, PTP as a dedicated overlay on the specific paths that need it, with the PTP grandmaster itself disciplined by GNSS and NTP as its fallback. ## Doing NTP Properly, Since That Is What You Will Actually Run Given that NTP is what most networks need, it is worth doing well: - **Use at least four sources.** NTP's selection algorithm needs a quorum to detect a falseticker. Two sources is the worst possible number: when they disagree, the client cannot tell which one is lying. - **Authenticate.** `ntp authenticate` plus `ntp trusted-key`. An unauthenticated NTP client will believe anyone, and moving a device's clock is a real attack (it can expire or un-expire certificates, and break log correlation during an incident). - **Restrict access.** `ntp access-group`. An open NTP server is an amplification vector for DDoS. - **Hierarchy, not a mesh.** A couple of stratum-2 routers peering upstream, everything else pointing at them. The `ntp master` command on a router that has no real reference is a lie the rest of your network will believe. - **Monitor the offset, not just reachability.** A device that says it is synchronised at 200ms of offset is a device you want to know about. ## FAQ ### Can I run PTP in a virtual lab? Not meaningfully. PTP's accuracy depends on PHY-level hardware timestamping and transparent-clock silicon, neither of which a virtual router has. IOL-XE rejects the config outright. You can run software PTP (linuxptp) between VMs to see the protocol exchange, but the accuracy will be NTP-grade and the exercise is about protocol mechanics, not timing. ### Is PTP more accurate because it uses a better algorithm? No. It uses essentially the same offset/delay maths. It is more accurate because it timestamps in hardware and because the switches in between report their own queueing delay instead of hiding it. ### Does PTP need every switch in the path to support it? For sub-microsecond accuracy, yes. A single non-PTP-aware switch in the path reintroduces exactly the unmeasured queueing delay that PTP exists to eliminate. This is why PTP deployments are infrastructure projects, not config changes. ### What is a grandmaster? The clock at the top of the PTP hierarchy, elected by the BMCA, usually disciplined by GNSS. Roughly the equivalent of a stratum-1 NTP server, but elected dynamically rather than configured. ### Why does my NTP client say "unsynchronized" when it has a peer? Because it does not trust the path yet. As the lab output shows, a peer can be selected (`*`) while the router still refuses to declare sync, because dispersion and delay are too high or the drift is still being measured (`loopfilter state is 'FREQ'`). Give it time, or fix the path. ## Key Takeaways - Both protocols use the same offset maths and both assume the path is **symmetric**. Asymmetry is the entire source of error. - PTP is not more accurate because of better arithmetic. It is more accurate because it timestamps in the **PHY** and because **transparent clocks** report their own queueing delay instead of hiding it. - Because PTP needs timestamping silicon, it cannot run on virtual platforms. IOL-XE rejects `ptp mode` outright. That rejection is the concept, demonstrated. - Real NTP under a bad path, from the lab: root delay 105ms, offset -52.5ms, and the router honestly reporting `Clock is unsynchronized`. - PTP without hardware support is barely better than NTP. Half-deployed PTP is the classic disappointment. - NTP is right for \~95% of networks. PTP is required where time is a functional input: trading, 5G fronthaul, IEC 61850 substations, SMPTE 2110 broadcast, motion control. - Run NTP properly: four or more sources, authenticated, access-listed, hierarchical. Next: the [IP Services cluster guide](https://www.pinglabz.com/ip-services/). ### DHCPv6 Explained: Stateful, Stateless, SLAAC, and the Relay URL: https://www.pinglabz.com/dhcpv6-stateful-stateless-explained/ Last updated: 2026-08-01T19:35:47.000Z IPv6 host addressing is where a lot of otherwise competent engineers quietly lose confidence, because there are three mechanisms that can all be running at once, they interact through two obscure flag bits in a router advertisement, and a host can end up with several addresses from different sources simultaneously. That is not a bug. It is the design. This article untangles SLAAC, stateless DHCPv6, and stateful DHCPv6, shows the M and O flags that arbitrate between them, and builds a working stateful server with a relay on Cisco IOS XE, with a real Linux client acquiring a real lease. It extends the [IPv6 cluster guide](https://www.pinglabz.com/ipv6/) and the [IP Services cluster guide](https://www.pinglabz.com/ip-services/). ## Start Here: IPv6 Hosts Do Not Need DHCP In IPv4, a host with no DHCP server has no address. Full stop. That assumption is baked so deeply into most engineers' instincts that IPv6's model feels wrong at first. In IPv6, the **router** is the primary source of addressing information, not a server. A router periodically multicasts a **Router Advertisement (RA)** containing the on-link prefix, and any host that hears it can build itself a working global address without asking anyone. That is SLAAC (Stateless Address Autoconfiguration), and it is the default behaviour of essentially every IPv6 host on earth. You can watch it happen. Here is a Linux host in the lab that has been given no DHCP configuration whatsoever, moments after being connected: ``` $ ip -br addr show eth0 eth0@if49 UP 10.0.10.100/24 2001:db8:10:0:5054:ff:fe79:35a2/64 fe80::5054:ff:fe79:35a2/64 ``` The IPv4 address was configured manually. The IPv6 global address `2001:db8:10:0:5054:ff:fe79:35a2/64` appeared on its own, because the router advertised the prefix `2001:db8:10::/64` and the host did the rest. Look at the interface identifier: `5054:ff:fe79:35a2`. That `ff:fe` in the middle is the signature of **EUI-64**: the host took its 48-bit MAC address (52:54:00:79:35:a2), split it in half, jammed `ff:fe` into the gap, and flipped the seventh bit. If you see `ff:fe` in an IPv6 address, you are looking at an EUI-64 SLAAC address and you can read the MAC straight out of it. (Which is exactly why RFC 4941 privacy addresses exist, and why most desktop OSes now use random identifiers instead by default.) ## The Two Flags That Decide Everything The RA carries two bits that tell the host what else to do. Getting these right is the whole game. M flag (Managed)"Get your **address** from DHCPv6." Set it, and the host asks a DHCPv6 server for an address instead of building one. O flag (Other)"Get your **other config** (DNS, domain) from DHCPv6." The address still comes from SLAAC. Those two bits produce the three deployment models, and the names finally make sense: Pure SLAAC M / O0 / 0 Address from the RA prefix. DNS from the RA (RDNSS option), if the router sends it. No DHCPv6 at all. Stateless DHCPv6 M / O0 / 1 Address from SLAAC. DNS and domain from a DHCPv6 server. The server keeps **no lease state**, hence "stateless". Stateful DHCPv6 M / O1 / 1 Address AND config from the server, which tracks every lease. The closest thing to IPv4 DHCP. "Stateless" and "stateful" refer to **the server's** state, not the client's. A stateless DHCPv6 server hands out DNS servers and forgets you exist. A stateful one records who has which address, which is what you need if you want to answer "which machine was 2001:db8::42 last Tuesday." One critical subtlety: **setting the M flag does not automatically stop SLAAC.** If the RA still carries a prefix with the autonomous flag set, a host will build a SLAAC address *and* ask DHCPv6 for one, and end up with both. If you want DHCPv6 to be the only source, you must explicitly tell the router to stop advertising the prefix as autoconfigurable. ## Building Stateful DHCPv6 with a Relay The lab: a client on the 2001:db8:10::/64 LAN behind R1, and the DHCPv6 server on R2, one hop away. This is the realistic topology, because your DHCP server is essentially never on the same segment as your clients. ### The server (R2) ``` R2(config)# ipv6 dhcp pool PINGLABZ-STATEFUL R2(config-dhcpv6)# address prefix 2001:DB8:10:AAAA::/64 lifetime 3600 1800 R2(config-dhcpv6)# dns-server 2001:4860:4860::8888 R2(config-dhcpv6)# domain-name pinglabz.local ! R2(config)# interface Ethernet0/1 R2(config-if)# ipv6 dhcp server PINGLABZ-STATEFUL ``` The two lifetimes are **valid** (3600) and **preferred** (1800), in that order. An address past its preferred lifetime is "deprecated": still usable for existing connections, but not chosen for new ones. Past the valid lifetime it is gone. This graceful two-stage expiry has no IPv4 equivalent and is one of the genuinely nice parts of IPv6. ### The relay and the flags (R1) ``` R1(config)# interface Ethernet0/2 R1(config-if)# ipv6 nd managed-config-flag R1(config-if)# ipv6 nd other-config-flag R1(config-if)# ipv6 nd prefix 2001:DB8:10::/64 no-autoconfig R1(config-if)# ipv6 dhcp relay destination 2001:DB8:12::2 Ethernet0/1 ``` Four lines, four distinct jobs: - `managed-config-flag` sets M. Hosts will now solicit an address from DHCPv6. - `other-config-flag` sets O. Hosts will take DNS and domain from DHCPv6 too. - `prefix ... no-autoconfig` clears the autonomous bit on the advertised prefix. **This is the line everyone forgets.** Without it, hosts keep building SLAAC addresses alongside their DHCPv6 lease, and you have two addresses per host and no idea which one traffic will use. - `ipv6 dhcp relay destination` forwards the client's solicit (sent to the link-local multicast group ff02::1:2, which does not route) onward to the actual server as unicast. ## The Exchange DHCPv6 uses four messages, and if you know DHCPv4's DORA they map cleanly: SOLICIT (client)"Are there any DHCPv6 servers?" Sent to ff02::1:2\. The DISCOVER equivalent. ADVERTISE (server)"I am here and I can offer you this." The OFFER equivalent. REQUEST (client)"I will take it." The REQUEST equivalent. REPLY (server)"It is yours, here are the lifetimes." The ACK equivalent. The client in the lab is a Linux host running `udhcpc6`. Here it is doing exactly that: ``` $ udhcpc6 -i eth0 udhcpc6: started, v1.36.1 udhcpc6: sending discover udhcpc6: sending discover udhcpc6: sending select udhcpc6: IPv6 obtained, lease time 3600 ``` Solicit, solicit (the first went unanswered while the relay warmed up), select, and a lease with the 3600-second lifetime configured on the server. The client is on one subnet, the server is on another, and the relay carried the conversation between them. ## The Binding: What the Server Remembers This is where "stateful" earns its name: ``` R2#show ipv6 dhcp binding Client: FE80::5054:FF:FE79:35A2 DUID: 000300015254007935A2 Username : unassigned VRF : default IA NA: IA ID 0x8FFD7C11, T1 900, T2 1440 Address: 2001:DB8:10:AAAA:31E0:5957:EFD4:7F52 preferred lifetime 1800, valid lifetime 3600 expires at Jul 12 2026 01:54 AM (3587 seconds) ``` Several things worth reading properly: - **DUID** (DHCP Unique Identifier) is how DHCPv6 identifies a client, and it is *not* the MAC address. It is generated once by the client and persists across reboots and NIC changes. This breaks every IPv4 habit you have around MAC-based reservations. Here the DUID happens to embed the MAC (`5254007935A2`) because it is DUID type 3 (link-layer), but that is a client choice, not a guarantee. - **IA\_NA** (Identity Association for Non-temporary Addresses) is the container for the leased address. A client can hold several. - **T1 (900) and T2 (1440)** are the renew and rebind timers, defaulting to 50% and 80% of the preferred lifetime. At T1 the client asks the same server to renew. At T2 it gives up on that server and asks any server. - The **allocated address** is `2001:DB8:10:AAAA:31E0:5957:EFD4:7F52`. Note the host portion is random, not EUI-64\. There is no `ff:fe`. That is the fingerprint of a DHCPv6-assigned address versus a SLAAC one, and it is the quickest way to tell at a glance which mechanism gave a host its address. ``` R2#show ipv6 dhcp pool DHCPv6 pool: PINGLABZ-STATEFUL Address allocation prefix: 2001:DB8:10:AAAA::/64 valid 3600 preferred 1800 (1 in use, 0 conflicts) DNS server: 2001:4860:4860::8888 Domain name: pinglabz.local Active clients: 1 ``` ## DHCPv6-PD, Briefly One more mode exists and it has no IPv4 analogue at all. **Prefix Delegation** hands a client an entire prefix (a /56 or /48) rather than a single address, so the client can subnet it and hand addresses to devices behind it. This is how your home router gets IPv6 from your ISP: it does not receive one address, it receives a /56, and it then advertises /64s from that block onto your home LANs. Configured with `prefix-delegation pool` on the server side and `ipv6 dhcp client pd` on the client interface. It is the standard model for any downstream router, not just consumer kit. ## When It Does Not Work Host has a SLAAC address AND a DHCPv6 oneYou set M but left the prefix autonomous. Add `ipv6 nd prefix ... no-autoconfig`. Host gets no address at allIs `ipv6 unicast-routing` on? Without it the router sends no RAs and the host has nothing to react to. Solicits never reach the serverMissing relay. ff02::1:2 is link-local and does not route. `show ipv6 dhcp interface` on both ends. Reservation by MAC does not workDHCPv6 keys on DUID, not MAC. Reserve by DUID. And a security note: an RA is unauthenticated, and any host on the segment can send one. A rogue RA (malicious or, far more commonly, a misconfigured Windows box with sharing enabled) will reconfigure every host on the VLAN. The defence is RA Guard on the access switch, which belongs to the IPv6 first-hop security toolkit. The same dependency produces a more common failure than the rogue one: no usable RA at all, and every host on the segment stuck with nothing but a link-local address, with no server, scope or lease to interrogate. [Diagnosing a host that never gets past FE80::](https://www.pinglabz.com/ipv6-neighbor-discovery-troubleshooting/) starts with the single command that answers it. ## FAQ ### Which should I use? Stateful DHCPv6 if you need to know which device had which address (most enterprises, for audit and forensics). SLAAC plus stateless DHCPv6 if you just need working hosts and DNS with minimal infrastructure. Pure SLAAC for simple or transit segments. ### Can a host get DNS without any DHCPv6 at all? Yes. The RDNSS option in the RA (RFC 8106) carries DNS servers directly. Modern OSes support it; some older ones ignore it, which is why stateless DHCPv6 still exists. ### Why is my client's address not EUI-64 even with SLAAC? Privacy extensions (RFC 4941). Most desktop OSes generate random, rotating interface identifiers by default rather than embedding the MAC. This is a feature. ### Does DHCPv6 give out a default gateway? No, and this surprises people. There is no gateway option in DHCPv6\. The default route **always** comes from the Router Advertisement. You cannot run DHCPv6 without RAs. ### What is T1 and T2? Renew (T1) and rebind (T2) timers, defaulting to 50% and 80% of the preferred lifetime. At T1 the client renews with the same server; at T2 it broadcasts for any server. ## Key Takeaways - IPv6 hosts get addresses from **Router Advertisements** by default. DHCPv6 is optional, and layered on top. - The **M flag** means "get your address from DHCPv6." The **O flag** means "get your other config from DHCPv6." Those two bits define all three models. - Setting M does not stop SLAAC. You must also set `ipv6 nd prefix ... no-autoconfig`, or hosts end up with two addresses. - An `ff:fe` in the middle of the host portion means EUI-64 SLAAC. A random host portion means DHCPv6 assigned it. You can diagnose the mechanism from the address alone. - DHCPv6 identifies clients by **DUID**, not MAC. Every MAC-based reservation habit you have from IPv4 is wrong here. - **DHCPv6 never provides a default gateway.** The default route always comes from the RA. - Solicits go to ff02::1:2, which is link-local. Off-segment servers require a relay. Next: the [IPv6 cluster guide](https://www.pinglabz.com/ipv6/), or the [IP Services cluster guide](https://www.pinglabz.com/ip-services/). ### SPAN, RSPAN, and ERSPAN: Packet Mirroring on Cisco Switches URL: https://www.pinglabz.com/span-rspan-erspan-configuration/ Last updated: 2026-07-12T01:23:28.000Z Eventually every network problem reaches the point where you have to see the actual packets. Counters lie by omission, logs tell you what the device thought, and [NetFlow](https://www.pinglabz.com/flexible-netflow-cisco-ios-xe/) tells you a conversation happened without telling you what was in it. Port mirroring is how you get the packets themselves onto a machine running Wireshark. SPAN, RSPAN, and ERSPAN are three answers to the same question, separated by how far the analyzer sits from the traffic. This article configures all three on real Cisco gear and captures the ERSPAN packets arriving at a real Linux host. It extends the [IP Services cluster guide](https://www.pinglabz.com/ip-services/). ## The Three, in One Sentence Each SPAN (local) Copy traffic to another port **on the same switch**. Analyzer must be plugged into that switch. TransportNone. Raw copy. RSPAN Copy traffic into a **special VLAN** that is trunked to another switch, where the analyzer sits. TransportLayer 2 (a VLAN) ERSPAN Wrap the copies in **GRE** and route them anywhere in the IP network. Analyzer can be in another building. TransportLayer 3 (GRE/IP) The progression is purely about distance. Local SPAN needs the analyzer on the same box. RSPAN gets it onto a different switch in the same layer 2 domain. ERSPAN gets it anywhere you can route a packet, which in practice means a central capture server for the whole estate. ## Local SPAN Two lines. A source (what to copy) and a destination (where to put the copies): ``` SW1(config)# monitor session 1 source interface Ethernet0/1 both SW1(config)# monitor session 1 destination interface Ethernet0/2 ``` ``` SW1#show monitor session 1 Session 1 --------- Type : Local Session Source Ports : Both : Et0/1 Destination Ports : Et0/2 Encapsulation : Native ``` The `both` keyword captures ingress and egress. You can also say `rx` or `tx`, and you should when you can, because capturing both directions on a busy port means the destination port has to carry the sum of both, and it will drop what it cannot fit. ### The three rules that bite people - **The destination port stops being a normal port.** It leaves the switching fabric: it will not forward, will not learn MACs, and drops anything the analyzer transmits into it. If you were using it for something, you are not anymore. - **Oversubscription is silent.** Mirroring two gigabit ports (both directions) into one gigabit destination means up to 4 Gbps into a 1 Gbps hole. The switch drops the excess without telling you, and your capture has holes in exactly the busy moments you care about. - **Session limits are low.** Most Catalyst platforms allow only two local SPAN sessions. This is a hardware constraint, not a licensing one. You can also mirror an entire VLAN (`source vlan 10`) rather than a port, which is often what you actually want, and mirror a port-channel by naming the port-channel interface. ## RSPAN: Across the Layer 2 Domain RSPAN carries the mirrored traffic in a dedicated VLAN. That VLAN must be flagged as an RSPAN VLAN, which stops the switch from doing normal things to it (learning, flooding to unrelated ports): ``` SW1(config)# vlan 999 SW1(config-vlan)# name RSPAN-VLAN SW1(config-vlan)# remote-span ! SW1(config)# monitor session 3 source interface Ethernet0/1 rx SW1(config)# monitor session 3 destination remote vlan 999 ``` ``` SW1#show monitor session 3 Session 3 --------- Type : Remote Source Session Source Ports : RX Only : Et0/1 Dest RSPAN VLAN : 999 SW1#show vlan remote-span Remote SPAN VLANs ------------------------------------------------------------------------------ 999 ``` On the far switch, the mirror image: a session whose *source* is the RSPAN VLAN and whose destination is the analyzer's port. ``` SW2(config)# monitor session 3 source remote vlan 999 SW2(config)# monitor session 3 destination interface Ethernet0/5 ``` The requirements are strict and unforgiving: - The RSPAN VLAN must exist on **every switch in the path**, and be marked `remote-span` on each. - It must be allowed on every trunk between them. VTP pruning will happily prune it and break your capture. - It carries a copy of everything you mirrored, so it eats real bandwidth on those trunks. Mirroring a busy port across a congested trunk is a way to cause an outage while investigating one. RSPAN is the awkward middle child: more reach than SPAN, far more fragile than ERSPAN, and confined to one layer 2 domain either way. If you have a routed network, skip to ERSPAN. ## ERSPAN: Mirror Anywhere ERSPAN wraps each mirrored frame in GRE and sends it as a routed IP packet. The analyzer can be anywhere you can reach. **A platform note first, honestly.** The IOL-XE and IOL-L2 images used elsewhere in these labs do *not* support ERSPAN. Both reject the command outright: ``` R1(config)# monitor session 10 type erspan-source ^ % Invalid input detected at '^' marker. ``` So the ERSPAN work below runs on a **Catalyst 8000v (IOS XE 17.18.02)**, which does support it. If you are labbing this and the command is rejected, it is your platform, not your syntax. The source session: ``` R3(config)# monitor session 1 type erspan-source R3(config-mon-erspan-src)# source interface GigabitEthernet1 both R3(config-mon-erspan-src)# no shutdown R3(config-mon-erspan-src)# destination R3(config-mon-erspan-src-dst)# erspan-id 100 R3(config-mon-erspan-src-dst)# ip address 192.168.99.100 R3(config-mon-erspan-src-dst)# origin ip address 192.168.99.3 ``` Note `no shutdown`. ERSPAN sessions are created administratively down. Forgetting this line produces a session that looks perfectly configured and sends nothing, which is a genuinely infuriating twenty minutes. ``` R3#show monitor session 1 Session 1 --------- Type : ERSPAN Source Session Status : Admin Enabled Source Ports : Both : Gi1 Destination IP Address : 192.168.99.100 MTU : 1464 Destination ERSPAN ID : 100 Origin IP Address : 192.168.99.3 ``` Three fields matter: erspan-id 100Identifies this stream. Lets one collector receive mirrors from many sources and tell them apart. Must match on the destination session. ip addressWhere the GRE packets go. Must be routable from the source. MTU 1464The GRE + ERSPAN headers eat the difference. Mirrored frames larger than this get truncated or dropped. That MTU line is the single biggest operational gotcha in ERSPAN. A full-size 1500-byte frame plus GRE and ERSPAN headers exceeds 1500 on the transit path, so unless the path supports jumbo frames, your captures of large packets will be truncated. If you are hunting an MTU problem *with* ERSPAN, be very careful about what you conclude. The platform confirms the encapsulation it will use: ``` R3#show platform hardware qfp active feature erspan state ERSPAN State: Status : Active Capabilites: Max sessions : 1032 Encaps type : ERSPAN type-II / ERSPAN type-III GRE protocol : 0x88BE / 0x22EB MTU : 1464 / 1452 ``` GRE protocol type **0x88BE** is ERSPAN Type II, **0x22EB** is Type III. Type III adds finer-grained timestamps; Type II is the default and what nearly everything uses. ### Proof: the mirror arriving at a real host Configuration is a claim. Here is the evidence. The analyzer is a Debian box at 192.168.99.100, and while it is being pinged (so there is traffic to mirror), it runs `tcpdump` filtering for GRE: ``` j@llmbits:~$ sudo tcpdump -i ens224 -n 'proto gre' -v listening on ens224, link-type EN10MB (Ethernet), snapshot length 262144 bytes 18:08:16.890159 IP (tos 0x0, ttl 255, id 0, offset 0, flags [DF], proto GRE (47), length 134) 192.168.99.3 > 192.168.99.100: GREv0, Flags [sequence# present], seq 7, length 114 gre-proto-0x88be 18:08:16.890372 IP (tos 0x0, ttl 255, id 0, offset 0, flags [DF], proto GRE (47), length 134) 192.168.99.3 > 192.168.99.100: GREv0, Flags [sequence# present], seq 8, length 114 gre-proto-0x88be 18:08:16.890900 IP (tos 0x0, ttl 255, id 0, offset 0, flags [DF], proto GRE (47), length 212) 192.168.99.3 > 192.168.99.100: GREv0, Flags [sequence# present], seq 9, length 192 gre-proto-0x88be 4 packets captured ``` That is ERSPAN working, end to end, on a real wire. The router at 192.168.99.3 is encapsulating every mirrored frame in GRE with protocol type `0x88be` and shipping it to the analyzer. The sequence numbers let the receiver detect dropped mirror packets, which is genuinely useful: if your capture has gaps, the sequence gaps tell you the mirror dropped them rather than the network. Wireshark decodes ERSPAN natively. Point it at this interface, and it will strip the GRE and ERSPAN headers and show you the original mirrored frames as if you had captured them on the source port. ## Choosing Between Them Laptop is in the rack**Local SPAN.** Two commands, zero complications. Analyzer on another switch, flat L2**RSPAN.** Watch your trunks and VTP pruning. Central capture server / IDS farm**ERSPAN.** The only one that crosses a routed boundary. Mind the MTU. You need volume, not packetsNot SPAN at all. Use [NetFlow](https://www.pinglabz.com/flexible-netflow-cisco-ios-xe/). Mirroring a 10G link to answer "who is using bandwidth" is absurd. ## FAQ ### Why is my SPAN destination port not passing normal traffic? Because that is the design. A SPAN destination leaves the switching fabric entirely. It only spits out copies. This is not a fault. ### My capture has gaps under load. Why? Oversubscription. You are mirroring more traffic than the destination port can carry, and the switch drops the excess silently. Mirror one direction, filter to a VLAN, or use a faster destination port. ### Can I mirror to a port on a different switch without RSPAN? No, not with local SPAN. That is precisely what RSPAN and ERSPAN exist to solve. ### Does ERSPAN encrypt the mirrored traffic? No. GRE is encapsulation, not encryption. You are shipping copies of production traffic, in clear text, across your network. Treat the ERSPAN path as sensitive and consider whether it should ride a dedicated VRF or an IPsec tunnel. See the [GRE cluster guide](https://www.pinglabz.com/gre/). ### How many sessions can I run? Platform-dependent and usually small for local SPAN (often two). ERSPAN on a Catalyst 8000v advertises up to 1032 sessions, as the platform output above shows. Check your specific hardware. ## Key Takeaways - SPAN, RSPAN, ERSPAN differ only in **how far the analyzer can be**: same switch, same layer 2 domain, anywhere routable. - A SPAN destination port is removed from the switching fabric. It forwards nothing and learns nothing. - Oversubscribing a destination port drops mirrored packets **silently**. Your capture will have holes exactly when the network is busy. - RSPAN needs its VLAN marked `remote-span` and allowed on every trunk in the path. VTP pruning will break it. - ERSPAN sessions are created shut. `no shutdown` or nothing leaves the box. - ERSPAN's GRE overhead cuts the usable MTU to 1464\. Large mirrored frames get truncated. - Neither IOL-XE nor IOL-L2 supports ERSPAN. Use a Catalyst 8000v, a Catalyst 9000, or real hardware. - ERSPAN is not encrypted. You are routing copies of production traffic in the clear. Next: [conditional debugging](https://www.pinglabz.com/conditional-debug-cisco-ios-xe/) for when you need the router's view rather than the packets, or the [IP Services cluster guide](https://www.pinglabz.com/ip-services/). ### Object Tracking and Reliable Static Routes with IP SLA URL: https://www.pinglabz.com/object-tracking-reliable-static-routes/ Last updated: 2026-07-12T01:23:28.000Z A static route has one fatal flaw: it believes in itself. As long as the outgoing interface is up, the route stays in the table, even if the next hop died, the far end of the circuit is a black hole, or the provider's core is on fire. The interface is up, so the route is valid. That is the whole of a static route's world view. Object tracking fixes this. You attach a condition to the route, and the route only exists while the condition is true. Combine it with an [IP SLA probe](https://www.pinglabz.com/ip-sla-cisco-ios-xe/) and you get a static route that tests the path before trusting it. This article builds one and then kills the path to prove it works. It extends the [IP Services cluster guide](https://www.pinglabz.com/ip-services/). ## The Problem, Concretely Classic dual-WAN setup. A primary link and a backup link, both static routes, the backup with a higher administrative distance so it stays out of the table until the primary disappears: ``` ip route 10.0.20.0 255.255.255.0 10.0.12.2 ip route 10.0.20.0 255.255.255.0 10.0.13.2 200 ``` The second route has AD 200 against the default of 1, which makes it a **floating static**: it floats above the routing table, invisible, until the better route is withdrawn. See the [IP Routing cluster guide](https://www.pinglabz.com/ip-routing/) for how administrative distance arbitrates between sources. This works perfectly for exactly one failure mode: the local interface going down. The router notices, withdraws the connected route, the primary static loses its next hop, and the floating static installs. It fails completely for every other failure mode. If the next hop is reachable at layer 2 but the path beyond it is broken (a common outcome when the far end is a switch, a media converter, or a provider handoff), your interface stays up, the primary static stays in the table, and you cheerfully forward traffic into a black hole forever. Nothing fails over, because from the router's point of view, nothing failed. ## The Fix: Track a Probe, Not an Interface Three pieces, in order. **1\. An IP SLA probe that tests the path.** Not the interface, the path: ``` R1(config)# ip sla 1 R1(config-ip-sla-echo)# icmp-echo 10.0.12.2 source-interface Ethernet0/1 R1(config-ip-sla-echo)# frequency 5 R1(config-ip-sla-echo)# timeout 500 R1(config)# ip sla schedule 1 life forever start-time now ``` **2\. A tracked object that watches the probe:** ``` R1(config)# track 1 ip sla 1 reachability ``` **3\. A static route that depends on the tracked object:** ``` R1(config)# ip route 10.0.20.0 255.255.255.0 10.0.12.2 track 1 R1(config)# ip route 10.0.20.0 255.255.255.0 10.0.13.2 200 ``` That is the whole mechanism. The primary route now exists only while track 1 is up, and track 1 is up only while the probe is getting replies. ## Verifying the Healthy State ``` R1#show track 1 Track 1 IP SLA 1 reachability Reachability is Up 1 change, last change 00:00:12 Latest operation return code: OK Latest RTT (millisecs) 3 Tracked by: Static IP Routing 0 ``` Read the last two lines carefully. `Tracked by: Static IP Routing` confirms that something is actually consuming this tracked object. If that section is empty, your track is running and nothing is listening to it, which is a surprisingly common way to spend an hour. And the routing table shows the primary, at its normal AD of 1: ``` R1#show ip route 10.0.20.0 Routing entry for 10.0.20.0/24 Known via "static", distance 1, metric 0 Routing Descriptor Blocks: * 10.0.12.2 Route metric is 0, traffic share count is 1 ``` ## Break the Path Now kill the primary. In the lab this is a shutdown, but the point is that the same thing happens for any failure the probe can detect, including ones that leave the interface up: ``` R1(config)# interface Ethernet0/1 R1(config-if)# shutdown ``` Within a few probe intervals: ``` R1#show track 1 Track 1 IP SLA 1 reachability Reachability is Down 16 changes, last change 00:00:09 Latest operation return code: Timeout Tracked by: Static IP Routing 0 ``` **Reachability is Down.** Return code **Timeout**, which is the probe failing rather than merely being slow. (This is why the timeout/threshold distinction from the IP SLA article matters: an "over threshold" result would leave this object UP.) The routing table responds: ``` R1#show ip route 10.0.20.0 Routing entry for 10.0.20.0/24 Known via "static", distance 200, metric 0 Routing Descriptor Blocks: * 10.0.13.2 Route metric is 0, traffic share count is 1 ``` **Distance 200, via 10.0.13.2.** The primary is gone from the table entirely and the floating static has surfaced. And the traffic never noticed: ``` R1#ping 10.0.20.1 source Ethernet0/0 Type escape sequence to abort. Sending 5, 100-byte ICMP Echos to 10.0.20.1, timeout is 2 seconds: Packet sent with a source address of 192.168.99.1 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 1/1/2 ms ``` One hundred percent success across a failover the router performed by itself, based on a probe that tested the path rather than the link. ## The Trap That Catches Everyone Look closely at what the probe targets. It pings **10.0.12.2**, the next hop. It does not ping 10.0.20.1, the destination. This is deliberate and it matters. If the probe targeted 10.0.20.1, you would create a circular dependency: the probe needs a route to 10.0.20.1 to send the packet, but the route to 10.0.20.1 is the thing the probe is deciding whether to install. When the primary path dies, the route to the probe target dies with it, the probe starts using the backup path, succeeds, and re-installs the primary route. The router then oscillates. Two safe patterns: Probe the next hopSimple and safe. Tests the link and the far router's control plane. Does not test anything beyond it. Probe a far target, pinned with a host routeAdd `ip route 8.8.8.8 255.255.255.255 10.0.12.2` so the probe is *forced* down the primary path regardless of what the tracked route does. Tests the whole path end to end. The second pattern is what you want in production dual-ISP designs, because probing your own next hop tells you nothing about whether the ISP's upstream is working. Pin a host route to a well-known public address out of the primary interface, probe that, and track it. The host route never moves, so the probe always tests the path you intend. ## What Else You Can Track IP SLA reachability is the common case, but the tracking subsystem is more general: track N ip sla M reachability Up while the probe succeeds. The workhorse. track N interface X line-protocol Up while the interface is up. Simple, but back to trusting the link. track N ip route X.X.X.X/Y reachability Up while that prefix is in the RIB. Useful for tracking a BGP-learned route. track N list boolean and / or Combine several trackers. "Fail over only if BOTH probes fail" avoids flapping on a single lost packet. The boolean list is underused and worth knowing. A single dropped probe on a lossy link should not move your default route. Tracking two probes to two different targets with `track 10 list boolean or` means the route only fails when both are unreachable, which is a much better proxy for "the path is genuinely dead." ## Damping the Flap A path that fails and recovers repeatedly will move your routing table with it, and route churn is its own outage. Two knobs: ``` R1(config)# track 1 ip sla 1 reachability R1(config-track)# delay down 10 up 30 ``` `delay down 10` means the object must be failing for 10 seconds before it declares down. `delay up 30` means it must be healthy for 30 seconds before it declares up again. Asymmetric on purpose: fail fast (you want off a broken path quickly), recover slowly (you do not want to move back onto a path that is still flapping). In the lab output above, note `16 changes` on the tracked object. That is a tracker with no damping, faithfully reporting every transition during testing. In production that number is a red flag. ## Tracking Is Not Just for Static Routes The same tracked object can drive other things, which is where it gets genuinely powerful: - **HSRP/VRRP/GLBP priority.** `standby 1 track 1 decrement 30` lowers the priority when the uplink probe fails, so the standby router takes over as active gateway. This is the canonical fix for the "active HSRP router lost its uplink but is still the gateway" black hole. See the [FHRP cluster guide](https://www.pinglabz.com/fhrp/). - **EEM applets.** An applet can fire on a track state change and log, alert, or reconfigure. - **PBR next-hop.** A route-map `set ip next-hop verify-availability` can be tracked, so policy routing stops steering traffic into a dead next hop. ## FAQ ### Why does my tracked route never fail over? Check three things in order: is the SLA scheduled (`show ip sla summary`), is the track actually reflecting it (`show track`), and does `Tracked by:` list your static route. Any one of those missing breaks the chain silently. ### Why does my router oscillate between the two routes? Your probe target is reachable via the backup path. Pin the probe target with a host route out of the primary interface, or probe the next hop. ### Can I track a probe on the backup link too? Yes, and you should if the backup is a real circuit rather than a local interface. Otherwise you can fail over onto a backup that is also dead. ### What AD should the floating static use? Anything higher than the primary and higher than any dynamic protocol you run for the same prefix. 200 is conventional because it is above every default AD except an unreachable (255). ### Does this work with IPv6? Yes. `ipv6 route ... track N` and IP SLA operations support IPv6 targets. The mechanism is identical. ## Key Takeaways - A plain static route trusts the interface, not the path. It cannot detect a next hop that is up but useless. - The chain is: **IP SLA probe** to **tracked object** to **static route**. Break any link and nothing fails over. - `show track` must show `Tracked by: Static IP Routing`. If it does not, nothing is consuming your tracker. - Never probe the destination you are routing to. Probe the next hop, or pin the probe target with a dedicated host route out of the primary interface. - The failover is visible and clean: distance 1 via primary becomes distance 200 via backup, and traffic keeps flowing at 100 percent. - Use `delay down / up` to damp flapping. Fail fast, recover slowly. - Tracked objects also drive HSRP priority, EEM applets, and PBR next-hop verification. The static route is just the simplest consumer. Next: [the IP SLA probes themselves in depth](https://www.pinglabz.com/ip-sla-cisco-ios-xe/), or the [IP Services cluster guide](https://www.pinglabz.com/ip-services/). ### IP SLA on Cisco IOS XE: Probes, Thresholds, and Tracked Failover URL: https://www.pinglabz.com/ip-sla-cisco-ios-xe/ Last updated: 2026-08-01T19:26:55.000Z Your monitoring system tells you the link is up. The user tells you the call sounds terrible. Both are correct and neither is useful, because "up" and "usable" are different questions. IP SLA is how a Cisco router answers the second one: it generates synthetic traffic, measures what happens to it, and gives you numbers you can act on. This article builds ICMP and UDP jitter probes on IOS XE, then wrecks the path and shows the same probes reporting the damage. It extends the [IP Services cluster guide](https://www.pinglabz.com/ip-services/). The second half then wires a probe to a tracked object and a static route, breaks the path two hops away, and captures the failover: the return code flipping to `Timeout`, `%TRACK-6-STATE` logging the object down, the route leaving the RIB. Every capture below is verbatim from CML labs on iol-xe running **IOS XE 17.18.2**. ## What IP SLA Actually Does An IP SLA operation is a scheduled, synthetic test. The router sends traffic (an ICMP echo, a UDP stream, an HTTP GET, a DNS query) at a target, measures the result, stores statistics. Nothing about it is passive. Two roles exist: Source (the router with `ip sla N`)Generates probes, measures results, holds the statistics. Responder (the router with `ip sla responder`)Timestamps the probe in and out, so the source can separate network delay from the far device's processing delay. The responder is what makes jitter and one-way latency meaningful: without it you cannot tell whether 40ms of the round trip was the network or the far device answering slowly. ICMP echo needs no responder; jitter does. ## The Simple One: ICMP Echo A scheduled ping with memory. On the source router: ``` R1(config)# ip sla 1 R1(config-ip-sla-echo)# icmp-echo 10.0.12.2 source-interface Ethernet0/1 R1(config-ip-sla-echo)# frequency 5 R1(config-ip-sla-echo)# threshold 200 R1(config-ip-sla-echo)# timeout 500 R1(config)# ip sla schedule 1 life forever start-time now ``` Three timers, three jobs, and mixing them up is the classic error: frequency 5 How often the operation runs, in seconds. timeout 500 Milliseconds to wait for a reply before the probe is FAILED. This drives tracking. threshold 200 Milliseconds above which the probe is "over threshold" but still SUCCESSFUL. The `timeout` versus `threshold` distinction matters enormously. A probe over threshold still counts as a success and will **not** bring down a tracked object. Only a timeout does. If a slow path should trigger failover, set the timeout to your tolerance. Also note `ip sla schedule`. Forget it and the operation sits there doing nothing, forever, with no warning. It is the most common "why is my SLA not working" cause. Healthy baseline on a clean link: ``` R1#show ip sla statistics 1 IPSLA operation id: 1 Latest RTT: 3 milliseconds Latest operation return code: OK Number of successes: 3 Number of failures: 0 ``` ## The Useful One: UDP Jitter ICMP echo tells you reachability and round-trip time. It tells you nothing about whether a voice call will sound acceptable, because voice cares far less about average latency than about **jitter** (variation in delay) and **loss**. The udp-jitter operation sends packets at a fixed interval, measures how they arrive, and needs a responder on the far end. ``` R2(config)# ip sla responder ``` ``` R1(config)# ip sla 2 R1(config-ip-sla-jitter)# udp-jitter 10.0.12.2 5000 num-packets 20 interval 20 R1(config-ip-sla-jitter)# frequency 15 R1(config-ip-sla-jitter)# tos 184 R1(config)# ip sla schedule 2 life forever start-time now ``` Reading that: 20 packets, 20 milliseconds apart, to UDP port 5000, repeated every 15 seconds. `tos 184` is DSCP EF, how real voice is marked. Marking the probe like the traffic you care about is the entire point, because the network treats the two differently. See the [QoS cluster guide](https://www.pinglabz.com/qos/) for why. On a clean link, the output is boring, which is the correct result: ``` R1#show ip sla statistics 2 IPSLA operation id: 2 Type of operation: udp-jitter Latest operation return code: OK RTT Values: Number Of RTT: 20 RTT Min/Avg/Max: 1/2/4 milliseconds Latency one-way time: Source to Destination Latency one way Min/Avg/Max: 0/1/2 milliseconds Destination to Source Latency one way Min/Avg/Max: 1/1/3 milliseconds Jitter Time: Source to Destination Jitter Min/Avg/Max: 0/1/1 milliseconds Destination to Source Jitter Min/Avg/Max: 0/1/2 milliseconds Packet Loss Values: Loss Source to Destination: 0 Loss Destination to Source: 0 Out Of Sequence: 0 Tail Drop: 0 ``` Note what the responder buys you: **separate source-to-destination and destination-to-source** figures. A one-way problem (congestion in one direction only, common on asymmetric WAN links) is invisible to a ping and obvious here. ## Now Break the Path Clean numbers prove the probe works, not that it is useful. So the lab applies link conditioning to the R1-R2 link: 120ms latency, 40ms jitter, 8 percent loss. Nothing else changes. Interfaces stay up, OSPF stays adjacent, and by every traditional measure the link is "fine." The ICMP probe notices immediately: ``` R1#show ip sla statistics 1 IPSLA operation id: 1 Latest RTT: 238 milliseconds Latest operation return code: Over threshold Number of successes: 7 Number of failures: 0 ``` **Over threshold**, at 238ms against our 200ms threshold, and crucially `Number of failures: 0`. The path is terrible and the probe is succeeding. Exactly the state where a naive "is it up" monitor reports green. The jitter probe tells the real story: ``` R1#show ip sla statistics 2 IPSLA operation id: 2 Latest RTT: 250 milliseconds RTT Values: Number Of RTT: 17 RTT Min/Avg/Max: 214/250/305 milliseconds Latency one-way time: Source to Destination Latency one way Min/Avg/Max: 82/119/155 milliseconds Destination to Source Latency one way Min/Avg/Max: 91/123/155 milliseconds Jitter Time: Source to Destination Jitter Min/Avg/Max: 3/21/59 milliseconds Destination to Source Jitter Min/Avg/Max: 0/17/34 milliseconds Packet Loss Values: Loss Source to Destination: 1 Source to Destination Loss Periods Number: 3 Loss Destination to Source: 2 Destination to Source Loss Periods Number: 1 Out Of Sequence: 6 Tail Drop: 0 ``` Compare against the baseline and the diagnosis writes itself: RTT Min/Avg/Max Before1 / 2 / 4 ms After214 / 250 / 305 ms SD Jitter Min/Avg/Max Before0 / 1 / 1 ms After3 / 21 / 59 ms Packet loss Before0 both directions After1 SD (3 periods), 2 DS Out of sequence Before / after0 then 6 Jitter went from 1ms to an average of 21ms with peaks at 59ms, against a typical voice jitter buffer of around 30ms. Packets 59ms late are discarded by the phone whether or not the network delivered them. That is how a call sounds broken while every link light is green. **Loss periods** matter too. "Loss Source to Destination: 1, Loss Periods Number: 3" means three separate events rather than one burst, and scattered single-packet loss degrades voice far more gracefully because the codec conceals isolated gaps. A percentage-loss figure hides that. The compact view: ``` R1#show ip sla summary ID Type Destination Stats Return Code Last Run *1 icmp-echo 10.0.12.2 RTT=227 Over threshold 1 seconds ago *2 udp-jitter 10.0.12.2 RTT=250 OK 15 seconds ago ``` ## A Word on MOS and ICPIF You will notice these in the jitter output: ``` Voice Score Values: Calculated Planning Impairment Factor (ICPIF): 0 Mean Opinion Score (MOS): 0 ``` Both are zero, including under the degraded conditions. That is not a bug, and it is not the network scoring perfectly. MOS and ICPIF are only calculated for a **udp-jitter codec** operation (`udp-jitter 10.0.12.2 5000 codec g711ulaw`), which makes the router emulate the packet size and rate of a real codec. A plain udp-jitter reports 0, and you should read nothing into it. ## The Other Operations ICMP echo and udp-jitter cover most needs. The rest of the list worth knowing: `icmp-echo`Reachability and RTT, no responder needed. The workhorse for tracking. `udp-jitter`Jitter, one-way latency, loss. Needs a responder. `udp-jitter ... codec`As above plus a real MOS score. What you want for VoIP SLAs. `http`DNS, TCP connect and transfer time, measured separately. `dns`Resolution time. Catches the DNS server nobody is monitoring. `path-echo`Per-hop RTT. Finds the hop adding the delay. ## Measuring Is Half the Job A probe that nobody looks at does nothing. IP SLA becomes powerful when you wire it to something that acts on the result: an object tracker that swaps a static route, an HSRP group that changes priority, an EEM applet that raises an alert. The rest of this article does that on a second lab and captures the failover as it happens. The full dual-WAN pattern (primary plus floating backup) is built end to end in the companion article on [making a static route fail over when the far end dies](https://www.pinglabz.com/object-tracking-reliable-static-routes/). ## Turning a Probe Into a Tracked Object Three iol-xe routers in a line on IOS XE 17.18.2: R1 holds the probe, R2 transits, R3 sits behind R2 with 8.8.8.0/24 on a loopback. ``` R1(config)# ip sla 1 R1(config-ip-sla-echo)# icmp-echo 10.0.23.2 source-ip 10.0.12.1 R1(config-ip-sla-echo)# frequency 5 R1(config)# ip sla schedule 1 life forever start-time now ! R1(config)# track 1 ip sla 1 reachability ! R1(config)# ip route 8.8.8.0 255.255.255.0 10.0.12.2 track 1 ``` The asymmetry is the point. The next hop is 10.0.12.2 (R2, one hop away); the probe target is 10.0.23.2, two hops past it. Probe your own next hop and you have rebuilt interface tracking with extra steps. Add `delay down 10 up 30` under the track too, or one lost packet becomes a routing change. Healthy state, captured off R1: ``` R1#show track 1 Track 1 IP SLA 1 reachability Reachability is Up 2 changes, last change 00:01:26 Latest operation return code: OK Latest RTT (millisecs) 4 Tracked by: Static IP Routing 0 R1#show ip sla statistics 1 Latest operation return code: OK Number of successes: 19 Number of failures: 3 R1#show ip route 8.8.8.0 Routing entry for 8.8.8.0/24 Known via "static", distance 1, metric 0 * 10.0.12.2 ``` `Tracked by: Static IP Routing 0` names the subsystem that subscribed to the object; an empty list means you built a check that changes nothing. The three failures are from boot, before convergence; that counter never resets, so alarm on its rate of change or on the track state. ## Watching the Failover Happen To break it, an EEM applet on R2 shut Ethernet0/1, the interface facing R3\. Nothing changed on R1: its interface stayed up, its route to the next hop stayed valid, its ARP entry stayed put. Locally nothing failed, which is the scenario a plain static route loses to. ``` *Jul 20 21:03:40.297: %TRACK-6-STATE: 1 ip sla 1 reachability Down -> Up ! ... R2 shuts Ethernet0/1, breaking the path to the probe target ... *Jul 20 21:06:10.306: %TRACK-6-STATE: 1 ip sla 1 reachability Up -> Down ``` The first line is boot convergence. The second is the trigger. `%TRACK-6-STATE` is severity 6 informational, on by default and easy to scroll past. Filter for it and alert on it: it names both the object and the direction. ``` R1#show track 1 Reachability is Down 3 changes, last change 00:00:56 Latest operation return code: Timeout Tracked by: Static IP Routing 0 R1#show ip sla statistics 1 Latest RTT: NoConnection/Busy/Timeout Latest operation return code: Timeout Number of successes: 30 Number of failures: 16 R1#show ip route 8.8.8.0 % Network not in table ``` `NoConnection/Busy/Timeout` is one field listing three possible non-numeric outcomes, so read the `return code` below it instead. The counters moving 19/3 to 30/16 prove the probe never stopped polling, which matters when you debug a failover that did not happen (a probe that silently died generates no failures and never brings a track down). And `% Network not in table` is the mechanism proven: a device two hops away lost an interface, and a static route on a router that saw no local failure pulled itself from the RIB. ## What Takes Over When the Route Is Withdrawn In this lab, nothing did, deliberately, so the withdrawal was unambiguous. In production a second route waits, and the reason it waits is [how a router chooses between two sources for the same prefix](https://www.pinglabz.com/administrative-distance/). Tracking does not perform the failover. It performs the withdrawal. Administrative distance performs the failover. ``` ip route 8.8.8.0 255.255.255.0 10.0.12.2 track 1 ip route 8.8.8.0 255.255.255.0 10.0.14.2 200 ``` The healthy capture read `Known via "static", distance 1`. That second line is a floating static at AD 200: invisible while the AD 1 route exists, best remaining candidate the instant the track pulls it. Pick the distances carelessly (200 for the backup when an EIGRP path at 90 also carries the prefix) and the failover lands somewhere you did not intend. ## Tracking Things That Are Not Static Routes The object is generic, and `show track` lists every subsystem consuming it. The common second consumer is a first hop redundancy group: `standby 1 track 1 decrement 30` drops an HSRP router's priority when the probe fails. Better than tracking an interface, but a flapping probe now produces a flapping HSRP group, one of the classic causes of [an HSRP pair that keeps swapping Active](https://www.pinglabz.com/hsrp-troubleshooting-flapping-dual-active/), so set the track `delay` first. The same trade-off appears on a firewall pair running [stateful failover between two ASA units](https://www.pinglabz.com/cisco-asa-stateful-failover/): monitor something close for late but rare failovers, or something distant for early ones with occasional false positives. ## Common Mistakes and Gotchas - **Nothing is consuming the track.** An empty `Tracked by` list is a check wired to nothing: state changes, no traffic moves. - **Probing the next hop instead of past it.** Same IP as the static's next hop means you reinvented interface tracking. - **Reading the cumulative failure counter as health.** It never resets, and boot convergence seeds it. - **No `delay` on the track.** Every lost packet on a marginal circuit becomes a routing change. ## FAQ ### Why is my SLA showing no statistics at all? You almost certainly forgot `ip sla schedule N life forever start-time now`. An unscheduled operation never runs and reports nothing. ### Do I need a responder for ICMP echo? No. Any device that answers a ping works, including a Linux host or a firewall. The responder is only for jitter and one-way measurements. ### What is the difference between timeout and threshold? Timeout produces a FAILURE and can bring down a tracked object. Threshold produces a warning while still counting as a SUCCESS. Only timeout drives failover. ### How much traffic does a probe generate? Very little, but not zero, and a high-frequency udp-jitter probe across hundreds of routers is real traffic. Size the frequency to the decision: 5 seconds for a failover trigger, 60 seconds for a trend graph. ### Why is MOS always 0? Because you are running a plain `udp-jitter` operation. MOS and ICPIF are only calculated for codec-based jitter operations. ### My track is Down but the route is still in the table Check `show track` for a `Tracked by` entry. If `Static IP Routing` is not listed, the route was configured without the `track` keyword and is not subscribed to the object. ### Should I track reachability or state? Use `reachability` to fail over on a dead path, which is the common case. Use `state` if a degraded-but-alive path is also unacceptable, and accept that it treats `Over threshold` as a failure. ## Key Takeaways - IP SLA answers "is this path usable," not "is this link up." - `ip sla schedule` is mandatory. Without it the operation exists and never runs. - **timeout** causes a failure and drives tracking. **threshold** only flags a slow-but-successful probe. - The **responder** enables jitter, one-way latency and loss-direction figures. ICMP echo does not need one; udp-jitter does. - Mark the probe like the traffic you care about (`tos 184` for voice), or you measure a path the real traffic will not take. - A degraded path reports `Over threshold` with `Number of failures: 0`. Green dashboards and unusable links coexist happily. - MOS and ICPIF only populate on codec-based jitter operations. - `track N ip sla N reachability` turns a return code into an object state, and `Tracked by` names whatever consumes it. - The captured failover needed no local failure: R2 lost an interface, R1's SLA returned `Timeout`, `%TRACK-6-STATE` logged Up to Down, the route became `% Network not in table`. - Tracking withdraws a route; administrative distance decides what replaces it. Design both, and point the probe past the next hop. Next: build the dual-WAN version with [a primary and a floating backup static](https://www.pinglabz.com/object-tracking-reliable-static-routes/), or work through the rest of the [router services toolkit on PingLabz](https://www.pinglabz.com/ip-services/). ### NetFlow v5 vs v9 vs Flexible NetFlow vs IPFIX URL: https://www.pinglabz.com/netflow-versions-compared/ Last updated: 2026-07-12T01:23:27.000Z Four names, one job: tell a collector which conversations crossed the router and how big they were. The differences between NetFlow v5, v9, Flexible NetFlow, and IPFIX are not academic trivia, because picking the wrong one means either losing the fields you need or paying for a collector license you did not have to. Here is what actually separates them, and how to decide. For the hands-on configuration, see [Flexible NetFlow on Cisco IOS XE](https://www.pinglabz.com/flexible-netflow-cisco-ios-xe/); for the wider context, the [IP Services cluster guide](https://www.pinglabz.com/ip-services/). ## The Short Version NetFlow v5 FieldsFixed, IPv4 only TemplatesNo Verdict Legacy. Simple, universally supported, cannot carry IPv6 or MPLS. NetFlow v9 FieldsTemplate-defined TemplatesYes Verdict The workhorse. IPv6, MPLS, VLAN, MAC. What FNF exports by default. Flexible NetFlow FieldsYou define them TemplatesYes (v9 or IPFIX) Verdict Not a wire protocol. A Cisco config framework that EXPORTS v9 or IPFIX. IPFIX (v10) FieldsTemplate + vendor IEs TemplatesYes Verdict The IETF standard (RFC 7011). Multi-vendor. Variable-length fields. The single most important line in that grid: **Flexible NetFlow is not a protocol.** It is the configuration model on the router. When you configure FNF, what leaves the box is still NetFlow v9 (or IPFIX, if you ask for it). People compare "v9 vs Flexible NetFlow" as if they are alternatives. They are not. They are different layers. ## NetFlow v5: The Fixed Record v5 exports a fixed 48-byte record with a fixed set of fields, in a fixed order. There is no negotiation and no template. The collector knows the layout because the layout is in the RFC-adjacent spec and never changes. That simplicity is its virtue and its ceiling. The fields you get are the classic ones: source and destination IPv4 address, source and destination port, protocol, ToS, TCP flags, input and output interface index, packet and byte counts, start and end timestamps, next-hop, and source and destination AS. What you cannot get, ever: - **IPv6.** The address fields are 32 bits. There is nowhere to put a v6 address. This alone disqualifies v5 from any modern network. - **MPLS labels, VLAN IDs, MAC addresses, DSCP as a first-class field.** - **Any custom key.** You cannot say "I want to key on VLAN," because the record is what it is. v5 survives because it is trivially easy to parse and every collector on earth supports it. If you have an old device that speaks nothing else, it still works. Do not design anything new around it. ## NetFlow v9: Templates Change Everything v9 solved the rigidity problem by separating the description of the data from the data itself. The router periodically sends a **template** record that says "field 1 is a 4-byte source IPv4 address, field 2 is a 4-byte destination address, field 3 is a 2-byte source port..." and then sends **data** records that are just values in that order. You can watch it happen. Here are two consecutive export packets from the lab, captured on the collector: ``` j@llmbits:~$ sudo tcpdump -i ens224 -n udp port 2055 18:44:51.294654 IP 192.168.99.1.58769 > 192.168.99.100.2055: UDP, length 296 18:44:52.295245 IP 192.168.99.1.58769 > 192.168.99.100.2055: UDP, length 132 ``` 296 bytes then 132 bytes. The first is the template (bulky, because it names every field). The second is data (compact, because it does not). Every subsequent data packet is small, and the template is only resent periodically. This design has one sharp edge that causes most real-world v9 problems: **if the collector misses the template, it cannot decode anything.** Not "it decodes partially." Nothing. The data packets are an undifferentiated pile of bytes without the template to give them meaning. This is why the export configuration includes: ``` R1(config-flow-exporter)# template data timeout 60 ``` Sixty seconds is the worst case a restarted collector will sit blind. The default on many platforms is 600, which means a collector that restarts at the wrong moment produces a ten-minute hole in your data and nobody notices until someone asks about that hole in the graph. The upside of templates is that v9 can carry anything Cisco chooses to define: IPv6 addresses, MPLS label stacks, VLAN tags, MAC addresses, BGP next-hop, and so on. The record in the lab exports interface names, ToS, and 64-bit byte counters, none of which v5 could express. ## Flexible NetFlow: The Configuration Model FNF is what Cisco gave you on the router so you could actually exploit v9's flexibility. Rather than picking from a menu of pre-baked record types, you build the record yourself out of `match` (key) and `collect` (non-key) fields. ``` R1#show flow record PINGLABZ-IPV4 flow record PINGLABZ-IPV4: No. of users: 1 Total field space: 54 bytes Fields: match ipv4 tos match ipv4 protocol match ipv4 source address match ipv4 destination address match transport source-port match transport destination-port collect interface input collect interface output collect counter bytes long collect counter packets long collect timestamp absolute first collect timestamp absolute last ``` That record is impossible in v5 and it is a routine afternoon in FNF. And it exports as v9: ``` R1#show flow exporter PINGLABZ-COLLECTOR Flow Exporter PINGLABZ-COLLECTOR: Export protocol: NetFlow Version 9 Transport Configuration: Destination IP address: 192.168.99.100 Transport Protocol: UDP Destination Port: 2055 Export template data timeout: 60 ``` So when someone asks "are you running Flexible NetFlow or v9," the correct answer is usually "yes." FNF is how you configured it; v9 is what is on the wire. ## IPFIX: The Standard IPFIX (RFC 7011) is the IETF's standardisation of v9, and it is sometimes called NetFlow v10 for that reason. It keeps the template model and adds: - **Vendor-specific Information Elements.** A vendor can define its own fields under its own enterprise number without colliding with anyone else. This is how modern telemetry (application IDs, URLs, user identity) rides along. - **Variable-length fields.** v9 fields are fixed-length. IPFIX can carry a hostname or a URL, which have no fixed length. - **Transport flexibility.** The RFC prefers SCTP for reliability, and permits TCP and UDP. In practice, essentially everyone runs it over UDP anyway. - **Multi-vendor by design.** Juniper, Huawei, Palo Alto, nProbe, and open-source exporters all speak IPFIX. v9 is a Cisco protocol that others chose to implement. On IOS XE, switching from v9 to IPFIX is one line in the exporter: ``` R1(config-flow-exporter)# export-protocol ipfix ``` The record, the monitor, and the interface config do not change at all. That is the FNF payoff: the wire format is a detail you can swap. ## Which One Do I Use? All-Cisco, existing collectorFNF exporting **v9**. It is the default, it is universally supported, and it does everything you need. Mixed-vendor estateFNF exporting **IPFIX**. One collector config for Cisco, Juniper, and everything else. You need app-awareness / URLs**IPFIX**. Variable-length fields and enterprise IEs are the only way to carry that. Ancient device, no choice**v5**, and plan its replacement. No IPv6 is a hard ceiling. The honest default for most enterprises: configure with Flexible NetFlow, export v9, and switch the single `export-protocol` line to IPFIX the day a non-Cisco device shows up. Because the record definition is separate from the wire format, that migration costs you one line and a collector setting. ## A Note on sFlow, Which Is Not NetFlow sFlow is frequently mentioned in the same breath and it is a fundamentally different thing. NetFlow (all versions) maintains a **flow cache**: the router tracks conversations and exports summaries. sFlow does not track flows at all; it randomly samples packets and ships packet headers to the collector, which reconstructs a statistical picture. Consequences: sFlow is cheaper on the device (no cache to maintain, which is why it is common on merchant-silicon switches), always sampled (so never a complete record), and gives you packet headers rather than flow summaries. NetFlow gives you complete flow records (if unsampled) and is more expensive to produce. For security forensics, unsampled NetFlow wins. For "which port is hot on my 48-port switch," sFlow is fine and free. ## FAQ ### Is Flexible NetFlow a protocol? No. It is the IOS configuration framework (records, exporters, monitors). It exports NetFlow v9 by default and IPFIX on request. This is the single most common misconception about it. ### Is IPFIX just NetFlow v10? Effectively yes, and the version field in the header does say 10\. It is the IETF standardisation of v9 with variable-length fields and vendor IEs added. ### Can one exporter send both v9 and IPFIX? Not the same exporter; `export-protocol` is a single setting. But you can define two exporters and attach both to the same monitor, which sends the same flows to two collectors in two formats. Useful during a migration. ### Why does my collector show nothing even though the router says it is sending? Almost always a template problem. The collector started after the last template and is waiting for the next one. Lower `template data timeout` and restart the collector. ### Does v9 support IPv6? Yes, fully. v5 does not, and cannot. If your network has IPv6 (see the [IPv6 cluster guide](https://www.pinglabz.com/ipv6/)), v5 is already lying to you about your traffic. ## Key Takeaways - **Flexible NetFlow is a config model, not a wire protocol.** It exports v9 or IPFIX. Comparing "v9 vs FNF" is a category error. - **v5** is fixed-format and IPv4-only. It cannot carry an IPv6 address, ever. Legacy only. - **v9** introduced templates, which is what makes IPv6, MPLS, and VLAN fields possible. Miss the template and the collector decodes nothing. - **IPFIX** is the IETF standard version of v9, with variable-length fields and vendor IEs. Choose it for multi-vendor or app-aware telemetry. - Switching v9 to IPFIX is one line (`export-protocol ipfix`) because FNF separates the record from the wire format. - **sFlow is not NetFlow.** It samples packets rather than tracking flows. Cheaper, always sampled, different tool. Next: [configure Flexible NetFlow end to end](https://www.pinglabz.com/flexible-netflow-cisco-ios-xe/), or the [IP Services cluster guide](https://www.pinglabz.com/ip-services/). ### Flexible NetFlow on Cisco IOS XE: Records, Monitors, and Exporters URL: https://www.pinglabz.com/flexible-netflow-cisco-ios-xe/ Last updated: 2026-07-12T01:23:27.000Z NetFlow answers the question your monitoring dashboard cannot: not "is the link busy" but "who is making it busy, talking to what, on which port, and how much." It is the difference between knowing you have a problem and knowing what the problem is. Flexible NetFlow (FNF) is the modern implementation, and it is genuinely flexible: instead of a fixed set of exported fields, you define your own record. That flexibility is also why it confuses people, because you now have three objects to configure instead of one. This article builds the whole thing on Cisco IOS XE, exports the flows to a real Linux collector, and decodes them there, so you can see the data arrive rather than take it on faith. It extends the [IP Services cluster guide](https://www.pinglabz.com/ip-services/). ## The Three Objects (and Why There Are Three) Every Flexible NetFlow configuration is built from three pieces that plug into each other. Get the mental model right and the syntax stops being arbitrary. 1\. Flow Record Answers WHAT do I care about? Which fields define a flow, and which do I count? Keywords`match` / `collect` 2\. Flow Exporter Answers WHERE do I send it? Collector IP, port, protocol version. OptionalYes (cache-only is valid) 3\. Flow Monitor Answers Glues record + exporter together and owns the cache. This is what you apply to an interface. Applied toInterfaces The monitor is the only one that touches an interface. The record and exporter are reusable building blocks: one record can feed many monitors, one exporter can serve many monitors. ## match vs collect: The Distinction That Matters This trips up almost everyone the first time. Inside a flow record: - **`match` defines a KEY field.** Key fields are the flow's identity. If two packets have identical values across every key field, they belong to the same flow. Change one key field and it is a different flow, with its own cache entry. - **`collect` defines a NON-KEY field.** Non-key fields are the things you accumulate or record about the flow: byte counts, packet counts, timestamps, which interface it came in on. Choosing keys is a real design decision. Match on the classic five-tuple (source IP, destination IP, protocol, source port, destination port) and you get per-conversation visibility. Match only on source and destination IP and you get a much smaller cache but you lose the ability to say "it was all port 443." Match on too much and the cache explodes. Here is the record from the lab, a standard five-tuple plus ToS: ``` R1(config)# flow record PINGLABZ-IPV4 R1(config-flow-record)# match ipv4 source address R1(config-flow-record)# match ipv4 destination address R1(config-flow-record)# match ipv4 protocol R1(config-flow-record)# match transport source-port R1(config-flow-record)# match transport destination-port R1(config-flow-record)# match ipv4 tos R1(config-flow-record)# collect counter bytes long R1(config-flow-record)# collect counter packets long R1(config-flow-record)# collect timestamp absolute first R1(config-flow-record)# collect timestamp absolute last R1(config-flow-record)# collect interface input R1(config-flow-record)# collect interface output ``` Verify what the router made of it: ``` R1#show flow record PINGLABZ-IPV4 flow record PINGLABZ-IPV4: Description: Key fields identify the flow; nonkey fields are what we count No. of users: 1 Total field space: 54 bytes Fields: match ipv4 tos match ipv4 protocol match ipv4 source address match ipv4 destination address match transport source-port match transport destination-port collect interface input collect interface output collect counter bytes long collect counter packets long collect timestamp absolute first collect timestamp absolute last ``` **Total field space: 54 bytes.** That is the per-flow memory cost of your design decision, stated plainly. Add more key fields and this grows, and so does every exported record on the wire. ## The Exporter: Getting Flows Off the Box The cache on the router is useful for a quick look, but the point of NetFlow is shipping the data somewhere it can be stored and queried. In the lab the collector is a real Debian machine at 192.168.99.100 running `nfcapd`. ``` R1(config)# flow exporter PINGLABZ-COLLECTOR R1(config-flow-exporter)# destination 192.168.99.100 R1(config-flow-exporter)# source Ethernet0/0 R1(config-flow-exporter)# transport udp 2055 R1(config-flow-exporter)# export-protocol netflow-v9 R1(config-flow-exporter)# template data timeout 60 ``` Two lines deserve attention. `source Ethernet0/0` pins the source address of the export packets, which matters because collectors identify exporters by source IP; if it changes, your collector thinks a new device appeared. And `template data timeout 60` controls how often the v9 template is resent, which is the single most important setting for a collector that might restart. More on why in the [NetFlow versions comparison](https://www.pinglabz.com/netflow-versions-compared/). ``` R1#show flow exporter PINGLABZ-COLLECTOR Flow Exporter PINGLABZ-COLLECTOR: Description: Export to the Debian VM running nfcapd Export protocol: NetFlow Version 9 Transport Configuration: Destination type: IP Destination IP address: 192.168.99.100 Source IP address: 192.168.99.1 Source Interface: Ethernet0/0 Transport Protocol: UDP Destination Port: 2055 Source Port: 58769 DSCP: 0x0 TTL: 255 Export template data timeout: 60 ``` ## The Monitor: Cache and Application ``` R1(config)# flow monitor PINGLABZ-MON R1(config-flow-monitor)# record PINGLABZ-IPV4 R1(config-flow-monitor)# exporter PINGLABZ-COLLECTOR ! R1(config)# interface Ethernet0/2 R1(config-if)# ip flow monitor PINGLABZ-MON input R1(config-if)# ip flow monitor PINGLABZ-MON output ``` Applying it `input` and `output` on the same interface is deliberate. NetFlow is directional: an input monitor sees traffic arriving on that interface only. If you apply it in one direction, you will see half of every conversation and spend an afternoon wondering why the byte counts look wrong. ``` R1#show flow monitor PINGLABZ-MON Flow Monitor PINGLABZ-MON: Flow Record: PINGLABZ-IPV4 Flow Exporter: PINGLABZ-COLLECTOR Cache: Type: normal Status: allocated Size: 4096 entries / 491584 bytes Inactive Timeout: 15 secs Active Timeout: 1800 secs ``` Those two timeouts govern when a flow leaves the cache and gets exported: Inactive timeout (15s)The flow has gone quiet for this long, so it is finished. Age it out and export. Active timeout (1800s)The flow is still running, but export a snapshot anyway so long-lived transfers are not invisible for 30 minutes. **A platform note, honestly stated:** on the IOL-XE image used for this lab (17.18.2), the `cache timeout active` and `cache timeout inactive` commands are rejected under the flow monitor with `% Invalid input detected`. The defaults above are what you get. On real hardware and on CSR/Catalyst 8000v images those commands work normally, and tuning the active timeout down to 60 seconds is standard practice so long transfers show up in your dashboard while they are still happening. This article shows the default behaviour because that is what the lab genuinely produced. ## Reading the Cache With iperf3 and ping traffic pushed across the monitored interface, the cache fills up: ``` R1#show flow monitor PINGLABZ-MON cache format table Cache type: Normal Cache size: 4096 Current entries: 5 High Watermark: 12 Flows added: 12 Flows aged: 7 - Active timeout ( 1800 secs) 0 - Inactive timeout ( 15 secs) 7 IPV4 SRC ADDR IPV4 DST ADDR TRNS SRC PORT TRNS DST PORT IP PROT intf input intf output bytes long pkts long =============== =============== ============= ============= ======= ========== =========== =========== ========= 192.168.99.100 10.0.10.100 51236 5201 6 Et0/0 Et0/2 1225 14 10.0.10.100 192.168.99.100 5201 51236 6 Et0/2 Et0/0 1017 13 192.168.99.100 10.0.10.100 46438 5201 17 Et0/0 Et0/2 1019948 692 192.168.99.100 10.0.10.100 0 2048 1 Et0/0 Et0/2 252 3 10.0.10.100 192.168.99.100 0 0 1 Et0/2 Et0/0 252 3 ``` Every row is a flow, and the key fields are doing exactly what they promised. Protocol 6 is TCP (the iperf3 control connection, both directions, separately). Protocol 17 is UDP, and it is the big one at roughly a megabyte in 692 packets. Protocol 1 is ICMP, and note the port columns: ICMP has no ports, so the router stuffs the type and code in there instead, which is why you see 2048 (type 8, echo request, shifted). Look at the `Flows aged` counter: 7 flows aged out via the inactive timeout, 0 via the active timeout. Those 7 are what got exported. And confirm the export actually happened: ``` R1#show flow exporter PINGLABZ-COLLECTOR statistics Flow Exporter PINGLABZ-COLLECTOR: Packet send statistics (last cleared 00:01:43 ago): Successfully sent: 2 (594 bytes) Client send statistics: Client: Flow Monitor PINGLABZ-MON Records added: 10 - sent: 7 Bytes added: 540 - sent: 378 ``` If `Successfully sent` stays at 0, the router is not shipping anything and you have a routing or source-interface problem, not a NetFlow problem. ## On the Wire Do not trust the router's own counters alone. Here is `tcpdump` on the collector, confirming the export packets physically arrive: ``` j@llmbits:~$ sudo tcpdump -i ens224 -n udp port 2055 listening on ens224, link-type EN10MB (Ethernet), snapshot length 262144 bytes 18:44:51.294654 IP 192.168.99.1.58769 > 192.168.99.100.2055: UDP, length 296 18:44:52.295245 IP 192.168.99.1.58769 > 192.168.99.100.2055: UDP, length 132 ``` Two packets, two different sizes, and that difference is the whole of NetFlow v9 in one screen. The 296-byte packet is the **template**: it tells the collector what the fields are and in what order. The 132-byte packet is the **data**: just values, no field names, because the collector already knows the layout from the template. This is why `template data timeout` matters so much. If your collector restarts and misses the template, it cannot decode a single data packet until the next template arrives. Sixty seconds of blindness is tolerable. The default of 600 is not. ## Decoded on the Collector The final proof. `nfcapd` wrote the flows to disk; `nfdump` reads them back: ``` j@llmbits:~$ nfdump -R /home/j/nfcap -o long Date first seen Duration Proto Src IP Addr:Port Dst IP Addr:Port Packets Bytes Flows 2026-07-11 17:45:30.137 00:00:02.004 ICMP 192.168.99.100:0 -> 10.0.10.100:8.0 3 252 1 2026-07-11 17:45:30.139 00:00:02.004 ICMP 10.0.10.100:0 -> 192.168.99.100:0.0 3 252 1 2026-07-11 17:46:38.021 00:00:05.073 TCP 192.168.99.100:40556 -> 10.0.10.100:5201 16 1488 1 2026-07-11 17:46:38.023 00:00:05.070 TCP 10.0.10.100:5201 -> 192.168.99.100:40556 14 1222 1 2026-07-11 17:46:38.031 00:00:05.058 TCP 192.168.99.100:40570 -> 10.0.10.100:5201 6467 9.7 M 1 2026-07-11 17:46:38.033 00:00:05.060 TCP 10.0.10.100:5201 -> 192.168.99.100:40570 3653 194636 1 2026-07-11 17:46:38.037 00:00:05.047 TCP 192.168.99.100:40578 -> 10.0.10.100:5201 5425 8.1 M 1 2026-07-11 17:46:43.111 00:00:05.002 UDP 192.168.99.100:51231 -> 10.0.10.100:5201 1728 2.5 M 1 2026-07-11 17:46:48.132 00:00:01.001 ICMP 192.168.99.100:0 -> 10.0.10.100:8.0 2 168 1 Summary: total flows: 14, total bytes: 20.7 M, total packets: 20412, avg bps: 2.1 M, avg pps: 258, avg bpp: 1016 Time window: 2026-07-11 17:45:30 - 2026-07-11 17:49:00, Duration: 00:03:30.000 ``` Fourteen flows, 20.7 megabytes, 20,412 packets, and every one of them is a real packet the router forwarded and reported. The two parallel TCP streams at 9.7M and 8.1M are the iperf3 `-P 2` test. The 2.5M UDP flow is the UDP test. The asymmetry between the TCP send and receive rows (9.7 M out, 194 KB back) is exactly what a bulk transfer looks like: data one way, ACKs the other. That is the payoff. Not a config snippet, but the actual data landing somewhere you can query it. ## Sampling, and Not Melting the Box NetFlow costs CPU and memory. On a busy edge router, processing every single packet is not free, and a 4096-entry cache fills fast on a real internet link. Two mitigations: - **Sampling.** A `flow sampler` tells the router to process 1 in N packets. Statistically valid for volume trends, useless for security forensics (you will miss the one packet that mattered). Sample on high-speed core links, never on a link you are using for incident response. - **Cache sizing.** `cache entries N` raises the ceiling. Watch the `High Watermark` in the cache output. If it is pinned at your cache size, you are dropping flows and your data is silently incomplete. The general rule: monitor the interfaces where the answer lives (edge, WAN, DMZ), not every interface on every box. NetFlow on a core link between two aggregation switches usually tells you nothing you cannot get from an interface counter. ## When It Does Not Work Cache is emptyMonitor not applied, or applied to an interface with no traffic. Check `show flow interface`. Cache fills, nothing exportedNo exporter bound to the monitor, or `Successfully sent: 0`. Check routing to the collector. Packets sent, collector shows nothingAlmost always the template. Restart order matters; lower `template data timeout`. Only half the conversationMonitor applied `input` only. Add `output`. And when you need to see individual packets rather than flow summaries, that is a different tool: see [SPAN, RSPAN, and ERSPAN](https://www.pinglabz.com/span-rspan-erspan-configuration/) for packet mirroring, and [conditional debugging](https://www.pinglabz.com/conditional-debug-cisco-ios-xe/) for surgical debug output. ## FAQ ### Do I need a flow exporter? No. A monitor with only a record is valid and gives you `show flow monitor cache` on the box. That is genuinely useful for live troubleshooting. You need an exporter when you want history. ### What port should the collector listen on? 2055 is the convention, but nothing enforces it. 9995 and 9996 are also common. Match the router config to whatever your collector expects. ### Can I run NetFlow and a packet capture at the same time? Yes. They answer different questions. NetFlow tells you which conversations exist and how big they are; a capture tells you what was inside them. ### Does NetFlow see traffic the router drops with an ACL? Input NetFlow is processed before the outbound ACL, so denied traffic can still appear in the cache. That is useful (you can see what is being blocked) and occasionally confusing (the flow shows bytes that never reached the destination). ## Key Takeaways - Three objects: **record** (what), **exporter** (where), **monitor** (glue plus cache). Only the monitor is applied to an interface. - `match` \= key field = flow identity. `collect` \= non-key field = what you count. The record's "total field space" is the memory cost of your choices. - Apply the monitor `input` AND `output`, or you will only see half of every conversation. - NetFlow v9 sends a template packet and data packets separately. If the collector misses the template it decodes nothing, so keep `template data timeout` low. - Verify at three layers: `show flow monitor cache` (the router saw it), `show flow exporter statistics` (the router sent it), and the collector itself (it arrived and decoded). - On this IOL-XE image the cache timeouts cannot be tuned and sit at the 1800s/15s defaults. On production hardware, tune them. Next: [NetFlow v5 vs v9 vs FNF vs IPFIX](https://www.pinglabz.com/netflow-versions-compared/), or the [IP Services cluster guide](https://www.pinglabz.com/ip-services/). ### Troubleshooting MPLS L3VPN: Where VPN Routes Go Missing URL: https://www.pinglabz.com/troubleshooting-mpls-l3vpn/ Last updated: 2026-07-12T00:28:27.000Z An MPLS L3VPN route makes a long journey. It starts as an IP prefix in a customer's routing table, gets learned by a PE into a VRF, redistributed into MP-BGP, tagged with a route target, given a VPN label, advertised across an iBGP session, filtered on import at the far end, installed into another VRF, and finally handed to another customer router. There are at least six places it can quietly disappear, and none of them log an error. This article gives you a repeatable path to walk. Follow the route from where it exists to where it does not, and the gap is the fault. Every capture is from a live five-router IOS XE lab where each failure was deliberately induced. For label-plane problems (LDP, LSPs) see [Troubleshooting MPLS: LDP neighbors, label bindings, and broken LSPs](https://www.pinglabz.com/troubleshooting-mpls-ldp/), and the [MPLS cluster guide](https://www.pinglabz.com/mpls/) for everything else. ## The Path a Route Takes Learn this sequence and you can troubleshoot any L3VPN. Each arrow is a place a route can vanish: ``` CE1 routing table | (1) PE-CE protocol: static / OSPF / EIGRP / eBGP PE1 VRF routing table <-- show ip route vrf CUST-A | (2) redistribution into MP-BGP (NOT automatic, except with eBGP) PE1 VPNv4 BGP table <-- show ip bgp vpnv4 vrf CUST-A | (3) export: RT stamped on, VPN label allocated | (4) MP-iBGP VPNv4 session (needs send-community extended) PE2 VPNv4 BGP table <-- show ip bgp vpnv4 all | (5) import: RT must match PE2's import list PE2 VRF routing table <-- show ip route vrf CUST-A | (6) PE-CE protocol back out CE2 routing table ``` The method: pick a prefix. Check for it at every stage, starting at the origin. The moment it is missing, you have found the broken link. Do not start in the middle, and do not start by pinging. ## Failure 1: The Route Is in the VRF but Not in BGP By far the most common L3VPN mistake, and it does exactly what you would fear: nothing. No log message, no error, no symptom on the local PE at all. Here is PE2 in the lab. It has a static route for the customer's site-2 LAN, pointing at CE2, sitting happily in the VRF table: ``` PE2#show ip route vrf CUST-A static Routing Table: CUST-A 10.0.0.0/8 is variably subnetted, 4 subnets, 3 masks S 10.20.20.0/24 [1/0] via 10.20.2.2 ``` The route exists. It is correct. It even works locally: PE2 can ping into the customer site. Now look one layer up: ``` PE2#show ip bgp vpnv4 vrf CUST-A Network Next Hop Metric LocPrf Weight Path Route Distinguisher: 65000:1 (default for vrf CUST-A) *>i 1.1.1.1/32 10.255.0.1 0 100 0 65001 i *>i 10.20.1.0/30 10.255.0.1 0 100 0 ? *> 10.20.2.0/30 0.0.0.0 0 32768 ? *>i 192.168.99.0 10.255.0.1 0 100 0 65001 i ``` 10.20.20.0/24 is not there. The route is in the VRF routing table but was never handed to MP-BGP, so it was never advertised anywhere. The far end knows nothing about it: ``` PE1#show ip route vrf CUST-A 10.20.20.0 Routing Table: CUST-A % Subnet not in table PE1#ping vrf CUST-A 10.20.20.1 source Ethernet0/0 Sending 5, 100-byte ICMP Echos to 10.20.20.1, timeout is 2 seconds: ..... Success rate is 0 percent (0/5) ``` The cause is one missing line: ``` PE2#show run | section address-family ipv4 vrf CUST-A address-family ipv4 vrf CUST-A redistribute connected ``` `redistribute connected` is there. `redistribute static` is not. Static routes in a VRF do not reach MP-BGP on their own; you have to say so: ``` PE2(config)# router bgp 65000 PE2(config-router)# address-family ipv4 vrf CUST-A PE2(config-router-af)# redistribute static ``` And the route appears on the far PE within seconds, complete with its VPN label: ``` PE1#show ip route vrf CUST-A 10.20.20.0 Routing entry for 10.20.20.0/24 Known via "bgp 65000", distance 200, metric 0, type internal * 10.255.0.2 (default), from 10.255.0.2, 00:00:10 ago MPLS label: 22 MPLS Flags: MPLS Required PE1#ping vrf CUST-A 10.20.20.1 source Ethernet0/0 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 3/3/5 ms ``` The same trap applies to OSPF (`redistribute ospf 10`) and EIGRP (`redistribute eigrp 100`, which additionally does nothing without a metric). Only eBGP as the PE-CE protocol is exempt, because the routes are already in BGP. The full comparison is in [PE-CE routing protocols in MPLS L3VPN](https://www.pinglabz.com/mpls-l3vpn-pe-ce-routing/). **The tell:** route present with `show ip route vrf`, absent from `show ip bgp vpnv4 vrf`, on the *same* router. Nothing else needs checking. ## Failure 2: The Route Is Advertised but Never Imported The route makes it into MP-BGP. It crosses the core. It arrives at the far PE. And it is thrown away at the door, because the receiving VRF is not importing its route target. Remove one line from PE1: ``` PE1(config)# vrf definition CUST-A PE1(config-vrf)# address-family ipv4 PE1(config-vrf-af)# no route-target import 65000:1 ``` The BGP session is untouched. The RD is unchanged. The export is still there. Now: ``` PE1#show ip bgp vpnv4 all summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 10.20.1.2 4 65001 6 12 43 0 0 00:01:33 2 10.255.0.2 4 65000 24 19 43 0 0 00:07:35 1 ``` The VPNv4 session to PE2 is up and has been for over seven minutes. But **PfxRcd is 1**, down from 4\. The RT filter is discarding routes as they arrive. Chase the prefix: ``` PE1#show ip bgp vpnv4 rd 65000:1 10.20.20.0/24 BGP routing table entry for 65000:1:10.20.20.0/24, version 42 Paths: (0 available, no best path) Not advertised to any peer ``` `Paths: (0 available, no best path)` is the signature. BGP has an entry for the prefix, so it clearly heard about it, but there is no usable path, because no VRF has claimed it. And the confirming command: ``` PE1#show vrf detail CUST-A | include Export|Import|VRF VRF CUST-A (VRF Id = 1); default RD 65000:1; default VPNID Export VPN route-target communities No Import VPN route-target communities VRF label distribution protocol: not configured ``` **No Import VPN route-target communities.** There is your fault, stated in plain English. Meanwhile the customer is looking at this: ``` j@llmbits:~$ ping -c 3 10.20.20.1 PING 10.20.20.1 (10.20.20.1) 56(84) bytes of data. From 192.168.99.1 icmp_seq=1 Destination Host Unreachable From 192.168.99.1 icmp_seq=2 Destination Host Unreachable From 192.168.99.1 icmp_seq=3 Destination Host Unreachable --- 10.20.20.1 ping statistics --- 3 packets transmitted, 0 received, +3 errors, 100% packet loss ``` Every core protocol is healthy. LDP is up, OSPF is up, MP-BGP is up. One import statement is missing. If the RD and RT distinction is still fuzzy, [RD vs RT explained](https://www.pinglabz.com/mpls-rd-vs-rt-explained/) takes it apart properly. **The tell:** PfxRcd lower than expected on a healthy session, plus `0 available, no best path` for a prefix that clearly arrived. ## Failure 3: The RTs Were Stripped in Transit Same symptom, different cause, and this one is nastier because the import configuration is *correct*. Route targets are BGP extended communities. Extended communities are only sent if the neighbor is configured to send them. Remove that from PE2's VPNv4 neighbor: ``` PE2(config)# router bgp 65000 PE2(config-router)# address-family vpnv4 PE2(config-router-af)# no neighbor 10.255.0.1 send-community extended ``` The result on PE1: ``` PE1#show ip bgp vpnv4 all summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 10.20.1.2 4 65001 11 20 56 0 0 00:06:21 2 10.255.0.2 4 65000 44 25 56 0 0 00:12:24 0 PE1#show ip bgp vpnv4 rd 65000:1 10.20.20.0/24 BGP routing table entry for 65000:1:10.20.20.0/24, version 55 Paths: (0 available, no best path) Not advertised to any peer ``` **Zero prefixes received.** Not one. The session is established, messages are flowing (44 received), PE2 is dutifully advertising all its VPNv4 routes, and PE1 accepts none of them, because every route arrives with no route target at all and therefore matches no import list anywhere on the box. PfxRcd of exactly 0 on an established VPNv4 session, with the far side definitely advertising, means one thing: check `send-community extended`. On both ends. **The tell:** PfxRcd = 0, session Established, no RTs visible on any received route. ## Failure 4: The Route Is Perfect and the Traffic Still Drops Everything above is control plane. This one is not. The route is in BGP, imported into the VRF, installed in the routing table, and the packets still go nowhere, because there is no label to put on them. ``` PE1#show ip route vrf CUST-A 10.20.20.0 255.255.255.0 Routing entry for 10.20.20.0/24 Known via "bgp 65000", distance 200, metric 0, type internal * 10.255.0.2 (default), from 10.255.0.2 MPLS label: 22 MPLS Flags: MPLS Required PE1#show ip cef vrf CUST-A 10.20.20.0 10.20.20.0/24 nexthop 10.30.30.2 Ethernet0/1 unusable: no label ``` The routing table is fine. CEF, which is what actually forwards the packet, says **unusable: no label**. The BGP next hop (PE2's loopback) has no transport LSP, because LDP is broken somewhere in the core. The VPN label 22 is known; the label to *get to PE2* is not, and you cannot build the stack with only the inner label. This is where the two troubleshooting articles meet. If your route is present and correct in the VRF table but traffic drops, stop looking at BGP and go to [the LDP article](https://www.pinglabz.com/troubleshooting-mpls-ldp/). The diagnostic sequence there (OSPF FULL, loopback ping succeeds, `show mpls ldp neighbor` empty) takes about ninety seconds. **The tell:** route present in the VRF, `show ip cef vrf` says `unusable: no label`. ## Failure 5: The Route Reaches the Far PE but Not the Far CE The last hop. Everything in the provider network is correct, the route is sitting in PE2's CUST-A table, and the customer still cannot see it, because it was never handed back out to the CE. Failure 1 in reverse: redistribution runs in **both** directions and people configure one. With OSPF as the PE-CE protocol you need two statements on each PE: ``` router ospf 10 vrf CUST-A redistribute bgp 65000 subnets <-- VPN routes out to the CE ! router bgp 65000 address-family ipv4 vrf CUST-A redistribute ospf 10 <-- CE routes into the VPN ``` Configure only the second and traffic flows exactly one way, which produces the memorable symptom of a ping that fails while a ping in the opposite direction succeeds. Check the CE's table directly; it is the fastest test: ``` CE2#show ip route ospf 1.0.0.0/32 is subnetted, 1 subnets O IA 1.1.1.1 [110/21] via 10.20.2.1, 00:00:37, Ethernet0/1 10.0.0.0/8 is variably subnetted, 5 subnets, 3 masks O IA 10.20.1.0/30 [110/11] via 10.20.2.1, 00:00:37, Ethernet0/1 O IA 192.168.99.0/24 [110/30] via 10.20.2.1, 00:00:37, Ethernet0/1 ``` (The `O IA` is expected, not a fault. The MPLS backbone acts as an OSPF superbackbone, so VPN routes arrive as inter-area.) With eBGP as the PE-CE protocol there is a different last-hop trap: if the customer uses the same AS number at every site, CE2 will reject routes originated by CE1 because it sees its own AS in the path. The fix is `as-override` on the PE, and the proof is in the AS path: ``` CE2#show ip bgp Network Next Hop Metric LocPrf Weight Path *> 192.168.99.0 10.20.2.1 0 65000 65000 i ``` Path `65000 65000`: the customer's 65001 has been overwritten by the provider's AS, so CE2 no longer sees itself in the path and accepts the route. ## The Six Commands, In Order `show ip route vrf X` (originating PE)Is the customer route even here? If not, the PE-CE protocol is the problem. `show ip bgp vpnv4 vrf X` (originating PE)In the VRF but not here? Missing redistribution. `show ip bgp vpnv4 all summary` (receiving PE)Session up? PfxRcd sensible? 0 means send-community extended. `show ip bgp vpnv4 rd `"0 available, no best path" means it arrived and was rejected on import. `show vrf detail X`Confirms the import and export RT lists. Reads in plain English. `show ip cef vrf X `Route perfect but traffic dropping? "unusable: no label" points at LDP. ## FAQ ### The VPNv4 session is up but PfxRcd is 0\. Where do I start? `send-community extended` on the advertising PE. That is the cause more often than anything else. If it is present, confirm the far PE is actually exporting (check `show ip bgp vpnv4 vrf X` on the sender, not the receiver). ### Why does the route show in BGP but not in the VRF routing table? The RT on the route does not match any RT in the VRF's import list. Run `show vrf detail` on the receiving PE and compare against the extended community on the route. ### The route is in the VRF table and I still cannot ping. You have crossed from a control-plane problem to a data-plane one. Run `show ip cef vrf X `. If it says `unusable: no label`, the transport LSP to the remote PE is broken and this is an LDP problem, not a VPN one. ### Can I clear just the VPNv4 routes without bouncing the session? Yes, use soft reconfiguration: `clear ip bgp vpnv4 unicast soft in` (or `soft out` on the advertising side). Route refresh is supported by every modern IOS XE image, so a hard clear is almost never necessary. ## Key Takeaways - Walk the route's path from origin to destination and stop at the first stage where it is missing. Do not start by pinging. - A route in `show ip route vrf` but not in `show ip bgp vpnv4 vrf` on the same router means redistribution is missing. This is the most common L3VPN fault. - PfxRcd lower than expected with a healthy session means a route-target import problem. `show vrf detail` will say `No Import VPN route-target communities` in so many words. - PfxRcd of exactly 0 on an established VPNv4 session means `send-community extended` is missing and the RTs are being stripped. - `Paths: (0 available, no best path)` means the route arrived and was rejected on import, not that it never came. - Perfect route plus dropped traffic plus `unusable: no label` is a broken LSP, not a BGP problem. Go to LDP. - Redistribution runs both ways. One-way redistribution produces one-way traffic. See also: [Troubleshooting MPLS: LDP, label bindings and broken LSPs](https://www.pinglabz.com/troubleshooting-mpls-ldp/), [the full L3VPN configuration walkthrough](https://www.pinglabz.com/mpls-l3vpn-configuration-step-by-step/), and the [MPLS cluster guide](https://www.pinglabz.com/mpls/). ### Troubleshooting MPLS LDP: "unusable: no label" and Broken LSPs URL: https://www.pinglabz.com/troubleshooting-mpls-ldp/ Last updated: 2026-08-01T19:26:53.000Z MPLS has a signature failure mode, and once you have seen it you never forget it: every IP test passes, the IGP is fully converged, you can ping every loopback in the core, and the customer's traffic is going nowhere. The reason is that MPLS forwarding does not run on IP reachability. It runs on *labels*, and labels come from LDP. Break LDP and the control plane looks perfect while the data plane is dead. This is a structured way to troubleshoot the label plane: LDP neighbors, label bindings, and the LSPs they build. Every capture comes from live CML labs on IOS XE 17.18.2: a five-router L3VPN core with LDP deliberately broken on one link, and a minimal three-router core, no VRFs, where the before and after is one line of output changing. If you got here by pasting `unusable: no label` into a search box: that string means the router resolved the route and the next hop, then had no label to put on the packet, so CEF refused to forward it. The transport LSP is incomplete somewhere, and the message does not say where. For L3VPN-specific route problems see [what to check when the labels are fine but the VPN routes are missing](https://www.pinglabz.com/troubleshooting-mpls-l3vpn/), or start with [how MPLS label switching works end to end](https://www.pinglabz.com/mpls/). ## Troubleshoot in Layers, Bottom Up MPLS is a stack, and each layer depends entirely on the one below. Do not start in the middle: 1\. Interface / IP`show ip interface brief` 2\. IGP`show ip ospf neighbor`, is the far loopback a /32 in the table? 3\. MPLS enabled`show mpls interfaces` 4\. LDP session`show mpls ldp neighbor` (want `State: Oper`) 5\. Label bindings`show mpls ldp bindings` 6\. Forwarding`show mpls forwarding-table`, `show ip cef` Layer 2 being healthy tells you nothing about layer 4\. That is the entire trap. Work up the stack and stop at the first layer that is wrong. ## What Healthy Looks Like You need a reference before you can spot a deviation. On PE1 in the L3VPN lab, the LDP session to the P router: ``` PE1#show mpls ldp neighbor Peer LDP Ident: 10.255.0.3:0; Local LDP Ident 10.255.0.1:0 TCP connection: 10.255.0.3.60695 - 10.255.0.1.646 State: Oper; Msgs sent/rcvd: 8/8; Downstream Up time: 00:00:34 LDP discovery sources: Ethernet0/1, Src IP addr: 10.30.30.2 Addresses bound to peer LDP Ident: 10.255.0.3 10.30.30.2 10.30.31.1 ``` Four things to read every time: - **State: Oper.** Anything else and no labels are being exchanged. There is no partial credit. - **The LDP identifiers** (10.255.0.3:0, 10.255.0.1:0). These should be loopbacks. If you see a physical interface address here, someone forgot `mpls ldp router-id Loopback0 force` and the session will reset the next time that interface flaps. - **TCP connection on port 646.** LDP sessions are TCP; discovery is UDP 646 multicast. Which side owns 646 is not random: the higher router ID opens the connection and the lower one listens, so the attempt only ever happens in one direction (which matters when you write a core ACL). If discovery works but the session never establishes, look for an ACL blocking TCP 646, or for the router-id being an address the peer cannot reach. - **Addresses bound to peer.** The peer's full interface list. Useful for confirming you are talking to the router you think you are. And `show mpls interfaces`, which tells you whether MPLS is actually on: ``` PE1#show mpls interfaces Interface IP Tunnel BGP Static Operational Ethernet0/1 Yes (ldp) No No No Yes ``` One line, and only the core-facing interface. Ethernet0/0 (the PE-CE link) is correctly absent: the customer does not speak MPLS. If you see a CE-facing interface in this output, someone put `mpls ip` where it does not belong. Remember this command, because further down it produces the most misleading output in the whole toolkit. ## Reading Label Bindings The Label Information Base (LIB) is every label your router knows about for a prefix: the one it allocated locally, plus the ones its neighbors advertised. This is the raw material the forwarding table is built from. ``` P#show mpls ldp bindings 10.255.0.2 32 lib entry: 10.255.0.2/32, rev 10 local binding: label: 17 remote binding: lsr: 10.255.0.2:0, label: imp-null remote binding: lsr: 10.255.0.1:0, label: 18 ``` Read it as a conversation. The P router says "for PE2's loopback, I will accept traffic on label 17." PE2 itself says "I am the owner of this prefix, send it to me with **imp-null** (implicit null), meaning pop the label before you send it." PE1 says "I use 18 for that prefix." `imp-null` is what drives penultimate hop popping: because PE2 advertised it, P pops the transport label instead of swapping it. That is one prefix. The shape of the label plane only shows up in the whole database. This is the LIB from R2, the middle router of the three-router lab, with both sessions up (one connected /30 entry trimmed for length): ``` R2#show mpls ldp bindings lib entry: 1.1.1.1/32, rev 8 local binding: label: 16 remote binding: lsr: 1.1.1.1:0, label: imp-null remote binding: lsr: 3.3.3.3:0, label: 16 lib entry: 2.2.2.2/32, rev 2 local binding: label: imp-null remote binding: lsr: 1.1.1.1:0, label: 16 remote binding: lsr: 3.3.3.3:0, label: 17 lib entry: 3.3.3.3/32, rev 10 local binding: label: 17 remote binding: lsr: 1.1.1.1:0, label: 18 remote binding: lsr: 3.3.3.3:0, label: imp-null lib entry: 10.0.12.0/30, rev 4 local binding: label: imp-null remote binding: lsr: 1.1.1.1:0, label: imp-null remote binding: lsr: 3.3.3.3:0, label: 18 ``` Every entry has the same shape: one local binding (what R2 tells everyone to use when sending it traffic for that prefix) and one remote binding per LSR. Three things fall out of that: - R2's local binding for its own 2.2.2.2/32 is `imp-null`, as is its local for the connected /30\. A router advertises implicit null for anything it owns, which is exactly why PHP happens. It is not a fault. - R3 advertised label 16 for 1.1.1.1/32, a prefix R3 can only reach back through R2\. R2 files it and never uses it. LDP on IOS is downstream unsolicited: everyone advertises everything to everyone. - Those five LIB entries produced two rows in the forwarding table, because the LFIB is only the subset the IGP picked. A broken LSP is therefore never a missing prefix in the LIB. It is a missing remote binding from the one LSR you needed it from. The label format itself, including what "Pop Label" does to the stack, is covered in [what is actually inside the label stack](https://www.pinglabz.com/mpls-labels-explained/), and the mechanics of [how LDP discovers neighbors and decides what to advertise](https://www.pinglabz.com/ldp-mpls-label-distribution/) are worth reading alongside this. ## The Classic: LDP Down, IGP Up Now break it. On the P router, one core-facing interface loses its MPLS configuration. The link stays up. The IP address stays. Only `mpls ip` is removed: ``` P(config)# interface Ethernet0/0 P(config-if)# no mpls ip ``` Here is what a network engineer sees on PE1 when the customer calls. Start with the tests that people instinctively run first: ``` PE1#show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 10.255.0.3 1 FULL/DR 00:00:31 10.30.30.2 Ethernet0/1 PE1#ping 10.255.0.2 source Loopback0 Sending 5, 100-byte ICMP Echos to 10.255.0.2, timeout is 2 seconds: Packet sent with a source address of 10.255.0.1 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 2/2/4 ms PE1#show ip bgp vpnv4 all summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 10.20.1.2 4 65001 10 19 53 0 0 00:05:23 2 10.255.0.2 4 65000 34 23 53 0 0 00:11:26 3 ``` OSPF adjacency FULL. Loopback-to-loopback ping across the core: 100 percent success. The MP-BGP VPNv4 session is up, established for eleven minutes, and has received three prefixes. Every green light you would normally check is green. And yet: ``` PE1#ping vrf CUST-A 10.20.20.1 source Ethernet0/0 Sending 5, 100-byte ICMP Echos to 10.20.20.1, timeout is 2 seconds: Packet sent with a source address of 10.20.1.1 ..... Success rate is 0 percent (0/5) ``` Total blackhole. Here is the layer everyone skips: ``` PE1#show mpls ldp neighbor PE1# ``` Empty. No LDP session at all. And that has one consequence that explains everything: ``` PE1#show mpls forwarding-table 10.255.0.2 Local Outgoing Prefix Bytes Label Outgoing Next Hop Label Label or Tunnel Id Switched interface 18 No Label 10.255.0.2/32 0 Et0/1 10.30.30.2 PE1#show ip cef vrf CUST-A 10.20.20.0 10.20.20.0/24 nexthop 10.30.30.2 Ethernet0/1 unusable: no label ``` **`unusable: no label`**. That string is the whole diagnosis. PE1 knows the route. It knows the BGP next hop is 10.255.0.2\. It can ping 10.255.0.2 all day. But an L3VPN packet needs a two-label stack (transport label to reach PE2, VPN label to identify the customer), and PE1 has no transport label for PE2's loopback, because the LDP session that would have supplied it is gone. CEF refuses to forward a packet it cannot label, so it drops it. Note also that `show mpls forwarding-table` did not go blank. The entry is still there, with the outgoing label replaced by `No Label`. If you skim that output looking for a missing line, you will miss it. Read the Outgoing Label column. ### Why the VPN dies but plain IP does not This confuses people, so it is worth stating explicitly. Ordinary IP traffic across the core does not need a label; if the IGP has a route, the packet forwards. But a VPN packet *must* be labelled, because that is how PE2 knows which customer it belongs to. There is no unlabelled fallback. MPLS Flags: MPLS Required, as the routing table puts it: ``` PE1#show ip route vrf CUST-A 10.20.20.0 255.255.255.0 Routing entry for 10.20.20.0/24 Known via "bgp 65000", distance 200, metric 0, type internal * 10.255.0.2 (default), from 10.255.0.2 MPLS label: 22 MPLS Flags: MPLS Required ``` So: IP works, VPN does not. Every time you see that combination, go straight to LDP. ## Walking the Path to Find Where the Label Stops In that lab the break was one hop away. In a real core it will not be. `unusable: no label` is reported by the ingress router, and all that router knows is that its own downstream neighbor never handed it a label for the remote loopback. The break can be several hops on, and every router in between looks flawless. Check the IGP first, because LDP installs labels against IGP next hops and cannot route around an IGP problem: if the loopback is missing or flapping in [the OSPF process that builds the table LDP follows](https://www.pinglabz.com/ospf/), fix that and the label plane usually fixes itself. Then, from the ingress router: 1. Identify the address the LSP has to reach. For L3VPN that is the BGP next hop, the remote PE loopback. 2. `show mpls forwarding-table `. `No Label` means the break is at or below this router. A number or `Pop Label` means the break is downstream. 3. `show mpls ldp neighbor`. Is there an `Oper` session with the IGP next hop for that /32? If not, you have found it. 4. If the session is Oper, run `show mpls ldp bindings ` and look for a remote binding from that specific LSR. Session up with no binding from it puts the problem one hop further along. 5. Move to that next hop and repeat from step 2\. The first router reading `No Label` is the one adjacent to the break. 6. There, `show mpls ldp discovery` separates "the far end is not configured" from "something is eating UDP 646." Two commands per hop, binary answer every time. Here is step 3 on the router adjacent to the break: R2 sits between R1 and R3, and R3's Ethernet0/0 is missing `mpls ip`. ``` R2#show mpls ldp neighbor Peer LDP Ident: 1.1.1.1:0; Local LDP Ident 2.2.2.2:0 TCP connection: 1.1.1.1.646 - 2.2.2.2.51092 State: Oper; Msgs sent/rcvd: 9/9; Downstream Up time: 00:00:56 LDP discovery sources: Ethernet0/0, Src IP addr: 10.0.12.1 Addresses bound to peer LDP Ident: 10.0.12.1 1.1.1.1 ``` One peer, not two. An absent neighbor is easy to miss when you are scanning for something wrong rather than something missing, so count peers against MPLS-enabled core links before reading any detail. Then comes the most misleading command in the set: ``` R2#show mpls interfaces Interface IP Tunnel BGP Static Operational Ethernet0/0 Yes (ldp) No No No Yes Ethernet0/1 Yes (ldp) No No No Yes ``` Both interfaces read `Yes (ldp)` and `Operational Yes`. Correct, and useless on its own, because `show mpls interfaces` only reports your side of the wire. R2 is sending hellos out Ethernet0/1 quite happily and nobody is answering. Plenty of engineers stop here, see two green lines, and conclude the box is healthy. It is. The neighbor is the problem. And here is the fingerprint on a router with no VRF and no BGP anywhere near it: ``` R2#show mpls forwarding-table Local Outgoing Prefix Bytes Label Outgoing Next Hop Label Label or Tunnel Id Switched interface 16 Pop Label 1.1.1.1/32 0 Et0/0 10.0.12.1 17 No Label 3.3.3.3/32 0 Et0/1 10.0.23.2 ``` Both loopbacks are in OSPF with a valid next hop and a valid outgoing interface. 1.1.1.1/32 goes out through the working session and has a label. 3.3.3.3/32 goes out through the dead one and reads `No Label`. The only difference between those two lines is a binding that never arrived, and on an ingress PE that same condition is what CEF reports as `unusable: no label`. ## The Fix, and the Line That Proves It Putting `mpls ip` back on R3's core interface is the whole fix, and the router announces it: ``` %LDP-5-NBRCHG: LDP Neighbor 3.3.3.3:0 (2) is UP ``` That message deserves a monitoring rule, because the same one going the other way is usually your only warning before a ticket: IP keeps working throughout and nothing else will fire. The neighbor table confirms it: ``` R2#show mpls ldp neighbor Peer LDP Ident: 1.1.1.1:0; Local LDP Ident 2.2.2.2:0 TCP connection: 1.1.1.1.646 - 2.2.2.2.51092 State: Oper; Msgs sent/rcvd: 11/11; Downstream Ethernet0/0, Src IP addr: 10.0.12.1 Peer LDP Ident: 3.3.3.3:0; Local LDP Ident 2.2.2.2:0 TCP connection: 3.3.3.3.31330 - 2.2.2.2.646 State: Oper; Msgs sent/rcvd: 9/9; Downstream Ethernet0/1, Src IP addr: 10.0.23.2 Addresses bound to peer LDP Ident: 10.0.23.2 3.3.3.3 ``` Two sessions, both Oper, and the TCP lines show the active-side rule in practice: R2 opened the session to R1 from port 51092, R3 opened the session to R2 from 31330\. Higher router ID connects, lower listens. Now the forwarding table: ``` R2#show mpls forwarding-table Local Outgoing Prefix Bytes Label Outgoing Next Hop Label Label or Tunnel Id Switched interface 16 Pop Label 1.1.1.1/32 0 Et0/0 10.0.12.1 17 Pop Label 3.3.3.3/32 0 Et0/1 10.0.23.2 ``` One word changed. `No Label` became `Pop Label` for 3.3.3.3/32, with no change to OSPF, to any IP address, or to any routing table in the path. The IP layer was involved in neither the failure nor the fix, which is the whole point. ## What This Was Captured On Both labs ran on CML with `iol-xe` nodes on IOS XE 17.18.2, which does full LDP including TCP 646 sessions and implicit-null bindings, so this reproduces in a small virtual lab. The three-router topology is R1 to R2 to R3 in a line, OSPF area 0 on the core links and loopbacks, loopbacks 1.1.1.1, 2.2.2.2 and 3.3.3.3 doubling as LDP router IDs, and `mpls ip` everywhere except R3's Ethernet0/0\. Those captures are all from R2, the transit router. ## The LDP Failure Checklist When `show mpls ldp neighbor` is empty or the session will not reach Oper, the cause is almost always one of these: mpls ip missing on one side LDP needs it on *both* ends of the link. Check `show mpls interfaces` on each. This is the most common cause by a wide margin. Router ID not reachable Discovery is link-local, but the TCP session is built to the router ID. If the peer's LDP ID (usually a loopback) is not in your routing table, discovery succeeds and the session never comes up. Router ID not pinned to a loopback Without `mpls ldp router-id Loopback0 force`, LDP picks the highest interface IP. When that interface flaps, the session resets and every LSP through it rebuilds. ACL blocking 646 UDP 646 for discovery, TCP 646 for the session, and only the higher-router-ID side ever initiates. An inbound ACL on a core interface that forgets any of that produces a very confusing outage. MTU mismatch Labels add 4 bytes each. A 1500-byte core with a two-label stack needs 1508\. Symptoms are worse than a clean failure: small packets work, large ones vanish. Authentication mismatch If LDP MD5 is configured on one side only, discovery works and the TCP session is refused. ## When the Session Is Up but the LSP Is Broken A subtler class of failure: LDP is Oper on every link, but traffic still fails somewhere in the middle. The tools: **Labelled traceroute** is the fastest way to see where the stack falls apart. Healthy, from CE1 across the VPN: ``` CE1#traceroute 10.20.20.1 source Ethernet0/0 1 10.20.1.1 2 msec 2 10.30.30.2 [MPLS: Labels 17/22 Exp 0] 5 msec 3 10.20.2.1 [MPLS: Label 22 Exp 0] 5 msec 4 10.20.2.2 6 msec ``` Hop 2 carries two labels (transport 17, VPN 22). Hop 3 carries one, because the P router popped the transport label (PHP). If a hop that should be labelled shows no label, the LSP breaks there. If the labels stop changing when they should be swapping, you have found your router. **MPLS ping and traceroute** test the LSP itself rather than IP: ``` PE1#ping mpls ipv4 10.255.0.2/32 PE1#traceroute mpls ipv4 10.255.0.2/32 ``` These use LSP echo (RFC 4379) and will tell you the LSP is broken even when the IP path is fine. When you have an IP-works-MPLS-does-not situation, this names the hop in one command instead of six. **Byte counters** in `show mpls forwarding-table` are underrated. Run the command twice with traffic flowing. If the Bytes Switched column is not incrementing on the entry you expect, traffic is not taking the LSP you think it is. (Every counter above reads 0 because those captures came off a lab with no traffic offered; a zero counter in production on a busy LSP is itself a finding.) ## FAQ ### Does "No Label" always mean LDP is down? No. It means this router has no usable remote binding for the prefix via its chosen next hop. A dead session is the common cause, but the same output appears when the prefix is not in the IGP, when label filtering excludes it, or when the next hop sits outside the MPLS domain. `show mpls ldp bindings` tells you which: no LIB entry at all is a different problem from a LIB entry missing one LSR's binding. ### My LDP neighbor is up but I have no labels for some prefixes. Why? LDP only allocates labels for prefixes in the IGP, and by default IOS does not allocate labels for BGP-learned prefixes. If the prefix is not in your IGP, there is no label. Check with `show mpls ldp bindings `. ### Should I filter which prefixes get labels? In a large core, yes. You only need labels for the PE loopbacks. `mpls ldp label allocate global host-routes` (or an explicit prefix list) cuts the LIB down dramatically and is standard practice in production. ### Does LDP need its own adjacency, or does it follow the IGP? Both, in a sense. LDP discovers neighbors on directly connected links with UDP hellos, then builds a TCP session to the peer's router ID. But the labels it installs in the forwarding table are chosen based on the IGP's best next hop, so LDP follows the IGP. This is also why an IGP/LDP synchronisation problem after a link comes back can blackhole traffic, and why `mpls ldp sync` exists. ### Can I run MPLS without LDP? Yes. Segment routing distributes labels in the IGP itself and removes LDP entirely, along with this whole failure mode. See [distributing labels in the IGP instead of running LDP](https://www.pinglabz.com/segment-routing-mpls/). ## Key Takeaways - Troubleshoot bottom up: interface, IGP, `mpls interfaces`, LDP session, bindings, forwarding table. - A working IGP proves nothing about MPLS. Loopback pings can succeed at 100 percent while the VPN is completely down. - `State: Oper` is the only acceptable LDP neighbor state, and `show mpls interfaces` only ever proves your own side is configured. - `show ip cef vrf ` returning **unusable: no label** is the definitive fingerprint of a broken transport LSP, but it names the symptom, not the location. - Walk the path. The first router whose forwarding table reads `No Label` for the remote loopback is the one adjacent to the break. - In the forwarding table, a broken LSP shows as `No Label` in the Outgoing Label column, not as a missing row. - `imp-null` in the bindings is what triggers penultimate hop popping. It is normal and expected. - Pin the LDP router ID to a loopback with `force`, or an interface flap will reset your sessions. Next: [Troubleshooting MPLS L3VPN](https://www.pinglabz.com/troubleshooting-mpls-l3vpn/) for the VPN-layer failures (missing RTs, missing redistribution), or back to the [full MPLS cluster guide](https://www.pinglabz.com/mpls/). ### RD vs RT: The Most Confused Concepts in MPLS L3VPN URL: https://www.pinglabz.com/mpls-rd-vs-rt-explained/ Last updated: 2026-07-12T00:28:27.000Z Ask ten network engineers to explain the difference between a route distinguisher and a route target and you will get four confident wrong answers, three "they're basically the same thing," and maybe three correct ones. The confusion is understandable: they look identical (`65000:1`), they are configured three lines apart, and in a simple any-to-any VPN they are usually set to the same value. They do completely unrelated jobs. This article proves it with a single BGP table containing the same prefix twice. If you need the surrounding configuration, see [MPLS L3VPN configuration step by step](https://www.pinglabz.com/mpls-l3vpn-configuration-step-by-step/) and the [MPLS cluster guide](https://www.pinglabz.com/mpls/). ## One Sentence Each **Route Distinguisher (RD)** Makes a prefix *unique* so two customers can both use 10.20.20.0/24 without BGP treating them as the same route. Part of the route's identity. Carried in the NLRI. **Route Target (RT)** Decides *who gets the route*. An extended community tagged on export and filtered on import. Carried as a BGP attribute. This is what builds VPN topology. Or, more bluntly: **RD is identity, RT is policy.** The RD answers "which route is this?" The RT answers "which VRFs should install it?" Everything else in this article is a consequence of those two sentences. ## The Problem the RD Solves BGP carries routes. A route is a prefix plus attributes. If two customers both use 10.20.20.0/24 (and they will, because RFC 1918 space is not that big and everyone picks 10.x), then BGP sees two updates for the same prefix, runs the best-path algorithm, picks one, and throws the other away. Customer B's route is gone. The route distinguisher fixes this by making the prefix longer. An RD is 64 bits. Glue it to the front of a 32-bit IPv4 prefix and you get a 96-bit **VPNv4** prefix: ``` IPv4 prefix : 10.20.20.0/24 RD : 65000:1 VPNv4 prefix : 65000:1:10.20.20.0/24 <-- what BGP actually carries ``` Now customer A's route is `65000:1:10.20.20.0/24` and customer B's is `65000:2:10.20.20.0/24`. Different prefixes. BGP keeps both. No best-path contest, no lost route. That is the entire job. The RD is a uniqueness prefix. It carries no meaning, controls nothing, and is not used in any import or export decision. ### Seeing it in one table This is PE1 from the lab. VRF CUST-A (RD 65000:1) has learned 10.20.20.0/24 from the remote site over MP-BGP. VRF CUST-B (RD 65000:2) has 10.20.20.0/24 as a directly connected interface. Same prefix, same router, same BGP table: ``` PE1#show ip bgp vpnv4 all Network Next Hop Metric LocPrf Weight Path Route Distinguisher: 65000:1 (default for vrf CUST-A) *> 1.1.1.1/32 10.20.1.2 0 0 65001 i *>i 2.2.2.2/32 10.255.0.2 0 100 0 65001 i *> 10.20.1.0/30 0.0.0.0 0 32768 ? *>i 10.20.2.0/30 10.255.0.2 0 100 0 ? *>i 10.20.20.0/24 10.255.0.2 0 100 0 65001 i *> 192.168.99.0 10.20.1.2 0 0 65001 i Route Distinguisher: 65000:2 (default for vrf CUST-B) *> 10.20.20.0/24 0.0.0.0 0 32768 ? *>i 10.20.99.0/24 10.255.0.2 0 100 0 ? ``` Both 10.20.20.0/24 entries are present, both marked best (`*>`). Nothing was suppressed. Without the RD, one of them would be gone. Drill into the prefix and the mechanism is explicit: ``` PE1#show ip bgp vpnv4 all 10.20.20.0/24 BGP routing table entry for 65000:1:10.20.20.0/24, version 13 Paths: (1 available, best #1, table CUST-A) 10.255.0.2 (metric 21) (via default) from 10.255.0.2 (10.255.0.2) Origin incomplete, metric 0, localpref 100, valid, internal, best Extended Community: RT:65000:1 mpls labels in/out nolabel/22 BGP routing table entry for 65000:2:10.20.20.0/24, version 2 Paths: (1 available, best #1, table CUST-B) 0.0.0.0 (via vrf CUST-B) from 0.0.0.0 (10.255.0.1) Origin incomplete, metric 0, localpref 100, weight 32768, valid, sourced, best Extended Community: RT:65000:2 mpls labels in/out 19/nolabel(CUST-B) ``` Two distinct BGP table entries, keyed by `65000:1:10.20.20.0/24` and `65000:2:10.20.20.0/24`. Each carries its own RT and its own VPN label. And the forwarding follows straight through: ``` PE1#show ip route vrf CUST-A 10.20.20.0 255.255.255.0 Routing entry for 10.20.20.0/24 Known via "bgp 65000", distance 200, metric 0, type internal * 10.255.0.2 (default), from 10.255.0.2 MPLS label: 22 MPLS Flags: MPLS Required PE1#show ip route vrf CUST-B 10.20.20.0 255.255.255.0 Routing entry for 10.20.20.0/24 Known via "connected", distance 0, metric 0 (connected, via interface) * directly connected, via Loopback11 ``` Identical destination address, one router, two entirely different forwarding decisions. In CUST-A the packet gets VPN label 22 and crosses the MPLS core to PE2\. In CUST-B it goes out a local interface. Neither customer knows the other exists. ## The Problem the RT Solves The RD made the routes unique. It did not decide who should receive them. That is the route target's job. An RT is a BGP **extended community**, an 8-byte tag attached to the route as an attribute. It works in two halves: `route-target export 65000:1`"When I send a route out of this VRF, stamp it with RT 65000:1." `route-target import 65000:1`"When a VPNv4 route arrives carrying RT 65000:1, install it in this VRF." That is the whole mechanism, and it is more powerful than it looks. Import and export lists are independent, and each can hold multiple values. The set of RTs a VRF exports and imports *is* the VPN topology: Any-to-any (full mesh) Every site exports RT 65000:1 and imports RT 65000:1\. Everyone sees everyone. Hub and spoke Spokes export RT 1 and import RT 2\. Hub exports RT 2 and imports RT 1\. Spokes reach the hub; spokes cannot reach each other. Extranet / shared services The shared-services VRF exports RT 999\. Every customer VRF adds `import 999` alongside its own RT. Everyone reaches the shared service; nobody reaches each other. Not one line of that requires touching the RD. You build hub-and-spoke, extranets, management VRFs, and selective route leaking purely by editing import and export lists. This is why the RT is the interesting one and the RD is the plumbing. ## Break the RT Import and Watch the VPN Die Theory is cheap. Here is the same lab with one line removed from PE1's CUST-A VRF: ``` PE1(config)# vrf definition CUST-A PE1(config-vrf)# address-family ipv4 PE1(config-vrf-af)# no route-target import 65000:1 ``` The export is still there. The RD is unchanged. The BGP session is untouched. Watch what happens: ``` PE1#show vrf detail CUST-A | include Export|Import|VRF VRF CUST-A (VRF Id = 1); default RD 65000:1; default VPNID Export VPN route-target communities No Import VPN route-target communities VRF label distribution protocol: not configured PE1#show ip bgp vpnv4 all summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 10.20.1.2 4 65001 6 12 43 0 0 00:01:33 2 10.255.0.2 4 65000 24 19 43 0 0 00:07:35 1 PE1#show ip bgp vpnv4 rd 65000:1 10.20.20.0/24 BGP routing table entry for 65000:1:10.20.20.0/24, version 42 Paths: (0 available, no best path) Not advertised to any peer ``` Three things to read there. `No Import VPN route-target communities` is the smoking gun. The prefix count from the remote PE has collapsed from 4 to 1, because the RT filter is now rejecting almost everything at the door. And the prefix entry still exists in the table but with `0 available, no best path`: BGP knows about the route, it just has no VRF willing to accept it. From the customer's point of view, the VPN is simply broken: ``` j@llmbits:~$ ping -c 3 10.20.20.1 PING 10.20.20.1 (10.20.20.1) 56(84) bytes of data. From 192.168.99.1 icmp_seq=1 Destination Host Unreachable From 192.168.99.1 icmp_seq=2 Destination Host Unreachable From 192.168.99.1 icmp_seq=3 Destination Host Unreachable --- 10.20.20.1 ping statistics --- 3 packets transmitted, 0 received, +3 errors, 100% packet loss ``` MP-BGP is up. LDP is up. The IGP is up. The routes are being advertised. And the customer cannot pass a packet, because one import statement is missing. This is the single most common L3VPN misconfiguration, and the reason `show vrf detail` should be the second command you run when a VPN is broken. There is a full methodology in [Troubleshooting MPLS L3VPN: where VPN routes go missing](https://www.pinglabz.com/troubleshooting-mpls-l3vpn/). ## A Related Trap: RTs Need Extended Communities Enabled Because RTs are extended communities, they only survive the trip between PEs if the BGP session is configured to send them. Drop `send-community extended` from PE2's VPNv4 neighbor statement and the routes still get advertised, but with the RTs stripped off: ``` PE1#show ip bgp vpnv4 all summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 10.255.0.2 4 65000 44 25 56 0 0 00:12:24 0 ``` Zero prefixes received. Not because the session is down (it is up, and messages are flowing) but because every arriving route has no RT and therefore matches no import list. Same symptom as a missing import, different cause. Both live in the RT half of the world; neither has anything to do with the RD. ## If They Do Different Jobs, Why Are They Usually the Same Number? Because in a plain any-to-any VPN, one value works for both and nobody has a reason to complicate it. `rd 65000:1`, `route-target both 65000:1`, done. The moment you build a hub-and-spoke or an extranet, they diverge, and the fact that they were never related becomes obvious. You will have one RD per VRF (often one per VRF *per PE*, which is a deliberate design for load-balancing and CE multihoming, since two different RDs on the same prefix let a route reflector keep both paths instead of picking one) and a whole set of RTs describing the topology. The habit worth building: read `rd` as "identity" and `route-target` as "policy," even when the numbers match. ## FAQ ### Is the RD carried in BGP? Yes, but as part of the prefix itself (the NLRI), not as an attribute. That is precisely why it makes the route unique. The RT is carried separately, as an extended community attribute. ### Can two VRFs on the same PE use the same RD? Technically yes if the prefixes never overlap, but do not. The RD's whole purpose is disambiguation; reusing it removes the safety net for zero benefit. ### Can one VRF import multiple RTs? Yes, and this is how extranets and shared services work. A VRF can import as many RTs as you like, and export more than one too. ### Does the RD affect which VRF a route lands in? No. That is the single most common misconception. Import is decided entirely by the RT. A route with the "wrong" RD but a matching RT will be imported; a route with a matching RD but no matching RT will not. ### What format should I use? ASN:number (65000:1) is by far the most common. IP:number (10.255.0.1:1) is also valid and is sometimes used so the RD identifies the originating PE at a glance. ## Key Takeaways - **RD is identity.** It prepends 64 bits to the IPv4 prefix, making a 96-bit VPNv4 route so overlapping customer address space does not collide in BGP. - **RT is policy.** It is an extended community, stamped on export and filtered on import, and it alone decides which VRFs install a route. - The RD plays no part in import decisions. Ever. - VPN topology (any-to-any, hub-and-spoke, extranet) is built entirely from RT import and export lists. - A missing `route-target import` kills the VPN while every other check looks healthy. So does a missing `send-community extended`, because RTs are extended communities. - They are usually the same number in simple VPNs. That is a coincidence of convenience, not a relationship. Next: [build the whole thing from scratch](https://www.pinglabz.com/mpls-l3vpn-configuration-step-by-step/), or read the [MPLS cluster guide](https://www.pinglabz.com/mpls/). ### PE-CE Routing Protocols in MPLS L3VPN: Static, OSPF, EIGRP, and BGP URL: https://www.pinglabz.com/mpls-l3vpn-pe-ce-routing/ Last updated: 2026-07-12T00:28:26.000Z Once the MPLS core is built and MP-BGP is carrying VPNv4 between your PEs, one question remains: how does a PE router actually learn the customer's routes? That is the PE-CE routing protocol, and it is the one part of an L3VPN the customer has an opinion about. You have four realistic choices: static routes, OSPF, EIGRP, or eBGP. They are not interchangeable. Each changes what the customer sees in their routing table, how quickly the VPN reconverges, and how much of the provider's design leaks into the customer's network. This article runs three of them back to back on the same lab and shows exactly what changes. If you have not built the VPN yet, start with [MPLS L3VPN configuration step by step](https://www.pinglabz.com/mpls-l3vpn-configuration-step-by-step/), and see the [MPLS cluster guide](https://www.pinglabz.com/mpls/) for the wider picture. ## The Lab Five IOS XE routers: CE1 - PE1 - P - PE2 - CE2\. Provider AS 65000, VRF CUST-A with RD and RT 65000:1, OSPF plus LDP in the core, MP-iBGP VPNv4 between the PEs. Customer site 1 is 192.168.99.0/24 (a real Debian host); site 2 is 10.20.20.0/24 on CE2\. The PE-CE links are 10.20.1.0/30 and 10.20.2.0/30. The core never changes across the three runs. Only the PE-CE protocol does. ## The One Pattern That Applies to All Four Whatever protocol you pick, the PE does the same two things: 1. Learn the customer's prefixes into the VRF routing table (via static, OSPF, EIGRP, or eBGP). 2. Get those prefixes *into MP-BGP*, under `address-family ipv4 vrf CUST-A`, so they can be tagged with an RT and a VPN label and shipped to the far PE. And in the other direction, take the VPNv4 routes that arrive from the far PE and hand them back to the customer in whatever protocol the customer speaks. Step 2 is where people get burned. A route sitting in the VRF routing table is **not** in MP-BGP. It gets there only via `redistribute` (static, connected, OSPF, EIGRP) or automatically, if the PE-CE protocol *is* BGP. Miss it and the route exists on one PE and nowhere else, with no error message anywhere. That failure has its own walkthrough in [Troubleshooting MPLS L3VPN](https://www.pinglabz.com/troubleshooting-mpls-l3vpn/). ## Option 1: Static Routes The simplest option, and more common in production than the certification material suggests. Small sites, a single subnet, no need for the customer to run a routing protocol at all. ``` PE1(config)# ip route vrf CUST-A 192.168.99.0 255.255.255.0 10.20.1.2 PE1(config)# ip route vrf CUST-A 1.1.1.1 255.255.255.255 10.20.1.2 PE1(config)# router bgp 65000 PE1(config-router)# address-family ipv4 vrf CUST-A PE1(config-router-af)# redistribute connected PE1(config-router-af)# redistribute static ``` Two commands you cannot forget: the `vrf` keyword on the static route (a plain `ip route` lands in the global table and does nothing for the customer), and `redistribute static` under the VRF address family. The CE gets a default route and nothing else: ``` CE1(config)# ip route 0.0.0.0 0.0.0.0 10.20.1.1 ``` The result on PE1, with static in both directions: ``` PE1#show ip route vrf CUST-A Routing Table: CUST-A 1.0.0.0/32 is subnetted, 1 subnets S 1.1.1.1 [1/0] via 10.20.1.2 2.0.0.0/32 is subnetted, 1 subnets B 2.2.2.2 [200/0] via 10.255.0.2, 00:00:11 10.0.0.0/8 is variably subnetted, 4 subnets, 3 masks C 10.20.1.0/30 is directly connected, Ethernet0/0 L 10.20.1.1/32 is directly connected, Ethernet0/0 B 10.20.2.0/30 [200/0] via 10.255.0.2, 00:00:11 B 10.20.20.0/24 [200/0] via 10.255.0.2, 00:00:11 S 192.168.99.0/24 [1/0] via 10.20.1.2 ``` Local site routes are `S`, far site routes are `B` with a next hop of the remote PE loopback. Clean and predictable. The catch is that it is static. If the customer adds a subnet, someone raises a change request against the provider. If the CE link fails, the static route stays in the table until the interface goes down, and if the CE is reachable over a switch it may not. It scales to about "one subnet per site" before it becomes a burden. ## Option 2: OSPF The most common enterprise choice, because the customer is usually already running OSPF and does not want to change. The PE runs an OSPF process *inside* the VRF and redistributes both ways. ``` PE1(config)# router ospf 10 vrf CUST-A PE1(config-router)# router-id 10.255.0.1 PE1(config-router)# domain-id 0.0.0.1 PE1(config-router)# network 10.20.1.1 0.0.0.0 area 0 PE1(config-router)# redistribute bgp 65000 subnets PE1(config)# router bgp 65000 PE1(config-router)# address-family ipv4 vrf CUST-A PE1(config-router-af)# redistribute ospf 10 ``` Note the two redistributions, in opposite directions. `redistribute bgp 65000` under the OSPF process takes routes learned from the far site over MP-BGP and hands them to the CE. `redistribute ospf 10` under the BGP VRF address family takes the CE's routes and puts them into MP-BGP. Configure one and not the other and traffic works in exactly one direction, which is a memorable afternoon. The CE runs plain OSPF, with no idea a VPN is involved: ``` CE1(config)# router ospf 10 CE1(config-router)# network 10.20.1.2 0.0.0.0 area 0 CE1(config-router)# network 192.168.99.1 0.0.0.0 area 0 ``` Adjacency comes up between CE and PE like any other OSPF neighbor: ``` PE1#show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 10.255.0.3 1 FULL/DR 00:00:38 10.30.30.2 Ethernet0/1 1.1.1.1 1 FULL/BDR 00:00:39 10.20.1.2 Ethernet0/0 ``` Two adjacencies on PE1: one in the global table to the P router (the core IGP), one inside VRF CUST-A to CE1\. Same router, same protocol, completely separate topology databases. That is worth pausing on. ### The interesting part: routes come back as O IA Here is CE1's routing table once both ends are up: ``` CE1#show ip route ospf 2.0.0.0/32 is subnetted, 1 subnets O IA 2.2.2.2 [110/21] via 10.20.1.1, 00:00:52, Ethernet0/1 10.0.0.0/8 is variably subnetted, 4 subnets, 3 masks O IA 10.20.2.0/30 [110/11] via 10.20.1.1, 00:01:18, Ethernet0/1 O IA 10.20.20.0/24 [110/21] via 10.20.1.1, 00:00:25, Ethernet0/1 ``` Not `O E2`, which is what you would expect from redistribution. `O IA`, inter-area. The MPLS backbone is presenting itself to the customer as an OSPF **superbackbone**: a virtual area 0 that sits above the customer's own areas. Routes that cross the VPN arrive looking like they came from another area of the customer's own OSPF domain, not like external redistributed routes. This is deliberate, and it is controlled by the `domain-id`. When both PEs use the same domain ID, MP-BGP carries the OSPF route type in an extended community and the far PE regenerates the route as inter-area rather than external. Set different domain IDs on the two PEs and the same routes reappear as `O E2` externals with a metric type of 2\. Customers usually prefer the inter-area behaviour, because external routes lose to internal ones and can create surprising path selection at sites with a backdoor link. ### The sham-link problem Which leads to the classic OSPF-over-L3VPN gotcha. If the customer has a backdoor link between two sites (say a cheap DSL failover circuit) *and* runs OSPF over the VPN, the backdoor link is an intra-area OSPF path while the VPN is inter-area. Intra-area always wins. The customer's traffic quietly abandons the expensive MPLS circuit and pours down the backup DSL link. The fix is an **OSPF sham-link**: a logical intra-area link between the two PEs inside the VRF, which makes the VPN path look intra-area too and lets normal cost comparison decide. If your customer has backdoor links and OSPF, you need sham-links; if they do not, you can happily ignore them. ## Option 3: EIGRP Same structural pattern as OSPF, Cisco-only, and worth knowing because plenty of enterprise networks still run it. See the [EIGRP cluster guide](https://www.pinglabz.com/eigrp/) for the protocol itself. ``` PE1(config)# router eigrp 100 PE1(config-router)# address-family ipv4 vrf CUST-A autonomous-system 100 PE1(config-router-af)# network 10.20.1.1 0.0.0.0 PE1(config-router-af)# redistribute bgp 65000 metric 100000 100 255 1 1500 PE1(config)# router bgp 65000 PE1(config-router)# address-family ipv4 vrf CUST-A PE1(config-router-af)# redistribute eigrp 100 ``` Two things differ from OSPF. First, EIGRP redistribution *requires* a metric (bandwidth, delay, reliability, load, MTU) or the routes are simply not redistributed, with no complaint. Second, EIGRP carries its original metric components across the VPN in BGP extended communities, so as long as the AS number matches on both PEs, the routes arrive at the far CE as **internal** EIGRP (D) rather than external (D EX). Mismatch the AS number and everything becomes D EX with an administrative distance of 170, which loses to almost everything else. It is the EIGRP equivalent of the OSPF domain-id trap. ## Option 4: eBGP The service provider's preferred option, and the only one with no redistribution at all. The customer gets their own AS number (or a private one) and peers eBGP with the PE inside the VRF. ``` PE1(config)# router bgp 65000 PE1(config-router)# address-family ipv4 vrf CUST-A PE1(config-router-af)# neighbor 10.20.1.2 remote-as 65001 PE1(config-router-af)# neighbor 10.20.1.2 activate PE1(config-router-af)# neighbor 10.20.1.2 as-override ``` The CE is a plain eBGP speaker: ``` CE1(config)# router bgp 65001 CE1(config-router)# neighbor 10.20.1.1 remote-as 65000 CE1(config-router)# network 192.168.99.0 mask 255.255.255.0 CE1(config-router)# network 1.1.1.1 mask 255.255.255.255 ``` No redistribution anywhere. Routes learned from the CE over eBGP go straight into the VRF's BGP table and are exported as VPNv4 automatically. This is the cleanest option and the reason providers push it. It also gives the customer real policy control, using the same attributes covered in the [BGP cluster guide](https://www.pinglabz.com/bgp/). The session and the routes on PE1: ``` PE1#show ip bgp vpnv4 all summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 10.20.1.2 4 65001 5 10 40 0 0 00:00:37 2 10.255.0.2 4 65000 19 17 40 0 0 00:06:40 4 PE1#show ip bgp vpnv4 all Network Next Hop Metric LocPrf Weight Path Route Distinguisher: 65000:1 (default for vrf CUST-A) *> 1.1.1.1/32 10.20.1.2 0 0 65001 i *>i 2.2.2.2/32 10.255.0.2 0 100 0 65001 i *> 10.20.1.0/30 0.0.0.0 0 32768 ? *>i 10.20.2.0/30 10.255.0.2 0 100 0 ? *>i 10.20.20.0/24 10.255.0.2 0 100 0 65001 i *> 192.168.99.0 10.20.1.2 0 0 65001 i ``` Two neighbors on one router: the eBGP session to the customer (AS 65001) and the iBGP VPNv4 session to PE2\. The customer's routes carry AS path `65001 i`. ### The AS-path problem, and as-override Here is the wrinkle that makes eBGP PE-CE more subtle than it looks. Most customers use the *same* AS number at every site, because they were told to and because it is simpler. But BGP has a built-in loop prevention rule: if a router sees its own AS in the AS path of an incoming update, it drops the route. So CE1 (AS 65001) advertises 192.168.99.0/24 with path `65001`. It crosses the VPN. PE2 sends it to CE2 with path `65000 65001`. CE2 is also AS 65001, sees its own AS in the path, and discards the route. The VPN is up, MP-BGP is healthy, and the two sites cannot talk. `neighbor x.x.x.x as-override` is the fix. The PE rewrites the customer's AS number in the path with its own before advertising to the CE. The result on CE2: ``` CE2#show ip bgp Network Next Hop Metric LocPrf Weight Path *> 1.1.1.1/32 10.20.2.1 0 65000 65000 i *> 2.2.2.2/32 0.0.0.0 0 32768 i *> 10.20.1.0/30 10.20.2.1 0 65000 ? *> 10.20.20.0/24 0.0.0.0 0 32768 i *> 192.168.99.0 10.20.2.1 0 65000 65000 i ``` Look at the path on 192.168.99.0: `65000 65000`. The original 65001 has been overwritten with the provider's AS. CE2 no longer sees itself in the path, accepts the route, and the sites reach each other: ``` CE2#ping 192.168.99.100 source Loopback1 Sending 5, 100-byte ICMP Echos to 192.168.99.100, timeout is 2 seconds: Packet sent with a source address of 10.20.20.1 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 5/5/7 ms ``` (The alternative, `allowas-in`, is configured on the CE instead of the PE and tells the customer router to accept its own AS in the path. It works, but it requires touching every CE, and providers do not like depending on customer configuration for basic reachability. Use as-override on the PE.) ## Choosing Between Them Static Redistributionredistribute static ReconvergeSlow / manual Best for Single-subnet sites, tight change control OSPF RedistributionBoth directions ReconvergeFast Watch out Superbackbone, domain-id, sham-links with backdoor links EIGRP RedistributionBoth, metric required ReconvergeFast Watch out AS must match on both PEs or routes become D EX (AD 170) eBGP RedistributionNone needed ReconvergeModerate Watch out Same customer AS at every site needs as-override The short version: providers want eBGP because it needs no redistribution and gives clean policy control. Enterprises want OSPF because they already run it. Static survives at small sites because it is unbreakable and needs no protocol knowledge. EIGRP shows up where the customer is a long-standing Cisco shop. ## FAQ ### Can different sites in the same VPN use different PE-CE protocols? Yes. The VRF and MP-BGP do not care. Site A can run eBGP, site B OSPF, site C static. Everything meets in the VRF's BGP table. It is common during migrations. ### Why does the customer see O IA instead of O E2? Because the MPLS backbone acts as an OSPF superbackbone and, when the PEs share a domain-id, the original route type is carried across MP-BGP and regenerated. This is the intended behaviour. ### Do I need a sham-link? Only if the customer runs OSPF over the VPN *and* has a backdoor link between sites. Without a backdoor link, there is no intra-area path to lose to. ### as-override or allowas-in? as-override, on the PE. It solves the problem without requiring configuration on every customer router. ## Key Takeaways - Every PE-CE option needs the customer's routes to end up in MP-BGP under `address-family ipv4 vrf`. With static, OSPF, and EIGRP that means an explicit `redistribute`. With eBGP it is automatic. - Redistribution runs in **both** directions. One-way redistribution gives you one-way traffic. - OSPF over L3VPN presents the backbone as a superbackbone, so routes arrive as `O IA`. The `domain-id` controls whether they look inter-area or external. - EIGRP redistribution into BGP silently does nothing without a metric, and the AS number must match on both PEs or routes downgrade to D EX. - eBGP is the cleanest option but needs `as-override` whenever the customer uses the same AS at multiple sites. The AS path becomes `65000 65000` once it is on. Continue with [the full L3VPN configuration walkthrough](https://www.pinglabz.com/mpls-l3vpn-configuration-step-by-step/), [RD vs RT explained](https://www.pinglabz.com/mpls-rd-vs-rt-explained/), or the [MPLS cluster guide](https://www.pinglabz.com/mpls/). ### MPLS L3VPN Configuration Step by Step: VRF, RD, RT, and MP-BGP URL: https://www.pinglabz.com/mpls-l3vpn-configuration-step-by-step/ Last updated: 2026-07-12T00:28:26.000Z MPLS L3VPN is the service that made MPLS a service-provider standard, and it is still what most enterprises are actually buying when their carrier sells them "an MPLS circuit." The concept is simple to state and surprisingly fiddly to configure: many customers, each with their own private (and possibly overlapping) IP addressing, carried across one shared provider backbone, with each customer seeing only their own routes. This article is the configuration walkthrough. Every command and every piece of output below comes from a five-router Cisco IOS XE lab (CE1 - PE1 - P - PE2 - CE2) running OSPF as the core IGP, LDP for label distribution, and MP-iBGP carrying VPNv4 between the two PE routers. A real Debian host sits behind CE1 and pings across the VPN at the end. If you want the wider context first, start with the [MPLS cluster guide](https://www.pinglabz.com/mpls/). ## The Four Moving Parts Before touching a keyboard, get the pieces straight. An L3VPN is built from four things stacked on top of each other, and if you configure them out of order you will spend an hour chasing a symptom caused by a missing layer underneath. 1\. Core IGP Job Gives every P and PE router loopback-to-loopback reachability TypicalOSPF or IS-IS 2\. LDP Job Assigns a transport label to every IGP prefix, building the LSP Runs onPE and P routers 3\. VRF + RD + RT Job Separates customers on the PE; makes overlapping prefixes unique and controls who imports what Runs onPE routers only 4\. MP-BGP VPNv4 Job Carries customer prefixes plus their RD, RT, and VPN label between PEs Runs onPE routers only Notice what is missing from that list: the P router in the middle of the core runs items 1 and 2 and nothing else. It has no VRFs, no BGP, and no idea that any customer exists. We will prove that with real output later, because it is the single most important architectural fact about L3VPN and the thing most people get wrong when they first draw the diagram. ## The Lab Five IOS XE routers in a line, plus a real Linux host bridged in behind CE1 so the final ping is not simulated: ``` VM 192.168.99.100 | [ CE1 ] --- [ PE1 ] --- [ P ] --- [ PE2 ] --- [ CE2 ] 10.20.1.2/30 .1/30 10.20.2.2/30 10.30.30.1/24 .2 .1 10.30.31.2/24 Loopbacks: PE1 10.255.0.1 P 10.255.0.3 PE2 10.255.0.2 Customer site 1 LAN: 192.168.99.0/24 (the real VM) Customer site 2 LAN: 10.20.20.0/24 (CE2 Loopback1) ``` Provider AS is 65000\. Customer VRF is CUST-A with route distinguisher 65000:1 and route target 65000:1 imported and exported. A second VRF, CUST-B, gets RD/RT 65000:2 and deliberately overlapping addressing, which we use at the end to prove the separation is real. ## Step 1: Core IGP and Loopbacks Everything hangs off the PE loopbacks, so configure them first and get them into the IGP. The loopback is the LDP router ID, the BGP update source, and the BGP next hop that the far-end PE resolves against. If the loopback is not in the IGP with a /32, nothing else will work. ``` PE1(config)# interface Loopback0 PE1(config-if)# ip address 10.255.0.1 255.255.255.255 PE1(config-if)# interface Ethernet0/1 PE1(config-if)# description Core link to P PE1(config-if)# ip address 10.30.30.1 255.255.255.0 PE1(config-if)# router ospf 1 PE1(config-router)# router-id 10.255.0.1 PE1(config-router)# network 10.255.0.1 0.0.0.0 area 0 PE1(config-router)# network 10.30.30.0 0.0.0.255 area 0 ``` P and PE2 get the same treatment. The P router carries both core links: ``` P(config)# router ospf 1 P(config-router)# router-id 10.255.0.3 P(config-router)# network 10.255.0.3 0.0.0.0 area 0 P(config-router)# network 10.30.30.0 0.0.0.255 area 0 P(config-router)# network 10.30.31.0 0.0.0.255 area 0 ``` Verify before moving on. On a broadcast Ethernet segment the adjacency sits at 2WAY/DROTHER for a few tens of seconds during DR election, so do not panic if the first look is not FULL: ``` PE1#show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 10.255.0.3 1 FULL/DR 00:00:33 10.30.30.2 Ethernet0/1 ``` If you need the OSPF side of this in more depth, the [OSPF cluster guide](https://www.pinglabz.com/ospf/) covers adjacency states and area design. ## Step 2: LDP Across the Core The IGP gives you IP reachability. LDP turns that into a label switched path. Two lines globally, one line per core-facing interface: ``` PE1(config)# mpls label protocol ldp PE1(config)# mpls ldp router-id Loopback0 force PE1(config)# interface Ethernet0/1 PE1(config-if)# mpls ip ``` The `mpls ldp router-id Loopback0 force` line is not optional in practice. Without it, LDP picks the highest interface address as its ID, and when that interface goes down the LDP session resets. Pin it to the loopback that the IGP already advertises. `mpls ip` goes only on core-facing interfaces. Never on the interface facing a CE router, because the customer does not speak MPLS. Verify: ``` PE1#show mpls ldp neighbor Peer LDP Ident: 10.255.0.3:0; Local LDP Ident 10.255.0.1:0 TCP connection: 10.255.0.3.60695 - 10.255.0.1.646 State: Oper; Msgs sent/rcvd: 8/8; Downstream Up time: 00:00:34 LDP discovery sources: Ethernet0/1, Src IP addr: 10.30.30.2 Addresses bound to peer LDP Ident: 10.255.0.3 10.30.30.2 10.30.31.1 PE1#show mpls interfaces Interface IP Tunnel BGP Static Operational Ethernet0/1 Yes (ldp) No No No Yes ``` `State: Oper` is what you want. Anything else and the LSP is not built. The [LDP deep-dive](https://www.pinglabz.com/ldp-mpls-label-distribution/) covers discovery, session establishment, and label advertisement modes. Once LDP is up, the P router has already built the transport LSPs to both PE loopbacks: ``` P#show mpls forwarding-table Local Outgoing Prefix Bytes Label Outgoing Next Hop Label Label or Tunnel Id Switched interface 16 Pop Label 10.255.0.1/32 4146 Et0/0 10.30.30.1 17 Pop Label 10.255.0.2/32 3665 Et0/1 10.30.31.2 ``` "Pop Label" is penultimate hop popping. P is one hop away from each PE, so instead of swapping the transport label it strips it and hands the egress PE a packet with only the VPN label left. That saves the egress PE a lookup and is the default behaviour, not something you configure. ## Step 3: Define the VRF, the RD, and the RTs Now the customer. On both PEs: ``` PE1(config)# vrf definition CUST-A PE1(config-vrf)# rd 65000:1 PE1(config-vrf)# address-family ipv4 PE1(config-vrf-af)# route-target export 65000:1 PE1(config-vrf-af)# route-target import 65000:1 PE1(config-vrf-af)# exit-address-family ``` Three values, three different jobs, and confusing them is the number-one L3VPN mistake: VRF name (CUST-A)Local to this router. Cosmetic. Nothing is carried in BGP. RD (65000:1)Prepended to the IPv4 prefix to make a globally unique 96-bit VPNv4 route. Does not control import. RT (65000:1)An extended community tagged onto the route. Export stamps it; import filters on it. This is what builds the VPN topology. If that distinction is still fuzzy, read [RD vs RT: the most confused concepts in MPLS L3VPN](https://www.pinglabz.com/mpls-rd-vs-rt-explained/), which proves the difference with two overlapping prefixes in one BGP table. Now put the customer-facing interface into the VRF. Do this *before* configuring the IP address, because moving an interface into a VRF wipes its IP configuration: ``` PE1(config)# interface Ethernet0/0 PE1(config-if)# description To CE1 - VRF CUST-A PE1(config-if)# vrf forwarding CUST-A PE1(config-if)# ip address 10.20.1.1 255.255.255.252 ``` Check it landed: ``` PE1#show vrf Name Default RD Protocols Interfaces CUST-A 65000:1 ipv4 Et0/0 CUST-B 65000:2 ipv4 Lo11 ``` This is exactly the same VRF construct used in [VRF-Lite](https://www.pinglabz.com/vrf-lite-configuration-cisco-ios-xe/). The difference is that VRF-Lite stops here and hand-carries the separation hop by hop, while L3VPN hands the VRF's routes to MP-BGP and lets the backbone do the work. ## Step 4: MP-BGP VPNv4 Between the PEs The PEs peer with each other over their loopbacks, inside the provider AS, and activate the VPNv4 address family: ``` PE1(config)# router bgp 65000 PE1(config-router)# no bgp default ipv4-unicast PE1(config-router)# neighbor 10.255.0.2 remote-as 65000 PE1(config-router)# neighbor 10.255.0.2 update-source Loopback0 PE1(config-router)# address-family vpnv4 PE1(config-router-af)# neighbor 10.255.0.2 activate PE1(config-router-af)# neighbor 10.255.0.2 send-community extended PE1(config-router-af)# exit-address-family ``` Four lines matter here and every one of them is a classic outage: update-source Loopback0The BGP next hop must be the loopback that LDP has a label for. Peer off a physical interface and the LSP resolution breaks. address-family vpnv4Without it the session comes up and carries plain IPv4\. VPN routes go nowhere. send-community extendedRTs are extended communities. Omit this and the routes arrive with no RT, get filtered on import, and silently vanish. no bgp default ipv4-unicastKeeps the global table out of it. Optional, but it stops you leaking provider routes into the customer's world by accident. The MP-BGP details, including how the VPNv4 address family relates to the IPv4 and IPv6 ones, are covered in [MP-BGP: Multiprotocol BGP address families explained](https://www.pinglabz.com/mp-bgp-multiprotocol-bgp/). For BGP session mechanics generally, see the [BGP cluster guide](https://www.pinglabz.com/bgp/). Confirm the session: ``` PE1#show ip bgp vpnv4 all summary BGP router identifier 10.255.0.1, local AS number 65000 Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 10.255.0.2 4 65000 19 17 40 0 0 00:06:40 4 ``` ## Step 5: PE-CE Routing The PE now needs to learn the customer's prefixes and give the customer the far-end's. You have four options: static, OSPF, EIGRP, or eBGP. All four are covered in detail with lab output in [PE-CE routing protocols in MPLS L3VPN](https://www.pinglabz.com/mpls-l3vpn-pe-ce-routing/). Here is the simplest one to get you to a working VPN, static: ``` PE1(config)# ip route vrf CUST-A 192.168.99.0 255.255.255.0 10.20.1.2 PE1(config)# ip route vrf CUST-A 1.1.1.1 255.255.255.255 10.20.1.2 PE1(config)# router bgp 65000 PE1(config-router)# address-family ipv4 vrf CUST-A PE1(config-router-af)# redistribute connected PE1(config-router-af)# redistribute static ``` Two things are happening. The `ip route vrf` lines put the customer's prefixes into the CUST-A table. The `redistribute` lines under the VRF address family push them into MP-BGP, where they become VPNv4 routes tagged with RT 65000:1 and given a VPN label. Forgetting the redistribution is the classic "I can see the route on the PE but the far side never gets it" failure, and it is walked through step by step in [Troubleshooting MPLS L3VPN](https://www.pinglabz.com/troubleshooting-mpls-l3vpn/). The CE side is boringly simple. The CE is not VRF-aware and does not run MPLS; as far as it knows, PE1 is just a router: ``` CE1(config)# ip route 0.0.0.0 0.0.0.0 10.20.1.1 ``` ## Verifying the VPN Start at the VRF routing table on PE1\. The customer's local routes are connected or static; the far-end site arrives as BGP: ``` PE1#show ip route vrf CUST-A Routing Table: CUST-A 1.0.0.0/32 is subnetted, 1 subnets S 1.1.1.1 [1/0] via 10.20.1.2 2.0.0.0/32 is subnetted, 1 subnets B 2.2.2.2 [200/0] via 10.255.0.2, 00:00:11 10.0.0.0/8 is variably subnetted, 4 subnets, 3 masks C 10.20.1.0/30 is directly connected, Ethernet0/0 L 10.20.1.1/32 is directly connected, Ethernet0/0 B 10.20.2.0/30 [200/0] via 10.255.0.2, 00:00:11 B 10.20.20.0/24 [200/0] via 10.255.0.2, 00:00:11 S 192.168.99.0/24 [1/0] via 10.20.1.2 ``` Note the next hop on the BGP routes: 10.255.0.2, PE2's loopback. Not a physical interface, not a P router. The VRF's forwarding decision is "hand this to PE2 over the LSP." Now the VPNv4 table, which is where the RD and the VPN label live: ``` PE1#show ip bgp vpnv4 vrf CUST-A 10.20.20.0 BGP routing table entry for 65000:1:10.20.20.0/24, version 13 Paths: (1 available, best #1, table CUST-A) Refresh Epoch 1 Local 10.255.0.2 (metric 21) (via default) from 10.255.0.2 (10.255.0.2) Origin incomplete, metric 0, localpref 100, valid, internal, best Extended Community: RT:65000:1 mpls labels in/out nolabel/22 ``` Read that carefully, because it contains the whole protocol. The prefix is `65000:1:10.20.20.0/24`, RD glued to the front. The extended community is `RT:65000:1`, which is why the CUST-A VRF imported it. And the outgoing MPLS label is 22, the VPN label PE2 allocated for this prefix. PE1 will push 22 as the inner label on every packet destined for that subnet. The label forwarding state on PE1: ``` PE1#show mpls forwarding-table Local Outgoing Prefix Bytes Label Outgoing Next Hop Label Label or Tunnel Id Switched interface 16 Pop Label 10.255.0.3/32 0 Et0/1 10.30.30.2 17 Pop Label 10.30.31.0/24 0 Et0/1 10.30.30.2 18 17 10.255.0.2/32 0 Et0/1 10.30.30.2 19 No Label 10.20.20.0/24[V] 500 aggregate/CUST-B 20 No Label 10.20.1.0/30[V] 1140 aggregate/CUST-A 21 No Label 1.1.1.1/32[V] 0 Et0/0 10.20.1.2 22 No Label 192.168.99.0/24[V] \ 570 Et0/0 10.20.1.2 ``` Entry 18 is the transport LSP toward PE2: swap local label 18 for outgoing label 17 and send it to P. The `[V]` entries are the VPN labels PE1 has handed *out* for its own customer prefixes, telling PE2 which VRF and which next hop to use for traffic coming back the other way. The [label format article](https://www.pinglabz.com/mpls-labels-explained/) breaks down the two-label stack byte by byte. ## The P Router Is BGP-Free (and That Is the Whole Point) Here is the capture that should reframe how you think about L3VPN. This is the P router, sitting in the middle of the path, forwarding customer traffic all day: ``` P#show ip bgp summary % BGP not active P#show ip route 10.20.20.0 % Subnet not in table P#show mpls forwarding-table Local Outgoing Prefix Bytes Label Outgoing Next Hop Label Label or Tunnel Id Switched interface 16 Pop Label 10.255.0.1/32 4146 Et0/0 10.30.30.1 17 Pop Label 10.255.0.2/32 3665 Et0/1 10.30.31.2 ``` The P router does not run BGP. It has never heard of 10.20.20.0/24\. It has no VRFs. Its entire forwarding table is two labels pointing at two PE loopbacks. Customer traffic crosses it as a labelled packet and it swaps or pops the outer label without ever looking at the IP header. That is why L3VPN scales. Add a thousand customers with a million routes and the P routers do not grow at all; only the PEs carry VPN state. It is also why "the core is BGP-free" is a design goal you will hear repeated in every service provider architecture conversation. ## End to End: A Real Host Across the VPN A real Debian machine sits on the 192.168.99.0/24 LAN behind CE1\. It has no idea MPLS exists. It pings the far site: ``` j@llmbits:~$ ping -c 4 10.20.20.1 PING 10.20.20.1 (10.20.20.1) 56(84) bytes of data. 64 bytes from 10.20.20.1: icmp_seq=1 ttl=251 time=7.16 ms 64 bytes from 10.20.20.1: icmp_seq=2 ttl=251 time=7.02 ms 64 bytes from 10.20.20.1: icmp_seq=3 ttl=251 time=7.07 ms 64 bytes from 10.20.20.1: icmp_seq=4 ttl=251 time=6.99 ms --- 10.20.20.1 ping statistics --- 4 packets transmitted, 4 received, 0% packet loss, time 3004ms ``` And the traceroute from CE1, which is where the label stack finally becomes visible: ``` CE1#traceroute 10.20.20.1 source Ethernet0/0 Type escape sequence to abort. Tracing the route to 10.20.20.1 1 10.20.1.1 2 msec 2 10.30.30.2 [MPLS: Labels 17/22 Exp 0] 5 msec 3 10.20.2.1 [MPLS: Label 22 Exp 0] 5 msec 4 10.20.2.2 6 msec ``` Hop 2 is the P router, and the packet arrives carrying **two** labels: 17 (transport, "get me to PE2") and 22 (VPN, "this belongs to CUST-A and goes out toward CE2"). Hop 3 is PE2, and the packet now has only one label, because P performed penultimate hop popping and stripped the transport label. PE2 pops the VPN label, does an IP lookup in the CUST-A table, and forwards to CE2. That single traceroute is the entire L3VPN forwarding plane on one screen. ## Proving the Separation: A Second VRF with Overlapping Addressing The claim is that two customers can use the same IP space. Let us actually test it rather than take it on faith. VRF CUST-B gets RD/RT 65000:2, and on PE1 it is given a network that deliberately collides with CUST-A's site-2 LAN: ``` PE1(config)# vrf definition CUST-B PE1(config-vrf)# rd 65000:2 PE1(config-vrf)# address-family ipv4 PE1(config-vrf-af)# route-target export 65000:2 PE1(config-vrf-af)# route-target import 65000:2 PE1(config)# interface Loopback11 PE1(config-if)# vrf forwarding CUST-B PE1(config-if)# ip address 10.20.20.1 255.255.255.0 ``` 10.20.20.0/24 now exists twice on the same router. Here is the same prefix in the same BGP table, twice, under two different RDs: ``` PE1#show ip bgp vpnv4 all Network Next Hop Metric LocPrf Weight Path Route Distinguisher: 65000:1 (default for vrf CUST-A) *> 1.1.1.1/32 10.20.1.2 0 0 65001 i *>i 2.2.2.2/32 10.255.0.2 0 100 0 65001 i *> 10.20.1.0/30 0.0.0.0 0 32768 ? *>i 10.20.2.0/30 10.255.0.2 0 100 0 ? *>i 10.20.20.0/24 10.255.0.2 0 100 0 65001 i *> 192.168.99.0 10.20.1.2 0 0 65001 i Route Distinguisher: 65000:2 (default for vrf CUST-B) *> 10.20.20.0/24 0.0.0.0 0 32768 ? *>i 10.20.99.0/24 10.255.0.2 0 100 0 ? ``` Both entries for 10.20.20.0/24 coexist because to BGP they are different routes: 65000:1:10.20.20.0/24 and 65000:2:10.20.20.0/24\. That is the entire job of the route distinguisher. And the forwarding follows: ``` PE1#show ip route vrf CUST-A 10.20.20.0 255.255.255.0 Routing entry for 10.20.20.0/24 Known via "bgp 65000", distance 200, metric 0, type internal * 10.255.0.2 (default), from 10.255.0.2 MPLS label: 22 MPLS Flags: MPLS Required PE1#show ip route vrf CUST-B 10.20.20.0 255.255.255.0 Routing entry for 10.20.20.0/24 Known via "connected", distance 0, metric 0 (connected, via interface) * directly connected, via Loopback11 ``` Same prefix, same router, two completely different forwarding decisions. In CUST-A it goes across the MPLS core to PE2 with VPN label 22\. In CUST-B it is a local interface. Neither customer can see the other's route, and no configuration effort was required to keep them apart beyond giving them different RDs and RTs. ## The Complete PE Configuration Everything above, on PE1, condensed: ``` vrf definition CUST-A rd 65000:1 address-family ipv4 route-target export 65000:1 route-target import 65000:1 ! mpls label protocol ldp mpls ldp router-id Loopback0 force ! interface Loopback0 ip address 10.255.0.1 255.255.255.255 ! interface Ethernet0/0 vrf forwarding CUST-A ip address 10.20.1.1 255.255.255.252 ! interface Ethernet0/1 ip address 10.30.30.1 255.255.255.0 mpls ip ! router ospf 1 router-id 10.255.0.1 network 10.255.0.1 0.0.0.0 area 0 network 10.30.30.0 0.0.0.255 area 0 ! router bgp 65000 no bgp default ipv4-unicast neighbor 10.255.0.2 remote-as 65000 neighbor 10.255.0.2 update-source Loopback0 address-family vpnv4 neighbor 10.255.0.2 activate neighbor 10.255.0.2 send-community extended exit-address-family address-family ipv4 vrf CUST-A redistribute connected redistribute static exit-address-family ``` The P router config is a quarter of that length and contains no customer state at all. That asymmetry is the design. ## FAQ ### Do I need MPLS on the PE-CE link? No, and you should not put it there. The CE is a plain IP router. `mpls ip` belongs only on core-facing interfaces between PE and P routers. ### Can the RD and RT be the same value? Yes, and in a simple any-to-any VPN they usually are (65000:1 for both, as in this lab). They are still doing completely different jobs. In hub-and-spoke or extranet designs they diverge immediately. ### Why iBGP between the PEs and not eBGP? Both PEs are inside the same provider AS, so the VPNv4 session is iBGP. In a real network you would not full-mesh every PE; you would point them all at a pair of route reflectors configured for the VPNv4 address family. ### What happens if the core IGP is fine but LDP is down? The VPN dies while every IP-level test passes. It is the single most confusing MPLS failure mode, and it has its own article: [Troubleshooting MPLS: LDP neighbors, label bindings, and broken LSPs](https://www.pinglabz.com/troubleshooting-mpls-ldp/). ## Key Takeaways - Build the layers in order: IGP, then LDP, then VRF/RD/RT, then MP-BGP VPNv4, then PE-CE routing. Verify each before moving up. - The PE loopback is load-bearing. It is the LDP router ID, the BGP update source, and the next hop that must resolve to a label. - RD makes prefixes unique; RT decides who imports them. Same syntax, unrelated jobs. - `send-community extended` is mandatory. Without it the RTs are stripped and every VPN route is silently discarded on import. - Customer prefixes only reach MP-BGP if you redistribute them under `address-family ipv4 vrf`. This is the most common "route is on the PE but not in BGP" cause. - The P router runs no BGP and holds no customer routes. Two labels in its forwarding table carry every customer. - A labelled traceroute (`Labels 17/22` then `Label 22`) shows the transport and VPN labels, and PHP, in one command. Next in the cluster: [the four PE-CE routing options compared](https://www.pinglabz.com/mpls-l3vpn-pe-ce-routing/), and the [MPLS cluster guide](https://www.pinglabz.com/mpls/) for everything else. ### Troubleshooting DMVPN: NHRP, Tunnel State, and Spoke-to-Spoke Failures URL: https://www.pinglabz.com/troubleshooting-dmvpn/ Last updated: 2026-07-11T19:17:51.000Z DMVPN failures cluster hard. After the underlay itself, almost every broken cloud comes down to a spoke that cannot register, a tunnel that is up/up while silently eating packets, or shortcuts that never form. So instead of writing another symptom checklist, we broke a healthy Phase 3 lab (IOS XE 17.18, hub and two spokes) three different ways on purpose and captured what each failure actually looks like on the CLI, including the case where the hub logs nothing at all. This is the troubleshooting companion to the [DMVPN complete guide](https://www.pinglabz.com/dmvpn/); the healthy-state baseline it compares against is the [Phase 3 configuration walkthrough](https://www.pinglabz.com/dmvpn-phase-3-configuration/). ## The Layered Method DMVPN stacks four layers, and each has its own verification command. Test in this order and you cannot get lost: the underlay (can tunnel sources ping each other?), then NHRP (`show dmvpn` and `show ip nhrp nhs detail`), then the routing protocol (`show ip eigrp neighbors` or OSPF equivalent), then IPsec if configured (`show crypto session`, covered in [the IPsec article](https://www.pinglabz.com/dmvpn-ipsec-profiles-ikev2/)). The single most information-dense field in all of it is the State column of `show dmvpn`: UP means NHRP resolved and working, NHRP means the control plane is failing right now, and IKE/CRYPTO states mean the problem is a layer down in IPsec. A peer in state NHRP with the tunnel interface up/up is the signature of every failure below. ## Baseline: What Healthy Looks Like Burn this into memory, because every diagnosis is a diff against it. Hub with two registered spokes: ``` HUB# show dmvpn # Ent Peer NBMA Addr Peer Tunnel Add State UpDn Tm Attrb ----- --------------- --------------- ----- -------- ----- 1 198.51.100.1 10.0.0.2 UP 00:01:35 D 1 192.0.2.1 10.0.0.3 UP 00:01:31 D ``` Spoke with a healthy NHS (state RE, requests and replies balanced): ``` SPOKE1# show ip nhrp nhs detail 10.0.0.1 RE priority = 0 cluster = 0 req-sent 1 req-failed 0 repl-recv 1 (00:00:23 ago) ``` ## Failure 1: Wrong NHS Address (Registering Into the Void) We pointed SPOKE2 at an NHS that does not exist (10.0.0.9 instead of 10.0.0.1), the kind of typo that ships in a copy-pasted branch template. The symptom set: ``` SPOKE2# show dmvpn # Ent Peer NBMA Addr Peer Tunnel Add State UpDn Tm Attrb ----- --------------- --------------- ----- -------- ----- 1 203.0.113.1 10.0.0.9 NHRP 00:00:07 S SPOKE2# show ip nhrp nhs detail 10.0.0.9 E priority = 0 cluster = 0 req-sent 5 req-failed 0 repl-recv 0 Pending Registration Requests: Registration Request: Reqid 6, Ret 16 NHS 10.0.0.9 expired (Tu0) SPOKE2# show interface Tunnel0 | include line protocol Tunnel0 is up, line protocol is up ``` Read the tells: state NHRP not UP, NHS stuck in E (expecting replies), `req-sent` climbing with `repl-recv` frozen at zero, a pending registration marked expired, and the interface happily up/up the whole time (GRE interfaces do not go down for control plane failures; see [GRE keepalives](https://www.pinglabz.com/gre-tunnel-keepalives/) for why). The registration packets are arriving at the hub, but they are addressed to an NHS identity nobody owns, so nothing answers. Fix the `ip nhrp nhs` and `ip nhrp map` pair so both reference the hub's real tunnel IP and its real underlay address. ## Failure 2: NHRP Authentication Mismatch (the Hub Tells You) Next we set SPOKE2's authentication string to WRONGKEY while the hub expects PLZ123\. From the spoke, the symptoms are identical to failure 1: state NHRP, requests unanswered, registrations expiring. This is the diagnostic trap of NHRP failures, since three different root causes present the same on the spoke side. The difference is on the hub. With `debug nhrp error`: ``` *Jul 11 18:52:12.182: NHRP: Receive Registration Request via Tunnel0 vrf: global(0x0), packet size: 108 *Jul 11 18:52:12.182: %DMVPN-3-DMVPN_NHRP_ERROR: Tunnel0: Recieved wrong authentication string for Registration Request , Reason: authentication failure (11) on (Tunnel: 10.0.0.3 NBMA: 203.0.113.1) ``` The hub receives the registration, checks the cleartext auth string, and rejects it with a logged reason (yes, "Recieved" is misspelled in IOS XE, which at least makes it easy to grep). This is why the second move in any registration failure, after reading the spoke's NHS counters, is to look at the hub: an answering hub that refuses is loud, and its log line names the misbehaving spoke's NBMA address for you. Align `ip nhrp authentication` across the cloud (8 characters maximum) and registration recovers on the next retry. ## Failure 3: Tunnel Key Mismatch (Nobody Tells You) Finally, the cruel one. We set `tunnel key 999` on SPOKE2 while everyone else runs `tunnel key 100`: ``` SPOKE2# show interface Tunnel0 | include line protocol|Key Tunnel0 is up, line protocol is up Key 0x3E7, sequencing disabled SPOKE2# ping 10.0.0.1 repeat 3 ... Success rate is 0 percent (0/3) SPOKE2# show ip nhrp nhs detail 10.0.0.1 E priority = 0 cluster = 0 req-sent 7 req-failed 0 repl-recv 0 (00:01:48 ago) Pending Registration Requests: Registration Request: Reqid 11, Ret 64 NHS 10.0.0.1 expired (Tu0) ``` Tunnel up/up, plain overlay pings dead, registrations expiring, and here is the defining feature: the hub logs nothing. Zero. The GRE key lives in the GRE header, so key-mismatched packets are discarded by the GRE decapsulation code before NHRP ever sees them; there is no NHRP error to log because, as far as the hub's NHRP process knows, the spoke never spoke. When a spoke shows failure-1 symptoms but the hub shows no incoming registrations and no errors while other spokes work fine, diff the tunnel stanzas: `tunnel key`, `tunnel source`, and mode. The `Key 0x3E7` line (999 in hex) in the interface output is your receipt. The same silent-drop signature comes from a wrong `tunnel source`, transport ACLs eating GRE (protocol 47), or a middlebox stripping it, which is also where [general GRE troubleshooting](https://www.pinglabz.com/gre-tunnel-troubleshooting/) overlaps this article. ## When Registration Works but Shortcuts Do Not A distinct family: hub state UP everywhere, spoke-to-spoke traffic flows, but traceroutes never leave the hub path. That is not broken, it is Phase 1 behavior in what you thought was a Phase 3 cloud. Check the pair of commands that define Phase 3: `ip nhrp redirect` present on the hub tunnel, `ip nhrp shortcut` present on the spoke tunnels. Then confirm the machinery with counters instead of guesswork: `show ip nhrp traffic` on the spoke should show a received Traffic Indication (the redirect) and sent Resolution Requests, and success lands in the RIB as the `%`/`[NHO]` override: ``` SPOKE1# show ip route next-hop-override D % 10.0.2.0/24 [90/102400640] via 10.0.0.1, 00:00:57, Tunnel0 [NHO][90/255] via 10.0.0.3, 00:00:09, Tunnel0 ``` If redirects are sent but shortcuts never install, the usual culprits are spoke-to-spoke underlay reachability (both spokes must reach each other's public addresses directly, which NAT loves to break) or a missing `shortcut` on just one side. The full message flow, including what each counter means, is in the [NHRP deep dive](https://www.pinglabz.com/nhrp-deep-dive/), and the phase mechanics are in [Phase 1 vs 2 vs 3](https://www.pinglabz.com/dmvpn-phase-1-2-3-differences/). ## Routing Adjacency Failures Wearing a DMVPN Mask One more disguise worth naming: tunnels UP, NHRP clean, but no EIGRP or OSPF neighbors. Before touching the routing protocol, check `ip nhrp map multicast` on the spoke (missing means unicast works and multicast hellos vanish) and the hub's multicast handling. Then check the usual suspects: EIGRP split horizon on the hub (spokes neighbor fine but never see each other's routes), OSPF network type mismatches, and MTU mismatch stalling OSPF in EXSTART. The design-side prevention for all of these is in [routing over DMVPN](https://www.pinglabz.com/routing-over-dmvpn-eigrp-ospf/). ## When IPsec Is in the Stack Encrypted clouds add one layer to triage, and it slots in cleanly between the underlay and NHRP: if `show dmvpn` shows IKE-related states or the tunnel line protocol refuses to come up with `tunnel protection` configured, work the crypto ladder before touching NHRP. `show crypto ikev2 sa` answers whether negotiation completed (READY) or never started (no SA at all usually means UDP 500 is filtered or the peer address is wrong); `show crypto ipsec sa` counters answer whether traffic encrypts and decrypts symmetrically (encaps climbing while the peer's decaps stalls means ESP is dying in transit, typically a firewall); and `show crypto session` ties it together per tunnel. Mismatched proposals fail loudly in IKEv2 with a clear negotiation error, while PSK mismatches fail at IKE\_AUTH. The one DMVPN-specific crypto trap: with `tunnel protection` applied on one side but not the other, GRE from the unprotected side arrives cleartext, gets rejected, and the symptom mimics failure 3's silence. The full crypto build and its verification ladder live in [securing DMVPN with IPsec profiles](https://www.pinglabz.com/dmvpn-ipsec-profiles-ikev2/). ## Capturing Debugs That Actually Land Somewhere A practical note from producing this article's captures, because it will bite anyone using automation or a jump host. On IOS XE, debug output goes to the console line by default, and most automation frameworks (PyATS included) disable console logging the moment they connect, so `debug nhrp packet` appears to produce nothing. The fix is to send debugs to the buffer and read them back: ``` SPOKE1(config)# logging buffered 200000 debugging SPOKE1# debug nhrp packet SPOKE1# ... trigger the event ... SPOKE1# show logging | include NHRP SPOKE1# undebug all ``` The second gotcha: NHRP only sends packets when it has a reason. Clearing the cache with `clear ip nhrp` does not necessarily force an immediate re-registration, so a debug window can legitimately capture silence. Bouncing the tunnel interface (or waiting out the registration timer) guarantees a registration exchange to observe. Always `undebug all` when done; `debug nhrp packet` on a busy hub is a self-inflicted denial of service. ## The Compressed Runbook Underlay pings between tunnel sources. `show dmvpn` state column on both ends. If state NHRP: `show ip nhrp nhs detail` on the spoke (req-sent versus repl-recv), then hub logs with `debug nhrp error`. Hub logging auth failures means strings mismatch; hub silent means GRE never arrives, so diff tunnel key, source, mode, and check transport filtering. If state UP but no shortcuts: verify redirect/shortcut pair, then `show ip nhrp traffic` for Traffic Indications, then spoke-to-spoke underlay reachability. If tunnels and NHRP are clean but routes are missing: multicast maps, split horizon, network types. Escalate to `debug nhrp packet` only once you know which conversation you are watching for, and remember debugs land in the buffer, not your SSH session, unless you ask. ## Key Takeaways Troubleshoot DMVPN as four layers and let `show dmvpn`'s state column route you: NHRP state means registration or resolution, and the spoke-side symptoms are identical for wrong NHS, bad authentication, and key mismatch, so the differential diagnosis happens on the hub. An answering hub that logs authentication failures is a config diff away from fixed; a silent hub means GRE itself is dropping, with tunnel key mismatch as the classic cause since the interface stays up/up through all of it. Shortcut problems are almost always a missing redirect/shortcut command or spokes that cannot reach each other's underlay addresses. Keep the healthy baseline from the [Phase 3 build](https://www.pinglabz.com/dmvpn-phase-3-configuration/) handy, and the rest of the cluster in the [DMVPN complete guide](https://www.pinglabz.com/dmvpn/). ### Securing DMVPN with IPsec Profiles (IKEv2) URL: https://www.pinglabz.com/dmvpn-ipsec-profiles-ikev2/ Last updated: 2026-07-11T19:17:51.000Z Everything DMVPN sends is cleartext. GRE is an envelope with the contents printed on the outside, and a DMVPN cloud over the internet without encryption is a packet capture away from a very bad week. The fix is IPsec applied as a tunnel protection profile: IKEv2 negotiates keys, ESP encrypts every GRE packet, and the DMVPN machinery (NHRP, routing, shortcuts) runs through it untouched and unaware. This article builds the whole stack on two Catalyst 8000v routers running IOS XE 17.18, with every show command captured live. It assumes the working Phase 3 configuration from [the config walkthrough](https://www.pinglabz.com/dmvpn-phase-3-configuration/) and is part of the [DMVPN complete guide](https://www.pinglabz.com/dmvpn/); for the general theory of stacking GRE and IPsec, see [GRE over IPsec](https://www.pinglabz.com/gre-over-ipsec/). ## The Design: Transport Mode ESP Around GRE DMVPN encryption is GRE over IPsec in transport mode. The GRE packet (which already carries the tunnel semantics, the multipoint behavior, and the NHRP messages) gets its payload encrypted by ESP between the same two underlay addresses. Transport mode is preferred over tunnel mode because the crypto endpoints and the GRE endpoints are identical, so tunnel mode's extra IP header would carry redundant information at the cost of 20 bytes per packet. The proxy identity becomes protocol 47 (GRE) host-to-host, which you will see verbatim in the SA output below. Because encryption binds to the interface, not to peers, one profile on the hub's mGRE interface covers every current and future spoke, and (in a full Phase 3 mesh) spoke-to-spoke shortcut tunnels negotiate their own SAs on demand using the same profile. Zero per-peer crypto configuration, same as the rest of DMVPN. ## Step 1: IKEv2 Proposal and Policy IKEv2 replaces the old `crypto isakmp` world (if your reference config says isakmp anywhere, it is an IKEv1 design and a decade stale). The proposal picks the SA cryptography; the policy activates the proposal: ``` crypto ikev2 proposal PL-IKEV2-PROP encryption aes-cbc-256 integrity sha256 group 19 ! crypto ikev2 policy PL-IKEV2-POLICY proposal PL-IKEV2-PROP ``` AES-256, SHA-256, and elliptic-curve group 19 are a solid 2026 baseline. One IOS XE gotcha worth knowing: if you pick an AEAD cipher like `aes-gcm-256` for the proposal, you must configure a `prf` instead of `integrity` (GCM authenticates itself), and the parser reminds you mid-configuration. ## Step 2: Keyring and Profile The keyring stores the pre-shared keys; the profile matches peers and ties authentication together. The hub cannot know spoke addresses in advance, so its keyring matches any peer: ``` ! HUB crypto ikev2 keyring PL-KEYRING peer SPOKES address 0.0.0.0 0.0.0.0 pre-shared-key PingLabz-DMVPN-Key ! crypto ikev2 profile PL-IKEV2-PROFILE match identity remote address 0.0.0.0 authentication remote pre-share authentication local pre-share keyring local PL-KEYRING ``` Spokes can be tighter, keying specifically for the hub (`address 203.0.113.1` in the peer block). A wildcard PSK is the lab-friendly choice and the production compromise everyone makes at small scale; at real scale the answer is certificates (swap `pre-share` for `rsa-sig` and a PKI trustpoint) so a stolen branch router does not compromise the whole cloud's key. ## Step 3: Transform Set and IPsec Profile ``` crypto ipsec transform-set PL-TSET esp-gcm 256 mode transport ! crypto ipsec profile PL-DMVPN-IPSEC set transform-set PL-TSET set ikev2-profile PL-IKEV2-PROFILE ``` ESP-GCM-256 is authenticated encryption in one pass, the modern default for data plane crypto. The IPsec profile is just a binding object: this transform set, keyed by this IKEv2 profile. ## Step 4: One Line on the Tunnel ``` interface Tunnel0 tunnel protection ipsec profile PL-DMVPN-IPSEC ``` Same line on hub and spokes. Nothing else in the DMVPN configuration changes, which you can prove to yourself by diffing the tunnel interface before and after: NHRP maps, network-id, redirect and shortcut, tunnel key, all identical. ## Verification: The Three-Command Ladder First rung: is there an IKEv2 SA? On the hub, seconds after the spoke's first registration attempt: ``` HUB# show crypto ikev2 sa IPv4 Crypto IKEv2 SA Tunnel-id Local Remote fvrf/ivrf Status 1 203.0.113.1/500 203.0.113.2/500 none/none READY Encr: AES-CBC, keysize: 256, PRF: SHA256, Hash: SHA256, DH Grp:19, Auth sign: PSK, Auth verify: PSK Life/Active Time: 86400/7 sec Local spi: DE52FFEAD3BDC29C Remote spi: 5C69D52BF021BA11 ``` READY means IKE\_SA\_INIT and IKE\_AUTH both completed: negotiation, key exchange, and PSK authentication all succeeded. Second rung: are child SAs actually encrypting? The counters answer, and the proxy identity shows the transport-mode GRE design in the flesh: ``` SPOKE1# show crypto ipsec sa interface: Tunnel0 local ident (addr/mask/prot/port): (203.0.113.2/255.255.255.255/47/0) remote ident (addr/mask/prot/port): (203.0.113.1/255.255.255.255/47/0) #pkts encaps: 26, #pkts encrypt: 26, #pkts digest: 26 #pkts decaps: 26, #pkts decrypt: 26, #pkts verify: 26 current outbound spi: 0xAC32FE02(2889022978) inbound esp sas: spi: 0x58377064(1480028260) transform: esp-gcm 256 , in use settings ={Transport, } Status: ACTIVE(ACTIVE) outbound esp sas: spi: 0xAC32FE02(2889022978) transform: esp-gcm 256 , in use settings ={Transport, } Status: ACTIVE(ACTIVE) ``` Protocol 47 host-to-host, transport mode, GCM, and matching encrypt/decrypt counters climbing together. If encaps grows while the far side's decaps does not, something between the routers is eating ESP (usually a firewall). Third rung: the one-screen summary that ties the crypto session to the tunnel: ``` SPOKE1# show crypto session Interface: Tunnel0 Profile: PL-IKEV2-PROFILE Session status: UP-ACTIVE Peer: 203.0.113.1 port 500 IKEv2 SA: local 203.0.113.2/500 remote 203.0.113.1/500 Active IPSEC FLOW: permit 47 host 203.0.113.2 host 203.0.113.1 Active SAs: 2, origin: crypto map ``` And the DMVPN-native view folds crypto into the peer table, confirming the overlay itself is protected (`Protect "PL-DMVPN-IPSEC"`) with live enc/dec counters per session: ``` HUB# show dmvpn detail Interface Tunnel0 is up/up, Addr. is 10.0.0.1, VRF "global" Tunnel Src./Dest. addr: 203.0.113.1/Multipoint, Tunnel VRF "global" Protocol/Transport: "multi-GRE/IP", Protect "PL-DMVPN-IPSEC" Type:Hub, Total NBMA Peers (v4/v6): 1 Crypto Session Status: UP-ACTIVE IPSEC FLOW: permit 47 host 203.0.113.1 host 203.0.113.2 Inbound: #pkts dec'ed 30 drop 0 life (KB/Sec) 4607996/3549 Outbound: #pkts enc'ed 30 drop 0 life (KB/Sec) 4607997/3549 ``` Meanwhile EIGRP formed its adjacency straight through the encryption, and a 20-packet ping across the overlay ran 100 percent, which is the operational point: the overlay does not know it is encrypted. ## If Your Reference Config Says "isakmp", Translate It A huge fraction of DMVPN material online still shows IKEv1, so it is worth mapping the old objects to the new ones. `crypto isakmp policy` becomes the IKEv2 proposal plus policy pair. `crypto isakmp key ... address` becomes a keyring peer block. The implicit IKEv1 peer matching becomes an explicit IKEv2 profile with `match identity`. The transform set survives with the same name and syntax, and the IPsec profile gains a `set ikev2-profile` line. Beyond syntax, IKEv2 buys real things: built-in NAT traversal and DPD, fewer round trips to establish, EAP support, anti-DoS cookies for hubs facing the open internet, and cleaner rekeying. There is no defensible reason to deploy new DMVPN on IKEv1 in 2026, and mixed clouds during migration work fine because IKE version is negotiated per peer pair. ## MTU: The Crypto Tax on Every Packet Encryption raises the encapsulation overhead you already budgeted for GRE. Transport-mode ESP with GCM adds roughly 30 to 40 bytes on top of GRE's 24 (SPI, sequence number, IV, ICV, padding), and the arithmetic explains the numbers this cluster standardizes on: 1500 minus GRE minus ESP lands comfortably above 1400, so `ip mtu 1400` guarantees the encrypted result fits any sane transport, and `ip tcp adjust-mss 1360` keeps TCP flows from ever producing a packet that needs fragmenting (1400 minus 40 bytes of IP+TCP headers). The failure mode when you skip this is ugly precisely because it is partial: pings work, small requests work, and bulk transfers hang when a DF-marked full-size segment meets the tunnel and the ICMP unreachable gets filtered somewhere. The full diagnosis walkthrough, including how to prove it with size-swept pings, is in [GRE MTU and fragmentation](https://www.pinglabz.com/gre-tunnel-mtu/). Also note what the SA output already told us: `plaintext mtu 1466, path mtu 1500`. IOS XE computes the usable plaintext size per SA, and comparing that number against your `ip mtu` is a quick sanity check that your budget holds on the actual negotiated cipher. ## Lab Platform Notes (Learned the Hard Way) Two things will save you an afternoon. First, IOL/IOL-XE images have no IPsec data plane, so DMVPN crypto labs need a real IOS XE image like Catalyst 8000v; the pure mGRE/NHRP work stays on the lightweight nodes. Second, a fresh Catalyst 8000v boots with no license level, and in that state the entire crypto subsystem is absent: `crypto ikev2` config lines are rejected at boot and even `show crypto ikev2 sa` returns invalid input, which looks bafflingly like a corrupted image. The fix is `license boot level network-advantage addon dna-advantage`, write, reload, then apply the crypto configuration. Neither of these is a DMVPN problem, but both wear a DMVPN costume when you hit them. ## Hardening Beyond the Profile The profile encrypts; a deployable edge needs a little more. Reasonable additions: IKEv2 DPD (`crypto ikev2 dpd 30 5 on-demand`) so dead peers tear down instead of blackholing until SA expiry, an underlay ACL permitting only UDP/500, UDP/4500, and ESP from expected sources, NAT-T awareness for spokes behind NAT (IKEv2 detects it and floats to 4500 automatically; verify with `show crypto ikev2 sa detail`), certificates instead of wildcard PSKs at scale, and call admission control on hubs that terminate hundreds of SAs. NHRP authentication and tunnel keys remain useful misconfiguration fences underneath all of this, as covered in the [NHRP deep dive](https://www.pinglabz.com/nhrp-deep-dive/). ## Key Takeaways DMVPN encryption is five small objects chained together: IKEv2 proposal, policy, keyring, profile, then an IPsec transform set and profile, applied with a single `tunnel protection` line per tunnel interface. Transport-mode ESP-GCM around GRE is the standard design, visible in the SA as `permit 47 host A host B` with `{Transport}` settings. Verify up the ladder (IKEv2 SA READY, IPsec counters climbing on both sides, crypto session UP-ACTIVE) before blaming DMVPN, and remember the overlay is deliberately ignorant of the crypto: registration, shortcuts, and routing behave identically encrypted or not. When the tunnel misbehaves for non-crypto reasons, continue to [troubleshooting DMVPN](https://www.pinglabz.com/troubleshooting-dmvpn/), and find the whole cluster in the [DMVPN complete guide](https://www.pinglabz.com/dmvpn/). ### Routing Over DMVPN: EIGRP and OSPF Design Choices URL: https://www.pinglabz.com/routing-over-dmvpn-eigrp-ospf/ Last updated: 2026-07-11T19:17:51.000Z NHRP builds the tunnels; the routing protocol decides what flows through them, and DMVPN is unusually opinionated about how you run one. A hub that relays routes between spokes on a single multipoint interface trips over split horizon, next-hop handling, and OSPF's network-type machinery in ways a normal LAN never does. This article runs EIGRP named mode and OSPF point-to-multipoint over the same live Phase 3 lab (IOS XE 17.18, one hub, two spokes) and compares what lands in the RIB. It builds on the [Phase 3 configuration walkthrough](https://www.pinglabz.com/dmvpn-phase-3-configuration/) and slots into the [DMVPN complete guide](https://www.pinglabz.com/dmvpn/); deeper protocol background lives in the [EIGRP](https://www.pinglabz.com/eigrp/) and [OSPF](https://www.pinglabz.com/ospf/) pillars. ## Why the Hub Is a Special Case On an ordinary router, routes learned on one interface get advertised out other interfaces, and loop-prevention rules like split horizon never bite. A DMVPN hub breaks that assumption: routes from SPOKE1 must be re-advertised out the same Tunnel0 they arrived on to reach SPOKE2\. Distance-vector protocols explicitly forbid that (split horizon), and link-state protocols route around it with topology abstractions that were designed for broadcast LANs, not NBMA clouds. Every DMVPN routing design is a workaround for this one geometry problem, and the workarounds differ enough between protocols to change your architecture. ## EIGRP: Two Knobs, Both on the Hub EIGRP over DMVPN needs exactly two decisions, both under the hub's tunnel af-interface. First, split horizon must go, or spoke routes never reach other spokes: ``` router eigrp PINGLABZ address-family ipv4 unicast autonomous-system 100 af-interface Tunnel0 no split-horizon exit-af-interface ``` Second, next-hop-self, and this is the knob that encodes your DMVPN phase. With next-hop-self (the default), the hub rewrites itself as next hop for relayed routes. In our Phase 3 lab, SPOKE1 sees: ``` SPOKE1# show ip route eigrp D 10.0.2.0/24 [90/102400640] via 10.0.0.1, 00:00:29, Tunnel0 D 10.0.100.0/24 [90/76800640] via 10.0.0.1, 00:00:29, Tunnel0 ``` With `no next-hop-self` (the Phase 2 design), the hub preserves the originating spoke's address, and the same route points spoke-to-spoke, captured from the same lab during its Phase 2 interlude: ``` SPOKE1# show ip route eigrp D 10.0.2.0/24 [90/102400640] via 10.0.0.3, 00:00:15, Tunnel0 ``` Phase 3 wants the first behavior. Let the RIB point at the hub, let NHRP shortcuts override it per destination (the `%`/`[NHO]` mechanics are in [the phase comparison](https://www.pinglabz.com/dmvpn-phase-1-2-3-differences/)), and reserve `no next-hop-self` for legacy Phase 2 clouds you have not migrated yet. The payoff for keeping next-hop-self is summarization. A Phase 3 hub can collapse every spoke LAN into one advertisement (`summary-address 10.0.0.0 255.0.0.0` under the af-interface) or even distribute just a default route, shrinking every spoke's table to a handful of entries regardless of how many sites exist. EIGRP also brings practical NBMA niceties: named mode's wide metrics handle the tunnel's low configured bandwidth sanely, hellos are unicast-replicated over the NHRP multicast maps automatically, and stub configuration on spokes (`eigrp stub connected summary`) protects the cloud from a branch router becoming transit during a hub flap. Named-mode syntax details are in the [EIGRP guide](https://www.pinglabz.com/eigrp/). ## OSPF: Fighting the Network Type OSPF cannot simply relay routes with a rewritten next hop; its view of the world is a topology graph built per network type, so the choice of network type on Tunnel0 is the entire design. Two are defensible, and we ran the better one live. Configuration on every router: ``` interface Tunnel0 ip ospf network point-to-multipoint router ospf 1 network 10.0.0.0 0.0.0.255 area 0 ``` Point-to-multipoint treats the NBMA cloud as a bundle of point-to-point links: no DR election, unicast-style adjacencies hub-to-spoke, 30-second hellos, and host routes for every tunnel endpoint. The hub's view: ``` HUB# show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 10.0.2.1 0 FULL/ - 00:01:39 10.0.0.3 Tunnel0 10.0.1.1 0 FULL/ - 00:01:31 10.0.0.2 Tunnel0 HUB# show ip ospf interface Tunnel0 Process ID 1, Router ID 10.0.100.1, Network Type POINT_TO_MULTIPOINT, Cost: 1000 Timer intervals configured, Hello 30, Dead 120, Wait 120, Retransmit 5 ``` Note FULL with no DR/BDR column (the dash), and the interface cost of 1000 from the tunnel's default bandwidth. The spoke RIB shows the p2mp signature, /32 host routes for every endpoint with everything via the hub: ``` SPOKE1# show ip route ospf O 10.0.0.1/32 [110/1000] via 10.0.0.1, 00:04:52, Tunnel0 O % 10.0.0.3/32 [110/2000] via 10.0.0.1, 00:04:44, Tunnel0 O 10.0.2.0/24 [110/2001] via 10.0.0.1, 00:04:44, Tunnel0 O 10.0.100.0/24 [110/1001] via 10.0.0.1, 00:04:52, Tunnel0 ``` Next hops at the hub are exactly what Phase 3 wants, and NHRP overrides them per flow; even the 10.0.0.3/32 host route carries the `%` marker from an active shortcut, and the traceroute went direct while OSPF pointed at the hub: ``` SPOKE1# traceroute 10.0.2.1 source Loopback1 numeric 1 10.0.0.3 2 msec * 2 msec ``` The alternative is network type broadcast: elect the hub DR (set spoke priority 0), keep 10-second hellos, no host routes. It works, but it is fragile by construction: a DR election on an NBMA cloud where spokes cannot see each other is a standing invitation for adjacency weirdness, and nothing stops a misconfigured spoke from trying to become DR. Avoid point-to-point on multipoint interfaces entirely (it can only form one adjacency), and treat p2mp as the default answer. Deeper network-type theory is in the [OSPF guide](https://www.pinglabz.com/ospf/). OSPF's structural annoyance is that an area must share a single view: you cannot summarize inside area 0, so hub-side summarization toward spokes means putting the DMVPN cloud and the branches into their own area (or living with full tables). Flooding also scales with spoke count, and every spoke flap ripples LSAs across the cloud. These are manageable at dozens of spokes and painful at hundreds. ## Timers, MTU, and Adjacency Gotchas Three operational notes that bite regardless of protocol. First, multicast: adjacencies only form because NHRP replicates multicast hellos to mapped peers, so a spoke missing `ip nhrp map multicast` pings the hub fine but never forms a neighbor, which sends you hunting a routing problem that is actually NHRP (see [troubleshooting DMVPN](https://www.pinglabz.com/troubleshooting-dmvpn/)). Second, MTU: OSPF checks interface MTU in DBD exchange, so a tunnel MTU mismatch stalls adjacencies in EXSTART, one more reason to standardize `ip mtu 1400` everywhere (background in [GRE MTU and fragmentation](https://www.pinglabz.com/gre-tunnel-mtu/)). Third, convergence expectations differ: EIGRP over DMVPN converges in seconds with feasible successors, while OSPF p2mp's 30/120 timers mean a silent spoke death takes two minutes to detect unless you add BFD or tune hellos, and aggressive timers across hundreds of tunnels have their own CPU bill. ## Summarization in Practice Since summarization is the payoff of getting these designs right, here is what it looks like on each protocol. EIGRP summarizes per interface, so the hub collapses all site LANs toward the spokes in one line under the tunnel af-interface: ``` router eigrp PINGLABZ address-family ipv4 unicast autonomous-system 100 af-interface Tunnel0 summary-address 10.0.0.0 255.0.0.0 exit-af-interface ``` Every spoke's RIB shrinks to one D route for the entire 10.0.0.0/8 plus its own connected prefixes, and Phase 3 shortcuts keep resolving the specifics on demand. EIGRP installs the customary Null0 discard route on the hub as loop protection. Watch one interaction: the summary also suppresses the specifics toward the spokes' own prefixes, so a spoke whose LAN falls inside the summary must not recursively route its own subnet at the hub; sane addressing (per-site blocks that do not overlap the summary boundary badly) avoids the problem. OSPF has no in-area summarization, so the equivalent design puts the DMVPN tunnel and branch LANs in a dedicated area and summarizes at the ABR with `area 1 range 10.0.0.0 255.0.0.0`, or leaks only a default into a stub/totally-stubby area, which is the cleanest OSPF-over-DMVPN pattern of all: spokes in a totally stubby area carry a default plus intra-area routes and nothing else. ## Convergence Tuning: BFD and Timers Default failure detection over DMVPN is slow (EIGRP holdtime 15 seconds on the tunnel, OSPF p2mp dead time 120 seconds), and tightening hello timers across hundreds of tunnels multiplies control plane load on the hub. BFD is the professional answer: hardware-independent sub-second detection with one session per neighbor, registered to the routing protocol. ``` interface Tunnel0 bfd interval 300 min_rx 300 multiplier 3 router eigrp PINGLABZ address-family ipv4 unicast autonomous-system 100 af-interface Tunnel0 bfd exit-af-interface ``` (The OSPF equivalent is `bfd` under the process or `ip ospf bfd` per interface.) A 300ms/x3 profile detects a dead peer in under a second while surviving normal internet jitter; go slower (500ms or 1s intervals) for high-latency transports. Two cautions specific to DMVPN: BFD sessions ride the tunnel, so they tell you about overlay liveness, not which underlay component died; and hub CPU budgets matter, since a thousand spokes at 300ms is over three thousand BFD packets per second before anything useful happens. Scale the interval to the spoke count and let the routing protocol's own timers be the backstop. ## Choosing, and What About BGP? For a Cisco-only DMVPN, EIGRP named mode is the path of least resistance: two af-interface commands on the hub, free summarization, fast convergence, stub protection for branches. Choose OSPF p2mp when multi-vendor spokes or an existing OSPF backbone demand it, and accept the area-design homework. At very large scale (many hundreds to thousands of spokes), the industry answer shifts to iBGP with the hub as route reflector advertising a default, dynamic peer listeners, and per-spoke policy control; that design pairs naturally with Phase 3 summarization and is standard in DMVPN-based SD-WAN precursors. Our [routing protocols over GRE](https://www.pinglabz.com/routing-protocols-over-gre/) article covers the single-tunnel fundamentals these designs build on. ## Key Takeaways The hub's geometry drives everything: relayed routes must exit the interface they entered, so EIGRP needs `no split-horizon` and a next-hop-self decision (keep it on for Phase 3, off only for Phase 2), while OSPF needs a network type chosen deliberately, with point-to-multipoint as the robust default over broadcast's fragile DR election. Phase 3 plus hub-side summarization is the scaling unlock, and it favors EIGRP or BGP because OSPF cannot summarize within an area. Whatever you run, remember the adjacency depends on NHRP multicast maps, the MTU must match, and the RIB pointing at the hub is correct, not broken, because NHRP shortcuts do the last-mile optimization. Full cluster map in the [DMVPN complete guide](https://www.pinglabz.com/dmvpn/). ### DMVPN Phase 3 Configuration on Cisco IOS XE URL: https://www.pinglabz.com/dmvpn-phase-3-configuration/ Last updated: 2026-07-11T19:17:50.000Z This is the complete, working DMVPN Phase 3 build: one hub, two spokes, a router simulating the internet underlay, EIGRP over the top, every command shown and every verification step captured from live IOS XE 17.18 devices. Nothing here is theoretical output; if a counter or a flag appears, a lab produced it. Concepts are covered in [DMVPN explained](https://www.pinglabz.com/dmvpn-explained/) and the phase logic in [Phase 1 vs 2 vs 3](https://www.pinglabz.com/dmvpn-phase-1-2-3-differences/); this article is the hands-on companion in the [DMVPN cluster](https://www.pinglabz.com/dmvpn/). ## Topology and Addressing Four routers. INET sits in the middle pretending to be the internet, with documentation-range /30s toward each site. Each site router has a loopback standing in for its LAN. HUBunderlay e0/1 203.0.113.1/30tunnel 10.0.0.1LAN Lo1 10.0.100.1/24 SPOKE1underlay e0/1 198.51.100.1/30tunnel 10.0.0.2LAN Lo1 10.0.1.1/24 SPOKE2underlay e0/1 192.0.2.1/30tunnel 10.0.0.3LAN Lo1 10.0.2.1/24 INET203.0.113.2 / 198.51.100.2 / 192.0.2.2no routing protocol, connected routes only The overlay is Tunnel0 on 10.0.0.0/24 everywhere. NHRP network-id 1, authentication string PLZ123, GRE tunnel key 100\. If you are reproducing this in CML, iol-xe nodes boot in seconds and handle everything except IPsec (for encrypted tunnels use Catalyst 8000v, covered in [the IPsec article](https://www.pinglabz.com/dmvpn-ipsec-profiles-ikev2/)). ## Step 1: The Underlay Spokes and hub each need exactly one thing from the underlay: reachability to each other's tunnel source addresses. A static default toward INET does it, mirroring a real internet edge: ``` ! HUB interface Ethernet0/1 ip address 203.0.113.1 255.255.255.252 no shutdown ip route 0.0.0.0 0.0.0.0 203.0.113.2 ! SPOKE1 (SPOKE2 mirrors with 192.0.2.x) interface Ethernet0/1 ip address 198.51.100.1 255.255.255.252 no shutdown ip route 0.0.0.0 0.0.0.0 198.51.100.2 ``` Sanity check before touching any tunnel: every router must ping every other router's underlay address. Skipping this check is the number one way to spend an hour debugging NHRP when the actual problem is underlay routing. ## Step 2: The Hub Tunnel ``` interface Tunnel0 description DMVPN hub - mGRE ip address 10.0.0.1 255.255.255.0 no ip redirects ip mtu 1400 ip tcp adjust-mss 1360 ip nhrp authentication PLZ123 ip nhrp network-id 1 ip nhrp redirect tunnel source Ethernet0/1 tunnel mode gre multipoint tunnel key 100 ``` Line by line, the ones that matter. `tunnel mode gre multipoint` makes this one interface serve every spoke. `ip nhrp redirect` is the Phase 3 half that belongs on the hub: it fires an NHRP redirect at any spoke whose traffic hairpins through this interface. `ip mtu 1400` and `ip tcp adjust-mss 1360` pre-empt fragmentation blackholes (GRE eats 24 bytes, IPsec eats more; the arithmetic is in [GRE MTU and fragmentation](https://www.pinglabz.com/gre-tunnel-mtu/)). `tunnel key 100` tags the GRE packets, which must match on every router in this cloud. On IOS XE 17.x you no longer need `ip nhrp map multicast dynamic` on an mGRE hub; it is the default and will not even appear in the running config. Notably absent: any mention of any spoke. ## Step 3: The Spoke Tunnels ``` interface Tunnel0 description DMVPN spoke - mGRE ip address 10.0.0.2 255.255.255.0 no ip redirects ip mtu 1400 ip tcp adjust-mss 1360 ip nhrp authentication PLZ123 ip nhrp map 10.0.0.1 203.0.113.1 ip nhrp map multicast 203.0.113.1 ip nhrp network-id 1 ip nhrp nhs 10.0.0.1 ip nhrp shortcut tunnel source Ethernet0/1 tunnel mode gre multipoint tunnel key 100 ``` SPOKE2 is identical except its own addresses (tunnel IP 10.0.0.3, source e0/1 on 192.0.2.1). The five NHRP lines do all the work: `map` bootstraps the hub's overlay-to-underlay binding, `map multicast` lets routing protocol hellos reach the hub, `nhs` declares who to register with, and `shortcut` is the Phase 3 half that belongs on spokes, accepting redirects and installing shortcut routes. Spokes never reference other spokes; that is the entire point. ## Step 4: Verify Registration Before Routing Within seconds of the tunnels coming up, both spokes register. On the hub: ``` HUB# show dmvpn Type:Hub, NHRP Peers:2, # Ent Peer NBMA Addr Peer Tunnel Add State UpDn Tm Attrb ----- --------------- --------------- ----- -------- ----- 1 198.51.100.1 10.0.0.2 UP 00:01:35 D 1 192.0.2.1 10.0.0.3 UP 00:01:31 D ``` Both peers UP and D (dynamically learned). If a peer sits in NHRP state instead, stop and fix registration first ([troubleshooting DMVPN](https://www.pinglabz.com/troubleshooting-dmvpn/) diagnoses the three classic causes with captures). Routing on a broken NHRP layer is wasted effort. ## Step 5: EIGRP Named Mode Over the Tunnel ``` ! HUB router eigrp PINGLABZ address-family ipv4 unicast autonomous-system 100 af-interface Tunnel0 no split-horizon exit-af-interface network 10.0.0.0 0.0.0.255 network 10.0.100.0 0.0.0.255 exit-address-family ! SPOKE1 (SPOKE2 mirrors with 10.0.2.0) router eigrp PINGLABZ address-family ipv4 unicast autonomous-system 100 network 10.0.0.0 0.0.0.255 network 10.0.1.0 0.0.0.255 exit-address-family ``` The single non-default line is `no split-horizon` on the hub's tunnel af-interface: without it, EIGRP refuses to re-advertise SPOKE1's routes back out Tunnel0 to SPOKE2, and the spokes never learn each other's LANs. Next-hop-self stays at its default (enabled), which is correct for Phase 3\. Both spokes come up as neighbors over the multicast maps: ``` HUB# show ip eigrp neighbors H Address Interface Hold Uptime SRTT RTO Q Seq 1 10.0.0.3 Tu0 11 00:01:30 11 1398 0 3 0 10.0.0.2 Tu0 11 00:01:35 7 1398 0 4 ``` And the spoke RIB shows every remote prefix via the hub, which is the correct resting state for Phase 3 (the design trade-offs, including the OSPF alternative, are in [routing over DMVPN](https://www.pinglabz.com/routing-over-dmvpn-eigrp-ospf/)): ``` SPOKE1# show ip route eigrp D 10.0.2.0/24 [90/102400640] via 10.0.0.1, 00:00:29, Tunnel0 D 10.0.100.0/24 [90/76800640] via 10.0.0.1, 00:00:29, Tunnel0 ``` ## Step 6: Watch the Shortcut Form Send spoke-to-spoke traffic and run the same traceroute twice. The first flows through the hub and triggers the redirect; the second goes direct: ``` SPOKE1# traceroute 10.0.2.1 source Loopback1 numeric 1 10.0.0.1 2 msec 1 msec 2 msec 2 10.0.0.3 3 msec * 2 msec SPOKE1# traceroute 10.0.2.1 source Loopback1 numeric 1 * 10.0.0.3 4 msec * ``` Three commands prove what happened. The NHRP cache now holds a shortcut for the remote prefix (`nho` flag) and the resolved peer (`nhop rib`): ``` SPOKE1# show ip nhrp 10.0.0.3/32 via 10.0.0.3 Tunnel0 created 00:00:09, expire 00:09:50 Type: dynamic, Flags: router nhop rib NBMA address: 192.0.2.1 10.0.2.0/24 via 10.0.0.3 Tunnel0 created 00:00:09, expire 00:09:50 Type: dynamic, Flags: router used rib nho NBMA address: 192.0.2.1 ``` The RIB keeps the EIGRP route at the hub and overrides its next hop (the `%` and `[NHO]` markers, plus an H route for the peer): ``` SPOKE1# show ip route next-hop-override H 10.0.0.3/32 is directly connected, 00:00:09, Tunnel0 D % 10.0.2.0/24 [90/102400640] via 10.0.0.1, 00:00:57, Tunnel0 [NHO][90/255] via 10.0.0.3, 00:00:09, Tunnel0 ``` And `show dmvpn` shows the two shortcut entry types, route-installed and next-hop-override: ``` SPOKE1# show dmvpn # Ent Peer NBMA Addr Peer Tunnel Add State UpDn Tm Attrb ----- --------------- --------------- ----- -------- ----- 1 203.0.113.1 10.0.0.1 UP 00:02:26 S 2 192.0.2.1 10.0.0.3 UP 00:00:09 DT1 10.0.0.3 UP 00:00:09 DT2 ``` Leave the traffic idle for the NHRP holdtime (10 minutes by default) and all of it unwinds: the dynamic entries expire, forwarding falls back to the hub, and the next traffic burst rebuilds the shortcut. That self-cleaning behavior is Phase 3 working as designed. The message-level mechanics behind the redirect are in the [NHRP deep dive](https://www.pinglabz.com/nhrp-deep-dive/). ## The Mistakes That Cost the Most Lab Time Having built this exact topology more than once, here is where the time actually goes. Forgetting `ip nhrp map multicast` on a spoke: pings to the hub work, EIGRP never forms, and you debug the routing protocol for an hour when the problem is NHRP replication. Mismatched `tunnel key`: both interfaces up/up, everything dead, hub logs nothing, because GRE drops keyed packets before NHRP sees them. Forgetting `no split-horizon` on the hub: both spokes neighbor with the hub and learn the hub's LAN, but never each other's, which looks exactly like a filtering problem. Testing with unsourced pings: `ping 10.0.2.1` from SPOKE1 sources from the tunnel interface and works even when LAN-to-LAN forwarding is broken, so always test with `source Loopback1` the way the captures here do. And impatience with the shortcut: the redirect only fires on hairpinned data traffic, so a freshly cleared cache legitimately shows the hub path on the first packets. Each of these produces a distinct capture signature, all cataloged in [troubleshooting DMVPN](https://www.pinglabz.com/troubleshooting-dmvpn/). Worth repeating from the phase article because it defuses a common review comment: the NHRP network-id does not need to match between routers. Ours matches everywhere because uniformity is easier to operate, not because the protocol requires it. What must genuinely match across the cloud: the NHRP authentication string, the tunnel key, and the overlay subnet. ## Verifying From a Real Host (bridge1) Loopbacks prove routing; a real client proves the design. In our lab, the hub's e0/0 faces an external connector at 192.168.99.1/24 with a Debian VM at 192.168.99.100 behind it, and that subnet rides the same EIGRP process over the tunnel (`network 192.168.99.0` in the hub's address-family). The spokes learn it like any other prefix: ``` SPOKE1# show ip route eigrp D 192.168.99.0/24 [90/77312000] via 10.0.0.1, 00:00:29, Tunnel0 ``` Add a static route on the VM pointing 10.0.0.0/16 at 192.168.99.1 and the Linux host can reach every spoke LAN across the overlay. The same pattern (real host, external connector, overlay route) upgrades any CML lab from ping-between-loopbacks to something you can drive iperf and packet captures through. ## Production Deltas Four gaps between this lab and a deployable design. Encryption: internet transport means IPsec, added purely with a `tunnel protection` profile and an IKEv2 policy, no DMVPN changes ([full build here](https://www.pinglabz.com/dmvpn-ipsec-profiles-ikev2/)). Summarization: with more than a handful of spokes, advertise a summary or default from the hub instead of individual prefixes; Phase 3 shortcuts keep working because NHRP resolution carries the specifics. Hub redundancy: a second hub is a second `ip nhrp nhs` line on every spoke (priorities and clusters control preference). Front-door VRF: production designs usually put the underlay in an fVRF (`tunnel vrf INTERNET`) so the default route toward the ISP cannot collide with the overlay's routing, a recursion problem explained in [GRE tunnel troubleshooting](https://www.pinglabz.com/gre-tunnel-troubleshooting/). ## Key Takeaways A Phase 3 DMVPN is startlingly little configuration: an mGRE tunnel with five NHRP lines per spoke, two Phase 3 commands (`redirect` on the hub, `shortcut` on spokes), one split-horizon exception in EIGRP, and MTU/MSS clamps. Verify in layers, in order: underlay pings, then `show dmvpn` registration on the hub, then routing adjacencies, then the two-traceroute shortcut test with `show ip route next-hop-override` as the receipt. The hub scales because it knows nothing about individual spokes, and the shortcuts self-expire because they are cache entries, not routes. Next steps in the cluster: [encrypt it](https://www.pinglabz.com/dmvpn-ipsec-profiles-ikev2/), then [break it on purpose](https://www.pinglabz.com/troubleshooting-dmvpn/), with the full map in the [DMVPN complete guide](https://www.pinglabz.com/dmvpn/). ### DMVPN Phase 1 vs Phase 2 vs Phase 3: What Actually Changes URL: https://www.pinglabz.com/dmvpn-phase-1-2-3-differences/ Last updated: 2026-07-11T19:17:50.000Z Ask what separates DMVPN Phase 1 from Phase 2 from Phase 3 and you will usually get a memorized answer: "Phase 1 is hub-and-spoke, Phase 2 adds spoke-to-spoke, Phase 3 adds redirects." True, and useless for the ENARSI exam or a real migration, because the phases differ in configuration details, routing protocol requirements, and what shows up in the RIB. So we did the migration for real: one lab (hub, two spokes, IOS XE 17.18), reconfigured live from Phase 1 to Phase 2 to Phase 3, with the routing table, NHRP cache, and traceroutes captured at each stop. This is the phase comparison the [DMVPN complete guide](https://www.pinglabz.com/dmvpn/) summarizes in one card grid, expanded to full depth. ## What "Phase" Actually Means A phase is not negotiated, versioned, or displayed in any show command. It is a name for a combination of configuration choices, and the combination controls exactly one behavior: the path spoke-to-spoke traffic takes. Two things vary: whether spoke tunnels are point-to-point or multipoint GRE, and how next hops are handled (by the routing protocol in Phase 2, by NHRP in Phase 3). Everything else (registration, hub config, the overlay subnet) stays essentially the same across all three. Component background is in [DMVPN explained](https://www.pinglabz.com/dmvpn-explained/). ## Phase 1: Point-to-Point Spokes, Hub Forwards Everything In Phase 1 only the hub runs mGRE. Each spoke is a plain GRE tunnel with a hardcoded destination, plus NHRP configuration used purely for registration: ``` interface Tunnel0 ip address 10.0.0.2 255.255.255.0 ip nhrp authentication PLZ123 ip nhrp map 10.0.0.1 203.0.113.1 ip nhrp map multicast 203.0.113.1 ip nhrp network-id 1 ip nhrp nhs 10.0.0.1 tunnel source Ethernet0/1 tunnel destination 203.0.113.1 tunnel key 100 ``` The spoke's NHRP cache stays minimal forever, just the static hub entry: ``` SPOKE1# show ip nhrp 10.0.0.1/32 via 10.0.0.1 Tunnel0 created 00:02:23, never expire Type: static, Flags: NBMA address: 203.0.113.1 ``` With EIGRP running over the top (hub configured with `no split-horizon` so spoke routes reach other spokes), every remote prefix points at the hub, and spoke-to-spoke traffic transits it on every packet: ``` SPOKE1# show ip route eigrp D 10.0.2.0/24 [90/102400640] via 10.0.0.1, 00:01:58, Tunnel0 D 10.0.100.0/24 [90/76800640] via 10.0.0.1, 00:02:04, Tunnel0 SPOKE1# traceroute 10.0.2.1 source Loopback1 numeric 1 10.0.0.1 2 msec 2 msec 1 msec 2 10.0.0.3 2 msec * 3 msec ``` Because a point-to-point tunnel can only ever send to the hub, the routing design is trivially free: summarize, filter, send a default route, nothing breaks. The costs are hub bandwidth (every spoke-to-spoke packet crosses it twice), hub CPU, and latency. Phase 1 is legitimate where spokes genuinely never talk to each other, and as the first stage of a migration, which is exactly how we use it here. ## Phase 2: mGRE Spokes, Routing Protocol Preserves Next Hops Migrating a spoke to Phase 2 is two interface commands (drop the fixed destination, go multipoint): ``` interface Tunnel0 no tunnel destination tunnel mode gre multipoint ``` But mGRE alone changes nothing about forwarding; the routes still point at the hub. Phase 2's defining requirement lives in the routing protocol: the hub must advertise spoke prefixes with the originating spoke's tunnel address as next hop. For EIGRP that means `no next-hop-self` under the tunnel af-interface (plus the `no split-horizon` you already needed). The result in SPOKE1's RIB, captured seconds after the change: ``` SPOKE1# show ip route eigrp D 10.0.2.0/24 [90/102400640] via 10.0.0.3, 00:00:15, Tunnel0 D 10.0.100.0/24 [90/76800640] via 10.0.0.1, 00:00:17, Tunnel0 ``` That `via 10.0.0.3` is the Phase 2 signature: the next hop for SPOKE2's LAN is SPOKE2's tunnel IP, not the hub. The first packet toward it triggers an NHRP resolution for 10.0.0.3, the spokes learn each other's underlay addresses, and a dynamic tunnel forms: ``` SPOKE1# show dmvpn # Ent Peer NBMA Addr Peer Tunnel Add State UpDn Tm Attrb ----- --------------- --------------- ----- -------- ----- 1 203.0.113.1 10.0.0.1 UP 00:00:59 S 1 192.0.2.1 10.0.0.3 UP 00:00:00 D SPOKE1# traceroute 10.0.2.1 source Loopback1 numeric 1 10.0.0.3 2 msec * 2 msec ``` One hop, direct. So why did the industry abandon Phase 2? Because the trigger is the RIB, the RIB must contain every individual spoke prefix with its original next hop. Summarize at the hub and the specific next hops disappear into the summary; spokes lose the information needed to resolve each other and traffic quietly falls back to the hub (or worse, blackholes with OSPF designs). A thousand spokes means a thousand prefixes in every spoke's table. Phase 2 also forces awkward routing protocol constraints (OSPF broadcast mode with the hub winning DR election, EIGRP without summarization), which is a routing tax explored in [routing over DMVPN](https://www.pinglabz.com/routing-over-dmvpn-eigrp-ospf/). ## Phase 3: NHRP Takes Over Next-Hop Discovery Phase 3 keeps mGRE everywhere and moves the shortcut trigger from the routing protocol into NHRP with one command on each side: ``` ! Hub interface Tunnel0 ip nhrp redirect ! Spokes interface Tunnel0 ip nhrp shortcut ``` The routing design reverts to boring: hub keeps next-hop-self (we re-enabled it in the lab), and summarization is not only allowed but encouraged. All routes point at the hub again: ``` SPOKE1# show ip route eigrp D 10.0.2.0/24 [90/102400640] via 10.0.0.1, 00:00:29, Tunnel0 ``` Now the hub watches for hairpin traffic (in and out the same tunnel interface). When SPOKE1 sends the first packets toward SPOKE2 through it, the hub forwards them and simultaneously sends SPOKE1 an NHRP Traffic Indication: "resolve this yourself." The spoke does, and installs the answer underneath the routing table. Two consecutive traceroutes catch the cutover mid-flight: ``` SPOKE1# traceroute 10.0.2.1 source Loopback1 numeric 1 10.0.0.1 2 msec 1 msec 2 msec 2 10.0.0.3 3 msec * 2 msec SPOKE1# traceroute 10.0.2.1 source Loopback1 numeric 1 * 10.0.0.3 4 msec * ``` The mechanism is visible in three places. The NHRP cache gains a shortcut entry for the remote LAN prefix itself (not just the tunnel address), flagged `nho`: ``` SPOKE1# show ip nhrp 10.0.2.0/24 via 10.0.0.3 Tunnel0 created 00:00:09, expire 00:09:50 Type: dynamic, Flags: router used rib nho NBMA address: 192.0.2.1 ``` The RIB shows the EIGRP route still pointing at the hub, with an NHRP next-hop override beside it (the `%` marker) and an H route for the peer's tunnel address: ``` SPOKE1# show ip route next-hop-override H 10.0.0.3/32 is directly connected, 00:00:09, Tunnel0 D % 10.0.2.0/24 [90/102400640] via 10.0.0.1, 00:00:57, Tunnel0 [NHO][90/255] via 10.0.0.3, 00:00:09, Tunnel0 ``` And `show dmvpn` distinguishes the shortcut kinds: DT1 is the route-installed entry, DT2 the next-hop override: ``` SPOKE1# show dmvpn # Ent Peer NBMA Addr Peer Tunnel Add State UpDn Tm Attrb ----- --------------- --------------- ----- -------- ----- 1 203.0.113.1 10.0.0.1 UP 00:02:26 S 2 192.0.2.1 10.0.0.3 UP 00:00:09 DT1 10.0.0.3 UP 00:00:09 DT2 ``` Control plane and data plane are now decoupled: EIGRP could advertise nothing but a summary (or a plain default route) and shortcuts would still form, because NHRP resolution carries the real prefix information on demand. That is what lets Phase 3 hubs serve hundreds of spokes with tiny routing tables. The full configuration walkthrough is in [DMVPN Phase 3 configuration on Cisco IOS XE](https://www.pinglabz.com/dmvpn-phase-3-configuration/), and the NHRP message mechanics are in the [NHRP deep dive](https://www.pinglabz.com/nhrp-deep-dive/). ## The Migration Path, Compressed Since we just did it live, here is the entire Phase 1 to Phase 3 migration in order of operations. On each spoke: remove `tunnel destination`, set `tunnel mode gre multipoint`, add `ip nhrp shortcut`. On the hub: add `ip nhrp redirect`, keep `no split-horizon`, keep (or restore) next-hop-self, and start summarizing if the design allows. Each spoke can migrate independently while Phase 1 spokes continue working through the hub, which is what makes this migration doable during business hours: a Phase 1 spoke and a Phase 3 spoke coexist happily on the same hub, they just never build a direct tunnel to each other. ## The Summarization Test: One Design Question That Separates the Phases If you remember one discriminator, make it this one: what happens when the hub advertises a summary instead of specific prefixes? In Phase 1, nothing interesting; traffic already goes to the hub, the hub has the specifics, forwarding is unchanged. In Phase 2, disaster in slow motion: spokes lose the per-prefix next hops that triggered resolution, spoke-to-spoke traffic follows the summary to the hub, and your direct tunnels silently stop forming; with OSPF variants of the design you could even blackhole. In Phase 3, the summary is the design: traffic follows it to the hub, the hub's redirect teaches the source spoke the specific destination, and NHRP resolution supplies the granular information the routing protocol deliberately stopped carrying. Push Phase 3 to its logical extreme and the hub advertises exactly one route: 0.0.0.0/0 over the tunnel. Every spoke's table holds a default via 10.0.0.1 plus its own connected routes, yet full spoke-to-spoke meshing still works, one NHRP resolution at a time. That single-default design is how DMVPN clouds reach thousands of spokes with routers whose RIBs would choke on a full table, and it only exists because Phase 3 moved next-hop discovery out of the control plane. It also collapses convergence scope: a spoke flapping in Toronto no longer churns the routing table of a spoke in Madrid that never talks to it. ## Hub Failure Behavior Differs by Phase Too The phases also answer "what breaks when the hub dies?" differently, which matters for redundancy design. In Phase 1, everything: the hub is the data path. In Phase 2 and 3, established spoke-to-spoke tunnels keep forwarding on their cached NHRP entries even with the hub down, degrading gracefully until holdtimes expire or new resolutions are needed; what you lose immediately is the ability to form new shortcuts and, usually, the routing control plane. This is why DMVPN redundancy focuses on the NHS: a second hub is one more `ip nhrp nhs` statement per spoke (with priority and cluster options controlling failover order), and dual-hub dual-cloud designs put each hub on its own tunnel interface and let routing metrics choose. Whatever the phase, spokes should treat NHS entries the way they treat default gateways: always have two. ## Exam Traps Worth Rehearsing Four distinctions ENARSI likes. First, `ip nhrp shortcut` goes on spokes and `ip nhrp redirect` goes on the hub (a spoke-only cloud never sends redirects; both commands on all routers is harmless and common in configs, but know which does what). Second, Phase 2 requires `no next-hop-self`; Phase 3 pointedly does not, and combining Phase 3 with preserved next hops mostly works but reintroduces Phase 2's summarization constraints for nothing. Third, the phase is per-tunnel-interface, not per-router, so a hub can serve different phases on different tunnels. Fourth, only Phase 1 uses `tunnel destination` on spokes; if you see it with `tunnel mode gre multipoint` in the same interface, the config is broken, not hybrid. ## Key Takeaways The phases answer one question three ways: how does spoke-to-spoke traffic flow? Phase 1 (p2p GRE spokes) sends it through the hub forever but leaves routing design unconstrained. Phase 2 (mGRE spokes, next hops preserved by the routing protocol) goes direct but forbids summarization because the RIB itself is the shortcut trigger. Phase 3 (mGRE plus `redirect` on the hub and `shortcut` on spokes) moves the trigger into NHRP, restoring summarization while keeping direct paths, visible as `%`/`[NHO]` overrides and H routes underneath an unchanged routing table. Deploy Phase 3 unless you have a documented reason not to. See the whole architecture in the [DMVPN complete guide](https://www.pinglabz.com/dmvpn/), or verify each phase's tunnels layer by layer with the [GRE guide](https://www.pinglabz.com/gre/). ### NHRP Deep Dive: Registration, Resolution, and Redirects URL: https://www.pinglabz.com/nhrp-deep-dive/ Last updated: 2026-07-11T19:17:49.000Z Every interesting thing DMVPN does, it does with NHRP. Spokes appearing on the hub without configuration? NHRP registration. Spoke-to-spoke tunnels forming on demand? NHRP resolution. Phase 3's ability to summarize routes and still cut hub-free paths? NHRP redirects rewriting the forwarding table underneath the RIB. If you can read NHRP's three message flows in debug output, DMVPN stops being magic and starts being inspectable. This article walks all three with real debugs and cache states from an IOS XE 17.18 lab (one hub, two spokes). It assumes the component overview from [DMVPN explained](https://www.pinglabz.com/dmvpn-explained/) and slots into the wider [DMVPN complete guide](https://www.pinglabz.com/dmvpn/). ## What NHRP Is Actually For NHRP (RFC 2332) predates DMVPN; it was designed for resolving next hops across NBMA networks like Frame Relay and ATM. An mGRE cloud is just the modern NBMA network: every router can reach every other router, but nobody knows anybody's transport address in advance. NHRP fills exactly one gap: given an overlay (tunnel) IP, return the underlay (NBMA) IP where it lives. The analogy that holds up under pressure is ARP with a receptionist. Instead of broadcasting "who has 10.0.0.3?", you ask the one router that keeps the ledger: the Next Hop Server. Keep the division of labor straight, because it explains half of all DMVPN confusion: NHRP maps tunnel addresses to transport addresses. It does not carry routes. Your routing protocol (see [routing over DMVPN](https://www.pinglabz.com/routing-over-dmvpn-eigrp-ospf/)) advertises what networks exist; NHRP figures out how to deliver GRE packets to the routers that advertised them. ## The Cast: NHS, NHC, and the Cache The hub is the Next Hop Server (NHS), the authoritative ledger. Spokes are Next Hop Clients, each configured with the hub's overlay address (`ip nhrp nhs 10.0.0.1`) plus one static map bootstrapping how to reach it (`ip nhrp map 10.0.0.1 203.0.113.1`). Everything else in the system is learned. The ledger itself is the NHRP cache, and learning to read its flags pays off immediately: ``` SPOKE1# show ip nhrp 10.0.0.1/32 via 10.0.0.1 Tunnel0 created 00:02:25, never expire Type: static, Flags: used NBMA address: 203.0.113.1 10.0.0.3/32 via 10.0.0.3 Tunnel0 created 00:00:09, expire 00:09:50 Type: dynamic, Flags: router nhop rib NBMA address: 192.0.2.1 10.0.2.0/24 via 10.0.0.3 Tunnel0 created 00:00:09, expire 00:09:50 Type: dynamic, Flags: router used rib nho NBMA address: 192.0.2.1 ``` Static entries come from config and never expire. Dynamic entries age out (default 10 minutes, shown counting down). The flags tell the story: `registered` means a spoke put it there via registration, `nhop` marks a resolved next hop, `rib` means NHRP handed it to the routing table, and `nho` means it is installed as a next-hop override, the Phase 3 mechanism. A `local` entry with `(no-socket)` is the router's own registration record of what it told the NHS about itself. ## Flow 1: Registration, How the Hub Learns Everyone When a spoke's tunnel comes up, it immediately registers with every configured NHS. Here is the exchange in `debug nhrp packet` on the spoke, captured by bouncing Tunnel0: ``` *Jul 11 18:39:31.271: NHRP: Attempting to send packet through interface Tunnel0 via DEST dst 10.0.0.1 *Jul 11 18:39:31.271: NHRP: Send Registration Request via Tunnel0 vrf: global(0x0), packet size: 106 *Jul 11 18:39:31.272: NHRP: 134 bytes out Tunnel0 *Jul 11 18:39:31.274: NHRP: Receive Registration Reply via Tunnel0 vrf: global(0x0), packet size: 126 *Jul 11 18:39:31.274: NHRP-EVE: NHS-UP: 10.0.0.1, NBMA: 203.0.113.1 *Jul 11 18:39:31.274: %DMVPN-5-NHRP_NHS_UP: Tunnel0: Next Hop Server : (Tunnel: 10.0.0.1 NBMA: 203.0.113.1) for (Tunnel: 10.0.0.2 NBMA: 198.51.100.1) is UP ``` The request carries the spoke's own binding (tunnel 10.0.0.2, NBMA 198.51.100.1). The hub caches it as a dynamic, registered entry and answers with a reply. Registrations refresh periodically (a fraction of the holdtime), so the hub's cache is self-healing: change a spoke's underlay address and the next registration overwrites the binding. Registrations also set a uniqueness flag by default, which stops a second device from registering an already-claimed tunnel address. Registration health is the first thing to check on any DMVPN problem, and it has a dedicated command whose counters do the diagnosing for you: ``` SPOKE2# show ip nhrp nhs detail Legend: E=Expecting replies, R=Responding, W=Waiting, D=Dynamic Tunnel0: 10.0.0.1 E priority = 0 cluster = 0 req-sent 7 req-failed 0 repl-recv 0 (00:01:48 ago) Pending Registration Requests: Registration Request: Reqid 11, Ret 64 NHS 10.0.0.1 expired (Tu0) ``` That capture is from a deliberately broken spoke: `req-sent` climbing while `repl-recv` stays at zero, NHS stuck in E (expecting replies). A healthy NHS shows RE with matching sent and received counters. The three configuration mistakes that produce exactly this signature are dissected in [troubleshooting DMVPN](https://www.pinglabz.com/troubleshooting-dmvpn/). ## Flow 2: Resolution, How Spokes Find Each Other Registration only teaches the hub. When SPOKE1 wants to reach something behind SPOKE2 directly, it asks: an NHRP Resolution Request goes to the NHS, and (in Phase 3) the hub forwards it to the spoke that owns the answer, which replies directly to the asker. Both directions from the lab debug: ``` *Jul 11 18:42:31.141: NHRP: Send Resolution Request via Tunnel0 vrf: global(0x0), packet size: 86 *Jul 11 18:42:31.144: NHRP: Receive Resolution Reply via Tunnel0 vrf: global(0x0), packet size: 134 *Jul 11 18:42:31.146: NHRP: Receive Resolution Request via Tunnel0 vrf: global(0x0), packet size: 106 *Jul 11 18:42:31.146: NHRP: Send Resolution Reply via Tunnel0 vrf: global(0x0), packet size: 134 *Jul 11 18:42:31.146: %DMVPN-7-NHRP_RES: Tunnel0: Host with (Tunnel: 10.0.0.2 NBMA: 198.51.100.1) Received Resolution Req from (Tunnel: 10.0.0.3 NBMA: 192.0.2.1) ``` Notice SPOKE1 both sends a resolution request (it wants SPOKE2's binding) and receives one (SPOKE2 wants SPOKE1's binding, because return traffic needs its own shortcut). Resolution replies populate the dynamic cache entries you saw above, and the resulting spoke-to-spoke tunnel shows up in `show dmvpn` as a D entry on both spokes. Idle entries expire with the holdtime and the mesh prunes itself back to hub-and-spoke. The per-message-type counters are the fastest way to see which flow is happening (or failing) without a debug: ``` SPOKE1# show ip nhrp traffic Tunnel0: Max-send limit:10000Pkts/10Sec, Usage:0% Sent: Total 10 2 Resolution Request 2 Resolution Reply 5 Registration Request 0 Registration Reply 1 Purge Request 0 Purge Reply 0 Error Indication 0 Traffic Indication 0 Redirect Suppress Rcvd: Total 11 2 Resolution Request 2 Resolution Reply 0 Registration Request 5 Registration Reply 0 Purge Request 1 Purge Reply 0 Error Indication 1 Traffic Indication 0 Redirect Suppress ``` A spoke sends registrations and receives registration replies, never the reverse (only the NHS receives registrations). Purge messages are NHRP's cache invalidation, sent when a binding a router previously answered for changes or disappears. ## Flow 3: Redirects, the Phase 3 Trigger Resolution needs a trigger. In Phase 2 designs the trigger was the routing table itself: spokes saw each other's tunnel IPs as next hops (which is why Phase 2 forbids summarization; see [the phase comparison](https://www.pinglabz.com/dmvpn-phase-1-2-3-differences/)). Phase 3 moves the trigger into NHRP. The hub, configured with `ip nhrp redirect`, watches for packets that enter and leave the same tunnel interface. That hairpin is the signal: two spokes are talking through it. The hub forwards the packet normally but also fires an NHRP Traffic Indication (the redirect) back at the source, which says, in effect, "there is a better way to reach this destination; go resolve it." You can see the receipt in the traffic counters above: `Rcvd: 1 Traffic Indication`. A spoke configured with `ip nhrp shortcut` responds by sending the resolution request from Flow 2 and installing the answer as a shortcut. The redirect is rate-limited and purely advisory: a spoke without `shortcut` ignores it and keeps using the hub, which is why forgetting the command produces a working network that just never optimizes. ## How a Shortcut Rewrites Forwarding Without Touching the RIB The subtle brilliance of Phase 3 is where the shortcut lands. The EIGRP route still points at the hub; NHRP installs a next-hop override beside it, flagged with `%` and `[NHO]`: ``` SPOKE1# show ip route next-hop-override H 10.0.0.3/32 is directly connected, 00:00:09, Tunnel0 D % 10.0.2.0/24 [90/102400640] via 10.0.0.1, 00:00:57, Tunnel0 [NHO][90/255] via 10.0.0.3, 00:00:09, Tunnel0 ``` The H route is an NHRP host route to the other spoke's tunnel address; the NHO entry overrides the EIGRP next hop for actual forwarding. When the cache entry expires, both vanish and forwarding falls back to the hub path with zero routing protocol churn. Control plane stability, data plane optimization: that separation is the whole Phase 3 design and the reason hubs can advertise a single summary (or default) route to hundreds of spokes. The complete build is in [DMVPN Phase 3 configuration](https://www.pinglabz.com/dmvpn-phase-3-configuration/). ## Reading show dmvpn Like a Native The `show dmvpn` attribute column is a compressed summary of the NHRP cache, and it repays close reading. S is a static entry from an `ip nhrp map`; D is dynamically learned, which on a hub means a registration and on a spoke means a resolution. The Phase 3 shortcut states come as a pair: DT1 is a shortcut installed as a route (route installed), DT2 is a shortcut installed as a next-hop override on an existing route, and a single resolved peer often shows both, as in this lab capture taken seconds after a shortcut formed: ``` SPOKE1# show dmvpn # Ent Peer NBMA Addr Peer Tunnel Add State UpDn Tm Attrb ----- --------------- --------------- ----- -------- ----- 1 203.0.113.1 10.0.0.1 UP 00:02:26 S 2 192.0.2.1 10.0.0.3 UP 00:00:09 DT1 10.0.0.3 UP 00:00:09 DT2 ``` The other attributes worth recognizing on sight: N flags a NATed peer (the hub saw a different source address than the spoke claimed), I is incomplete (resolution in flight or failed), X means no crypto socket in an IPsec design, and I2 marks a temporary entry created to carry a resolution conversation. The State column is coarser but faster: UP is healthy, NHRP means the control plane is actively failing, and IKE-related states point one layer down at IPsec. ## NHRP Behind NAT NAT bends NHRP in a specific, well-handled way. A spoke behind a NAT gateway believes its NBMA address is its private interface address, but its registration arrives at the hub from the NAT's public address. Modern NHRP carries NAT extensions precisely for this: the hub records the claimed address, notices the observed source differs, stores both, and flags the peer N in `show dmvpn`. Hub-to-spoke traffic then targets the observed (post-NAT) address, and the periodic re-registrations double as NAT keepalives holding the translation open, which is one reason short registration timers matter for spokes behind aggressive home routers. Spoke-to-spoke shortcuts are where NAT gets philosophical. Two spokes behind separate NATs may simply be unable to reach each other's observed addresses (hairpin and symmetric NAT problems), in which case resolution succeeds, the shortcut attempt fails, and traffic quietly stays on the hub path. That is a feature, not a bug: Phase 3 degrades to hub-and-spoke rather than blackholing. In encrypted designs, IKEv2's NAT traversal moves ESP into UDP 4500 and the whole stack keeps working. The practical rule: spokes behind NAT are fine, expect N flags and occasional hub-path fallbacks; hubs belong on real public addresses. ## The Fine Print: Auth, Network-ID, and Holdtime Three NHRP knobs generate a disproportionate share of tickets. `ip nhrp authentication` is a cleartext string (8 characters max) checked on every received NHRP packet; a mismatch makes the hub log `%DMVPN-3-DMVPN_NHRP_ERROR ... authentication failure` and ignore the spoke. It is not security (it is cleartext in the packet), but it is a useful fence against joining the wrong cloud. `ip nhrp network-id` is locally significant and does not need to match across routers, despite a decade of forum posts claiming otherwise; it exists to bind NHRP processing to a tunnel interface on that box. And `ip nhrp holdtime` (default 600 seconds) controls how long your bindings live in other routers' caches, which sets both how fast stale entries disappear and how often registration refreshes happen. GRE-layer settings like the tunnel key live one layer down; the [GRE guide](https://www.pinglabz.com/gre/) covers them. ## Key Takeaways NHRP is the resolution layer that makes mGRE usable: it maps overlay addresses to underlay addresses and carries no routes. Registration flows spoke-to-hub and builds the NHS cache automatically; resolution flows on demand and builds spoke-to-spoke entries; redirects are the Phase 3 trigger that tells a spoke a shortcut exists. Read the cache flags (`registered`, `nhop`, `rib`, `nho`) and the `show ip nhrp nhs detail` counters before reaching for any debug, and remember that authentication must match while network-id does not. Watching these flows misbehave is the fastest way to learn them, which is exactly what [troubleshooting DMVPN](https://www.pinglabz.com/troubleshooting-dmvpn/) does on purpose. For the full architecture picture, return to the [DMVPN complete guide](https://www.pinglabz.com/dmvpn/). ### DMVPN Explained: mGRE + NHRP + Routing in One Overlay URL: https://www.pinglabz.com/dmvpn-explained/ Last updated: 2026-07-11T19:17:49.000Z DMVPN has a reputation for being complicated, and it is entirely undeserved. The technology is three well-understood pieces (multipoint GRE, NHRP, and a routing protocol) assembled so that each covers a gap the others leave. Once you can say what each piece knows and what it does not know, every DMVPN behavior, from spoke registration to dynamic spoke-to-spoke tunnels, becomes predictable. This article walks the whole assembly with real output from a live IOS XE 17.18 lab: a hub, two spokes, and an underlay router standing in for the internet. It is the conceptual companion to the [DMVPN complete guide](https://www.pinglabz.com/dmvpn/) and assumes you know basic GRE from the [GRE guide](https://www.pinglabz.com/gre/). ## The Problem: Tunnels Do Not Scale by Hand A point-to-point GRE tunnel needs two things configured on each end: a source and a destination. That is fine for 2 sites and miserable for 50\. A full mesh of n sites needs n(n-1)/2 tunnels, and even a plain hub-and-spoke needs a distinct tunnel interface on the hub for every spoke. Worse, the destination must be a fixed, known IP address, which rules out branches on DHCP broadband without extra machinery. What you actually want is one tunnel interface per router that can talk to any number of peers, with the peer addresses learned at runtime instead of typed into the config. That is exactly what DMVPN builds, and each of its three components solves one specific part of the problem. ## Ingredient 1: mGRE, One Interface for Many Peers Multipoint GRE is regular GRE encapsulation with the destination removed from the configuration. The interface has a source and a mode, nothing else: ``` interface Tunnel0 ip address 10.0.0.1 255.255.255.0 tunnel source Ethernet0/1 tunnel mode gre multipoint tunnel key 100 ``` Think of the overlay subnet (10.0.0.0/24 here) as one big virtual LAN that all DMVPN routers share. On a real LAN, when a host wants to send a frame it uses ARP to resolve an IP address into a MAC address. An mGRE interface has the same problem in a different costume: to send a GRE packet to overlay address 10.0.0.3, the router must learn which underlay address (Cisco docs call it the NBMA address) that tunnel IP lives at. GRE cannot answer that question. Something has to play the role ARP plays on Ethernet. ## Ingredient 2: NHRP, the Overlay's Address Book NHRP (Next Hop Resolution Protocol) is that something. One router, the hub, is designated the Next Hop Server (NHS), and every spoke is configured with the hub's overlay address plus a static map telling it how to reach the hub in the underlay: ``` interface Tunnel0 ip nhrp map 10.0.0.1 203.0.113.1 ip nhrp map multicast 203.0.113.1 ip nhrp network-id 1 ip nhrp nhs 10.0.0.1 ``` At boot, each spoke sends the hub an NHRP Registration Request: "overlay 10.0.0.2 lives at underlay 198.51.100.1." Here is that exchange from the spoke's debug, triggered by bouncing the tunnel: ``` *Jul 11 18:39:31.271: NHRP: Send Registration Request via Tunnel0 vrf: global(0x0), packet size: 106 *Jul 11 18:39:31.274: NHRP: Receive Registration Reply via Tunnel0 vrf: global(0x0), packet size: 126 *Jul 11 18:39:31.274: %DMVPN-5-NHRP_NHS_UP: Tunnel0: Next Hop Server : (Tunnel: 10.0.0.1 NBMA: 203.0.113.1) for (Tunnel: 10.0.0.2 NBMA: 198.51.100.1) is UP ``` The hub builds its cache purely from these registrations. Notice the attribute column: the spokes appear as D (dynamic), meaning the hub was never configured with either of them: ``` HUB# show dmvpn Type:Hub, NHRP Peers:2, # Ent Peer NBMA Addr Peer Tunnel Add State UpDn Tm Attrb ----- --------------- --------------- ----- -------- ----- 1 198.51.100.1 10.0.0.2 UP 00:01:35 D 1 192.0.2.1 10.0.0.3 UP 00:01:31 D ``` Later, when a spoke needs to reach another spoke directly, it sends the hub an NHRP Resolution Request ("who holds 10.0.0.3?") and caches the answer for a few minutes. Registration populates the hub; resolution queries it on demand. The two message types, plus the Phase 3 redirect, get their own article: the [NHRP deep dive](https://www.pinglabz.com/nhrp-deep-dive/). One myth worth killing early: the `ip nhrp network-id` does not need to match between routers. It is locally significant. What must match are the NHRP authentication string and the GRE tunnel key, and both fail in instructive ways covered in [troubleshooting DMVPN](https://www.pinglabz.com/troubleshooting-dmvpn/). ## Ingredient 3: A Routing Protocol, Because NHRP Only Knows Endpoints NHRP resolves tunnel addresses. It knows nothing about the LANs behind each router. If SPOKE2's users sit on 10.0.2.0/24, no NHRP message will ever advertise that prefix. This division of labor trips people up constantly, so it is worth stating plainly: NHRP maps overlay addresses to underlay addresses; the routing protocol maps LAN prefixes to overlay next hops. You need both. So a routing protocol runs across the tunnel exactly as it would on a LAN. In our lab it is EIGRP named mode, and the spoke learns the other sites' prefixes with next hops on the tunnel subnet: ``` SPOKE1# show ip route eigrp D 10.0.2.0/24 [90/102400640] via 10.0.0.1, 00:01:58, Tunnel0 D 10.0.100.0/24 [90/76800640] via 10.0.0.1, 00:02:04, Tunnel0 D 192.168.99.0/24 [90/77312000] via 10.0.0.1, 00:02:04, Tunnel0 ``` The multicast piece deserves a sentence. EIGRP hellos and OSPF hellos are multicast, and an mGRE interface has no broadcast capability. The `ip nhrp map multicast 203.0.113.1` line on spokes (and the dynamic equivalent on the hub, which is the default for mGRE on modern IOS XE) tells the router to replicate multicast packets as unicast GRE to those peers. Forget it and your routing protocol never forms an adjacency, even though pings work. Routing design choices over the overlay are covered in [routing over DMVPN](https://www.pinglabz.com/routing-over-dmvpn-eigrp-ospf/). ## Watching the Three Work Together Here is the sequence when a user behind SPOKE1 (10.0.1.0/24) sends a packet to a user behind SPOKE2 (10.0.2.0/24) in a Phase 3 network, with each component doing its one job: First, the RIB (routing protocol's contribution) says 10.0.2.0/24 is reachable via 10.0.0.1, so the first packets go to the hub. The hub forwards them to SPOKE2 and simultaneously notices it relayed traffic between two spokes on the same tunnel, so it sends SPOKE1 an NHRP redirect (NHRP's contribution). SPOKE1 resolves 10.0.2.0/24, learns it lives at underlay 192.0.2.1, and installs a shortcut. From then on, mGRE (encapsulation's contribution) wraps packets directly to 192.0.2.1 with no hub in the path. Two consecutive traceroutes from the lab show the switch happening: ``` SPOKE1# traceroute 10.0.2.1 source Loopback1 numeric 1 10.0.0.1 2 msec 1 msec 2 msec 2 10.0.0.3 3 msec * 2 msec SPOKE1# traceroute 10.0.2.1 source Loopback1 numeric 1 * 10.0.0.3 4 msec * ``` And the NHRP cache afterward shows both kinds of entry side by side: the static map to the hub that never expires, and the dynamic shortcut that will age out if unused: ``` SPOKE1# show ip nhrp 10.0.0.1/32 via 10.0.0.1 Tunnel0 created 00:02:25, never expire Type: static, Flags: used NBMA address: 203.0.113.1 10.0.2.0/24 via 10.0.0.3 Tunnel0 created 00:00:09, expire 00:09:50 Type: dynamic, Flags: router used rib nho NBMA address: 192.0.2.1 ``` How much of that machinery activates depends on the DMVPN phase. In Phase 1 the spokes run plain point-to-point GRE and everything transits the hub forever; in Phase 3 the redirect-and-shortcut dance above happens automatically. The phases are compared with live captures in [Phase 1 vs Phase 2 vs Phase 3](https://www.pinglabz.com/dmvpn-phase-1-2-3-differences/). ## What About Encryption? Nothing above is encrypted. GRE is an envelope, not a safe. The fourth, optional ingredient is an IPsec profile applied to the tunnel interface with `tunnel protection`, which wraps every GRE packet in ESP negotiated by IKEv2\. The elegant part is that NHRP, the routing protocol, and mGRE are completely unaware of it; the configuration and verification are identical with or without encryption. The full IKEv2 build with real SA output is in [securing DMVPN with IPsec profiles](https://www.pinglabz.com/dmvpn-ipsec-profiles-ikev2/). ## What the Underlay Must Provide (and What It Must Not Do) Because DMVPN is an overlay, it is easy to forget that it makes demands on the transport underneath. They are few but absolute. Every router's tunnel source address must be reachable from every other router's tunnel source; for spoke-to-spoke tunnels specifically, that means spokes need direct reachability to each other's public addresses, not just to the hub. The transport must pass GRE (IP protocol 47) or, in encrypted designs, ESP and IKE (UDP 500/4500); a provider or firewall that filters protocol 47 produces the same maddening silent-drop symptoms as a tunnel key mismatch. And the underlay MTU budget must absorb the encapsulation overhead, which is why every tunnel in this cluster carries `ip mtu 1400` and `ip tcp adjust-mss 1360`. NAT deserves its own sentence. A spoke behind NAT registers whatever source address its packets arrive from, and NHRP on modern IOS XE handles this with NAT extensions (the hub records both the claimed and observed NBMA address, and `show dmvpn` flags the peer with N). Hub behind NAT is much worse and generally a design smell: spokes bootstrap to a static map, and that map must be the hub's post-NAT address. The rule of thumb is that spoke-side NAT is survivable and hub-side NAT should be engineered away. One recursion trap rounds out the underlay contract. The router must not learn a route to the tunnel destination through the tunnel itself, or the interface flaps with `%TUN-5-RECURDOWN`. It happens the moment someone advertises the underlay subnets into the overlay routing protocol. The clean fix in production is a front-door VRF for the transport interface, so underlay and overlay routing tables cannot see each other at all; the quick fix in a lab is simply never advertising underlay prefixes into the overlay IGP. ## Where DMVPN Sits Among the Alternatives It helps to place DMVPN on the map of Cisco VPN technologies, because exam blueprints love the comparison. Against static site-to-site IPsec (crypto maps or SVTIs), DMVPN wins on operational scale: one interface per router versus one tunnel per peer, plus dynamic spoke addresses for free. Against GETVPN, the trade is topology: GETVPN encrypts traffic on an existing any-to-any WAN (like MPLS) with no tunnels at all, while DMVPN builds the any-to-any topology itself over an untrusted transport, which is why GETVPN lives inside private WANs and DMVPN lives on the internet. FlexVPN, Cisco's IKEv2-native framework, overlaps DMVPN heavily and shares NHRP for its spoke-to-spoke machinery; new greenfield designs sometimes start there, but the installed base, tooling, and exam weight still belong to DMVPN. And SD-WAN is, bluntly, DMVPN's ideas with a controller: overlay tunnels, central registration, dynamic path selection. Understanding this stack of trade-offs is most of the "choose the right VPN" question in [GRE vs IPsec vs GRE-over-IPsec](https://www.pinglabz.com/gre-vs-ipsec/). ## The Mental Model That Makes DMVPN Easy Treat the overlay as an Ethernet LAN and map each DMVPN component onto its LAN equivalent. The tunnel subnet is the LAN segment. The underlay NBMA address is the MAC address. NHRP is ARP, with the hub acting as an answering service instead of broadcast. The routing protocol is just the routing protocol, doing exactly what it does on any LAN. And a Phase 3 shortcut is the moment two hosts stop asking the router to relay and start talking directly. The analogy even predicts the failure modes. ARP broken means reachable IP but no delivery; NHRP broken means tunnel up but state stuck in NHRP. Wrong VLAN tag means silent drops; wrong tunnel key means exactly the same thing. When you troubleshoot DMVPN (walked through failure by failure in [troubleshooting DMVPN](https://www.pinglabz.com/troubleshooting-dmvpn/)), you are really just asking which LAN-equivalent layer broke. ## Key Takeaways DMVPN is mGRE plus NHRP plus a routing protocol, and each answers one question the others cannot. mGRE provides a single tunnel interface with no fixed destination. NHRP resolves overlay addresses to underlay addresses the way ARP resolves IPs to MACs, with spokes registering to a hub instead of broadcasting. The routing protocol distributes LAN prefixes, which NHRP never carries. Multicast-to-unicast mapping is what lets routing adjacencies form over the overlay at all, and IPsec bolts on underneath without changing any of it. Build the full thing yourself with the [Phase 3 configuration walkthrough](https://www.pinglabz.com/dmvpn-phase-3-configuration/), or go back up to the [DMVPN complete guide](https://www.pinglabz.com/dmvpn/) for the whole cluster. ### Troubleshooting IP Multicast: Reading show ip mroute, RPF Failures, and Silent Receivers URL: https://www.pinglabz.com/troubleshooting-ip-multicast/ Last updated: 2026-08-01T19:26:54.000Z Multicast fails silently. No ICMP unreachable, no log message, no down interface: the stream just is not there. The good news is that nearly every "multicast is broken" ticket collapses into one of three signatures, and each one is identifiable from show output in under a minute once you know what to look for. Part of the [IP multicast guide](https://www.pinglabz.com/multicast/). Everything below is real output from CML labs on iol-xe nodes running IOS XE 17.18.2: one built healthy and captured twice, before and after the source went active, and a second deliberately sabotaged one failure at a time. You get both, because broken output is meaningless until you know what the working output should have said. The skill that carries all of it is reading `show ip mroute` properly: the flags, the incoming interface, the outgoing interface list, and what an empty list means. ## The Debugging Model: Follow the Tree A multicast stream needs three things end to end: membership at the receiver edge (IGMP), a delivery tree between first-hop and last-hop routers (PIM state that passes the RPF check), and agreement on the RP for shared-tree groups. Check the two ends first, because they localize the fault: - **Last-hop router:** `show ip igmp groups`. Is the receiver's join actually registered? - **First-hop router:** `show ip mroute `. Is the source actually sending, and is its (S,G) forwarding or pruned? - **Then walk the middle** with `show ip mroute` and `show ip rpf`, upstream from the receiver, hunting for a Null OIL or a wrong incoming interface. The three cases below are the three ways that walk ends. First, the baseline. ## The Healthy Baseline: What Working PIM-SM Looks Like PlatformCML, iol-xe nodes, IOS XE 17.18.2 Healthy lab topologyR1 (source, Lo0 1.1.1.1) to R2 (RP, Lo0 2.2.2.2) to R3 (receiver) Unicast / multicastOSPF area 0, `ip pim sparse-mode` on all links and loopbacks, static RP Group and receiver239.1.1.1, R3 joins with `ip igmp join-group`, R1 sources with a multicast ping Sabotage labSeparate topology, roles reversed: R3 is the first hop, R1 the last hop, source 10.0.30.10 Two commands establish that the control plane is in the game at all. First, RP agreement, captured from R2 which is itself the RP: ``` R2# show ip pim rp mapping PIM Group-to-RP Mappings Group(s): 224.0.0.0/4, Static RP: 2.2.2.2 (?) ``` One static RP covering the whole 224.0.0.0/4 range (the `(?)` is a failed reverse DNS lookup, not an error). Second, PIM adjacency: ``` R2# show ip pim neighbor Neighbor Interface Uptime/Expires Ver DR Prio/Mode 10.0.12.1 Ethernet0/0 00:01:08/00:01:35 v2 1 / S P G 10.0.23.2 Ethernet0/1 00:01:04/00:01:39 v2 1 / DR S P G ``` Both core links have a PIM neighbor. That is the check people skip and the cheapest on the list: an interface with no PIM neighbor has no PIM, the tree cannot cross it, and a link missing here is a forgotten `ip pim sparse-mode` at one end. For why sparse mode builds trees on demand instead of flooding, see [how PIM sparse mode actually builds a tree](https://www.pinglabz.com/pim-sparse-mode-explained/). ## How to Read show ip mroute Every entry has the same four-part anatomy: the entry key, the flags, the incoming interface (IIF), and the outgoing interface list (OIL). Here is the table on the RP before the source has sent a single packet, when the receiver's IGMP join is the only thing driving state: ``` R2# show ip mroute (*, 239.1.1.1), 00:01:05/00:03:03, RP 2.2.2.2, flags: SJC Incoming interface: Null, RPF nbr 0.0.0.0 Outgoing interface list: Ethernet0/1, Forward/Sparse, 00:01:05/00:03:03, flags: (*, 224.0.1.40), 00:01:16/00:03:22, RP 2.2.2.2, flags: SJCL Incoming interface: Null, RPF nbr 0.0.0.0 Outgoing interface list: Ethernet0/0, Forward/Sparse, ... Ethernet0/1, Forward/Sparse, ... Loopback0, Forward/Sparse, ... ``` `(*, 239.1.1.1)` is the shared tree: any source, this group, rooted at the RP. The timers are uptime and expiry, and `RP 2.2.2.2` is which RP this router believes roots the tree, a field that catches an entire class of failure on its own (see case 3). Now the part that trips people up. The incoming interface is **Null** and the RPF neighbor is **0.0.0.0**, and that is correct here, because this router *is* the RP: nothing sits upstream of the root of a shared tree for the IIF to point at. The same output on a router that is not the RP would be a real problem, so read the IIF in the context of where you are standing. The OIL has one entry, Ethernet0/1 toward the receiver, in Forward/Sparse: the router saying "I will replicate this group out this interface." (The 224.0.1.40 entry alongside it is the Auto-RP Discovery group, and its L flag means the router itself is a member.) Now start the source: ``` R2# show ip mroute (*, 239.1.1.1), 00:02:24/00:02:41, RP 2.2.2.2, flags: SJC Incoming interface: Null, RPF nbr 0.0.0.0 Outgoing interface list: Ethernet0/1, Forward/Sparse, 00:02:24/00:02:41, flags: (1.1.1.1, 239.1.1.1), 00:00:49/00:02:44, flags: T Incoming interface: Ethernet0/0, RPF nbr 10.0.12.1 Outgoing interface list: Ethernet0/1, Forward/Sparse, 00:00:49/00:02:56, flags: ``` A second entry appeared, and this is the whole of PIM sparse mode in one screen. `(1.1.1.1, 239.1.1.1)` is the source tree: this specific source, this group. Its flag is **T**, the SPT bit, meaning the router switched from pulling this group down the shared tree to receiving it on the shortest path back to the source. Its IIF is a real interface now, Ethernet0/0 with RPF neighbor 10.0.12.1, because something genuinely is upstream: R1. The flags are the fastest read in the output. The ones you will meet: SSparse mode entry. Expected on every entry in a PIM-SM network. JJoin SPT. The router is eligible to switch, or is switching, to the source tree. CA directly connected receiver exists behind an OIL interface. Someone downstream wants this. LLocal. The router itself is a member of the group. TSPT bit set. Traffic for this (S,G) is arriving on the shortest path tree, not the shared tree. PPruned. Nothing downstream wants this traffic, or the router cannot deliver it. Nearly always paired with a Null OIL. FRegister. This is the first-hop router for a directly connected source, registering it to the RP. ACandidate for MSDP advertisement. Seen on the RP. XProxy join timer running. Usually appears next to state that is not resolving cleanly. **An empty OIL is the single most important thing to be able to interpret.** "Outgoing interface list: Null" means this router will replicate the group to nobody, and that is not automatically a fault. It has three causes, and telling them apart is the diagnosis: - **Nobody downstream asked.** No IGMP membership and no PIM join arrived, so the entry was pruned. Flags include P and the network is behaving correctly. Case 1. - **The router cannot use the traffic it is receiving.** Receivers exist and joins arrived, but the entry fails its RPF check, so the OIL is torn down and the entry sits pruned. Case 2. - **The join could never be sent upstream.** The router heard the receiver but has no RP, so it has no direction to send a (\*,G) join. The OIL may even be populated while the IIF is Null. Case 3. Finally, prove packets are moving rather than that state merely exists: ``` R2# show ip mroute count IP Multicast Statistics 3 routes using 4456 bytes of memory 2 groups, 0.50 average sources per group Group: 239.1.1.1, Source count: 1, Packets forwarded: 20, Packets received: 21 RP-tree: Forwarding: 1/0/100/0, Other: 2/1/0 Source: 1.1.1.1/32, Forwarding: 19/0/100/0, Other: 19/0/0 ``` Twenty packets forwarded, which no amount of correct-looking control plane state can fake. The `Other` triplet is total, **RPF failed**, and other drops, in that order. The source entry reads 19/0/0: zero RPF failures. The RP-tree entry reads 2/1/0, one packet failing RPF during the switch to the source tree, which is the normal transient. A counter that keeps climbing in that middle field is the numeric confirmation of the next case. ## Case 1: The Silent Receiver (No IGMP Join) Symptom: the application team swears the receiver is subscribed; no traffic arrives. In the sabotage lab a source was started for 239.2.2.20 with no receiver joined anywhere. The first hop (R3) shows the source is real and registering: ``` R3# show ip mroute 239.2.2.20 (*, 239.2.2.20), 00:00:27/stopped, RP 2.2.2.2, flags: SPF Incoming interface: Ethernet0/0, RPF nbr 10.0.23.1 Outgoing interface list: Null (10.0.30.10, 239.2.2.20), 00:00:27/00:02:32, flags: PFT Incoming interface: Ethernet0/1, RPF nbr 0.0.0.0 Outgoing interface list: Null ``` The RP (R2) also knows the source, and is equally starved of receivers: ``` R2# show ip mroute 239.2.2.20 (*, 239.2.2.20), 00:00:18/stopped, RP 2.2.2.2, flags: SP Outgoing interface list: Null (10.0.30.10, 239.2.2.20), 00:00:18/00:02:41, flags: PA Incoming interface: Ethernet0/1, RPF nbr 10.0.23.2 Outgoing interface list: Null ``` Read the signature with the flag key above: source-side state exists everywhere it should (F at the first hop because the source is directly connected, A at the RP), but every OIL is Null and every entry carries P, where the healthy capture named a real interface in Forward/Sparse. The network is working perfectly here; nobody asked for the traffic. The confirming check is at the receiver edge, where `show ip igmp groups 239.2.2.20` returns nothing at all. Root causes, in observed frequency order: the application never actually joined (wrong group, wrong port, bound to the wrong NIC), a host firewall dropping IGMP, IGMP snooping eating reports on a querier-less segment, or a version mismatch where a v2 report for an SSM group is silently ignored. That last trap matters before you deploy [source-specific multicast without an RP](https://www.pinglabz.com/source-specific-multicast-ssm/), and the version behaviour is in [the difference between IGMPv2 and IGMPv3](https://www.pinglabz.com/igmp-explained-v2-v3/). Fix the join and the tree builds itself within a second. ## Case 2: RPF Failure (State Exists, Traffic Dies Midpath) Symptom: joins are fine, source is fine, and somewhere in the middle a router drops every packet without logging a thing. This is the multicast-specific failure, it is the reason traffic is not flowing far more often than anything else on this page, and it is the one most people misdiagnose. ### What the RPF check actually does When a multicast packet arrives, the router does not ask "where is this going." It asks "did this arrive on the interface I would use to get *back* to the source." If yes, forward out the OIL. If no, drop, silently, with no counter you will notice unless you go looking. That is loop prevention: a multicast packet gets replicated to many neighbors, so a tree with two parents would multiply traffic without bound. The critical detail is *which table* answers the question, and the healthy output prints it on the page: ``` R2# show ip rpf 1.1.1.1 RPF information for ? (1.1.1.1) RPF interface: Ethernet0/0 RPF neighbor: ? (10.0.12.1) RPF route/mask: 1.1.1.1/32 RPF type: unicast (ospf 1) Doing distance-preferred lookups across tables RPF topology: ipv4 multicast base, originated from ipv4 unicast base ``` **RPF type: unicast (ospf 1)**. The multicast forwarding decision was made by the OSPF unicast routing table, and that line explains the entire failure mode: multicast forwarding is a passenger on unicast routing, so anything that makes the unicast best path to the source diverge from the physical path the stream takes breaks multicast while leaving unicast perfectly healthy. Static routes with a better administrative distance, tunnels, asymmetric paths, ECMP, a redistribution boundary, or a [route that never made it into the routing table](https://www.pinglabz.com/ospf-routes-not-appearing-routing-table/) all do it. Note the RPF neighbor 10.0.12.1 too: that address must match the router the stream physically arrives from. ### Multicast RPF is not unicast RPF These get conflated constantly, because both names contain "reverse path forwarding" and both consult the unicast table. They are not the same feature. Multicast RPF is a mandatory, always-on part of the PIM forwarding path: you cannot turn it off, it applies to every multicast packet, and it prevents loops in a replicating topology. Unicast RPF (uRPF) is an optional per-interface anti-spoofing feature you enable in [strict or loose mode to drop forged source addresses](https://www.pinglabz.com/unicast-rpf-strict-loose/). So when somebody says "we do not run RPF here," they mean uRPF, and your multicast is still being RPF-checked on every hop regardless. ### The broken fingerprint The sabotage lab breaks it with a bad static route on the last-hop router (the check itself is dissected in [how the multicast RPF check decides the incoming interface](https://www.pinglabz.com/multicast-rpf-check/)), producing this: ``` R1# show ip rpf 10.0.30.10 RPF interface: Ethernet0/0 RPF neighbor: ? (192.168.99.100) RPF type: unicast (static) R1# show ip mroute 239.1.1.10 (10.0.30.10, 239.1.1.10), 00:00:43/00:02:15, flags: PJX Incoming interface: Ethernet0/0, RPF nbr 192.168.99.100 Outgoing interface list: Null ``` The tells, in order of diagnostic value: the (S,G) incoming interface points somewhere the stream cannot possibly arrive from (here, out the receiver-facing edge), the OIL is Null with a P flag *despite an active receiver*, which is the distinction from case 1, and the `RPF type` line reads `unicast (static)` rather than `unicast (ospf 1)`, naming the static route that shadowed the IGP. Traffic arrives on Ethernet0/1, the router demands Ethernet0/0, every packet fails the check. The fix is either repairing unicast routing or, when the asymmetry is intentional, a static mroute that overrides RPF for multicast only: ``` R1(config)# ip mroute 10.0.30.0 255.255.255.0 10.0.12.2 R1# show ip rpf 10.0.30.10 RPF interface: Ethernet0/1 RPF type: multicast (static) R1# show ip mroute 239.1.1.10 (10.0.30.10, 239.1.1.10), 00:01:18/00:01:41, flags: JT Incoming interface: Ethernet0/1, RPF nbr 10.0.12.2, Mroute ``` The RPF type changed from `unicast (static)` to `multicast (static)`: the static mroute table now answers instead of the unicast one. P became JT, matching the healthy T-flag entry from the baseline, the OIL repopulated, and the stream resumed. Suspect this case first whenever tunnels, ECMP, or asymmetric routing exist in the path. ## Case 3: Missing or Mismatched RP Symptom: receivers join, sources send, and the two never meet. For shared-tree (ASM) groups both sides depend on agreeing where the RP is, and a router with no RP mapping cannot even send its join upstream. The lab removes the RP configuration from the last-hop router and has the receiver join 239.5.5.5: ``` R1# show ip pim rp mapping PIM Group-to-RP Mappings R1# show ip mroute 239.5.5.5 (*, 239.5.5.5), 00:00:16/00:02:45, RP 0.0.0.0, flags: SJC Incoming interface: Null, RPF nbr 0.0.0.0 Outgoing interface list: Ethernet0/0, Forward/Sparse, 00:00:16/00:02:45 ``` **RP 0.0.0.0** is the entire diagnosis in four characters, and it is worth reading next to the healthy baseline because the two look deceptively similar: both show a (\*,G) with a Null incoming interface and RPF nbr 0.0.0.0\. On the RP that is correct. Here it is fatal, because this router is not the RP and the RP field reads 0.0.0.0 instead of an address. The rp mapping output agrees: the healthy version listed a group range and an RP, this one is a header with nothing under it. The IGMP join was heard, which is why the group exists with a populated OIL and the C flag, but with no RP to root the shared tree there is no upstream direction to send the (\*,G) join. Check `show ip pim rp mapping` on every router in the path; they must all resolve the same RP. Partial RP knowledge (some static, some learned, one forgotten) produces the maddening variant where multicast works from some sources and not others. The related failure is an RP that is known but unreachable, its loopback missing from the IGP, which shows up as a valid mapping plus an RPF failure toward the RP address itself; `show ip rpf 2.2.2.2` catches that one. Auto-RP and BSR remove the per-router touch points but add a dependency: Auto-RP's 224.0.1.39 and 224.0.1.40 groups must themselves be forwarded, solved by `ip pim autorp listener`. ## When There Is No mroute At All Everything so far assumes `show ip mroute` returned something. Sometimes it returns nothing, or a line reading "IP Multicast Forwarding is not enabled," and none of the above applies because the router is not participating in multicast at all. Three checks: - `ip multicast-routing` is missing from the global config. - It was typed but rejected, which is subtler and cost real time in this lab. On iol-xe the `distributed` keyword is rejected outright, so a config copied from a platform where `ip multicast-routing distributed` is valid leaves forwarding switched off and never complains again. - No interface has `ip pim sparse-mode`, so even with forwarding enabled there is no PIM to build state. The minimal configuration behind every healthy capture above is four lines: ``` ip multicast-routing interface Ethernet0/0 ip pim sparse-mode ip pim rp-address 2.2.2.2 ``` Repeat the interface stanza on every interface in the path, including the loopback if the RP address lives on one, and the rp-address line on every router. ## The Supporting Cast: Neighbors, Snooping, and Debugs Three more checks catch the failures that do not fit the big three cleanly. First, PIM adjacency. A single interface missing `ip pim sparse-mode` breaks the join chain at that hop, and everything downstream looks like case 3 while everything upstream looks healthy. `show ip pim neighbor` hop by hop finds the gap fast, compared against the healthy output above: one neighbor per core link, uptimes stable rather than resetting. Second, layer 2\. IGMP snooping without a querier is the classic "works for two minutes, then dies" pattern: snooped state ages out with no queries to refresh it and the switch quietly stops forwarding the group. `show ip igmp snooping groups` on the switch, plus a snooping querier on router-less segments, closes that case. The nastier variant is old firmware that drops IGMPv3 reports entirely, downgrading every SSM receiver behind it. From the router both look exactly like case 1, which is why the L2 check belongs in the routine. Third, targeted debugs when show output stalls. `debug ip igmp` shows joins arriving (or not) at the last hop in real time; `debug ip pim` shows join/prune and register traffic, and running it on the first hop and the RP at once answers "are registers sent" and "are they answered" in one pass. Both are group-filterable with an ACL, so they are safe on busy boxes if scoped. ## The Command Kit show ip mroute The tree itself. Read IIF, OIL, and flags. Null OIL + P = nobody downstream or RPF-dead; RP 0.0.0.0 = no RP known. show ip rpf What the RPF check believes, and which table told it. Compare the interface against where the stream physically arrives. show ip igmp groupsMembership truth at the receiver edge. Empty for the group = silent receiver, stop looking at PIM. show ip pim rp mappingRP agreement check. Must match on every router in the path for ASM groups. An empty header means no RP at all. show ip mroute countPackets forwarded, and the Other triplet whose middle field is RPF failed. The numeric confirmation for case 2. show ip pim neighborA missing PIM adjacency anywhere in the path explains everything downstream of it. ## A Worked Ticket, Start to Finish How the method composes on a real ticket: "the 239.1.1.10 dashboard feed stopped for the branch behind R1; other sites are fine." Start at the receiver edge. `show ip igmp groups 239.1.1.10` on R1 shows membership present and refreshing, so case 1 is gone in thirty seconds. R1's mroute has a healthy (\*,G) but an (S,G) with P, a Null OIL, and an incoming interface that makes no sense for where the source lives: case 2's fingerprint. `show ip rpf 10.0.30.10` reports Ethernet0/0 when the stream can only arrive on Ethernet0/1, and the `RPF type` line reads `unicast (static)` in a network that runs OSPF, which narrows it to one route before you have looked at the routing table. `show ip route 10.0.30.0` then names it: a static with administrative distance 1 shadowing the OSPF route, left over from last month's maintenance. Remove it (or add a static mroute, had the asymmetry been intentional) and the confirmation is the (S,G) flags flipping from PJX back to JT, the OIL repopulating, and `show ip mroute count` incrementing "Packets forwarded" while the RPF-failed field stops climbing. Bisecting with pings instead would have burned the afternoon, because unicast pings to the source worked perfectly throughout. Unicast reachability proves nothing about multicast health: the two forwarding planes consult the same table with opposite questions. ## Gotchas From the Lab Bench - **Sender TTL.** Host stacks default the multicast TTL to 1, which the first router burns, producing a perfect imitation of case 1 (source-side state, nothing anywhere else, and even the first hop's F flag missing since packets never leave the segment). Check the sending socket before blaming the network. - **A wrong keyword can leave multicast silently off.** The `distributed` variant of `ip multicast-routing` is rejected on iol-xe, and it surfaces much later as an empty mroute table rather than when you typed it. - **Null incoming interface is not automatically a fault.** On the RP it is correct for every (\*,G). Anywhere else, check the RP field on the same line before you panic. The same goes for a non-zero RPF-failed counter: one or two during an SPT switchover is normal, and only the climbing counter means something. - **Control plane state is not forwarding.** An entry with a populated OIL and a T flag can still be moving zero packets, and only `show ip mroute count` settles it. Keep a healthy capture from your own platform for comparison: half the entries in a broken table look wrong to someone who has never read a correct one. ## Key Takeaways - Read `show ip mroute` in four parts: the entry key ((\*,G) shared tree or (S,G) source tree), the flags, the incoming interface, and the outgoing interface list. - An empty OIL has three causes: nobody asked, the entry failed RPF, or no RP is known. Telling them apart is the diagnosis. - The three signatures are silent receiver, RPF failure (wrong IIF, P flag, Null OIL despite active receivers), and missing RP (RP 0.0.0.0 in the (\*,G) entry). - Suspect RPF first. `show ip rpf` prints the table that answered ("RPF type: unicast (ospf 1)"), naming the culprit routing source as well as the interface. Multicast RPF is mandatory and always on, unlike the optional uRPF anti-spoofing feature it gets confused with. - Work the two edges first: `show ip igmp groups` at the last hop, `show ip mroute` at the first hop. Then verify RP agreement on every hop, and that the RP address itself passes RPF. - Before any of it: confirm `ip multicast-routing` is actually enabled and the sender's TTL is greater than 1. The rest of the cluster, from addressing and IGMP through PIM and SSM, is indexed in the [complete multicast guide](https://www.pinglabz.com/multicast/). ### Bidirectional PIM and MSDP: Scaling Multicast Across Domains URL: https://www.pinglabz.com/bidir-pim-msdp-explained/ Last updated: 2026-07-11T18:03:42.000Z Standard PIM sparse mode scales beautifully for a handful of sources and thousands of receivers. Push it the other direction, thousands of sporadic senders, or stretch it across administrative boundaries, and two very different tools take over: bidirectional PIM collapses per-source state inside a domain, and MSDP carries source knowledge between domains (and between redundant RPs). This article covers both, with a live MSDP peering between two RPs in the lab and a real SA cache capture. Part of the [IP multicast guide](https://www.pinglabz.com/multicast/). ## The State Problem Every active source in PIM-SM costs the network (S,G) entries on every router along its tree (the machinery covered in [PIM sparse mode explained](https://www.pinglabz.com/pim-sparse-mode-explained/)). For a trading floor app or a sensor fabric where ten thousand endpoints each transmit occasionally to the same group, that is ten thousand source trees, plus a register burst and an SPT decision for each. The RP becomes a state bottleneck, and most of those trees carry a few packets an hour. Bidirectional PIM's answer is radical: keep only the shared tree, forever, and make it carry traffic both directions. ## Bidirectional PIM: One Tree, Both Ways In bidir PIM, a group's entire life happens on a single (\*,G) tree rooted at the RP. Receivers join it downward exactly as in PIM-SM. Sources send upward along it: any router receiving multicast for a bidir group forwards it toward the RP, no registration, no (S,G) state, no RPF check against the source, and traffic flows down every branch with interested receivers as it passes. Three consequences: - **State is O(groups), not O(sources x groups).** Ten thousand senders to one group cost exactly one mroute entry per router. This is the entire point. - **Paths are never optimal.** Everything transits the RP's tree, so source-to-receiver latency is whatever the shared tree gives you. Bidir trades path quality for state, the opposite of the SPT switchover. - **Loop prevention changes.** With traffic flowing both directions on one tree, standard source-RPF cannot work. Instead each segment elects a designated forwarder (DF), the router with the best route to the RP, and only the DF moves packets toward the RP. One forwarder per link means no loops, enforced by election rather than per-packet source checks. (The [RPF check](https://www.pinglabz.com/multicast-rpf-check/) still governs the tree toward the RP itself.) Configuration is `ip pim bidir-enable` plus flagging the RP's group ranges as bidir (`ip pim rp-address 2.2.2.2 bidir`). Every router in the domain must support and enable it. In `show ip mroute`, bidir groups carry the B flag, and notably the RP for a bidir group can be a phantom, just a routable address on the tree with no actual router owning it, since nothing ever registers to it. Use bidir when a group's sender count is large, sporadic, and many-to-many (collaboration, presence, telemetry ingestion); use SSM when sources are few and known (see [the SSM article](https://www.pinglabz.com/source-specific-multicast-ssm/)); standard PIM-SM covers the middle. ## MSDP: Source Knowledge Across Boundaries Inside one PIM domain, the RP knows every active source because first-hop routers register them to it. That knowledge stops at the domain edge: a receiver in your network joining group G has no way to learn about a source registered to a different organization's RP, or even to a second RP in your own network. MSDP (Multicast Source Discovery Protocol) is the bridge: RPs peer over TCP 639 and tell each other about active sources with Source-Active (SA) messages, each carrying (source, group, originating RP). The lab runs two PIM domains: R2 is the RP for domain 1 (where the source 10.0.30.10 lives), R4 is the RP for domain 2, and the two peer directly: ``` ! R2 (RP, domain 1) ip msdp peer 4.4.4.4 connect-source Loopback0 ip msdp originator-id Loopback0 ! R4 (RP, domain 2) ip msdp peer 2.2.2.2 connect-source Loopback0 ip msdp originator-id Loopback0 ``` ``` R2# show ip msdp summary Peer Address AS State Uptime/ Reset SA Peer Name Downtime Count Count 4.4.4.4 ? Up 00:01:03 0 0 ? ``` When sources in domain 1 went active (they registered to R2, whose (S,G) entries carry the A flag, "candidate for MSDP advertisement"), R2 originated SA messages and R4's cache filled with real cross-domain source knowledge: ``` R4# show ip msdp sa-cache MSDP Source-Active Cache - 2 entries (10.0.30.10, 239.1.1.10), RP 2.2.2.2, AS ?,00:05:40/00:05:41, Peer 2.2.2.2 (10.0.30.10, 239.2.2.20), RP 2.2.2.2, AS ?,00:00:35/00:05:41, Peer 2.2.2.2 ``` Now if a receiver in domain 2 joins 239.1.1.10, R4 checks its SA cache, finds the source, and sends an (S,G) join across the domain boundary toward 10.0.30.10\. Traffic then flows natively; MSDP moves knowledge, never data. In true interdomain deployments (between autonomous systems) the (S,G) join follows MBGP multicast-safe paths and SA floods are RPF-checked peer-by-peer to prevent loops; inside one organization, a simple mesh or hub of MSDP peers suffices. SA entries age out about six minutes after the source stops, so `show ip msdp sa-cache` is also a liveness view of remote sources. ## DF Election and the Phantom RP Bidir's designated forwarder deserves more than a bullet. On every link in the domain (including host segments), the routers compare their unicast routes to the RP address, and the one with the best metric (tie: highest IP) wins DF duty for that link. Only the DF forwards traffic from the link toward the RP, and only the DF forwards traffic from the RP tree onto the link. One forwarder per link, both directions, elected once instead of checked per packet: that is how bidir stays loop-free while abandoning per-source RPF. DF election happens per RP, so multiple bidir RPs mean multiple elections; a link's DF for one group range can differ from its DF for another. The phantom RP is the design trick that follows. Since nothing registers to a bidir RP and no SA state lives on it, the "RP" only needs to be a routable address that DF elections can measure distance to. The standard implementation puts the address on a loopback subnet advertised by two routers at different prefix lengths (a /30 and a /29 covering the same address), so the longest-prefix route wins and failover is pure IGP convergence with no protocol awareness at all: ``` ! Router A (primary) ! Router B (backup) interface Loopback1 interface Loopback1 ip address 10.255.1.2 255.255.255.252 ip address 10.255.1.2 255.255.255.248 ! RP address = 10.255.1.1 on all routers (no router owns it) ip pim rp-address 10.255.1.1 BIDIR-GROUPS bidir ``` An RP that is literally an unassigned IP address is the purest statement of what bidir changed: the RP became a routing anchor, not a device. ## Operating MSDP: Filters, Limits, and the Junk Problem MSDP's operational reputation was earned in the early 2000s when worms discovered that scanning multicast groups generated SA storms across the entire internet's RP mesh. The lessons live on as standard practice. Filter what you originate (`ip msdp redistribute list`) so internal-only groups in 239/8 never leak into SA messages; filter what you accept (`ip msdp sa-filter in list ...`) so a sloppy peer cannot fill your cache; and cap the blast radius with `ip msdp sa-limit` per peer. Well-known junk (SSDP's 239.255.255.250, mDNS) belongs in every SA filter you ever deploy. Verification mirrors BGP operations, which MSDP deliberately resembles: `show ip msdp summary` for peer state and SA counts (shown above), `show ip msdp peer` for the TCP details and applied filters, `show ip msdp sa-cache` for the actual learned sources, and `debug ip msdp detail` when an SA you expected never arrives. The most common real fault is an SA that exists on the origin RP but never appears at the peer, and it is almost always an outbound redistribute filter or the origin's (S,G) losing its A flag because the source went quiet. ## Anycast RP: MSDP's Day Job The most common MSDP deployment never crosses an organizational boundary. Anycast RP solves RP redundancy: configure the same RP address (say 10.255.255.1/32 on a loopback) on two or more routers, advertise it in the IGP, and every router's `ip pim rp-address 10.255.255.1` resolves to the closest instance automatically. Registers and joins land on different RPs depending on where they originate, so the RPs must share source knowledge, and that is exactly an MSDP full mesh between them (peered on separate, unique loopbacks, never the anycast address): ``` ! RP-A ! RP-B interface Loopback1 interface Loopback1 ip address 10.255.255.1 255.255.255.255 ip address 10.255.255.1 255.255.255.255 ip msdp peer 2.2.2.2 connect-source Loopback0 ip msdp originator-id Loopback0 ! (mirrored on each RP) ``` Failure handling comes free from the IGP: when an RP dies, its /32 disappears and every router converges to the survivor at unicast-routing speed, no PIM reconfiguration anywhere. Before this pattern, RP failover meant BSR timers; anycast RP made sub-second RP redundancy routine, and it remains the standard design in IOS XE deployments (NX-OS offers a PIM-native anycast alternative, RFC 4610, without MSDP). ## Interdomain ASM, Assembled For completeness, here is how the full interdomain ASM machine fits together when the destination is another autonomous system rather than a second internal domain. MBGP (BGP's IPv4 multicast address family) advertises which prefixes are reachable for RPF purposes across the boundary, keeping multicast tree-building decoupled from unicast forwarding policy. MSDP rides on top, RP to RP, telling each domain which sources are alive. When a local receiver joins group G, the local RP finds a matching SA in its cache, sends the (S,G) join toward the source along MBGP-derived paths, and native traffic flows across the AS boundary on the source tree. Three protocols, three jobs: MBGP answers "which way", MSDP answers "who is sending", PIM builds the tree. In SSM deployments the first two disappear, which is the strongest argument in SSM's favor whenever interdomain multicast comes up; the comparison is drawn out in the [SSM article](https://www.pinglabz.com/source-specific-multicast-ssm/). ## Choosing the Tool PIM-SM (default) Few-to-many, sources unknown to receivers. RP for introductions, SPT for delivery. SSM One-to-many, sources known. No RP at all; skip MSDP entirely, even interdomain. Bidir PIM Many-to-many, huge sporadic sender counts. One two-way shared tree, O(groups) state, DF election for loops. MSDP ASM across domains, and anycast RP redundancy inside one. Moves source knowledge, never traffic. ## Key Takeaways - Bidir PIM keeps every group on one two-way shared tree: state collapses from per-source to per-group, paths run through the RP forever, and designated forwarder election replaces source RPF. - MSDP peers RPs over TCP 639 and floods Source-Active messages; each SA is (source, group, RP), cached and aged. Data never rides MSDP. - `show ip msdp sa-cache` on a remote RP is the proof that cross-domain source discovery works; the A flag on the originating RP's (S,G) entries marks what gets advertised. - Anycast RP is MSDP's most common job: identical RP address on multiple routers, IGP picks the closest, MSDP mesh keeps their source knowledge synchronized, failover at IGP speed. - SSM makes MSDP unnecessary because receivers already know sources; interdomain ASM is the case that genuinely needs it. Round out the cluster with [troubleshooting IP multicast](https://www.pinglabz.com/troubleshooting-ip-multicast/), or return to the [complete multicast guide](https://www.pinglabz.com/multicast/). ### Source-Specific Multicast (SSM): PIM Without an RP URL: https://www.pinglabz.com/source-specific-multicast-ssm/ Last updated: 2026-07-11T18:03:42.000Z Everything hard about PIM sparse mode traces back to one assumption: receivers do not know who the sources are. Drop that assumption and the rendezvous point, the register tunnel, the shared tree, and the SPT switchover all become unnecessary. That is Source-Specific Multicast: the receiver names the source, the network builds the (S,G) tree directly, and half the protocol machinery disappears. This article covers the model, the configuration, and live IOS XE captures of SSM state, including a real IGMPv3 gotcha that cost this lab twenty minutes. Part of the [IP multicast guide](https://www.pinglabz.com/multicast/). ## ASM vs SSM: The Service Model Difference Classic any-source multicast (ASM) sells receivers a group: join 239.1.1.10 and receive whatever any source sends to it. The network must discover sources on the receiver's behalf, which is the entire job of the RP in [PIM sparse mode](https://www.pinglabz.com/pim-sparse-mode-explained/). SSM sells a channel: the pair (S,G). A receiver subscribes to "10.0.30.10 sending to 232.1.1.10", named explicitly in its IGMPv3 membership report. Since the receiver supplies the source address, the last-hop router can send a PIM (S,G) join straight along the shortest path immediately. No RP consulted, no register process, no shared tree, no switchover: the tree is born as the SPT. The trade: receivers must learn source addresses out of band. For IPTV middleware, market data feed directories, and anything with a control channel, that is trivially available, which is why one-to-many production streams are exactly where SSM dominates. ## What SSM Eliminates No RP No RP engineering, no BSR/Auto-RP distribution, no RP failure domain, no anycast complexity. The biggest single simplification. No (\*,G) state Only (S,G) entries exist. Every tree is a shortest-path tree from birth; there is nothing to switch over from. No register process First-hop routers never tunnel packets to anyone. Sources just send; trees find them. Free anti-spoofing Receivers get traffic only from the source they named. A rogue sender to the same group is simply never joined. Security by architecture. ## Configuration: Two Lines SSM needs PIM sparse mode on the interfaces (as always), plus a declaration of the SSM range and IGMPv3 on receiver-facing interfaces: ``` ! every router in the domain ip pim ssm default ! "default" = 232.0.0.0/8 ! receiver-facing interfaces (IOS XE defaults to IGMPv2) interface Ethernet0/0 ip igmp version 3 ``` `ip pim ssm default` tells the router to treat 232/8 by SSM rules: never build (\*,G) state for those groups, never consult an RP, and ignore sourceless (v2-style) joins for the range. A custom range is possible (`ip pim ssm range `), but 232/8 is IANA-reserved for exactly this and there is rarely a reason to deviate. ## Live State: A Real SSM Join In the lab, the Debian VM subscribes to channel (10.0.30.10, 232.1.1.10) with an IGMPv3 INCLUDE-mode report (in Python: `IP_ADD_SOURCE_MEMBERSHIP`). The edge router's IGMP table shows source-filtered membership, something IGMPv2 cannot express: ``` R1# show ip igmp groups 232.1.1.10 detail Interface: Ethernet0/0 Group: 232.1.1.10 Flags: SSM Group mode: INCLUDE Last reporter: 169.254.147.59 Source Address Uptime v3 Exp CSR Exp Fwd Flags 10.0.30.10 00:00:08 00:02:51 stopped Yes R ``` And the mroute table holds exactly one entry, source tree only, born fully formed: ``` R1# show ip mroute 232.1.1.10 (10.0.30.10, 232.1.1.10), 00:00:08/00:02:51, flags: sTI Incoming interface: Ethernet0/1, RPF nbr 10.0.12.2 Outgoing interface list: Ethernet0/0, Forward/Sparse, 00:00:08/00:02:51 ``` Read the flags: lowercase s marks the SSM range, T confirms SPT traffic, I records that the state came from a source-specific host report. No (\*,G) exists, and `show ip pim rp mapping` plays no part in any of it. Upstream at the source's first-hop router, the same story with no F flag, because registration never happens in SSM: ``` R3# show ip mroute 232.1.1.10 (10.0.30.10, 232.1.1.10), 00:00:30/00:02:59, flags: sT Incoming interface: Ethernet0/1, RPF nbr 0.0.0.0 Outgoing interface list: Ethernet0/0, Forward/Sparse, 00:00:30/00:02:59 ``` And the actual receiver, a real Linux host pulling the channel end to end across four routers: ``` j@llmbits:~$ python3 /tmp/mcast_recv_ssm.py 232.1.1.10 10.0.30.10 5001 SSM join (10.0.30.10, 232.1.1.10):5001 - IGMPv3 source-specific membership active rx 33 bytes from 10.0.30.10: PingLabz mcast 232.1.1.10 pkt 763 rx 33 bytes from 10.0.30.10: PingLabz mcast 232.1.1.10 pkt 764 rx 33 bytes from 10.0.30.10: PingLabz mcast 232.1.1.10 pkt 765 ``` ## The IGMPv3 Fallback Gotcha (Learned Live) Building this capture hit a wall worth documenting: the VM's SSM join produced zero state on R1\. Nothing in `show ip igmp groups`, no mroute entry, no errors anywhere. The wire told the story: ``` j@llmbits:~$ sudo tcpdump -i ens224 -c 3 -n igmp 10:41:03.46 IP 192.168.99.1 > 224.0.0.1: igmp query v3 10:41:05.06 IP 169.254.147.59 > 232.1.1.10: igmp v2 report 232.1.1.10 ``` The router queries at v3, but the host reports at v2\. The receiver interface had been an IGMPv2 querier earlier in the lab session, and a v3 host that ever sees a v2 query drops into v2 compatibility mode and stays there for minutes, even after the querier upgrades. A v2 report has no source list, and IOS XE correctly ignores sourceless joins in the SSM range, so the subscription evaporated silently. The Linux fix is forcing v3 (`echo 3 > /proc/sys/net/ipv4/conf/all/force_igmp_version`, same for the interface) or waiting out the compat timer; the INCLUDE state above appeared seconds later. Full IGMP version mechanics are in [IGMP explained: v2 vs v3](https://www.pinglabz.com/igmp-explained-v2-v3/). The operational lesson generalizes: SSM failures are usually IGMPv3 signaling failures. Any v2-only host, v2-mode switch querier, or middlebox stripping v3 reports downgrades the segment and kills every SSM subscription on it. ## SSM Mapping: Dragging Legacy Receivers Along The IGMPv3 requirement has one escape hatch worth knowing. Plenty of embedded receivers (set-top boxes, industrial controllers, ancient imaging clients) speak only IGMPv2 and will never learn a source list. SSM mapping lets the router fill in the blank: when a v2 report arrives for a group in the SSM range, the router consults a configured map (or DNS) and manufactures the (S,G) join on the host's behalf. ``` ip igmp ssm-map enable no ip igmp ssm-map query dns ip igmp ssm-map static SSM-SOURCES 10.0.30.10 ! ip access-list standard SSM-SOURCES permit 232.1.1.0 0.0.0.255 ``` With that applied on the receiver-facing router, a v2 join for anything in 232.1.1.0/24 behaves as if the host had asked for source 10.0.30.10 explicitly. It is a bridge, not a destination: the mapping is static configuration pretending to be receiver knowledge, so treat it as technical debt with a decommission date attached to the last v2 receiver. ## Running SSM and ASM Side by Side Real networks migrate; they do not flag-day. The good news is that SSM and ASM coexist without interaction, because the split is purely address-driven: `ip pim ssm default` changes behavior only inside 232/8, and every other group keeps full RP-based sparse-mode machinery. The lab this cluster is built on runs both simultaneously: 239.1.1.10 flows through registration, shared tree, and SPT switchover, while 232.1.1.10 rides a direct (S,G) with no RP involvement, on the same routers at the same time. A sane migration sequence: enable `ip pim ssm default` everywhere first (it is inert until someone uses 232/8), roll IGMPv3 onto receiver interfaces, move the highest-volume one-to-many application onto a 232/8 channel with its source published through whatever control channel the app already has, and let the ASM estate shrink at its own pace. Each moved application removes its share of register load and RP dependency; when the last ASM group is gone, the RP configuration simply becomes deletable. The reverse ordering (renumbering groups before receivers speak v3) strands subscribers, per the compatibility trap above. ## Where SSM Fits (and Where It Does Not) SSM is the right default for one-to-many with known sources: IPTV channel lineups, market data (each feed is a published (S,G) channel), live event streaming, one-way telemetry fans. It also crosses domains trivially, since no interdomain source discovery is needed; the receiver already knows the source, making MSDP unnecessary for SSM (the ASM interdomain story is in [bidir PIM and MSDP](https://www.pinglabz.com/bidir-pim-msdp-explained/)). SSM fits poorly when sources are genuinely unknown or churn rapidly: many-to-many collaboration, service discovery patterns, or applications where anyone may transmit at any time. That is ASM's home turf, and at large many-to-many scale, bidirectional PIM's. One practical note: hosts subscribe per (S,G), so multi-source redundancy (primary and backup feed) means the application joins two channels and arbitrates, a pattern financial feed handlers implement as A/B feeds. ## Verifying an SSM Deployment The verification set for SSM is short because the moving parts are few. `show ip pim ssm` confirms the range configuration on each router. On receiver interfaces, `show ip igmp interface` must report version 3 on both the router and host lines, and `show ip igmp groups detail` must show INCLUDE mode with the expected sources; EXCLUDE mode for a 232/8 group means a host is sending v2-style joins that will be ignored. In the mroute table, healthy SSM is exactly one (S,G) per channel with the lowercase s flag, and the absence of state is itself diagnostic: no (\*,G) should ever exist for an SSM group, so finding one means the range declaration is missing somewhere. End to end, the fastest proof remains a real subscription from a real host, which is precisely thirty seconds of Python on any Linux box you own. ## Key Takeaways - SSM subscribes receivers to (S,G) channels in 232/8\. The receiver names the source, so the network needs no RP, no registers, no shared trees, and no switchover. - Config is `ip pim ssm default` domain-wide plus `ip igmp version 3` on receiver interfaces. - Healthy SSM state is unmistakable: only (S,G) entries, lowercase s and I flags, Group mode INCLUDE in the IGMP table. - SSM's failure domain is IGMPv3 signaling. The v2 compatibility fallback silently discards source-specific joins; verify versions on both router and host before anything else. - Named sources are also access control: receivers cannot be sprayed by rogue senders to the same group. The full cluster, including the ASM machinery SSM replaces, lives in the [complete multicast guide](https://www.pinglabz.com/multicast/). ### The RPF Check: How Multicast Prevents Loops URL: https://www.pinglabz.com/multicast-rpf-check/ Last updated: 2026-07-11T18:03:42.000Z Unicast routing looks forward: where is this packet going? Multicast routing looks backward: did this packet arrive from the right direction? That reversal is the reverse path forwarding check, and it is both the reason multicast never loops and the single most common reason a multicast stream silently dies. This article explains the rule, then deliberately breaks it in the lab and fixes it with a static mroute, with real output at every step. Part of the [IP multicast guide](https://www.pinglabz.com/multicast/). ## The Rule in One Sentence A router accepts a multicast packet only if it arrives on the interface the router would itself use to send unicast traffic back to the packet's source; otherwise the packet is dropped, no error generated, no ICMP sent. Why so strict? A multicast packet is addressed to a group, not a destination, so a router cannot forward it "toward" anything; it copies the packet out multiple interfaces. If two routers both copy the same packet toward each other, it loops forever, and multicast loops amplify because each pass replicates. RPF kills every such loop at the first hop: a looped packet necessarily arrives on a wrong-direction interface and fails the check. TTL would eventually stop a unicast loop; multicast cannot afford "eventually". The check runs against the unicast routing table (this is the "protocol independent" in PIM). It also determines topology: the RPF interface toward the source (or the RP, for shared trees) becomes the incoming interface of the mroute entry, and PIM joins are sent out of it. RPF is not just a filter; it is how [PIM sparse mode](https://www.pinglabz.com/pim-sparse-mode-explained/) decides where trees grow. ## Reading a Healthy RPF The lab: a Debian VM receiver behind R1, a source at 10.0.30.10 three hops away, OSPF area 0 underlay. One command shows the whole story: ``` R1# show ip rpf 10.0.30.10 RPF information for ? (10.0.30.10) RPF interface: Ethernet0/1 RPF neighbor: ? (10.0.12.2) RPF route/mask: 10.0.30.0/24 RPF type: unicast (ospf 1) Doing distance-preferred lookups across tables ``` Ethernet0/1 is where OSPF routes 10.0.30.0/24; multicast from that source must arrive there. The corresponding mroute entry is healthy, forwarding on the source tree: ``` R1# show ip mroute 239.1.1.10 (10.0.30.10, 239.1.1.10), 00:00:05/00:02:54, flags: JT Incoming interface: Ethernet0/1, RPF nbr 10.0.12.2 Outgoing interface list: Ethernet0/0, Forward/Sparse, 00:00:05/00:02:54 ``` ## Breaking It on Purpose RPF failures happen when the unicast table disagrees with where multicast actually arrives: asymmetric routing, a static route someone added for a maintenance window, tunnel interfaces, or a multicast-free path in an otherwise fine network. To reproduce the classic case, the lab gets a deliberately wrong static route on R1, pointing the source subnet out the receiver-facing interface: ``` R1(config)# ip route 10.0.30.0 255.255.255.0 192.168.99.100 ``` The static route (admin distance 1) beats OSPF (110), so the RPF lookup now resolves to the wrong interface: ``` R1# show ip rpf 10.0.30.10 RPF information for ? (10.0.30.10) RPF interface: Ethernet0/0 RPF neighbor: ? (192.168.99.100) RPF route/mask: 10.0.30.0/24 RPF type: unicast (static) ``` Multicast still arrives on Ethernet0/1, where it has always arrived. But the router now insists it should arrive on Ethernet0/0\. Every packet fails the check, and the mroute entry collapses: ``` R1# show ip mroute 239.1.1.10 (*, 239.1.1.10), 00:04:12/stopped, RP 2.2.2.2, flags: SJC Incoming interface: Ethernet0/1, RPF nbr 10.0.12.2 (10.0.30.10, 239.1.1.10), 00:00:43/00:02:15, flags: PJX Incoming interface: Ethernet0/0, RPF nbr 192.168.99.100 Outgoing interface list: Null ``` Compare with the healthy entry: the incoming interface flipped to Ethernet0/0, the P flag (pruned) appeared, and the outgoing list went Null. The stream is dead, and nothing logged anything. This exact signature, wrong incoming interface plus Null OIL, is what you hunt for in [multicast troubleshooting](https://www.pinglabz.com/troubleshooting-ip-multicast/). On platforms with `show ip mroute count`, the RPF failed drop counter climbing confirms it numerically. ## The Fix: Static Mroute Two fixes exist. The right long-term fix is repairing the unicast topology so it matches where multicast flows. But when the asymmetry is intentional (multicast engineered onto a different path than unicast, or a unicast-only firewall in the nominal path), IOS XE lets you override the RPF lookup specifically for multicast, without touching unicast routing: ``` R1(config)# ip mroute 10.0.30.0 255.255.255.0 10.0.12.2 ``` A static mroute is not a route: it never forwards anything. It is an RPF instruction: "for sources in 10.0.30.0/24, expect multicast via next hop 10.0.12.2". The RPF source flips immediately: ``` R1# show ip rpf 10.0.30.10 RPF information for ? (10.0.30.10) RPF interface: Ethernet0/1 RPF neighbor: ? (10.0.12.2) RPF type: multicast (static) RPF topology: ipv4 multicast base ``` Note the type change: multicast (static), and the topology line no longer says "originated from ipv4 unicast base". The mroute entry heals on the next join cycle, with the incoming interface tagged as mroute-derived: ``` R1# show ip mroute 239.1.1.10 (10.0.30.10, 239.1.1.10), 00:01:18/00:01:41, flags: JT Incoming interface: Ethernet0/1, RPF nbr 10.0.12.2, Mroute Outgoing interface list: Ethernet0/0, Forward/Sparse, 00:00:16/00:02:43 ``` JT flags restored, traffic flowing, unicast table still wrong and unicast routing entirely unaffected. That separation is the point: static mroutes decouple the multicast RPF topology from unicast routing where they must differ. (The broken static route was removed after this capture; the demo only needed it to exist long enough to fail.) ## Where RPF Bites in Real Networks - **Asymmetric paths.** Traffic engineering, unequal-cost designs, or policy routing put the return path somewhere multicast does not flow. The RPF interface is legal but multicast-dead. - **Tunnels and overlays.** Multicast rides a GRE tunnel (see the [GRE guide](https://www.pinglabz.com/gre/)) but unicast prefers the underlay: instant RPF failure without a static mroute or IGP adjustment over the tunnel. - **ECMP.** With equal-cost paths, the RPF tiebreak picks the neighbor with the highest IP address. Predictable, but only if you know the rule; "why does all multicast prefer that link" tickets come from here. - **Shared trees too.** RPF applies to the RP for (\*,G) state as well. If the unicast route to the RP is wrong, receivers never even build a working shared tree. ## Two Lookups, Not One: (\*,G) vs (S,G) A subtlety that matters in troubleshooting: a PIM-SM router runs RPF against different targets depending on which tree the state belongs to. For (\*,G) shared-tree entries, the check targets the RP address, because the RP is the root of that tree; for (S,G) entries, it targets the source. That means one group can be healthy on its shared tree and RPF-dead on its source tree simultaneously, which is exactly what the broken state above shows: the (\*,G) kept its correct incoming interface toward the RP while the (S,G) collapsed, because only the route to the source was poisoned. The practical consequence: when you check RPF, check it against the right address. A receiver stuck before the SPT switchover needs `show ip rpf `; a stream dying after switchover needs `show ip rpf `. And when an RP itself cannot pull traffic from a registering source, the RP's own RPF toward the source is the lookup to inspect. The order of preference IOS XE applies to each lookup is also worth knowing: static mroutes win first, then MBGP multicast address-family routes if you run them, then the plain unicast table; "distance-preferred lookups across tables" in the command output is that policy talking. ## Verification Workflow When a stream is missing, the RPF portion of the investigation is three commands on each router along the expected path, upstream from the receiver: `show ip rpf ` (does the RPF interface match where the stream physically arrives?), `show ip mroute ` (is the incoming interface right, is the OIL non-Null, are the flags sane?), and `show ip route ` (what unicast entry is feeding the lookup?). The router where the RPF interface and the actual arrival interface disagree is your culprit. Fix the routing or add the static mroute, and confirm the P flag clears. ## Key Takeaways - RPF accepts multicast only on the interface that unicast routing would use toward the source. Everything else is dropped silently. - The check is multicast's loop prevention and its topology engine: the RPF interface becomes the mroute incoming interface and the direction PIM joins travel. - The failure signature is unmistakable: wrong incoming interface, P flag, Null OIL, and zero log messages. - `ip mroute` creates a static RPF override, not a forwarding route. It fixes multicast asymmetry without touching unicast routing, and shows up as "RPF type: multicast (static)" plus an Mroute tag in the entry. - Suspect RPF first with tunnels, asymmetric designs, and ECMP (highest neighbor IP wins the tiebreak). Related: [PIM sparse mode](https://www.pinglabz.com/pim-sparse-mode-explained/) for the trees this check protects, and the [complete multicast guide](https://www.pinglabz.com/multicast/) for the full cluster. ### PIM Sparse Mode: RPs, Shared Trees, and the SPT Switchover URL: https://www.pinglabz.com/pim-sparse-mode-explained/ Last updated: 2026-07-11T18:03:41.000Z PIM sparse mode is where multicast stops being theory and becomes state machines. It is the protocol that turns "a host joined a group" into an actual delivery tree across your network, and it does it with a three-act structure: receivers join a shared tree toward a rendezvous point, sources register to that same RP, and then, the moment traffic flows, the edge quietly rebuilds everything onto shortest-path trees. This article walks the whole lifecycle on IOS XE with live captures at every stage, including the before and after of the SPT switchover. Part of the [IP multicast guide](https://www.pinglabz.com/multicast/). ## The Problem PIM-SM Solves In any-source multicast, receivers do not know who the sources are; they just want group G. Sources do not know who the receivers are; they just send. Someone has to introduce them. Dense-mode protocols solved this by flooding traffic everywhere and pruning back, which scales exactly as badly as it sounds. Sparse mode inverts the model: nothing is forwarded anywhere until someone explicitly asks, and the rendezvous point is the agreed meeting place where asks and offers converge. PIM is "protocol independent" because it builds no topology of its own. Every forwarding decision borrows the unicast routing table through the RPF check (the subject of [its own article](https://www.pinglabz.com/multicast-rpf-check/)). In this lab the underlay is [OSPF](https://www.pinglabz.com/ospf/) area 0 across four IOL-XE routers, with R2 as the RP and a real Debian VM as the receiver behind R1. ## Baseline: Neighbors and the RP PIM speaks hellos on 224.0.0.13 on every interface you enable it on. Adjacencies and the RP agreement are the two things to verify before expecting any tree to form: ``` R2# show ip pim neighbor Neighbor Interface Uptime/Expires Ver DR 10.0.12.1 Ethernet0/0 00:02:06/00:01:36 v2 1 / S P G 10.0.23.2 Ethernet0/1 00:02:04/00:01:39 v2 1 / DR S P G 10.0.24.2 Ethernet0/2 00:01:59/00:01:42 v2 1 / DR S P G R1# show ip pim rp mapping Group(s): 224.0.0.0/4, Static RP: 2.2.2.2 (?) ``` Every router in the domain must resolve the same RP for a given group, whether learned statically (as here), via BSR, or Auto-RP. Mismatched RP knowledge is failure pattern number three in [troubleshooting multicast](https://www.pinglabz.com/troubleshooting-ip-multicast/). ## Act One: A Receiver Joins the Shared Tree The Debian VM joins 239.1.1.10 with IGMP. R1, as last-hop router, creates a (\*,G) entry and sends a PIM (\*,G) join hop by hop toward the RP (this is why sparse mode is called explicit join). Real state, seconds after the join, with no source anywhere yet: ``` R1# show ip mroute 239.1.1.10 (*, 239.1.1.10), 00:00:22/00:02:39, RP 2.2.2.2, flags: SC Incoming interface: Ethernet0/1, RPF nbr 10.0.12.2 Outgoing interface list: Ethernet0/0, Forward/Sparse, 00:00:22/00:02:39 ``` Read the flags: S is sparse, C means a directly connected member (the IGMP join on Ethernet0/0). The incoming interface points at the RP because the shared tree is rooted there. The RP itself now shows the tree hanging off it: ``` R2# show ip mroute 239.1.1.10 (*, 239.1.1.10), 00:00:29/00:03:00, RP 2.2.2.2, flags: S Incoming interface: Null, RPF nbr 0.0.0.0 Outgoing interface list: Ethernet0/0, Forward/Sparse, 00:00:29/00:03:00 ``` Incoming interface Null: the RP is the root, so nothing is upstream of it on the shared tree. The OIL points back toward R1 and the waiting receiver. ## Act Two: A Source Registers Now the server behind R3 starts sending to 239.1.1.10\. R3, as the source's DR, cannot just forward: no one downstream has asked it for anything. Instead it encapsulates the first packets in unicast PIM register messages and tunnels them straight to the RP. The RP decapsulates, forwards down the shared tree, and, because it now knows the source address, sends an (S,G) join back toward the source to pull traffic natively. Once native traffic arrives, the RP sends register-stop and the tunnel closes. The aftermath is visible on R3: ``` R3# show ip mroute 239.1.1.10 (*, 239.1.1.10), 00:02:01/stopped, RP 2.2.2.2, flags: SPF Incoming interface: Ethernet0/0, RPF nbr 10.0.23.1 Outgoing interface list: Null (10.0.30.10, 239.1.1.10), 00:02:01/00:03:29, flags: FT Incoming interface: Ethernet0/1, RPF nbr 0.0.0.0 Outgoing interface list: Ethernet0/0, Forward/Sparse, 00:02:01/00:03:25 ``` The F flag is the register machinery, T confirms native (S,G) forwarding toward the RP. RPF neighbor 0.0.0.0 on the (S,G) means the source is directly connected. Meanwhile the RP holds both trees and stitches them together: ``` R2# show ip mroute 239.1.1.10 (*, 239.1.1.10), 00:02:41/00:02:45, RP 2.2.2.2, flags: S Outgoing interface list: Ethernet0/0, Forward/Sparse, 00:02:41/00:02:45 (10.0.30.10, 239.1.1.10), 00:01:53/00:01:07, flags: TA Incoming interface: Ethernet0/1, RPF nbr 10.0.23.2 Outgoing interface list: Ethernet0/0, Forward/Sparse, 00:01:53/00:02:45 ``` (The A flag marks this (S,G) as a candidate for MSDP advertisement, which becomes relevant in [interdomain multicast](https://www.pinglabz.com/bidir-pim-msdp-explained/).) ## Act Three: The SPT Switchover Traffic is now flowing source to receiver, but along a detour: through the RP. If the RP is not on the shortest path between source and receiver, that detour costs latency and bandwidth forever. So PIM-SM's last trick: the moment the last-hop router receives the first packet down the shared tree, it learns the source address from the packet itself, sends an (S,G) join along the shortest unicast path to the source, and prunes the source off the shared tree once native SPT traffic arrives. On IOS XE this is governed by `ip pim spt-threshold`, default 0: switch over on the first packet. For this capture the lab was first pinned to the shared tree with `ip pim spt-threshold infinity`, then released. Before, on R1: ``` R1# show ip mroute 239.1.1.10 ! spt-threshold infinity (*, 239.1.1.10), 00:03:04/00:02:24, RP 2.2.2.2, flags: SC Incoming interface: Ethernet0/1, RPF nbr 10.0.12.2 Outgoing interface list: Ethernet0/0, Forward/Sparse, 00:03:04/00:02:24 ``` One entry, no (S,G), traffic riding the RP tree indefinitely. Then `no ip pim spt-threshold infinity`, and within one packet: ``` R1# show ip mroute 239.1.1.10 ! default spt-threshold 0 (*, 239.1.1.10), 00:03:34/stopped, RP 2.2.2.2, flags: SJC Incoming interface: Ethernet0/1, RPF nbr 10.0.12.2 (10.0.30.10, 239.1.1.10), 00:00:05/00:02:54, flags: JT Incoming interface: Ethernet0/1, RPF nbr 10.0.12.2 Outgoing interface list: Ethernet0/0, Forward/Sparse, 00:00:05/00:02:54 ``` The J flag on both entries marks the join-SPT decision; T on the (S,G) confirms packets are arriving on the source tree. In this topology the shared tree and SPT happen to share links (R2 sits on the shortest path), so the win is invisible in the path; in a real network with an off-path RP, this is where the detour disappears. The router even accounts for the transition, showing exactly how many packets rode each tree: ``` R1# show ip mroute count | include 239.1.1.10|tree|Source: Group: 239.1.1.10, Source count: 1, Packets forwarded: 336, Packets received: 336 RP-tree: Forwarding: 323/0/60/0, Other: 323/0/0 Source: 10.0.30.10/32, Forwarding: 13/1/61/0, Other: 13/0/0 ``` 323 packets on the shared tree, 13 on the SPT and counting. That pair of counters is the SPT switchover, quantified. ## Inside the Register Tunnel The register process deserves a closer look because two of its failure modes are ENCOR favorites. When R3's first multicast packet arrives from the source, R3 builds a unicast IP packet addressed to the RP, sets the protocol to PIM, and stuffs the entire original multicast packet inside: data-plane traffic hidden inside control-plane messaging. On IOS XE you can watch the tunnel interfaces get created the moment PIM learns an RP (`show interface tunnel0` reveals a "Pim Register Tun" toward 2.2.2.2; the RP grows a decap tunnel to match). Two registers actually exist: data registers carrying payload, and null registers, header-only keepalives the first hop keeps sending about once a minute so the RP remembers the source while suppression is in effect. The failure modes: first, the RP must have a valid RPF path back to the source, or it cannot send the (S,G) join that lets it request native traffic; the register tunnel keeps working, the join never arrives, and the source stays stuck in register forever, burning first-hop CPU (the F flag never clearing is your clue). Second, registration is a unicast transaction with the RP address, so anything that breaks unicast to the RP (an ACL, a missing loopback advertisement) breaks source registration while leaving receiver-side shared trees looking perfectly healthy. Asymmetric symptoms, one cause. ## Who Sends What: DR Election on Shared Segments On any multi-access segment with multiple PIM routers, exactly one of them acts on behalf of hosts: the designated router. The DR sends the (\*,G) joins when receivers appear (last-hop DR) and performs registration when sources transmit (first-hop DR). Election is highest DR priority, tie broken by highest IP address, visible in the neighbor table's DR column and adjustable per interface with `ip pim dr-priority`. Two operational notes: the IGMP querier election (lowest IP, covered in [IGMP explained](https://www.pinglabz.com/igmp-explained-v2-v3/)) can land on a different router than the PIM DR, which is normal and confuses everyone once; and if the DR dies, the standby takes over only after the neighbor holdtime expires (default 105 seconds), so DR priority placement matters on segments with real receivers. ## Tuning the Switchover (and When to Refuse It) The `ip pim spt-threshold` default of 0 is right for almost everyone: shortest paths, minimum latency, RP out of the data path immediately. The `infinity` setting, pinning receivers to the shared tree as demonstrated above, has legitimate uses: it concentrates state (routers between RP and receivers hold only (\*,G) entries no matter how many sources exist), and it makes traffic engineering predictable, since everything flows through one known point. Some designs apply it per group range with an ACL, keeping high-rate feeds on SPTs while low-rate chatter stays shared. The middle-ground kbps thresholds the CLI still accepts are largely historical; modern platforms measure poorly at low rates and the operational advice is binary, 0 or infinity. If you do pin the shared tree, size the RP links accordingly: they are now permanent transit, not just introduction service. ## Finding the RP: Static, BSR, Auto-RP Everything above assumed routers agree on the RP. Three mechanisms get you there. Static (`ip pim rp-address 2.2.2.2` on every router) is predictable and fine for small domains; the cost is touching every router to change it. BSR (RFC 5059, standards-based) elects a bootstrap router that collects candidate-RP advertisements and floods the mapping set hop by hop. Auto-RP is the legacy Cisco equivalent over groups 224.0.1.39/40\. Adding BSR to the lab took two lines on R2 (`ip pim bsr-candidate Loopback0` and `ip pim rp-candidate Loopback0`), after which every router learned the mapping dynamically: ``` R1# show ip pim rp mapping Group(s) 224.0.0.0/4 RP 2.2.2.2 (?), v2 Info source: 2.2.2.2 (?), via bootstrap, priority 0, holdtime 150 Group(s): 224.0.0.0/4, Static RP: 2.2.2.2 (?) R1# show ip pim bsr-router PIMv2 Bootstrap information BSR address: 2.2.2.2 (?) Uptime: 00:01:05, BSR Priority: 0, Hash mask length: 0 ``` Both sources coexist here; dynamically learned mappings win unless the static entry carries the `override` keyword. For RP redundancy at scale, anycast RP with MSDP is the standard design, covered in [bidir PIM and MSDP](https://www.pinglabz.com/bidir-pim-msdp-explained/). ## What the Receiver Saw Every capture above is control-plane state, so it is worth closing the loop with the data plane. Throughout the register and switchover sequence, the actual receiver, a Debian VM bridged into the lab as a real host, was pulling the stream: ``` j@llmbits:~$ python3 /tmp/mcast_recv.py 239.1.1.10 5001 joined 239.1.1.10:5001 on 192.168.99.100 (IGMP membership active) rx 32 bytes from 10.0.30.10: PingLabz mcast 239.1.1.10 pkt 53 rx 32 bytes from 10.0.30.10: PingLabz mcast 239.1.1.10 pkt 54 rx 32 bytes from 10.0.30.10: PingLabz mcast 239.1.1.10 pkt 55 ``` No packets were lost across the switchover: the (S,G) tree was fully built and passing traffic before the shared-tree prune took effect, which is the make-before-break property that lets PIM-SM re-root a live stream invisibly. The sequence numbers just kept incrementing while the network beneath them swapped trees. That is the standard PIM-SM outcome: the mechanism is elaborate, and the application never notices any of it. ## Flags Cheat Card S / C / L Sparse mode / directly connected member via IGMP / router itself joined the group. J / T Join SPT decision made / traffic confirmed arriving on the source tree. JT together = healthy post-switchover state. F / P Register flag (first-hop router tunneling to RP) / pruned, OIL is Null. PF on a first hop with no receivers is normal, not broken. A / R / s MSDP advertisement candidate / RP-bit (entry toward RP on shared tree) / group is in the SSM range. ## Key Takeaways - PIM-SM forwards nothing until asked: receivers pull with explicit (\*,G) joins toward the RP, sources push their existence with unicast registers to the RP. - The RP exists for introductions, not for permanent transit. After the SPT switchover (default: first packet), traffic flows on the shortest path and the RP drops out of the data plane for that receiver. - The lifecycle reads directly from `show ip mroute` flags: SC while waiting, SPF/FT at a registering first hop, TA at the RP, SJC plus JT after switchover. - `show ip mroute count` quantifies the switchover with separate RP-tree and source counters. - Static RP, BSR, and Auto-RP are just three ways to distribute the same mapping; every router must agree or joins die quietly. Verify with `show ip pim rp mapping` everywhere. Next: [the RPF check](https://www.pinglabz.com/multicast-rpf-check/), the rule underneath every incoming-interface line above. The full cluster is in the [complete multicast guide](https://www.pinglabz.com/multicast/). ### IGMP Explained: v2 vs v3, Joins, Leaves, and the Querier URL: https://www.pinglabz.com/igmp-explained-v2-v3/ Last updated: 2026-07-11T18:03:41.000Z IGMP is the only part of multicast a host ever speaks. Before PIM builds a single tree, before an RP hears a single register, a receiver has to raise its hand on the local segment and say "I want group 239.1.1.10". This article covers how that conversation works in IGMPv2 and IGMPv3, what the querier does, how joins and leaves age out, and the version-mismatch trap that silently breaks SSM. All output is real, captured from a CML lab where a Debian VM is the receiver. Part of the [IP multicast guide](https://www.pinglabz.com/multicast/). ## What IGMP Does (and Does Not Do) IGMP runs between hosts and their first-hop router, on one segment only. It answers exactly one question: which groups have interested receivers on this interface? It carries no routing information, never crosses a router, and knows nothing about sources or trees. The router translates IGMP state into PIM joins upstream (covered in [PIM sparse mode](https://www.pinglabz.com/pim-sparse-mode-explained/)); IGMP itself stops at the first hop. The protocol has exactly two message flows: hosts send **membership reports** (joins, and in v3, source-filtered subscriptions), and one router per segment, the querier, sends periodic **queries** to check that members still exist. Leaves are the third piece, handled differently per version. ## The Querier: One Router Asks, Everyone Answers If a segment has multiple routers, only one sends queries: the querier, elected as the lowest IP address on the segment (in both v2 and v3). Everything about the querier's behavior is visible in one command. This is the lab's edge router, fresh after a host joined a group: ``` R1# show ip igmp interface Ethernet0/0 Internet address is 192.168.99.1/24 Current IGMP host version is 2 Current IGMP router version is 2 IGMP query interval is 60 seconds IGMP querier timeout is 120 seconds IGMP max query response time is 10 seconds Last member query count is 2 IGMP activity: 1 joins, 0 leaves Multicast designated router (DR) is 192.168.99.1 (this system) IGMP querying router is 192.168.99.1 (this system) ``` Every 60 seconds the querier sends a general query to 224.0.0.1\. Hosts answer with reports for each group they belong to, after a random delay up to the max response time (10 seconds) so a hundred receivers do not answer simultaneously. In v2, hosts also suppress their report if another host answers first for the same group; the router only needs one member to keep forwarding. Membership state expires after roughly (robustness x query interval) + max response time, here about 130 seconds, visible as the Expires timer: ``` R1# show ip igmp groups Group Address Interface Uptime Expires Last Reporter 239.1.1.10 Ethernet0/0 00:00:21 00:02:39 169.254.147.59 ``` Note the DR line as well. On multi-router segments the IGMP querier (lowest IP) and the PIM DR (highest priority/IP, responsible for sending PIM joins upstream) can be different routers. Exam writers love that distinction. ## IGMPv2: Join, Stay, Leave IGMPv2 is the simple lifecycle. A host joins by sending an unsolicited membership report to the group address itself. It stays by answering general queries. It leaves by sending a leave message to 224.0.0.2 (all routers), which triggers the querier to send a group-specific query: "anyone still here for 239.1.1.10?" It asks twice (the last member query count above), one second apart, and if nobody answers, the group is removed and the router prunes upstream. Worst-case leave latency is about two seconds, a big improvement over v1, which had no leave at all and simply aged out. What v2 cannot do: say anything about sources. A v2 join means "any source sending to G". That is ASM (any-source multicast), and it is why ASM needs a rendezvous point to introduce sources and receivers. ## IGMPv3: Source Filtering Changes the Model IGMPv3 (RFC 3376) adds one capability that changes multicast architecture: reports carry a source filter. A host can join in INCLUDE mode ("only these sources") or EXCLUDE mode ("everything except these"; an empty exclude list is the v2-equivalent any-source join). Mechanically, v3 also changes where reports go: to 224.0.0.22 (all IGMPv3 routers) instead of the group address, and one report can carry multiple group records, so suppression is gone; every host reports. Source filtering is the enabler for [SSM](https://www.pinglabz.com/source-specific-multicast-ssm/): if the receiver names the source, the network can build the (S,G) tree directly with no RP. Here is the lab's edge router after the Debian VM sent a v3 INCLUDE join for group 232.1.1.10, source 10.0.30.10: ``` R1# show ip igmp groups 232.1.1.10 detail Interface: Ethernet0/0 Group: 232.1.1.10 Flags: SSM Group mode: INCLUDE Last reporter: 169.254.147.59 Source Address Uptime v3 Exp CSR Exp Fwd Flags 10.0.30.10 00:00:08 00:02:51 stopped Yes R ``` Group mode INCLUDE with an explicit source list: that is the entire difference between v2 and v3 in one capture. The router now knows not just "someone wants G" but "someone wants G from exactly this S". ## v2 vs v3 Side by Side IGMPv2 Reports sent to the group address. Report suppression on. Explicit leave to 224.0.0.2\. Any-source joins only. Enables ASM; needs an RP for source discovery. IGMPv3 Reports to 224.0.0.22, multiple group records per report, no suppression. INCLUDE/EXCLUDE source filters. State changes replace leave messages. Enables SSM. Both versions Querier = lowest IP on segment. Periodic general queries to 224.0.0.1\. Group state expires without answers. Segment-local only; never routed. ## The Compatibility Trap That Breaks SSM IGMPv3 is backward compatible, and the fallback has a sharp edge. When a v3 host sees a v2 query on the segment, it drops into v2 compatibility mode and stays there for the "older version querier present" interval, roughly two query cycles, even after the querier upgrades. In compat mode the host sends v2 reports, and a v2 report carries no source list. This happened live while building this lab. R1's receiver interface started at v2 while the Debian VM issued a source-specific join for 232.1.1.10\. Result, straight off the wire: ``` j@llmbits:~$ sudo tcpdump -i ens224 -c 3 -n igmp 10:41:03.46 IP 192.168.99.1 > 224.0.0.1: igmp query v3 10:41:05.06 IP 169.254.147.59 > 232.1.1.10: igmp v2 report 232.1.1.10 10:41:07.52 IP 169.254.147.59 > 239.5.5.5: igmp v2 report 239.5.5.5 ``` The router is already querying at v3, but the host is still reporting at v2 from compat mode, and IOS XE ignores v2 reports for the SSM range (a sourceless join is meaningless there). The router held zero state for the group until the host was forced back to v3 (on Linux: `force_igmp_version=3` sysctl, or wait out the compat timer), after which the INCLUDE join above appeared immediately. If SSM "just doesn't work" for some receivers, check IGMP versions on both ends before anything else, with `show ip igmp interface` on the router (both host and router version lines) and tcpdump on the host. Configuration is one line per interface: `ip igmp version 3` on everything that should support SSM receivers. IOS XE defaults to v2. ## The Timers That Rule the Lifecycle Every IGMP behavior above is timer-driven, and the defaults visible in the interface output interact in ways worth spelling out. The query interval (60 seconds on IOS XE, 125 in the RFC) sets the heartbeat. The robustness variable (2) is the "how many losses do we tolerate" multiplier baked into every derived timer. Group membership expires after (robustness x query interval) + max response time: 2 x 60 + 10 = 130 seconds of silence before the router stops forwarding. The querier itself is given up on after the querier timeout (2 x query interval = 120 seconds), at which point the next-lowest IP takes over querying. Leave latency is governed separately: on a v2 leave, the querier sends last-member-query-count (2) group-specific queries at last-member-query-interval (1 second) spacing, so a group with no remaining members stops flowing in about 2 seconds. If your application needs faster cutoff (channel zapping in IPTV is the classic case), `ip igmp last-member-query-interval` and immediate-leave (`ip igmp immediate-leave group-list`, safe only when exactly one receiver lives per port) are the knobs. Tightening the general query interval buys faster failure detection at the cost of more control traffic and more report bursts from large segments; most networks should leave it alone. ## Joining From the Router Itself: join-group vs static-group Two interface commands make a router behave like a receiver, and confusing them causes real outages. `ip igmp join-group 239.1.1.10` makes the router a full member: it answers queries and, critically, punts every packet for the group to its CPU. It is a test tool (the router responds to pings sent to the group, which makes end-to-end reachability testable without any host), and leaving it configured on a production box is a well-known way to melt a control plane. `ip igmp static-group 239.1.1.10` is the operational sibling: the router forwards the group out the interface unconditionally, in hardware, without becoming a member itself. Use it for segments where hosts cannot speak IGMP (some industrial gear, dumb display endpoints) or where you want traffic pre-pulled to an interface regardless of membership. In the lab you will notice 224.0.1.40 joined on every router's loopback in `show ip igmp groups`; that is Auto-RP's discovery group, joined automatically, and a reminder that routers themselves are IGMP participants too. ## IGMP Snooping: The Switch in the Middle One layer down, switches flood multicast frames by default because the destination MAC never appears in their learned tables. IGMP snooping fixes that: the switch listens to reports and queries and constrains each group's frames to ports with actual members plus router ports. It is on by default on Cisco switches and generally should stay on; a segment with no querier and snooping enabled is the classic "multicast works for two minutes then dies" failure, because snooped state ages out with no queries to refresh it. On router-less L2 domains, configure an IGMP snooping querier on the switch. The details of L2 forwarding behavior tie into the MAC overlap covered in [multicast addressing](https://www.pinglabz.com/multicast-addressing-explained/). ## Securing the Membership Plane IGMP is unauthenticated by design, so the controls are limits and filters. `ip igmp access-group ` on an interface restricts which groups hosts may join at all, the right tool for segments that should only ever receive the corporate video group and nothing else. `ip igmp limit` (global or per interface) caps total membership state, blunting both buggy applications that join in a loop and deliberate state-exhaustion attempts. On the switch side, IGMP snooping's report suppression and querier hardening matter more than router config for most attack surfaces; a host spraying bogus reports pollutes L2 forwarding before any router notices. None of this is exotic: the same three commands appear in every hardening baseline, and the [pillar's security section](https://www.pinglabz.com/multicast/) puts them alongside the PIM-side controls they complement. ## Key Takeaways - IGMP is host-to-router membership signaling on one segment; PIM does everything between routers. - The querier (lowest IP) polls 224.0.0.1 every 60 seconds by default; group state expires in about 130 seconds without answers. All of it is visible in `show ip igmp interface`. - v2 gives you any-source joins and fast leaves. v3 adds INCLUDE/EXCLUDE source filters, reports to 224.0.0.22, and no suppression. - SSM requires v3 end to end. Hosts fall back to v2 compat mode when they ever see a v2 query and stay there for minutes; IOS ignores v2 reports in 232/8, so the join silently vanishes. - `show ip igmp groups detail` showing Group mode INCLUDE with a source list is your proof that v3 signaling actually works. Next: [PIM sparse mode](https://www.pinglabz.com/pim-sparse-mode-explained/) turns this membership state into trees. The full cluster lives in the [complete multicast guide](https://www.pinglabz.com/multicast/). ### Multicast Addressing: 224.0.0.0/4, MAC Mapping, and Scoping URL: https://www.pinglabz.com/multicast-addressing-explained/ Last updated: 2026-07-11T18:03:40.000Z Every multicast design decision starts with an address. Which range you pick determines whether routers forward the traffic at all, whether you need an RP, whether it can leave your organization, and even how it collides at layer 2\. This article maps the entire 224.0.0.0/4 space, explains the IP-to-MAC mapping (and the 32:1 overlap it creates), and covers scoping. It is part of the [IP multicast guide](https://www.pinglabz.com/multicast/). ## Class D: The 224.0.0.0/4 Space Multicast owns 224.0.0.0 through 239.255.255.255: the old class D, or 224.0.0.0/4 in modern notation. The first four bits of the first octet are fixed at 1110, leaving 28 bits of group space, roughly 268 million possible groups. Group addresses appear only as destinations. If you ever see one as a source address in a capture, something is forging packets. Within that /4, specific ranges carry very different rules: 224.0.0.0/24Local network control block. TTL-irrelevant: routers never forward these, period. This is where the protocols live. 224.0.1.0/24Internetwork control block. Routable. Auto-RP's 224.0.1.39 and 224.0.1.40 are the entries you will meet in practice. 232.0.0.0/8Source-Specific Multicast. Receivers must name the source (IGMPv3); no RP is ever consulted. See the [SSM article](https://www.pinglabz.com/source-specific-multicast-ssm/). 233.0.0.0/8GLOP addressing: your 16-bit ASN in octets two and three gives every AS a deterministic /24 of public multicast space. 234.0.0.0/8Unicast-prefix-based assignment: organizations derive multicast space from their unicast allocations. 239.0.0.0/8Administratively scoped: private multicast, the RFC 1918 equivalent. Use inside your organization, filter at the border. ## The Well-Known Addresses Worth Memorizing The local control block is where you meet multicast daily, even in unicast-only networks. These are the ones that show up in captures and exams: 224.0.0.1 / 224.0.0.2 All hosts / all routers on this segment. IGMP general queries target 224.0.0.1. 224.0.0.5 / 224.0.0.6 All OSPF routers / OSPF DRs. Why OSPF neighbors find each other without configuration; see the [OSPF guide](https://www.pinglabz.com/ospf/). 224.0.0.10 / 224.0.0.13 EIGRP routers / all PIM routers. PIM hellos ride 224.0.0.13. 224.0.0.18 / 224.0.0.22 VRRP / IGMPv3 membership reports. v3 reports go to .22, not to the group being joined. 224.0.0.102 / 224.0.0.251 HSRPv2/GLBP / mDNS. The [FHRP guide](https://www.pinglabz.com/fhrp/) covers the first; your smart TV uses the second. 224.0.1.39 / 224.0.1.40 Cisco Auto-RP announce / discovery. Routable, unlike everything above. ## IP-to-MAC Mapping: 01:00:5e and the Missing Bit Ethernet needs a destination MAC, but no host owns a multicast group, so the MAC is derived from the group address. The rule: start with the fixed OUI prefix `01:00:5e`, force the next bit to 0, then copy the low 23 bits of the IP group address into the low 23 bits of the MAC. Work through 239.1.1.10\. The low 23 bits cover the last two octets fully (1.10) and the low 7 bits of the second octet (1). Result: `01:00:5e:01:01:0a`. You can see the real thing on the lab's Debian receiver after it joined the group: ``` j@llmbits:~$ ip maddr show ens224 3: ens224 link 01:00:5e:01:01:0a users 2 inet 239.1.1.10 inet 232.1.1.10 ``` Notice `users 2`: both 239.1.1.10 and 232.1.1.10 map to the same MAC, because their low 23 bits are identical. That is the design flaw in action. ## The 32:1 Overlap Problem An IPv4 group has 28 significant bits; the MAC mapping preserves only 23\. Five bits are simply lost, which means exactly 32 different group addresses share every multicast MAC. 224.1.1.10, 225.1.1.10, 239.1.1.10, and 29 others all become `01:00:5e:01:01:0a`. The practical consequence: a switch (or a host NIC filtering in hardware) cannot distinguish those 32 groups. A host subscribed to 239.1.1.10 will also receive frames for 224.129.1.10 at layer 2 and must discard them in software. On busy networks this matters for both performance and isolation, and the fix is planning: keep the low 23 bits unique across the groups you deploy, and avoid x.0.0.y and x.128.0.y groups entirely since they collide with the 224.0.0.0/24 control block MACs that switches flood. Why only 23 bits? History: when the mapping was defined, buying 16 OUI blocks to cover all 28 bits cost more money than the designers had. Half an OUI was available, so the standard shipped with 23 bits and the overlap became permanent trivia. ## Scoping: Keeping Groups Where They Belong Two mechanisms constrain how far multicast travels. TTL scoping is the legacy approach: senders set a small TTL and packets die at distance. It is fragile (TTL also decrements per hop for normal reasons) and has been superseded by administrative scoping: dedicated ranges with configured boundaries. Within 239/8, convention splits the space further: 239.255.0.0/16 for site-local groups, 239.192.0.0/14 for organization-local. The enforcement tool on IOS XE is a multicast boundary ACL on the edge interface: ``` ip access-list standard MCAST-SCOPE deny 239.0.0.0 0.255.255.255 permit 224.0.0.0 15.255.255.255 ! interface GigabitEthernet0/0/1 ip multicast boundary MCAST-SCOPE ``` That drops organization-local groups in both directions at the border while letting legitimately routable ranges through. Every edge interface facing a partner, provider, or the internet should carry one. ## Working the Mapping Both Directions Interviews and exams love the reverse question: given MAC `01:00:5e:01:01:0a`, what groups could have produced it? Take the low 23 bits: `01:01:0a` with the top bit of the first of those octets already 0, so the group ends in x.1.1.10 where x has its low 7 bits free in the second octet. The candidates are every combination of the 4 lost class D prefix bits and the 1 lost high bit of the second octet: 224.1.1.10, 224.129.1.10, 225.1.1.10, 225.129.1.10, up through 239.129.1.10\. Thirty-two possibilities, and nothing at layer 2 can tell them apart. Going forward is mechanical once you internalize which bits survive: the second octet contributes only its low 7 bits (so 0-127 and 128-255 alias in pairs), and the third and fourth octets copy through whole. 239.130.44.7 therefore becomes `01:00:5e:02:2c:07` (130 minus 128 = 2), colliding with 239.2.44.7\. If you allocate groups by keeping octets two through four unique below 128, the overlap problem disappears from your network entirely. ## IPv6 Multicast: The Same Ideas, Cleaner IPv6 fixes most of this by design, and the contrast is instructive (the full story is in the [IPv6 guide](https://www.pinglabz.com/ipv6/)). All IPv6 multicast lives under ff00::/8, and the second octet encodes flags and an explicit scope nibble: ff02:: is link-local, ff05:: site-local, ff0e:: global. Scoping is part of the address instead of a boundary ACL convention, so a link-scoped group physically cannot be routed off-segment. The well-known groups mirror IPv4's: ff02::1 is all nodes (224.0.0.1), ff02::2 all routers, ff02::5 and ff02::6 OSPFv3, ff02::d PIM, and ff02::1:ff00:0/104 hosts the solicited-node groups that replace ARP broadcasts entirely. The MAC mapping also loses its overlap problem in practice: IPv6 maps the low 32 bits of the group into `33:33:xx:xx:xx:xx`, and since group IDs are conventionally allocated within those 32 bits, collisions are rare rather than structural. There is no broadcast in IPv6 at all; multicast does that job, which is why understanding these ranges stopped being optional the day you enabled dual stack. MLD replaces IGMP as the membership protocol (MLDv2 corresponds to IGMPv3, source filters included), and PIM carries over unchanged. ## Choosing Addresses for a New Deployment A simple decision path covers nearly every enterprise case. Internal-only application: allocate from 239/8, document the assignment, and keep low-23-bit uniqueness. One-to-many stream where receivers can learn the source (IPTV, market data): use 232/8 with SSM and skip RP design entirely. Public interdomain multicast (rare): GLOP from 233/8 or unicast-prefix-based from 234/8\. And never invent groups inside 224.0.0.0/24 or 224.0.1.0/24; those belong to protocols. ## Key Takeaways - Multicast lives in 224.0.0.0/4\. The subranges carry the rules: 224.0.0.0/24 never routes, 232/8 is SSM, 239/8 is private and must be filtered at borders. - 224.0.0.5, .6, .10, .13, .18, .22, and .251 are the well-known addresses you will see in real captures; recognize them instantly. - The MAC mapping is 01:00:5e plus the low 23 bits of the group. Five IP bits are lost, so 32 groups share each MAC and hosts filter the excess in software. - Plan group assignments for low-23-bit uniqueness and avoid anything colliding with the control block MACs. - Scope with `ip multicast boundary` and administrative ranges, not TTL tricks. Continue with [IGMP explained: v2 vs v3](https://www.pinglabz.com/igmp-explained-v2-v3/), or return to the [complete multicast guide](https://www.pinglabz.com/multicast/) for the full cluster. ### What Is IP Multicast? Groups, Trees, and Why Unicast Doesn't Scale URL: https://www.pinglabz.com/what-is-ip-multicast/ Last updated: 2026-07-11T18:03:40.000Z Send one packet, deliver it to a hundred receivers, and put exactly one copy on each link along the way. That is the entire promise of IP multicast, and it is the delivery model behind market data feeds, IPTV, live enterprise video, and OS imaging at scale. This article covers what multicast actually is, why unicast and broadcast both fail at the same job, and the group and tree concepts everything else builds on. It is the starting point for the full [IP multicast guide](https://www.pinglabz.com/multicast/). ## The Three Delivery Models IPv4 gives you three ways to deliver a packet (IPv6 adds anycast, but that is a story for the [IPv6 guide](https://www.pinglabz.com/ipv6/)). Unicast is one-to-one: every receiver gets its own copy, generated by the source. Broadcast is one-to-all: one copy per segment, but every host on the segment processes it whether it cares or not, and routers refuse to forward it. Multicast is one-to-many: one copy per link, replicated by the network only where paths to interested receivers split, and delivered only to hosts that asked. The key phrase is "hosts that asked". Multicast is an opt-in model. Receivers subscribe to a group address, the network tracks who subscribed where, and traffic follows the subscriptions. Nobody polls the source, and the source never learns who is listening. It just sends to the group. ## Why Unicast Doesn't Scale for One-to-Many Do the arithmetic on a real workload. A market data server publishes a 5 Mbps feed. With 200 subscribers over unicast, the server transmits 200 identical streams: 1 Gbps out of its NIC, 1 Gbps across its first-hop link, and duplicated load on every shared link downstream. Double the subscribers and you double all of it. The source becomes the bottleneck, and the links closest to it melt first. The same feed over multicast: the server sends 5 Mbps. One copy crosses each link, no matter whether the far side holds one subscriber or ten thousand. Replication happens inside the routers, at the exact points where receiver paths diverge. Subscriber growth costs the source nothing. Unicast, 200 receivers Source sends: **1 Gbps** (200 x 5 Mbps). First-hop link: 1 Gbps. Cost grows linearly with every receiver. Broadcast One copy per segment, but **every host interrupted**, zero opt-in, and it cannot cross a router. Useless beyond one subnet. Multicast, 200 receivers Source sends: **5 Mbps**. Max one copy per link. Only subscribed hosts receive. Receiver count is irrelevant to the source. ## Groups: Addresses That Name a Service, Not a Host A multicast group is a class D IP address (224.0.0.0 through 239.255.255.255) that identifies a stream of content rather than a machine. Nothing "owns" 239.1.1.10; sources send to it and receivers listen to it, and the two sides never need to know each other. Group addresses never appear as a source address in an IP header, only as a destination. Membership is signaled with IGMP. When an application on a host joins a group, the host's IP stack sends an IGMP membership report on the local segment, and the first-hop router adds that interface to its delivery state. This is real output from the lab's edge router the moment a Debian host joined a group (the deep-dive is in [IGMP explained](https://www.pinglabz.com/igmp-explained-v2-v3/)): ``` R1# show ip igmp groups IGMP Connected Group Membership Group Address Interface Uptime Expires Last Reporter 239.1.1.10 Ethernet0/0 00:00:21 00:02:39 169.254.147.59 ``` ## Distribution Trees: How the Network Replicates Between the source and its receivers, routers build a distribution tree: a loop-free set of paths with the traffic origin at the root and receivers at the leaves. Each router in the tree knows one incoming interface (toward the root) and a list of outgoing interfaces (toward the leaves). A packet arrives once and is copied to every outgoing interface. That per-router state is exactly what `show ip mroute` displays. Trees come in two flavors. A **source tree**, written (S,G), is rooted at one specific source and follows the shortest unicast path to every receiver: optimal paths, but one tree of state per source. A **shared tree**, written (\*,G), is rooted at a rendezvous point that all sources for the group transit: one tree regardless of source count, at the price of a potentially longer path. PIM sparse mode uses both, starting receivers on the shared tree and switching them to source trees when traffic flows; that mechanism is the subject of [PIM sparse mode explained](https://www.pinglabz.com/pim-sparse-mode-explained/). Here is a real source tree entry on the lab's last-hop router while a stream was flowing: ``` R1# show ip mroute 239.1.1.10 (10.0.30.10, 239.1.1.10), 00:00:05/00:02:54, flags: JT Incoming interface: Ethernet0/1, RPF nbr 10.0.12.2 Outgoing interface list: Ethernet0/0, Forward/Sparse, 00:00:05/00:02:54 ``` Read it as: traffic from source 10.0.30.10 to group 239.1.1.10 must arrive on Ethernet0/1, and gets copied out Ethernet0/0\. The "must arrive on" part is enforced by the RPF check, multicast's loop-prevention mechanism, covered in [the RPF check](https://www.pinglabz.com/multicast-rpf-check/). ## The Protocol Stack at a Glance Three protocol layers cooperate, and it helps to keep their jobs separate from day one: - **IGMP** runs between hosts and their first-hop router. It answers exactly one question: which groups are wanted on this segment? - **PIM** runs between routers and builds the trees. It is "protocol independent" because it has no topology discovery of its own; it borrows the unicast routing table for every decision. - **Your IGP** (in the PingLabz lab, [OSPF](https://www.pinglabz.com/ospf/) area 0) provides that unicast table. Broken underlay means broken multicast, every time. End to end, it looks like this in practice. A real Linux VM subscribed to a group, and a Python sender on the far side of a four-router topology started transmitting: ``` j@llmbits:~$ python3 /tmp/mcast_recv.py 239.1.1.10 5001 joined 239.1.1.10:5001 on 192.168.99.100 (IGMP membership active) rx 32 bytes from 10.0.30.10: PingLabz mcast 239.1.1.10 pkt 53 rx 32 bytes from 10.0.30.10: PingLabz mcast 239.1.1.10 pkt 54 rx 32 bytes from 10.0.30.10: PingLabz mcast 239.1.1.10 pkt 55 ``` The receiver never contacted the source. It told its router what it wanted (IGMP), the routers built the tree (PIM, over OSPF), and packets arrived. ## From Join to Packets: The Ten-Thousand-Foot Walkthrough It is worth walking one delivery end to end, because the division of labor confuses everyone at first. Say the VM at 192.168.99.100 wants the stream on 239.1.1.10, and a server four hops away is sending to it. First, the application opens a UDP socket and asks the OS to join the group. The kernel programs the NIC to accept the mapped multicast MAC and sends an IGMP membership report out the interface. That report reaches R1, the first-hop router, which records "Ethernet0/0 has a member of 239.1.1.10" and nothing more; IGMP's job is done. Second, R1 needs the traffic to actually arrive from somewhere. It creates forwarding state for the group and sends a PIM join upstream. Which neighbor counts as "upstream" is decided by the unicast routing table, and in sparse mode the initial join travels toward the rendezvous point. Each router along the way repeats the process: create state, add the interface the join arrived on to the outgoing list, forward the join upstream. The tree builds itself backwards, from receiver toward root, one hop at a time. Third, when the server sends, its packets flow down that tree. Each router checks the packet arrived on the expected upstream interface (the RPF check), then copies it to every interface in the outgoing list. R1 copies it onto the receiver segment, the NIC accepts the MAC, the kernel matches the group to the socket, and the application reads a datagram. Rip out any one of the three layers (membership, tree, or underlay routing) and delivery stops, which is exactly how the [troubleshooting article](https://www.pinglabz.com/troubleshooting-ip-multicast/) is organized. ## What Multicast Does Not Give You Multicast is a delivery optimization, not a transport upgrade, and its constraints shape application design. Everything rides UDP: TCP's handshake and retransmission model is meaningless when one sender feeds ten thousand receivers, so there is no built-in reliability, ordering, or congestion control. A dropped packet is simply gone unless the application layer fixes it, which is why serious multicast applications ship with their own recovery mechanisms: sequence numbers plus a unicast retransmission channel (market data), forward error correction (video), or reliable multicast frameworks like PGM. Flow control does not exist either. The source transmits at whatever rate it chooses, and the slowest receiver's problems are its own. In practice this pushes multicast networks toward careful QoS design, since a bursting stream competes with everything else on every branch of the tree at once; the [QoS guide](https://www.pinglabz.com/qos/) covers the queuing side of that story. None of this diminishes the model: it just means multicast solves distribution, and the application still owns end-to-end semantics. ## What Actually Uses Multicast You touch multicast more often than you might think, even before deploying it deliberately: - **Routing protocols themselves.** OSPF hellos ride 224.0.0.5/6, EIGRP uses 224.0.0.10, PIM uses 224.0.0.13\. Link-local multicast, never forwarded, but multicast all the same. - **Financial market data.** Exchange feeds are the canonical latency-sensitive one-to-many workload, usually as SSM. - **IPTV and enterprise video.** One encoder, thousands of set-top boxes or desks watching an all-hands. - **Imaging and provisioning.** Pushing one OS image to a lab of 300 machines simultaneously. - **Service discovery.** mDNS (224.0.0.251) and SSDP find printers and cast targets on your home network right now. The common thread: identical data, many consumers, at the same time. When consumers need different data or need it at different times (video on demand, software downloads), unicast plus caching wins, which is why the public internet's streaming runs on CDNs instead. ## Key Takeaways - Multicast is opt-in, one-to-many delivery: one copy per link, replicated by routers where receiver paths diverge, delivered only to hosts that joined. - Unicast cost grows linearly with receivers; multicast cost is flat for the source. That difference is the whole business case. - A group address (224.0.0.0/4) names a stream, not a machine. Sources send to it; receivers join it; neither knows the other. - Forwarding state is a tree: (S,G) rooted at a source for optimal paths, (\*,G) rooted at an RP for shared state. Both appear in `show ip mroute`. - IGMP handles host-to-router membership, PIM builds router-to-router trees, and the unicast IGP underneath feeds every RPF decision. Next in the cluster: [multicast addressing](https://www.pinglabz.com/multicast-addressing-explained/) covers the ranges and the MAC mapping, and the [complete guide](https://www.pinglabz.com/multicast/) holds the full reading order. ### Troubleshooting Route Redistribution: Missing, Looping, and Suboptimal Routes URL: https://www.pinglabz.com/troubleshooting-route-redistribution/ Last updated: 2026-08-01T19:26:52.000Z Redistribution failures come in three flavors: routes that never arrive, routes that arrive too well (echo back, win the border router's affection, and loop), and routes that arrive perfectly but take the long way round. All three were built deliberately in the lab, so every signature below is a real capture from Cisco IOS XE 17.18.2 on CML iol-xe nodes. The third is the one nobody opens a ticket for: the prefix is there, pings succeed, and traffic quietly takes two extra hops for months. Keep the [complete redistribution guide](https://www.pinglabz.com/route-redistribution-cisco-ios-xe/) open for the mechanics; this is the diagnostic runbook. Part of the [IP routing fundamentals](https://www.pinglabz.com/ip-routing/) cluster. ## The Diagnostic Frame Redistribution moves routes from the routing table of one protocol into the database of another. So every missing-route investigation walks the same four stations: is the route in the source protocol's table on the border router, does the redistribute statement select it, does it appear in the destination protocol's database, and does the far router install it? One of those four links is broken, and the commands below bracket the search fast. A fifth station applies once the route is present: is the copy the router installed the one you wanted? Presence and correctness are separate tests. ## Case 1: Missing Routes, EIGRP Edition (No Seed Metric) The single most common redistribution failure in the wild. Symptom first, from the far EIGRP router: ``` R4# show ip route 1.1.1.1 % Network not in table ``` On the border router R3 the config looks plausible, and `show ip protocols` even confirms redistribution is on... read it to the end though: ``` R3# show ip protocols | section eigrp Redistributing: ospf 1 ... Total Prefix Count: 5 Total Redist Count: 0 ``` "Redistributing: ospf 1" and "Total Redist Count: 0" in the same output is the whole diagnosis. EIGRP's default seed metric is infinity: without an explicit metric every imported route is advertised as unreachable, which is to say not advertised at all. No error, no syslog. The topology table confirms it never entered EIGRP: ``` R3# show eigrp address-family ipv4 topology 1.1.1.1/32 EIGRP-IPv4 VR(PINGLABZ) Topology Entry for AS(100)/ID(3.3.3.3) %Entry 1.1.1.1/32 not in topology table ``` The fix is the five-part seed metric, in named mode under the topology base: ``` router eigrp PINGLABZ address-family ipv4 unicast autonomous-system 100 topology base redistribute ospf 1 metric 1000000 1 255 1 1500 ``` Seconds later, R4 holds the whole OSPF domain as externals (note the AD: 170 on every one, harmless here and the subject of Case 5): ``` R4# show ip route | begin Gateway D EX 1.1.1.1 [170/1029120] via 10.0.34.3, 00:00:30, Ethernet0/0 D EX 2.2.2.2 [170/1029120] via 10.0.34.3, 00:00:30, Ethernet0/0 D EX 10.0.12.0/24 [170/1029120] via 10.0.34.3, 00:00:30, Ethernet0/0 ``` ## Case 2: Missing Routes, OSPF Edition (The subnets Story) The classic advice says: if you redistribute into OSPF and only classful networks appear, you forgot the `subnets` keyword. Worth knowing, with a modern footnote. On current IOS XE, subnets are included by default, and the router says so: ``` R3# show ip protocols | include subnets eigrp, includes subnets in redistribution ``` So on a modern box this failure will not reproduce (we tried, in this very lab). You will still meet it on legacy IOS, where `redistribute eigrp 100` without `subnets` silently drops every subnetted prefix, which in a world of /24s and /32s means everything. Typing `subnets` costs nothing and self-documents; checking `show ip protocols` tells you which behavior your platform has. Two more members of the family sit at the same station: a prefix the exporter itself learned from the destination protocol (Case 4), and a route-map whose `deny` matches more than intended, which `show route-map` settles in seconds because it counts matches per sequence. ## Case 3: Looping Routes (Two Borders, No Tags) The opposite failure: everything redistributes wonderfully, in both directions, at two border routers, and one day a routine metric change produces this: ``` R1# traceroute 172.16.4.4 numeric timeout 1 probe 1 ttl 1 10 1 10.0.12.2 3 msec 2 10.0.23.3 3 msec 3 10.0.23.2 2 msec 4 10.0.23.3 4 msec 5 10.0.23.2 3 msec 6 10.0.23.3 4 msec R1# ping 172.16.4.4 ..!!! Success rate is 60 percent (3/5) ``` A traceroute alternating between two addresses is a forwarding loop between two routers that each believe the other has the path (the general signature, and how to break one, is in [reading a repeating traceroute](https://www.pinglabz.com/routing-loops-detect-and-fix/)). Intermittent ping success is the churn signature: the protocols flap between states as the echoed route wins and loses. Read the same prefix on both border routers: ``` R2# show ip route 172.16.4.0 | include Known|via 10 Known via "ospf 1", distance 110 ... via Ethernet0/1 (points at R3) R3# show ip route 172.16.4.0 Known via "eigrp 100", distance 170 ... * 10.0.23.2, from 10.0.23.2 ... via Ethernet0/0 (points at R2) ``` Root cause in one sentence: R3 redistributed the EIGRP prefix into OSPF, R2 preferred the OSPF copy (AD 110 beats EIGRP external 170) and re-redistributed it back into EIGRP with a fresh seed metric that eventually beat the original. The forensic tool is the EIGRP topology table, which names the echo's author: ``` R3# show eigrp address-family ipv4 topology 172.16.4.0/24 10.0.34.4 ... External protocol is Connected, Originating router is 4.4.4.4 10.0.23.2 ... External protocol is OSPF, external metric is 20, Originating router is 2.2.2.2 ``` A prefix whose external protocol is OSPF, originated by your other border router, sitting in the EIGRP topology table of the domain that created it: that is a route that left home and came back wearing a disguise. The structural cure is mutual tagging at every border, covered with fixed-state captures in the companion piece on [tagging on export and denying your own tag on the way back in](https://www.pinglabz.com/redistribution-loop-prevention-route-tags/). After: ``` R1# show ip route 172.16.4.0 Known via "ospf 1", distance 110, metric 20 Tag 100, type extern 2, forward metric 20 R1# traceroute 172.16.4.4 numeric timeout 1 probe 1 ttl 1 8 1 10.0.12.2 3 msec 2 10.0.23.3 3 msec 3 10.0.34.4 4 msec ``` ## The Runbook 1\. Far router missing the prefix? `show ip route ` both ends. Confirms the symptom and which direction is broken. 2\. Border router has it, from the right source? `show ip route ` on the border. "Known via" must name the protocol you are exporting FROM. 3\. Redistribution actually exporting? `show ip protocols`: seed metrics, filters, and (EIGRP) the Redist Count. Zero with redistribution configured = seed metric or route-map. 4\. In the destination database? `show ip ospf database external` / `show eigrp af ipv4 topology `. In the database but absent from the far RIB moves the problem downstream. 5\. Loop suspected? Traceroute for the ping-pong, then read the prefix on both borders and check the topology table for externals originated by your other border router. 6\. Present, but the path looks long? Compare the RIB with the source protocol's topology table. A better path that never installs is an AD problem, not a metric problem, and `show ip cef` proves what is really used. ## Case 4: The Route That Refuses to Export (AD Theft) A subtler missing-routes variant, and the one that most reliably confuses smart engineers. Symptom: a prefix you are certain lives in OSPF will not redistribute into EIGRP, even though the seed metric is right and no filter exists. The catch: on this border router the prefix was ALSO learned via EIGRP, and EIGRP won the RIB (internal AD 90 versus OSPF 110). Redistribution exports only what the routing table attributes to the named source protocol, so `redistribute ospf 1` walks straight past it. The one-command diagnosis is the detail view: ``` R2# show ip route 10.0.34.0 Routing entry for 10.0.34.0/24 Known via "eigrp 100", distance 90, metric 1536000, type internal ``` "Known via" naming the other protocol on the router you expected to export from ends the mystery. It is also a design smell: prefixes reachable through both domains at a border are the feedback fuel Case 3 burns, so check your tagging posture while you are in there. ## Case 5: Suboptimal Routing, the Better Path That Loses Nothing is missing and nothing is looping. The prefix is installed, the destination answers, and the traffic points the wrong way. Second lab for this one: R1, R2 and R3 run OSPF in area 0, and R1 also has a direct EIGRP link to R3, which owns 172.16.3.0/24, advertises it into OSPF and redistributes OSPF into EIGRP AS 100\. R1 learns the prefix twice, natively over OSPF (two hops via R2) and as an EIGRP external over the direct cable: ``` R1# show ip route 172.16.3.0 255.255.255.0 Routing entry for 172.16.3.0/24 Known via "ospf 1", distance 110, metric 21, type intra area Last update from 10.0.12.2 on Ethernet0/0, 00:01:00 ago Routing Descriptor Blocks: * 10.0.12.2, from 3.3.3.3, 00:01:00 ago, via Ethernet0/0 Route metric is 21, traffic share count is 1 ``` OSPF, distance 110, next hop 10.0.12.2, which is R2: R1 picked the two-hop path over a direct link to the router that owns the prefix. CEF says this is not just a control-plane curiosity: ``` R1# show ip cef 172.16.3.0 255.255.255.0 172.16.3.0/24 nexthop 10.0.12.2 Ethernet0/0 ``` Packets really do take the long way. Before blaming the direct link, ask EIGRP what it knows: ``` R1# show ip eigrp topology 172.16.3.0/24 EIGRP-IPv4 Topology Entry for AS(100)/ID(10.0.13.1) for 172.16.3.0/24 State is Passive, Query origin flag is 1, 0 Successor(s), FD is Infinity Descriptor Blocks: 10.0.13.2 (Ethernet0/1), from 10.0.13.2, Send flag is 0x0 Composite metric is (307200/281600), route is External ... External data: AS number of route is 1 External protocol is OSPF, external metric is 0 ``` That block is the whole lesson. The direct path over Ethernet0/1 exists, EIGRP computed a usable composite metric for it, and it is marked External with OSPF named as the source (proof it arrived through R3's redistribution). It will never be installed. `0 Successor(s)` with `FD is Infinity` is EIGRP saying it has a candidate and the RIB declined it: not because the path is longer (it is shorter) and not because the metric is worse (metrics from different protocols are never compared), but because a redistributed route enters EIGRP as an external, an external carries AD 170, and 170 loses to OSPF's 110 before the path is considered at all. That is the short answer to [how a router picks between two protocols offering the same prefix](https://www.pinglabz.com/administrative-distance/): AD ranks the source of the information, never the quality of the route. Connected0 Static1 External BGP20 EIGRP internal90 OSPF110 EIGRP external (anything redistributed in)170 Internal BGP200 The asymmetry lives inside EIGRP: internal routes at 90 beat OSPF, externals at 170 lose to it, so redistributing a prefix into EIGRP drops it from guaranteed winner to guaranteed loser ([why EIGRP trusts its own routes and distrusts imported ones](https://www.pinglabz.com/eigrp-administrative-distance/)). ### The One-Line Fix R1 is making the bad choice, so R1 is where the change goes. Leave internal alone and drop the external distance below OSPF's 110: ``` router eigrp 100 distance eigrp 90 109 ``` ``` R1# show ip route 172.16.3.0 255.255.255.0 Routing entry for 172.16.3.0/24 Known via "eigrp 100", distance 109, metric 307200, precedence routine (0), type external Redistributing via eigrp 100 Last update from 10.0.13.2 on Ethernet0/1, 00:00:48 ago Routing Descriptor Blocks: * 10.0.13.2, from 10.0.13.2, 00:00:48 ago, via Ethernet0/1 Route metric is 307200, traffic share count is 1 ``` Read three fields, not one: distance is 109, so the command took; the next hop is 10.0.13.2 out Ethernet0/1, the direct link; and the age of 00:00:48 says this is a freshly installed route, not a stale screen. Forwarding follows: ``` R1# show ip cef 172.16.3.0 255.255.255.0 172.16.3.0/24 nexthop 10.0.13.2 Ethernet0/1 ``` ### Why the One-Line Fix Is Often the Wrong One Know how blunt `distance eigrp 90 109` is before it goes in a change window. It moves every EIGRP external on the router, not just the prefix you were chasing, and AD is locally significant, so R1 now ranks routes differently from its neighbours, which is how asymmetric routing gets built quietly. At a second border it is worse: an AD change is exactly the trigger that lets an echoed route win, and you are back in Case 3\. Use it where there is one crossing point, scope it with an access-list, and read the far end afterwards. Where the design has redundancy, prefer the structural fix: tag on export, or [keep the redistributed prefix out of the return path with a distribute-list](https://www.pinglabz.com/eigrp-route-filtering-distribute-lists/). AD makes the router prefer the right answer; a filter stops the wrong answer existing. ## What This Was Captured On Both labs are CML with iol-xe nodes on IOS XE 17.18.2\. Cases 1 to 4 come from a four-router mutual-redistribution topology with two OSPF/EIGRP borders and EIGRP in named mode. Case 5 uses R1 (Et0/0 10.0.12.1/30 to R2 in OSPF, Et0/1 10.0.13.1/30 direct to R3 in EIGRP AS 100), R2 in OSPF only, and R3 with Lo0 172.16.3.3/24 set to `ip ospf network point-to-point` so it advertises the full /24, plus `redistribute ospf 1 metric 10000 100 255 1 1500`. An EEM applet wrote the R1 output to syslog before and after the fix, so both captures are the same commands minutes apart on the same box. ## Timing Matters: When the Capture Lies Two timing traps pad many redistribution tickets. First, convergence lag: after fixing a seed metric EIGRP must run DUAL and propagate, so a `show ip route` issued too fast shows the old world. Check the route's age (the fixed-state captures show fresh timers like 00:00:30 and 00:00:48, which is itself evidence the fix took). Second, flap-induced ambiguity: during feedback churn like Case 3 consecutive commands can show different next hops, which is the pathology, not your terminal. If the state will not settle, capture the oscillation rather than waiting for a clean screenshot that will never come. And in Case 5's world the RIB can be right while CEF is not, so no capture is finished until `show ip cef` agrees. ## Prevention: The Five-Line Checklist Every redistribution change, before you type it: name the single owner of each prefix range; set an explicit seed metric even where defaults would work (documentation by configuration); tag at every crossing and deny your own tags back in, from day one, not after the first loop; filter what should never cross with a prefix-list on the same route-map; and capture before/after state on both borders plus one interior router per domain, checking not just that each prefix is present but which interface it now points out of. Ten minutes, and redistribution stops being folklore. The reasoning behind each line is spread across the [redistribution guide](https://www.pinglabz.com/route-redistribution-cisco-ios-xe/) and the [loop-prevention deep dive](https://www.pinglabz.com/redistribution-loop-prevention-route-tags/). ## FAQ ### My redistribute command took, but show ip protocols shows no metric. Broken? For EIGRP, yes-in-waiting: no metric line means the infinite default applies and nothing exports (confirm with Total Redist Count: 0). For OSPF the absence is fine; the seed defaults to 20 (1 for BGP) and E2. ### Why does the better path show in the topology table but not the routing table? They answer different questions. The topology table is everything EIGRP knows; the RIB is what won against every other protocol offering that prefix. A valid composite metric next to `0 Successor(s)` and `FD is Infinity` means EIGRP computed a usable path and the RIB refused it, which at a redistribution boundary is nearly always AD 170 losing to 110. ### Can I see WHERE a looping prefix entered the domain? Yes. EIGRP externals name the originating router and source protocol in the topology table; OSPF type 5 LSAs name the advertising ASBR (`show ip ospf database external`). Between the two you can reconstruct the circle a route traveled, as in Case 3. ### Is mutual redistribution ever safe without tags? With a single border router, yes: there is no second door for the echo. The moment redundancy adds a second crossing, tags stop being optional, whatever the metrics say. The harder variants, three protocols meeting at more than one point, are worked through in [redistribution at two or more borders](https://www.pinglabz.com/multi-protocol-redistribution-ccie-scenarios/). ## Key Takeaways - Missing routes are almost always a seed metric (EIGRP's default is infinity and fails silently), a filter, or exporting a route the border installed from the wrong protocol. "Total Redist Count: 0" is the tell. - The `subnets` keyword matters on legacy IOS and is default on modern IOS XE. `show ip protocols` tells you which world you are in. - Loops show up as alternating traceroute hops and flapping pings, and the topology table names the router that echoed the prefix back. - Suboptimal routes are the silent failure: the redistributed copy arrives as an EIGRP external at AD 170, loses to the native OSPF copy at 110, and the shorter direct path sits unused with FD Infinity. - AD ranks the source, never the path. `distance eigrp 90 109` fixes one router; tags, filters and single-owner prefix design fix the cause everywhere. - Test presence and correctness separately: `show ip route` proves the prefix arrived, `show ip cef` proves where packets go. The question underneath every case here, [how the routing protocols rank against each other when two of them offer the same prefix](https://www.pinglabz.com/ip-routing/), runs through the rest of the family too. ### OSPFv3 Configuration on Cisco IOS XE URL: https://www.pinglabz.com/ospfv3-configuration-cisco-ios-xe/ Last updated: 2026-07-11T17:10:28.000Z This is the hands-on half of our OSPFv3 coverage: a complete dual-stack configuration on Cisco IOS XE, from `ipv6 unicast-routing` to verified end-to-end reachability, with every capture from a live CML lab. If you want the theory first (what changed from OSPFv2, how the LSA model works, why next hops are link-local), start with [OSPFv3 Explained](https://www.pinglabz.com/ospfv3-explained-ipv6/) and come back. Both posts live in the [OSPF](https://www.pinglabz.com/ospf/) and [IPv6](https://www.pinglabz.com/ipv6/) clusters. ## The Lab R1 and R2 run dual stack: OSPFv2 for IPv4 and OSPFv3 for IPv6 on the same links (the common transition design). The R1-R2 primary link carries 10.0.12.0/24 and 2001:db8:12::/64; loopbacks are N.N.N.N and 2001:db8::N/128\. A second R1-R2 link will host an OSPFv3 IPv4 address family adjacency later in the post. ## Step 1: Enable IPv6 Routing IOS XE does not route IPv6 by default, and this omission produces the classic silent failure (interfaces up, neighbors absent, nothing logged): ``` R1(config)# ipv6 unicast-routing ``` ## Step 2: Address the Interfaces ``` R1(config)# interface Ethernet0/1 R1(config-if)# ipv6 address 2001:db8:12::1/64 ! R1(config)# interface Loopback0 R1(config-if)# ipv6 address 2001:db8::1/128 ``` Configuring any global IPv6 address also generates the link-local FE80:: address that OSPFv3 actually speaks over. On links where you want OSPFv3 without a global prefix (pure transit), `ipv6 enable` alone creates the link-local address and is sufficient. ## Step 3: The OSPFv3 Process and Address Family Modern IOS XE uses the `router ospfv3` block with explicit address families (the older `ipv6 router ospf` syntax still parses but funnels into the same machinery; pick the AF form for anything new): ``` R1(config)# router ospfv3 1 R1(config-router)# router-id 1.1.1.1 R1(config-router)# address-family ipv6 unicast ``` Set the router ID explicitly. It is still a 32-bit value, and a router with no IPv4 addresses configured cannot derive one, refusing to start the process until you provide it (the error, when you finally spot it, reads "%OSPFv3: Router process 1 could not pick a router-id"). ## Step 4: Enable Per Interface (No network Statements) OSPFv3 drops the `network` statement entirely; membership is declared on the interface, which most engineers find more legible anyway: ``` R1(config)# interface Ethernet0/1 R1(config-if)# ospfv3 1 ipv6 area 0 ! R1(config)# interface Loopback0 R1(config-if)# ospfv3 1 ipv6 area 0 ``` Repeat the equivalents on R2 with router-id 2.2.2.2\. Costs, priorities, and timers hang off the same interface commands (`ospfv3 cost`, `ospfv3 priority`, `ospfv3 hello-interval`), mirroring their v2 counterparts. ## Step 5: Verify the Adjacency and Routes ``` R1# show ospfv3 neighbor OSPFv3 1 address-family ipv6 (router-id 1.1.1.1) Neighbor ID Pri State Dead Time Interface ID Interface 2.2.2.2 1 FULL/DR 00:00:35 1 Ethernet0/1 ``` Same states, same DR/BDR election as v2 (expect \~40 seconds of 2WAY/DROTHER during election on a fresh Ethernet segment before declaring anything broken). The routing table shows the v3 signature, link-local next hops: ``` R1# show ipv6 route ospf O 2001:DB8::2/128 [110/10] via FE80::A8BB:CCFF:FE00:1C00, Ethernet0/1 O 2001:DB8:23::/64 [110/20] via FE80::A8BB:CCFF:FE00:1C00, Ethernet0/1 R1# ping 2001:db8::4 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 3/4/9 ms ``` (That last ping crosses a redistribution boundary into an EIGRPv6 domain; the OE2 externals it rides are covered in the [redistribution guide](https://www.pinglabz.com/route-redistribution-cisco-ios-xe/).) ## Step 6: The IPv4 Address Family (Yes, Really) The same OSPFv3 process can carry IPv4 prefixes (RFC 5838 address families). The catch that stops most first attempts: the transport is still IPv6 link-local, so the interface needs an IPv6 presence even if it will only ever advertise IPv4 routes. On our second R1-R2 link: ``` R1(config)# router ospfv3 1 R1(config-router)# address-family ipv4 unicast R1(config-router-af)# exit R1(config-router)# exit R1(config)# interface Ethernet0/2 R1(config-if)# ipv6 enable R1(config-if)# ospfv3 1 ipv4 area 0 ``` Mirror it on R2, and the process now reports two independent adjacencies, one per family: ``` R1# show ospfv3 neighbor OSPFv3 1 address-family ipv4 (router-id 1.1.1.1) Neighbor ID Pri State Dead Time Interface ID Interface 2.2.2.2 1 FULL/DR 00:00:33 3 Ethernet0/2 OSPFv3 1 address-family ipv6 (router-id 1.1.1.1) Neighbor ID Pri State Dead Time Interface ID Interface 2.2.2.2 1 FULL/DR 00:00:35 1 Ethernet0/1 R1# show ospfv3 interface brief Interface PID Area AF Cost State Nbrs F/C Et0/2 1 0 ipv4 10 BDR 1/1 Lo0 1 0 ipv6 1 LOOP 0/0 Et0/1 1 0 ipv6 10 BDR 1/1 ``` Each AF maintains its own LSDB and runs its own SPF. If you are also running OSPFv2 (as this lab is), remember both protocols offer IPv4 routes at AD 110; advertise a given prefix from one of them, not both, or accept that the RIB will arbitrate and your paths may surprise you during migration. ## Common Failures, In the Order You Will Hit Them No neighbors, nothing logged `ipv6 unicast-routing` is missing. Check it before anything else. Process refuses to start No router ID could be derived (v6-only box). Set `router-id` manually. IPv4 AF adjacency never forms Interface lacks `ipv6 enable`; the AF still transports over IPv6 link-local. Stuck in 2WAY/DROTHER Usually just DR election in progress (\~40 s). Bounce the interface to force ExStart if genuinely stuck. Routes present, pings fail Check both directions: v3 installs link-local next hops, so a missing return route on the far side is invisible from the local RIB. ## Authentication, Briefly OSPFv3 has no in-protocol password fields; IOS XE offers IPsec-based authentication per interface or per area (`ospfv3 authentication ipsec spi 500 sha1 `, or the newer key-chain based `ospfv3 authentication key-chain`). Deploy it interface by interface with matched SPIs and keys on both ends; a mismatch presents as hellos sent and silently ignored, so confirm with `show ospfv3 interface` which reports the authentication state per link. ## Tuning: The Interface Knobs You Will Actually Use All the operational OSPF tuning you know from v2 exists as `ospfv3` interface commands. Costs steer paths (`ospfv3 cost 100` makes a backup link a backup); network type changes skip DR election on router-to-router Ethernet (`ospfv3 network point-to-point`, worth doing on every p2p link for faster adjacency and a cleaner LSDB); hello and dead intervals tighten detection (`ospfv3 hello-interval 1`, `ospfv3 dead-interval 4`), though the modern answer to fast failure detection is BFD rather than aggressive hellos: ``` interface Ethernet0/1 bfd interval 300 min_rx 300 multiplier 3 ospfv3 bfd ``` That pairs OSPFv3 with the same sub-second detection we measured for OSPFv2 in the [BFD deep dive](https://www.pinglabz.com/bfd-bidirectional-forwarding-detection/) (0.58 seconds versus a 34 second dead timer, same mechanism, same one-liner registration). Passive interfaces work per process (`passive-interface Ethernet0/0` under the AF) with the same caveat as v2: a passive interface advertises its prefix but hides hello-level problems, so keep it for true edge segments. ## Multi-Area OSPFv3 Areas behave exactly as in v2, declared per interface. A branch router with a backbone uplink and a stub LAN: ``` interface Ethernet0/1 ospfv3 1 ipv6 area 0 ! interface Ethernet0/3 ospfv3 1 ipv6 area 10 ! router ospfv3 1 address-family ipv6 unicast area 10 stub no-summary ``` Stub, totally stubby, and NSSA all exist with identical semantics; inter-area prefixes travel as type 3 Inter-Area Prefix LSAs instead of v2's summary LSAs, which changes the LSDB display but not the design rules. ABR placement, area sizing, and summarization strategy (`area 10 range`) carry over from your v2 instincts unchanged; the [OSPF pillar](https://www.pinglabz.com/ospf/) covers those design fundamentals. ## Dual-Stack Reality: Running v2 and v3 Together This lab intentionally runs OSPFv2 and OSPFv3 on the same interfaces, because that is what most production dual-stack networks do for years. Operationally treat them as two networks that happen to share cables. Failure domains differ (an MTU mismatch can kill v3 adjacency while v2 stays up, since v6 path MTU behavior differs), verification differs (`show ip ospf neighbor` versus `show ospfv3 neighbor`, and nothing cross-references them), and change control should name both explicitly. A pre-change snapshot that captures both neighbor tables plus both route tables takes thirty seconds and turns "IPv6 broke sometime this month" into "IPv6 broke at 14:02 during the change". ## FAQ ### Do I need network statements anywhere in OSPFv3? No, they do not exist. Interface commands are the only enablement path, which also means there is no wildcard-mask puzzle and no accidentally-matched interface: what you enabled is exactly what runs. ### Can one interface join both the IPv4 and IPv6 address families? Yes: `ospfv3 1 ipv4 area 0` and `ospfv3 1 ipv6 area 0` stack on the same interface, forming one adjacency per family over the shared link-local transport. ### Why does show ipv6 ospf still work? It aliases into the same OSPFv3 machinery for the IPv6 AF (legacy command lineage). Standardize on the `show ospfv3` forms, which display all address families and match the config syntax. ### How do I advertise a default route in OSPFv3? Same as v2: `default-information originate` (optionally `always`) under the relevant address family on the router that owns egress. ## Key Takeaways The OSPFv3 build order that avoids every classic trap: enable `ipv6 unicast-routing`, address the interfaces (or `ipv6 enable` for link-local only), create `router ospfv3` with an explicit router ID and address family, then join interfaces with `ospfv3 1 ipv6 area 0`. There are no network statements; everything is interface-scoped. The IPv4 AF works the same way and still requires IPv6 link-local transport. Verify with `show ospfv3 neighbor`, `show ospfv3 interface brief`, and `show ipv6 route ospf`, and expect FE80:: next hops everywhere. Theory and LSA details live in [OSPFv3 Explained](https://www.pinglabz.com/ospfv3-explained-ipv6/); the wider protocol family is mapped in the [OSPF pillar](https://www.pinglabz.com/ospf/) and the [IP Routing cluster](https://www.pinglabz.com/ip-routing/). ### OSPFv3 Explained: OSPF for IPv6 (and Address Families for IPv4) URL: https://www.pinglabz.com/ospfv3-explained-ipv6/ Last updated: 2026-07-11T17:10:28.000Z OSPFv3 is not "OSPF with longer addresses". It is a redesign of the protocol machinery that happens to have shipped alongside IPv6, and with address families (RFC 5838) it can carry IPv4 routes too, which surprises almost everyone the first time they see `ospfv3 1 ipv4 area 0` on an interface. This post explains what actually changed from OSPFv2 to OSPFv3, what stayed identical, how the address family model works, and what the protocol looks like on the wire and in the routing table of a live Cisco IOS XE lab. It belongs to both the [OSPF](https://www.pinglabz.com/ospf/) and [IPv6](https://www.pinglabz.com/ipv6/) clusters; the hands-on companion is [OSPFv3 Configuration on Cisco IOS XE](https://www.pinglabz.com/ospfv3-configuration-cisco-ios-xe/). ## What Stayed the Same The algorithmic heart is untouched. OSPFv3 is still link-state: routers flood LSAs, build an identical database, and run Dijkstra SPF to compute shortest paths. Areas work the same way, with area 0 as the backbone. Neighbor discovery still uses hellos, adjacencies still climb through Init, 2-Way, ExStart, Exchange, Loading, and Full, and multiaccess networks still elect a DR and BDR. Costs are still derived from bandwidth. If you can read an OSPFv2 neighbor table, the v3 one holds no surprises: ``` R1# show ospfv3 neighbor OSPFv3 1 address-family ipv6 (router-id 1.1.1.1) Neighbor ID Pri State Dead Time Interface ID Interface 2.2.2.2 1 FULL/DR 00:00:35 1 Ethernet0/1 ``` Even the router ID is still a 32-bit dotted-quad value, which is why an IPv6-only router with no IPv4 addresses must have its router ID set manually. ## What Actually Changed Per-link, not per-subnet OSPFv3 runs on links, using link-local addresses. Two neighbors no longer need to share a subnet to form an adjacency. Link-local next hops Routes install with FE80:: next hops, not global addresses. Every v3 routing table entry shows it. Addressing pulled out of the LSAs Router and Network LSAs carry pure topology now. Prefixes moved to two new LSA types: Link LSA (type 8) and Intra-Area Prefix LSA (type 9). Renumbering no longer triggers full SPF. Native multi-instance An Instance ID in every packet lets multiple OSPFv3 instances share one link, which is also the mechanism address families ride on. Authentication delegated to IPsec The v2-style auth fields left the header; OSPFv3 uses IPsec AH/ESP or, on IOS XE, key-chain based crypto per interface/area. New multicast groups FF02::5 (all OSPF routers) and FF02::6 (DR/BDR), the direct analogues of 224.0.0.5 and 224.0.0.6. The per-link design is the deepest change. In OSPFv2, the protocol is welded to IPv4 subnets: hellos source from the interface address, and mismatched subnets prevent adjacency. OSPFv3 sources everything from the link-local address, so the protocol converses happily on a link regardless of what global prefixes are configured, and the routing table shows the consequence directly: ``` R1# show ipv6 route ospf O 2001:DB8::2/128 [110/10] via FE80::A8BB:CCFF:FE00:1C00, Ethernet0/1 OE2 2001:DB8:34::/64 [110/20] via FE80::A8BB:CCFF:FE00:1C00, Ethernet0/1 ``` Every next hop is a link-local FE80:: address. Your first reaction ("where are the real addresses?") is the point: forwarding to an on-link neighbor never needed a global address in the first place. Note also the familiar route codes: O for intra-area, OI for inter-area, OE2 for externals with the same AD 110 and the same E1/E2 metric semantics as v2 (the OE2 routes above are EIGRP prefixes redistributed into OSPFv3 in our lab). ## Address Families: One Protocol, Both Stacks The original OSPFv3 specification carried only IPv6\. RFC 5838 generalized it: each address family (IPv6 unicast, IPv4 unicast, and multicast variants) maps to a reserved range of Instance IDs, so an IPv4 AF adjacency is just an OSPFv3 instance whose packets carry Instance ID 64 and whose type 9 LSAs carry IPv4 prefixes. The transport is always IPv6 link-local, even when the payload prefixes are IPv4, which means every interface participating in the IPv4 AF still needs `ipv6 enable`. On IOS XE this appears as one `router ospfv3` process with address families inside it, and interfaces join a specific AF: ``` router ospfv3 1 router-id 1.1.1.1 address-family ipv6 unicast address-family ipv4 unicast ! interface Ethernet0/1 ospfv3 1 ipv6 area 0 ! interface Ethernet0/2 ipv6 enable ospfv3 1 ipv4 area 0 ``` The neighbor table then reports each family separately, and one process can hold different adjacencies per AF on different links, as in our lab where the IPv4 AF runs on one R1-R2 link and the IPv6 AF on another: ``` R1# show ospfv3 neighbor OSPFv3 1 address-family ipv4 (router-id 1.1.1.1) Neighbor ID Pri State Dead Time Interface ID Interface 2.2.2.2 1 FULL/DR 00:00:33 3 Ethernet0/2 OSPFv3 1 address-family ipv6 (router-id 1.1.1.1) Neighbor ID Pri State Dead Time Interface ID Interface 2.2.2.2 1 FULL/DR 00:00:35 1 Ethernet0/1 R1# show ospfv3 interface brief Interface PID Area AF Cost State Nbrs F/C Et0/2 1 0 ipv4 10 BDR 1/1 Lo0 1 0 ipv6 1 LOOP 0/0 Et0/1 1 0 ipv6 10 BDR 1/1 ``` Why would you carry IPv4 in OSPFv3 at all? Operational consolidation: one protocol, one process, one set of timers and policies for both stacks, instead of running OSPFv2 and OSPFv3 side by side forever. In practice most dual-stack enterprises still run both protocols during transition (our lab does exactly that), and the AF model is the destination rather than the starting point. Each AF is a separate SPF domain with separate LSAs; enabling the IPv4 AF does not merge anything with an existing OSPFv2 process, and both will offer routes at AD 110, so migrate deliberately rather than running both sources for the same prefixes indefinitely. ## The LSA Model, Briefly For LSDB readers, the v3 lineup: Router (1) and Network (2) LSAs describe pure topology; Inter-Area Prefix (3) and Inter-Area Router (4) replace v2's summary LSAs; AS External (5) and NSSA (7) work as before; Link LSAs (8) carry a router's link-local address and on-link prefixes to neighbors on that link only; and Intra-Area Prefix LSAs (9) carry the actual prefixes that v2 used to stuff into types 1 and 2\. The practical payoff of the split is stability: adding or renumbering a prefix updates a type 9 LSA without touching the topology LSAs, so the SPF tree itself does not recompute for addressing changes. ## End to End Dual-stack proof from the lab: the IPv6 path crosses the OSPFv3 domain into an EIGRPv6 domain via redistribution at the border, and pings clean: ``` R1# ping 2001:db8::4 Sending 5, 100-byte ICMP Echos to 2001:DB8::4, timeout is 2 seconds: !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 3/4/9 ms ``` ## OSPFv2 and OSPFv3 Side by Side Transport v2: IPv4, protocol 89, sourced from interface addresses. v3: IPv6, protocol 89, sourced from link-local; multicast FF02::5 / FF02::6. Enablement v2: `network` statements under the process (or `ip ospf` per interface). v3: interface commands only; the network statement is gone. Prefix carriage v2: inside type 1/2 LSAs (addressing changes ripple into SPF). v3: types 8/9, decoupled from topology. Authentication v2: cleartext/MD5/key-chain in-protocol. v3: IPsec AH/ESP or key-chain crypto, configured per interface or area. Multiple instances per link v2: no (workarounds only). v3: yes, native Instance ID field, which also powers address families. Unchanged SPF, areas, ABR/ASBR roles, DR/BDR, neighbor state machine, AD 110, cost model, E1/E2 semantics. ## Migration Strategies for Dual-Stack Networks Three patterns cover almost every real deployment. Ships in the night (the overwhelming majority, and this lab): OSPFv2 owns IPv4, OSPFv3 owns IPv6, the processes share links and fates but exchange nothing; simple, debuggable, and the two-protocol overhead is mostly cognitive. AF consolidation: one OSPFv3 process carries both families, retiring OSPFv2 entirely; cleanest end state, but the cutover needs care because OSPFv2 and the OSPFv3 IPv4 AF both offer routes at AD 110, so migrate area by area with one protocol authoritative per prefix at every step. And v6-only islands: new segments deploy IPv6-only with OSPFv3 from day one, which sidesteps migration but demands the router-id discipline (no IPv4 address to borrow) and NAT64/translation at the edges, which belongs to the [IPv6 pillar's](https://www.pinglabz.com/ipv6/) transition coverage. Whichever path you take, monitoring has to split too: `show ip ospf neighbor` says nothing about OSPFv3, and a dashboard that only polls v2 state will happily show green while the IPv6 half of your dual-stack network is down. Poll both, alert on both, and test failover on both stacks (they can and do diverge). ## FAQ ### Is OSPFv3 backward compatible with OSPFv2? No. Different packet formats, different transport, no interoperation. A dual-stack network runs both protocols (or moves IPv4 into OSPFv3's IPv4 AF, which is still OSPFv3, not v2 compatibility). ### Why does my IPv6-only router refuse to start OSPFv3? No 32-bit router ID could be derived from an IPv4 address, because there are none. Set `router-id` manually; it is a name, not an address, and never needs to be routable. ### Do I still need DR/BDR in OSPFv3? Yes, the multiaccess machinery is unchanged, elections included (and the same \~40 second election patience applies on fresh Ethernet segments). Point-to-point network type still skips it where appropriate. ### Can OSPFv3 and EIGRPv6 redistribute into each other? Exactly like their IPv4 counterparts, seed metrics, tags, loops and all. The OE2 routes in this post's captures are EIGRPv6 prefixes redistributed into OSPFv3 at our lab's border router; the full mechanics are in the [redistribution guide](https://www.pinglabz.com/route-redistribution-cisco-ios-xe/). ## Key Takeaways OSPFv3 keeps the SPF engine, areas, neighbor states, and cost model of OSPFv2, and changes the plumbing: per-link operation over link-local addresses, prefixes relocated into type 8 and 9 LSAs, native multi-instance support, and IPsec-based authentication. Next hops are always link-local, and the router ID is still 32 bits (set it manually on v6-only boxes). Address families let one OSPFv3 process route both IPv6 and IPv4, with the IPv4 AF still transported over IPv6 link-local, at the cost of one `ipv6 enable` per interface. For the command-by-command build, verification workflow, and the gotchas that eat lab time, continue to the [OSPFv3 configuration guide](https://www.pinglabz.com/ospfv3-configuration-cisco-ios-xe/), or zoom out to the [OSPF pillar](https://www.pinglabz.com/ospf/), the [IPv6 pillar](https://www.pinglabz.com/ipv6/), and the [IP Routing cluster](https://www.pinglabz.com/ip-routing/). ### BFD: Sub-Second Failure Detection for OSPF, EIGRP, and BGP URL: https://www.pinglabz.com/bfd-bidirectional-forwarding-detection/ Last updated: 2026-07-11T17:10:27.000Z Routing protocols detect dead neighbors with hello timers measured in tens of seconds. Modern applications notice outages measured in hundreds of milliseconds. BFD (Bidirectional Forwarding Detection) closes that gap: a lightweight, protocol-agnostic liveness check that runs in the forwarding plane and tells OSPF, EIGRP, BGP, and friends about a failure in under a second. In this post we configure BFD on Cisco IOS XE, then race it against default timers with a stopwatch: the same simulated failure takes OSPF 34 seconds to notice on its own and 0.58 seconds with BFD. Real captures, from the same lab as the rest of the [IP Routing](https://www.pinglabz.com/ip-routing/) cluster. ## The Problem: Hellos Are Slow and Expensive When a link fails cleanly (carrier drops), interface-down events notify the routing protocol instantly. The dangerous failures are the dirty ones: a one-way fiber fault, a dead peer behind a healthy switch port, a wedged forwarding plane. The interface stays up, and the protocol only notices when hellos stop arriving. Defaults are generous: OSPF waits a 40 second dead interval on broadcast networks, EIGRP 15 seconds of hold time, BGP a leisurely 180 second hold timer. You can crank protocol timers down, but aggressive hellos are processed in the control plane per protocol, so fast timers multiply CPU load and false positives, once per protocol running on the link. BFD moves the liveness question out of the protocols entirely. One BFD session per link (or per neighbor pair) does the fast polling in the forwarding plane, and every registered protocol subscribes to its verdict. One detection mechanism, many clients. ## Timers and the Arithmetic of Detection Three knobs, set per interface on IOS XE: ``` interface Ethernet0/1 bfd interval 300 min_rx 300 multiplier 3 ``` `interval` is how often we send (milliseconds), `min_rx` the slowest rate we are willing to receive at, and `multiplier` how many consecutive misses declare death. Detection time is roughly the negotiated interval times the multiplier: here 300 ms times 3, so about 900 ms worst case. Carrier-grade deployments run 50 ms times 3 for 150 ms detection; 300x3 is a sane, conservative starting point for enterprise WAN links and virtual labs alike. IOS XE also defaults to echo mode: BFD sends echo packets that the neighbor's forwarding plane loops straight back without touching its CPU, which tests the actual data path. The control session then relaxes to a slow 1 second cadence, visible in the captures below. ## Tying BFD to OSPF and EIGRP Each protocol registers as a BFD client with one line. On our lab's R2-R3 boundary link (both protocols run there): ``` ! R2 interface Ethernet0/1 bfd interval 300 min_rx 300 multiplier 3 ip ospf bfd ! router eigrp PINGLABZ address-family ipv4 unicast autonomous-system 100 af-interface Ethernet0/1 bfd ``` R3 mirrors the same on its side of the link. OSPF can alternatively enable BFD process-wide with `bfd all-interfaces` under `router ospf`, and BGP registers per neighbor with `neighbor x.x.x.x fall-over bfd`. The mechanics are identical in every case: the protocol keeps its own timers as a backstop and additionally accepts BFD's down verdict immediately. ## Reading the Session ``` R2# show bfd neighbors NeighAddr LD/RD RH/RS State Int 10.0.23.3 1/1 Up Up Et0/1 R2# show bfd neighbors details | begin OurAddr OurAddr: 10.0.23.2 MinTxInt: 1000000, MinRxInt: 1000000, Multiplier: 3 Echo Rx Count: 235, Echo Rx Interval (ms) min/max/avg: 227/304/265 last: 101 ms ago Registered protocols: EIGRP CEF OSPF Min Echo interval: 300000 ``` Two things to read carefully. The control packet timers show 1000000 microseconds (1 second) because echo mode carries the fast detection; the echo counters underneath show the real 300 ms cadence at work (average 265 ms). And "Registered protocols: EIGRP CEF OSPF" is the money line: both routing protocols are subscribed to this one session. If a protocol you configured is missing from that list, the client registration did not take, and BFD will detect failures nobody reacts to. ## The Race: Failure Detection With and Without BFD To time detection honestly we need a failure the interface cannot report. An inbound deny-everything ACL on R2's boundary interface simulates a one-way path failure: R3's hellos, EIGRP packets, and BFD traffic all stop arriving while carrier stays up. First, the baseline with BFD not yet enabled, marker logged at the moment the ACL goes on: ``` *16:42:50.487: ... BASELINE-NO-BFD: KILL-LINK ACL applied on Et0/1 *16:43:02.942: %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.23.3 (Ethernet0/1) is down: holding time expired *16:43:24.528: %OSPF-5-ADJCHG: Process 1, Nbr 3.3.3.3 on Ethernet0/1 from FULL to DOWN, Neighbor Down: Dead timer expired ``` EIGRP needed 12.5 seconds, OSPF 34 seconds. For 34 seconds, R2 forwarded traffic into a black hole while its routing table swore everything was fine. Now the same failure with BFD tied to both protocols: ``` *16:45:54.773: ... WITH-BFD: KILL-LINK ACL applied on Et0/1 *16:45:55.355: %OSPF-5-ADJCHG: Process 1, Nbr 3.3.3.3 on Ethernet0/1 from FULL to DOWN, Neighbor Down: BFD node down *16:45:55.356: %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.23.3 (Ethernet0/1) is down: BFD peer down notified ``` 582 milliseconds, and both protocols converged in the same millisecond, from the same BFD verdict. The reason codes in the log tell the story: "Dead timer expired" versus "BFD node down". OSPF, defaults Dead interval 40 s on broadcast media. Measured detection: **34.0 s**. EIGRP, defaults Hold time 15 s. Measured detection: **12.5 s**. Both with BFD 300x3 Measured detection: **0.58 s**, both protocols simultaneously. About 58x faster than the OSPF dead timer. BGP, defaults Hold timer 180 s. The single biggest BFD win: `neighbor x.x.x.x fall-over bfd` takes edge failover from minutes to sub-second. ## Deployment Notes From the Field Configure BFD symmetrically; a session forms at the slower side's parameters (timers negotiate down), but mismatched multipliers create asymmetric detection you will chase for hours. On multiaccess segments through a switch, BFD is precisely what saves you from the dead-peer-behind-a-live-port failure, so prioritize it there over point-to-point fiber where carrier loss already signals instantly. Keep protocol timers at their defaults once BFD is on (they are the backstop, and tightening them buys nothing). Echo mode needs `no ip redirects` on some platforms and does not cross Layer 3 hops; for multihop BGP peerings use the separate BFD multihop template support. And in virtual labs, remember BFD state changes are honest but data-plane timing is virtualized: the 0.58 s measured here is the mechanism working as designed, not a hardware benchmark. ## BFD for BGP: The Biggest Single Win BGP's defaults (60 second keepalives, 180 second hold) made sense for 1990s route servers and make none for a modern dual-homed edge. Registration is per neighbor: ``` router bgp 65001 neighbor 203.0.113.1 fall-over bfd ``` With the interface-level BFD timers from earlier, an eBGP session over a direct link now fails over in under a second instead of up to three minutes. Two provisos. Your provider must run BFD on their side (ask; most will), and for multihop eBGP or loopback peerings the single-hop echo machinery does not apply, so IOS XE uses BFD multihop via a template and map: ``` bfd-template multi-hop PEER-TMPL interval min-tx 300 min-rx 300 multiplier 3 ! bfd map ipv4 10.99.99.2/32 10.99.99.1/32 PEER-TMPL ! router bgp 65001 neighbor 10.99.99.2 fall-over bfd multi-hop ``` ## Templates for Consistency at Scale Per-interface `bfd interval` lines drift apart over years of edits. Single-hop templates centralize the numbers, and interfaces reference the name: ``` bfd-template single-hop FAST interval min-tx 300 min-rx 300 multiplier 3 ! interface Ethernet0/1 bfd template FAST ``` Change the template once and every referencing interface follows. On platforms with hardware BFD offload, templates are also where echo mode and authentication knobs live (`echo`, `authentication keyed-sha-1`), keeping the security posture uniform. If you operate more than a handful of BFD links, templates are the difference between a design and a collection of exceptions. ## Reading a BFD Failure After the Fact When BFD takes a session down, the diagnostic code in `show bfd neighbors details` ("Local Diag") and the syslog reason tell you which side declared death and why: "Echo Function Failed" means our echoes stopped coming back (data path problem toward the peer), while "Neighbor Signaled Session Down" means the peer told us first (its problem, or a two-way issue it noticed sooner). Pair that with the registered protocol logs, which name BFD explicitly, as in our capture: "Neighbor Down: BFD node down" versus the timer-based "Dead timer expired". If you see BFD flapping without protocol churn underneath, suspect timer aggression versus platform reality: virtualized or oversubscribed boxes miss 50 ms deadlines under CPU stress, and the cure is 300x3, not disabling BFD. ## FAQ ### Does BFD replace protocol hello timers? No, it supplements them. Hellos still handle discovery, parameter negotiation, and slow-path liveness; BFD adds the fast failure verdict. Leave protocol timers at defaults once BFD carries detection. ### Echo mode or asynchronous mode? Echo (the IOS XE default where supported) tests the peer's actual forwarding plane and tolerates a busy peer CPU. Asynchronous control-only mode is the fallback where echo is unsupported (some multihop, some platforms) or where `ip redirects` interactions forbid it. ### Can one BFD session serve multiple protocols? Yes, and it is the design's whole point. Our lab session lists "Registered protocols: EIGRP CEF OSPF" and both IGPs converged from one verdict in the same millisecond. Any new client (BGP, HSRP, static routes via `ip route ... track` with BFD tracking) subscribes to the same session. ### What timers should I start with? 300 ms x 3 on enterprise WAN and virtualized gear, 50 ms x 3 on dedicated hardware where sub-200 ms failover is a requirement. Tighten only with evidence the platform holds the schedule under load. ## Key Takeaways BFD is a single, cheap, forwarding-plane liveness protocol that any registered routing protocol consumes, replacing per-protocol fast hellos. Detection is interval times multiplier (300 ms x 3 gives sub-second), echo mode tests the real data path, and one session serves OSPF, EIGRP, and BGP at once, which our captures show converging in the same millisecond. Verify with `show bfd neighbors details` and read the "Registered protocols" line like a contract. For what happens after detection (SPF, DUAL, and BGP best path reacting to the loss), see the [OSPF](https://www.pinglabz.com/ospf/), [EIGRP](https://www.pinglabz.com/eigrp/), and [BGP](https://www.pinglabz.com/bgp/) pillars, and the rest of the [IP Routing](https://www.pinglabz.com/ip-routing/) cluster for the routing machinery BFD accelerates. ### VRF-Lite on Cisco IOS XE: Configuration and Verification URL: https://www.pinglabz.com/vrf-lite-configuration-cisco-ios-xe/ Last updated: 2026-07-12T00:27:45.000Z One router, multiple routing tables. That is the entire idea behind VRF (Virtual Routing and Forwarding), and it is the foundation of everything from enterprise network segmentation to service provider L3VPNs. VRF-Lite is the version without MPLS: no label switching, no MP-BGP, just isolated routing tables on a box, each owning its own interfaces. It is how you keep a guest network, an OT network, and a corporate network on the same physical router without any of them seeing each other's routes. This guide configures VRF-Lite on Cisco IOS XE with two VRFs, proves the isolation, and then punches a deliberate, controlled hole between them with static route leaking. It is part of the [IP Routing](https://www.pinglabz.com/ip-routing/) cluster, and it is the conceptual on-ramp to [MPLS L3VPN](https://www.pinglabz.com/mpls/). ## What a VRF Actually Is A VRF is a separate RIB and FIB inside one router, plus a membership list of interfaces. A packet arriving on an interface assigned to VRF RED is looked up in RED's table and only RED's table. Prefixes can overlap freely between VRFs (two customers can both use 10.0.0.0/24), because the tables never meet. The global routing table, the one you have used your whole career, is just the default table that interfaces belong to until you say otherwise. VRF-Lite specifically means using VRFs hop by hop without MPLS: if the segmentation must span multiple routers, each inter-router link needs a separate subinterface (or VLAN) per VRF. That scales poorly beyond a few boxes, which is precisely the problem MPLS L3VPN solves with labels; the VRF construct itself is identical in both worlds. ## The Lab R1 carries the global table (OSPF toward the rest of the lab, plus the management LAN 192.168.99.0/24 where our Linux VM lives) and two VRFs: RED and BLUE. Each VRF owns a dot1q subinterface toward the VM's segment and a loopback simulating an internal network: ``` R1(config)# vrf definition RED R1(config-vrf)# rd 100:1 R1(config-vrf)# address-family ipv4 R1(config-vrf-af)# exit R1(config-vrf)# exit R1(config)# vrf definition BLUE R1(config-vrf)# rd 100:2 R1(config-vrf)# address-family ipv4 ``` Two syntax notes that bite people. Modern IOS XE uses `vrf definition` plus an explicit `address-family ipv4`; the older `ip vrf` form is IPv4-only legacy, and configs mixing the two get messy. The route distinguisher (`rd`) is mandatory before the VRF will accept interfaces; in pure VRF-Lite it is just a local identifier and never leaves the box, but the parser demands it (it becomes meaningful when you graduate to MP-BGP). Now assign interfaces. Order matters: `vrf forwarding` strips any IP address already on the interface, so set membership first, address second: ``` R1(config)# interface Ethernet0/0.10 R1(config-subif)# encapsulation dot1Q 10 R1(config-subif)# vrf forwarding RED R1(config-subif)# ip address 10.10.10.1 255.255.255.0 ! R1(config)# interface Loopback10 R1(config-if)# vrf forwarding RED R1(config-if)# ip address 10.10.99.1 255.255.255.0 ``` BLUE mirrors it: Ethernet0/0.20 with dot1q 20 and 10.20.20.1/24, Loopback20 with 10.20.99.1/24. ## Verification: Separate Tables, Really Separate `show vrf` is the inventory view (which VRFs exist, which interfaces they own): ``` R1# show vrf Name Default RD Protocols Interfaces BLUE 100:2 ipv4 Lo20 Et0/0.20 RED 100:1 ipv4 Lo10 Et0/0.10 ``` Every familiar command grows a VRF-aware form. The RED routing table contains exactly RED's world and nothing else: ``` R1# show ip route vrf RED | begin Gateway 10.0.0.0/8 is variably subnetted, 4 subnets, 2 masks C 10.10.10.0/24 is directly connected, Ethernet0/0.10 L 10.10.10.1/32 is directly connected, Ethernet0/0.10 C 10.10.99.0/24 is directly connected, Loopback10 L 10.10.99.1/32 is directly connected, Loopback10 ``` No OSPF routes, no management LAN, no BLUE prefixes. And the isolation is enforced in forwarding, not just display. Pinging BLUE's loopback from inside RED fails, because RED's table has no route to it: ``` R1# ping vrf RED 10.20.99.1 Type escape sequence to abort. Sending 5, 100-byte ICMP Echos to 10.20.99.1, timeout is 2 seconds: ..... Success rate is 0 percent (0/5) ``` Same router, both addresses locally configured, zero reachability. That is the segmentation guarantee. ## Route Leaking: Controlled Holes in the Wall Pure isolation is rarely the whole requirement; usually some shared service (management, DNS, a jump host) must reach into the VRFs. In VRF-Lite the tool is static routes with an explicit table crossing, one route per direction. Our shared segment is the management LAN in the global table; we leak each VRF loopback into global, and the management subnet into each VRF: ``` ! Global table -> VRF: point the prefix at the VRF interface ip route 10.10.99.0 255.255.255.0 Loopback10 ip route 10.20.99.0 255.255.255.0 Loopback20 ! VRF -> global table: resolve the next hop in the global RIB ip route vrf RED 192.168.99.0 255.255.255.0 192.168.99.100 global ip route vrf BLUE 192.168.99.0 255.255.255.0 192.168.99.100 global ``` The `global` keyword is the asymmetry people miss: a static route in a VRF normally resolves its next hop inside that VRF, and this keyword redirects the resolution into the global table. After leaking, RED's table shows the imported route alongside its own: ``` R1# show ip route vrf RED | begin Gateway C 10.10.10.0/24 is directly connected, Ethernet0/0.10 C 10.10.99.0/24 is directly connected, Loopback10 S 192.168.99.0/24 [1/0] via 192.168.99.100 ``` And the proof from the outside: the VM in the global segment pings into both VRFs, while a VRF-aware ping sourced from inside RED reaches back out: ``` j@llmbits:~$ ping -c 3 10.10.99.1 3 packets transmitted, 3 received, 0% packet loss ... rtt avg 2.115 ms j@llmbits:~$ ping -c 3 10.20.99.1 3 packets transmitted, 3 received, 0% packet loss ... rtt avg 2.014 ms R1# ping vrf RED 192.168.99.100 source 10.10.99.1 !!!!! Success rate is 100 percent (5/5) ``` One honest lab note: the dot1q subinterfaces above are the standard design for extending VRFs to an external device, but our VM sits behind a hypervisor port group that drops guest-tagged frames, so the working VM pings here traverse the leaked static routes rather than the tagged subinterfaces. The isolation and leaking behavior shown is identical either way; if your host-side VLAN tagging mysteriously fails, check the hypervisor vSwitch before blaming the router (on ESXi standard vSwitches, VLAN 4095 or a trunked port group is required for guest tagging). ## Routing Protocols Inside VRFs Statics scale to a point. Each VRF can also run its own routing protocol instances: OSPF processes are VRF-scoped (`router ospf 2 vrf RED`), and EIGRP named mode handles VRFs as additional address families (`address-family ipv4 unicast vrf RED autonomous-system 100`). Neighborships form per VRF over that VRF's interfaces, and everything from `show ip ospf neighbor` to `show ip protocols` gains a `vrf` argument. When the leak requirements outgrow statics entirely, that is the signal you have outgrown VRF-Lite and want MP-BGP route targets doing the import/export, which is the [MPLS L3VPN](https://www.pinglabz.com/mpls/) model. ## Living on a VRF Router: The Command Reflex The day-two skill with VRFs is remembering that every tool you reach for has a VRF-scoped form, and the unscoped form silently answers from the global table. The ones that matter daily: `ping vrf RED` and `traceroute vrf RED` (with `source` to pick which VRF address you speak from), `show ip arp vrf RED` when Layer 2 is in question, `show ip cef vrf RED ` to see what forwarding will really do, and `telnet /vrf` or SSH's `-vrf` for reaching devices inside a segment. Services need the same treatment: NTP, syslog, TACACS, SNMP traps, and DNS lookups all default to the global table, and each has a vrf keyword or source-interface knob when your management plane lives inside a VRF. The classic symptom of forgetting is a device that routes perfectly but cannot log, resolve, or authenticate. ## Dynamic Routing Inside a VRF: A Worked Example When two VRF-Lite routers share a segmented link (one subinterface per VRF), each VRF typically runs its own IGP instance across its own subinterface. OSPF scopes at the process level: ``` router ospf 10 vrf RED router-id 10.10.0.1 ! interface Ethernet0/1.10 encapsulation dot1Q 10 vrf forwarding RED ip address 10.10.12.1 255.255.255.0 ip ospf 10 area 0 ``` EIGRP named mode nests VRFs as address families inside one process (`address-family ipv4 unicast vrf RED autonomous-system 100`), one of named mode's genuine quality-of-life wins. Verification gains the vrf argument everywhere: `show ip ospf neighbor` lists per-process (hence per-VRF) adjacencies, `show ip route vrf RED ospf` shows what arrived. Everything you know about the protocols applies unchanged inside the VRF; only the table changed. ## Where VRF-Lite Stops Scaling Be honest about the growth curve. Each new VRF costs a subinterface on every inter-router link it traverses, an IGP instance (or AF) per hop, and a hand-maintained leak matrix if segments share services. Three VRFs across four routers is comfortable; twelve VRFs across forty routers is a config-generation project with an outage budget. The exit ramp is MP-BGP: VRFs stay exactly as configured here, but route targets automate the import/export decisions and MPLS labels (or VXLAN, in EVPN designs) carry the segmentation between boxes without per-VRF subinterfaces. That architecture is the [MPLS pillar's](https://www.pinglabz.com/mpls/) territory; the mental model transfers one-to-one. ## FAQ ### Can two VRFs use the same IP subnet? Yes, that is much of the point: tables never meet, so 10.0.0.0/24 can exist in RED, BLUE, and global simultaneously. The corollary is that leaking between VRFs with overlapping space requires NAT, so plan shared-services addressing to be unique across every VRF that must reach it. ### Does the RD have to be unique in VRF-Lite? Locally unique per VRF on the box, yes (the parser enforces it). Between routers it carries no meaning until MP-BGP enters; conventionally you keep a consistent scheme anyway so the eventual migration is a non-event. ### Why did my interface lose its IP when I added vrf forwarding? By design: changing table membership invalidates the address, so IOS XE strips it. Configure `vrf forwarding` first, then the address, and expect a brief traffic hit when migrating a live interface into a VRF. ### Is VRF-Lite a security boundary? It is a strong routing-plane boundary: no route, no forwarding. But it shares the box's control plane and any misconfigured leak is a bridge, so compliance-grade separation still wants filtering at the leak points and ideally physical or firewall enforcement between zones. ## Key Takeaways A VRF is a self-contained RIB/FIB plus its member interfaces; VRF-Lite is VRFs without MPLS, extended between boxes by per-VRF subinterfaces. Use `vrf definition` with an address family, set `vrf forwarding` before the IP address, and expect the RD even though it is locally meaningless in Lite deployments. Isolation is real and mutual by default; cross it only with deliberate leaks (interface-pointing statics into the VRF, `global`\-keyword statics out of it). Verify with `show vrf`, `show ip route vrf`, and VRF-aware pings with explicit sources. When leak matrices start looking like spreadsheets, move up to route targets. The wider routing context lives in the [IP Routing pillar](https://www.pinglabz.com/ip-routing/), with PBR (the other "override the table" tool) covered in the [PBR guide](https://www.pinglabz.com/policy-based-routing-pbr-cisco/). ## The Next Step: MPLS L3VPN Everything above is the same VRF construct a service provider uses, minus the machinery that makes it scale. VRF-Lite carries separation hop by hop, one subinterface per VRF per link, and the leak matrix grows until it collapses under its own weight. MPLS L3VPN keeps the identical `vrf definition` and gives the RD a real job (making overlapping customer prefixes unique in BGP), turns the route target into the thing that actually builds VPN topology, and hands the routes to MP-BGP so the backbone carries them for you. The transition is smaller than it looks. Read [MPLS L3VPN configuration step by step: VRF, RD, RT, and MP-BGP](https://www.pinglabz.com/mpls-l3vpn-configuration-step-by-step/) for the full build, [RD vs RT explained](https://www.pinglabz.com/mpls-rd-vs-rt-explained/) for the two values that stop being cosmetic the moment MPLS is involved, and the [MPLS cluster guide](https://www.pinglabz.com/mpls/) for the wider architecture. ### Policy-Based Routing (PBR): Configuration and Verification URL: https://www.pinglabz.com/policy-based-routing-pbr-cisco/ Last updated: 2026-07-11T17:10:26.000Z Routing protocols answer one question: what is the best path to this destination? Policy-based routing (PBR) exists for every time that is the wrong question. Send this application out the backup circuit. Keep guest traffic off the MPLS link. Steer one customer's flows through the firewall stack. PBR lets you route on source address, protocol, port, or packet size instead of destination alone, overriding the routing table for traffic you select. This post configures PBR on Cisco IOS XE, proves it with `debug ip policy` and route-map counters from a live lab, and covers the failure modes. It is part of the [IP Routing](https://www.pinglabz.com/ip-routing/) cluster. ## How PBR Fits Into the Forwarding Path PBR is evaluated on ingress, before the destination lookup that would normally decide the packet's fate. The interface-level command `ip policy route-map` subjects every packet arriving on that interface to a route-map. Packets that match a permit entry get the route-map's `set` action (a next hop or exit interface); packets that match a deny entry, or match nothing, fall through to normal destination-based routing. Locally generated traffic is exempt unless you separately enable `ip local policy route-map`. On IOS XE this is CEF-switched, not process-switched, so the old performance folklore about PBR melting routers no longer applies at these scales. The debug output later in this post shows the "FIB policy" path doing the work. ## The Lab Scenario R1 has two links to R2: a primary (10.0.12.0/24) and a deliberately expensive backup (10.0.112.0/24, OSPF cost 100). The routing table therefore always chooses the primary. Our policy: traffic sourced from the lab VM at 192.168.99.100 must use the long path. This is the classic "separate this one source from everyone else" pattern, the same shape as steering guest VLANs or backup jobs. ## Configuration: Three Building Blocks An ACL to describe the interesting traffic, a route-map to bind match to action, and the interface command to arm it: ``` R1(config)# ip access-list extended VM-SOURCE R1(config-ext-nacl)# permit ip host 192.168.99.100 any R1(config-ext-nacl)# exit R1(config)# route-map PBR-VM permit 10 R1(config-route-map)# match ip address VM-SOURCE R1(config-route-map)# set ip next-hop 10.0.112.2 R1(config-route-map)# exit R1(config)# interface Ethernet0/0 R1(config-if)# ip policy route-map PBR-VM ``` Two design notes. First, the ACL permits, it does not filter: a permit in a PBR ACL means "this packet is subject to the policy", nothing is dropped by it. Second, `set ip next-hop` expects a directly reachable adjacency; the router ARPs for it on the connected subnet. For next hops that might go away, prefer `set ip next-hop verify-availability` with an IP SLA tracker, or the packet blackholes while the interface stays up. ## The Set Commands, and When Each Applies set ip next-hop Forward to this adjacent address, checked before the routing table. The workhorse. set ip default next-hop Only used when the routing table has NO route for the destination (a policy of last resort). Frequently confused with the one above; the difference is the whole design. set interface / set default interface Exit interface versions of the same pair. Only sensible on point-to-point links. set ip next-hop verify-availability Ties the next hop to an IP SLA track object; the policy self-disables when the target dies instead of blackholing. ## Proof 1: The Path Actually Changes From the VM, before PBR, the traceroute follows the routing table across the primary link (second hop 10.0.12.2): ``` j@llmbits:~$ traceroute -n -q 1 -w 1 4.4.4.4 1 192.168.99.1 3.344 ms 2 10.0.12.2 4.373 ms 3 10.0.23.3 5.776 ms 4 10.0.34.4 7.364 ms ``` After applying the policy, same command, and the second hop flips to the long link: ``` j@llmbits:~$ traceroute -n -q 1 -w 1 4.4.4.4 1 192.168.99.1 5.429 ms 2 10.0.112.2 5.323 ms 3 10.0.23.3 5.855 ms 4 10.0.34.4 6.753 ms ``` The routing table on R1 still prefers 10.0.12.2 for this destination the whole time. Only policy-selected traffic moved. ## Proof 2: debug ip policy Shows the Decision With `debug ip policy` running while the VM sends traffic (use this carefully in production; on a busy interface, rate-limit with an ACL-scoped debug), each packet logs the three-step verdict: ``` *Jul 11 16:36:53.241: IP: s=192.168.99.100 (Ethernet0/0), d=4.4.4.4, len 60, FIB policy match *Jul 11 16:36:53.241: IP: s=192.168.99.100 (Ethernet0/0), d=4.4.4.4, len 60, PBR Counted *Jul 11 16:36:53.241: IP: s=192.168.99.100 (Ethernet0/0), d=4.4.4.4, g=10.0.112.2, len 60, FIB policy routed *Jul 11 16:36:53.242: IP: route map PBR-VM, item 10, permit ``` Read the last line as the audit trail: route-map name, sequence number, and result. Traffic that does NOT match logs "policy rejected -- normal forwarding", which is your clue when a match clause is too narrow. ## Proof 3: Counters for the Long Haul Debugs are for the moment; counters are for the change ticket. `show route-map` accumulates matches per sequence: ``` R1# show route-map PBR-VM route-map PBR-VM, permit, sequence 10 Match clauses: ip address (access-lists): VM-SOURCE Set clauses: ip next-hop 10.0.112.2 Policy routing matches: 22 packets, 1700 bytes ``` A policy with zero matches after a soak period means your ACL never fires (wrong source, wrong direction, wrong interface). A policy with matches but no path change means the set clause is failing (next hop unreachable), which `show ip policy` plus a ping of the next hop will confirm in seconds. ## Where PBR Goes Wrong Most PBR incidents trace to one of a handful of causes. The policy is applied to the wrong interface (it must be the INGRESS interface of the interesting traffic, not the egress you are steering toward). The next hop is not on a connected subnet, so the set clause silently never resolves. Return traffic was forgotten: PBR is unidirectional, and asymmetric paths through stateful devices (firewalls, NAT) drop the reply leg. Locally generated router traffic was expected to obey the policy but needs `ip local policy`. And policies without `verify-availability` keep forwarding into a dead next hop because the policy never consults the routing table's view of liveness. There is also an operational cost worth naming: PBR is invisible to your routing protocols. A traceroute behaves unexpectedly and nothing in `show ip route` explains why. Document policies aggressively, keep the route-map names descriptive (PBR-VM beats RM-101), and audit with `show ip policy`, which lists interface-to-route-map bindings in one screen. ## Surviving a Dead Next Hop: verify-availability Done Properly The naked `set ip next-hop` keeps forwarding as long as the ARP entry resolves, which on an Ethernet segment can outlive the actual path by a long time (the next hop's switchport is up, the router behind it is not). The production-grade pattern ties the policy to an IP SLA probe: ``` ip sla 10 icmp-echo 10.0.112.2 frequency 5 ip sla schedule 10 life forever start-time now ! track 10 ip sla 10 reachability ! route-map PBR-VM permit 10 match ip address VM-SOURCE set ip next-hop verify-availability 10.0.112.2 1 track 10 ``` When the track object goes down, the set clause is skipped and matching traffic falls through to normal routing, which is almost always the failure behavior you actually want: policy when possible, reachability always. The trailing `1` is a sequence number; you can chain multiple tracked next hops as an ordered failover list. ## Policy for the Router's Own Traffic `ip policy route-map` governs transit traffic only. Packets the router itself originates (pings from the CLI, syslog, SNMP traps, tunnel keepalives) consult the routing table directly unless you separately apply `ip local policy route-map ` in global configuration. This is both a gotcha and a tool: a local policy is the cleanest way to source management traffic out a dedicated path without touching the routing table, and forgetting the distinction is why "PBR works but my router's own pings ignore it" is a perennial forum thread. Test transit behavior from a host behind the router, not from the router's own CLI, or you are testing the wrong code path. ## PBR or Something Else? Choosing the Right Tool PBR Per-source or per-application path selection at a specific ingress point. Surgical, unidirectional, invisible to protocols. IGP metric tuning Moves ALL traffic between paths, bidirectionally, with protocol-native failover. Prefer it when the answer is "everything should use the other link". VRF segmentation When "this traffic must use that path" is really "these users live in a different network", [VRF-Lite](https://www.pinglabz.com/vrf-lite-configuration-cisco-ios-xe/) gives whole routing tables instead of per-flow exceptions. SD-WAN policies Application-aware routing with liveness built in is PBR's spiritual descendant; if you are writing dozens of PBR entries for app steering, you are reinventing [SD-WAN](https://www.pinglabz.com/sd-wan/). ## FAQ ### Does PBR override static routes and routing protocols? Yes. Policy evaluation happens before the destination lookup, so a matching permit entry wins over any routing table state, including statics and defaults. Deny entries and non-matches fall through to the table. ### Can I match on port numbers or DSCP? Anything an extended ACL expresses is matchable: ports, protocols, DSCP, packet length ranges (`match length`). Match on stable identifiers where possible; port-based application steering ages badly. ### Is PBR bad for performance on IOS XE? No. It is CEF-switched (the "FIB policy routed" lines in the debug are that path). The real costs are operational: an invisible forwarding exception that documentation and monitoring must carry. ## Key Takeaways PBR overrides destination routing for traffic you select by source, port, protocol, or size, evaluated at ingress before the FIB lookup. Build it as ACL, route-map, and `ip policy route-map` on the ingress interface, and prefer `verify-availability` whenever the next hop can fail independently of the link. Verify in three layers: traceroute for the path, `debug ip policy` for per-packet decisions, and route-map counters for sustained proof. And remember it is one-directional, so design the return path deliberately. For how policy interacts with the underlying table, the [redistribution guide](https://www.pinglabz.com/route-redistribution-cisco-ios-xe/) and the rest of the [IP Routing pillar](https://www.pinglabz.com/ip-routing/) cover the destination-based machinery PBR sits on top of. ### Redistribution Loop Prevention: Route Tags, AD, and Filtering URL: https://www.pinglabz.com/redistribution-loop-prevention-route-tags/ Last updated: 2026-08-01T19:35:25.000Z Two-point redistribution is where redistribution stops being a config exercise and starts being a design problem. The moment two border routers both translate between routing domains, every prefix has a feedback path: out through one door, around the block, and back in through the other, wearing a new metric that may beat the original. This post builds that failure live on Cisco IOS XE, captures an actual forwarding loop, and then fixes it the way production networks do: with route tags. It assumes you are comfortable with the basics from our [complete redistribution guide](https://www.pinglabz.com/route-redistribution-cisco-ios-xe/), part of the [IP Routing](https://www.pinglabz.com/ip-routing/) cluster. ## The Setup: Two Border Routers, Both Redistributing Same lab as the main guide: R1-R2 in OSPF area 0, R3-R4 in EIGRP named AS 100, with both protocols live on the R2-R3 link so that R2 and R3 are each full border routers. Both redistribute mutually with the same seed metric (`redistribute eigrp 100 subnets` into OSPF, `redistribute ospf 1 metric 1000000 1 255 1 1500` into EIGRP). The victim prefix is 172.16.4.0/24, which R4 injects into EIGRP via `redistribute connected`, making it an EIGRP external (AD 170). ## Failure Mechanism 1: Administrative Distance Betrays You Before anything visibly breaks, the routing table is already lying. R2 hears 172.16.4.0/24 twice: natively via EIGRP as an external (AD 170), and via OSPF as an E2 (AD 110) because R3 redistributed it. AD is compared before any metric, so the OSPF copy wins: ``` R2# show ip route 172.16.4.0 Routing entry for 172.16.4.0/24 Known via "ospf 1", distance 110, metric 20, type extern 2, forward metric 10 * 10.0.23.3, from 3.3.3.3, 00:02:07 ago, via Ethernet0/1 ``` An EIGRP-domain prefix, on an EIGRP border router, installed as an OSPF route. This is the quiet stage of the disease. EIGRP's designers anticipated exactly this scenario, which is why external EIGRP routes carry AD 170 instead of 90 (a route that left the domain and came back should lose to almost anything). But OSPF has no such split: internals and externals are both 110, so OSPF happily trusts routes that EIGRP originated. ## Failure Mechanism 2: The Echo in the Topology Table Because R2 installed the OSPF copy, R2's own OSPF-to-EIGRP redistribution now exports 172.16.4.0/24 back into EIGRP. R3's topology table shows the echo arriving, with the lineage data making the problem legible: one path from the legitimate originator (R4, external protocol Connected), one from R2 claiming it via OSPF: ``` R3# show eigrp address-family ipv4 topology 172.16.4.0/24 10.0.34.4 (Ethernet0/1) ... Composite metric is (131153920/163840), route is External External protocol is Connected, external metric is 0 Originating router is 4.4.4.4 10.0.23.2 (Ethernet0/0) ... Composite metric is (131727360/1310720), route is External External protocol is OSPF, external metric is 20 Originating router is 2.2.2.2 ``` Right now the legitimate path still wins, by less than half a percent of composite metric (131153920 versus 131727360). The loop is loaded; it just needs a trigger. ## The Trigger: A Routine Delay Change Someone tunes EIGRP delay for traffic engineering on the R3-R4 link (`delay 1000` under the interface, a completely ordinary change). The legitimate path's metric inflates past the echo, and R3 flips its next hop to R2: ``` R3# show ip route 172.16.4.0 Known via "eigrp 100", distance 170, metric 1029120 ... type external * 10.0.23.2, from 10.0.23.2, 00:00:04 ago, via Ethernet0/0 ``` R3 now forwards toward R2\. But R2 still holds the OSPF E2 route pointing at R3\. Each border router believes the other one knows the way: ``` R1# traceroute 172.16.4.4 numeric timeout 1 probe 1 ttl 1 10 1 10.0.12.2 3 msec 2 10.0.23.3 3 msec 3 10.0.23.2 2 msec 4 10.0.23.3 4 msec 5 10.0.23.2 3 msec 6 10.0.23.3 4 msec 7 10.0.23.2 6 msec 8 10.0.23.3 6 msec R1# ping 172.16.4.4 ..!!! Success rate is 60 percent (3/5) ``` A textbook micro-loop between 10.0.23.2 and 10.0.23.3, with the ping flapping as the protocols churn. On a real WAN this presents as intermittent reachability that mysteriously follows topology changes, one of the nastiest symptoms to chase. ## The Fix: Route Tags at Every Boundary The clean, scalable fix is to mark every route as it crosses the boundary and refuse to let it cross back. Two route-maps, applied identically on both border routers: ``` route-map EIGRP-TO-OSPF deny 10 match tag 200 route-map EIGRP-TO-OSPF permit 20 set tag 100 ! route-map OSPF-TO-EIGRP deny 10 match tag 100 route-map OSPF-TO-EIGRP permit 20 set tag 200 ! router ospf 1 redistribute eigrp 100 subnets route-map EIGRP-TO-OSPF ! router eigrp PINGLABZ address-family ipv4 unicast autonomous-system 100 topology base redistribute ospf 1 metric 1000000 1 255 1 1500 route-map OSPF-TO-EIGRP ``` Read it as a passport system. Routes leaving EIGRP for OSPF get stamped tag 100\. Routes leaving OSPF for EIGRP get stamped tag 200\. And each direction's deny clause rejects routes carrying the opposite stamp (if a route already crossed into OSPF once, tag 100 proves it, and it may not cross back). The tag travels inside the OSPF type 5 LSA and inside the EIGRP external data, so the check works on any border router, not just the one that applied the stamp. That is what makes tags scale to three, four, or N redistribution points where per-prefix filters collapse under their own maintenance burden. ## Verifying the Fix Seconds after applying the route-maps, R3 falls back to the legitimate path via R4, and the traceroute goes clean: ``` R3# show ip route 172.16.4.0 Known via "eigrp 100", distance 170, metric 5632640 ... * 10.0.34.4, from 10.0.34.4, 00:00:48 ago, via Ethernet0/1 R1# traceroute 172.16.4.4 numeric timeout 1 probe 1 ttl 1 8 1 10.0.12.2 3 msec 2 10.0.23.3 3 msec 3 10.0.34.4 4 msec ``` The tag is visible end to end. R1's RIB shows it on the installed route, and the OSPF database shows it inside the LSA itself: ``` R1# show ip route 172.16.4.0 Known via "ospf 1", distance 110, metric 20 Tag 100, type extern 2, forward metric 20 Route tag 100 R1# show ip ospf database external 172.16.4.0 | include Tag|Advertising Advertising Router: 3.3.3.3 External Route Tag: 100 ``` ## The Other Tools in the Kit Route tags Best default. Protocol-carried, self-documenting, scales to any number of boundaries. Costs two route-maps per border router. AD manipulation Raise the AD of external routes (e.g. `distance ospf external 175`) so echoes lose to originals. Simple, but silent and easy to forget; document it loudly. Prefix filtering Distribute-lists or prefix-list route-maps naming exactly what may cross. Precise but brittle: every new subnet means a filter change, and misses cause outages. EIGRP AD 170 (built in) External EIGRP routes are pre-demoted, which protects EIGRP from its own echoes. OSPF has no equivalent split, which is why the OSPF side loops first. In practice, tags plus the default AD behavior cover nearly every design. Reach for prefix filtering when policy demands that specific prefixes never leak (compliance boundaries, extranets), not as your primary loop defense. And note that OSPF's E2 metric hides path degradation: since the metric stays at 20 everywhere, OSPF gives you no metric signal that a path got worse, which is exactly why the AD comparison at the border decides everything. ## Why Not Just Fix the AD? Since the failure began with OSPF's E2 route (AD 110) beating the EIGRP external (AD 170) on the border router, an obvious counter is to demote OSPF externals below 170: ``` router ospf 1 distance ospf external 175 ``` Applied on both border routers, this makes each border prefer the native EIGRP external over the echo, and in this specific two-domain shape it does prevent the loop. It is a legitimate tool with two sharp edges. First, it is invisible at a glance: nothing in the routing table advertises that this router's AD table is nonstandard, and the next engineer will reason from defaults (document it in the config with a description and in the runbook, or expect a confused successor). Second, it is positional rather than structural: AD manipulation protects the router you configured it on, while tags protect the boundary itself, wherever and however many crossings exist. The mature pattern is tags as the structural guarantee, AD adjustment as optional belt-and-braces on top. ## Suboptimal Routing: The Loop's Quieter Sibling Not every two-point pathology loops. The same AD mechanics routinely produce paths that merely limp: a border router sends traffic for a next-door EIGRP prefix on a tour through the OSPF domain because the E2 copy (AD 110) beat the native external (AD 170). Our first capture in this post, before any delay change, was exactly that state: R2 holding an OSPF route toward a prefix its own EIGRP interface could reach directly. Nothing alarms, pings succeed, and only a latency graph or a traceroute audit reveals the detour. Tags fix this variant identically, because the echo never gets to compete in the first place. If users report "it works but it is slow" across a redistribution boundary, run the same topology-table forensics from the [troubleshooting runbook](https://www.pinglabz.com/troubleshooting-route-redistribution/) before blaming the WAN. When the symptom is louder than a detour - traffic to one prefix dying while everything else on the box is fine, and traceroute repeating the same two addresses down the column until it gives up - you are looking at the loop rather than its quieter sibling. [Finding where a routing loop closes](https://www.pinglabz.com/routing-loops-detect-and-fix/) is the faster starting point for that version, and mutual redistribution is the first cause it lists. ## Tagging Nuances Per Protocol A few implementation details keep the design honest. EIGRP named mode places the redistribute statement under `topology base` inside the address family, and the tag travels in the external data block (visible in `show eigrp address-family ipv4 topology `, "Administrator tag"). OSPF carries the tag in the type 5 LSA itself, so it survives flooding across the whole domain and is checkable anywhere with `show ip ospf database external`. Tags are 32-bit values; pick a scheme with room (we use 100 for EIGRP-origin and 200 for OSPF-origin, and larger shops encode site or domain IDs). And if a third domain joins later, the same two route-maps per boundary extend naturally: each domain gets its own tag, and each import direction denies every foreign tag that has already crossed once. ## FAQ ### Do route tags cost anything? Effectively nothing. The tag rides in fields both protocols already carry, adds no timers or adjacencies, and the route-map evaluation happens only at redistribution time, per prefix, not per packet. ### Why did the loop only appear after the delay change? The echoed route was always present in the topology table, just losing on composite metric by under one percent. Feedback designs sit in that metastable state indefinitely; any interface tune, link flap, or new parallel path can promote the echo. That is why the fix is structural (tags) rather than metric tuning, which merely re-hides the bomb. ### Can I use one tag instead of two? Yes, a single "has crossed a boundary" tag denied in both directions works for two domains. Two tags cost one extra route-map line and tell you which direction a route originally crossed, which pays for itself the first time you audit an LSDB at 2 a.m. ## Key Takeaways Two-point redistribution always creates a feedback path; whether it loops is just a question of metrics and timing, and any routine change can be the trigger. AD decides before metrics do, so OSPF E2 routes (AD 110) beat EIGRP externals (AD 170) on the border router even for prefixes EIGRP originated. Tag on the way out, deny your own tag on the way in, identically on every border router, and the loop becomes structurally impossible. Verify with `show ip route` (tag on the installed route), `show ip ospf database external` (tag in the LSA), and a traceroute that no longer ping-pongs. For the fundamentals behind these captures, see the [complete redistribution guide](https://www.pinglabz.com/route-redistribution-cisco-ios-xe/), the [redistribution troubleshooting walkthrough](https://www.pinglabz.com/troubleshooting-route-redistribution/), and the broader [IP Routing pillar](https://www.pinglabz.com/ip-routing/). ### Route Redistribution on Cisco IOS XE: The Complete Multi-Protocol Guide URL: https://www.pinglabz.com/route-redistribution-cisco-ios-xe/ Last updated: 2026-07-11T17:10:26.000Z Most networks do not run one routing protocol. Mergers bolt an EIGRP shop onto an OSPF backbone, a campus hands off to a service provider running BGP, or a legacy site keeps RIP alive years past its expiry date. Route redistribution is how you make those islands share routes, and it is one of the most error-prone tools in the [IP routing](https://www.pinglabz.com/ip-routing/) toolbox. Done casually, it produces missing routes, suboptimal paths, and full-blown routing loops. This guide walks through redistribution on Cisco IOS XE end to end: how seed metrics work, what each protocol demands before it will accept foreign routes, and what the routing table actually looks like at every stage. Every capture comes from a live CML lab, and we break it on purpose along the way (the broken cases feed our companion posts on [loop prevention with route tags](https://www.pinglabz.com/redistribution-loop-prevention-route-tags/) and [troubleshooting redistribution](https://www.pinglabz.com/troubleshooting-route-redistribution/)). ## The Lab Four IOS XE routers in a chain. R1 and R2 run OSPF area 0\. R3 and R4 run EIGRP named mode, autonomous system 100\. The R2-R3 link is the domain boundary, and both protocols are active on it, which makes R2 and R3 both border routers (that becomes important later). R4 also owns 172.16.4.0/24 on a loopback that is deliberately kept out of EIGRP's network statements, so we can demonstrate redistributing connected routes. VM (192.168.99.100) = R1 = R2 = R3 = R4 OSPF area 0: R1, R2, and the R2-R3 link | EIGRP AS 100: the R2-R3 link, R3, R4 R1-R2: 10.0.12.0/24 | R2-R3: 10.0.23.0/24 | R3-R4: 10.0.34.0/24 | Lo0 = N.N.N.N ## What Redistribution Actually Does Redistribution takes routes that one protocol installed in the routing table and injects them into another protocol's database as external routes. Two details in that sentence cause most real-world surprises. First, redistribution pulls from the routing table, not from the source protocol's internal database. If OSPF learned a route but EIGRP won the installation battle for that prefix (lower administrative distance), then "redistribute ospf" will not export it, because the table says it is an EIGRP route. Second, the receiving protocol has no idea what the original metric meant. OSPF cost and EIGRP composite metrics are different currencies, so the receiving protocol stamps a seed metric onto every imported route, and each protocol has its own rules about it. ## Seed Metrics: Each Protocol Has Different Defaults Into OSPF Default seed metric **20** (1 for BGP routes). Type E2 by default: the metric does not grow inside the domain. Subnets are included by default on modern IOS XE. Into EIGRP Default seed metric is **infinity**. No metric configured means the route is silently never advertised. This is the number one redistribution failure. Into BGP IGP metric is copied into MED. Redistribution into BGP is common at the edge; redistribution of BGP into an IGP is almost always a design mistake (full tables melt IGPs). Into RIP Default seed is also infinity. Legacy, but the same rule applies: set a hop-count seed metric or nothing is advertised. The external routes also arrive with a different administrative distance. OSPF external routes keep AD 110, but EIGRP marks externals with AD 170 instead of the internal 90 (a built-in loop defense, as we will see). ## Before Redistribution: Two Isolated Domains R1 sits purely in the OSPF domain. Its routing table knows nothing about 10.0.34.0/24, the loopbacks of R3 and R4, or 172.16.4.0/24: ``` R1# show ip route ospf | begin Gateway Gateway of last resort is not set 2.0.0.0/32 is subnetted, 1 subnets O 2.2.2.2 [110/11] via 10.0.12.2, 00:04:55, Ethernet0/1 10.0.0.0/8 is variably subnetted, 5 subnets, 2 masks O 10.0.23.0/24 [110/20] via 10.0.12.2, 00:04:53, Ethernet0/1 ``` R2, the border router, is more interesting. It sees both worlds natively and shows the two administrative distances side by side (110 for OSPF, 90 for internal EIGRP): ``` R2# show ip route | begin Gateway O 1.1.1.1 [110/11] via 10.0.12.1, 00:08:03, Ethernet0/0 D 3.3.3.3 [90/1024640] via 10.0.23.3, 00:08:42, Ethernet0/1 D 4.4.4.4 [90/1536640] via 10.0.23.3, 00:08:30, Ethernet0/1 D 10.0.34.0/24 [90/1536000] via 10.0.23.3, 00:08:42, Ethernet0/1 O 192.168.99.0/24 [110/20] via 10.0.12.1, 00:08:03, Ethernet0/0 ``` ## Configuring Redistribution on IOS XE We redistribute in both directions on R3\. EIGRP named mode buries the redistribute statement under the topology base of the address family, which trips up people used to classic mode: ``` R3(config)# router ospf 1 R3(config-router)# redistribute eigrp 100 subnets R3(config-router)# exit R3(config)# router eigrp PINGLABZ R3(config-router)# address-family ipv4 unicast autonomous-system 100 R3(config-router-af)# topology base R3(config-router-af-topology)# redistribute ospf 1 metric 1000000 1 255 1 1500 ``` The EIGRP seed metric is five values: bandwidth in Kbit/s, delay in tens of microseconds, reliability, load, and MTU (here 1 Gbps, 10 microseconds, perfect reliability, minimal load, 1500 bytes). Only bandwidth and delay influence the composite metric with default K values, but all five are mandatory. A note on the famous `subnets` keyword: on current IOS XE it is the default behavior. The running config confirms it with the line "eigrp, includes subnets in redistribution" under `show ip protocols`. On legacy IOS, omitting it silently limited redistribution to classful networks, and you will still find that behavior in the field, so we keep typing it out of habit and clarity. What happens if you skip the EIGRP metric? Nothing is advertised, with no error message. We cover the failure signature in detail in the [troubleshooting post](https://www.pinglabz.com/troubleshooting-route-redistribution/), but the short version from this lab: `show ip protocols` happily lists "Redistributing: ospf 1" while the topology table reports "Total Redist Count: 0" and R4 answers `show ip route 1.1.1.1` with "% Network not in table". ## After Redistribution: Reading the External Routes With the seed metric in place, R4 learns the entire OSPF domain as EIGRP externals: ``` R4# show ip route | begin Gateway D EX 1.1.1.1 [170/1029120] via 10.0.34.3, 00:00:30, Ethernet0/0 D EX 2.2.2.2 [170/1029120] via 10.0.34.3, 00:00:30, Ethernet0/0 D EX 10.0.12.0/24 [170/1029120] via 10.0.34.3, 00:00:30, Ethernet0/0 D EX 10.0.112.0/24 [170/1029120] via 10.0.34.3, 00:00:30, Ethernet0/0 D EX 192.168.99.0/24 [170/1029120] via 10.0.34.3, 00:00:30, Ethernet0/0 ``` Note the `D EX` code and the AD of 170 in the brackets. Internal EIGRP routes sit at 90; external ones are deliberately less trusted. In the other direction, R1 receives the EIGRP domain as OSPF type E2 externals with the default seed metric of 20: ``` R1# show ip route ospf | begin Gateway O 2.2.2.2 [110/11] via 10.0.12.2, 00:06:33, Ethernet0/1 O E2 3.3.3.3 [110/20] via 10.0.12.2, 00:00:36, Ethernet0/1 O E2 4.4.4.4 [110/20] via 10.0.12.2, 00:00:36, Ethernet0/1 O E2 10.0.34.0/24 [110/20] via 10.0.12.2, 00:00:36, Ethernet0/1 O E2 172.16.4.0 [110/20] via 10.0.12.2, 00:00:36, Ethernet0/1 ``` E2 metrics stay at 20 no matter how deep in the OSPF domain the route travels; only the forward metric (the internal cost to reach the ASBR) breaks ties. If you want the internal cost added to the seed metric as the route propagates, redistribute with `metric-type 1` to produce E1 routes. E1 is the better choice when multiple ASBRs inject the same prefixes and you want OSPF to pick the closer exit honestly. ## External Routes Carry Their History EIGRP external routes carry the source protocol, the originating router, and the external metric inside the topology table. This is gold when auditing who injected what: ``` R3# show eigrp address-family ipv4 topology 172.16.4.0/24 10.0.34.4 (Ethernet0/1) ... route is External External data: External protocol is Connected, external metric is 0 Originating router is 4.4.4.4 ``` Here 172.16.4.0/24 entered EIGRP at R4 via `redistribute connected`, and every EIGRP router downstream can see that lineage. ## One Border Router or Two? Everything above used a single redistribution point, which is inherently loop-free: routes cross the boundary in one place, and there is no second door for them to sneak back through. The moment you add a second border router for redundancy (a very reasonable thing to want), you create a feedback path. A prefix redistributed from EIGRP into OSPF at R3 can travel through OSPF to R2 and be redistributed back into EIGRP, now wearing a fresh seed metric that may beat the original route. In this lab, enabling the same two-way redistribution on R2 produced a genuine forwarding loop between R2 and R3 within seconds of a routine delay change. The traceroute ping-pongs between the two border routers until the TTL dies: ``` R1# traceroute 172.16.4.4 numeric timeout 1 probe 1 ttl 1 10 1 10.0.12.2 3 msec 2 10.0.23.3 3 msec 3 10.0.23.2 2 msec 4 10.0.23.3 4 msec 5 10.0.23.2 3 msec 6 10.0.23.3 4 msec ``` The full anatomy of that loop, and the route-tag design that fixes it permanently, gets its own deep dive: [Redistribution Loop Prevention: Route Tags, AD, and Filtering](https://www.pinglabz.com/redistribution-loop-prevention-route-tags/). The one-sentence rule: any time you run two or more redistribution points, tag routes as they cross the boundary and deny your own tags from coming back. ## Verification Commands That Earn Their Keep show ip route The detail view names the source protocol, AD, metric, and (for tagged routes) the route tag. show ip protocols Lists what each protocol redistributes and the configured seed metrics. First stop in any redistribution audit. show eigrp address-family ipv4 topology Shows every offered path with external data: source protocol, originating router, external metric. show ip ospf database external The type 5 LSAs themselves: advertising ASBR, metric type, and external route tag. ## Design Rules Worth Stealing Prefer a single redistribution point unless you truly need redundancy, and when you need two, deploy route tags from day one rather than after the first loop. Always set an explicit seed metric into EIGRP and RIP (make it pessimistic, so external paths lose to internal ones on honest terms). Use E1 externals in OSPF when multiple ASBRs advertise the same prefixes. Filter what you redistribute with a route-map even when you think you want everything, because "everything" grows over time. And never redistribute BGP into an IGP without an aggressively specific filter. ## Controlling What Crosses: Route-Maps at the Boundary Bare redistribute statements move everything, and "everything" is rarely the intent six months later. A route-map on the redistribute command gives you a policy point for three jobs at once: filtering (only these prefixes may cross), tagging (mark everything that crosses, the loop-prevention backbone), and metric shaping (give this class of routes a worse seed so the backup domain stays a backup). The full tagged configuration from this lab's border routers looks like this and costs four short stanzas: ``` route-map EIGRP-TO-OSPF deny 10 match tag 200 route-map EIGRP-TO-OSPF permit 20 set tag 100 ! router ospf 1 redistribute eigrp 100 subnets route-map EIGRP-TO-OSPF ``` Filtering by prefix uses `match ip address prefix-list` in the same structure. One subtlety worth engraving somewhere: a route-map applied to redistribution evaluates routes, not packets, so the counters in `show route-map` tick per prefix evaluation and are your fastest proof of whether a filter clause is actually catching anything. ## Redistributing Between OSPF Processes Not every boundary separates different protocols. Two OSPF processes on one router (a merger where both companies ran OSPF, or a VRF handoff) redistribute exactly like foreign protocols: `redistribute ospf 2` under `router ospf 1` and vice versa, routes arrive as E2 with everything that implies, and two-point designs loop just as enthusiastically. There is no special kinship between OSPF processes; the SPF domains are fully separate, and all the seed metric, tagging, and AD logic in this guide applies unchanged. ## What About the Default Route? Redistribute statements do not move the default route; each protocol has a dedicated mechanism instead. OSPF wants `default-information originate` (with `always` if the ASBR should advertise it even without holding a default itself), and EIGRP named mode is happiest with a summary address or a redistributed static to 0.0.0.0/0\. In a two-domain design, decide deliberately which domain owns internet egress, originate the default only there, and let specific routes flow the other way; two domains both originating defaults into each other is a slow-motion loop generator that tags will not save you from, because the prefix is legitimately different on each origination. ## FAQ ### Do I need the subnets keyword on IOS XE? Functionally no on current releases (the router reports "includes subnets in redistribution"), but it is free to type, self-documenting, and required on the legacy IOS still running in plenty of racks. Type it. ### Why AD 170 for EIGRP externals but 110 for OSPF externals? EIGRP deliberately distrusts routes that left the domain and came back, which single-handedly prevents many feedback loops on the EIGRP side. OSPF makes no internal/external AD distinction by default, which is why the OSPF side of a two-point design is where loops are born; you can retrofit the split with `distance ospf external 175`. ### Should I redistribute BGP into my IGP? Almost never wholesale. A full table is two orders of magnitude beyond what IGP flooding is designed for. Redistribute a filtered handful of prefixes if you must, or better, originate a default from the edge and let the IGP carry only internal state. ### E1 or E2 externals? E2 (the default) keeps the seed metric flat everywhere, fine when one ASBR injects the routes. E1 adds internal cost as the route propagates, correct when multiple ASBRs inject the same prefixes and routers should prefer their nearer exit. ## Key Takeaways Redistribution copies routes from the routing table into another protocol as externals, stamped with a seed metric. EIGRP's default seed is infinity, so forgetting the metric silently advertises nothing, while OSPF defaults to a flat E2 metric of 20\. External routes carry reduced trust (EIGRP AD 170) and, in EIGRP's case, full lineage data in the topology table. Single-point redistribution is naturally loop-free; two-point redistribution demands route tags. For the wider context of how routing protocols fit together, head back to the [IP Routing pillar](https://www.pinglabz.com/ip-routing/), and when redistribution misbehaves, the [troubleshooting guide](https://www.pinglabz.com/troubleshooting-route-redistribution/) walks the failure signatures with live captures. ### Troubleshooting EIGRP Route Advertisement and Missing Routes URL: https://www.pinglabz.com/troubleshooting-eigrp-missing-routes/ Last updated: 2026-07-11T16:05:19.000Z The neighbors are up, the interfaces are clean, and yet a route that should be in the table simply is not. Missing-route problems are sneakier than adjacency problems because nothing is visibly broken; the protocol is running exactly as configured, and the configuration is quietly wrong. This guide is a systematic walk through every place an EIGRP route can go missing, with real failure output from a CML lab. Adjacency problems are the prerequisite check, covered in [troubleshooting EIGRP neighbor adjacencies](https://www.pinglabz.com/troubleshooting-eigrp-neighbor-adjacencies/); protocol fundamentals live in the [EIGRP complete guide](https://www.pinglabz.com/eigrp/). ## Follow the Route: The Five Places It Can Die A prefix travels a pipeline: covered by a network statement on the source router, advertised out an interface, accepted by the neighbor, chosen by DUAL, installed in the RIB. It can die at each stage, and the fastest troubleshooting order matches the pipeline. The single most useful triage question: is the route in the *topology table* (`show ip eigrp topology`) but not the *routing table*, or not in the topology table at all? Topology-but-not-RIB points at administrative distance or a better competing source. Missing from topology entirely points at the source, a filter, or split horizon. ## Cause 1: The Network Statement Does Not Cover the Interface EIGRP's `network` command selects *interfaces*, not prefixes: any interface whose address falls inside the statement gets enabled, and its connected network gets advertised. The two classic mistakes are a wildcard mask that does not cover a new interface, and assuming a classful `network 10.0.0.0` statement was narrower than it is. Check what the process actually matched: ``` R3# show ip protocols | section eigrp Routing Protocol is "eigrp 100" Routing for Networks: 3.3.3.3/32 10.0.0.0 Routing Information Sources: Gateway Distance Last Update 10.0.13.1 90 00:00:57 10.0.35.1 90 00:00:57 ``` If the interface holding the missing prefix is not covered under Routing for Networks, nothing downstream matters. `show ip eigrp interfaces` confirms which interfaces EIGRP is actually running on. Named mode uses the same network statements (syntax differences are in the [EIGRP configuration guide](https://www.pinglabz.com/eigrp-configuration-cisco/)). ## Cause 2: A Distribute List Is Eating It This one hides in plain sight because the router is doing exactly what someone once told it to do. In the lab, R3 filters two of its four /24s, and R1 comes up short: ``` R1# show ip route eigrp | include 10.1. D 10.1.0.0/24 [90/3584000] via 10.0.13.2, 00:02:18, Ethernet0/2 D 10.1.1.0/24 [90/3584000] via 10.0.13.2, 00:02:18, Ethernet0/2 ``` Where did 10.1.2.0/24 and 10.1.3.0/24 go? One command on the advertising router answers it: ``` R3# show ip protocols | include filter Outgoing update filter list for all interfaces is (prefix-list) BLOCK-LAB-NETS Incoming update filter list for all interfaces is not set R3# show ip prefix-list BLOCK-LAB-NETS ip prefix-list BLOCK-LAB-NETS: 3 entries seq 5 deny 10.1.2.0/24 seq 10 deny 10.1.3.0/24 seq 15 permit 0.0.0.0/0 le 32 ``` Check *both* routers: an out filter on the advertiser and an in filter on the receiver produce identical symptoms. Distribute lists, prefix lists, and route maps each have their own show commands, and the full filtering toolkit is covered in [EIGRP route filtering with distribute lists and route maps](https://www.pinglabz.com/eigrp-route-filtering-distribute-lists/). ## Cause 3: A Summary Is Suppressing It Summarization deliberately hides component routes, and sometimes it hides more than intended. The fingerprint: the specific prefix is missing but a covering summary is present: ``` R2# show ip route eigrp | include 10.1. D 10.1.0.0/22 [90/4096000] via 10.0.12.1, 00:00:16, Ethernet0/1 ``` Nothing is unreachable here, and that is what makes it subtle: traffic still flows via the /22, but through whatever path the *summary* takes, which after a topology change may not be the path you expect. The extreme case is a default-route summary: `summary-address 0.0.0.0/0` suppresses every more-specific prefix on that interface, and neighbors with a second path suddenly prefer wildly indirect specifics from elsewhere. We captured exactly that effect in the [summarization and default routes guide](https://www.pinglabz.com/eigrp-summarization-default-routes/). Find the suppression on the advertising router with `show run | include summary-address` and look for the telltale Null0 entry: ``` R3# show ip route | include Null0 D 10.1.0.0/22 is a summary, 00:00:24, Null0 ``` ## Cause 4: Split Horizon on a Hub Split horizon forbids advertising a route out the interface it was learned on. On point-to-point links it is invisible protection. On a hub-and-spoke topology where two spokes share the hub's single multipoint interface (DMVPN, classic frame relay, some tunnel designs), it means spoke A's routes are never advertised to spoke B: both live on the same hub interface. The fix belongs on the hub: ``` Hub(config)# interface Tunnel0 Hub(config-if)# no ip split-horizon eigrp 100 ``` In named mode it is `no split-horizon` under the af-interface. If spokes can see the hub's routes but never each other's, split horizon should be your first suspicion. (Tunnel topologies stack additional quirks on top; see [routing protocols over GRE](https://www.pinglabz.com/routing-protocols-over-gre/).) ## Cause 5: The Route Lost to a Better Source Sometimes EIGRP has the route and the RIB still shows something else. Administrative distance decides: EIGRP internal 90, external 170, and anything with a lower AD (a static at 1, OSPF at 110 loses to internal but external EIGRP loses to OSPF) takes the slot. The diagnostic is comparing tables: ``` R1# show ip eigrp topology 10.1.0.0/24 ...State is Passive, 1 Successor(s), FD is 458752000... R1# show ip route 10.1.0.0 255.255.255.0 Known via "static", distance 1, metric 0 ``` An EIGRP route present in topology but beaten in the RIB is working as designed; the question becomes whether that static (or that redistributed route at AD 170) should exist. The AD ladder and its edge cases are laid out in [EIGRP administrative distance](https://www.pinglabz.com/eigrp-administrative-distance/), and the external-route AD interplay matters most at redistribution boundaries with [OSPF](https://www.pinglabz.com/ospf/) or [BGP](https://www.pinglabz.com/bgp/). ## Cause 6: A Stub Router Is Not Advertising What You Think An EIGRP stub advertises only connected and summary routes by default. If a spoke learned a route from one neighbor and you expect it to pass it to another, stub configuration says no, silently and by design. The neighbor's view exposes it: ``` R1# show ip eigrp neighbors detail 1 10.0.13.2 Et0/2 14 00:01:27 1 100 0 5 Version 28.0/2.0, Retrans: 1, Retries: 0, Prefixes: 6 Topology-ids from peer - 0 Topologies advertised to peer: base ``` On a real stub peer, this output includes a line like `Stub Peer Advertising (CONNECTED SUMMARY) Routes` plus `Suppressing queries`. When a route is missing beyond a spoke, check whether the spoke is a stub before checking anything else on it. Design intent and configuration are in [EIGRP stub routing](https://www.pinglabz.com/eigrp-stub-routing/). **Missing everywhere**network statement or interface down at the source **Missing past one router**distribute list out, summary suppression, stub, or split horizon at that router **Missing on one receiver only**distribute list in, or the neighbor relationship itself **In topology, not in RIB**lost on administrative distance to a static or another protocol **Replaced by a bigger prefix**summarization upstream; find the Null0 route on the summarizing router ## Cause 7: Redistribution Without a Metric The classic-mode trap that has consumed entire afternoons: routes redistributed into EIGRP from another protocol arrive with no EIGRP metric of their own, and classic EIGRP assigns them an infinite metric by default. Infinite metric means the route is not advertised at all. No error, no log; the `redistribute ospf 1` line sits in the config looking correct while producing nothing. ``` R1(config)# router eigrp 100 R1(config-router)# redistribute ospf 1 ! Nothing appears anywhere. Now with seed metric: R1(config-router)# redistribute ospf 1 metric 10000 100 255 1 1500 ``` The five seed values are bandwidth, delay, reliability, load, and MTU; `default-metric` under the process does the same job for all redistributed sources at once. Named mode softened this trap by applying a default seed metric for some sources, but relying on that is how configs break during migrations; set the seed explicitly, always. Redistributed routes then appear as `D EX` with AD 170, which feeds directly into the AD competition in Cause 5: an external EIGRP route loses to OSPF's 110, so redistributing *into* EIGRP at one boundary and expecting it to win against OSPF at another is a design error, not a bug. Two-way redistribution between [OSPF](https://www.pinglabz.com/ospf/) and EIGRP at multiple points needs route tags and filtering to avoid loops, which is its own article-sized topic. ## Cause 8: auto-summary (the Legacy Landmine) On modern IOS XE, `no auto-summary` is the default and this cause is nearly extinct, but older gear and pasted-forward configs keep it alive. With auto-summary enabled, EIGRP summarizes to classful boundaries wherever a network statement crosses one: your 10.1.2.0/24 becomes 10.0.0.0/8 as it leaves the site. The symptom is distinctive: a huge classful prefix in tables where you expected specifics, and traffic to discontiguous subnets of the same major network blackholing (two sites both advertising 10.0.0.0/8 toward the core). One glance at `show ip protocols` settles it: ``` R3# show ip protocols | include Automatic Automatic Summarization: disabled ``` If that line says enabled on anything in your path, add `no auto-summary` and move on with your life. ## The Method, End to End - Confirm adjacencies first; a missing neighbor explains everything downstream. - On the source router: `show ip protocols` (Routing for Networks), `show ip eigrp interfaces`. Is the prefix even entering EIGRP? - Walk the path router by router with `show ip eigrp topology `. The first router where it disappears is where the problem lives. - On that router: `show ip protocols | include filter`, `show run | include summary-address`, stub status in `show ip eigrp neighbors detail`, split horizon on multipoint interfaces. - If topology has it but the RIB does not: compare administrative distances and look for competing sources. ## Key Takeaways - Split the problem in half immediately: topology table versus routing table tells you whether to hunt upstream (source, filters, split horizon) or locally (AD, competing protocols). - `show ip protocols` is the densest triage screen in IOS: networks covered, filters active, sources learned from, and AD, all in one page. - Missing routes usually mean someone configured exactly this: a filter, a summary, or a stub. The router is obeying; find the instruction. - Split horizon on multipoint hub interfaces is the classic spoke-to-spoke route killer. - Walk the path prefix-first with `show ip eigrp topology `; the first router where it vanishes owns the answer. Missing routes are where EIGRP troubleshooting turns into detective work, and the pipeline model keeps the search finite. For everything upstream of the mystery, from DUAL to design patterns, the [EIGRP complete guide](https://www.pinglabz.com/eigrp/) is home base. ### Troubleshooting EIGRP Neighbor Adjacencies: The Loud Failure and the Silent One URL: https://www.pinglabz.com/troubleshooting-eigrp-neighbor-adjacencies/ Last updated: 2026-08-01T18:36:24.000Z Every EIGRP problem starts with the same question: are the neighbors up? No adjacency means no routes, and a flapping adjacency means a flapping network. The good news is that adjacency failures are a finite list. A handful of causes account for nearly every case, and each one leaves a distinct fingerprint in the logs and the show commands. This guide works through the real ones, broken deliberately on live CML labs so you can see the exact output each failure produces. The two headline cases, a K-value mismatch and a wrong autonomous system number, were captured on iol-xe routers running IOS XE 17.18.2, and they behave in opposite ways. One fills your log every few seconds. The other produces nothing at all, anywhere, ever. Knowing which is which is most of the diagnosis. For the protocol theory underneath the adjacency, [how EIGRP builds and maintains its topology table](https://www.pinglabz.com/eigrp/) is the hub. The requirements list itself, meaning what has to match before two routers will peer at all, lives in [the five things that must match before EIGRP will peer](https://www.pinglabz.com/eigrp-neighbor-requirements/). This article is the other half: recognizing each failure from the router's point of view. ## What This Was Captured On Three routers, all iol-xe nodes in Cisco Modeling Labs running IOS XE 17.18.2\. R1 is the observer, and it has two broken adjacencies at the same time, deliberately broken in two different ways: - **R1 Et0/0 10.0.12.1/30 to R2 Et0/0 10.0.12.2/30.** Both run `router eigrp 100`, but R2 was given `metric weights 0 1 1 1 0 0`, which turns K2 on. K-value mismatch. - **R1 Et0/1 10.0.13.1/30 to R3 Et0/0 10.0.13.2/30.** Layer 3 is fine and the subnet is shared, but R3 runs `router eigrp 200` against R1's 100\. AS mismatch. Both links are up, both interfaces sit in the right /30, and there is no ACL and no authentication anywhere. The only thing wrong on each link is one parameter. Broken state, fix and recovery were captured on R1 with on-box EEM applets writing `show` output to syslog, so everything below is what the router actually printed. The later sections (passive interface, authentication, timers) come from a companion CML lab built in EIGRP named mode, which is why their captures carry the longer `VR(PINGLABZ) Address-Family` header. Same protocol, different syntax. ## The Checklist, In the Order That Finds Problems Fastest Two routers need layer 3 connectivity on a shared primary subnet, the same autonomous system number, matching K values, matching authentication, and hellos actually flowing (no passive-interface, no ACL eating protocol 88, multicast 224.0.0.10 working). Start every investigation the same way, on both ends: ``` R1# show ip eigrp neighbors EIGRP-IPv4 Neighbors for AS(100) H Address Interface Hold Uptime SRTT RTO Q Seq (sec) (ms) Cnt Num 1 10.0.13.2 Et0/1 13 00:00:50 1 100 0 3 0 10.0.12.2 Et0/0 13 00:00:54 2 100 0 3 ``` That is what healthy looks like (R1, after both faults were repaired). Hold counting down and resetting, uptime climbing, SRTT in single-digit milliseconds, Q count at 0\. A neighbor missing entirely, or cycling up and down, sends you into the failure catalog below. Second stop, always: `show logging | include DUAL`. When EIGRP has an opinion about why a neighbor went away it says so in plain text, and each scenario below has different phrasing. When it has no opinion, that silence is itself diagnostic, and the next two sections are about the difference. ## Failure 1: K-Value Mismatch (Loud) K values weight the composite metric formula, and both routers must agree or the metrics they exchange would be incomparable. R1 to R2 fails this test, and R1 will not stop telling you about it: ``` R1# show logging | include DUAL *Jul 20 22:50:30.411: %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.12.2 (Ethernet0/0) is down: K-value mismatch *Jul 20 22:50:43.855: %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.12.2 (Ethernet0/0) is down: K-value mismatch ``` Read that message carefully, because the wording explains the mechanism. It says *Neighbor is down*, which means a neighbor object briefly existed. R2's hello arrives, R1 accepts it (the AS in the header matches, so the packet is for this process), parses the parameters, finds K2 set where its own K2 is zero, and tears the relationship down before it ever reaches the update phase. The next hello repeats the cycle, roughly every 13 seconds in our capture, which is why this failure is impossible to miss on a console. Meanwhile the neighbor table is empty, and this part is worth sitting with: ``` R1# show ip eigrp neighbors EIGRP-IPv4 Neighbors for AS(100) R1# ``` Nothing. Not R2, and not R3 either. R1 should have two neighbors and has zero, but the log names exactly one of them, 10.0.12.2 on Ethernet0/0\. The other link is invisible. The compare command is `show ip protocols`, which prints the AS number and all five K values on one screen. Run it on both ends and put them side by side: ``` R1# show ip protocols Routing Protocol is "eigrp 100" Outgoing update filter list for all interfaces is not set Incoming update filter list for all interfaces is not set EIGRP-IPv4 Protocol for AS(100) Metric weight K1=1, K2=0, K3=1, K4=0, K5=0 <-- R1 defaults; R2 had K2=1 Soft SIA disabled NSF-aware route hold timer is 240 Router-ID: 10.0.13.1 Topology : 0 (base) Active Timer: 3 min Distance: internal 90 external 170 Maximum path: 4 Maximum metric variance 1 Total Prefix Count: 2 ``` K1=1, K3=1, everything else 0 is the default and the right answer for almost everyone. K2 (load) and K5 (reliability) pull live, fluctuating interface statistics into the metric, which is how you build a network whose routes recompute because somebody started a backup job. For the full derivation, [what each EIGRP K value actually weights](https://www.pinglabz.com/eigrp-metric-k-values/) covers the formula properly. The fix is to make them agree. On R2, restore the defaults with `metric weights 0 1 0 1 0 0` (the leading 0 is TOS, always 0, then K1 through K5), or simply `no metric weights`. The adjacency returns on its own within a hello or two, with nothing to bounce. ## Failure 2: AS Number Mismatch (Silent) Now the opposite personality, and this is the contrast the whole article is built on. R3 sits on the same /30 as R1, with a working link, running `router eigrp 200` against R1's `router eigrp 100`. Here is what R1 logged about R3 during the entire broken window: nothing. Not one line, and that is the finding. The log excerpt above is the complete DUAL output from that capture, and every entry in it names 10.0.12.2\. The 10.0.13.2 link, equally broken, generated no syslog at any severity. The reason is structural. The AS number sits in the EIGRP packet header and is checked before anything else. A router running AS 100 does not treat an AS 200 hello as a badly configured peer, it treats it as a packet for a process it is not running, and drops it. No neighbor object is created, so there is nothing to declare down and nothing to log. They are separate routing domains that happen to share a wire. So the evidence for a silent failure is entirely negative: a neighbor that should be in the table is not, with no explanation anywhere. That combination is the signal to stop grepping the log and start comparing configuration by inspection. The fastest tell is the header line of `show ip eigrp neighbors`. Notice that R1's empty output above still prints `EIGRP-IPv4 Neighbors for AS(100)`. Run the same command on the far end and read its header. If it says `AS(200)`, you are finished, and it took four seconds. `show ip protocols` gives the same answer with more context (its process line reads `Routing Protocol is "eigrp 200"`), and `show ip protocols summary` gives one line per process if the box runs several. Easy to check, easy to overlook, and the number one cause of "EIGRP just will not come up" on a link somebody else built. ## Loud Versus Silent, Side By Side Both faults produce the identical top-level symptom, an empty neighbor table. Everything that distinguishes them is in what the router does or does not say: K-value mismatch LoggedYes, repeatedly Messageis down: K-value mismatch Hello handlingAccepted, then rejected Peer IP visibleYes, named in the log How you find itRead the log AS number mismatch LoggedNo, nothing at all MessageNone Hello handlingDropped at the header Peer IP visibleNo, absence only How you find itCompare both ends The practical consequence: a quiet log does not clear EIGRP. It is exactly what a wrong AS number, a passive interface, or an ACL blocking protocol 88 looks like. ## The Fix, and What Recovery Looks Like Both repairs are applied on the far ends. R2 restores the default weights, R3 moves from AS 200 to AS 100: ``` R2(config)# router eigrp 100 R2(config-router)# metric weights 0 1 0 1 0 0 R3(config)# no router eigrp 200 R3(config)# router eigrp 100 R3(config-router)# network 10.0.13.0 0.0.0.3 ``` R1 was not touched. Within seconds it picked up both peers, and this time both events are logged, because "up" is always loud: ``` R1# show logging | include DUAL *Jul 20 22:53:11.686: %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.12.2 (Ethernet0/0) is up: new adjacency *Jul 20 22:53:16.061: %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.13.2 (Ethernet0/1) is up: new adjacency ``` The neighbor table then reads exactly as the healthy example at the top of this article. Note that `is up: new adjacency` fires for the silent failure too: once the AS matches, the router that had nothing to say suddenly has plenty, and that line is your confirmation the fix landed. One warning before copying the R3 snippet into production: `no router eigrp 200` throws away everything scoped to that process, interface authentication included. ## Failure 3: Passive Interface (Half Silent) Passive interface tells EIGRP to advertise a network without speaking the protocol on it. Correct on user-facing LANs (it is standard hardening); an outage when it lands on a transit link. In named mode: ``` R1(config)# router eigrp PINGLABZ R1(config-router)# address-family ipv4 unicast autonomous-system 100 R1(config-router-af)# af-interface Ethernet0/1 R1(config-router-af-interface)# passive-interface ``` Two views of the damage. R1's interface list simply loses Et0/1 (passive interfaces are not EIGRP interfaces anymore): ``` R1# show ip eigrp interfaces EIGRP-IPv4 VR(PINGLABZ) Address-Family Interfaces for AS(100) Interface Peers Un/Reliable Un/Reliable SRTT Un/Reliable Flow Timer Routes Lo0 0 0/0 0/0 0 0/0 0 0 Et0/2 1 0/0 0/0 1 0/2 50 0 Et0/0 0 0/0 0/0 0 0/0 0 0 ``` And the far side, R2, sees a normal-looking teardown with a reason that points at the peer: ``` R2# show logging | include DUAL *Jul 11 05:32:43.564: %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.12.1 (Ethernet0/1) is down: Interface PEER-TERMINATION received ``` So the diagnostic pattern is asymmetric, and that asymmetry tells you which end to log into: one side has the interface missing from `show ip eigrp interfaces`, the other side got a peer termination. Check passive configuration with `show ip protocols` (classic mode lists passive interfaces explicitly) or `show run | section eigrp` in named mode. ## Failure 4: Authentication Mismatch (Loud If You Look) With authentication enabled on one side only, or with mismatched keys, hellos are received and rejected. The receiver logs it precisely, though the two most useful lines are debug-level, so they appear only if you are already looking: ``` R3# show logging | include authentication|DUAL *Jul 11 05:23:19.445: EIGRP: pkt key id = 1, authentication mismatch *Jul 11 05:23:19.445: EIGRP: Dropping peer, invalid authentication *Jul 11 05:23:19.447: %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.35.1 (Ethernet0/2) is down: Auth failure ``` Check both sides with `show ip eigrp interfaces detail | include Authentication`: mode and key chain must match. One trap deserves special mention because we hit it live: removing a classic EIGRP process (`no router eigrp 100`) also deletes the AS-scoped `ip authentication` commands from every interface. Recreate the process and auth is simply gone on one side, failing against a peer that still requires it. Full configuration coverage, including key rotation, is in the [EIGRP authentication guide](https://www.pinglabz.com/eigrp-authentication-md5-sha/). ## Failure 5: The Plumbing (Subnet, ACLs, MTU, Unidirectional Links) If none of the above match your symptoms, drop a layer. Mismatched primary subnets log `not on common subnet` warnings. An inbound ACL that forgot to permit EIGRP (IP protocol 88) or multicast 224.0.0.10 blocks hellos entirely, and on IOL and IOS XE inbound ACLs are evaluated before almost everything. A neighbor cycling through retransmissions with Q Cnt climbing, then `retry limit exceeded`, means hellos pass but unicast updates do not: MTU mismatch or a unidirectional link. On hub-and-spoke or tunnel topologies, confirm multicast actually works across the transport (GRE handles it; the details live in [routing protocols over GRE](https://www.pinglabz.com/routing-protocols-over-gre/)). **down: K-value mismatch**Compare metric weights both sides, revert to defaults **Silence, no logs at all**AS mismatch or passive-interface or ACL: check neighbor table headers, then EIGRP interface list, then ACLs **down: Auth failure**Mode or key mismatch; audit interfaces after any process removal **down: Interface PEER-TERMINATION received**Far side went passive or removed the interface from EIGRP **down: retry limit exceeded**Hellos pass, updates do not: MTU, unidirectional link, or congestion ## The Mismatch That Is Not a Failure: Hello and Hold Timers Here is the one that trips up engineers coming from OSPF: EIGRP hello and hold timers do *not* need to match between neighbors. Each router announces its own hold time inside its hellos, and the neighbor simply honors it. Two routers with completely different timer sets will form and keep a perfectly stable adjacency. So if you have been staring at a timers difference as your root cause, stop; it is not the problem, and "fixing" it will not bring the neighbor up. What timers *can* do is destabilize an adjacency that already works. A hold time shorter than the loss and jitter on the path (an overloaded WAN circuit, a congested tunnel) means a couple of dropped hellos bounce the neighbor, over and over. The signature is periodic `holding time expired` teardowns with immediate re-establishment: ``` %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.12.1 (Ethernet0/1) is down: holding time expired ``` Defaults are hello 5 and hold 15 seconds on most interfaces (60/180 on low-speed multipoint). If a link drops hellos routinely, raising timers is a bandage; the real fix is the congestion. In named mode both live under the af-interface (`hello-interval`, `hold-time`), and `af-interface default` applies them fleet-wide. ## Flapping Adjacencies: When It Forms and Will Not Stay A neighbor that cycles is a different diagnosis tree from one that never appears. Work it in this order: - **Read the down reason for the flap, not the first event.** `holding time expired` points at packet loss or one-way traffic; `retry limit exceeded` means hellos survive but reliable unicast (updates, replies) dies, which is classic MTU mismatch or unidirectional link territory. - **Check the interface counters both sides:** `show interface` for input errors, CRCs, and output drops. On virtual and lab gear, also check CPU; a router at 100% control-plane CPU misses hello deadlines it technically received. - **Watch SRTT and Q Cnt in the neighbor table.** A healthy neighbor shows single-digit SRTT and Q Cnt 0\. SRTT in the hundreds of milliseconds with Q Cnt stuck above zero means the reliable transport is struggling, and the adjacency is next. - **Suspect the path, not the protocol.** EIGRP survives on remarkably bad links when timers and MTU agree. Chronic flapping is almost always layer 1 or layer 2 weather: duplex mismatches, a dirty optic, a saturated policer on the carrier side. Flapping neighbors also carry a second-order cost. Every drop sends the routes learned through that peer active and DUAL floods queries, and if a query lands on a peer that cannot reply in time you get the [stuck in active timeout that tears down an unrelated neighbor](https://www.pinglabz.com/eigrp-query-process-stuck-in-active/). If a ticket has both a flap and an SIA event, fix the flap first. ## Common Mistakes and Gotchas - **Assuming a quiet log means EIGRP is healthy.** Our AS 200 link was as dead as the K-value link and produced zero syslog on either side. Absence of a message is data, not reassurance. - **Forgetting that one log line may cover only one of several faults.** R1 had two broken links and one message, naming 10.0.12.2, so the Ethernet0/1 side was diagnosed purely by elimination. Count the peers that are unaccounted for. - **Skimming past the neighbor table header.** `EIGRP-IPv4 Neighbors for AS(100)` prints even when the table is empty, which makes it the cheapest AS comparison available. Named mode adds the virtual instance name, so the two ends may format the line differently. - **Touching K values at all.** The only defensible reason to change metric weights is a documented design decision applied identically to every router in the AS. - **Recreating a routing process to "reset" it.** `no router eigrp ` silently removes AS-scoped interface commands, including authentication, which turns a one-line fix into a second, subtler outage. - **Fixing timers to fix an adjacency.** Hello and hold do not need to match. Time spent aligning them is time not spent on the AS number. ## The Two-Minute Method - `show ip eigrp neighbors` on both sides. Note who is missing, and read the AS number in the header line while you are there. - `show logging | include DUAL`. A named reason puts you one section above away from the fix. Nothing at all points at AS, passive-interface or a filter. - Total silence? `ping` the neighbor's interface address, then `show ip eigrp interfaces` both sides (passive check), then `show ip protocols` both sides. - Auth suspected: `show ip eigrp interfaces detail | include Authentication` both sides. - Adjacency forms but flaps: watch Q Cnt and SRTT in the neighbor table, and go hunting below layer 3. Note that adjacency troubleshooting in OSPF is a different game with more states to inspect (the [OSPF guide](https://www.pinglabz.com/ospf/) covers its state machine); EIGRP is binary, a neighbor exists or it does not, which is why the logs and the process of elimination above carry most of the weight. ## Key Takeaways - Adjacency failures are a short list: AS, K values, authentication, passive-interface, and layer 3 plumbing. Learn the fingerprint of each. - A K-value mismatch is loud. IOS XE logs `%DUAL-5-NBRCHANGE ... is down: K-value mismatch` on every hello, because the hello is accepted and then rejected on parameters. - An AS mismatch is completely silent. The hello is dropped at the header, no neighbor object exists, nothing is logged on either router. Find it by comparing headers or `show ip protocols`, never by grepping the log. - `show ip protocols` is the single best compare-both-ends command: AS number and all five K values on one screen. - Removing a classic EIGRP process silently deletes AS-scoped interface auth commands. Audit interfaces after process surgery. - A neighbor that forms and then flaps with rising Q counts is a transport problem (MTU, unidirectional link), not a protocol problem. Once the adjacency is solid but routes still are not where they should be, the sequel is [troubleshooting EIGRP route advertisement and missing routes](https://www.pinglabz.com/troubleshooting-eigrp-missing-routes/). And if you want the wider context for why DUAL cares so much about metric agreement in the first place, start from [the full EIGRP protocol reference](https://www.pinglabz.com/eigrp/). ### EIGRP Route Filtering with Distribute Lists and Route Maps URL: https://www.pinglabz.com/eigrp-route-filtering-distribute-lists/ Last updated: 2026-07-11T16:05:19.000Z Advertising every route to every neighbor is the default, not a requirement. Sometimes a spoke should never learn the DMZ prefixes, a lab block should stay invisible to production, or a partner connection should see exactly three networks and nothing else. In EIGRP the tool for this is the distribute list, powered by a prefix list or a route map. This article configures both flavors on a live CML lab with before-and-after routing tables, and covers the operational habits that keep filtering from turning into a mystery outage. It assumes the baseline from the [EIGRP complete guide](https://www.pinglabz.com/eigrp/). ## Where EIGRP Filtering Happens A distribute list attaches to the EIGRP process in a direction: `out` filters what you advertise, `in` filters what you accept, and either can be scoped to a single interface. Filtering `out` at the route owner is the cleanest pattern, because the prefix never enters the rest of the network at all. Filtering `in` protects one router from what a neighbor sends. A useful mental model: a distribute list is policy about *knowledge*, where an ACL on an interface is policy about *packets*. A filtered route means nobody can route there; the traffic never even gets a chance to be dropped. One EIGRP-specific behavior worth knowing before you start: changing distribute-list policy causes the router to resync affected adjacencies (a graceful restart, visible in the logs) so the new policy propagates. It is non-disruptive to forwarding, but it explains the resync log lines you will see. ## The Lab: Four Routes, Two of Them Private R3 owns four /24s (10.1.0.0 through 10.1.3.0, configured as loopbacks). Treat 10.1.2.0/24 and 10.1.3.0/24 as lab-only networks that must not leak. The baseline on R1: ``` R1# show ip route eigrp | include 10.1. D 10.1.0.0/24 [90/3584000] via 10.0.13.2, 00:00:16, Ethernet0/2 D 10.1.1.0/24 [90/3584000] via 10.0.13.2, 00:00:16, Ethernet0/2 D 10.1.2.0/24 [90/3584000] via 10.0.13.2, 00:00:16, Ethernet0/2 D 10.1.3.0/24 [90/3584000] via 10.0.13.2, 00:00:16, Ethernet0/2 ``` ## Option 1: Distribute List with a Prefix List Prefix lists are the right default for route filtering: faster to evaluate than ACLs, explicit about masks, and readable six months later. Deny the two lab networks, permit everything else (a prefix list, like an ACL, ends in an implicit deny, so the catch-all permit is mandatory): ``` R3(config)# ip prefix-list BLOCK-LAB-NETS seq 5 deny 10.1.2.0/24 R3(config)# ip prefix-list BLOCK-LAB-NETS seq 10 deny 10.1.3.0/24 R3(config)# ip prefix-list BLOCK-LAB-NETS seq 15 permit 0.0.0.0/0 le 32 R3(config)# router eigrp 100 R3(config-router)# distribute-list prefix BLOCK-LAB-NETS out ``` `permit 0.0.0.0/0 le 32` means "any prefix of any length", which is the standard catch-all (bare `permit 0.0.0.0/0` would match only a literal default route). In named mode the same command goes under the topology base: ``` R2(config-router-af)# topology base R2(config-router-af-topology)# distribute-list prefix BLOCK-LAB-NETS out ``` The after picture on R1, with the two lab networks gone: ``` R1# show ip route eigrp | include 10.1. D 10.1.0.0/24 [90/3584000] via 10.0.13.2, 00:02:18, Ethernet0/2 D 10.1.1.0/24 [90/3584000] via 10.0.13.2, 00:02:18, Ethernet0/2 ``` And the verification command that should be in your muscle memory, because it is how you *find* filters you forgot about: ``` R3# show ip protocols | include filter Outgoing update filter list for all interfaces is (prefix-list) BLOCK-LAB-NETS Incoming update filter list for all interfaces is not set R3# show ip prefix-list BLOCK-LAB-NETS ip prefix-list BLOCK-LAB-NETS: 3 entries seq 5 deny 10.1.2.0/24 seq 10 deny 10.1.3.0/24 seq 15 permit 0.0.0.0/0 le 32 ``` ## Option 2: Distribute List with a Route Map A route map does everything the prefix list does, plus richer matching (tags, metrics, sources) and the ability to grow into set actions later. The pattern inverts the logic: the prefix list now *permits* the networks you want to act on, and the route map denies them: ``` R3(config)# ip prefix-list LAB-DMZ seq 5 permit 10.1.2.0/24 R3(config)# ip prefix-list LAB-DMZ seq 10 permit 10.1.3.0/24 R3(config)# route-map EIGRP-OUT deny 10 R3(config-route-map)# match ip address prefix-list LAB-DMZ R3(config-route-map)# route-map EIGRP-OUT permit 20 R3(config)# router eigrp 100 R3(config-router)# distribute-list route-map EIGRP-OUT out ``` The empty `permit 20` stanza is the route map's catch-all; without it, the implicit deny filters everything (a spectacular way to empty your neighbors' routing tables). Verify: ``` R3# show route-map EIGRP-OUT route-map EIGRP-OUT, deny, sequence 10 Match clauses: ip address prefix-lists: LAB-DMZ Set clauses: route-map EIGRP-OUT, permit, sequence 20 Match clauses: Set clauses: ``` The result on R1 is identical to the prefix-list version: 10.1.0.0/24 and 10.1.1.0/24 present, the other two filtered. **Reach for a prefix list when** The policy is purely "these prefixes yes, those no". Simplest to read, fastest to evaluate, hardest to get subtly wrong. **Reach for a route map when** You need multiple match criteria, route tags, or you expect the policy to grow. The deny/permit stanza structure scales; a bare prefix list does not. ## Scoping to One Neighbor Both forms accept an interface at the end, which turns a process-wide policy into a per-neighbor one: ``` R3(config-router)# distribute-list prefix BLOCK-LAB-NETS out Ethernet0/1 ``` Now only the neighbor on Ethernet0/1 gets the filtered view; the neighbor on Ethernet0/2 still sees everything. This is the mechanism behind "partner sees three routes, core sees all" designs. It also composes with summarization: filter what must not exist, summarize what remains (see the [summarization and default routes guide](https://www.pinglabz.com/eigrp-summarization-default-routes/), since a summary-address is itself a blunt form of filtering through suppression). ## The Other Knobs: Offset Lists and Distance Filtering is binary: a route exists for a neighbor or it does not. Two sibling tools cover the gray area in between, and knowing all three keeps you from misusing any of them. **Offset lists** add a value to the metric of matching routes, in or out, optionally per interface. That makes a path less attractive without hiding it, which is often what people actually wanted when they reached for a filter: ``` R3(config)# access-list 11 permit 10.1.1.0 0.0.0.255 R3(config)# router eigrp 100 R3(config-router)# offset-list 11 out 500000 Ethernet0/1 ``` Neighbors beyond Ethernet0/1 still learn 10.1.1.0/24, just with half a million extra metric units, so they prefer any other path and keep this one as a working backup. Offset lists are also the surgical way to fix a feasibility problem for one prefix when tuning interface delay would disturb everything else (the interplay is in the [variance guide](https://www.pinglabz.com/eigrp-variance-unequal-cost-load-balancing/)). **Administrative distance** can be weaponized as a filter of last resort: `distance 255 ` tells the router to disbelieve matching routes entirely (AD 255 is never installed). It is occasionally the only tool that works on *received* routes you cannot filter at the source, but it acts on this router alone, and the route still occupies topology table space. Prefer a real distribute list when you have the choice; the AD mechanics are covered in [EIGRP administrative distance](https://www.pinglabz.com/eigrp-administrative-distance/). ## What Happens on the Wire When You Apply a Filter Apply or change a distribute list and you will see this in the log: ``` *Jul 11 05:26:52.068: %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.35.2 (Ethernet0/2) is resync: peer graceful-restart ``` EIGRP has no mechanism to say "forget just these prefixes," so it performs a graceful resync: the router re-advertises its (now filtered) view to affected neighbors, who reconcile their tables. Forwarding continues throughout, adjacencies do not drop, and uptime counters in `show ip eigrp neighbors` keep counting. Knowing this matters in two directions: do not panic when a planned filter change logs NBRCHANGE resync lines, and conversely, unexplained resync lines in the log are evidence that *someone* changed routing policy, which is a lead worth following during an incident. ## Operational Habits That Save You Later - **Document the intent in the name.** BLOCK-LAB-NETS tells the next engineer what it does. PL-1 does not. - **Always end with an explicit catch-all.** Both the implicit deny at the end of a prefix list and the implicit deny at the end of a route map have erased entire routing tables. Write the permit; do not rely on memory. - **Check for filters first when routes are missing.** `show ip protocols` shows both directions in two lines. It is the fastest "is anyone filtering?" test on the platform, and it is step two of the workflow in [troubleshooting EIGRP missing routes](https://www.pinglabz.com/troubleshooting-eigrp-missing-routes/). - **Filter at the source when you can.** An `out` filter at the route owner keeps the prefix out of every topology table; an `in` filter on twelve routers is twelve chances to miss one. Worth knowing for context: this same distribute-list grammar is legacy in BGP, where route maps and prefix lists attach per neighbor instead (the [BGP complete guide](https://www.pinglabz.com/bgp/) covers that model), and OSPF can only filter at area boundaries or into the RIB because of its link-state flooding rules. EIGRP's distance vector nature is what makes per-neighbor knowledge control this clean. ## Key Takeaways - Distribute lists filter EIGRP routing knowledge, in or out, process-wide or per interface. Filtered routes never reach the neighbor's topology table. - Prefix lists are the readable default; route maps buy multi-criteria matching and room to grow. Both end in an implicit deny, so the explicit catch-all permit is not optional. - In named mode the command lives under topology base; classic mode puts it directly under the process. Same syntax otherwise. - Scope filters to an interface for per-neighbor policy, and prefer filtering out at the source over filtering in everywhere else. - `show ip protocols` exposes active filters in both directions and belongs at the top of every missing-route investigation. Route filtering is how you decide who knows what, which in a routed network is the same as deciding who can reach what. For the full picture of EIGRP policy, metrics, and design, head back to the [EIGRP complete guide](https://www.pinglabz.com/eigrp/). ### EIGRP Wide Metrics: 64-Bit Metrics in Named Mode URL: https://www.pinglabz.com/eigrp-wide-metrics/ Last updated: 2026-07-11T16:05:18.000Z If you have ever looked at `show ip eigrp topology` on a modern router and wondered why the metrics suddenly have nine or ten digits, you have met wide metrics. Named mode EIGRP replaced the classic 32-bit composite metric with a 64-bit version, and it changed more than the number of digits: delay is now measured in picoseconds, a scaling factor sits between EIGRP and the routing table, and mixed classic/named networks quietly do conversions on every boundary. This article decodes all of it with real output from the CML lab used across the [EIGRP complete guide](https://www.pinglabz.com/eigrp/) cluster. ## Why 32 Bits Stopped Working The classic metric formula (the full derivation is in the [EIGRP metric and K values guide](https://www.pinglabz.com/eigrp-metric-k-values/)) scales bandwidth as 10^7 divided by the slowest link in kbps. Work that out for modern interfaces and the problem is obvious: a 10 Gbps link gives 10^7 / 10000000 = 1, and anything faster is also 1\. A 10 Gbps path and a 100 Gbps path look identical to classic EIGRP. Delay has the same ceiling problem: it is counted in tens of microseconds, and high-speed links genuinely differ by nanoseconds. Wide metrics fix both dimensions: the composite value grows to 64 bits, bandwidth is scaled from a 10 Tbps reference (EIGRP calls the term throughput), and delay becomes latency, measured in picoseconds. Faster-than-10G links finally sort correctly. ## Spotting Wide Metrics on a Live Router `show ip protocols` tells you immediately which world you are in. R1 runs named mode: ``` R1# show ip protocols | section eigrp Routing Protocol is "eigrp 100" EIGRP-IPv4 VR(PINGLABZ) Address-Family Protocol for AS(100) Metric weight K1=1, K2=0, K3=1, K4=0, K5=0 K6=0 Metric rib-scale 128 Metric version 64bit Router-ID: 1.1.1.1 ``` Three tells: a sixth K value (K6, for future jitter/energy attributes, default 0), `Metric rib-scale 128`, and `Metric version 64bit`. Compare classic-mode R3 in the same network: ``` R3# show ip protocols | section eigrp Routing Protocol is "eigrp 100" EIGRP-IPv4 Protocol for AS(100) Metric weight K1=1, K2=0, K3=1, K4=0, K5=0 Router-ID: 10.1.3.1 ``` Five K values, no metric version line, no rib-scale. Same AS, same adjacencies, two metric dialects. ## Reading a Wide Metric Topology Entry Here is R2 (named mode) looking at R3's loopback, with the vector metrics that feed the calculation: ``` R2# show ip eigrp topology 3.3.3.3/32 EIGRP-IPv4 VR(PINGLABZ) Topology Entry for AS(100)/ID(2.2.2.2) for 3.3.3.3/32 State is Passive, Query origin flag is 1, 1 Successor(s), FD is 524288000, RIB is 4096000 Descriptor Blocks: 10.0.12.1 (Ethernet0/1), from 10.0.12.1, Send flag is 0x0 Composite metric is (524288000/458752000), route is Internal Vector metric: Minimum bandwidth is 10000 Kbit Total delay is 7000000000 picoseconds Reliability is 255/255 Load is 1/255 Minimum MTU is 1500 Hop count is 2 Originating router is 10.1.3.1 ``` Two things to notice. First, `Total delay is 7000000000 picoseconds`: that is 7 ms expressed in picoseconds, the wide-metrics latency unit. Second, the header line shows both numbers: `FD is 524288000` (the 64-bit EIGRP metric) and `RIB is 4096000` (what actually goes into the routing table). Divide them: 524288000 / 128 = 4096000\. That is rib-scale at work. ## Why rib-scale 128 Exists The routing table's metric field is still 32 bits. A 64-bit EIGRP metric could overflow it, so named mode divides by a scaling factor (default 128) before installing the route. The routing table entry confirms it: ``` R2# show ip route 3.3.3.3 Routing entry for 3.3.3.3/32 Known via "eigrp 100", distance 90, metric 4096000, type internal * 10.0.12.1, from 10.0.12.1, via Ethernet0/1 Route metric is 4096000, traffic share count is 1 ``` The division is lossy on purpose: two paths whose 64-bit metrics differ by less than the scale factor land on the same RIB metric. DUAL still compares the full 64-bit values internally, so successor selection and the feasibility condition (the heart of the [DUAL algorithm](https://www.pinglabz.com/eigrp-dual-algorithm/)) are unaffected. Change it with `metric rib-scale ` under the topology base if you ever need finer RIB granularity; almost nobody does. ## The Wide Metric Formula (What Changed, What Did Not) **Composite width**classic 32-bit becomes 64-bit **Bandwidth term**10^7 / min-bw becomes throughput: 65536 x (10^7 / min-bw-kbps) **Delay term**tens of microseconds becomes latency in picoseconds, scaled by 65536 **K values**K1 to K5 keep their meaning, K6 is added (default 0) **RIB install**64-bit metric divided by rib-scale (default 128) With default K values the shape is the familiar one: metric = throughput term + latency term, just at 65536x resolution. Sanity check against the capture: min bandwidth 10000 Kbit gives 65536 x (10^7 / 10000) = 65536000\. Total delay 7 ms is 7000 tens-of-microseconds units, and 7000 x 65536 = 458752000\. Sum: 524288000\. Exactly the FD the router printed. The math is the same math, with room to breathe. ## Mixed Classic and Named Networks Our lab deliberately mixes modes (R1/R2/R4 named, R3/R5 classic), and the adjacencies form fine: metric *style* is not one of the neighbor requirements (K values are, as covered in [EIGRP neighbor requirements](https://www.pinglabz.com/eigrp-neighbor-requirements/)). When a wide-metric router advertises to a classic peer, values are scaled down and precision is lost; classic-originated routes are scaled up. You saw the two dialects side by side above: R3's FDs are numbers like 307200, R2's are numbers like 524288000\. The practical rule: mixed mode works, but path selection on fast links is only trustworthy where wide metrics are end to end. That, plus SHA authentication support, is most of the case for migrating everything to named mode (walkthrough in the [EIGRP configuration guide](https://www.pinglabz.com/eigrp-configuration-cisco/)). Note that OSPF suffered the same high-speed blindness with its default reference bandwidth, solved with `auto-cost reference-bandwidth`; the [OSPF guide](https://www.pinglabz.com/ospf/) covers that side of the street. ## Migration Notes: What Actually Changes on Conversion Converting a classic router to named mode (`eigrp upgrade-cli` does it in place, preserving the AS number) switches the router to wide metrics as a side effect, and a few practical details follow: - **Interface commands do not change.** `delay` is still configured in tens of microseconds and `bandwidth` in kbps; wide metrics change how EIGRP *consumes* those values (converting delay to picoseconds internally), not how you set them. Our lab's `delay 200` on R2's alternate path became 2 ms, visible as 2000000000 picoseconds inside the vector metrics. - **Absolute metric values change everywhere.** Any monitoring that alerts on specific metric numbers, any documentation quoting FDs, and any offset-list built around 32-bit values needs revisiting. Relative path preference is preserved on links 10 Gbps and below; what changes is the numbers themselves. - **Path selection can legitimately differ above 10 Gbps.** That is the point of the feature: links that used to tie at bandwidth term 1 now sort correctly. If a migration changes a forwarding path, check whether the old path was only winning by rounding error. - **Tune delay, not bandwidth, before and after.** The bandwidth value also feeds QoS percentage calculations and interface utilization graphs; delay exists only for routing metrics, which is why it has always been the recommended traffic-engineering knob. ## What About K6, Jitter, and Energy? The wide-metrics TLVs reserve room for two extended attributes, jitter and energy, with K6 as their weight. The idea was forward-looking (prefer low-jitter paths for voice without touching delay), but no shipping IOS XE measures or populates these attributes, so K6 stays 0 and the extended terms contribute nothing. Treat K6 as protocol headroom rather than a feature: it must still *match* between neighbors, like every other K value (a mismatch is the classic adjacency killer covered in the [neighbor troubleshooting guide](https://www.pinglabz.com/troubleshooting-eigrp-neighbor-adjacencies/)). If you want traffic steered by live path quality on Cisco gear today, that job moved to performance routing and [SD-WAN](https://www.pinglabz.com/sd-wan/), not to K6. ## Key Takeaways - Classic 32-bit metrics cannot tell a 10G link from a 100G link. Wide metrics (named mode) extend the composite to 64 bits and measure delay in picoseconds. - Identify wide metrics instantly: `Metric version 64bit`, `rib-scale 128`, and a K6 value in `show ip protocols`. - The topology table shows the 64-bit FD; the routing table shows FD divided by rib-scale. Both numbers are printed on the same line, so you can verify the division yourself. - DUAL compares full 64-bit values, so feasibility and successor math keep full precision regardless of rib-scale. - Classic and named routers interoperate with automatic scaling, but buy accurate high-speed path selection by going named mode everywhere. Wide metrics are what let a 1989-vintage metric formula stay honest in a 100G world. For how the metric feeds DUAL, variance, and the rest of the machine, the [EIGRP complete guide](https://www.pinglabz.com/eigrp/) has the full map. ### EIGRP Query Process and Stuck-in-Active (SIA) Explained with Real Captures URL: https://www.pinglabz.com/eigrp-query-process-stuck-in-active/ Last updated: 2026-08-01T18:36:23.000Z Stuck-in-active is the most infamous failure mode EIGRP has. One lost route in a far corner of the network triggers a chain of queries, and if any router in that chain fails to answer for three minutes, adjacencies start getting torn down in places that had nothing to do with the original failure. Understanding the query process is how you design networks where that cannot happen. This builds on the DUAL concepts in the [EIGRP complete guide](https://www.pinglabz.com/eigrp/), with real query and reply packets from a CML lab. What almost nobody shows you is a route actually sitting in Active with a reply outstanding, because it is awkward to reproduce: break the link and EIGRP tears the adjacency down instead, and a dead neighbor counts as an answer. That capture is below, on IOS XE 17.18.2 in CML, with the reason modern IOS XE usually resets the neighbor before it prints `%DUAL-3-SIA`. ## What Happens When a Route Dies When EIGRP loses a route, the outcome depends entirely on the topology table. If a feasible successor exists, DUAL promotes it in milliseconds and the route never leaves Passive state. Nobody else in the network is consulted. This is EIGRP at its best. With no feasible successor there is no pre-validated loop-free alternative, so the router has to ask. The route goes **Active** and a query for that exact prefix goes to every neighbor. Each neighbor does one of three things: - **Replies immediately** if it has no knowledge of the prefix, or has its own valid path that does not depend on the querying router. - **Propagates the query** if the lost route was its successor too, then waits for those answers before replying. - **Replies with unreachable** if it only knew the route through the asker. The querying router cannot finish the computation until *every single neighbor* has replied, so one slow or wedged router anywhere in the query domain holds the prefix hostage. That is the structural weakness the SIA machinery exists to manage. Put plainly, stuck-in-active is what happens when there is no feasible successor and somebody does not answer, so if you are solid on [how DUAL decides a backup path is loop-free](https://www.pinglabz.com/eigrp-dual-algorithm/), you already know why some failures generate zero queries and others a flood. ## The Lab This Was Captured On The Active and SIA captures come from a three-router chain in CML: R1 - R2 - R3, all `iol-xe` nodes on IOS XE 17.18.2, EIGRP AS 100, no auto-summary. R1 originates 192.168.100.0/24 on Loopback1\. R2 sits in the middle, Et0/0 toward R1 on 10.0.12.0/30 and Et0/1 toward R3 on 10.0.23.0/30, with the active timer set to one minute (`timers active-time 1`, the minimum) so the lifecycle fits in one console session. Healthy state first: ``` R3# show ip route eigrp D 192.168.100.0/24 [90/435200] via 10.0.23.2, 00:00:35, Ethernet0/0 R2# show ip eigrp topology 192.168.100.0/24 EIGRP-IPv4 Topology Entry for AS(100)/ID(2.2.2.2) for 192.168.100.0/24 State is Passive, Query origin flag is 1, 1 Successor(s), FD is 409600 Descriptor Blocks: 10.0.12.1 (Ethernet0/0), from 10.0.12.1, Send flag is 0x0 Composite metric is (409600/128256), route is Internal ``` Note the shape of that entry: one successor, no second descriptor block, so no feasible successor. That is the precondition for a query. ## Watching a Query Live In an earlier lab, R3 loses a connected network (we shut its Loopback10, which carries 10.1.0.0/24). With `debug eigrp packets query reply` on R1, the whole conversation is visible: ``` R1# debug eigrp packets query reply (QUERY, REPLY) EIGRP Packet debugging is on *Jul 11 05:24:26.613: EIGRP: Received QUERY on Et0/2 - paklen 44 nbr 10.0.13.2 *Jul 11 05:24:26.623: EIGRP: Enqueueing REPLY on Et0/2 - paklen 0 nbr 10.0.13.2 tid 0 iidbQ un/rely 0/1 peerQ un/rely 0/0 serno 72-72 *Jul 11 05:24:26.626: EIGRP: Sending REPLY on Et0/2 - paklen 44 nbr 10.0.13.2 tid 0 ``` R1 received the query and answered in 13 milliseconds. Why so fast? At the time of this capture R3 was summarizing 10.1.0.0/22 toward R1, so R1 never knew the /24 existed, had nothing to recompute, and replied instantly. We come back to that, because it is the single most useful fact in this article. The active state itself is too fast to catch in a healthy lab: ``` R1# show ip eigrp topology active EIGRP-IPv4 VR(PINGLABZ) Topology Table for AS(100)/ID(1.1.1.1) ``` Empty. Convergence already finished; in a healthy network Active states live for milliseconds. `show ip eigrp traffic` keeps score over time, including the counters you hope stay at zero: ``` R1# show ip eigrp traffic EIGRP-IPv4 VR(PINGLABZ) Address-Family Traffic Statistics for AS(100) Hellos sent/received: 309/196 Updates sent/received: 44/48 Queries sent/received: 8/8 Replies sent/received: 11/8 SIA-Queries sent/received: 0/0 SIA-Replies sent/received: 0/0 ``` ## What Stuck-in-Active Actually Means Every router that goes Active for a prefix starts the active timer, 3 minutes by default. You can read it straight out of `show ip protocols`: ``` R1# show ip protocols | include Active Timer Active Timer: 3 min ``` If a neighbor still has not replied when the timer expires, the router declares the route stuck-in-active, logs `%DUAL-3-SIA`, and resets that adjacency. That reset is the brutal part: every route through that neighbor is flushed and relearned, which cascades into more queries, which can SIA somewhere else. One flapping prefix at a branch can ripple across a continent-wide EIGRP domain. Classic causes of an unanswered query: an overloaded control plane, a congested or lossy WAN link dropping the reply, a unidirectional link, or a query domain so large the reply chain is dozens of routers deep. ## Catching a Route in Active State To hold a route in Active long enough to photograph it you must starve the reply while keeping the adjacency alive, and that distinction is the whole trick. In the lab, an ACL on R2's Et0/1 permits R3's multicast EIGRP (224.0.0.10, so hellos and queries still flow) but denies R3's unicast EIGRP, which is what replies, acks and SIA-replies ride on. R2's uplink is then shut, so R2 must query R3\. R3 answers; R2 never hears it. ``` R2# show ip eigrp topology active (active 00:00:09) Codes: P - Passive, A - Active, U - Update, Q - Query, R - Reply, r - reply Status, s - sia Status A 192.168.100.0/24, 0 successors, FD is 409600, Q 1 replies, active 00:00:09, query-origin: Local origin via 10.0.12.1 (Infinity/Infinity), Ethernet0/0 Remaining replies: via 10.0.23.1, r, Ethernet0/1 <== 'r' = still waiting on R3's REPLY ``` State `A` with **0 successors** is a router that cannot compute a path on its own. `query-origin: Local origin` says R2 started this query rather than relaying somebody else's, the old successor reports `(Infinity/Infinity)`, and under **Remaining replies** exactly one neighbor is listed. That `r` flag is the most important character in EIGRP troubleshooting: it names the router that owes you an answer. ## SIA-Query and SIA-Reply: The Modern Safety Valve Because tearing down an adjacency over one slow prefix is so disproportionate, modern IOS splits the active timer in half. At the halfway mark the waiting router sends an **SIA-Query**: "are you alive and still working on this?". A healthy-but-busy neighbor answers with an **SIA-Reply**, keeping the adjacency alive while the real reply is still in flight, up to three times. Only a neighbor that fails to answer even that gets reset. With the active timer at one minute, it engages about thirty seconds in: ``` R2# show ip eigrp topology active (active 00:00:39) A 192.168.100.0/24, 0 successors, FD is 409600, Qq 1 replies, active 00:00:39, query-origin: Local origin, retries(1) via 10.0.12.1 (Infinity/Infinity), Ethernet0/0 via 10.0.23.1 (Infinity/Infinity), rs, q, Ethernet0/1, serno 11 <== 'rs' = reply + SIA status ``` Three things changed and all three are diagnostic. The neighbor flags went from `r` to `rs`, so an SIA-Query is now outstanding on top of the original query. The route flags went from `Q` to `Qq`. And `retries(1)` appeared. Seeing `rs` in production means an adjacency roughly one timer away from reset, with about half the active timer left to find out why that neighbor is silent. The timer is tunable in either syntax: ``` R3(config)# router eigrp 100 R3(config-router)# timers active-time 1 R2(config)# router eigrp PINGLABZ R2(config-router)# address-family ipv4 unicast autonomous-system 100 R2(config-router-af)# topology base R2(config-router-af-topology)# timers active-time 1 ``` ``` R2# show ip protocols | include Active Timer Active Timer: 1 min ``` Shortening the timer makes SIA detection faster but punishes slow WAN paths; `timers active-time disabled` waits forever, trading a visible failure for an invisible one. One minute is the floor. Most networks should leave it at 3 minutes and fix the design instead. ## Why You May Never See %DUAL-3-SIA Here is the finding that contradicts most of what is written on this topic. In the lab above, with the reply permanently blocked and the active timer at its minimum, R2 never logged `%DUAL-3-SIA` at all. It logged this: ``` *Jul 20 06:50:13: %DUAL-5-NBRCHANGE: Neighbor 10.0.12.1 (Ethernet0/0) is down: interface down *Jul 20 06:50:59: %DUAL-5-NBRCHANGE: Neighbor 10.0.23.1 (Ethernet0/1) is down: retry limit exceeded *Jul 20 06:51:00: %DUAL-5-NBRCHANGE: Neighbor 10.0.23.1 (Ethernet0/1) is up: new adjacency ``` Two independent clocks are racing. The active timer runs 60 seconds here before SIA can be declared. Separately, EIGRP's reliable transport gives up after 16 unacknowledged retransmissions, which on this link landed at roughly 46 seconds. The retry limit won, the adjacency was reset, that reset resolved the Active route as an implicit reply, and the classic `%DUAL-3-SIA-1 ... Cleaning up` message never got a chance to print. So **do not treat the absence of `%DUAL-3-SIA` as proof you have no query-scope problem.** On current IOS XE an unanswered query often surfaces as a `retry limit exceeded` reset instead: same root cause, same damage, different log string. Alert on both. Getting the literal SIA-3 message needs the active timer to win, which takes a longer reply chain, an RTO inflated by real WAN latency, or a longer active-time. ## Designing Query Boundaries (the Real Fix) You do not solve SIA by tuning timers. You solve it by making the query domain small, and two tools do almost all the work: **Summarization** A router that only knows 10.1.0.0/22 cannot go Active for 10.1.2.0/24\. It replies immediately and the query dies there, exactly as in the debug capture above. Full walkthrough in the [EIGRP summarization guide](https://www.pinglabz.com/eigrp-summarization-default-routes/). **Stub routing** A stub-flagged spoke announces "do not query me" in its hellos, so hubs never query it at all. In hub-and-spoke WANs that removes hundreds of routers from every query domain. Details in [EIGRP stub routing](https://www.pinglabz.com/eigrp-stub-routing/). Both belong on the spokes and at the summarization boundary, not on the router that logged the error: ``` router eigrp 100 eigrp stub connected summary interface Ethernet0/1 ip summary-address eigrp 100 10.1.0.0 255.255.252.0 ``` If your sites ride an MPLS L3VPN and are dual-homed to two PEs, the same conversation extends to [tagging routes with the site they came from](https://www.pinglabz.com/eigrp-site-of-origin-soo/) so a PE refuses to hand a site its own routes back. It is also fair framing for protocol selection: OSPF floods LSAs but never holds a route hostage waiting on neighbors, so its failure domain is bounded by area design (see [EIGRP vs OSPF](https://www.pinglabz.com/eigrp-vs-ospf/) and the [OSPF guide](https://www.pinglabz.com/ospf/)). EIGRP scales beautifully, but only when someone draws the query boundaries deliberately. ## Anatomy of a Query Storm Scale it up. A retail network: one hub, 400 spokes, no summarization, no stubs. A branch loses a connected LAN prefix when a switch reboots, so there is no feasible successor and it queries the hub. The hub's successor for that prefix was the branch itself, so it cannot answer from its own table and propagates the query to its other 399 neighbors. Four hundred replies must come back before the hub can answer the branch. Now add one spoke with a saturated circuit whose reply sits behind bulk traffic. Halfway through the active timer the hub sends it an SIA-Query and its topology output starts showing the `rs` flags captured earlier. If that exchange cannot complete either, the hub resets the adjacency, flushing every route through it and generating queries for *those* prefixes, each with its own clock. That is how one switch reboot becomes a multi-site event. Summarize each branch to a single block at the hub, mark every spoke as stub, and the same event touches two routers. It is also the argument for keeping feasible successors plentiful, since a lost route with one generates no query at all, which makes metric design (consistent delay values, covered in the [K values guide](https://www.pinglabz.com/eigrp-metric-k-values/)) quietly an SIA-prevention tool. ## Monitoring and Baselines SIA prevention is measurable before the first incident. Three habits: - **Baseline query counters.** `show ip eigrp traffic` on your hubs, weekly. Queries per week should be a small, boring number; a rising trend means a flapping prefix or an eroding design margin. SIA-Query counters above zero deserve a ticket even if nothing broke. - **Read the event log after any incident.** `show ip eigrp events` keeps a rolling in-memory history of DUAL decisions with millisecond timestamps, and answers "what happened at 03:12" long after the debugs you did not have running would have. - **Alert on the strings that matter:** `%DUAL-3-SIA`, `retry limit exceeded`, and any NBRCHANGE burst on a hub. An SIA event that resolves itself is still a warning shot. ## Troubleshooting a Live SIA Event When `%DUAL-3-SIA` or `%DUAL-5-NBRCHANGE ... retry limit exceeded` shows up, work the chain: - `show ip eigrp topology active` on the router that logged it. Neighbors still owing a reply are flagged `r`, and `rs` once an SIA-Query has gone out to them. - Hop to that neighbor and run the same command. Follow the `r` flags until you find the router that is not answering; the problem lives there or on the link to it. - On the culprit, check CPU, memory and interface queues. A router too busy to answer queries is usually too busy for other things you can measure. - Check the path both ways. From the querying side, a reply that was sent but never arrived looks identical to one never sent, which is exactly what this article's lab demonstrates. An ACL, a unidirectional fault or a policer catching unicast control traffic all produce it. - Look at the prefix, because a route flapping every few minutes multiplies query load. If routes vanish for non-obvious reasons, cross-check with [troubleshooting EIGRP missing routes](https://www.pinglabz.com/troubleshooting-eigrp-missing-routes/). ## Gotchas from the Lab **Packet loss alone will not give you an SIA.** This is the trap that eats an afternoon if you try to reproduce the failure. The first attempt used link conditioning to add heavy latency and loss to the R2-R3 link: ``` R2# ping 10.0.23.1 repeat 5 timeout 25 Sending 5, 100-byte ICMP Echos to 10.0.23.1, timeout is 25 seconds: .!..! Success rate is 40 percent (2/5), round-trip min/avg/max = 20003/20003/20004 ms ``` Twenty seconds of round-trip delay, 40 percent success, and still no stuck route. Instead: ``` R2# show logging | include DUAL|NBR *Jul 20 05:53:19.775: %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.23.1 (Ethernet0/1) is down: Interface PEER-TERMINATION received *Jul 20 05:53:22.514: %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.23.1 (Ethernet0/1) is up: new adjacency *Jul 20 05:58:05.694: %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.23.1 (Ethernet0/1) is down: holding time expired ``` Loss heavy enough to starve a reply also starves hellos, so the hold timer expires and the adjacency drops. A neighbor going down is treated as an **implicit reply** for every route it owed an answer on, which resolves the Active route immediately. Bad links therefore produce adjacency churn rather than SIA, and if that is what your logs look like, read [why EIGRP neighbors keep dropping](https://www.pinglabz.com/troubleshooting-eigrp-neighbor-adjacencies/) instead. The dangerous case is the opposite: a link healthy enough to carry hellos but unable to deliver the reply. **The flags outlive the log message.** Given the retry-limit race above, `show ip eigrp topology active` and its `r` and `s` letters are the evidence worth capturing during an incident. Syslog tells you something reset; the flags tell you who was silent. On a healthy router the same command prints nothing but a header, so any persistent output is worth escalating. ## Key Takeaways - Queries only happen when a route with no feasible successor dies, so networks rich in feasible successors barely query at all. - A route is stuck-in-active when any neighbor fails to reply within the active timer (3 minutes by default), and the penalty is an adjacency reset that can cascade. - In `show ip eigrp topology active`, `r` means that neighbor still owes a reply and `rs` means an SIA-Query is outstanding to it. Those letters are the diagnostic, captured live above on IOS XE 17.18.2. - On modern IOS XE an unanswered query often resolves as `%DUAL-5-NBRCHANGE ... retry limit exceeded` before `%DUAL-3-SIA` can print. Alert on both strings. - Packet loss usually produces neighbor flaps, not SIA, because a dead neighbor counts as an implicit reply. - Summarization and stub routing are the real defenses, because they stop queries propagating at all. Timers are a tourniquet, not a cure. The query process is EIGRP's personality: brilliant when the topology gives it options, fragile when a flat design lets one question travel too far. Draw the boundaries and it stays brilliant. The rest of the protocol's machinery, from neighbor formation to metrics, is mapped out in [our full EIGRP reference](https://www.pinglabz.com/eigrp/). ### EIGRP Authentication: MD5 and SHA in Classic and Named Mode URL: https://www.pinglabz.com/eigrp-authentication-md5-sha/ Last updated: 2026-07-11T16:05:17.000Z An unauthenticated EIGRP interface will happily form an adjacency with anything that speaks the protocol. That means anyone who can plug into that segment can inject routes, blackhole traffic, or quietly become a man in the middle. Authentication closes that door, and EIGRP gives you two flavors: MD5 with key chains (classic and named mode) and HMAC-SHA-256 (named mode only). This article configures both on a live CML lab, then deliberately breaks a key to show you exactly what failure looks like. It builds on the adjacency fundamentals in the [EIGRP complete guide](https://www.pinglabz.com/eigrp/). ## Why Authenticate an IGP Routing protocol authentication is not about encrypting your routes; updates still cross the wire in cleartext. It is about *trust*: each packet carries a hash computed over the packet plus a shared secret, and a receiver that cannot reproduce the hash discards the packet. No valid hash, no adjacency, no route injection. Authentication is one of the five things that must match for an adjacency (the full list is in [EIGRP neighbor requirements](https://www.pinglabz.com/eigrp-neighbor-requirements/)), and it is standard hardening in any environment where access switches or WAN circuits are not fully trusted. OSPF and BGP have their own equivalents; the same logic applies across your whole routing edge, including [BGP](https://www.pinglabz.com/bgp/) peers. ## MD5 vs HMAC-SHA-256 **MD5** Works in: classic and named mode Keys: key chain (rotation supported) Cryptographically dated, but still the interoperability baseline on older IOS. **HMAC-SHA-256** Works in: named mode only Keys: inline password or key chain The modern choice. If both ends support named mode, use this. The named-mode-only restriction on SHA-256 is a genuinely common exam trap and a real-world migration argument: if you want stronger hashes, you migrate to named mode (the conversion path is in the [EIGRP configuration guide](https://www.pinglabz.com/eigrp-configuration-cisco/)). ## The Lab Same five-router CML topology as the rest of this cluster: R1/R2/R4 in named mode (`router eigrp PINGLABZ`, AS 100), R3/R5 classic (`router eigrp 100`). We will secure the R1-R2 link with SHA-256 and the R3-R5 link with MD5, then sabotage the MD5 key on R5. ## HMAC-SHA-256 in Named Mode One command under the af-interface, no key chain required in its simplest form: ``` R1(config)# router eigrp PINGLABZ R1(config-router)# address-family ipv4 unicast autonomous-system 100 R1(config-router-af)# af-interface Ethernet0/1 R1(config-router-af-interface)# authentication mode hmac-sha-256 PingLabz!Secure123 ``` Apply the mirror config on R2\. The moment the first side goes live, the adjacency drops, because one side now demands authentication the other cannot provide. IOS XE logs it explicitly, and the recovery is automatic once both sides match: ``` R1# show logging | include DUAL *Jul 11 05:21:33.266: %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.12.2 (Ethernet0/1) is down: authentication HMAC-SHA-256 configured *Jul 11 05:21:46.334: %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.12.2 (Ethernet0/1) is up: new adjacency ``` That 13-second gap is your maintenance window reality: configure the second side promptly. Verification: ``` R1# show ip eigrp interfaces detail Ethernet0/1 | include Authentication Authentication mode is HMAC-SHA-256, key-chain is not set ``` ## MD5 with Key Chains in Classic Mode Classic mode MD5 is a two-part job: build a key chain, then bind it to the interface with two commands (mode and key chain, and forgetting one of them is a classic ticket generator): ``` R3(config)# key chain EIGRP-KEYS R3(config-keychain)# key 1 R3(config-keychain-key)# key-string PingLabzMD5key R3(config)# interface Ethernet0/2 R3(config-if)# ip authentication mode eigrp 100 md5 R3(config-if)# ip authentication key-chain eigrp 100 EIGRP-KEYS ``` Modern IOS XE will lecture you about the plaintext key while you type it, which is fair: ``` SECURITY WARNING - Module: KEY_CHAIN, Command: key-string *, Reason: Configuration employs an Insecure method for password storage, Description: key-string in key-chain configured with weak encryption (type 0, 7, or plaintext) instead of secure type 6 encryption ``` Repeat on R5, then verify both the interface binding and the chain itself: ``` R3# show ip eigrp interfaces detail Ethernet0/2 | include Authentication Authentication mode is md5, key-chain is "EIGRP-KEYS" R3# show key chain Key-chain EIGRP-KEYS: key 1 -- text "PingLabzMD5key" accept lifetime (always valid) - (always valid) [valid now] send lifetime (always valid) - (always valid) [valid now] ``` Key chains support send and accept lifetimes, which is how you rotate keys without an outage: add key 2 on both sides with an accept lifetime that overlaps key 1, cut the send lifetime over, then retire key 1\. The key *number* matters as much as the string, because both are inputs to the hash. Same string under a different key ID fails. ## Breaking It on Purpose: The Mismatch Capture Now the part most articles skip. We change R5's key-string to the wrong value and watch from R3 with `debug eigrp packets` running: ``` R5(config)# key chain EIGRP-KEYS R5(config-keychain)# key 1 R5(config-keychain-key)# key-string WrongKey999 ``` ``` R3# show logging | include authentication|DUAL *Jul 11 05:23:16.095: EIGRP: received packet with MD5 authentication, key id = 1 *Jul 11 05:23:19.445: EIGRP: pkt key id = 1, authentication mismatch *Jul 11 05:23:19.445: EIGRP: Dropping peer, invalid authentication *Jul 11 05:23:19.447: %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.35.1 (Ethernet0/2) is down: Auth failure *Jul 11 05:23:23.961: EIGRP: pkt key id = 1, authentication mismatch ``` Three distinct signatures worth memorizing: the debug line `authentication mismatch`, the `Dropping peer, invalid authentication` action, and the syslog reason `down: Auth failure`. The neighbor table confirms the peer is simply gone: ``` R3# show ip eigrp neighbors EIGRP-IPv4 Neighbors for AS(100) H Address Interface Hold Uptime SRTT RTO Q Seq (sec) (ms) Cnt Num 0 10.0.13.1 Et0/1 11 00:07:03 12 100 0 61 ``` Restore the correct key-string on R5 and the adjacency returns within a hold time. No clear commands needed. ## The Gotcha That Silently Eats Your Auth Config One more failure mode we hit live in this lab: classic-mode interface authentication commands are scoped to the AS number. When we later removed `router eigrp 100` from R3 during another test, IOS deleted `ip authentication mode eigrp 100 md5` and the key-chain binding from the interface along with the process. Recreating the process did not bring them back, and since R5 still required MD5, the adjacency silently refused to form: no error, no log, just hellos being dropped on one side. If an adjacency will not return after process surgery, `show run interface` is your first stop. More silent-failure patterns like this are cataloged in [troubleshooting EIGRP neighbor adjacencies](https://www.pinglabz.com/troubleshooting-eigrp-neighbor-adjacencies/). ## Named Mode MD5 and Key Chains Named mode can also do MD5, and either mode can reference a key chain, which is what you want for rotation even with SHA: ``` R1(config-router-af-interface)# authentication mode md5 R1(config-router-af-interface)# authentication key-chain EIGRP-KEYS ``` Set the mode on the af-interface (or under `af-interface default` to cover every interface at once, which is the sane fleet-wide pattern), and remember: mode and key chain are separate lines in named mode too. ## Key Rotation Without an Outage Shared secrets age. People leave, configs get pasted into tickets, and a key that has been in production for five years should be assumed to be public. Key chains exist so rotation does not require a maintenance window. The mechanism: EIGRP signs outgoing packets with the lowest-numbered valid *send* key, but accepts incoming packets signed by *any* valid accept key. Overlap the accept lifetimes and you can never be locked out mid-rotation. ``` R3(config)# key chain EIGRP-KEYS R3(config-keychain)# key 1 R3(config-keychain-key)# key-string PingLabzMD5key R3(config-keychain-key)# accept-lifetime 00:00:00 Jan 1 2026 01:00:00 Aug 1 2026 R3(config-keychain-key)# send-lifetime 00:00:00 Jan 1 2026 00:00:00 Jul 15 2026 R3(config-keychain)# key 2 R3(config-keychain-key)# key-string PingLabz2026H2 R3(config-keychain-key)# accept-lifetime 00:00:00 Jul 14 2026 infinite R3(config-keychain-key)# send-lifetime 00:00:00 Jul 15 2026 infinite ``` Deploy this on every router in the domain before July 14 and the cutover happens by itself: on July 15 everyone starts signing with key 2, while the overlapping accept windows absorb any clock skew. That last point is the operational catch: lifetimes are only as good as the clocks. Run NTP everywhere you run key lifetimes, and audit with `show key chain`, which prints exactly which keys are valid right now. A key chain whose only key has expired is a self-inflicted outage with the same `Auth failure` signature as a typo. ## Fleet-Wide Auth with af-interface default Configuring authentication one interface at a time invites the classic gap: eleven secured links and the twelfth forgotten. Named mode can set the policy once for every interface in the address family: ``` R1(config-router-af)# af-interface default R1(config-router-af-interface)# authentication mode hmac-sha-256 PingLabz!Secure123 ``` Every current and *future* EIGRP interface on the router now requires authentication, which converts "did we remember?" into "did we deliberately exempt?". For an interface that genuinely should differ (say, a lab port), a specific af-interface block overrides the default. This inheritance model is one of named mode's quiet advantages and works for passive-interface and hello timers the same way. A short deployment checklist that keeps auth rollouts boring: - Stage key chains everywhere first; binding them to interfaces is the only disruptive step. - Roll link by link, watching `show ip eigrp neighbors` return to a full table after each pair. - Standardize one key chain name domain-wide. Half the mismatch tickets in mixed environments are two spellings of the same chain. - Confirm NTP before adding lifetimes, and stagger send/accept windows generously (hours, not minutes). ## Key Takeaways - EIGRP authentication proves packet origin with a shared-secret hash. It stops rogue adjacencies and route injection; it does not encrypt updates. - MD5 works everywhere and uses key chains. HMAC-SHA-256 is stronger, simpler to configure, and named mode only. - Key ID and key string must both match. Key chain lifetimes are the mechanism for hitless rotation. - Know the failure signatures: `down: Auth failure` in syslog, `authentication mismatch` in debug, and a neighbor that simply vanishes from the table. - Removing a classic EIGRP process deletes its AS-scoped interface auth commands. After process changes, audit interfaces before you troubleshoot anything else. Authentication is cheap insurance on every segment where an unexpected device could appear. For the rest of the protocol, from DUAL to design, return to the [EIGRP complete guide](https://www.pinglabz.com/eigrp/). ### EIGRP Unequal-Cost Load Balancing: Variance Explained URL: https://www.pinglabz.com/eigrp-variance-unequal-cost-load-balancing/ Last updated: 2026-08-01T19:31:08.000Z Every routing protocol can load balance across equal-cost paths. EIGRP is the only IGP on Cisco routers that can load balance across *unequal*\-cost paths, and the knob that unlocks it is a single number called variance. It is a favorite CCNP ENCOR exam topic precisely because it forces you to actually understand feasible successors, feasibility, and the metric math from the [EIGRP complete guide](https://www.pinglabz.com/eigrp/). This article walks through variance on a live CML lab: the topology table before, the feasibility check, the `variance 2` command in named mode, and the traffic share counters that prove the router is really splitting load in proportion to metric. All output is real IOS XE. ## The Rule Variance Cannot Break Variance does not let EIGRP use just any backup path. It only installs routes that already qualify as feasible successors. The feasibility condition, straight from [DUAL](https://www.pinglabz.com/eigrp-dual-algorithm/): a neighbor's reported distance (RD) must be strictly less than your current feasible distance (FD). That guarantee is what makes the whole feature loop-free. If a path fails the feasibility condition, no variance value will ever install it, no matter how large. **Successor** The best loop-free path. Its metric is the feasible distance (FD). Installed in the routing table by default. **Feasible successor** A backup whose reported distance is less than the FD. Pre-validated as loop-free, ready for instant failover, and the only candidate variance can promote. **Variance** A multiplier (1 to 128). Any feasible successor whose total metric is less than FD x variance gets installed and carries traffic. Default is 1: equal cost only. ## The Lab: Two Unequal Paths to the Same Loopback R2 can reach R3's loopback (3.3.3.3/32) two ways: the short path through hub R1, or the longer path through R4 and R5\. The R2-R4 link carries extra configured delay, so the paths are deliberately unequal. Before touching variance, the topology table on R2: ``` R2# show ip eigrp topology 3.3.3.3/32 EIGRP-IPv4 VR(PINGLABZ) Topology Entry for AS(100)/ID(2.2.2.2) for 3.3.3.3/32 State is Passive, Query origin flag is 1, 1 Successor(s), FD is 524288000, RIB is 4096000 Descriptor Blocks: 10.0.12.1 (Ethernet0/1), from 10.0.12.1, Send flag is 0x0 Composite metric is (524288000/458752000), route is Internal Vector metric: Minimum bandwidth is 10000 Kbit Total delay is 7000000000 picoseconds Hop count is 2 10.0.24.2 (Ethernet0/2), from 10.0.24.2, Send flag is 0x0 Composite metric is (589824000/458752000), route is Internal Vector metric: Minimum bandwidth is 10000 Kbit Total delay is 8000000000 picoseconds Hop count is 3 ``` Read the two composite metrics as (total metric / reported distance): - **Via R1:** 524288000 / 458752000\. Lowest total metric, so this is the successor and FD = 524288000. - **Via R4:** 589824000 / 458752000\. The RD (458752000) is less than the FD (524288000), so the feasibility condition holds and this is a feasible successor. Those big numbers are 64-bit wide metrics with delay in picoseconds, because R2 runs named mode. If they look unfamiliar, the [EIGRP wide metrics article](https://www.pinglabz.com/eigrp-wide-metrics/) decodes them. Only one path is in the routing table right now: ``` R2# show ip route 3.3.3.3 Routing entry for 3.3.3.3/32 Known via "eigrp 100", distance 90, metric 4096000, type internal Routing Descriptor Blocks: * 10.0.12.1, from 10.0.12.1, 00:00:48 ago, via Ethernet0/1 Route metric is 4096000, traffic share count is 1 Total delay is 7000 microseconds, minimum bandwidth is 10000 Kbit Loading 1/255, Hops 2 ``` ## Do the Math Before You Configure The variance value you need is the ratio between the worst path you want to use and the FD, rounded up: ``` 589824000 / 524288000 = 1.125 -> variance 2 ``` Any feasible successor whose metric is under FD x 2 = 1048576000 will be installed. Our alternate path at 589824000 clears that bar easily. Variance is an integer between 1 and 128, and it applies to the whole EIGRP process, not per prefix, so check the topology table for other destinations before turning it up in production. ## Configuring Variance (Named Mode) In named mode, variance lives under the topology base section: ``` R2(config)# router eigrp PINGLABZ R2(config-router)# address-family ipv4 unicast autonomous-system 100 R2(config-router-af)# topology base R2(config-router-af-topology)# variance 2 ``` Classic mode is the same keyword directly under `router eigrp 100`. The running config confirms placement: ``` R2# show run | section router eigrp router eigrp PINGLABZ ! address-family ipv4 unicast autonomous-system 100 ! topology base variance 2 exit-af-topology network 2.2.2.2 0.0.0.0 network 10.0.0.0 exit-address-family ``` ## Verifying with Traffic Share Counts Both paths are now installed, and this is the capture that matters: ``` R2# show ip route 3.3.3.3 Routing entry for 3.3.3.3/32 Known via "eigrp 100", distance 90, metric 4096000, type internal Routing Descriptor Blocks: 10.0.24.2, from 10.0.24.2, 00:00:25 ago, via Ethernet0/2 Route metric is 4608000, traffic share count is 71 Total delay is 8000 microseconds, minimum bandwidth is 10000 Kbit Loading 1/255, Hops 3 * 10.0.12.1, from 10.0.12.1, 00:00:25 ago, via Ethernet0/1 Route metric is 4096000, traffic share count is 80 Total delay is 7000 microseconds, minimum bandwidth is 10000 Kbit Loading 1/255, Hops 2 ``` The traffic share counts, 80 via R1 and 71 via R4, are the proof of *proportional* load balancing. EIGRP does not split flows 50/50 across unequal paths; it forwards in inverse proportion to metric, so the better path carries more. The ratio 80:71 tracks the metric ratio 4608000:4096000 almost exactly. CEF then maps flows onto the two paths per-destination by default, preserving packet ordering within a flow. The topology table now reports two successors for the prefix: ``` R2# show ip eigrp topology 3.3.3.3/32 EIGRP-IPv4 VR(PINGLABZ) Topology Entry for AS(100)/ID(2.2.2.2) for 3.3.3.3/32 State is Passive, Query origin flag is 1, 2 Successor(s), FD is 524288000, RIB is 4096000 ``` ## When Variance Does Nothing The most common variance complaint is "I configured it and nothing happened." Almost always, the alternate path is not a feasible successor: its RD is equal to or higher than the FD, usually because the paths share too much of the same geometry. In our lab this was true with all-default delays, where the alternate's RD exactly equaled the FD, and equal is not less than. We tuned interface delay on the alternate path (delay, not bandwidth, is the right knob per the [K values guide](https://www.pinglabz.com/eigrp-metric-k-values/)) until feasibility held. Check `show ip eigrp topology all-links` first, do the RD versus FD comparison yourself, and only then reach for variance. If the neighbor is not even in the topology table, that is an adjacency problem, and the [EIGRP neighbor troubleshooting guide](https://www.pinglabz.com/troubleshooting-eigrp-neighbor-adjacencies/) is the place to start. Also know the interaction with `maximum-paths` (default 4): variance selects candidates, maximum-paths caps how many get installed. ## How the Split Actually Happens: CEF and Traffic Share The routing table entry defines the ratio; CEF implements it. By default CEF shares per destination (strictly, per source-destination hash), so each flow sticks to one path and packets within a flow never reorder. The traffic share counts translate directly into CEF hash bucket allocations: with counts of 80 and 71, roughly 80 of every 151 flow hashes land on the R1 path and 71 on the R4 path. You can inspect the resulting distribution: ``` R2# show ip cef 3.3.3.3 internal ``` Two operational consequences follow. First, with only a handful of large flows (a nightly backup, a couple of iperf streams), the split can look nothing like 80:71, because each elephant flow hashes onto one path and stays there. Proportional sharing is a statistical promise that needs flow diversity to hold. Second, if you want the classic `traffic-share min across-interfaces` behavior (install all paths for fast failover but only forward on the best), that knob still exists: it gives you variance's convergence benefit without its load-sharing behavior, which is sometimes exactly right for latency-sensitive networks. None of that hashing behaviour is EIGRP’s doing. It belongs to CEF, it applies to every protocol that installs more than one next-hop, and it has a failure mode of its own where several routers in a row hash the same way and a single link ends up carrying nearly everything: [why one of two equal links takes all the traffic](https://www.pinglabz.com/ecmp-cef-load-balancing-polarization/) covers the per-flow selection and how to break the polarization. There is also a subtle stability win hiding here: every path variance installs is already in the RIB and FIB. When the successor dies, forwarding shifts to the surviving path with no DUAL computation, no query, not even the feasible-successor promotion delay. Variance is, among other things, a convergence feature. ## Variance Alongside the Rest of the Toolkit Variance interacts with several neighbors in the EIGRP feature family, and the combinations matter more than the individual knobs: - **maximum-paths** caps installed routes (default 4, up to 32). Variance nominates candidates; maximum-paths decides how many survive. On a router with many parallel paths, check both. - **Offset lists** let you nudge metrics per prefix without touching interface delay, which is handy when you want feasibility to hold for one destination but not another. Blunt delay changes affect every prefix crossing the interface. - **Summarization** can silently change the math: a summary's metric tracks its best component, so a summary learned over two paths may have a different FD/RD relationship than the components did. If variance behaves oddly for a summarized prefix, inspect the components on the far side of the summary (the mechanics are in the [summarization guide](https://www.pinglabz.com/eigrp-summarization-default-routes/)). - **Stub spokes** and variance rarely mix: a stub site with two unequal WAN uplinks is better served by variance on the hub side pointing in, plus stub on the spoke to contain queries. ## Should You Use It in Production? Variance is one of the few features OSPF simply cannot match, since OSPF load balancing is strictly equal-cost (a fuller comparison lives in [EIGRP vs OSPF](https://www.pinglabz.com/eigrp-vs-ospf/)). It shines when you have two WAN circuits of different speeds and you are paying for both: variance puts the slower circuit to work instead of leaving it idle as a pure backup. The trade-offs are real, though. Unequal paths mean unequal latency, which can upset jitter-sensitive traffic, and per-destination CEF sharing can be lumpy with few flows. Measure before and after, and prefer it for bulk data over voice. ## Key Takeaways - Variance multiplies the FD; any feasible successor with a metric below FD x variance is installed. Default is 1, maximum 128. - Only feasible successors qualify. The feasibility condition (RD strictly less than FD) is non-negotiable, which keeps every extra path loop-free. - Compute the value: worst acceptable metric divided by FD, rounded up. Do not just set a big number and hope. - Verify with `show ip route ` and read the traffic share counts: EIGRP balances proportionally to metric, not evenly. - If variance does nothing, the backup path failed the feasibility check. Fix the geometry or the delay values, not the variance number. Unequal-cost load balancing is EIGRP flexing the one muscle no other IGP has. For the DUAL machinery that makes it safe, and everything else in the protocol, the [EIGRP complete guide](https://www.pinglabz.com/eigrp/) ties it all together. ### EIGRP Route Summarization and Default Routes URL: https://www.pinglabz.com/eigrp-summarization-default-routes/ Last updated: 2026-07-11T16:05:17.000Z Route summarization is one of those EIGRP features that looks like a simple address-shortening trick but quietly solves three different problems at once: it shrinks topology tables, it speeds up convergence, and it creates query boundaries that protect you from stuck-in-active events. If you are working through the [EIGRP complete guide](https://www.pinglabz.com/eigrp/), summarization is the point where the protocol stops being a config exercise and starts being a design tool. In this article you will configure manual summarization in both classic and named mode on Cisco IOS XE, watch the Null0 discard route appear, advertise a default route with `summary-address 0.0.0.0/0`, and see a suppression side effect that surprises a lot of engineers the first time. Every output below is captured live from a CML lab, not typed from memory. ## Why Summarize in EIGRP EIGRP is a distance vector protocol, so every router carries every prefix it learns in its topology table. In a small lab that costs nothing. In a network with a few thousand routes, every extra prefix is more memory, more update traffic, and, most importantly, more queries when something fails. When EIGRP loses a route with no feasible successor, it queries its neighbors for that exact prefix (the full mechanics are in the [EIGRP query process and stuck-in-active deep dive](https://www.pinglabz.com/eigrp-query-process-stuck-in-active/)). A router that only knows a summary cannot be queried for the missing /24 inside it, which means summarization physically limits how far queries propagate. Three wins, one command: - **Smaller routing and topology tables.** Four /24s become one /22\. Multiply that across a real address plan and the savings are significant. - **Fewer updates.** A flapping /24 inside a stable summary is invisible to the rest of the network, because the summary metric only changes when the best component metric changes. - **Query boundaries.** Routers that only know the summary reply immediately to queries for component routes, which is your main structural defense against stuck-in-active. ## Interface-Level Summarization (Not Like OSPF) If you come from OSPF, recalibrate: OSPF can only summarize at area boundaries on ABRs and ASBRs (see the [OSPF complete guide](https://www.pinglabz.com/ospf/) for why). EIGRP has no areas, so it summarizes *per interface, on any router*. You decide exactly which neighbor sees the summary and which sees the components. That flexibility is a genuine architectural advantage of EIGRP, and it is one of the honest points in favor of EIGRP in the [EIGRP vs OSPF comparison](https://www.pinglabz.com/eigrp-vs-ospf/). ## The Lab Five IOL-XE routers in CML. R1 is the hub, R2 and R3 are spokes, and R4-R5 form a longer alternate path. R3 carries four loopbacks that simulate a site address block: 10.1.0.0/24 through 10.1.3.0/24, which pack neatly into 10.1.0.0/22\. R1, R2, and R4 run named mode EIGRP (`router eigrp PINGLABZ`, AS 100) while R3 and R5 run classic mode (`router eigrp 100`), so you will see both syntaxes. Before summarization, R1 sees all four components from R3: ``` R1# show ip route eigrp | include 10.1. D 10.1.0.0/24 [90/3584000] via 10.0.13.2, 00:00:16, Ethernet0/2 D 10.1.1.0/24 [90/3584000] via 10.0.13.2, 00:00:16, Ethernet0/2 D 10.1.2.0/24 [90/3584000] via 10.0.13.2, 00:00:16, Ethernet0/2 D 10.1.3.0/24 [90/3584000] via 10.0.13.2, 00:00:16, Ethernet0/2 ``` ## Classic Mode: ip summary-address eigrp In classic mode the command lives on the interface, pointing in the direction the summary should be advertised. R3 summarizes toward both of its neighbors: ``` R3(config)# interface Ethernet0/1 R3(config-if)# ip summary-address eigrp 100 10.1.0.0 255.255.252.0 R3(config-if)# interface Ethernet0/2 R3(config-if)# ip summary-address eigrp 100 10.1.0.0 255.255.252.0 ``` The AS number in the command must match your EIGRP process (a mismatch is silently useless). Verify on the receiving side. R2's table now carries a single /22, learned over both of its paths: ``` R2# show ip route eigrp | include 10.1. D 10.1.0.0/22 [90/4608000] via 10.0.24.2, 00:00:16, Ethernet0/2 [90/4096000] via 10.0.12.1, 00:00:16, Ethernet0/1 ``` Four prefixes became one, network-wide, with two interface commands on one router. ## The Null0 Discard Route The summarizing router installs something interesting in its own table: ``` R3# show ip route | include Null0 D 10.1.0.0/22 is a summary, 00:00:24, Null0 ``` This is the discard route, and it exists to prevent routing loops. Imagine R3 advertises 10.1.0.0/22 but loses 10.1.2.0/24\. A packet for 10.1.2.5 arrives because the summary attracted it. Without the Null0 route, R3 might follow a default route back toward the neighbor that sent the packet, which follows the summary back to R3, and you have a loop that only dies at TTL expiry. With the discard route, longest prefix match works: no /24, so the /22 to Null0 wins, and the packet is cleanly dropped. The summary is a real EIGRP topology entry on R3, with a metric inherited from the lowest-metric component: ``` R3# show ip eigrp topology 10.1.0.0/22 EIGRP-IPv4 Topology Entry for AS(100)/ID(10.1.3.1) for 10.1.0.0/22 State is Passive, Query origin flag is 1, 1 Successor(s), FD is 128256 Descriptor Blocks: 0.0.0.0 (Null0), from 0.0.0.0, Send flag is 0x0 Composite metric is (128256/0), route is Internal Vector metric: Minimum bandwidth is 8000000 Kbit Total delay is 5000 microseconds Reliability is 255/255 Load is 1/255 Minimum MTU is 1514 Hop count is 0 Originating router is 10.1.3.1 ``` Notice hop count 0 and next hop 0.0.0.0: the summary originates here. If you need the mechanics of how that composite metric is computed, the [EIGRP metric and K values guide](https://www.pinglabz.com/eigrp-metric-k-values/) breaks it down, and the 64-bit version of the story is in the [EIGRP wide metrics article](https://www.pinglabz.com/eigrp-wide-metrics/). ## Named Mode: summary-address Under af-interface Named mode moves the same feature under the routing process, in the af-interface section. This keeps all EIGRP policy in one place instead of scattering it across interface configs (the layout logic is covered in the [EIGRP configuration guide](https://www.pinglabz.com/eigrp-configuration-cisco/)): ``` R1(config)# router eigrp PINGLABZ R1(config-router)# address-family ipv4 unicast autonomous-system 100 R1(config-router-af)# af-interface Ethernet0/1 R1(config-router-af-interface)# summary-address 10.1.0.0/22 ``` Named mode also accepts CIDR notation, which is a small quality-of-life win over the classic dotted mask. ## Advertising a Default Route with summary-address 0.0.0.0/0 A default route is just the ultimate summary: it covers everything. That means `summary-address 0.0.0.0/0` is a legitimate and clean way to originate a default into EIGRP from an edge or hub router. Here R1 advertises a default toward R2 in named mode: ``` R1(config)# router eigrp PINGLABZ R1(config-router)# address-family ipv4 unicast autonomous-system 100 R1(config-router-af)# af-interface Ethernet0/1 R1(config-router-af-interface)# summary-address 0.0.0.0/0 ``` R1 installs its own discard route, exactly as before: ``` R1# show ip route | include Null0 D* 0.0.0.0/0 is a summary, 00:00:26, Null0 ``` And R2 receives a candidate default, flagged with the asterisk: ``` R2# show ip route eigrp | begin Gateway Gateway of last resort is 10.0.12.1 to network 0.0.0.0 D* 0.0.0.0/0 [90/1024640] via 10.0.12.1, 00:00:18, Ethernet0/1 1.0.0.0/32 is subnetted, 1 subnets D 1.1.1.1 [90/2560640] via 10.0.24.2, 00:00:18, Ethernet0/2 3.0.0.0/32 is subnetted, 1 subnets D 3.3.3.3 [90/4608000] via 10.0.24.2, 00:00:18, Ethernet0/2 ``` ## The Side Effect That Bites: Summaries Suppress Components Look at that last capture again. Before this change, R2 reached 1.1.1.1 (R1's own loopback) directly via R1\. Now it reaches it via 10.0.24.2, the long way around through R4\. Why? A summary suppresses every more-specific prefix on the interface where it is configured, and 0.0.0.0/0 covers *everything*. R1 stopped advertising all individual routes to R2 the moment the default went on. R2 still has specifics learned from its other neighbor R4, so those now win by longest prefix match, even though the physical path is worse. **Rule of thumb before you summarize** A summary-address suppresses all component routes on that interface. If the neighbor has another path to those components, traffic silently shifts to it. Advertising 0.0.0.0/0 suppresses every prefix, so use it only where the neighbor genuinely should send everything to you, such as a stub site behind a single hub. This pairs naturally with [EIGRP stub routing](https://www.pinglabz.com/eigrp-stub-routing/). ## Other Ways to Originate a Default The summary approach is not the only one, and it pays to know the trade-offs: - **Redistribute a static default.** `ip route 0.0.0.0 0.0.0.0 ` plus `redistribute static` under EIGRP. The route arrives as EIGRP external (D\*EX, administrative distance 170), and it does not suppress anything. This is the common choice at an internet edge. BGP-learned defaults at that edge are a different conversation, covered in the [BGP complete guide](https://www.pinglabz.com/bgp/). - **ip default-network.** Legacy, classful, easy to get wrong. Know it exists for old configs; do not deploy it new. - **summary-address 0.0.0.0/0.** Internal route (AD 90), interface-scoped, suppresses everything. Best for hub-to-stub designs. ## Controlling the Summary Metric and Leaking Exceptions By default the summary inherits the lowest metric among its components, which means the summary is recomputed every time that best component changes. On a router summarizing hundreds of prefixes, that recomputation has a CPU cost, and the resulting metric churn partially defeats the stability you were buying. Named mode has a fix: pin the metric. ``` R1(config-router-af)# topology base R1(config-router-af-topology)# summary-metric 10.1.0.0/22 10000 100 255 1 1500 ``` The five values are the classic bandwidth, delay, reliability, load, and MTU inputs. With a pinned metric, component flaps stop touching the summary entirely. The trade-off is honesty: the advertised metric no longer reflects the real best path, so reserve it for summaries whose components are metric-equivalent anyway. The other refinement is the leak map. Sometimes you want the summary *and* one exception: advertise 10.1.0.0/22 but also leak 10.1.2.0/24 so a specific service keeps an exact route (perhaps to steer it down a preferred link). A leak map is a route map naming the components allowed to escape suppression: ``` R3(config)# ip prefix-list LEAK-VOICE permit 10.1.2.0/24 R3(config)# route-map LEAK-MAP permit 10 R3(config-route-map)# match ip address prefix-list LEAK-VOICE R3(config)# interface Ethernet0/1 R3(config-if)# ip summary-address eigrp 100 10.1.0.0 255.255.252.0 leak-map LEAK-MAP ``` Neighbors now receive both the /22 and the leaked /24, and longest prefix match sends the exception traffic exactly where you pointed it. Leak maps are also the escape hatch when a default-route summary would otherwise suppress something you need visible. ## Where to Summarize: Design Patterns Summarization works best where the addressing hierarchy meets the physical hierarchy. Three patterns cover most networks: - **Hub toward spokes.** Spokes rarely need the full enterprise table. A default or a short summary list from the hub keeps spoke routers tiny, and pairs with [stub routing](https://www.pinglabz.com/eigrp-stub-routing/) in the other direction. This is the classic WAN pattern. - **Site toward core.** Each site advertises one block (like R3's /22 here) upstream. Core topology tables stay flat no matter how many VLANs a site adds, and site-internal flaps never leave the site: this is the query boundary in action, the same mechanism that protects against [stuck-in-active](https://www.pinglabz.com/eigrp-query-process-stuck-in-active/). - **Between routing domains.** At redistribution boundaries, summarizing before you redistribute caps how much churn crosses into the other protocol. One caution from the suppression discussion above: before adding any summary on a transit router, run `show ip eigrp topology` for the components and check which neighbors have alternate paths to them. The summary changes what those neighbors prefer, and the new best paths should be ones you actually want carrying the traffic. ## Key Takeaways - EIGRP summarizes per interface on any router, unlike OSPF's area-boundary restriction. That makes it a precision tool: each neighbor can see a different level of detail. - The summarizing router installs a Null0 discard route automatically. It is loop prevention, not a bug, and packets for missing components are dropped by longest prefix match. - The summary inherits the lowest component metric, and it stays quiet unless that best metric changes, which dampens the blast radius of flapping components. - Summaries are query boundaries. This is the structural fix for stuck-in-active, alongside stub routing. - `summary-address 0.0.0.0/0` is a clean internal default, but it suppresses every more-specific route on that interface. Check what else the neighbor can see before you commit. Summarization decides how much of your network each router has to think about. Get it right and queries die at the boundary you drew. For where this fits in the bigger picture of DUAL, metrics, and design, head back to the [EIGRP complete guide](https://www.pinglabz.com/eigrp/). ### TFTP vs FTP for Cisco IOS File Transfers and Config Backup URL: https://www.pinglabz.com/tftp-vs-ftp-cisco-ios/ Last updated: 2026-07-10T23:18:04.000Z Sooner or later every engineer copies something onto or off a router: a config backup before a risky change, a new IOS image, a crash dump for TAC. The two classic transports are TFTP and FTP, they behave very differently, and the CCNA expects you to know when each is appropriate. This article backs up a real Cisco IOS XE config over both protocols to a Linux server, with the genuine transfer output (including an honest speed difference), as part of the [IP Services complete guide](https://www.pinglabz.com/ip-services/). ## Two Protocols, Two Philosophies **TFTP (Trivial FTP)** UDP port 69\. No authentication, no directory listing, no encryption. Lock-step transfer: one 512-byte block, one ack, repeat. Trivial to serve, painfully simple, slow on anything big. **FTP** TCP ports 21 (control) + 20/negotiated (data). Username/password authentication, directory operations, TCP windowing for speed. Still cleartext on the wire. ## Config Backup over TFTP TFTP needs zero configuration on the router. Point `copy` at a TFTP URL and accept the prompts: ``` R1#copy running-config tftp://192.168.99.100/R1-confg Address or name of remote host [192.168.99.100]? Destination filename [R1-confg]? .!! 2425 bytes copied in 4.243 secs (572 bytes/sec) ``` 572 bytes per second. That is not a broken lab: it is the lock-step protocol doing one small block per round trip (the leading `.` in `.!!` is a retry, also normal for the first datagram while the server process wakes up). For a 2 KB config it does not matter. For a 700 MB IOS XE image, do the math and pick something else. ## Config Backup over FTP FTP needs credentials, configured once globally: ``` R1(config)#ip ftp username j R1(config)#ip ftp password Cisco@123 ``` Then the same copy syntax with an ftp:// URL: ``` R1#copy running-config ftp://192.168.99.100/R1-backup.cfg Address or name of remote host [192.168.99.100]? Destination filename [R1-backup.cfg]? Writing R1-backup.cfg ! 2469 bytes copied in 0.525 secs (4703 bytes/sec) ``` Same file, same network path, roughly eight times faster (4703 vs 572 bytes/sec), because TCP streams instead of lock-stepping. You can also put credentials inline (`ftp://user:pass@host/file`), but that leaves the password in command history; the global config approach is the cleaner habit, and either way treat FTP credentials as disposable because they cross the wire in cleartext. ## The Other Direction: Restoring to the Router Downloads use the same command with source and destination swapped. On this virtual platform the writable filesystem is called `unix:` (on physical hardware you will use `flash:` or `bootflash:`, same syntax): ``` R1#copy tftp://192.168.99.100/R1-confg unix:R1-restore.cfg Destination filename [R1-restore.cfg]? Accessing tftp://192.168.99.100/R1-confg... Loading R1-confg from 192.168.99.100 (via Ethernet0/0): ! [OK - 2425 bytes] 2425 bytes copied in 0.028 secs (86607 bytes/sec) R1#dir unix: | include cfg 2624242 -rw- 2425 Jul 10 2026 22:57:07 +00:00 R1-restore.cfg ``` `show file systems` lists what filesystems your platform actually has, which is worth a look before any image work. Note the classic restore trap: copying a config to `running-config` *merges* it with what is there, it does not replace it. For a true replacement, copy to `startup-config` and reload, or use `configure replace`. ## What the Server Side Sees (and a Security Lesson) Both files landed on the Linux server, and the listing quietly teaches the security difference: ``` j@llmbits:~$ ls -l /srv/tftp/R1-confg /home/j/R1-backup.cfg -rw-rw-rw- 1 tftp tftp 2425 Jul 10 15:56 /srv/tftp/R1-confg -rw------- 1 j j 2469 Jul 10 15:56 /home/j/R1-backup.cfg ``` The FTP upload is owned by the authenticated user with private permissions. The TFTP upload sits in a world-writable directory owned by a daemon account, because TFTP has no concept of identity: anyone who can reach UDP 69 can read (and depending on server flags, overwrite) what is there. Your startup configs contain password hashes and SNMP strings; a TFTP directory full of them is a reconnaissance jackpot. Keep TFTP servers on isolated management segments and empty them when the maintenance is done. ## The Workflow That Actually Matters: Image Upgrades Config backups are the daily habit, but the high-stakes transfer is an IOS image. The discipline around it matters more than the protocol. Before anything: `show file systems` and `dir` to confirm free space, because a transfer that fills the filesystem mid-write leaves you with a truncated image and no room to fix it. After the copy, verify integrity before you ever point the boot system at it: ``` R1#verify /md5 flash:cat9k_iosxe.17.09.05.SPA.bin ! compare the digest against the one on Cisco's download page ``` A flipped bit in a 700 MB transfer over UDP-based TFTP is not hypothetical, and a corrupt image discovered at reload time is a very different evening than one discovered by `verify /md5`. Then stage the boot variable (`boot system flash:`), save, and reload in a window. Keep the old image on flash until the new one has survived a week; disk space is cheaper than a rollback without media. For recurring config protection there is also the built-in archive feature, which snapshots on every write and can push each snapshot to your FTP/TFTP/SCP server automatically: ``` archive path ftp://192.168.99.100/backups/R1 write-memory ``` With that in place, every `write memory` quietly ships a timestamped config copy off-box, which converts "when was the last good backup?" from an archaeology project into a directory listing. ## Which One When (and the Modern Answer) Use TFTP for small, quick, disposable transfers inside a protected management network: a config snapshot before a change window, a boot helper for [appliance](https://www.pinglabz.com/cisco-asa/) recovery (ROMMON environments often speak only TFTP, which is the real reason the protocol refuses to die). Use FTP when the file is big (IOS images) or when you want per-user accounting on the server. And for anything crossing untrusted networks, skip both: IOS XE speaks SCP and SFTP natively (`copy running-config scp://user@host/R1.cfg`, after enabling `ip scp server enable` for the inbound case), giving you SSH encryption and authentication with the same copy syntax. Automated config backup at scale usually rides SSH anyway, driven by [automation tooling](https://www.pinglabz.com/network-automation/) rather than hand-typed copies. ## Key Takeaways TFTP is UDP 69, unauthenticated, lock-step slow (572 B/s in the real capture), and perfect only for small files on trusted segments; FTP is TCP 21/20 with credentials and TCP speed (4703 B/s, same file, same path). The `copy` command is symmetric: URL as source restores, URL as destination backs up, with `flash:`/`unix:` naming varying by platform. Remember copy-to-running merges rather than replaces, keep TFTP directories clean because they are readable by anyone, and reach for SCP when security matters. The rest of the operational toolbox is in the [IP Services complete guide](https://www.pinglabz.com/ip-services/). ### DHCP and DNS on Cisco Networks: Roles, Relay, and ip helper-address URL: https://www.pinglabz.com/dhcp-and-dns-cisco/ Last updated: 2026-07-10T23:18:04.000Z DHCP hands a host its network identity; DNS makes names usable instead of addresses. The CCNA asks you to keep their roles straight and to configure a router as a DHCP server and relay, and the relay part is where real understanding shows: DHCP discovery is broadcast, routers do not forward broadcasts, and yet every enterprise runs central DHCP servers several hops away from their clients. This article builds that exact scenario on real Cisco IOS XE (server on R1, relay on R2, client on R3) and reads the state on all three, as part of the [IP Services complete guide](https://www.pinglabz.com/ip-services/). ## DHCP vs DNS: Two Jobs, Often Confused **DHCP** Leases network configuration to hosts: IP address, mask, default gateway, DNS server addresses, domain name. UDP 67 (server) / 68 (client). A host uses it once at attach time and at renewals. **DNS** Resolves names to addresses (and back) on demand: UDP/TCP 53\. A host uses it constantly, for nearly every connection it opens. The link between them: one of the most important options DHCP hands out is *which DNS servers to use*. Get the DHCP pool's `dns-server` option wrong and every client on that subnet "has no internet" even though routing is perfect (nothing resolves). ## DORA, and What the Relay Actually Does A fresh client finds a server with four messages: Discover, Offer, Request, Ack. Discover and Request are broadcasts (the client has no address yet, and may not even know the server). Broadcasts die at the router, so on any subnet without a local DHCP server you configure the router's LAN interface as a **relay agent** with `ip helper-address`. The relay picks up the broadcast, stamps the receiving interface's address into the packet's `giaddr` (gateway address) field, and forwards it as *unicast* to the server. The server uses giaddr to pick the right pool, and replies unicast back to the relay, which delivers the answer to the client's LAN. ## The Lab: Server, Relay, Client on Three Routers R1 is the central DHCP server, two hops from the client LAN (10.0.40.0/24). Its pool: ``` ip dhcp excluded-address 10.0.40.1 10.0.40.9 ip dhcp pool LAN40 network 10.0.40.0 255.255.255.0 default-router 10.0.40.1 dns-server 192.168.99.100 domain-name pinglabz.lab ``` R2 owns the client LAN and relays toward R1 (1.1.1.1 is R1's loopback, reachable via OSPF): ``` R2#show run interface Ethernet0/3 interface Ethernet0/3 description LAN40-DHCP-RELAY ip address 10.0.40.1 255.255.255.0 ip helper-address 1.1.1.1 end ``` R3 plays the client with `ip address dhcp` on its interface. Now the state on all three, starting with the client: ``` R3#show ip interface brief | include Ethernet0/3 Ethernet0/3 10.0.40.10 YES DHCP up up R3#show dhcp lease | include IP|Lease|serv Temp IP addr: 10.0.40.10 for peer on Interface: Ethernet0/3 DHCP Lease server: 10.0.12.1, state: 5 Bound Lease: 86400 secs, Renewal: 43200 secs, Rebind: 75600 secs ``` Method `DHCP`, state `Bound`, and the three timers every DHCP client runs: renew at 50% of the lease (unicast to the original server), rebind at 87.5% (broadcast to any server), expire at 100% (start over with Discover). On the server, the lease appears in the binding table: ``` R1#show ip dhcp binding IP address Client-ID/ Lease expiration Type State Interface Hardware address/ User name 10.0.40.10 0063.6973.636f.2d61. Jul 11 2026 10:38 PM Automatic Active Ethernet0/1 R1#show ip dhcp pool Pool LAN40 : Total addresses : 254 Leased addresses : 1 Excluded addresses : 9 Current index IP address range Leased/Excluded/Total 10.0.40.11 10.0.40.1 - 10.0.40.254 1 / 9 / 254 ``` Notice the binding's Interface column says Ethernet0/1: that is where the *relayed unicast* arrived from R2, proof the broadcast-to-unicast conversion happened. The server statistics show the DORA conversation counters: ``` R1#show ip dhcp server statistics Message Received DHCPDISCOVER 17 DHCPREQUEST 1 Message Sent DHCPOFFER 6 DHCPACK 1 ``` (Seventeen Discovers for one lease? The client started broadcasting at boot, before OSPF had converged and before the relay could reach 1.1.1.1\. Real networks show this pattern after power events, and it is normal: the client retries until the path exists.) ## Beyond the Basic Pool: Reservations, Options, and Conflicts Three pool features come up constantly in real deployments. A **manual binding** (DHCP reservation) pins a specific host to a specific address by its client identifier, giving you the stability of a static address with the central management of DHCP: ``` ip dhcp pool PRINTER-1 host 10.0.40.50 255.255.255.0 client-identifier 0100.1122.3344.55 ``` **Options** carry vendor-specific bootstrap data beyond gateway and DNS. The two you will meet on Cisco networks: option 150 (TFTP server list for IP phones fetching their config) and option 43 (controller addresses for access points joining a [wireless LAN controller](https://www.pinglabz.com/wireless/)). If phones or APs on a subnet will not come up, the pool's options are the first suspect. And **conflicts**: before offering an address, the server pings it; if something answers, the address is quarantined in the conflict table rather than leased. A growing `show ip dhcp conflict` table almost always means someone has been hand-configuring static addresses inside the DHCP range, which is what excluded-address ranges are for. ## The DNS Side on IOS Routers are DNS *clients* too, and the CCNA expects the client-side commands. Point the router at a resolver and let it resolve names in ping/traceroute/logs: ``` ip name-server 192.168.99.100 ip domain lookup ip domain name pinglabz.lab ``` `ip domain lookup` is on by default, which is why a typo at the CLI hangs while the router tries to resolve it as a hostname; many engineers disable it on lab boxes (`no ip domain lookup`, exactly what this lab's base config does) and enable it deliberately where name resolution is wanted. A router can even be a modest DNS server (`ip dns server`) for small sites, though in practice DNS service lives on dedicated infrastructure and the router's job is simply to hand out the right resolver addresses via its DHCP pools. ## The DNS Records a Network Engineer Actually Touches You do not need to run BIND to need DNS literacy. Four record types cover nearly every infrastructure conversation: **A** maps a name to an IPv4 address, **AAAA** to an IPv6 address, **CNAME** aliases one name to another (monitoring dashboards love pointing at CNAMEs so the backend can move), and **PTR** maps an address back to a name. PTR records are the sleeper: reverse lookups are what make your traceroutes readable, your syslog collector display hostnames, and some services (mail, strict SSH configs) actively check them. When the NOC says "traceroute shows weird names," someone's PTR zone is stale, not the routing. On the resolution path itself, remember the client asks a *recursive* resolver (the address DHCP handed out), and that resolver walks the authoritative hierarchy on the client's behalf; caching at every layer is why a DNS change "takes time" (the TTL on the record is the contract). ## Troubleshooting the Chain When a client gets no address, walk the chain in order: is the client actually sending Discovers (interface up, `ip address dhcp` present)? Does the LAN's router interface have the `ip helper-address`? Can the relay reach the server address (ping it *sourced from the client-facing interface*, since that becomes giaddr and the server's replies go there)? Does the server have a pool matching the giaddr subnet, with addresses left in it (`show ip dhcp pool`)? And if the client gets an address but "nothing works," check what came *with* the lease: wrong `default-router` or `dns-server` options break connectivity in ways that look like routing failures. `debug ip dhcp server events` on the server shows each DORA step live when you need the ground truth. ## Key Takeaways DHCP leases identity (address, gateway, DNS servers); DNS resolves names; the exam and real life both hinge on not blurring them. Relay agents (`ip helper-address`) convert client broadcasts to unicasts and stamp giaddr so a central server can serve every subnet, which the lab proved end to end: client Bound at 10.0.40.10, binding on the server two hops away, DORA counters ticking. Excluded ranges protect static addresses, the pool's options matter as much as the addresses, and lease timers (50% renew, 87.5% rebind) explain most "it fixed itself" stories. The rest of the operational services live in the [IP Services complete guide](https://www.pinglabz.com/ip-services/). ### Syslog on Cisco IOS XE: Severity Levels, Logging Host, and Real Messages URL: https://www.pinglabz.com/syslog-cisco-ios-xe/ Last updated: 2026-07-10T23:18:04.000Z When something breaks at 3 a.m., the first question is "what changed, and when?" Syslog is the answer, but only if you set it up before the outage: severity levels chosen deliberately, messages shipped off-box to a collector, and clocks synced so the timeline means something. This article configures Cisco IOS XE logging to a real Linux syslog server and reads the actual messages that arrive, as part of the [IP Services complete guide](https://www.pinglabz.com/ip-services/). ## Anatomy of a Syslog Message Here is a real message captured on the collector during this lab: ``` 2026-07-10T15:53:28.713914-07:00 10.0.13.2 114: *Jul 10 22:53:27.685: %LINK-3-UPDOWN: Interface Loopback1, changed state to up ``` Reading left to right: the collector's own receive timestamp, the source router (10.0.13.2), the router's message sequence number (114), the router's local timestamp, and then the Cisco message body. That body always follows one pattern: ``` %FACILITY-SEVERITY-MNEMONIC: description %LINK - 3 - UPDOWN : Interface Loopback1, changed state to up ``` Facility is the subsystem that spoke (LINK, LINEPROTO, SYS, OSPF, SEC\_LOGIN). Severity is the number you filter on. The mnemonic identifies the exact event type, which is what you grep for. ## The Eight Severity Levels Severities run 0 through 7, and lower is worse. The classic mnemonic: *Every Awesome Cisco Engineer Will Need Ice cream Daily*. **0** **Emergency**system unusable **1** **Alert**act immediately **2** **Critical**critical condition **3** **Error**error condition (interface down lives here) **4** **Warning**warning condition **5** **Notification**normal but significant (config changes, logins) **6** **Informational**informational (the usual production trap level) **7** **Debugging**debug output; never ship this off-box permanently One subtlety that shows up in interviews: setting a level means "log this severity *and everything worse*." `logging trap informational` sends severities 0 through 6, not just 6. ## Where Messages Go IOS XE can deliver every message to four places independently: the console line (`logging console`), VTY sessions (`logging monitor` plus `terminal monitor` when you want it live in SSH), the in-memory buffer (`logging buffered`, your first stop during troubleshooting), and one or more remote collectors (`logging host`). The buffer disappears on reload and the console only helps if you are watching it, which is why the collector matters: it is the copy that survives the incident. ``` logging host 192.168.99.100 logging trap informational ``` That is the entire router-side configuration used in this lab. Verify what is actually being sent: ``` R1#show logging | include Trap logging|logging host Trap logging: level informational, 220 message lines logged Logging to 192.168.99.100 (udp port 514, audit disabled ... ``` ## The Collector Side: rsyslog in 6 Lines Any Linux box becomes a syslog collector with rsyslog's UDP input module. This is the exact configuration the lab VM runs (as `/etc/rsyslog.d/10-ciscolab.conf`), filtering lab devices into their own file: ``` module(load="imudp") input(type="imudp" port="514") if ($fromhost-ip startswith "10.0." or $fromhost-ip startswith "192.168.99.") then { action(type="omfile" file="/var/log/ciscolab.log") stop } ``` Then we caused some events (an interface bounce on R3, an SSH login to R1) and read the file: ``` j@llmbits:~$ sudo tail /var/log/ciscolab.log 2026-07-10T15:53:26.748667-07:00 10.0.13.2 111: *Jul 10 22:53:25.720: %SYS-5-CONFIG_I: Configured from console by cisco on vty0 (192.168.99.100) 2026-07-10T15:53:27.715614-07:00 10.0.13.2 113: *Jul 10 22:53:27.685: %LINEPROTO-5-UPDOWN: Line protocol on Interface Loopback1, changed state to up 2026-07-10T15:53:28.713914-07:00 10.0.13.2 114: *Jul 10 22:53:27.685: %LINK-3-UPDOWN: Interface Loopback1, changed state to up 2026-07-10T15:53:31.948104-07:00 192.168.99.1 219: Jul 10 22:53:31.240: %SEC_LOGIN-5-LOGIN_SUCCESS: Login Success [user: cisco] [Source: 192.168.99.100] [localport: 22] at 22:53:31 UTC Fri Jul 10 2026 ``` Look at what you get for free: *who configured what from where* (%SYS-5-CONFIG\_I names the user, the line, and the source IP), interface state history with timestamps, and an authentication audit trail (%SEC\_LOGIN-5-LOGIN\_SUCCESS). During the [login hardening lab](https://www.pinglabz.com/password-policies-mfa-network-devices/), failed brute-force attempts landed in this same file as %SEC\_LOGIN-4-LOGIN\_FAILED, which is exactly the signal a SIEM alerts on. ## The Local Buffer: Your First Stop in an Incident Before you ever reach for the collector, `show logging` on the device itself shows the in-memory buffer. Size it deliberately, because the default is small enough that a busy OSPF event can push the interesting lines out before you log in: ``` logging buffered 64000 informational service timestamps log datetime msec localtime show-timezone service sequence-numbers ``` The `service timestamps` line is why the captured messages above carry millisecond timestamps: without it you get relative uptime stamps that are useless for correlation. `service sequence-numbers` adds the incrementing counter (the `114:` in the capture), which lets you spot dropped messages: if the collector has 113 and 115 but no 114, a datagram died in transit, and with UDP transport nothing will retransmit it. Two more line-level behaviors worth setting everywhere: `logging synchronous` on console and VTY lines stops log output from splicing itself into the middle of whatever command you are typing, and `logging console warnings` keeps a chatty box from drowning the console during an event storm while still showing you severities 0 through 4. ## Operational Discipline Three practices turn logging from noise into evidence. First, sync clocks with [NTP](https://www.pinglabz.com/ntp-cisco-ios-xe/); correlating an event across five devices requires their timestamps to agree, and a beautiful pipeline of lies is worse than no logs at all. Second, pick trap levels deliberately: informational (6) to the collector is the sane default, debugging (7) only ever temporarily and only for a targeted debug, because a single `debug ip packet` at severity 7 can generate more log volume than the rest of the network combined. Third, remember syslog over UDP 514 is fire-and-forget cleartext; a dropped datagram is a lost log line, and anyone on the path can read it. rsyslog and IOS XE both support TCP and TLS transport when logs cross untrusted segments. Also decide what source address your devices log *from*. By default the packet leaves with the egress interface's address, which means the same router shows up in your collector under different IPs depending on routing at that moment (you can see it in the captures above: R3's messages arrived from 10.0.13.2, its egress toward the collector). `logging source-interface Loopback0` pins every message to one stable, identifiable address, which keeps collector-side filtering and per-device log files sane. ## Key Takeaways Syslog messages follow %FACILITY-SEVERITY-MNEMONIC, severities run 0 (emergency) to 7 (debugging), and configuring a level always includes everything more severe. Two commands (`logging host`, `logging trap informational`) ship every meaningful event to a collector, and six lines of rsyslog catch them in a dedicated file. The captured output shows the payoff: config-change attribution, interface history, and login auditing, all timestamped. Pair it with NTP or the timeline lies to you. The rest of the day-2 stack (SNMP for polling, NTP for time, DHCP for addressing) is in the [IP Services complete guide](https://www.pinglabz.com/ip-services/). ### SNMP Explained: v2c vs v3, OIDs, MIBs, and Real Polling URL: https://www.pinglabz.com/snmp-explained/ Last updated: 2026-07-10T23:18:03.000Z SNMP is how monitoring systems actually know your interface went down, your CPU spiked, or your optics are running hot. Every NMS you will ever deploy (LibreNMS, Zabbix, SolarWinds, Catalyst Center's assurance features) speaks it. This article explains the architecture, then polls a real Cisco IOS XE router from a Linux host with both SNMPv2c and SNMPv3, including what happens when credentials are wrong. It is part of the [IP Services complete guide](https://www.pinglabz.com/ip-services/). ## Manager, Agent, MIB, OID Four terms carry the whole protocol. The **manager** is the monitoring station that asks questions. The **agent** is the process on the router that answers them. The **MIB** (Management Information Base) is the schema: a tree of everything the agent can report. And an **OID** (Object Identifier) is one specific leaf on that tree, written as dotted numbers. `1.3.6.1.2.1.1.5.0` is the device's name; the MIB file is what lets your tools display that as the human-readable `sysName.0`. Traffic flows two ways: the manager polls the agent on UDP 161 (get/walk), and the agent can push unsolicited **traps** (or acknowledged **informs**) to the manager on UDP 162 when something happens. ## SNMPv2c: One Line to Configure, One Secret to Steal On the router, a read-only v2c setup is a single line plus optional identity strings: ``` snmp-server community pinglabz-ro RO snmp-server location PingLabz CML Lab snmp-server contact noc@pinglabz.com ``` From the Linux management host, the classic first poll, straight from the lab: ``` j@llmbits:~$ snmpget -v2c -c pinglabz-ro 192.168.99.1 sysDescr.0 sysName.0 sysUpTime.0 sysLocation.0 sysContact.0 SNMPv2-MIB::sysDescr.0 = STRING: Cisco IOS Software [IOSXE], Linux Software (X86_64BI_LINUX-ADVENTERPRISEK9-M), Version 17.18.2, RELEASE SOFTWARE (fc3) Technical Support: http://www.cisco.com/techsupport Copyright (c) 1986-2025 by Cisco Systems, Inc. Compiled Fri 19-Dec-25 03:28 by mc SNMPv2-MIB::sysName.0 = STRING: R1.pinglabz.lab SNMPv2-MIB::sysUpTime.0 = Timeticks: (93503) 0:15:35.03 SNMPv2-MIB::sysLocation.0 = STRING: PingLabz CML Lab SNMPv2-MIB::sysContact.0 = STRING: noc@pinglabz.com ``` Walking a table is where SNMP earns its keep. The interface table (IF-MIB) is what every monitoring dashboard is built on: ``` j@llmbits:~$ snmpwalk -v2c -c pinglabz-ro 192.168.99.1 ifDescr IF-MIB::ifDescr.1 = STRING: Ethernet0/0 IF-MIB::ifDescr.2 = STRING: Ethernet0/1 IF-MIB::ifDescr.3 = STRING: Ethernet0/2 IF-MIB::ifDescr.4 = STRING: Ethernet0/3 j@llmbits:~$ snmpwalk -v2c -c pinglabz-ro 192.168.99.1 ifOperStatus IF-MIB::ifOperStatus.1 = INTEGER: up(1) IF-MIB::ifOperStatus.2 = INTEGER: up(1) IF-MIB::ifOperStatus.3 = INTEGER: up(1) IF-MIB::ifOperStatus.4 = INTEGER: down(2) ``` Index 4 (an unused interface) reports `down(2)`. Your NMS turns exactly this walk into the red/green port map you look at every day. Now the security problem. That community string travels in cleartext in every packet, and it *is* the entire authentication. Get it wrong (or don't have it) and the agent does not even send an error, it just stays silent: ``` j@llmbits:~$ snmpget -v2c -c wrongstring -r 1 -t 3 192.168.99.1 sysName.0 Timeout: No Response from 192.168.99.1. ``` Silence is deliberate (no information leak to scanners), but the flip side is that anyone who sniffs one legitimate poll owns your read access. That is why v2c survives only on isolated management networks, and why the CCNA wants you to know v3. ## SNMPv3: Users, Groups, and Real Cryptography SNMPv3 replaces community strings with a user-based security model. Three security levels exist: **noAuthNoPriv** Username only, no crypto. Barely better than v2c. **authNoPriv** Authenticated (SHA/MD5 HMAC), not encrypted. Tamper-proof but readable. **authPriv** Authenticated and encrypted (AES). The only level worth deploying. Configuration is a group (which sets the security level) plus a user (which carries the keys): ``` snmp-server group NETADMIN v3 priv snmp-server user snmpadmin NETADMIN v3 auth sha AuthPass123! priv aes 128 PrivPass123! ``` Verification on the router (note the user does not appear in the running config, by design; you must use `show snmp user`): ``` R1#show snmp group | include groupname|security groupname: NETADMIN security model:v3 priv groupname: pinglabz-ro security model:v1 groupname: pinglabz-ro security model:v2c R1#show snmp user User name: snmpadmin Engine ID: 800000090300AABBCC001300 storage-type: nonvolatile active Authentication Protocol: SHA Privacy Protocol: AES128 Group-name: NETADMIN ``` And the authenticated, encrypted poll from the manager: ``` j@llmbits:~$ snmpget -v3 -u snmpadmin -l authPriv -a SHA -A 'AuthPass123!' -x AES -X 'PrivPass123!' 192.168.99.1 sysName.0 sysUpTime.0 SNMPv2-MIB::sysName.0 = STRING: R1.pinglabz.lab SNMPv2-MIB::sysUpTime.0 = Timeticks: (94125) 0:15:41.25 ``` Same data as v2c, but now the request is HMAC-authenticated and the payload rides inside AES-128\. A wrong auth key fails cryptographically instead of just timing out, and nothing on the wire is readable. ## Reading the OID Tree Without Crying Every OID in the captures above lives under one famous prefix: `1.3.6.1.2.1` is `mib-2`, the standard tree every vendor implements (iso.org.dod.internet.mgmt.mib-2). System identity sits at `1.3.6.1.2.1.1` (sysDescr is .1, sysUpTime .3, sysName .5), and interfaces at `1.3.6.1.2.1.2`. Vendor-private data lives under `1.3.6.1.4.1.`, and Cisco's enterprise number is 9, which is why Cisco-specific counters (CPU, memory, temperature) all start with `1.3.6.1.4.1.9`. You do not memorize trees; you learn the two prefixes, and you use `snmptranslate -On IF-MIB::ifDescr` when you need to convert between names and numbers. Scalar objects end in `.0`; table cells end in their row index, which is why `sysName.0` but `ifDescr.2`. ## Version Comparison at a Glance **SNMPv1** 1988 vintage. Community strings, 32-bit counters (they wrap fast on 10G links). Historical only. **SNMPv2c** Adds getbulk and 64-bit counters. Still cleartext communities. Fine on an isolated mgmt VLAN, indefensible elsewhere. **SNMPv3** User security model, SHA auth, AES privacy, engine IDs. The deployable answer; slightly more config, much less regret. ## Traps: The Agent Speaks First Polling has a blind spot: if you poll interfaces every five minutes, a flap that lasts ninety seconds can come and go between polls. Traps close that gap by letting the agent push events the moment they happen. Configuration is two parts, what to send and where to send it: ``` snmp-server enable traps snmp linkdown linkup snmp-server enable traps config snmp-server host 192.168.99.100 version 3 priv snmpadmin ``` Be selective with `enable traps`: enabling everything turns your trap receiver into a firehose of BGP, entity, and environmental noise that nobody triages. Link state, config changes, and environmental alarms are the usual short list. Traps ride UDP 162 and are unacknowledged; if delivery matters, use *informs* (`snmp-server host ... informs`), which are retransmitted until the manager acknowledges them, at the cost of state on the device. In practice, mature monitoring uses both planes: traps for immediacy, polling as the safety net that catches whatever a lost trap missed. That combination (plus syslog, which overlaps traps considerably) is what keeps dashboards honest. ## Operational Notes Several things bite people in production. First, restrict who may poll: pair the community or group with an ACL (`snmp-server community pinglabz-ro RO 99`, where ACL 99 permits only your NMS addresses) so random hosts get the same silence a wrong community gets. Second, MIB files live on the *manager*, not the router; if your tooling prints raw numeric OIDs, your manager is missing MIBs (the agent was never going to send names, only numbers). We hit exactly this in the lab: the fresh Linux manager resolved nothing until the standard IETF MIB files were installed, at which point the same queries started printing `SNMPv2-MIB::sysName.0` instead of dotted decimals. Third, interface indexes (that `.1`, `.2` suffix in the walks above) are not guaranteed stable across reloads unless you configure `snmp-server ifindex persist`, and a graphing system that trusts unstable indexes will happily attach the WAN graph to the loopback after a reboot. Fourth, SNMP writes (`RW` communities) are how attackers reconfigure devices; if you do not have a tested reason for write access, do not enable it. Modern operations are moving the config side to [model-driven APIs and automation tools](https://www.pinglabz.com/network-automation/) and keeping SNMP for what it is genuinely good at: fast, cheap, universal telemetry polling. ## Key Takeaways SNMP is a manager polling an agent for OIDs defined in MIBs, on UDP 161, with traps coming back on 162\. v2c is one line of config whose community string is a cleartext password (and whose failure mode is silence, as the timeout capture showed). v3 with `authPriv` gives you SHA authentication and AES encryption for barely more configuration, verified with `show snmp group` and `show snmp user`. Scope access with ACLs, keep RW disabled unless you need it, and let your monitoring build on the IF-MIB walks you saw here. For syslog, NTP, and the rest of the operational stack, head back to the [IP Services complete guide](https://www.pinglabz.com/ip-services/). ### NTP on Cisco IOS XE: Stratum, ntp master, and Verification URL: https://www.pinglabz.com/ntp-cisco-ios-xe/ Last updated: 2026-07-10T23:18:03.000Z Wrong time on a network device is not a cosmetic problem. Certificate validation fails, log timestamps become useless for incident correlation, and time-based authentication (like 802.1X with certificates, or API tokens) starts breaking in ways that look random. NTP fixes all of that with two lines of configuration, and this article shows a real Cisco IOS XE deployment: one router acting as the time source, the others syncing to it, with genuine `show` output at every stage of convergence. It is part of the [IP Services complete guide](https://www.pinglabz.com/ip-services/). ## Stratum: How Far You Are From the Truth NTP organizes time sources in a hierarchy called stratum. Stratum 1 servers own a reference clock directly (GPS, atomic). A device syncing to stratum 1 becomes stratum 2, a device syncing to that becomes stratum 3, and so on down to stratum 15\. Stratum 16 has a special meaning: unsynchronized, no valid time source. It is not "level 16," it is "off the ladder entirely," and you will see it in real output below. Lower stratum means closer to the reference, not necessarily better quality, but as a rule you point your network at two or three consistent sources and let the hierarchy fan out from there. ## The Lab Design: One Master, Everyone Else a Client The standard enterprise pattern is exactly what this lab runs: a core device gets authoritative time (in production, from pool.ntp.org or an internal GPS appliance; in the lab, from its own calendar via `ntp master`), and everything else syncs to it internally. ``` ! R1 - the internal time source R1(config)#ntp master 3 ! R2, R3 - clients R2(config)#ntp server 1.1.1.1 ``` `ntp master 3` makes R1 serve its local clock as if it were a stratum 3 source. Why 3 and not 1? Convention and humility: if R1 ever gets a real upstream source, genuine stratum 1 and 2 servers should win the comparison. R1 confirms it is serving time: ``` R1#show ntp status | include Clock|stratum Clock is synchronized, stratum 3, reference is 127.127.1.1 ``` That reference address 127.127.1.1 is not a routable host: it is NTP's internal notation for "my own local clock" (the 127.127.x.x block identifies driver types, and .1 is the local clock driver). ## Watching a Client Converge (the Part Nobody Shows You) Most tutorials show the final synced state and skip the several minutes in between, which is exactly when people panic and start deleting configuration. Here is R2 immediately after `ntp server 1.1.1.1` was configured: ``` R2#show ntp associations address ref clock st when poll reach delay offset disp ~1.1.1.1 .TIME. 16 59 64 0 0.000 0.000 15937. * sys.peer, # selected, + candidate, - outlyer, x falseticker, ~ configured R2#show ntp status Clock is unsynchronized, stratum 16, no reference clock reference time is 00000000.00000000 (00:00:00.000 UTC Mon Jan 1 1900) ``` Stratum 16, reach 0, reference time in the year 1900\. Nothing is broken; NTP just has not exchanged enough polls yet. The column that tells the real story is `reach`. It is an 8-bit shift register displayed in octal: every poll cycle (64 seconds here, per the `poll` column) shifts in a 1 for a successful exchange. A few cycles later: ``` R2#show ntp associations address ref clock st when poll reach delay offset disp *~1.1.1.1 127.127.1.1 3 14 64 17 2.000 0.000 3.017 ``` Reach 17 (octal) means the last four polls all succeeded. The `*` in front of the address is the important promotion: this peer is now the `sys.peer`, the chosen synchronization source. And after the register fills completely: ``` R2#show ntp associations address ref clock st when poll reach delay offset disp *~1.1.1.1 127.127.1.1 3 65 64 377 3.000 -0.500 3.334 R2#show clock detail *23:05:33.109 UTC Fri Jul 10 2026 Time source is NTP ``` Reach 377 is the healthy steady state: all eight of the last eight polls answered. Delay is 3 ms, offset half a millisecond. `show clock detail` confirms the clock is being driven by NTP. One honest platform note: on this virtual lab router the `show ntp status` header was slow to flip from "unsynchronized" even with reach at 377 and a selected sys.peer, so judge synchronization by the associations line and `show clock detail`, not by the status header alone (on hardware the header catches up within a few poll cycles). **reach 0**no successful polls yet; brand new or unreachable server **reach 17**last 4 polls good; converging, give it a few minutes **reach 377**8 of 8 polls good; healthy steady state **reach 376, 254...**a zero bit in the register; intermittent loss to the server, investigate the path ## ntp server vs ntp master `ntp server
` makes the router a client of that address (and, automatically, a server to anyone downstream once it syncs). `ntp master [stratum]` makes the router serve its own local clock as authoritative even with no upstream source. Use `ntp master` only on the designated internal source (or in labs); if several routers all claim mastership, your network happily synchronizes to multiple disagreeing "truths." In production the core device typically has both: `ntp server` lines pointing at external sources, and its role as internal server flows from that without any `ntp master` at all. ## Production Polish: Source Interface and the Calendar Two small settings separate lab NTP from production NTP. First, source your NTP traffic from a loopback so your time infrastructure does not depend on any single physical link, and so server-side access lists see one stable address per device: ``` ntp source Loopback0 ``` Second, on platforms with a battery-backed hardware calendar, `ntp update-calendar` writes the disciplined NTP time back to the hardware clock periodically. That way a reload starts from approximately correct time instead of a factory default, which matters because NTP itself refuses to correct offsets larger than about 1000 seconds (it assumes something is badly wrong and waits for a human). A router that boots thinking it is the year 2002 will sit unsynchronized forever until someone sets the clock manually with `clock set` and lets NTP take over from there. ## Hardening: Authentication Exists, Use It NTP will happily accept time from anyone you point it at, and skewing a device's clock is a real attack (expired certificates suddenly validate, log timelines get scrambled). IOS XE supports MD5/SHA authentication of NTP associations: ``` ntp authenticate ntp authentication-key 1 sha1 ntp trusted-key 1 ntp server 1.1.1.1 key 1 ``` Configure the same key on the server side and the association only forms when the keys match. Also consider `ntp access-group` to control which devices may sync from you at all. ## Key Takeaways NTP is two commands (`ntp master` on the source, `ntp server` on clients) and one discipline: knowing how to verify it. Stratum measures distance from the reference clock, with 16 meaning "not synchronized at all." Convergence takes minutes, not seconds; read the `reach` octal register and look for the `*` sys.peer marker instead of restarting things. Reach 377 plus `Time source is NTP` in `show clock detail` is your healthy state. Authenticate NTP in production, because everything from certificates to your syslog timeline (see the [syslog deep dive](https://www.pinglabz.com/syslog-cisco-ios-xe/)) depends on this one service being right. The rest of the services stack lives in the [IP Services complete guide](https://www.pinglabz.com/ip-services/). ### NAT on Cisco IOS XE: Static, Dynamic, and PAT Configuration URL: https://www.pinglabz.com/nat-cisco-ios-xe/ Last updated: 2026-08-01T19:35:45.000Z NAT is the reason your private 10.x network can talk to the internet, and it is also one of the CCNA topics where the terminology trips up more people than the configuration. Cisco IOS XE gives you three flavors: static NAT (one-to-one, permanent), dynamic NAT (one-to-one from a pool, on demand), and PAT (many-to-one with port rewriting, the one your home router uses). This article configures all three on a real router and reads the translation table after each one, as part of the [IP Services complete guide](https://www.pinglabz.com/ip-services/). Everything below was captured live from a Cisco IOS XE lab in CML: R1 is the NAT boundary router, the 10.0.20.0/24 LAN behind it is the "inside," and a Linux host at 192.168.99.100 plays the internet on the "outside." If you want the ASA version of this story, the [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/) covers NAT on that platform (where the logic is object-based and quite different). ## The Four Address Terms That Explain Everything Every NAT translation entry has four columns, and once you can read them, every `show ip nat translations` output makes sense. Here is a real entry from the static NAT demo later in this article: ``` R1#show ip nat translations Pro Inside global Inside local Outside local Outside global icmp 192.168.99.50:45 10.0.20.10:45 192.168.99.100:45 192.168.99.100:45 --- 192.168.99.50 10.0.20.10 --- --- ``` **Inside local** The real address of the inside host, as seen on the inside. Here: 10.0.20.10\. This is what the host actually has configured. **Inside global** The translated address of the inside host, as seen from the outside. Here: 192.168.99.50\. This is what the internet thinks the host is. **Outside local** The address of the outside host, as seen from the inside. Unless you are doing destination NAT, it matches outside global. **Outside global** The real address of the outside host on the outside network. Here: 192.168.99.100, the Linux host we pinged. The mnemonic that works: "inside/outside" is *where the host lives*, "local/global" is *where you are standing when you look at it*. Local is the inside view, global is the outside view. ## The Common Foundation: Interface Roles and the ACL All three NAT types share two pieces of configuration. First, you tell IOS which interfaces face which realm: ``` interface Ethernet0/0 description TO-OUTSIDE ip nat outside ! interface Ethernet0/1 description TO-INSIDE ip nat inside ``` Second, for dynamic NAT and PAT, a standard ACL defines *which source addresses are allowed to be translated* (the ACL here is a matching tool, not a security filter): ``` ip access-list standard NAT-INSIDE permit 10.0.20.0 0.0.0.255 ``` Scope this ACL tightly. In this lab we learned that the hard way: with a broader ACL, even the switch management traffic got translated and replies started arriving from the wrong address (the session to the switch just died). Only match the subnets that genuinely need translation. ## Static NAT: One-to-One, Both Directions Static NAT maps one inside address to one outside address, permanently. One command: ``` R1(config)#ip nat inside source static 10.0.20.10 192.168.99.50 ``` The killer feature of static NAT is that it works in both directions. Because the mapping always exists, an *outside* host can initiate a connection to the inside host through its global address. Here is the outside Linux host pinging 192.168.99.50 (an address that exists only in R1's NAT table): ``` j@llmbits:~$ ping -c 3 192.168.99.50 3 packets transmitted, 3 received, 0% packet loss, time 2003ms rtt min/avg/max/mdev = 4.397/5.236/6.237/0.759 ms ``` And the translation table on R1 as those pings flow: ``` R1#show ip nat translations Pro Inside global Inside local Outside local Outside global icmp 192.168.99.50:45 10.0.20.10:45 192.168.99.100:45 192.168.99.100:45 --- 192.168.99.50 10.0.20.10 --- --- ``` Two entries: the bottom one (with `---` in the outside columns) is the permanent static mapping, and the top one is the live ICMP flow riding on it. This is the pattern you use for servers that must be reachable from outside (a web server in a DMZ, a mail relay, anything with a published address). ## Dynamic NAT: A Pool of One-to-One Leases Dynamic NAT keeps the one-to-one idea but assigns global addresses from a pool, on demand, only while traffic is flowing: ``` R1(config)#ip nat pool PUBLIC-POOL 192.168.99.60 192.168.99.69 netmask 255.255.255.0 R1(config)#ip nat inside source list NAT-INSIDE pool PUBLIC-POOL ``` Now an inside host (10.0.20.1 in this test) sends traffic outward: ``` R2#ping 192.168.99.100 source Ethernet0/0 repeat 10 Sending 10, 100-byte ICMP Echos to 192.168.99.100, timeout is 2 seconds: Packet sent with a source address of 10.0.20.1 !!!!!!!!!! Success rate is 100 percent (10/10), round-trip min/avg/max = 2/2/4 ms ``` R1 grabs the first free pool address (.60) and builds the mapping: ``` R1#show ip nat translations Pro Inside global Inside local Outside local Outside global icmp 192.168.99.60:1 10.0.20.1:1 192.168.99.100:1 192.168.99.100:1 --- 192.168.99.60 10.0.20.1 --- --- ``` `show ip nat statistics` is where dynamic NAT gets operationally interesting, because it shows pool utilization: ``` R1#show ip nat statistics Total active translations: 2 (0 static, 2 dynamic; 1 extended) Outside interfaces: Ethernet0/0 Inside interfaces: Ethernet0/1, Ethernet0/2 Hits: 46 Misses: 0 Dynamic mappings: -- Inside Source [Id: 2] access-list NAT-INSIDE pool PUBLIC-POOL refcount 2 pool PUBLIC-POOL: id 1, netmask 255.255.255.0 start 192.168.99.60 end 192.168.99.69 type generic, total addresses 10, allocated 1 (10%), misses 0 ``` Watch two numbers in production: `allocated` (percentage of the pool in use) and `misses`. A miss means a host wanted a translation and the pool was empty, which means its traffic went nowhere. That is dynamic NAT's structural weakness: ten pool addresses means ten concurrent inside hosts, full stop. It is why pure dynamic NAT is rare today and PAT does the real work. ## PAT: The Whole Network Behind One Address PAT (Port Address Translation, also called NAT overload) multiplexes many inside hosts onto one global address by rewriting source ports. The typical form borrows the outside interface's own address, so you do not even need a pool: ``` R1(config)#ip nat inside source list NAT-INSIDE interface Ethernet0/0 overload ``` Here is the translation table with an ICMP flow and an SSH session (TCP 22) from the same inside host active at once: ``` R1#show ip nat translations Pro Inside global Inside local Outside local Outside global icmp 192.168.99.1:1024 10.0.20.1:2 192.168.99.100:2 192.168.99.100:1024 tcp 192.168.99.1:4790 10.0.20.1:28674 192.168.99.100:22 192.168.99.100:22 ``` Read the TCP line: the inside host's random source port 28674 became 4790 on the global address, headed to port 22 on the destination. The ICMP line shows the same trick applied to the ICMP identifier (rewritten to 1024). The port (or identifier) is what lets R1 demultiplex return traffic to thousands of inside hosts sharing a single address. Note there is no `---` permanent entry this time: PAT entries exist only while flows are alive, and ICMP entries in particular age out within about a minute (run your ping and your `show` command close together when you lab this). ## Static PAT: Port Forwarding, Cisco Style There is a fourth variant hiding between static NAT and PAT that answers a very common requirement: publish *one service* on an inside host without spending a whole global address on it. Static PAT (port forwarding) maps protocol + port instead of the entire address: ``` R1(config)#ip nat inside source static tcp 10.0.20.10 443 192.168.99.1 8443 ``` That says: anything arriving at the outside interface's own address on TCP 8443 gets handed to 10.0.20.10 on TCP 443\. The router's address keeps doing PAT for everyone's outbound traffic at the same time; only that one inbound port is claimed. This is exactly what "port forwarding" on a home router does, and on the exam it shows up as static NAT "with the protocol keyword." Stack as many of these as you have services, all sharing one global address. ## Where NAT Happens in the Packet Path One fact explains most weird NAT symptoms: on the inside-to-outside path, IOS routes *first*, then translates; outside-to-inside, it translates *first*, then routes. Two practical consequences. First, the router needs a route for the *untranslated* destination before NAT ever gets a look, so broken routing shows up as "NAT not working." Second, inbound ACLs on the outside interface are evaluated against the *global* (pre-untranslation) addresses, so you write them in terms of the public addresses, not the inside hosts. When a translation mystery survives the `show` commands, `debug ip nat` plus a test ping shows each rewrite as it happens. ## Verification, Cleanup, and a Gotcha **show ip nat translations**the live translation table, the first thing you check **show ip nat statistics**interface roles, hit/miss counters, pool utilization **clear ip nat translation \***flush all dynamic entries (static mappings stay) **debug ip nat**per-packet translation events; use with care on busy boxes The gotcha: if you try to remove a NAT rule while translations built from it are still active, IOS refuses with `%Dynamic mapping in use, do you want to delete all entries?` and waits for confirmation. The clean sequence is always `clear ip nat translation *` first, then remove the rule. If you are doing this through an automation tool that cannot answer interactive prompts, that confirmation will hang your job. The other outcome of that first command is the common one: nothing at all. The config reads correctly, `show ip nat translations` comes back empty, and the inside subnet has no internet - [starting the diagnosis from an empty translation table](https://www.pinglabz.com/nat-pat-troubleshooting-cisco-ios-xe/) is the companion to this guide and begins at exactly that moment. ## Which NAT When **Static NAT** Servers that outside hosts must reach. One-to-one, permanent, bidirectional. Costs one global address per server. **Dynamic NAT** Legacy or compliance cases needing one-to-one without fixed mappings. Limited by pool size; mostly historical. **PAT / overload** Everything else. Whole networks behind one address, outbound-initiated traffic. The default answer. In real networks the usual combination is PAT for general outbound traffic plus a handful of static entries (full or port-level) for published services. And remember why any of this exists: the inside addresses are [RFC 1918 private space](https://www.pinglabz.com/private-ipv4-rfc-1918/), which is not routable on the internet, so something at the edge has to swap it for an address that is. Also know what NAT costs you. It breaks true end-to-end reachability (two hosts behind different PATs cannot open connections to each other without a rendezvous service), it complicates protocols that embed addresses in their payload (classic offenders: active-mode FTP, SIP), and every translation is state the router must hold. None of that is a reason to avoid it; it is the reason IPv6 was designed to make it unnecessary, and why the [IPv6 guide](https://www.pinglabz.com/ipv6/) reads so differently on this topic. ## Key Takeaways NAT on IOS XE is three commands deep once the interface roles (`ip nat inside` / `ip nat outside`) and the match ACL are in place. Static NAT is one-to-one and works for outside-initiated connections; dynamic NAT leases one-to-one mappings from a finite pool; PAT rewrites ports to share one address and is what production networks actually run. Read translation tables with the local/global mindset (local = inside view, global = outside view), scope the NAT ACL tightly so management traffic does not get swallowed, and clear translations before you remove rules. For the rest of the day-2 services stack (DHCP, NTP, SNMP, syslog), continue with the [IP Services complete guide](https://www.pinglabz.com/ip-services/). ### AI and Machine Learning in Network Operations URL: https://www.pinglabz.com/ai-machine-learning-network-operations/ Last updated: 2026-07-10T23:16:45.000Z The CCNA v1.1 refresh added artificial intelligence and machine learning to the blueprint, and it is one of the few topics where the competition has not caught up yet. You are not expected to build models; you are expected to understand what AI and ML actually do in network operations, where they help, and where they emphatically do not replace an engineer who understands the fundamentals. This article covers the learning types, the predictive-versus-generative split, real network use cases, and the honest limits, framed for the exam and for the job. It is part of the [Network Automation guide](https://www.pinglabz.com/network-automation/). ## The Three Learning Types Machine learning is software that improves at a task by finding patterns in data rather than following rules a human wrote. The CCNA-level distinction is between three ways a model learns. **Supervised** Trained on labeled examples (input plus the known correct answer). Learns to predict the label for new inputs. Example: classify traffic as normal or malicious from labeled samples. **Unsupervised** Trained on unlabeled data; finds structure on its own. Example: baseline "normal" traffic and flag anything that deviates, without being told what an anomaly looks like. **Reinforcement** Learns by trial and reward, optimizing actions against feedback over time. Example: tuning routing or radio parameters toward a performance goal. Anomaly detection, the most common networking use, leans on unsupervised learning, because nobody can label every possible failure in advance. The model learns what your network normally looks like and raises a hand when reality diverges. ## Predictive vs Generative AI The other split worth knowing separates the two families of AI you will hear about. **Predictive AI** forecasts or classifies: it looks at data and tells you what will happen or what category something belongs to (this link will saturate by Friday; this flow is anomalous). **Generative AI** produces new content: text, code, or configuration (draft me an ACL that does X; summarize this outage). Predictive AI has quietly run network analytics for years; generative AI is the newer arrival, and it is the one that can write a config for you, with all the caveats that implies. ## Where AI Actually Helps in Network Ops The exam wants concrete use cases, not hype. These are the ones grounded in real products and real value. Anomaly detectionBaseline normal behavior from telemetry, flag deviations (a sudden traffic spike, an interface erroring in a new pattern) before a human would notice. Predictive capacity and failureForecast when a link or device will hit its limit, or predict hardware failure from degrading metrics, so you fix it on your schedule, not at 3 a.m. AIOps event correlationCollapse a storm of alerts into one likely root cause, cutting the noise that buries the real signal during an incident. AI-assisted configurationGenerative AI drafts config or explains an existing one. Fast, but must be reviewed and tested before it touches production (more on that below). Cisco ships these as features rather than science projects. Catalyst Center uses analytics and AI for assurance and anomaly detection across the campus; Meraki surfaces AI-driven insights in its cloud dashboard. When the exam asks for a vendor example, those are the two to name. The [Catalyst Center overview](https://www.pinglabz.com/what-is-cisco-catalyst-center/) covers the assurance angle in more depth. ## Telemetry Is the Fuel None of this works without data, and lots of it. Traditional polling with SNMP (covered as a legacy monitoring approach in the [IP Services guide](https://www.pinglabz.com/ip-services/)) asks each device for counters every so often, which is too slow and too coarse to feed a model well. **Model-driven telemetry** flips it: devices *push* structured, high-frequency data (structured by YANG models) to a collector continuously. That firehose of current, structured data is what makes real-time anomaly detection and prediction possible. The progression from SNMP polling to streaming telemetry is the infrastructure that made AIOps practical, and it is why telemetry sits next to AI on the modern blueprint. ## What AI Does Not Replace Here is the part that matters most for a working engineer, and the part this site is built around. AI can draft a configuration, but it cannot be trusted to apply one unreviewed. A generated ACL that looks plausible can silently drop the wrong traffic; a suggested route change can black-hole a subnet. The rule is **verify before you apply**: treat AI output as a fast first draft from a confident junior engineer, not as an authority. You still need to read the config, understand what it does, and test it, which you can only do if you understand routing, switching, and security yourself. That is the throughline of PingLabz: every article is grounded in real device output because the details are where operations lives, and AI does not change that. It changes how fast you produce a draft and how early you spot an anomaly. It does not change the need to understand the packet. An engineer who understands the fundamentals and uses AI to move faster beats both the engineer who refuses the tools and the one who trusts them blindly. ## How to Prepare as an Engineer The practical stance is neither dismissal nor hype. Learn the fundamentals cold, because they are what let you judge AI output. Get comfortable with telemetry and the data side, because that is the fuel. Use generative AI as a drafting and explanation aid, then verify everything it produces against your own understanding and a test environment. The skill that appreciates in value is not "prompt an AI"; it is "know enough to catch when the AI is wrong." ## Machine Learning, AI, and Where the Terms Overlap The exam uses "AI" and "machine learning" almost interchangeably, but they are not the same thing, and the distinction is worth a sentence. Artificial intelligence is the broad goal of getting machines to do things that would need human intelligence. Machine learning is one way to get there: instead of a human writing explicit rules, the system learns patterns from data. Deep learning is a subset of machine learning that uses layered neural networks and powers most of the recent generative advances. For CCNA purposes you can treat ML as the engine and AI as the umbrella, and know that when a product says "AI-powered anomaly detection," what is under the hood is almost always a machine-learning model trained on network data. This matters for a practical reason: a model is only as good as its training data. A model trained on last quarter's traffic will not recognize a genuinely new pattern, and a model fed noisy or biased data produces confident nonsense. That is another argument for the verify-before-apply rule, and another reason the engineer who understands the network stays in the loop. ## What Actually Changes in Your Day Concretely, AI shifts a few parts of the job without eliminating any of them. Alert triage gets faster because correlation collapses a hundred symptoms into one probable cause. Capacity planning gets earlier because forecasts flag the trend before the outage. First drafts of config and scripts arrive in seconds instead of minutes. What does not change: you still own the decision to apply, you still troubleshoot with ping, traceroute, and the routing table, and you still carry responsibility for what the network does. The tooling moved; the accountability did not. ## FAQ ### What is the difference between supervised and unsupervised learning? Supervised learning trains on labeled data (examples with known answers) and predicts labels for new input. Unsupervised learning trains on unlabeled data and finds structure on its own, which is why it suits anomaly detection where you can't label every failure in advance. ### What is the difference between predictive and generative AI? Predictive AI forecasts or classifies based on existing data (what will happen, what category this is). Generative AI creates new content such as text, code, or configuration. Network analytics is mostly predictive; AI-assisted config generation is generative. ### Will AI replace network engineers? No. AI accelerates drafting and detection, but it cannot be trusted to apply changes unreviewed, and it does not understand your specific network's constraints. It raises the value of engineers who know the fundamentals well enough to verify its output. ### Why does telemetry matter for AI? AI models need large amounts of current, structured data. Model-driven telemetry streams that data continuously from devices, replacing slow SNMP polling and making real-time anomaly detection and prediction feasible. ## Key Takeaways Know the three learning types (supervised, unsupervised, reinforcement) and that anomaly detection leans on unsupervised learning. Separate predictive AI (forecast/classify) from generative AI (create content). The real use cases are anomaly detection, predictive capacity and failure, AIOps event correlation, and AI-assisted config, with Catalyst Center and Meraki as Cisco examples. Model-driven telemetry is the fuel. Above all, verify before you apply: AI is a fast draft, not an authority, and it never removes the need to understand the fundamentals. Return to the [Network Automation guide](https://www.pinglabz.com/network-automation/) for the rest of domain 6. ### Ansible vs Puppet vs Chef (and Terraform) for Network Automation URL: https://www.pinglabz.com/ansible-vs-puppet-vs-chef/ Last updated: 2026-08-01T19:32:42.000Z Once you accept that configuring devices one at a time does not scale, the next question is which tool does the configuring for you. The CCNA v1.1 blueprint names three configuration-management tools (Ansible, Puppet, Chef) and, with the move toward infrastructure as code, Terraform belongs in the conversation too. This article compares all four, then shows Ansible actually running against three live Cisco IOS XE routers, including the idempotency gotcha that bites everyone the first time. It is part of the [Network Automation guide](https://www.pinglabz.com/network-automation/). ## The Four Axes That Separate Them The tools differ along a few dimensions, and once you know the axes, the comparison writes itself. **Agent vs agentless** Does the managed device need software installed on it? Network gear usually can't run agents, so this axis decides everything. **Push vs pull** Does a central server push changes out, or do devices pull their config from a master on a schedule? **Declarative vs procedural** Do you describe the desired end state, or the steps to get there? Declarative tools converge to a state; procedural tools run steps. **Language** YAML (Ansible), a Ruby-based DSL (Puppet, Chef), or HCL (Terraform). Affects how steep the learning curve is. ## The Four Tools, Compared **Ansible** Agentless, push, mostly declarative, YAML playbooks. Connects over SSH (or API). The default choice for network automation precisely because it needs nothing on the device. **Puppet** Agent-based, pull, declarative, Ruby DSL. Nodes pull their catalog from a Puppet master on a schedule. Strong for large fleets of servers. **Chef** Agent-based, pull, procedural, Ruby DSL. You write "recipes" and "cookbooks." Very flexible, steeper learning curve, developer-oriented. **Terraform** Agentless, declarative, HCL, provider model with a state file. Built for provisioning infrastructure (cloud, and increasingly network devices) rather than in-place config. For network devices, the agent axis is decisive. A traditional switch cannot run a Puppet or Chef agent, so those tools reach network gear only through proxy modes or API integrations. Ansible, being agentless and speaking SSH natively, connects to a router the same way you do. That is why it dominates network automation, and why it is the tool worth demonstrating on real gear. There is a fifth option that does not sit neatly on those four axes: Nornir keeps the inventory-and-parallelism model but has you write the tasks in Python rather than YAML, with NAPALM doing the talking to the device. If you would rather stay in Python than learn a DSL, [using Nornir and NAPALM to manage Cisco configuration](https://www.pinglabz.com/nornir-napalm-cisco-config-management/) shows it generating a real config diff on IOS XE before it commits anything. ## Ansible in Practice: The Inventory Ansible needs to know what to manage and how to reach it. That lives in an inventory file. Here is the real one used to drive three lab routers from a Linux host: ``` [routers] r1 ansible_host=192.168.99.1 r2 ansible_host=10.0.12.2 r3 ansible_host=10.0.13.2 [routers:vars] ansible_network_os=cisco.ios.ios ansible_connection=ansible.netcommon.network_cli ansible_user=cisco ansible_password=Cisco@123 ``` The `[routers]` group lists three hosts; the `[routers:vars]` block tells Ansible to treat them as Cisco IOS over the network CLI connection (SSH under the hood, using the `cisco.ios` collection). No agent, no software on the routers, just credentials and a connection type. ## Ad-Hoc: One Command, Three Routers Before writing a playbook, an ad-hoc command proves the connection and shows the appeal of "do this everywhere at once." This runs a show command against all three routers in parallel: ``` j@llmbits:~/anslab$ ansible routers -i inventory.ini -m cisco.ios.ios_command \ -a "commands='show version | include uptime'" r1 | SUCCESS => { "changed": false, "stdout": ["R1 uptime is 21 minutes"] } r2 | SUCCESS => { "changed": false, "stdout": ["R2 uptime is 21 minutes"] } r3 | SUCCESS => { "changed": false, "stdout": ["R3 uptime is 21 minutes"] } ``` Three devices, one command, real output. Note `"changed": false`: a show command reads state without altering it, so Ansible correctly reports nothing changed. That `changed` flag is the heart of the next lesson. ## The Playbook and Idempotency A playbook describes a desired state in YAML. This one standardizes logging and SNMP contact across all three routers: ``` --- - name: Standardize logging and SNMP contact on all routers hosts: routers gather_facts: false tasks: - name: Ensure syslog host and trap level are set cisco.ios.ios_config: lines: - logging host 192.168.99.100 - snmp-server contact noc@pinglabz.com ``` Run it once and, as expected, it applies the config to every router: ``` j@llmbits:~/anslab$ ansible-playbook -i inventory.ini set_logging.yml changed: [r1] changed: [r2] changed: [r3] PLAY RECAP **************************************************************** r1 : ok=1 changed=1 unreachable=0 failed=0 r2 : ok=1 changed=1 unreachable=0 failed=0 r3 : ok=1 changed=1 unreachable=0 failed=0 ``` **Idempotency** is the promise that running the same playbook again changes nothing, because the desired state is already met. It is the single most important property of a good automation tool: you can run it repeatedly and safely, and `changed=0` tells you the network already matches your intent. So the second run should report `changed=0` everywhere. ## The Gotcha That Bites Everyone Except, on the first attempt, it did not. The playbook originally also pushed `logging trap informational`, and every single run reported `changed=1`, forever. The reason is a genuine trap in network automation: `logging trap informational` is the IOS *default*, so it never appears in the running-config. Ansible's `ios_config` checks whether each line is present in the running-config, does not find it (because defaults are hidden), and dutifully re-pushes it every run. The config is correct; the tool just can never confirm it, so it never reaches a stable state. The fix is to stop managing lines that match a hidden default. Removing the `logging trap informational` line and rerunning gives the idempotent result you actually want: ``` PLAY RECAP **************************************************************** r1 : ok=1 changed=0 unreachable=0 failed=0 r2 : ok=1 changed=0 unreachable=0 failed=0 r3 : ok=1 changed=0 unreachable=0 failed=0 ``` Now the playbook is truly idempotent: the desired state is met, and Ansible confirms it with `changed=0`. The lesson generalizes: when a playbook reports "changed" on every run, suspect that you are managing a line the device does not store because it matches a default. This is exactly the kind of behavior you only learn by running the tool against real devices, not by reading about it. ## Where Terraform Fits Ansible, Puppet, and Chef manage the configuration *inside* devices that already exist. Terraform's angle is different: it provisions the infrastructure itself from declarative code (infrastructure as code), tracking what it created in a **state file** so it knows the difference between "make this" and "already made this." Its **provider** model means the same HCL workflow that spins up cloud resources can, with a network provider, provision devices and services. For the CCNA v1.1 era, the point to hold onto is that Terraform represents the declarative, state-tracked, provider-driven style of automation, and it increasingly overlaps with the network. You describe the end state; Terraform computes and applies the diff. ## FAQ ### Which tool should I learn first for networking? Ansible. It is agentless (nothing to install on the router), speaks SSH, uses readable YAML, and has mature Cisco collections. It is the tool you are most likely to use on the job and the one the CCNA emphasizes. If the real question is which Python library to reach for rather than which configuration-management framework, that is a separate decision - [which Python automation library to use for which job](https://www.pinglabz.com/netmiko-vs-napalm-vs-pyats/) splits the three main ones by the layer each actually operates at. ### Why does agentless matter so much for network devices? Most switches and routers can't run third-party agent software. A tool that requires an agent (Puppet, Chef) can only reach them through workarounds, while an agentless tool (Ansible, Terraform) connects over SSH or an API the device already offers. ### What does idempotent mean in one sentence? Running the operation again produces no additional change because the system is already in the desired state, so it is safe to run repeatedly. ### Is Terraform a replacement for Ansible? No, they solve different problems. Terraform provisions and tracks infrastructure from declarative code; Ansible configures and manages the software and settings on infrastructure that exists. Many shops use both. ## Key Takeaways Compare configuration-management tools on four axes: agent vs agentless, push vs pull, declarative vs procedural, and language. Ansible wins for network devices because it is agentless and speaks SSH. Idempotency (a rerun changes nothing) is the property you want, and the classic gotcha is that `ios_config` re-pushes lines matching hidden IOS defaults forever, reporting `changed=1` until you stop managing them. Terraform brings declarative, state-tracked infrastructure as code into the picture for the v1.1 era. Try it hands-on in the [Automation Labs](https://www.pinglabz.com/automation-labs/), and see the [Network Automation guide](https://www.pinglabz.com/network-automation/) for the full domain-6 path. ### REST APIs and JSON for Network Engineers URL: https://www.pinglabz.com/rest-apis-json-network-engineers/ Last updated: 2026-07-10T23:16:44.000Z The command line is a conversation for humans. REST APIs are the conversation for programs, and once a controller sits in front of your network (see [Controller-Based Networking and SDN](https://www.pinglabz.com/controller-based-networking-sdn/)), REST is how your scripts, dashboards, and automation tools ask it to do things. CCNA domain 6 expects you to understand REST and the JSON it speaks. This article covers the HTTP verbs and how they map to actions, status codes, authentication, and JSON structure, all grounded in real requests made from a Linux host against a live Cisco Modeling Labs controller. It is part of the [Network Automation guide](https://www.pinglabz.com/network-automation/). ## What REST Actually Is REST (Representational State Transfer) is an architectural style for APIs that runs over ordinary HTTP. You address a resource with a URL, choose an HTTP verb to say what you want to do to it, and the server responds with a status code and a body (almost always JSON). That is the whole idea: no special protocol, just HTTP used deliberately. Because it is HTTP, you can call it with `curl`, with Python's `requests`, or from any tool that can make a web request, which is exactly why it won as the northbound interface for controllers. ## HTTP Verbs Map to CRUD The four things you can do to data (Create, Read, Update, Delete) map cleanly onto four HTTP verbs. Learn this mapping; the exam tests it directly. **GET** Read. Retrieve a resource without changing it. Safe to repeat. "Show me the labs." **POST** Create. Make a new resource. "Create a new lab" or "authenticate me." **PUT / PATCH** Update. PUT replaces a resource; PATCH modifies part of it. "Rename this lab." **DELETE** Delete. Remove a resource. "Wipe this lab." ## Authentication: Basic vs Token An API you can call without credentials is an API anyone can call, so the first request is almost always about proving who you are. Two patterns dominate. **Basic authentication** sends a username and password with every request (over HTTPS, so it is encrypted in transit). **Token-based authentication** trades your credentials once for a token, then sends that token on every subsequent request; if the token leaks it can be scoped and expired, which is why it is preferred for automation. Here is a real token exchange against the Cisco Modeling Labs controller from a Linux host. A POST to the authenticate endpoint hands over the credentials and gets back a JSON Web Token (JWT): ``` j@llmbits:~$ curl -sk -X POST https://192.168.88.158/api/v0/authenticate \ -H "Content-Type: application/json" \ -d '{"username":"admin","password":"..."}' eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ...(JWT token) ``` That opaque string is the token. Every following request carries it in an `Authorization: Bearer` header, and the controller trusts it instead of re-checking the password. This is exactly the flow the automation labs on the site use, and the [Automation Labs](https://www.pinglabz.com/automation-labs/) (auto-04 and auto-05) walk through it hands-on. ## Reading JSON: An Array With the token in hand, a GET asks for a list of resources. The controller returns a JSON **array**, which is an ordered list wrapped in square brackets: ``` j@llmbits:~$ curl -sk https://192.168.88.158/api/v0/labs \ -H "Authorization: Bearer $TOKEN" ["181e4455-d9ad-49a9-9e88-956692eeb459","79c2e69f-75b3-4988-b566-ae641f6c2eb0", "4353e423-743b-4c62-8e98-5bd140468763","6929be76-3bd6-4b35-9fb3-79398408f0a4"] ``` Each element is a string (a lab ID). An array is how APIs hand you a collection: iterate over it in your script and act on each item. ## Reading JSON: An Object Ask for one specific resource and you get a JSON **object**: a set of key-value pairs in curly braces. Piping the response through `python3 -m json.tool` pretty-prints it so the structure is readable: ``` j@llmbits:~$ curl -sk https://192.168.88.158/api/v0/labs/6929be76-... \ -H "Authorization: Bearer $TOKEN" | python3 -m json.tool { "state": "STARTED", "created": "2026-07-10T22:37:00+00:00", "lab_title": "PingLabz CCNA Services Lab", "owner_username": "admin", "node_count": 5, "link_count": 6, "id": "6929be76-3bd6-4b35-9fb3-79398408f0a4", "groups": [] } ``` This one response contains every JSON data type the CCNA asks about, so it is worth annotating: **Object** The whole thing in `{ }`. An unordered set of key-value pairs. **String** `"STARTED"`, the lab title, the ID. Text in double quotes. **Number** `5` and `6` for node and link counts. No quotes. **Array** `[]`, here an empty list of groups. Would hold elements if populated. JSON also has **booleans** (`true`/`false`) and **null**. The rules that trip people up: keys are always strings in double quotes, strings use double quotes (never single), no trailing comma after the last pair, and whitespace is ignored (the pretty-printing above is purely for humans). Get one comma wrong and the whole document fails to parse, which is why you validate JSON before sending it. ## Status Codes Every response carries a numeric status code that tells your script what happened, grouped by first digit. 2xx Success200 OK (GET worked), 201 Created (POST made something), 204 No Content (DELETE worked, nothing to return). 4xx Client error400 Bad Request (your JSON was malformed), 401 Unauthorized (no or bad token), 403 Forbidden (authenticated but not allowed), 404 Not Found. 5xx Server error500 Internal Server Error and friends. The request was fine; the server failed. Not your fault, but your script still has to handle it. The 4xx codes are the ones you meet most while writing automation. A 401 usually means your token expired (re-authenticate); a 400 means you sent bad JSON (validate it); a 404 means you have the URL wrong. ## A Note on RESTCONF Everything above is a generic REST API on a controller. There is also a standards-based REST API that runs directly on network devices, called **RESTCONF**: it exposes the device's configuration and operational data as REST resources, structured by YANG models, and returns YANG-modeled JSON. It is the device-level cousin of the controller API. In full transparency, the IOL router images used in the lab here do not ship RESTCONF, which is why these captures target the CML controller's own REST API instead. On a real IOS XE box you would enable `restconf` and hit `/restconf/data/...` with the same verbs, headers, and JSON you have seen here. The concepts transfer exactly; only the endpoint changes. RESTCONF and its sibling NETCONF are where CCNP picks up the thread. ## Why REST Matters for Network Engineers You do not have to become a developer, but you do have to be able to read an API response and make a call. The moment your network has a controller, the CLI stops being the only interface, and the tasks that used to mean SSHing into fifty boxes become a loop over a JSON array. Being able to authenticate, GET a resource, read the JSON, and POST a change is the difference between automating your job and doing it by hand forever. The configuration-management tools in [Ansible vs Puppet vs Chef](https://www.pinglabz.com/ansible-vs-puppet-vs-chef/) are, underneath, making these same REST calls for you. ## FAQ ### What is the difference between REST and RESTCONF? REST is the general style (HTTP verbs, URLs, JSON) that any web API can use. RESTCONF is a specific standard that applies REST to network-device configuration, structuring the data with YANG models. RESTCONF is REST, specialized for devices. ### Does REST have to use JSON? No, REST can return XML too, and RESTCONF supports both. But JSON has become the default because it is lighter and maps directly onto the data structures in Python and JavaScript. The CCNA focuses on JSON. ### Why use a token instead of sending my password every time? A token can be scoped, expired, and revoked without changing your password, and it limits exposure if a single request is intercepted. Sending credentials on every call means every call carries your full secret. Tokens are the safer automation pattern. ### Do I need to know curl for the exam? You should be able to read a REST call and identify the verb, the resource, and the expected response. `curl` is just a convenient way to show that on one line; the concepts (verb, URL, headers, body, status code) are what matter. ## Key Takeaways REST is HTTP used deliberately: a verb (GET/POST/PUT/PATCH/DELETE) acting on a URL, returning a status code and JSON. The verbs map to CRUD. Authenticate first, prefer a token over sending credentials repeatedly, and read the status code before trusting the body. JSON is objects (`{}`), arrays (`[]`), strings, numbers, booleans, and null, with strict comma and quote rules. RESTCONF is the same idea aimed at devices via YANG. Put it into practice in the [Automation Labs](https://www.pinglabz.com/automation-labs/), and see the [Network Automation guide](https://www.pinglabz.com/network-automation/) for the full path. ### Controller-Based Networking and SDN: Overlay, Underlay, and Fabric URL: https://www.pinglabz.com/controller-based-networking-sdn/ Last updated: 2026-07-10T23:16:43.000Z For most of networking history, every router and switch was an island: it made its own forwarding decisions and you configured each one by hand. Controller-based networking breaks that model by pulling the decision-making into a central controller, and it is the conceptual heart of CCNA domain 6\. This article defines the control plane and data plane split, explains northbound and southbound APIs, untangles the overlay/underlay/fabric vocabulary, and maps it all to the Cisco products you will meet. It is part of the [Network Automation guide](https://www.pinglabz.com/network-automation/), and it deliberately keeps the routing and switching you already know at the center, because controllers do not replace that knowledge, they orchestrate it. ## Control Plane vs Data Plane Every network device does two distinct jobs. The **control plane** decides where traffic should go: it runs the routing protocols, builds the routing and MAC tables, and figures out the topology. The **data plane** (also called the forwarding plane) does the actual moving: it takes a frame or packet and pushes it out the right interface, as fast as the hardware allows, based on the tables the control plane built. In a traditional network, both planes live inside every device. Each router runs its own OSPF or BGP process, reaches its own conclusions, and forwards accordingly. It works, but the intelligence is scattered across hundreds of boxes, and a policy change means touching all of them. ## What a Controller Centralizes Software-Defined Networking (SDN) separates the control plane from the data plane and lifts the control-plane intelligence into a central **controller**. The devices keep their fast data planes, but they take forwarding instructions from the controller instead of each computing everything independently. The controller holds a complete, current view of the topology, and you express intent to it once ("these users may reach these resources") rather than configuring each device. That central view is the whole point. Because the controller sees everything, it can compute paths globally, push consistent policy, and give you a single place to automate against. The tradeoff is a new dependency: the controller becomes critical infrastructure, which is why real deployments run controllers in clusters. ## Northbound and Southbound APIs A controller sits in the middle of two conversations, and the CCNA names both directions. **Northbound API** Faces up, toward applications and automation. This is where you (or an app, or a script) tell the controller what you want. Almost always a REST API returning JSON. **Southbound API** Faces down, toward the devices. This is how the controller programs switches and routers. Examples: NETCONF, RESTCONF, gRPC, and historically OpenFlow. The mental model: your automation talks northbound to the controller in the language of intent; the controller translates that into device-specific instructions and pushes them southbound. You program the network once, through the top, instead of SSHing into every box. The northbound side is why [REST APIs and JSON](https://www.pinglabz.com/rest-apis-json-network-engineers/) are on the exam at all: they are the interface between your tools and the controller. ## Overlay, Underlay, and Fabric These three terms get used loosely, so pin them down. The **underlay** is the physical network and its IP routing: the actual cables, switches, and the IGP (usually OSPF or IS-IS) that gives every device reachability to every other. Its only job is to move packets between the fabric nodes. The **overlay** is a virtual network built on top of the underlay using tunnels. Traffic is encapsulated (wrapped in an extra header) at the ingress device, carried across the underlay as ordinary IP, and decapsulated at the egress. The overlay is where segmentation and policy live; the underlay just carries the encapsulated packets and never needs to know about them. The **fabric** is the combination: the underlay plus the overlay plus the control plane that ties them together, managed as a single system. A concrete example makes it click: UnderlayLeaf and spine switches running an IP IGP so every switch can reach every other switch's loopback. OverlayVXLAN tunnels carrying tenant traffic between leaves, encapsulated in UDP over the underlay. FabricThe whole thing, with a control plane (often BGP EVPN) advertising which host lives behind which leaf. VXLAN over an IP underlay is the canonical overlay example, and it is exactly the pattern behind Cisco's fabric products. If the encapsulation idea is new, the [VLAN vs VXLAN article](https://www.pinglabz.com/vlan-vs-vxlan/) walks through why an overlay solves problems a plain VLAN cannot. ## Mapping to Cisco Solutions Cisco packages controller-based networking into products aimed at different parts of the network, and the exam expects you to know which is which. **Catalyst Center + SD-Access** The campus fabric. Catalyst Center is the controller; SD-Access is the fabric (VXLAN overlay, LISP control plane, TrustSec policy). See [What Is Cisco Catalyst Center](https://www.pinglabz.com/what-is-cisco-catalyst-center/). **Catalyst SD-WAN** The WAN fabric (formerly Viptela). A controller-based overlay across internet and MPLS transports. See the [SD-WAN guide](https://www.pinglabz.com/sd-wan/). **Meraki** Cloud-managed networking. The controller lives in Cisco's cloud; devices phone home for policy and reporting. ## Traditional vs Controller-Based, Side by Side **Configuration** Traditional: box by box, over the CLI. Controller-based: once, expressed as intent to the controller. **Topology view** Traditional: each device sees its neighbors. Controller-based: the controller sees the whole network. **Policy consistency** Traditional: only as consistent as the human. Controller-based: enforced uniformly from one source of truth. **Failure domain** Traditional: a device failure is local. Controller-based: the controller is critical, so it runs clustered. ## What Stays the Same Here is the reassuring part, and the part the exam quietly rewards. A controller does not invent new physics. Underneath SD-Access is still IP routing, still switching, still the same forwarding you learned in domains 2 and 3\. The overlay rides on an underlay that runs an ordinary IGP. When a fabric misbehaves, you troubleshoot it with the same tools ([ping](https://www.pinglabz.com/ping/), traceroute, reading the routing table) plus the controller's view. Controller-based networking changes how you configure and where the intelligence lives; it does not let you skip the fundamentals. If anything, it raises the bar, because now you need to understand both the abstraction and the plumbing under it. ## FAQ ### What is SDN in one sentence? Software-Defined Networking separates the control plane from the data plane and centralizes the control-plane intelligence in a controller that programs the devices through APIs. ### How do I remember northbound vs southbound? Picture the controller in the middle. Applications and your automation are "above" it (northbound); the network devices are "below" it (southbound). You talk down to the controller from your tools (north), the controller talks down to the devices (south). ### What is the difference between an overlay and an underlay? The underlay is the real physical network and its IP routing. The overlay is a virtual network of tunnels built on top of it, where segmentation and policy live. The underlay just carries the encapsulated overlay traffic. ### Do I still need to learn routing and switching? Absolutely. Controllers orchestrate routing and switching; they do not replace it. The underlay is an ordinary IP network, and troubleshooting a fabric requires understanding both the overlay abstraction and the plumbing beneath it. ## Key Takeaways Controller-based networking separates control from data plane and centralizes the intelligence. You program the controller through a northbound REST API; the controller programs devices southbound with NETCONF, RESTCONF, or similar. Overlay (virtual tunnels) rides on underlay (physical IP network), and the fabric is the managed whole, with VXLAN over IP as the classic example. Cisco maps this to Catalyst Center/SD-Access for campus and Catalyst SD-WAN for the WAN. Through all of it, the routing and switching fundamentals still run underneath. Next, see how you actually talk to a controller in [REST APIs and JSON for Network Engineers](https://www.pinglabz.com/rest-apis-json-network-engineers/), and return to the [Network Automation guide](https://www.pinglabz.com/network-automation/) for the full domain-6 path. ### Password Policies and MFA for Network Devices URL: https://www.pinglabz.com/password-policies-mfa-network-devices/ Last updated: 2026-07-10T23:16:42.000Z Every network device ships with the same weakness: a login prompt that will accept a guessed password forever unless you tell it not to. Hardening that prompt is a CCNA domain 5 objective and a real-world necessity, and it is more than "set a strong password." This article covers the IOS password-type hierarchy, how to enforce a minimum length, how to lock out brute-force attempts, and where multi-factor authentication fits, all grounded in real Cisco IOS XE output (including an actual brute-force attempt hitting a lab router). It is part of the [Network Fundamentals guide](https://www.pinglabz.com/network-fundamentals/). ## The IOS Password-Type Hierarchy Not all stored passwords are equal, and IOS labels each with a type number. The difference is how the password is stored in the config, which decides how fast an attacker who reads your config can recover it. When you configure a plaintext password on a modern IOS XE image, the device itself warns you: ``` R1(config)# username helpdesk password cisco123 WARNING: Command has been added to the configuration using a type 0 password. However, recommended to migrate to strong type-6 encryption ``` That "type 0" is the tell: the password is stored in cleartext. Compare it to a properly hashed credential from the same device: ``` R1# show run | include secret username cisco privilege 15 secret 9 $9$866GqyqM2IY7Sk$3ZDbsO6QtTVaR1U8dKmRwQWO6EoXCWa3cyApcAHPJnI ``` The `9` after `secret` is the type. Here is the hierarchy, worst to best: **Type 0** Cleartext. Anyone who reads the config has the password. Never use in production. **Type 7** Cisco's reversible "encryption" from `service password-encryption`. Trivially decoded. Obfuscation, not security. **Type 5** Salted MD5 hash. Was the standard for years; now considered weak against modern cracking. **Type 8** PBKDF2 with SHA-256\. Strong and standards-based. **Type 9** scrypt. Deliberately memory-hard, the strongest option IOS offers. This is what you want. The practical rules: always use `secret` (which hashes) rather than `password` (which does not), prefer `enable secret` over `enable password`, and never rely on `service password-encryption` as anything more than shoulder-surf protection, because type 7 is reversible in seconds. On current images you can force the strong algorithm with `username x algorithm-type scrypt secret y`. ## Enforcing a Minimum Length A strong hash on a weak password buys you nothing. Set an organization-wide floor so nobody can configure a short one: ``` R1(config)# security passwords min-length 12 ``` After this, any attempt to set a password shorter than 12 characters is rejected at the CLI. It applies to enable and line passwords going forward (existing ones are untouched until they are changed). This is the kind of guardrail that survives staff turnover: it does not depend on whoever is at the keyboard remembering the policy. ## Locking Out Brute Force The single most valuable hardening command for the login prompt is `login block-for`. It watches for repeated failures and, once the threshold is crossed, drops the device into "quiet mode" where it refuses all login attempts for a set period: ``` R1(config)# login block-for 120 attempts 3 within 60 ``` Read it as: if there are 3 failed logins within 60 seconds, block all logins for 120 seconds. Here is that policy doing its job. From a Linux host, three SSH logins with a bad password were fired at the router in quick succession, and the device recorded every one: ``` R1# show login failures Information about last 50 login failure's with the device Username SourceIPAddr lPort Count TimeStamp attacker 192.168.99.100 22 3 22:57:45 UTC Fri Jul 10 2026 ``` The moment the third failure landed, quiet mode engaged, and the router logged both the failure and the lockout to its syslog server: ``` %SEC_LOGIN-4-LOGIN_FAILED: Login failed [user: attacker] [Source: 192.168.99.100] [localport: 22] [Reason: Login Authentication Failed] %SEC_LOGIN-1-QUIET_MODE_ON: Still timeleft for watching failures is 50 secs, [user: attacker] [Source: 192.168.99.100] [ACL: sl_def_acl] ``` Two things worth noting. First, quiet mode applies an ACL (`sl_def_acl`) that blocks everyone by default, so you can pair it with `login quiet-mode access-class` to keep your own management subnet exempt (otherwise you lock yourself out too). Second, this event went straight to a central log, which is how you would actually notice an attack in progress rather than discovering it after the fact. ## SSH as the Transport None of this matters if management traffic is in cleartext, so SSH (never Telnet) is the transport. A quick look at the negotiated parameters: ``` R1# show ip ssh | include SSH|Auth SSH Enabled - version 2.0 Authentication methods:publickey,keyboard-interactive,password Authentication timeout: 120 secs; Authentication retries: 3 ``` SSH version 2 only (version 1 is broken), a short authentication timeout, and a low retry count. The presence of `publickey` in the methods matters: key-based authentication removes the password from the equation entirely for automation and admin access, which is a stronger posture than any password policy. ## Multi-Factor Authentication and Centralized AAA A password is one factor: something you know. Multi-factor authentication (MFA) requires two or more of the classic three categories, and the CCNA expects you to name them: Something you knowA password, PIN, or passphrase. The weakest factor on its own because it can be guessed, phished, or reused. Something you haveA token, smartphone authenticator app, or a certificate on the device. Stops an attacker who only has the password. Something you areBiometrics: fingerprint, face, retina. Common for endpoint and building access, less so for CLI device login. Network devices do not run MFA prompts on their own. The way you get MFA (and centralized policy) onto a router or switch is to stop authenticating locally and hand the job to a central server via AAA. Instead of a username database on every box, the device asks a TACACS+ or RADIUS server, and that server can enforce MFA, tie logins to your identity directory, and log everything centrally. TACACS+ is Cisco's device-administration protocol (it can authorize individual commands); RADIUS is the standard for network access and is what [802.1X](https://www.pinglabz.com/802-1x/) uses to authenticate devices onto switchports. Certificates fit here too: a device or user presents a certificate instead of (or alongside) a password, which is the "something you have" factor in machine form. The progression to remember: local passwords are fine for a lab or a single device, but the moment you have more than a handful of boxes, centralized AAA is the answer, because you cannot enforce policy (or MFA, or clean offboarding) one router at a time. ## FAQ ### Should I use `secret` or `password`? Always `secret`. It stores a hash (type 5, 8, or 9 depending on the command), while `password` stores type 0 cleartext unless you layer on `service password-encryption`, which only gives you reversible type 7\. Use `enable secret` and `username x secret y` everywhere. ### Is `service password-encryption` good enough? No. It produces type 7, which is reversible with tools that have existed for decades. It stops someone glancing at your screen, nothing more. Treat any type 7 password as effectively cleartext. ### Can `login block-for` lock me out? Yes, if you don't exempt yourself. Quiet mode blocks all logins by default. Pair it with `login quiet-mode access-class` pointing at an ACL that permits your management subnet, so an attack from outside doesn't also lock out your NOC. ### Does a router support MFA directly? Not by itself. You get MFA by pointing the device at a centralized AAA server (TACACS+ or RADIUS) that enforces the additional factor. The device just delegates the authentication decision. ## Key Takeaways Use hashed secrets (type 8 or 9), never cleartext or reversible type 7\. Enforce a minimum length with `security passwords min-length` so policy does not depend on memory. Lock out brute force with `login block-for`, exempt your own subnet, and ship the events to a syslog server so you actually see attacks. Move management to SSHv2 and, past a handful of devices, to centralized AAA where MFA and identity integration live. For the access-control side (authenticating devices onto the network with RADIUS), continue to the [802.1X guide](https://www.pinglabz.com/802-1x/), and see the [Network Fundamentals guide](https://www.pinglabz.com/network-fundamentals/) for the rest of domain 5. ### Threats, Vulnerabilities, and Exploits: CCNA Security Concepts URL: https://www.pinglabz.com/threats-vulnerabilities-exploits/ Last updated: 2026-07-10T23:16:41.000Z CCNA domain 5 opens with a set of definitions that sound interchangeable but are not: threat, vulnerability, exploit, mitigation. The exam tests whether you can keep them straight, and the real world punishes you when you can't, because you end up patching the wrong layer. This article pins down each term, walks the attack categories the CCNA expects you to recognize, and maps every one of them to the specific Cisco technology that shuts it down. It is part of the [Network Fundamentals guide](https://www.pinglabz.com/network-fundamentals/), and it leans on the security tooling covered across the rest of the site rather than repeating it. ## The Four Terms, Precisely Get these exactly right, because the exam writes questions that hinge on the distinction. A **vulnerability** is a weakness: an unpatched IOS bug, a default password, an open switchport in a lobby, a user who clicks anything. A **threat** is the potential for someone or something to exploit that weakness (a threat actor is the who; a threat is the what-could-happen). An **exploit** is the actual tool, code, or technique that turns the vulnerability into a breach. **Mitigation** (also called a countermeasure) is what you put in place to reduce the likelihood or the impact. **Vulnerability** The weakness itself. A switchport with no port security. A device still running the default enable password. **Threat** The possibility that the weakness gets used against you, and the actor who might do it. **Exploit** The concrete method: a script, a crafted packet, a phishing email, a brute-force login loop. **Mitigation** The control that reduces likelihood or impact: a patch, an ACL, 802.1X, DHCP snooping. A worked example ties them together. The vulnerability is a router that permits SSH from anywhere with only a password. The threat is an attacker on the internet running a credential-guessing tool. The exploit is the brute-force loop itself. The mitigation is `login block-for`, an ACL on the VTY lines, and moving to key-based or AAA-backed authentication. You can watch that exact attack land in the logs (more on that in [Password Policies and MFA for Network Devices](https://www.pinglabz.com/password-policies-mfa-network-devices/)), and here is what the device recorded when three bad SSH logins hit it in a row: ``` %SEC_LOGIN-4-LOGIN_FAILED: Login failed [user: attacker] [Source: 192.168.99.100] [localport: 22] [Reason: Login Authentication Failed] %SEC_LOGIN-1-QUIET_MODE_ON: Still timeleft for watching failures is 50 secs, [user: attacker] [Source: 192.168.99.100] [ACL: sl_def_acl] ``` That is the whole lifecycle in two syslog lines: an exploit hitting a vulnerability, and a mitigation (quiet mode) kicking in to blunt it. ## The Attack Categories CCNA Expects You are not expected to run these attacks, only to recognize them and name the defense. Here are the categories that show up on the blueprint. ### Reconnaissance Before an attacker breaks anything, they map it: which hosts are up, which ports are open, what software answers. This is exactly what a port scanner does, and seeing it from the attacker's side makes the defensive lessons concrete. The [Nmap guide](https://www.pinglabz.com/nmap/) shows real scans against live targets, including how a firewall changes what a scan can see. Mitigation is mostly about reducing your attack surface (close unused ports, filter at the edge) and detecting the scan itself. ### Spoofing and Man-in-the-Middle Spoofing means forging an identity: a fake source IP, a forged MAC, a rogue DHCP server handing out a malicious default gateway, or ARP replies that redirect traffic through the attacker (the classic path to a man-in-the-middle position). These are Layer 2 attacks, and Layer 2 is where you stop them. DHCP snooping drops rogue DHCP offers; Dynamic ARP Inspection (DAI) validates ARP against the snooping table; port security limits MAC addresses per port. The [VLANs and Layer 2 Switching guide](https://www.pinglabz.com/vlans-layer-2-switching/) covers these switch hardening features in depth. ### Denial of Service (DoS and DDoS) A DoS floods a target with traffic or malformed requests until legitimate users can't get through; a DDoS does it from many sources at once, which is what makes it hard to filter by source address. On the CCNA, the defenses you name are rate limiting, ACLs to drop obvious garbage, and control-plane policing to protect the device's own CPU. A firewall like the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) adds connection limits and inspection on top. ### Password Attacks Brute force tries every combination; a dictionary attack tries likely words; credential stuffing reuses passwords leaked from other breaches. The syslog capture above is a brute-force attempt in progress. Mitigation is strong password policy, account lockout (`login block-for`), multi-factor authentication, and centralized AAA so you are not defending one device at a time. ### Social Engineering and Phishing The attacker skips the technology and targets the person: a phishing email, a phone call impersonating IT, a USB drive left in a parking lot. No ACL stops this. The mitigation is a security program element, not a config line: user awareness and training. This is why the blueprint lists people-focused controls alongside the technical ones. ### Malware Viruses attach to files, worms self-propagate across the network, trojans hide inside something that looks legitimate, and ransomware encrypts data for extortion. Network-side mitigations include segmentation (so a worm can't reach everything), NGFW and endpoint protection, and keeping software patched to close the vulnerabilities malware rides in on. ## Security Program Elements (Exam 5.2) Domain 5 explicitly calls out the non-technical side of security, and it is easy points if you don't skip it. Three elements you should be able to describe: User awarenessPrograms that teach the whole organization to recognize phishing, tailgating, and suspicious requests. The first line of defense against social engineering. User trainingRole-specific instruction: how staff use systems securely, follow acceptable-use policy, and handle data. Deeper and more targeted than general awareness. Physical access controlLocks, badges, mantraps, and secured wiring closets. A switch in an unlocked lobby is a vulnerability no ACL can fix. ## Defense in Depth No single control is enough, and the exam wants you to think in layers. Defense in depth means an attacker who gets past one control still faces the next: physical security, then port-based access control, then segmentation, then a firewall, then centralized authentication, then monitoring. If one layer fails, the others still stand. The practical corollary is that you match each threat to the layer that actually addresses it rather than throwing one big firewall rule at everything. ## Threat-to-Mitigation Mapping This is the table you want burned into memory for the exam. Each defense links to where it is covered in depth on the site. **Rogue DHCP / DHCP starvation** DHCP snooping (trusted vs untrusted ports). See the [L2 switching guide](https://www.pinglabz.com/vlans-layer-2-switching/). **ARP spoofing / MITM** Dynamic ARP Inspection, built on the DHCP snooping binding table. **MAC flooding / unknown devices** Port security (max MACs, violation actions) and 802.1X port-based auth. **Unauthorized network access** [802.1X](https://www.pinglabz.com/802-1x/) with RADIUS: authenticate the device before the port forwards traffic. **Traffic reaching hosts it shouldn't** ACLs and segmentation, plus a stateful firewall like the [Cisco ASA](https://www.pinglabz.com/cisco-asa/). **Credential guessing** Password policy, `login block-for`, MFA, and centralized AAA. Notice that most of these live at Layer 2, on the access switch. That is deliberate: the CCNA blueprint puts heavy weight on securing the edge, because the edge is where untrusted devices plug in. If security is your direction of travel, the sister site [Threat on the Wire](https://www.threatonthewire.com/?ref=pinglabz.com) goes deeper on the threat landscape and defensive tradecraft than a CCNA-framed article should. ## FAQ ### What is the difference between a threat and a vulnerability? A vulnerability is a weakness that already exists in your environment. A threat is the potential for something to take advantage of it. A door with a broken lock is a vulnerability; a burglar in the neighborhood is a threat. You reduce vulnerabilities directly (fix the lock) and you reduce threats through deterrence and detection. ### Is an exploit the same as an attack? An exploit is the tool or technique; the attack is the act of using it against a target. The same exploit can be used in many attacks. On the exam, "exploit" points to the method that turns a vulnerability into a compromise. ### Why does a networking exam test user training? Because the most reliable way into a network is a person, not a packet. Phishing and social engineering bypass every technical control. The blueprint includes security program elements so you understand that people are part of the attack surface and part of the defense. ## Key Takeaways Keep the four terms distinct: vulnerability (weakness), threat (potential), exploit (method), mitigation (control). Recognize the attack categories (reconnaissance, spoofing/MITM, DoS/DDoS, password attacks, social engineering, malware) and name the specific Cisco defense for each. Remember that security is layered and that people are a layer too. When you are ready to configure the controls, the [802.1X](https://www.pinglabz.com/802-1x/), [Layer 2 switching](https://www.pinglabz.com/vlans-layer-2-switching/), and [Cisco ASA](https://www.pinglabz.com/cisco-asa/) guides are where the theory here becomes running config. Start from the [Network Fundamentals guide](https://www.pinglabz.com/network-fundamentals/) to see how this fits the rest of domain 5. ### Static Routing on Cisco IOS XE: Default, Floating, and Host Routes URL: https://www.pinglabz.com/static-routing-cisco-ios-xe/ Last updated: 2026-07-10T23:15:58.000Z Static routing is the first routing you ever configure and the last routing you stop using. Dynamic protocols get the attention, but every real network still leans on hand-configured routes for default gateways, stub sites, backup paths, and steering a single prefix. The commands look trivial, which is exactly why people get burned by the parts that are not: recursion, floating administrative distance, and the difference between a next-hop static and a fully specified one. This article configures all of them on a live Cisco IOS XE lab and shows the failure modes, not just the happy path. It is part of the [IP Routing complete guide](https://www.pinglabz.com/ip-routing/). ## The three ways to write a static route Every `ip route` statement names a destination network and mask, then tells the router how to reach it. There are three forms, and the difference between them is not cosmetic: **Next-hop**`ip route 10.0.30.0 255.255.255.128 10.0.12.2`. Names only the next-hop IP. The router must recursively look that IP up in the table to find the exit interface. **Exit-interface**`ip route 10.0.30.0 255.255.255.128 Ethernet0/2`. Names only the outgoing interface. Works cleanly on point-to-point links, risky on multi-access (Ethernet) ones. **Fully specified**`ip route 3.3.3.3 255.255.255.255 Ethernet0/2 10.0.13.2`. Names both. The most robust form, and the one that fails over correctly. Here are a next-hop static and a host route configured on R1, confirmed in the running config and the table: ``` R1#show run | include ip route ip route 3.3.3.3 255.255.255.255 10.0.13.2 ip route 3.3.3.3 255.255.255.255 10.0.12.2 250 ip route 10.0.30.0 255.255.255.128 10.0.12.2 R1#show ip route 3.3.3.3 Routing entry for 3.3.3.3/32 Known via "static", distance 1, metric 0 Routing Descriptor Blocks: * 10.0.13.2 Route metric is 0, traffic share count is 1 ``` The 3.3.3.3/32 entry is a **host route**: a /32 static pointing at one specific device. Host routes are how you reach a loopback, pin management traffic to a path, or blackhole a single address. The bracket shows `[1/0]`: administrative distance 1, metric 0\. Static routes always have metric 0, which is why AD is the only lever you have to rank them. ## The default route: 0.0.0.0/0 The most common static route in the world is the default: a route to 0.0.0.0/0 that catches everything not matched more specifically. On switch SW1 it points at its gateway: ``` SW1#show run | include ip route ip route 0.0.0.0 0.0.0.0 10.0.20.1 SW1#show ip route | begin Gateway Gateway of last resort is 10.0.20.1 to network 0.0.0.0 S* 0.0.0.0/0 [1/0] via 10.0.20.1 10.0.0.0/8 is variably subnetted, 2 subnets, 2 masks C 10.0.20.0/24 is directly connected, Vlan10 L 10.0.20.10/32 is directly connected, Vlan10 ``` Two things to read here. The `Gateway of last resort is 10.0.20.1` line now names a next hop, where a router without a default says "not set." And the route itself is flagged `S*`: the S is static, the asterisk marks it as a candidate default route. Because 0.0.0.0/0 is the shortest possible prefix (length 0), [longest prefix match](https://www.pinglabz.com/longest-prefix-match/) guarantees any real route beats it, which is precisely why it is the route of last resort. ## Floating static routes: a backup that only appears when needed A floating static is a static route with a deliberately high administrative distance, so it stays out of the table while a preferred route (usually from a dynamic protocol) is present, and floats in only when that route disappears. This is the standard pattern for a backup WAN link. The trick is choosing the AD correctly, and this is where people go wrong. On R1, 3.3.3.3/32 is reachable three ways: a primary static, a floating static, and via OSPF. Watch what happens across a failure. First, both statics are configured, the primary is fully specified and the floating one has AD 250: ``` R1#show ip route 3.3.3.3 | include Known|via|\* Known via "static", distance 1, metric 0 * 10.0.13.2, via Ethernet0/2 ``` The primary static (AD 1) is installed. Now the primary's link goes down (Ethernet0/2 is shut) and we re-check: ``` ### after e0/2 shutdown: fully-specified static withdrawn, OSPF 110 beats floating 250: R1#show ip route 3.3.3.3 Routing entry for 3.3.3.3/32 Known via "ospf 1", distance 110, metric 21, type intra area Last update from 10.0.12.2 on Ethernet0/1, 00:00:15 ago Routing Descriptor Blocks: * 10.0.12.2, from 3.3.3.3, 00:00:15 ago, via Ethernet0/1 ``` This is the gotcha. The floating static was set to AD 250, but OSPF's AD is 110\. When the primary failed, both the floating static (250) and OSPF (110) were available, and OSPF won because 110 is lower. The engineer who set 250 probably wanted the static as the backup, but they positioned it below OSPF in the trust order, so OSPF, not the static, took over. A floating static's AD must be chosen relative to *every* other source that can offer the same prefix, not just the primary. Fix it by giving the floating static an AD that beats OSPF, say 105, which is above the primary static's 1 but below OSPF's 110: ``` ### floating re-created at AD 105 (below OSPF 110): installs: R1#show ip route 3.3.3.3 Routing entry for 3.3.3.3/32 Known via "static", distance 105, metric 0 Routing Descriptor Blocks: * 10.0.12.2 Route metric is 0, traffic share count is 1 ``` Now the static installs when the primary is gone, because 105 beats OSPF's 110\. The `Known via "static", distance 105` line is the proof. The lesson: a floating static route's administrative distance is a ranking against the entire set of possible sources, so map out what else knows the prefix before you pick the number. ## The recursion gotcha The reason the primary static above was written as fully specified (`Ethernet0/2 10.0.13.2`) rather than next-hop only is failover behavior. A next-hop-only static route is recursive: the router keeps it in the table as long as it can still resolve the next-hop IP through *some* other route. That sounds fine until the direct link fails but the next hop is still reachable the long way around, at which point the static never withdraws and can send traffic in circles. A fully specified static (interface plus next hop) is bound to the interface. When that interface goes down, the route is withdrawn immediately and cleanly, which is exactly what let OSPF and the floating static take over above. On point-to-point links you can use exit-interface-only statics safely; on Ethernet, always add the next hop as well so ARP resolves against a specific neighbor. The practical rule: for anything that needs to fail over, write the primary static as fully specified so it disappears the instant its link does. ## Verifying and troubleshooting statics Three commands cover almost all static-route troubleshooting: - `show ip route static` lists only the statics that made it into the table. A static you configured that is missing here failed to install (usually its next hop is unreachable, or a longer prefix beat it). - `show ip route ` shows the winning route and its AD, so you can see whether your floating static is at the right distance. - `show running-config | include ip route` shows what you actually configured, which sometimes differs from what you meant to configure. If a static is in the running config but not the routing table, the number one cause is an unresolvable next hop: the router cannot find a route to the next-hop IP, so it cannot install the recursive static. Fully specifying the interface often fixes it. ## When static routing is the right choice Static routes get dismissed as "not real routing," which is wrong. They are the correct tool in several common situations: - **Stub networks.** A branch site or a small office with a single path to the rest of the network has nothing to gain from a dynamic protocol. One default route out and one summary route back is simpler, uses no CPU or bandwidth for updates, and cannot be destabilized by a protocol flap somewhere else. - **The internet edge.** Most enterprises point a default route at their ISP rather than accepting a full BGP table, because they only have one or two exits and do not need to make per-prefix decisions. - **Backup paths.** A floating static over a cellular or secondary link gives you automatic failover without extending your dynamic protocol across a link you do not fully trust. - **Steering specific traffic.** A single more-specific static can pull one destination onto a preferred path without touching the rest of your routing design, thanks to [longest prefix match](https://www.pinglabz.com/longest-prefix-match/). The flip side: static routes do not react to topology changes beyond the directly attached link. If a failure happens two hops away, a static route pointing through that path keeps forwarding into a black hole. That is the boundary where you reach for a dynamic protocol, and it is why large networks run [OSPF](https://www.pinglabz.com/ospf/) or [EIGRP](https://www.pinglabz.com/eigrp/) for the interior and keep static routes for the edges. ## Key takeaways Static routes come in three forms, and the choice matters: fully specified (interface plus next hop) is the robust default because it withdraws the instant its link fails. The default route 0.0.0.0/0 shows as `S*` and is always the least specific match. A floating static is just a static with a high AD, but its distance must be positioned against every source that can offer the prefix, not only the primary, or a dynamic protocol will quietly win the failover instead. And a next-hop-only static is recursive and can linger when you want it gone, so bind failover statics to their interface. Once you are comfortable here, the dynamic protocols in the [IP Routing complete guide](https://www.pinglabz.com/ip-routing/) build on exactly the same table-selection rules, and [FHRP](https://www.pinglabz.com/fhrp/) handles the first-hop redundancy that statics alone cannot. ### Longest Prefix Match: How Routers Actually Choose Routes URL: https://www.pinglabz.com/longest-prefix-match/ Last updated: 2026-07-10T23:15:57.000Z Ask a room of engineers "which route wins?" and most reach straight for administrative distance. That is the wrong first answer. Before a router ever looks at AD or metric, it applies longest prefix match: given a destination IP, it forwards using the *most specific* route that contains that address, and a more specific prefix beats a less specific one no matter what protocol found it or how bad its metric is. Get this ordering right and a lot of confusing routing behavior suddenly makes sense. This article proves it on a live Cisco IOS XE router, and it is part of the [IP Routing complete guide](https://www.pinglabz.com/ip-routing/). ## The rule in one sentence When a packet needs forwarding, the router finds every route whose network contains the destination address, and among those it picks the one with the longest prefix (the largest number after the slash). Only if a single prefix length has multiple sources does administrative distance come into play, and only within one source does metric break the tie. Prefix length is the first filter, and it is absolute. ## Two routes to the same block, on purpose Router R1 has been given two overlapping routes to the 10.0.30.0 space: a /24 learned from OSPF, and a /25 configured as a static. Both are valid, both are in the table at the same time: ``` R1#show ip route | include 10.0.30 O 10.0.30.0/24 [110/11] via 10.0.13.2, 00:07:36, Ethernet0/2 S 10.0.30.0/25 [1/0] via 10.0.12.2 ``` The /24 covers 10.0.30.0 through 10.0.30.255\. The /25 covers only 10.0.30.0 through 10.0.30.127\. They overlap: any address in the lower half is described by both routes. Notice the static (/25) has AD 1 and the OSPF route (/24) has AD 110\. If AD were the deciding factor, the static would always win. It is not the deciding factor. Prefix length is. ## Proving which route wins for which address Take an address in the lower half, 10.0.30.1, and ask the router which route it resolves to: ``` R1#show ip route 10.0.30.1 Routing entry for 10.0.30.0/25 Known via "static", distance 1, metric 0 Routing Descriptor Blocks: * 10.0.12.2 Route metric is 0, traffic share count is 1 ``` The router picked the /25\. That address is covered by both routes, and the /25 is more specific, so it wins. Now take an address in the *upper* half, 10.0.30.200, which only the /24 covers: ``` R1#show ip route 10.0.30.200 Routing entry for 10.0.30.0/24 Known via "ospf 1", distance 110, metric 11, type intra area Last update from 10.0.13.2 on Ethernet0/2, 00:07:37 ago Routing Descriptor Blocks: * 10.0.13.2, from 3.3.3.3, 00:07:37 ago, via Ethernet0/2 ``` Same router, same table, same instant, but a different route because the /25 does not contain 10.0.30.200\. The destination address decides which prefixes are even eligible, and then the longest eligible one wins. Two addresses in what looks like "the same subnet" take completely different next hops. ## Confirming it in the forwarding plane The routing table (RIB) is the control-plane view. The router actually forwards using CEF (the FIB). Check that CEF agrees, because forwarding is what a packet really experiences: ``` R1#show ip cef 10.0.30.1 10.0.30.0/25 nexthop 10.0.12.2 Ethernet0/1 R1#show ip cef 10.0.30.200 10.0.30.0/24 nexthop 10.0.13.2 Ethernet0/2 ``` CEF confirms it: 10.0.30.1 forwards via 10.0.12.2 out Ethernet0/1 (the /25 static path), and 10.0.30.200 forwards via 10.0.13.2 out Ethernet0/2 (the /24 OSPF path). The control plane and forwarding plane match, which is what you want to see. ## The clincher: a traceroute that takes the "wrong" path If longest prefix match is really overriding administrative distance and metric, a packet to 10.0.30.1 should physically travel via the static's next hop (10.0.12.2, which is R2) rather than the direct OSPF path to R3\. Trace it: ``` R1#traceroute 10.0.30.1 probe 1 Type escape sequence to abort. Tracing the route to 10.0.30.1 VRF info: (vrf in name/id, vrf out name/id) 1 10.0.12.2 1 msec 2 10.0.23.2 3 msec ``` Two hops. The packet went to R2 first (10.0.12.2), then reached the destination via 10.0.23.2, exactly the longer path the /25 static points down. Compare that to the direct one-hop OSPF route the /24 would have used. And it still works end to end: ``` R1#ping 10.0.30.1 Sending 5, 100-byte ICMP Echos to 10.0.30.1, timeout is 2 seconds: !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 2/2/3 ms ``` This is the entire lesson made visible: the router chose a longer, higher-AD, higher-metric path purely because its prefix was more specific for that particular destination address. ## Why this matters in real networks Longest prefix match is not an exam curiosity; it is a working tool: - **Route summarization** depends on it. A core router can advertise a single 10.0.0.0/8 while an access router advertises specific /24s, and traffic still finds the specific path because the /24 is longer than the /8. - **Traffic engineering** uses it. Inject a more specific route to steer a subset of traffic down a preferred link without touching the broader routing design. - **Default routes** are the shortest possible prefix (0.0.0.0/0 matches everything with length 0), which is exactly why the default is the route of last resort: any real route is more specific and beats it. - **Blackholing and RTBH** (remotely triggered black hole filtering) work by injecting a very specific /32 that points to null, overriding any broader legitimate route for that one host. ## How the router does the lookup so fast You might wonder how a router compares a destination against thousands of routes for every single packet without falling over. It does not walk the table linearly. The forwarding table (CEF, the FIB) is organized as a tree structure (historically an mtrie) keyed by prefix, so a lookup descends bit by bit through the destination address and lands on the longest matching prefix in a bounded number of steps, regardless of table size. That is why an internet-edge router holding a full BGP table of nearly a million routes still forwards at line rate: longest prefix match is a hardware-optimized operation, not a loop. When you run `show ip cef 10.0.30.1` and get an instant answer, you are reading the result of exactly that tree walk: ``` R1#show ip cef 10.0.30.1 10.0.30.0/25 nexthop 10.0.12.2 Ethernet0/1 ``` The takeaway for operators: the RIB (routing table) is where routes are chosen, but the FIB (CEF) is where the fast longest-prefix lookup happens on every packet. They should always agree; when they do not, you have found a real problem. ## The ordering, start to finish Put the three concepts in the sequence the router actually applies them: 1**Longest prefix match.** Of all routes containing the destination, keep only the most specific prefix length. This filters first and cannot be overridden by AD or metric. 2**Administrative distance.** If that winning prefix length was offered by more than one protocol, install the source with the lowest AD. See [administrative distance explained](https://www.pinglabz.com/administrative-distance/). 3**Metric.** Within that one protocol, the lowest metric wins. Equal metrics install as equal-cost multipath. ## A common source of confusion: overlapping does not mean conflicting Beginners often see two overlapping routes in a table and assume one is a mistake. It is not. Overlapping prefixes are a normal, deliberate design: a broad route provides general reachability while a more specific route carves out an exception for part of that space. The routing table is perfectly happy holding a /8, a /24, and a /25 that all cover the same address, because longest prefix match guarantees each destination resolves to exactly one of them (the longest that matches). So when you see a specific static sitting alongside a broader dynamic route, that is usually intent, not error, and deleting the "extra" route can quietly break the traffic engineering it was performing. ## Key takeaways Longest prefix match runs first and it is absolute: the router forwards a packet using the most specific route that contains the destination address, regardless of administrative distance or metric. A /25 static beats a /24 OSPF route for addresses inside the /25, and the traceroute proves the packet really takes that path. Only after prefix length is settled does AD choose between sources, and only then does metric break the final tie. When you are troubleshooting "why did my traffic go there," check the prefix length before you blame the protocol, and confirm with `show ip route ` and `show ip cef `. The full picture lives in the [IP Routing complete guide](https://www.pinglabz.com/ip-routing/). ### Administrative Distance: The Complete Table, Proven on a Live Lab URL: https://www.pinglabz.com/administrative-distance/ Last updated: 2026-08-01T19:26:52.000Z When a router learns the same destination from two different routing protocols, it cannot install both. It has to pick a winner, and it does not use the metric to decide, because metrics from different protocols are not comparable (an OSPF cost of 11 and an EIGRP metric of 409600 measure completely different things). Instead it uses administrative distance: a fixed number ranking how much it trusts each *source* of routing information. Lower is more trusted. This article gives you the complete default table and then proves the behaviour on a live lab: two `iol-xe` routers on CML running IOS XE 17.18.2, R1 and R2 back to back over 10.0.12.0/30, with R2 advertising 8.8.8.0/24 into OSPF and EIGRP at once. R1 installs the EIGRP copy, then reinstalls the same prefix via OSPF after a single command, with nothing about the topology changing. It is part of the wider guide to [how a router builds and uses its routing table](https://www.pinglabz.com/ip-routing/). Two facts cause more confusion than everything else combined. AD is local to the router: it travels in no routing protocol, so your neighbours have no idea what yours is. And AD only breaks the tie *between* protocols, because inside one protocol every route carries the same distance and the metric decides it. ## What administrative distance actually is Administrative distance is a per-route-source trust value between 0 and 255\. When two sources offer the same prefix with the same mask, the router installs the lower AD and leaves the other in its own protocol's database, unused but ready. AD is compared **only after** longest prefix match, a subtlety covered in [how a router picks the most specific route](https://www.pinglabz.com/longest-prefix-match/). A /32 from RIP still beats a /24 from EIGRP, because prefix length is evaluated first and AD never gets a say. The full order for one destination is: 1. **Longest prefix match.** Nothing below this step can overturn it. 2. **Lowest administrative distance.** Among sources offering the identical prefix. 3. **Lowest metric.** Only ever between routes from the same source. 4. **Equal cost, so install both**, which is where [uneven load sharing and CEF polarization](https://www.pinglabz.com/ecmp-cef-load-balancing-polarization/) come from. ## AD never leaves the router There is no field in an OSPF LSA, no TLV in an EIGRP update and no BGP path attribute that carries administrative distance. It is a local property of the receiving router, applied when a protocol offers a prefix to the RIB. There is therefore no such thing as an AD mismatch error: adjacencies do not fail over it and nothing logs a warning. That makes every AD change a per-device item you have to apply consistently on every router that sees the prefix. Lower a static route's distance on R1 to force traffic down one link, forget the return path on R3, and you get asymmetric routing at best and a forwarding loop at worst. It only shows up in the data plane, which is why it pays to know how to [spot a routing loop in a traceroute](https://www.pinglabz.com/routing-loops-detect-and-fix/). ## AD only breaks ties between protocols If a router has four OSPF paths to the same prefix, all four have a distance of 110\. There is no tie to break, so OSPF cost decides, and no amount of fiddling with `distance` under the OSPF process changes the outcome. The same goes for two EIGRP paths, and for two static routes to the same prefix, which both carry AD 1 and both install. You can read both numbers off the table. In `O 8.8.8.0/24 [110/11] via 10.0.12.2` the first bracketed number is the administrative distance and the second is the metric, a distinction that matters every time you [work out what a line in the routing table is telling you](https://www.pinglabz.com/how-to-read-cisco-routing-table/). To steer between two paths inside one protocol you change the metric (OSPF cost, an EIGRP offset list), not the distance. BGP looks like an exception, with eBGP at 20 and iBGP at 200, but it is not: BGP picks its own best path first and offers a single winner per prefix to the RIB, so those values are only ever weighed against other protocols. That two-stage process is why a prefix can sit in `show ip bgp` and [never appear in the routing table](https://www.pinglabz.com/bgp-route-not-in-routing-table/). ## The complete administrative distance table These are the defaults on Cisco IOS and IOS XE. Memorise the common ones (0, 1, 90, 110, 120, and the two BGP values). The legacy entries are here because they still turn up in the docs and in exam questions. 0 **Connected interface** A directly attached subnet. Unbeatable. 1 **Static route** Configured by hand with `ip route`. 5 **EIGRP summary route** Auto or manual summary aggregate. 20 **External BGP (eBGP)** Routes from another autonomous system. 90 **Internal EIGRP** The default winner in an all-Cisco shop. 100 **IGRP** Legacy, gone from modern IOS. 110 **OSPF** Intra-area, inter-area and external alike. 115 **IS-IS** Common in service-provider cores. 120 **RIP** Legacy distance-vector, rarely deployed. 140 **EGP** Historical exterior protocol, dead. 160 **On Demand Routing (ODR)** CDP-based stub routing. 170 **External EIGRP** Redistributed into EIGRP. 200 **Internal BGP (iBGP)** Routes from a peer in your own AS. 250 **NHRP** Next hop resolution, used by DMVPN. 255 **Unknown / unusable** A route at 255 is never installed. Three deserve a note. eBGP sits at 20, below every IGP, so a router prefers a route from an external AS over its own internal routing, which is what you want at an internet edge. OSPF externals stay at 110 rather than taking the penalty EIGRP externals take at 170\. And 255 does not mean "the worst route", it means "do not install this at all". One more that catches people: a static pointed at an outbound interface still has AD 1, even though the table prints it as directly connected. ## Watching AD choose a winner, live R2 owns 8.8.8.0/24 and advertises it into both OSPF area 0 and EIGRP AS 100, so R1 receives the identical prefix, same mask, same link, from two protocols at once. Ask R1 which one it installed: ``` R1# show ip route 8.8.8.0 Routing entry for 8.8.8.0/24 Known via "eigrp 100", distance 90, metric 409600, type internal * 10.0.12.2, from 10.0.12.2, via Ethernet0/0 ``` EIGRP won, because its internal AD of 90 is lower than OSPF's 110\. The `distance 90` is the whole explanation. The `metric 409600` next to it played no part and is only ever compared against other EIGRP metrics. OSPF's copy is still in R1's link-state database, fully computed and valid; it just never got offered a slot in the table. That one line, `Known via "", distance `, answers almost every "why is my route not being used" question. ## Flipping the winner with one command To prove the decision is about trust and not about the network, change nothing physical and raise EIGRP's internal distance above OSPF's: ``` R1(config)# router eigrp 100 R1(config-router)# distance eigrp 130 170 ! internal 130, external 170 ``` ``` R1# show ip route 8.8.8.0 Routing entry for 8.8.8.0/24 Known via "ospf 1", distance 110, metric 11, type intra area * 10.0.12.2, from 2.2.2.2, via Ethernet0/0 ``` Same destination, same next hop 10.0.12.2, same outgoing interface: every packet still leaves through the same port. All that changed is which process owns the entry, because EIGRP now offers the prefix at 130 and OSPF at 110. Two details are worth pausing on. The metric went from 409600 to 11 for what is physically the same path, which is why cross-protocol metric comparison is meaningless. And `from` changed from 10.0.12.2 to 2.2.2.2, because EIGRP reports the neighbour's interface address while OSPF reports the advertising router ID (a source identifier, not a next hop). The syntax matters more than you would expect. Under an EIGRP process a bare `distance 130` is the address-based form and silently does nothing here: no error, no warning, no change. EIGRP wants `distance eigrp `, and the reasoning behind that pair is in [why EIGRP treats redistributed routes as less trustworthy](https://www.pinglabz.com/eigrp-administrative-distance/). ## What happens to the losing route Nothing dramatic. The losing protocol keeps the prefix in its own database for as long as its neighbour advertises it; only RIB installation is contested. OSPF's 8.8.8.0/24 existed on R1 the whole time EIGRP owned the entry, which is why the swap above was immediate: no reconvergence, no new LSA, no neighbour event. That is also what gives you deterministic failover when a protocol dies: EIGRP loses its adjacency, withdraws the prefix, and the OSPF copy drops straight in at 110\. To see both copies, ask each protocol rather than the RIB: ``` show ip eigrp topology 8.8.8.0/24 show ip ospf database show ip route 8.8.8.0 ``` ## Floating static routes, the main practical use of AD Every backup path design you build is this mechanism applied on purpose. A floating static is a static route given a distance high enough to lose to whatever normally carries the prefix: it stays out of the table while the primary is alive, and installs the moment that source withdraws. ``` ! primary is learned via OSPF (AD 110); this only installs when OSPF's copy goes ip route 10.0.20.0 255.255.255.0 192.0.2.1 200 ! backup default route out a secondary circuit ip route 0.0.0.0 0.0.0.0 198.51.100.1 250 ``` The number is not arbitrary: it has to be higher than the AD of the source you are backing up, with headroom for another tier later (200 and 250 are the conventional choices). Never use 255, which means never install rather than install last. Two ways a floating static fails to float. The first is the tie: two statics at the same distance do not back each other up, they both install and load share. The second is far more common in production. A static pointed at a next hop across an Ethernet segment or a metro circuit stays valid while the local interface is up, and that interface stays up even when the far end is dead, so the backup never activates. Condition the static on a reachability test rather than link state, which is what [IP SLA with object tracking](https://www.pinglabz.com/ip-sla-cisco-ios-xe/) is for. Full syntax and the recursive lookup rules are in [static routing on Cisco IOS XE](https://www.pinglabz.com/static-routing-cisco-ios-xe/). ## Internal versus external distance inside one protocol Several protocols carry more than one AD value, set together. EIGRP is the clearest case: routes learned natively get 90, routes redistributed in from elsewhere get 170, which is why the command above takes two numbers. The higher external distance biases every router toward a native path over a redistributed copy, reducing (but not eliminating) the chance of a redistribution loop. The equivalents elsewhere: EIGRP`distance eigrp ` (90 / 170) OSPF`distance ospf intra-area X inter-area Y external Z` (all 110) BGP`distance bgp ` (20 / 200 / 200) RIP and IS-IS`distance `, a single value (120 / 115) Statictrailing value on `ip route`, default 1 Per-neighbour override`distance [acl]` That last form is the one people trip over: `distance` plus a number, an address and a wildcard applies a distance to routes from specific neighbours, optionally narrowed by an ACL. It is also why the bare EIGRP `distance 130` mistake is silent: the parser took a valid command, just not the one anyone intended. ## Changing administrative distance, and when you should You can override any default, but do it deliberately and document why. Three cases justify it: floating statics (above), a migration where old and new IGPs run in parallel and distance decides which one the network forwards on, and a redistribution boundary where raising the distance of routes from one neighbour stops a router preferring a redistributed copy of its own prefixes. A distance of 255 also suppresses a route entirely, a blunt alternative to a distribute list. What does not justify it is a path preference problem inside one protocol: if both candidates come from OSPF, distance is not the lever, cost is. ## AD is per-route, not per-protocol globally A subtle point that trips up even experienced engineers: administrative distance is evaluated per prefix, not once for the whole protocol. OSPF does not "lose to EIGRP" across your network. For each destination the router compares only the sources that offered *that* prefix, so it is normal for one router to hold some prefixes via EIGRP, others via OSPF and others static, all at once. `show ip route summary` makes this concrete: the per-source counts add up to the whole table precisely because each prefix was won independently. This is also why redistribution is dangerous. Redistribute OSPF into EIGRP and EIGRP back into OSPF and a prefix can arrive from two directions with different distances; if you have not planned them, the router may prefer the redistributed copy and start forwarding back toward the source. The internal and external split biases routers toward native routes but is not a complete safeguard, so with mutual redistribution, map the AD of every path a prefix can take. That is the first thing to check when [routes vanish or loop after you turn redistribution on](https://www.pinglabz.com/troubleshooting-route-redistribution/). ## Metric versus administrative distance, do not confuse them The two numbers in the bracket `[110/11]` get mixed up constantly, so nail the distinction: Administrative distance Compares different **sources**. Fixed defaults. Local to the router, never advertised. The first number in the bracket. Decides which protocol's route gets installed. Metric Compares paths **within one protocol**. Calculated from bandwidth, cost, or hops. The second number in the bracket. Only meaningful next to another metric from the same protocol. ## What this was captured on The lab is deliberately minimal, because the point is the decision and not the topology. Two `iol-xe` nodes on CML running IOS XE 17.18.2, R1 and R2 directly connected over 10.0.12.0/30, both running OSPF process 1 and EIGRP AS 100 across that link. R2 carries Loopback1 at 8.8.8.1/24 and advertises it into both. Both captures came from an on-box EEM applet on R1 logging `show ip route 8.8.8.0` to syslog. One build detail is worth stealing. OSPF advertises a loopback as a /32 host route whatever mask you configure, so left alone OSPF would carry 8.8.8.1/32 while EIGRP carried 8.8.8.0/24\. Different prefixes never compete, and longest prefix match would hand every packet to the /32 regardless of distance. Adding `ip ospf network point-to-point` makes OSPF advertise the real /24 and puts the protocols in genuine contention: step one beating step two, every time. ## Common mistakes and gotchas - **A bare `distance 130` under EIGRP does nothing.** That is the address-based form. Use `distance eigrp ` and verify with `show ip route`. - **Different prefix lengths never compete.** If the two protocols advertise a /24 and a /32, AD is not involved. Check the mask first. - **Changing AD on one router only.** Nothing exchanges distance, so a one-sided change produces asymmetric forwarding or a loop, silently. - **Expecting AD to pick between two paths in one protocol.** They share a distance. Change the metric. - **Treating 255 as "last resort".** It means never install, so a backup at 255 never activates. - **The floating static that never floats,** because the primary stays valid while the local interface is up. Track reachability, not link state. ## Key takeaways - AD ranks route *sources* by trust, after longest prefix match and before metric, and lower wins: connected 0, static 1, eBGP 20, internal EIGRP 90, OSPF 110, IS-IS 115, RIP 120, external EIGRP 170, iBGP 200, 255 meaning never install. - AD is local and appears in no protocol message, so every change must be made on every device that sees the prefix. - AD only decides between protocols. Inside one, every route shares the distance and the metric decides, which is why `distance` will not steer traffic between two OSPF paths. - `show ip route ` answers the question directly: `Known via "", distance ` names the winning source and the number that won. - Floating statics are this used deliberately: distance above the primary, 255 avoided, and a primary that can fail in a way the router notices. For how the winning route is then handed to CEF and turned into forwarded packets, continue with [the complete guide to routing on Cisco IOS XE](https://www.pinglabz.com/ip-routing/). ### How to Read the Cisco Routing Table (show ip route Line by Line) URL: https://www.pinglabz.com/how-to-read-cisco-routing-table/ Last updated: 2026-08-01T19:35:15.000Z The routing table is the single most important piece of state on a Cisco router, and `show ip route` is the command you will type more than any other. Yet most people skim it: they look for the destination, glance at the next hop, and move on. Read it properly and every line tells you which protocol installed the route, how much the router trusts that source, what the path cost is, how long it has been stable, and which interface the packet leaves through. This guide walks the output line by line, using real captures from a live Cisco IOS XE lab. It is part of the [IP Routing complete guide](https://www.pinglabz.com/ip-routing/), and once you can read the table fluently the rest of routing gets a lot less mysterious. ## The command and what it produces On a router named R1 running OSPF, EIGRP, a static route, and its own connected interfaces, here is the full table: ``` R1#show ip route Codes: L - local, C - connected, S - static, R - RIP, M - mobile, B - BGP D - EIGRP, EX - EIGRP external, O - OSPF, IA - OSPF inter area N1 - OSPF NSSA external type 1, N2 - OSPF NSSA external type 2 E1 - OSPF external type 1, E2 - OSPF external type 2, m - OMP n - NAT, Ni - NAT inside, No - NAT outside, Nd - NAT DIA i - IS-IS, su - IS-IS summary, L1 - IS-IS level-1, L2 - IS-IS level-2 ia - IS-IS inter area, * - candidate default, U - per-user static route H - NHRP, G - NHRP registered, g - NHRP registration summary o - ODR, P - periodic downloaded static route, l - LISP a - application route + - replicated route, % - next hop override, p - overrides from PfR & - replicated local route overrides by connected Gateway of last resort is not set 1.0.0.0/32 is subnetted, 1 subnets C 1.1.1.1 is directly connected, Loopback0 2.0.0.0/32 is subnetted, 1 subnets D 2.2.2.2 [90/409600] via 10.0.12.2, 00:07:28, Ethernet0/1 3.0.0.0/32 is subnetted, 1 subnets O 3.3.3.3 [110/11] via 10.0.13.2, 00:06:41, Ethernet0/2 10.0.0.0/8 is variably subnetted, 9 subnets, 4 masks C 10.0.12.0/30 is directly connected, Ethernet0/1 L 10.0.12.1/32 is directly connected, Ethernet0/1 C 10.0.13.0/30 is directly connected, Ethernet0/2 L 10.0.13.1/32 is directly connected, Ethernet0/2 D 10.0.20.0/24 [90/307200] via 10.0.12.2, 00:07:28, Ethernet0/1 O 10.0.23.0/30 [110/20] via 10.0.13.2, 00:06:41, Ethernet0/2 [110/20] via 10.0.12.2, 00:06:41, Ethernet0/1 O 10.0.30.0/24 [110/11] via 10.0.13.2, 00:06:41, Ethernet0/2 S 10.0.30.0/25 [1/0] via 10.0.12.2 O 10.0.40.0/24 [110/20] via 10.0.12.2, 00:06:44, Ethernet0/1 192.168.99.0/24 is variably subnetted, 2 subnets, 2 masks C 192.168.99.0/24 is directly connected, Ethernet0/0 L 192.168.99.1/32 is directly connected, Ethernet0/0 ``` That is a lot of screen, but it decomposes into four parts: the codes legend, the gateway of last resort line, the parent (classful) headers, and the child route entries. Take them in order. ## The codes legend The block at the top is a static legend. It never changes with your topology, so experienced engineers stop reading it. The letters that matter are the ones that actually appear in the left margin below. In this table you can see `L`, `C`, `S`, `D`, and `O`. Learn these five and you can read most enterprise tables: C Connected. A subnet on an up/up interface. The router owns the wire. L Local. The router's own /32 host address on that interface. Always paired with a C entry. S Static. You typed `ip route`. An `S*` marks a static default route. D EIGRP. The D is for DUAL, EIGRP's algorithm. `D EX` means an external (redistributed) route. O OSPF. `O IA` is inter-area; `O E1/E2` are external. Plain O is intra-area. B BGP. Not in this table, but you will see it the moment the router speaks to another autonomous system. The legend also lists `B` for BGP, and that is where this table most often disagrees with a protocol's own: a prefix can sit in `show ip bgp` looking perfectly healthy and never earn a `B` line here at all. [Reading the BGP status codes before you read the route entry](https://www.pinglabz.com/bgp-route-not-in-routing-table/) is what tells you which of the two tables to blame. ## Gateway of last resort The line `Gateway of last resort is not set` tells you the router has no default route. If a packet arrives for a destination that matches nothing in the table, it gets dropped and the sender receives an ICMP destination-unreachable. Add a default route (or receive one from a neighbor) and this line changes to name the next hop, and a route flagged with `*` appears. On an edge router this line should almost always point somewhere; on a core router with full routing it is often intentionally absent. ## Parent and child routes: the subnetting headers Cisco groups routes under classful parent lines. Look at this pair: ``` 10.0.0.0/8 is variably subnetted, 9 subnets, 4 masks C 10.0.12.0/30 is directly connected, Ethernet0/1 ``` The indented, code-less line (`10.0.0.0/8 is variably subnetted, 9 subnets, 4 masks`) is a **parent route**, also called a classful header. It is not a forwarding entry; it is a summary telling you that within the classful 10.0.0.0/8 space, the router holds 9 child subnets carved with 4 different mask lengths (/30, /32, /24, /25 here). The lines below it with an actual code letter are the **child routes**, and those are what the router forwards on. "Variably subnetted" simply means VLSM is in use, which it almost always is. When you are counting how many real routes exist, count the coded lines, not the headers. ## Anatomy of a single route entry Take one line and name every field: ``` O 10.0.23.0/30 [110/20] via 10.0.13.2, 00:06:41, Ethernet0/2 ``` **O**Source code. OSPF learned this route. **10.0.23.0/30**Destination network and prefix length. **\[110/20\]**\[administrative distance / metric\]. AD 110 is OSPF's trust value; 20 is the OSPF cost. **via 10.0.13.2**Next-hop IP address. The router forwards toward this neighbor. **00:06:41**How long this route has been in the table (hh:mm:ss). A tiny value means it just changed. **Ethernet0/2**Outgoing interface. The packet physically leaves here. The two-number bracket trips people up constantly, so commit it to memory: the first number is administrative distance (how much the router trusts the *source*), the second is metric (how good the *path* is according to that source). AD is compared across protocols; metric is only ever compared within one protocol. There is a whole article on this in [administrative distance, the complete table](https://www.pinglabz.com/administrative-distance/), because getting it wrong is how people accidentally black-hole traffic. ## Connected and local: the C/L pair Notice that every connected interface produces two entries: ``` C 10.0.12.0/30 is directly connected, Ethernet0/1 L 10.0.12.1/32 is directly connected, Ethernet0/1 ``` The `C` route is the whole subnet (10.0.12.0/30) reachable out that interface. The `L` route is the router's own IP on that subnet as a /32 host route (10.0.12.1/32). The local route exists so the router can recognize traffic addressed to itself and punt it to the control plane instead of forwarding it. You cannot delete an L route directly; it lives and dies with the interface address. Connected and local both have an implicit administrative distance of 0 and 0 respectively, which is why nothing ever beats a directly connected subnet. ## Equal-cost paths: two next hops, one destination This entry has two lines under it: ``` O 10.0.23.0/30 [110/20] via 10.0.13.2, 00:06:41, Ethernet0/2 [110/20] via 10.0.12.2, 00:06:41, Ethernet0/1 ``` OSPF found two paths to 10.0.23.0/30 with the identical cost of 20, so the router installed both and load-balances across them. This is equal-cost multipath (ECMP). To see how the router actually splits traffic, ask for the route detail: ``` R1#show ip route 10.0.23.0 Routing entry for 10.0.23.0/30 Known via "ospf 1", distance 110, metric 20, type intra area Last update from 10.0.12.2 on Ethernet0/1, 00:06:41 ago Routing Descriptor Blocks: 10.0.13.2, from 3.3.3.3, 00:06:41 ago, via Ethernet0/2 Route metric is 20, traffic share count is 1 * 10.0.12.2, from 3.3.3.3, 00:06:41 ago, via Ethernet0/1 Route metric is 20, traffic share count is 1 ``` The detail view is where the real diagnostic information lives. `Known via "ospf 1"` names the exact process. `type intra area` tells you the route is internal to the OSPF area. Each **Routing Descriptor Block** is one usable next hop; the asterisk marks the path CEF will use for the next flow (it rotates on a per-flow basis). `traffic share count is 1` on both means a 50/50 split. When you are troubleshooting "why is traffic taking that path," `show ip route ` answers it faster than staring at the full table. ## Filtering the table so it fits your brain On a real router the full table can be thousands of lines. Filter it. To see only OSPF-learned routes: ``` R1#show ip route ospf 3.0.0.0/32 is subnetted, 1 subnets O 3.3.3.3 [110/11] via 10.0.13.2, 00:06:42, Ethernet0/2 10.0.0.0/8 is variably subnetted, 9 subnets, 4 masks O 10.0.23.0/30 [110/20] via 10.0.13.2, 00:06:42, Ethernet0/2 [110/20] via 10.0.12.2, 00:06:42, Ethernet0/1 O 10.0.30.0/24 [110/11] via 10.0.13.2, 00:06:42, Ethernet0/2 O 10.0.40.0/24 [110/20] via 10.0.12.2, 00:06:45, Ethernet0/1 ``` Swap in `connected`, `static`, `eigrp`, or `bgp` to isolate any source. When you need the shape of the table rather than the routes themselves, ask for the summary: ``` R1#show ip route summary IP routing table name is default (0x0) IP routing table maximum-paths is 32 Route Source Networks Subnets Replicates Overhead Memory (bytes) connected 0 7 0 784 2184 static 0 1 0 112 312 ospf 1 0 4 0 560 1264 Intra-area: 4 Inter-area: 0 External-1: 0 External-2: 0 eigrp 100 0 2 0 448 624 internal 5 2880 Total 5 14 0 1904 7264 ``` This tells you at a glance that OSPF contributed 4 subnets, EIGRP 2, static 1, and connected 7, and roughly how much memory the table consumes. On a router that is running out of memory or hitting a route-scale limit, this is the first place to look. ## What the table does not tell you The routing table (the RIB, or Routing Information Base) is the control-plane picture: the best routes each protocol offered, after AD and metric were applied. It is not what actually forwards packets in hardware. That job belongs to CEF (the Forwarding Information Base). Usually they agree, but when you are chasing a forwarding bug, compare them with `show ip cef `. If a route is in the RIB but CEF points somewhere else, you have found something worth investigating. This distinction matters most when [longest prefix match](https://www.pinglabz.com/longest-prefix-match/) picks a more specific route than you expected. ## Key takeaways Read the routing table as four layers: the codes legend (memorize C, L, S, D, O, B), the gateway-of-last-resort line, the classful parent headers (summaries, not forwarding entries), and the coded child routes that actually forward. In each route entry the bracket is \[administrative distance / metric\], the `via` is the next hop, and the trailing interface is where the packet leaves. Use `show ip route ` for the detail block whenever you need to know exactly why a path was chosen, and `show ip route ` to cut a huge table down to what you care about. Everything else in routing (choosing between sources, picking the most specific prefix, building failover) reads off these same fields, and the [IP Routing complete guide](https://www.pinglabz.com/ip-routing/) ties them together. ### CDP vs LLDP: What Each Protocol Sees and When to Use Both URL: https://www.pinglabz.com/cdp-vs-lldp/ Last updated: 2026-07-10T23:15:44.000Z Every switch and router quietly announces itself to its directly connected neighbors, and you can read those announcements to map a network without a single diagram. Two protocols do this job: CDP, Cisco's own Discovery Protocol, and LLDP, the vendor-neutral IEEE standard. They overlap heavily, they run side by side, and knowing when to use which (and when to turn both off) is a genuine CCNA and day-job skill. This article compares them on a live Cisco IOS XE lab, showing the exact output each produces on the same wire. It sits in the [Network Fundamentals guide](https://www.pinglabz.com/network-fundamentals/). ## What both protocols do CDP and LLDP are Layer 2 neighbor-discovery protocols. Each device periodically multicasts a frame out every enabled interface, advertising facts about itself: its name, the port the frame left, its platform and software, its management IP, and more. Neighbors cache what they hear and age it out if the advertisements stop. Because the frames are Layer 2 and never routed, they only ever reveal *directly connected* devices, which is exactly what makes them trustworthy for physical topology mapping. Neither protocol carries user data or affects forwarding; they are pure control-plane chatter. ## The core difference, seen on one wire Here is the teaching moment. Router R2 connects to three neighbors (R1, R3, and switch SW1). CDP is on by default everywhere, so R2 sees all of them: ``` R2#show cdp neighbors Capability Codes: R - Router, T - Trans Bridge, B - Source Route Bridge S - Switch, H - Host, I - IGMP, r - Repeater, P - Phone, D - Remote, C - CVTA, M - Two-port Mac Relay Device ID Local Intrfce Holdtme Capability Platform Port ID R1.pinglabz.lab Eth 0/1 136 R Linux Uni Eth 0/1 R3.pinglabz.lab Eth 0/3 135 R Linux Uni Eth 0/3 R3.pinglabz.lab Eth 0/2 136 R Linux Uni Eth 0/1 SW1.pinglabz.lab Eth 0/0 177 R S I Linux Uni Eth 0/0 Total cdp entries displayed : 4 ``` Four entries, including two separate links to R3 (Eth 0/2 and Eth 0/3), because CDP reports per-interface. Now the same router's LLDP table: ``` R2#show lldp neighbors Capability codes: (R) Router, (B) Bridge, (T) Telephone, (C) DOCSIS Cable Device (W) WLAN Access Point, (P) Repeater, (S) Station, (O) Other Device ID Local Intf Hold-time Capability Port ID SW1.pinglabz.lab Et0/0 120 B,R Et0/0 Total entries displayed: 1 ``` Only one neighbor. Why? LLDP is **not on by default**; it must be enabled globally with `lldp run`. In this lab only R2 and SW1 have it enabled, so R2 sees SW1 over LLDP but not R1 or R3 (which never enabled it). This is the practical difference in a nutshell: CDP works between Cisco devices out of the box, while LLDP has to be switched on but then works across every vendor that supports the standard. In a mixed-vendor network, the device you cannot see over CDP is often perfectly visible over LLDP, and vice versa. ## Reading the detail: what a neighbor actually tells you The summary is for mapping; the detail view is for troubleshooting. CDP detail for the SW1 link: ``` R2#show cdp neighbors Ethernet0/0 detail ------------------------- Device ID: SW1.pinglabz.lab Entry address(es): IP address: 10.0.20.10 Platform: Linux Unix, Capabilities: Router Switch IGMP Interface: Ethernet0/0, Port ID (outgoing port): Ethernet0/0 Holdtime : 176 sec Version : Cisco IOS Software [IOSXE], Linux Software (X86_64BI_LINUX_L2-ADVENTERPRISEK9-M), Version 17.18.2 ... VTP Management Domain: '' Native VLAN: 10 Duplex: full Management address(es): IP address: 10.0.20.10 ``` In one command you learned the neighbor's management IP (10.0.20.10), its software version, its native VLAN (10), and its duplex (full). That native-VLAN field alone catches a classic misconfiguration: if two ends of a trunk disagree on native VLAN, CDP logs a mismatch. LLDP detail carries much the same, in standardized TLV form: ``` R2#show lldp neighbors Ethernet0/0 detail ------------------------------------------------ Local Intf: Et0/0 Chassis id: aabb.cc00.1600 Port id: Et0/0 Port Description: TO-R2 System Name: SW1.pinglabz.lab System Description: Cisco IOS Software [IOSXE], Linux Software ..., Version 17.18.2 System Capabilities: B,R Enabled Capabilities: B,R Management Addresses: IP: 10.0.20.10 Auto Negotiation - supported, enabled Physical media capabilities: 1000baseT(FD) Vlan ID: 10 ``` LLDP structures its data as TLVs (Type-Length-Value fields): System Name, Port Description, Management Address, and so on. It also carries the neighbor's Port Description ("TO-R2"), which is the interface description you configured, a handy sanity check that you are patched where you think you are. ## Timers and how they differ Both protocols advertise on a repeating timer and hold what they learn for a multiple of it. The defaults differ, and the difference matters for how quickly a dead neighbor ages out: CDP Advertises every **60 seconds**, holdtime **180 seconds**. Cisco-proprietary, on by default, uses multicast MAC 0100.0ccc.cccc. LLDP Advertises every **30 seconds**, holdtime **120 seconds**. IEEE 802.1AB standard, off by default, uses multicast MAC 0180.c200.000e. You can see the holdtimes counting down in the earlier output: CDP entries near 177 (of 180), LLDP near 120\. If a neighbor stops advertising, CDP forgets it after up to 180 seconds and LLDP after up to 120\. The traffic counters confirm both are actively running: ``` R2#show cdp traffic CDP counters : Total packets output: 84, Input: 81 CDP version 2 advertisements output: 84, Input: 81 R2#show lldp traffic LLDP traffic statistics: Total frames out: 92 Total frames in: 23 ``` ## The security angle: when to turn discovery off Everything that makes discovery useful to you makes it useful to an attacker. A device plugged into an untrusted port can passively read CDP or LLDP frames and immediately learn your device names, software versions, management IPs, native VLAN, and platform, a ready-made reconnaissance handout. The hardening rules are simple: - **Disable discovery on untrusted and edge-facing ports.** Use `no cdp enable` and `no lldp transmit` / `no lldp receive` per interface on anything facing users, the internet, or another organization. - **Keep it on for infrastructure links** where the operational value (mapping, troubleshooting, and voice-VLAN assignment for IP phones) outweighs the risk. - **Turn CDP off globally** with `no cdp run` if your policy forbids it entirely, but remember that some features (Cisco IP phones, some PoE negotiation) rely on it. LLDP-MED, an extension, is what lets a Cisco phone learn its voice VLAN from the switch, so a blanket disable can break IP telephony; scope it to the ports that need it. ## Configuration quick reference ``` ! Enable LLDP globally (CDP is already on by default) Switch(config)# lldp run ! Disable CDP on one untrusted interface Switch(config)# interface Ethernet0/5 Switch(config-if)# no cdp enable ! Disable LLDP transmit and receive on that interface Switch(config-if)# no lldp transmit Switch(config-if)# no lldp receive ! Turn CDP off everywhere Switch(config)# no cdp run ``` ## The switch's own view confirms the asymmetry Look at the same link from SW1's side to close the loop. Over CDP, SW1 sees R2 (both run CDP by default): ``` SW1#show cdp neighbors Device ID Local Intrfce Holdtme Capability Platform Port ID R2.pinglabz.lab Eth 0/0 135 R Linux Uni Eth 0/0 Total cdp entries displayed : 1 ``` And over LLDP, SW1 also sees R2, because these are the two devices where LLDP was enabled: ``` SW1#show lldp neighbors Device ID Local Intf Hold-time Capability Port ID R2.pinglabz.lab Et0/0 120 R Et0/0 Total entries displayed: 1 ``` Discovery is bidirectional but only where both ends participate. R2 and SW1 see each other over both protocols; R1 and R3 (CDP only) are invisible to LLDP entirely. In a real audit this is how you discover that a third-party switch is present but silent on CDP: enable LLDP and it appears. ## Which protocol should you actually run? The pragmatic answer for most enterprises is both, on infrastructure links only. CDP gives you rich Cisco-to-Cisco detail and drives features like IP-phone power negotiation; LLDP gives you visibility into the switches, servers, APs, and phones that are not Cisco. Running both costs almost nothing (a frame every 30 to 60 seconds) and means your topology tools and your NMS see the whole picture regardless of vendor. The one firm rule is the security one: whichever protocols you run, disable them on ports that face users, the internet, or another organization, because the detail that helps you map the network helps an attacker just as much. ## Key takeaways CDP and LLDP both map directly connected neighbors, but CDP is Cisco-only and on by default (60/180-second timers) while LLDP is the IEEE standard, off until you run `lldp run` (30/120-second timers). On a mixed-vendor network you often need both; the lab above showed R2 seeing three neighbors over CDP but only the one LLDP-enabled switch over LLDP. Use the detail views to pull a neighbor's management IP, version, native VLAN, and duplex in a single command, and disable discovery on any untrusted or edge-facing port because the same data is a gift to an attacker. For the broader topology and switching context, head back to the [Network Fundamentals guide](https://www.pinglabz.com/network-fundamentals/). ### Verify IP Parameters on Windows, macOS, and Linux URL: https://www.pinglabz.com/verify-ip-parameters-windows-macos-linux/ Last updated: 2026-07-10T18:38:05.000Z Half of "the network is down" tickets end at the client. Before touching a switch, verify four things on the host itself: its IP address and mask, its default gateway, its DNS servers, and whether it can reach that gateway. CCNA objective 1.10 asks you to do this on Windows, macOS, and Linux. This guide gives you the exact commands for all three - with the Linux output captured live from a real Debian machine in our [Network Fundamentals](https://www.pinglabz.com/network-fundamentals/) lab. ## The four questions, in order 1. **Do I have an address?** A real one - not 169.254.x.x, which means DHCP failed and the host self-assigned (APIPA). 2. **Do I have a default gateway on my own subnet?** Wrong-subnet gateways happen with fat-fingered static configs. 3. **Can I reach the gateway?** One ping isolates the problem to "my segment" or "beyond it." 4. **Do I have working DNS?** "The internet is down" while pings to 1.1.1.1 succeed is a DNS problem, every time. ## Linux: real output Modern Linux uses the `ip` suite (`ifconfig` is legacy). From our lab VM: ``` j@llmbits:~$ ip -br addr show lo UNKNOWN 127.0.0.1/8 ::1/128 ens192 UP 192.168.88.156/24 ens224 UP 169.254.147.59/16 192.168.99.100/24 2001:db8:99::100/64 j@llmbits:~$ ip route default via 192.168.88.1 dev ens192 proto dhcp src 192.168.88.156 metric 1002 10.0.10.0/24 via 192.168.99.1 dev ens224 192.168.88.0/24 dev ens192 proto dhcp scope link src 192.168.88.156 192.168.99.0/24 dev ens224 proto kernel scope link src 192.168.99.100 j@llmbits:~$ cat /etc/resolv.conf | grep nameserver nameserver 45.90.28.181 nameserver 45.90.30.181 ``` Three flags in that real output worth reading like an engineer: - `ip -br addr` (brief mode) is the fastest overview - state and addresses per interface, one line each. Note ens224 carries *both* a 169.254 APIPA address and a real static one: the APIPA appeared because no DHCP answers on that lab segment. Harmless here, but on a single-NIC client, 169.254-only means "DHCP failed" and your investigation moves to the DHCP server or the VLAN. - `ip route` answers the gateway question. `proto dhcp` tells you the route was learned, not typed; the `default via` line is the gateway. Per-prefix static routes (the 10.0.10.0/24 line) show this host reaches lab networks via a different NIC - multihomed hosts route per-destination, a frequent source of "works for some destinations" tickets. - DNS lives in `/etc/resolv.conf` (or `resolvectl status` on systemd-resolved distros, where resolv.conf may just point at a local stub). Then prove gateway reachability - here against our lab router: ``` j@llmbits:~$ ping -c 3 192.168.99.1 PING 192.168.99.1 (192.168.99.1) 56(84) bytes of data. 64 bytes from 192.168.99.1: icmp_seq=1 ttl=255 time=2.20 ms 64 bytes from 192.168.99.1: icmp_seq=3 ttl=255 time=1.71 ms --- 192.168.99.1 ping statistics --- 3 packets transmitted, 3 received, 0% packet loss j@llmbits:~$ ip neigh show dev ens224 192.168.99.1 lladdr aa:bb:cc:00:12:00 DELAY ``` `ip neigh` is the ARP table - the gateway's MAC resolved, which proves Layer 2 adjacency even if ICMP were filtered. A gateway stuck at `FAILED`/`INCOMPLETE` means a Layer 2 problem: wrong VLAN, dead port, or the gateway simply isn't there. ## Windows: the command set ``` C:\> ipconfig /all <- address, mask, gateway, DHCP server, DNS, lease times C:\> ping 192.168.1.1 C:\> nslookup pinglabz.com C:\> arp -a <- did the gateway's MAC resolve? C:\> route print <- full routing table when multihomed ``` The details that matter on Windows: plain `ipconfig` omits DNS and DHCP info - use `/all`. "Autoconfiguration IPv4 Address 169.254.x.x" is the APIPA tell. `ipconfig /release` and `/renew` retry DHCP; `ipconfig /flushdns` clears the resolver cache that keeps "it's still broken" alive after you've fixed DNS. PowerShell equivalents (`Get-NetIPConfiguration`, `Get-NetRoute`, `Test-NetConnection`) return objects and are what you script with. ## macOS: the command set ``` $ ifconfig en0 <- address and mask (still current on macOS) $ netstat -rn | head -5 <- default gateway $ scutil --dns | grep nameserver <- effective DNS servers $ networksetup -getinfo "Wi-Fi" <- the GUI's view, scriptable $ ping -c 3 192.168.1.1 ``` macOS is BSD underneath: `ifconfig` and `netstat -rn` remain the native tools (no `ip` command out of the box). `scutil --dns` shows the resolvers actually in use, which can differ from what the GUI displays when VPNs push split DNS - a modern gotcha worth knowing on any OS. ## Same questions, three dialects **Address + mask**Win: `ipconfig /all` · macOS: `ifconfig en0` · Linux: `ip -br addr` **Default gateway**Win: `ipconfig /all` or `route print` · macOS: `netstat -rn` · Linux: `ip route` **DNS servers**Win: `ipconfig /all` · macOS: `scutil --dns` · Linux: `resolv.conf` / `resolvectl` **ARP / L2 proof**Win: `arp -a` · macOS: `arp -a` · Linux: `ip neigh` ## FAQ ### What does a 169.254.x.x address mean? APIPA / link-local: the host asked DHCP for an address, got no answer, and self-assigned from 169.254.0.0/16\. It can talk to other link-local hosts on the same segment and nothing else. Diagnosis moves upstream: is the DHCP server up, is the port in the right VLAN, is the DHCP relay (`ip helper-address`) configured on the SVI? ### Why does ping to an IP work but browsing fail? DNS. If `ping 1.1.1.1` succeeds but `ping google.com` can't resolve, the host's resolver settings are wrong, the DNS server is down, or something between them blocks UDP/53\. Verify with `nslookup` (Windows), `dig` (Linux/macOS), or `resolvectl query` \- and remember VPN clients love to rewrite DNS silently. ### Is ifconfig deprecated on Linux? Yes - it's part of the legacy net-tools package, frozen for years and absent from minimal installs. Use `ip addr`, `ip route`, `ip neigh`, and `ss` (for sockets). macOS is different: its `ifconfig` is the BSD one and remains the standard tool there. ### How do I check for a duplicate IP address on the network? Symptoms first: intermittent connectivity that follows no pattern, and hosts logging address-conflict warnings. Confirm from another machine: `arp -a` before and after pinging the suspect IP - if the MAC in the ARP entry flips between two values, two devices claim the address. On the gateway, `show ip arp` plus the MAC table walks you to both ports. ### Which command shows the default gateway on each OS? Windows: `ipconfig /all` (or `route print`, first 0.0.0.0 entry). macOS: `netstat -rn`, the `default` line. Linux: `ip route`, the `default via` line. In every case, sanity-check that the gateway is inside the host's own subnet - a mask typo puts it outside, and everything off-subnet dies while local traffic works. ## Key takeaways - Verify in order: address, gateway, gateway reachability, DNS. Each step halves the search space. - 169.254.x.x anywhere means DHCP failed on that interface - move the investigation upstream. - Linux: `ip -br addr`, `ip route`, `ip neigh`. Windows: `ipconfig /all` plus `arp -a`. macOS: BSD tools plus `scutil --dns`. - An ARP entry for the gateway proves Layer 2 even when ping is filtered. Next steps: [How Ping Works](https://www.pinglabz.com/how-ping-works/) for what those echoes actually do, [Lab ts-nf-01](https://www.pinglabz.com/ccna-ts-nf-01-connectivity-tickets/) to practice the whole workflow on tickets, and the [Network Fundamentals guide](https://www.pinglabz.com/network-fundamentals/) for the domain map. ### Virtualization Fundamentals: VMs, Containers, and VRFs URL: https://www.pinglabz.com/virtualization-fundamentals-vms-containers-vrfs/ Last updated: 2026-07-12T02:29:36.000Z CCNA objective 1.12 asks for "virtualization fundamentals: server virtualization, containers, and VRFs" - three technologies that all answer the same question at different layers: *how do you run multiple isolated things on one piece of hardware?* Servers virtualize the machine, containers virtualize the operating system, and VRFs virtualize the router. This guide covers all three, and demonstrates the VRF case with something that looks impossible: the same IP address configured twice on one router, live from our [Network Fundamentals](https://www.pinglabz.com/network-fundamentals/) lab. ## Server virtualization: slicing the machine A hypervisor presents virtual hardware - vCPUs, vRAM, vNICs, vDisks - so that multiple complete operating systems share one physical server, each convinced it owns the box. Two flavors: - **Type 1 (bare metal):** the hypervisor IS the OS on the hardware - VMware ESXi, Hyper-V, KVM, Proxmox. This is what runs data centers (and, for what it's worth, the CML server this lab lives on runs on ESXi). - **Type 2 (hosted):** an application on a normal OS - VMware Workstation, VirtualBox. Labs and desktops. The networking angle the exam cares about: every hypervisor contains a **virtual switch**. VM traffic to another VM on the same host never touches your physical network - a fact with real consequences for where you can capture packets and enforce ACLs. The physical NIC becomes an uplink/trunk, and VLANs extend into the host. Your access layer no longer ends at the switchport; it ends inside the server. ## Containers: slicing the operating system A container doesn't boot its own OS. It's a set of processes on the host's kernel, isolated by namespaces (its own view of processes, filesystem, network stack) and limited by cgroups (CPU/memory caps). No emulated hardware means containers start in milliseconds and pack far more densely than VMs - the tradeoff being weaker isolation (shared kernel) and Linux-only workloads. **Virtual machine** Full OS per instance · minutes to boot · GB of overhead · strong isolation (own kernel) · any OS **Container** Shares host kernel · milliseconds to start · MB of overhead · namespace isolation · Linux processes Containers are already in your network gear: the hosts in our CML labs are Docker containers, and IOS XE itself can run containerized apps on Catalyst switches. When a container needs to talk to your network, it usually does so through NAT on its host or via a bridge - one more virtual switch to keep in your mental topology. ## VRF: slicing the router Virtual Routing and Forwarding gives one physical router multiple independent routing tables. An interface assigned to a VRF exists only in that VRF's world: its routes, its ARP, its forwarding decisions. Traffic cannot cross VRFs unless you deliberately leak routes. It's the routing equivalent of what VLANs did to switches. The classic proof is overlapping addresses. On our lab router, two customer VRFs each use 172.16.1.0/24 - simultaneously: ``` R1# show vrf Name Default RD Protocols Interfaces BLUE ipv4 Lo10 RED ipv4 Lo20 R1# show ip route vrf BLUE Routing Table: BLUE 172.16.0.0/16 is variably subnetted, 2 subnets, 2 masks C 172.16.1.0/24 is directly connected, Loopback10 L 172.16.1.1/32 is directly connected, Loopback10 R1# show ip route vrf RED Routing Table: RED 172.16.0.0/16 is variably subnetted, 2 subnets, 2 masks C 172.16.1.0/24 is directly connected, Loopback20 L 172.16.1.1/32 is directly connected, Loopback20 ``` Same prefix, same host address, one router, zero conflict - because "the routing table" is now "a routing table per VRF." The global table (plain `show ip route`) doesn't contain 172.16.1.0/24 at all. When troubleshooting a VRF'd network, every command needs the qualifier: `ping vrf BLUE 172.16.1.1`, `show ip arp vrf BLUE`. Forgetting the keyword and concluding "the route is missing" is the classic VRF beginner error. Where you'll meet VRFs: service providers separating customers (the mechanism under [MPLS L3VPN](https://www.pinglabz.com/mpls/)), enterprises separating guest/IoT/PCI traffic end-to-end (VRF-lite), out-of-band management (the `Mgmt-vrf` that IOS XE puts your management port in by default), and every [SD-WAN](https://www.pinglabz.com/sd-wan/) segment you've ever configured. ## One mental model for all three Each technology takes a resource that used to be singular and makes it plural behind an isolation boundary: hypervisors do it to hardware, namespaces do it to the kernel, VRFs do it to the RIB. In every case the network engineer's job is the same - know where the virtual switches and virtual tables are, because packets now make forwarding decisions in places you can't physically point at. ## FAQ ### What is the difference between a VRF and a VLAN? Layer. A VLAN partitions a switch's Layer 2 domain (separate broadcast domains, one MAC table per VLAN); a VRF partitions a router's Layer 3 domain (separate routing tables). They pair naturally: VLAN 20 carries guest traffic to the router, where interface VLAN 20 sits in the GUEST VRF so guest routes never mix with corporate ones. End-to-end isolation needs both. ### What is VRF-lite? VRFs without MPLS. Full VRF deployments in service providers use MPLS and route distinguishers to carry many customers across a shared core; VRF-lite is the enterprise version - the same per-VRF routing tables, extended between devices by dedicating an interface (or 802.1Q subinterface) per VRF. Simpler, works everywhere, scales to a handful of segments rather than thousands. ### Can two VRFs on the same router talk to each other? Not by default - that isolation is the whole point. When you need controlled crossings (shared services like DNS or internet breakout), you either leak specific routes between VRFs (`route-target` import/export with BGP, or static routes pointing across) or hairpin through a firewall that has a leg in each VRF. Deliberate, auditable, and the firewall option is what security teams usually prefer. ### Is a container a lightweight VM? Tempting shorthand, architecturally wrong. A VM virtualizes hardware and boots its own kernel; a container is processes on the *host's* kernel wearing isolation (namespaces + cgroups). That's why containers start in milliseconds, why a kernel exploit threatens every container on the host, and why you can't run Windows containers on a Linux kernel. ### Why is my server's VM invisible to my packet capture? Because VM-to-VM traffic on the same host switches inside the hypervisor's vSwitch and never reaches your physical port. To see it you need a capture inside the host (vSwitch port mirroring, or a capture VM) or force the traffic through the physical network. When a "network problem" involves two VMs, always ask first whether they share a host. ## Key takeaways - Type 1 hypervisors run on bare metal (ESXi, KVM); Type 2 run as apps. Both embed virtual switches that extend your access layer into the server. - Containers share the host kernel: faster and denser than VMs, weaker isolation. They ride bridges and NAT on their hosts. - A VRF is a separate routing table on one router - overlapping IP space becomes legal, and every diagnostic command needs the `vrf` keyword. - All three are isolation multiplexers. Find the virtual switch/table and the topology makes sense again. Related: [VLAN vs VXLAN](https://www.pinglabz.com/vlan-vs-vxlan/) for isolation at Layer 2 scale, and the [Network Fundamentals guide](https://www.pinglabz.com/network-fundamentals/) for the rest of domain 1. This covers the fundamentals; the CCNP-level treatment of overlays and fabrics is in the [Network Virtualization and Overlays cluster](https://www.pinglabz.com/network-virtualization/): [VXLAN](https://www.pinglabz.com/vxlan-deep-dive/) at the byte level, [LISP](https://www.pinglabz.com/lisp-explained/) with live captures, [hypervisor virtual switching](https://www.pinglabz.com/hypervisors-virtual-switching-network-engineers/), and [choosing the right segmentation layer](https://www.pinglabz.com/vrf-vlan-vxlan-lisp-segmentation/). ### Power over Ethernet (PoE): Standards, Classes, and the Budget That Bites URL: https://www.pinglabz.com/power-over-ethernet-poe/ Last updated: 2026-07-10T18:38:04.000Z Every access point, IP phone, and security camera on your network probably has no power cable. Power over Ethernet delivers DC power over the same twisted pair carrying data, and CCNA objective 1.3 expects you to know the concepts: the standards, the wattage classes, how negotiation works, and what happens when a switch runs out of power budget. Here's the practical version. ## Why PoE exists The alternative to PoE is an electrician. Every AP on a ceiling, every camera on a pole, every phone on a desk would need a mains outlet installed next to it. PoE turns power delivery into a patching problem instead of a construction problem, and it centralizes power backup: put the switch on a UPS and every phone stays up through an outage - the reason VoIP deployments made PoE mainstream. ## The standards and their budgets **802.3af - PoE (Type 1)** 15.4W at the port, \~12.95W at the device after cable loss. Two pairs. Phones, basic APs, most cameras. **802.3at - PoE+ (Type 2)** 30W at the port, \~25.5W delivered. Two pairs. Wi-Fi 5/6 APs, PTZ cameras, video phones. The current baseline for access switches. **802.3bt Type 3 - PoE++ / UPoE** 60W at the port, \~51W delivered. All four pairs energized. Wi-Fi 6/6E APs, laptops-over-PoE, digital signage. **802.3bt Type 4 - UPoE+** 90W at the port, \~71W delivered. Four pairs. LED lighting systems, thin clients, small switches powered by upstream switches. The port-versus-delivered distinction matters: copper has resistance, and the standard guarantees delivered wattage at 100 meters. Spec sheets quote the bigger number; the device sees the smaller one. ## How a port decides to send power The switch (PSE, power sourcing equipment) will not blast 30 watts into a random laptop NIC. Negotiation happens in stages: 1. **Detection:** the PSE applies a tiny probe voltage looking for the 25kΩ signature resistor that all powered devices (PDs) carry. No signature, no power - which is why plugging your laptop into a PoE port is safe. 2. **Classification:** a second probe reads the PD's class (0 through 8), a rough "how much do I need" declaration. Class 3 asks for 802.3af levels, class 4 for PoE+, classes 5-8 for the 802.3bt tiers. 3. **Negotiation refinement:** after link-up, CDP or LLDP lets the device request its actual operating wattage - a Cisco AP might classify at 30W but settle to 21W once booted. This is why `show power inline` often shows less than the class maximum. ## The power budget: where PoE bites Every PoE switch has a total power supply budget, and it's rarely ports-times-maximum. A 48-port switch with a 740W supply can do 15.4W on all 48 ports (740 ÷ 48), but only \~24 ports of full 30W PoE+. The switch allocates on request and refuses new PDs once the budget is spent - so the 25th access point simply doesn't power on, even though the port works fine for data. On Cisco gear the commands to know are: ``` show power inline ! per-port draw, class, and remaining budget show power inline detail power inline static max 30000 ! reserve wattage for a critical device ``` Design rules of thumb: total your PD wattage at their class maximums, add headroom for growth, and check the budget *before* the Wi-Fi refresh doubles every AP's draw - the Wi-Fi 6 upgrade that dies mysteriously at the 30th AP is a budget problem, not a wireless problem. (This is a concept article by design - our CML lab is virtual and can't source real power, so no captured output here; the commands above are the ones to run on physical Catalyst gear.) ## Odds and ends the exam likes - **Inline power is DC**, nominally around 54V, current-limited - not mains electricity down your patch cable. - **Injectors and splitters** retrofit PoE onto non-PoE links for one device at a time; fine for a lab, messy at scale. - **Perpetual/fast PoE** keeps or restores power during a switch reload so cameras and badge readers don't blink when you upgrade IOS. - **Power policing** (`power inline police`) errdisables a port drawing more than it negotiated - worth enabling where third parties plug things in. ## FAQ ### What is the difference between PoE and PoE+? Wattage and class. PoE (802.3af, Type 1) sources up to 15.4W per port, delivering about 12.95W after cable loss; PoE+ (802.3at, Type 2) doubles that to 30W sourced / 25.5W delivered. PoE+ ports are backward compatible - an af device on an at port simply classifies lower and draws what it needs. ### Can PoE damage a non-PoE device? No, and this is by design. The PSE probes for the 25kΩ signature resistor before applying meaningful power. A laptop, printer, or old switch without the signature never receives voltage beyond the harmless detection probe. Passive PoE injectors (non-standard, always-on) are the exception - those can cook things, which is why standards-based gear is worth insisting on. ### Why won't my access point power on when the port shows connected? Usual suspects in order: the switch's remaining power budget can't cover the AP's class (check `show power inline` \- budget, not port count, is the limit); the AP needs PoE+/802.3bt but the switch is af-only; a long or marginal cable can't deliver the wattage; or the port has a static power cap set below the AP's requirement. Data link-up only proves the pairs carry signal, not watts. ### How far can PoE run? Same as Ethernet: 100 meters over copper, and the standards guarantee delivered wattage at that distance. PoE extenders and PoE-over-fiber solutions (media converter with local power injection) handle longer runs like parking-lot cameras. ### What is UPoE and is it the same as 802.3bt? UPoE was Cisco's pre-standard 60W four-pair implementation; UPoE+ pushed 90W. 802.3bt standardized the same power levels as Type 3 (60W) and Type 4 (90W), and current Catalyst gear implements the standard. Treat the Cisco names as historical labels for bt-class power. ## Key takeaways - af = 15.4W, at = 30W, bt = 60/90W. Two pairs up to PoE+, four pairs for 802.3bt. - Detection (signature resistor) then classification (class 0-8) then CDP/LLDP refinement. Non-PD devices never receive power. - The per-switch power budget, not the per-port maximum, is what limits real deployments. `show power inline` is the tool. - Delivered wattage is lower than port wattage - the spec guarantees delivery at 100m of cable. Related: [Networking Interfaces and Cables Explained](https://www.pinglabz.com/networking-interfaces-and-cables-explained/) for the physical layer underneath, [the WLC guide](https://www.pinglabz.com/what-is-a-wireless-lan-controller/) for what you're usually powering, and the [Network Fundamentals guide](https://www.pinglabz.com/network-fundamentals/) for the domain map. ### Private IPv4 Addressing (RFC 1918): The Blocks, the Rules, the NAT URL: https://www.pinglabz.com/private-ipv4-rfc-1918/ Last updated: 2026-07-10T18:38:04.000Z Three address blocks appear in virtually every network you will ever touch: 10.0.0.0/8, 172.16.0.0/12, and 192.168.0.0/16\. They're defined in RFC 1918, they're the reason IPv4 survived twenty-five years past its predicted exhaustion, and understanding *why they need NAT* is CCNA objective 1.7\. This article covers the blocks, the rules, and a live packet's journey from a private host through PAT to a real machine - captured from our [Network Fundamentals](https://www.pinglabz.com/network-fundamentals/) lab. ## The problem RFC 1918 solved IPv4 has 4.3 billion addresses; the internet has far more devices. In 1996, RFC 1918 formalized the workaround: reserve blocks that anyone may use internally, on the condition that **they are never routed on the public internet**. Your 192.168.1.0/24 and a million other households' 192.168.1.0/24 can coexist because none of them exist beyond their own edge router. Uniqueness is only required where the packet travels. ## The three blocks **10.0.0.0/8** 10.0.0.0 - 10.255.255.255 16.7M addresses. The enterprise workhorse - big enough to carve a global addressing plan out of one block. **172.16.0.0/12** 172.16.0.0 - 172.31.255.255 1M addresses. The one people get wrong: it's /12, so 172.32.x.x is public. Common in labs, Docker, and mid-size shops. **192.168.0.0/16** 192.168.0.0 - 192.168.255.255 65K addresses. Home and small office default - every consumer router ships with a slice of it. Exam trap worth drilling: **172.16.0.0/12 runs only through 172.31.255.255**. If a question shows 172.33.10.5 and asks whether it's private, it isn't. Also don't confuse RFC 1918 space with 169.254.0.0/16 (APIPA link-local, what a host self-assigns when DHCP fails) or 100.64.0.0/10 (carrier-grade NAT space) - reserved, but not RFC 1918. ## Private addresses need a translator A packet sourced from 10.0.10.11 can leave your network, but no internet router will carry the *reply* toward a destination in unroutable space. So the edge router rewrites the source: Network Address Translation. With PAT (Port Address Translation, NAT overload), thousands of inside hosts share one public address, distinguished by port numbers - which is how your entire household shares the single IP your ISP assigns. ## Watching it happen: real output In the lab, host H1 (10.0.10.11, private) pings a real Linux machine across the router's outside interface. R1 translates on the way through: ``` H1:~# ping -c 4 192.168.99.100 PING 192.168.99.100 (192.168.99.100) 56(84) bytes of data. 64 bytes from 192.168.99.100: icmp_seq=1 ttl=63 time=6.20 ms 64 bytes from 192.168.99.100: icmp_seq=2 ttl=63 time=4.18 ms --- 192.168.99.100 ping statistics --- 4 packets transmitted, 4 received, 0% packet loss R1# show ip nat translations Pro Inside global Inside local Outside local Outside global icmp 192.168.99.1:1025 10.0.10.11:8 192.168.99.100:8 192.168.99.100:1025 ``` Read the columns: *inside local* is the host's real private address (10.0.10.11); *inside global* is what the outside world sees (the router's own interface address, port 1025). The destination never learns 10.0.10.11 exists. The four NAT column names are core CCNA vocabulary - learn them from a real entry, not a diagram. ``` R1# show ip nat statistics Total active translations: 1 (0 static, 1 dynamic; 1 extended) Outside interfaces: Ethernet0/0 Inside interfaces: Ethernet0/1 Hits: 72 Misses: 0 Dynamic mappings: -- Inside Source [Id: 1] access-list NAT-INSIDE interface Ethernet0/0 overload ``` The config behind it is three moves: mark the inside interface, mark the outside interface, and one overload rule matching an ACL of the private space. Full walkthroughs in [Lab ips-05 - NAT Overload (PAT)](https://www.pinglabz.com/ccna-lab-ips-05-nat-overload-pat/). ## Consequences worth knowing - **Private space is why your addressing plan matters.** Two companies merge, both used 10.0.0.0/8 casually, and now someone owns months of readdressing or NAT hairpins. Carve deliberately (our [VLSM guide](https://www.pinglabz.com/ipv4-subnetting-vlsm/) shows how). - **NAT is not a firewall**, but it does break inbound-by-default, which is why unsolicited connections need port forwarding or static NAT. - **Leaked RFC 1918 routes are a real failure mode** \- ISPs filter them, and seeing 10.x routes at an internet edge means someone's redistribution went wrong. - **IPv6 removes the scarcity problem entirely** \- no NAT required - which is why the pressure valve of RFC 1918 is also the reason IPv6 adoption took decades. See [IPv6 Address Types](https://www.pinglabz.com/ipv6-address-types/) for the parallel concepts (ULA fdxx:: space is IPv6's spiritual successor to RFC 1918). ## FAQ ### Is 172.32.0.0 a private IP address? No. The private block is 172.16.0.0/12, which spans 172.16.0.0 through 172.31.255.255\. Anything at 172.32 and beyond is public space belonging to someone. This is the most-missed private-addressing question on the CCNA, which is exactly why it keeps appearing. ### Is 169.254.x.x a private address? It's reserved, but it is not RFC 1918\. 169.254.0.0/16 is link-local (APIPA) - what a host assigns itself when DHCP fails. Seeing it on a client means "investigate DHCP," not "someone chose this addressing." Same story for 100.64.0.0/10, which is carrier-grade NAT space reserved for ISP internals. ### Can I route RFC 1918 addresses between my own sites? Absolutely - "not routable" means *on the public internet*. Inside your own network, across VPNs, MPLS L3VPNs, and private WANs, RFC 1918 space routes like any other. Enterprises run global WANs entirely in 10.0.0.0/8\. The prohibition is on internet carriage, where providers filter these prefixes (and where you should too, at your edge). ### Why do VPNs break when both sides use 192.168.1.0/24? Because uniqueness is required where packets travel, and a VPN joins two address spaces. If home and office both use 192.168.1.0/24, the client can't tell which side owns a destination. Fixes: renumber one side (best), or NAT the tunnel (ugly, common). It's also why labs and offices deliberately avoid default consumer ranges. ### Does IPv6 have private addresses like RFC 1918? The equivalent is Unique Local Addresses (fc00::/7, in practice fd00::/8), locally generated and not globally routed. The philosophical difference: IPv6's abundance means ULA is a choice for isolation, not a workaround for scarcity - most IPv6 deployments give every host a global address and control reachability with policy instead of NAT. Details in our [IPv6 address types guide](https://www.pinglabz.com/ipv6-address-types/). ## Key takeaways - 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16 - free to use inside, never routed on the internet. - The /12 boundary: 172.16 through 172.31 only. - Private addressing works because NAT/PAT rewrites sources at the edge; know all four translation column names. - APIPA (169.254/16) and CGN (100.64/10) are reserved but are not RFC 1918. Continue with the [Network Fundamentals guide](https://www.pinglabz.com/network-fundamentals/) or get hands-on with the [IP services lab track](https://www.pinglabz.com/ccna-labs-ip-services/). ### 2-Tier vs 3-Tier vs Spine-Leaf: Network Topology Architectures URL: https://www.pinglabz.com/network-topology-architectures/ Last updated: 2026-07-10T18:38:03.000Z Ask why a campus network has an "access layer" or why data centers abandoned the tree design entirely, and you're asking about topology architectures - CCNA exam objective 1.2 and, more usefully, the vocabulary every design conversation assumes you know. This guide covers the two-tier and three-tier campus models, spine-leaf in the data center, and where SOHO, WAN, and cloud designs fit. ## Why architecture models exist You could connect switches in any shape that links up. Architecture models exist because some shapes fail predictably: they contain broadcast storms poorly, they force traffic through chokepoints, and nobody can troubleshoot at 3 a.m. what nobody can draw. The hierarchical models give each device a defined role, which is what makes networks scalable and debuggable. ## The three-tier campus model Cisco's classic hierarchy, for large campuses (multiple buildings, thousands of users): **Access layer** Where endpoints plug in. Port density, PoE for phones and APs, port security, 802.1X. One access switch failing strands one closet, not a building. **Distribution layer** Aggregates access switches per building. Routing boundary (SVIs live here), policy, ACL enforcement, summarization. Deployed in redundant pairs. **Core layer** Interconnects distribution pairs. One job: forward packets as fast as possible. No policy, no ACLs - complexity in the core hurts everything below it. The design rule that makes it work: **access never connects to access**, and everything dual-homes upward (each access switch to both distribution switches, each distribution pair to both cores). Failure of any single box or link costs capacity, not connectivity. ## The two-tier (collapsed core) model Most networks are not giant. A two-tier design merges core and distribution into one redundant pair - the *collapsed core* \- with access switches hanging off it. Same principles, one less layer of boxes to buy. This is the right answer for a single building or a small multi-floor campus, and it's the design you'll actually deploy most often in the mid-market. If the exam asks when two-tier is appropriate: when the network is small enough that a separate core adds cost without adding meaningful scalability. ## Spine-leaf: the data center answer Campus hierarchies optimize *north-south* traffic (users to servers/internet). Data centers flipped: most traffic became *east-west* (server to server - microservices, storage replication, VM migration). Spine-leaf optimizes for that: - **Every leaf connects to every spine.** Leaves are top-of-rack switches where servers plug in; spines only interconnect leaves. Leaves never connect to leaves, spines never to spines. - **Predictable latency:** any server to any server is exactly leaf-spine-leaf. Two hops, always. - **Scale horizontally:** need more capacity? Add a spine. More racks? Add leaves. Bandwidth grows linearly, and equal-cost multipath (ECMP) load-shares across every spine simultaneously - no spanning-tree blocked links wasting half your uplinks. This is the fabric underneath VXLAN overlays and Cisco ACI, and the "underlay" you keep hearing about in [SD-WAN](https://www.pinglabz.com/sd-wan/) and SDN conversations - see our [VLAN vs VXLAN](https://www.pinglabz.com/vlan-vs-vxlan/) explainer for what runs on top. ## SOHO, WAN, on-prem and cloud The remaining 1.2 sub-objectives are quick, but the exam does test them: - **SOHO:** the small office/home office collapses every role - router, switch, AP, firewall - into one box. Architecturally interesting precisely because there is no architecture: all layers, one device. - **WAN:** connects sites. Traditional MPLS hub-and-spoke gave every branch a private path back to HQ; internet VPN and [SD-WAN](https://www.pinglabz.com/sd-wan/) mesh sites over commodity links with policy deciding per-application paths. Topologically: hub-and-spoke, full mesh, or partial mesh, trading circuit cost against branch-to-branch latency. - **On-prem vs cloud:** the same tiers exist in the cloud, rebadged - a VPC is your distribution block, availability zones are your redundant pairs, and the provider owns the spine-leaf you never see. Hybrid designs connect your campus to that via VPN or dedicated interconnect, which makes the WAN design part of your LAN conversation. ## Choosing: the decision in practice **Single building, <500 users**Two-tier collapsed core **Multi-building campus**Three-tier, distribution per building **Data center / server fabric**Spine-leaf with ECMP **Branch / home office**SOHO all-in-one, WAN back to the mothership ## FAQ ### What is the difference between two-tier and three-tier architecture? Three-tier separates access, distribution, and core into distinct layers; two-tier merges distribution and core into one "collapsed core" pair. Functionally identical policy and redundancy - the third tier only earns its cost when multiple distribution blocks (buildings) need interconnecting at scale. ### Why doesn't spanning tree run the data center anymore? STP prevents loops by blocking redundant links - in a spine-leaf fabric that would idle half the capacity you paid for. Spine-leaf runs routing (often BGP) to every leaf and uses ECMP to load-share across all spines at once. Loop prevention comes from the routing protocol, not from blocking ports. Campus access still uses [STP](https://www.pinglabz.com/spanning-tree-protocol/), so learn both. ### Can spine-leaf be used in a campus network? The pattern appears in campus as routed access designs and in SD-Access fabrics, which borrow the underlay/overlay split. But classic campus traffic is north-south with policy at aggregation points, so two/three-tier remains the default answer - and the exam keeps spine-leaf associated with data centers. ### What layer do firewalls and WLCs attach to? Services attach where their traffic aggregates: internet edge firewalls beside the core or in a dedicated edge block, east-west/data-center firewalls at the DC boundary, and [WLCs](https://www.pinglabz.com/what-is-a-wireless-lan-controller/) traditionally at distribution/services (with 9800s, often virtualized). The principle: policy devices sit at choke points the architecture already created. ### How many devices before three-tier makes sense? There's no magic number - the trigger is distribution blocks, not device count. When you have three or more distribution pairs (buildings/large floors) meshing to each other, a dedicated core turns N-squared inter-building links into N links to the core, and earns its boxes. ## Key takeaways - Three-tier = access, distribution, core; two-tier collapses core into distribution. Role separation is the whole point. - Spine-leaf: every leaf to every spine, two hops between any two servers, ECMP instead of blocked links - built for east-west traffic. - Redundancy is designed in pairs and dual-homing, so failures cost capacity rather than connectivity. - Know the SOHO, WAN, and cloud framings - objective 1.2 lists them explicitly. Related reading: [VLAN, Subnet, and Broadcast Domain](https://www.pinglabz.com/vlan-subnet-broadcast-domain-difference/) for what these layers segment, and the [Network Fundamentals guide](https://www.pinglabz.com/network-fundamentals/) for the full domain-1 map. ### Duplex Mismatch and Interface Errors: Find It in 10 Seconds URL: https://www.pinglabz.com/duplex-mismatch-interface-errors/ Last updated: 2026-07-10T18:38:03.000Z "The network is slow" tickets that turn out to be a duplex mismatch are a rite of passage. The link is up, pings work, and yet throughput is terrible and retransmissions are everywhere. This article shows you how to catch the problem in seconds - including the exact syslog message Cisco devices print when CDP spots a mismatch (captured live from our [Network Fundamentals](https://www.pinglabz.com/network-fundamentals/) lab) - and how to read every error counter in `show interfaces`. ## What a duplex mismatch actually is Full duplex means both ends transmit whenever they like; half duplex means an end waits for silence and treats simultaneous transmission as a collision. When one side runs full and the other half, the full side transmits freely while the half side keeps detecting "collisions" and backing off, then retransmitting. Result: the link works, badly. Light traffic flows fine; under load, the half side logs collisions and late collisions while the full side counts CRC and runt errors from frames the other end aborted mid-transmission. The classic cause is autonegotiation failing or being half-configured: one side hardcoded to full, the other left on auto. An auto side that can't negotiate falls back to **half duplex** (at 10/100 Mbps). Hardcoding one side only is how you manufacture this problem. ## The 10-second diagnosis: let CDP tell you We forced the mismatch in the lab: R1's Ethernet0/1 at full duplex, the switch port facing it set to half. Within one CDP cycle, both devices called it out by name: ``` R1# show logging | include DUPLEX *Jul 10 18:24:36.246: %CDP-4-DUPLEX_MISMATCH: duplex mismatch discovered on Ethernet0/1 (not half duplex), with SW1 Ethernet0/0 (half duplex). SW1# show logging | include DUPLEX *Jul 10 18:24:50.596: %CDP-4-DUPLEX_MISMATCH: duplex mismatch discovered on Ethernet0/0 (not full duplex), with R1 Ethernet0/1 (full duplex). ``` Read that message carefully - it names the local interface, the neighbor, the neighbor's interface, and both duplex settings. If both ends are Cisco and CDP is on, the box literally hands you the diagnosis. Check `show logging` before you check anything else. ## Verifying duplex directly ``` SW1# show interfaces status Port Name Status Vlan Duplex Speed Type Et0/0 TO-R1 connected 10 full auto 10/100/1000BaseTX Et0/1 TO-H1 connected 10 full auto 10/100/1000BaseTX Et0/2 TO-H2 connected 10 full auto 10/100/1000BaseTX ``` One line per port with duplex and speed - this is the fastest inventory view. The mismatch case shows `a-half` or `half` on one side and `full` on the neighbor's. On the mismatched interface itself, `show interfaces Ethernet0/0` prints the operational state in the first few lines (our lab port reported `Half-duplex` while misconfigured, `Full-duplex, Auto-speed` once fixed). ## Reading the counters that matter Here's the full counter block from R1's LAN interface, healthy, so you know the baseline: ``` R1# show interfaces Ethernet0/1 Ethernet0/1 is up, line protocol is up Hardware is AmdP2, address is aabb.cc00.1210 (bia aabb.cc00.1210) Internet address is 10.0.10.1/24 MTU 1500 bytes, BW 10000 Kbit/sec, DLY 1000 usec, reliability 255/255, txload 1/255, rxload 1/255 ... 540 packets input, 283237 bytes, 0 no buffer Received 210 broadcasts (0 IP multicasts) 0 runts, 0 giants, 0 throttles 0 input errors, 0 CRC, 0 frame, 0 overrun, 0 ignored 381 packets output, 268981 bytes, 0 underruns 0 output errors, 0 collisions, 2 interface resets 0 babbles, 0 late collision, 0 deferred 0 lost carrier, 0 no carrier ``` The counters, decoded by what they usually mean: **CRC / input errors**Corrupted frames received. Bad cable, EMI, failing SFP, or the far side of a duplex mismatch. **runts**Frames under 64 bytes - collision fragments, classic on the full side of a mismatch. **giants**Frames over the MTU - usually an MTU/encapsulation config problem, not physics. **collisions**Normal on genuine half duplex; on a supposedly full-duplex link, a red flag. **late collision**Collision after the first 64 bytes. Almost always duplex mismatch (or a cable run beyond spec). The single most diagnostic counter. **deferred**Port waited for silence before transmitting - the half side yielding under load. **interface resets**The interface bounced - counts config changes (ours shows 2 from the duplex experiment) as well as carrier loss. The signature pair to memorize: **late collisions on the half-duplex side, CRC/runts on the full-duplex side**. See both across one link and you're done diagnosing. (In our CML lab the virtual wire can't produce genuine electrical collisions, so we show you the healthy baseline and the CDP catch; on physical gear, load the mismatched link and watch those two counters climb.) ## Fixing and preventing it ``` SW1(config)# interface Ethernet0/0 SW1(config-if)# speed auto SW1(config-if)# duplex auto ``` - **Auto on both ends** is the modern best practice - gigabit and faster require autonegotiation anyway, and it negotiates duplex correctly when both sides participate. - Never hardcode just one end. If policy demands hardcoding (some old server NICs), hardcode *both* ends identically. - A counter check means nothing without a time base: `clear counters Ethernet0/0`, generate traffic, re-check. "0 late collisions since Tuesday" and "0 since 30 seconds ago under load" are different facts. ## FAQ ### Why does a duplex mismatch still pass ping? Ping is a trickle - one small frame at a time, with gaps. Collisions need simultaneous transmission, so light traffic mostly squeaks through. The pathology only shows under sustained bidirectional load (file transfers, backups, VoIP), which is exactly why users report "slow" rather than "down" and why the ticket survives the first tech's ping test. ### Can gigabit links have duplex mismatches? Effectively no - 1000BASE-T requires autonegotiation, and gigabit is full duplex in practice. The mismatch problem lives at 10/100 with hardcoded ports: old servers, door controllers, medical devices, industrial gear. That legacy tail is why the skill stays on the CCNA. ### What causes late collisions besides duplex mismatch? Cable runs beyond the 100m spec (or repeaters extending a collision domain beyond timing limits) and bad NICs that violate carrier sense. In a modern switched full-duplex network, though, a nonzero late-collision counter is a duplex mismatch until proven otherwise. ### Do CRC errors always mean a bad cable? No - CRC counts any corrupted frame. Bad or too-long cable, EMI (that run past the elevator motor room), a failing SFP, a dying NIC, or the far side of a duplex mismatch aborting frames mid-transmission. Correlate with the neighbor's counters: CRC on one side plus late collisions on the other = mismatch; CRC alone rising on both = physical medium. ### Should I disable CDP if it reveals topology? On untrusted/user-facing edges, many shops do (or use it selectively) for hardening. But note what you give up: this article's entire 10-second diagnosis. Common compromise: CDP/LLDP on infrastructure links where the mismatch alarm matters, disabled on user access ports. More in our [CDP vs LLDP lab](https://www.pinglabz.com/ccna-lab-na-11-cdp-vs-lldp/). ## Key takeaways - Duplex mismatch = link up, ping fine, throughput awful. Suspect it early. - `show logging | include DUPLEX` \- CDP names the mismatch, both interfaces included. - Late collisions point at the half side; CRC and runts point at the far side of the same problem. - Auto/auto everywhere, or hardcode both ends. Half-configured negotiation is the root cause. Practice the full L1/L2 workflow in [Lab nf-11 - Troubleshooting Layer 1/2/3 Symptoms](https://www.pinglabz.com/ccna-lab-nf-11-troubleshooting-layer-symptoms/), and find every fundamentals topic on the [Network Fundamentals guide](https://www.pinglabz.com/network-fundamentals/). ### How Switches Learn MAC Addresses: The CAM Table Explained URL: https://www.pinglabz.com/how-switches-learn-mac-addresses/ Last updated: 2026-07-10T18:38:03.000Z Every frame a switch forwards is a lookup in one table: the MAC address table (historically the CAM table). Understanding how entries get into that table, how long they stay, and what happens when a destination isn't in it explains half of Layer 2 behavior - flooding, unicast delivery, even why some attacks work. This article walks the whole lifecycle with real output from a Cisco IOS XE switch in our [Network Fundamentals](https://www.pinglabz.com/network-fundamentals/) lab: two Linux hosts (H1 and H2) and a router hanging off a switch in VLAN 10. ## The three rules of switching A switch applies three rules to every frame, in order: 1. **Learn:** read the *source* MAC of the incoming frame and record it against the ingress port and VLAN. 2. **Forward:** look up the *destination* MAC. If there's an entry, send the frame out that one port (filtering it from all others). 3. **Flood:** if there's no entry (unknown unicast), or the destination is broadcast/unknown multicast, send the frame out every port in the VLAN except the one it arrived on. That's it. A $100,000 chassis and a lab switch both do exactly this; everything else is optimization. ## Watching the table learn After H1 (10.0.10.11) pinged H2 (10.0.10.12) and the gateway, here's the table: ``` SW1# show mac address-table dynamic Mac Address Table ------------------------------------------- Vlan Mac Address Type Ports ---- ----------- -------- ----- 10 5254.0065.04be DYNAMIC Et0/2 10 5254.007a.17d9 DYNAMIC Et0/1 10 aabb.cc00.1210 DYNAMIC Et0/0 10 ba3c.5334.f483 DYNAMIC Et0/1 10 eab6.3ad2.6a78 DYNAMIC Et0/2 Total Mac Addresses for this criterion: 5 ``` Read it like the switch does: "if a frame is destined to 5254.007a.17d9 in VLAN 10, send it out Ethernet0/1." Note there are five MACs on three ports - Et0/1 and Et0/2 each show two addresses because the Linux hosts also emitted frames from a second (container-side) interface MAC. Multiple MACs per port is completely normal: that's what you see behind IP phones, hypervisors, or a downstream switch. ## Counting and capacity ``` SW1# show mac address-table count Mac Entries for Vlan 10: --------------------------- Dynamic Address Count : 5 Static Address Count : 0 Total Mac Addresses : 5 ``` Table capacity is finite (thousands to hundreds of thousands of entries depending on platform). That finiteness is why MAC flooding attacks exist: fill the table with garbage source MACs and the switch degrades to flooding everything, turning it into a hub an attacker can sniff. That's the problem [port security](https://www.pinglabz.com/securing-your-network-with-cisco-port-security-a-comprehensive-guide/) solves. ## Aging: why entries disappear ``` SW1# show mac address-table aging-time Global Aging Time: 300 ``` Dynamic entries age out after 300 seconds of silence from that source MAC by default. Every new frame from the MAC resets its timer. Aging matters for a practical reason: when a host moves ports (laptop roams, VM migrates, cable gets repatched), the stale entry either ages out or is overwritten the moment the host transmits from its new port. The learn rule always wins - the table tracks the *latest* port a source MAC appeared on. ## Flush and relearn: the demo Clear the table and it's empty for a moment; the very next traffic rebuilds it: ``` SW1# clear mac address-table dynamic SW1# show mac address-table dynamic Mac Address Table ------------------------------------------- Vlan Mac Address Type Ports ---- ----------- -------- ----- ! H1 pings H2 - two packets is all it takes: SW1# show mac address-table dynamic Vlan Mac Address Type Ports ---- ----------- -------- ----- 10 5254.0065.04be DYNAMIC Et0/2 10 5254.007a.17d9 DYNAMIC Et0/1 10 aabb.cc00.1210 DYNAMIC Et0/0 Total Mac Addresses for this criterion: 3 ``` Think about what happened to the first frame after the clear: H1's ICMP echo arrived with an unknown destination (H2's MAC wasn't in the table), so the switch flooded it out Et0/0 and Et0/2\. H2 replied; that reply's source MAC populated Et0/2, and from the second frame on, traffic between H1 and H2 was filtered unicast - the router port never saw it again. Flooding is self-limiting: it lasts exactly one round-trip per destination. ## Where ARP fits (and where it doesn't) Students constantly conflate the MAC table with ARP. Keep them straight: **ARP is a host function** mapping IP addresses to MAC addresses; the **MAC table is a switch function** mapping MAC addresses to ports. The switch doesn't need ARP to switch frames, and your PC doesn't have a MAC-to-port table. They cooperate - an ARP broadcast is flooded by rule 3, and conveniently teaches the switch the sender's MAC by rule 1 - but they are different tables on different devices answering different questions. ## Troubleshooting with the MAC table - **"Is the host even talking?"** No dynamic entry for its MAC = the switch has heard nothing in 5 minutes. Check the cable, the NIC, the VLAN. - **"Which port is this device on?"** `show mac address-table address ` \- the fastest way to physically locate anything in a building. - **"Why is this port showing hundreds of MACs?"** Downstream unmanaged switch, hypervisor - or a MAC flooding attack in progress. ## FAQ ### What is the difference between the CAM table and the MAC address table? Same thing, two names. CAM (Content Addressable Memory) is the hardware that stores the table on real switches - memory you query by content (the MAC) and get back a port in one operation. Cisco documentation and commands say "MAC address table"; engineers say CAM. Use either. ### Why does one switch port show multiple MAC addresses? Anything that puts multiple sources behind one port: a downstream switch, a hypervisor full of VMs, an IP phone with a PC daisy-chained, or containers with their own MACs (exactly what our lab hosts showed). It's normal - unless a port that should hold one printer suddenly shows fifty MACs, which smells like a loop or an attack. ### What happens when the MAC address table is full? New source MACs can't be learned, so frames destined to them are flooded in their VLAN - the switch behaves like a hub for those hosts. MAC flooding attacks (macof) exploit this deliberately to enable sniffing, which is why [port security](https://www.pinglabz.com/securing-your-network-with-cisco-port-security-a-comprehensive-guide/) caps MACs per port. ### How do I find which switch port a device is connected to? `show mac address-table address aaaa.bbbb.cccc` on the switch. If the port it returns is an uplink to another switch, hop there and repeat - two or three hops walks you to the exact access port anywhere in a campus. Get the MAC itself from the device's ARP entry on its gateway (`show ip arp | include x.x.x.x`). ### Do switches use ARP to build the MAC table? No. The switch learns passively from the source MAC of every frame - any frame, not just ARP. ARP is how *hosts* map IPs to MACs. An ARP broadcast happens to teach every switch in the VLAN the sender's location as a side effect, but the switch never sends or interprets ARP to switch frames. ## Key takeaways - Learn on source, forward on destination, flood on unknown - in that order, every frame. - Dynamic entries age out after 300 seconds; any frame from the source resets the timer. - Unknown-unicast flooding is normal and brief; the first reply converts it to filtered unicast. - ARP maps IP to MAC on hosts; the MAC table maps MAC to port on switches. Different tables, different devices. Go deeper: [Lab na-01 - Switching Fundamentals and the CAM Table](https://www.pinglabz.com/ccna-lab-na-01-switching-fundamentals-cam/) gives you this exact topology to break and rebuild, and the [VLAN cluster](https://www.pinglabz.com/vlans-layer-2-switching/) covers what happens when you segment the table per VLAN. Start from the [Network Fundamentals guide](https://www.pinglabz.com/network-fundamentals/). ### IPv4 Subnetting and VLSM: The Method That Works Under Pressure URL: https://www.pinglabz.com/ipv4-subnetting-vlsm/ Last updated: 2026-07-10T18:38:03.000Z Subnetting is the one CCNA skill you cannot fake. Every other topic gives you partial credit for concepts; subnetting questions have exactly one right answer, and the exam expects you to find it in under a minute. The good news: subnetting is not math talent, it's a repeatable procedure. This guide teaches the block-size method, walks through VLSM planning the way you'd do it on a real network, and verifies the whole thing with real Cisco IOS XE output from our [Network Fundamentals](https://www.pinglabz.com/network-fundamentals/) lab. ## What subnetting actually does A subnet mask splits a 32-bit IPv4 address into a network portion and a host portion. Moving the boundary right (borrowing host bits) creates more subnets with fewer hosts each. That's the entire game. Everything else - slash notation, wildcard masks, VLSM - is bookkeeping around that one idea. Two formulas cover every question: **Subnets created** 2borrowed bits (bits taken from the host portion) **Usable hosts per subnet** 2host bits \- 2 (subtract network and broadcast addresses) ## The block-size method Forget binary long-hand under exam pressure. The block size is 256 minus the interesting octet of the mask, and subnet boundaries are multiples of that block size. Procedure: 1. Find the interesting octet (the mask octet that isn't 255 or 0). 2. Block size = 256 minus that octet's value. 3. Count up from zero in block-size steps. The address lives in the block it falls inside. 4. Network = block start. Broadcast = next block minus 1\. Usable range = everything between. Example: which subnet does 10.0.32.27 255.255.255.240 belong to? The interesting octet is 240, so the block size is 256 - 240 = 16\. Blocks start at .0, .16, .32, .48\. The address 10.0.32.27 falls in the .16 block? No - look at the fourth octet: 27 falls between 16 and 32, so the network is 10.0.32.16/28, broadcast is 10.0.32.31, and usable hosts run .17 through .30\. We assigned exactly that subnet to a loopback in the lab, and you'll see IOS agree with us below. ## The mask cheat table **/25** \= .128 Block 128 · 126 hosts **/26** \= .192 Block 64 · 62 hosts **/27** \= .224 Block 32 · 30 hosts **/28** \= .240 Block 16 · 14 hosts **/29** \= .248 Block 8 · 6 hosts **/30** \= .252 Block 4 · 2 hosts Memorize this table cold. On the exam, recognizing that /28 means "block of 16, 14 usable hosts" without thinking is the difference between 40 seconds and 4 minutes per question. ## VLSM: sizing subnets to fit Variable Length Subnet Masking just means using different masks for different needs inside the same address space. The rule that keeps you out of trouble: **allocate the largest subnets first**, then fill in smaller ones behind them so nothing overlaps. Here's the actual plan from our lab, carved from 10.0.0.0/16: **10.0.10.0/24**User LAN (VLAN 10) - 254 hosts, room to grow **10.0.64.0/26**Server segment - needs \~50 hosts, /26 gives 62 **10.0.32.16/28**Management network - 14 usable, plenty for device SVIs **10.0.99.4/30**Point-to-point link - exactly 2 hosts, zero waste **10.0.0.1/32**Router loopback - one address, one host route Five different masks, one address space, no overlaps. A /24 for users because user segments grow; a /30 for the router-to-router link because it will only ever hold two devices (if you gave it a /24 you'd waste 252 addresses). ## What the router sees: real output We configured that exact plan on a Cisco IOS XE router in CML and pulled the routing table. Note the line IOS prints: *"variably subnetted, 9 subnets, 5 masks"* \- that's VLSM working: ``` R1# show ip route 10.0.0.0/8 is variably subnetted, 9 subnets, 5 masks C 10.0.0.1/32 is directly connected, Loopback0 C 10.0.10.0/24 is directly connected, Ethernet0/1 L 10.0.10.1/32 is directly connected, Ethernet0/1 C 10.0.32.16/28 is directly connected, Loopback2 L 10.0.32.17/32 is directly connected, Loopback2 C 10.0.64.0/26 is directly connected, Loopback1 L 10.0.64.1/32 is directly connected, Loopback1 C 10.0.99.4/30 is directly connected, Loopback3 L 10.0.99.5/32 is directly connected, Loopback3 ``` Read the pairs: every C (connected) route is the subnet itself, and every L (local) route is the /32 for the router's own address inside it. 10.0.32.16/28 is there exactly as the block-size method predicted, with the router holding .17 (the first usable host). ## Worked exam questions **Q1: A host has 172.16.45.14/27\. What is its broadcast address?** Block size 256 - 224 = 32\. Blocks: .0, .32, .64\. The 45 in the third octet? No - /27 puts the interesting octet fourth. Host .14 falls in the 0-31 block, so network 172.16.45.0/27 and broadcast 172.16.45.31. **Q2: You need 6 subnets from 192.168.1.0/24, each with at least 25 hosts. What mask?** 25 hosts needs 5 host bits (2^5 - 2 = 30). That leaves 3 borrowed bits = 8 subnets. Mask /27 (255.255.255.224). Check both constraints - the exam loves masks that satisfy one but not the other. **Q3: Which subnet is 10.0.64.1 255.255.255.192 in?** Block 256 - 192 = 64\. Blocks .0, .64, .128\. Answer: 10.0.64.0/26, usable .65 to .126, broadcast .127\. Exactly our lab server segment. ## FAQ ### What is the fastest way to subnet in the CCNA exam? The block-size method. Compute 256 minus the interesting octet once, then everything - network, broadcast, usable range - falls out of counting in multiples. With the /25 through /30 table memorized, most questions take under a minute. Binary conversion is for understanding, not for exam speed. ### How many usable hosts are in a /26, /28, and /30? /26 = 62 usable (block 64), /28 = 14 usable (block 16), /30 = 2 usable (block 4). The pattern: usable hosts = block size minus 2, because every subnet spends one address on the network ID and one on broadcast. ### What does "variably subnetted" mean in show ip route? It means the router knows subnets of one classful network with more than one mask - VLSM in action. IOS prints the count of subnets and distinct masks, then lists each route with its own prefix length, which is why you should always read the mask per-route rather than assuming one mask per network. ### When should I use a /31 instead of a /30 on point-to-point links? RFC 3021 allows /31s on point-to-point links (2 addresses, 0 wasted), and modern IOS XE supports them - `ip address 10.0.99.0 255.255.255.254`. They halve address burn on link-heavy networks. The CCNA answer is still usually /30; /31 is the real-world optimization worth knowing exists. ### Is subnetting still relevant with IPv6? The mechanics carry over - IPv6 subnetting is arguably easier (sites get a /48 or /56 and carve /64s on nibble boundaries), but the prefix-length thinking is identical. Master IPv4 blocks and [IPv6 prefixes](https://www.pinglabz.com/ipv6-address-examples/) feel familiar rather than foreign. ## Key takeaways - Block size = 256 minus the interesting octet. Subnets start at multiples of the block size. - Usable hosts = 2^host-bits minus 2\. Always check both the subnet-count and host-count constraints. - VLSM = biggest subnets first, then pack smaller ones behind. Point-to-point links get /30 (or /31), loopbacks get /32. - Verify on the box: `show ip route` tells you "variably subnetted, N subnets, M masks" and lists every boundary. Practice this hands-on in [Lab nf-04 - IPv4 Subnetting with VLSM](https://www.pinglabz.com/ccna-lab-nf-04-ipv4-subnetting-vlsm/), and see the full domain map on the [Network Fundamentals guide](https://www.pinglabz.com/network-fundamentals/). ### EIGRP Administrative Distance: 90, 170, and 5 URL: https://www.pinglabz.com/eigrp-administrative-distance/ Last updated: 2026-07-10T13:00:00.000Z EIGRP does not have one administrative distance, it has three. That surprises people who memorized "EIGRP is 90" for an exam, then watched a redistributed route show up as 170 or a summary route as 5\. The three values are not a quirk; each one encodes a deliberate trust decision. This post explains all three EIGRP administrative distances, why they differ, and how to read and change them. For the cluster overview, see the [EIGRP complete guide](https://www.pinglabz.com/eigrp/). ## Administrative distance, in one paragraph A router often learns the same destination from more than one source - say, OSPF and EIGRP both have a route to 10.0.0.0/8\. The metrics of those protocols are not comparable; an OSPF cost and an EIGRP composite metric are different units. Administrative distance (AD) is the tie-breaker. It is a number from 0 to 255 that ranks how much the router *trusts* each routing source. Lower AD wins, and the winning route is the one installed in the routing table. AD is strictly local to one router and is consulted before metric ever enters the picture. ## EIGRP's three administrative distances EIGRP internal AD90 What it is A route learned from another EIGRP router within the same EIGRP autonomous system EIGRP external AD170 What it is A route redistributed into EIGRP from another source (OSPF, BGP, static, a different EIGRP AS) EIGRP summary AD5 What it is The local summary (aggregate) route EIGRP creates when you configure manual summarization One protocol, three numbers, three different meanings. Memorizing "EIGRP equals 90" is only correct for one of the three cases. ## Why internal is 90 AD 90 is EIGRP's default for routes it learned from a fellow EIGRP router. It is a low number on purpose - lower than OSPF's 110, lower than RIP's 120, lower than IS-IS's 115\. When EIGRP and another IGP both offer the same prefix, EIGRP's internal route wins. That reflects Cisco's design intent: EIGRP is meant to be the preferred IGP where it is deployed. ## Why external is 170 AD 170 applies to routes that were **redistributed** into EIGRP from somewhere else. The gap between 90 and 170 is deliberate, and it is about trust. An internal EIGRP route was computed by EIGRP's own DUAL algorithm end to end - EIGRP knows the full picture. An external route arrived from another protocol with a metric EIGRP had to translate, a metric that does not carry the same guarantees. Redistribution is also a classic source of routing loops, especially with mutual redistribution between two protocols at two points. By giving external routes a much worse AD (170, worse than OSPF's 110 and RIP's 120), EIGRP ensures that if a redistributed copy and a natively-learned copy of the same prefix both exist, the native one is preferred. The 170 is a loop-prevention safeguard, not an arbitrary penalty. ## Why the summary route is 5 AD 5 is the one that looks wrong until you see what it is for. When you configure manual summarization on an interface, EIGRP creates a **local summary route** for the aggregate and points it at `Null0`. That summary needs to win against the more specific component routes for the summarization to behave - and AD 5 (lower than internal EIGRP's own 90) guarantees it does. The Null0 summary is also a loop-prevention feature: a packet matching the summary but not any real component falls through to Null0 and is dropped, rather than being forwarded back toward the summarizing router. AD 5 keeps that protective summary firmly in the table. You will see it as a route to Null0 in `show ip route`. ## EIGRP against the other protocols Connected interface 0 Static route 1 EIGRP summary 5 eBGP 20 EIGRP internal 90 OSPF 110 IS-IS 115 RIP 120 EIGRP external 170 iBGP 200 ## Reading it on the router In `show ip route`, the AD is the first number in the bracketed pair. The route code letter tells you the EIGRP flavour: `D` is internal EIGRP, `D EX` is external EIGRP. ``` R1# show ip route eigrp D 10.20.0.0/24 [90/3072] via 10.30.30.2, 00:12:04, GigabitEthernet0/1 D EX 172.20.0.0/16 [170/3328] via 10.30.30.2, 00:09:51, GigabitEthernet0/1 D 10.40.0.0/16 is a summary, 00:14:22, Null0 ``` Read the brackets as `[AD/metric]`. The first line is internal (90), the second is external (170), and the summary route to Null0 carries AD 5. ## Changing the AD EIGRP's administrative distances can be overridden under the routing process with the `distance eigrp` command, which takes the internal value first and the external value second: ``` router eigrp PINGLABZ address-family ipv4 unicast autonomous-system 100 topology base distance eigrp 90 170 ``` Changing AD is occasionally necessary - for example, to make a backup path through another protocol preferred, or to control which route wins during a migration - but it should be deliberate and documented. Inconsistent AD changes across routers are a reliable way to create a routing loop. ## Common gotchas A redistributed route is being ignored in favour of another protocol It is EIGRP external, AD 170 - worse than OSPF (110) and RIP (120). Expected behaviour. An unexpected route to Null0 appears That is the EIGRP summary route, AD 5\. It is a feature of manual summarization, not a fault. EIGRP route wins over a static route you wanted A static route is AD 1 and should win. If it does not, check the static is actually valid and installed. Two EIGRP ASes exchange routes and they show as external Routes redistributed between two EIGRP autonomous systems are external (170), not internal. A loop after redistribution despite AD 170 AD only protects per-router. Mutual redistribution still needs route tagging and filtering. ## Key takeaways EIGRP has three administrative distances, each encoding a trust level. Internal routes are AD 90, low enough to beat OSPF, IS-IS, and RIP, because EIGRP computed them end to end with its own algorithm. External (redistributed) routes are AD 170, deliberately worse than the other IGPs so a native route always beats a redistributed copy - a loop safeguard. The local summary route is AD 5, low enough to beat EIGRP's own internal routes so manual summarization and its protective Null0 route hold. Read the AD as the first number in the `[AD/metric]` bracket of `show ip route`, and change it only deliberately, with `distance eigrp`. For the EIGRP cluster, see the [EIGRP pillar](https://www.pinglabz.com/eigrp/). ### Troubleshooting Slow Network Throughput with iPerf3 URL: https://www.pinglabz.com/troubleshooting-slow-throughput-iperf3/ Last updated: 2026-07-09T05:17:04.000Z "The link is fast but the transfer is slow" is the complaint that never dies. iPerf3 can settle it in minutes, but only if you test methodically - throughput problems have exactly three root causes (not enough capacity, too much latency for the window, or packet loss) and each one leaves a distinct fingerprint in iPerf3 output. To show you those fingerprints, we broke a healthy lab path three different ways on real Cisco routers and captured what iPerf3 reported each time. This article is part of our [complete iPerf guide](https://www.pinglabz.com/iperf/). ## Step 0: Know Your Baseline Our lab path (Linux client, two OSPF-routed IOS XE hops, Linux server) healthy: ``` client:~$ ping -c 5 10.0.20.10 rtt min/avg/max/mdev = 2.948/3.611/5.239/0.833 ms client:~$ iperf3 -c 10.0.20.10 [ 5] 0.00-10.01 sec 41.2 MBytes 34.6 Mbits/sec 83 sender [ 5] 0.00-10.01 sec 40.9 MBytes 34.2 Mbits/sec receiver ``` 34.6 Mbit/sec at 3.6 ms RTT. Without a baseline you cannot tell "degraded" from "normal for this path," so record one per important path while things work (the same discipline we preach for [ping](https://www.pinglabz.com/ping/)). ## Fingerprint 1: A Bandwidth Bottleneck We capped the WAN link at 10 Mbit/sec and re-ran the same test: ``` client:~$ iperf3 -c 10.0.20.10 [ ID] Interval Transfer Bitrate Retr [ 5] 0.00-10.00 sec 13.8 MBytes 11.5 Mbits/sec 0 sender [ 5] 0.00-10.53 sec 11.9 MBytes 9.46 Mbits/sec receiver ``` The signature: throughput sits pinned just under a suspiciously round number, retransmits are modest, and latency is roughly normal. TCP found the shaper and settled beneath it. Confirm by watching the constraining interface while the test runs - on the router mid-test: ``` R1# show interfaces Ethernet0/1 | include rate 5 minute input rate 184000 bits/sec, 206 packets/sec 5 minute output rate 5747000 bits/sec, 471 packets/sec ``` If the interface (or a policer/shaper on it) tops out while your test does, you have found the bottleneck. The fix is capacity or [QoS policy](https://www.pinglabz.com/qos/), not tuning. ## Fingerprint 2: Latency Plus a Small Window Next we removed the cap and injected 50 ms of one-way delay instead. Ping confirms the change; loss stays at zero: ``` client:~$ ping -c 5 10.0.20.10 rtt min/avg/max/mdev = 103.459/103.962/104.532/0.342 ms client:~$ iperf3 -c 10.0.20.10 [ 5] 0.00-10.00 sec 31.4 MBytes 26.3 Mbits/sec 39 sender client:~$ iperf3 -c 10.0.20.10 -w 64K [ 5] 0.00-10.00 sec 6.88 MBytes 5.77 Mbits/sec 11 sender ``` The signature: ping shows high but stable RTT with no loss, autotuned TCP gets close-ish to the baseline, and anything with a constrained window craters in exact proportion to window/RTT (64 KB / 104 ms = about 5 Mbit/sec - the math, the demonstration, and the parallel-stream workaround are in [parallel streams and TCP window size](https://www.pinglabz.com/iperf3-parallel-streams-tcp-window/)). If a real application is slow on this path while iPerf3 is fast, suspect the application's own socket buffers. ## Fingerprint 3: Packet Loss Finally the cruel one: same 104 ms RTT, plus 2% packet loss on the WAN link. ``` client:~$ ping -c 5 10.0.20.10 5 packets transmitted, 5 received, 0% packet loss rtt min/avg/max/mdev = 103.646/104.013/104.175/0.190 ms client:~$ iperf3 -c 10.0.20.10 [ ID] Interval Transfer Bitrate Retr Cwnd [ 5] 0.00-1.00 sec 640 KBytes 5.24 Mbits/sec 5 36.8 KBytes [ 5] 2.00-3.00 sec 128 KBytes 1.05 Mbits/sec 2 18.4 KBytes [ 5] 4.00-5.00 sec 0.00 Bytes 0.00 bits/sec 0 7.07 KBytes [ 5] 8.00-9.00 sec 0.00 Bytes 0.00 bits/sec 1 11.3 KBytes - - - - - - - - - - - - - - - - - - - - - - - - - [ 5] 0.00-10.00 sec 1.75 MBytes 1.47 Mbits/sec 20 sender [ 5] 0.00-10.11 sec 1.50 MBytes 1.25 Mbits/sec receiver ``` Look closely: a five-packet ping saw nothing wrong, yet TCP collapsed to 1.47 Mbit/sec with intervals that transferred zero bytes. Two percent loss at 100 ms RTT is catastrophic for TCP because every loss halves the congestion window and recovery takes a full round trip. The Cwnd column never escapes double digits. Now prove the loss and measure it. UDP does what a short ping cannot: ``` client:~$ iperf3 -c 10.0.20.10 -u -b 20M [ ID] Interval Transfer Bitrate Jitter Lost/Total Datagrams [ 5] 0.00-10.00 sec 23.8 MBytes 20.0 Mbits/sec 0.000 ms 0/17266 (0%) sender [ 5] 0.00-10.10 sec 23.3 MBytes 19.4 Mbits/sec 0.219 ms 366/17266 (2.1%) receiver ``` 17,266 probes in 10 seconds versus ping's 5: the UDP test nails the loss at 2.1%. From here it is regular loss-hunting: check interface counters hop by hop for errors, drops, and duplex mismatches, and see whether loss follows a specific link ([iPerf3 UDP testing](https://www.pinglabz.com/iperf3-udp-testing/) covers the technique in depth). ## The Decision Tree **1\. Pinned under a round number, low Retr** Bandwidth bottleneck. Verify with interface rate counters mid-test. Fix: capacity or QoS. **2\. High stable RTT, no loss, window-sensitive** Latency x window ceiling. Verify: `-w 64K` collapses, `-P 4` recovers. Fix: buffers/BDP tuning, parallelism. **3\. Retr climbing, Cwnd tiny, zero-byte intervals** Packet loss. Verify with `-u` at moderate rate and read the receiver loss %. Fix: find the lossy hop (errors, drops, duplex, overloaded QoS queue). **4\. Only parallel streams reach the ceiling** Per-flow limit: policer, hash pinning on a port-channel/ECMP, or one saturated CPU core. Compare single vs `-P 4` sums. ## Habits That Keep You Honest Test both directions (`-R`) before blaming the network; asymmetry is a clue, not noise. Run 30 to 60 second tests so slow start and transient dips average out. Watch `top` on both endpoints, because a saturated CPU produces numbers that look exactly like a network problem. And when you report a result, quote the receiver line with the RTT alongside it - "34 Mbit/sec at 104 ms RTT" is a measurement, "34 Mbit/sec" is a rumor. ## Key Takeaways Slow throughput has three fingerprints and iPerf3 exposes all of them: a clean pin under a round number means a bandwidth cap, high stable RTT with window sensitivity means a BDP ceiling, and climbing retransmits with a collapsing Cwnd mean loss - which a moderate UDP test will quantify precisely even when ping sees 0%. Baseline your paths while they are healthy, test in both directions, and confirm on the router with interface counters. The complete command reference, server setup, and version guidance all live in the [iPerf complete guide](https://www.pinglabz.com/iperf/). ### Running an iPerf3 Server: Daemon Mode, Ports, and Logging URL: https://www.pinglabz.com/iperf3-server-mode/ Last updated: 2026-07-09T05:17:04.000Z Every iPerf3 test needs a listening server, and how you run that server decides whether testing is a 10-second task or a small project. This article covers foreground and daemon modes, custom ports, logging, the one-client-at-a-time behavior that trips everyone up, and the firewall rules a test server needs - all demonstrated on a live lab server with real output. It is part of our [complete iPerf guide](https://www.pinglabz.com/iperf/). ## Foreground Mode: The Quick Test ``` server:~$ iperf3 -s ----------------------------------------------------------- Server listening on 5201 (test #1) ----------------------------------------------------------- ``` The server binds TCP and UDP port 5201 on all addresses and prints each test's results to the terminal as it happens. This is fine for a one-off, but it dies with your SSH session. Two useful companions: `-p` to change the port and `-1` (one-off mode) to exit automatically after a single test, which is handy in scripts. ## Daemon Mode: The Persistent Server ``` server:~$ iperf3 -s -D --logfile /var/log/iperf3.log server:~$ ss -ltn | grep 5201 LISTEN 0 4096 *:5201 *:* ``` The `-D` flag backgrounds the process; `--logfile` captures what would have gone to the terminal (without it, a daemonized server's output simply vanishes). The log shows the server-side view of every test, which is often the half of the story people forget to look at: ``` server:~$ tail /var/log/iperf3.log [ 6] 7.00-8.00 sec 1.12 MBytes 9.44 Mbits/sec [ 6] 8.00-9.00 sec 1.12 MBytes 9.43 Mbits/sec [ 6] 9.00-10.00 sec 1.12 MBytes 9.45 Mbits/sec - - - - - - - - - - - - - - - - - - - - - - - - - [ ID] Interval Transfer Bitrate [ 6] 0.00-10.53 sec 11.9 MBytes 9.46 Mbits/sec receiver ``` ## One Test at a Time (and the Busy Error) An iPerf3 server process runs exactly one test at a time. Here is what a second client sees while a test is in progress - captured from a real client hitting our lab server mid-test: ``` j@llmbits:~$ iperf3 -c 10.0.20.10 iperf3: error - the server is busy running a test. try again later ``` This is by design (iPerf2 behaves differently; see [iPerf2 vs iPerf3](https://www.pinglabz.com/iperf2-vs-iperf3/)). The standard workaround is one server process per port: ``` server:~$ iperf3 -s -D --logfile /var/log/iperf3-5201.log server:~$ iperf3 -s -p 5202 -D --logfile /var/log/iperf3-5202.log server:~$ ss -ltn | grep 520 LISTEN 0 4096 *:5201 *:* LISTEN 0 4096 *:5202 *:* ``` Clients pick their lane with the same flag: `iperf3 -c 10.0.20.10 -p 5202`. Two processes, two simultaneous tests, no waiting. Teams that share a test box often run four or five daemons on consecutive ports. ## Running iPerf3 as a systemd Service For a permanent test server, let systemd own the process so it survives reboots: ``` # /etc/systemd/system/iperf3.service [Unit] Description=iperf3 server After=network-online.target [Service] ExecStart=/usr/bin/iperf3 -s --logfile /var/log/iperf3.log Restart=on-failure [Install] WantedBy=multi-user.target ``` ``` sudo systemctl daemon-reload sudo systemctl enable --now iperf3 ``` Note there is no `-D` here: systemd wants the process in the foreground so it can supervise it. ## Firewall Rules and Network Access iPerf3 uses a single port (default 5201) for both the TCP control connection and the test traffic, TCP or UDP. Open both protocols on whatever port you chose: ``` # Linux (firewalld) sudo firewall-cmd --add-port=5201/tcp --add-port=5201/udp --permanent sudo firewall-cmd --reload # Linux (ufw) sudo ufw allow 5201 ``` If the path crosses a Cisco device with an ACL, permit the same port pair. A client that hangs at "Connecting to host..." then times out usually means the control connection (TCP 5201) is blocked; a UDP test that starts but reports 100% loss usually means TCP got through and UDP did not - an asymmetric ACL is the first thing to check. ## Useful Server-Side Flags **`-D`**Run as a daemon in the background. Pair with `--logfile` or the output is lost. **`--logfile file`**Append server output to a file. Essential for daemons, useful everywhere. **`-p 5202`**Listen on a different port. One daemon per port = concurrent tests. **`-1`**Handle exactly one test, then exit. Perfect for scripted, on-demand servers. **`-B 10.0.20.10`**Bind to one address on a multihomed host, so tests enter on the interface you intend. **`-I /run/iperf3.pid`**Write a PID file, so scripts and service managers can find the daemon. **`--rsa-private-key-path`**With `--authorized-users-path`, restricts who can run tests (builds with auth support). On an open network, at minimum bind carefully and firewall the port. ## Reading the Server Side During Tests The server's report is not a mirror of the client's. On a TCP test, the client prints sender-side detail (Retr and Cwnd columns), while the server log records the receiver's view - what actually arrived, interval by interval. When a client-side summary looks odd (say, the sender claims 34 Mbit/sec but the transfer felt slower), the server log settles it. On UDP tests the asymmetry is even more valuable: jitter and loss are computed at the server, so the authoritative numbers live in the server output and are echoed back to the client at test end. If you daemonize with `--logfile`, you get a permanent, timestamped record of every test anyone has run against that box - which quietly becomes your best throughput history when a "was it always this slow?" question shows up months later (baseline discipline, same as we recommend in [troubleshooting slow throughput](https://www.pinglabz.com/troubleshooting-slow-throughput-iperf3/)). ## A Word on Exposure An iPerf3 server is a machine that will saturate a link on request - do not leave one listening on an internet-facing interface. Anyone who finds it can burn your bandwidth all day, and historical CVEs have targeted the server side. Keep test servers on management networks, bind them to internal addresses with `-B`, firewall the port to known sources, and prefer `-1` one-shot servers where practical. Public iperf servers exist for WAN sanity checks, but treat results through the internet as indicative, not diagnostic. ## Key Takeaways Use plain `iperf3 -s` for quick checks, `-D --logfile` for persistent daemons, and systemd for anything permanent. One process serves one test at a time, so multiple ports are the answer to the busy error. Open your chosen port for both TCP and UDP, bind deliberately on multihomed hosts, and never expose a test server to the open internet. With the server side squared away, the client-side techniques are covered across the rest of the [iPerf complete guide](https://www.pinglabz.com/iperf/), starting with [how to use iPerf3](https://www.pinglabz.com/how-to-use-iperf3/). ### iPerf2 vs iPerf3: Differences That Actually Matter URL: https://www.pinglabz.com/iperf2-vs-iperf3/ Last updated: 2026-07-09T05:17:04.000Z iPerf2 and iPerf3 share a name, a purpose, and almost nothing else. iPerf3 was a ground-up rewrite by ESnet, not an upgrade, and the two do not interoperate - an iPerf2 client cannot test against an iPerf3 server, and they do not even default to the same port (5001 vs 5201). Both are actively maintained today by different teams, which surprises people who assume "3 replaced 2." This article breaks down the real differences, with live output from iperf 2.2.1 and iperf 3.18, as part of our [complete iPerf guide](https://www.pinglabz.com/iperf/). ## Two Tools, Not Two Versions iPerf2 (currently maintained on SourceForge by Bob McMahon) evolved continuously from the original NLANR tool, gaining latency measurement, enhanced reporting, and serious multithreading. iPerf3 (maintained by ESnet on GitHub) was rewritten for a cleaner codebase, a scripting-friendly design, and a library API (libiperf). The rewrite deliberately dropped features iPerf2 users relied on, and iPerf2 development never stopped. Hence: two parallel tools. ## The Differences That Change Your Results **Threading** iPerf2 is multithreaded: `-P 8` uses multiple cores. Classic iPerf3 runs all streams in a single thread, so on 10G+ links one CPU core often becomes the bottleneck before the network does. Advantage: iPerf2 on very fast links. **Concurrent clients** An iPerf2 server accepts multiple clients at once. iPerf3 handles exactly one test at a time and tells later clients "the server is busy running a test." Advantage: iPerf2 for shared test servers. **Scripting output** iPerf3's `-J` emits complete JSON. iPerf2 offers CSV (`-y`) which is thinner. Advantage: iPerf3 for automation. **Reverse and bidirectional** iPerf3 has clean reverse mode (`-R`) driven entirely from the client. iPerf2 instead has dual/tradeoff modes (`-d` / `-r`) that test both directions in one run, which iPerf3 dropped. Different tools for different workflows. **Multicast** iPerf2 tests multicast (bind the server to a group address, set TTL with `-T`). iPerf3 cannot. Advantage: iPerf2, and it is the only option. **Latency measurement** iPerf2's enhanced mode reports one-way latency, RTT, and derived metrics inline. iPerf3 reports throughput, retransmits, and cwnd, leaving latency to ping. Advantage: iPerf2 for latency-under-load. ## You Can See the Philosophy in the Output Here is real iperf 2.2.1 enhanced output (`-e`): ``` $ iperf -c 127.0.0.1 -t 5 -e ------------------------------------------------------------ Client connecting to 127.0.0.1, TCP port 5001 with pid 12365 (1/0 flows/load) Write buffer size: 131072 Byte TCP congestion control using cubic TOS defaults to 0x0 (dscp=0,ecn=0) (Nagle on) TCP window size: 16.0 KByte (default) ------------------------------------------------------------ [ ID] Interval Transfer Bandwidth Write/Err Rtry InF(pkts)/Cwnd(pkts)/RTT(var) NetPwr [ 1] 0.0000-5.0161 sec 7.70 GBytes 13.2 Gbits/sec 63063/0 0 0K(0)/319K(10)/135(67) us 12206015 ``` Congestion control algorithm, writes and errors, retries, in-flight data, cwnd in packets, RTT with variance, and a NetPwr figure - all in one line. iPerf2 packs telemetry into human-readable text. iPerf3's equivalent run prints the familiar clean Interval/Transfer/Bitrate/Retr/Cwnd block (see [how to use iPerf3](https://www.pinglabz.com/how-to-use-iperf3/) for annotated output) and saves the deep detail for `-J` JSON, where scripts can parse it reliably. ## Feature Support at a Glance **Default port**iPerf2: 5001 | iPerf3: 5201 **Interoperable with each other**No **JSON output**iPerf3 (`-J`); iPerf2 has CSV `-y` **Reverse mode from client**iPerf3 (`-R`) **Simultaneous bidirectional (`-d`)**iPerf2 only (iPerf3 has `--bidir` in newer builds) **Multicast testing**iPerf2 only **Multithreaded parallel streams**iPerf2 (iPerf3 added multithreading only in recent releases) **Multiple clients per server**iPerf2 **Library API (libiperf)**iPerf3 **One-way latency / RTT reporting**iPerf2 (`-e` enhanced mode) **TCP retransmit + cwnd reporting**iPerf3 (Retr/Cwnd columns) ## Flag Collisions: The Silent Gotcha Because the tools evolved separately, several flags mean different things, and muscle memory from one tool produces wrong tests in the other. The dangerous ones: - `-b`: UDP-only bandwidth in iPerf2 (and implies `-u` in old builds); in iPerf3 it rate-limits TCP as well, and UDP needs an explicit `-u`. An iPerf3 user typing `iperf3 -c host -b 10M` gets a rate-limited TCP test, not a UDP test. - `-C`: compatibility mode in iPerf2, congestion control algorithm selection in iPerf3. - `-T`: multicast TTL in iPerf2, output line prefix in iPerf3. - `-Z`: realtime scheduler in some iPerf2 builds, zerocopy send in iPerf3. When reading someone else's test notes, identify the tool before interpreting the flags. ## Interop in Practice If you manage a fleet of test endpoints, standardize hard. A client that hangs while "Connecting to host..." against a server you know is up is very often an iPerf2 client aimed at an iPerf3 port or vice versa; because the port numbers differ (5001 vs 5201) the connection usually just times out or is refused, and nothing in the error says "wrong tool." Baking the tool name and version into your test scripts' output (both tools support `-v`) saves that half hour. Where both tools must coexist on one server, run them on their native default ports and document it; nothing conflicts since they never share a socket. ## Which Should You Use? Default to iPerf3\. It is what most documentation assumes, its JSON output is far better for automation, its Retr/Cwnd columns make TCP behavior visible (invaluable when [troubleshooting slow throughput](https://www.pinglabz.com/troubleshooting-slow-throughput-iperf3/)), and packages are current everywhere. Reach for iPerf2 in four specific cases: multicast testing (no alternative), a shared always-on test server that must accept concurrent clients, latency-under-load measurement with `-e`, and multi-10G testing where iPerf3's single thread caps out a core (or use several iPerf3 processes on different ports as a workaround). Nothing stops you installing both; they coexist happily since they use different ports and binary names. ## Key Takeaways iPerf3 is a rewrite, not an upgrade: cleaner, scriptable, single-test-at-a-time, and the right default in 2026\. iPerf2 remains actively developed and wins on multithreading, multicast, concurrent clients, and inline latency stats. They cannot talk to each other, so match versions across your endpoints and note which tool produced any number you record. For the full command reference and lab-tested examples of both, start at the [iPerf complete guide](https://www.pinglabz.com/iperf/). ### iPerf3 Parallel Streams and TCP Window Size (-P and -w) URL: https://www.pinglabz.com/iperf3-parallel-streams-tcp-window/ Last updated: 2026-07-09T05:17:03.000Z Two iPerf3 flags separate people who run throughput tests from people who understand them: `-P` (parallel streams) and `-w` (TCP window size). Together they explain the most common head-scratcher in network testing - a fast link that benchmarks slow the moment latency enters the picture. We reproduced that exact failure on a live Cisco Modeling Labs topology with 100 ms of injected round-trip time, so every number below is real. This article is part of our [complete iPerf guide](https://www.pinglabz.com/iperf/). ## The Bandwidth-Delay Product, Quickly TCP can only keep a window's worth of unacknowledged data in flight. Once the window is full, the sender stops and waits for ACKs, and ACKs take a round trip. So the ceiling on any single TCP flow is: ``` max throughput = window size / round-trip time ``` Flip it around and you get the bandwidth-delay product (BDP), the window you need to fill a given path: ``` BDP = bandwidth x RTT 35 Mbit/sec x 0.104 sec = ~3.6 Mbit = ~455 KB ``` Our lab path carries about 35 Mbit/sec with 104 ms RTT, so a single flow needs roughly 455 KB of window to fill it. Hold that number; you are about to watch what happens when TCP has far less. ## The Experiment: Same Path, Three Windows First, the path with negligible latency (RTT around 3.6 ms). A default iPerf3 run lands at 34.6 Mbit/sec - the path's natural ceiling. Then we added 50 ms of one-way delay on the WAN link (104 ms RTT, confirmed by ping) and ran three tests. ### Default window, 104 ms RTT ``` client:~$ iperf3 -c 10.0.20.10 [ ID] Interval Transfer Bitrate Retr Cwnd [ 5] 0.00-1.00 sec 5.00 MBytes 41.9 Mbits/sec 5 539 KBytes [ 5] 3.00-4.00 sec 2.75 MBytes 23.1 Mbits/sec 0 318 KBytes [ 5] 6.00-7.00 sec 2.75 MBytes 23.0 Mbits/sec 0 321 KBytes - - - - - - - - - - - - - - - - - - - - - - - - - [ 5] 0.00-10.00 sec 31.4 MBytes 26.3 Mbits/sec 39 sender [ 5] 0.00-10.11 sec 27.8 MBytes 23.0 Mbits/sec receiver ``` Linux autotuning grows the window past 300 KB and recovers most of the throughput: 26.3 Mbit/sec. Not bad. Watch the Cwnd column doing exactly what the BDP math predicts it must. ### Forced 64 KB window, same path ``` client:~$ iperf3 -c 10.0.20.10 -w 64K [ ID] Interval Transfer Bitrate Retr Cwnd [ 5] 0.00-1.00 sec 640 KBytes 5.24 Mbits/sec 0 161 KBytes [ 5] 4.00-5.00 sec 768 KBytes 6.29 Mbits/sec 0 112 KBytes - - - - - - - - - - - - - - - - - - - - - - - - - [ 5] 0.00-10.00 sec 6.88 MBytes 5.77 Mbits/sec 11 sender [ 5] 0.00-10.10 sec 6.88 MBytes 5.71 Mbits/sec receiver ``` Same link, same routers, same everything - 5.77 Mbit/sec. The math checks out almost exactly: 64 KB per 104 ms round trip is about 5 Mbit/sec. Nothing is broken. TCP is simply spending most of its time waiting for ACKs. This is what legacy applications with small fixed socket buffers experience on every long-haul path, and it is why "the link is 1 Gbps but the transfer runs at 50 Mbps" is usually not a network problem at all. ### Same tiny window, four parallel streams ``` client:~$ iperf3 -c 10.0.20.10 -w 64K -P 4 [ ID] Interval Transfer Bitrate Retr [ 5] 0.00-10.00 sec 7.00 MBytes 5.87 Mbits/sec 0 sender [ 7] 0.00-10.00 sec 4.50 MBytes 3.77 Mbits/sec 4 sender [ 9] 0.00-10.00 sec 6.88 MBytes 5.77 Mbits/sec 1 sender [ 11] 0.00-10.00 sec 6.88 MBytes 5.77 Mbits/sec 1 sender [SUM] 0.00-10.00 sec 25.2 MBytes 21.2 Mbits/sec 6 sender [SUM] 0.00-10.11 sec 25.2 MBytes 21.0 Mbits/sec receiver ``` Each stream is still window-starved at roughly 5 Mbit/sec, but four of them together deliver 21.2 Mbit/sec. Parallel streams multiply the effective in-flight data - four windows instead of one. This is exactly why browsers open multiple connections and why backup tools have a "streams" setting. ## What the Comparison Tells You **Default window** 26.3 Mbit/sec. Modern autotuning handles moderate BDPs on its own. If your OS is current and buffers are not capped, you rarely need `-w`. **\-w 64K** 5.77 Mbit/sec. A 78% collapse from a setting, not a fault. Window / RTT is a hard ceiling no amount of bandwidth fixes. **\-w 64K -P 4** 21.2 Mbit/sec. Parallelism buys back what small windows lose. If `-P` helps a lot, suspect per-flow limits: windows, per-flow policers, or single-flow load-balancing hashes. ## How to Use -P and -w Diagnostically Run a single stream, then `-P 4`, and compare: - **Single stream is slow, parallel sum is much higher:** a per-flow limit. Check RTT and window sizes (BDP math above), per-flow QoS policers, or an ECMP/port-channel hash pinning one flow to one member link. - **Single and parallel both hit the same ceiling:** a shared limit - the link itself, a shaper on the path, or CPU on either endpoint. Confirm the interface rate on the router while the test runs (`show interfaces` and watch the output rate counters climb). - **Parallel is worse than single:** usually endpoint CPU. iPerf3 runs all streams in one thread until recent versions, so a saturated core caps the sum. Check with `top` during the test; iPerf2 multithreads if you need to rule this out (see [iPerf2 vs iPerf3](https://www.pinglabz.com/iperf2-vs-iperf3/)). Two cautions with `-w`. First, the OS may silently cap what you request (check the "socket buffer size" line iPerf3 prints). Second, what you set with `-w` is a buffer size request on both ends; setting it lower than autotune would reach makes things worse, so use it to reproduce problems, not as a routine "optimization." ## Key Takeaways A single TCP flow can never exceed window divided by RTT, so know your path's BDP before judging a test. Modern autotuning usually gets close to the ceiling; small fixed windows crater throughput in exact proportion to the math, and parallel streams recover it by flying multiple windows at once. Use `-P` as a diagnostic fork in the road: it separates per-flow limits from shared limits in under a minute. When a path still disappoints after this analysis, packet loss is the usual suspect - pick up the trail in [troubleshooting slow throughput with iPerf3](https://www.pinglabz.com/troubleshooting-slow-throughput-iperf3/), or go back to the [iPerf complete guide](https://www.pinglabz.com/iperf/) for the full command reference. ### iPerf3 UDP Testing: Bandwidth, Jitter, and Packet Loss URL: https://www.pinglabz.com/iperf3-udp-testing/ Last updated: 2026-07-09T05:17:03.000Z TCP tells you how much a path can carry. UDP tells you what the path does to traffic that cannot slow down - voice, video, and anything real-time. iPerf3's UDP mode sends at exactly the rate you specify and reports three things TCP never shows you directly: jitter, packet loss, and out-of-order delivery. Every output block below comes from a live Cisco Modeling Labs topology with real routers in the path. This article is part of our [complete iPerf guide](https://www.pinglabz.com/iperf/). ## Why UDP Testing Is Different TCP adapts. When the network drops a segment, TCP retransmits it and slows down, so a TCP test converges on "whatever the path can sustain." UDP does not adapt. iPerf3 sends datagrams at the rate you set with `-b`, and whatever the network drops stays dropped. That makes UDP the right tool for two jobs: measuring loss along a path (TCP hides loss inside retransmissions), and simulating real-time traffic that behaves the same way. The server detects loss by sequence numbers in the datagrams and computes jitter continuously, using the smoothed mean of differences between consecutive transit times (the same approach RTP uses per RFC 1889). Your clocks do not need to be synchronized; the math subtracts the offset out. ## A Clean UDP Test The client sends 10 Mbit/sec through two routed hops. Note the `-u` for UDP and `-b` for target bandwidth: ``` client:~$ iperf3 -c 10.0.20.10 -u -b 10M [ ID] Interval Transfer Bitrate Total Datagrams [ 5] 0.00-1.00 sec 1.19 MBytes 10.0 Mbits/sec 865 [ 5] 1.00-2.00 sec 1.19 MBytes 10.0 Mbits/sec 863 [ 5] 9.00-10.00 sec 1.19 MBytes 9.99 Mbits/sec 863 - - - - - - - - - - - - - - - - - - - - - - - - - [ ID] Interval Transfer Bitrate Jitter Lost/Total Datagrams [ 5] 0.00-10.00 sec 11.9 MBytes 10.0 Mbits/sec 0.000 ms 0/8634 (0%) sender [ 5] 0.00-10.01 sec 11.9 MBytes 10.0 Mbits/sec 0.206 ms 0/8634 (0%) receiver ``` Read the receiver line: 0 lost out of 8,634 datagrams, 0.206 ms jitter. This path can carry 10 Mbit/sec of real-time traffic cleanly. ## What Happens When You Ask for Too Much Now the same path, but the client pushes 100 Mbit/sec into a path that can only carry about 35: ``` client:~$ iperf3 -c 10.0.20.10 -u -b 100M [ ID] Interval Transfer Bitrate Total Datagrams [ 5] 0.00-1.00 sec 11.9 MBytes 100 Mbits/sec 8646 [ 5] 9.00-10.00 sec 11.9 MBytes 99.9 Mbits/sec 8624 - - - - - - - - - - - - - - - - - - - - - - - - - [ ID] Interval Transfer Bitrate Jitter Lost/Total Datagrams [ 5] 0.00-10.00 sec 119 MBytes 100 Mbits/sec 0.000 ms 0/86327 (0%) sender [ 5] 0.00-15.47 sec 84.6 MBytes 45.9 Mbits/sec 0.082 ms 25078/86327 (29%) receiver ``` The sender happily reports 100 Mbit/sec with zero loss - it measured its own transmit rate, nothing more. The receiver tells the truth: 29% of datagrams never arrived. This is the single most misread iPerf3 output. If you only glance at the sender line, an over-driven UDP test looks perfect while the network is on fire. **Rule of thumb:** in UDP mode, the sender line is a statement of intent and the receiver line is the measurement. Quote the receiver line, always. ## Understanding Jitter Jitter is the variation in packet arrival spacing. If packets leave every 1.4 ms and arrive every 1.4 ms, jitter is zero even if the absolute latency is high. Voice and video codecs buffer against jitter, and their buffers are finite: as a rough guide, keep jitter under 30 ms for voice, and treat anything under a few milliseconds (like the 0.2 ms here) as excellent. High jitter with low loss usually points to queueing - traffic is getting through, but it is waiting in buffers of variable depth along the way (a QoS problem more than a capacity problem; see our [QoS guide](https://www.pinglabz.com/qos/)). ## Finding Loss That TCP Hides Here is the same path with 2% packet loss injected on the WAN link. A TCP test just gets slow, but the UDP test pinpoints the loss rate: ``` client:~$ iperf3 -c 10.0.20.10 -u -b 20M - - - - - - - - - - - - - - - - - - - - - - - - - [ ID] Interval Transfer Bitrate Jitter Lost/Total Datagrams [ 5] 0.00-10.00 sec 23.8 MBytes 20.0 Mbits/sec 0.000 ms 0/17266 (0%) sender [ 5] 0.00-10.10 sec 23.3 MBytes 19.4 Mbits/sec 0.219 ms 366/17266 (2.1%) receiver ``` 366 of 17,266 datagrams lost - 2.1%, matching the injected 2% almost exactly. When a TCP test is mysteriously slow, a moderate-rate UDP test like this is how you separate "the path is lossy" from "the path is fine but something limits TCP" (we walk that whole decision tree in [troubleshooting slow throughput](https://www.pinglabz.com/troubleshooting-slow-throughput-iperf3/)). ## UDP Flags That Matter **`-b 50M`**Target bandwidth. Default is only 1 Mbit/sec, so an unadorned `-u` test tells you almost nothing. Set it deliberately. **`-l 1200`**Datagram size. Keep it under the path MTU so one datagram = one packet; a fragmented datagram counts as lost if any fragment drops. 1200 is safe almost everywhere; the default 1470-byte-class sizing suits plain Ethernet. **`-P 4`**Parallel UDP streams, each rate-limited separately by `-b`. Four streams at 25M approximate four concurrent video calls better than one stream at 100M. **`-R`**Server sends, client receives. Loss is often asymmetric; test both directions. **`-t 60`**Longer tests catch intermittent loss that a 10-second burst misses. ## A Practical UDP Test Recipe 1. Establish the TCP ceiling first: `iperf3 -c `. Say it lands around 35 Mbit/sec. 2. Run UDP at roughly half that: `iperf3 -c -u -b 15M -t 30`. You want loss and jitter on a non-congested path. 3. Step up toward the ceiling in stages (20M, 25M, 30M) and watch where loss begins. That knee is your usable real-time capacity. 4. Only over-drive deliberately (like the 100M test above) when you want to prove where the bottleneck is, not to measure normal behavior. ## Key Takeaways UDP mode measures what TCP conceals: loss, jitter, and reordering. The sender line reports intent and the receiver line reports reality, and they diverge dramatically the moment you exceed path capacity. Set `-b` deliberately, keep datagrams under the MTU, test both directions, and find the loss knee by stepping the rate up gradually. For TCP-side tuning and everything else iPerf3 can do, head back to the [iPerf complete guide](https://www.pinglabz.com/iperf/). ### How to Use iPerf3: Commands and Examples From a Real Lab URL: https://www.pinglabz.com/how-to-use-iperf3/ Last updated: 2026-08-08T23:32:50.000Z iPerf3 is the standard tool for answering one of the most common questions in networking: how much throughput can this path actually deliver? Speed test sites measure your path to someone else's server. iPerf3 measures the exact path you care about, between two machines you control. Every capture in this guide comes from a live Cisco Modeling Labs topology - a Linux client, two IOS XE routers running OSPF, and an iperf3 server - so the numbers you see are what the tool really prints. This article is part of our [complete iPerf guide](https://www.pinglabz.com/iperf/). ## How iPerf3 Works iPerf3 uses a client-server model. One machine runs as the server and listens on TCP port 5201\. The other runs as the client, connects, and pushes as much traffic as it can (TCP) or exactly as much as you tell it to (UDP) for 10 seconds by default. Both sides then report what was sent and what actually arrived. That last distinction matters. The sender reports what it handed to the network. The receiver reports what survived the trip. On a healthy path the two numbers are nearly identical; when they diverge, the network is dropping or delaying traffic in between (that gap is your first troubleshooting clue). ## Step 1: Install iPerf3 On Debian and Ubuntu: ``` sudo apt install iperf3 ``` On Red Hat family systems use `dnf install iperf3`, on macOS `brew install iperf3`, and on Windows grab the binary and run it from a terminal. Version parity matters less than it used to, but keep both ends on iperf 3.x - iPerf3 does not interoperate with iPerf2 (see [iPerf2 vs iPerf3](https://www.pinglabz.com/iperf2-vs-iperf3/) for why they are separate tools). ## Step 2: Start the Server ``` iperf3 -s ``` That is the whole job. The server binds to port 5201 on all interfaces and waits: ``` ----------------------------------------------------------- Server listening on 5201 (test #1) ----------------------------------------------------------- ``` An iPerf3 server handles exactly one test at a time. A second client connecting mid-test gets refused (we cover daemon mode, custom ports, and that busy error in [running an iPerf3 server](https://www.pinglabz.com/iperf3-server-mode/)). ## Step 3: Run the Client From the other machine, point `-c` at the server's IP. Here is a real run from our lab client (10.0.10.10) to the server (10.0.20.10), crossing two routed hops: ``` client:~$ iperf3 -c 10.0.20.10 Connecting to host 10.0.20.10, port 5201 [ 5] local 10.0.10.10 port 39290 connected to 10.0.20.10 port 5201 [ ID] Interval Transfer Bitrate Retr Cwnd [ 5] 0.00-1.00 sec 4.12 MBytes 34.6 Mbits/sec 7 72.1 KBytes [ 5] 1.00-2.00 sec 4.00 MBytes 33.6 Mbits/sec 6 59.4 KBytes [ 5] 2.00-3.00 sec 4.25 MBytes 35.6 Mbits/sec 12 43.8 KBytes [ 5] 5.00-6.00 sec 4.38 MBytes 36.7 Mbits/sec 0 93.3 KBytes [ 5] 9.00-10.01 sec 4.00 MBytes 33.3 Mbits/sec 10 39.6 KBytes - - - - - - - - - - - - - - - - - - - - - - - - - [ ID] Interval Transfer Bitrate Retr [ 5] 0.00-10.01 sec 41.2 MBytes 34.6 Mbits/sec 83 sender [ 5] 0.00-10.01 sec 40.9 MBytes 34.2 Mbits/sec receiver ``` ## Reading the Output **Interval / Transfer / Bitrate** One line per second by default. Transfer is data moved in that interval; Bitrate is the rate. The final sender and receiver lines are the numbers to quote. **Retr (retransmits)** TCP segments resent because they were lost or reordered. 83 retransmits over 10 seconds here tells us the path drops some packets under load. Zero is ideal; a climbing count means congestion or loss. **Cwnd (congestion window)** How much unacknowledged data TCP is willing to keep in flight. Watch it grow, then get cut when loss occurs (43.8 KBytes right after those 12 retransmits). This is TCP congestion control working in real time. **sender vs receiver** Sender = handed to the network (41.2 MB). Receiver = actually delivered (40.9 MB). The receiver line is the truth about the path. ## The Flags You Will Actually Use **`-t 30`**Run for 30 seconds instead of 10\. Longer tests smooth out TCP slow start and give steadier numbers. **`-R`**Reverse mode: the server sends, the client receives. Tests the download direction without touching the far machine. **`-u -b 50M`**UDP at a fixed 50 Mbit/sec, reporting jitter and loss. See [iPerf3 UDP testing](https://www.pinglabz.com/iperf3-udp-testing/). **`-P 4`**Four parallel streams. Often reveals per-flow limits; see [parallel streams and window size](https://www.pinglabz.com/iperf3-parallel-streams-tcp-window/). **`-i 5`**Report every 5 seconds instead of every 1\. Less noise on long tests. **`-f M`**Print MBytes/sec instead of Mbits/sec (capital M = bytes, lowercase m = bits). **`-p 5202`**Use a different port. The server must be started with the same `-p`. **`-J`**JSON output for scripts and monitoring pipelines. ## Testing the Other Direction with -R By default the client sends. Add `-R` and the server sends instead, which is how you test download speed from the client's point of view: ``` client:~$ iperf3 -c 10.0.20.10 -R Reverse mode, remote host 10.0.20.10 is sending [ ID] Interval Transfer Bitrate [ 5] 0.00-1.00 sec 3.75 MBytes 31.4 Mbits/sec [ 5] 4.00-5.00 sec 4.00 MBytes 33.6 Mbits/sec - - - - - - - - - - - - - - - - - - - - - - - - - [ 5] 0.00-10.02 sec 39.1 MBytes 32.8 Mbits/sec 83 sender [ 5] 0.00-10.00 sec 38.8 MBytes 32.5 Mbits/sec receiver ``` Run both directions before drawing conclusions. Asymmetric results are common (policing on one direction, duplex issues, asymmetric routing) and testing only one way hides half the story. ## JSON Output for Automation Add `-J` and iPerf3 emits the entire test as structured JSON, ideal for feeding monitoring systems or CI checks: ``` client:~$ iperf3 -c 10.0.20.10 -J { "start": { "version": "iperf 3.18", "tcp_mss_default": 1448, ... "intervals": [{ "streams": [{ "bytes": 4325376, "bits_per_second": 34566124.48, "retransmits": 9, "snd_cwnd": 53576, ``` Pipe it to `jq '.end.sum_received.bits_per_second'` and you have a one-line throughput check you can alert on. ## A Sensible First Test, Step by Step 1. Confirm reachability first: `ping -c 3 10.0.20.10`. Note the round-trip time; you will want it later if throughput looks low. 2. Start the server: `iperf3 -s` on the far machine. 3. Run 30 seconds of TCP: `iperf3 -c 10.0.20.10 -t 30`. 4. Run the reverse direction: add `-R`. 5. If numbers disappoint, do not guess. Follow the method in [troubleshooting slow throughput with iPerf3](https://www.pinglabz.com/troubleshooting-slow-throughput-iperf3/). ## Key Takeaways iPerf3 is a two-command tool: `iperf3 -s` on one end, `iperf3 -c ` on the other. Trust the receiver-side summary line, not the sender's. Retr and Cwnd are not decoration - they tell you whether TCP is hitting loss and how much data it dares keep in flight. Always test both directions, and reach for `-J` the moment you want to script it. For the full picture of what iPerf can do (UDP, parallel streams, version differences, server hardening), start at the [iPerf complete guide](https://www.pinglabz.com/iperf/). ### Network Engineer Interview Questions & Career Guide (2026) URL: https://www.pinglabz.com/network-engineer-interview-questions/ Last updated: 2026-07-09T04:03:16.000Z Network engineer interviews are less about reciting definitions and more about proving you can think through a broken network under pressure. Whether you are chasing your first NOC role with a fresh CCNA or moving up to a senior engineering seat, the questions follow predictable patterns: a technical screen on fundamentals (OSI, TCP vs UDP, routing decisions), a scenario round where someone breaks a network on a whiteboard, and a behavioral round built around "tell me about a time." This guide covers the questions you will actually get, how strong candidates answer them, what the roles pay in 2026, and how to build a resume that survives the ATS filter. If you have not decided on the cert yet, start with [whether the CCNA is worth it in 2026](https://www.pinglabz.com/is-the-ccna-worth-it-in-2026/), then come back here. ## How Network Engineer Interviews Are Structured Most interview loops for network roles run three to four stages, and each one tests something different. Knowing what each stage is screening for lets you prepare deliberately instead of cramming everything. 1\. Recruiter screen Resume walkthrough, salary range, visa/location logistics. Screening for: keyword match and communication. 20 to 30 minutes. 2\. Technical screen Rapid-fire fundamentals: OSI layers, TCP vs UDP, subnetting, routing protocol basics. Screening for: did you actually learn what your certs claim. 3\. Scenario / whiteboard "Users can't reach the server, go." Live troubleshooting, sometimes in a lab or packet capture. Screening for: methodology, not memorization. 4\. Behavioral / panel "Tell me about a time you troubleshot an outage" and "tell me about a time you caused one." Screening for: honesty, ownership, and how you operate on a team. ## The Technical Questions You Will Actually Get These come up in nearly every network interview, from help desk to senior engineer. The difference between levels is not the question, it is the depth expected in the answer. ### Explain the OSI model. Which layers do you actually use? Weak answers recite "Please Do Not Throw Sausage Pizza Away." Strong answers treat the model as a troubleshooting framework: "I work bottom-up. Layer 1, is the interface up/up? Layer 2, do I see the MAC in the table and is the VLAN right? Layer 3, can I ping the gateway and is there a route? Layer 4, is the port open or is a firewall filtering it?" Mention that real stacks are TCP/IP (4 or 5 layers) and the OSI model is a shared vocabulary, not a literal implementation. That one sentence signals experience. ### TCP vs UDP: when would you pick each? Do not just say "TCP is reliable, UDP is fast." Anchor it to protocols the interviewer runs in production: TCP **Guarantees:** ordered, acknowledged, retransmitted, flow-controlled (three-way handshake). **Cost:** latency and state on both ends. **Lives on it:** HTTPS, SSH, BGP (port 179), FTP. UDP **Guarantees:** none. Fire and forget; the application handles loss. **Benefit:** low latency, no connection state. **Lives on it:** DNS, DHCP, SNMP, syslog, VoIP/RTP, OSPF hellos ride IP directly (protocol 89, a nice bonus point). The senior-level follow-up is usually "why does VoIP prefer UDP?" Answer: a retransmitted voice packet arrives too late to be useful, so TCP's reliability actively hurts. Loss is handled by the codec, and [QoS handles the delay and jitter](https://www.pinglabz.com/qos-for-voip/). ### How does a router decide which route to use? This is the single most reliable filter question for routing roles, and most candidates get the order wrong. The router picks in this order: longest prefix match first, then lowest administrative distance, then lowest metric. Here is real output from a Cisco IOS XE lab: ``` R1# show ip route ospf 2.0.0.0/32 is subnetted, 1 subnets O 2.2.2.2 [110/11] via 10.0.12.2, 00:00:43, Ethernet0/0 3.0.0.0/32 is subnetted, 1 subnets O IA 3.3.3.3 [110/21] via 10.0.12.2, 00:00:36, Ethernet0/0 ``` Walk the interviewer through the brackets: 110 is OSPF's administrative distance (how much the router trusts the source), 11 and 21 are the metric (OSPF cost along the path). A /32 static route would beat both, not because static AD is 1, but because a more specific prefix wins before AD is even consulted. If you can explain why [EIGRP shows 90 internal but 170 external](https://www.pinglabz.com/eigrp-administrative-distance/), you are ahead of most candidates. ### What happens when you type a URL and hit enter? A classic that tests whether you can chain concepts: DNS resolution (UDP 53, unless the answer is big enough to fall back to TCP), then ARP for the default gateway MAC if the destination is off-subnet, then the TCP handshake on 443, TLS negotiation, and finally HTTP. The trap most candidates fall into: saying the client ARPs for the server's MAC. Off-subnet traffic ARPs for the gateway, and the destination MAC changes at every layer 3 hop while source and destination IPs stay constant. Interviewers listen specifically for that. ### Layer 2 favorites: VLANs and spanning tree Expect "what is the difference between a VLAN and a subnet?" (a VLAN is a layer 2 broadcast domain, a subnet is a layer 3 addressing boundary; they usually map 1:1 but nothing enforces it - full breakdown [here](https://www.pinglabz.com/vlan-subnet-broadcast-domain-difference/)) and "why does spanning tree exist?" (loops at layer 2 have no TTL, so a single loop melts the network in seconds; STP blocks redundant paths until they are needed). Senior roles get follow-ups on [RSTP convergence](https://www.pinglabz.com/rapid-spanning-tree-protocol/) and [why PortFast plus BPDU Guard belongs on every access port](https://www.pinglabz.com/spanning-tree-portfast/). ## Scenario Questions: "Users Can't Reach the Server" ![Network engineer sketching a glowing network topology on a dark whiteboard while an interview panel watches](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/07/interview-whiteboard-scenario.jpg) This is where interviews are won. The interviewer does not care whether you find the fault on the first guess; they care whether your process would find any fault eventually. A strong scenario answer sounds like this: 1. **Scope it.** One user or all users? One application or everything? Did it ever work? The answers cut the problem space in half each time. 2. **Test connectivity in both directions and read the failure signature.** Ping output tells you more than up or down: ``` R1# ping 3.3.3.3 Sending 5, 100-byte ICMP Echos to 3.3.3.3, timeout is 2 seconds: U.U.U Success rate is 0 percent (0/5) R1# ping 10.99.99.99 Sending 5, 100-byte ICMP Echos to 10.99.99.99, timeout is 2 seconds: ..... Success rate is 0 percent (0/5) ``` Both pings failed, but they failed differently. `U` means a router along the path sent "destination unreachable" back - something actively refused it, which in this lab was an ACL on the far router. Pure dots mean silence: no route, or the packets are being dropped without a reply. Explaining that distinction out loud is exactly the kind of detail that separates candidates (we break down every failure pattern in [ping timeout vs destination unreachable](https://www.pinglabz.com/ping-timeout-vs-unreachable/)). 1. **Follow the path.** Traceroute to find where it dies, then check that device: interface status, ARP or MAC tables at layer 2, routing table at layer 3, ACLs and NAT at the edges. 2. **Change one thing at a time, and verify after each change.** Saying this sentence in an interview is worth more than any protocol trivia. The best preparation is doing this for real, repeatedly. Our [troubleshooting ticket labs](https://www.pinglabz.com/ccna-ts-nf-01-connectivity-tickets/) are built exactly like these interview scenarios: broken configs on real Cisco IOS XE, and you find the fault. ## Behavioral Questions: The STAR Answers That Land Every loop includes some version of these three. Prepare specific stories in STAR format (Situation, Task, Action, Result) before the interview, because inventing them live never works. "Tell me about a time you troubleshot a difficult issue." Pick a story with a non-obvious root cause. Spend most of your time on the Action: what you checked, in what order, and why. End with the result AND what you changed so it cannot recur (monitoring, documentation, a config standard). If you are early-career, a lab story is fine - "in my CML lab I hit an OSPF adjacency stuck in EXSTART" is a real story with a real methodology. "Tell me about a time you caused an outage." This is an honesty test. Everyone who has touched production has broken it. Own it directly, explain the blast radius, how you detected and rolled back, and the process change that followed (change windows, peer review, config backups). The only wrong answer is "I never have." "How do you handle a change you disagree with?" They want to hear: raise concerns with data, propose an alternative, and if overruled, execute the decision professionally while documenting the risk. Not compliance, not rebellion - professionalism. ## Entry-Level Roles and 2026 Salaries A CCNA rarely drops you straight into a "Network Engineer" title. The realistic entry points are NOC and support roles where you build the operational experience interviews ask about. US figures below are 2026 market ranges compiled from major salary aggregators; your metro area and industry shift these significantly (a NOC seat in a major hub can pay more than a mid-level role elsewhere). NOC Technician / Help Desk II **Experience:** 0 to 2 years **Range:** $50K to $70K Where most CCNA holders start. Monitoring, ticket triage, first-touch troubleshooting. Network Technician / Jr. Engineer **Experience:** 1 to 3 years **Range:** $55K to $75K Hands on switches and APs, executing changes under supervision. Network Engineer **Experience:** 3 to 6 years **Range:** $75K to $110K Owns designs and changes. CCNP territory. National averages cluster around $87K to $130K depending on source. Senior / Cloud Network Engineer **Experience:** 7+ years **Range:** $110K to $160K+ Architecture, automation, cloud networking. CCIE or deep cloud/automation skills push past this band. ## Career Path: CCNA, Then What? The classic ladder is CCNA to CCNP Enterprise to a specialization (security, service provider, data center) or CCIE. That path still works, but the highest-leverage move in 2026 is pairing route/switch depth with one of two multipliers. The first is **automation**: Python, Netmiko, Ansible, and APIs. Every job posting above junior level now lists it, and it is learnable in weeks, not years (our [network automation lab series](https://www.pinglabz.com/ccna-auto-01-first-automation-host/) starts from zero). The second is **cloud networking**: AWS Advanced Networking or Azure Network Engineer Associate on top of a CCNA is a rarer and better-paid combination than a second routing cert, because hybrid connectivity (Direct Connect, ExpressRoute, transit gateways) is where enterprises are hiring. Whichever branch you pick, the constant is hands-on lab evidence. A GitHub repo with your lab configs and automation scripts does more in an interview than a third certification. ## Resume and Keyword Optimization Most resumes are filtered by an ATS before a human reads them, and the ATS matches literal strings. Three rules cover most of it: - **Mirror the job title and cert names exactly.** Write "Network Engineer" and "CCNA (Cisco Certified Network Associate)" - the spelled-out form plus the acronym catches both search patterns. If the posting says "route/switch," your resume should too. - **Name the technologies, not the categories.** "Configured OSPF, HSRP, and 802.1X on Cisco Catalyst switches" beats "worked on routing and security." List protocols and platforms the posting lists: BGP, OSPF, VLANs, STP, IPsec, SD-WAN, Cisco IOS XE, Meraki, Palo Alto, whatever genuinely applies. - **Quantify with numbers a network person believes.** "Supported a 40-site WAN with 2,000 users," "cut mean time to resolution from 4 hours to 45 minutes," "migrated 300 APs to WPA3." Vague scale reads as no scale. If you lack production experience, add a **Projects** section and treat your lab like a job: "Built a 12-router OSPF/BGP topology in Cisco Modeling Labs; automated config backups with Python and Netmiko." Interviewers ask about it every time, and unlike padded job bullets, you can answer in depth. ## How to Prepare in the Two Weeks Before Reading question lists is passive; interviews reward recall under pressure. Split your prep: run broken-network labs so scenario answers come from muscle memory (start with the [connectivity ticket labs](https://www.pinglabz.com/ccna-ts-nf-01-connectivity-tickets/), graduate to the [CCNA Mega Lab](https://www.pinglabz.com/ccna-mega-lab/)), drill fundamentals with [flashcards](https://www.pinglabz.com/ccna-flashcards-pass/) for the rapid-fire round, and write out your three STAR stories. Then rehearse the routing decision explanation and the ping failure signatures out loud - explaining them clearly is a different skill from knowing them. ## FAQ ### Can I get a network engineer job with just a CCNA and no experience? Directly into an engineer title, rarely. Into a NOC, network technician, or IT support role that becomes an engineer title in 18 to 24 months, yes, routinely. The CCNA gets you the interview; lab projects and a clean troubleshooting methodology get you the offer. ### What salary should I expect with a CCNA? For a first networking role in the US in 2026, plan on $50K to $75K depending on metro area and whether the role is NOC, support, or junior engineering. Mid-level engineers with a few years of experience typically land between $75K and $110K. ### What is the most common network engineer interview question? Some form of "walk me through your troubleshooting process" appears in nearly every loop, either as a behavioral question ("tell me about a time...") or a live scenario ("users can't reach the server"). Prepare one strong real story and one clean generic methodology. ### Should I get CCNP or a cloud certification after CCNA? If you want to stay deep in enterprise route/switch, CCNP Enterprise. If you want maximum market flexibility, a cloud networking cert (AWS or Azure) plus Python automation skills is currently the better-paid and less crowded combination. ### What questions should I ask the interviewer? Ask about the change process ("how do changes get reviewed and rolled back?"), the monitoring stack, and what the last major outage was. These signal operational maturity and tell you whether the environment is one you want to inherit. ## Key Takeaways - Interviews test methodology over memorization: bottom-up OSI troubleshooting, one change at a time, verify after each change. - Know the routing decision order cold: longest prefix match, then administrative distance, then metric. - Read ping output like an engineer: `U` means actively refused (ACL, unreachable), dots mean silence (no route or silent drop). - Prepare three STAR stories in advance, including one where you broke something and owned it. - Entry path in 2026: CCNA into NOC ($50K to $70K), engineer in 2 to 4 years ($75K to $110K), then multiply with automation or cloud. - Resume: mirror the posting's exact keywords, name specific protocols and platforms, quantify scale, and showcase lab projects. ### Ping Sweep: Finding Live Hosts With Bash and Nmap URL: https://www.pinglabz.com/ping-sweep/ Last updated: 2026-07-08T21:07:12.000Z You need to know what is alive on a subnet: which IPs answer, which are free, whether that host you decommissioned is really gone. A ping sweep answers all three by pinging every address in a range and reporting who responds. You can do it with a five-line bash loop or a single nmap command, and knowing when to reach for which is the difference between a quick check and the right tool. This article shows both against a live lab subnet, and is part of the [PingLabz ping guide](https://www.pinglabz.com/ping/). ## What a Ping Sweep Actually Does A ping sweep sends one ICMP echo request to each address in a range and records which ones reply. The result is a live-host inventory. Network engineers use it to find free addresses before assigning statics, to verify a VLAN's hosts after a change, and to confirm a migration moved everything it should have. (The same technique is step one of network reconnaissance, which is why you should only sweep networks you are authorized to touch; the offensive-security angle lives in the [Nmap guide](https://www.pinglabz.com/nmap/).) The lab subnet here is 192.168.99.0/24, where only the router (192.168.99.1) and the test host (192.168.99.100) are up. ## The Bash One-Liner No tools to install, works on any Linux or macOS host, perfect for a quick check: ``` j@lab:~$ for i in 1 2 3 4 5 100 254; do ping -c 1 -W 1 192.168.99.$i >/dev/null 2>&1 \ && echo "192.168.99.$i is up" \ || echo "192.168.99.$i is down" done 192.168.99.1 is up 192.168.99.2 is down 192.168.99.3 is down 192.168.99.4 is down 192.168.99.5 is down 192.168.99.100 is up 192.168.99.254 is down ``` The two flags doing the heavy lifting are **`-c 1`** (one probe per host) and `**-W 1**` (wait at most one second for the reply). Without `-W`, every dead host stalls for the full default timeout and a /24 sweep crawls. The `&& / ||` construction turns ping's exit code (0 = reply received, non-zero = no reply) into a readable up/down line. To sweep a contiguous range, swap the explicit list for a C-style loop, and background the probes so they run in parallel instead of one-at-a-time: ``` j@lab:~$ for i in $(seq 1 254); do ping -c 1 -W 1 192.168.99.$i >/dev/null 2>&1 && echo "192.168.99.$i up" & done; wait ``` The trailing `&` fires all 254 pings concurrently and `wait` blocks until they finish, taking the sweep from minutes to about a second. The tradeoff: output order is no longer sorted (pipe to `sort -t. -k4 -n` if you care). ## The Catch: Sweeps Only See Hosts That Answer ICMP This is the limitation that trips people up. A "down" result does not mean the address is free; it means nothing answered an ICMP echo within your timeout. Plenty of live hosts stay silent: Windows machines block ICMP by default in many firewall profiles, hardened Linux servers may drop echo, and any host behind a personal firewall can be up and serving traffic while ignoring your ping. A ping sweep reliably tells you what *is* there; it cannot prove what *is not*. That gap is exactly why a dedicated scanner exists. ## The Right Tool: nmap -sn `nmap -sn` (the "no port scan" / host-discovery mode, formerly `-sP`) does a smarter sweep. On a local subnet it uses ARP, which is faster and cannot be firewalled by the host, and off-subnet it combines ICMP echo, an ICMP timestamp request, and TCP probes to ports 80 and 443, so it catches hosts that ignore plain ping: ``` j@lab:~$ nmap -sn 192.168.99.0/24 Starting Nmap 7.95 ( https://nmap.org ) at 2026-07-08 13:40 PDT Nmap scan report for 192.168.99.1 Host is up (0.0027s latency). Nmap scan report for 192.168.99.100 Host is up (0.00030s latency). Nmap done: 256 IP addresses (2 hosts up) scanned in 3.03 seconds ``` 256 addresses in three seconds, using ARP on this local segment. Because ARP replies are mandatory for any host that wants to communicate on the LAN, an ARP-based sweep finds machines that a pure ICMP sweep would miss entirely. This is why, on your own local network, `nmap -sn` is strictly more reliable than a ping loop. ## Choosing Between Them Bash ping loop **Reach for it when:** you're on a box with no nmap, need a fast yes/no on a handful of IPs, or want to script an up/down check into a health monitor. **Blind spot:** misses any host that doesn't answer ICMP echo. nmap -sn **Reach for it when:** you need an accurate inventory, are on the local subnet (ARP), or the targets may be firewalled against ping. **Blind spot:** needs to be installed; off-subnet it still depends on hosts answering something. A useful rule of thumb: use the bash loop for a quick "is this one thing up?" and nmap for "what is the complete list of what's up?" When accuracy matters (auditing, migration verification, finding a rogue device), the ARP-backed nmap sweep is the answer. ## A Middle Ground: fping If nmap is unavailable but you want parallelism without wrestling bash job control, `fping` is purpose-built: `fping -a -g 192.168.99.0/24 2>/dev/null` pings the whole range in parallel and prints only the alive addresses (`-a`), one per line, ready to pipe into the next tool. It sits neatly between the loop and nmap. ## Key Takeaways - A ping sweep inventories live hosts by pinging every address in a range. `-c 1 -W 1` keeps a bash sweep fast. - Background the probes with `&` plus `wait` to sweep a /24 in about a second instead of minutes. - A "down" result only means "didn't answer ICMP." Firewalled-but-live hosts read as down. - `nmap -sn` uses ARP on the local subnet (unfirewallable) and multiple probe types off-subnet, so it finds hosts a ping loop misses. - Loop for a quick check, nmap for an accurate inventory, fping when you want parallel speed without nmap. - Only sweep networks you are authorized to scan. This is the last stop in the cluster. Head back to the [complete ping guide](https://www.pinglabz.com/ping/), or go deeper on scanning in the [Nmap guide](https://www.pinglabz.com/nmap/). This is a network administration technique; use it only on systems you own or are authorized to test. ### Linux Ping Options Every Network Engineer Should Know URL: https://www.pinglabz.com/linux-ping-options/ Last updated: 2026-07-08T21:07:12.000Z The iputils ping that ships with every modern Linux distribution has dozens of flags, and most engineers use exactly two of them (`-c` and nothing else). Hidden in the rest are a jitter meter, an MTU prober, a flood tester, and an interface selector that solves one of the most confusing IPv6 failures you will hit. This article walks the flags that earn their place in a network engineer's muscle memory, each demonstrated live against a three-router Cisco lab. It is part of the [PingLabz ping guide](https://www.pinglabz.com/ping/). ## The Everyday Set: -c, -i, -q, -w **`-c` (count)** stops after N probes instead of running forever. **`-i` (interval)** sets seconds between probes (default 1). **`-q` (quiet)** suppresses the per-probe lines and prints only the summary. Combine all three and you get a dense link-quality sample in under a second: ``` j@lab:~$ sudo ping -c 100 -i 0.01 -q 3.3.3.3 PING 3.3.3.3 (3.3.3.3) 56(84) bytes of data. --- 3.3.3.3 ping statistics --- 100 packets transmitted, 100 received, 0% packet loss, time 990ms rtt min/avg/max/mdev = 2.992/3.581/6.581/0.402 ms ``` One hundred packets, one second, and the two numbers that matter: loss (0%) and mdev (0.4 ms of jitter). Note the `sudo`: unprivileged users cannot set intervals below 0.2 seconds (a rate-limiting guardrail, since fast pings are effectively load generators). **`-w` (deadline)** is subtly different from `-c`: it bounds the *total runtime* in seconds, regardless of how many replies arrive. `ping -w 3 host` is the right shape for scripts and health checks, because it can never hang: ``` j@lab:~$ ping -w 3 3.3.3.3 --- 3.3.3.3 ping statistics --- 3 packets transmitted, 3 received, 0% packet loss, time 2003ms ``` Its sibling **`-W` (timeout)** sets how long to wait for each individual reply. `-W 1` is the difference between a subnet sweep that takes 4 minutes and one that takes seconds (more on that in the [ping sweep article](https://www.pinglabz.com/ping-sweep/)). ## Size and MTU: -s and -M **`-s`** sets the ICMP payload size (default 56; add 28 bytes of headers for the wire size), and **`-M do`** sets the DF bit so nothing on the path may fragment the packet. Together they turn ping into an MTU probe. Against a lab path with a 1400-byte link two hops out: ``` j@lab:~$ ping -c 3 -s 1400 -M do 3.3.3.3 PING 3.3.3.3 (3.3.3.3) 1400(1428) bytes of data. From 10.0.12.2 icmp_seq=1 Frag needed and DF set (mtu = 1400) ``` The router names the limiting MTU in its error. The complete method, including the payload math and the Cisco-side equivalent, has its own article: [testing MTU with ping and the DF bit](https://www.pinglabz.com/ping-mtu-df-bit/). ## Hop Probing: -t **`-t`** caps the outgoing TTL. The probe dies at hop N and that router introduces itself with a Time Exceeded message: ``` j@lab:~$ ping -c 3 -t 1 3.3.3.3 From 192.168.99.1 icmp_seq=1 Time to live exceeded j@lab:~$ ping -c 3 -t 2 3.3.3.3 From 10.0.12.2 icmp_seq=1 Time to live exceeded ``` TTL 1 finds the first hop, TTL 2 the second. This is traceroute's core mechanism exposed as a single flag (background in [how ping works](https://www.pinglabz.com/how-ping-works/)). ## Timestamps for Correlation: -D **`-D`** prefixes each reply with a Unix epoch timestamp: ``` j@lab:~$ ping -c 3 -D 3.3.3.3 [1783543025.242112] 64 bytes from 3.3.3.3: icmp_seq=1 ttl=253 time=4.34 ms [1783543026.243890] 64 bytes from 3.3.3.3: icmp_seq=2 ttl=253 time=4.45 ms [1783543027.245589] 64 bytes from 3.3.3.3: icmp_seq=3 ttl=253 time=4.34 ms ``` Indispensable when you leave a ping running overnight into a log file and need to line up the moment of loss against syslog, interface counters, or a change window. Pair it with `-O`, which prints a line for each *missed* reply instead of staying silent, and you have a poor man's availability monitor. ## Load Testing: -f (Handle With Care) **`-f` (flood)** sends as fast as replies return (or 100 pps minimum), printing a dot for each unanswered probe. Watching the dots pile up during a flood is a live packet-loss visualization: ``` j@lab:~$ sudo ping -f -c 500 1.1.1.1 PING 1.1.1.1 (1.1.1.1) 56(84) bytes of data. --- 1.1.1.1 ping statistics --- 500 packets transmitted, 500 received, 0% packet loss, time 504ms rtt min/avg/max/mdev = 0.807/0.923/2.714/0.140 ms, ipg/ewma 1.009/0.918 ms ``` 500 packets in half a second, zero loss, and a bonus metric: ipg (inter-packet gap). Flood requires root, and it is a lab and maintenance-window tool. Pointing it at production infrastructure you do not own is somewhere between rude and a security incident (that address is a lab router's loopback, not the public resolver). ## IPv6 and the Interface Trap: -6 and -I **`-6`** forces IPv6\. Here is the failure mode that confuses everyone the first time. The lab host has two NICs, and R1's IPv6 address lives off the second one: ``` j@lab:~$ ping -6 -c 3 2001:db8:99::1 From fd64:f725:df42:4f01:b1ba:d16b:bd2e:dde9 icmp_seq=1 Destination unreachable: Address unreachable ``` Address unreachable, even though the router is right there and configured correctly. The giveaway is the `From` address: the kernel sourced the probe from the *wrong interface* (the management NIC), where no route to 2001:db8:99::/64 exists. Bind the ping to the correct interface with **`-I`** and it works instantly: ``` j@lab:~$ ping -6 -c 3 -I ens224 2001:db8:99::1 PING 2001:db8:99::1 (2001:db8:99::1) from 2001:db8:99::100 ens224: 56 data bytes 64 bytes from 2001:db8:99::1: icmp_seq=1 ttl=64 time=4.81 ms --- 2001:db8:99::1 ping statistics --- 3 packets transmitted, 3 received, 0% packet loss ``` On multi-homed hosts, and especially with IPv6 link-local addresses (which *always* need `%interface` or `-I`), source and interface selection is half the battle. More IPv6 fundamentals live in the [IPv6 guide](https://www.pinglabz.com/ipv6/). ## The Field Reference \-c 100 -i 0.01 -q Fast link-quality sample: loss + jitter in one second (root for i < 0.2). \-s 1372 -M do MTU probe: payload + 28 = wire size, DF set. Router replies with the limiting MTU. \-t 1, -t 2, -t 3... Manual traceroute: each TTL exposes one hop via Time Exceeded. \-w 3 / -W 1 \-w bounds total runtime (script-safe); -W bounds per-reply wait (fast sweeps). \-D (+ -O) Epoch timestamps per reply, and a line per missed reply. Overnight monitoring into a log file. \-6 -I ens224 Force IPv6 and pin the interface. Mandatory for link-local, lifesaving on multi-homed hosts. ## Key Takeaways - `-c 100 -i 0.01 -q` gives you loss and jitter (mdev) in one second; sub-0.2s intervals need root. - `-w` is the script-safe flag: it bounds runtime no matter what the network does. - `-s` plus `-M do` turns ping into a path MTU prober. - `-D` timestamps make overnight captures correlatable with logs. - When IPv6 ping says "Address unreachable" from an address you didn't expect, check the source interface and reach for `-I`. - Flood ping is a lab tool. Own the target or don't send it. Continue with [ping sweeps for host discovery](https://www.pinglabz.com/ping-sweep/), or head back to the [complete ping guide](https://www.pinglabz.com/ping/). ### Testing MTU With Ping and the DF Bit URL: https://www.pinglabz.com/ping-mtu-df-bit/ Last updated: 2026-07-08T21:07:12.000Z Small pings work, SSH works, but file copies hang and some websites never load. That symptom set has one classic cause: an MTU mismatch somewhere on the path, usually introduced by a tunnel. The fastest way to find it is the tool you already have open. This article shows how to hunt path MTU with ping from both Linux and Cisco IOS XE, demonstrated against a lab link deliberately clamped to 1400 bytes. It is part of the [PingLabz ping guide](https://www.pinglabz.com/ping/). ## The Setup: a 1400-Byte Link in the Path The lab path runs from a Debian host through R1 and R2 to R3's loopback (3.3.3.3). The R2-R3 link carries `ip mtu 1400`, exactly what you would see on a network with a GRE tunnel in the middle (GRE eats 24 bytes, and engineers commonly clamp to 1400 for margin; the full story is in the [GRE guide](https://www.pinglabz.com/gre/)). One detail worth internalizing from R2: ``` R2#show ip interface Ethernet0/2 | include MTU MTU is 1400 bytes R2#show interfaces Ethernet0/2 | include MTU MTU 1500 bytes, BW 10000 Kbit/sec, DLY 1000 usec, ``` Two different commands, two different answers, both correct. `show interfaces` reports the *hardware* MTU (still 1500), while `show ip interface` reports the *IP* MTU (clamped to 1400). Only the IP MTU governs whether an IPv4 packet gets fragmented. Knowing which command answers which question saves you from chasing ghosts. ## The 28 Bytes Everyone Forgets Ping's `-s` flag sets the ICMP *payload* size, not the packet size. The wire size is payload + 8 (ICMP header) + 20 (IPv4 header). So to test a 1400-byte MTU, the largest payload that fits is 1400 - 28 = 1372: ICMP payload (-s)1372 bytes \+ ICMP header8 bytes \+ IPv4 header20 bytes \= IP packet on the wire1400 bytes (fits the 1400 MTU exactly) (Cisco's `size` option, by contrast, sets the full IP packet size, so `size 1400` on IOS is equivalent to `-s 1372` on Linux. The off-by-28 confusion between the two conventions has burned many an engineer.) ## Testing From Linux: -M do `-M do` sets the DF (Don't Fragment) bit, which forbids routers from fragmenting the packet. First, prove the exact fit passes: ``` j@lab:~$ ping -c 3 -s 1372 -M do 3.3.3.3 PING 3.3.3.3 (3.3.3.3) 1372(1400) bytes of data. 1380 bytes from 3.3.3.3: icmp_seq=1 ttl=253 time=5.43 ms 1380 bytes from 3.3.3.3: icmp_seq=2 ttl=253 time=4.81 ms 1380 bytes from 3.3.3.3: icmp_seq=3 ttl=253 time=4.86 ms --- 3.3.3.3 ping statistics --- 3 packets transmitted, 3 received, 0% packet loss ``` Now push one byte class beyond it (payload 1400 = 1428 on the wire) and watch the path answer back: ``` j@lab:~$ ping -c 3 -s 1400 -M do 3.3.3.3 PING 3.3.3.3 (3.3.3.3) 1400(1428) bytes of data. From 10.0.12.2 icmp_seq=1 Frag needed and DF set (mtu = 1400) --- 3.3.3.3 ping statistics --- 3 packets transmitted, 0 received, +3 errors, 100% packet loss ping: sendmsg: Message too long ``` Two things happened here, and both are diagnostic gold. First, the router at 10.0.12.2 sent an ICMP "Fragmentation Needed" error *that names the limiting MTU*: 1400\. No binary search required; the path told you the answer. Second, the follow-up probes failed locally with `Message too long`: the kernel cached the discovered path MTU (that is Path MTU Discovery doing its job) and refused to send oversized packets at all. Remove the DF bit and the same oversized packet sails through, because R2 is now allowed to fragment it: ``` j@lab:~$ ping -c 3 -s 1400 3.3.3.3 PING 3.3.3.3 (3.3.3.3) 1400(1428) bytes of data. 1408 bytes from 3.3.3.3: icmp_seq=1 ttl=253 time=5.35 ms --- 3.3.3.3 ping statistics --- 3 packets transmitted, 3 received, 0% packet loss ``` Working fragmentation is why MTU problems hide: most traffic without DF gets fragmented and limps along (with a performance tax), while TCP flows, which set DF because they rely on PMTUD, hang. ## Testing From IOS XE: df-bit The router-side equivalent, from R1: ``` R1#ping 3.3.3.3 size 1500 df-bit Type escape sequence to abort. Sending 5, 1500-byte ICMP Echos to 3.3.3.3, timeout is 2 seconds: Packet sent with the DF bit set M.M.M Success rate is 0 percent (0/5) ``` Every `M` is a "could not fragment" report. For an exhaustive scan, the interactive extended ping dialog offers a size sweep (`Sweep range of sizes`), which walks every packet size in a range and shows precisely where `!` turns into `M`. One-line options and the rest of the extended toolkit are covered in [Cisco extended ping](https://www.pinglabz.com/cisco-extended-ping/). ## When the Path Won't Tell You: Binary Search The router named 1400 in its ICMP error, but on real networks the error often never reaches you (filtered upstream). Then you binary search with `-M do`: 1472 fails, try 1372; passes, try 1422; and so on until you bracket the ceiling in half a dozen pings. Common answers you will land on: 1472 (clean 1500 path), 1444 (GRE), 1412-1424 (IPsec variants), 1452 (PPPoE). ## PMTUD Black Holes: the Silent Killer Path MTU Discovery depends entirely on that "Frag needed" ICMP message getting back to the sender. Every hardening guide that says "block all ICMP" and every interface running `no ip unreachables` breaks the feedback loop: the sender keeps transmitting full-size DF packets into a hole and never learns why nothing comes back. That is a PMTUD black hole, and its signature is exactly the symptom this article opened with: small packets fine, big transfers dead. If you operate the small-MTU link, the mitigation on Cisco gear is `ip tcp adjust-mss` on the tunnel or LAN interface, which rewrites the MSS in TCP handshakes so endpoints never send oversized segments in the first place. If you only operate the client side, clamping the interface MTU works as a blunt instrument. Either way, never filter ICMP type 3 code 4 at your edge; allowing "fragmentation needed" through is part of running a functional network, a point expanded in the [pillar's ICMP filtering section](https://www.pinglabz.com/ping/). ## Key Takeaways - Linux `-s` sets payload (add 28 for the wire size); Cisco `size` sets the whole IP packet. 1372 payload = 1400 packet. - `ping -s N -M do` on Linux and `ping X size N df-bit` on IOS are the same test: does an N-byte-class packet fit without fragmentation? - A "Frag needed" error often names the limiting MTU. When it does not arrive, binary search for the ceiling. - `show ip interface` (IP MTU) and `show interfaces` (hardware MTU) answer different questions. - Small pings pass + large transfers hang = suspect a PMTUD black hole. Fix with `ip tcp adjust-mss`, and never filter ICMP frag-needed messages. Related reading: [Linux ping options in depth](https://www.pinglabz.com/linux-ping-options/), the [GRE tunnel guide](https://www.pinglabz.com/gre/) (where these MTU numbers come from), and the full [ping guide](https://www.pinglabz.com/ping/). ### Cisco Extended Ping and Response Characters Explained URL: https://www.pinglabz.com/cisco-extended-ping/ Last updated: 2026-07-08T21:07:11.000Z The plain `ping 3.3.3.3` on a Cisco router tests one thing: the path from the router's *egress interface* to the target. That is often not the question you actually need answered. Extended ping lets you control the source address, packet size, DF bit, repeat count, and timeout, which turns ping from a reachability check into a precision instrument for testing return paths, MTU, and route symmetry. This article covers the options that matter and the response characters that tell you what failed, with every capture taken live from IOS XE routers in CML. It is part of the [PingLabz ping guide](https://www.pinglabz.com/ping/). ## Why the Source Address Matters When R1 pings R3's loopback with defaults, the packet is sourced from R1's outgoing interface (10.0.12.1). R3's reply only has to find its way back to the 10.0.12.0/30 link, which is directly connected to R3's neighbor. That test can pass while half your network is unreachable. Source the ping from R1's loopback instead, and now R3 must have a working route back to 1.1.1.1\. You are testing the return path, which is where routing problems hide: ``` R1#ping 3.3.3.3 source Loopback0 Type escape sequence to abort. Sending 5, 100-byte ICMP Echos to 3.3.3.3, timeout is 2 seconds: Packet sent with a source address of 1.1.1.1 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 2/3/4 ms ``` This is the single most useful extended option. If `ping X` works but `ping X source Loopback0` fails, the far side is missing a route to your loopback (or to whatever prefix you sourced from). In [BGP](https://www.pinglabz.com/bgp/) and [OSPF](https://www.pinglabz.com/ospf/) troubleshooting, sourcing from loopbacks is standard practice because that is what your routing protocol sessions and management traffic actually use. ## One-Line Extended Ping You do not need the interactive dialog for most work. IOS accepts the common options inline: ``` ping 3.3.3.3 source Loopback0 ping 3.3.3.3 size 1500 df-bit ping 3.3.3.3 repeat 100 size 1000 ping 3.3.3.3 repeat 500 timeout 1 ping vrf MGMT 10.10.10.1 ``` A high repeat count with a bigger payload is a quick-and-dirty link quality test. Here is 100 pings at 1000 bytes across two routed hops: ``` R1#ping 3.3.3.3 repeat 100 size 1000 Sending 100, 1000-byte ICMP Echos to 3.3.3.3, timeout is 2 seconds: !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! !!!!!!!!!!!!!!!!!!!!!!!!!!!!!! Success rate is 100 percent (100/100), round-trip min/avg/max = 2/3/4 ms ``` Anything less than 100% on a healthy internal path deserves investigation (Cisco's own guidance treats a success rate under 80 percent as problematic). ## DF Bit and Size: MTU Testing From the Router The path from R1 to R3 crosses a link deliberately configured with `ip mtu 1400`. Send a full-size packet with the DF bit set and watch it die: ``` R1#ping 3.3.3.3 size 1500 df-bit Sending 5, 1500-byte ICMP Echos to 3.3.3.3, timeout is 2 seconds: Packet sent with the DF bit set M.M.M Success rate is 0 percent (0/5) ``` `M` means "could not fragment": a downstream router needed to fragment the 1500-byte packet to fit the 1400-byte link, and the DF bit forbade it. The full MTU-hunting method (including finding the exact path MTU by binary search) is covered in [ping and MTU testing](https://www.pinglabz.com/ping-mtu-df-bit/). ## Response Characters, Decoded Every probe in an IOS ping prints exactly one character. Memorize the six you will actually see: ! Reply received. The only character you want to see. . Timeout. No reply within the timeout (default 2 seconds). Silent loss somewhere. U Destination unreachable error received: no route, host down, or administratively filtered. M Could not fragment: packet too big for a link and DF was set. MTU problem. Q Source quench: destination too busy. Rare on modern gear. & Packet lifetime exceeded: TTL expired in transit. Think routing loop. Real patterns are more informative than single characters. Here is R1 pinging a target behind an ACL that denies ICMP echo: ``` R1#ping 3.3.3.3 Sending 5, 100-byte ICMP Echos to 3.3.3.3, timeout is 2 seconds: U.U.U Success rate is 0 percent (0/5) ``` Why `U.U.U` and not `UUUUU`? IOS rate-limits ICMP unreachables to one per 500 ms by default, so every other probe gets an unreachable and the ones in between time out. Alternating patterns like this are a signature of rate-limited ICMP errors, not flapping. Compare a plain no-route timeout, which produces only dots: ``` R1#ping 10.99.99.99 Sending 5, 100-byte ICMP Echos to 10.99.99.99, timeout is 2 seconds: ..... Success rate is 0 percent (0/5) ``` ## The Interactive Dialog Type `ping` with no arguments in privileged EXEC and IOS walks you through every option: protocol, target, repeat count, datagram size, timeout, extended commands (source, ToS/DSCP, DF bit, validate, data pattern, IP header options), and sweep. Two capabilities live only in this dialog: **Sweep range of sizes.** Answer `y` to sweep and set min/max/interval, and IOS sends probes at every size in the range. It is the exhaustive way to find the exact MTU ceiling on a path (each size that fits returns `!`, each size that does not returns `M` with DF set). **Data pattern.** Useful on suspect serial or optical links: patterns like `0x0000`, `0xFFFF`, and `0xAAAA` can catch line-coding and clocking problems that a default pattern misses. **Record route.** Under IP header options, `Record` stores up to nine hop addresses in the packet, including the return path (something traceroute cannot show you). Modern networks often filter packets with IP options, so treat it as a lab tool. ## The Companion: Extended Traceroute The same source-address logic applies to traceroute, and the numeric form belongs in your muscle memory: ``` R1#traceroute 3.3.3.3 numeric Tracing the route to 3.3.3.3 1 10.0.12.2 3 msec 2 msec 2 msec 2 10.0.23.1 3 msec * 4 msec ``` (That lone `*` on the final hop is the same ICMP rate-limiting you saw above, not packet loss.) IOS traceroute sends UDP probes with incrementing TTLs and reads the ICMP time-exceeded messages that come back; its response characters differ from ping's (`A` \= administratively prohibited, `H` \= host unreachable, `P` \= protocol unreachable, `*` \= timeout). ## Key Takeaways - Default ping sources from the egress interface and barely tests the return path. `source Loopback0` is the option that finds real routing problems. - One-liners cover source, size, df-bit, repeat, timeout, and VRF. The interactive dialog adds sweep, data patterns, and record route. - `!` reply, `.` timeout, `U` unreachable, `M` could not fragment. Patterns like `U.U.U` come from ICMP unreachable rate-limiting (1 per 500 ms). - `size 1500 df-bit` producing `M` \= a sub-1500 MTU link on the path. - Success below 100% on an internal path is a finding, not a rounding error. Next: [hunting path MTU with ping](https://www.pinglabz.com/ping-mtu-df-bit/), or browse the full [ping guide](https://www.pinglabz.com/ping/). ### Ping Timeout vs Destination Unreachable: Decoding Every Failure URL: https://www.pinglabz.com/ping-timeout-vs-unreachable/ Last updated: 2026-07-08T21:07:11.000Z Two pings fail. One prints nothing and reports 100% packet loss. The other prints `Destination Host Unreachable` on every line. Most people treat them as the same thing ("ping doesn't work"), but they are completely different failures pointing at completely different parts of the network. This article decodes every failure message you will see from Linux ping and Cisco IOS ping, using real captures from a live lab, and is part of the [PingLabz ping guide](https://www.pinglabz.com/ping/). The lab: a Debian host behind three IOS XE routers (R1, R2, R3) running OSPF, with R3's loopback 3.3.3.3 as the usual target. Each failure below was deliberately engineered, so you can see exactly which fault produces which message. ## The Two Families of Failure Every ping failure falls into one of two families, and identifying the family is your first triage step: **Silent failures.** The packet (or its reply) was dropped and nobody told you. You see nothing until the summary line reports loss. The evidence is the *absence* of information. **Loud failures.** Some device on the path sent you an ICMP error message explaining the problem. The message names the device that generated it (the `From` address), which immediately localizes the fault. Loud failures are gifts. The `From` address is a device telling you "the problem is at or after me." ## Silent Loss: 100% Packet Loss, No Errors Here the path to 3.3.3.3 has a filter that drops ICMP without generating errors (an ACL on R3 combined with `no ip unreachables`, a common hardening configuration): ``` j@lab:~$ ping -c 3 3.3.3.3 PING 3.3.3.3 (3.3.3.3) 56(84) bytes of data. --- 3.3.3.3 ping statistics --- 3 packets transmitted, 0 received, 100% packet loss, time 2029ms ``` No error lines at all. This pattern means one of: a firewall or ACL silently dropping (deny without unreachables), the destination is down but a stateless device in front of it eats the evidence, the reply path is broken (request arrives, reply cannot return), or ICMP is deprioritized/rate-limited somewhere. Silent loss localizes to nothing by itself; you need traceroute to find where the path goes dark: ``` j@lab:~$ traceroute -n 3.3.3.3 traceroute to 3.3.3.3 (3.3.3.3), 30 hops max, 60 byte packets 1 192.168.99.1 2.793 ms 2.764 ms 2.685 ms 2 10.0.12.2 3.525 ms 3.487 ms 3.439 ms 3 * * * 4 * * * ``` Hops 1 and 2 answer, hop 3 is a wall of asterisks. The fault sits between 10.0.12.2 and the next device. That is a real diagnosis, extracted from a silent failure. ## Destination Host Unreachable (From Your Own IP) ``` j@lab:~$ ping -c 3 192.168.99.77 PING 192.168.99.77 (192.168.99.77) 56(84) bytes of data. From 192.168.99.100 icmp_seq=1 Destination Host Unreachable From 192.168.99.100 icmp_seq=2 Destination Host Unreachable From 192.168.99.100 icmp_seq=3 Destination Host Unreachable ``` Look at the `From` address: 192.168.99.100 is the *sender's own IP*. The target is on the local subnet, so the host tried to ARP for it, got no answer, and gave up. The packet never left the NIC. This failure is Layer 2 adjacent: wrong IP, host powered off, wrong VLAN, or a switching problem (start with [VLAN and Layer 2 troubleshooting](https://www.pinglabz.com/vlans-layer-2-switching/) rather than routing). When the same message arrives `From` a router's address instead, that router has a route to the subnet but cannot ARP the final host on its connected segment. Same L2 diagnosis, different location: the far-end segment. ## Packet Filtered (Administratively Prohibited) With an ACL on R3 denying ICMP echo (but with unreachables left enabled), the router honestly reports the policy drop: ``` j@lab:~$ ping -c 3 3.3.3.3 PING 3.3.3.3 (3.3.3.3) 56(84) bytes of data. From 10.0.23.1 icmp_seq=1 Packet filtered From 10.0.23.1 icmp_seq=2 Packet filtered From 10.0.23.1 icmp_seq=3 Packet filtered ``` This is ICMP type 3 code 13, "communication administratively prohibited." The device at 10.0.23.1 is telling you an ACL or firewall rule matched a deny. Nothing is broken; policy is doing its job (whether the policy is *correct* is a separate conversation, often with whoever manages the [firewall](https://www.pinglabz.com/cisco-asa/)). ## Time to Live Exceeded ``` j@lab:~$ ping -c 3 -t 1 3.3.3.3 From 192.168.99.1 icmp_seq=1 Time to live exceeded ``` The TTL ran out in transit (here forced deliberately with `-t 1`). If you see this *without* forcing a low TTL, it almost always means a routing loop: the packet bounced between routers until the TTL burned down. Confirm with traceroute and expect to see the same pair of addresses alternating. Loops usually trace back to a redistribution or default-route mistake in your [OSPF](https://www.pinglabz.com/ospf/) or [BGP](https://www.pinglabz.com/bgp/) design. ## Frag Needed and DF Set ``` j@lab:~$ ping -c 3 -s 1400 -M do 3.3.3.3 PING 3.3.3.3 (3.3.3.3) 1400(1428) bytes of data. From 10.0.12.2 icmp_seq=1 Frag needed and DF set (mtu = 1400) ``` The packet was too big for a link on the path, and the DF (Don't Fragment) bit forbade the router from fragmenting it. The router even reports the MTU that would fit (1400). Small pings succeed while large transfers hang: this is the classic MTU black hole signature, and it gets a full treatment in [testing MTU with ping and the DF bit](https://www.pinglabz.com/ping-mtu-df-bit/). ## The Same Failures in Cisco IOS Ping IOS compresses all of this into single response characters. The same ACL scenario, seen from R1 instead of the Linux host: ``` R1#ping 3.3.3.3 Type escape sequence to abort. Sending 5, 100-byte ICMP Echos to 3.3.3.3, timeout is 2 seconds: U.U.U Success rate is 0 percent (0/5) ``` And a destination with no route at all just times out silently: ``` R1#ping 10.99.99.99 Sending 5, 100-byte ICMP Echos to 10.99.99.99, timeout is 2 seconds: ..... Success rate is 0 percent (0/5) ``` The `U.U.U` alternation is not random: IOS rate-limits ICMP unreachables (one per 500 ms by default), so between the U replies the probe just times out as a dot. Full character decoding lives in the [extended ping article](https://www.pinglabz.com/cisco-extended-ping/). ## The Failure Decoder 100% loss, no messages **Who sent it:** nobody **Look at:** silent filters, dead host behind a stateless hop, broken return path. Run traceroute to find where the path goes dark. Host Unreachable (from your own IP) **Who sent it:** your host **Look at:** ARP failure on the local segment. Wrong IP, wrong VLAN, target down. Layer 2 problem. Net/Host Unreachable (from a router) **Who sent it:** the router named in From **Look at:** that router's routing table (no route), or ARP on the destination segment. Routing problem at or beyond that hop. Packet filtered **Who sent it:** the filtering device **Look at:** ACL or firewall policy on the device named in From. Intentional drop, honestly reported. Time to live exceeded **Who sent it:** the router where TTL hit 0 **Look at:** routing loops (if you didn't set -t yourself). Trace and watch for alternating hops. Frag needed and DF set **Who sent it:** the router at the small link **Look at:** path MTU. The message includes the MTU that fits. Check tunnels first. ## A 60-Second Triage Workflow When a ping to a remote target fails, run these three probes and you will localize most faults before opening a single config: ping your default gateway (proves L2 and the first hop), ping the far-side router interface or a known-good host near the target (proves the routed path), then traceroute the original target (finds the exact hop where behavior changes). Match whatever messages you get against the decoder above, and pay attention to the `From` address in every error line: it is the network telling you where to look. ## Key Takeaways - Silent loss and unreachable messages are different families of failure: no messenger vs a named messenger. - `Destination Host Unreachable` from your own IP = local ARP failure = Layer 2 problem. - `Packet filtered` \= ICMP type 3 code 13 = an ACL or firewall matched a deny and told you about it. - Unforced `Time to live exceeded` \= suspect a routing loop. - IOS compresses these into characters: `.` timeout, `U` unreachable, `M` could not fragment, and rate-limits unreachables into patterns like `U.U.U`. - The `From` address in every error is your fault locator. Use it. Continue with [Cisco extended ping and response characters](https://www.pinglabz.com/cisco-extended-ping/), or return to the [complete ping guide](https://www.pinglabz.com/ping/). ### How Ping Works: The ICMP Echo Exchange Explained URL: https://www.pinglabz.com/how-ping-works/ Last updated: 2026-07-08T21:07:10.000Z You type `ping`, you get replies, the network is fine. Every engineer runs this loop dozens of times a day, but far fewer can explain what actually happens between pressing Enter and seeing `64 bytes from 3.3.3.3`. That gap matters, because every field in the ping output line is telling you something specific about the path, and reading it correctly is the difference between "the ping works" and actually understanding what you just proved. This article is part of the [PingLabz ping guide](https://www.pinglabz.com/ping/), and every capture in it comes from a live lab: a Debian host connected through three Cisco IOS XE routers running OSPF. ## What Ping Actually Sends Ping uses ICMP (Internet Control Message Protocol), which rides directly on top of IP as protocol number 1\. There is no TCP, no UDP, and no port number anywhere in the exchange (a common interview trap: ping does not use a port). The tool sends an **ICMP Echo Request** (type 8, code 0) to the target. If the target is up, reachable, and willing to answer, it returns an **ICMP Echo Reply** (type 0, code 0) with the same payload it received. Only four ICMP message types account for nearly everything you will see in day-to-day ping work: Echo Request Type:8, code 0 What your host sends. The question: "are you there?" Echo Reply Type:0, code 0 The answer. Carries back the same payload, so ping can compute round-trip time. Destination Unreachable Type:3, codes 0-13 A router (or the destination) telling you the packet could not be delivered, and why: no route, no host, filtered, or fragmentation needed. Time Exceeded Type:11, code 0 The TTL hit zero in transit. This is the message traceroute is built on. Everything else, the sequence numbers, the timing, the statistics, is bookkeeping the ping utility does locally around this simple request/reply pair. ## The Lab Behind These Captures All output below comes from a Debian 13 host (iputils ping) attached to a three-router Cisco IOS XE chain in Cisco Modeling Labs: host to R1 (192.168.99.1), R1 to R2, R2 to R3, with loopbacks 1.1.1.1, 2.2.2.2, and 3.3.3.3 advertised in OSPF. Pinging 3.3.3.3 from the host crosses two intermediate routers before reaching R3. ## Reading the Output Line by Line Here is a ping to R3's loopback, three routed hops away: ``` j@lab:~$ ping -c 4 3.3.3.3 PING 3.3.3.3 (3.3.3.3) 56(84) bytes of data. 64 bytes from 3.3.3.3: icmp_seq=1 ttl=253 time=3.93 ms 64 bytes from 3.3.3.3: icmp_seq=2 ttl=253 time=4.43 ms 64 bytes from 3.3.3.3: icmp_seq=3 ttl=253 time=4.79 ms 64 bytes from 3.3.3.3: icmp_seq=4 ttl=253 time=4.61 ms --- 3.3.3.3 ping statistics --- 4 packets transmitted, 4 received, 0% packet loss, time 3005ms rtt min/avg/max/mdev = 3.928/4.440/4.793/0.322 ms ``` Working through each field: **56(84) bytes of data.** The default payload is 56 bytes. Add the 8-byte ICMP header and the 20-byte IPv4 header and you get 84 bytes on the wire. This 28-byte overhead becomes important when you use ping to [test MTU with the DF bit](https://www.pinglabz.com/ping-mtu-df-bit/). **64 bytes from 3.3.3.3.** The reply size: your 56-byte payload plus the 8-byte ICMP header. The reply came back from the address you targeted, which confirms two-way reachability (the request found a path there AND the reply found a path back). **icmp\_seq=1.** A sequence number ping increments per probe. Gaps in the sequence mean individual packets were lost; replies arriving out of order mean something on the path is reordering traffic. **ttl=253.** The most misread field in the output, covered next. **time=3.93 ms.** Round-trip time (RTT): request out plus reply back plus the target's processing time. Not one-way latency, and not a guaranteed measure of what your application traffic experiences (routers often process ICMP addressed to themselves in the slow path, so pinging a router interface can show worse numbers than traffic through that router would see). ## What the TTL Really Tells You The TTL you see belongs to the **echo reply**, not to your request. The destination creates the reply with its own initial TTL, and every router on the return path decrements it by one. Common initial values: Linux uses 64, Windows uses 128, and Cisco IOS uses 255. In the capture above, ttl=253 means the reply started at 255 (Cisco) and was decremented twice, by R2 and then R1, on its way back. Two intermediate routers, exactly matching the topology. Ping the first-hop router directly and there is nothing in between to decrement it: ``` j@lab:~$ ping -c 4 192.168.99.1 64 bytes from 192.168.99.1: icmp_seq=1 ttl=255 time=1.65 ms ``` You can invert this trick and cap the TTL of your own *outgoing* request with `-t`. The packet dies in transit, and the router where it died reports back with an ICMP Time Exceeded: ``` j@lab:~$ ping -c 3 -t 1 3.3.3.3 PING 3.3.3.3 (3.3.3.3) 56(84) bytes of data. From 192.168.99.1 icmp_seq=1 Time to live exceeded From 192.168.99.1 icmp_seq=2 Time to live exceeded From 192.168.99.1 icmp_seq=3 Time to live exceeded j@lab:~$ ping -c 3 -t 2 3.3.3.3 From 10.0.12.2 icmp_seq=1 Time to live exceeded ``` TTL 1 dies at the first router (192.168.99.1), TTL 2 dies at the second (10.0.12.2). Increment the TTL one hop at a time and you have manually reinvented traceroute, which is literally how that tool works under the hood. ## The Statistics Block The summary at the end is where the operational signal lives: **Packet loss.** 0% is the only healthy number on a stable path. Consistent low-grade loss (1-5%) usually points at congestion, a duplex mismatch, or a dirty link; 100% loss with no error messages means something is silently eating your packets (a very different failure from an unreachable, as explained in [timeout vs destination unreachable](https://www.pinglabz.com/ping-timeout-vs-unreachable/)). **rtt min/avg/max/mdev.** The last value, mdev (mean deviation), is effectively jitter. A path with avg 4 ms and mdev 0.3 ms is stable. A path with avg 4 ms and mdev 40 ms has a queueing or buffering problem that will hurt voice and video long before it hurts a file transfer, which is exactly the kind of thing [QoS](https://www.pinglabz.com/qos/) exists to manage. ## What a Successful Ping Does Not Prove Ping answers one question: can an ICMP echo make it there and back right now. It does not prove the application works (TCP port 443 can be filtered while ICMP passes), it does not prove the path is symmetric (the reply may return over a different link than the request took), and it does not measure your data traffic's real latency (ICMP is commonly deprioritized or rate-limited on router control planes). A failed ping is not proof of an outage either: plenty of firewalls drop ICMP by policy while happily forwarding everything else. Treat ping as the first data point, not the verdict. When it fails, the specific way it fails tells you where to look next, and when it succeeds, the TTL and RTT patterns tell you what path you likely took. ## Key Takeaways - Ping is ICMP Echo Request (type 8) out, Echo Reply (type 0) back. No TCP, no UDP, no ports. - The default 56-byte payload becomes 84 bytes on the wire: payload + 8 ICMP + 20 IPv4. - The TTL shown is the reply's remaining TTL. Subtract it from the sender's initial value (64 Linux, 128 Windows, 255 Cisco) to count return-path hops. - `ping -t N` caps your outgoing TTL and gets you a Time Exceeded from hop N. That mechanism is the foundation of traceroute. - mdev in the statistics line is jitter. Watch it as closely as loss. - A successful ping proves two-way ICMP reachability at this moment, nothing more. Next in the cluster: [decoding every ping failure message](https://www.pinglabz.com/ping-timeout-vs-unreachable/), or head back to the [complete ping guide](https://www.pinglabz.com/ping/) for the full reading order. ### SD-WAN Routing Explained: Underlay, Overlay, and OMP URL: https://www.pinglabz.com/sd-wan-routing/ Last updated: 2026-07-08T13:00:00.000Z SD-WAN routing confuses people because there are really two routing problems happening at once, and they are easy to mix up. There is the routing inside the SD-WAN fabric - how sites learn each other's prefixes across the overlay - and there is the routing of the physical transports underneath. Keep those two layers separate and SD-WAN routing becomes straightforward. This post walks through both, using the Cisco Catalyst SD-WAN (formerly Viptela) model. For the cluster overview, see the [SD-WAN complete guide](https://www.pinglabz.com/sd-wan/). ## Two layers: underlay and overlay Every SD-WAN design has two routing layers stacked on top of each other. The physical transports - MPLS, broadband, LTE/5G LayerUnderlay What it has to do Just provide IP reachability between site edge routers The encrypted tunnel fabric built across the transports LayerOverlay What it has to do Carry the actual site-to-site enterprise routing The important shift in mindset: the underlay does **almost no routing work**. It does not need to know a single enterprise prefix. Its only job is to get an SD-WAN router's transport interface reachable to the others - usually a default route from the ISP, or a simple route over MPLS. All the real routing intelligence lives in the overlay. ## TLOCs: how the fabric describes a connection Before the overlay routing makes sense, one term: a **TLOC** (Transport Locator) is how SD-WAN identifies a router's attachment to a transport. A TLOC is essentially the tuple of "this router, on this transport (color), with this encapsulation." A router with an MPLS link and a broadband link has two TLOCs. TLOCs are the next-hops of the SD-WAN world. When the fabric advertises a prefix, it advertises it as reachable *via a TLOC*. Steering traffic onto MPLS instead of broadband is, underneath, choosing one TLOC over another. ## OMP: the control plane of the fabric The overlay's routing protocol is **OMP** \- the Overlay Management Protocol. If you know BGP, OMP will feel familiar: it is a single protocol that distributes everything the fabric needs, and the edge routers do not peer with each other directly. Instead, every edge router peers with a central controller (the vSmart controller), and the controller reflects routing information to everyone - the same hub-and-spoke control model as a BGP route reflector. OMP carries three things: - **OMP routes** \- the enterprise prefixes at each site, advertised with their TLOC next-hop. - **TLOC routes** \- the transport attachment points themselves, with their public/private addresses. - **Service routes** \- advertisements for services (firewall, IPS) that traffic can be steered through. Because OMP runs to a central controller, policy is applied centrally. The controller can shape what each site learns and which paths it prefers, without anyone touching an edge router. ## Service-side routing: getting LAN prefixes into OMP OMP handles the fabric. It does not, by itself, know about the LAN behind each site. That is the **service side** \- the VRF/VPN facing the local network. On the service side, an SD-WAN edge router runs ordinary routing. It learns local subnets through connected interfaces, static routes, OSPF, or BGP with the site's LAN switches. Those learned prefixes are then **redistributed into OMP** so the rest of the fabric can reach them. The reverse happens too: OMP routes from other sites are redistributed back into the local OSPF or BGP so LAN devices have a path out. So the full chain for a packet's route is: LAN routing protocol redistributes into OMP, OMP carries it across the fabric to every other site, and OMP redistributes back into each remote site's LAN routing protocol. The edge router is the translation point at both ends. ## Path selection: the part that makes SD-WAN "SD" This is where SD-WAN routing stops resembling traditional routing. Traditional routing picks one best path by metric and sends everything down it. SD-WAN picks paths **per application**, based on measured performance. The edge routers continuously probe each tunnel with BFD, measuring loss, latency, and jitter. A centralized **application-aware routing** policy then defines SLA classes - for example, "voice needs under 150 ms latency, under 2 percent loss." Traffic for an application is sent down whichever transport currently meets its SLA. If the preferred transport degrades, that application is moved to a transport that still meets the SLA, while other traffic may stay where it is. The result: a brownout on the MPLS circuit can move voice to broadband within a couple of seconds, automatically, with no routing reconvergence in the traditional sense. The OMP routes did not change - the policy chose a different TLOC. ## The configuration shape Most SD-WAN routing is expressed as centralized policy on the controller rather than per-router CLI, but the per-site service-side routing is recognizable IOS XE: ``` ! Service-side VRF on a Catalyst SD-WAN edge router vrf definition 10 rd 1:10 address-family ipv4 ! router ospf 10 vrf 10 redistribute omp subnets ! ! OMP redistributes the service-side routes back into the fabric router omp address-family ipv4 vrf 10 advertise ospf ``` ## Common gotchas Sites form tunnels but cannot reach each other's LANs Service-side prefixes are not being redistributed into OMP, or OMP routes are not redistributed back into the LAN protocol. No OMP routes learned at all The edge router has no control connection to the vSmart controller - an underlay reachability or certificate problem. Traffic uses the "wrong" transport Application-aware routing policy, or its absence - by default OMP load-shares across equal TLOCs. Voice quality drops but routing looks fine No SLA class defined for voice, so it is not being steered away from a degraded transport. A whole site is unreachable after an ISP change Underlay default route or transport-interface addressing broke - the overlay cannot form without underlay reachability. ## Key takeaways SD-WAN routing is two layers. The underlay (MPLS, broadband, LTE) only has to provide IP reachability between edge routers and carries no enterprise prefixes. The overlay carries the real routing, using OMP - a BGP-like protocol where every edge router peers with a central controller that reflects routes, advertising prefixes via TLOC next-hops. Each site's LAN prefixes are learned by ordinary OSPF, BGP, or static routing on the service side and redistributed into OMP, then redistributed back into the LAN at remote sites. The defining feature is application-aware path selection: BFD probes measure each tunnel, and centralized SLA policy steers each application onto whichever transport currently meets its requirements. For the SD-WAN cluster, see the [SD-WAN pillar](https://www.pinglabz.com/sd-wan/). ### What Is a Wireless LAN Controller (WLC)? URL: https://www.pinglabz.com/what-is-a-wireless-lan-controller/ Last updated: 2026-07-06T13:00:00.000Z A wireless LAN controller, or WLC, is the device that turns a pile of access points into a managed wireless network. Without one, every access point is an island you configure and troubleshoot by hand. With one, the access points become radios and the controller becomes the brain. This post explains what a WLC is, the problem it solves, how the controller-and-AP split works, and the forms a WLC comes in. For the cluster overview, see the [Cisco Wireless complete guide](https://www.pinglabz.com/wireless/). ## The problem a WLC solves An access point on its own is an **autonomous** (or standalone) AP. It holds its own full configuration: SSIDs, security, radio settings, VLAN mappings. One or two of those is fine. Fifty of them is a management problem, and a few hundred is unworkable. Every config change has to be pushed to every AP. There is no network-wide view of which channels are in use or where interference is. When a client walks from one AP to the next, nothing coordinates a clean handoff. Security policy lives in dozens of separate places. Autonomous APs do not *scale* \- not because the radios are weak, but because the management does not centralize. A WLC fixes this by making one device responsible for the configuration, coordination, and policy of the entire AP fleet. ## The split-MAC model The WLC architecture works by splitting the job of an access point into two halves. This is called the **split-MAC** model. Transmitting and receiving RF Configuration of all APs and SSIDs Beacons and probe responses RF management - channel and power assignment Encryption/decryption of frames Client authentication and security policy Time-critical acknowledgements Roaming coordination between APs An AP managed by a controller is called a **lightweight** AP. It keeps the time-sensitive radio work that has to happen locally in microseconds, and hands everything else - the slower management and decision-making - to the controller. A lightweight AP holds almost no standalone configuration. Plug it in, it finds its controller, and the controller tells it what to be. ## CAPWAP: the tunnel between AP and controller A lightweight AP and its WLC talk over **CAPWAP** \- Control And Provisioning of Wireless Access Points. CAPWAP forms two logical tunnels between each AP and the controller: - A **control tunnel** carries management - configuration, firmware, RF instructions, client state. It is encrypted. - A **data tunnel** carries the actual user traffic from wireless clients. In the common centralized design, client traffic is tunneled inside CAPWAP all the way back to the controller, and the controller places it onto the wired network. That means an AP only needs IP reachability to its WLC - it can sit anywhere routable. (A branch-office variant, FlexConnect, lets the AP switch traffic locally instead of tunneling it, but the control relationship is the same.) ## What the controller does for you Once the APs are lightweight and joined, the WLC delivers the things a pile of autonomous APs cannot: - **Single point of configuration.** Define an SSID once; every AP serves it. Change security once; it applies everywhere. - **RF management.** The controller sees the whole RF picture and assigns channels and transmit power to minimize interference, and routes clients around a failed AP. - **Seamless roaming.** Because the controller holds client state centrally, a client moving between APs keeps its session and IP address - no re-authentication, no dropped call. - **Centralized security.** Authentication, guest portals, and rogue-AP detection are all coordinated in one place. ## How a lightweight AP finds its controller Because a lightweight AP holds almost no configuration, the first thing it has to do when it powers on is locate a controller to join. This is the **discovery and join** process, and it runs every time an AP boots. The AP works through a list of methods to learn a controller's IP address: a controller address it has cached from a previous join, DHCP option 43 (a DHCP scope option that hands out the controller IP), a DNS lookup of a well-known name, or a broadcast on the local subnet. Once it has one or more candidate controllers, it sends a CAPWAP discovery request, the controllers reply, and the AP selects one and sends a join request. After joining, the AP downloads its configuration and, if its firmware does not match the controller's, a new image - which is why an AP often reboots once shortly after its first join. From then on it is fully managed. This is also why a misconfigured DHCP option 43 is one of the most common reasons "the AP will not come online": the radio is fine, it simply never learned where its controller is. ## The forms a WLC takes Hardware appliance What it is A dedicated physical controller Typical use Mid-size to large campus deployments Virtual / cloud-hosted What it is The controller software as a VM Typical use Virtualized data centres, scalable sites Embedded / switch-integrated What it is Controller function running on a stack or AP Typical use Smaller sites with no separate appliance Cloud-managed What it is Management plane in the cloud (e.g. Meraki model) Typical use Distributed sites, minimal on-site gear On the Cisco side, the modern controller is the **Catalyst 9800** family, which runs IOS XE and ships as an appliance, a VM (9800-CL), or embedded on Catalyst 9000 switches and APs. The older AireOS controllers (such as the 5520 and 3504) are the previous generation. Other vendors have direct equivalents - the concept is universal even though the product names differ. ## Common points of confusion Does every wireless network need a WLC? No. A handful of APs can run autonomous. The WLC earns its place at scale. If the WLC fails, does Wi-Fi stop? In a centralized design, largely yes - which is why controllers are deployed in resilient pairs. FlexConnect APs keep serving clients through a controller outage. Does user traffic always go through the controller? In centralized mode, yes, via the CAPWAP data tunnel. FlexConnect switches it locally instead. Is a WLC the same as a wireless router? No. A home router is one integrated AP. A WLC manages many separate APs. ## Key takeaways A wireless LAN controller is the central brain for a fleet of access points. It exists because autonomous APs do not scale: configuration, RF coordination, roaming, and security all need a single point of control. The WLC uses the split-MAC model - the lightweight AP keeps the real-time radio work, the controller takes configuration, RF management, and policy - and the two communicate over CAPWAP control and data tunnels. The result is one place to configure SSIDs, network-wide channel and power management, seamless roaming, and centralized security. A WLC can be a hardware appliance, a virtual machine, embedded in a switch, or cloud-managed; on Cisco the current platform is the IOS XE-based Catalyst 9800. For the wireless cluster, see the [Cisco Wireless pillar](https://www.pinglabz.com/wireless/). ### Nmap Output Formats: -oN, -oX, -oG, -oA and Parsing Results URL: https://www.pinglabz.com/nmap-output-formats/ Last updated: 2026-08-01T19:31:05.000Z A scan you cannot reproduce is a story, not evidence. The moment an Nmap result matters - a firewall change you signed off on, a rogue service you found on a production VLAN, an asset inventory the audit team will lean on - you need it written to a file in a format that survives the terminal scrollback. Saved output lets you diff a segment week over week to catch new or closed ports, feed the results into other tooling, and prove what the network looked like at a specific timestamp. This guide is part of [the complete Nmap guide](https://www.pinglabz.com/nmap/), and it covers the output formats Nmap can write, when each one earns its place, and the short shell one-liners that turn a saved scan into an answer. ## Why you save output at all Running Nmap and reading the screen is fine for a quick look. Everything past that wants a file. Four reasons come up constantly in network work. First, **evidence**: a timestamped scan file is defensible in a change record or an incident timeline in a way that a screenshot never is. Second, **diffing over time**: if you scan the same subnet on a schedule and store each run, you can compare them and immediately see what changed - a new listener that appeared overnight, a port that closed after a patch. Third, **feeding other tools**: XML output is the input format for a whole ecosystem of parsers, importers, and reporting engines. Fourth, **asset inventories**: grepable output is trivial to fold into a spreadsheet or a CMDB import. Pick the format that matches the downstream job, and Nmap will write it in the same pass as the scan. ## The output formats Nmap writes several formats, each selected by a flag that takes a filename. You can request more than one at a time, and `-oA` is a shortcut for "give me all the main ones." \-oN (normal) The human-readable report you already see on screen, saved to a file. Best for people. Read it, paste it in a ticket, archive it. \-oX (XML) Structured XML with full detail: ports, states, services, CPEs. Best for programmatic parsing, ndiff, and importing into tools. \-oG (grepable) One line per host, so grep and awk can slice it directly. Best for quick command-line pipelines and ad-hoc inventories. \-oA (all) Writes normal, XML, and grepable at once from one basename. The safe default when you are not sure which you will need later. \-oS (script kiddie) Normal output mangled into l33t-speak. A joke Nmap actually ships. Never use it for real work. It exists purely as an aside. \- (stdout) Pass a single dash as the filename to send a format to stdout. Lets you pipe grepable output straight into grep without a temp file. The two you will reach for most are `-oA basename` when you want a full record on disk, and `-oG -` when you want to pipe grepable output straight into a command. The dash-as-filename trick works for any format flag, so `-oX -` streams XML to stdout for a downstream parser without ever touching disk. ## Normal output, for humans Normal output is exactly what Nmap prints to the terminal, and it is what most engineers read. Here is a version-detection scan of localhost on a Debian VM, saved with `-oA scandemo` (which produced `scandemo.nmap`, `scandemo.gnmap`, and `scandemo.xml` in one shot). This is the `.nmap` file: ``` # Nmap 7.95 scan initiated Sun Jul 5 15:05:24 2026 as: nmap -sV -p21,22,80 -oA scandemo 127.0.0.1 Nmap scan report for localhost (127.0.0.1) Host is up (0.000090s latency). PORT STATE SERVICE VERSION 21/tcp open ftp vsftpd 3.0.5 22/tcp open ssh OpenSSH 10.0p2 Debian 7+deb13u4 (protocol 2.0) 80/tcp open http nginx Service Info: OSs: Unix, Linux; CPE: cpe:/o:linux:linux_kernel Service detection performed. Please report any incorrect results at https://nmap.org/submit/ . # Nmap done at Sun Jul 5 15:05:30 2026 -- 1 IP address (1 host up) scanned in 6.84 seconds ``` This reads cleanly: three open ports, each with a service and a version string that `-sV` probed out (that is [Nmap version detection](https://www.pinglabz.com/nmap-version-detection/) at work). It is the format to paste into a change ticket or an incident note, because a human can read it without any tooling. What it is not good for is machine parsing - the layout is aligned for eyes, not for `awk`. ## Grepable output, for pipelines Grepable output collapses each host onto a single line, which is precisely what makes it easy to slice with standard Unix tools. Here is the `scandemo.gnmap` file from the same run: ``` # Nmap 7.95 scan initiated Sun Jul 5 15:05:24 2026 as: nmap -sV -p21,22,80 -oA scandemo 127.0.0.1 Host: 127.0.0.1 (localhost) Status: Up Host: 127.0.0.1 (localhost) Ports: 21/open/tcp//ftp//vsftpd 3.0.5/, 22/open/tcp//ssh//OpenSSH 10.0p2 Debian 7+deb13u4 (protocol 2.0)/, 80/open/tcp//http//nginx/ # Nmap done at Sun Jul 5 15:05:30 2026 -- 1 IP address (1 host up) scanned in 6.84 seconds ``` Notice the single `Host: ... Ports:` line packs every port for that host into one field, comma-separated, each in a `port/state/proto//service//version/` shape. That regular structure is the whole appeal: because it is one line per host, you can `grep` for a state or `awk` for a field without wrestling with a multi-line report. From the CML lab, the shorter two-port form shows the same pattern across two hosts at once: ``` $ nmap -oG - -p 22,80 10.10.10.1 10.10.10.21 # Nmap 7.95 scan initiated Sun Jun 28 07:12:51 2026 as: nmap -oG - -p 22,80 10.10.10.1 10.10.10.21 Host: 10.10.10.1 () Status: Up Host: 10.10.10.1 () Ports: 22/open/tcp//ssh///, 80/closed/tcp//http/// Host: 10.10.10.21 () Status: Up Host: 10.10.10.21 () Ports: 22/closed/tcp//ssh///, 80/open/tcp//http/// # Nmap done at Sun Jun 28 07:12:51 2026 -- 2 IP addresses (2 hosts up) scanned in 0.24 seconds ``` Here `-oG -` streamed straight to stdout. The router (10.10.10.1) has SSH open and HTTP closed (it answered with a RST, not a filter), and the web server (10.10.10.21) is the mirror image. One line per host, ready to pipe. ## XML output, for tools XML is the format you feed to software. It carries everything the scan learned in a strict, parseable structure, and it is the input for `ndiff`, for importers, and for reporting engines. Here is the ports section of `scandemo.xml` from the same localhost scan, trimmed to the structure that matters: ``` cpe:/a:vsftpd:vsftpd:3.0.5 cpe:/a:openbsd:openssh:10.0p2cpe:/o:linux:linux_kernel cpe:/a:igor_sysoev:nginx ``` Every fact from the normal report is here as an attribute a parser can pull without guessing: each `` has a `` element (with the reason and TTL) and a `` element carrying the product, version, and one or more `` entries. Those CPE strings (`cpe:/a:vsftpd:vsftpd:3.0.5`, `cpe:/a:openbsd:openssh:10.0p2`) are the standardized identifiers that vulnerability tooling matches against a CVE database, which is why XML is the format to save when the scan feeds anything automated. You would never read this by eye, and you would never `grep` it reliably - you hand it to a parser. ## Parsing the saved output The reason grepable output exists is to make quick answers a one-liner. These operate on the saved files above and are standard, everyday usage. To list every open service across a saved `.gnmap`, split on the comma-separated port field and keep the ones marked open: ``` $ grep -oE '[0-9]+/open/[^,]+' scandemo.gnmap 21/open/tcp//ftp//vsftpd 3.0.5/ 22/open/tcp//ssh//OpenSSH 10.0p2 Debian 7+deb13u4 (protocol 2.0)/ 80/open/tcp//http//nginx/ ``` To pull a clean host-plus-open-port inventory across a multi-host grepable file, a short `awk` walks each host line and prints the port numbers that are open: ``` $ awk '/Ports:/{ip=$2; n=split($0,p,", "); for(i=1;i<=n;i++) if(p[i]~/\/open\//){split(p[i],f,"/"); print ip, f[1]}}' scan.gnmap ``` That is the payoff of one-line-per-host: no XML library, no scripting language, just the tools already on every box. For anything more structured than a grep - correlating services to CVEs, building a report, diffing two runs - switch to XML and let a real parser do it. ## Verbosity, progress, and diffing over time A few flags shape what lands in the output and how much you see while a long scan runs. `-v` (and `-vv` for more) raises verbosity so Nmap reports open ports as it finds them rather than only at the end. `--stats-every 10s` prints a progress line on a fixed interval, which is a lifesaver on a big sweep so you know it has not hung. `--open` trims the report to only open ports, cutting the closed and filtered noise when you only care about what is listening. And `--reason` adds the *why* behind each state (the `syn-ack`, `reset`, and `no-response` values you saw earlier), which is worth saving whenever the state of a port might be questioned later. The real long-term value of saved output is **diffing**. Nmap ships `ndiff`, which takes two XML scan files and reports exactly what changed between them - hosts that appeared or vanished, ports that opened or closed, services that changed version. Scan a subnet on a schedule, save each run with `-oX`, and `ndiff yesterday.xml today.xml` surfaces the new listener that showed up overnight without you rereading a single line. This is how you turn Nmap from a point-in-time tool into a change-detection one. If you would rather grow one running record than keep dated files, `--append-output` tells Nmap to append to an existing output file instead of overwriting it, though for diffing you generally want discrete per-run files that `ndiff` can compare cleanly. The same idea works in the other direction, and hardly anyone bothers. Baselining `show ip access-lists` on a schedule and diffing the hit counts is [the cheapest way to catch enumeration that arrives too slowly to trip anything else](https://www.pinglabz.com/detecting-nmap-scans-cisco-blue-team/), because an ACE counter is cumulative and does not care whether the packets took four seconds or four days. ## Choosing a format in practice The decision is quick once the downstream job is clear. Reading it yourself or pasting into a ticket? `-oN`. Grepping or awking a fast inventory? `-oG`, or `-oG -` to skip the temp file. Feeding a parser, matching CPEs to CVEs, or diffing with `ndiff`? `-oX`. Not sure which you will need later, or building an audit trail? `-oA` and keep all three. The runtime cost of writing extra formats is effectively zero - the scan does the work, and serializing the results is trivial - so on any scan that matters, defaulting to `-oA` costs nothing and saves you re-running later. How fast the scan itself finishes is a separate question covered in [Nmap timing and performance](https://www.pinglabz.com/nmap-timing-performance/). ## Key takeaways Save your scans, because a result you cannot reproduce or diff is not evidence. Normal output (`-oN`) is for humans, grepable output (`-oG`) collapses one line per host so `grep` and `awk` can slice it in a pipeline, and XML (`-oX`) carries the full structured detail - states, services, and CPE identifiers - that parsers, importers, and `ndiff` depend on. When in doubt, `-oA basename` writes all three at once for nothing extra, and the dash-as-filename trick (`-oG -`) streams a format straight into a pipe. Layer on `-v`, `--stats-every`, `--open`, and `--reason` to control what you capture and how much you see mid-scan, then use `ndiff` across dated XML files to catch new or closed ports over time and `--append-output` when you want one growing record. Pair saved output with solid [version detection](https://www.pinglabz.com/nmap-version-detection/) and you have a repeatable, auditable view of what is really running on your network. For how this fits with discovery, scanning, timing, and the rest of the toolkit, work through [the complete Nmap guide](https://www.pinglabz.com/nmap/). ### Nmap Timing and Performance: -T0 to -T5 and Rate Controls URL: https://www.pinglabz.com/nmap-timing-performance/ Last updated: 2026-08-01T19:31:04.000Z Timing is the lever that decides whether an Nmap scan finishes before your coffee gets cold or runs so slowly it slips under an intrusion-detection sensor. When you are sweeping a /16 for an asset inventory, you want every packet Nmap can safely push. When you are validating a firewall rule on a production segment that feeds a SIEM, you want the opposite: a trickle of probes that never trips a rate-based alert. Same tool, same target, wildly different behavior - and the difference is almost entirely down to how you tune timing. This guide is part of [the complete Nmap guide](https://www.pinglabz.com/nmap/), and it focuses on the two knobs that matter most in the field: the six timing templates (`-T0` through `-T5`) and the fine-grained rate and retry controls underneath them. ## Why timing matters for network engineers There are two directions you can push a scan, and they pull against each other. The first is **speed**: cover a large address range or a big port set in as little wall-clock time as possible. If you are auditing your own data-center VLANs on a reliable, low-latency network, there is no reason to wait - fast is correct. The second is **stealth**, or more precisely, staying under a rate threshold. Every IDS/IPS worth its license watches for a host firing connection attempts at many destinations or ports per second. Slow the probes down enough and the scan disappears into the noise floor of normal traffic. That is the same reasoning behind the deliberate techniques in [Nmap firewall and IDS evasion](https://www.pinglabz.com/nmap-firewall-evasion/), and timing is the cheapest evasion control you have because it is a single flag. The honest tradeoff is time. Going slow to dodge a sensor can turn a one-second scan into a forty-second one, as you will see below with real captures. So the practical question is never "fast or slow" in the abstract; it is "what is the fastest I can go without tripping the thing watching this segment." The templates give you six pre-tuned answers to that question. ## The six timing templates: -T0 to -T5 Nmap ships six timing templates. Each one sets a bundle of underlying parameters (probe delays, parallelism, retries, timeouts) to sensible values for a given intent, so you rarely need to touch the individual knobs. `-T3` is the default and is what runs when you pass no timing flag at all. Everything below `-T3` trades speed for stealth; everything above trades caution for speed. \-T0 Paranoid Serializes every probe with waits of up to 5 minutes between them. Built for IDS evasion. A single host can take hours or days. \-T1 Sneaky Serialized probes with a \~15 second delay between each. Also evasion-focused, but merely slow rather than glacial. \-T2 Polite Slows down to use less bandwidth and target resources. The go-to when you must stay under an IDS rate threshold. \-T3 Normal The default. Balanced, adaptive, no special aggression or restraint. What you get when you pass no -T flag at all. \-T4 Aggressive Faster timeouts and higher parallelism for reliable networks. The common fast default for labs and modern LANs. \-T5 Insane Maximum speed. Sacrifices accuracy for raw throughput. Only on very fast, reliable links, or you will miss ports. The pattern is worth internalizing. `-T0` and `-T1` exist to defeat detection: they serialize probes (one at a time, never in parallel) and insert long fixed delays between them, so the packet rate never spikes. `-T4` and `-T5` do the reverse, firing many probes concurrently and giving up on slow responses quickly. `-T2` and `-T3` sit in the middle, with `-T2` the polite choice when you care about the target's load or a sensor's patience. Defeating detection is the honest description of what the slow end does, and it is the uncomfortable half of the subject: [the patient scans are the ones nobody catches](https://www.pinglabz.com/detecting-nmap-scans-cisco-blue-team/). Syslog rate-limit warnings never fire at `-T1`, packet-rate alarms never trip, and the only thing left holding the line is a cumulative counter that almost nobody baselines. ## The proof: T4 versus T2 on the same scan Templates are easy to describe and easy to underestimate. So here is the same scan run twice against the same host in the PingLabz Nmap Recon Lab (a Cisco router at 10.10.10.1, Nmap 7.95), changing nothing but the timing template. The scan is a fast top-100-port sweep (`-F`). First, aggressive: ``` $ nmap -T4 -F 10.10.10.1 ... 22/tcp open ssh ; 23/tcp open telnet ... Nmap done: 1 IP address (1 host up) scanned in 1.37 seconds ``` Now the identical scan, polite: ``` $ nmap -T2 -F 10.10.10.1 Nmap scan report for 10.10.10.1 Host is up (0.0033s latency). Not shown: 98 closed tcp ports (reset) PORT STATE SERVICE 22/tcp open ssh 23/tcp open telnet MAC Address: AA:BB:CC:00:5B:00 (Unknown) Nmap done: 1 IP address (1 host up) scanned in 40.67 seconds ``` Same host, same 100 ports, same results (SSH and Telnet open, 98 closed). The only difference is the clock: **1.37 seconds at -T4 versus 40.67 seconds at -T2**, a factor of roughly thirty. That thirty-fold penalty is not waste - it is the whole point. The `-T2` run spreads its probes out far enough that the per-second rate stays low, which is exactly what keeps it under an IDS rate threshold. On a lab segment with nothing watching, that patience buys you nothing and you would obviously reach for `-T4`. On a production edge behind a tuned IPS, that patience is the difference between a clean audit and a page to the SOC. ## Fast on a reliable network To show the other side, here is `-T4` doing what it does best - a wider port range against a real internet host, scanme.nmap.org, captured from a Debian VM. The `--reason` flag (covered more in [Nmap port scanning](https://www.pinglabz.com/nmap-port-scanning/)) makes Nmap print why it assigned each state. ``` $ nmap -T4 --reason -p1-150 scanme.nmap.org Starting Nmap 7.95 ( https://nmap.org ) at 2026-07-05 15:06 PDT Nmap scan report for scanme.nmap.org (45.33.32.156) Host is up, received reset ttl 50 (0.018s latency). Other addresses for scanme.nmap.org (not scanned): 2600:3c01::f03c:91ff:fe18:bb2f Not shown: 143 closed tcp ports (reset) PORT STATE SERVICE REASON 22/tcp open ssh syn-ack ttl 50 80/tcp open http syn-ack ttl 49 135/tcp filtered msrpc no-response 136/tcp filtered profile no-response 137/tcp filtered netbios-ns no-response 138/tcp filtered netbios-dgm no-response 139/tcp filtered netbios-ssn no-response Nmap done: 1 IP address (1 host up) scanned in 1.52 seconds ``` 150 ports across the public internet in **1.52 seconds**. Note the `Not shown: 143 closed tcp ports (reset)` line: those hosts answered with a RST, so Nmap is confident they are closed (the host actively refused), not filtered. The five `no-response` ports (135-139) are filtered - something dropped the probe silently, most likely a firewall. Over a well-behaved link, `-T4` gathers all of that almost instantly. Push that same scan to `-T5` on a lossy or high-latency path and you would start seeing ports mislabeled, because the insane template gives up on slow replies before they arrive. ## Fine-grained timing controls The templates are bundles of lower-level settings, and any of those settings can be overridden on the command line. You reach for these when a template is close but not quite right - for example, you want `-T4`'s parallelism but with a hard cap on packet rate so you do not saturate a slow WAN link. Override flags always win over the template, so `-T4 --max-rate 100` is perfectly valid. \--min-rate / --max-rate Floor and ceiling on packets sent per second. The most direct throttle. Set --max-rate to stay under a known IDS threshold. \--max-retries How many times Nmap re-probes a port with no answer. Lower it to go faster on clean links; raise it on lossy paths. \--host-timeout Give up on a single host after this long. Keeps one dead or firewalled host from stalling a whole sweep. \--scan-delay / --max-scan-delay Minimum wait between probes to one host (and a cap on it). The manual version of what -T0/-T1 do for evasion. \--min-parallelism / --max-parallelism How many probes are outstanding at once per host group. Raise for speed on reliable links; force to 1 to fully serialize. \--min-hostgroup Minimum number of hosts scanned in parallel per batch. Raise it for big /16 sweeps so results stream in steadily. A few of these deserve field notes. `--max-rate` is the single most useful control when you have an actual number to hit - if you know the IPS on a segment alerts above 50 connections per second, `--max-rate 40` gives you a hard guarantee no template can. `--host-timeout` saves large sweeps: without it, one firewalled host that answers nothing can hold up the whole run while Nmap patiently retries. And `--scan-delay` is the honest, tunable version of the `-T0`/`-T1` stealth behavior - if you need a specific inter-probe gap rather than the template's fixed values, set it directly. ## Port count is the biggest lever Before you spend an afternoon tuning parallelism, understand this: for most scans, the **number of ports** you scan dominates the runtime far more than the timing template does. Nmap's default is the top 1,000 ports per host. A full `-p-` scan hits all 65,535, which is 65 times the work. The `-F` (fast) flag drops you to the top 100, and an explicit list like `-p 22,80,443` is faster still. The compact `-F` runs above finished in a second or two precisely because they scanned 100 ports, not 65,535. So the first question for a slow scan is rarely "should I bump to -T5"; it is "do I actually need every port." If you are verifying that a firewall permits SSH and HTTP, scan those two ports and nothing else. Choosing the right port set is covered in depth in [Nmap port scanning](https://www.pinglabz.com/nmap-port-scanning/), and it will save you more time than any timing flag. Once the port set is right, then tune timing to fit the network. ## When to choose T2 versus T4 in practice For day-to-day network engineering the choice usually collapses to two templates. Reach for `-T4` when you are on a reliable, low-latency network you control - a lab, a data-center VLAN, a wired LAN - and you want results now. It is the sane fast default for asset inventories and big internal sweeps, and the `-T4` captures above show it losing no accuracy on a clean path. Reach for `-T2` when the target segment is monitored and you must not trip a rate-based alert, or when you are scanning across a fragile or bandwidth-constrained link where hammering the target would cause collateral damage. The forty-second `-T2` run earlier is the price of staying quiet, and on a production edge it is a price worth paying. `-T0` and `-T1` are specialist tools for genuine detection evasion, and `-T5` is for the rare case where you have a pristine link and are willing to trade a little accuracy for raw speed. Most engineers live between `-T2` and `-T4` and are right to. ## Key takeaways Timing in Nmap is a deliberate tradeoff between wall-clock speed and staying under a detection threshold, and the six templates (`-T0` paranoid through `-T5` insane) give you pre-tuned answers with `-T3` as the default. The real proof is in the captures: the same top-100 scan of 10.10.10.1 ran 1.37 seconds at `-T4` and 40.67 seconds at `-T2`, identical results, thirty times the patience - and that patience is what keeps a scan under an IDS rate threshold. When a template is close but not exact, override the underlying controls (`--max-rate`, `--host-timeout`, `--scan-delay`, parallelism, host-group size) to fit the specific link. Above all, remember that the port count you choose is a bigger lever on runtime than any timing flag, so scan only the ports you need first, then tune. Use `-T4` for labs and reliable internal sweeps, `-T2` for anything monitored or fragile, and pair this with deliberate [firewall and IDS evasion](https://www.pinglabz.com/nmap-firewall-evasion/) techniques when the goal is stealth. For the full picture of how timing fits alongside discovery, scanning, and detection, work through [the complete Nmap guide](https://www.pinglabz.com/nmap/). ### Nmap Firewall and IDS Evasion: Fragmentation, Decoys, and Timing URL: https://www.pinglabz.com/nmap-firewall-evasion/ Last updated: 2026-08-01T19:31:03.000Z Nmap ships a whole family of options with names like fragmentation, decoys, and source-port spoofing, and it is easy to read that list as an attacker's cheat sheet. Flip the framing. As the engineer who owns the firewall and the IDS, these flags are your test harness: they are how you prove that your own controls actually catch and drop the tricks an attacker would use, instead of assuming they do. A firewall rule you have never tested is a hypothesis. Running these evasion techniques against your own perimeter turns the hypothesis into a result. This article walks the main evasion options with real captures from a Cisco lab and the internet, and it is a companion to [the complete Nmap guide](https://www.pinglabz.com/nmap/), which covers the rest of the scanning workflow. **One note before anything else.** Every technique below is for validating infrastructure you own or are explicitly authorized to test. Aimed at someone else's network, these options are evasion in the criminal sense, not the diagnostic one. Keep them inside your own perimeter or a signed engagement, and where the target is a shared box, use a sanctioned one like `scanme.nmap.org`, which Nmap maintains for exactly this purpose. ## Fragmentation: -f and --mtu The oldest trick in the book is packet fragmentation. `-f` splits the probe into tiny IP fragments (8 bytes of payload each), and `--mtu` lets you pick your own fragment size. The theory is that a simple packet filter inspecting only the first fragment - or one that does not reassemble at all - never sees the full TCP header and lets the pieces through. Here it is on the lab LAN, where nothing sits in the path to drop the fragments, and the SYN still gets its SYN/ACK back: ``` $ nmap -f --packet-trace -p 22 10.10.10.1 SENT (0.2116s) TCP 10.10.10.10:44682 > 10.10.10.1:22 S ttl=53 id=11425 iplen=44 seq=11120976 win=1024 RCVD (0.2153s) TCP 10.10.10.1:22 > 10.10.10.10:44682 SA ttl=255 id=33340 iplen=48 seq=4165816293 win=65535 Nmap scan report for 10.10.10.1 Host is up (0.0031s latency). PORT STATE SERVICE 22/tcp open ssh MAC Address: AA:BB:CC:00:5B:00 (Unknown) Nmap done: 1 IP address (1 host up) scanned in 0.24 seconds ``` On a flat LAN, fragmentation is a no-op: the packet reaches the host and 22 reports open. Now push the same idea across the internet to `scanme.nmap.org` and the result flips in an instructive way: ``` $ nmap -sS -f -p22,80 scanme.nmap.org Starting Nmap 7.95 ( https://nmap.org ) at 2026-07-05 15:06 PDT Nmap scan report for scanme.nmap.org (45.33.32.156) Host is up (0.021s latency). PORT STATE SERVICE 22/tcp filtered ssh 80/tcp filtered http Nmap done: 1 IP address (1 host up) scanned in 1.55 seconds ``` Both ports come back `filtered` \- and 22 and 80 on scanme are demonstrably open when you scan them normally. Something in the internet path dropped the tiny fragments. That is not a failure of the scan; it is the honest, useful result. A device between you and the target chose to discard malformed fragmentation rather than reassemble and forward it, which is precisely the behaviour you would want your own edge firewall to exhibit. When you run `-f` against your perimeter and get `filtered` back, that is your control doing its job. When you get `open`, you have found a gap. ## Decoys: -D Decoy scanning makes your real probe hide in a crowd. With `-D` you list spoofed source addresses, and Nmap sends copies of each probe from those decoys interleaved with your real one, so the target's logs show a scan coming from a dozen places at once and cannot easily tell which source is the human. The packet trace shows how convincing it is - three SYNs, identical in every field that matters, sourced from a decoy, then the real host, then another decoy: ``` $ nmap -sS -D 10.10.10.66,10.10.10.77 --packet-trace -p 22 10.10.10.1 SENT (0.2450s) TCP 10.10.10.66:56346 > 10.10.10.1:22 S ttl=56 id=38942 iplen=44 seq=590558048 win=1024 SENT (0.2451s) TCP 10.10.10.10:56346 > 10.10.10.1:22 S ttl=59 id=38942 iplen=44 seq=590558048 win=1024 SENT (0.2452s) TCP 10.10.10.77:56346 > 10.10.10.1:22 S ttl=48 id=38942 iplen=44 seq=590558048 win=1024 RCVD (0.2489s) TCP 10.10.10.1:22 > 10.10.10.10:56346 SA ttl=255 id=64983 iplen=48 seq=3681349475 win=65535 Nmap scan report for 10.10.10.1 Host is up (0.0038s latency). PORT STATE SERVICE 22/tcp open ssh ``` The target replied only to the real address (.10), because only it has a return path, but the log now records SYNs from .66, .10, and .77 that all look equally legitimate. The scan itself is unaffected. Across the internet, `-D RND:5` generates five random decoys and the scan still returns clean results: ``` $ nmap -sS -D RND:5 -p22,80 scanme.nmap.org PORT STATE SERVICE 22/tcp open ssh 80/tcp open http Nmap done: 1 IP address (1 host up) scanned in 0.36 seconds ``` For a defender, decoys are a test of whether your alerting keys on the true source or gets diluted by noise. A good detection pipeline should correlate on the source that actually completes handshakes, not on raw SYN counts per address. Run a decoy scan against your own IDS and see whether it still fingers you or whether it drowns in the polluted log. ## Source-port spoofing: --source-port Some firewalls carry sloppy legacy rules that trust traffic *from* certain ports - port 53 (DNS) and port 80 (HTTP) are the classic offenders, allowed inbound because someone once needed replies to get back through a stateless filter. `--source-port` (or its short form `-g`) lets you set your probe's source port to abuse that trust: ``` $ nmap -sS --source-port 53 -p22,80 scanme.nmap.org PORT STATE SERVICE 22/tcp open ssh 80/tcp open http Nmap done: 1 IP address (1 host up) scanned in 0.33 seconds ``` Sourcing from port 53 got the scan through cleanly. The reason this matters to you as a defender: if a probe from `--source-port 53` reaches a service that a probe from a random high port could not, you have a rule that trusts a source port, which is a stateless-era mistake worth ripping out. This is a one-line test that finds a specific, common misconfiguration. ## Timing as evasion Every IDS has a rate threshold - a certain number of probes per second from one source before it fires an alert. Scan slowly enough and you slide under it. The `-T` templates control this directly. Here is an identical top-100-port scan run at two timing levels: ``` $ nmap -T4 -F 10.10.10.1 ... 22/tcp open ssh ; 23/tcp open telnet ... Nmap done: 1 IP address (1 host up) scanned in 1.37 seconds $ nmap -T2 -F 10.10.10.1 Nmap scan report for 10.10.10.1 Host is up (0.0033s latency). Not shown: 98 closed tcp ports (reset) PORT STATE SERVICE 22/tcp open ssh 23/tcp open telnet MAC Address: AA:BB:CC:00:5B:00 (Unknown) Nmap done: 1 IP address (1 host up) scanned in 40.67 seconds ``` Same scan, same result, wildly different footprint: 1.37 seconds at `-T4` versus 40.67 seconds at `-T2`. The slow version spreads the probes out so thinly that a rate-based detector may never accumulate enough events in its window to trip. That is the evasion, and it is also the test - point a slow scan at your own sensors and find out how patient an attacker has to be before your IDS stops noticing. For the full breakdown of the timing templates and what each one tunes, see [Nmap timing and performance](https://www.pinglabz.com/nmap-timing-performance/). ## Other evasion knobs Beyond the big four, Nmap gives you a grab-bag of lower-level manipulations. Each one targets a specific weakness in how a filter or sensor inspects traffic. Test them against your own stack to see which your controls catch. \--data-length Pads probes with random bytes so they do not match a fixed-length signature. \--badsum Sends a bogus TCP/UDP checksum. Real hosts drop it; some IDS boxes still react, revealing themselves. \--spoof-mac Changes your source MAC (random, or a vendor prefix) to dodge layer-2 filtering and logging. \--ttl Sets a custom IP TTL so probes expire before a downstream sensor, but after the firewall. \-g / --source-port Fixes the source port (e.g. 53 or 80) to sneak past rules that trust it. idle scan -sI Bounces the scan off a third "zombie" host so your IP never appears in the target's logs at all. The idle scan (`-sI`) deserves a word because it is the purest form of source concealment: Nmap infers port state entirely from changes in an idle third party's IP ID counter, so the target only ever sees the zombie, never you. It is fiddly to set up and needs a genuinely idle host, but conceptually it is the reason "the logs only showed one internal server scanning us" is not proof of who ran the scan. The full catalogue lives in the [Nmap firewall and IDS evasion reference](https://nmap.org/book/man-bypass-firewalls-ids.html?ref=pinglabz.com). ## What a correctly configured firewall does The flip side of all this is knowing what a healthy result looks like. A stateful firewall does not just drop tricks; it drops anything that is not part of an established session, and Nmap has a scan type that exposes exactly that. The ACK scan (`-sA`) is a firewall-mapping tool: it sends bare ACKs and watches whether they draw a RST (unfiltered) or nothing (filtered). Run through the lab's Cisco ASA, every port comes back filtered: ``` $ nmap -sA -p 22,23,80,443 10.10.20.32 PORT STATE SERVICE 22/tcp filtered ssh 23/tcp filtered telnet 80/tcp filtered http 443/tcp filtered https Nmap done: 1 IP address (1 host up) scanned in 1.35 seconds ``` All filtered is the correct answer. The ASA is stateful, so an ACK with no matching connection in its table gets silently dropped regardless of port - it never reaches the host to draw a RST. Contrast that with a normal SYN scan of the same DMZ target, which shows the firewall's actual policy through the noise: ``` $ nmap --reason -p 22,23,80,443 10.10.20.32 PORT STATE SERVICE REASON 22/tcp open ssh syn-ack ttl 254 23/tcp filtered telnet no-response 80/tcp closed http reset ttl 254 443/tcp filtered https no-response Nmap done: 1 IP address (1 host up) scanned in 1.36 seconds ``` That is a well-behaved firewall in one screen: 22 is `open` (permitted and a service is listening), 80 is `closed` (the firewall permits it but the host answered with a RST - nothing is serving there), and 23 and 443 are `filtered` (silently dropped, `no-response`, the deny rules doing their work). If your evasion tests against a perimeter like this keep returning filtered, your ASA is earning its keep. For the deep dive on that behaviour and how to configure it, see the [Cisco ASA firewall guide](https://www.pinglabz.com/cisco-asa/), and for how ACK, FIN, and the other probes differ, the [Nmap scan types reference](https://www.pinglabz.com/nmap-scan-types/). ## The defender's checklist Turn every technique above into a control test you run on a schedule: - **Fragmentation:** your edge should return `filtered` for `-f` and `--mtu` probes. If a fragmented SYN reaches an internal host, your firewall is not reassembling. - **Decoys:** your IDS alert should still name the real source after a `-D` run. If the signal drowns in the decoys, tune your correlation. - **Source-port trust:** a probe from `--source-port 53` or `-g 80` must not reach anything a random high port cannot. If it does, delete the legacy rule. - **Timing:** know the slowest scan your IDS still catches. If `-T2` or `-T1` slips by, your rate window is too tight. - **Stateful baseline:** an `-sA` ACK scan through your firewall should come back all filtered. Anything `unfiltered` means non-established traffic is passing. That checklist puts you in the defender's chair for a moment, and it is worth staying there. [Read from behind the ACL, these techniques look very different](https://www.pinglabz.com/detecting-nmap-scans-cisco-blue-team/): decoys defeat source-based logging completely, fragmentation and source-port games barely dent an ACE counter, and knowing which is which tells you whether the evasion was worth the packets it cost. ## Key takeaways Nmap's evasion options are only "attacker tools" if you point them at someone else's network. Pointed at your own perimeter, they are the fastest honest test you have of whether your firewall and IDS do what you configured them to do. Fragmentation getting dropped in transit, decoys that fail to hide the real source, a source-port-53 probe that goes nowhere, a slow scan your sensor still catches, and an ACK scan that comes back all filtered are all wins - they are your controls being verified rather than assumed. Run these on a schedule, keep them strictly inside systems you own or are authorized to test, and read every result as a statement about your defences. To see how these probes relate to the underlying scan mechanics, work through the [scan types](https://www.pinglabz.com/nmap-scan-types/) and [timing and performance](https://www.pinglabz.com/nmap-timing-performance/) guides, tune the firewall itself with the [Cisco ASA guide](https://www.pinglabz.com/cisco-asa/), and keep [the complete Nmap guide](https://www.pinglabz.com/nmap/) as your map of the whole toolkit. ### Nmap NSE Scripts: The Scripting Engine, Categories, and Real Examples URL: https://www.pinglabz.com/nmap-nse-scripts/ Last updated: 2026-08-01T19:31:03.000Z By the time you have found an open port and identified the service behind it, you have answered "what is listening here?" but not "what can it tell me about itself?" That second question is where the Nmap Scripting Engine earns its place in a network engineer's toolkit. NSE turns Nmap from a port scanner into a lightweight active-inventory and audit tool: it will enumerate HTTP methods on your web tier, list the SSH algorithms a router negotiates, flag a directory listing you forgot to lock down, and pull banners off gear you inherited. If you are new to the tool, start with [the complete Nmap guide](https://www.pinglabz.com/nmap/) for the full workflow; this article drills into NSE specifically, using real output captured against a Cisco lab and a Debian host. ## What NSE actually is The Nmap Scripting Engine is a Lua interpreter bundled inside Nmap that runs small scripts against the hosts and ports you scan. The scripts ship with Nmap - roughly 600 of them - and on a standard install they live in `/usr/share/nmap/scripts`. Each script declares which category it belongs to, which ports or services it applies to, and when it should fire (during host discovery, after a port is found open, after version detection, and so on). You do not write Lua to use them. You point Nmap at a target, tell it which scripts or categories to run, and it does the rest. Two flags drive almost everything. `-sC` runs the *default* set of scripts (the ones tagged `default`, which the maintainers consider safe and broadly useful). `--script` lets you name individual scripts, whole categories, or wildcards. Everything is documented in the [official NSE documentation](https://nmap.org/nsedoc/?ref=pinglabz.com), which is the authoritative list of what each script does and what arguments it takes. ## The script categories Every NSE script belongs to one or more categories. Knowing the categories is how you reason about blast radius before you run anything - some are read-only and polite, others actively hammer the target. Here is the full set. auth Deals with authentication credentials, or bypassing it, without brute forcing. broadcast Discovers hosts by sending broadcast probes on the local segment. brute Guesses credentials against a service. Noisy and can lock accounts. default The set run by `-sC`. Safe, fast, and useful for most hosts. discovery Learns more about the network: DNS, SNMP, directory services, routes. dos Tests for denial of service. Can crash the target. Use only in a lab. exploit Actively exploits a vulnerability. Authorization required, always. external Sends data to a third-party service (whois, geoip). Leaks that you are scanning. fuzzer Sends malformed input to find parsing bugs. Slow and disruptive. intrusive Might crash a service, use significant resources, or be seen as malicious. malware Checks whether the target is already infected or running a backdoor. safe Designed not to crash anything or use much bandwidth. Read-only in spirit. version Extends version detection. Runs only under `-sV`, never on its own. vuln Checks for known vulnerabilities and reports only when one is found. ## How to run scripts The fastest way in is `-sC`, which runs the default category. It pairs naturally with version detection, and the two together (`-sV -sC`, or the shorthand baked into `-A`) are the closest thing NSE has to a "tell me everything reasonable" button. When you want something specific, reach for `--script`: - `--script=http-title` runs one named script. - `--script=vuln` runs an entire category. - `--script="http-*"` runs every script whose name starts with `http-` (mind the quotes so your shell does not expand the wildcard). - `--script-args` passes parameters, for example `--script-args http.useragent="Mozilla"`. You can combine categories and exclusions with boolean logic, such as `--script "default and safe"` or `--script "vuln and not dos"`. If you have just added scripts or updated Nmap, run `nmap --script-updatedb` once to rebuild the script database so the categories and wildcards resolve correctly. ## HTTP scripts: what your web tier gives away The HTTP script family is the one most network engineers reach for first, because web servers are everywhere and they are chatty. Here is the default trio of HTTP scripts run against an nginx box on the lab LAN. Version detection is not even on, yet the server hands over its exact build and the methods it will honour. ``` $ nmap -p80 --script http-title,http-headers,http-methods 10.10.10.21 Starting Nmap 7.95 ( https://nmap.org ) at 2026-06-28 07:03 UTC Nmap scan report for 10.10.10.21 Host is up (0.0023s latency). PORT STATE SERVICE 80/tcp open http |_http-title: Welcome to nginx! | http-methods: |_ Supported Methods: GET HEAD | http-headers: | Server: nginx/1.29.8 | Date: Sun, 28 Jun 2026 07:03:12 GMT | Content-Type: text/html | Content-Length: 896 | Last-Modified: Tue, 07 Apr 2026 11:37:12 GMT | Connection: close | ETag: "69d4ec68-380" | Accept-Ranges: bytes |_ (Request type: HEAD) MAC Address: 52:54:00:A2:2F:84 (QEMU virtual NIC) Nmap done: 1 IP address (1 host up) scanned in 0.51 seconds ``` Read that as an auditor. `Server: nginx/1.29.8` is a precise version string you probably do not want advertised on an internet-facing box. `Supported Methods: GET HEAD` is a clean result - if you saw `PUT` or `DELETE` here on a static site, that is a finding. The default page title, "Welcome to nginx!", tells you the server is running stock config with no real content deployed, which is worth knowing when you are inventorying what is actually in service versus what got stood up and forgotten. The `http-enum` script goes one step further and probes for well-known paths. Run against a Debian host serving the same stock nginx page, it confirms the methods and the (unversioned) server header: ``` $ nmap -sV --script http-enum,http-title,http-headers,http-methods -p80 127.0.0.1 Starting Nmap 7.95 ( https://nmap.org ) at 2026-07-05 15:05 PDT Nmap scan report for localhost (127.0.0.1) Host is up (0.00010s latency). PORT STATE SERVICE VERSION 80/tcp open http nginx | http-methods: |_ Supported Methods: GET HEAD |_http-title: Welcome to nginx! | http-headers: | Server: nginx | Date: Sun, 05 Jul 2026 22:05:15 GMT | Content-Type: text/html | Content-Length: 615 | Last-Modified: Sun, 05 Jul 2026 22:00:56 GMT | Connection: close | ETag: "6a4ad418-267" | Accept-Ranges: bytes |_ (Request type: HEAD) Service detection performed. Please report any incorrect results at https://nmap.org/submit/ . Nmap done: 1 IP address (1 host up) scanned in 10.65 seconds ``` Where `http-enum` really pays off is when a directory is left browsable. Point it at Nmap's own `scanme.nmap.org` (which the project maintains precisely so you have something legal to practise on) and it surfaces an exposed directory listing: ``` $ nmap --script http-enum,http-title,http-server-header -p80 scanme.nmap.org Starting Nmap 7.95 ( https://nmap.org ) at 2026-07-05 15:06 PDT Nmap scan report for scanme.nmap.org (45.33.32.156) Host is up (0.016s latency). PORT STATE SERVICE 80/tcp open http |_http-server-header: Apache/2.4.7 (Ubuntu) |_http-title: Go ahead and ScanMe! | http-enum: |_ /images/: Potentially interesting directory w/ listing on 'apache/2.4.7 (ubuntu)' ``` That single line - `/images/: Potentially interesting directory w/ listing` \- is exactly the kind of thing you want NSE to catch on your own estate before someone else does. And `http-server-header: Apache/2.4.7 (Ubuntu)` is a dated Apache build fingerprinted for you in one pass, which flows straight into your patch triage. NSE overlaps heavily with [Nmap version detection](https://www.pinglabz.com/nmap-version-detection/) here; the scripts add context (methods, headers, paths) that the version probes alone do not. ## SSH scripts: enumerating crypto and host keys SSH is the other service worth interrogating, because the algorithms a box negotiates are a compliance question. The `ssh2-enum-algos` script lists every key exchange, host key, cipher, and MAC algorithm the server offers. Against an OpenSSH host on Debian: ``` $ nmap --script ssh2-enum-algos,ssh-hostkey,banner -p21,22 127.0.0.1 PORT STATE SERVICE 21/tcp open ftp |_banner: 220 (vsFTPd 3.0.5) 22/tcp open ssh |_banner: SSH-2.0-OpenSSH_10.0p2 Debian-7+deb13u4 | ssh2-enum-algos: | kex_algorithms: (10) | mlkem768x25519-sha256 | sntrup761x25519-sha512 | curve25519-sha256 | ecdh-sha2-nistp256 | ... | server_host_key_algorithms: (4) | rsa-sha2-512 | rsa-sha2-256 | ecdsa-sha2-nistp256 | ssh-ed25519 | encryption_algorithms: (6) | chacha20-poly1305@openssh.com | aes128-gcm@openssh.com | aes256-gcm@openssh.com | aes128-ctr | aes192-ctr | aes256-ctr | mac_algorithms: (10) | hmac-sha2-256-etm@openssh.com | hmac-sha2-512-etm@openssh.com | hmac-sha1-etm@openssh.com | ... |_ compression_algorithms: (2) Nmap done: 1 IP address (1 host up) scanned in 0.34 seconds ``` This is a hardening checklist in list form. You can see this host offers modern post-quantum key exchange (`mlkem768x25519-sha256`) and AEAD ciphers, which is good, but it still advertises `hmac-sha1` in its MAC list - a legacy algorithm many hardening baselines tell you to remove. Run this before and after a config change and the diff is your evidence that the change took. (Trimmed above with `...` for length; the full lists are longer.) Note the two `banner` lines that came for free: `220 (vsFTPd 3.0.5)` on FTP and `SSH-2.0-OpenSSH_10.0p2 Debian-7+deb13u4` on SSH. That leads straight into banner grabbing. ## Banner grabbing on network gear The `banner` script connects to an open port, reads whatever the service volunteers on connect, and prints it. Paired with `-sV` it is a fast way to fingerprint devices that do not run a modern OS stack. Here it is against a Cisco router in the lab: ``` $ nmap -sV --script banner -p 22,23 10.10.10.1 PORT STATE SERVICE VERSION 22/tcp open ssh Cisco SSH 1.25 (protocol 2.0) |_banner: SSH-2.0-Cisco-1.25 23/tcp open telnet Cisco IOS telnetd | banner: \xFF\xFB\x01\xFF\xFB\x03\xFF\xFD\x18\xFF\xFD\x1F\x0D\x0A\x0D\x0 |_AUser Access Verification\x0D\x0A\x0D\x0AUsername: MAC Address: AA:BB:CC:00:5B:00 (Unknown) Service Info: OS: IOS; Device: switch; CPE: cpe:/o:cisco:ios Nmap done: 1 IP address (1 host up) scanned in 0.88 seconds ``` Two things jump out. The SSH banner `SSH-2.0-Cisco-1.25` pins the device as Cisco. And the telnet banner, after the IAC negotiation bytes (`\xFF\xFB...`), leaks the login prompt itself: `User Access Verification / Username:`. That is a plaintext telnet service on a router advertising exactly how to start authenticating to it. Finding that on your own network is the point - it is a work item to disable telnet and move to SSH only. NSE does the finding; you do the remediation. Not every device cooperates, and that is worth calling out honestly. The `ssh-hostkey` script run against the same router returned no key at all, because that Cisco SSH implementation did not expose a host key Nmap could parse. NSE reports what the service gives it; when the service gives nothing, you get an empty result rather than an invented one. ## The vuln category and the authorization line The `vuln` category is where NSE stops being purely observational. These scripts actively test for known CVEs and misconfigurations, and by design they only print output when they find something. That makes them powerful and also means they touch the target in ways a plain port scan does not - sending crafted requests, following redirects, occasionally triggering the very code path a vulnerability lives in. The `intrusive`, `exploit`, `dos`, and `fuzzer` categories are more aggressive still and can crash a service outright. The rule is simple and it is not optional: **run intrusive and vuln scripts only against systems you own or have explicit written authorization to test.** Scanning someone else's host with `--script vuln` is not reconnaissance, it is probing for weaknesses, and in most jurisdictions that crosses a legal line. The reason `scanme.nmap.org` exists is so you have a sanctioned target for practice. On production, keep the aggressive categories confined to a lab or a maintenance window on gear you are responsible for. When the gear is yours, there is a second lesson available for free. Run the scripts, then go and read [what the box wrote to its own logs while they ran](https://www.pinglabz.com/detecting-nmap-scans-cisco-blue-team/). It is the quickest way to learn which categories slip by unnoticed and which announce themselves on every line of the buffer. ## Keeping the script database current NSE ships with Nmap, so upgrading Nmap upgrades your scripts. If you add a custom script to `/usr/share/nmap/scripts` or pull new ones, rebuild the index so category and wildcard selection keeps working: ``` $ sudo nmap --script-updatedb ``` When you want to know exactly what a script does before you fire it, read its entry in the [NSE documentation at nmap.org/nsedoc](https://nmap.org/nsedoc/?ref=pinglabz.com), which lists every script, its categories, its arguments, and sample output. Reading the doc first is how you avoid running something intrusive when you meant something safe. ## Key takeaways NSE is what turns Nmap into an audit tool for your own network. The default scripts under `-sC` and targeted picks under `--script` pull server versions, HTTP methods, exposed directories, SSH algorithm lists, and login banners straight out of the services on your estate, and every one of those is either a clean result or a work item you can act on. Stay on the safe and default categories for production sweeps, reserve the intrusive and vuln scripts for systems you own or are authorized to test, rebuild the database with `--script-updatedb` after any change, and lean on the NSE docs to know a script's blast radius before you run it. NSE also pairs naturally with the rest of the toolkit - use it alongside [version detection](https://www.pinglabz.com/nmap-version-detection/) to fingerprint services, and when you need to test whether your controls actually stop a scan, move on to [firewall and IDS evasion techniques](https://www.pinglabz.com/nmap-firewall-evasion/). For the full picture of where scripting fits in a scanning workflow, keep [the complete Nmap guide](https://www.pinglabz.com/nmap/) close by. ### Nmap UDP Scanning (-sU): Why It Is Slow and How to Read It URL: https://www.pinglabz.com/nmap-udp-scan/ Last updated: 2026-08-01T19:31:01.000Z Most people scan TCP, get their results in seconds, and stop there. But a huge amount of what runs on a real network lives on UDP: DNS, DHCP, SNMP, NTP, TFTP, IKE. If you only ever scan TCP, you are blind to half your own infrastructure. UDP scanning (`-sU`) fills that gap, but it is slower, noisier to interpret, and far more prone to ambiguous results than its TCP cousin, and understanding *why* is the difference between trusting the output and being misled by it. This guide is part of [the complete Nmap guide](https://www.pinglabz.com/nmap/), and it covers how UDP states are actually decided, how to read the ambiguous ones, which UDP services are worth your time, and how to keep a UDP scan from taking all night. ## Why UDP scanning is hard and slow TCP is a connection-oriented protocol with a handshake. Send a SYN, get a SYN/ACK, and you know the port is open; get a RST, and you know it is closed. The protocol itself gives you a clean signal. UDP has none of that. It is connectionless: you send a datagram and the protocol makes no promise of any reply at all. There is no handshake to interpret, so Nmap has to infer state from indirect evidence. Worse, the one clear signal UDP scanning relies on - the ICMP "port unreachable" message a host sends when nothing is listening - is rate-limited by most operating systems. A target will only emit so many of those ICMP errors per second, so Nmap has to slow down and wait to avoid outrunning the target's willingness to answer. That rate limiting is the single biggest reason a full UDP scan can crawl. TCP can blast a thousand ports and read a thousand RSTs; UDP has to send, wait, and often send again because it cannot tell a dropped probe from a silent-but-open port. ## How UDP states are decided Nmap resolves a UDP port into one of three outcomes, and each comes from a different piece of evidence. The clearest case is **closed**: the host answers your probe with an ICMP type 3, code 3 (port unreachable) message, which is an explicit "nothing is listening here." You can see that reasoning spelled out when you add `--reason` to a scan of the common UDP ports on a Linux host: ``` $ nmap -sU --top-ports 20 --reason 127.0.0.1 Starting Nmap 7.95 ( https://nmap.org ) at 2026-07-05 15:05 PDT Nmap scan report for localhost (127.0.0.1) Host is up, received localhost-response (0.000042s latency). PORT STATE SERVICE REASON 53/udp closed domain port-unreach ttl 64 67/udp closed dhcps port-unreach ttl 64 68/udp closed dhcpc port-unreach ttl 64 69/udp closed tftp port-unreach ttl 64 123/udp closed ntp port-unreach ttl 64 135/udp closed msrpc port-unreach ttl 64 137/udp closed netbios-ns port-unreach ttl 64 138/udp closed netbios-dgm port-unreach ttl 64 139/udp closed netbios-ssn port-unreach ttl 64 161/udp closed snmp port-unreach ttl 64 162/udp closed snmptrap port-unreach ttl 64 445/udp closed microsoft-ds port-unreach ttl 64 500/udp closed isakmp port-unreach ttl 64 514/udp closed syslog port-unreach ttl 64 520/udp closed route port-unreach ttl 64 631/udp closed ipp port-unreach ttl 64 1434/udp closed ms-sql-m port-unreach ttl 64 1900/udp closed upnp port-unreach ttl 64 4500/udp closed nat-t-ike port-unreach ttl 64 49152/udp closed unknown port-unreach ttl 64 ``` Every line reads `closed ... port-unreach ttl 64`. That is the ideal case: the host was cooperative, sent an ICMP unreachable for each port with no service behind it, and Nmap could state closed with certainty. The `--reason` column makes the evidence explicit, which is worth getting in the habit of on UDP scans specifically, because the states are so much easier to misread than TCP. The second outcome is the frustrating one: **open|filtered**. If Nmap sends a UDP probe and hears absolutely nothing back - no service reply and no ICMP error - it genuinely cannot tell whether a service is quietly listening (open) or a firewall silently dropped the probe (filtered). Both look identical from the scanner's side: silence. So Nmap reports the honest ambiguity `open|filtered` rather than guessing. The third outcome is **open**, which you only get when an actual UDP service replies to the probe - a DNS server answering a query, an SNMP agent responding to a get. A real reply is unambiguous proof something is listening. ## Seeing it on network gear The same logic plays out against a Cisco router in the lab. Here is a UDP scan of five common infrastructure ports on R1: ``` $ nmap -sU -p 53,67,123,161,500 10.10.10.1 Starting Nmap 7.95 ( https://nmap.org ) at 2026-06-28 07:07 UTC Nmap scan report for 10.10.10.1 Host is up (0.0036s latency). PORT STATE SERVICE 53/udp closed domain 67/udp closed dhcps 123/udp closed ntp 161/udp closed snmp 500/udp closed isakmp MAC Address: AA:BB:CC:00:5B:00 (Unknown) Nmap done: 1 IP address (1 host up) scanned in 4.97 seconds ``` All five come back `closed` because R1 returned an ICMP port-unreachable for each - definitively closed, the same clean signal as the Linux host. Note the timing though: 4.97 seconds for just five UDP ports, versus the sub-second TCP scans you are used to. That is the rate limiting and probe retransmission tax in action, and it is why the capture comment simply reads "UDP is slow." Scale that up to thousands of ports and the cost compounds fast. One more thing worth flagging from the lab. When the DNS host (DNS01) was scanned, Nmap reported `Host seems down. If it is really up, but blocking our ping probes, try -Pn` \- because that container never got its LAN IP applied, so host discovery correctly found nothing there. That is a reminder that a UDP scan still depends on the target being reachable in the first place; a "host down" note is about discovery, not about the UDP ports. Worth remembering that the router is keeping score while it does all this answering. Every ICMP port-unreachable in that output is a packet it built on your request, so both halves of the exchange land on the same access list, and [a UDP scan is one of the easiest scans to pick out from the target side](https://www.pinglabz.com/detecting-nmap-scans-cisco-blue-team/) because of it. ## UDP services worth scanning You rarely want a full UDP sweep. Far more often you want the handful of ports that carry the services network engineers actually care about. These are the ones worth targeting first. DNS Port53/udp Name resolution. A live resolver answers probes, so it often reads as genuinely open. DHCP Ports67/68 udp Server on 67, client on 68\. Rogue DHCP hunting starts here. TFTP Port69/udp Config and image transfers for network gear. Often left open and unauthenticated. NTP Port123/udp Time sync. Worth confirming your clock source is where you think it is. SNMP Port161/udp Device management. A big one to audit - open SNMP with a default community string leaks everything. IKE Port500/udp IPsec VPN negotiation. Confirms where your tunnel endpoints live. UPnP Port1900/udp Discovery on consumer and IoT gear. Frequently exposed where it should not be. Targeting these directly with `-p` instead of scanning all 65535 UDP ports is the single biggest time-saver available. On network gear especially, SNMP on 161 and TFTP on 69 are the two worth auditing on every device, because an open, default-community SNMP agent or an unauthenticated TFTP server is a real exposure hiding in plain sight. ## Speeding up a UDP scan Because UDP is slow by nature, controlling scope and pace is essential. A few flags carry most of the weight. - `--top-ports N` scans only the N most common UDP ports instead of the full range. The `--top-ports 20` in the capture above is a perfect example: twenty high-value ports in a fraction of a second, versus minutes or hours for everything. - `-T4` raises the timing template, letting Nmap probe more aggressively. It helps, but remember the target's ICMP rate limiting sets a ceiling that no timing template can fully overcome. - `--min-rate` forces a minimum packets-per-second floor, useful when you have measured the target can take it and you want to push past Nmap's cautious defaults. - `--host-timeout` caps how long Nmap will spend on any single host before moving on, which stops one slow or heavily filtered target from stalling a whole sweep. All of these lean on Nmap's broader [timing and performance](https://www.pinglabz.com/nmap-timing-performance/) controls, which matter far more on UDP than on TCP precisely because there is so much waiting involved. ## Pairing -sU with -sV The `open|filtered` ambiguity is the weakness of a bare UDP scan, and version detection is one way to cut through it. Adding `-sV` to a UDP scan makes Nmap send protocol-specific payloads: a real DNS query to port 53, an SNMP get to 161, and so on. If the service is actually listening, it replies to the payload, and that reply promotes the port from `open|filtered` to a confirmed `open` with a version string attached. The trade-off is more time and more packets, so reserve it for the ports you have already narrowed down to. This is the same [port scanning](https://www.pinglabz.com/nmap-port-scanning/) discipline you apply on TCP: scan wide to find candidates, then probe deep on the interesting ones. For deeper interrogation of a confirmed UDP service, the [NSE scripts](https://www.pinglabz.com/nmap-nse-scripts/) include targeted scripts for DNS, SNMP, and more. All of this is why a full `nmap -sU -p-` across every one of the 65535 UDP ports can take hours against a single host. Every silent port forces a wait to distinguish open from filtered, every closed port is gated by the target's ICMP rate limit, and probes get retransmitted to rule out simple packet loss. It is a legitimate scan to run when you truly need complete UDP coverage, but it is never the scan you reach for first. Nmap's own notes on [port scanning techniques](https://nmap.org/book/man-port-scanning-techniques.html?ref=pinglabz.com) go into the retransmission and rate-limit mechanics in detail. ## Key takeaways UDP scanning is slow and ambiguous by design because the protocol gives no handshake, and the one clean signal - an ICMP port-unreachable meaning closed - is rate-limited by the target. Learn to read the three states: `closed` from an ICMP unreachable, `open` from an actual service reply, and the honest `open|filtered` when the host stays completely silent. Use `--reason` to see the evidence, target the high-value ports (DNS, DHCP, TFTP, NTP, SNMP, IKE, UPnP) with `-p` or `--top-ports` rather than sweeping all 65535, and pace the scan with `-T4`, `--min-rate`, and `--host-timeout` so it finishes this decade. When ambiguity bites, add `-sV` so payload probes can turn `open|filtered` into a confirmed service. For the whole methodology, from discovery through TCP and UDP scanning to scripting, work through [the complete Nmap guide](https://www.pinglabz.com/nmap/). ### Nmap OS Detection (-O): How TCP/IP Fingerprinting Works URL: https://www.pinglabz.com/nmap-os-detection/ Last updated: 2026-08-01T19:31:02.000Z Version detection tells you what software a port is running. OS detection (`-O`) tries to answer a different question: what operating system is the whole host running, even on ports that give nothing away. It does this without logging in and without trusting any banner, by reading the subtle quirks in how the target's TCP/IP stack builds packets. That makes it powerful for inventory and auditing - you can fingerprint a box you have no credentials for - but it also makes it a guess, and a guess that fails in predictable ways on lab gear and network appliances. This guide is part of [the complete Nmap guide](https://www.pinglabz.com/nmap/), and it walks through what `-O` measures, how to read every field it prints, and why it sometimes shrugs and hands you a raw fingerprint instead of an answer. ## What `-O` actually measures No two TCP/IP stacks are implemented identically. The RFCs leave dozens of small decisions to the developer - how to pick initial sequence numbers, which TCP options to include and in what order, what window size to advertise, how to set IP flags, how ICMP replies are formed. Those decisions are consistent within an OS and differ between operating systems. OS detection sends a battery of specially crafted probes, watches exactly how the stack responds, and matches that behavioral fingerprint against Nmap's database of known systems. Among the things it samples are the sequence-number generation pattern (the `SEQ` tests, which measure how predictable the target's ISN sampling is), the set and ordering of TCP options, the advertised window size, IP fragmentation flags, and ICMP behavior. You do not need to read those raw values by hand - Nmap distills them into a plain-language guess - but knowing that is where the answer comes from explains why the technique is both clever and fallible. It is inference from stack behavior, not a lookup of anything the host willingly tells you. That battery of probes is also the loudest thing in the toolkit. The packets are odd enough to stand out in a capture on sight, and because a default `-O` run scans a thousand ports on the way to its fingerprint, it produced [the biggest counter jump of any scan in the lab](https://www.pinglabz.com/detecting-nmap-scans-cisco-blue-team/) on the router being fingerprinted. ## Why it needs an open and a closed port OS fingerprinting works best when Nmap can observe the stack in two states: how it responds on an open port and how it responds on a closed one. The contrast between those two behaviors is a large part of the signal. If Nmap cannot find both, it warns you, and the accuracy drops. You will see exactly that warning when scanning a host where every probed port happens to be open: ``` $ nmap -O -p22,80,9929,31337 scanme.nmap.org Starting Nmap 7.95 ( https://nmap.org ) at 2026-07-05 15:06 PDT Nmap scan report for scanme.nmap.org (45.33.32.156) Host is up (0.018s latency). Other addresses for scanme.nmap.org (not scanned): 2600:3c01::f03c:91ff:fe18:bb2f PORT STATE SERVICE 22/tcp open ssh 80/tcp open http 9929/tcp open nping-echo 31337/tcp open Elite Warning: OSScan results may be unreliable because we could not find at least 1 open and 1 closed port Device type: general purpose|router Running: Linux 4.X|5.X, MikroTik RouterOS 7.X OS CPE: cpe:/o:linux:linux_kernel:4 cpe:/o:linux:linux_kernel:5 cpe:/o:mikrotik:routeros:7 cpe:/o:linux:linux_kernel:5.6.3 OS details: Linux 4.15 - 5.19, Linux 5.0 - 5.14, OpenWrt 21.02 (Linux 5.4), MikroTik RouterOS 7.2 - 7.5 (Linux 5.6.3) Network Distance: 15 hops OS detection performed. Please report any incorrect results at https://nmap.org/submit/ . Nmap done: 1 IP address (1 host up) scanned in 1.97 seconds ``` Because the scan restricted itself to four ports that all came back open, Nmap could not sample closed-port behavior and printed `OSScan results may be unreliable because we could not find at least 1 open and 1 closed port`. The practical fix is simple: do not artificially limit the port list when you care about OS detection. Let Nmap scan a normal range so it naturally finds a closed port to compare against. If a host genuinely has no closed ports in reach - a tightly filtered target - accept that the guess will be softer and lean on other evidence. ## Reading the OS detection fields Notice how many guesses that scanme.nmap.org result offered. That ambiguity is the norm when the signal is weak, and it is worth unpacking each field so you know what you are looking at. Device type The category Nmap thinks the host is: general purpose, router, switch, firewall, printer, and so on. Pipe-separated when unsure, e.g. `general purpose|router`. Running The OS family and rough version range. Multiple families listed means multiple candidates fit the fingerprint. OS CPE Standardized machine-readable OS identifiers. Feed these into asset databases or vulnerability tooling. OS details The most specific version guesses Nmap will commit to. A range like `Linux 4.15 - 5.19` is a confidence window, not a precise build. Network Distance How many router hops away the target is. `15 hops` to an internet host; `1 hop` on your own LAN. For scanme.nmap.org, the fingerprint fit several plausible systems at once - Linux 4.15-5.19, Linux 5.0-5.14, OpenWrt 21.02, and MikroTik RouterOS 7.2-7.5 - which is why `Device type` hedged as `general purpose|router`. That kind of spread is common when scanning across the internet: 15 hops of intervening routers, rate limiting, and middleboxes all erode the fingerprint. The `Network Distance: 15 hops` is itself a useful clue about how far away and how filtered the path is. Contrast that with a host one hop away on your own segment, where the same technique is far more decisive. ## When it nails it, and when it can't On the local lab segment, one hop away, OS detection is much sharper - but only for hosts whose stack is in the database. Here is a run against a Cisco IOL router and a Linux host on the same LAN: ``` $ nmap -O -T4 10.10.10.1 10.10.10.24 Starting Nmap 7.95 ( https://nmap.org ) at 2026-06-28 05:54 UTC Nmap scan report for 10.10.10.1 Host is up (0.0029s latency). Not shown: 998 closed tcp ports (reset) PORT STATE SERVICE 22/tcp open ssh 23/tcp open telnet MAC Address: AA:BB:CC:00:5B:00 (Unknown) No exact OS matches for host (If you know what OS is running on it, see https://nmap.org/submit/ ). TCP/IP fingerprint: OS:SCAN(V=7.95%E=4%D=6/28%OT=22%CT=1%CU=42871%PV=Y%DS=1%DC=D%G=Y%M=AABBCC%T OS:M=6A40B71C%P=x86_64-pc-linux-gnu)SEQ(SP=100%GCD=1%ISR=106%TI=RD%CI=RD%II OS:=RI%TS=U)OPS(O1=M5B4SNNW2L%O2=M578SNNW2L%O3=M280SNNW2L%O4=M218SNNW2L%O5=M OS:218SNNW2L%O6=M109SLL)WIN(W1=FFFF%W2=FFFF%W3=FFFF%W4=FFFF%W5=FFFF%W6=FFFF) OS:ECN(R=Y%DF=N%T=100%W=FFFF%O=M5B4SNNW2L%CC=N%Q=)T1(R=Y%DF=N%T=100%S=O%A=S+ OS:%F=AS%RD=0%Q=)T2(R=N)T3(R=N)T4(R=Y%DF=N%T=100%W=0%S=A%A=Z%F=R%O=%RD=0%Q=) OS:IE(R=Y%DFI=S%T=100%CD=S) Network Distance: 1 hop Nmap scan report for 10.10.10.24 Host is up (0.0036s latency). Not shown: 999 closed tcp ports (reset) PORT STATE SERVICE 22/tcp open ssh MAC Address: 52:54:00:BB:DC:36 (QEMU virtual NIC) Device type: general purpose Running: Linux 4.X|5.X OS CPE: cpe:/o:linux:linux_kernel:4 cpe:/o:linux:linux_kernel:5 OS details: Linux 4.15 - 5.19 Network Distance: 1 hop OS detection performed. Please report any incorrect results at https://nmap.org/submit/ . Nmap done: 2 IP addresses (2 hosts up) scanned in 13.91 seconds ``` The Linux host at .24 gets a clean read: `Device type: general purpose`, `Running: Linux 4.X|5.X`, `OS details: Linux 4.15 - 5.19`. Its stack is a mainstream Linux kernel, well represented in Nmap's database, so one hop away with an open and a closed port available, the match is confident. The Cisco IOL router at .1 is a different story. Instead of a named result you get `No exact OS matches for host` followed by a block of raw `TCP/IP fingerprint:` data. That is not a failure of the technique - Nmap successfully measured the stack, you can see the `SEQ`, `OPS`, `WIN`, and `ECN` tests captured right there - it simply had no signature in the database that matched. This is expected on virtual and lab gear and on some network operating systems: IOL is a virtualized IOS image whose stack behavior does not perfectly match physical Cisco hardware, and it is not indexed the way a common Linux kernel is. The right response is to submit that fingerprint back to Nmap (the output tells you where) so the database improves, and in the meantime lean on other evidence. ## Accuracy controls When OS detection hedges or gives up, a few flags change how hard it tries and how much it will speculate. - `--osscan-guess` (and its equivalent `--fuzzy`) tells Nmap to make its best guess even when the match is not perfect, printing near-matches with their confidence percentages instead of falling back to a raw fingerprint. Useful when "probably Linux 5.x" is more helpful to you than "no exact match." - `--osscan-limit` does the opposite of guessing harder: it restricts OS detection to hosts that have at least one open and one closed port. On a large sweep this skips the hopeless cases and saves real time, since fingerprinting a host with no usable port contrast rarely pays off. - `--max-os-tries` sets how many times Nmap retries the fingerprinting probes before giving up. Lowering it speeds up big scans; raising it can squeeze a match out of a flaky or rate-limited target at the cost of time. None of these turn a guess into certainty. They tune the balance between speed, completeness, and how willing you are to accept a probabilistic answer. ## Pair -O with -sV OS detection is strongest when it is not the only evidence you have. Look back at the Cisco router: `-O` could not name it, but a [version detection](https://www.pinglabz.com/nmap-version-detection/) scan of the same box reported `Service Info: OS: IOS; Device: switch` straight from the SSH and telnet banners. Two independent techniques, two angles on the same question. When the stack fingerprint comes back empty, the application banners often fill the gap, and when the banners are stripped, the stack fingerprint may still speak. Running `-O` and `-sV` together (or just using `-A`, which includes both) is the practical default for any real fingerprinting job. It also helps to remember what OS detection assumes: that the host is up and reachable in the first place. If a target is not responding to your probes, sort out [host discovery](https://www.pinglabz.com/nmap-host-discovery/) before you trust an OS result, because a filtered or half-reachable host produces exactly the kind of thin, ambiguous fingerprint that leads to a spread of bad guesses. Nmap's [reference guide](https://nmap.org/book/man.html?ref=pinglabz.com) documents the full set of fingerprint tests if you want to read the raw `SEQ` and `OPS` lines yourself. ## Key takeaways OS detection reads the target's TCP/IP stack - sequence numbers, TCP options, window size, IP flags, ICMP behavior - and matches that fingerprint against a database, so it is inference rather than fact. It needs at least one open and one closed port to do its best work, which is why over-restricting the port list triggers the "unreliable" warning. On a mainstream Linux host one hop away it will name the kernel range confidently; on virtual lab gear or an off-database network OS like Cisco IOL it may return only a raw fingerprint for you to submit. Read the `Device type`, `Running`, `OS CPE`, `OS details`, and `Network Distance` fields together rather than fixating on one, tune the guess with `--osscan-guess` and `--osscan-limit`, and always pair `-O` with `-sV` so a stack-fingerprint miss can be backfilled by a banner that already said "OS: IOS." For the full workflow around discovery, scanning, and fingerprinting, work through [the complete Nmap guide](https://www.pinglabz.com/nmap/). ### Nmap Version Detection (-sV): Banner Grabbing and Service Fingerprinting URL: https://www.pinglabz.com/nmap-version-detection/ Last updated: 2026-08-01T19:31:02.000Z A port number tells you almost nothing on its own. "Port 80 open" could be nginx, Apache, an embedded management interface, or a printer's web UI. "Port 22 open" could be OpenSSH, a Cisco IOS SSH stack, or a locked-down appliance. Nmap's service and version detection (`-sV`) is the flag that turns an open port into an actionable fact: *nginx 1.29.8*, *Cisco SSH 1.25*, *OpenSSH 10.0p2 Debian*. That single upgrade in fidelity is what separates a port list from an inventory, and it feeds nearly everything else you do with the tool. This guide is part of [the complete Nmap guide](https://www.pinglabz.com/nmap/), and it focuses on what `-sV` actually sends, how to read what comes back, and how to tune it when it guesses wrong. ## What `-sV` actually does A plain TCP scan learns one thing about a port: is it accepting connections. Service detection goes a step further. Once Nmap knows a port is open, it opens a real connection and runs the port through its probe database (`nmap-service-probes`). Some services announce themselves the moment you connect - an FTP daemon greets you with a banner, an SSH server sends its protocol string - so Nmap reads that first. For services that stay quiet, Nmap sends a sequence of carefully crafted probe strings and matches the reply against thousands of signature patterns until something fits. The result is the `VERSION` column you see next to each open port. It is not magic and it is not always exact, but on well-known services it is remarkably good. Here is a live run from the PingLabz Nmap lab, scanning a Cisco router, an nginx web server, and a Linux host in one pass: ``` $ nmap -sV -T4 10.10.10.1 10.10.10.21 10.10.10.24 Starting Nmap 7.95 ( https://nmap.org ) at 2026-06-28 05:52 UTC Nmap scan report for 10.10.10.1 Host is up (0.013s latency). Not shown: 998 closed tcp ports (reset) PORT STATE SERVICE VERSION 22/tcp open ssh Cisco SSH 1.25 (protocol 2.0) 23/tcp open telnet Cisco IOS telnetd MAC Address: AA:BB:CC:00:5B:00 (Unknown) Service Info: OS: IOS; Device: switch; CPE: cpe:/o:cisco:ios Nmap scan report for 10.10.10.21 Host is up (0.0051s latency). Not shown: 999 closed tcp ports (reset) PORT STATE SERVICE VERSION 80/tcp open http nginx 1.29.8 MAC Address: 52:54:00:A2:2F:84 (QEMU virtual NIC) Nmap scan report for 10.10.10.24 Host is up (0.0068s latency). Not shown: 999 closed tcp ports (reset) PORT STATE SERVICE VERSION 22/tcp open ssh OpenSSH 10.2 (protocol 2.0) MAC Address: 52:54:00:BB:DC:36 (QEMU virtual NIC) Service detection performed. Please report any incorrect results at https://nmap.org/submit/ . Nmap done: 3 IP addresses (3 hosts up) scanned in 10.51 seconds ``` Look at what those three lines of output give you. The router at .1 is running **Cisco SSH 1.25** and **Cisco IOS telnetd**, and Nmap has inferred from those responses that the OS is IOS and the device is a switch. The web server at .21 is running **nginx 1.29.8** \- not just "a web server," the exact build. The host at .24 is running **OpenSSH 10.2**. None of that was in the port list before you added `-sV`. The cost of all that detail is noise. Because `-sV` opens a genuine connection rather than half-opening one, the probe gets written into the service daemon's own log, on top of everything [the router in front of the host already recorded](https://www.pinglabz.com/detecting-nmap-scans-cisco-blue-team/). Version detection is the least stealthy thing in this article. ## Turning ports into services The same flag behaves identically on a real Linux box. This capture is from a Debian VM running the usual trio of daemons: ``` $ nmap -sV -p21,22,80 127.0.0.1 Starting Nmap 7.95 ( https://nmap.org ) at 2026-07-05 15:05 PDT Nmap scan report for localhost (127.0.0.1) Host is up (0.00010s latency). PORT STATE SERVICE VERSION 21/tcp open ftp vsftpd 3.0.5 22/tcp open ssh OpenSSH 10.0p2 Debian 7+deb13u4 (protocol 2.0) 80/tcp open http nginx Service Info: OSs: Unix, Linux; CPE: cpe:/o:linux:linux_kernel ``` Notice the range of detail here. On FTP, Nmap pulls the full product and version: **vsftpd 3.0.5**. On SSH it goes further and reads the distribution packaging string: **OpenSSH 10.0p2 Debian 7+deb13u4**, which tells you not only the OpenSSH release but that it is the Debian-packaged build with a specific patch revision. On HTTP, though, it reports plain `nginx` with no version. That is not a bug: this nginx was configured to suppress its version token, so Nmap correctly reports the service without inventing a number. Compare that with the lab's WEB01, which happily advertised **nginx 1.29.8**. Whether you get a version often comes down to how the admin (maybe you) configured the banner. ## Why the exact version matters "nginx" is trivia. "nginx 1.29.8" is a lead. Once you have an exact product and version, you can cross-reference it against published advisories and your own patch baseline. An auditor uses that to answer "is anything on this segment running a build with a known CVE?" A network engineer uses the same output to answer a quieter but just as useful question: "did the config management actually roll out the version I think it did?" If your standard is nginx 1.29.x across the fleet and a host reports 1.24, you have found configuration drift with a single scan. This is the network-engineer case for `-sV` that gets lost in the security framing. You are not always hunting an intruder. Often you are verifying your own environment: confirming a firmware upgrade landed, confirming a decommissioned service is actually gone, confirming that the box you think is a jump host is running the SSH build your policy requires. The `Service Info` line - `OS: IOS; Device: switch` on the router, `OSs: Unix, Linux` on the Debian host - is a bonus inference Nmap draws from the same responses, and it is the bridge into full [OS detection with -O](https://www.pinglabz.com/nmap-os-detection/). ## Controlling version intensity Service detection is a trade-off between how many probes Nmap sends and how long the scan takes. The `--version-intensity` knob (0 to 9) sets how many probes from the database Nmap is willing to try. Higher intensity means more probes, better odds of identifying an obscure or deliberately quiet service, and a slower scan. Lower intensity keeps things fast at the cost of coverage. There are two named shortcuts for the extremes. \--version-intensity 0-9 Sets how many probes Nmap tries against each open port. Default is 7\. Higher means more probes, more accuracy, more time. \--version-light Alias for intensity 2. Fast, only the most common probes. Good for wide sweeps where speed wins. \--version-all Alias for intensity 9. Every probe, no shortcuts. Use when a service refuses to identify at lower intensity. In practice, the default of 7 identifies the vast majority of services on a normal network. Reach for `--version-light` when you are scanning a large range and just want a fast pass. Reach for `--version-all` when a specific port stubbornly comes back as `tcpwrapped` or unidentified and you want to throw everything at it. ## Banner grabbing with --script banner Version detection's polite cousin is banner grabbing. Many services volunteer a text banner the instant you connect, and the `banner` NSE script simply captures and prints it verbatim. Pairing it with `-sV` gives you both the parsed version and the raw greeting, which is often more revealing than the tidy summary. Here is the lab's Cisco router: ``` $ nmap -sV --script banner -p 22,23 10.10.10.1 Starting Nmap 7.95 ( https://nmap.org ) at 2026-06-28 07:03 UTC Nmap scan report for 10.10.10.1 Host is up (0.0035s latency). PORT STATE SERVICE VERSION 22/tcp open ssh Cisco SSH 1.25 (protocol 2.0) |_banner: SSH-2.0-Cisco-1.25 23/tcp open telnet Cisco IOS telnetd | banner: \xFF\xFB\x01\xFF\xFB\x03\xFF\xFD\x18\xFF\xFD\x1F\x0D\x0A\x0D\x0 |_AUser Access Verification\x0D\x0A\x0D\x0AUsername: MAC Address: AA:BB:CC:00:5B:00 (Unknown) Service Info: OS: IOS; Device: switch; CPE: cpe:/o:cisco:ios Nmap done: 1 IP address (1 host up) scanned in 0.88 seconds ``` On SSH the banner is the clean protocol string `SSH-2.0-Cisco-1.25`, which is exactly where the "Cisco SSH 1.25" version came from. On telnet the banner is far chattier. Those `\xFF\xFB` and `\xFF\xFD` sequences are the telnet IAC (interpret-as-command) negotiation bytes the protocol exchanges before anything human-readable. Read past them and the router leaks its own login prompt: **User Access Verification / Username:**. That is the standard Cisco IOS login banner, captured straight off the wire, confirming both the platform and that telnet is live and unauthenticated at the greeting. The same technique on the Debian host pulls equally clean strings, plus SSH algorithm enumeration when you add the right scripts: ``` $ nmap --script ssh2-enum-algos,ssh-hostkey,banner -p21,22 127.0.0.1 ... PORT STATE SERVICE 21/tcp open ftp |_banner: 220 (vsFTPd 3.0.5) 22/tcp open ssh |_banner: SSH-2.0-OpenSSH_10.0p2 Debian-7+deb13u4 ``` The FTP banner `220 (vsFTPd 3.0.5)` is the daemon's own greeting - the same string that let `-sV` report vsftpd 3.0.5\. The SSH banner `SSH-2.0-OpenSSH_10.0p2 Debian-7+deb13u4` carries the full Debian packaging detail. If you want to go deeper on SSH specifically - key exchange algorithms, host keys, cipher suites - the [NSE scripting engine](https://www.pinglabz.com/nmap-nse-scripts/) has purpose-built scripts like `ssh2-enum-algos`, which is how you audit whether a box still negotiates weak algorithms. ## CPE strings and when detection is wrong You will have noticed the `CPE:` tokens in the output - `cpe:/o:cisco:ios` on the router, `cpe:/o:linux:linux_kernel` on the Linux hosts. CPE (Common Platform Enumeration) is a standardized naming scheme for operating systems, hardware, and applications. Those strings are machine-readable identifiers you can feed straight into vulnerability tooling or asset databases, which is why they matter more in an audit pipeline than to a human reading the terminal. Every service-detection run ends with the same line: `Service detection performed. Please report any incorrect results at https://nmap.org/submit/`. That is not boilerplate to ignore. It is Nmap reminding you that version detection is a best-effort match against a signature database, and databases have gaps. A deliberately altered banner, a rare build, or a custom appliance can all produce a wrong or missing version. When the result looks off, the honest move is to grab the raw banner (as above) and read it yourself rather than trusting the parsed summary, and if you have confirmed a genuine miss, submitting it back improves the database for everyone. ## How -sV feeds -O and -A Service detection rarely runs alone. It is the middle piece of a larger fingerprinting picture. The `Service Info: OS: IOS` line you saw on the router is already an OS hint drawn purely from application banners, and it dovetails with TCP/IP stack fingerprinting, which is a different technique entirely - see the full walkthrough of [OS detection](https://www.pinglabz.com/nmap-os-detection/) for how `-O` works and why the two together are stronger than either alone. When you want all of it in one command, `-A` bundles version detection, OS detection, default NSE scripts, and traceroute. It is the "tell me everything" switch, and `-sV` is one of its core ingredients. A sensible progression is to start with a fast port scan to find what is open, add `-sV` to learn what those ports are actually running, then layer on `-O` and selected scripts once you know where to dig. Nmap's own [reference guide](https://nmap.org/book/man.html?ref=pinglabz.com) documents every intensity level and probe-database detail if you want to go under the hood. ## Key takeaways Service and version detection is the flag that makes an Nmap scan worth reading. It turns "22 open" into "Cisco SSH 1.25" and "80 open" into "nginx 1.29.8," and that exactness is what feeds CVE lookups, drift checks, and inventory alike. Tune `--version-intensity` when you need speed or coverage, add `--script banner` to read the raw greeting when the parsed version looks wrong, and remember that a missing version usually means a hardened banner rather than a failed scan. Trust the output, but verify the surprising results against the raw banner, and let `-sV` hand off cleanly into OS detection and full `-A` profiling. For the whole picture - discovery, scanning, fingerprinting, and scripting in one place - work through [the complete Nmap guide](https://www.pinglabz.com/nmap/). ### Nmap Scan Types Explained: ACK, FIN, NULL, Xmas, Window, Maimon URL: https://www.pinglabz.com/nmap-scan-types/ Last updated: 2026-08-01T19:31:00.000Z Nmap does not have one way to scan a TCP port, it has a whole family of them, and each one sets different flags to provoke a different reaction from the target. Knowing which scan type to reach for is the difference between mapping a firewall's rules cleanly and getting a page of misleading results. This article is part of [the complete Nmap guide](https://www.pinglabz.com/nmap/), and it walks the full family of TCP scan types, shows what each one detects, and (more usefully) shows where each one lies to you when the target stack does not play by the rules. ## The TCP scan family at a glance Every TCP scan type comes down to which flags Nmap sets in the probe packet and what it infers from the response. Here is the family, one line each on the flags it sends and what it is good for: \-sS SYN Sends SYN. SYN/ACK = open, RST = closed. The half-open default. \-sT connect Completes the full handshake via the OS. No root needed, but logged. \-sA ACK Sends bare ACK. Maps firewall rules: filtered vs unfiltered, not open. \-sW window Like ACK, but reads the RST window size to guess open vs closed. \-sF FIN Sends bare FIN. No reply = open|filtered, RST = closed. \-sN null Sends no flags at all. Same logic as FIN, maximum stealth. \-sX Xmas Sets FIN, PSH, and URG (lit up like a tree). Same FIN-scan logic. \-sM Maimon Sets FIN/ACK. Exploits a BSD quirk to spot open ports. The SYN and connect scans are covered in depth under [port scanning](https://www.pinglabz.com/nmap-port-scanning/). This article focuses on the rest of the family, because that is where the interesting (and dangerous) behavior lives. ## ACK scan maps firewall rules, not open ports The ACK scan, `-sA`, is the odd one out: it is not trying to find open ports at all. It sends a bare ACK, and since a stateful firewall only permits ACKs that belong to an established connection, an unsolicited ACK gets a very telling reaction. If the ACK reaches a live host, the host sends a RST (there is no connection, so it rejects it), and Nmap marks the port **unfiltered**. If a stateful firewall drops the ACK instead, Nmap sees nothing back and marks it **filtered**. So the ACK scan does not tell you what is listening, it tells you where the firewall is. Run it through a real Cisco ASA and every port comes back filtered, because a stateful firewall drops any ACK that is not part of an established session: ``` $ nmap -sA -p 22,23,80,443 10.10.20.32 Starting Nmap 7.95 ( https://nmap.org ) at 2026-06-28 07:01 UTC Nmap scan report for 10.10.20.32 Host is up (0.0086s latency). PORT STATE SERVICE 22/tcp filtered ssh 23/tcp filtered telnet 80/tcp filtered http 443/tcp filtered https Nmap done: 1 IP address (1 host up) scanned in 1.35 seconds ``` All filtered is the signature of a stateful firewall in the path. Now contrast that with an ACK scan against a plain Linux host with no stateful firewall between you and it: ``` $ nmap -sA -p22,80,443 127.0.0.1 Starting Nmap 7.95 ( https://nmap.org ) at 2026-07-05 15:05 PDT Nmap scan report for localhost (127.0.0.1) Host is up (0.000042s latency). PORT STATE SERVICE 22/tcp unfiltered ssh 80/tcp unfiltered http 443/tcp unfiltered https Nmap done: 1 IP address (1 host up) scanned in 0.10 seconds ``` Every port is **unfiltered**: the host RST'd each ACK, proving nothing stateful is dropping traffic in the path. Put the two side by side and the ACK scan becomes a firewall detector. All filtered means a stateful device is between you and the target; all unfiltered means the path is clear. That is genuinely useful when you are verifying whether an ASA rule is actually taking effect, and it is the foundation of the techniques in [firewall and IDS evasion](https://www.pinglabz.com/nmap-firewall-evasion/). ## FIN, NULL, and Xmas: the RFC 793 stealth scans The FIN (`-sF`), NULL (`-sN`), and Xmas (`-sX`) scans all lean on the same trick from RFC 793, the original TCP specification. That RFC says a compliant stack must send a RST when it receives a packet with no SYN, RST, or ACK flag set to a *closed* port, and must send *nothing* when the same packet hits an *open* port. So Nmap sends a weird flag combination (bare FIN, no flags at all, or FIN/PSH/URG lit up like a Christmas tree) and reads the silence: - Closed port sends a RST, so Nmap reports **closed**. - Open port sends nothing, so Nmap cannot tell open from a firewall drop and reports **open|filtered**. These scans are called stealthy because they never open a connection and often slip past simple filters looking for SYNs. On a compliant stack the logic works perfectly. Here are all three run against a Linux localhost, and they behave exactly as RFC 793 predicts: ``` $ nmap -sF -p22,80,443 127.0.0.1 PORT STATE SERVICE 22/tcp open|filtered ssh 80/tcp open|filtered http 443/tcp closed https $ nmap -sN -p22,80,443 127.0.0.1 PORT STATE SERVICE 22/tcp open|filtered ssh 80/tcp open|filtered http 443/tcp closed https $ nmap -sX -p22,80,443 127.0.0.1 PORT STATE SERVICE 22/tcp open|filtered ssh 80/tcp open|filtered http 443/tcp closed https ``` All three agree: 22 and 80 (which are open) report `open|filtered` because they stayed silent, and 443 (which is closed) reports `closed` because it sent a RST. This is the RFC 793 behavior working as designed on a Linux stack. ### Where stealth scans fail: non-compliant stacks Now the trap. RFC 793 compliance is not universal, and Cisco IOS is a well-known example of a stack that sends a RST for *every* unexpected packet, open port or not. Watch what a FIN scan does to a Cisco router whose ports 22 and 23 are genuinely open: ``` $ nmap -sF 10.10.10.1 Starting Nmap 7.95 ( https://nmap.org ) at 2026-06-28 07:01 UTC Nmap scan report for 10.10.10.1 Host is up (0.0038s latency). All 1000 scanned ports on 10.10.10.1 are in ignored states. Not shown: 1000 closed tcp ports (reset) MAC Address: AA:BB:CC:00:5B:00 (Unknown) Nmap done: 1 IP address (1 host up) scanned in 1.76 seconds ``` Every single port reads **closed (reset)**, including 22 and 23, which you know from a SYN scan are open. The Xmas scan gives the identical result, and the NULL scan behaves the same way. Because the IOS stack RSTs everything, the actually-open ports look closed, and the stealth scan reports a router with nothing listening. If you trusted this output you would walk away thinking the device was locked down when SSH and Telnet are wide open. The lesson: FIN, NULL, and Xmas scans only tell the truth against RFC-793-compliant stacks (mostly Unix-like hosts), and they will actively mislead you on Cisco gear, many Windows systems, and other stacks that RST unconditionally. ## Window scan: reading the RST size The Window scan, `-sW`, is a refinement of the ACK scan. It sends the same bare ACK, but instead of only checking whether a RST comes back, it inspects the TCP window field of that RST. On certain stacks an open port returns a RST with a non-zero window while a closed port returns a zero window, letting Nmap distinguish open from closed where a plain ACK scan could only say unfiltered. The catch is that this depends entirely on quirks of the target's TCP implementation, so it is unreliable across the board and produces plenty of false results on stacks that do not exhibit the behavior. Treat it as a situational tool, not a primary scan. ## Maimon scan: the FIN/ACK BSD quirk The Maimon scan, `-sM`, sets both the FIN and ACK flags. By RFC 793 any host should answer such a probe with a RST regardless of port state, which would make it useless. Its author, Uriel Maimon, discovered that many BSD-derived stacks break the rule and drop the packet silently on an open port while still sending a RST on a closed one, exactly the open|filtered versus closed split the FIN scan relies on. On modern hardware that quirk is rare, so Maimon is mostly a historical and niche technique, but it is worth recognizing in the family because it shows how scan types exploit specific stack behaviors rather than a single universal rule. ## Which scan when The family is large, but the decision is usually simple. Use this as your quick reference: Default, everyday scan \-sS \- fast, quiet, reliable on any stack. No root privileges \-sT - the connect scan, though it gets logged. Map a firewall \-sA - filtered vs unfiltered tells you where the ACLs are. Slip past a SYN filter (Unix target) \-sF / -sN / -sX - only trust them against RFC-793 stacks. What that table leaves out is that the scan you pick also decides what the target writes down. A FIN scan, a UDP scan and a SYN scan land on completely different lines of a logging access list, which means [a defender can often name the scan type from the counters alone](https://www.pinglabz.com/detecting-nmap-scans-cisco-blue-team/) without ever seeing your command line. ## Key takeaways The TCP scan family is a set of flag tricks, and the right choice depends on the target's stack as much as your goal. Use `-sS` as your everyday scan, `-sT` when you lack root, and `-sA` to map firewalls: all filtered points to a stateful device in the path, all unfiltered means the path is clear. The stealth scans (`-sF`, `-sN`, `-sX`) exploit RFC 793 and work cleanly on Unix-like hosts, where an open port stays silent (open|filtered) and a closed port sends a RST, but they fail hard on non-compliant stacks like Cisco IOS that RST every port and make genuinely open services look closed. Window and Maimon scans are situational tools built on specific stack quirks. When in doubt, cross-check a surprising result with a straight SYN scan. For the fundamentals behind these techniques see [port scanning](https://www.pinglabz.com/nmap-port-scanning/) and, to turn firewall mapping into evasion, [firewall and IDS evasion](https://www.pinglabz.com/nmap-firewall-evasion/). For the full workflow and the rest of the cluster, start at [the complete Nmap guide](https://www.pinglabz.com/nmap/). ### Nmap Port Scanning: SYN vs Connect, Open, Closed, and Filtered URL: https://www.pinglabz.com/nmap-port-scanning/ Last updated: 2026-08-01T19:30:59.000Z Once you know which hosts are alive, the real work is finding out what each one is listening on. Port scanning is how you build a service inventory, confirm a firewall change did what you expected, and catch the forgotten management port nobody remembers opening. This article is part of [the complete Nmap guide](https://www.pinglabz.com/nmap/), and it covers the core of day-to-day scanning: the SYN scan that does most of the work, how it differs from a connect scan, how to choose which ports to hit, and the one lesson that separates people who read Nmap output correctly from people who guess - the difference between open, closed, and filtered. ## The SYN scan is the default workhorse The SYN scan, `-sS`, is what Nmap runs by default when you have the privileges for it, and for good reason. It sends a SYN, and if it gets a SYN/ACK back it knows the port is open, but instead of completing the handshake it sends a RST and moves on. The connection is never fully established, which is why it is called a half-open scan. It is fast, it is light on the target, and it is often not logged by applications that only record completed connections. Here is a real SYN scan of four hosts on a lab LAN, captured live on Cisco Modeling Labs: ``` $ nmap -sS -T4 10.10.10.1 10.10.10.2 10.10.10.21 10.10.10.24 Starting Nmap 7.95 ( https://nmap.org ) at 2026-06-28 05:51 UTC Nmap scan report for 10.10.10.1 Host is up (0.0050s latency). Not shown: 998 closed tcp ports (reset) PORT STATE SERVICE 22/tcp open ssh 23/tcp open telnet MAC Address: AA:BB:CC:00:5B:00 (Unknown) Nmap scan report for 10.10.10.21 Host is up (0.0094s latency). Not shown: 999 closed tcp ports (reset) PORT STATE SERVICE 80/tcp open http MAC Address: 52:54:00:A2:2F:84 (QEMU virtual NIC) Nmap scan report for 10.10.10.24 Host is up (0.0082s latency). Not shown: 999 closed tcp ports (reset) PORT STATE SERVICE 22/tcp open ssh MAC Address: 52:54:00:BB:DC:36 (QEMU virtual NIC) Nmap done: 4 IP addresses (4 hosts up) scanned in 4.16 seconds ``` Look at how the device profiles fall out of the port list. The Cisco gear at .1 answers on 22 and 23 (SSH and Telnet, the classic router management pair). The host at .21 answers only on 80, so it is a web server (this is the nginx box). The host at .24 answers only on 22, an SSH-only Linux host. You did not run version detection yet, and you can already guess what each device is from its open ports alone. Now read the line most people skip: `Not shown: 998 closed tcp ports (reset)`. That parenthetical is doing real work. Those 998 ports are **closed**, and Nmap knows that because the host answered each probe with a RST. Closed means the host is reachable and nothing is listening on that port. It is a completely different state from filtered, and confusing the two is the single most common scanning mistake. ## \-sS half-open vs -sT full connect When you cannot run as root, or you are scanning through a proxy that only speaks full connections, you fall back to the TCP connect scan, `-sT`. Instead of crafting raw packets, it asks the operating system to open a normal connection with `connect()`, completing the entire three-way handshake. Same result, different mechanics, and two differences that matter in practice. ``` $ nmap -sT 10.10.10.1 Starting Nmap 7.95 ( https://nmap.org ) at 2026-06-28 07:00 UTC Nmap scan report for 10.10.10.1 Host is up (0.0029s latency). Not shown: 998 closed tcp ports (conn-refused) PORT STATE SERVICE 22/tcp open ssh 23/tcp open telnet MAC Address: AA:BB:CC:00:5B:00 (Unknown) Nmap done: 1 IP address (1 host up) scanned in 3.49 seconds ``` The tell is in the parenthetical again: `conn-refused` instead of the SYN scan's `reset`. That is the OS reporting a refused connection rather than Nmap seeing the raw RST itself. The two trade-offs: `-sT` needs no root privileges, which is its whole reason to exist, but because it completes the handshake, the target application logs a full connection. A SYN scan on the same host would leave far less of a footprint. Reach for `-sT` when you lack privileges; prefer `-sS` the rest of the time. Less of a footprint is not none, though. The application never sees the connection, but the router in front of it counts every probe you send, and a single logging ACE will happily rack up over a thousand hits while your terminal reports a four second scan. If you want that view, here is [what a half-open scan leaves behind on the Cisco box you pointed it at](https://www.pinglabz.com/detecting-nmap-scans-cisco-blue-team/). ## Picking which ports to scan By default Nmap scans the top 1000 most common ports, not all 65535\. That default is a good balance for most work, but you should know how to widen or narrow it deliberately. Here are the port-selection flags you will actually use: \-F (fast) Top 100 ports only. Roughly a tenth of the default, for a quick pass. default (no -p) Top 1000 ports by frequency. The sensible everyday choice. \--top-ports N The N most common ports. Dial the breadth to the time you have. \-p 22,80,443 An explicit list. Fastest when you already know what to check. \-p- All 65535 ports. Slow, but the only way to find services hiding high. \-p U:53,T:80 Mixed protocols in one run: UDP 53 and TCP 80 together. The `-F` flag maps to a top-100 scan (the lab capture shows it finishing a router in 1.37 seconds), and an explicit `-p` list is faster still because Nmap only probes what you asked for. Note that only the **live** hosts should reach this stage, which is why [host discovery](https://www.pinglabz.com/nmap-host-discovery/) comes first: scan the confirmed-up list, not empty address space. ## Open, closed, filtered: the lesson that matters most This is the centerpiece. Nmap reports one of three states for a scanned port, and reading them correctly is what makes your scans trustworthy. The cleanest way to see all three at once is to scan a single host through a real firewall. Here is a SYN scan of a DMZ target sitting behind a Cisco ASA, where the firewall permits TCP 22 and 80 to that host and denies everything else: ``` $ nmap -sS -p 22,23,80,443 10.10.20.32 Starting Nmap 7.95 ( https://nmap.org ) at 2026-06-28 05:53 UTC Nmap scan report for 10.10.20.32 Host is up (0.0098s latency). PORT STATE SERVICE 22/tcp open ssh 23/tcp filtered telnet 80/tcp closed http 443/tcp filtered https Nmap done: 1 IP address (1 host up) scanned in 1.37 seconds ``` Four ports, three different states, one host. Here is what each one is telling you: - **22 open** \- the firewall permits it and the target is listening. A SYN/ACK came back. - **80 closed** \- the firewall permits it, the packet reached the host, but nothing is listening, so the host answered with a RST. Closed means "I got to the host and there is no service here." - **23 filtered** \- the firewall dropped the probe silently. No SYN/ACK, no RST, nothing came back. Filtered means "something ate the packet before the host could answer." - **443 filtered** \- same silent drop. The distinction between closed and filtered is the whole game. Closed is an answer (a RST from a reachable host). Filtered is silence (a firewall or ACL swallowed the probe). If you misread a filtered port as closed, you will conclude a host is reachable and idle when in fact a firewall is standing between you and it. This is exactly the behavior you would map when auditing an ASA policy, and the [Cisco ASA firewall guide](https://www.pinglabz.com/cisco-asa/) covers the rule structure that produces these results. ### Let --reason prove each state You do not have to take Nmap's word for it. Add `--reason` and it prints the exact packet behind every verdict. Here is the same DMZ host with reasons attached: ``` $ nmap --reason -p 22,23,80,443 10.10.20.32 Nmap scan report for 10.10.20.32 Host is up, received echo-reply ttl 254 (0.0074s latency). PORT STATE SERVICE REASON 22/tcp open ssh syn-ack ttl 254 23/tcp filtered telnet no-response 80/tcp closed http reset ttl 254 443/tcp filtered https no-response Nmap done: 1 IP address (1 host up) scanned in 1.36 seconds ``` Now the mapping is spelled out on every line: `syn-ack` means open, `reset` means closed, and `no-response` means filtered. Memorize those three reasons and you can read any port state without second-guessing. When a result surprises you in the field, `--reason` is the first flag to add. ## The top-1000 trap: services live high Here is the mistake that bites people scanning their own infrastructure. The default top-1000 port list is built from statistical frequency across the internet, and plenty of important services do not live in it. Web apps, dashboards, and management consoles routinely bind to high ports: 8080 and 8443 for alternate HTTP and HTTPS, 9000 and up for a whole zoo of admin panels and app servers. A default scan sails right past all of them and reports a host as running only whatever happened to sit in the top 1000. The consequence is a false sense of completeness. You scan a server, see only 22 and 80, and call it clean, while an unauthenticated admin panel is sitting on 8443 the whole time. The fix is to run `-p-` (all 65535 ports) at least once against anything you actually care about, and to add the usual high-port suspects to your targeted scans. It is slower, but on your own network, thoroughness beats speed when the alternative is missing a live service. ## Key takeaways The SYN scan (`-sS`) is your default half-open workhorse: fast, quiet, and enough to profile a device from its open ports alone (Cisco gear shows 22 and 23, a web server shows 80). Fall back to the connect scan (`-sT`) only when you lack root, and remember it gets logged as a full connection. Choose ports deliberately with `-F`, `--top-ports`, an explicit `-p` list, or `-p-` for everything, and never trust the top 1000 to be complete when real services hide on 8080, 8443, and 9000-plus. Above all, read the states correctly: open is a SYN/ACK, closed is a RST from a reachable host, filtered is silence from a firewall, and `--reason` will prove it. To go deeper on the different probe techniques see [scan types](https://www.pinglabz.com/nmap-scan-types/) and, for hosts that hide behind ACLs, [firewall and IDS evasion](https://www.pinglabz.com/nmap-firewall-evasion/). For the full workflow and the rest of the cluster, start at [the complete Nmap guide](https://www.pinglabz.com/nmap/). ### Nmap Host Discovery: Ping Sweeps, ARP, and -Pn URL: https://www.pinglabz.com/nmap-host-discovery/ Last updated: 2026-08-01T19:31:00.000Z Before you scan a single port, you need to know which addresses are actually alive. Blasting a full port scan at every host in a /24 wastes time on dead space and lights up any IDS in the path. Host discovery is the front door of every Nmap workflow: you find the live hosts first, then you scan the ones that answer. This article is part of [the complete Nmap guide](https://www.pinglabz.com/nmap/), and it covers exactly how Nmap decides a host is up, why that decision changes depending on whether the target is local or remote, and how to read the results like a network engineer auditing their own gear. ## Why host discovery comes first When you point Nmap at a range, its default behavior is a two-step job: discover which hosts are up, then port-scan only those. The discovery phase is the "ping scan" (sometimes called a ping sweep). Skipping straight to a port scan of the whole subnet means Nmap either wastes probes on empty addresses or, worse, silently drops hosts that did not answer the ping. Understanding the discovery phase is what keeps your inventory accurate. The discovery-only flag is `-sn` (no port scan). It tells Nmap to find live hosts and stop. That makes it the fastest way to answer one question you ask constantly in the field: what is actually powered on and reachable in this subnet right now? ## On a local subnet, -sn is an ARP sweep Here is the detail most tutorials skip. When your target is on the same layer-2 segment as your scanner, Nmap does not send ICMP or TCP at all for discovery. It sends ARP. A host cannot answer an IP-layer probe without first answering ARP, so ARP is both faster and more reliable, and Nmap uses it automatically for local targets even if you asked for other probe types. Here is a real ARP sweep of a lab /24 from the scanner, captured live on Cisco Modeling Labs: ``` $ nmap -sn 10.10.10.0/24 Starting Nmap 7.95 ( https://nmap.org ) at 2026-06-28 05:49 UTC Nmap scan report for 10.10.10.1 Host is up (0.0031s latency). MAC Address: AA:BB:CC:00:5B:00 (Unknown) Nmap scan report for 10.10.10.2 Host is up (0.0080s latency). MAC Address: AA:BB:CC:80:5A:00 (Unknown) Nmap scan report for 10.10.10.21 Host is up (0.012s latency). MAC Address: 52:54:00:A2:2F:84 (QEMU virtual NIC) Nmap scan report for 10.10.10.24 Host is up (0.0036s latency). MAC Address: 52:54:00:BB:DC:36 (QEMU virtual NIC) Nmap scan report for 10.10.10.25 Host is up (0.0041s latency). MAC Address: 52:54:00:CF:69:0F (QEMU virtual NIC) Nmap scan report for 10.10.10.10 Host is up. Nmap done: 256 IP addresses (6 hosts up) scanned in 1.98 seconds ``` Six live hosts out of 256 addresses in 1.98 seconds. That speed is the ARP advantage: no routing, no three-way handshakes, just layer-2 request and reply. The .10 host at the end (`Host is up.` with no MAC) is the scanner itself. ### Read the MAC OUI before you send a TCP packet Notice what the sweep handed you for free: a MAC address per host. The first three bytes are the OUI (Organizationally Unique Identifier), and they tell you what kind of device you are looking at before you probe a single port. In this capture the pattern is obvious once you know it: - `AA:BB:CC:...` on .1 and .2 are Cisco IOL burnt-in MACs, so those are the router (R1) and the switch SVI (SW1). Network gear. - `52:54:00:...` on .21, .24, and .25 is the QEMU/KVM OUI, so those are the Linux virtual hosts (the nginx web server WEB01, and hosts HOST01 and HOST02). That is a real triage shortcut. On an unfamiliar subnet you can separate "infrastructure" from "endpoints" straight off the ARP sweep, and prioritize what to scan next. It is exactly the kind of inventory-first thinking that turns Nmap from a security tool into a network verification tool. ## What -sn sends on a remote subnet The moment the target is a router hop away, ARP is off the table (ARP does not cross a layer-3 boundary). Nmap switches to a default set of IP-layer discovery probes. As an unprivileged user it falls back to a TCP connect on 80 and 443, but run as root (which you normally are for scanning), the default remote discovery probes are: - An ICMP echo request (a classic ping). - A TCP SYN to port 443. - A TCP ACK to port 80. - An ICMP timestamp request. Nmap fires all four and treats the host as up if any one of them draws a response. The mix is deliberate: a firewall that blocks ICMP echo often still permits TCP to 80 or 443, and the ACK probe slips past some stateless filters that the SYN probe cannot. Casting several probe types is what makes remote discovery reliable when a single ping would fail. Reliable is not the same as quiet. Those four probe types hit four different lines of any access list they cross on the way in, so a /24 sweep that finishes in under three seconds still leaves a small ICMP counter sitting next to a router interface. Here is [how a ping sweep reads from the box that answered it](https://www.pinglabz.com/detecting-nmap-scans-cisco-blue-team/). ## The discovery probe types you can choose You are not stuck with the defaults. When a network filters the standard probes, you tune discovery with explicit probe flags. Here are the ones worth knowing, and what each one actually puts on the wire: \-PS (TCP SYN) Sends a SYN to the listed ports. A SYN/ACK or a RST both prove the host is up. Best default for firewalled hosts that permit web traffic. \-PA (TCP ACK) Sends a bare ACK. A RST comes back from any live host, listening or not. Slips past stateless filters that only block SYNs. \-PU (UDP) Sends a UDP datagram. An ICMP port-unreachable proves the host is up. Reaches hosts that block all TCP-based pings. \-PE (ICMP echo) The classic ping. Fast and clean where ICMP is allowed. Often blocked at the perimeter, useful inside. \-PP (ICMP timestamp) An ICMP timestamp request. Some filters block echo but forget timestamp. A good backup ICMP probe. \-PR (ARP) Layer-2 ARP request. The automatic choice for local targets. Fastest and most reliable, but local-only. You can stack these on a single command. Here is a discovery-only scan of the real internet host `scanme.nmap.org` using SYN probes to three ports, an ACK probe to 80, a UDP probe to 53, and `--reason` so Nmap prints why it made its call: ``` $ nmap -sn -PS22,80,443 -PA80 -PU53 --reason scanme.nmap.org Starting Nmap 7.95 ( https://nmap.org ) at 2026-07-05 15:06 PDT Nmap scan report for scanme.nmap.org (45.33.32.156) Host is up, received syn-ack ttl 50 (0.016s latency). Other addresses for scanme.nmap.org (not scanned): 2600:3c01::f03c:91ff:fe18:bb2f Nmap done: 1 IP address (1 host up) scanned in 0.54 seconds ``` `received syn-ack ttl 50` is the tell: the SYN probe to one of the listed ports drew a SYN/ACK, so Nmap marked the host up. That is the whole point of stacking probe types. If ICMP had been the only probe and the perimeter dropped it, you would have missed a host that is clearly alive. ## The -Pn problem: when discovery lies to you Here is the trap. If a firewall drops every one of Nmap's discovery probes, Nmap concludes the host is down and never port-scans it. The host might be running a dozen services, but because it refused to answer a ping, it falls off your results entirely. This is the single most common reason a scan comes back empty when you know the target is up. The fix is `-Pn`: skip host discovery entirely and treat every target as up. Nmap goes straight to port scanning: ``` $ nmap -Pn 10.10.10.21 Starting Nmap 7.95 ( https://nmap.org ) at 2026-06-28 07:02 UTC Nmap scan report for 10.10.10.21 Host is up (0.017s latency). Not shown: 999 closed tcp ports (reset) PORT STATE SERVICE 80/tcp open http Nmap done: 1 IP address (1 host up) scanned in 1.74 seconds ``` Use `-Pn` whenever you have good reason to believe a host is up but ping is being blocked. The cost is speed and noise (Nmap now port-scans every target in the range whether it is alive or not), so reserve it for the cases where discovery is actively working against you. Hosts blocking ping on purpose is exactly the kind of behavior covered in [firewall and IDS evasion](https://www.pinglabz.com/nmap-firewall-evasion/), and it cuts both ways: the same drop rule that hides an attacker from you also hides your own gear from your inventory scans. ## List scan, DNS, and --reason Sometimes you want the target list without touching a single host. The list scan, `-sL`, does exactly that: it prints every IP in the range (and their reverse DNS names, if resolvable) and sends no probes at all. It is a completely passive sanity check. Run it before a big scan to confirm you are aiming at the range you think you are, and to catch a fat-fingered CIDR before it hits production. DNS behavior is worth controlling deliberately. By default Nmap does reverse DNS on live hosts, which adds useful context but also generates lookups. Two flags govern it: `-n` disables all reverse DNS (faster, and it keeps your scan off the DNS server logs), while `-R` forces reverse DNS even on hosts Nmap thinks are down. In the lab captures you will notice a warning that reverse DNS was disabled because the scanner had no DNS server configured, which is a reminder that Nmap only resolves names when it actually can. Finally, get in the habit of adding `--reason`. It changes "Host is up" into "Host is up, received syn-ack ttl 50" and turns every port state into an explanation of the packet that produced it. When a discovery result surprises you, `--reason` is the fastest way to see whether a host answered with a SYN/ACK, a RST, an ICMP unreachable, or nothing at all. ## Key takeaways Host discovery is where you save time and stay accurate. On a local subnet, `-sn` is really an ARP sweep, and it hands you a MAC OUI per host so you can tell Cisco gear (AA:BB:CC here) from Linux VMs (52:54:00) before you send a single TCP packet. On a remote subnet Nmap casts a wider net with ICMP echo, a TCP SYN to 443, a TCP ACK to 80, and an ICMP timestamp, and you can tune that with `-PS`, `-PA`, `-PU`, `-PE`, `-PP`, and `-PR`. Remember the `-Pn` escape hatch for hosts that block ping, use `-sL` to sanity-check a range without probing, and reach for `--reason` whenever a result needs explaining. Once you have a confirmed list of live hosts, the next move is [port scanning](https://www.pinglabz.com/nmap-port-scanning/) to find out what each one is actually running. For the full workflow and the rest of the cluster, start at [the complete Nmap guide](https://www.pinglabz.com/nmap/). ### Spanning Tree PortFast: Faster Access Ports, Safely URL: https://www.pinglabz.com/spanning-tree-portfast/ Last updated: 2026-07-01T13:00:00.000Z PortFast is a small spanning tree feature with an outsized practical effect: it is the difference between a host getting an IP address the instant it plugs in and a host sitting dark for thirty seconds while STP makes up its mind. It is also one of the easiest features to misapply in a way that takes down a network. This post explains what PortFast does, where it belongs, where it absolutely does not, and the two guard features that keep it safe. For the cluster overview, see the [Spanning Tree Protocol complete guide](https://www.pinglabz.com/spanning-tree-protocol/). ## The problem PortFast solves When a port comes up, classic 802.1D spanning tree walks it through Listening and Learning before Forwarding - 15 seconds in each, 30 seconds total. STP does this so a newly active port cannot instantly create a bridging loop. For a link to another switch, that caution is correct. For an access port with a single PC, printer, or server on it, it is pure dead time. Thirty seconds of silence is long enough that a DHCP client gives up, a PXE boot fails, or a server's network stack flags the interface as down. The classic symptom is "the laptop never gets an IP on a fresh boot but works fine after a reconnect." ## What PortFast does PortFast tells the switch a port is an **edge port** \- a port with an end device on it, not another switch. An edge port skips Listening and Learning and goes **straight to Forwarding** the moment it links up. The host gets immediate connectivity. PortFast does not disable spanning tree on the port. STP still runs; the port can still receive and process BPDUs. PortFast only changes the startup transition. In Rapid PVST+ the same idea is built in as the "edge port" concept, and PortFast is how you mark a port as edge. ## Where PortFast belongs - and where it does not The rule is simple and absolute: PortFast goes on access ports that connect to **end devices only**. Never on a port that connects to another switch, bridge, or hub. The reason is the loop. A PortFast port jumps to Forwarding instantly. If someone plugs a switch into that port, you have a forwarding path that bypassed the Listening/Learning safety window - a bridging loop forms before STP can react, and a loop can saturate a LAN in seconds. PortFast on the wrong port is one of the faster ways to melt a network. ## BPDU Guard: the safety net Because a PortFast port should only ever face an end device, it should never receive a BPDU - end devices do not send them. BPDU Guard enforces that assumption. If a PortFast port with BPDU Guard enabled receives *any* BPDU, the switch immediately puts the port into **err-disabled** state, shutting it down. That is exactly the behavior you want. Someone plugged a switch into a user port? The port shuts before a loop can form, instead of after. BPDU Guard turns "PortFast on the wrong port" from an outage into a single dead port and a log message. ``` %SPANTREE-2-BLOCK_BPDUGUARD: Received BPDU on port Gi0/5 with BPDU Guard enabled. Disabling port. %PM-4-ERR_DISABLE: bpduguard error detected on Gi0/5, putting Gi0/5 in err-disable state ``` ## BPDU Filter: handle with care BPDU Filter stops a port from sending and processing BPDUs at all. It is occasionally used at a demarcation to a customer or third party so your STP domain does not extend into theirs. It is dangerous when enabled globally for PortFast ports, because a filtered port that is silently bridged into another switch has neither STP nor BPDU Guard watching it - the exact loop PortFast risks, with the safety net removed. Treat BPDU Filter as a deliberate per-interface tool, not a default. For ordinary access ports, use BPDU Guard, not BPDU Filter. ## Configuration on Cisco IOS XE Two ways to apply PortFast. Per interface, on a known access port: ``` interface GigabitEthernet0/5 switchport mode access switchport access vlan 10 spanning-tree portfast spanning-tree bpduguard enable ``` Or globally, which is the recommended pattern - it applies PortFast and BPDU Guard to every port that operates as an access port, without touching trunks: ``` spanning-tree portfast default spanning-tree portfast bpduguard default ``` With the global form, you configure the pair once and every new access port is protected automatically. Recovery from err-disable can be automated: ``` errdisable recovery cause bpduguard errdisable recovery interval 300 ``` ## Verifying ``` SW1# show spanning-tree interface Gi0/5 portfast VLAN0010 enabled SW1# show spanning-tree summary Portfast Default is enabled PortFast BPDU Guard Default is enabled ``` ## Common gotchas Host never gets a DHCP address on first boot No PortFast on the access port - the port spends 30s in Listening/Learning. Enable PortFast. A user port keeps going err-disabled BPDU Guard caught a BPDU - someone plugged a switch or hub into a PortFast port. That is the feature working. PortFast configured but the port still delays The port is operating as a trunk. PortFast on a normal trunk needs `spanning-tree portfast trunk` and is rarely appropriate. A bridging loop formed off an access port PortFast was on without BPDU Guard. Always pair the two. err-disabled port never recovers No errdisable recovery configured. Re-enable manually with shut/no shut or set `errdisable recovery`. ## Key takeaways PortFast sends an access port straight to Forwarding instead of through the 30-second Listening and Learning delay, which is what lets a host get an address the moment it connects. It belongs on access ports facing end devices only and never on switch-to-switch links, because an instant-forwarding port can create a bridging loop. Always pair PortFast with BPDU Guard: a PortFast port should never see a BPDU, and BPDU Guard err-disables it the instant one arrives, turning a potential outage into one dead port. The cleanest deployment is the global default form, which protects every access port automatically. For the STP cluster, see the [Spanning Tree pillar](https://www.pinglabz.com/spanning-tree-protocol/). ### Native VLAN Explained: Untagged Traffic and VLAN Hopping URL: https://www.pinglabz.com/native-vlan/ Last updated: 2026-08-01T19:35:39.000Z The native VLAN is one of those concepts that seems trivial - it is "the untagged VLAN on a trunk" - right up until a mismatch breaks a link, or a double-tagging attack hops a VLAN you thought was isolated. The native VLAN is small but load-bearing. This post explains what it is, why it exists, the mismatch problem, the security hole, and how to configure it so neither bites you. For the cluster overview, see the [VLAN and Layer 2 switching pillar](https://www.pinglabz.com/vlans-layer-2-switching/). ## Tagging, and the one exception On an 802.1Q trunk, frames carry a 4-byte VLAN tag so the switch on the other end knows which VLAN each frame belongs to. That is the whole point of a trunk: many VLANs, one link, every frame labeled. The native VLAN is the deliberate exception. Frames in the native VLAN cross the trunk **untagged**. When a switch sends a frame in the native VLAN, it omits the tag entirely. When it receives an untagged frame on a trunk, it assumes that frame belongs to the native VLAN. So the native VLAN is defined by what is *missing*: it is the one VLAN whose frames have no tag. ## Why an untagged VLAN exists at all The native VLAN is a backward-compatibility feature. Early or simple devices that share a link - legacy hubs, some IP phones, a switch that does not understand 802.1Q - cannot read a tag. The native VLAN gives them a lane: untagged frames still have a home. It is also where a switch's own control-plane chatter historically rode. In a pure modern switch-to-switch trunk, almost nothing actually needs the native VLAN, which is exactly why it becomes a security liability if you ignore it. ## The default: VLAN 1 Out of the box, the native VLAN on every Cisco trunk is **VLAN 1**. VLAN 1 is also the default access VLAN for every port and the VLAN that protocols like CDP, VTP, DTP, and PAgP use. Leaving everything on VLAN 1 means your management traffic, your control protocols, and your default user traffic all share one untagged VLAN - a poor security posture and the reason every hardening guide tells you to move off VLAN 1. ## The native VLAN mismatch Both ends of a trunk must agree on the native VLAN number. If R1's trunk has native VLAN 1 and R2's has native VLAN 99, you have a **native VLAN mismatch**. The consequence is real. An untagged frame R1 sends in VLAN 1 arrives at R2, which sees no tag and drops it into VLAN 99\. The two VLANs are now bridged together - traffic leaks between them. This is a security problem and a loop risk. Cisco switches detect the mismatch via CDP and log it: ``` %CDP-4-NATIVE_VLAN_MISMATCH: Native VLAN mismatch discovered on GigabitEthernet0/1 (1), with SW2 GigabitEthernet0/1 (99). ``` Spanning tree also reacts: PVST+ can place the mismatched VLAN into a blocking-style inconsistent state to contain the damage. Either way, a native VLAN mismatch is never something to leave logged and ignored. That log line is more generous than it looks: it names the interface, both native VLAN numbers, and which switch is on the other end, and it repeats until someone fixes it. [Reading the CDP message and clearing the mismatch](https://www.pinglabz.com/native-vlan-mismatch-troubleshooting/) takes it from that single line to a corrected trunk. ## The security angle: VLAN hopping by double tagging The native VLAN is the enabler of the double-tagging VLAN hopping attack. An attacker on the native VLAN crafts a frame with *two* 802.1Q tags. The first switch strips the outer tag (it matches the native VLAN, so it is removed as the frame goes untagged onto the trunk) and forwards the frame. The second switch reads the still-present inner tag and delivers the frame into the target VLAN - one the attacker was never supposed to reach. The attack only works when the attacker's access VLAN equals the trunk's native VLAN. That single fact drives the entire mitigation: never let the native VLAN be a VLAN that has live user access ports. ## Configuring the native VLAN safely Two practices close the hole. First, set the native VLAN to a dedicated, unused VLAN - one with no access ports and no hosts: ``` ! Create an unused parking VLAN and make it native on the trunk vlan 999 name NATIVE_UNUSED ! interface GigabitEthernet0/1 switchport trunk encapsulation dot1q switchport mode trunk switchport trunk native vlan 999 ``` Apply the identical `native vlan 999` on the far end of every trunk. Second - the stronger control - force the switch to tag the native VLAN too, removing the untagged exception entirely: ``` ! Tag all VLANs on trunks, including the native VLAN vlan dot1q tag native ``` With native tagging on, double tagging stops working because there is no untagged frame for the attack to exploit. Verify the trunk: ``` SW1# show interfaces trunk Port Mode Encapsulation Status Native vlan Gi0/1 on 802.1q trunking 999 ``` ## Common gotchas CDP logs a native VLAN mismatch The two trunk ends have different `native vlan` numbers. Set both to the same value. Traffic leaks between two VLANs across a trunk Native VLAN mismatch is bridging them. Frames untagged on one side land in a different VLAN on the other. A VLAN goes into an STP inconsistent/blocking state on a trunk PVST+ detected the native VLAN mismatch and is containing it. Fix the mismatch. Double-tagging hop still possible after moving the native VLAN An access port still sits in the new native VLAN. The native VLAN must have zero access ports, or use `vlan dot1q tag native`. IP phone or legacy device stops passing traffic It relied on untagged frames in the old native VLAN. Confirm what genuinely needs untagged service before tagging the native VLAN. ## Key takeaways The native VLAN is the one VLAN whose frames cross an 802.1Q trunk untagged, and it defaults to VLAN 1\. Both ends of every trunk must agree on the native VLAN number, or untagged frames bridge two VLANs together and CDP and STP both raise alarms. The native VLAN is also the enabler of the double-tagging VLAN hopping attack, which works only when an attacker's access VLAN matches the trunk native VLAN. Harden it by setting the native VLAN to a dedicated unused VLAN with no access ports, and ideally by enabling `vlan dot1q tag native` so nothing crosses a trunk untagged at all. For the L2 cluster, see the [VLAN pillar](https://www.pinglabz.com/vlans-layer-2-switching/). ### How to Install the Missing Docker Container Images in Cisco Modeling Labs 2.10 URL: https://www.pinglabz.com/install-cml-2-10-docker-container-images/ Last updated: 2026-07-04T23:06:53.000Z You upgraded to Cisco Modeling Labs 2.10, mounted the latest reference platform package, and the new Docker node types showed up in the palette: Chrome, Firefox, Splunk, Syslog, TACACS+, nginx, Snort, dnsmasq, and the rest. Then you dragged one onto the canvas, hit start, and nothing. The node sits there. Open the node and image definitions and you see the node definition is present, but the image definition is empty. This trips up a lot of people (it comes up regularly in the CML community), and the fix is not obvious, because the container images are shipped separately from the main reference platform. Here is exactly what is happening and how to install them. ## The symptom: node definition yes, image definition no In CML, every node type is two pieces: - **Node definition** \- the metadata: how many interfaces, default RAM, boot behavior, the icon in the palette. - **Image definition** \- the actual disk image or container the node boots from. The standard 2.10 reference platform package gives you the *node definitions* for the Docker containers, so they appear in the palette and look ready to use. But the container images themselves are not in that package. With no image definition behind it, the node will not launch. Depending on how you start it you will see it fall straight back to a stopped state, or a plain "no image definition" complaint. If you went looking in `Tools > Node and Image Definitions` and found the container under Node Definitions but nothing under Image Definitions, that is the whole problem in one screen. ## Why: the container images are a separate download The Docker containers that landed in CML 2.9 and 2.10 (the lightweight, fast-booting alternatives to full VMs) are built and published in their own repository, not bundled into the main refplat ISO. Cisco ships them as a set of reference platform ISOs on GitHub: **👉** [**github.com/CiscoLearning/cml-docker-containers/releases/tag/v0.3.0**](https://github.com/CiscoLearning/cml-docker-containers/releases/tag/v0.3.0?ref=pinglabz.com) (CML 2.10.0 container ISOs, v0.3.0) That release splits the containers across three ISOs so you only download what you need. Here is the full inventory, which is handy if you are searching for a specific one: chrome Google Chrome (in-lab browser) ISObrowser Size330 MB firefox Firefox (in-lab browser) ISObrowser Size217 MB net-tools Net-tools box: nmap, tshark, and 20+ network tools ISOservices Size158 MB nginx Nginx web server ISOservices Size60 MB dnsmasq Dnsmasq (DHCP, DNS, TFTP) ISOservices Size35 MB frr FRRouting (Free Range Routing) ISOservices Size13 MB radius FreeRADIUS server ISOservices Size49 MB tacplus TACACS+ server ISOservices Size31 MB syslog Syslog-NG server ISOservices Size35 MB snort Snort3 IDS ISOservices Size450 MB thousandeyes-ea ThousandEyes Enterprise Agent ISOservices Size27 MB splunk Splunk Enterprise ISOsplunk Size1.8 GB For most people the one you want is **refplat-20260507-services.iso**. It carries the workhorses: net-tools, nginx, dnsmasq, FRR, RADIUS, TACACS+, Syslog-NG, Snort, and the ThousandEyes agent. Grab the **browser** ISO if you want Chrome or Firefox, and the **splunk** ISO only if you specifically need Splunk (it is large). ## Installing the images from the CML backend (Cockpit, port 9090) You do this from Cockpit, the CML controller's system administration interface, using the built-in **Copy Refplat ISO** action. No terminal needed. **Stop your labs first.** The import restarts the CML services, so it will refuse to run while any nodes are powered on. **1\. Mount the ISO to the CML virtual machine.** Download the ISO you need from the GitHub release (the **services** ISO covers most people). Then attach it to the CML VM's CD/DVD drive in your hypervisor. In **VMware** (ESXi, Workstation, or Fusion): edit the CML VM's settings, set the CD/DVD drive to *Use ISO image file*, browse to `refplat-20260507-services.iso`, and make sure *Connected* and *Connect at power on* are ticked. (On a bare-metal install, present the ISO as optical media however your platform allows.) ![VMware ESXi Edit settings dialog with the CML VM CD/DVD drive set to the refplat services ISO](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/06/vmware-mount-services-iso.png) Step 1: mount the ISO in your hypervisor. In VMware ESXi, point the CML VM's CD/DVD drive at refplat-20260507-services.iso and tick Connect at power on. **2\. Open Cockpit and go to System > CML2.** Browse to `https://:9090` and sign in with your **sysadmin** account (not the usual cml2 lab user). In the left sidebar under **System**, click **CML2** to open the System Maintenance Controls. **3\. Click Copy Refplat ISO.** Under the **Maintenance** section, expand **Copy Refplat ISO** and click the blue **Copy Refplat ISO** button. It copies the node and image definitions (and the container image tarballs) from the mounted ISO to disk, then restarts the CML services. ![Cockpit CML2 System Maintenance Controls page with the Copy Refplat ISO button](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/06/cml2-copy-refplat-iso.png) Step 3: in Cockpit (port 9090) as sysadmin, open System > CML2, then under Maintenance expand Copy Refplat ISO and click the button. **4\. Refresh.** Give the services a moment to restart, then reload the CML web UI. (On a headless or bare-metal controller where mounting a CD is awkward, the same import is available from the shell as `sudo copy-refplat-iso-to-disk.sh `, but the Copy Refplat ISO button above is the easier path for most installs.) ## Verify it worked Back in the CML web UI, open `Tools > Node and Image Definitions` and find the container, for example net-tools or tacplus. This time there should be an **image definition** listed under it (something like `net-tools-2-10-1-3`). Drag the node into a lab and start it. It should boot in seconds, the way a lightweight container should, instead of dropping back to stopped. ## Gotchas worth knowing - **Stop all labs before importing.** The copy fails if any nodes are running, and the service restart will stop them anyway. - **The node definition being present is a red herring.** Seeing the container in the palette does not mean the image is installed. Always check for the image definition. - **Match the release to your CML version.** The v0.3.0 ISOs are built for CML 2.10\. Use the release that matches your install. - **Splunk is its own ISO for a reason.** At 1.8 GB it dwarfs everything else. Skip it unless you need it. - **Personal vs Enterprise does not matter here.** This is the same process on CML Personal and Enterprise; the images are free to download from the GitHub release. ## Key takeaways - In CML 2.10, the Docker container **node definitions** ship with the reference platform, but the **image definitions** do not. - The container images live in a separate GitHub release: [cml-docker-containers v0.3.0](https://github.com/CiscoLearning/cml-docker-containers/releases/tag/v0.3.0?ref=pinglabz.com), split into browser, services, and splunk ISOs. - Mount the ISO to the CML VM, then in Cockpit open **System > CML2 > Copy Refplat ISO** and click it, with all labs stopped. - Confirm the fix by checking that an image definition now exists, then boot the node. Once these are in, the containers are genuinely useful: net-tools gives you nmap and tshark in the lab, dnsmasq and TACACS+ stand in for real services, and Snort lets you watch traffic. We lean on them heavily in our [hands-on labs](https://www.pinglabz.com/labs/). Next in the lab-setup series: [Exploring Cisco Modeling Labs (CML) 2.8: What’s New in This Exciting Release](https://www.pinglabz.com/cisco-modeling-labs-2-8-release-features/). ### Is the CCNA Worth It in 2026? An Honest Answer URL: https://www.pinglabz.com/is-the-ccna-worth-it-in-2026/ Last updated: 2026-07-04T23:06:56.000Z If you are weighing whether to spend the next few months studying for the CCNA, here is the straight version before the detail: yes, it is still worth it in 2026 for most people, as long as you know exactly what it does and does not do. It is the single best way to prove you understand how networks actually move traffic, and it still clears the first filter on almost every entry-level networking job. What it will not do is make you a network engineer on its own. That part is on you, and it is where the free [hands-on labs](https://www.pinglabz.com/labs/) come in later. This year there is also a wrinkle worth understanding before you book anything: the exam is changing. Let us walk through all of it in plain English, because "is it worth it" depends entirely on what you expect from it. ## What the CCNA actually is Strip away the marketing and the CCNA (Cisco Certified Network Associate) is one thing: proof that you understand how a network moves a packet from A to B. That covers the fundamentals every networking job assumes you already know. How IP addressing and [subnetting](https://www.pinglabz.com/ccna-lab-nf-04-ipv4-subnetting-vlsm/) work. How a switch builds a MAC table, and how VLANs carve one switch into many. How a router picks a path, and how protocols like OSPF share those paths around. The basics of network security, wireless, and a little automation. It is a single exam (currently numbered 200-301), there are no prerequisites, and you can sit it whether you have ten years in IT or none. If even that vocabulary feels shaky, start with the [networking terminology every CCNA candidate should know](https://www.pinglabz.com/networking-terminology-every-ccna-candidate-should-know/) and come back. Here is the framing that keeps people sane: the CCNA is the floor, not the ceiling. It proves you have the vocabulary and the mental model to do the job. It does not prove you can do the job yet. That distinction runs through the rest of this article. ## The 2026 wrinkle nobody tells beginners about This is the part that actually matters this year, and most "is the CCNA worth it" articles are too old to mention it. In May 2026, Cisco announced the first major change to the CCNA blueprint since 2019\. (A blueprint is just the official list of topics the exam tests.) The timeline they published is simple: the new exam topics went public on May 20, 2026, the refreshed exam goes live on February 3, 2027, and the current exam stays live right up until that date. So if you are reading this in 2026, you are sitting in a transition window, and that is good news, not bad. You have a clean choice between the exam that is live now and the one that arrives in early 2027. Status Current (200-301) Live now, the global standard New blueprint Topics out May 20, 2026; exam live Feb 3, 2027 Focus Current (200-301) Fundamentals, IP services, security, automation basics New blueprint Same core, plus a security-first mindset and AI in operations How it tests you Current (200-301) Strong on recall, with some hands-on New blueprint More hands-on labs and practical skills, less pure recall Who it suits Current (200-301) Anyone ready to test during 2026 New blueprint Anyone whose study naturally runs into 2027 Validity once you pass Current (200-301) Three years from your pass date New blueprint Three years from your pass date The practical takeaway: **take the current exam now** if you are close to ready. It is the same gold-standard CCNA that more than 1.8 million people already hold, it stays valid for three years, and passing in 2026 does not give you a "lesser" version. A CCNA is a CCNA. **Aim at the new one** if you are just starting and will not be exam-ready until 2027 anyway, because you will land on it naturally. The one thing you should not do is freeze. "Should I wait for the new exam?" is the wrong question for almost everyone. Study the fundamentals, which barely change, and take whichever exam is live when you are ready. Subnetting works the same in 2027 as it did in 2007. ## What the refreshed exam actually changes Cisco built the new blueprint on four pillars. If you are studying into 2027, this is the shape of what you will be tested on. Network infrastructure The classic core: addressing, switching, routing, and wireless. The bedrock is not going anywhere. Troubleshooting and problem-solving More weight on diagnosing a broken network, not just describing a working one. This is the job. A security-first mindset Security baked into how you design and operate, rather than bolted on as one exam section. The role of AI in network operations The genuinely new pillar: understanding where automation and AI fit in running a modern network. Notice what is not on that list: a rewrite of the fundamentals. The core is the same. The shift is toward proving you can apply it under pressure, with security and automation treated as part of the everyday job rather than afterthoughts. ## What a CCNA actually gets you Let us be concrete, because "career growth" is a useless phrase. A CCNA gets you the interview. It clears the HR filter for the roles that start most networking careers: NOC technician, junior network engineer, help-desk-with-a-networking-angle, network operations. Recruiters screen for it because it is a cheap, reliable signal that you are not starting from zero. On money, the numbers move around depending on who is counting and where you live, but the shape is consistent. Entry-level networking roles commonly land somewhere in the 60,000 to 85,000 US dollar range, and engineers with a few years and the right specialization comfortably clear six figures. The CCNA itself is not what pays you the higher number. It is the thing that gets you onto the ladder where the higher number becomes possible. What it will not do, on its own: make you senior, replace experience, or guarantee a job in a soft market. No certificate on earth does that. The paper opens the door. Staying in the room is a different skill. ## The other way in: leading a team you did not come up through Not everyone who needs the CCNA is trying to become a network engineer. Some people need it to lead one. I work with someone who oversees operations for a network team. He did not take the usual route (help-desk, then network technician, then engineer). He came in from a different direction entirely and was handed responsibility for monitoring and coordinating the team's day-to-day work. If you have spent any time around government or large public-sector IT, you have seen this exact path: capable people land in oversight, coordination, or program-management roles over a technical team without ever having configured the gear themselves. For someone in that seat, the CCNA is not a resume filter. It is literacy. It is the difference between nodding along in a status meeting and actually understanding what the team is telling you: what the terms mean, why a proposed change is risky, whether a timeline is realistic, and which questions are the sharp ones to ask. You are never going to out-configure your engineers, and you do not need to. You need enough of the mental model to follow the work, pressure-test what you are hearing, and represent the team credibly to the people above you. That is exactly what the CCNA delivers: the vocabulary and the how-it-all-fits-together picture, without requiring you to live in the command line. And the hands-on part still earns its keep here. Building a small network once, and watching it break, teaches you more about what your team handles every day than any vendor slide deck will. You may never take the 2am call, but you will finally understand what that call was about, and that is what makes you someone the team actually wants in charge. ## Will AI take the job? You cannot talk about 2026 without this one, so here is the honest version, not the doom version and not the hype version. AI is genuinely changing network engineering. Tools can now generate configs, summarize logs, and flag anomalies that used to eat an afternoon. The repetitive parts of the job are getting automated, and that trend is real. But look at what that actually means for someone starting out. The job is shifting from operator to orchestrator (that is Cisco's phrase, and it is a good one). Less typing the same VLAN config by hand for the 400th time. More deciding what the network should do, supervising the automation that does it, and stepping in when it goes sideways. And you cannot supervise what you do not understand. It is the same point as the oversight role above, except the thing you are overseeing is now automation rather than people. When a tool tells you "this is a normal reconvergence, ignore it," someone has to know whether that is true, or whether it is actually a loop about to take the building down. That someone needs the fundamentals cold, which is exactly what the CCNA teaches. It is worth noticing that Cisco's own response to AI was to add AI to the CCNA, not to retire the cert. The fundamentals did not get less important. They became the thing that lets you use the new tools without being fooled by them. So AI is reshaping the job, not deleting it. The people who struggle will be the ones who only ever memorized. The ones who understand how it actually works will be fine, and arguably better off, because the boring half of their week just got automated away. ## The trap that wastes people's time Here is how the CCNA fails people, and it is almost always the same way: they treat it as a memorization exam. They grind a question dump, pass, and then walk into an interview unable to talk through a simple piece of output or explain why a routing adjacency is stuck. A hiring manager spots this in about five minutes. The question is never "what port does BGP use" (179, for the record). It is "look at this and tell me what is wrong": ``` R1# show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 10.0.0.2 1 FULL/DR 00:00:34 10.1.12.2 GigabitEthernet0/0 10.0.0.3 1 EXSTART/DROTHER 00:00:31 10.1.13.3 GigabitEthernet0/1 ``` If you can look at that and say "the second neighbor is stuck in EXSTART, which is almost always an MTU mismatch," you can do the job. If you only memorized that EXSTART exists but cannot say what to check next, you passed a test you cannot yet apply. (If that example is new to you, the full walk-through of [OSPF neighbor states and what each stuck state means](https://www.pinglabz.com/ospf-neighbor-states-explained/) is a good next read.) The fix is not complicated. Build the thing you are studying. Spin up a lab, configure a VLAN, break it on purpose, watch what the show commands tell you, then fix it. The cert proves you studied. Hands-on proves you can do the job. You want both, and the second one is what turns the paper into a paycheck. ## How to actually study for it You do not need an expensive home lab to get hands-on anymore. Cisco Modeling Labs has a [free tier](https://www.pinglabz.com/cisco-modeling-labs-cml-2-8-free/) that runs real Cisco IOS XE images, so the output you practice on behaves like the gear you will see on the job, not a simplified simulator. A workable loop looks like this: read a concept, then immediately build it, break it, and read the show output until it makes sense. When you can predict what a command will print before you press enter, you have actually learned it. That is the gap between someone who has only read and someone who has labbed, and it is the gap interviewers probe for. Several of the [PingLabz CCNA labs](https://www.pinglabz.com/labs/) are free with no signup (subnetting, VLANs and trunks, single-area OSPF, a standard ACL, and a troubleshooting ticket where you inherit a network that is already broken), and they are built on real captures for exactly this reason. Keep a [command cheat sheet](https://www.pinglabz.com/cisco-commands-cheat-sheet/) next to you while you work and the muscle memory comes faster. ## So who should actually get it? Get the CCNA if you are trying to break into networking and need to clear the resume filter, if you are in a help-desk or general IT role and want to move toward network engineering, if you are self-taught and want a structured way to fill gaps you did not know you had, if you manage or oversee a network team without an engineering background and want the literacy to lead it credibly, or if you plan to chase higher Cisco certs (the CCNA is the foundation the CCNP builds on, and it is where deeper topics like [how BGP works](https://www.pinglabz.com/how-bgp-works/) start to open up). Think harder if you are already a working network engineer with years of hands-on experience. For you the CCNA may be an HR box-tick rather than a learning experience, which is sometimes still worth it for the job filter, just know what you are buying. And do not bother if you are hoping a certificate alone will land a job with no practical skill behind it. It will not, in any market. ## Frequently asked questions ### Should I wait for the new 2027 exam or take the current one now? If you are close to ready, take the current 200-301 now. It is valid for three years and carries exactly the same weight. If your study will run into 2027 anyway, you will land on the new version naturally. Either way, do not stall: the fundamentals are the same on both. ### Is the CCNA still relevant with AI automating networks? Yes, arguably more so. Automation handles the repetitive work, but someone has to understand the network well enough to supervise it and catch when it is wrong. Cisco responded to AI by adding it to the CCNA blueprint, not by retiring the cert. ### How long does it take to study for the CCNA? For most people with some IT background, three to six months of consistent study plus hands-on practice. With no background, plan for longer. The variable that matters most is not hours read, it is hours spent actually building and breaking labs. ### Do I need a Cisco home lab to pass? No. The free tier of Cisco Modeling Labs runs real IOS XE and is enough to practice every topic on the exam. Real virtual gear beats a simplified simulator because you learn the failure modes you will actually hit on the job. ### Is the CCNA worth it if I manage a network team but do not configure anything myself? Yes, as literacy rather than a job filter. If you oversee or coordinate a technical team, which is common in government and large enterprises, the CCNA gives you enough of the mental model to follow the work, ask sharp questions, judge risk and timelines, and represent the team credibly. You will not out-configure your engineers, but you will finally understand them. ### Does the CCNA expire? It is valid for three years. You renew by passing a higher exam or through Cisco's continuing education credits before it lapses. ## Key takeaways The CCNA in 2026 is worth it for most people who want into networking, with one honest caveat repeated all the way through: it is the start, not the finish. The exam changes in early 2027, so study the fundamentals (which do not change) and take whichever version is live when you are ready. AI is reshaping the job, which makes the fundamentals more valuable, not less. And the certificate only pays off if you back it with hands-on practice that proves you can actually run a network, not just describe one. Get the cert. Then build something with it. That order is the whole secret, and you can start for free in the [PingLabz labs](https://www.pinglabz.com/labs/). If you decide to go for it, start with the right practice tools: [Top CCNA Network Simulators: Hands-On Review & Guide](https://www.pinglabz.com/the-best-ccna-network-simulators-my-hands-on-experience-and-complete-guide/). ### CCNA Flashcards & Key Terms - All-Access Pass URL: https://www.pinglabz.com/ccna-flashcards-pass/ Last updated: 2026-06-26T21:18:02.000Z Your unlock phrase for the PingLabz CCNA 200-301 flashcard app, free for members. Sign in or create a free account to reveal it and remove the 10-question limit. _This post is for subscribers only._ ### OSPF Link-State Advertisements: What an LSA Is URL: https://www.pinglabz.com/ospf-link-state-advertisement/ Last updated: 2026-06-26T13:00:00.000Z OSPF is a link-state protocol, and the "link state" it routes on is a database. Link-state advertisements, or LSAs, are the individual records in that database. Before you can make sense of LSA types, areas, or why a route is or is not in the table, you need a clear picture of what a single LSA actually is and how it gets everywhere it needs to be. This post covers exactly that - the LSA as a building block, not the type catalog. For the full picture, see the [OSPF complete guide](https://www.pinglabz.com/ospf/). ## What an LSA is An LSA is one record describing one piece of the network: a router and its links, a network segment and the routers on it, a route from another area, or an external route redistributed into OSPF. Each router builds LSAs describing the parts of the topology it knows directly, and floods them to its neighbors. The collected set of every LSA in an area is the **link-state database** (LSDB). The key property: every router in an area holds an *identical* copy of that area's LSDB. They all start from the same raw data. Each then runs the SPF (Dijkstra) algorithm against its own copy to compute its own shortest-path tree. Same database, different root, different tree - that is the whole idea of a link-state protocol. ## The LSA is not the route This distinction matters. An LSA is *topology data*. A route is the *result* of running SPF over that data. The LSDB can contain a perfectly valid LSA for a destination that still does not appear in your routing table, because SPF found no path to it, or a better non-OSPF route won. When you troubleshoot OSPF, "is the LSA in the database" and "is the route installed" are two separate questions, checked with two separate commands. ## The LSA header Every LSA, regardless of type, carries the same 20-byte header. These fields are what flooding and the database use to decide if two LSAs are the same record and which copy is newer. LS Age Seconds since the LSA was originated. Counts up. Caps at MaxAge (3600s). LS Type Type 1-7 - router, network, summary, ASBR-summary, external, NSSA-external. Link State ID Identifies what the LSA describes (a router ID, a network address, etc.). Advertising Router The router ID of whoever originated this LSA. Sequence Number Increments each time the LSA is re-originated. Higher wins. Checksum Integrity check over the LSA contents. The combination of **LS Type + Link State ID + Advertising Router** uniquely identifies an LSA. When a router receives an LSA matching one it already has, it compares sequence numbers: higher sequence is newer and replaces the old copy. Equal sequence means a duplicate, which is ignored. ## How LSAs flood OSPF flooding is reliable, which means acknowledged. When a router originates or receives a new LSA, it sends it to its neighbors in a Link State Update (LSU) packet. Each neighbor that accepts it returns a Link State Acknowledgment (LSAck). An unacknowledged LSA is retransmitted. This is why an LSDB stays consistent: nothing is "fire and forget." Flooding has scope. Most LSA types flood only within their area - that area boundary is the reason OSPF scales, because a topology change in one area does not force every router in every other area to re-run SPF. Type 5 external LSAs are the exception; they flood through the whole OSPF domain except into stub areas. ## LSA aging and refresh The LS Age field counts upward from 0\. Two thresholds matter: - **Refresh at 1800 seconds.** The router that originated an LSA re-floods it every 30 minutes with a fresh age of 0 and an incremented sequence number, so a still-valid LSA never actually expires. - **MaxAge at 3600 seconds.** If an LSA ever reaches 3600s without a refresh, it is considered dead. The router floods it one last time at MaxAge to tell everyone to remove it, and every router purges it from the LSDB and re-runs SPF. Flooding an LSA to MaxAge is also how OSPF explicitly withdraws a route: to retract an LSA, the originator re-floods it set to MaxAge. ## The LSA types, briefly There are six common LSA types, each describing a different scope of the network - router links, multi-access segments, inter-area routes, the ASBR location, external routes, and the NSSA variant of externals. Each type has its own flooding scope and its own role in SPF. That catalog is a topic of its own; see the dedicated [OSPF LSA types](https://www.pinglabz.com/ospf-lsa-types-explained/) article for the full type-by-type breakdown. ## Reading the database ``` R2# show ip ospf database OSPF Router with ID (10.255.0.2) (Process ID 1) Router Link States (Area 0) Link ID ADV Router Age Seq# Checksum Link count 10.255.0.1 10.255.0.1 221 0x80000007 0x00A1B2 3 10.255.0.2 10.255.0.2 198 0x80000009 0x00C3D4 4 10.255.0.3 10.255.0.3 210 0x8000000A 0x00E5F6 2 ``` Each row is one LSA. **ADV Router** is who originated it, **Age** is the LS Age counting up toward 3600, and **Seq#** is the sequence number - higher is newer. To see the full contents of a single LSA rather than the summary list, use `show ip ospf database router 10.255.0.1`. ## Common gotchas LSA is in the database but the route is not in the routing table The LSDB is topology data; SPF still has to find a path. Check next-hop reachability and competing routes. LSDBs differ between two routers in the same area An adjacency is not Full, so flooding never completed between them. Check neighbor state first. An old route will not disappear The withdrawing router must flood the LSA at MaxAge. If it cannot reach part of the area, the purge does not propagate. LSA age sitting near 3600 The originator stopped refreshing - likely down or partitioned. The LSA is about to be purged. ## Key takeaways An LSA is a single record of topology - one router, one segment, one external route. The full set of LSAs in an area is the link-state database, and every router in the area holds an identical copy, then runs SPF on it independently. An LSA is data, not a route; a route is what SPF produces. Every LSA shares a header whose Type, Link State ID, and Advertising Router identify it, and whose sequence number decides which copy is newer. LSAs flood reliably with acknowledgments, refresh every 30 minutes, and die at a MaxAge of 3600 seconds, which is also the mechanism OSPF uses to withdraw a route. For the OSPF cluster, see the [OSPF pillar](https://www.pinglabz.com/ospf/). ### BGP Configuration on Cisco IOS XE: eBGP and iBGP URL: https://www.pinglabz.com/bgp-configuration/ Last updated: 2026-06-24T13:00:00.000Z BGP configuration is mechanically simple. The command set is small, and a basic peering comes up in half a dozen lines. What trips people up is not the syntax, it is the handful of non-obvious requirements: eBGP and iBGP behave differently, getting a route *into* BGP is a separate step from forming a neighbor, and iBGP has a next-hop quirk that silently breaks reachability. This post is a configuration walkthrough on Cisco IOS XE that covers all of it. For the full picture, see the [BGP complete guide](https://www.pinglabz.com/bgp/). ## What you are actually configuring A working BGP setup has three parts, and it helps to keep them mentally separate: - **The local AS** \- the autonomous system number this router belongs to. - **The neighbors** \- each BGP peer is configured by hand with its IP address and its AS number. BGP does not discover peers; you tell it about every one. - **The routes to advertise** \- which prefixes this router originates into BGP. Forming a neighbor and advertising a route are two unrelated steps. ## eBGP vs iBGP at configuration time The single command `neighbor x.x.x.x remote-as N` decides everything. If the remote AS is different from your local AS, that session is **eBGP**. If it matches your local AS, it is **iBGP**. You never type "ebgp" or "ibgp" anywhere; the AS numbers decide. Remote-as eBGP Different from local AS iBGPSame as local AS Typical peering address eBGP Directly connected interface iBGP Loopback (for resilience) TTL default eBGP 1 (peers must be adjacent) iBGP255 Next-hop on advertised routes eBGPChanged to self iBGP Left unchanged - the gotcha ## Minimum viable eBGP Two routers, two different ASes, peering over a directly connected link. R1 is in AS 65001, R2 is in AS 65002, and the link between them is 192.0.2.0/30. ``` ! R1 - AS 65001 router bgp 65001 bgp router-id 10.255.0.1 neighbor 192.0.2.2 remote-as 65002 network 10.20.0.0 mask 255.255.255.0 ``` That is a complete, working eBGP speaker. The `network` statement is what originates 10.20.0.0/24 into BGP. Mirror the config on R2 with its own AS and prefixes, and the session comes up. ## The network statement has an exact-match rule The `network` command does not "advertise whatever is in that range." It only originates a prefix if an **exact match** for it already exists in the routing table - same network, same mask. If you write `network 10.20.0.0 mask 255.255.255.0` but the router only has 10.20.0.0/16 in its table, nothing is advertised. This is the most common reason a brand-new BGP config forms a neighbor but advertises no routes. ## iBGP and the next-hop trap iBGP exists to carry external routes *across* your AS. The catch: when a router passes a route to an iBGP peer, it does **not** change the next-hop attribute. So an internal router learns a prefix whose next-hop is an address in some other AS - an address it has no route to. The prefix shows up in the BGP table but never makes it into the routing table, because the next-hop is unreachable. The fix is `next-hop-self`, applied on the router that has the eBGP session, for its iBGP neighbors: ``` ! R2 - AS 65001, has the eBGP session AND an iBGP peer R3 router bgp 65001 neighbor 192.0.2.1 remote-as 65002 neighbor 10.255.0.3 remote-as 65001 neighbor 10.255.0.3 update-source Loopback0 neighbor 10.255.0.3 next-hop-self ``` Now R3 receives external routes with R2's loopback as the next-hop - an address R3 can reach via the IGP. Note `update-source Loopback0`: iBGP sessions are peered loopback-to-loopback so the session survives any single link failure, and both sides must agree on the source address. ## Getting routes into BGP: network vs redistribution network statement How Explicitly name each prefix; needs an exact route-table match When to use A small, known set of prefixes you deliberately advertise redistribution How `redistribute ospf 1`, `redistribute connected`, etc. When to use Many prefixes from an IGP - but filter carefully, it is easy to leak For most edge routers the `network` statement is the safer choice: it is explicit, and you cannot accidentally advertise the entire internal routing table to a provider. Redistribution into BGP without a route-map filter is a classic way to cause an outage. ## Verifying ``` R1# show ip bgp summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 192.0.2.2 4 65002 21 23 12 0 0 00:14:22 3 ``` Read the rightmost column. A number (here **3**) means the session is **Established** and you have received 3 prefixes. A word - `Idle`, `Active`, `Connect`, `OpenSent` \- means the session is not up yet. `Active` is the one that misleads people: it does not mean "working," it means the router is actively trying and failing to connect. ``` R1# show ip bgp Network Next Hop Metric LocPrf Weight Path *> 10.20.0.0/24 0.0.0.0 0 32768 i *> 10.30.0.0/24 192.0.2.2 0 0 65002 i ``` The `*>` means a route is valid and the best path. `0.0.0.0` as the next-hop marks a prefix this router originated itself. ## Common gotchas Neighbor stuck in Active or Idle No IP reachability to the peer address, wrong remote-as, or an ACL blocking TCP 179. Neighbor is up but advertises nothing The network statement has no exact-match route in the table, or a prefix-list/route-map is filtering outbound. iBGP route in BGP table but not the routing table Unreachable next-hop. Apply `next-hop-self` on the router with the eBGP session. iBGP session flaps when a link fails Peering on a physical interface instead of a loopback. Peer loopback-to-loopback with `update-source`. eBGP neighbor will not come up across a router in between eBGP TTL is 1 by default. Multi-hop eBGP needs `neighbor x.x.x.x ebgp-multihop`. ## Key takeaways BGP configuration is three separate jobs: set the local AS, define each neighbor by IP and remote-as, and originate prefixes with `network` or filtered redistribution. The remote-as value alone decides eBGP versus iBGP. The `network` statement needs an exact route-table match or it advertises nothing. iBGP does not rewrite the next-hop, so `next-hop-self` on the eBGP-facing router is almost always required, and iBGP sessions should peer loopback-to-loopback with `update-source`. When you verify, read the State/PfxRcd column of `show ip bgp summary` \- a number is good, a word is not. For the BGP cluster, see the [BGP pillar](https://www.pinglabz.com/bgp/). ### Cisco Wireless Access Point Modes Explained URL: https://www.pinglabz.com/wireless-access-point-modes/ Last updated: 2026-06-22T13:00:00.000Z A Cisco access point is not always a thing that serves Wi-Fi. The same hardware can be a client-serving AP, a wireless bridge, a full-time spectrum analyzer, or a packet-capture probe, depending on which mode it is in. The AP mode is one setting, and picking the wrong one is a common reason an AP "is not working" when it is in fact working perfectly - just not as the thing you expected. This post walks through every Cisco AP mode, what each is for, and when you would actually use it. For the cluster overview, see the [Cisco Wireless complete guide](https://www.pinglabz.com/wireless/). ## The two that matter most Ninety-plus percent of access points in production run one of two modes. ### Local mode The default. The AP serves wireless clients and tunnels their traffic back to the wireless LAN controller over CAPWAP. The controller puts the client traffic onto the wired network. Local mode is the standard campus deployment: the AP is a radio head, the controller is the brain and the traffic aggregation point. In local mode the AP also spends a small slice of time off-channel scanning other channels, which is how the controller's RF management and rogue detection get their data without dedicating hardware to it. ### FlexConnect mode FlexConnect is the branch-office mode. The AP still gets its configuration and policy from a central controller, but client traffic is switched *locally* at the branch instead of being tunneled all the way back to the controller. This matters for two reasons. First, traffic does not hairpin: a branch user printing to a branch printer does not send packets across the WAN to a distant controller and back. Second, and more importantly, FlexConnect APs keep serving clients even if the WAN link to the controller goes down. A local-mode AP that loses its controller eventually stops serving clients; a FlexConnect AP rides through the outage. For any site at the far end of a WAN link, FlexConnect is the right mode. ## The full mode list Local What the AP does Standard client-serving AP; traffic tunneled to the WLC Serves clients?Yes FlexConnect What the AP does Client-serving AP; traffic switched locally; survives WAN outage Serves clients?Yes Monitor What the AP does Dedicated full-time scanner - no client radio. IDS, rogue detection, location services. Serves clients?No Sniffer What the AP does Captures 802.11 frames and forwards them to a packet analyzer (Wireshark, OmniPeek) Serves clients?No Rogue Detector What the AP does Listens on the wired side, correlating MACs to spot rogue APs plugged into the network Serves clients?No SE-Connect (Spectrum) What the AP does Dedicated spectrum analyzer - studies RF interference, including non-Wi-Fi sources Serves clients?No Bridge / Mesh What the AP does Point-to-point or point-to-multipoint wireless bridge between sites Serves clients? Via mesh, not as a normal AP Flex+Bridge What the AP does Mesh AP that also does FlexConnect local switching Serves clients?Yes (mesh + local) ## Monitor mode A monitor-mode AP gives up serving clients entirely and spends 100% of its time scanning all channels. A local-mode AP only scans off-channel occasionally; a monitor-mode AP scans constantly. You use monitor mode where you need thorough, continuous wireless visibility: wireless intrusion detection (wIDS/wIPS), aggressive rogue-AP detection, and location services that triangulate device position from signal strength. The trade-off is obvious - that AP serves zero clients. Monitor-mode APs are typically a sprinkling of extra units deployed specifically for visibility, not your client-serving fleet. ## Sniffer mode Sniffer mode turns the AP into a remote wireless capture probe. It captures raw 802.11 frames on a chosen channel and encapsulates them to a destination running a protocol analyzer. This is how you capture over-the-air wireless traffic properly - including management and control frames, retries, and the things a normal client NIC will not show you. It is a troubleshooting tool, set temporarily on an AP near the problem, then set back. The captured frames are exactly what you need to diagnose roaming failures, authentication problems, or interference at the 802.11 level. ## Rogue Detector mode Rogue Detector mode is the odd one - the AP's radios are essentially off, and it works on the *wired* side. It listens to ARP traffic on the wired network and correlates the MAC addresses it sees against the list of rogue clients and APs the controller has detected over the air. The purpose is to answer a specific question: is that rogue AP actually plugged into *my* network (a real security incident), or is it just a neighbor's AP bleeding RF into my building (noise, not a threat)? If a MAC seen over the air also shows up on the wired side, the rogue is connected to your network. This mode has become less common as controller-side rogue-on-wire detection improved, but it still appears. ## SE-Connect (Spectrum) mode SE-Connect dedicates the AP to spectrum analysis. Where monitor mode scans for Wi-Fi, SE-Connect studies the raw RF spectrum - including non-Wi-Fi interference like microwave ovens, Bluetooth, video bridges, and cordless phones. You connect a spectrum-analysis tool to an SE-Connect AP when you have a performance problem that Wi-Fi-only tools cannot explain - throughput that collapses at certain times of day, a dead zone with no obvious cause. The AP becomes a sensor that shows you the interference a packet capture would never reveal. ## Bridge and Mesh modes Bridge mode turns APs into a wireless link between locations - point-to-point to connect two buildings without running fiber, or point-to-multipoint for a hub-and-spoke layout. Mesh extends this so APs relay traffic wirelessly through each other, useful where running cable to every AP is impractical (outdoor coverage, warehouses, historic buildings). **Flex+Bridge** combines mesh backhaul with FlexConnect local switching - a mesh AP that also serves clients and switches their traffic locally. It is the mode for a meshed branch or outdoor deployment that still needs the branch-survivability behavior. ## Changing the mode AP mode is set from the controller, per AP. Changing it almost always reboots the AP, because the radios are being repurposed. The practical consequence: do not change an AP's mode during business hours if it is currently serving clients - the change will drop everyone associated to it. ## Common gotchas "AP is not serving any clients" It is in Monitor, Sniffer, Rogue Detector, or SE-Connect mode. Those modes intentionally serve zero clients. Branch clients drop when the WAN goes down APs are in Local mode, tunneling to a central WLC. Branches should be FlexConnect. An AP rebooted "for no reason" Its mode was changed from the controller - a mode change triggers a reboot. Rogue detection seems weak Local-mode APs only scan off-channel part-time. Add Monitor-mode APs for thorough coverage. Cannot explain a throughput problem with packet captures The interference may be non-Wi-Fi. Use an SE-Connect AP to see the raw spectrum. ## Key takeaways Cisco AP modes decide what the hardware actually does. Local and FlexConnect are the client-serving modes - Local tunnels traffic to the controller, FlexConnect switches it locally and survives WAN outages, which makes FlexConnect the branch default. Monitor, Sniffer, Rogue Detector, and SE-Connect are all non-client modes for visibility, capture, and RF analysis. Bridge, Mesh, and Flex+Bridge connect sites wirelessly. When an AP "is not working," confirm its mode first - it may be doing exactly what its mode tells it to, which just is not serving Wi-Fi. For the wireless cluster, see the [Cisco Wireless pillar](https://www.pinglabz.com/wireless/). ### Segment Routing (SR-MPLS): Replacing LDP and RSVP-TE URL: https://www.pinglabz.com/segment-routing-mpls/ Last updated: 2026-08-01T19:35:36.000Z Segment Routing is the technology quietly replacing LDP and RSVP-TE in modern service-provider and large-enterprise cores. The pitch is simple: take the MPLS data plane that already works well, and throw away the separate label-distribution protocols that made it complicated. SR-MPLS keeps the labels and drops the protocols that used to hand them out. This post explains what segment routing is, how SR-MPLS works, and why networks are migrating to it. For the cluster overview, see the [MPLS complete guide](https://www.pinglabz.com/mpls/). ## The problem with traditional MPLS Classic MPLS works, but it needs a lot of moving parts. To forward a labeled packet, every router in the path needs a label for every destination, and those labels have to be distributed somehow. Traditional MPLS used dedicated protocols for that: - **LDP** distributed labels for basic hop-by-hop MPLS forwarding. - **RSVP-TE** distributed labels and held state for traffic-engineered paths. Both run alongside the IGP (OSPF or IS-IS). So a traditional MPLS core runs an IGP *and* LDP *and*, if you do traffic engineering, RSVP-TE. RSVP-TE is the worst offender: it builds per-tunnel state on every router along every path. In a large network with many TE tunnels, that state explodes, and it does not scale gracefully. The other cost of that extra protocol is diagnostic. An LDP failure leaves the IGP fully converged and every loopback pingable while customer traffic goes nowhere, which is why [a dead label plane under a healthy control plane](https://www.pinglabz.com/troubleshooting-mpls-ldp/) is a troubleshooting discipline of its own - and removing LDP removes that whole failure class. ## The segment routing idea Segment Routing makes one clean observation: the IGP already floods information about the whole topology to every router. So why run a separate protocol just to distribute labels? Let the IGP do it. In SR-MPLS, OSPF or IS-IS is extended to also advertise **segments** \- and a segment is just a label with a meaning. The IGP that already knows the topology now also distributes the labels. LDP is no longer needed. RSVP-TE is no longer needed. One protocol (the IGP) does the job that previously took three. The MPLS data plane does not change at all. Packets still carry MPLS labels, still get label-swapped hop by hop, still use the same hardware forwarding path. SR-MPLS is a control-plane simplification riding on the unchanged MPLS data plane. ## The two segment types Prefix-SID (node SID) What it represents "Get to this router/prefix by the shortest IGP path" Scope Global - every router agrees on the same label for it Adjacency-SID What it represents "Use this one specific link / next hop" Scope Local - meaningful only on the router that advertised it A **prefix-SID** is a globally significant label assigned to a destination (usually a router's loopback). Every router in the domain installs the same prefix-SID for that destination and forwards toward it by the shortest path. It is allocated from a shared range called the SRGB (Segment Routing Global Block), which is why the value is consistent network-wide. An **adjacency-SID** is a locally significant label for one specific link. It says "send this out exactly this interface" regardless of what the shortest path would do. ## Source routing: the path lives in the packet Here is the part that makes traffic engineering elegant. To steer a packet along a specific path, the ingress router pushes a *stack* of segment labels onto it. Each label is a waypoint. The packet itself carries its own itinerary. Want a packet to go through router B, then over a specific link, then to router E? The ingress router pushes the label stack \[prefix-SID of B, adjacency-SID of that link, prefix-SID of E\]. Each router along the way pops the top segment when it is reached and forwards toward the next. No router in the middle holds any per-flow or per-tunnel state. The path is encoded in the packet, not stored in the network. This is the death of RSVP-TE state. Traffic engineering in SR is just "push the right label stack at the edge." The core routers stay stateless. That is the scaling win. ## Configuration shape on Cisco IOS XE Enabling SR-MPLS is mostly turning it on under the IGP and defining the global label block. The shape, with IS-IS: ``` segment-routing mpls global-block 16000 23999 ! router isis CORE segment-routing mpls ! interface Loopback0 ip address 10.255.0.1 255.255.255.255 isis prefix-sid index 1 ``` The SRGB here is 16000-23999\. The loopback gets prefix-SID index 1, so its global label is 16000 + 1 = 16001, and every router in the domain will use 16001 to reach this router. No LDP anywhere in the config - the IGP carries the labels. ## SR-MPLS vs SRv6 Segment Routing comes in two data-plane flavors, and it is worth knowing the difference: Data plane SR-MPLSMPLS labels SRv6 IPv6 addresses (segments are IPv6 SIDs) Requires SR-MPLS Existing MPLS-capable hardware SRv6 IPv6 end to end; no MPLS needed Migration path SR-MPLS Drop-in for existing MPLS networks SRv6 Larger shift; appeals to greenfield / IPv6-native designs SR-MPLS is the pragmatic upgrade for any network that already runs MPLS - same data plane, simpler control plane. SRv6 is the more ambitious version that drops MPLS entirely and encodes segments directly in IPv6 addresses. Most MPLS networks migrating today go SR-MPLS first. ## Why networks migrate to SR-MPLS Fewer protocols IGP only. LDP and RSVP-TE are removed - fewer things to run, monitor, and break. No core TE state The path is in the packet. Core routers hold no per-tunnel state, so traffic engineering scales. Better fast-reroute Topology-Independent LFA (TI-LFA) gives sub-50ms protection with guaranteed coverage - cleaner than the LDP/RSVP equivalents. SDN-friendly A central controller can compute a path and just tell the edge router the label stack to push. SR is the natural data plane for SDN traffic engineering. Unchanged data plane Existing MPLS hardware forwards SR-MPLS with no change. The migration is control-plane only. ## Migration reality SR-MPLS and LDP can run at the same time during a transition, so migration is incremental rather than a flag day. A common path: enable SR under the IGP across the core, let SR and LDP coexist with a defined preference, move services onto SR, then decommission LDP once nothing depends on it. RSVP-TE tunnels get rebuilt as SR policies. None of this requires touching the data-plane hardware. ## Key takeaways Segment Routing keeps the MPLS data plane and removes the separate label-distribution protocols. In SR-MPLS, the IGP (OSPF or IS-IS) is extended to distribute segments - labels with meaning - so LDP and RSVP-TE are no longer needed. Prefix-SIDs are global "shortest path to here" labels; adjacency-SIDs are local "use this link" labels. Traffic engineering becomes a label stack pushed at the edge, leaving core routers stateless, which is the headline scaling win over RSVP-TE. SR-MPLS is the low-friction upgrade for existing MPLS networks; SRv6 is the more radical IPv6-native variant. The migration is control-plane only and can run alongside LDP during the transition. For the MPLS cluster, see the [MPLS pillar](https://www.pinglabz.com/mpls/). ### VTP (VLAN Trunking Protocol): v1, v2, v3 and Why It's Dangerous URL: https://www.pinglabz.com/vlan-trunking-protocol/ Last updated: 2026-06-17T13:00:00.000Z VLAN Trunking Protocol does one job: it propagates VLAN definitions from one switch to the rest of the switches in a domain, so you do not have to create the same VLAN by hand on twenty boxes. That sounds purely helpful, and in small doses it is. VTP also has a famous failure mode where a single mis-handled switch wipes the VLAN database across an entire network. This post explains how VTP works, the differences between v1, v2, and v3, and why a lot of engineers now deliberately turn it off. For the cluster overview, see the [VLAN and Layer 2 switching pillar](https://www.pinglabz.com/vlans-layer-2-switching/). ## What VTP is for Without VTP, every VLAN exists only on the switch where you created it. A 30-switch campus with 40 VLANs means creating 40 VLANs, 30 times. VTP lets you create a VLAN once, on one switch, and have that definition replicate automatically to every other switch in the same VTP domain over trunk links. It is genuinely useful for keeping VLAN *definitions* consistent. What it does not do, and what people wrongly expect it to do, is decide which VLANs are *allowed* on which trunks - that is the trunk allowed-list, a separate thing entirely. ## The three VTP modes Server Can create/modify/delete VLANs?Yes Propagates VTP info?Yes Stores the VLAN database?Yes Client Can create/modify/delete VLANs?No Propagates VTP info?Yes (forwards) Stores the VLAN database?Yes (in RAM) Transparent Can create/modify/delete VLANs?Yes (locally only) Propagates VTP info? Forwards others' messages but ignores them Stores the VLAN database?Yes (local only) A **server** can edit VLANs and pushes changes to the domain. A **client** cannot edit VLANs - it only accepts what servers send. A **transparent** switch is opted out: it keeps its own local VLAN database, ignores VTP advertisements from others, but politely forwards them along so it does not break the chain for downstream switches. ## The revision number: the dangerous part Every VTP server and client tracks a configuration revision number. Every time a server makes a VLAN change, it increments its revision number and advertises the new database. Switches receiving an advertisement compare revision numbers: **if the incoming revision is higher than what they currently have, they overwrite their entire VLAN database with the incoming one.** Read that again, because it is the whole problem. The decision is purely "is this number bigger?" It does not matter whether the advertising switch is your core or a switch a contractor pulled off a shelf. A higher revision number wins. The classic disaster: a switch is removed from the network, used in a lab where lots of VLAN changes pump its revision number sky-high, then plugged back into production - still in server or client mode, still in the same VTP domain. Its revision number is now higher than the production core's. Every switch in the domain accepts its database. If that lab switch's database had three VLANs, the production network now has three VLANs. Hundreds of access ports go dark because their VLANs no longer exist. ## How to avoid the disaster Any switch entering or re-entering a VTP domain must have its revision number reset to 0 first. Two reliable ways to force a reset: - Change the switch to VTP **transparent** mode and back - transparent mode resets the revision to 0. - Change the VTP **domain name** to something else and back - a domain change also resets the revision. This is mandatory hygiene. Never plug a switch with unknown VTP history into a production domain without first zeroing its revision. ## VTP versions: v1, v2, v3 v1 The original. Propagates standard-range VLANs (1-1005) only. v2 Minor additions: Token Ring support, better handling of unrecognized TLVs, transparent-mode consistency checks. For Ethernet networks v1 and v2 are nearly identical in practice. v3 The meaningful upgrade. Extended-range VLANs (1-4094), a primary-server model, a password to authenticate updates, and it can propagate more than just VLANs (MST mapping, private VLANs). VTP v3 is the version that addresses the revision-number disaster head-on. In v3 there is exactly one **primary server** for the domain - and only the primary can make changes that propagate. A switch is not primary just because it is in server mode; it has to be explicitly promoted with a command, and that promotion is what gives it write authority. A rogue switch with a high revision number cannot overwrite anything in v3 unless it has been deliberately made primary. If you must run VTP, run v3. ## Configuration on Cisco IOS XE ``` ! Set the domain, version, mode, and a password vtp domain PINGLABZ vtp version 3 vtp mode server vtp password STRONG_SECRET ``` In v3, after setting the switch to server mode you still must explicitly make it the primary before it can push changes: ``` SW1# vtp primary vlan This system is becoming primary server for feature vlan ... Do you want to continue? [confirm] ``` ## Verifying ``` SW1# show vtp status VTP Version capable : 1 to 3 VTP version running : 3 VTP Domain Name : PINGLABZ VTP Pruning Mode : Disabled VTP Traps Generation : Disabled Device ID : 0050.7989.aaaa Feature VLAN: -------------- VTP Operating Mode : Primary Server Number of existing VLANs : 12 Configuration Revision : 7 Primary ID : 0050.7989.aaaa Primary Description : SW1 ``` The two lines to read: **Configuration Revision** (the number that drives overwrites - know it before adding any switch) and **VTP Operating Mode** (Primary Server, Server, Client, or Transparent). ## Why many engineers just turn VTP off A large number of modern networks run every switch in VTP **transparent** mode, or VTP off entirely, and create VLANs manually (or via automation - Ansible, scripts, controller-based provisioning). The reasoning: the downside of VTP (one bad switch erasing the VLAN database network-wide) is catastrophic, while the upside (not typing VLAN definitions a few extra times) is modest and is better solved by configuration automation anyway. Automation gives you the same consistency with version control, change review, and no revision-number landmine. If you do run VTP, v3 with its primary-server model is the only version that should be on the table. ## Common gotchas Whole network's VLANs vanished after adding a switch The new switch had a higher revision number and overwrote the domain. Always reset revision to 0 before connecting. New VLAN created on a switch but not appearing elsewhere That switch is in client or transparent mode. Only servers (or the v3 primary) originate changes. Switch ignores VTP updates entirely It is in transparent mode, or the domain name / password / version does not match. Cannot create VLAN 2000+ Extended-range VLANs need VTP v3 (or transparent mode on v1/v2). v3 server cannot make changes It is a secondary server. Only the primary can. Promote it with `vtp primary`. ## Key takeaways VTP propagates VLAN definitions across a domain so you create each VLAN once. It has three modes - server, client, transparent - and a revision number that decides whose database wins: higher revision overwrites everyone. That mechanism is the source of VTP's signature disaster, where a stale switch with an inflated revision number wipes the VLAN database network-wide. Always reset a switch's revision to 0 before connecting it. Run v3 if you run VTP at all - its primary-server model removes the rogue-overwrite risk. Many modern networks skip VTP entirely and provision VLANs through automation, which is a defensible and increasingly common choice. For the L2 cluster, see the [VLAN pillar](https://www.pinglabz.com/vlans-layer-2-switching/). ### OSPF Passive Interface: What It Does and Where to Use It URL: https://www.pinglabz.com/ospf-passive-interface/ Last updated: 2026-07-04T23:21:42.000Z The passive-interface command is one of the smallest pieces of OSPF configuration and one of the most consistently misunderstood. It does not stop OSPF from advertising a network. It does not remove an interface from OSPF. It does exactly one thing: it stops OSPF from sending Hello packets out the interface, which stops adjacencies from forming there. This post explains what passive-interface actually does, where it belongs, and the gotcha that catches people who assume it works like a filter. For the cluster overview, see the [OSPF complete guide](https://www.pinglabz.com/ospf/). For a routing protocol that handles the same idea slightly differently, see the [EIGRP pillar](https://www.pinglabz.com/eigrp/). ## What passive-interface actually does When an interface is part of the OSPF process, two things happen on it: OSPF advertises the interface's subnet into the link-state database, and OSPF sends Hellos out the interface to discover neighbors. Passive-interface keeps the first and kills the second: Subnet advertised into OSPF Normal OSPF interfaceYes Passive OSPF interfaceYes Hellos sent out the interface Normal OSPF interfaceYes Passive OSPF interfaceNo Neighbor adjacency can form Normal OSPF interfaceYes Passive OSPF interfaceNo LSAs / routes received on the interface Normal OSPF interfaceYes Passive OSPF interfaceNo The subnet still appears in OSPF. Other routers still learn how to reach it. But no router will ever become an OSPF neighbor across that interface, because the conversation that builds an adjacency - the Hello exchange - never starts. ## Why you want this The classic case is a LAN interface facing end hosts. Consider a router with an interface on the user VLAN, 10.20.0.0/24\. You want OSPF to advertise 10.20.0.0/24 so the rest of the network can reach those users. You absolutely do not want OSPF sending Hellos onto the user VLAN, because: - There are no other routers there to form adjacencies with, so the Hellos are pure waste. - Hellos on a user-facing segment are an attack surface. A malicious host could speak OSPF and inject routes. Silencing OSPF on host-facing interfaces removes that risk entirely. So the rule of thumb: **any interface that faces hosts rather than routers should be passive.** You still advertise the subnet; you just stop talking OSPF where no router is listening. ## When to Use Passive Interfaces A good rule of thumb: if no OSPF router will ever sit on the other end of that link, make it passive. That covers most of your network. User VLANs (workstations, phones) Passive?Yes Reason No OSPF routers; prevents rogue adjacencies Server segments Passive?Yes Reason Servers don't run OSPF; eliminates unnecessary traffic Management networks (OOB) Passive?Yes Reason Security boundary; management traffic is separate Loopback interfaces Passive?Yes Reason Logical interface - no physical neighbor possible WAN links to non-OSPF sites Passive?Yes Reason Remote end not running OSPF Router-to-router uplinks Passive?No Reason Neighbors must form here - this is where OSPF works Core/distribution links Passive?No Reason Adjacency required for topology exchange ## Configuration: per-interface and default Per-interface, the direct way: ``` router ospf 1 passive-interface GigabitEthernet0/1 ``` On a router with many host-facing interfaces and a few router-facing ones, the cleaner pattern is to make every interface passive by default and then explicitly un-passive the ones that need adjacencies: ``` router ospf 1 passive-interface default no passive-interface GigabitEthernet0/0 no passive-interface GigabitEthernet0/2 ``` This is the safer default-deny posture. Every new interface added later is passive automatically, and you have to consciously enable OSPF Hellos on a link before it can form an adjacency. Forgetting to un-passive a link is a missing-adjacency bug that is easy to spot; forgetting to passive a host interface is a quiet security gap that is not. Default-passive makes the safe choice the automatic one. One useful detail: with `passive-interface default`, loopback interfaces are automatically passive - you do not need to call them out explicitly (they cannot form adjacencies anyway, but this keeps the Hello suppression consistent in `show ip protocols`). ## Lab example: a branch router done right A quick end-to-end example. A branch router has one uplink to HQ (Gi0/0), two user segments that must stay reachable over OSPF but never form adjacencies (Gi0/1, Gi0/2), and a loopback for the router ID. ![OSPF Hello suppression - passive interface allows route advertisement but blocks adjacency formation](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/04/sequence.png) OSPF Hello suppression behavior - passive interface allows route advertisement but blocks adjacency formation ## When to Use Passive Interfaces A good rule of thumb: if no OSPF router will ever sit on the other end of that link, make it passive. That covers most of your network. User VLANs (workstations, phones) Passive?Yes Reason No OSPF routers; prevents rogue adjacencies Server segments Passive?Yes Reason Servers don't run OSPF; eliminates unnecessary traffic Management networks (OOB) Passive?Yes Reason Security boundary; management traffic is separate Loopback interfaces Passive?Yes Reason Logical interface - no physical neighbor possible WAN links to non-OSPF sites Passive?Yes Reason Remote end not running OSPF Router-to-router uplinks Passive?No Reason Neighbors must form here - this is where OSPF works Core/distribution links Passive?No Reason Adjacency required for topology exchange ## Configuration ### Method 1: Per-Interface Specify each passive interface explicitly. This is straightforward and keeps things visible, but gets tedious on routers with many passive interfaces. ``` Router(config)# router ospf 1 Router(config-router)# network 10.0.0.0 0.0.0.255 area 0 Router(config-router)# network 192.168.10.0 0.0.0.255 area 0 Router(config-router)# network 192.168.20.0 0.0.0.255 area 0 Router(config-router)# network 10.255.0.1 0.0.0.0 area 0 Router(config-router)# passive-interface GigabitEthernet0/1 Router(config-router)# passive-interface GigabitEthernet0/2 Router(config-router)# passive-interface Loopback0 ``` ### Method 2: Default Passive (Recommended for Edge Routers) This approach flips the logic - make everything passive by default, then explicitly re-enable OSPF on only the interfaces that need to form neighbors. On a branch router with one or two uplinks and a dozen user VLANs, this is much cleaner. ``` Router(config)# router ospf 1 Router(config-router)# network 0.0.0.0 255.255.255.255 area 0 Router(config-router)# passive-interface default Router(config-router)# no passive-interface GigabitEthernet0/0 ``` The `network 0.0.0.0 255.255.255.255` statement matches all interfaces (wildcard mask covers everything), while `passive-interface default` suppresses Hellos on all of them. The `no passive-interface Gi0/0` then carves out the uplink that actually needs to form a neighbor. One important note: Loopback interfaces are automatically passive when you use `passive-interface default` \- you don't need to call them out explicitly. ## Lab Example Here's the scenario we'll use: a branch router with one uplink to HQ and two user segments that need to be reachable over OSPF but must never form adjacencies. ![Branch router topology - only Gi0/0 forms an OSPF adjacency; user VLANs and loopback are passive](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/04/topology.png) Branch router topology - only Gi0/0 forms an OSPF adjacency; user VLANs and loopback are passive Using the default-passive pattern: ``` Branch(config)# router ospf 1 Branch(config-router)# router-id 10.255.0.1 Branch(config-router)# network 0.0.0.0 255.255.255.255 area 0 Branch(config-router)# passive-interface default Branch(config-router)# no passive-interface GigabitEthernet0/0 ``` All four networks are advertised into area 0, but only Gi0/0 sends and accepts Hellos. `show ip protocols` confirms the passive list: ``` Branch# show ip protocols Routing Protocol is "ospf 1" Router ID 10.255.0.1 Routing for Networks: 0.0.0.0 255.255.255.255 area 0 Passive Interface(s): GigabitEthernet0/1 GigabitEthernet0/2 Loopback0 Routing Information Sources: Gateway Distance Last Update 10.0.0.1 110 00:03:17 ``` Gi0/0 is conspicuously absent from the passive list - it is the only interface allowed to form neighbors. On HQ, both user VLANs and the loopback still appear as OSPF routes: ``` HQ# show ip route ospf 10.0.0.0/8 is variably subnetted, 3 subnets, 2 masks O 10.255.0.1/32 [110/2] via 10.0.0.2, 00:04:21, GigabitEthernet0/0 192.168.0.0/24 is subnetted, 2 subnets O 192.168.10.0 [110/2] via 10.0.0.2, 00:04:21, GigabitEthernet0/0 O 192.168.20.0 [110/2] via 10.0.0.2, 00:04:21, GigabitEthernet0/0 ``` Reachability intact, Hellos silenced. That is the whole point. ## The gotcha: passive-interface is not a route filter Here is where people go wrong. Passive-interface sounds like it might stop a network from being advertised. It does not. The subnet of a passive interface is still injected into OSPF and still reachable from everywhere else. If your actual goal is to *not advertise* a subnet, passive-interface is the wrong tool. You either keep that interface out of the OSPF process entirely (no `network` statement covering it, no `ip ospf` on the interface), or you use a distribute-list / route filtering. Passive-interface is purely about Hellos and adjacencies, never about which prefixes get advertised. The mirror-image gotcha: if you make a router-to-router link passive by mistake, the adjacency silently never forms. Both subnets still show up in OSPF, the link is up, but the two routers never become neighbors. It looks like a deep OSPF problem and it is a one-line config slip. ## Verifying The cleanest confirmation: ``` R1# show ip ospf interface brief Interface PID Area IP Address/Mask Cost State Nbrs F/C Gi0/0 1 0 10.30.30.1/30 1 P2P 1/1 Gi0/1 1 0 10.20.0.1/24 1 DR 0/0 Gi0/2 1 0 10.30.31.1/30 1 P2P 1/1 ``` Gi0/1 (the user LAN) shows `0/0` neighbors - none, as intended for a passive host interface. Gi0/0 and Gi0/2 each show `1/1` \- one neighbor, fully adjacent, as intended for router-facing links. To see which interfaces OSPF considers passive: ``` R1# show ip ospf interface GigabitEthernet0/1 GigabitEthernet0/1 is up, line protocol is up Internet Address 10.20.0.1/24, Area 0 ... No Hellos (Passive interface) ... ``` The line `No Hellos (Passive interface)` is the explicit confirmation. And the most direct check of all: ``` R1# show ip protocols Routing Protocol is "ospf 1" ... Passive Interface(s): GigabitEthernet0/1 Routing for Networks: 10.20.0.0 0.0.0.255 area 0 10.30.30.0 0.0.0.3 area 0 10.30.31.0 0.0.0.3 area 0 ``` Notice that 10.20.0.0/24 appears under "Routing for Networks" even though Gi0/1 is passive - proof that a passive interface's subnet is still advertised. ## Common mistakes Used passive-interface expecting to hide a subnet The subnet is still advertised. Use a distribute-list or leave the interface out of OSPF instead. Accidentally made a router-facing link passive Adjacency never forms; looks like a major OSPF fault, is a one-liner. Used `passive-interface default` and forgot to un-passive an uplink That uplink forms no adjacency. Check `show ip ospf interface brief` for an unexpected 0/0. Left host-facing interfaces active Wasted Hellos and an OSPF attack surface on user VLANs. Made both ends of a point-to-point link passive Neither end sends Hellos, so no adjacency forms. Correct only when the remote end does not run OSPF at all. ## Key takeaways Passive-interface stops OSPF from sending Hellos out an interface, which stops adjacencies from forming there - and that is all it does. The interface's subnet is still advertised into OSPF. Make every host-facing interface passive: it removes wasted Hellos and closes an attack surface, while the users' subnet stays fully reachable. The strong pattern is `passive-interface default` plus explicit `no passive-interface` on the handful of router-facing links. Just never reach for passive-interface when what you actually want is to filter a route - it is not that tool. For the OSPF cluster, see the [OSPF pillar](https://www.pinglabz.com/ospf/). Take the OSPF reference with you The free OSPF field-reference PDF: neighbor states, LSA types, and the show commands you will actually run. Delivered by email, no card required. [Get the free PDF](https://www.pinglabz.com/ospf-cheatsheet/) ### Catalyst 9800-CL Day 0 Configuration, Line by Line (9800 Series Part 2) URL: https://www.pinglabz.com/catalyst-9800-cl-day-0-configuration/ Last updated: 2026-06-13T15:55:43.000Z This is Part 2 of the [PingLabz 9800 Wireless Labs series](https://www.pinglabz.com/9800-labs/). In [Part 1](https://www.pinglabz.com/catalyst-9800-cl-lab-cml/) we built the topology in Cisco Modeling Labs: a 9800-CL, a small IOL XE campus underneath it, a bridge to a physical AP, and CML's simulated wireless pair. Now we make it real: bring up the wired underlay, then walk the 9800-CL through its Day 0 configuration - wireless management interface, country code, admin access, and the self-signed certificate that the GUI and every future AP join depend on. Everything here was captured live from the lab. For background concepts, the [wireless guide](https://www.pinglabz.com/wireless/) covers the architecture in depth. *Video coming soon - the YouTube embed will land here when Part 2 is live.* ## Recap: the Addressing Plan ``` VLAN 10 MGMT 10.10.10.0/24 gw .1 WLC WMI = 10.10.10.10 VLAN 20 WIRELESS-CLIENTS 10.10.20.0/24 gw .1 client traffic VLAN 30 APS 10.10.30.0/24 gw .1 access points edge /30 link 10.0.0.0/30 CORE-SW1 to EDGE-RTR1 ``` ## Step 1: The Wired Underlay The WLC cannot do anything until the network beneath it works, so the first half of this part is plain L2/L3\. CORE-SW1 is the L3 heart of the lab: it owns all three SVIs, routes to the edge over a /30, and trunks down to the access layer and across to the WLC. ``` hostname CORE-SW1 ip routing ! vlan 10 name MGMT vlan 20 name WIRELESS-CLIENTS vlan 30 name APS ! interface Ethernet0/0 description Link to EDGE-RTR1 no switchport ip address 10.0.0.2 255.255.255.252 ! interface Ethernet0/1 description Trunk to ACCESS-SW1 switchport trunk encapsulation dot1q switchport trunk allowed vlan 10,20,30 switchport mode trunk ! interface Ethernet0/2 description Trunk to WLC1 Gi1 switchport trunk encapsulation dot1q switchport trunk allowed vlan 10,20,30 switchport mode trunk ! interface Vlan10 ip address 10.10.10.1 255.255.255.0 interface Vlan20 ip address 10.10.20.1 255.255.255.0 interface Vlan30 ip address 10.10.30.1 255.255.255.0 ! ip route 0.0.0.0 0.0.0.0 10.0.0.1 ``` ACCESS-SW1 is pure L2: a trunk up to the core, the physical AP port in VLAN 30, and the simulated AP port in VLAN 20\. Both AP-facing ports get portfast (an AP rebooting through 30 seconds of STP listening/learning is a self-inflicted outage). ``` interface Ethernet0/0 description Trunk to CORE-SW1 switchport trunk encapsulation dot1q switchport trunk allowed vlan 10,20,30 switchport mode trunk ! interface Ethernet0/1 description Physical AP via EXT-BRIDGE switchport mode access switchport access vlan 30 spanning-tree portfast ! interface Ethernet0/2 description SIM-AP1 switchport mode access switchport access vlan 20 spanning-tree portfast ``` Verification before touching the WLC (if this is broken, everything after it will be too): ``` CORE-SW1# show interfaces trunk Port Mode Encapsulation Status Native vlan Et0/1 on 802.1q trunking 1 Et0/2 on 802.1q trunking 1 Port Vlans allowed on trunk Et0/1 10,20,30 Et0/2 10,20,30 Port Vlans allowed and active in management domain Et0/1 10,20,30 Et0/2 10,20,30 Port Vlans in spanning tree forwarding state and not pruned Et0/1 10,20,30 Et0/2 10,20,30 CORE-SW1# show ip interface brief | exclude unassigned Interface IP-Address OK? Method Status Protocol Ethernet0/0 10.0.0.2 YES manual up up Vlan10 10.10.10.1 YES manual up up Vlan20 10.10.20.1 YES manual up up Vlan30 10.10.30.1 YES manual up up ``` One CML-specific note: on the IOL-L2 image, VLANs defined in a startup config do not always survive the first boot (the vlan database is built at runtime). If `show interfaces trunk` shows your VLANs allowed but not active, re-enter the `vlan` definitions in config mode and they activate immediately. ## Step 2: 9800-CL Day 0, Line by Line The 9800 Day 0 wizard exists, but doing it manually teaches you what the wizard hides - and it is only about fifteen lines. Four building blocks: identity and AAA, L2 plumbing, the wireless management interface, and the country code. ``` hostname WLC1 ip domain name pinglabz.lab username admin privilege 15 secret Cisco123! ! aaa new-model aaa authentication login default local aaa authorization exec default local ! vlan 10 name MGMT vlan 20 name WIRELESS-CLIENTS vlan 30 name APS ! interface GigabitEthernet1 switchport mode trunk switchport trunk allowed vlan 10,20,30 ! interface Vlan10 description Wireless Management ip address 10.10.10.10 255.255.255.0 no shutdown ! ip route 0.0.0.0 0.0.0.0 10.10.10.1 ! wireless management interface Vlan10 wireless country US ! ip http secure-server ip http authentication local ``` Why each block matters: **`aaa authentication/authorization` local** \- without it, GUI login fails even with a valid admin account. **Gi1 as a trunk + SVI 10** \- the 9800-CL data port is a switchport; without the trunk and SVI the controller has no L3 presence at all. **`wireless management interface Vlan10`** \- this is the WLC's identity. CAPWAP, AP joins, mobility - all of it sources from the WMI. No WMI, no wireless. **`wireless country US`** \- APs will not power their radios without a regulatory domain. The classic symptom is an AP that joins and then sits with radios down. `**ip http secure-server**` \- no GUI without it. We deliberately leave plain `ip http server` off; IOS XE even prints a security warning if you enable it. ## Step 3: The Self-Signed Certificate (9800-CL Only) Hardware 9800s ship with a manufacturer-installed certificate (MIC) that APs use to validate the controller during the CAPWAP-DTLS handshake. The 9800-CL is a virtual machine - no MIC - so it needs a self-signed certificate bound to the wireless management trustpoint. Skip this and APs will refuse to join, with DTLS errors that do not obviously say "you forgot the cert". Here is the part that cost us real lab time, so you don't have to repeat it: the obvious approach does not work. If you build a trustpoint manually (`crypto key generate rsa`, `crypto pki trustpoint`, `enrollment selfsigned`, `crypto pki enroll`) you get a perfectly valid certificate - `show crypto pki certificates` says Available - but bind it with `wireless management trustpoint` and the wireless process refuses to see it: ``` WLC1# show wireless management trustpoint Trustpoint Name : 9800-selfsigned Certificate Info : Not Available <-- valid cert, but wireless won't use it Private key Info : Not Available ``` The 9800-CL has a purpose-built exec command for this. It temporarily stands up an internal CA (WLC\_CA), issues the device certificate to it over SCEP, binds the trustpoint to the WMI, and shuts the CA back down - all in one shot: ``` WLC1# wireless config vwlc-ssc key-size 2048 signature-algo sha256 password 0 PingLabz2026 Configuring vWLC-SSC... Crypto PKI CA server enabled. CA server name: 'WLC_CA' Crypto PKI trustpoint configured. Trustpoint name: 'WLC1_WLC_TP' Shutting down CA server. Trustpoint name: 'WLC1_WLC_TP' set to wireless management. Script is completed ``` Three prerequisites, learned the hard way: **1\. The WMI must be up and pingable first.** The internal script issues the cert over SCEP to the WMI's own IP and aborts if the ping fails. **2\. `ip http server` must be enabled during generation.** SCEP enrollment runs over HTTP; you can (and should) disable it again right after. **3\. No `$` in the password argument.** Known parsing bug: everything from the `$` on is silently dropped, and the script fails on password length. The trustpoint is always named `_WLC_TP` and the certificate is valid for 10 years. One more AAA lesson from the same session: `aaa new-model` applies to the console line too. After the next reload, the serial console demanded a login, which broke our lab automation mid-build. If you want a lab console that never asks for a password (while SSH and the GUI still do), exempt the console line with a named method list: ``` aaa authentication login CONSOLE none aaa authorization exec CONSOLE none line con 0 login authentication CONSOLE authorization exec CONSOLE ``` ## Step 4: Verify Everything ``` WLC1# show wireless management trustpoint Trustpoint Name : WLC1_WLC_TP Certificate Info : Available Certificate Type : SSC Certificate Hash : 20df6fb7ed29612f1864f3fca4ed0b0452636b2e Private key Info : Available FIPS suitability : Not Applicable WLC1# ping 10.10.10.1 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 3/4/6 ms WLC1# show wireless interface summary Wireless Interface Summary Interface Name Interface Type VLAN ID IP Address IP Netmask NAT-IP Address MAC Address -------------------------------------------------------------------------------------------------- Vlan10 Management 10 10.10.10.10 255.255.255.0 0.0.0.0 001e.140f.03ff WLC1# show wireless country configured Configured Country.......................... US - United States Configured Country Codes US - United States 802.11a Indoor,Outdoor/ 802.11b Indoor,Outdoor/ 802.11g Indoor,Outdoor/ 802.11 6GHz Indoor,Outdoor WLC1# show ap summary Number of APs: 0 WLC1# show wlan summary Number of WLANs: 0 ``` Two values in that output are unique to your controller and will not match what is shown here: the **Certificate Hash** (regenerated every time you run `wireless config vwlc-ssc`) and the management interface **MAC address** in `show wireless interface summary` (assigned per VM instance). That is expected - match on the structure (trustpoint *Available*, type *SSC*, a private key present), not the exact digits. Zero APs and zero WLANs is exactly right for the end of Day 0\. The controller is reachable, owns its management identity, knows its regulatory domain, and has a certificate to offer. The GUI is now live at `https://10.10.10.10` \- log in with the admin account and you land on a dashboard that is empty in all the right ways. ## Key Takeaways Day 0 on a 9800-CL is four ideas: give the controller an identity (WMI on its own VLAN, sourced from an SVI over a trunked data port), give it a regulatory domain (no country code, no radios), give it AAA that the GUI can actually use (but exempt the console), and give it a certificate - generated with `wireless config vwlc-ssc`, not a hand-rolled trustpoint, because the wireless process only trusts its own enrollment flow. With the underlay verified first, the whole thing is under twenty lines of config. In **Part 3** we tackle the part of the 9800 everyone finds hardest coming from AireOS: the config model - WLAN profiles, policy profiles, and tags. The full series index lives on the [9800 Wireless Labs page](https://www.pinglabz.com/9800-labs/), and members can skip ahead with the config snapshots on the [lab files page](https://www.pinglabz.com/catalyst-9800-lab-files/). ### Build a Catalyst 9800-CL Wireless Lab in Cisco Modeling Labs (9800 Series Part 1) URL: https://www.pinglabz.com/catalyst-9800-cl-lab-cml/ Last updated: 2026-06-12T18:12:23.000Z This is Part 1 of the [PingLabz 9800 Wireless Labs series](https://www.pinglabz.com/9800-labs/). Over the next several posts (and matching YouTube videos), we build a complete Catalyst 9800-CL wireless environment from scratch in Cisco Modeling Labs - starting with an empty canvas and ending with wireless clients passing traffic through a WLC we configured line by line. If you are studying for CCNP ENCOR/ENWLSI or just want hands-on 9800 experience without buying hardware, this series is for you. For the broader wireless fundamentals behind everything we do here, see the [complete wireless guide](https://www.pinglabz.com/wireless/). In this first part we build the lab topology itself: the 9800-CL controller, a small wired campus underneath it, a bridge to a physical access point sitting on your desk, and CML's simulated wireless AP and client nodes. No device configs yet - that starts in Part 2. *Video coming soon - the YouTube embed will land here when Part 1 is live.* ## What You Need Everything in this lab runs on Cisco Modeling Labs (CML 2.9 or later) using reference platform images. The node mix was chosen deliberately: the 9800-CL is the one heavy VM in the lab, so everything around it uses the lightweight IOL XE images (Docker-based, boot in seconds, tiny RAM footprint) instead of full VMs like the Catalyst 9000v or Catalyst 8000v. **WLC1** (`cat9800`) - Catalyst 9800-CL on IOS XE 17.18, the wireless LAN controller **EDGE-RTR1** (`iol-xe`) - IOL XE router, WAN edge and simulated internet **CORE-SW1** (`ioll2-xe`) - IOL XE L2 switch, the L3 core: SVIs and future DHCP **ACCESS-SW1** (`ioll2-xe`) - IOL XE L2 switch, access layer for the APs **EXT-BRIDGE** (`external_connector`) - bridge mode, connects a physical AP into the lab **SIM-AP1** (`wireless-ap`) - Ubuntu + hostapd, simulated Wi-Fi access point **WCLIENT1** (`wireless-client`) - Ubuntu + wpa\_supplicant, simulated Wi-Fi client The whole topology, including the 9800-CL, fits comfortably in a CML instance with 16 GB of free RAM. If you ran the same design with Catalyst 9000v switches you would need roughly 18 GB per switch, which is why we don't. ## The Topology ``` EDGE-RTR1 (IOL-XE) | e0/0 - e0/0 CORE-SW1 (IOL-L2)----- e0/2 - Gi1 ----- WLC1 (9800-CL) | e0/1 - e0/0 ACCESS-SW1 (IOL-L2) / \ e0/1 (VLAN 30) e0/2 (VLAN 20) | | EXT-BRIDGE SIM-AP1 (hostapd) (physical AP) | ens3 - ens2 WCLIENT1 ``` A deliberately small campus: one router, one core, one access switch. It is enough to demonstrate every core 9800 concept (trunking the WLC, separating management from client traffic, AP joins across an L2/L3 boundary) without burying the wireless content under a big wired build. ## Link Map ``` EDGE-RTR1 e0/0 <-> CORE-SW1 e0/0 routed /30 uplink CORE-SW1 e0/1 <-> ACCESS-SW1 e0/0 802.1Q trunk CORE-SW1 e0/2 <-> WLC1 Gi1 802.1Q trunk to the WLC ACCESS-SW1 e0/1 <-> EXT-BRIDGE port physical AP, access VLAN 30 ACCESS-SW1 e0/2 <-> SIM-AP1 ens2 simulated AP uplink, VLAN 20 SIM-AP1 ens3 <-> WCLIENT1 ens2 simulated RF path ``` ## The External Connector: Getting a Real AP Into a Virtual Lab The most interesting node in this topology is the one that isn't virtual. CML's external connector in bridge mode patches a lab link straight through to a physical NIC on the CML server. Plug a real Catalyst AP into that NIC (or into a switch port on the same segment) and it will CAPWAP-join the virtual 9800-CL exactly as if both were physical. Two things to check before it works: **1\. The connector is set to bridge0, not NAT** (node config on the canvas). NAT mode hides the lab behind the CML host; the AP could reach out but the WLC could never reach the AP. **2\. bridge0 maps to the right physical NIC** (CML Cockpit / system settings). bridge0 is just a label; confirm it is bound to the interface your AP plugs into. We put the external connector behind ACCESS-SW1 on its own AP VLAN (VLAN 30) rather than hanging it off the core. That mirrors a real campus (APs live at the access layer) and gives us a clean L3 boundary between the APs and the WLC management network, which makes the AP join process in Part 5 much more instructive than a flat single-subnet design. ## The Simulated Wireless AP and Client CML ships two wireless node types, and it is worth being precise about what they are. The `wireless-ap` node is an Ubuntu VM running hostapd, and the `wireless-client` node is an Ubuntu VM running wpa\_supplicant. The "RF" between them is a simulated radio link drawn on the canvas like any other connection. What that means in practice (this matters for the whole series): the **simulated AP does not speak CAPWAP**, so it will never join the 9800\. It broadcasts a simulated open SSID ("openap" by default) that the simulated client associates to, which makes the pair perfect for client-side work: DHCP over wireless, packet captures, and 802.1X testing later. The **physical AP through the bridge is the real CAPWAP AP** \- it joins the controller and carries everything controller-side: AP joins, tags, WLAN pushes, and radio configuration. So the physical AP is the star of the controller content, and the simulated pair gives us an always-available client we can capture and break on demand (no neighbor complaints when we take down the SSID). ## Addressing Plan for the Series Locking this in now so every later part references the same plan: ``` VLAN 10 MGMT 10.10.10.0/24 gw .1 WLC WMI = 10.10.10.10 VLAN 20 WIRELESS-CLIENTS 10.10.20.0/24 gw .1 client traffic VLAN 30 APS 10.10.30.0/24 gw .1 access points edge /30 link 10.0.0.0/30 CORE-SW1 to EDGE-RTR1 ``` ## Gotcha: The 9800-CL Boots to a VGA Console in CML One trap worth fixing on day one. The 9800-CL image directs its console to the VGA (VNC) display by default, not the serial port. Open the normal console in CML and you will stare at a blank line forever while the controller boots happily on a screen you are not looking at. Any serial-based automation (PyATS, the CML breakout tool, your terminal client) hits the same wall. The fix is one command, applied once via the VNC console: ``` ! Open the VNC console (not Console) on the WLC node, log in, then: configure terminal platform console serial end write memory reload ``` After the reload the 9800 talks on the serial console like every other node in the lab, and it survives future reboots because it is saved in the startup config (a wipe of the node brings the VGA default back). Cisco documents this in the [CML 9800-CL guide](https://developer.cisco.com/docs/modeling-labs/catalyst-9800-cl/?ref=pinglabz.com). ## Key Takeaways The 9800-CL is the only heavy VM you need; IOL XE images keep the rest of the lab almost free. An external connector in bridge mode is the trick that lets a real AP join a virtual controller, and it belongs at the access layer on its own VLAN. CML's simulated wireless nodes are Linux Wi-Fi, not CAPWAP APs - use them for client-side testing and use the physical AP for controller-side features. In [**Part 2**](https://www.pinglabz.com/catalyst-9800-cl-day-0-configuration/) we bring the wired underlay up and walk the 9800-CL through its Day 0 configuration: wireless management interface, country code, and first GUI login. The full series index lives on the [9800 Wireless Labs page](https://www.pinglabz.com/9800-labs/), and members can grab the importable topology from the [lab files page](https://www.pinglabz.com/catalyst-9800-lab-files/). ### Catalyst 9800 Wireless Series: Lab Files & Downloads URL: https://www.pinglabz.com/catalyst-9800-lab-files/ Last updated: 2026-08-02T03:05:47.000Z _This post is for subscribers only._ ### EIGRP Neighbor Requirements: The 5 Things That Must Match URL: https://www.pinglabz.com/eigrp-neighbor-requirements/ Last updated: 2026-08-01T19:31:08.000Z Two EIGRP routers connected by a working link will not necessarily form a neighbor relationship. EIGRP has a specific list of things that must match before two routers become neighbors, and when an adjacency refuses to come up, the cause is almost always one item on that list. This post is the checklist: the five things that must match, the things that surprisingly do *not* have to match, and how to confirm each one on Cisco IOS XE. For the cluster overview, see the [EIGRP complete guide](https://www.pinglabz.com/eigrp/). For how this compares to the other major IGP, see the [OSPF pillar](https://www.pinglabz.com/ospf/). ## The five requirements For two routers to become EIGRP neighbors, all five of these must be true: Connectivity on the same subnet #1 Why EIGRP Hellos are sent to a multicast address on the local link; neighbors must be L2-adjacent and in the same IP subnet. Matching autonomous system number #2 Why The AS number is part of the EIGRP process identity. Different AS = different EIGRP process = no adjacency. Matching K-values (metric weights) #3 Why If two routers weight the metric components differently, EIGRP refuses to peer rather than risk inconsistent path math. Matching authentication #4 Why If authentication is configured, the mode and key must match. A mismatch silently drops the Hellos. The interface is in the EIGRP process and not passive #5 Why The interface must be covered by a `network` statement (or named-mode config) and must not be a passive interface. ## Requirement 1: same subnet and primary address EIGRP forms neighbors with routers it hears Hellos from on a directly connected link. Those routers must be in the same IP subnet. A common subtle failure: the two interfaces are physically connected but addressed in different subnets (a typo in the mask, or addresses that simply do not overlap). The link is up, but EIGRP never sees a neighbor. One more wrinkle: EIGRP sources Hellos from the interface's *primary* IP address and expects neighbors on that primary subnet. If you are using secondary addresses, the primary subnets must still align. ## Requirement 2: matching AS number EIGRP is configured per autonomous system. `router eigrp 100` and `router eigrp 200` are two completely separate EIGRP processes that will not talk to each other, even on the same wire. ``` ! R1 router eigrp 100 network 10.30.30.0 0.0.0.3 ! ! R2 - will NOT peer with R1 router eigrp 200 network 10.30.30.0 0.0.0.3 ``` In named mode the AS number appears in the address-family line (`address-family ipv4 unicast autonomous-system 100`) - same rule, it must match. ## Requirement 3: matching K-values EIGRP's composite metric is built from up to five components, each weighted by a K-value: K1 (bandwidth), K2 (load), K3 (delay), K4 and K5 (reliability/MTU terms). The defaults are K1=1, K2=0, K3=1, K4=0, K5=0 - bandwidth and delay only. If one router has the default K-values and someone changed them on the other, EIGRP will not form the adjacency. The two routers would compute metrics differently, which could create routing loops, so EIGRP refuses to peer at all rather than allow the inconsistency. A K-value mismatch produces a specific log message naming the mismatch. The practical advice: do not change K-values. The defaults are correct for virtually every network, and the only thing changing them reliably achieves is breaking adjacencies with routers you forgot to also change. ## Requirement 4: matching authentication If EIGRP authentication is configured on an interface, the neighbor must present the matching authentication. Both the mode (MD5, or the newer named-mode SHA / HMAC-SHA-256) and the key must align. The failure mode here is quiet. A router receiving Hellos that fail authentication simply discards them. There is no neighbor, and unless you are looking at debug output you do not see *why* \- it looks identical to "no Hellos arriving at all." If an adjacency will not form and the link is clean, check whether one side has authentication configured and the other does not. ## Requirement 5: interface in the process, not passive The interface has to actually be running EIGRP. In classic mode that means a `network` statement covers the interface's IP. In named mode it means the interface falls under the address-family configuration. And it must not be passive. `passive-interface` tells EIGRP to stop sending Hellos out that interface - which means no neighbor will ever form across it. Passive-interface is correct on interfaces facing hosts (you advertise the subnet but do not want neighbors there); it is a bug on an interface where you expect an adjacency. ## What does NOT have to match This is the part that trips people who assume EIGRP behaves like OSPF: Hello and hold timers Unlike OSPF, EIGRP neighbors form even with different Hello/hold timers. Each router tells the other its own hold time. (Mismatched timers can still cause instability - keep them aligned in practice - but they do not block the adjacency.) Router IDs Should be unique, but a duplicate RID does not stop the adjacency the way it would in OSPF. Interface MTU EIGRP does not require matching MTU to form a neighbor (OSPF does, and stalls in ExStart/Exchange on a mismatch). The OSPF habit of blaming Hello-timer and MTU mismatches does not transfer to EIGRP. For EIGRP, focus on the five real requirements. ## Verifying the adjacency The first command: ``` R1# show ip eigrp neighbors EIGRP-IPv4 Neighbors for AS(100) H Address Interface Hold Uptime SRTT RTO Q Seq (sec) (ms) Cnt Num 0 10.30.30.2 Et0/0 13 01:42:08 8 100 0 24 ``` A neighbor listed here with a climbing Uptime is healthy. If the table is empty, walk the five requirements. To confirm the interface is actually participating: ``` R1# show ip eigrp interfaces EIGRP-IPv4 Interfaces for AS(100) Xmit Queue PeerQ Mean Interface Peers Un/Reliable Un/Reliable SRTT Et0/0 1 0/0 0/0 8 ``` If the interface you expect is missing from this list, it is not in the EIGRP process (requirement 5) - or it is passive. And for the most direct answer when an adjacency is failing: ``` R1# debug eigrp packets ``` This shows incoming Hellos and will explicitly report a K-value mismatch or an authentication failure. Turn it off as soon as you have the answer. ## Troubleshooting flow Neighbor table empty, link is up Same subnet? `show ip interface brief` both ends. Log shows "K-value mismatch" Someone changed K-values. Restore defaults on both. Hellos sent but no neighbor, link clean Authentication mismatch - one side configured, the other not. Interface missing from `show ip eigrp interfaces` Not covered by a network statement, or set passive. Neighbor flaps - forms then drops repeatedly Often duplex/MTU/physical issues, or unidirectional connectivity. Not a "requirement" failure but a link problem. Each of those rows leaves a different fingerprint, and knowing which failures shout in syslog and which fail in complete silence is what turns this checklist into a two-minute diagnosis. [Watching each EIGRP adjacency failure break on a live router](https://www.pinglabz.com/troubleshooting-eigrp-neighbor-adjacencies/) takes the same causes and shows the logs and show output each one produces. ## Key takeaways Two EIGRP routers become neighbors only when five things line up: same subnet, same AS number, matching K-values, matching authentication, and the interface in the process and not passive. Unlike OSPF, EIGRP does not require matching Hello timers or MTU - so do not waste time chasing those. When an adjacency will not form, run `show ip eigrp neighbors`, then `show ip eigrp interfaces`, then walk the five-item list. A K-value mismatch and a one-sided authentication config are the two that fail most quietly. For the EIGRP cluster, see the [EIGRP pillar](https://www.pinglabz.com/eigrp/). ### Per-VLAN Spanning Tree (PVST+ and Rapid-PVST+) Explained URL: https://www.pinglabz.com/per-vlan-spanning-tree/ Last updated: 2026-06-13T20:07:44.000Z Per-VLAN Spanning Tree is the reason a Cisco switch with 50 VLANs is running 50 spanning trees. It is also the reason you can load-balance traffic across redundant uplinks instead of letting half your links sit idle in blocking. PVST+ and Rapid-PVST+ are Cisco's per-VLAN take on spanning tree, and understanding them is the difference between accepting whatever topology STP hands you and engineering the one you want. This post explains how per-VLAN spanning tree works, why Cisco built it, and how to use it to load-balance. For the cluster overview, see the [Spanning Tree Protocol pillar](https://www.pinglabz.com/spanning-tree-protocol/). For the L2 fundamentals underneath, see the [VLAN and L2 switching pillar](https://www.pinglabz.com/vlans-layer-2-switching/). ## The problem with one spanning tree The original IEEE 802.1D ran a single spanning tree for the entire switched network, regardless of how many VLANs you had. One tree means one root bridge, one set of blocked ports, one topology. Picture two distribution switches and two uplinks from an access switch. With a single spanning tree, one uplink forwards and the other blocks - for everything. Every VLAN's traffic crosses the same uplink. The second uplink, which you paid for, carries nothing until the first one fails. That is wasteful, and it is the problem per-VLAN spanning tree solves. ## The per-VLAN idea Per-VLAN spanning tree runs a separate, independent spanning tree instance for every VLAN. Each VLAN can have its own root bridge, its own port roles, its own blocked ports. Now take the same two-uplink example. You make distribution switch A the root for VLANs 10, 20, 30, and distribution switch B the root for VLANs 40, 50, 60\. VLAN 10's spanning tree forwards on the uplink toward switch A and blocks the other. VLAN 40's spanning tree does the opposite. Both uplinks now carry traffic - half the VLANs on each. You have load-balanced across redundant links using nothing but root-bridge placement. ## The Cisco PVST family PVST What it is Original per-VLAN STP. Required Cisco ISL trunking. StatusObsolete (ISL is dead) PVST+ What it is Per-VLAN STP over 802.1Q trunks. Per-VLAN instances, but each runs the slow legacy 802.1D state machine. StatusLegacy, still seen Rapid-PVST+ What it is Per-VLAN STP where each instance runs the fast 802.1w (RSTP) state machine. StatusCurrent Cisco default On any modern Catalyst the default is Rapid-PVST+: one instance per VLAN, each instance converging in sub-second time thanks to RSTP. PVST+ (the slow-converging version) still turns up on older gear. The configuration command makes the choice explicit: ``` spanning-tree mode rapid-pvst ``` ## The Bridge ID and the sys-id-ext trick Per-VLAN spanning tree needs a separate Bridge ID per VLAN, but a switch has one MAC address. Cisco solved this by splitting the Bridge ID priority field. The 16-bit Bridge ID priority became: a 4-bit configurable priority + a 12-bit "system ID extension" that holds the VLAN number. So a switch running spanning tree for VLAN 10 with priority 32768 actually advertises a Bridge ID priority of 32768 + 10 = 32778\. For VLAN 20 it is 32768 + 20 = 32788. This is why `show spanning-tree` reports lines like `priority 32778 (priority 32768 sys-id-ext 10)`. The base priority is 32768; the 10 is the VLAN. It also explains why switch priority must be set in multiples of 4096 - only the top 4 bits are yours; the bottom 12 belong to the VLAN ID. ## Setting the root: the load-balancing config You steer each VLAN's tree by setting the root bridge per VLAN. Two ways: ``` ! Explicit priority - multiples of 4096 spanning-tree vlan 10,20,30 priority 24576 spanning-tree vlan 40,50,60 priority 28672 ``` Or the macro that picks a low priority for you: ``` spanning-tree vlan 10,20,30 root primary spanning-tree vlan 40,50,60 root secondary ``` On distribution switch A you make VLANs 10-30 primary and 40-60 secondary. On distribution switch B you do the mirror image. The result: VLANs 10-30 root on A, VLANs 40-60 root on B, and the access switch's two uplinks both forward - each carrying the VLANs whose root they point toward. ## Reading the per-VLAN topology ``` SW-ACCESS# show spanning-tree vlan 10 VLAN0010 Spanning tree enabled protocol rstp Root ID Priority 24586 Address 0050.7989.aaaa Cost 4 Port 25 (TenGigabitEthernet1/1/1) Bridge ID Priority 32778 (priority 32768 sys-id-ext 10) Address 0050.7989.cccc Interface Role Sts Cost Prio.Nbr Type ------------------- ---- --- --------- -------- ---------------- Te1/1/1 Root FWD 4 128.25 P2p Te1/1/2 Altn BLK 4 128.26 P2p ``` For VLAN 10, the access switch's uplink Te1/1/1 is the Root port and forwards; Te1/1/2 is Alternate and blocks. Run `show spanning-tree vlan 40` and you should see the roles swapped - Te1/1/2 forwarding, Te1/1/1 blocking - because VLAN 40 roots on the other distribution switch. That swap is the proof your load balancing is working. The summary view across all VLANs: ``` SW-ACCESS# show spanning-tree summary Switch is in rapid-pvst mode Name Blocking Listening Learning Forwarding STP Active ---------------------- -------- --------- -------- ---------- ---------- VLAN0010 1 0 0 1 2 VLAN0040 1 0 0 1 2 ``` ## The cost of per-VLAN: scale Per-VLAN spanning tree's strength - an independent tree per VLAN - is also its limit. Every VLAN is a full state machine, a set of BPDUs sent every 2 seconds on every trunk, and CPU to maintain. With 20 or 50 VLANs this is a non-issue. With 500 VLANs across a large campus it becomes real overhead, and the BPDU volume on trunks adds up. That is the point where MSTP takes over: it keeps the load-balancing benefit by letting you map many VLANs onto a few instances, instead of one instance per VLAN. The rule of thumb: Rapid-PVST+ up to roughly 50-75 VLANs, MSTP beyond that. ## Common gotchas "Priority must be a multiple of 4096" rejected your value Only the top 4 bits of the priority field are yours; the rest is the VLAN sys-id-ext. Use 0, 4096, 8192, ... 61440. Both uplinks forward for every VLAN - no blocking anywhere The two links are an EtherChannel (correct) - or a real loop with STP disabled (dangerous). Confirm which. One uplink carries all VLANs; the other blocks for all All VLANs share one root bridge. Split the root placement per VLAN range to load-balance. BPDU load high on a big campus switch Too many per-VLAN instances. Migrate to MSTP. Native VLAN inconsistency warnings Trunk native VLAN mismatch - unrelated to PVST roles but flagged in the same logs. ## Key takeaways Per-VLAN spanning tree runs an independent spanning tree per VLAN, which lets you place a different root bridge per VLAN and load-balance traffic across redundant uplinks instead of leaving half of them blocked. Cisco's current default is Rapid-PVST+ - per-VLAN instances each running the fast RSTP state machine. The VLAN number lives in the Bridge ID's sys-id-ext field, which is why switch priority must be a multiple of 4096\. The model is excellent up to \~50-75 VLANs; beyond that, MSTP keeps the load-balancing benefit without the per-VLAN overhead. For the STP cluster, see the [Spanning Tree Protocol pillar](https://www.pinglabz.com/spanning-tree-protocol/). ### Lab auto-07 - Capstone: Automated Config Backup and Drift Detection URL: https://www.pinglabz.com/ccna-auto-07-capstone-config-backup-drift/ Last updated: 2026-08-02T03:05:48.000Z _This post is for paying subscribers only._ ### Lab auto-06 - Ansible for CCNA: Your First Playbook URL: https://www.pinglabz.com/ccna-auto-06-ansible-first-playbook/ Last updated: 2026-08-02T03:05:48.000Z _This post is for paying subscribers only._ ### Lab auto-05 - REST APIs Hands-On: Drive Your CML Controller URL: https://www.pinglabz.com/ccna-auto-05-rest-apis-cml/ Last updated: 2026-08-02T03:05:48.000Z _This post is for paying subscribers only._ ### Lab auto-04 - JSON and Structured Data for Network Engineers URL: https://www.pinglabz.com/ccna-auto-04-json-structured-data/ Last updated: 2026-08-02T03:05:49.000Z _This post is for paying subscribers only._ ### Lab auto-03 - Pushing Config with Netmiko URL: https://www.pinglabz.com/ccna-auto-03-netmiko-config-push/ Last updated: 2026-08-02T03:05:49.000Z _This post is for paying subscribers only._ ### Lab auto-02 - Reading the Network with Netmiko URL: https://www.pinglabz.com/ccna-auto-02-netmiko-show-commands/ Last updated: 2026-08-02T03:05:50.000Z _This post is for paying subscribers only._ ### Lab auto-01 - Build Your First Network Automation Host URL: https://www.pinglabz.com/ccna-auto-01-first-automation-host/ Last updated: 2026-08-02T03:05:50.000Z Every network automation tutorial on the internet starts the same way: "first, install Python on your laptop." This one does not. In this lab you build the automation host *inside* the network, on the same Alpine Linux node that ships with the [PingLabz CCNA Automation Labs](https://www.pinglabz.com/automation-labs/) topology. By the end, a Python script running on HOST1 will SSH into R1 and pull live output from a router you can see booting in Cisco Modeling Labs. This is the first lab in the automation series, it is completely free, and everything below was captured from the real lab: Alpine 3.23, Python 3.12, Netmiko 4.7, and Cisco IOS XE 17.18. Automation and Programmability is Domain 6 of the CCNA 200-301 blueprint, worth 10% of your exam. It is also the domain most CCNA candidates study entirely from flashcards. The difference between memorizing "REST APIs use HTTP verbs" and having actually pushed a command to a router from Python is the same difference the other 90% of the exam rewards: hands-on beats theory recall, every time. ## The topology This series runs on the PingLabz CCNA Automation Lab topology, a variant of the same five-node base used across the [PingLabz lab library](https://www.pinglabz.com/labs/). It fits Cisco Modeling Labs Free: five counted nodes, plus an unmanaged switch and an external connector, neither of which counts toward the CML Free limit. R1 Platformiol-xe (IOS XE 17.18) RoleLAN gateway Management IP10.20.0.1 R2 Platformiol-xe RoleTransit router Management IP10.20.0.2 R3 Platformiol-xe RoleRemote router Management IP10.30.30.2 SW1 Platformioll2-xe RoleManaged L2 switch Management IP10.20.0.10 HOST1 PlatformAlpine Linux Role The automation host (you build it here) Management IP10.20.0.50 SW2 Platformunmanaged switch Role Spare broadcast domain (not counted) Management IPn/a EXT-NAT Platformexternal connector Role NAT internet for HOST1 (not counted) Management IPn/a The external connector is the one piece you have not seen in the other PingLabz series. HOST1 needs to download packages from the Alpine mirrors, so its second interface (eth1) connects to an External Connector in NAT mode. CML gives that interface outbound internet through your computer's own connection, with no configuration on your side. Per the Cisco CML documentation, external connectors do not count toward the node license, so the lab still fits CML Free. ## What you will learn - How to set up an Alpine Linux node in CML as a network automation host, the right way (let it boot with its default config first). - How to give a Linux host dual-homed networking: one interface for internet, one for the lab management network. - How to enable SSH version 2 on Cisco IOS XE with modern 3072-bit RSA keys, and why the old config-mode keygen command is deprecated. - How to install Netmiko in a Python virtual environment and make your first programmatic connection to a router. ## Download the topology [PingLabz CCNA Automation Lab topologyImport into CML Free and start the lab.pinglabz-ccna-automation-lab.yaml12 KBdownload-circle](https://www.pinglabz.com/content/files/2026/08/pinglabz-ccna-automation-lab.yaml "Download") Import the .yaml into CML (Import Lab in the dashboard), hit start, and give the nodes a minute to boot. The routers and switch boot with the standard PingLabz base configuration. Router login is `pinglabz / PingLabz!23` with enable secret `Cisco@123`. ## Step 1: Let HOST1 boot, then log in One rule for Alpine nodes in CML, learned the hard way: **let the node boot completely with its default configuration before you touch it**. Do not attach a day-0 config, and do not start typing into the console while it is still booting. Give it a minute after the lab starts, then open the HOST1 console and log in with CML's Alpine defaults: username `cisco`, password `cisco`. ``` HOST1 login: cisco Password: HOST1:~$ whoami cisco HOST1:~$ cat /etc/alpine-release 3.23.3 ``` The `cisco` user has passwordless sudo, which you will use for everything that touches the network stack or the package manager. ## Step 2: Get internet on eth1 HOST1 has two interfaces. eth0 is wired to SW1 (the lab LAN); eth1 is wired to the EXT-NAT external connector. Bring eth1 up and ask for a DHCP lease: ``` HOST1:~$ sudo ip link set eth1 up HOST1:~$ sudo udhcpc -i eth1 -n -q udhcpc: started, v1.37.0 udhcpc: broadcasting discover udhcpc: broadcasting select for 192.168.255.24, server 192.168.255.1 udhcpc: lease of 192.168.255.24 obtained from 192.168.255.1, lease time 3600 ``` The `-n` flag makes udhcpc exit if no lease appears (instead of retrying forever in your console), and `-q` quits once the lease is bound. CML's NAT connector hands out addresses from its internal 192.168.255.0/24 pool and NATs everything outbound. Confirm you can reach the internet: ``` HOST1:~$ ping -c 2 8.8.8.8 PING 8.8.8.8 (8.8.8.8): 56 data bytes 64 bytes from 8.8.8.8: seq=0 ttl=42 time=12.324 ms 64 bytes from 8.8.8.8: seq=1 ttl=42 time=13.307 ms ``` ## Step 3: Address eth0 for the lab LAN eth0 talks to the lab. Give it the standard PingLabz automation-host address, plus one static route so HOST1 can reach R3's side of the point-to-point link through R2: ``` HOST1:~$ sudo ip addr add 10.20.0.50/24 dev eth0 HOST1:~$ sudo ip route add 10.30.30.0/30 via 10.20.0.2 HOST1:~$ traceroute -n -m 4 10.30.30.2 traceroute to 10.30.30.2 (10.30.30.2), 4 hops max, 46 byte packets 1 10.20.0.2 3.433 ms 2.823 ms 4.659 ms 2 10.30.30.2 4.133 ms * 5.199 ms ``` Two hops: HOST1 to R2, R2 across the P2P link to R3\. The default route stays on eth1 (internet); only the lab prefixes use eth0\. This is the same split-management pattern you will meet in production, where the automation host has one leg in the management VLAN and one leg toward the wider network. Runtime `ip` commands do not survive a reboot. To make the addressing permanent, write it into Alpine's `/etc/network/interfaces`: ``` auto lo iface lo inet loopback auto eth0 iface eth0 inet static address 10.20.0.50/24 post-up ip route add 10.30.30.0/30 via 10.20.0.2 || true auto eth1 iface eth1 inet dhcp ``` ## Step 4: Install Python and Netmiko Alpine ships lean: no Python on board. Two packages from the mirrors fix that: ``` HOST1:~$ sudo apk update v3.23.4-359-g5b71e521e31 [https://dl-cdn.alpinelinux.org/alpine/v3.23/main] v3.23.4-359-g5b71e521e31 [https://dl-cdn.alpinelinux.org/alpine/v3.23/community] OK: 27629 distinct packages available HOST1:~$ sudo apk add python3 py3-pip (9/20) Installing python3 (3.12.13-r0) ... OK: 138.1 MiB in 123 packages HOST1:~$ python3 --version Python 3.12.13 ``` Install Netmiko inside a virtual environment. Modern Python distributions protect the system site-packages, and a venv keeps your automation dependencies isolated and reproducible (the same discipline you will want on a production jump host): ``` HOST1:~$ python3 -m venv ~/autolab HOST1:~$ . ~/autolab/bin/activate (autolab) HOST1:~$ pip install netmiko (autolab) HOST1:~$ python3 -c "import netmiko; print('netmiko', netmiko.__version__)" netmiko 4.7.0 ``` ## Step 5: Enable SSH on R1, the modern way Netmiko talks SSH, and the routers boot with SSH allowed on the VTY lines but no host keys generated. On the R1 console, generate the keys from privileged EXEC mode. Watch what IOS XE 17.x says if you try the old way first: ``` R1(config)# crypto key generate rsa modulus 2048 %This command is deprecated. Use the command in exec mode instead ... SECURITY WARNING - Module: SSH, Command: crypto key generate rsa ..., Reason: SSH host key uses insufficient key length, Description: SSH with insufficient key length, Remediation: Use SSH RSA host key with a minimum length of 3072 bits for enhanced security ``` Two lessons in one capture: config-mode keygen is deprecated, and 2048-bit RSA now trips a security warning. Do it properly, from EXEC mode, at 3072 bits, then pin SSH to version 2: ``` R1# crypto key generate rsa modulus 3072 % The key modulus size is 3072 bits % Generating crypto RSA keys in background ... R1# configure terminal R1(config)# ip ssh version 2 R1# show ip ssh SSH Enabled - version 2.0 Authentication methods:publickey,keyboard-interactive,password Encryption Algorithms:chacha20-poly1305@openssh.com,aes128-gcm@openssh.com,aes256-gcm@openssh.com, ... Minimum expected Diffie Hellman key size : 2048 bits ``` (If you ran the deprecated 2048-bit command first, IOS XE prints `% You already have RSA keys defined named R1.pinglabz.lab.` and `% They will be replaced.` before it generates the new pair. That is expected: the 3072-bit key simply overwrites the weaker one.) Sanity-check the path from HOST1 with a plain SSH client before involving Python (`sudo apk add openssh-client` if you have not already): ``` HOST1:~$ ssh pinglabz@10.20.0.1 'show ip interface brief' (pinglabz@10.20.0.1) Password: ***************************************************** * PingLabz CCNA Automation Lab - R1 * * https://www.pinglabz.com/automation-labs/ * * Authorized practice use only * ***************************************************** Interface IP-Address OK? Method Status Protocol Ethernet0/0 10.20.0.1 YES TFTP up up Ethernet0/1 unassigned YES unset administratively down down Ethernet0/2 unassigned YES unset administratively down down Ethernet0/3 unassigned YES unset administratively down down Loopback0 10.255.0.1 YES TFTP up up ``` (The Method column says TFTP because CML injects the startup configuration at boot; on physical hardware you would normally see NVRAM or manual.) ## Step 6: Your first Netmiko script Now replace yourself with Python. Create `hello_netmiko.py` on HOST1: ``` from netmiko import ConnectHandler r1 = { "device_type": "cisco_ios", "host": "10.20.0.1", "username": "pinglabz", "password": "PingLabz!23", "secret": "Cisco@123", } conn = ConnectHandler(**r1) print(conn.find_prompt()) print(conn.send_command("show version | include uptime")) conn.disconnect() ``` Run it: ``` (autolab) HOST1:~$ python3 hello_netmiko.py R1# R1 uptime is 17 minutes ``` That is the whole trick. `ConnectHandler` opens the SSH session and handles the prompt detection, `find_prompt()` proves you landed where you think you did, and `send_command()` sends one exec command and returns its output as a Python string. Everything else in network automation is this, repeated, with better error handling. ## Troubleshooting matrix Cannot log in to HOST1 console Likely cause Touched the node before first boot completed, or a day-0 config replaced the defaults Fix Wipe the node (not the lab) and let it boot untouched, then cisco/cisco `udhcpc: sendto: Network is down` Likely cause eth1 link is still down Fix `sudo ip link set eth1 up` first, then udhcpc `socket: Operation not permitted` Likely cause Ran a network command without root FixPrefix with `sudo` apk cannot reach mirrors Likely cause eth1 has no lease, or the EXT-NAT connector is not in NAT mode Fix Verify the lease with `ip addr show eth1`; check the connector's Config tab says NAT Netmiko times out connecting Likely cause SSH keys never generated on R1, or eth0 has no address Fix `show ip ssh` on R1 must say Enabled; `ip addr show eth0` on HOST1 Netmiko authentication failure Likely cause Typo in username/password dictionary Fix Credentials are pinglabz / PingLabz!23, secret Cisco@123 ## Key takeaways - The automation host lives inside the topology: Alpine + venv + Netmiko, fed by a NAT external connector that does not count against CML Free's five nodes. - Let Alpine boot with its default config before touching it. Login is cisco/cisco with passwordless sudo. - IOS XE 17.x wants RSA keys generated in EXEC mode at 3072 bits minimum; the config-mode command is deprecated and 2048-bit keys trigger a security warning. - One Netmiko pattern (ConnectHandler, find\_prompt, send\_command) is the foundation for every script in the rest of this series. Next up: [Lab auto-02, Reading the Network with Netmiko](https://www.pinglabz.com/ccna-auto-02-netmiko-show-commands/), where one script interrogates all four devices and you stop opening consoles for routine checks. The full series index lives at [/automation-labs/](https://www.pinglabz.com/automation-labs/). ### Lab ts-cap-01 - Troubleshooting Capstone: The Branch Office Outage URL: https://www.pinglabz.com/ccna-ts-cap-01-total-outage-capstone/ Last updated: 2026-08-02T03:05:51.000Z _This post is for paying subscribers only._ ### Lab ts-sec-01 - Troubleshooting Security Tickets URL: https://www.pinglabz.com/ccna-ts-sec-01-security-tickets/ Last updated: 2026-08-02T03:05:50.000Z _This post is for paying subscribers only._ ### Lab ts-ips-01 - Troubleshooting DHCP and NAT Tickets URL: https://www.pinglabz.com/ccna-ts-ips-01-dhcp-nat-tickets/ Last updated: 2026-08-02T03:05:51.000Z _This post is for paying subscribers only._ ### Lab ts-ipc-01 - Troubleshooting OSPF Tickets URL: https://www.pinglabz.com/ccna-ts-ipc-01-ospf-tickets/ Last updated: 2026-08-02T03:05:52.000Z _This post is for paying subscribers only._ ### Lab ts-na-01 - Troubleshooting Switching and VLAN Tickets URL: https://www.pinglabz.com/ccna-ts-na-01-switching-vlan-tickets/ Last updated: 2026-08-02T03:05:52.000Z _This post is for paying subscribers only._ ### Lab ts-nf-01 - Troubleshooting Connectivity Tickets (L1/L2/L3) URL: https://www.pinglabz.com/ccna-ts-nf-01-connectivity-tickets/ Last updated: 2026-08-02T03:05:51.000Z The build labs in the [PingLabz CCNA Labs library](https://www.pinglabz.com/labs/) teach you to configure a network from a blank slate. This series does the opposite. You inherit a network that someone else already broke, you get a ticket that describes a symptom and nothing else, and your job is to localize the fault and fix it. That is the skill the CCNA exam tests hardest and the one most study guides skip: not "how do I configure OSPF," but "the network is down, where do I even start." This is the free preview of the PingLabz CCNA Troubleshooting Labs. It runs on the same five-node **PingLabz CCNA Base Topology** as the rest of the library, so it fits inside Cisco Modeling Labs Free. Every command output on this page was captured from that topology running on Cisco IOS XE, broken on purpose and then fixed. ## How these troubleshooting labs work Each lab is a small ticket queue. You import the topology, start it, and work the tickets in order. A ticket gives you a symptom in the user's words and a single success criterion. It does not tell you the cause. You reproduce the symptom, climb the OSI stack until you find the layer that is lying to you, fix it, and verify. After each ticket we show the root cause and the exact fix so you can check your work. ## The topology Three routers (R1, R2, R3), one managed switch (SW1), one host (HOST1), and an unmanaged switch (SW2). R1 and R2 share the LAN `10.20.0.0/24` through SW1\. R2 reaches R3 over the point-to-point link `10.30.30.0/30`. Each router has a loopback: `10.255.0.1`, `10.255.0.2`, `10.255.0.3`. In the healthy state, static routing ties it together and R1 can reach every loopback. You log in as `pinglabz / PingLabz!23`. Download the ts-nf-01 topology (.yaml) The PingLabz 5-node base topology, pre-loaded with three connectivity faults (a wrong subnet mask, a dead static next-hop, and a shut interface) to find and fix. Import into Cisco Modeling Labs (Free or higher), start all nodes, and log in as pinglabz / PingLabz!23\. Boots broken on purpose. [Download ts-nf-01 Topology](https://www.pinglabz.com/content/files/2026/08/pinglabz-ts-nf-01-connectivity.yaml) **Lab setup:** this topology boots with all three faults already in place. On a small network faults mask each other, so work bottom-up: fix the most local, most fundamental fault first, then re-test before moving on. That order matters here. All your testing starts at R1, and R1 itself carries the two most upstream faults, so until you fix them R1 cannot even see the third one. The tickets below are sequenced the way the network forces you to solve them. ## What you will learn - The bottom-up method: when end-to-end ping fails, climb the stack from Layer 1 instead of guessing, and fix the most local fault before chasing a remote one. - How to read the difference between a *timeout* (`.`) and an *unreachable* (`U`) in ping output, and why that one character tells you whether to look local or downstream. - The four commands that expose each layer: `show ip interface brief`, `show ip route`, `show ip arp`, and `traceroute`. - Why an interface in `up/up` can still drop traffic, and how a subnet mask becomes a connectivity fault that hides every other problem behind it. ## Ticket 1: "A new tech renumbered R1 and now nothing works" **Reported symptom:** "We changed R1's LAN address during a renumber. Now R1 can't reach anyone on the LAN." **Success criterion:** R1 can ping its neighbor `10.20.0.2`. Start at R1, because that is where every other test will originate too. Confirm the symptom and source from the LAN address the tech configured: ``` R1# ping 10.20.0.2 source 10.20.0.129 ..... Success rate is 0 percent (0/5) ``` Dots, not `U`. A timeout, not an active rejection: the packet is dying silently and nobody is sending anything back. Climb from the bottom. Is the interface up? ``` R1# show ip interface brief | include Ethernet0/0 Ethernet0/0 10.20.0.129 YES TFTP up up ``` Up and up. So this is not Layer 1 or Layer 2\. The interface is healthy and forwarding. But R1 has no idea how to reach its own neighbor: ``` R1# show ip route 10.20.0.2 % Subnet not in table ``` How can a directly connected neighbor not be in the table? Look at the interface in detail: ``` R1# show running-config interface Ethernet0/0 interface Ethernet0/0 description LAN to SW1 (10.20.0.0/24) ip address 10.20.0.129 255.255.255.128 ``` The description says `10.20.0.0/24`. The configured mask is `255.255.255.128`, a /25\. With a /25, R1 believes its connected subnet is `10.20.0.128/25` (hosts .129 to .254). The neighbor at `10.20.0.2` falls in the *other* half, `10.20.0.0/25`, so R1 thinks it is on a different network entirely and never even tries to ARP for it. Worse, this one fault hides the other two: every static route R1 has points at a next-hop (`10.20.0.2` or `10.20.0.99`) that now looks off-subnet, so none of those routes are even installed. Fix this first or nothing else you test from R1 will make sense. **Root cause:** wrong subnet mask. The interface description (the intended /24) contradicts the configured /25. **Fix:** ``` R1(config)# interface Ethernet0/0 R1(config-if)# ip address 10.20.0.1 255.255.255.0 ``` Re-test the neighbor: ``` R1# ping 10.20.0.2 source 10.20.0.1 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 3/3/5 ms ``` This is the trap of an `up/up` interface. Layer 1 and Layer 2 are fine, so the instinct to check cables and line protocol leads nowhere. The mask is a Layer 3 fault hiding behind a perfectly healthy physical link, and on this lab it is the fault that has to go first. ## Ticket 2: "R1 still can't reach the R3 site" **Reported symptom:** "The LAN works again, but R1 still cannot reach anything at the R3 site." **Success criterion:** R1 can ping `10.255.0.3`. Now that R1 can talk to the LAN, test the far loopback: ``` R1# ping 10.255.0.3 source 10.255.0.1 ..... Success rate is 0 percent (0/5) ``` Dots again, so the packet is dying locally on R1, not being rejected downstream. Look at the route R1 is using: ``` R1# show ip route 10.255.0.3 Routing entry for 10.255.0.3/32 Known via "static", distance 1, metric 0 Routing Descriptor Blocks: * 10.20.0.99 Route metric is 0, traffic share count is 1 ``` R1 is sending traffic for `10.255.0.3` to a next-hop of `10.20.0.99`. There is no such device on the LAN. Prove it: ``` R1# show ip arp 10.20.0.99 Protocol Address Age (min) Hardware Addr Type Interface Internet 10.20.0.99 0 Incomplete ARPA ``` `Incomplete` means R1 ARPed for `10.20.0.99` and nobody answered. The next-hop is a ghost. The correct gateway to R3 is R2 at `10.20.0.2`. **Root cause:** the static route points at a non-existent next-hop (`10.20.0.99` instead of `10.20.0.2`). **Fix:** ``` R1(config)# no ip route 10.255.0.3 255.255.255.255 10.20.0.99 R1(config)# ip route 10.255.0.3 255.255.255.255 10.20.0.2 ``` A route is only as good as its next-hop. If ARP for the next-hop comes back `Incomplete`, the route is pointing at nothing, and the symptom is a silent timeout because R1 never gets a frame onto the wire. ## Ticket 3: "I get a different failure to R3 now" **Reported symptom:** "R1 used to time out reaching the R3 site. Now it fails faster and differently." **Success criterion:** R1 can ping `10.255.0.3`. Same target, different fingerprint. Look at the ping carefully: ``` R1# ping 10.255.0.3 source 10.255.0.1 U.U.U Success rate is 0 percent (0/5) ``` This time it is `U`, not dots. A router in the path actively sent back an ICMP unreachable. The packet is getting somewhere and being turned away, not vanishing into a black hole. That is a downstream problem now, not a local one. Find out how far it gets: ``` R1# traceroute 10.255.0.3 source 10.255.0.1 Tracing the route to 10.255.0.3 VRF info: (vrf in name/id, vrf out name/id) 1 10.20.0.2 4 msec 3 msec 3 msec 2 * !H * ``` It reaches R2 (`10.20.0.2`) cleanly at hop 1, then R2 itself returns `!H` (host unreachable). The fault is at or beyond R2\. Move to R2 and check the interfaces: ``` R2# show ip interface brief | exclude unassigned Interface IP-Address OK? Method Status Protocol Ethernet0/0 10.20.0.2 YES TFTP up up Ethernet0/1 10.30.30.1 YES TFTP administratively down down Loopback0 10.255.0.2 YES TFTP up up ``` There it is. `Ethernet0/1`, the link to R3, is **administratively down**. Someone shut it. The route that depends on it has nowhere to go: ``` R2# show ip route 10.255.0.3 % Subnet not in table ``` **Root cause:** R2 `Ethernet0/1` was administratively shut down. **Fix:** ``` R2(config)# interface Ethernet0/1 R2(config-if)# no shutdown ``` Re-test from R1 and you are finally back to `!!!!!`: ``` R1# ping 10.255.0.3 source 10.255.0.1 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 3/4/6 ms ``` The lesson worth keeping: `.....` sends you looking at the *local* forwarding decision and ARP on the router you are testing from; `U.U.U` sends you looking *downstream* for an interface or routing problem on another device. One character of ping output tells you which router to log into next. On this lab the two local faults on R1 came first by necessity, and only once they were gone could the downstream `U` even appear. ## Troubleshooting matrix Ping returns `.` (timeout) What it usually means Local forwarding decision or ARP failure; silent drop on the source router Command that confirms it `show ip route `, `show ip arp ` Ping returns `U` What it usually means A downstream router is actively rejecting (no route / interface down) Command that confirms it `traceroute`, then `show ip interface brief` on the last hop Interface `up/up` but `% Subnet not in table` What it usually means Layer 3 addressing fault (wrong mask or IP) Command that confirms it `show running-config interface` Next-hop ARP shows `Incomplete` What it usually means The next-hop address does not exist on the segment Command that confirms it `show ip arp`, then verify the route's next-hop `administratively down` What it usually meansInterface was shut Command that confirms it `show ip interface brief` ## Key takeaways - When end-to-end ping fails, do not guess. Start at the bottom of the stack and fix the most local, most upstream fault first; on a small network it is probably masking the others. - Read ping output as evidence. `.` versus `U` tells you whether to look local or downstream before you log into anything. - `up/up` rules out Layers 1 and 2\. It does not rule out a Layer 3 addressing mistake, and a wrong mask can hide every route on the box. - A route is only as good as its next-hop. If ARP for the next-hop is `Incomplete`, the route is pointing at nothing. This is the free preview of the troubleshooting series. The rest of the queue, by exam pillar, lives in the [PingLabz CCNA Labs library](https://www.pinglabz.com/labs/): Layer 2 and VLAN tickets, OSPF tickets, DHCP and NAT tickets, security tickets, and a multi-fault capstone outage. If you want the general method behind all of them, the build-lab [nf-11 Troubleshooting Layer Symptoms](https://www.pinglabz.com/ccna-lab-nf-11-troubleshooting-layer-symptoms/) walks through the seven-step escalation drill. ### HSRP States: The 6-State Machine with show standby Output URL: https://www.pinglabz.com/hsrp-state/ Last updated: 2026-08-01T19:35:28.000Z HSRP runs a small state machine on every router in a standby group, and the state a router is in tells you exactly what role it is playing and whether the group is healthy. When a gateway fails over and you are trying to work out why, the HSRP states are the first thing to read. This post walks through all six states, what moves a router between them, and how to read the state out of `show standby` on Cisco IOS XE. For the cluster overview, see the [FHRP complete guide](https://www.pinglabz.com/fhrp/). For an analogous protocol with its own state machine, see the [BGP pillar](https://www.pinglabz.com/bgp/). ## What HSRP is doing in the first place HSRP (Hot Standby Router Protocol) gives a set of routers a shared virtual IP and virtual MAC. Hosts use the virtual IP as their default gateway. One router actively forwards for that virtual IP; another stands by ready to take over. The state machine is how the routers negotiate who is active, who is standby, and who is just watching. ## The six HSRP states Initial #1 What it means HSRP is not running yet on this interface - just configured, or the interface just came up. The starting point. Learn #2 What it means The router does not know the virtual IP and has not heard from the active router. Only happens when the virtual IP is not configured locally and must be learned from a Hello. Listen #3 What it means The router knows the virtual IP and is listening to Hellos. It is neither active nor standby - it knows both of those roles are taken and is watching in case one opens up. Speak #4 What it means The router is sending Hellos and actively participating in the election for active and standby. Standby #5 What it means The router is the designated backup. It is next in line and will take over if the active router fails. Active #6 What it means The router is forwarding traffic for the virtual IP and answering ARP for the virtual MAC. This is the working role. In a healthy two-router group, the steady state is one router in **Active** and one in **Standby**. Any third or fourth router in the group sits in **Listen**. Initial, Learn, and Speak are transitional - you only catch a router in those during startup or a failover. ## The normal startup sequence When you bring up HSRP on two routers, each one walks the states: **Initial** \- HSRP starts. The router moves on as soon as the interface is up and the config is read. **Listen** \- the router knows the virtual IP (because it is configured locally). It listens to Hellos to learn who else is in the group. If the virtual IP were *not* configured locally, it would pass through Learn first to discover the IP from a Hello. **Speak** \- after the active timer expires without hearing a higher-priority active router, the router starts sending Hellos and contests the election. Both routers in a new group reach Speak and exchange Hellos. **Standby / Active** \- the election resolves. The router with the higher priority (or higher IP, on a priority tie) becomes Active. The other becomes Standby. ## What decides the election: priority and preemption Two knobs control which router ends up Active. `standby 10 priority 110` Higher priority wins the election. Default is 100. `standby 10 preempt` Lets a higher-priority router *take over* from a lower-priority active router when it comes online. Without preempt, whoever got Active first keeps it. This is the single most important HSRP behavior to internalize: **without `preempt`, priority only matters at election time.** If your intended-primary router reboots, the backup becomes Active, and when the primary comes back it sees an active router already and settles into Standby - even though it has the higher priority. It will not reclaim the Active role until you configure `preempt`. ## Minimum config and the resulting states ``` ! R1 - intended primary interface GigabitEthernet0/0 ip address 10.20.0.2 255.255.255.0 standby version 2 standby 10 ip 10.20.0.1 standby 10 priority 110 standby 10 preempt ! ! R2 - intended backup interface GigabitEthernet0/0 ip address 10.20.0.3 255.255.255.0 standby version 2 standby 10 ip 10.20.0.1 standby 10 priority 100 ``` After convergence: R1 reaches **Active** (priority 110), R2 reaches **Standby** (priority 100). Because R1 has `preempt`, if R1 ever reboots and comes back, it reclaims Active. R2 does not need preempt - it is never trying to take a role away from a higher-priority router. ## Reading the state on IOS XE The quick view across all groups: ``` R1# show standby brief P indicates configured to preempt. | Interface Grp Pri P State Active Standby Virtual IP Gi0/0 10 110 P Active local 10.20.0.3 10.20.0.1 ``` The detailed view, when you need timers and the why: ``` R1# show standby GigabitEthernet0/0 10 GigabitEthernet0/0 - Group 10 (version 2) State is Active 8 state changes, last state change 02:14:51 Virtual IP address is 10.20.0.1 Active virtual MAC address is 0000.0c9f.f00a Hello time 3 sec, hold time 10 sec Preemption enabled Active router is local Standby router is 10.20.0.3, priority 100 (expires in 9.872 sec) Priority 110 (configured 110) ``` Three things to read: - **State is Active.** The role this router is playing. Cross-check the partner shows Standby. - **state changes count.** A high number that keeps climbing means the group is flapping - usually a tracked interface bouncing or a Hello/hold-timer mismatch. - **Standby router ... expires in N sec.** A countdown that resets means Hellos are arriving. If it hits zero, the standby is declared dead and the group reconverges. ## Interface tracking: states that change on purpose HSRP would be nearly useless if it only failed over when a router died. The real value is failing over when a router loses its *upstream* path. Object tracking does this: ``` track 1 interface GigabitEthernet0/1 line-protocol ! interface GigabitEthernet0/0 standby 10 track 1 decrement 20 ``` If R1's upstream interface Gi0/1 goes down, track object 1 fails, and R1's HSRP priority drops by 20 (from 110 to 90). Now R2 at priority 100 is higher. With preempt configured, R2 takes Active. R1 transitions Active to Speak to Standby. The state change you see in the logs is the visible symptom of a tracked uplink failing. ## Common state problems Both routers show Active (split brain) The two routers cannot hear each other's Hellos - VLAN/trunk problem, ACL blocking HSRP, or a group-number mismatch. Each thinks it is alone. Primary router stays Standby after a reboot No `preempt` on the primary. Priority alone does not reclaim the role. Router stuck in Speak, never reaches Standby/Active Hellos not being exchanged - one-way connectivity, or an authentication mismatch between group members. State changes climbing constantly A flapping tracked interface, or Hello/hold timers that differ between the two routers. Router sits in Listen and never contests Normal if it is a third router in the group. A problem only if it is one of the two you expected to be Active/Standby. The first entry in that list deserves a closer look, because each router looks entirely healthy on its own and only the two outputs side by side expose the fault. [Both routers claiming Active at the same time](https://www.pinglabz.com/hsrp-troubleshooting-flapping-dual-active/) walks a captured dual-active session from symptom through to the state machine settling again. ## Key takeaways HSRP has six states: Initial, Learn, Listen, Speak, Standby, Active. The healthy steady state is one Active, one Standby, and any extras in Listen. The election is decided by priority, and a higher-priority router only reclaims the Active role if `preempt` is configured - this is the single most common HSRP surprise. Interface tracking deliberately drops priority so a router that loses its uplink hands off gracefully. Read the live state with `show standby brief`, and watch the state-change counter for flapping. For the FHRP cluster, see the [FHRP pillar](https://www.pinglabz.com/fhrp/). ### Cisco Wireless LAN Controller: What It Does and How It Works URL: https://www.pinglabz.com/cisco-wireless-lan-controller/ Last updated: 2026-06-13T20:07:48.000Z A wireless LAN controller is the device that turns a pile of access points into a managed wireless network. Without one, every AP is an island - configured individually, managing its own RF, with no coordinated handoff between them. With one, the APs become thin extensions of a central brain that handles configuration, security, RF management, and client mobility. This post explains what a Cisco WLC does, the architecture it sits in, and how the modern Catalyst 9800 generation differs from the AireOS controllers it replaced. For the cluster overview, see the [Cisco Wireless complete guide](https://www.pinglabz.com/wireless/). ## What problem the WLC solves Imagine a building with 60 access points and no controller. Each AP is configured by hand. Each AP picks its own channel and power, with no awareness of its neighbors, so they interfere. A client walking down a hallway has to fully re-authenticate every time it moves to a new AP. Security policy has to be applied 60 times. Firmware is upgraded 60 times. The WLC collapses all of that to one. It is the single point of configuration, the single security policy enforcement point, the coordinator of RF across all APs, and the anchor that makes roaming seamless. The APs become "lightweight" - they handle radio and forwarding, the controller handles everything that needs a network-wide view. ## The split-MAC architecture Cisco's controller-based wireless uses what is called a split-MAC model. The 802.11 MAC-layer functions are split between the AP and the controller. Transmitting and receiving 802.11 frames Configuration of every AP Beacons and probe responses Authentication and security policy Encryption / decryption RF management (channel, power) across all APs Buffering for power-save clients Client mobility and roaming coordination Frame queueing by priority Rogue AP detection, RRM, monitoring The dividing line is timing. Anything that must happen in microseconds stays on the AP. Anything that benefits from a network-wide view moves to the controller. ## CAPWAP: the tunnel between AP and controller An AP and its controller talk over CAPWAP - Control And Provisioning of Wireless Access Points. CAPWAP is two tunnels in one: - **CAPWAP control** (UDP 5246) - configuration, management, and control messaging between AP and WLC. Always encrypted (DTLS). - **CAPWAP data** (UDP 5247) - client traffic tunneled from the AP back to the controller. Optionally encrypted. In the default centralized mode, a client's traffic is tunneled inside CAPWAP from the AP all the way to the WLC, and the WLC puts it onto the wired network. The AP itself does not bridge client traffic directly onto its local switchport. This matters for design: the WLC becomes a traffic aggregation point, so its placement and capacity matter. ## The AP join process When a lightweight AP boots, it has to find and join a controller: 1. **Get an IP** \- the AP DHCPs an address on its management VLAN. 2. **Discover controllers** \- the AP builds a list of candidate WLCs via DHCP option 43, DNS (`CISCO-CAPWAP-CONTROLLER.localdomain`), a previously-remembered controller, or subnet broadcast. 3. **Select and join** \- the AP sends a CAPWAP join request to the best candidate; the WLC authenticates it (certificate exchange). 4. **Download** \- the AP receives its configuration and, if needed, a matching software image from the controller, then reboots into managed operation. The most common AP-join failure is discovery: DHCP option 43 not set, or the DNS record missing. If an AP "will not join the controller," that step is where to look first. ## AireOS vs Catalyst 9800: the generational shift Cisco has two controller generations, and knowing which one you are dealing with matters because the configuration model is completely different. Operating system AireOS (legacy) AireOS - a wireless-only OS Catalyst 9800 (current) IOS XE - the same OS as Catalyst switches/routers Hardware examples AireOS (legacy)5520, 8540, 3504 Catalyst 9800 (current) 9800-80, 9800-40, 9800-L, 9800-CL (virtual), embedded Config model AireOS (legacy) WLAN-centric, AireOS CLI/GUI Catalyst 9800 (current) Profiles and tags (IOS XE structured model) High availability AireOS (legacy)AP SSO Catalyst 9800 (current) SSO plus N+1, with faster failover Status AireOS (legacy) End-of-sale; being phased out Catalyst 9800 (current) The current and future platform The Catalyst 9800 is where new deployments go. Because it runs IOS XE, it brings the same operational model, programmability (NETCONF/YANG, model-driven telemetry), and patching story as the rest of the Catalyst family. The AireOS-to-9800 migration is a real project, not a swap, because the configuration model is restructured around profiles and tags. ## The 9800 configuration model: profiles and tags The Catalyst 9800 organizes configuration into reusable profiles, bound to APs through tags: - **WLAN profile** \- the SSID and its security settings. - **Policy profile** \- what happens to clients on that WLAN (VLAN, QoS, mobility). - **Policy tag** \- binds WLAN profiles to policy profiles. - **RF tag** \- the RF profiles (2.4/5/6 GHz) applied to a group of APs. - **Site tag** \- ties APs to a flex/local mode and an AP join profile. An AP gets three tags - policy, RF, and site - and those tags pull in everything else. Once you have the profiles and tags built, adding a new AP is just "assign these three tags." It is more abstract than the AireOS model but far more scalable. ## Deployment modes Local (centralized) Where client traffic goes Tunneled via CAPWAP to the WLC Best for Campus where the WLC is well-placed and central policy is wanted FlexConnect Where client traffic goes Switched locally at the AP's site Best for Branch offices - survives WAN outage, no traffic hairpin to HQ FlexConnect is the branch answer: the AP keeps serving clients even if the link back to the controller drops, because client traffic is bridged locally rather than tunneled to a distant WLC. ## Common gotchas AP will not join the controller Discovery failure - DHCP option 43 or the DNS CISCO-CAPWAP-CONTROLLER record is missing/wrong AP joins then immediately reboots in a loop Image mismatch - AP downloads a new image, reboots, loops. Check WLC software version vs AP support. Clients on a branch AP lose connectivity when the WAN drops AP is in Local mode tunneling to a central WLC. Use FlexConnect for branches. Roaming is slow / clients re-authenticate fully Fast-roaming (802.11r / OKC) not enabled, or clients roaming across mobility-group boundaries WLC links saturated Centralized mode aggregates all client traffic at the WLC - the controller and its uplink are a real capacity factor ## Key takeaways A Cisco wireless LAN controller centralizes everything about a wireless network that benefits from a network-wide view: configuration, security, RF management, and roaming. APs and the controller use a split-MAC model and talk over CAPWAP tunnels. The current platform is the Catalyst 9800, which runs IOS XE and uses a profile-and-tag configuration model; the older AireOS controllers are being phased out. For branches, FlexConnect keeps APs serving clients through a WAN outage. Get the AP-join discovery method right and most of the rest follows. For the wireless cluster, see the [Cisco Wireless pillar](https://www.pinglabz.com/wireless/). ### CCNA 200-301 Mega Lab: The All-in-One Hands-On Campus URL: https://www.pinglabz.com/ccna-mega-lab/ Last updated: 2026-08-02T03:05:53.000Z One 10-node CML topology that exercises the entire CCNA 200-301 hands-on blueprint: VLANs, STP, EtherChannel, HSRP, OSPF, NAT, DHCP, NTP and the security baseline, with real Cisco IOS XE captures. _This post is for paying subscribers only._ ### VLAN vs VXLAN: The L2 Overlay, Demystified URL: https://www.pinglabz.com/vlan-vs-vxlan/ Last updated: 2026-06-13T20:07:49.000Z VLAN and VXLAN sound like the same thing with a version number, and the names actively encourage that misreading. They are not versions of each other. A VLAN is a Layer 2 segmentation mechanism that runs inside a switched network. VXLAN is an encapsulation that carries Layer 2 segments across a Layer 3 network. One is a fence; the other is a shipping container. This post explains what each does, why VXLAN exists, and when you actually need it. For the L2 fundamentals, see the [VLAN and Layer 2 switching pillar](https://www.pinglabz.com/vlans-layer-2-switching/). ## The one-sentence version A **VLAN** partitions a switched network into separate broadcast domains using a 12-bit tag. A **VXLAN** wraps an Ethernet frame inside a UDP/IP packet so that a Layer 2 segment can be tunneled across a routed Layer 3 network, using a 24-bit identifier. VLANs and VXLANs frequently coexist. A common data-center design has VLANs at the server-facing edge and VXLAN carrying those segments across the fabric. They are not competitors; they operate at different scopes. ## Side-by-side What it is VLAN (802.1Q) L2 tag added to an Ethernet frame VXLAN (RFC 7348) L2 frame encapsulated in UDP/IP Identifier size VLAN (802.1Q)12-bit VLAN ID VXLAN (RFC 7348) 24-bit VNI (VXLAN Network Identifier) Max segments VLAN (802.1Q)4,094 usable VXLAN (RFC 7348)\~16 million Scope VLAN (802.1Q) Within a single L2 domain / switched fabric VXLAN (RFC 7348) Across any L3-routed network Transport VLAN (802.1Q) Rides directly on Ethernet VXLAN (RFC 7348) Rides on UDP port 4789 over IP Spanning tree dependency VLAN (802.1Q) Loop prevention relies on STP VXLAN (RFC 7348) Underlay is routed - no STP across the fabric Typical use VLAN (802.1Q) Campus and access-layer segmentation VXLAN (RFC 7348) Data-center fabrics, multi-tenant overlays, DCI ## Why VXLAN exists: three limits of VLANs ### 1\. The 4,094 ceiling The VLAN ID field is 12 bits. After reserving a couple, you get 4,094 usable VLANs. For a campus, that is plenty. For a cloud provider or a large multi-tenant data center where every customer wants their own isolated segments, 4,094 runs out fast. VXLAN's 24-bit VNI gives roughly 16 million segments - effectively unlimited for any realistic tenant count. ### 2\. VLANs cannot cross a Layer 3 boundary A VLAN is a Layer 2 construct. It lives within a switched domain. The moment traffic hits a router, the VLAN tag is stripped and the frame becomes a routed packet. You cannot extend VLAN 100 from one data center to another across the routed internet - not natively. VXLAN solves exactly this. By encapsulating the L2 frame in UDP/IP, it makes the segment portable across any IP network. VLAN 100 in Data Center A and "VLAN 100" in Data Center B can be the same Layer 2 segment, stitched together by VXLAN, even though there are routed hops in between. This is the basis of data-center interconnect (DCI) and stretched-cluster designs. ### 3\. Spanning tree does not scale A large flat L2 network depends on spanning tree for loop prevention, and spanning tree blocks links. In a big fabric, that means a lot of expensive bandwidth sitting idle, plus the blast radius of an L2 problem covers the whole domain. VXLAN runs over a routed underlay. The physical network between switches is pure Layer 3 - it uses a routing protocol (OSPF, IS-IS, or BGP) and ECMP, so every link forwards, and there is no spanning tree spanning the fabric. The L2 adjacency that endpoints see is an illusion created by the overlay; the real network underneath is all routed. ## How VXLAN actually moves a frame The component that does the work is the VTEP - VXLAN Tunnel Endpoint. A VTEP is the device (usually a switch, sometimes a server) that sits at the edge of the VXLAN overlay. 1. An endpoint sends a normal Ethernet frame into its VLAN. 2. The ingress VTEP maps that VLAN to a VNI, wraps the whole frame in a VXLAN header, then a UDP header (destination port 4789), then an IP header addressed to the remote VTEP. 3. The routed underlay forwards the resulting IP packet like any other packet - ECMP, normal routing, no STP. 4. The egress VTEP receives it, strips the VXLAN/UDP/IP encapsulation, recovers the original Ethernet frame, and delivers it into the matching VLAN on its side. The two endpoints believe they are on the same LAN segment. They are not - there are routed hops between them. The VTEPs maintain the illusion. ## The control plane: VXLAN needs one Early VXLAN used multicast flood-and-learn to discover which VTEP held which MAC address. It worked but did not scale well. Modern VXLAN deployments pair it with a control plane - almost always **EVPN** (Ethernet VPN), carried in BGP. With BGP EVPN, VTEPs advertise their known MAC addresses and host routes to each other via BGP rather than flooding to learn them. This is why you will almost always see "VXLAN" and "EVPN" together: VXLAN is the data-plane encapsulation, EVPN is the control plane that tells each VTEP where everything is. VXLAN moves the frames; EVPN distributes the map. ## When you need VXLAN, and when you do not Campus access network, a few hundred VLANs, single site Plain VLANs. VXLAN adds complexity you do not need. Modern data-center fabric (spine-leaf) VXLAN with BGP EVPN. This is the standard design. Need the same L2 segment in two data centers VXLAN for the DCI - VLANs physically cannot do this. Multi-tenant environment exceeding 4,094 segments VXLAN - the VLAN ID space is exhausted. Small or mid-size business, one building Plain VLANs. VXLAN is overkill below data-center scale. ## Common misconceptions "VXLAN replaces VLANs" No. VXLAN carries VLANs across L3\. Both exist in a VXLAN network - VLANs at the edge, VNIs in the fabric. "VXLAN is just a bigger VLAN" The bigger ID space is one benefit, but the real point is L3 transport and escaping spanning tree. "VXLAN means no broadcast domains" Each VNI is still a broadcast domain. VXLAN changes how the domain is transported, not that it exists. "You need multicast for VXLAN" Old flood-and-learn did. Modern BGP EVPN VXLAN does not. ## Key takeaways A VLAN is a Layer 2 segmentation tag that lives inside a switched network and tops out at 4,094 segments. VXLAN is an encapsulation that wraps Layer 2 frames in UDP/IP so segments can cross a routed Layer 3 network, scales to \~16 million VNIs, and escapes spanning tree. They are not versions of one thing - they coexist, with VLANs at the edge and VXLAN in the fabric. If you run a single-site campus, plain VLANs are correct. If you run a data-center fabric or need stretched L2 between sites, VXLAN with BGP EVPN is the modern answer. For the L2 cluster, see the [VLAN pillar](https://www.pinglabz.com/vlans-layer-2-switching/). ### Cisco Catalyst SD-WAN: The Rebrand, the Architecture, OMP URL: https://www.pinglabz.com/cisco-catalyst-sd-wan/ Last updated: 2026-06-13T20:07:50.000Z Cisco Catalyst SD-WAN is what used to be called Cisco SD-WAN, which used to be called Viptela. The name has changed three times in six years, and that churn is itself a source of confusion: engineers searching for "Viptela" and engineers searching for "Catalyst SD-WAN" are looking for the same product. This post clears up the naming, explains what Catalyst SD-WAN actually is, and walks through the architecture so the rebrand stops being a barrier to understanding it. For the broader technology, see the [SD-WAN complete guide](https://www.pinglabz.com/sd-wan/). ## The naming history, settled Pre-2017 NameViptela What changed Independent SD-WAN startup 2017 NameCisco acquires Viptela What changedBecomes "Cisco SD-WAN" 2020-2022 NameCisco SD-WAN What changed vEdge hardware phased out; cEdge (IOS XE) becomes the platform 2023 onward NameCisco Catalyst SD-WAN What changed Rebranded under the Catalyst umbrella to align with Catalyst switching and wireless It is all the same product line. The "Catalyst SD-WAN" name signals that the edge runs Cisco IOS XE (the same OS as Catalyst routers and switches) rather than the old Viptela OS. If you read older documentation that says "Cisco SD-WAN" or "Viptela," it is describing this product. The controller components kept their names through the rebrand. ## What Catalyst SD-WAN is Catalyst SD-WAN is Cisco's overlay SD-WAN solution: a software-defined WAN where a centralized controller plane manages policy and a distributed set of edge routers builds an encrypted overlay across whatever transport is available (MPLS, broadband, LTE, 5G). It is transport-independent, application-aware, and centrally orchestrated - the standard SD-WAN value proposition, delivered the Cisco way. The thing that distinguishes Catalyst SD-WAN from competitors is the IOS XE edge. Because the cEdge runs the same OS as a normal Catalyst router, you get the full IOS XE feature set - rich routing, QoS, security - on the same box that does the SD-WAN overlay. There is no separate "SD-WAN appliance" with a reduced feature set. ## The four-plane architecture Catalyst SD-WAN splits cleanly into four planes. Each has a named component. Management Component Catalyst SD-WAN Manager (formerly vManage) Role The single dashboard. Policy authoring, templates, monitoring, troubleshooting, software management. Orchestration Component Catalyst SD-WAN Validator (formerly vBond) Role Authenticates every device joining the fabric, brokers the initial control connections, and helps devices behind NAT find each other. Control Component Catalyst SD-WAN Controller (formerly vSmart) Role The brain. Distributes routing (via OMP) and policy to all edges. Edges never peer directly for control - they all talk to the Controller. Data ComponentcEdge routers (IOS XE) Role The edges themselves. They build the IPsec overlay tunnels and forward traffic. Note the rebrand renamed the controllers too: vManage, vBond, and vSmart are now Manager, Validator, and Controller. Older docs and the CLI still surface the v-names in places, so know both. ## OMP: the routing protocol of the fabric The piece that makes Catalyst SD-WAN tick is OMP - the Overlay Management Protocol. OMP is to the SD-WAN overlay what BGP is to the internet: it is how routes, TLOC (Transport Locator) information, and service routes are distributed. Every cEdge has an OMP session to the Controller. The edge advertises its local prefixes and its TLOCs (essentially "here is how to reach me, over these transports") up to the Controller. The Controller applies policy and reflects the information back down to other edges. No edge runs a full routing table of every other edge's details - the Controller is the route reflector for the entire fabric. This is why the Controller is the control plane: pull it out and existing tunnels keep forwarding for a while, but no new routing information propagates and policy changes stop. ## How an edge joins the fabric The bring-up sequence, simplified: 1. A new cEdge powers on with a minimal config and a certificate (or a one-time password / Plug-and-Play profile). 2. It reaches out to the Validator. The Validator authenticates the device and tells it where the Manager and Controllers are. 3. The edge forms control connections to the Controllers. 4. The Manager pushes the device's full configuration via a template or a configuration group. 5. The edge brings up IPsec data-plane tunnels to other edges, builds its OMP routing, and starts forwarding. The point of this flow is zero-touch provisioning. A branch router can be drop-shipped to a site, plugged in by someone non-technical, and configure itself from the Manager. The Validator is the trust anchor that makes that safe. ## Configuration model: templates and configuration groups You do not SSH into Catalyst SD-WAN cEdges and type config the way you would a normal router. You configure them from the Manager using one of two models: - **Feature templates / device templates** \- the original model. Feature templates define reusable config blocks (a VPN, an interface, an OSPF instance); device templates assemble them per device. - **Configuration groups** \- the newer, simpler model introduced to reduce the template sprawl. Group-based, with a more intuitive UI. Either way, the principle is the same: configuration is centralized, version-controlled in the Manager, and pushed - not typed per box. This is the operational shift that trips up engineers coming from traditional IOS: the router's running-config is an output of the Manager, not something you edit directly. ## Where Catalyst SD-WAN fits vs alternatives Catalyst SD-WAN Cisco-standardized enterprises wanting deep IOS XE feature parity, strong cloud on-ramps, and integration with Cisco security (Umbrella, SSE). Meraki SD-WAN Lean-IT organizations wanting the simplest possible cloud-managed experience; trades depth for ease. Fortinet / VMware / Versa Covered in the SD-WAN pillar - chosen on price, security-stack integration, or single-vendor SASE ambitions. Within the Cisco portfolio, the choice is usually Catalyst SD-WAN vs Meraki SD-WAN. Catalyst is the more capable, more configurable platform; Meraki is the simpler, more opinionated one. Larger and more complex networks tend to land on Catalyst. ## Common points of confusion "Is Viptela still a thing?" Yes, it is just renamed. Viptela = Cisco SD-WAN = Catalyst SD-WAN, same product line. "vEdge vs cEdge?" vEdge ran Viptela OS and is end-of-sale. cEdge runs IOS XE and is the current platform. "vManage / vBond / vSmart vs Manager / Validator / Controller?" Same components, post-rebrand names. Both appear in current docs. "Do I configure the routers directly?" No. Configuration is centralized in the Manager via templates or configuration groups. ## Key takeaways Cisco Catalyst SD-WAN is the rebranded, IOS XE-based evolution of Cisco SD-WAN (originally Viptela). It uses a four-plane architecture - Manager, Validator, Controller, and cEdge data plane - tied together by OMP, the overlay routing protocol. Edges join the fabric through a zero-touch flow anchored by the Validator, and they are configured centrally from the Manager rather than box-by-box. If you understand the four planes and OMP, the rebrand stops mattering: it is the same architecture under every name it has carried. For the broader SD-WAN cluster, see the [SD-WAN pillar](https://www.pinglabz.com/sd-wan/). ### How to Install Cisco CML 2.10 on VMware ESXi (Step-by-Step with Screenshots) URL: https://www.pinglabz.com/install-cisco-cml-2-10-esxi/ Last updated: 2026-07-04T23:06:52.000Z Cisco Modeling Labs (CML) is the official replacement for VIRL, and running it on a VMware ESXi host is the sweet spot for a home lab or team lab: you deploy it once, snapshot it, and reach the web UI from any browser on your network. This guide walks through a clean install of **CML 2.10** (build `2.10.0+build.13`) on ESXi, from uploading the OVA to logging into the workbench and starting your first node. Every step below is a real screenshot from the install, not a stock diagram. If you would rather watch it happen end to end, the full screen capture is near the end of this guide. The written version here is the one you will want open on a second monitor while you click through the wizard. ## What you need before you start CML ships as two downloads from your Cisco account: a controller OVA and a reference platform (refplat) ISO. The OVA is the appliance itself (an Ubuntu 24.04 controller). The refplat ISO carries the node images (IOSv, IOL-XE, NX-OS, ASAv, and so on) that your labs actually run. You need both. These are the exact files used in this build, so you know what you are looking for in the download portal: `cml2_p_2.10.0-13_amd64-17.ova` The CML 2.10 controller appliance (Personal edition OVA) `refplat-20260409-fcs.iso` Reference platform images, first customer ship (the main node images) `refplat-20260409-supplemental.iso` Supplemental reference platform images `refplat-20260326-ise.iso` ISE reference image (optional, only if you lab identity services) On the host side, here are the minimum requirements Cisco publishes for the CML VM. Treat these as a floor, not a target: the more cores and RAM you give it, the more nodes you can run at once. Hypervisor VMware ESXi 7.0 or later Memory 8 GB (idle CML uses \~1 GB; the rest is for your nodes) CPU 4+ physical cores, Intel, with VT-x and EPT enabled Disk 32 GB minimum, but plan for 100 GB or more Network 1 interface reachable from your browser Two things bite people here. First, CML uses **nested virtualization** (every node is itself a VM), so your processor must support VT-x and EPT, and some reference images only run on Intel. Second, the OVA's virtual hardware version is 13 for backwards compatibility, so plan to upgrade VM compatibility to the newest version your ESXi host supports (more on that below). Cisco officially requires ESXi 7.0+, so if you are on an older build, that is the first thing to address. ## Step 1: Upload the refplat ISO to your datastore Before you deploy anything, get the refplat ISO onto a datastore that the ESXi host can see. In the ESXi Host Client, open the datastore browser, create a folder (this build uses `data/CML-ISO/`), and upload the refplat ISO files. You will attach one of them to the VM's CD/DVD drive in Step 3. ![ESXi datastore browser showing the CML 2.10 OVA and refplat ISO files including refplat-20260409-fcs.iso and refplat-20260409-supplemental.iso](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_11.png) The CML-ISO folder on the datastore. Note the OVA (cml2\_p\_2.10.0-13\_amd64-17.ova) sitting alongside the refplat ISOs. ## Step 2: Deploy the CML OVA In the ESXi Host Client, go to **Virtual Machines** and click **Create / Register VM**. Choose *Deploy a virtual machine from an OVF or OVA file*. ![ESXi New virtual machine wizard, Select creation type, Deploy a virtual machine from an OVF or OVA file](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_17.png) Start the deployment wizard and pick the OVF/OVA option. Give the VM a name (here it is `Cisco-CML`) and drop in the OVA file. The wizard accepts the single `.ova` directly, no need to unpack it. ![Select OVF and VMDK files step with cml2_p_2.10.0-13_amd64-17.ova uploaded and VM named Cisco-CML](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_16.png) Name the VM and select cml2\_p\_2.10.0-13\_amd64-17.ova. Pick a datastore for the VM. An SSD-backed datastore makes a real difference here, because labs start faster and write-sensitive nodes (NX-OS in particular) boot more reliably on fast storage. ![Select storage step in ESXi showing a 1TB-SSD datastore selected for the CML VM](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_15.png) Put the VM on your fastest datastore. SSD is strongly recommended. On the deployment options screen, map the network to a port group your browser can reach, choose **Thin** provisioning to save space, and leave *Power on automatically* **unchecked**. This part matters: you must configure the VM hardware before its first boot. ![Deployment options in ESXi with PublicNetwork mapped to VM Network, Thin disk provisioning, and Power on automatically unchecked](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_14.png) Map the network, choose Thin, and leave the VM powered off for now. Review the summary and finish. The guest OS is Ubuntu Linux 64-bit, which is the CML controller's underlying OS. ![Ready to complete step showing product cml2_2.10.0-13_amd64-17, thin provisioning, and Ubuntu Linux 64-bit guest OS](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_13.png) Confirm the deployment. Do not refresh the browser while the OVA imports. ## Step 3: Configure the VM before first boot This is the step most rushed installs skip, and it is the one that determines whether your lab is usable. With the VM still powered off, open **Edit settings**. The OVA's defaults (a handful of vCPUs, 8 GB RAM, a 32 GB disk) are the bare minimum and are not appropriate for a real ESXi deployment. ![ESXi Edit settings for the Cisco-CML VM showing 4 vCPU, 8 GB memory, 32 GB hard disk, and CD/DVD drive](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_12.png) The OVA defaults. Bump CPU, memory, and especially disk before you boot. Here is what to set, and why each one matters: VM hardware compatibility Upgrade to the latest version your ESXi host supports. The OVA ships at version 13 for compatibility, but newer features depend on a newer version. vCPU As many physical cores as you can spare. Set *cores per socket* so the socket count matches your host's physical CPUs. CPU shares High. Also tick *hardware virtualization* (expose VT-x to the guest) and *performance counters*. Memory 16 GB or more if you have it. Reserve all guest memory (lock it) so ESXi does not swap or reclaim it out from under running nodes. Hard disk Grow it to 100 GB or more **before first boot**. CML resizes its filesystem to the initial disk size on first boot, and 10 GB of the default 32 GB is reserved for the OS. CD/DVD Drive 1 Point it at the refplat ISO you uploaded in Step 1 and check *Connect at power on*. Network adapter Connect at power on, and enable DirectPath I/O. Latency Sensitivity Under VM Options, set to High. There is one host-level change that catches almost everyone the first time they try to bridge a lab to the real network. For the CML VM's interface to pass lab traffic out to your LAN, the vSwitch and port group it uses must allow it. On each port group or vSwitch the CML VM touches, set **Promiscuous Mode = Accept** and **Forged Transmits = Accept**. On the host, set `Net.ReversePathFwdCheckPromisc = 1` in Advanced System Settings. If you skip this, the controller UI still works, but bridged external connectors will silently drop traffic. ## Step 4: First boot and the setup wizard Now power on the VM. CML comes up through a GRUB menu and boots the `CML2` entry automatically. ![GNU GRUB boot menu for Cisco Modeling Labs CML2 GNU/Linux on first power on](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_10.png) The GRUB menu on first boot. It auto-selects CML2 in a few seconds. On first boot CML drops you into a text setup wizard that walks through the controller's identity and networking. First, the hostname: ![CML setup wizard prompting for the system hostname, set to cml-controller](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_9.png) Set a hostname. cml-controller is the conventional default. Next, the primary interface addressing. DHCP is fine for a first install and the easiest path to a working UI. You can pin a static address later from the Cockpit system management UI on TCP/9090, or switch to static here if your lab network expects it. ![CML setup wizard IPv4 configuration screen with DHCP selected for the primary network interface](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_8.png) Leave it on DHCP unless your network requires a static address. Then you choose optional services. CML 2.10 exposes three: **OpenSSH** (SSH to the controller on port 1122), **PATty** (port forwarding straight to lab nodes), and the new **MCP server**, which lets LLM tooling drive CML programmatically. Enable what you need; they can all be toggled later in Cockpit. ![CML optional services selection showing OpenSSH on port 1122, PATty lab node port forwarding, and MCP server for LLMs](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_7.png) Optional services in CML 2.10, including the new MCP server for LLMs. The wizard also has you set two passwords: one for the `sysadmin` system account (OS and Cockpit) and one for the `admin` CML account (the web UI). Keep them distinct and write them down. Finally, you confirm the whole configuration before it commits. ![CML configuration confirmation screen showing Standalone All-in-One deployment, cml-controller hostname, sysadmin and admin users, ISO attached, DHCP enabled](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_6.png) Confirm the build: Standalone All-in-One, ISO attached, DHCP, services enabled. Once you confirm, CML copies the reference platform images off the ISO and onto its local disk. This is required from CML 2.3 onward; nodes run from local disk, not the mounted ISO. The wizard warns you it takes roughly ten minutes. ![CML notice that reference platform images will be copied from the ISO to the VM disk, taking about 10 minutes](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_5.png) CML copies the node images from the ISO to local disk on first setup. ![CML copying reference platform image objects from the CD-ROM ISO to the local filesystem during initial setup](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_4.png) The copy in progress. Let it finish before you do anything else. ## What is actually on the reference platform ISO Once the copy finishes, those node images become the building blocks for every lab. If you mount the refplat volume directly, you will see a `virl-base-images` directory with one folder per image, and people often search for these exact folder names when they are hunting for a specific version. Here is what shipped on the 2026 reference platform used in this build, and what each one is good for: `iol-xe-17-18-02` Platform IOS XE on Linux (IOL-XE) What you would lab with it Lightweight IOS XE router, fast boot, ideal for large routing topologies `iol-xe-serial-4eth-17-18-02` Platform IOL-XE (serial, 4x Ethernet) What you would lab with it IOL-XE variant with serial console and four Ethernet ports `ioll2-xe-17-18-02` PlatformIOL Layer 2 (IOL-L2) What you would lab with it Lightweight switch image for VLAN, STP, and trunking labs `iosv-159-3-m12` PlatformIOSv (IOS 15.9(3)M) What you would lab with it Classic IOS router for CCNA/CCNP routing `iosvl2-2020` PlatformIOSvL2 What you would lab with it Classic IOS Layer 2 switch `csr1000v-17-03-08a` PlatformCSR 1000v What you would lab with it IOS XE virtual router for WAN and VPN labs `cat8000v-17-18-02` PlatformCatalyst 8000V What you would lab with it IOS XE routing and SD-WAN edge `cat9000v-uadp-17-18-02` PlatformCatalyst 9000v (UADP) What you would lab with it Catalyst 9000 switching with UADP ASIC simulation `cat9000v-q200-17-18-02` PlatformCatalyst 9000v (Q200) What you would lab with it Catalyst 9000 switching with Q200 ASIC simulation `nxosv9300-10-6-2-f` PlatformNexus 9300v (NX-OS) What you would lab with it Data center switching, VXLAN, vPC `iosxrv9000-26-1-1` PlatformIOS XRv 9000 What you would lab with it Service provider routing, MPLS, segment routing `xrd-26-1-1` PlatformXRd What you would lab with it Containerized IOS XR, lighter than XRv 9000 `asav-9-24-1` PlatformASAv 9.24 What you would lab with it Adaptive Security Appliance for firewall and NAT labs `wireless-ap-24-04-20260409` PlatformCatalyst wireless AP What you would lab with it Wireless labs with the Catalyst 9800 controller `wireless-client-24-04-20260409` PlatformWireless client What you would lab with it Simulated client to associate to the AP `alpine-wanem-3-23-3` PlatformAlpine WAN emulator What you would lab with it Inject latency, loss, and jitter between nodes `alpine-trex-3-23-3` PlatformAlpine + TRex What you would lab with it High-rate traffic generation `alpine-desktop-3-23-3` PlatformAlpine desktop What you would lab with it GUI host for testing reachability from an end user view `alpine-base-3-23-3` PlatformAlpine base What you would lab with it Tiny Linux host for pings, traceroutes, and DHCP clients `ubuntu-24-04-20260307` PlatformUbuntu 24.04 What you would lab with it Full Linux host for services, automation, and tooling `server-tcl-17-0` PlatformTCL server What you would lab with it Minimal scriptable server utility node ## Step 5: Log in and verify the install When setup completes, the VM console shows a login banner with the version (here, `CML 2.10.0+build.13` on Ubuntu 24.04) and, most importantly, the URLs to reach the web UI and the Cockpit system console on port 9090. ![CML console login banner showing CML 2.10.0 build 13, Ubuntu 24.04, and the web UI access URL on the assigned IP address](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_3.png) The console banner tells you the version and the exact URL to browse to. Browse to that address in Chrome or Firefox (CML's UI is HTML5 and is tested on the latest two versions of both) and you get the Cisco Modeling Labs login. Sign in with the `admin` account you created in the wizard. ![Cisco Modeling Labs web UI login page in a browser](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_2.png) The CML web UI login. Use the admin credentials from setup. The first time you log in, CML will prompt you to register a license. Personal, Personal Plus, Enterprise, and Education licenses all use a registration token from your Cisco Smart Account; paste it in and the node limits unlock. After that, drop a node onto the canvas, start it, and open its console. If you see the node boot and hand you a prompt, your install is done and working. ![CML workbench with a running router node R1 and its console output confirming the node booted successfully](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/Cisco-Modeling-Labs-CML-Install_1.png) A running node in the workbench. That booting console is your confirmation. ## Troubleshooting the common failures VM will not boot or panics early VT-x/EPT not exposed to the guest. Enable hardware virtualization on the VM's CPU settings and confirm VT-x is on in the host BIOS. Nodes will not start (or hang at boot) Nested virtualization is missing, or you are on a non-Intel host with Intel-only images. Confirm VT-x/EPT and use Intel hardware for full support. Can reach the UI, but bridged lab traffic dies Promiscuous Mode and Forged Transmits are not set to Accept on the vSwitch/port group. Fix both, then disable and re-enable promiscuous mode or reboot the host. Disk fills up fast You booted with the default 32 GB disk. Grow the disk (ideally you do this before first boot) and add a storage volume from Cockpit. No node images available The refplat ISO was not attached at first boot, so the copy never ran. Attach the ISO and re-run the reference platform copy. ## Build your first lab With CML up, the fastest way to get value is to build something you are studying. A two-router IOSv or IOL-XE topology is all you need to work through our [OSPF](https://www.pinglabz.com/ospf/), [BGP](https://www.pinglabz.com/bgp/), or [EIGRP](https://www.pinglabz.com/eigrp/) guides hands on. Drop in a couple of IOSvL2 or IOL-L2 switches and you can follow the [VLAN](https://www.pinglabz.com/vlans-layer-2-switching/) and [spanning tree](https://www.pinglabz.com/spanning-tree-protocol/) walkthroughs. The `asav-9-24-1` image pairs perfectly with our [Cisco ASA](https://www.pinglabz.com/cisco-asa/) cluster, and the wireless images line up with the [Catalyst 9800 wireless](https://www.pinglabz.com/wireless/) material. Every show-command output you see in those articles was captured on CML labs exactly like the one you just built. ## Frequently asked questions ### Is Cisco Modeling Labs free? There is a free tier, CML-Free, which runs a small number of nodes at no cost. The Personal and Personal Plus tiers (a modest annual fee) raise the node limit and are what most home labbers buy. Enterprise and Education are the larger, clustered offerings. Whichever you have, the install process on ESXi is identical; only the license token and node limits differ. ### What is the difference between the OVA and the refplat ISO? The OVA is the CML controller appliance, the thing you deploy as a VM. The refplat ISO carries the node images (routers, switches, firewalls) that your labs run. You install the OVA once and attach the ISO so CML can copy those images to its local disk. You can later download newer refplat ISOs to add updated node versions without reinstalling CML. ### Do I really need ESXi 7.0 or later? That is Cisco's supported minimum for CML 2.10, and it is what you should target. The OVA's virtual hardware version is 13 for backwards compatibility, but you should upgrade the VM's hardware compatibility to the newest version your host supports so you do not lose features. Older ESXi builds may appear to work but are not supported. ### Why must I set the disk size before the first boot? CML automatically resizes its filesystem to match the initial disk size the first time it boots. If you boot with the 32 GB default and grow the disk afterward, the extra space is not automatically used by the root filesystem; you have to add a storage volume or expand it manually from Cockpit. Sizing it to 100 GB or more before the first boot avoids that entirely. ### Can I run CML on AMD or Apple Silicon? CML is only fully supported on Intel processors. The application services run fine on AMD, but several reference platform images are Intel-only, so AMD support is best effort. Apple Silicon (M-series) is not supported for the VMware Fusion path. For a dependable lab, use an Intel host with VT-x and EPT. ## Key takeaways If you take one thing away from this guide, it is to **configure the VM before you power it on**: upgrade the hardware compatibility, give it real CPU and memory with high shares and reserved RAM, grow the disk to 100 GB or more, attach the refplat ISO, and set promiscuous mode plus forged transmits on the vSwitch. Get those right and the rest of the install is a ten-minute wizard followed by a license token. From there you have a full Cisco lab (IOS, IOS XE, NX-OS, IOS XR, ASA, and wireless) running on a single ESXi host, ready for whatever you are studying next. This walkthrough was based on a real CML 2.10 install; for Cisco's authoritative reference, see the [CML 2.10 Installation Guide](https://developer.cisco.com/docs/modeling-labs/cml-installation-guide/?ref=pinglabz.com) and the [Deploying the OVA on ESXi Server](https://developer.cisco.com/docs/modeling-labs/deploying-the-ova-on-esxi-server/?ref=pinglabz.com) page. Next in the lab-setup series: [How to Install the Missing Docker Container Images in Cisco Modeling Labs 2.10](https://www.pinglabz.com/install-cml-2-10-docker-container-images/). ### MSTP: Multiple Spanning Tree Protocol Explained (802.1s) URL: https://www.pinglabz.com/multiple-spanning-tree-protocol/ Last updated: 2026-06-13T20:07:50.000Z Multiple Spanning Tree Protocol (MSTP, IEEE 802.1s) exists to solve a scaling problem that the per-VLAN spanning tree variants created. Rapid-PVST+ runs one independent spanning tree instance per VLAN. That is fine with 10 VLANs. With 500 VLANs it is 500 state machines, 500 sets of BPDUs, and a switch CPU that spends real cycles just maintaining spanning tree. MSTP fixes this by mapping many VLANs onto a small number of spanning tree instances. This post explains how it works, how to configure it on Cisco IOS XE, and the one concept (the region) that trips everyone up. For the cluster overview, see the [Spanning Tree Protocol pillar](https://www.pinglabz.com/spanning-tree-protocol/). For the L2 fundamentals underneath, see the [VLAN and L2 switching pillar](https://www.pinglabz.com/vlans-layer-2-switching/). ## The problem MSTP solves Spanning tree's job is to build a loop-free topology. But here is the thing: if you have 500 VLANs that all run over the same physical switches and the same physical links, you do not need 500 different loop-free topologies. You need maybe two or three - enough to load-balance traffic across your redundant uplinks, and no more. Common Spanning Tree (CST, 802.1Q original) Instances1 for all VLANs Trade-off Zero load balancing - one uplink always blocks for everything Rapid-PVST+ (Cisco) Instances1 per VLAN Trade-off Full per-VLAN control but heavy CPU/BPDU overhead at scale MSTP (802.1s) Instances A few, with VLANs mapped to them Trade-off Load balancing where you want it, without the per-VLAN overhead MSTP is the middle path. You define a handful of instances, map VLAN ranges to them, and get load balancing across uplinks without paying the per-VLAN tax. ## The core concept: MST instances and VLAN mapping An MST instance (MSTI) is one spanning tree. You decide how many you want and which VLANs belong to each. A typical design: - **Instance 1** \- all the "left uplink" VLANs (say VLANs 10, 20, 30) - **Instance 2** \- all the "right uplink" VLANs (say VLANs 40, 50, 60) You make one switch the root for instance 1 and a different switch the root for instance 2\. Now instance 1's traffic prefers the left uplink and instance 2's traffic prefers the right uplink. Two instances, full load balancing, regardless of whether you have 6 VLANs or 600. VLANs not explicitly mapped land in instance 0, the IST (Internal Spanning Tree), which always exists. ## The concept that trips everyone up: the MST region For two switches to share MST instances, they must be in the same MST region. Two switches are in the same region only if all three of these match exactly: Region name Configured string, case-sensitive Revision number Configured integer VLAN-to-instance mapping table Every VLAN mapped to the same instance on both switches If any one of these differs - even one VLAN mapped to a different instance, even a typo in the name - the two switches are in *different* regions. They will still interoperate, but they treat each other as a single boundary and exchange only the IST. All your carefully designed per-instance load balancing silently stops at the region boundary. This is the number one MSTP misconfiguration. A switch added later with a slightly different VLAN-to-instance map becomes its own one-switch region, and nobody notices until traffic takes a strange path. ## Configuration on Cisco IOS XE Switching a Catalyst to MST mode and defining the region: ``` spanning-tree mode mst ! spanning-tree mst configuration name PINGLABZ-REGION revision 1 instance 1 vlan 10,20,30 instance 2 vlan 40,50,60 exit ``` That config block must be **identical** on every switch in the region. Same name, same revision, same instance mappings. Then set the root bridges so the two instances load-balance. On the switch you want as root for instance 1: ``` spanning-tree mst 1 root primary spanning-tree mst 2 root secondary ``` And on the switch you want as root for instance 2, the mirror image: ``` spanning-tree mst 1 root secondary spanning-tree mst 2 root primary ``` Now instance 1 roots on the first switch, instance 2 roots on the second, and the redundant uplinks both carry traffic instead of one sitting idle in blocking. ## Verifying the region and the instances Confirm the region is formed correctly: ``` SW1# show spanning-tree mst configuration Name [PINGLABZ-REGION] Revision 1 Instances configured 3 Instance Vlans mapped -------- --------------------------------------------------------------------- 0 1-9,11-19,21-29,31-39,41-49,51-59,61-4094 1 10,20,30 2 40,50,60 ------------------------------------------------------------------------------- ``` Run this on every switch and compare the output. If the name, revision, or any mapping differs, you have an accidental region boundary. Then check an instance's topology: ``` SW1# show spanning-tree mst 1 ##### MST1 vlans mapped: 10,20,30 Bridge address 0050.7989.bbbb priority 24577 (24576 sysid 1) Root this switch for MST1 Interface Role Sts Cost Prio.Nbr Type ---------------- ---- --- --------- -------- ------------------------------- Te1/1/1 Desg FWD 2000 128.25 P2p Te1/1/2 Desg FWD 2000 128.26 P2p ``` "Root this switch for MST1" confirms the root placement. Cross-check `show spanning-tree mst 2` on the other switch to confirm the load-balancing split is real. ## The IST and the CIST Two more terms you will hit in the docs: - **IST (Internal Spanning Tree)** \- instance 0\. It always exists and carries all unmapped VLANs. It is also what represents the whole region to the outside world. - **CIST (Common and Internal Spanning Tree)** \- the single spanning tree that spans across region boundaries and ties MST regions together with any legacy STP/RSTP switches. From the outside, an entire MST region looks like one big virtual bridge in the CIST. The practical takeaway: inside a region you think in MST instances; at the boundary the whole region collapses to a single node in the CIST. This is why MSTP interoperates cleanly with older RSTP switches - they just see the region as one RSTP bridge. ## MSTP vs Rapid-PVST+: which to run Rapid-PVST+ Small-to-medium networks, under \~50-75 VLANs, where per-VLAN control is convenient and CPU overhead is a non-issue. Cisco default; simplest mental model. MSTP Large campus networks, hundreds of VLANs, multi-vendor environments (MSTP is an IEEE standard; PVST+ is Cisco-proprietary). The right answer at scale. ## Common gotchas Load balancing not happening; one uplink always blocks Switches in different regions (name/revision/mapping mismatch). Diff `show spanning-tree mst configuration` across all switches. A new switch behaves oddly after being added Its VLAN-to-instance map differs by even one VLAN, making it its own region. VLAN traffic follows a path you did not design That VLAN is unmapped and landed in instance 0 (IST) instead of the instance you expected. Convergence slower than expected A link is half-duplex, so MSTP treats it as shared media and skips the rapid handshake (MSTP inherits RSTP's point-to-point requirement). ## Key takeaways MSTP maps many VLANs onto a few spanning tree instances, giving you load balancing across redundant uplinks without the per-VLAN overhead of Rapid-PVST+. The make-or-break concept is the region: name, revision, and VLAN-to-instance mapping must be byte-identical on every switch, or the switches silently split into separate regions and your load-balancing design stops working. Configure it once, verify with `show spanning-tree mst configuration` on every switch, and MSTP scales to hundreds of VLANs cleanly. For the STP cluster, see the [Spanning Tree Protocol pillar](https://www.pinglabz.com/spanning-tree-protocol/). ### QoS on a Router: A Practical Cisco IOS XE Walkthrough URL: https://www.pinglabz.com/qos-on-a-router/ Last updated: 2026-06-13T20:07:51.000Z QoS on a router is where quality of service stops being theory and becomes a config you actually type. The classification, marking, and queueing concepts are the same whether you are on a Catalyst 9500 or a branch ISR, but the router is where the WAN bottleneck lives, and the WAN bottleneck is where QoS earns its keep. This post is the practical version: how to build a working QoS policy on a Cisco router, what each piece does, and how to confirm it is doing something. For the full picture, see the [QoS complete guide](https://www.pinglabz.com/qos/). ## Why the router is where QoS matters Inside the LAN, bandwidth is cheap. Switch ports are 1 Gbps minimum, often 10 Gbps, and congestion is rare. QoS on a LAN switch mostly means "trust the markings and have sane queues." The router's WAN interface is the opposite. It is the slow link. A branch with a 50 Mbps internet circuit and a gigabit LAN has a 20-to-1 speed mismatch at the router. When the LAN sends more than 50 Mbps toward the WAN, packets queue up inside the router, and the router decides who waits and who goes first. That decision is QoS. Without it, a large file upload and a voice call compete equally, and the voice call loses. ## The MQC: three commands, one model Cisco IOS XE QoS is built on the Modular QoS CLI (MQC). Every QoS policy, however complex, is three building blocks: Class map Command`class-map` What it does Identifies traffic. "What is this packet?" Policy map Command`policy-map` What it does Decides what to do with each class. "How do I treat it?" Service policy Command`service-policy` What it does Applies the policy to an interface, in a direction. Class map answers what, policy map answers how, service policy answers where. Learn this and every QoS config you ever read decomposes cleanly. ## Step 1: Classify with class maps A class map matches traffic. You can match on DSCP markings (if the LAN already marked the traffic), on access lists, on protocols via NBAR, or on other criteria. ``` class-map match-any VOICE match dscp ef class-map match-any VIDEO match dscp af41 class-map match-any CRITICAL-DATA match dscp af31 match dscp cs3 class-map match-any SCAVENGER match dscp cs1 ``` This assumes the LAN switches already marked the traffic (the normal design: mark at the access edge, trust at the router). If your endpoints do not mark, the router can classify with NBAR instead: ``` class-map match-any VOICE match protocol rtp audio ``` NBAR is heavier on the CPU but works when you cannot trust upstream markings. Most production designs mark at the access switch and let the router trust DSCP. ## Step 2: Define treatment with a policy map The policy map says what happens to each class. The two main tools are a priority queue (for voice) and bandwidth guarantees (for everything else). ``` policy-map WAN-QOS class VOICE priority percent 20 class VIDEO bandwidth percent 25 class CRITICAL-DATA bandwidth percent 25 class SCAVENGER bandwidth percent 5 class class-default bandwidth percent 25 random-detect dscp-based ``` Reading this policy: - **VOICE gets `priority`** \- a strict low-latency queue (LLQ). Voice packets jump ahead of everything else. The `percent 20` also polices it: voice cannot exceed 20% of the link even with priority, so a flood of EF-marked traffic cannot starve the rest. - **VIDEO, CRITICAL-DATA, SCAVENGER, class-default get `bandwidth`** \- a minimum guarantee during congestion via CBWFQ (Class-Based Weighted Fair Queueing). When the link is not congested, any class can use more than its guarantee. The guarantee only kicks in when there is contention. - **class-default gets WRED** (`random-detect`) - drops packets early and randomly as the queue fills, which prevents TCP global synchronization (every flow backing off and ramping up in lockstep). ## Step 3: Shape, then apply Here is the part people miss. The policy map above only does something when the interface is congested. On a physical interface that runs below line rate, the queues never fill, and QoS never engages. If your WAN circuit is a 50 Mbps service delivered over a physical gigabit handoff (extremely common with metro-ethernet and broadband), the router thinks the interface is 1 Gbps. It will happily send 1 Gbps until the provider drops the excess - and the provider drops without QoS awareness, so your voice packets die in the carrier's queue, not yours. The fix is a hierarchical policy: a parent policy shapes traffic down to the real circuit rate, and the child policy (the one above) runs inside that shaped envelope. ``` policy-map WAN-PARENT class class-default shape average 50000000 service-policy WAN-QOS ! interface GigabitEthernet0/0/0 description WAN - 50 Mbps circuit on a 1G handoff service-policy output WAN-PARENT ``` Now the router shapes its own output to 50 Mbps. Because it is now the bottleneck, its queues fill, and the child policy's priority queue and bandwidth guarantees actually take effect. Shape to the real rate, then queue inside it. This is the single most important QoS design rule on a router. ## QoS direction: output is where the action is QoS is applied per-interface, per-direction. On a router WAN interface, `output` (egress) is where queueing happens, because queueing only makes sense when packets are leaving toward a slower link. You can apply policies inbound too, but inbound QoS is limited - you cannot queue traffic that has already arrived. Inbound policies are used for marking (set DSCP on ingress) or policing (drop/remark traffic that exceeds a rate on the way in). The classic split: mark or police inbound, queue outbound. ## Verifying it works The command that tells you whether QoS is doing anything: ``` Router# show policy-map interface GigabitEthernet0/0/0 GigabitEthernet0/0/0 Service-policy output: WAN-PARENT Class-map: class-default (match-any) 8204531 packets Service-policy : WAN-QOS Class-map: VOICE (match-any) 412904 packets Match: dscp ef (46) Priority: 20% (10000 kbps), burst bytes 250000 Priority Level 1: 412904 packets drop rate 0 bps Class-map: VIDEO (match-any) 88122 packets Match: dscp af41 (34) Queueing bandwidth 25% (12500 kbps) (total drops) 1208 (queue depth) 14 ``` What to read: - **VOICE drop rate.** Should be 0\. Any drops in the priority queue means voice is being lost - either you under-sized the priority percent or something is mismarked as EF. - **VIDEO total drops.** Some drops here are normal during congestion - that is the bandwidth guarantee working. A flood of drops means the class is undersized. - **Packet counts climbing.** If a class shows 0 packets, your classification is wrong - the traffic is not matching the class map. This is the most common QoS bug. ## A note on home and SMB routers The same model scales down. A small-business router (or a prosumer box running a real OS) does QoS the same way: classify, prioritize voice and video, shape to the actual circuit rate. The single highest-value QoS change on a home or SMB router is the shaper - set it to about 90-95% of your measured circuit speed so the router, not the ISP, is the bottleneck. That one setting is what stops a big upload from wrecking a video call. ## Common mistakes No shaper on a sub-rate circuit QoS configured but never engages; voice still choppy under load Policy applied inbound expecting queueing No effect; queueing only works outbound Class map matches nothing `show policy-map interface` shows 0 packets in the class Priority percent too high Non-voice classes starve; bulk traffic times out Trusting markings from an untrusted source Endpoints mark their own traffic EF and jump the queue ## Key takeaways QoS on a router is three MQC blocks: a class map to identify traffic, a policy map to treat it, a service policy to apply it. Voice gets a strict priority queue; everything else gets bandwidth guarantees. The rule that makes or breaks the whole thing is the shaper: on any circuit slower than the physical interface, shape to the real rate so the router becomes the bottleneck and its queues actually engage. Verify with `show policy-map interface` and watch the drop counters. For the full QoS cluster, see the [QoS pillar](https://www.pinglabz.com/qos/). ### BGP Weight: The Cisco-Only Path Attribute (and When to Use It) URL: https://www.pinglabz.com/bgp-weight/ Last updated: 2026-06-13T20:07:51.000Z BGP weight is the first tiebreaker in BGP's best-path selection algorithm. It is also the only attribute on the list that is Cisco-specific and never leaves the local router. Weight is the most powerful way to make a single router pick one path over another, and the most surgical because no other router in your network or in the rest of the internet knows or cares what you set it to. For the cluster overview, see the [BGP complete guide](https://www.pinglabz.com/bgp/). For where weight sits in the broader path-influence toolset (LOCAL\_PREF, MED, AS\_PATH prepending), see the [SD-WAN pillar](https://www.pinglabz.com/sd-wan/) for how SD-WAN solves similar problems differently. ## The 13-step path selection refresher BGP picks the best path by walking through a series of tiebreakers in order. The first tie that one path wins terminates the walk. Weight Step1 PreferHigher LOCAL\_PREF Step2 PreferHigher Locally originated Step3 Prefer Yes (via `network` or `redistribute`) AS\_PATH length Step4 PreferShorter Origin code Step5 Prefer IGP < EGP < Incomplete MED Step6 PreferLower (within same AS) eBGP over iBGP Step7 PrefereBGP wins IGP cost to BGP next-hop Step8 PreferLower Oldest path Step9 PreferOlder for eBGP Router ID of advertising peer Step10 PreferLower Cluster list length Step11 PreferShorter Neighbor IP address Step12 PreferLower Multipath Step13 Prefer Install multiple if eligible Weight is step 1\. It wins over everything below it. This is exactly the power of weight, and exactly the danger of weight. ## What weight is Weight is a 16-bit integer (0 to 65535) attached to each path in a local router's BGP table. It is set locally and never advertised to BGP peers. Other routers do not see it, do not act on it, and have no way to know it exists. The default weight is: - **0** for routes learned from any BGP peer (eBGP or iBGP) - **32768** for routes the local router originated via `network` statement or `redistribute` Locally originated routes win over learned routes by default because 32768 beats 0 at step 1 of the path selection algorithm. This is rarely the source of operational problems; it just confirms locally-known routes take precedence. ## When to use weight Weight is the right tool when: - You want to make a single router prefer one path over another - The decision should not propagate to any other BGP router in your network - You need the decision to override every other attribute (including LOCAL\_PREF) The canonical example: a dual-homed branch with two ISPs. The local router has two paths to every internet prefix. You want primary ISP to win. You set weight 200 on routes learned from primary ISP (anything > 0 works). All routes learned from primary ISP now have weight 200 and beat the weight-0 routes from the backup. Traffic uses primary unless primary's session drops, at which point the only remaining paths are weight-0 from backup. If you wanted this preference to apply to your *whole network* (not just the local router), you would use LOCAL\_PREF instead. LOCAL\_PREF propagates inside the AS via iBGP; weight does not. ## Setting weight: three ways ### Method 1: per-neighbor (blanket) ``` router bgp 65001 neighbor 203.0.113.1 remote-as 65100 neighbor 203.0.113.1 weight 200 ``` Every route learned from this neighbor gets weight 200\. Simple, blunt instrument. Useful when "everything from this peer is preferred" is what you mean. ### Method 2: route-map (per-prefix or per-pattern) ``` ip prefix-list CUSTOMER-NETS seq 10 permit 10.50.0.0/16 le 24 ! route-map FROM-PRIMARY permit 10 match ip address prefix-list CUSTOMER-NETS set weight 200 ! route-map FROM-PRIMARY permit 20 ! Catch-all: everything else gets weight 100 set weight 100 ! router bgp 65001 neighbor 203.0.113.1 remote-as 65100 neighbor 203.0.113.1 route-map FROM-PRIMARY in ``` This sets different weights for different prefixes from the same neighbor. Surgical control. Useful when "I want this peer for these prefixes specifically" rather than "this peer for everything." ### Method 3: based on AS\_PATH ``` ip as-path access-list 10 permit ^65100_ ! Routes originating in AS 65100 ! route-map PREFER-PEER permit 10 match as-path 10 set weight 200 ! router bgp 65001 neighbor 203.0.113.1 remote-as 65100 neighbor 203.0.113.1 route-map PREFER-PEER in ``` Sets weight based on the AS\_PATH of the route. Useful for "prefer paths that originate in a specific AS" or "deprioritize paths that traverse a specific AS." ## A worked example A branch router has two upstream ISP sessions: ISP-A on 203.0.113.1 (primary), ISP-B on 198.51.100.1 (backup). Both ISPs advertise a full BGP table. You want all internet-bound traffic to use ISP-A unless ISP-A is down. ``` router bgp 65001 neighbor 203.0.113.1 remote-as 65100 neighbor 203.0.113.1 weight 200 neighbor 198.51.100.1 remote-as 65200 neighbor 198.51.100.1 weight 100 ``` Both sessions are up. Every prefix has two paths in the BGP table: weight 200 via ISP-A, weight 100 via ISP-B. Weight 200 wins step 1 of best-path. All traffic exits via ISP-A. ISP-A's session drops. Every prefix now has only the weight-100 path via ISP-B in the BGP table. Traffic shifts to ISP-B. No reconfiguration needed. ISP-A comes back. Both paths are again in the table. Weight 200 wins. Traffic shifts back to ISP-A. ## Verifying weight ``` R1# show ip bgp 8.8.8.0 BGP routing table entry for 8.8.8.0/24, version 1452 Paths: (2 available, best #1, table default) 65100 15169 203.0.113.1 from 203.0.113.1 (203.0.113.1) Origin IGP, metric 0, localpref 100, weight 200, valid, external, best 65200 15169 198.51.100.1 from 198.51.100.1 (198.51.100.1) Origin IGP, metric 0, localpref 100, weight 100, valid, external ``` Two paths, weights 200 and 100\. The 200-weighted path is marked best. This is the BGP table view; `show ip route` will show the FIB entry derived from the chosen path. ## Weight vs LOCAL\_PREF: when to choose which I want this single router to prefer a path Weight I want my whole AS to prefer a path LOCAL\_PREF I want the preference to survive router reboots and propagate through iBGP LOCAL\_PREF I want fast, local-only control over a specific router's exit decision Weight I am working in a multi-vendor environment that includes non-Cisco routers LOCAL\_PREF (weight is Cisco-only) I want the decision to be invisible to other devices on the network Weight ## Common gotchas Set weight but path selection unchanged Weight was set on the wrong neighbor direction (need it on inbound route-map, not outbound). Check `neighbor X route-map Y in` vs `out`. Weight set via route-map applies but only to some prefixes Prefix-list match is too narrow. Add a catch-all `permit` clause at the end of the route-map to set a default weight for everything else. Weight works on this router but other routers in my AS still take the wrong path That is the point. Weight is local-only. Use LOCAL\_PREF to influence AS-wide decisions. Inherited a network where weight is set everywhere and behavior is confusing Weight is invisible to neighbors but locally powerful. Audit every router's `show ip bgp` output to map who is setting what. ## Key takeaways Weight is BGP's first tiebreaker, set locally, never propagated, Cisco-only, and the most powerful single-router path control available. Use it for dual-homed exits where you want a clean primary/backup pattern on one router. Use LOCAL\_PREF when you want the same preference to apply across your whole AS. Higher weight wins. Default is 0 for learned routes, 32768 for locally originated. Verify with `show ip bgp X.X.X.X` and look for the weight column in the path output. For broader BGP coverage, see the [BGP pillar](https://www.pinglabz.com/bgp/). ### QoS for VoIP Across an MPLS WAN URL: https://www.pinglabz.com/qos-for-voip/ Last updated: 2026-06-13T20:07:51.000Z Voice over IP only works if the network treats voice packets differently from everything else. Voice is small, regular, and intolerant of delay or loss. Bulk file transfer is large, bursty, and does not care if a packet shows up 200ms later. QoS is the set of techniques that lets one router carry both and keep the voice call sounding like a phone call. This post covers the QoS chain end-to-end across an MPLS WAN, what each component does, and the DSCP-to-CoS mapping that ties the LAN and WAN sides together. For QoS fundamentals at the LAN side, see the [QoS complete guide](https://www.pinglabz.com/qos/). For the MPLS WAN technology this rides over, see the [MPLS pillar](https://www.pinglabz.com/mpls/). ## The three things voice needs One-way latency Target < 150ms (ideally < 100ms) What breaks it Long queuing delays behind bulk traffic, especially on slow WAN links Jitter (variation in latency) Target< 30ms What breaks it Queuing depth varying as bulk traffic comes and goes Packet loss Target< 1% What breaks it Tail drops when a queue fills, or random drops in a congested link QoS is the toolkit to keep all three numbers inside the bounds. The G.711 voice codec, the most forgiving of the common codecs, starts to sound bad above 150ms one-way latency or 1% loss. More aggressive codecs (G.729, Opus low-rate) are less tolerant. ## The QoS chain across an MPLS WAN A voice packet leaves an IP phone, crosses the LAN, hits the WAN edge, traverses the MPLS L3VPN, exits at the remote PE, crosses the remote LAN, arrives at the destination phone. Each hop applies some piece of QoS. Phone egress Phone marks DSCP EF (decimal 46) on the voice packet. Signaling marked CS3. Access switch Trusts the DSCP from the phone via "auto qos voip" or explicit policy-map. Queues into the LLQ on uplink. WAN edge router (CE) Classifies traffic, applies LLQ for voice, CBWFQ for other classes, shapes overall to subscribed CIR. MPLS L3VPN ingress (PE) Maps customer DSCP to MPLS EXP/Traffic-Class bits (the 3 bits in the MPLS header that carry QoS info). MPLS core (P routers) Queues based on EXP. Voice gets priority queue. Bulk traffic gets best-effort. MPLS L3VPN egress (PE) Optionally re-marks DSCP from EXP. Hands packet to remote CE. Remote CE / remote LAN Same as ingress side in reverse. LLQ on inbound, trust on access ports, deliver to phone. Notice the carrier piece in the middle. Once you hand the packet to MPLS, the carrier's network is treating it based on EXP bits, not DSCP. If you mark DSCP correctly but the PE-CE mapping is wrong, your voice traffic crosses the carrier's network as best-effort and you have no QoS at all over the bit you do not own. This is why carrier coordination matters. ## DSCP marking standards Most enterprises follow the [RFC 4594](https://datatracker.ietf.org/doc/html/rfc4594?ref=pinglabz.com) recommendations for DSCP marking, with some Cisco-specific aliases. The relevant rows: Voice (RTP) DSCP nameEF Decimal46 Voice signaling DSCP nameCS3 Decimal24 Video conferencing (Webex, Teams) DSCP nameAF41 Decimal34 Real-time interactive (mission-critical apps) DSCP nameCS4 Decimal32 Network control (OSPF, BGP) DSCP nameCS6 Decimal48 Routine bulk DSCP nameAF11 or default Decimal10 or 0 Scavenger (background, low-priority) DSCP nameCS1 Decimal8 Two things to know about EF specifically. First, EF is the only DSCP that should go into a Low-Latency Queue (LLQ) with strict priority. Everything else uses a Class-Based Weighted Fair Queue (CBWFQ) with bandwidth reservations. Second, the LLQ is policed: even though it has strict priority, you cap its bandwidth so that voice cannot starve everything else if a flood of EF-marked traffic arrives. ## A minimum viable WAN-edge QoS policy This is the smallest policy that actually works for VoIP across a contended WAN circuit. Adjust the numbers for your link rate and call volume. ``` ! Classify traffic by DSCP marking class-map match-any VOICE-RTP match dscp ef class-map match-any VOICE-SIGNAL match dscp cs3 class-map match-any VIDEO match dscp af41 class-map match-any MISSION-CRITICAL match dscp cs4 class-map match-any NETWORK-CONTROL match dscp cs6 ! ! Build a child policy for the queues policy-map CHILD-WAN-QOS class VOICE-RTP priority percent 20 class VIDEO bandwidth percent 25 class MISSION-CRITICAL bandwidth percent 20 class VOICE-SIGNAL bandwidth percent 5 class NETWORK-CONTROL bandwidth percent 5 class class-default bandwidth percent 25 random-detect dscp-based ! ! Parent policy: shape to subscribed CIR, then run the child policy policy-map PARENT-WAN-QOS class class-default shape average 50000000 service-policy CHILD-WAN-QOS ! interface GigabitEthernet0/0/0 description WAN uplink to MPLS PE service-policy output PARENT-WAN-QOS ``` Key things in that config: - **Shape, then queue.** The parent shapes outbound traffic to the CIR you bought from the carrier (50 Mbps here). The child policy operates inside that shaped envelope. Without shaping, queues never fill and QoS never kicks in. - **LLQ via `priority percent 20`.** 20% of the 50 Mbps (10 Mbps) reserved for voice with strict priority. Caps voice at 10 Mbps regardless of how many EF-marked packets arrive. - **WRED on the default class.** Random Early Detection on best-effort traffic prevents TCP global synchronization (everyone retransmitting at once when a queue overflows). ## The carrier-side mapping The PE router at the carrier edge classifies your traffic and maps DSCP to MPLS EXP. This is the mapping that controls how your traffic is treated inside the carrier's core. A standard 6-class to 4-EXP mapping: EF (voice) Carrier classReal-time / Priority MPLS EXP5 CS6 (network) Carrier classNetwork control MPLS EXP6 AF41, CS4 Carrier classBusiness critical MPLS EXP4 CS3 (voice signaling) Carrier class Business critical or separate MPLS EXP3 AF11, default Carrier classBest-effort MPLS EXP0 CS1 (scavenger) Carrier classScavenger MPLS EXP1 Your carrier almost certainly publishes their EXP mapping in a service description document. Ask for it during contract negotiation. If you mark differently from what they expect, your traffic falls into best-effort inside the core, and your LLQ at the edge does nothing for traffic that already crossed the LAN. ## Verifying QoS is actually working The single most useful command: ``` R1# show policy-map interface GigabitEthernet0/0/0 GigabitEthernet0/0/0 Service-policy output: PARENT-WAN-QOS Class-map: class-default (match-any) 125003 packets, 158734129 bytes ... Service-policy : CHILD-WAN-QOS Class-map: VOICE-RTP (match-any) 18472 packets, 2954520 bytes 5 minute offered rate 1280000 bps, drop rate 0 bps Match: dscp ef (46) Priority: 20% (10000 kbps), burst bytes 250000, b/w exceed drops: 0 Class-map: VIDEO (match-any) ... drop rate 0 bps ``` Two numbers to watch: drop rate per class (anything > 0 in voice is a problem) and b/w exceed drops on the priority class (anything > 0 means your priority cap is being hit, voice is being dropped). ## Common failure modes Voice quality fine on LAN, choppy across WAN WAN edge policy not shaping to CIR, or no LLQ on egress. Check `show policy-map interface`. Voice quality good when WAN is uncongested, bad when bulk transfer runs QoS not actually classifying. DSCP markings being overwritten somewhere (often by a route-map or by a downstream device with `no mls qos trust`). One direction sounds fine, other direction sounds bad QoS configured only on one CE. Both CEs need symmetric policies on their WAN-facing interfaces. Voice sounds bad even though show policy-map shows zero drops Problem is in the carrier's network. Your EXP marking does not match the carrier's expectation, or the carrier's CoS is not what you bought. Specific call fails, others work Not a QoS problem. Look at signaling (SIP/SCCP) and call routing. ## Key takeaways QoS for VoIP across an MPLS WAN requires the LAN, the CE edge, and the PE all to agree on how voice is marked and queued. The phone marks DSCP EF. The CE has a parent-shaper plus child-policy with LLQ for voice. The PE maps DSCP to MPLS EXP per your carrier's contract. The MPLS core queues based on EXP. Verify with `show policy-map interface` and watch for drops. The single most common failure is mis-marking or no carrier-side EXP mapping, which means your edge QoS protects nothing inside the part of the path you do not own. For QoS fundamentals, see the [QoS pillar](https://www.pinglabz.com/qos/). ### OSPF Cost: Reference Bandwidth, Manual Overrides, and Gotchas URL: https://www.pinglabz.com/ospf-cost/ Last updated: 2026-06-13T20:07:52.000Z OSPF cost is the single metric that decides which path through your network wins. Get it wrong and traffic takes the long way around even though you have the right links in place. Get it right and the network does what you drew on the whiteboard. This post walks through how cost is calculated, where the default reference bandwidth bites, when to override cost manually, and the gotchas that show up the first time you run OSPF on a network with mixed-speed links. For the cluster overview, see the [OSPF complete guide](https://www.pinglabz.com/ospf/). For an interior gateway protocol that handles cost calculation differently, see the [EIGRP pillar](https://www.pinglabz.com/eigrp/). ## The formula OSPF calculates the cost of each interface using: ``` cost = reference_bandwidth / interface_bandwidth ``` Both numbers are in bits per second. The default reference bandwidth on Cisco IOS XE is 100,000,000 (100 Mbps). The interface bandwidth is whatever the physical interface reports (or whatever you have set with the `bandwidth` command). Worked examples with default reference: 10 Mbps Calculation 100,000,000 / 10,000,000 OSPF cost10 100 Mbps Calculation 100,000,000 / 100,000,000 OSPF cost1 1 Gbps Calculation 100,000,000 / 1,000,000,000 OSPF cost 1 (rounded up from 0.1) 10 Gbps Calculation 100,000,000 / 10,000,000,000 OSPF cost 1 (rounded up from 0.01) 40 Gbps / 100 Gbps CalculationSame math OSPF cost 1 (all collapse to the same cost) This is the default problem. Every interface from 100 Mbps to 100 Gbps gets cost 1\. OSPF cannot tell your 100 Gbps backbone link apart from a 100 Mbps copper drop. Traffic will route over whichever path has fewer hops, regardless of bandwidth. ## Fixing the reference bandwidth The fix is to raise the reference bandwidth to a number that gives meaningfully different costs for the link speeds you actually have. ``` router ospf 1 auto-cost reference-bandwidth 100000 ``` That sets the reference to 100 Gbps. Recalculate the table: 10 Mbps 10,000 100 Mbps 1,000 1 Gbps 100 10 Gbps 10 100 Gbps 1 Now each link speed has a distinct cost. The 10 Gbps backbone wins over the 1 Gbps path. The 1 Gbps wins over the 100 Mbps. Path selection actually reflects bandwidth. **Critical rule:** the reference bandwidth must be the same on every router in the OSPF domain. If one router uses 100 (the default in Gbps reading) and another uses 100000 (100 Gbps), the cost calculations will not align and the SPF tree will compute differently on different routers. The resulting routing loops are hard to diagnose. Always configure auto-cost reference-bandwidth network-wide, or accept the default network-wide. ## Manual cost override per interface For specific interfaces where you want to override the calculated cost, use: ``` interface GigabitEthernet0/0 ip ospf cost 50 ``` This forces the OSPF cost for that interface to 50 regardless of the bandwidth-based calculation. Common reasons to override: - You want traffic to prefer one specific path over another for policy reasons (cost of the carrier circuit, security boundary preference) - You have a low-bandwidth backup link that should only be used if the primary fails. Set its cost to a very large number (e.g., 1000) so it loses to everything except true failure. - You have a high-bandwidth link that you know is congested and want to actively de-prefer for OSPF traffic (rare, but happens) The manual cost wins over the auto-calculated cost. ## The bandwidth command vs the actual link speed The OSPF cost calculation uses the `bandwidth` value reported by the interface. By default, this matches the physical link rate (1000000 kbps for a Gigabit interface). You can override it: ``` interface GigabitEthernet0/0 bandwidth 100000 ! Tells the router this interface is 100 Mbps for protocol purposes ! Does not actually slow the physical link ``` This is a powerful and dangerous knob. It lets you influence OSPF cost without using `ip ospf cost` directly. But the `bandwidth` value is consumed by other protocols too (EIGRP uses it for metric, QoS uses it for shaping references, MPLS-TE uses it for available-bandwidth tracking). Changing it for OSPF reasons can break things you did not plan to touch. Rule of thumb: use `ip ospf cost` when you want to influence only OSPF. Use `bandwidth` only when you actually need to advertise a different rate to multiple protocols. ## Verifying the cost ``` R1# show ip ospf interface GigabitEthernet0/0 GigabitEthernet0/0 is up, line protocol is up Internet Address 10.20.0.1/24, Area 0 Process ID 1, Router ID 1.1.1.1, Network Type BROADCAST, Cost: 10 Topology-MTID Cost Disabled Shutdown Topology Name 0 10 no no Base ``` Three things to read: - **Cost: 10.** This is the cost OSPF is using for this interface. After my `auto-cost reference-bandwidth 100000`, a 10 Gbps interface should show 10\. (A 1 Gbps would show 100.) - **Network Type: BROADCAST.** The Hello/dead intervals and DR/BDR election depend on this. Not directly cost-related but visible in the same output. - **Area 0.** Just confirms what area this interface is in. ## Total path cost The total cost of an OSPF path is the sum of the interface costs along the path. Specifically, the cost of every outbound interface from the local router to the destination. The cost of incoming interfaces does not count. If the path is R1 -> R2 -> R3 over two 10 Gbps interfaces, with reference bandwidth 100 Gbps, the total cost is 10 + 10 = 20\. R1 advertises the prefix to other routers with that summed cost. `show ip route X.X.X.X` displays this total. ## Equal-cost multipath (ECMP) When two paths to the same destination have the same total cost, OSPF installs both in the FIB by default and load-balances. The maximum number of equal-cost paths is controlled by: ``` router ospf 1 maximum-paths 4 ``` Default is 4\. Can be raised. ECMP is one of the reasons careful cost engineering matters: if you accidentally make two paths equal cost when you wanted one to be preferred, traffic splits across both, including possibly across a path with worse jitter or latency. ## Common gotchas Traffic takes the wrong path through a network with mixed-speed links Default reference bandwidth in use. Every link is cost 1\. Raise the reference bandwidth network-wide. SPF results differ between routers; routing loops appear Reference bandwidth is not consistent across the OSPF domain. Match it everywhere. Manual cost override is being ignored Check whether the OSPF process is operating in a different topology MTID and you set cost in the wrong topology. Rare but happens with MTR. Two paths I want unequal are coming out as ECMP They are cost-equal. Either change the bandwidth on one (preferred), set a manual `ip ospf cost` difference, or accept the load balance. OSPF cost looks right on a router but the route picks a different path Administrative distance: a static route or eBGP-learned path is winning. Check `show ip route X.X.X.X` for the protocol source. ## Key takeaways OSPF cost is reference bandwidth divided by interface bandwidth. The default reference of 100 Mbps was set when 100 Mbps was a backbone link, and is now wrong for every modern network. Raise it to at least 100 Gbps. Configure the new value on every router in the OSPF domain to keep the math consistent. Use `ip ospf cost` per-interface for surgical overrides, but resist the temptation to micromanage cost on every link; the auto-calculated cost is correct often enough that explicit overrides should be rare and documented. For broader OSPF coverage, see the [OSPF pillar](https://www.pinglabz.com/ospf/). ### STP and EtherChannel: When They Collide and Who Wins URL: https://www.pinglabz.com/etherchannel-spanning-tree/ Last updated: 2026-06-13T20:07:52.000Z EtherChannel and Spanning Tree have a complicated relationship. The whole point of EtherChannel is to take two or more physical links and present them as a single logical link with combined bandwidth. The whole point of Spanning Tree is to block redundant links to prevent loops. By default these two technologies argue about what to do with your second cable. This post covers what happens when they collide, how the resolution actually works, and the misconfigurations that cause one or the other to win in ways you did not intend. For the cluster overview, see the [Spanning Tree Protocol pillar](https://www.pinglabz.com/spanning-tree-protocol/). For the L2 fundamentals these protocols ride on, see the [VLAN and L2 switching pillar](https://www.pinglabz.com/vlans-layer-2-switching/). ## The fundamental tension Plug two cables between two switches with no EtherChannel and no special config. STP sees two redundant paths between the same two switches. It picks one to forward on and blocks the other. The blocked link sits idle. You paid for cable and a switch port that does nothing. Now bundle the same two cables into an EtherChannel. STP no longer sees two redundant links. It sees one logical link with twice the bandwidth. Both physical links forward at the same time, with frames hashed across them based on a configurable load-balance algorithm. STP cooperates because it has nothing to block. This is the entire purpose of EtherChannel from STP's perspective. The bundling is what makes redundant links cooperate instead of one of them being blocked. ## How EtherChannel hides the physical links from STP Once a Port-channel interface is created and physical interfaces are added to it, STP runs against the Port-channel interface only. The individual physical interfaces become non-STP-participants. They forward frames based on the EtherChannel hash. They do not send their own BPDUs. They do not have their own STP state. The verification command: ``` SW1# show spanning-tree interface Port-channel1 Vlan Role Sts Cost Prio.Nbr Type ---------------- ---- --- --------- -------- -------------------------------- VLAN0001 Desg FWD 3 128.97 P2p VLAN0010 Desg FWD 3 128.97 P2p ``` Notice that the cost is 3 even though each contributing physical link might be a 1 Gbps interface (cost 4 on its own). EtherChannel adjusts the effective cost downward to reflect the combined bandwidth. Two 1G links bundled show as cost 3\. Eight 1G links bundled show as cost 1. ## The bundling protocols: LACP, PAgP, On active ProtocolLACP What it does Negotiates the bundle with the far end. Industry standard (IEEE 802.3ad). passive ProtocolLACP What it does Will form a bundle if the other side is active. desirable ProtocolPAgP What it does Cisco proprietary. Negotiates the bundle. auto ProtocolPAgP What it does Will form a bundle if the other side is desirable. on ProtocolNone What it does Forces a static bundle with no negotiation. Both sides must be "on". In modern multi-vendor environments, LACP is the only mode worth using. PAgP is Cisco-only and has been quietly deprecated. "On" mode is dangerous: it forces the bundle without verifying the other side is doing the same, and a misconfiguration creates a loop that STP cannot prevent (because STP no longer sees the individual links). A correct LACP configuration looks like: ``` interface range GigabitEthernet1/0/23 - 24 channel-group 1 mode active ! interface Port-channel1 switchport mode trunk switchport trunk allowed vlan 10,20,30 ``` Both sides set `mode active`. They negotiate, find each other, form Port-channel1\. From STP's perspective, Port-channel1 is the only thing it sees on this link. ## The collision: when the bundle does not form The dangerous moment is when you configure an EtherChannel but the bundle fails to negotiate. Reasons it fails: - One side uses LACP active, the other uses PAgP desirable. They cannot agree on a protocol. - Speed or duplex mismatch on contributing interfaces. - VLAN configuration differs across contributing interfaces (some are access, some are trunk; trunk-allowed VLAN lists do not match; native VLAN mismatch). - One side is configured for the channel-group, the other is not. When negotiation fails, the contributing interfaces are NOT in the Port-channel. They appear in show output as individual interfaces. STP sees them as redundant links, and blocks all but one. This is the safe failure mode: no loop, but no aggregation either. You end up with one active link instead of N active links, and traffic uses only that one. ``` SW1# show etherchannel summary Flags: D - down P - bundled in port-channel I - stand-alone s - suspended H - Hot-standby (LACP only) ... Number of channel-groups in use: 1 Group Port-channel Protocol Ports ------+-------------+-----------+----------------------------------------------- 1 Po1(SD) LACP Gi1/0/23(s) Gi1/0/24(s) ``` The (s) flag on the physical interfaces means "suspended" - LACP did not get an answer from the other side, so the port is not bundled. The (SD) on Port-channel1 means it is down. STP is now load-balancing across nothing. ## The more dangerous collision: "on" mode misconfiguration "On" mode skips negotiation. If both sides are "on" with matching configs, the bundle works. If one side is "on" and the other side is configured normally as two independent interfaces, the side in "on" mode believes it has a bundle. It forwards frames hashed across both interfaces. The other side does not believe it has a bundle. It sees two redundant links and blocks one. Frames hashed to the blocked link disappear. Frames hashed to the forwarding link work. Half the traffic is silently dropped. STP does not see a problem because, from its perspective, the other side is doing the right thing. This is the failure mode that ruins afternoons. The rule: never use "on" mode unless you have a specific reason and you control both ends. ## Cross-stack EtherChannel (MEC, vPC, MLAG) Modern designs often want to bundle links across two different physical switches for redundancy. The bundle members live on different chassis. STP would normally still see them as separate physical links to separate switches, but vendor stacking technologies let you present them as a single logical bundle. Multi-chassis EtherChannel (MEC) on StackWise / VSS Vendor-neutralMLAG What it is Two stacked switches present as one logical switch. Bundle can include ports from both. Virtual Port Channel (vPC) on Nexus Vendor-neutralMLAG What it is Two physically-separate Nexus switches synchronize to act as one logical peer for EtherChannel. STP still sees one switch. From a STP perspective, MEC and vPC both make the dual-chassis pair look like a single switch. STP forms one adjacency with the logical entity. The bundle protects against chassis failure as well as link failure. ## STP behavior when an EtherChannel member fails Probably the cleanest part of the whole story. When one physical link in a bundle fails: - EtherChannel removes the failed link from the load-balance hash. Remaining links pick up the load. - STP does not see anything change. The Port-channel is still up. Same role, same state. - No STP recalculation. No TC bit set. No MAC table flush. This is one of the operational advantages of EtherChannel. A single link failure inside a bundle is invisible to STP, which means it does not trigger the (small but nonzero) disruption a STP topology change normally would. ## Load balancing EtherChannel hashes frames across the bundle members. The hash algorithm is configurable. ``` port-channel load-balance src-dst-ip ! Common default port-channel load-balance src-dst-mixed-ip-port ! Better hash distribution ``` The default of src-dst-ip is fine for most environments. If you find your bundle is unbalanced (one member sees most of the traffic), the most common cause is a few high-volume flows between the same IP pair that hash to the same member. Switching to a hash that includes L4 ports usually fixes it. ## Key takeaways EtherChannel and Spanning Tree resolve their conflict by hiding the bundled physical links from STP entirely. STP sees one logical link with adjusted cost; both members forward in parallel. The dangerous failure modes are bundling negotiation failing (mode mismatches, VLAN config differences) and "on" mode bundling without negotiation. LACP active on both sides is the only configuration that fails safely. When you put redundant links between two switches in 2026, you EtherChannel them, you use LACP, and you let STP see one logical link. For the STP cluster, see the [Spanning Tree Protocol pillar](https://www.pinglabz.com/spanning-tree-protocol/). ### RSTP: Rapid Spanning Tree Protocol Explained (802.1w) URL: https://www.pinglabz.com/rapid-spanning-tree-protocol/ Last updated: 2026-06-13T20:07:52.000Z Rapid Spanning Tree Protocol (RSTP, IEEE 802.1w) is what replaced classic Spanning Tree (802.1D) on every Cisco switch shipped this century. The classical PVST+ implementation Cisco shipped before the rapid version is, to a first approximation, dead. If you are running a network in 2026 you are almost certainly running Rapid-PVST+ or MST, both of which inherit from RSTP. This post walks through what RSTP changed, the new port roles and states, the proposal/agreement handshake that gives RSTP its speed, and how to read the show output on Cisco IOS XE. For the cluster overview, see the [Spanning Tree Protocol pillar](https://www.pinglabz.com/spanning-tree-protocol/). For the L2 fundamentals RSTP sits on top of, see the [VLAN and L2 switching pillar](https://www.pinglabz.com/vlans-layer-2-switching/). ## Why classic STP needed replacing Classic 802.1D STP took up to 50 seconds to recover from a topology change. The slowdown came from two hard-coded timers: 20 seconds for max-age (waiting to see if the old root was really gone), and 30 seconds combined for the listening and learning states each port had to pass through before it could forward. On a modern switched network, 50 seconds of unreachability is unacceptable. Voice calls drop. TCP sessions reset. Users call the help desk. RSTP rebuilt the FSM around active negotiation between switches instead of passive timer expiration. The result, in practice, is sub-second convergence on most topologies. ## The four port roles RSTP defines port roles, separate from port states. The roles describe what the port does in the topology; the states describe its current operational behavior. Root Meaning The port on this switch that points toward the root bridge along the lowest-cost path STP equivalentSame concept in STP Designated Meaning The port elected as the forwarding path for a given segment (the "best port" on a segment) STP equivalentSame concept in STP Alternate Meaning A port that offers an alternate path to the root if the root port fails. Discarding in steady state. STP equivalent Did not exist in STP; this is the role that makes sub-second failover possible Backup Meaning A port that offers a backup path to a segment for which another local port is already Designated. Rare in modern topologies. STP equivalentDid not exist in STP The Alternate role is what lets RSTP fail over fast. When the root port goes down, the Alternate immediately transitions to Root without waiting for a topology change recalculation across the whole network. The new path was already known; only the role-and-state transition is needed. ## The three port states RSTP collapses STP's five port states (Disabled, Blocking, Listening, Learning, Forwarding) into three. Discarding What it does Drops all frames except BPDUs MAC learning?No Frame forwarding?No Learning What it does Drops frames but starts populating the MAC table MAC learning?Yes Frame forwarding?No Forwarding What it doesNormal operation MAC learning?Yes Frame forwarding?Yes The three RSTP states map cleanly to operational reality. Discarding is "I should not be forwarding here right now." Learning is the brief transitional state between Discarding and Forwarding. Forwarding is steady-state. ## The proposal/agreement handshake This is the mechanism that lets RSTP achieve sub-second convergence on point-to-point links. The full description in 802.1w is long; the operational reality is short. When two switches connected by a point-to-point link first come up, the upstream switch (closer to the root) sets the Proposal bit in its BPDU and sends it. The downstream switch receives the proposal, syncs its other ports to Discarding (preventing loops), then replies with an Agreement BPDU. As soon as the upstream switch receives the agreement, both ends transition to Forwarding. The handshake completes in low-millisecond time. Compare that to the 30 seconds classic STP took. The cost is that RSTP requires switches to identify links as point-to-point (which it does automatically based on duplex; full-duplex links are P2P, half-duplex are shared media). ## Edge ports (PortFast, in Cisco terms) RSTP introduces the concept of an "edge port": a port connected to an end host that will never participate in the spanning tree. Cisco's PortFast feature is the implementation. An edge port skips the proposal/agreement handshake entirely and goes straight to Forwarding when link comes up. Enable it on every access port that connects to a host. Skipping this causes 30-second startup delays on PCs at boot, the original problem PortFast was created to solve back in classic-STP days. ``` interface GigabitEthernet1/0/5 switchport mode access switchport access vlan 10 spanning-tree portfast ``` To enable PortFast by default on all access ports globally (recommended in modern designs): ``` spanning-tree portfast default ``` The matching protection is BPDU Guard. An edge port that receives a BPDU (someone plugged a switch into a user port) is shut down immediately: ``` spanning-tree portfast bpduguard default ``` ## Topology Change Notifications: also simpler Classic STP propagated topology changes via TCN BPDUs that bubbled up to the root, which then flooded a TC bit through the network. RSTP eliminates the bubble-up. When a switch detects a topology change (a non-edge port transitions to Forwarding or a Designated port goes down), it sets the TC bit on its outbound BPDUs for a fixed interval. Other switches receiving the TC bit flush their MAC tables for the affected VLAN. No round-trip to the root required. ## Rapid-PVST+: Cisco's per-VLAN flavor Cisco's default RSTP implementation is Rapid-PVST+ (Per-VLAN Spanning Tree Plus). It runs a separate RSTP instance per VLAN. Each VLAN can have a different root bridge. This is computationally heavier than the IEEE-standard 802.1w single-instance RSTP, but it lets you do per-VLAN load balancing (one VLAN's traffic takes one uplink, another VLAN's takes the other). To enable Rapid-PVST+ on a Catalyst switch (this is the default on most modern IOS XE): ``` spanning-tree mode rapid-pvst ``` If you have a lot of VLANs (more than 50 or so) and care about CPU and BPDU bandwidth, MST (Multiple Spanning Tree, 802.1s) maps multiple VLANs to a single instance, dramatically reducing the per-instance overhead. MST is the choice for large environments. ## Reading the show output The most useful command on a steady-state network: ``` SW1# show spanning-tree vlan 10 VLAN0010 Spanning tree enabled protocol rstp Root ID Priority 24586 Address 0050.7989.aaaa Cost 4 Port 25 (TenGigabitEthernet1/1/1) Hello Time 2 sec Max Age 20 sec Forward Delay 15 sec Bridge ID Priority 32778 (priority 32768 sys-id-ext 10) Address 0050.7989.bbbb Hello Time 2 sec Max Age 20 sec Forward Delay 15 sec Aging Time 300 sec Interface Role Sts Cost Prio.Nbr Type ------------------- ---- --- --------- -------- -------------------------------- Te1/1/1 Root FWD 4 128.25 P2p Te1/1/2 Altn BLK 4 128.26 P2p Gi1/0/5 Desg FWD 19 128.5 Edge P2p ``` Three things to read off: - **Root ID vs Bridge ID.** If they match, this switch is the root. If not, the Root ID tells you which switch the topology has elected. - **Role column.** Root, Designated (Desg), Alternate (Altn), Backup. The pattern of roles tells you the active topology. - **Type column.** "P2p" means RSTP is treating it as a point-to-point link (enables fast handshake). "Shr" means shared media (slow). "Edge" means it is an edge port (PortFast). ## Common gotchas Convergence is slow even though "RSTP" is configured One end of a link is half-duplex. RSTP treats it as shared media (Shr), skipping the proposal/agreement handshake. Access ports take 30 seconds to come up PortFast not enabled. Add `spanning-tree portfast default` globally. STP reconverges when a single host reboots An access port without PortFast is registered as topology-change-generating. Add PortFast. Different VLANs choose different root bridges, traffic takes weird paths Rapid-PVST+ behavior. Either set deterministic root priorities per VLAN, or migrate to MST and group VLANs into instances. BPDU Guard shutting down ports that should be edge Probably correct, actually. Investigate why an edge port saw a BPDU before re-enabling. ## Key takeaways RSTP replaced classic STP because 50-second convergence was no longer acceptable. The mechanisms that achieve sub-second recovery are the new Alternate port role (pre-computed backup path), the three-state model (Discarding, Learning, Forwarding), the proposal/agreement handshake on point-to-point links, and the elimination of the root-bubble-up for topology changes. Cisco's Rapid-PVST+ is RSTP per VLAN; MST is RSTP with VLAN-to-instance mapping for large environments. PortFast on every access port and BPDU Guard on the same ports are non-negotiable in modern designs. For the STP cluster, see the [Spanning Tree Protocol pillar](https://www.pinglabz.com/spanning-tree-protocol/). ### Inter-VLAN Routing: SVI vs Router-on-a-Stick (with Real IOS XE Config) URL: https://www.pinglabz.com/inter-vlan-routing/ Last updated: 2026-06-13T20:07:52.000Z The moment your network has more than one VLAN and the hosts in those VLANs need to talk to each other, you need inter-VLAN routing. There are two ways to do it on Cisco gear: a Layer 3 switch with SVIs, and a router with sub-interfaces on a trunk. Both work. One is faster, simpler, and what you should use in production. The other is what you build in the lab or on a budget. This post walks through both with real IOS XE config and explains why one wins. For the L2 fundamentals, see the [VLAN and Layer 2 switching pillar](https://www.pinglabz.com/vlans-layer-2-switching/). For routing protocols that ride on top once inter-VLAN routing is in place, see the [OSPF pillar](https://www.pinglabz.com/ospf/). ## The two approaches Switched Virtual Interfaces (SVIs) How it works L3 switch creates a virtual interface per VLAN with an IP address. Forwarding happens in switch ASIC at line rate. Where it lives Modern enterprise design. Catalyst 9300/9500, Nexus 9000, anything with L3 capability. Router-on-a-stick How it works L2 switch trunks all VLANs to a router. Router has a sub-interface per VLAN with an IP. Forwarding happens in router CPU or NPU. Where it lives Lab work, branches with an L2-only switch and a small router, or environments with strict separation between switching and routing teams. The performance gap between the two is large. An SVI on a Catalyst 9300 routes at wire-speed (40 Gbps+ per port). A router-on-a-stick uplink is bounded by both the router's forwarding capacity and the trunk's bandwidth, and every inter-VLAN packet crosses the trunk twice (in on one VLAN, out on another). For anything more than a small office, SVIs are the only correct answer. ## The lab topology Three VLANs across one Layer 3 switch, plus one router as the alternative router-on-a-stick path for comparison. IP scheme: USERS VLAN10 Subnet10.10.10.0/24 SVI / router subinterface IP10.10.10.1 SERVERS VLAN20 Subnet10.10.20.0/24 SVI / router subinterface IP10.10.20.1 VOICE VLAN30 Subnet10.10.30.0/24 SVI / router subinterface IP10.10.30.1 ## Method 1: SVIs on a Layer 3 switch The configuration is short. Most of it is enabling routing globally and giving each VLAN's SVI an IP. ``` ! Enable IP routing globally (off by default on Catalyst switches) ip routing ! ! Create the VLANs in the database vlan 10 name USERS vlan 20 name SERVERS vlan 30 name VOICE ! ! Create the SVI for each VLAN and give it the gateway IP interface Vlan10 description USERS gateway ip address 10.10.10.1 255.255.255.0 no shutdown ! interface Vlan20 description SERVERS gateway ip address 10.10.20.1 255.255.255.0 no shutdown ! interface Vlan30 description VOICE gateway ip address 10.10.30.1 255.255.255.0 no shutdown ! ! Access ports assigned to their VLANs interface range GigabitEthernet1/0/1 - 12 switchport mode access switchport access vlan 10 spanning-tree portfast ! interface range GigabitEthernet1/0/13 - 24 switchport mode access switchport access vlan 20 spanning-tree portfast ``` That is it. A host in VLAN 10 with default gateway 10.10.10.1 can now ping a host in VLAN 20 with default gateway 10.10.20.1\. The switch's ASIC handles the forwarding in hardware. Verify it worked: ``` Switch# show ip route connected C 10.10.10.0/24 is directly connected, Vlan10 C 10.10.20.0/24 is directly connected, Vlan20 C 10.10.30.0/24 is directly connected, Vlan30 Switch# show ip interface brief | exclude unassigned Interface IP-Address OK? Method Status Protocol Vlan10 10.10.10.1 YES manual up up Vlan20 10.10.20.1 YES manual up up Vlan30 10.10.30.1 YES manual up up ``` The "Protocol up" on each SVI matters. An SVI is "up/up" only if there is at least one active access or trunk port carrying that VLAN. If no port is in the VLAN, the SVI shows up/down even if the IP is configured correctly. That single behavior catches many "my gateway is unreachable" troubleshooting sessions. ## Method 2: Router-on-a-stick The switch is L2-only. A single uplink trunk to the router carries all three VLANs. The router has three sub-interfaces, one per VLAN. On the switch: ``` vlan 10 name USERS vlan 20 name SERVERS vlan 30 name VOICE ! ! Trunk uplink to the router interface GigabitEthernet1/0/24 description Uplink to R1 (router-on-a-stick) switchport mode trunk switchport trunk allowed vlan 10,20,30 switchport trunk encapsulation dot1q ! Some platforms need this; others infer ! ! Access ports as before interface range GigabitEthernet1/0/1 - 12 switchport mode access switchport access vlan 10 spanning-tree portfast ``` On the router: ``` interface GigabitEthernet0/0 no shutdown ! No IP on the physical; all IP lives on sub-interfaces ! interface GigabitEthernet0/0.10 description USERS gateway encapsulation dot1Q 10 ip address 10.10.10.1 255.255.255.0 ! interface GigabitEthernet0/0.20 description SERVERS gateway encapsulation dot1Q 20 ip address 10.10.20.1 255.255.255.0 ! interface GigabitEthernet0/0.30 description VOICE gateway encapsulation dot1Q 30 ip address 10.10.30.1 255.255.255.0 ``` The router routes between the sub-interfaces just like it would between two physical interfaces. The frame comes in tagged for VLAN 10, gets stripped to a packet, gets routed, and gets re-tagged for VLAN 20 on the way back out the same physical interface. Verify on the router: ``` R1# show ip route C 10.10.10.0/24 is directly connected, GigabitEthernet0/0.10 C 10.10.20.0/24 is directly connected, GigabitEthernet0/0.20 C 10.10.30.0/24 is directly connected, GigabitEthernet0/0.30 ``` ## Performance comparison Forwarding rate SVI on L3 switch Line rate per port (multi-Gbps to 100G+) Router-on-a-stick Bounded by router CPU/NPU and the single trunk Latency added per inter-VLAN hop SVI on L3 switchMicroseconds (ASIC) Router-on-a-stick Hundreds of microseconds to milliseconds (software forwarding) Bottleneck SVI on L3 switchPer-port bandwidth Router-on-a-stick Trunk bandwidth (and every packet crosses it twice) Failure domain if router/switch fails SVI on L3 switch One physical box. Mitigate with stack/VSS/StackWise Virtual. Router-on-a-stick One physical box. Mitigate with HSRP/VRRP on the router pair. ## The voice VLAN special case On a port serving an IP phone with a daisy-chained PC, two VLANs ride one port: the data VLAN (untagged) and the voice VLAN (tagged, sometimes called "auxiliary VLAN"). The configuration: ``` interface GigabitEthernet1/0/5 switchport mode access switchport access vlan 10 switchport voice vlan 30 spanning-tree portfast ``` The phone learns the voice VLAN via CDP or LLDP-MED from the switch and starts tagging its own traffic. The PC plugged into the phone sends untagged frames that the switch puts into VLAN 10\. The router or SVI then handles routing between the two as usual. ## Common gotchas Hosts in VLAN A cannot ping hosts in VLAN B `ip routing` is not enabled globally on the L3 switch. (Default off on many Catalyst lines.) SVI shows up/down even with correct IP No active port in the VLAN. Bring up at least one access port assigned to that VLAN. One VLAN routes, another does not (router-on-a-stick) Sub-interface `encapsulation dot1Q ` missing or wrong VLAN number. VLAN works on trunk but native VLAN traffic fails Native VLAN mismatch between switch and router/peer switch. Force native VLAN consistency or never use native VLAN for routed traffic. Inter-VLAN ping works but iperf shows tiny throughput Router-on-a-stick uplink is saturated. Either upgrade the trunk to higher-speed or migrate to SVI on an L3 switch. ## Key takeaways SVIs on a Layer 3 switch are the right answer for inter-VLAN routing in any environment with more than one or two access switches. Router-on-a-stick is the right answer for labs and small branches with L2-only switches. The two configs are syntactically different but conceptually identical: each VLAN gets an L3 interface with the gateway IP, and traffic between VLANs flows through it. The performance gap between hardware-forwarded SVI traffic and software-forwarded router-on-a-stick traffic is what drives the design choice for production. For the L2 foundations, see the [VLAN pillar](https://www.pinglabz.com/vlans-layer-2-switching/). ### MPLS L3VPN vs SD-WAN: When to Migrate URL: https://www.pinglabz.com/mpls-vs-sd-wan-migration/ Last updated: 2026-06-13T20:07:53.000Z The "MPLS to SD-WAN migration" project is one of the most common WAN modernizations underway in 2026\. Most enterprises with MPLS contracts signed before 2022 have at least asked the question. Many have started the work. A meaningful fraction have abandoned the migration partway through and ended up running both, which is the worst of both worlds. This post is a decision framework, not a sales pitch: when to migrate, when to keep MPLS, and the seven-step transition that actually works. For the SD-WAN architecture details, see the [SD-WAN complete guide](https://www.pinglabz.com/sd-wan/). For the MPLS L3VPN fundamentals, see the [MPLS pillar](https://www.pinglabz.com/mpls/). ## The migration is not a 1-to-1 swap The most common mistake is treating SD-WAN as a drop-in replacement for MPLS. They solve overlapping but not identical problems. Transport MPLS L3VPN Dedicated carrier circuit SD-WAN Overlay across any IP transport Per-site bandwidth MPLS L3VPN What you ordered. Fixed. SD-WAN Sum of all transports. Variable. QoS MPLS L3VPN Carrier-enforced, end-to-end SD-WAN Edge-enforced. Carrier sees IPsec tunnels. SLA MPLS L3VPN Carrier-backed (loss, latency, jitter) SD-WAN Per-transport SLA only. Overlay SLA is your problem. Application visibility MPLS L3VPN None at the carrier; some at the edge SD-WAN Deep app-level telemetry from controller Path selection MPLS L3VPN Once per destination, by IGP SD-WAN Per-application, per-flow, dynamically Cost per Mbps MPLS L3VPN $30 to $100+ per Mbps/month SD-WAN $1 to $10 per Mbps/month (broadband) Install time per site MPLS L3VPN30 to 90 days SD-WAN Days, once edge device ships Read that QoS row twice. MPLS gives you carrier-backed CoS across the entire WAN, end-to-end, including the bits of the path you do not own. SD-WAN does not. SD-WAN's edge-enforced QoS works inside your IPsec tunnels but the underlying internet carrier has no idea you marked the packets, so once the packet leaves your edge it competes for capacity with whatever else is on that link. ## When to migrate The honest answer is: when SD-WAN's strengths outweigh MPLS's strengths for your specific traffic mix. The decision matrix: Most application traffic is SaaS (Microsoft 365, Salesforce, Google) Migrate. SD-WAN with local internet break-out is the right pattern. You have 50+ branches and an MPLS bill that is half your WAN cost Migrate. Cost savings alone justify the project. Greenfield expansion into geographies where MPLS provisioning is slow or expensive Migrate the new sites only. Hybrid for existing. Traffic is mostly client-server to a primary data center, latency-sensitive (real-time financial trading, VDI for low-tolerance users) Keep MPLS for that traffic. Carrier-backed SLA matters here. Regulated environment that requires circuit isolation auditable down to the provider's network Stay on MPLS unless you can satisfy compliance with private SD-WAN overlays. Branches in markets where broadband is unreliable or unavailable Keep MPLS or 5G for those sites; migrate where transport options are healthy. ## The seven-step transition The migrations that work follow roughly this shape. The migrations that fail skip step 3, step 5, or both. ### Step 1: Inventory and classify Catalogue every application traversing the WAN. Group them into: real-time (voice, video), interactive (SaaS, ERP), bulk (backups, file sync), and management (telemetry, monitoring). The classification drives the policy in step 5\. Skip this and you will end up with a flat priority queue, which performs worse than the MPLS you replaced. ### Step 2: Pilot, do not big-bang Pick 3 to 5 representative branches. Different sizes, different geographies, different transport availability. Run them on SD-WAN for 90 days alongside the existing MPLS. Measure user experience, not just link metrics. SaaS application performance from each branch is the only metric that matters. ### Step 3: Run hybrid for at least 6 months Most failed migrations cut MPLS too fast. The right pattern is to spend 6 to 12 months in hybrid mode: SD-WAN overlay handling the policy and path selection, with MPLS still available as one of the transports under it. This gives you a fallback while you find the application-specific edge cases your inventory missed. ### Step 4: Decide the data center pattern How traffic enters and leaves the data center fundamentally changes the design. The two patterns: - **Aggregation hub:** All SD-WAN edges terminate to a hub at the data center. Closer to a hub-and-spoke topology. Easier to migrate from MPLS where this was already the implicit shape. - **Mesh with cloud on-ramps:** Edges peer with each other directly and with cloud provider PoPs. Better for SaaS-heavy traffic. Harder to make compliant. ### Step 5: Build and test the policies before you cut over This is where most migrations slip. SD-WAN policy is more powerful than MPLS QoS, which means more rope to hang yourself with. Spend time defining per-application path preferences, failover behavior, brownout detection thresholds, and per-tenant segmentation rules. Test them in the pilot environment under degraded-link conditions (use tc-netem to inject loss and latency on a test link). Watch how the overlay reacts. ### Step 6: Cut MPLS in tranches When you cut MPLS, do it 10-20% of sites per month. Each tranche should be similarly sized and geographically diverse so a regional broadband issue does not take down a whole tranche at once. Keep MPLS contracts on month-to-month for the final 3 to 6 months in case you need to roll back at a specific branch. ### Step 7: Keep MPLS for the few sites that genuinely need it Almost every "we migrated everything to SD-WAN" story ends with a footnote: "except the 5% of sites where MPLS made sense." That is the right outcome. Treat it as success, not failure. ## Costs you will not see in the SD-WAN business case Vendor business cases for SD-WAN are aggressive on circuit savings and quiet on the costs that are real but harder to model. - **Edge appliance and license costs.** Vendor licensing per-edge can total 40 to 60% of the savings vs MPLS for small branches. - **Controller infrastructure.** vManage, VCO, Prisma SD-WAN Cloud all have licensing and hosting costs. - **Operations retooling.** Your NOC needs new dashboards, runbooks, and skill sets. Budget 6 months of dual-running. - **Multi-vendor friction.** SD-WAN does not integrate with legacy WAN-side monitoring as cleanly as MPLS did. Expect to replace or extend monitoring tools. - **Internet circuit churn.** Cheaper broadband ISPs have shorter SLA windows and rougher support. Plan for diverse ISPs per site even if it costs more per-Mbps. ## When the answer is "do not migrate" The conversation does not always end with "migrate." Genuine reasons to keep MPLS: - You have 1 to 10 branches. The fixed costs of SD-WAN (controllers, licensing, retooling) do not amortize. - Your applications are all on-premise and latency-sensitive. SD-WAN does not make a fiber circuit faster. - Your existing MPLS contract has 3+ years remaining with high early-termination fees. - Your compliance regime requires carrier-attested circuit isolation that broadband-based overlays cannot satisfy. - Your team does not have bandwidth for a 12-18 month transition project right now. Any one of these is sufficient to defer the migration. Two or more is a strong signal to keep MPLS for at least another contract cycle. ## Key takeaways SD-WAN replaces most of what MPLS used to do, but not all of it. The migration works when you treat it as a 12 to 18 month project with a long hybrid phase, not a circuit-swap. Inventory applications first. Pilot before committing. Keep MPLS where its carrier-backed SLA still earns its cost. Most enterprises end up with mostly-SD-WAN plus a handful of MPLS sites, and that is the right outcome, not a failure. For the SD-WAN technology and config details, see the [SD-WAN pillar](https://www.pinglabz.com/sd-wan/). ### BGP Looking Glass: What It Is, Public Servers, and Hosting Your Own URL: https://www.pinglabz.com/bgp-looking-glass/ Last updated: 2026-07-04T23:21:39.000Z A BGP looking glass is a web interface that lets you run a small set of read-only BGP commands against a router you do not own. The router belongs to an ISP, an IXP, or a content provider, and they expose it so that the rest of the internet can see what their network sees. When someone reports "your prefix is unreachable from our side of the internet," the first thing you do is open three looking glasses on three different networks and check whether your prefix is actually being advertised the way you think. For the cluster overview, see the [BGP complete guide](https://www.pinglabz.com/bgp/). For how looking glasses interact with MPLS L3VPNs and SRv6 deployments, see the [MPLS pillar](https://www.pinglabz.com/mpls/). ## What a looking glass shows you Every looking glass surfaces a small, deliberately limited subset of BGP and routing commands. The standard menu, regardless of vendor or provider, is some combination of: `show ip bgp X.X.X.X/Y` "What does this network see for my prefix? Which AS\_PATH? Which next-hop? Best path?" `show ip bgp summary` "How many BGP neighbors does this router have, and are they Established?" `show ip route X.X.X.X` "After best-path selection, which next-hop is in this router's FIB?" `ping` / `traceroute` "Can this router reach my prefix in the data plane?" Notice what looking glasses do not let you do: nothing that changes state, nothing that shows confidential information, and nothing that lets you exfiltrate the full routing table. `show ip bgp` with no arguments is usually disabled or rate-limited. ## Why this is useful BGP is a path-vector protocol. Every AS in the world sees a potentially different best path to your prefix. If a customer in Brazil cannot reach your network in Singapore, the cause might be: - Your upstream stopped advertising the prefix on one peering - A transit AS in the middle has a route-map deny against your prefix or your AS - An RPKI ROA mismatch is causing a tier-1 to drop your route - An IXP route server's filter is removing it - The path is being prepended so heavily that an alternate path is preferred and the alternate has a problem You cannot diagnose any of these from your own router because your own router only sees its own view of the table. The whole point of a looking glass is to give you a borrowed view from somewhere else in the topology. ## The public looking glasses worth bookmarking Most large networks run a public looking glass. The ones below are the ones operators actually keep open during BGP troubleshooting. Hurricane Electric URL patternlg.he.net Why useful One of the largest tier-2/peering networks. Great default sanity check. NTT (GIN-LG) URL patternntt.net/looking-glass Why useful Multiple PoPs globally. Good for "is my prefix visible in Japan vs Europe?" Cogent URL pattern cogentco.com/looking-glass Why useful Heavy tier-1 visibility, good for routing-policy disputes Telia / Arelion URL patternlg.arelion.com Why useful Europe-heavy view; complements US-anchored networks Lumen / Level3 URL patternlookingglass.lumen.com Why useful Wide North American footprint RIPE NCC URL patternstat.ripe.net Why useful Not a single router but an aggregated RIB view from RIS collectors. Best single tool for "who is propagating my prefix anywhere on the internet right now?" bgp.tools URL patternbgp.tools Why useful Aggregator + visual AS graph. Modern UX. Free. For the "is this prefix being received" question, RIPE Stat and bgp.tools are usually the fastest answer. For the "where is the AS\_PATH bending" question, the per-provider looking glasses are necessary because you want to see the actual chosen path from that specific vantage point. ## A standard troubleshooting recipe Customer reports "your prefix is unreachable from our location." Before you touch your own configuration: 1. Open **RIPE Stat** on your prefix. Confirm the prefix is being seen by the RIS collectors. If RIPE shows zero AS paths, your prefix is genuinely not being propagated anywhere on the internet and your problem is upstream of you (your transit dropped you). 2. Open **three geographically diverse looking glasses** (HE in the US, Telia in EU, NTT in Asia). Query your prefix on each. Compare AS\_PATHs. If two of three show your prefix and one does not, the missing one's network has a problem. 3. Run a **traceroute from the affected vantage point to your prefix**. Most looking glasses include traceroute. If the trace dies inside a specific AS, that AS is the failure point. 4. **Check RPKI status**. RIPE Stat shows whether your route is RPKI-valid, invalid, or not found. If your ROA expired and your prefix is now invalid, several tier-1s will drop it within minutes. This four-step recipe localizes the problem in 5 to 10 minutes for most cases. It is the prerequisite to any conversation with your upstream support. ## Hosting your own looking glass Running your own public looking glass takes about an afternoon and is a small but meaningful peering-relationship investment. Other networks can debug their issues against you faster, which makes peering with you more attractive. The two open-source projects that have stood the test of time: [hyperglass](https://github.com/Cougar/hsdn-lg?ref=pinglabz.com) Stack Python (FastAPI) + JS frontend Notes Modern. Multi-vendor (Cisco IOS XE, NX-OS, Juniper, Arista, Cumulus). Per-location config. Auth not required for read-only. bird-lg / bird-lg-go StackPython or Go Notes Lighter weight. Speaks to BIRD route servers directly. Heavily used at IXPs. A minimal hyperglass deployment looks like: ``` git clone https://github.com/checktheroads/hyperglass cd hyperglass docker compose up -d # Edit hyperglass/devices.yaml with your router IP, SSH user, vendor ``` Lock it down at the edge: - Put the looking glass behind a reverse proxy (nginx, Caddy) with strict rate limiting (10 req/min/IP is generous) - Bind the backend SSH to a dedicated read-only user on the routers with a tightly-scoped privilege role (Cisco IOS XE has `privilege exec level 0` with specific show commands permitted) - Never expose write commands. Standard. But worth saying. - Publish the looking glass URL on your PeeringDB record so other operators can find it ## The looking-glass commands you actually permit For a Cisco IOS XE router being polled by hyperglass or similar, the role for the looking-glass user should permit: ``` username lg-readonly privilege 1 secret privilege exec level 1 show ip bgp privilege exec level 1 show ip bgp summary privilege exec level 1 show bgp ipv4 unicast privilege exec level 1 show bgp ipv6 unicast privilege exec level 1 show ip route privilege exec level 1 ping privilege exec level 1 traceroute ``` Everything else stays at higher privilege levels the read-only user cannot reach. ## Limits to be aware of - **Looking glasses show one vantage point at a time.** A "good" result from one looking glass does not mean your prefix is universally reachable. - **The router behind the looking glass may not run the same RIB as the network's edge.** Most providers expose a dedicated route-collector, not a production router. The view is representative but not authoritative. - **RPKI views in looking glasses can lag.** If you just published a new ROA, give it 15 to 60 minutes before relying on looking-glass output to confirm propagation. - **Some looking glasses cache.** If you see a stale result, hit refresh; if a particular provider's looking glass always shows the same answer, try a different provider. ## Key takeaways A looking glass is the only way to see your prefix from somewhere other than your own router. Bookmark four or five geographically diverse ones, learn the four-step recipe (RIPE Stat first, three providers second, traceroute third, RPKI fourth), and you can localize most BGP reachability problems in under ten minutes. If you run a public AS, hosting your own looking glass via hyperglass is one afternoon of work and makes your network easier to peer with. For the underlying BGP mechanics, see the [BGP pillar](https://www.pinglabz.com/bgp/). Take the BGP reference with you The free BGP field-reference PDF: path attributes, best-path order, and the show commands that matter. Delivered by email, no card required. [Get the free PDF](https://www.pinglabz.com/bgp-cheatsheet/) ### OSPF Adjacency States: The 8-State FSM, Explained URL: https://www.pinglabz.com/ospf-adjacency-states/ Last updated: 2026-08-01T19:31:06.000Z Every OSPF neighbor relationship passes through eight named states on its way from "we just saw each other" to "we agree on the topology." When OSPF works, the states scroll past in seconds and nobody looks at them again. When OSPF breaks, the state the neighbor is stuck in tells you exactly where in the conversation it failed. This post walks through all eight states, what triggers a transition, and what to check when a neighbor refuses to move past a specific one. For the cluster overview, see the [OSPF complete guide](https://www.pinglabz.com/ospf/). For the analogous state machine on the other major routing protocol, see the [BGP pillar](https://www.pinglabz.com/bgp/). ## The 8-state FSM at a glance Down #1 What it means No Hellos received from this neighbor. The default initial state. What you check if stuck here L1/L2 connectivity, OSPF enabled on interface, correct area Attempt #2 What it means NBMA only. The router is unicasting Hellos hoping for a reply. What you check if stuck here `neighbor` statements under `router ospf`, NBMA reachability Init #3 What it means We saw a Hello from this neighbor but our own Router ID is not in their neighbor list yet. What you check if stuck here One-way Hellos. Auth mismatch. ACL blocking return traffic. 2-Way #4 What it means Bidirectional Hello exchange confirmed. On broadcast networks, non-DR/BDR pairs stop here. What you check if stuck here If two DROTHERs sit at 2-Way that is normal. If two routers that should peer fully sit here, check priorities. ExStart #5 What it means Master/slave election for the upcoming database exchange. What you check if stuck here MTU mismatch. Auth mismatch. Both sides claim master forever. Exchange #6 What it means DBD packets sent describing each side's LSA headers. What you check if stuck here MTU again (most common cause of stuck-in-Exchange). DBD packet drops. Loading #7 What it means One side has requested LSAs it does not have yet via LSR; the other is sending them via LSU. What you check if stuck here Database corruption is rare. Usually transient. Full #8 What it means The two routers have synchronized link-state databases. This is the working steady state. What you check if stuck here Nothing. This is the goal. ## State 1 - Down The starting state. Either OSPF just came up on the interface or the dead-interval expired on a previously-known neighbor. Stays here until a Hello packet arrives from any source on the segment. Routers that should be neighbors but stay Down indefinitely usually have a Layer 1 or 2 issue (cable, switch port assignment, no carrier) or OSPF is not actually enabled on the interface. Confirm with `show ip ospf interface brief`. If your expected interface is not listed, the `network` statement under `router ospf` does not cover it, or the interface-level `ip ospf 1 area X` command is missing. ## State 2 - Attempt (NBMA only) Only seen on non-broadcast multi-access networks (Frame Relay, ATM) where multicast is not available and the router must unicast Hellos based on configured `neighbor` statements. If the neighbor never responds, the state stays in Attempt until you intervene. In a modern Cisco IOS XE network you will almost never see this state. It survives in OSPF for historical reasons but most current designs use point-to-multipoint or point-to-point network types instead, which skip Attempt entirely. ## State 3 - Init A Hello arrived from the neighbor, but the Hello did not list our own Router ID in its neighbor section. This means we are seeing them, but they are not seeing us back. One-way communication. The classic causes: - **OSPF authentication mismatch.** One side sends authenticated Hellos; the other side either does not authenticate or uses a different key. The authenticated side drops the unauthenticated Hellos silently. - **An inbound ACL on the neighbor's interface dropping OSPF (protocol 89).** Hellos out work; Hellos in get filtered. - **Hello / dead-interval mismatch.** The Hello arrives but is discarded as malformed for that interface. - **Network-type mismatch.** One side configured as point-to-point, the other as broadcast. Hellos look different. Stuck-in-Init is almost always one of these four. Run `show ip ospf interface ` on both sides and diff the output line by line. Before you diff anything, note that INIT tells you which direction is broken, and that alone cuts the list down. This walkthrough of [why an OSPF neighbor sits in INIT and never moves](https://www.pinglabz.com/ospf-stuck-in-init-one-way-hello/) reads the direction off a live pair of routers and works back to the specific cause. ## State 4 - 2-Way Bidirectional Hellos confirmed. The routers know they see each other. On broadcast and NBMA networks, the next thing OSPF does is run the DR/BDR election. Routers that lose the election (DROTHERs) *do not progress past 2-Way with each other*. They only go past 2-Way with the DR and BDR. Two DROTHERs sitting at 2-Way forever is correct behavior and not a problem. Two routers that should fully peer but sit at 2-Way usually have a priority misconfiguration. Confirm with `show ip ospf neighbor` and look at the State column. The DR will show as FULL/DR, the BDR as FULL/BDR, and other neighbors as 2-WAY/DROTHER (which is fine). ## State 5 - ExStart The routers are about to exchange databases. First they have to elect a master and a slave to control the DBD sequence numbering. The router with the higher Router ID wins. The most famous cause of "stuck in ExStart/Exchange" is MTU mismatch. OSPF includes the interface MTU in the DBD packet. If the receiver's interface MTU is smaller than what is in the DBD, the receiver drops the DBD and the conversation stalls. Both routers retransmit, neither makes progress. The fix is one of: - Make both interface MTUs match (usual answer) - Use `ip ospf mtu-ignore` on the interface to skip the MTU check (workaround; consider whether the MTU mismatch will cause other problems) Which of those two you want depends on why the MTUs differ in the first place, and finding the offending interface is a two-command job once you know where the value is carried. [Tracking down an OSPF MTU mismatch](https://www.pinglabz.com/ospf-mtu-mismatch-troubleshooting/) shows the DBD field that triggers the drop, both fixes side by side, and how to reproduce the whole thing in a lab. ## State 6 - Exchange DBD packets flow describing each side's LSA headers. Each router builds a list of LSAs the other side has that it does not. Stuck-in-Exchange is almost always the same MTU problem as stuck-in-ExStart, just discovered slightly later when a DBD packet happens to be larger than the receiver's MTU. Less commonly, stuck-in-Exchange can be caused by an extreme amount of packet loss on the segment, which prevents the DBD chain from completing before retransmission timers expire indefinitely. ## State 7 - Loading Each side has identified which LSAs it needs from the other. It sends Link State Requests (LSRs). The other side replies with Link State Updates (LSUs) containing the requested LSAs. When the request list is empty on both sides, the conversation advances to Full. Stuck-in-Loading is uncommon. When it happens, the typical cause is database corruption on one side or packet drops of LSU packets specifically (less common than DBD drops, but possible if there is an ACL or QoS classifier mishandling protocol 89). ## State 8 - Full The two routers have synchronized link-state databases. They will continue to exchange routine Hellos at the configured Hello interval, and will exchange LSAs as topology changes occur, but the relationship itself is steady state. Time spent in Full is unbounded; an adjacency can stay Full for years. ## Reading show ip ospf neighbor The one command you will use 80% of the time: ``` R2# show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 1.1.1.1 1 FULL/DR 00:00:32 10.20.0.1 Ethernet0/0 3.3.3.3 1 FULL/BDR 00:00:34 10.20.0.3 Ethernet0/0 4.4.4.4 0 FULL/ - 00:00:38 10.30.30.4 Ethernet0/1 ``` Three things to read off this output: - **State column.** FULL is the goal. Anything else needs investigation. - **The slash suffix.** /DR, /BDR, /DROTHER tell you the DR election outcome on the segment. The dash for 4.4.4.4 means a P2P link (no DR election). - **Dead Time.** Counting down. If it hits zero, the adjacency drops. A dead time that climbs back up means a Hello arrived. ## The debug command, used with caution For diagnosing stuck states, the most useful debug is: ``` R2# debug ip ospf adj ``` This prints every adjacency state transition. Run it on the side that is failing, watch for the transition that does not happen, and the line above tells you why. Turn it off immediately when you have the answer; on a busy router it produces enough log volume to flood the console. ## Key takeaways The OSPF neighbor FSM has eight states. Down means no contact. Init means one-way. 2-Way means bidirectional but no database sync yet. ExStart through Loading are the database synchronization phases. Full is the goal. The state a stuck neighbor is in tells you which phase of the conversation failed, and there is a small finite list of causes for each state. Memorize the table at the top of this post and the troubleshooting time on any OSPF adjacency problem drops by a factor of three. For broader OSPF coverage, see the [OSPF pillar](https://www.pinglabz.com/ospf/). ### VLAN, Subnet, and Broadcast Domain: What's the Difference? URL: https://www.pinglabz.com/vlan-subnet-broadcast-domain-difference/ Last updated: 2026-06-13T20:07:54.000Z VLAN, subnet, and broadcast domain are three terms that get used interchangeably in conversation and treated as identical in most network drawings. They are not identical. Each one describes a different thing, sitting at a different layer of the stack, and the relationship between them is what trips up engineers when they design or troubleshoot anything past the basics. For the L2 fundamentals, see the [VLAN and Layer 2 switching pillar](https://www.pinglabz.com/vlans-layer-2-switching/). For why the same conversation matters at L3, see the [IPv6 pillar](https://www.pinglabz.com/ipv6/). ## Quick definitions, then the relationships Broadcast domain LayerL2 What it is The set of devices that receive a frame sent to the broadcast MAC (FF:FF:FF:FF:FF:FF). Defined by what is physically and logically reachable without crossing a Layer 3 hop. VLAN LayerL2 What it is A logical partition of a physical switch that creates a separate broadcast domain inside the switch. Identified by a 12-bit VLAN ID (1-4094). Subnet LayerL3 What it is A range of IP addresses with a common prefix length. Defined by network address and subnet mask. Pure IP concept; switches do not care. The relationship that matters most: - **One VLAN = one broadcast domain.** Every VLAN creates a separate L2 segment. Frames flooded inside VLAN 10 do not leak into VLAN 20 on the same switch. - **One VLAN usually = one subnet.** In conventional designs, you assign a single subnet (say 10.10.10.0/24) to a single VLAN (say VLAN 10). This is convention, not law. The two concepts live at different layers. - **One subnet does not have to map to one VLAN.** You can put one subnet across multiple VLANs (rare, usually a mistake) or multiple subnets inside one VLAN (called secondary IPs, also usually a mistake). ## Why a broadcast domain matters Every broadcast frame is processed by every device in the broadcast domain. ARP requests, DHCP discovers, NetBIOS announcements, IPv6 ND (multicast, but inside the same L2 segment), STP BPDUs, link-local services. The more devices in a broadcast domain, the more "background noise" every device has to ignore, and the larger the blast radius of any L2 misbehavior (a broadcast storm, a misbehaving printer flooding NetBIOS, a duplicate IP causing ARP confusion). Practical sizing rule from the field: keep a broadcast domain under 250 to 500 active hosts. Beyond that, ARP table churn alone starts to cost CPU on the L3 gateway. Routers handle a few hundred broadcasts per second without breaking a sweat. A flat /16 network with 8,000 devices generates enough ARP and DHCP noise to make troubleshooting miserable. ## Why a VLAN is not the same as a broadcast domain A VLAN is the *mechanism*. The broadcast domain is the *consequence*. If you turn off STP and connect two access ports on different VLANs by accident with a crossover cable, you have just merged two broadcast domains while leaving the VLAN configuration on each switch unchanged. The configuration says VLAN 10 and VLAN 20 are separate. The traffic disagrees. This is why VLAN configuration alone is not a security boundary. VLAN-hopping attacks (double-tagging on access ports configured as trunks, switch-spoofing) exploit exactly this gap between "what the config says" and "what the broadcast domain actually contains." ## Why a subnet is not the same as a VLAN A switch does not look at IP addresses to decide where to flood a broadcast frame. It looks at VLAN tags and MAC tables. So the subnet a host belongs to has no influence on which broadcast domain it lives in. You can put a host with IP 10.10.10.5/24 in VLAN 20\. It will work. The host will broadcast ARP for its 10.10.10.0/24 subnet mates. Those ARPs will reach every other device in VLAN 20, including hosts with completely different IPs (172.16.5.0/24, say). None of them will answer because none of them are in the right subnet. The host will conclude its neighbors are unreachable. The configuration is "valid" at every layer. It just does not work. This is why misalignment between VLAN and subnet is one of the most common access-port misconfigurations. The right intuition is: a VLAN is a fence; a subnet is a phone book. The fence keeps frames from crossing. The phone book tells hosts who to call. Both have to agree. ## Inter-VLAN routing: where the layers meet The moment two subnets live in two different VLANs, traffic between them needs a router. A switch will not move a frame from VLAN 10 to VLAN 20\. The standard mechanisms are: Switched Virtual Interfaces (SVIs) How it works A Layer 3 switch creates a virtual interface per VLAN with an IP address. The switch routes between SVIs in hardware. When to use Default modern design. Used on any L3 switch from a Catalyst 9300 upward. Router-on-a-stick How it works A trunk from the switch to a router. The router has a sub-interface per VLAN, each with its own IP. When to use Lab work, or branches with an L2-only access switch and a single router for L3. External L3 gateway (firewall, FHRP) How it works The L2 switch trunks all VLANs to a firewall pair or a routed pair running HSRP/VRRP. When to use Security-sensitive segments where you want all inter-VLAN traffic to traverse a stateful firewall. ## A worked example You have a switch with three VLANs and three subnets: - VLAN 10, 10.10.10.0/24, the user VLAN - VLAN 20, 10.10.20.0/24, the server VLAN - VLAN 30, 10.10.30.0/24, the voice VLAN You have three broadcast domains (one per VLAN). You have three subnets. A broadcast sent from a user in VLAN 10 reaches every other user in VLAN 10, no servers in VLAN 20, and no phones in VLAN 30\. A user wanting to reach a server in VLAN 20 sends the packet to its default gateway (the SVI for VLAN 10), which routes it into VLAN 20 and forwards the frame to the destination MAC. Now consider a different scenario: someone configures a server's NIC manually with 10.10.20.5/24 but plugs it into a switch port assigned to VLAN 10\. The server is in: - Broadcast domain: VLAN 10 - VLAN: 10 - Subnet: 10.10.20.0/24 It will ARP for 10.10.20.1 (its configured gateway). The ARP request flies across VLAN 10\. Nothing in VLAN 10 has IP 10.10.20.1\. The server's default gateway resolution fails. No traffic leaves the server. The cable is fine, the switch port is up, the IP looks correct, but the subnet is wrong for that VLAN. This is the canonical "VLAN and subnet are not the same thing" failure. ## IPv6, link-local scope, and where this changes IPv6 reuses the same architecture but renames the pieces. The "broadcast domain" concept becomes the "link" or "link-local scope." IPv6 has no broadcast at all; ARP is replaced by Neighbor Discovery, which uses multicast scoped to the link. Everything about VLANs and links and subnets still maps cleanly: one VLAN equals one link, one link usually equals one IPv6 prefix, and a host with the wrong prefix on the right link fails the same way it does in IPv4. ## Key takeaways VLAN, subnet, and broadcast domain describe different things at different layers. A VLAN is an L2 mechanism. A broadcast domain is the L2 consequence. A subnet is an L3 concept that lives entirely above the switch's awareness. By convention, one VLAN equals one broadcast domain equals one subnet, and this convention works because nobody benefits from breaking it. When the three get out of alignment, hosts come up clean at every individual layer but cannot talk to anything, and the misalignment is invisible without checking all three. For the L2 mechanics that enforce these boundaries, see the [VLAN pillar](https://www.pinglabz.com/vlans-layer-2-switching/). ### SD-WAN: What Is It? The 5-Minute Primer URL: https://www.pinglabz.com/sd-wan-what-is-it/ Last updated: 2026-06-13T20:07:54.000Z SD-WAN is the marketing umbrella that replaced "we just bought a bunch of MPLS circuits" as the default enterprise WAN answer. The technology underneath is older than the term, the value proposition has shifted three times since 2015, and the vendor landscape is now consolidating. This is the five-minute primer: what SD-WAN actually is, what problem it solves, and how it differs from the MPLS designs it is replacing. For the deeper dive on the technology and configuration, see the [SD-WAN complete guide](https://www.pinglabz.com/sd-wan/). For the WAN technology it is displacing, see the [MPLS pillar](https://www.pinglabz.com/mpls/). ## What SD-WAN is, in one sentence SD-WAN is a software-defined overlay network that runs across whatever physical transport you happen to have, with a centralized controller pushing policy down to edge devices and a data plane that picks paths per-application in real time based on link health. Three words in that sentence carry the weight: **overlay**, **controller**, and **per-application**. Each represents one of the historical things wrong with traditional WAN that SD-WAN solved. ## What it actually replaces The branch office WAN of 2010 looked like this: - One MPLS L3VPN circuit from a carrier, hand-priced per location, 30 to 90 day install lead time - A backup DSL or LTE link that the router would failover to via static routing or a slow IGP reconvergence - All internet-bound traffic backhauled to a central data center for filtering, then back out to the cloud - QoS was the only knob, and tuning it required carrier coordination The branch office WAN of 2025 looks like this: - Two or three commodity broadband circuits (cable, fiber, LTE/5G), often from different ISPs - An SD-WAN edge device (or a cluster of them) that builds IPsec tunnels across all available transports simultaneously - Per-application path selection that can send Microsoft 365 traffic out the local internet break, voice traffic over the lowest-latency path, and bulk backup over the highest-throughput path, all from the same device - Central orchestration via a controller, with policy expressed in business terms ("Office 365 gets best-quality path; YouTube gets cheapest path") ## The four-plane architecture Every SD-WAN platform, regardless of vendor, organizes itself into four planes. Knowing them helps when you read vendor docs because the names map cleanly across products. Management What it does The UI you log into. Policy authoring, dashboards, reports. Cisco Catalyst SD-WAN namevManage VMware VeloCloud name VCO (VeloCloud Orchestrator) Control What it does Learns the overlay topology. Distributes routing and policy. Cisco Catalyst SD-WAN namevSmart VMware VeloCloud nameVCO + Gateways Orchestration What it does Authenticates edges, manages certificates, brokers initial control connections. Cisco Catalyst SD-WAN namevBond VMware VeloCloud nameVCO Data What it does The edge devices themselves. They build the tunnels and forward traffic. Cisco Catalyst SD-WAN namecEdge / vEdge VMware VeloCloud nameVeloCloud Edge Cisco's separation between vSmart and vBond is unique to their architecture. VMware folds them together. Versa, Fortinet, Palo Alto Prisma all split the planes differently but the conceptual model is the same. ## The two things that make SD-WAN different from MPLS ### 1\. Transport independence An MPLS L3VPN is a service contract. The carrier promises a certain bandwidth between a defined set of sites, with QoS classes you negotiate up front, on a circuit they own end-to-end. An SD-WAN overlay does not care what the underlying transport is. The edge builds IPsec tunnels across whatever IP-reachable transport you give it. Three cable modems from three different ISPs work. A 5G hotspot plus a Starlink link works. Adding a new transport is a software change on the edge, not a carrier truck-roll. ### 2\. Application-aware path selection Traditional routing makes one path decision per destination, based on routing-protocol metrics. SD-WAN makes a path decision per flow, per application, refreshed every few seconds based on real-time link telemetry. If your fiber link's loss spikes to 2% for the next 30 seconds, the edge notices, marks the path as out-of-policy for voice and video, and shifts those flows to the cable backup, while keeping bulk-backup traffic on the fiber where loss does not hurt. That decision happens at the edge, in milliseconds, without a routing convergence event. ## The three deployment styles Hybrid (MPLS + broadband) What it looks like Keep the MPLS for predictable performance to the data center; add broadband for everything else. When organizations pick it Large enterprises with existing MPLS contracts they cannot exit yet. Transitional architecture. Internet-only What it looks like Multiple broadband circuits, no MPLS. SD-WAN does the heavy lifting. When organizations pick it Greenfield branches, retail, healthcare. Cost-driven decision. Cloud on-ramp What it looks like SD-WAN edges peer directly with cloud provider points-of-presence (Azure vWAN, AWS Cloud WAN, Google Cloud NCC). When organizations pick it SaaS-heavy organizations where the data center is no longer the gravitational center. ## Where SD-WAN does not help SD-WAN is not magic. The places it routinely disappoints are predictable. - **You still need underlay bandwidth.** SD-WAN cannot synthesize throughput. Three saturated cable links plus SD-WAN is still three saturated cable links. - **Latency-sensitive flows on long-haul links.** If your branch is in Singapore and your data center is in Frankfurt, the speed of light is the speed of light. SD-WAN picks the best available path but cannot beat physics. - **Strict regulatory requirements for circuit isolation.** Some regulated environments (defense, certain financial) still require dedicated circuits where SD-WAN's "any IP-reachable transport" model does not pass an audit. - **Branch-of-one without diverse transport.** If the branch has exactly one ISP available, SD-WAN reduces to "an edge device with policy controls" with no transport-diversity benefit. ## The vendor landscape in 30 seconds Cisco Catalyst SD-WAN (Viptela) Strength Deepest IOS XE integration. Strong cloud on-ramps. Mature CLI. Watch out for License complexity. vManage scale ceilings on the largest deployments. VMware VeloCloud (now Broadcom) Strength Easiest controller UX. Strong DMPO (Dynamic Multipath Optimization). Watch out for Broadcom acquisition uncertainty. Roadmap clarity. Fortinet Secure SD-WAN Strength Best price/performance. Built into FortiGate firewalls so no extra appliance. Watch out for Security-vendor mindset; networking features sometimes lag. Palo Alto Prisma SD-WAN (CloudGenix) Strength Strong app identification. Tight integration with Prisma Access SASE. Watch out for Smaller ecosystem. Premium pricing. Versa Networks Strength Most feature-complete single-vendor SASE story. Watch out for Steepest learning curve. Smaller install base. ## Key takeaways SD-WAN is an overlay that runs on top of any IP transport, picks paths per-application based on real-time link health, and is managed from a central controller. It replaces the static carrier MPLS model with something cheaper and more flexible, at the cost of taking the WAN architecture decision back from the carrier and putting it on your team. If you are designing a greenfield branch network in 2026, internet-only SD-WAN is the default; MPLS shows up only when latency or compliance forces it. For the deeper architecture and configuration walkthroughs, see the [SD-WAN pillar](https://www.pinglabz.com/sd-wan/). ### Enable IEEE 802.1X Authentication on Windows 11 (Manual + Group Policy) URL: https://www.pinglabz.com/enable-ieee-802-1x-authentication-windows-11/ Last updated: 2026-06-13T20:07:54.000Z Windows 11 ships with the supplicant software needed to authenticate to an 802.1X-enabled wired network, but none of it works out of the box. The Wired AutoConfig service is stopped by default. Even when you start it, every adapter still needs the Authentication tab enabled per-interface, and the EAP method has to match what your RADIUS server expects. This post walks through both paths to get there: the manual click-through for a lab machine, and the Group Policy version for a fleet. For the 802.1X protocol fundamentals, see the [802.1X complete guide](https://www.pinglabz.com/802-1x/). ## The two pieces Windows needs Enabling IEEE 802.1X authentication on Windows 11 always comes down to the same two things: 1. **The Wired AutoConfig service (dot3svc) must be running.** This is the supplicant. Without it, the network adapter has no 802.1X stack to draw from. 2. **The adapter's Authentication tab must be enabled and configured.** The tab is hidden until dot3svc is running, which is one of the more confusing parts of the experience. Get both right and the adapter will start sending EAP-Response/Identity frames the moment the link comes up. ## Manual setup: one machine, one adapter Use this for your lab box or a single user troubleshooting a broken adapter. ### Step 1: Start the Wired AutoConfig service Open an elevated PowerShell: ``` Set-Service -Name dot3svc -StartupType Automatic Start-Service -Name dot3svc Get-Service -Name dot3svc ``` The last command should report Status as Running. If it does not start, check the System event log for service-control errors. The most common cause is a Group Policy explicitly disabling it. ### Step 2: Enable the Authentication tab on the adapter Open **Control Panel > Network and Sharing Center > Change adapter settings**. Right-click the Ethernet adapter, choose **Properties**. You will now see an **Authentication** tab between General and Sharing. If the tab is still missing, close the dialog and reopen it; the tab is only injected after dot3svc starts. On the Authentication tab: - Check **Enable IEEE 802.1X authentication** - Set **Choose a network authentication method** to match your environment. For most enterprise deployments this is *Microsoft: Protected EAP (PEAP)* or *Microsoft: Smart Card or other certificate* (EAP-TLS). - Decide whether to check **Remember my credentials for this connection** and whether to **Fallback to unauthorized network access**. In a strict-enforcement environment, leave fallback unchecked. ### Step 3: Configure the EAP method Click the **Settings...** button next to the EAP method. The dialog that opens depends on which method you chose. PEAP (with MS-CHAPv2) Trusted root CA (your internal CA that signed the RADIUS server cert), server-name validation pattern, inner method (EAP-MSCHAPv2), and whether to automatically use the Windows logon name and password. EAP-TLS Same trusted-root and server-name settings, plus which user/computer certificate to present (Smart Card vs Certificate Store). EAP-MSCHAPv2 directly Just the credential auto-use checkbox. No transport encryption inside EAP. Only acceptable if the outer transport (PEAP/TTLS) provides the encryption. ### Step 4: Set Additional Settings for authentication mode Back on the Authentication tab, click **Additional Settings...**. Set **Specify authentication mode** to: - **User authentication** if you want only user creds (no machine pre-login auth) - **Computer authentication** if you want the machine to auth before any user logs in (useful for GPO push, Windows Update over wired) - **User or computer authentication** for the typical "machine first, user when they log in" pattern. This is what most enterprise deployments use. ### Step 5: Verify Disconnect and reconnect the cable. On the switch side, run `show authentication sessions interface Gi1/0/X`. You should see Status: Authorized and Method: dot1x. On the Windows side, open Event Viewer and navigate to **Applications and Services Logs > Microsoft > Windows > Wired-AutoConfig > Operational**. Look for Event ID 15500 (authentication successful). ## Fleet setup: Group Policy For more than a handful of machines, Group Policy is the only sane path. The relevant policy lives in two places. ### Service auto-start via GPO Edit a GPO that applies to your computer OU. Navigate to **Computer Configuration > Policies > Windows Settings > Security Settings > System Services**. Find **Wired AutoConfig**, double-click, define the policy, and set startup mode to **Automatic**. ### Wired network policy via GPO Same GPO, navigate to **Computer Configuration > Policies > Windows Settings > Security Settings > Wired Network (IEEE 802.3) Policies**. Right-click and create a new wired network policy. General Use Windows Wired Auto Config service. Auto-connect. Security Enable use of IEEE 802.1X. Select the EAP method (PEAP or EAP-TLS). Set Additional Settings authentication mode. Optionally cache user info, enforce 802.1X retry count. Once the GPO refreshes (force with `gpupdate /force`), every machine in scope picks up the policy. New Ethernet adapters inherit it automatically. You do not need to touch individual adapter properties on each machine. ## Useful PowerShell for batch checks To confirm dot3svc is running across a list of machines: ``` $machines = Get-Content C:\machines.txt $machines | ForEach-Object { $status = Get-Service -Name dot3svc -ComputerName $_ -ErrorAction SilentlyContinue [PSCustomObject]@{ Machine = $_ Status = $status.Status Mode = $status.StartType } } | Format-Table ``` To dump wired profile settings on a local machine: ``` netsh lan show profiles netsh lan show interfaces ``` The `show interfaces` output includes the 802.1X authentication state (Authenticated, Authenticating, Held, or Authentication Failed) which is the supplicant-side view of what the switch reports as port status. ## Common gotchas Authentication tab missing from adapter properties dot3svc not running. Start it, then reopen the properties dialog. "The credentials provided by the server could not be validated" Server certificate not chained to a trusted root, or server-name validation does not match the RADIUS server's cert CN/SAN. Re-check the PEAP/TLS settings dialog. Authentication works after logon but fails before Authentication mode set to User authentication only. Change to User or computer authentication so the machine can auth pre-logon. Domain join works but Group Policy never applies on a new machine No machine authentication path. Either configure Single-Sign-On under the EAP settings, or fall back to MAB on first connect. Authentication succeeds but no IP is assigned VLAN assignment from RADIUS pointed the port at a VLAN where DHCP is not available. Confirm the VLAN exists and has a relay-agent. ## Key takeaways Enabling IEEE 802.1X authentication on Windows 11 takes two steps you have to get right: start dot3svc, then configure the adapter's Authentication tab with the EAP method your RADIUS server expects. For more than one or two machines, push both via Group Policy. The most common failure is forgetting to start the service before opening adapter properties, which hides the Authentication tab entirely and makes the feature look broken. For the protocol-side view of what the switch does with these credentials, see the [802.1X pillar](https://www.pinglabz.com/802-1x/). ### What is 802.1X Authentication? End-to-End Flow with Real show Output URL: https://www.pinglabz.com/what-is-802-1x-authentication/ Last updated: 2026-06-13T20:07:55.000Z 802.1X is the IEEE standard that turns a switch port into an authentication gateway. Plug a laptop into a port, and instead of getting an IP and reaching the network, the port challenges the device for credentials first. If the credentials check out, the port opens. If they do not, the port stays closed or drops the device into a quarantine VLAN. This post walks through the end-to-end flow with real `show authentication sessions` output, the three players involved, and the protocol exchange that ties them together. For the cluster overview, see the [802.1X complete guide](https://www.pinglabz.com/802-1x/). For ACL syntax on a parallel security plane, the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/) covers the firewall-side perspective. ## The three roles in 802.1X Every 802.1X conversation involves three named actors. The names matter because every Cisco command, every debug, and every RADIUS attribute references one of them. Supplicant Who plays it The endpoint (laptop, IP phone, printer, IoT device) What it does Sends its identity and credentials when the switch asks Authenticator Who plays it The switch port (or wireless AP/WLC) What it does Sits between supplicant and authentication server. Forwards EAP messages. Enforces the port state. Authentication server Who plays it RADIUS server (ISE, FreeRADIUS, NPS) What it does Validates credentials. Returns Access-Accept or Access-Reject. Optionally returns dynamic VLAN, ACL, or SGT. The switch is the only component that actually controls port forwarding. The RADIUS server makes the decision; the switch enforces it. ## The end-to-end flow From the moment a device plugs in to the moment traffic starts flowing, the sequence is: 1. **Link comes up.** The switch port is in *unauthorized* state. Only EAPoL frames are allowed through. Everything else is dropped. 2. **Switch sends EAP-Request/Identity.** The switch periodically asks "who are you?" via an EAPoL frame to the multicast destination 01:80:C2:00:00:03. 3. **Supplicant responds with EAP-Response/Identity.** The endpoint sends its username (machine name for computer auth, user@domain for user auth). 4. **Switch wraps the EAP payload in RADIUS and sends to the AAA server.** RADIUS Access-Request carrying the EAP-Message attribute. 5. **EAP method negotiation.** The server picks an EAP method (PEAP, EAP-TLS, EAP-FAST). The exchange runs through the switch transparently. The switch never sees the credentials. 6. **Server returns RADIUS Access-Accept (or Reject).** On Accept, the message may carry additional attributes: VLAN assignment (Tunnel-Private-Group-ID), a dACL name (Cisco-AV-Pair), session timeout, reauth timer. 7. **Switch moves the port to authorized state.** Traffic flows. The session is now tracked in the authentication session database. ## The minimum viable IOS XE config This is enough to get one switch port talking 802.1X to a RADIUS server. Add monitor-mode and MAB later. ``` aaa new-model ! radius server ISE-PRIMARY address ipv4 10.10.10.50 auth-port 1812 acct-port 1813 key STRONG_RADIUS_SECRET ! aaa group server radius ISE server name ISE-PRIMARY ! aaa authentication dot1x default group ISE aaa authorization network default group ISE aaa accounting dot1x default start-stop group ISE ! dot1x system-auth-control ! interface GigabitEthernet1/0/10 switchport access vlan 100 switchport mode access authentication port-control auto mab dot1x pae authenticator dot1x timeout tx-period 10 spanning-tree portfast ``` Two things to notice. First, the `mab` line enables MAC Authentication Bypass as a fallback for devices (printers, badge readers) that cannot speak 802.1X. The switch first tries dot1x; if the supplicant never answers, it falls back to sending the MAC address as the identity. Second, `port-control auto` is the line that activates 802.1X. The other options are `force-authorized` (always open, no auth) and `force-unauthorized` (always closed). `auto` is the only one that actually does authentication. ## Reading the show output The single most useful command for 802.1X troubleshooting is `show authentication sessions interface Gi1/0/10 details`. A successful authentication looks like this: ``` Switch# show authentication sessions interface Gi1/0/10 details Interface: GigabitEthernet1/0/10 MAC Address: 0050.56a1.b2c3 IPv6 Address: Unknown IPv4 Address: 10.10.100.45 User-Name: alex@pinglabz.local Status: Authorized Domain: DATA Oper host mode: multi-auth Oper control dir: both Session timeout: 3600s (server), Remaining: 3127s Common Session ID: 0A0A0A0100000123456789AB Acct Session ID: 0x0000018C Handle: 0x4D000001 Current Policy: POLICY_Gi1/0/10 Local Policies: Service Template: DEFAULT_LINKSEC_POLICY_SHOULD_SECURE Server Policies: Vlan Group: Vlan: 100 ACS ACL: xACSACLx-IP-CORP_DEFAULT_ACL-5d4f3a2b Method status list: Method State dot1x Authc Success ``` Two lines tell you everything. `Status: Authorized` means the port is open. `Method: dot1x, State: Authc Success` means 802.1X authentication (not MAB) succeeded. If you see `Method: mab, State: Authc Success`, the supplicant either could not speak 802.1X or was too slow, and MAB took over. ## Host modes: who can plug in The interface command `authentication host-mode ` controls how many devices the port allows after a single successful auth. single-host (default) What it allows One MAC per port. Any second MAC triggers a violation. When to use Lockdown deployments. Rare in modern networks. multi-host What it allows One authenticated MAC opens the port for everyone else (no auth required for the rest). When to use Avoid. Defeats most of the point of 802.1X. multi-domain What it allows One device in the DATA domain plus one in the VOICE domain. Both authenticate. When to use Standard for desk phones with PCs plugged through them. multi-auth What it allows Every MAC authenticates independently. Each can get its own VLAN/dACL. When to use Most enterprise deployments. The right default if you do not have a specific reason to choose another mode. ## Common failure modes When 802.1X breaks, it almost always breaks in one of these ways. The fix and the show command to confirm it are in the same row. Port stays in unauthorized state forever Likely cause Supplicant has 802.1X disabled (Windows: Wired AutoConfig service stopped) Confirm with `show authentication sessions interface X` shows no method attempted Authentication Failed in logs, supplicant gets reject Likely cause Bad credentials, wrong EAP method, certificate trust failure Confirm with `debug radius authentication` on switch; ISE Live Logs Port authorizes but device cannot reach DHCP Likely cause Wrong VLAN assigned, dACL blocking DHCP, or switch in monitor-mode Confirm with `show authentication sessions interface X details` for assigned VLAN; `show access-list` for dACL Auth works for one device, breaks when second device plugs in via the same port Likely cause Host-mode set to single-host or multi-domain when you needed multi-auth Confirm with `show running-config interface X` Re-authentication kills active sessions Likely cause Server returning a short session-timeout. Periodic re-auth is hitting an unreachable AAA server. Confirm with `show authentication sessions interface X details` for timeout value ## Open vs closed mode (monitor mode) Production rollouts almost never flip 802.1X to fully enforced on day one. The intermediate state is *open mode*, also called monitor mode. ``` interface GigabitEthernet1/0/10 authentication open ``` With `authentication open`, the port forwards traffic *regardless* of authentication status. The switch still runs the 802.1X exchange, still talks to RADIUS, still logs the result, but never blocks anyone. This lets you discover every device on the network, see which ones speak 802.1X, see which fall back to MAB, see which fail entirely, and fix the misbehaving 5% before flipping to closed mode. Most rollouts spend 30 to 90 days in monitor mode. ## Key takeaways 802.1X is conceptually simple. A switch port asks for credentials, forwards them to RADIUS, and enforces RADIUS's verdict. The complexity is in the configurations that surround the simple core: which host mode, what fallback for non-supplicant devices, what VLAN assignments, what dACLs, and how to stage the rollout without bricking your network. The minimum viable config above gets you talking to ISE. Everything beyond that is policy. For the full cluster, see the [802.1X pillar](https://www.pinglabz.com/802-1x/). ### OSPF vs BGP Redistribution: Which Protocol Wins Which Tie URL: https://www.pinglabz.com/ospf-vs-bgp-redistribution/ Last updated: 2026-07-04T23:21:40.000Z Mutual redistribution between OSPF and BGP is the one design discussion that splits a room of network engineers faster than almost any other. The mechanics are simple. The consequences of getting it wrong are not. This post walks through what redistribution actually does to a route, why OSPF and BGP make different tiebreaker decisions when they both know about the same prefix, and how to decide which protocol should own which decision. If you want the foundational reading first, start with the [OSPF complete guide](https://www.pinglabz.com/ospf/) and the [BGP complete guide](https://www.pinglabz.com/bgp/). This post assumes you already know what an LSA is and what an AS\_PATH is. ## What redistribution actually does When you type `redistribute ospf 1 subnets` under `router bgp 65001`, the router walks its RIB, finds every route whose source is OSPF process 1, and injects a copy into BGP's local routing table. The same prefix is now known to two protocols. Each protocol keeps its own copy with its own metric and its own next-hop selection logic. When the router decides which copy to install in the forwarding table, it uses **administrative distance** as the tiebreaker. This is where it starts to matter. ## The administrative distance tiebreaker Cisco IOS XE ships with the following defaults. Lower wins. Connected 0 Static 1 eBGP 20 EIGRP (internal) 90 OSPF (all types) 110 IS-IS 115 RIP 120 EIGRP (external) 170 iBGP 200 Read that table twice. eBGP at 20 beats OSPF at 110\. iBGP at 200 loses to OSPF. That single asymmetry is the source of most redistribution accidents you will ever see. ## The two redistribution directions There are two flows worth understanding separately because they fail in different ways. ### OSPF into BGP This is the common case. You have an internal OSPF domain, and you want some of those internal prefixes to be reachable from an external BGP peer. You have two options. `network` statement under router bgp What it does Advertises a single prefix that exactly matches a route in the local routing table (any source) When to use You know the exact prefixes you want to advertise. Safer. Recommended for production. `redistribute ospf 1 subnets` What it does Advertises every OSPF route the router knows. The `subnets` keyword is required or you only get classful summaries. When to use Lab work, or when the OSPF prefix list is large enough that maintaining a network-statement list is operationally painful. Always pair with a route-map filter. A minimal OSPF-into-BGP example with a safety filter: ``` ip prefix-list OSPF-TO-BGP-OK seq 10 permit 10.10.0.0/16 le 24 ! route-map OSPF-TO-BGP permit 10 match ip address prefix-list OSPF-TO-BGP-OK ! router bgp 65001 redistribute ospf 1 route-map OSPF-TO-BGP subnets ``` BGP picks up an OSPF-sourced route with origin code `?` (incomplete) and a MED equal to the OSPF cost at the redistribution point. That MED carryover is something you almost always want to override with a route-map `set metric`, because OSPF cost values are meaningful to OSPF, not to your eBGP neighbor. ### BGP into OSPF This is the dangerous one. The simple version looks like this: ``` router ospf 1 redistribute bgp 65001 subnets ``` You just told OSPF to take every BGP route the router knows and inject it as an OSPF external (Type 5 LSA, or Type 7 if the redistributing router sits in an NSSA). If you peer with a full BGP table, that is roughly 950,000 prefixes you just dumped into your OSPF link-state database. Your switches will reload. The rule is: **filter ruthlessly when you redistribute BGP into OSPF**. Use a route-map. Cap the count. Tag the routes so you can identify them later. A working pattern looks like this. ``` ip prefix-list BGP-TO-OSPF-OK seq 10 permit 192.0.2.0/24 ip prefix-list BGP-TO-OSPF-OK seq 20 permit 198.51.100.0/24 ! route-map BGP-TO-OSPF permit 10 match ip address prefix-list BGP-TO-OSPF-OK set tag 65001 ! router ospf 1 redistribute bgp 65001 subnets route-map BGP-TO-OSPF metric-type 1 ``` The `set tag` is not cosmetic. It is your loop-prevention mechanism on the return path. If the same prefix tries to come back in via another redistribution point, you match the tag in a deny clause and drop it. ## OSPF external types: E1 vs E2 When you inject a route into OSPF via redistribution, you choose its external type. The default is E2. Metric stays constant across the OSPF domain. Receivers see the same cost regardless of how far they are from the ASBR. TypeE2 (default) When to use You want all internal routers to treat the external destination as equally expensive. Most common. Metric = redistributed cost + internal OSPF cost to the ASBR. Closer ASBRs win. TypeE1 When to use You have multiple ASBRs advertising the same prefix and you want traffic to hit the closest one. If you redistribute BGP into OSPF at two routers without thinking about this, you get a 50/50 split based on hash, not based on proximity. Use E1 and let OSPF do the proximity math for you. ## The tiebreaker game in production You have a prefix, 10.50.0.0/24, that is reachable two ways: - Via OSPF, learned from an internal area, cost 100 - Via eBGP, learned from a transit peer eBGP wins. Admin distance 20 beats OSPF's 110\. The router installs the BGP-learned next-hop in the forwarding table. Internal traffic destined for 10.50.0.0/24 leaves the router via the eBGP next-hop, not via the OSPF path. If 10.50.0.0/24 is actually inside your own network, this is a routing loop in slow motion. The traffic exits to the transit provider, who has no idea where 10.50.0.0/24 is, and either drops it or sends it right back to you. The fix is to never accept your own internal prefixes from an external peer. Use a route-map on the eBGP inbound direction with an as-path filter or a prefix-list deny. ## Verifying redistribution worked The two commands worth memorizing: ``` show ip route 10.50.0.0 show ip bgp 10.50.0.0 ``` `show ip route` tells you which copy the RIB installed. The first line shows the source code (B for BGP, O for intra-area OSPF, O E2 for OSPF external type 2, etc.). If you redistributed BGP into OSPF and the route shows up as `O E2` on remote routers, the redistribution succeeded. `show ip bgp` tells you what the BGP table looks like. After OSPF-into-BGP redistribution, you should see your internal prefixes with origin code `?` and the local AS prepended. ## When to skip redistribution entirely Most modern designs prefer not to mutually redistribute at all. The alternatives are worth considering before you write a single redistribute command. - **Static routes pointed at a service IP.** If the only thing crossing the boundary is a handful of prefixes, static plus a network statement under BGP is simpler than redistribution and impossible to loop. - **BGP everywhere.** Enterprises with consistent multi-vendor infrastructure increasingly run iBGP as the internal IGP-equivalent and skip OSPF entirely for prefix distribution, using OSPF only for next-hop reachability. - **Route leaking with explicit filters.** EIGRP and BGP support route-leaking constructs (route-target import / export in VRFs) that give you per-prefix control without the redistribution blast radius. ## Key takeaways OSPF and BGP redistribute in both directions, but they fail in different ways. OSPF into BGP usually just leaks a few prefixes too many and is easy to filter. BGP into OSPF can take down your entire OSPF domain if you forget the filter. Always use a route-map. Always tag injected routes. Always set the metric explicitly rather than relying on the cross-protocol metric translation, because it never says what you think it says. For deeper coverage on either side, see the [OSPF pillar](https://www.pinglabz.com/ospf/) and the [BGP pillar](https://www.pinglabz.com/bgp/). Take the BGP reference with you The free BGP field-reference PDF: path attributes, best-path order, and the show commands that matter. Delivered by email, no card required. [Get the free PDF](https://www.pinglabz.com/bgp-cheatsheet/) ### Lab sec-10 - Wireless Security WPA2 vs WPA3 (Concept Lab) URL: https://www.pinglabz.com/ccna-lab-sec-10-wireless-security-wpa2-vs-wpa3/ Last updated: 2026-06-13T20:07:55.000Z _This post is for paying subscribers only._ ### Lab sec-09 - IPsec Site-to-Site Overview (Concept Lab) URL: https://www.pinglabz.com/ccna-lab-sec-09-ipsec-vpn-site-to-site-overview/ Last updated: 2026-06-13T20:07:55.000Z _This post is for paying subscribers only._ ### Lab sec-05 - DHCP Snooping and Dynamic ARP Inspection URL: https://www.pinglabz.com/ccna-lab-sec-05-dhcp-snooping-and-dai/ Last updated: 2026-08-02T03:05:54.000Z _This post is for paying subscribers only._ ### Lab sec-08 - 802.1X Port-Based Authentication (Switch Side) URL: https://www.pinglabz.com/ccna-lab-sec-08-802-1x-port-based-auth/ Last updated: 2026-08-02T03:05:55.000Z _This post is for paying subscribers only._ ### Lab sec-07 - AAA New-Model with Local Fallback URL: https://www.pinglabz.com/ccna-lab-sec-07-aaa-with-local-fallback/ Last updated: 2026-08-02T03:05:55.000Z _This post is for paying subscribers only._ ### Lab sec-06 - SSH and Disable Telnet URL: https://www.pinglabz.com/ccna-lab-sec-06-ssh-and-disable-telnet/ Last updated: 2026-08-02T03:05:54.000Z _This post is for paying subscribers only._ ### Lab ips-10 - QoS LLQ + CBWFQ on WAN Egress URL: https://www.pinglabz.com/ccna-lab-ips-10-qos-llq-cbwfq-wan-egress/ Last updated: 2026-08-02T03:05:54.000Z _This post is for paying subscribers only._ ### Lab sec-04 - Port Security and MAC Pinning URL: https://www.pinglabz.com/ccna-lab-sec-04-port-security-mac-pinning/ Last updated: 2026-08-02T03:05:56.000Z _This post is for paying subscribers only._ ### Lab sec-03 - Extended ACL (Named) URL: https://www.pinglabz.com/ccna-lab-sec-03-extended-acl-named/ Last updated: 2026-08-02T03:05:57.000Z _This post is for paying subscribers only._ ### Lab sec-01 - Line and Enable Passwords (Modern Best Practice) URL: https://www.pinglabz.com/ccna-lab-sec-01-line-and-enable-passwords/ Last updated: 2026-08-02T03:05:55.000Z _This post is for paying subscribers only._ ### Lab ips-06 - NTP Server and Client URL: https://www.pinglabz.com/ccna-lab-ips-06-ntp-server-client/ Last updated: 2026-08-02T03:05:56.000Z _This post is for paying subscribers only._ ### Lab ips-09 - QoS Classification and Marking URL: https://www.pinglabz.com/ccna-lab-ips-09-qos-classification-marking/ Last updated: 2026-08-02T03:05:58.000Z _This post is for paying subscribers only._ ### Lab ips-08 - SNMPv2c vs SNMPv3 URL: https://www.pinglabz.com/ccna-lab-ips-08-snmpv2c-vs-snmpv3/ Last updated: 2026-08-02T03:05:59.000Z _This post is for paying subscribers only._ ### Lab ips-07 - Syslog and Buffer Sizing URL: https://www.pinglabz.com/ccna-lab-ips-07-syslog-and-buffer-sizing/ Last updated: 2026-08-02T03:05:57.000Z _This post is for paying subscribers only._ ### Lab ips-03 - Static NAT Inside/Outside URL: https://www.pinglabz.com/ccna-lab-ips-03-static-nat-inside-outside/ Last updated: 2026-08-02T03:05:57.000Z _This post is for paying subscribers only._ ### Lab ips-02 - DHCP Relay (ip helper-address) URL: https://www.pinglabz.com/ccna-lab-ips-02-dhcp-relay-ip-helper-address/ Last updated: 2026-08-02T03:05:58.000Z _This post is for paying subscribers only._ ### Lab ips-05 - NAT Overload (PAT) URL: https://www.pinglabz.com/ccna-lab-ips-05-nat-overload-pat/ Last updated: 2026-08-02T03:06:00.000Z _This post is for paying subscribers only._ ### Lab ips-04 - Dynamic NAT with Pool URL: https://www.pinglabz.com/ccna-lab-ips-04-dynamic-nat-with-pool/ Last updated: 2026-08-02T03:06:00.000Z _This post is for paying subscribers only._ ### Lab ipc-13 - VRRP vs HSRP Comparison URL: https://www.pinglabz.com/ccna-lab-ipc-13-vrrp-vs-hsrp/ Last updated: 2026-08-02T03:05:59.000Z _This post is for paying subscribers only._ ### Lab ipc-12 - HSRP Active-Standby URL: https://www.pinglabz.com/ccna-lab-ipc-12-hsrp-active-standby/ Last updated: 2026-08-02T03:06:00.000Z _This post is for paying subscribers only._ ### Lab ips-01 - IOS XE as a DHCP Server URL: https://www.pinglabz.com/ccna-lab-ips-01-ios-xe-as-dhcp-server/ Last updated: 2026-08-02T03:06:01.000Z _This post is for paying subscribers only._ ### Lab ipc-14 - GLBP Active-Active Load Balancing URL: https://www.pinglabz.com/ccna-lab-ipc-14-glbp-load-balancing/ Last updated: 2026-08-02T03:06:01.000Z _This post is for paying subscribers only._ ### Lab ipc-09 - Static Route to a Remote Loopback (with Loopback Source) URL: https://www.pinglabz.com/ccna-lab-ipc-09-static-default-with-loopback-source/ Last updated: 2026-08-02T03:06:02.000Z _This post is for paying subscribers only._ ### Lab ipc-08 - OSPF MD5 / SHA Authentication URL: https://www.pinglabz.com/ccna-lab-ipc-08-ospf-md5-sha-authentication/ Last updated: 2026-08-02T03:06:01.000Z _This post is for paying subscribers only._ ### Lab ipc-11 - EIGRP Feasible Successor and DUAL URL: https://www.pinglabz.com/ccna-lab-ipc-11-eigrp-feasible-successor-dual/ Last updated: 2026-08-02T03:06:02.000Z _This post is for paying subscribers only._ ### Lab ipc-10 - EIGRP Named-Mode AS 100 URL: https://www.pinglabz.com/ccna-lab-ipc-10-eigrp-named-mode-classic/ Last updated: 2026-08-02T03:06:03.000Z _This post is for paying subscribers only._ ### Lab ipc-05 - OSPF Multi-Area (ABR + Type 3 LSAs) URL: https://www.pinglabz.com/ccna-lab-ipc-05-ospf-multi-area/ Last updated: 2026-08-02T03:06:02.000Z _This post is for paying subscribers only._ ### Lab ipc-03 - RIPv2 Essentials URL: https://www.pinglabz.com/ccna-lab-ipc-03-rip-v2-essentials/ Last updated: 2026-08-02T03:06:03.000Z _This post is for paying subscribers only._ ### Lab ipc-07 - OSPF DR/BDR Election URL: https://www.pinglabz.com/ccna-lab-ipc-07-ospf-dr-bdr-election/ Last updated: 2026-08-02T03:06:04.000Z _This post is for paying subscribers only._ ### Lab ipc-06 - OSPF Network Types URL: https://www.pinglabz.com/ccna-lab-ipc-06-ospf-network-types/ Last updated: 2026-08-02T03:06:04.000Z _This post is for paying subscribers only._ ### Lab na-14 - Wireless Architecture Overview (Concept Lab) URL: https://www.pinglabz.com/ccna-lab-na-14-wireless-architecture-overview/ Last updated: 2026-06-13T20:08:04.000Z _This post is for paying subscribers only._ ### Lab na-13 - SVI on L3 Switch (Inter-VLAN Routing) URL: https://www.pinglabz.com/ccna-lab-na-13-svi-on-l3-switch/ Last updated: 2026-08-02T03:06:04.000Z _This post is for paying subscribers only._ ### Lab ipc-02 - Floating Static Routes (AD Manipulation) URL: https://www.pinglabz.com/ccna-lab-ipc-02-floating-static-routes/ Last updated: 2026-08-02T03:06:06.000Z _This post is for paying subscribers only._ ### Lab ipc-01 - Default Static Route + Floating Backup URL: https://www.pinglabz.com/ccna-lab-ipc-01-default-static-route/ Last updated: 2026-08-02T03:06:05.000Z _This post is for paying subscribers only._ ### Lab na-10 - Root Guard URL: https://www.pinglabz.com/ccna-lab-na-10-rapid-pvst-root-guard/ Last updated: 2026-08-02T03:06:05.000Z _This post is for paying subscribers only._ ### Lab na-09 - PortFast and BPDU Guard URL: https://www.pinglabz.com/ccna-lab-na-09-rapid-pvst-portfast-bpduguard/ Last updated: 2026-08-02T03:06:06.000Z _This post is for paying subscribers only._ ### Lab na-12 - Router-on-a-Stick Inter-VLAN Routing URL: https://www.pinglabz.com/ccna-lab-na-12-router-on-stick/ Last updated: 2026-08-02T03:06:06.000Z _This post is for paying subscribers only._ ### Lab na-11 - CDP vs LLDP URL: https://www.pinglabz.com/ccna-lab-na-11-cdp-vs-lldp/ Last updated: 2026-08-02T03:06:07.000Z _This post is for paying subscribers only._ ### Lab na-05 - Voice VLAN on Access Ports URL: https://www.pinglabz.com/ccna-lab-na-05-voice-vlan-on-access-ports/ Last updated: 2026-08-02T03:06:07.000Z _This post is for paying subscribers only._ ### Lab na-08 - Configure Rapid-PVST URL: https://www.pinglabz.com/ccna-lab-na-08-configure-rapid-pvst/ Last updated: 2026-08-02T03:06:08.000Z _This post is for paying subscribers only._ ### Lab na-07 - PAgP vs LACP URL: https://www.pinglabz.com/ccna-lab-na-07-etherchannel-pagp-vs-lacp/ Last updated: 2026-08-02T03:06:08.000Z _This post is for paying subscribers only._ ### Lab na-06 - LACP Active EtherChannel URL: https://www.pinglabz.com/ccna-lab-na-06-etherchannel-lacp-active/ Last updated: 2026-08-02T03:06:08.000Z _This post is for paying subscribers only._ ### Lab nf-12 - OSI vs TCP/IP in the Cisco IOS CLI URL: https://www.pinglabz.com/ccna-lab-nf-12-osi-vs-tcpip-in-cli/ Last updated: 2026-06-13T20:08:07.000Z _This post is for paying subscribers only._ ### Lab na-04 - DTP and Static Trunking URL: https://www.pinglabz.com/ccna-lab-na-04-dtp-and-static-trunking/ Last updated: 2026-08-02T03:06:09.000Z _This post is for paying subscribers only._ ### Lab na-02 - VLANs and Access Ports URL: https://www.pinglabz.com/ccna-lab-na-02-vlans-and-access-ports/ Last updated: 2026-08-02T03:06:09.000Z _This post is for paying subscribers only._ ### Lab na-01 - Switching Fundamentals and the CAM Table URL: https://www.pinglabz.com/ccna-lab-na-01-switching-fundamentals-cam/ Last updated: 2026-08-02T03:06:10.000Z _This post is for paying subscribers only._ ### Lab nf-08 - Configure IPv6 Static Routes URL: https://www.pinglabz.com/ccna-lab-nf-08-configure-ipv6-static-routes/ Last updated: 2026-06-13T20:08:09.000Z _This post is for paying subscribers only._ ### Lab nf-11 - Troubleshooting Layer 1 / 2 / 3 Symptoms URL: https://www.pinglabz.com/ccna-lab-nf-11-troubleshooting-layer-symptoms/ Last updated: 2026-06-13T20:08:09.000Z _This post is for paying subscribers only._ ### Lab nf-10 - GRE Tunnel Between Two Routers URL: https://www.pinglabz.com/ccna-lab-nf-10-gre-tunnel-between-two-routers/ Last updated: 2026-06-13T20:08:09.000Z _This post is for paying subscribers only._ ### Lab nf-09 - Configure PPP and CHAP on a Serial Link URL: https://www.pinglabz.com/ccna-lab-nf-09-configure-ppp-and-chap/ Last updated: 2026-06-13T20:08:09.000Z _This post is for paying subscribers only._ ### Lab nf-03 - IPv4 Addressing Essentials URL: https://www.pinglabz.com/ccna-lab-nf-03-ipv4-addressing-essentials/ Last updated: 2026-06-13T20:08:10.000Z _This post is for paying subscribers only._ ### Lab nf-07 - Static Routes: Next-Hop vs Exit-Interface URL: https://www.pinglabz.com/ccna-lab-nf-07-static-routes-next-hop-vs-exit/ Last updated: 2026-06-13T20:08:10.000Z _This post is for paying subscribers only._ ### Lab nf-06 - Configure ARP and Static ARP URL: https://www.pinglabz.com/ccna-lab-nf-06-configure-arp-and-static-arp/ Last updated: 2026-06-13T20:08:10.000Z _This post is for paying subscribers only._ ### Lab nf-05 - IPv6 Addressing and EUI-64 URL: https://www.pinglabz.com/ccna-lab-nf-05-ipv6-addressing-and-eui64/ Last updated: 2026-08-01T19:35:48.000Z _This post is for paying subscribers only._ ### Lab nf-02 - IOS XE CLI Survival URL: https://www.pinglabz.com/ccna-lab-nf-02-ios-xe-cli-survival/ Last updated: 2026-06-13T20:08:11.000Z _This post is for paying subscribers only._ ### Lab sec-02 - Standard ACL (Numbered) URL: https://www.pinglabz.com/ccna-lab-sec-02-standard-acl-numbered/ Last updated: 2026-08-02T03:06:10.000Z Standard access lists filter traffic based on source IP only. Numbered standard ACLs use IDs 1-99 and 1300-1999\. They are the simplest filter on a Cisco router and a great introduction to the ACL concept: a list of permits and denies, evaluated top-down, with an implicit deny at the end. This lab configures a standard ACL on R1 that denies one specific host and permits the rest of the LAN. This is the fifth free preview lab in the library. ## What you will learn - Numbered vs named ACLs (we use numbered here; named is next lab) - Standard vs extended (standard = source-IP only) - How to apply an ACL inbound or outbound on an interface - The implicit deny at the end of every ACL - How to read `show ip access-lists` and the line numbers ## What this lab does NOT cover - Extended ACLs (next lab, [sec-03](https://www.pinglabz.com/ccna-lab-sec-03-extended-acl-named/)) - Time-based ACLs - IPv6 ACLs ## Topology Download the CCNA Base Topology .yaml 3 iol-xe routers + 1 alpine + 1 ioll2-xe managed switch + 1 unmanaged switch. [Download CCNA Base Topology](https://www.pinglabz.com/content/files/2026/08/pinglabz-ccna-base-topology-2.yaml) ## Step 1: Create a numbered standard ACL on R1 ``` R1#configure terminal R1(config)#access-list 10 deny 10.20.0.99 0.0.0.0 R1(config)#access-list 10 permit 10.20.0.0 0.0.0.255 R1(config)#access-list 10 deny any log ``` Three lines, evaluated top-down: - Line 10: deny host 10.20.0.99 (wildcard 0.0.0.0 = exactly this host) - Line 20: permit 10.20.0.0/24 (wildcard 0.0.0.255 = any host in the subnet) - Line 30: explicit deny + log (matches anything else, including 0.0.0.0/0 traffic) The explicit deny-log is a hardening pattern - it generates a log message when something gets blocked, useful for audit. Without it, the implicit deny at the end of every ACL silently drops traffic. ## Step 2: Apply the ACL inbound on Et0/0 ``` R1(config)#interface Ethernet0/0 R1(config-if)#ip access-group 10 in ``` The ACL is applied INBOUND on Ethernet0/0\. R1 evaluates the ACL against every packet ENTERING that interface. ## Step 3: Verify with show ip access-lists ``` R1#show ip access-lists Standard IP access list 10 10 deny 10.20.0.99 20 permit 10.20.0.0, wildcard bits 0.0.0.255 30 deny any log ``` IOS auto-numbers lines in increments of 10 (10, 20, 30). This lets you insert lines later without renumbering everything. ## Step 4: Insert a new ACE (access-control entry) To add a permit for 10.20.0.50 before the deny line: ``` R1(config)#ip access-list standard 10 R1(config-std-nacl)#15 permit 10.20.0.50 ``` Line 15 is inserted between line 10 and line 20. ``` R1#show ip access-lists 10 Standard IP access list 10 10 deny 10.20.0.99 15 permit 10.20.0.50 20 permit 10.20.0.0, wildcard bits 0.0.0.255 30 deny any log ``` ## Where standard ACLs should be applied Best practice for standard ACLs (filtering by source only): apply them CLOSE TO THE DESTINATION. The reason: a standard ACL has no idea where the packet is going. If you apply close to the source, you might block traffic that should reach some destinations. Extended ACLs (filtering by source AND destination AND protocol) are applied close to the SOURCE - they have all the info to make the right decision. ## Wildcard mask gotcha ACL wildcard masks are INVERSE of subnet masks. The bit positions tell you what to MATCH (0 = exact, 1 = don't care): 255.255.255.255 (/32) Wildcard mask0.0.0.0 What it matchesOne host exactly 255.255.255.0 (/24) Wildcard mask0.0.0.255 What it matchesWhole /24 subnet 255.255.0.0 (/16) Wildcard mask0.0.255.255 What it matchesWhole /16 n/a Wildcard mask0.0.0.255 What it matches Same as `any` within source bits Cisco shorthand: `host 10.20.0.99` is equivalent to `10.20.0.99 0.0.0.0`. `any` is equivalent to `0.0.0.0 255.255.255.255`. ## Verification - `show ip access-lists` shows the ACL with line numbers and counters - 10.20.0.99 cannot communicate through R1; other hosts in 10.20.0.0/24 can - The deny-log line generates syslog messages when triggered ## Troubleshooting matrix ACL doing the opposite of what you expect Likely cause Wildcard mask confusion (zeros match) Fix Re-check; remember `0.0.0.0` \= exact host Permit + deny + permit order wrong Likely cause ACLs are top-down with first-match Fix Reorder; specific permits before general denies Standard ACL not blocking traffic to a specific destination Likely cause Standard ACL has no destination match Fix Use extended ACL (sec-03) ACL applied but no counters Likely cause Wrong direction (in vs out) Fix Flip with `ip access-group 10 in` or `out` ## Engineer's note: production reality Standard ACLs are rare in modern production. Extended ACLs cover everything standard ACLs do and more, with finer-grained control. The remaining use cases for standard ACLs: route-map prefix-matching, redistribution filtering, distribute-lists in RIP/OSPF. Otherwise prefer extended. ## Related reading on PingLabz - [Lab sec-03: Extended ACL](https://www.pinglabz.com/ccna-lab-sec-03-extended-acl-named/) - [CCNA Labs: Security Fundamentals](https://www.pinglabz.com/ccna-labs-security-fundamentals/) ## Key takeaways - Standard ACLs match source IP only. Numbered 1-99 or 1300-1999. - Top-down first-match. Implicit deny at the end. - Apply standard ACLs close to the destination. - Wildcard masks are the INVERSE of subnet masks. - Use `any` and `host X` shortcuts for clarity. ## Up next [Lab sec-03: Extended ACL (named)](https://www.pinglabz.com/ccna-lab-sec-03-extended-acl-named/) ### Lab ipc-04 - OSPF Single-Area URL: https://www.pinglabz.com/ccna-lab-ipc-04-ospf-single-area/ Last updated: 2026-08-02T03:06:10.000Z OSPF (Open Shortest Path First) is the link-state protocol that powers most enterprise networks. Single-area OSPF puts every router in Area 0 - the backbone - and is the simplest production-grade dynamic routing protocol you can deploy. This lab configures OSPF single-area on R1, R2, R3 of the base topology, watches the DR/BDR election, and confirms end-to-end reachability across the loopbacks. This is the third of five free preview labs. ## What you will learn - The `router ospf` configuration syntax with router-id and network statements - OSPF Area 0 (backbone) - every multi-area network has one - How OSPF elects a Designated Router (DR) and Backup DR (BDR) on broadcast segments - How to read `show ip ospf neighbor` and the FULL state - How OSPF routes appear: `O` for intra-area ## What this lab does NOT cover - Multi-area OSPF (ABRs, areas other than 0) - that is the next lab, [ipc-05](https://www.pinglabz.com/ccna-lab-ipc-05-ospf-multi-area/) - OSPF authentication - covered in [ipc-08](https://www.pinglabz.com/ccna-lab-ipc-08-ospf-md5-sha-authentication/) - OSPF in non-broadcast network types - covered in [ipc-06](https://www.pinglabz.com/ccna-lab-ipc-06-ospf-network-types/) ## Topology Download the CCNA Base Topology .yaml The PingLabz CCNA Base Topology - 3 iol-xe routers + 1 alpine + 1 ioll2-xe switch. [Download CCNA Base Topology](https://www.pinglabz.com/content/files/2026/08/pinglabz-ccna-base-topology-2.yaml) ## Step 1: Configure OSPF on R1 ``` R1#configure terminal R1(config)#router ospf 100 R1(config-router)#router-id 10.255.0.1 R1(config-router)#network 10.20.0.0 0.0.0.255 area 0 R1(config-router)#network 10.255.0.1 0.0.0.0 area 0 R1(config-router)#passive-interface Loopback0 ``` Five commands: - `router ospf 100` \- process ID 100\. Locally significant; does not have to match between routers. - `router-id 10.255.0.1` \- explicit router-id. Pin it to the loopback for stability. - `network ... area 0` \- enable OSPF on matching interfaces and place them in Area 0\. The wildcard mask is the inverse of the subnet mask. - `passive-interface Loopback0` \- do not send Hellos out the loopback (no neighbors there). ## Step 2: Same pattern on R2 and R3 ``` R2(config)#router ospf 100 R2(config-router)#router-id 10.255.0.2 R2(config-router)#network 10.20.0.0 0.0.0.255 area 0 R2(config-router)#network 10.30.30.0 0.0.0.3 area 0 R2(config-router)#network 10.255.0.2 0.0.0.0 area 0 R2(config-router)#passive-interface Loopback0 R3(config)#router ospf 100 R3(config-router)#router-id 10.255.0.3 R3(config-router)#network 10.30.30.0 0.0.0.3 area 0 R3(config-router)#network 10.255.0.3 0.0.0.0 area 0 R3(config-router)#passive-interface Loopback0 ``` ## Step 3: Verify the neighbor state on R1 ``` R1#show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 10.255.0.2 1 FULL/DR 00:00:38 10.20.0.2 Ethernet0/0 ``` One neighbor - R2 (10.255.0.2) - in FULL state. R2 is the DR (Designated Router) on this broadcast segment. R1 lost the DR election because R2 has a higher IP-address tiebreaker on default priorities. R1 is therefore the BDR. ## Step 4: Look at OSPF on each interface ``` R1#show ip ospf interface brief Interface PID Area IP Address/Mask Cost State Nbrs F/C Lo0 100 0 10.255.0.1/32 1 LOOP 0/0 Et0/0 100 0 10.20.0.1/24 10 BDR 1/1 ``` Lo0 is in "LOOP" state - it advertises but does not form adjacencies (passive). Et0/0 is BDR with 1 neighbor that is FULL (the F/C column means Full / Configured). ## Step 5: Look at the routing table ``` R1#show ip route ospf Gateway of last resort is not set 10.0.0.0/8 is variably subnetted, 6 subnets, 3 masks O 10.30.30.0/30 [110/20] via 10.20.0.2, 00:00:10, Ethernet0/0 O 10.255.0.2/32 [110/11] via 10.20.0.2, 00:00:19, Ethernet0/0 O 10.255.0.3/32 [110/21] via 10.20.0.2, 00:00:05, Ethernet0/0 ``` O-coded routes (OSPF intra-area). `[110/N]` is AD/cost. R3's loopback cost is 21 (10 for the LAN + 10 for the P2P + 1 for the loopback). The path is R1 -> R2 -> R3. ## Step 6: End-to-end test ``` R1#ping 10.255.0.3 source Loopback0 repeat 3 Type escape sequence to abort. Sending 3, 100-byte ICMP Echos to 10.255.0.3, timeout is 2 seconds: Packet sent with a source address of 10.255.0.1 !!! Success rate is 100 percent (3/3), round-trip min/avg/max = 3/3/4 ms ``` ## Step 7: Look at the LSA database (briefly) ``` R1#show ip ospf database router self-originate | begin LS age LS age: 11 LS Type: Router Links Link State ID: 10.255.0.1 Advertising Router: 10.255.0.1 LS Seq Number: 80000005 Number of Links: 2 Link connected to: a Stub Network (Link ID) Network/subnet number: 10.255.0.1 (Link Data) Network Mask: 255.255.255.255 TOS 0 Metrics: 1 Link connected to: a Transit Network (Link ID) Designated Router address: 10.20.0.2 (Link Data) Router Interface address: 10.20.0.1 TOS 0 Metrics: 10 ``` R1's own Router LSA. Two links: the stub (loopback) and the transit (LAN segment via DR 10.20.0.2). ## Verification - `show ip ospf neighbor` shows R2 in FULL state as DR - `show ip ospf interface brief` shows R1 as BDR on Et0/0 - `show ip route ospf` shows three O-coded routes with \[110/cost\] pairs - Ping from R1 to R3's loopback returns 3/3 success ## Troubleshooting matrix Neighbor never goes FULL Likely cause Hello timer or area mismatch Confirm with `show ip ospf interface` on both ends Fix Match Hello/Dead timers and area Neighbor stuck in EXSTART Likely cause MTU mismatch on the segment Confirm with `show interfaces | include MTU` on both Fix Match MTU or use `ip ospf mtu-ignore` Router-id keeps changing on reload Likely cause No explicit router-id; OSPF picks highest active loopback Confirm with `show ip ospf | include Router ID` Fix Set `router-id` explicitly under `router ospf 100` Wrong router elected DR Likely cause Default priority 1 + IP tiebreaker Confirm with `show ip ospf interface | include DR|Priority` Fix Set `ip ospf priority 100` on the intended DR (covered in ipc-07) ## Engineer's note: production reality Single-area OSPF is the right design for small networks - sub-50 routers, all in one campus or data center. Beyond that scale, multi-area design becomes necessary to reduce LSA flooding scope and SPF recalculation cost. The dividing line moves over time as hardware gets faster, but the principle stays. Almost every modern enterprise routing deployment is OSPF (or sometimes IS-IS, which is similar conceptually). EIGRP exists in Cisco-only environments and is fine. BGP is for the WAN edge and the cloud-routing interface. Inside the LAN, OSPF dominates. ## Related reading on PingLabz - [OSPF: The Complete Guide](https://www.pinglabz.com/ospf/) - [Lab ipc-05: OSPF multi-area](https://www.pinglabz.com/ccna-lab-ipc-05-ospf-multi-area/) - [Lab ipc-07: OSPF DR/BDR election](https://www.pinglabz.com/ccna-lab-ipc-07-ospf-dr-bdr-election/) ## Key takeaways - Single-area OSPF: every router in Area 0. - Configuration: `router ospf N` \+ `router-id X` \+ `network ... area 0` \+ `passive-interface` on loopbacks. - DR/BDR election happens on broadcast segments. Default priority 1, IP-address tiebreaker. - OSPF routes are O-coded with cost in `[110/cost]`. - FULL state = adjacency complete; LSAs are flowing. ## Up next [Lab ipc-05: OSPF multi-area](https://www.pinglabz.com/ccna-lab-ipc-05-ospf-multi-area/) \- add a second area and an ABR. ### Lab na-03 - VLANs, Trunks, and VTP URL: https://www.pinglabz.com/ccna-lab-na-03-vlans-trunks-vtp/ Last updated: 2026-08-02T03:06:11.000Z VLANs work fine on a single switch. The minute you have two switches and want VLAN 10 on both, you need a trunk: a port that carries traffic for multiple VLANs and tags each frame with its VLAN ID so the receiving switch knows which broadcast domain it belongs to. This lab takes the three-switch triangle, looks at the existing 802.1Q trunks between them, and walks through VTP, the protocol Cisco invented to synchronize VLAN configurations across switches. This is the second of five free preview labs in the library, and the highest-leverage Pillar 2 search target. ## What you will learn - What an 802.1Q trunk is and how it differs from an access port - The role of the native VLAN and why moving it off VLAN 1 is a hardening best practice - How to configure a trunk on Cisco IOSvL2 switches - The output of `show interfaces trunk` and how to read all four columns - What VTP is, the three modes (server, client, transparent), and why transparent is the modern default ## What this lab does NOT cover - DTP (Dynamic Trunking Protocol) - that is the next lab, [na-04](https://www.pinglabz.com/ccna-lab-na-04-dtp-and-static-trunking/) - VLAN pruning beyond a quick mention - VTP version 3 in depth - we show version 1, the version on by default ## Topology Download the STP+VLAN Reference Lab .yaml Drop this into CML's Import dialog. Three IOSvL2 switches in a triangle with VLANs 10/20/99, dot1q trunks, rapid-PVST root election, and an LACP EtherChannel between SW1 and SW2. [Download STP+VLAN Reference Lab](https://www.pinglabz.com/content/files/2026/08/pinglabz-stp-vlan-reference-2.yaml) ## Step 1: Examine an existing trunk The three switches are already trunked. Look at SW1's view: ``` SW1#show interfaces trunk Port Mode Encapsulation Status Native vlan Gi0/0 on 802.1q trunking 99 Gi0/1 on 802.1q trunking 99 Po1 on 802.1q trunking 99 Port Vlans allowed on trunk Gi0/0 10,20,99 Gi0/1 10,20,99 Po1 10,20,99 Port Vlans allowed and active in management domain Gi0/0 10,20,99 Gi0/1 10,20,99 Po1 10,20,99 Port Vlans in spanning tree forwarding state and not pruned Gi0/0 10,20,99 Gi0/1 10,20,99 Po1 10,20 ``` This output has four sub-tables. Read them carefully: - **Top** \- Mode, Encapsulation, Status, Native VLAN. Native VLAN 99 (not VLAN 1 - that is the hardening choice). - **Vlans allowed on trunk** \- what you configured with `switchport trunk allowed vlan` - **Vlans allowed and active in management domain** \- VLANs from the allowed list that also exist on this switch - **Vlans in spanning tree forwarding state and not pruned** \- VLANs that are actually forwarding right now. The Po1 column shows VLAN 99 missing - STP is blocking VLAN 99 on Po1 because Gi0/0 is the preferred path for that VLAN. This is normal multipath behavior. ## Step 2: Configure a trunk from scratch To prove you can configure trunking yourself, take Gi0/2 (currently access mode) and convert it to a trunk: ``` SW1#configure terminal SW1(config)#interface GigabitEthernet0/2 SW1(config-if)#switchport trunk encapsulation dot1q SW1(config-if)#switchport mode trunk SW1(config-if)#switchport trunk native vlan 99 SW1(config-if)#switchport trunk allowed vlan 10,20,99 SW1(config-if)#end ``` Five commands - all important: 1. `switchport trunk encapsulation dot1q` \- explicitly 802.1Q (older switches negotiated ISL or dot1q; modern is always dot1q) 2. `switchport mode trunk` \- hardcode trunk mode, do not negotiate 3. `switchport trunk native vlan 99` \- hardening: never use the default VLAN 1 as native 4. `switchport trunk allowed vlan 10,20,99` \- explicit allow-list - only these VLANs cross the trunk ## Step 3: Native VLAN - the hardening point Frames that arrive on a trunk WITHOUT a VLAN tag get placed in the native VLAN. Default native is VLAN 1\. The hardening best practice is to: 1. Change the native VLAN to something other than VLAN 1 (we use VLAN 99 throughout PingLabz labs) 2. Never use the native VLAN as a data VLAN 3. Match the native VLAN on both ends of every trunk Why? Because a malicious host on the native VLAN could send untagged frames that the trunk treats as native, bypassing tag-based security. Native VLAN 99 with no hosts on it eliminates that attack surface. ## Step 4: VTP - the VLAN database synchronization protocol VTP (VLAN Trunking Protocol) lets a switch act as a "server" advertising its VLAN database to "clients" that copy it. Sounds useful. In practice it has caused so many production outages (a wiped VTP server can wipe every client's VLAN database) that modern best practice is to use **VTP transparent mode** on every switch, which means each switch maintains its own VLANs locally and ignores incoming VTP advertisements. ``` SW1#show vtp status VTP Version capable : 1 to 3 VTP version running : 1 VTP Domain Name : VTP Pruning Mode : Disabled VTP Traps Generation : Disabled Device ID : 5254.00a7.8000 Configuration last modified by 0.0.0.0 at 0-0-00 00:00:00 Feature VLAN: -------------- VTP Operating Mode : Transparent Maximum VLANs supported locally : 1005 Number of existing VLANs : 8 Configuration Revision : 0 MD5 digest : 0xC1 0xB6 0x8B 0x58 0x57 0x8A 0xBC 0xCB 0x47 0x47 0xFE 0x01 0x4D 0x6C 0x92 0x31 ``` Two lines tell you what mode we are in: - **VTP version running: 1** \- the lab is on VTP version 1 (the most common) - **VTP Operating Mode: Transparent** \- this switch ignores VTP advertisements from neighbors and only respects local VLAN configuration The Configuration Revision number is the dangerous part of VTP server/client mode. Higher revision wins. If you ever plug in a switch with a higher revision number and the same domain, it overwrites the VLAN database on the rest of the network. Transparent mode avoids this entirely. ## Step 5: VTP server/client mode (educational only) If you want to see VTP server/client in action, change SW1 to server and SW2 to client: ``` SW1(config)#vtp domain PINGLABZ SW1(config)#vtp mode server SW1(config)#vtp version 2 SW2(config)#vtp domain PINGLABZ SW2(config)#vtp mode client ``` Note that SW2 does NOT get a `vtp version 2` command. A VTP client cannot have its version set directly - if you try, IOS rejects it with "Cannot modify version in VTP client mode unless the system is in VTP version 3". The client inherits version 2 from the server's advertisements automatically; you will see "VTP version running : 2" on SW2 within one advertisement interval. (The Device ID above is the switch's base MAC and differs per import.) Then create a VLAN on SW1 - it propagates to SW2 automatically. ``` SW1(config)#vlan 40 SW1(config-vlan)#name TEMPORARY-DEMO ``` Wait 10 seconds, then on SW2: ``` SW2#show vlan brief | include 40 40 TEMPORARY-DEMO active ``` VLAN 40 appears on SW2 even though we never configured it there. That is VTP. **For production, never do this.** Set everyone to transparent mode and configure VLANs locally on each switch. ## Verification - `show interfaces trunk` on SW1 shows Gi0/0, Gi0/1, Po1 as 802.1Q trunks with native VLAN 99 - All trunks allow VLANs 10, 20, 99 - `show vtp status` shows VTP Operating Mode: Transparent - If you went through Step 5, VLAN 40 propagates from SW1 (server) to SW2 (client) - then revert ## Troubleshooting matrix Trunk shows "off" or "not-trunking" Likely cause `switchport mode trunk` missing on one end Confirm with `show interfaces switchport` on both ends Fix Hardcode `switchport mode trunk` on both "Native VLAN mismatch" CDP log Likely cause One end has native VLAN 99, the other has native VLAN 1 Confirm with Compare `show interfaces trunk` on both ends Fix Match the native VLAN on both ends VLAN exists on the switch but not in the trunk's "active in management domain" Likely cause VLAN not added to allowed list Confirm with `show interfaces trunk` "Vlans allowed on trunk" Fix `switchport trunk allowed vlan add N` VLAN database wiped after restart Likely cause VTP client copied an empty server's database Confirm with `show vtp status` on every switch Fix Set everyone to transparent; rebuild VLANs locally ## Engineer's note: production reality The single biggest cause of VLAN-related production outages is VTP. The protocol was designed for an era when networks were small and stable. Modern networks are large and dynamic; VTP server mode is a footgun. Every enterprise that has been bitten by a "rogue VTP server wiped our VLAN database" incident moves to transparent mode and stays there. Trunks themselves are stable and well-understood. The most common operational issue is "allowed VLAN list drift" - someone adds a new VLAN to one switch but forgets to add it to the trunk's allowed list, so traffic for that VLAN never crosses. Automation that derives the allowed-VLAN list from the VLAN-to-port mapping prevents this. ## Related reading on PingLabz - [VLANs and Layer 2 Switching: The Complete Guide](https://www.pinglabz.com/vlans-layer-2-switching/) - [Lab na-04: DTP and static trunking](https://www.pinglabz.com/ccna-lab-na-04-dtp-and-static-trunking/) - [Lab na-12: Router-on-a-stick (inter-VLAN routing over a trunk)](https://www.pinglabz.com/ccna-lab-na-12-router-on-stick/) ## Key takeaways - A trunk carries multiple VLANs and tags each frame with its VLAN ID (using 802.1Q). - Configure with `switchport mode trunk` \+ `switchport trunk allowed vlan N,M,...` \+ `switchport trunk native vlan X`. - Native VLAN should never be VLAN 1 in production. Match the native VLAN on both ends. - VTP synchronizes VLAN databases. Use **transparent mode** in production. Server/client mode is a known footgun. - `show interfaces trunk` has four sub-tables; the last (forwarding state) tells you what VLANs are actually crossing right now. ## Up next [Lab na-04: DTP and static trunking](https://www.pinglabz.com/ccna-lab-na-04-dtp-and-static-trunking/) looks at the protocol that NEGOTIATES trunking (and why you should turn it off). ### Lab nf-04 - IPv4 Subnetting with VLSM URL: https://www.pinglabz.com/ccna-lab-nf-04-ipv4-subnetting-vlsm/ Last updated: 2026-06-13T20:08:12.000Z Variable-Length Subnet Masking is the skill that turns a CCNA candidate into someone who can stand at a whiteboard with five rectangles and turn them into a working network. The math is not complex. The discipline is. This lab walks you through carving a single /16 parent block into four right-sized subnets, configuring them on a real router, and proving the result with `show ip route`. You will work on R1 from the [PingLabz CCNA Base Topology](https://www.pinglabz.com/lab-ip-scheme/), using loopback interfaces so you can run the entire exercise on one router without disturbing the rest of the lab. ## What you will learn - How to translate "I need a subnet for N hosts" into "use a /X mask" - How to carve a parent block (10.50.0.0/16 here) into multiple right-sized children without wasting addresses - How to configure those subnets on a Cisco router and verify them with `show ip route` - How to spot the canonical IOS clue that VLSM is happening ("variably subnetted, N subnets, M masks") - The two common mistakes that trip up engineers - overlap and unintended summarization ## What this lab does NOT cover - Route summarization for redistribution between protocols (covered in [IP Connectivity](https://www.pinglabz.com/ccna-labs-ip-connectivity/) labs) - IPv6 subnetting (covered in [nf-05](https://www.pinglabz.com/ccna-lab-nf-05-ipv6-addressing-and-eui64/)) - Private vs. public address policy (covered in [nf-03](https://www.pinglabz.com/ccna-lab-nf-03-ipv4-addressing-essentials/)) ## The scenario You have one parent block: **10.50.0.0/16**. That is 65,536 addresses to spend. You need to allocate four subnets: HQ-LAN Hosts needed500 Subnet size 512 (next power of 2 that fits) Required mask/23 (510 usable) Branch-LAN Hosts needed100 Subnet size128 Required mask/25 (126 usable) DMZ Hosts needed14 Subnet size16 Required mask/28 (14 usable) WAN-P2P Hosts needed2 (one each end) Subnet size4 Required mask/30 (2 usable) Total addresses needed: 512 + 128 + 16 + 4 = 660\. Well under the 65,536 the parent /16 gives you. The challenge is to allocate them efficiently and without overlap. ## Step 1: The hosts-to-mask rule For a subnet that needs N usable hosts, you need (N + 2) addresses minimum (network + broadcast are not usable). Round up to the next power of 2\. The mask is whatever number of host bits gives you that power of 2. 4 Host bits2 Usable hosts2 Mask/30 (255.255.255.252) 8 Host bits3 Usable hosts6 Mask/29 (255.255.255.248) 16 Host bits4 Usable hosts14 Mask/28 (255.255.255.240) 32 Host bits5 Usable hosts30 Mask/27 (255.255.255.224) 64 Host bits6 Usable hosts62 Mask/26 (255.255.255.192) 128 Host bits7 Usable hosts126 Mask/25 (255.255.255.128) 256 Host bits8 Usable hosts254 Mask/24 (255.255.255.0) 512 Host bits9 Usable hosts510 Mask/23 (255.255.254.0) 1024 Host bits10 Usable hosts1022 Mask/22 (255.255.252.0) 500 hosts? Need at least 502 addresses. Closest power of 2 that fits is 512\. That is 9 host bits, leaving 23 network bits. /23. 100 hosts? Need at least 102\. Closest power of 2 is 128\. That is 7 host bits, leaving 25 network bits. /25. ## Step 2: Allocate from the parent block (biggest first) The discipline that prevents overlap: **allocate biggest first, contiguous from the parent block.** If you start with the smallest and try to fit the biggest at the end, you waste space. Starting at 10.50.0.0: 1. **HQ-LAN (/23, 512 addresses).** Starts at 10.50.0.0\. Ends at 10.50.1.255\. Next free address: 10.50.2.0. 2. **Branch-LAN (/25, 128 addresses).** Starts at 10.50.2.0\. Ends at 10.50.2.127\. Next free address: 10.50.2.128. 3. **DMZ (/28, 16 addresses).** Starts at 10.50.2.128\. Ends at 10.50.2.143\. Next free address: 10.50.2.144. 4. **WAN-P2P (/30, 4 addresses).** Starts at 10.50.2.144\. Ends at 10.50.2.147\. Next free address: 10.50.2.148. Done. Plenty of /16 left over for future allocations. HQ-LAN CIDR10.50.0.0/23 Network10.50.0.0 First host10.50.0.1 Last host10.50.1.254 Broadcast10.50.1.255 Branch-LAN CIDR10.50.2.0/25 Network10.50.2.0 First host10.50.2.1 Last host10.50.2.126 Broadcast10.50.2.127 DMZ CIDR10.50.2.128/28 Network10.50.2.128 First host10.50.2.129 Last host10.50.2.142 Broadcast10.50.2.143 WAN-P2P CIDR10.50.2.144/30 Network10.50.2.144 First host10.50.2.145 Last host10.50.2.146 Broadcast10.50.2.147 ## Step 3: Configure the four subnets on R1 Console into R1 and configure four loopback interfaces, one per subnet, taking the first usable host address in each: ``` R1# configure terminal R1(config)# interface Loopback1 R1(config-if)# description HQ-LAN (needs 500 hosts -> /23 = 510 usable) R1(config-if)# ip address 10.50.0.1 255.255.254.0 R1(config-if)# no shutdown R1(config-if)# interface Loopback2 R1(config-if)# description Branch-LAN (needs 100 hosts -> /25 = 126 usable) R1(config-if)# ip address 10.50.2.1 255.255.255.128 R1(config-if)# no shutdown R1(config-if)# interface Loopback3 R1(config-if)# description DMZ (needs 14 hosts -> /28 = 14 usable) R1(config-if)# ip address 10.50.2.129 255.255.255.240 R1(config-if)# no shutdown R1(config-if)# interface Loopback4 R1(config-if)# description WAN-P2P (needs 2 hosts -> /30 = 2 usable) R1(config-if)# ip address 10.50.2.145 255.255.255.252 R1(config-if)# no shutdown R1(config-if)# end ``` Each loopback gets the FIRST USABLE host in its subnet. That is the convention - the router itself takes .1 (or the lowest available), and remaining addresses go to hosts. ## Step 4: Verify with show ip interface brief Real capture from the lab after running the config above: ``` R1# show ip interface brief Interface IP-Address OK? Method Status Protocol Ethernet0/0 10.20.0.1 YES TFTP up up Ethernet0/1 unassigned YES unset administratively down down Ethernet0/2 unassigned YES unset administratively down down Ethernet0/3 unassigned YES unset administratively down down Loopback0 10.255.0.1 YES TFTP up up Loopback1 10.50.0.1 YES manual up up Loopback2 10.50.2.1 YES manual up up Loopback3 10.50.2.129 YES manual up up Loopback4 10.50.2.145 YES manual up up ``` Five loopbacks total now - Loopback0 is the base topology's router-ID, Loopback1-4 are the VLSM allocations we just made. The Method column distinguishes them: `TFTP` for the configs loaded from the CML startup-config, `manual` for the changes you just typed. ## Step 5: The "variably subnetted" line is the proof This is the signature line that tells you VLSM is happening. Run `show ip route connected` on R1: ``` R1# show ip route connected <...routing protocol codes legend omitted...> Gateway of last resort is not set 10.0.0.0/8 is variably subnetted, 11 subnets, 6 masks C 10.20.0.0/24 is directly connected, Ethernet0/0 L 10.20.0.1/32 is directly connected, Ethernet0/0 C 10.50.0.0/23 is directly connected, Loopback1 L 10.50.0.1/32 is directly connected, Loopback1 C 10.50.2.0/25 is directly connected, Loopback2 L 10.50.2.1/32 is directly connected, Loopback2 C 10.50.2.128/28 is directly connected, Loopback3 L 10.50.2.129/32 is directly connected, Loopback3 C 10.50.2.144/30 is directly connected, Loopback4 L 10.50.2.145/32 is directly connected, Loopback4 C 10.255.0.1/32 is directly connected, Loopback0 ``` Read it carefully: - **"10.0.0.0/8 is variably subnetted, 11 subnets, 6 masks"**. This is the giveaway. IOS prints this line whenever a single parent network has subnets of different mask lengths inside it. The "11 subnets" counts every connected route under 10/8 (5 C entries plus 6 L entries because IOS also shows each interface's /32 local route). The "6 masks" counts the unique prefix lengths: /8 (the parent), /23 (HQ-LAN), /24 (LAN), /25 (Branch-LAN), /28 (DMZ), /30 (WAN-P2P), /32 (the local routes and the loopback). VLSM in action. - **Connected vs. local routes.** Each interface with an IP creates two routing-table entries: a `C` for the subnet ("everything in 10.50.0.0/23 is reachable via Loopback1") and an `L` for the interface address itself as a /32 ("10.50.0.1 specifically is me"). The L entries are why a /23 subnet shows up as two route entries. ## Step 6: Common mistakes Allocating smallest first What happens Big subnets do not fit on a boundary, you end up "wasting" the gap and reusing addresses How to detect The math fails: you allocate /28 at 10.50.0.0, then need /23 starting at 10.50.0.16 which is not a /23 boundary Two subnets that overlap What happens 10.50.2.0/25 and 10.50.2.128/25 do NOT overlap (good). But 10.50.0.0/23 and 10.50.1.0/24 DO overlap. The /24 is inside the /23. How to detect `show ip route` shows one of them as "longer match"; the broader one is masked by the more specific. Confusing routing decisions. Wrong mask on the router interface What happens You meant /25 but typed 255.255.255.0 (which is /24) How to detect `show ip interface brief` shows the address; `show ip interface eth-or-loopback` shows the /xx Asymmetric masks on a P2P link What happens R1 is /30 (255.255.255.252) but R2 is /29 (255.255.255.248). Reachable in one direction only. How to detect Ping R2 from R1 fails or returns asymmetric replies; mask check on both sides reveals the issue ## Verification - You can take "I need 50 hosts" and immediately reach for /26. - You allocate biggest first, contiguous, from the parent block. - `show ip interface brief` on R1 lists Loopback1-4 with the four /23, /25, /28, /30 addresses, all up. - `show ip route connected` shows "10.0.0.0/8 is variably subnetted, 11 subnets, 6 masks" followed by the C and L entries for every interface. ## Troubleshooting matrix "% Bad mask /29 for address 10.50.0.1" Likely cause The mask you typed does not align the address to a subnet boundary Confirm with The error message shows the address and mask Fix Check that the address is the network address or a usable host in the subnet that mask defines `show ip route` shows fewer subnets than you configured Likely cause One of your `ip address` commands was overwritten by a later one on the same interface Confirm with Re-check with `show running-config interface Loopback1` etc. Fix Reconfigure the interface with the right address and mask "variably subnetted" line is missing Likely cause All your subnets have the same mask, so it is not actually VLSM Confirm with `show ip route` uses a single-mask format Fix Not a problem unless you specifically intended different mask lengths Subnets overlap accidentally Likely cause Allocation math was wrong Confirm with Check that each subnet's address range does not intersect another's Fix Re-derive the allocation table; biggest first prevents this ## Engineer's note: production reality Real address plans live in spreadsheets, IPAM tools (Infoblox, BlueCat, NetBox), or YAML files in version control. You do not usually do VLSM math at a whiteboard - you do it once, capture it in IPAM, and the rest of the team consumes the plan. The skill the math teaches you is what to do when IPAM is wrong, when someone's documentation lies, or when you have to absorb a new acquisition's address space into your own. The math is fast once the discipline is muscle memory: hosts -> bits -> mask, biggest first, biggest first, biggest first. Modern best practice for new designs: use /16 or larger per site, leave room to grow, document hierarchically (region.site.purpose), and never let two subnets touch even when they could be carved into one. The address space is cheap; the cognitive overhead of overlap is expensive. ## Related reading on PingLabz - [Lab nf-03: IPv4 addressing essentials](https://www.pinglabz.com/ccna-lab-nf-03-ipv4-addressing-essentials/) \- the foundation this lab builds on - [Lab nf-07: Static routes - next-hop vs exit-interface](https://www.pinglabz.com/ccna-lab-nf-07-static-routes-next-hop-vs-exit/) \- use these subnets in routing decisions - [CCNA Labs: IP Connectivity](https://www.pinglabz.com/ccna-labs-ip-connectivity/) \- where the routing happens ## Key takeaways - For N hosts, find the smallest power of 2 that is greater than or equal to (N + 2). That tells you the host bits. Subtract from 32 to get the prefix length. - Allocate biggest first, contiguous from the parent block. This is the single discipline that prevents waste and overlap. - `show ip route` prints "variably subnetted, N subnets, M masks" whenever you have VLSM in a single parent network. That line is the proof. - Each interface creates two routing-table entries: a `C` (the subnet) and an `L` (the /32 for the interface itself). - The two common mistakes are smallest-first allocation and overlapping subnets. Biggest-first contiguous allocation prevents both. ## Up next [Lab nf-05: IPv6 addressing and EUI-64](https://www.pinglabz.com/ccna-lab-nf-05-ipv6-addressing-and-eui64/) takes the same address-and-mask logic into the 128-bit world. Same math, more bits, slightly different conventions. ### Lab nf-01 - Cisco Modeling Labs Free Quick Start URL: https://www.pinglabz.com/ccna-lab-nf-01-cml-quick-start/ Last updated: 2026-08-02T03:06:11.000Z This is the first lab in the [PingLabz CCNA Labs library](https://www.pinglabz.com/ccna-labs-network-fundamentals/). There is no networking in it. The goal is to get Cisco Modeling Labs Free installed on your computer, import the PingLabz CCNA Base Topology, boot it, and log in to a router and a switch. Once you have that working, every other lab in the library starts the same way: download the .yaml, import, boot, type along. The whole thing should take about 30 minutes the first time. Every subsequent lab in the library will take you under 90 seconds to bring up. ## What you will learn - What Cisco Modeling Labs Free is and how it differs from the paid Personal and Enterprise tiers - How to download, install, and license CML Free on your computer - How to import a PingLabz lab .yaml file into CML - How to boot the lab and open a console to a router or switch - The first set of CLI commands every CCNA lab assumes you can run on both a router and a managed switch ## What this lab does NOT cover - Networking. There is no routing, no addressing, no protocol configuration here. We are setting up the tooling. - Building topologies from scratch. You will do that in later labs. Here you just import ours. - The CML Personal or Enterprise editions. Their license tiers and node caps are different. CML Free is what we target across the entire library. ## System requirements CML Free is shipped as an OVA file that runs as a virtual appliance on VMware Workstation, VMware Fusion, VMware ESXi, or a Linux KVM host. You install it once and it lives on your machine until you uninstall it. The VM itself runs a Linux controller that orchestrates the network nodes. CPU Minimum 4 cores with VT-x or AMD-V Recommended for the labs library6+ cores RAM Minimum8 GB allocated to CML Recommended for the labs library16 GB allocated to CML Disk Minimum20 GB free Recommended for the labs library 50 GB free (room for multiple labs) Hypervisor Minimum VMware Workstation Pro / Fusion / ESXi, or Linux KVM Recommended for the labs librarySame Browser Minimum Modern Chrome, Firefox, Safari, or Edge Recommended for the labs librarySame CML Free is capped at 5 nodes per running lab. Unmanaged switches do not count toward the 5\. Every PingLabz CCNA lab is sized to fit inside that limit. ## Step 1: Create a Cisco account and download CML Free 1. Go to [the CML Free landing page](https://mkto.cisco.com/cml-free.html?ref=pinglabz.com). 2. If you do not already have a Cisco account, create one. It is free. 3. Sign in and click the download link for the current CML Free release. You will get an OVA file (around 5 GB) and a small refplat.iso file for the node images. 4. Save both files somewhere you can find them again. You will not need to download them a second time. ## Step 2: Import the OVA into your hypervisor Open VMware Workstation (or Fusion on Mac, or ESXi if you are running a homelab server) and import the OVA. The defaults are sensible. The one setting you should adjust is the RAM allocation - bump it from 8 GB to 16 GB if you have it, and the controller will boot faster and run more nodes comfortably. Once the VM is created, attach the refplat.iso as a CD/DVD to the VM. CML reads its node images off this disc the first time it boots. Without it, you will boot to a working controller but with no usable node types. Power the VM on. The first boot takes a few minutes while the controller initializes. When you see a login prompt on the VM console, the controller is ready. ## Step 3: First login to CML The CML console will print the controller's IP address once it has finished booting. Open that URL in your browser: ``` https:/// ``` You will get a self-signed-certificate warning. Click through it. The login page comes up. The default username is `admin` and the password is whatever you set during OVA import (Cisco's default suggestion is `1234QWer!`). You should land on the CML Workbench. It is empty - no labs yet. ## Step 4: Download the PingLabz CCNA Base Topology .yaml This is the reusable six-node topology that powers most labs in the library: three routers, one host, one managed switch (so you can run real CCNA switch commands), and one unmanaged switch as a spare L2 broadcast domain. Fully configured with the canonical [PingLabz IP scheme](https://www.pinglabz.com/lab-ip-scheme/) so labs reinforce each other across the library. iol-xe (router) NodeR1 Role LAN gateway and primary loopback router iol-xe (router) NodeR2 Role Transit router (LAN side and P2P to R3) iol-xe (router) NodeR3 Role Remote router across the point-to-point link ioll2-xe (managed L2 switch) NodeSW1 Role LAN broadcast domain plus a management SVI at 10.20.0.10 alpine (Linux host) NodeHOST1 Role LAN client (assigned an IP per-lab) unmanaged switch NodeSW2 Role Spare LAN broadcast domain for dual-LAN labs (does not count toward the CML Free 5-node cap) Download the CCNA Base Topology .yaml Drop this into CML's Import dialog. The reader needs no networking configuration to bring it up - the routers and SW1 come pre-configured with hostnames, IP addressing, and a vty user. [Download pinglabz-ccna-base-topology.yaml](https://www.pinglabz.com/content/files/2026/08/pinglabz-ccna-base-topology-2.yaml) ## Step 5: Import the topology into CML 1. In the CML Workbench, click the menu icon at the top left and pick **Import Lab**. 2. Drop the .yaml file you just downloaded into the import dialog (or click and browse to it). 3. Give the lab whatever title you like (the default "PingLabz CCNA Base Topology" is fine). Click **Import**. 4. The Workbench opens with six nodes on the canvas: R1, R2, R3, SW1, HOST1, SW2\. The routers, SW1, and host are unstarted (gray). SW2 is pre-started (it is unmanaged, no boot needed). ## Step 6: Start the lab Click the play icon at the top of the Workbench to start every node. The router and SW1 nodes show a "BOOTING" state for about 60 seconds while IOS XE comes up. Once the icons turn green, the lab is running. ## Step 7: Open a console to R1 1. Click the R1 node icon on the canvas. 2. In the right-hand panel, click **Console**. A serial console opens in a new tab. 3. You will see the boot scroll, then a PingLabz MOTD banner, then a username prompt. 4. Log in with the PingLabz canonical credentials: **username** `pinglabz`, **password** `PingLabz!23`. 5. Because `pinglabz` is a privilege-15 user, you land directly in privileged-exec mode (the prompt ends in `#`) - no `enable` needed. *Heads up:* the .yaml also includes `admin` and `sysadmin` users with the password `Cisco@123`. Those are for PyATS automation tooling - you can ignore them for hands-on lab work. `pinglabz / PingLabz!23` is the user you log in as. ## Step 8: First commands on R1 These are the show commands you will run at the start of every lab in this library. Run them now and confirm the output matches what you see below. (Your ping times and timestamps will obviously differ.) ``` R1#show ip interface brief Interface IP-Address OK? Method Status Protocol Ethernet0/0 10.20.0.1 YES TFTP up up Ethernet0/1 unassigned YES unset administratively down down Ethernet0/2 unassigned YES unset administratively down down Ethernet0/3 unassigned YES unset administratively down down Loopback0 10.255.0.1 YES TFTP up up R1#ping 10.20.0.2 repeat 3 Type escape sequence to abort. Sending 3, 100-byte ICMP Echos to 10.20.0.2, timeout is 2 seconds: !!! Success rate is 100 percent (3/3), round-trip min/avg/max = 2/3/4 ms R1#ping 10.20.0.10 repeat 3 Type escape sequence to abort. Sending 3, 100-byte ICMP Echos to 10.20.0.10, timeout is 2 seconds: !!! Success rate is 100 percent (3/3), round-trip min/avg/max = 1/2/3 ms R1#show version | include Cisco IOS Cisco IOS Software [IOSXE], Linux Software (X86_64BI_LINUX-ADVENTERPRISEK9-M), Version 17.18.2, RELEASE SOFTWARE (fc3) ``` The ping to R2 (10.20.0.2) works because R1 and R2 are on the same LAN via SW1\. The ping to 10.20.0.10 reaches SW1's management SVI - confirming the switch is L3-reachable. No routing protocol is needed for any of this. This is the connectivity baseline every other lab builds on top of. The `Method` column shows `TFTP` for the configured interfaces - that is how CML pushes the startup config to the node at boot. It is normal and only shows up in CML, not on physical hardware. ## Step 9: First commands on SW1 (the managed switch) Open a console to SW1 the same way you did R1\. Log in with the same credentials (`pinglabz` / `PingLabz!23`). Then run the L2-side equivalents of the verification commands you just ran on R1: ``` SW1#show ip interface brief Interface IP-Address OK? Method Status Protocol Ethernet0/0 unassigned YES unset up up Ethernet0/1 unassigned YES unset up up Ethernet0/2 unassigned YES unset up up Ethernet0/3 unassigned YES unset up up Vlan1 10.20.0.10 YES TFTP up up SW1#show vlan brief VLAN Name Status Ports ---- -------------------------------- --------- ------------------------------- 1 default active Et0/0, Et0/1, Et0/2, Et0/3 1002 fddi-default act/unsup 1003 token-ring-default act/unsup 1004 fddinet-default act/unsup 1005 trnet-default act/unsup SW1#show mac address-table Mac Address Table ------------------------------------------- Vlan Mac Address Type Ports ---- ----------- -------- ----- 1 5254.0064.48d9 DYNAMIC Et0/2 1 aabb.cc00.1800 DYNAMIC Et0/1 1 aabb.cc00.1a00 DYNAMIC Et0/0 Total Mac Addresses for this criterion: 3 SW1#show spanning-tree vlan 1 VLAN0001 Spanning tree enabled protocol rstp Root ID Priority 32769 Address aabb.cc00.0500 This bridge is the root Hello Time 2 sec Max Age 20 sec Forward Delay 15 sec Bridge ID Priority 32769 (priority 32768 sys-id-ext 1) Address aabb.cc00.0500 Hello Time 2 sec Max Age 20 sec Forward Delay 15 sec Aging Time 300 sec Interface Role Sts Cost Prio.Nbr Type ------------------- ---- --- --------- -------- -------------------------------- Et0/0 Desg FWD 100 128.1 P2p Edge Et0/1 Desg FWD 100 128.2 P2p Edge Et0/2 Desg FWD 100 128.3 P2p Edge Et0/3 Desg FWD 100 128.4 P2p Edge SW1#show interfaces status Port Name Status Vlan Duplex Speed Type Et0/0 LAN access port connected 1 full auto 10/100/1000BaseTX Et0/1 LAN access port connected 1 full auto 10/100/1000BaseTX Et0/2 LAN access port connected 1 full auto 10/100/1000BaseTX Et0/3 LAN access port connected 1 full auto 10/100/1000BaseTX ``` One heads-up before you compare: the exact MAC addresses in this output vary with every CML import. iol-xe MAC addresses (the `aabb.cc00.xxxx` values) are assigned dynamically by the CML server, and HOST1's KVM-style `5254.00xx.xxxx` MAC is randomized too. Match the structure and the port mappings, not the digits. This is what a managed switch actually exposes. `show vlan brief` tells you which ports belong to which VLAN - all four Ethernet ports sit in VLAN 1 because they are access mode by default. `show mac address-table` shows the L2 forwarding table: SW1 has learned the MAC addresses of R1 (Et0/0, prefix `aabb.cc00` is the iol-xe convention), R2 (Et0/1), and HOST1 (Et0/2, prefix `5254.00` is the KVM-derived alpine convention). `show spanning-tree` confirms rapid-pvst is running and SW1 has elected itself the root bridge with priority 32769 - all ports are designated forwarding in P2p Edge mode (the result of `spanning-tree portfast` on the access ports). Every Pillar 2 lab in this library leans on these commands. You will see them a lot. ## Verification: what success looks like If you can read this output yourself on your R1 and SW1 consoles, you are done with the quick-start: - R1 boots without errors - SW1 boots without errors - You can log in to both with `pinglabz / PingLabz!23` and land directly at the `#` prompt - On R1, `show ip interface brief` shows Loopback0 at 10.255.0.1 and Ethernet0/0 at 10.20.0.1, both "up / up" - On R1, `ping 10.20.0.2` and `ping 10.20.0.10` both return 3/3 successful - On SW1, `show vlan brief` lists VLAN 1 with all four Ethernet ports - On SW1, `show mac address-table` shows at least the R1 and R2 MAC addresses That is the green light for the rest of the library. ## Troubleshooting matrix OVA import says "no node images available" Likely cause refplat.iso not mounted to the CML VM Confirm with VM settings -> CD/DVD Fix Mount the refplat.iso, restart the CML VM CML boots but cannot start nodes Likely cause Not enough RAM allocated to the VM Confirm with CML Workbench -> system status Fix Power off CML VM, raise RAM to 16 GB, power on R1 boots but the console shows garbled text Likely cause Wrong terminal type, common with some browsers Confirm with Open the console in another browser Fix Try Chrome or Firefox; clear browser cache Login prompt rejects pinglabz / PingLabz!23 Likely cause You are typing the password in console-disabled state, or caps-lock is on Confirm with Watch for "% Login invalid" message Fix Verify caps-lock; the password is case-sensitive ping 10.20.0.2 fails with all dots Likely cause R2 has not finished booting yet Confirm with Check R2 console for the login prompt Fix Wait 30 more seconds, retry the ping SW1 MAC address table is empty Likely cause R1 and R2 have not generated any traffic yet Confirm with Run a ping between them, then re-check Fix Ping 10.20.0.2 from R1, then "show mac address-table" on SW1 First ping packet of every test shows as "." Likely cause ARP resolution for the destination - the first packet is dropped while ARP completes Confirm with Subsequent pings show all "!" Fix Expected behavior; not a fault. Run the ping a second time for a clean 3/3. Lab shows "License: 6 nodes" warning when starting Likely cause You added an extra node by accident Confirm with Look at the node count in the Workbench header Fix Remove the extra node; CML Free is hard-capped at 5 counted nodes (unmanaged switches do not count) ## Beyond CML Free CML Free is the perfect starting point for these labs, but it is not the only tier Cisco offers. CML Personal lifts the cap to 20 nodes for a one-time fee; CML Enterprise scales further and is licensed per-instance. Most engineers running these labs never need anything beyond Free - all PingLabz CCNA labs are sized to fit. If you find yourself drawn to topology designs with more than 5 routers, the Personal tier is worth considering. One real-world note: the IOS XE in iol-xe and ioll2-xe runs as a userspace process inside the CML controller, not as a full virtual machine. That keeps boot time fast (60 seconds vs several minutes for a full VM image), and lets you run multiple devices comfortably inside a single laptop's worth of RAM. The behaviour is identical to full virtual IOS XE images for the operations that matter to a CCNA lab. The one caveat: ioll2-xe (the switch) does not expose 802.1X exec commands like `show authentication sessions`. Our 802.1X lab uses a dedicated reference topology with a different image to work around this. ## Related reading on PingLabz - [The PingLabz lab IP scheme](https://www.pinglabz.com/lab-ip-scheme/) \- the canonical addressing used across every lab - [CCNA Labs: Network Fundamentals](https://www.pinglabz.com/ccna-labs-network-fundamentals/) \- the cluster this lab opens - [Lab nf-04: IPv4 subnetting with VLSM](https://www.pinglabz.com/ccna-lab-nf-04-ipv4-subnetting-vlsm/) \- the next free lab, takes 45 minutes ## Key takeaways - CML Free is capped at 5 counted nodes per lab. Unmanaged switches do not count. Every PingLabz CCNA lab fits this limit. - The whole library uses a downloadable .yaml as the lab's starting point. Import once per lab, run. - The PingLabz IP scheme is consistent across the library so labs reinforce each other. - R1 and SW1 share the same documented login: `pinglabz / PingLabz!23`. Privilege-15 user lands straight at the `#` prompt. - SW1 is a managed L2 switch with a working CLI - you can `show vlan`, `show mac address-table`, `show spanning-tree`, configure trunks, run port security, all the real CCNA L2 commands. ## Up next Now that the tooling works, the rest of the library is networking. The natural next lab is [Lab nf-04: IPv4 subnetting with VLSM](https://www.pinglabz.com/ccna-lab-nf-04-ipv4-subnetting-vlsm/) \- which uses the same topology you just imported. ### Cisco ASA Field Reference - Free 9-Page Cheat Sheet URL: https://www.pinglabz.com/cisco-asa-cheatsheet/ Last updated: 2026-08-02T03:06:11.000Z _This post is for subscribers only._ ### Migration Considerations from ASA to Secure Firewall (FTD) URL: https://www.pinglabz.com/cisco-asa-to-ftd-migration/ Last updated: 2026-06-13T20:08:13.000Z If you are managing Cisco ASA hardware in 2026, you are managing inherited gear. Cisco's strategic firewall is now Secure Firewall Threat Defense (FTD, formerly known as Firepower Threat Defense). The platform is not just a UI rebrand: FTD runs a different management model (FMC or FDM), a different config plane (object-driven, GUI-first), and a different feature posture (next-gen IPS, URL filtering, SSL decryption native). This article covers the practical question every team running ASA hits eventually: when do we migrate, what tooling does the migration well, what does the migration not move automatically, and what do we verify on the FTD side before declaring the migration done. The answer length here is 1,000-1,500 words because there is no live lab walkthrough; the procedural steps are the value. This is the Migration article in the [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/) reading order. ## When to Migrate (and When to Stay) You need next-gen IPS, URL filtering, AMP, or SSL decryption native in the firewall. You have a stable ASA with all-static-policy that does exactly what it needs to do. Your hardware is end-of-sale (ASA 5500-X) and the FTD-capable replacement is the natural successor. You have heavy AnyConnect or custom MPF investment that does not migrate cleanly. Your security org has standardized on FMC for centralized management and policy. You are running multi-context (which has a counterpart in FTD called Multi-Instance, but the migration is non-trivial). You are buying new firewalls anyway and want to land on Cisco's strategic platform. You operate a small fleet that is a long way from the next refresh and the migration cost outweighs the upside. Most enterprises migrate when they refresh hardware. That gives you fresh boxes that are FTD-capable from day one and removes the awkward "convert in place" path. If your existing ASA hardware is FTD-capable (most ASA 5500-X models are), in-place reimaging is supported but is operationally messy: it is a wipe-and-reload, the device is offline during the install, and the new image needs FMC registration before policy can be pushed. Hard to do without a maintenance window and a proven backup. ## The Cisco Firewall Migration Tool Cisco publishes a free tool called the **Cisco Secure Firewall Migration Tool** (formerly the Firepower Migration Tool) that takes an ASA running-config and converts it into an FTD policy plus an upload to FMC. The tool runs as a Windows or macOS application; the workflow is approximately: 1. Export the ASA running-config (`more system:running-config` output, saved to a text file). 2. Open the migration tool, select "Cisco ASA to Threat Defense." 3. Upload the running-config and the show-tech output (some platform info comes from show-tech, not running-config). 4. The tool parses ACLs, NAT rules, network and service objects, interface IPs, static routes, NAT pools, and AAA references. It surfaces a per-object review screen. 5. You triage: accept clean translations, manually fix or skip items the tool flags as "needs review." 6. The tool can either generate a policy file you push to FMC manually, or push directly to a target FMC (and then to a target FTD device). 7. You bring up the FTD on the same network position, fail traffic over, and verify. The tool's parser has been refined over many releases and is reliable for straightforward ASA configs. Where it struggles is custom MPF, AnyConnect profiles, multi-context, and any feature that has no direct FTD equivalent. ## What Migrates Cleanly Network and service objects, object-groups FTD equivalent Network and Port objects, Object Groups Migration status Clean. The naming model is similar enough that the tool maps these directly. Extended ACLs (interface-bound) FTD equivalent FTD Access Control Policy rules Migration status Mostly clean. Each ACL line becomes one ACP rule; ordering preserved. Auto NAT and Manual NAT FTD equivalent FTD NAT rules (with the same Section 1/2/3 model) Migration status Mostly clean. The tool preserves order and section. Interface IPs, security levels, subinterfaces FTD equivalent FTD Interface configuration Migration status Clean for routed mode. Transparent has caveats. Static routes, default route, AD, tracked routes FTD equivalent FTD Routing -> Static Route Migration statusClean. Site-to-site IKEv2 VPN with PSK FTD equivalent FTD Site-to-Site VPN policy Migration status Mostly clean for IKEv2 + PSK. IKEv1 rarely. AAA server-group definitions (RADIUS/LDAP/TACACS) FTD equivalent FTD AAA / Identity / Realm Migration status Definitions move; the references on tunnel-groups need rework. Default inspection (DNS, FTP, ICMP) FTD equivalentFTD default inspection Migration status Clean - FTD has equivalent built-ins. Logging destinations (host, severity) FTD equivalent FTD Platform Settings -> Syslog Migration statusClean. NTP, SNMPv3, banners FTD equivalentFTD Platform Settings Migration statusClean. For the 80%-case ASA (perimeter firewall doing routed-mode NAT, ACL, a couple of S2S tunnels, syslog, NTP, AAA), the migration tool produces a usable FTD config that needs only minor cleanup. ## What Does Not Migrate (Or Migrates Painfully) AnyConnect / Remote-Access VPN profiles FTD equivalent FTD RA VPN with AnyConnect profile XML Migration friction Heavy. Profile XMLs do not auto-migrate. Plan to rebuild. Dynamic Access Policies (DAP) FTD equivalent FTD has identity / posture but the DAP record model is different Migration friction Manual rebuild against the new policy model. Custom MPF inspection (HTTP, ESMTP custom maps) FTD equivalent FTD inspection / file policy / SSL policy Migration friction Conceptual rebuild. The capabilities are deeper but the config is different. Multi-context FTD equivalent FTD Multi-Instance (different model) Migration friction The Cisco Migration Tool has limited support; large multi-context migrations are usually professional services engagements. Transparent mode FTD equivalent FTD transparent / inline mode Migration friction Conceptually maps but requires manual setup. Custom EtherType ACLs FTD equivalent Not all EtherType ACL features are present in FTD Migration frictionReview case by case. Cluster (chassis cluster) FTD equivalent FTD HA + cluster supported on Firepower 4100/9300 Migration friction Re-architecture rather than migration. Platform-specific knobs (jumbo frames, custom TCP timeout overrides) FTD equivalent Some are exposed in FlexConfig only Migration friction FlexConfig is FTD's escape hatch for ASA-only commands; expect to use it. The single biggest migration pain is AnyConnect. The ASA's tunnel-group + group-policy + DAP + AAA + cert-trustpoint model is intricate, and FTD reimplements the same outcomes with a different config shape. Plan for a side-by-side build of the new RA VPN on the FTD, a pilot user group, and a phased cutover rather than a flag-day switch. ## Post-Migration Verification Checklist After the FTD goes live in the same network position, verify: 1. **Interface IPs and zones.** Every interface in FMC matches the ASA-side configuration. Security zones (FTD's equivalent of named interfaces) are assigned correctly. 2. **Routing.** Default route present, static routes intact, ECMP working if you had multi-WAN. 3. **NAT rules.** Run packet-tracer (FTD has a similar tool) for representative outbound and inbound flows. Verify the same source IP gets PATed to the same outside IP, and that the DMZ static NAT publishes the correct public IP. 4. **ACL hit counters.** Generate test traffic for each high-value rule. Watch the FTD ACP rule hit counter climb. 5. **S2S tunnels.** Each VPN peer should re-establish. `show crypto ipsec sa` equivalent in FTD. Encaps/decaps counters climbing in both directions. 6. **Logging.** FMC's Connection Events table fills as traffic flows. Confirm syslog still flows to your SIEM. 7. **NTP.** Both FMC and FTD synced; clock skew breaks cert validation. 8. **HA pair.** If you migrated a failover pair, verify FTD HA peering, sync state, and fail-back behavior. 9. **Performance.** Connection table size and CPU under load. FTD with all features on uses more CPU than ASA software with the same throughput; size accordingly. 10. **Rollback path.** Keep the ASA available as a hot spare for at least one to two weeks. Document the rollback procedure (restore config, re-IP, redirect upstream) before you need it. ## Hidden Costs to Plan For Migrations always cost more than the engineering scope predicts. Three line items that consistently get under-budgeted: - **FMC is its own platform.** If you do not already operate Firepower Management Center, you need an FMC instance (virtual or hardware), licensing, integration into your monitoring, and an admin learning curve. - **License rework.** ASA Smart Licensing is different from FTD Smart Licensing. Bulk migration generally needs a license-conversion exercise with your Cisco partner. - **Operational learning.** Your NOC and SecOps teams need to learn FMC's policy model, which is more object-and-rule-driven than the ASA CLI they have been writing for years. Budget training, not just deployment time. ## Key Takeaways The Cisco Secure Firewall Migration Tool covers the 80% case (objects, ACLs, NAT, interfaces, static routes, basic S2S VPN) cleanly. The remaining 20% (AnyConnect profiles, custom MPF inspection, multi-context, DAP, cluster) is where the real migration time goes. Plan for a side-by-side build of those features on FTD, not a one-click migration. Most teams migrate during a hardware refresh because that gives the cleanest cutover - new FTD-capable boxes, a parallel build, a planned switchover. In-place reimage of FTD-capable ASA hardware works but requires a real maintenance window and a proven backup. Either path, plan for FMC, license rework, and a learning curve as fixed costs that are independent of the technical migration. For the rest of the cluster, the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/) has the full reading order, including the comparison article [ASA vs FTD vs Firepower](https://www.pinglabz.com/asa-vs-ftd-firepower/) for a more direct feature-by-feature look at the two platforms. ### Cisco ASA Multiple Context Mode URL: https://www.pinglabz.com/cisco-asa-multiple-context-mode/ Last updated: 2026-06-13T20:08:13.000Z Multiple context mode turns a single Cisco ASA into a hypervisor for several virtual firewalls. Each context has its own configuration, interfaces, ACLs, NAT rules, routing table, and even its own administrators. The classic use cases are service providers (one customer per context), enterprises with strict network segmentation (separate contexts per business unit or per regulatory boundary), and data centers running tenant isolation in hardware. This article covers what multi-context mode is, what it costs, how the system / admin / user-context split works, and the specific configuration you write to enable it. The lab examples are reference-only because switching the live PingLabz ASA Reference Lab to multi-context would destroy the entire single-context running config we built across Sessions 1-3. This is a Fundamentals article in the [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/). If you are deciding whether multi-context fits your environment, the punchline is that it usually does not - most enterprises use separate physical firewalls or virtual ASAv instances on a hypervisor instead. Read on for when multi-context is still the right answer. ## What Multi-Context Mode Is (and Is Not) Multiple independent firewalls on one chassis Multiple Cisco software images running side-by-side Each context with its own admin, config, interfaces Each context with its own kernel or platform Routed or transparent per-context (your choice) Mixed firewall-mode in transparent contexts that share an interface Resource quotas via class definitions Hard isolation between contexts at the hypervisor level Storage of per-context config in flash or off-box Separate licensing per context (ASA license is per-chassis) The mental model: imagine the ASA's data plane as one switching fabric, divided into N virtual firewalls, each with its own forwarding table and policy plane. The hardware is shared; the configuration and the policy are isolated. ## The Three Context Types Every multi-context ASA has three kinds of contexts: The execution space that boots the box, defines all other contexts, allocates interfaces and resources. Has no traffic-forwarding policy of its own. Type**System** How manyExactly 1 (built-in) One user-context designated as the management entry point. Logging in to the box on the management interface lands you in the admin context. Type**Admin** How many Exactly 1 (you designate) The actual virtual firewalls. Each has its own config, interfaces, policy. Type**User context** How many 1 to 250 depending on platform license The admin context is also a user context for forwarding purposes - it can have data interfaces, ACLs, NAT, the works - but it is the one you SSH into to manage the box. From the admin context you run `changeto context NAME` to drop into another user context's CLI for ad-hoc work. ## What You Give Up Multi-context mode does not support every feature single-context does. The list of unsupported features changes per release; check the configuration guide for your target version, but as of recent ASA software: - Dynamic routing protocols are supported but are licensed per-context (some early releases had it as system-only). - VPN was historically not supported in multi-context; recent releases support site-to-site VPN per context, but remote-access VPN (AnyConnect) is more limited. - Threat detection is limited. - QoS is configured per-context but the underlying queues are platform-wide. - Multicast routing is not supported (only stub multicast forwarding). The biggest practical loss for a perimeter firewall is the historic VPN restriction. Confirm against the release notes before committing to multi-context for a VPN-heavy use case. ## Enabling Multi-Context Mode (Destructive) The single command that enables it: ``` ciscoasa# changeto context system ! (in case you are not in system context) ciscoasa(config)# mode multiple WARNING: This command will change the behavior of the device WARNING: This command will initiate a Reboot Proceed with change mode? [confirm] Convert the system configuration? [confirm] The old running configuration file will be written to flash The admin context configuration will be written to flash The new running configuration file was written to flash Security context mode: multiple *** --- SHUTDOWN NOW --- ... [reload] ... ``` Two important notes from that warning. First, the box reloads. Second, the existing single-context config is converted: a slimmed-down system config (defining basic chassis-level stuff) and an admin context that inherits most of the rest. The conversion is best-effort and rarely produces a working multi-context layout straight away; expect to rework the config after the conversion. The reverse (`mode single`) is also destructive: every user context is dropped, and the box reloads to single-context mode. Treat the mode change as a one-way commitment for production purposes. ## In the System Context: Define Resources and Contexts The system context is where you define interfaces, allocate them to contexts, and create the contexts themselves. ``` ! From the system context: ASA-MULTI(config)# admin-context ADMIN-CTX ASA-MULTI(config)# context ADMIN-CTX ASA-MULTI(config-ctx)# config-url disk0:/admin-ctx.cfg ASA-MULTI(config-ctx)# allocate-interface GigabitEthernet0/0 ASA-MULTI(config-ctx)# allocate-interface Management0/0 ! ASA-MULTI(config)# context CUSTOMER-A ASA-MULTI(config-ctx)# description Tenant A perimeter firewall ASA-MULTI(config-ctx)# config-url disk0:/customer-a.cfg ASA-MULTI(config-ctx)# allocate-interface GigabitEthernet0/1 ASA-MULTI(config-ctx)# allocate-interface GigabitEthernet0/2 ! ASA-MULTI(config)# context CUSTOMER-B ASA-MULTI(config-ctx)# description Tenant B perimeter firewall ASA-MULTI(config-ctx)# config-url disk0:/customer-b.cfg ASA-MULTI(config-ctx)# allocate-interface GigabitEthernet0/3 ASA-MULTI(config-ctx)# allocate-interface GigabitEthernet0/4 ``` Walkthrough: `admin-context ADMIN-CTX` picks which named context plays the admin role. `context NAME` creates a new context. `config-url` points at the on-flash file that holds that context's config (or a remote URL like ftp://). `allocate-interface` hands a physical interface (or a subinterface, or a port-channel) to the context. From this point, the context can use that interface as if it were the only firewall on the box. Subinterface allocation is also supported, which is the high-density model: ``` ASA-MULTI(config)# context CUSTOMER-A ASA-MULTI(config-ctx)# allocate-interface GigabitEthernet0/0.10 ASA-MULTI(config-ctx)# allocate-interface GigabitEthernet0/0.20 ``` Customer A gets two VLAN-tagged subinterfaces on the shared physical Gi0/0; customer B gets two different VLANs on the same physical interface. The chassis runs one trunk to a common upstream switch and the multi-context model isolates each customer at Layer 3. Shared interfaces (one physical interface allocated to multiple contexts at the same time) are also supported in routed mode but require careful packet-classifier configuration. Stick with non-shared interface allocation unless you have a specific reason and have read the platform-specific multi-context docs. ## Inside a User Context Once the context exists and has interfaces allocated, drop into it and configure as if it were a normal ASA: ``` ASA-MULTI# changeto context CUSTOMER-A ASA-MULTI/CUSTOMER-A# configure terminal ASA-MULTI/CUSTOMER-A(config)# hostname CUSTOMER-A CUSTOMER-A(config)# interface GigabitEthernet0/1 CUSTOMER-A(config-if)# nameif outside CUSTOMER-A(config-if)# security-level 0 CUSTOMER-A(config-if)# ip address 198.51.100.10 255.255.255.0 CUSTOMER-A(config-if)# no shutdown CUSTOMER-A(config)# interface GigabitEthernet0/2 CUSTOMER-A(config-if)# nameif inside CUSTOMER-A(config-if)# security-level 100 CUSTOMER-A(config-if)# ip address 10.30.0.1 255.255.255.0 CUSTOMER-A(config-if)# no shutdown CUSTOMER-A(config)# route outside 0.0.0.0 0.0.0.0 198.51.100.1 1 CUSTOMER-A(config)# write memory ``` Within a user context, every command works the way it does in a single-context ASA. NAT, ACL, VPN (where supported), inspection - everything. The prompt prefix `ASA-MULTI/CUSTOMER-A#` reminds you which context you are in. Return to the system context with `changeto system`. ## Resource Classes Without limits, one rogue context can exhaust the box's conn table or NAT pool. Define resource classes from the system context and assign each user context to one: ``` ! From system: ASA-MULTI(config)# class GOLD-CLASS ASA-MULTI(config-class)# limit-resource conns 200000 ASA-MULTI(config-class)# limit-resource xlates 100000 ASA-MULTI(config-class)# limit-resource hosts 50000 ASA-MULTI(config-class)# limit-resource asdm 10 ASA-MULTI(config-class)# limit-resource ssh 10 ! ASA-MULTI(config)# class SILVER-CLASS ASA-MULTI(config-class)# limit-resource conns 50000 ASA-MULTI(config-class)# limit-resource xlates 25000 ASA-MULTI(config-class)# limit-resource hosts 10000 ! ASA-MULTI(config)# context CUSTOMER-A ASA-MULTI(config-ctx)# member GOLD-CLASS ASA-MULTI(config)# context CUSTOMER-B ASA-MULTI(config-ctx)# member SILVER-CLASS ``` Each context's traffic is now capped at its class's limits. `show resource usage` in the system context tells you per-context utilization vs cap. ## Reference Show Output From a Cisco-published multi-context reference: ``` ASA-MULTI# show context Context Name Class Interfaces Mode URL *ADMIN-CTX default GigabitEthernet0/0, Routed disk0:/admin-ctx.cfg Management0/0 CUSTOMER-A GOLD GigabitEthernet0/1, Routed disk0:/customer-a.cfg GigabitEthernet0/2 CUSTOMER-B SILVER GigabitEthernet0/3, Routed disk0:/customer-b.cfg GigabitEthernet0/4 Total active Security Contexts: 3 ``` The asterisk on ADMIN-CTX marks the admin context. Total active contexts is 3 (one admin + two customer); platform license caps would limit this to 2, 5, 20, 50, 100, 250 etc. depending on the chassis tier. ``` ASA-MULTI# show resource usage context CUSTOMER-A CPU usage 5 sec, 1 min, 5 min: 12.5%, 13.1%, 13.0% Used Total %Use Limit Class Conns 42184 n/a 21% 200000 GOLD Xlates 18430 n/a 18% 100000 GOLD Hosts 4221 n/a 8% 50000 GOLD SSH 2 n/a 20% 10 GOLD ASDM 1 n/a 10% 10 GOLD ``` Customer A is using 21% of its conn cap and 18% of its xlate cap. Steady-state utilization in healthy ranges. If "Used" approached the limit, customer A traffic would start getting denied at the resource layer (separate from any policy denial), and you would either bump the class or move customer A to GOLD-PLUS. ## Failover Across Contexts Failover in multi-context mode is configured at the system context and applies across all contexts. The two units must run the same number and shape of contexts; configurations replicate from active to standby per-context. The failover commands themselves (`failover lan unit primary`, etc.) are system-context-only. One mode-specific subtlety: active/active failover (where the active role is split per-context, half the contexts run on unit A and the other half on unit B) is only supported in multi-context mode. It is one of the few hard reasons to choose multi-context: when you want to use both physical units actively rather than the active/standby pattern documented in [active/standby failover](https://www.pinglabz.com/cisco-asa-active-standby-failover/). ## When Multi-Context Fits, and When It Does Not Fits when: - You are a service provider running tenant isolation in shared hardware. - You have hard regulatory boundaries (PCI, government tenant separation) and the audit framework wants firewall isolation in addition to logical separation. - You want active/active failover (each unit actively forwards for some contexts while standing by for others). - You have a high-density data center with limited rack space and many small firewall responsibilities to consolidate. Does not fit when: - You need full RA VPN per business unit (historic limit; check the release notes). - Your fleet is hybrid with FTD: FTD has a similar feature called Multi-Instance, but it is not the same as ASA multi-context, and migrating between the two is non-trivial. - You can solve the same problem with separate physical or virtual ASA instances and the operational complexity of multi-context is hard to justify. For most enterprises, separate ASAv virtual instances on a hypervisor (or separate hardware ASAs in remote sites) win on operational simplicity even at the cost of more boxes. ## Key Takeaways Multi-context mode partitions a single Cisco ASA into multiple virtual firewalls, each with its own configuration, interfaces, and admin scope. The system context defines the contexts and allocates resources; the admin context is the management entry point; user contexts run the actual firewall policy. Switching between single and multi-context is destructive (the box reloads and the existing config is converted on a best-effort basis), so plan a maintenance window and a config rebuild rather than treating it as a config tweak. Resource classes are mandatory in any production multi-context deployment to prevent one context from starving the others. Active/active failover is one of the few features unique to multi-context. Most enterprises end up using ASAv virtual instances or separate physical units rather than multi-context; multi-context is the right answer for service providers and shared-hardware tenant isolation but rarely for a typical enterprise edge. Next: [Migration Considerations from ASA to Secure Firewall (FTD)](https://www.pinglabz.com/cisco-asa-to-ftd-migration/) covers how to move from ASA software to FTD and what does and does not migrate cleanly. The full reading order is on the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). ### Cisco ASA Inspection Engines: MPF and Custom Maps URL: https://www.pinglabz.com/cisco-asa-inspection-engines/ Last updated: 2026-06-13T20:08:14.000Z The Cisco ASA Modular Policy Framework (MPF) is the configuration model that decides which traffic gets which kind of treatment as it crosses the firewall. It is what binds inspection engines to specific traffic, what applies QoS or connection limits, and what sets timeouts that diverge from the platform defaults. The pieces are three: **class-maps** (matching traffic), **policy-maps** (the actions to apply), and **service-policies** (binding the policy-map to an interface or globally). This article walks the model, the default global policy that ships on every ASA, and how to add custom HTTP, FTP, or ESMTP inspection without breaking what is already working. This is a Fundamentals article in the [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/). The companion article on the data-plane side is [Cisco ASA Packet Flow](https://www.pinglabz.com/cisco-asa-packet-flow/), which covers where the inspection engines actually fire in the per-packet pipeline. ## The Three MPF Pieces `class-map NAME` What it does Matches traffic. ACL-style criteria, or built-in `match default-inspection-traffic`. Example `class-map WEB-CLASS` \+ `match access-list WEB-ACL` `policy-map NAME` What it does Lists actions to apply to matched traffic. One or more class entries, each with one or more action lines. Example `policy-map WEB-POLICY` \+ `class WEB-CLASS` \+ `inspect http` `service-policy NAME {global | interface NAME}` What it does Activates the policy-map. Either globally (one policy across all interfaces) or per-interface. Example `service-policy WEB-POLICY interface outside` One policy-map can have many class entries, and each class can have many actions. Activating the policy-map binds the whole thing at once. ## The Default Global Policy Every fresh ASA ships with a global service-policy already in place. It catches traffic with a built-in class and applies inspection for the standard application-layer protocols. ``` ASA-PERIM# show running-config policy-map global_policy policy-map global_policy class inspection_default inspect dns preset_dns_map inspect ftp inspect h323 h225 inspect h323 ras inspect ip-options inspect netbios inspect rsh inspect rtsp inspect skinny inspect esmtp inspect sqlnet inspect sunrpc inspect tftp inspect sip inspect xdmcp class class-default user-statistics accounting ! ASA-PERIM# show running-config service-policy service-policy global_policy global ``` Read it as: a class called `inspection_default` matches a built-in set of "ports we know about for the listed protocols" (DNS on 53, FTP on 21, etc.). The actions are `inspect ` for each application-layer engine. The service-policy line binds it globally. The ASA does not include `inspect icmp` by default on every release. The lab's running config has it explicitly added (which is why pings come back). Confirm: ``` ASA-PERIM# show service-policy global | include icmp Inspect: icmp, packet 1244, lock fail 0, drop 0, reset-drop 0 ``` If the line is missing, add it: ``` ASA-PERIM(config)# policy-map global_policy ASA-PERIM(config-pmap)# class inspection_default ASA-PERIM(config-pmap-c)# inspect icmp ``` Without `inspect icmp`, ICMP echo requests leave the firewall but echo replies have no matching conn entry and get dropped. ICMP is the easiest "is the ASA letting traffic through" test, so getting inspection on for it early is worth doing. ## What an Inspection Engine Actually Does Three things, depending on the protocol: 1. **Stateful tracking for connectionless protocols.** ICMP is the canonical example: the ASA pretends ICMP is stateful so echo replies match echo requests. 2. **Pinhole opening for protocols with negotiated secondary ports.** FTP active mode negotiates a data port over the control channel. Without inspection, the data channel hits the perimeter ACL and gets denied; with inspection, the ASA reads the PORT command, opens a temporary pinhole for the negotiated port, and closes it when the session ends. 3. **Application-layer security checks.** The HTTP inspection engine can enforce HTTP method allow-lists, header sanity, URL length limits, and reject malformed requests. The ESMTP engine masks server banners, blocks dangerous SMTP verbs, and enforces RFC compliance on email transactions. The first category is essential and almost always wanted. The second is required for any application that uses negotiated secondary ports. The third is application-layer firewalling that overlaps with what the next-gen successor (FTD) does in a more flexible model. ASA's application-layer inspection still works but is rarely the central piece of a modern security strategy. ## Custom Class-Maps: Match More Specifically Sometimes the built-in `inspection_default` class is too broad and you want to apply an action to a narrower slice of traffic. Three matching styles cover almost everything: ``` ! Match by ACL (the most common) ASA-PERIM(config)# access-list WEB-ACL extended permit tcp any object DMZ-WEB eq https ASA-PERIM(config)# class-map WEB-CLASS ASA-PERIM(config-cmap)# match access-list WEB-ACL ! ! Match by port (lighter than an ACL) ASA-PERIM(config)# class-map SSH-CLASS ASA-PERIM(config-cmap)# match port tcp eq 22 ! ! Match by destination IP only ASA-PERIM(config)# class-map DMZ-CLASS ASA-PERIM(config-cmap)# match destination-address dmz 192.168.50.0 255.255.255.0 ``` One class-map can have one match line. For more complex matching (multiple conditions ANDed or ORed), build the matching into an ACL and use `match access-list`. ## Worked Example: Custom HTTP Inspection for the DMZ Apply HTTP inspection only to traffic destined for the DMZ web server, with custom restrictions (only GET/POST allowed, no PUT/DELETE, max URI length 1024 bytes). ``` ! 1. Define what custom HTTP inspection should do ASA-PERIM(config)# policy-map type inspect http DMZ-HTTP-MAP ASA-PERIM(config-pmap)# parameters ASA-PERIM(config-pmap-p)# protocol-violation action drop-connection log ASA-PERIM(config-pmap)# match request method PUT ASA-PERIM(config-pmap-c)# drop-connection log ASA-PERIM(config-pmap)# match request method DELETE ASA-PERIM(config-pmap-c)# drop-connection log ASA-PERIM(config-pmap)# match request uri length gt 1024 ASA-PERIM(config-pmap-c)# drop-connection log ! ! 2. Define the class-map that matches DMZ-bound web traffic ASA-PERIM(config)# access-list DMZ-WEB-ACL extended permit tcp any object DMZ-WEB eq http ASA-PERIM(config)# access-list DMZ-WEB-ACL extended permit tcp any object DMZ-WEB eq https ASA-PERIM(config)# class-map DMZ-WEB-CLASS ASA-PERIM(config-cmap)# match access-list DMZ-WEB-ACL ! ! 3. Build the policy-map with the inspection action referencing the inspect map ASA-PERIM(config)# policy-map DMZ-POLICY ASA-PERIM(config-pmap)# class DMZ-WEB-CLASS ASA-PERIM(config-pmap-c)# inspect http DMZ-HTTP-MAP ! ! 4. Bind it to the outside interface ASA-PERIM(config)# service-policy DMZ-POLICY interface outside ``` The new service-policy on the outside interface coexists with the global one. Per-interface policies evaluate before global; if a packet matches an interface policy class, the interface policy wins and the global is skipped for that packet. If the packet does not match any class on the interface policy, the global policy is consulted next. Verify the new policy is in place: ``` ASA-PERIM# show service-policy interface outside Interface outside: Service-policy: DMZ-POLICY Class-map: DMZ-WEB-CLASS Inspect: http DMZ-HTTP-MAP, packet 87, lock fail 0, drop 0, reset-drop 0 ``` The packet count climbing means the inspection engine is firing. The drop and reset-drop counters tell you whether any real packets are being rejected by the application-layer rules. ## Connection Limits and Timeouts MPF also drives per-flow connection limits and timeouts. Useful for limiting a high-fan-out source from exhausting the conn table, or for shortening the idle timeout on protocols you know are bursty. ``` ASA-PERIM(config)# policy-map global_policy ASA-PERIM(config-pmap)# class inspection_default ASA-PERIM(config-pmap-c)# set connection conn-max 100000 ASA-PERIM(config-pmap-c)# set connection embryonic-conn-max 1000 ASA-PERIM(config-pmap-c)# set connection per-client-max 5000 ASA-PERIM(config-pmap-c)# set connection timeout idle 1:00:00 ASA-PERIM(config-pmap-c)# set connection timeout half-closed 0:10:00 ``` Walk through: - `conn-max 100000`: cap the total simultaneous conns matching this class. Useful for keeping a single application from hogging the table. - `embryonic-conn-max 1000`: cap on TCP half-open SYNs. The ASA's classic SYN-flood mitigation; below this the ASA proxies the SYN/ACK to validate the source. - `per-client-max 5000`: cap per source IP. Stops a single noisy client from hogging. - `timeout idle`: how long an idle conn lives before the ASA tears it down. Default is 1 hour; tune up for long-lived idle protocols. - `timeout half-closed`: TCP FIN/FIN-WAIT timeout. Default is 10 minutes; tune down to reclaim conns faster after one side has gone away. ## QoS via MPF The ASA's QoS support is narrower than IOS - it does priority queuing and policing on selected traffic, not full diffserv-style queuing. ``` ASA-PERIM(config)# class-map VOICE-CLASS ASA-PERIM(config-cmap)# match dscp ef ! ASA-PERIM(config)# policy-map QOS-POLICY ASA-PERIM(config-pmap)# class VOICE-CLASS ASA-PERIM(config-pmap-c)# priority ! ASA-PERIM(config)# service-policy QOS-POLICY interface outside ``` Voice traffic (DSCP EF) gets the priority queue on the outside interface. Most ASA deployments do not need this because the upstream router or WAN edge handles QoS; the ASA's job is forwarding and security, not shaping. ## Verification Commands `show service-policy global` What classes are on the global policy and their per-class hit/drop counters. `show service-policy interface NAME` The interface-bound service-policy and its counters. `show running-config policy-map` All policy-map definitions in the running config. `show running-config class-map` All class-map definitions. `show running-config service-policy` The bindings (which policy-map is bound where). `show conn protocol tcp` Active TCP conns; useful to see what inspection has tracked. `show inspect http` Per-engine internal counters and stats (replace `http` with the engine you care about). ## Common Traps **Removing the global policy without replacing it.** If you do `no service-policy global_policy global` without binding a replacement, you lose all default inspection. ICMP stops working, FTP active stops working, your DNS bonus features evaporate. If you are migrating to per-interface policies, build the new ones first and remove the global last - or keep the global and let per-interface policies layer on top. **Conflicting per-interface and global classes.** Per-interface wins over global, even for the same class name. If you build an interface policy that covers `inspection_default`, you have just disabled the global engines for traffic on that interface. Be explicit about which interface policies cover which classes. **HTTP inspection breaking valid traffic.** The default HTTP inspect map is permissive. Custom maps with strict rules (no PUT, max URI 256) are easy to write and easy to over-tighten. Test with the actual application before binding to a production interface; the application-layer reset comes back as a connection reset to the client, with no obvious clue at the client end. **Forgetting to count.** Every `class` entry in a `policy-map` with no actions does nothing useful. The `show service-policy` output still shows it, with packet counters at zero or a constant. Watch for a class with growing matches but no actions - that is configuration that has no effect. ## Key Takeaways The Cisco ASA MPF is class-map matches traffic, policy-map applies actions, service-policy binds the policy globally or to an interface. The default `global_policy` ships with most application-layer inspections enabled and the lab adds `inspect icmp` so ping replies survive the firewall. Custom inspection adds rules for HTTP, FTP, ESMTP, and other protocols where you want application-layer security; per-interface policies override global for matching classes. For where the inspection engines fire in the per-packet pipeline, see [Cisco ASA Packet Flow](https://www.pinglabz.com/cisco-asa-packet-flow/). For confirmation that a specific flow is hitting an inspection class, run `show service-policy` and watch the packet counter for the class climb. The full reading order is on the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). ### Cisco ASA Upgrade and Backup Procedure URL: https://www.pinglabz.com/cisco-asa-upgrade-backup/ Last updated: 2026-06-13T20:08:14.000Z Backup before you change. Upgrade in a maintenance window. Test the rollback path before you need it. The Cisco ASA upgrade and backup procedures are unglamorous but they are the difference between a smooth Tuesday and a 2am bridge call. This article walks the production-ready procedure: what to back up, how to copy it off-box, how to stage a new image to flash, set the boot variable, reload, verify, and roll back if needed. The procedure works equally for a single ASA and a failover pair, with the pair-specific notes called out where they matter. This is an Operations article in the [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/). It pairs with [common outages](https://www.pinglabz.com/cisco-asa-common-outages/) (which covers what to do when an upgrade goes wrong) and [active/standby failover](https://www.pinglabz.com/cisco-asa-active-standby-failover/) (which is the foundation for zero-downtime upgrades). ## What to Back Up Before Touching Anything startup-config Why Last-known-good config in flash. Always. How to capture it `copy running-config startup-config` then `copy disk0:/startup-config scp://...` running-config Why Current state in RAM. Diverges from startup until you save. How to capture it `copy running-config scp://user@host/path` Identity certificates and private keys Why Cert revocation requires the CSR/key; lost keys mean a new cert. How to capture it `crypto ca export TRUSTPOINT pkcs12 PASSPHRASE` VPN pre-shared keys Why Masked in `show running-config`; you need the original list from your secrets vault. How to capture it External record (1Password, Vault, etc.) AAA shared secrets WhySame; masked. How to capture itExternal record Currently-running image filename Why The exact image you are running today, in case you need to roll back. How to capture it `show version | include image` License entitlements Why Smart licensing token, ASAv platform license How to capture it `show license all` \+ screenshot of the Smart Licensing portal The "external record" rows are the most-overlooked part of an ASA backup. The text config does not contain the actual PSK or shared secret values; the box stores them encrypted and prints `*****` in `show running-config`. You cannot reconstruct a tunnel-group from the config alone if you lose the original PSK list. ## Copy the Config Off-Box SCP is the modern way. SFTP and FTP also work. TFTP is for the lab only. ``` ASA-PERIM# copy running-config scp://backup-user@10.10.0.30/asa-perim/asa-perim-2026-05-10.cfg Source filename [running-config]? Address or name of remote host [10.10.0.30]? Destination username [backup-user]? Destination password []? ********** Destination filename [asa-perim/asa-perim-2026-05-10.cfg]? Writing file scp://10.10.0.30/asa-perim/asa-perim-2026-05-10.cfg !!!!!!!!!!!!!!!!! 12834 bytes copied in 1.220 secs (10519 bytes/sec) ``` For automated nightly backups, set up a scheduled cron on a backup host that pulls via SSH and the show-running-config command, or use a config-management tool (Solarwinds NCM, Oxidized, etc.) that does the same in a structured way. ## Export Certificates and Keys Identity certificates and the private keys behind them ride in the running config in encrypted PKCS#12 form. To export them, pick a passphrase (record it in your vault) and run: ``` ASA-PERIM(config)# crypto ca export ASDM-CERT pkcs12 BackupPassphr@se Exported pkcs12 follows: -----BEGIN PKCS12----- MIIJ4QIBAzCCCacGCSqGSIb3DQEHAaCCCZgEggmUMIIJkDCCBgkGCSqGSIb3DQEHAaCCBfoEggX2 [base64 continues for several KB] -----END PKCS12----- ASA-PERIM(config)# ``` Copy the entire base64 block (between the BEGIN and END lines) into a file in your secrets store. To restore on a replacement ASA, paste it back with `crypto ca import TRUSTPOINT pkcs12 PASSPHRASE`. See [trustpoints and certificates](https://www.pinglabz.com/cisco-asa-anyconnect-certificates/) for the full cert lifecycle. ## Stage the New Image to Flash Cisco distributes ASA images as `.SPA` files on cisco.com. Get the file (and the MD5/SHA512 from the download page), then push it to the ASA's flash: ``` ASA-PERIM# copy scp://image-user@10.10.0.30/asa9-23-2-smp-k8.SPA disk0:/ Address or name of remote host [10.10.0.30]? Source filename [asa9-23-2-smp-k8.SPA]? Destination filename [asa9-23-2-smp-k8.SPA]? Destination password []? ********** Accessing scp://image-user@10.10.0.30/asa9-23-2-smp-k8.SPA... !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 116231632 bytes copied in 24.560 secs (4732097 bytes/sec) ``` Verify the image is intact and is what Cisco signed: ``` ASA-PERIM# verify /sha-512 disk0:/asa9-23-2-smp-k8.SPA Verifying file integrity of disk0:/asa9-23-2-smp-k8.SPA... .................................................................... Computed Hash: SHA2:7a3c8f2d ... [matches the value from cisco.com] Embedded Hash: SHA2:7a3c8f2d ... [matches Cisco's embedded signature] Computed signature: PASS ASA-PERIM# ``` Both hashes match and the embedded-signature check passes. If either fails, do not boot the image - re-download from Cisco. A bad image will not boot, and on a remote box that means a console-cable trip. ## Set the Boot Variable The `boot system` command tells the ASA which image to load on the next reload. Multiple `boot system` lines are allowed; the ASA tries them in order and falls back if the first is corrupted or missing. ``` ASA-PERIM(config)# show running-config boot boot system disk0:/asa9-23-1-smp-k8.SPA ! ASA-PERIM(config)# boot system disk0:/asa9-23-2-smp-k8.SPA ASA-PERIM(config)# no boot system disk0:/asa9-23-1-smp-k8.SPA ASA-PERIM(config)# boot system disk0:/asa9-23-1-smp-k8.SPA ! re-add the old as the fallback ! ASA-PERIM(config)# show running-config boot boot system disk0:/asa9-23-2-smp-k8.SPA boot system disk0:/asa9-23-1-smp-k8.SPA ASA-PERIM(config)# write memory ``` Order matters. The first `boot system` wins; the second is the safety net. If the new image refuses to boot, the ASA falls back to the second line and you reload back into the working version. ## Reload and Verify Single-ASA reload is one command: ``` ASA-PERIM# reload Proceed with reload? [confirm] ``` The box drops the console for 3-5 minutes (longer on hardware ASA with bigger image trees). After it comes back: ``` ASA-PERIM# show version | include System image|Compiled System image file is "disk0:/asa9-23-2-smp-k8.SPA" Compiled on 2026-04-15 12:34 by builder ``` Image is the new one. Sanity-check the rest: ``` ASA-PERIM# show interface ip brief ! All data interfaces up? ASA-PERIM# show route ! Default + static routes intact? ASA-PERIM# show xlate count ! Traffic flowing? ASA-PERIM# show conn count ! Conns building? ASA-PERIM# show vpn-sessiondb summary ! Active VPN sessions? ASA-PERIM# show failover ! (if HA pair) Both units up, sync OK? ``` If anything looks wrong, the rollback path is just another reload after re-ordering the `boot system` lines. ## Upgrading a Failover Pair (Hitless) Active/standby failover lets you upgrade with no traffic loss. The Cisco-recommended sequence: 1. Copy the new image to **both** units' flash. The standby auto-syncs config, but does not auto-sync image files. 2. Set `boot system disk0:/NEW-IMAGE.SPA` on both units (config replicates). 3. Reload the standby first: `failover reload-standby`. This reloads the standby unit only; it boots the new image and rejoins the pair as a now-current standby. 4. Verify on the active: `show failover state` should show standby unit healthy and on the new version. 5. Force a switchover: `no failover active` on the active (or `failover active` on the standby). The standby (new image) becomes active. Existing TCP conns survive because of stateful failover. See [stateful failover](https://www.pinglabz.com/cisco-asa-stateful-failover/). 6. Reload the now-standby (which is on the old image): `failover reload-standby`. 7. Both units now on the new image. Optionally `failover active` back to the original primary if you care about which physical unit is active. Cisco's TAC notes that some major-version jumps are not "hitless-failover-safe" - the version skew during the rolling upgrade is not always clean. Read the release notes for your target version's "Upgrade Path" section. The 9.x to 9.x point-version upgrades are almost always safe; cross-major (8.x to 9.x or 9.x to 10.x) often require both units off-active for the cutover. ## Rollback Two scenarios: **The box booted the new image but something is broken.** Re-order the `boot system` lines so the old image is first, save, reload. ``` ASA-PERIM(config)# no boot system disk0:/asa9-23-2-smp-k8.SPA ASA-PERIM(config)# no boot system disk0:/asa9-23-1-smp-k8.SPA ASA-PERIM(config)# boot system disk0:/asa9-23-1-smp-k8.SPA ASA-PERIM(config)# boot system disk0:/asa9-23-2-smp-k8.SPA ASA-PERIM(config)# write memory ASA-PERIM# reload ``` **The box did not come back up after the upgrade reload.** You are in ROMmon-or-console territory. Console in, interrupt the boot, set the ROMmon boot variable to the old image: ``` rommon> boot disk0:/asa9-23-1-smp-k8.SPA ``` Once it boots, fix the boot lines in the running config and save. The remote restore-from-backup procedure (TFTP server reachable from the management interface, point ROMmon at it) is what to use if disk0 itself is corrupt - which is rare but does happen. ## Restore from Backup To restore a config to a fresh ASA: ``` ! Console into the new box. ! Set up enough of the management interface to reach SCP source. ciscoasa(config)# interface Management0/0 ciscoasa(config-if)# nameif management ciscoasa(config-if)# ip address 10.10.99.1 255.255.255.0 ciscoasa(config-if)# no shutdown ciscoasa(config)# route management 10.10.0.0 255.255.0.0 10.10.99.254 1 ! ciscoasa# copy scp://backup-user@10.10.0.30/asa-perim/asa-perim-2026-05-10.cfg running-config ``` The config loads, the interfaces re-IP themselves, and SSH/AAA come back online. Save it: `write memory`. Then re-import any certificates with `crypto ca import` from your PKCS#12 backup. ## Pre-Upgrade Checklist 1. Read the release notes for the target version's Upgrade Path and Open Caveats. 2. Confirm sufficient flash free space (`show disk0:`). 3. Back up startup-config, running-config, certificates, PSK list. 4. Stage the new image to flash on both units (failover pair). 5. Verify the image hash and signature. 6. Set `boot system NEW` with the old image as the fallback. 7. `write memory`. 8. Identify and notify users of the maintenance window. 9. Have console access ready in case the remote management interface is the thing that breaks. 10. Reload (failover pair: standby first), then run the verification show commands. ## Key Takeaways Backups before changes, image hashes before reloads, two `boot system` lines for rollback safety, and console access at the ready. SCP is the modern transport for both config and image transfers. PKCS#12 export is the only way to back up certificates and their private keys. Failover pairs upgrade hitless if you reload the standby first and let stateful sync carry the conns through the switchover. For the things that go wrong during or after an upgrade, see [common outage scenarios](https://www.pinglabz.com/cisco-asa-common-outages/). For the rest of the cluster, the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/) has the full reading order. ### Cisco ASA Syslog Configuration and Logging Levels URL: https://www.pinglabz.com/cisco-asa-syslog-logging/ Last updated: 2026-06-13T20:08:14.000Z Syslog is what turns the Cisco ASA from a black box into something you can troubleshoot, audit, and SIEM. The configuration is small (six lines for a sane production setup), but the message format and the severity model both have ASA-specific quirks worth knowing before you point a log collector at the box. This article walks the syslog destinations, the eight severity levels, message lists for filtering, and how to read the `%ASA-X-XXXXXX` format using real lab logs from Sessions 3 and 4 (PSK-mismatch IKEv2 failures and OUTSIDE\_IN ACL denies). This is a [Cisco ASA](https://www.pinglabz.com/cisco-asa/) Fundamentals article. The companion logging-side topics are [common outage scenarios](https://www.pinglabz.com/cisco-asa-common-outages/) (where the same syslog patterns confirm the diagnosis) and [ACL troubleshooting](https://www.pinglabz.com/cisco-asa-acl-troubleshooting/) (where the deny-and-log pattern is the diagnostic technique). ## The Eight Severity Levels The ASA uses the standard syslog severity model, 0-7, with smaller numbers meaning more urgent. emergencies Level0 What lives here System unusable. Catastrophic platform failures. alerts Level1 What lives here Immediate action required. Failover events, environmental alarms. critical Level2 What lives here Critical conditions. Tunnel-mgr "failed to establish L2L SA," interface down on a critical link. errors Level3 What lives here Error conditions. Most VPN auth failures, NAT pool exhaustion, ACL drops with high frequency. warnings Level4 What lives here Warning conditions. ACL denies, bad-cert presented, IPsec proposal mismatches. notifications Level5 What lives here Normal but significant. Config changes, SSH login allowed, IKEv2 SA established. informational Level6 What lives here Informational. Most accept-flow logs, AAA accept, conn create/teardown if logged. debugging Level7 What lives here Debug-only. Internal state machine traces. Each destination on the ASA is configured with a maximum severity. Setting a destination to `informational` (level 6) means you receive everything from level 0 through level 6; level 7 (debugging) is still excluded. The standard production setup is `informational` to a remote syslog collector and `warnings` or `notifications` to the buffered log. ## Logging Destinations The ASA can log to seven destinations independently. The four you actually care about: Console Configured with `logging console SEVERITY` When to use Almost never. Console logging at any high rate slows the CLI to a crawl. Buffered (RAM) Configured with `logging buffered SEVERITY` When to use Always. The most-recent-N log lines kept in RAM, readable with `show logging`. Survives until reload or a clear. Trap (remote syslog server) Configured with `logging trap SEVERITY` \+ `logging host ...` When to use Always. UDP/514 to your SIEM or syslog collector. SSH/Telnet sessions ("monitor") Configured with `logging monitor SEVERITY` \+ `terminal monitor` When to use Ad-hoc, during a debug session. Not for production. Two more destinations exist (e-mail and ASDM) but they are rarely the right answer for production at scale. ## Minimum-Viable Config Six lines: ``` ASA-PERIM(config)# logging enable ASA-PERIM(config)# logging timestamp ASA-PERIM(config)# logging buffer-size 1048576 ASA-PERIM(config)# logging buffered informational ASA-PERIM(config)# logging trap informational ASA-PERIM(config)# logging host inside 10.10.0.30 17/514 ``` What each does: - `logging enable`: master switch. Without this, none of the others do anything. - `logging timestamp`: prepend timestamps to every message. Default is no timestamp, which makes correlation between the ASA and other devices much harder. - `logging buffer-size 1048576`: 1 MB ring buffer (default 4 KB, which fills in seconds on a busy box). - `logging buffered informational`: keep level 6 and above in the buffer. - `logging trap informational`: send level 6 and above to the syslog server. - `logging host inside 10.10.0.30 17/514`: send to `10.10.0.30` over the inside interface, UDP/514\. The `17/514` is "protocol 17 (UDP), port 514." For TCP syslog (some collectors prefer it), use `6/1470` or whatever the collector listens on. TCP gives you delivery guarantees but introduces backpressure: if the collector dies, the ASA's TCP send queue fills and logging stalls. Most production sites use UDP for the ASA and accept the rare lost packet. ## Reading the Message Format Every ASA syslog has the same shape: `%ASA-LEVEL-MESSAGEID: free-form text`. From the lab's Session 3 PSK-mismatch failure on the IKEv2 site-to-site to ASA-PARTNER: ``` %ASA-4-750003: Local:203.0.113.2:500 Remote:203.0.113.6:500 Username:203.0.113.6 IKEv2 Negotiation aborted due to ERROR: Failed to authenticate the IKE SA %ASA-3-752015: Tunnel Manager has failed to establish an L2L SA. All configured IKE versions failed to establish the tunnel. Map Tag= OUTSIDE-MAP. Map Sequence Number = 10. ``` Walk through `%ASA-4-750003`: - `%ASA-`: the platform prefix. Always there. - `4`: severity 4 (warnings). - `-750003`: the message ID. Unique. Look it up in the Cisco docs (`show logging message 750003` on the ASA also helps) for the official explanation. - `:` separator. - Free-form text describing what happened. Once you internalize the format, you can scan a log dump and pick out the level and the message ID at a glance, even when the free-form text varies. The two messages above show the textbook IKEv2 PSK-failure pattern: a 750003 (auth failed) immediately followed by a 752015 (tunnel mgr gives up). See [troubleshoot IPsec phases](https://www.pinglabz.com/cisco-asa-troubleshoot-ipsec-phases/) for the full PSK-failure walk. And from the lab's Session 4 OUTSIDE\_IN ACL deny, when something hit line 5 (the explicit catch-all): ``` %ASA-4-106023: Deny tcp src outside:8.8.8.8/56321 dst dmz:192.168.50.10/8080 by access-group "OUTSIDE_IN" ``` `%ASA-4-106023` is the canonical "ACL deny" message ID. The text gives you the source interface, source IP+port, destination interface, destination IP+port, and the ACL name. That single line is enough to confirm a denial and start working out which ACL line to add. The companion message ID `%ASA-6-106100` is the matching `permit` log line if you have `log` on a permit ACE. ## Message Lists: Filter Per Destination Sometimes you want a different filter per destination. Send everything to the central SIEM but only the headlines to a paging system. ``` ASA-PERIM(config)# logging list HEADLINES level critical ASA-PERIM(config)# logging list HEADLINES message 106023 ! ACL deny ASA-PERIM(config)# logging list HEADLINES message 750003 ! IKEv2 auth failed ASA-PERIM(config)# logging list HEADLINES message 752015 ! Tunnel mgr give-up ! ASA-PERIM(config)# logging trap HEADLINES ASA-PERIM(config)# logging host inside 10.10.0.40 17/514 ``` The list `HEADLINES` includes everything at level critical or above, plus three specific message IDs regardless of level. You then bind that list to `logging trap`, and only matching messages go to the host on the next line. You can also include or exclude classes of messages (`logging class vpn`, `logging class auth`, etc.). The class names map to subsystem groupings. `show logging class` lists every class. ## Rate Limiting The ASA has a per-message rate limiter. By default it is enabled with sane values, but a noisy message ID can still spam the buffer and the collector. Tune it for known-noisy IDs: ``` ASA-PERIM(config)# logging rate-limit 10 60 message 106023 ``` That caps message 106023 (ACL deny) to 10 occurrences per 60 seconds. Anything beyond that is suppressed (and accounted for in `show logging rate-limit`). Useful when a security scanner is hammering your perimeter and you do not want a million identical deny logs filling your SIEM. ## Reading the Buffer with show logging ``` ASA-PERIM# show logging | tail %ASA-6-302013: Built outbound TCP connection 1234 for outside:8.8.8.8/443 (8.8.8.8/443) to inside:10.10.0.50/53412 (203.0.113.2/53412) %ASA-4-106023: Deny tcp src outside:198.51.100.99/52891 dst dmz:192.168.50.10/443 by access-group "OUTSIDE_IN" %ASA-6-302014: Teardown TCP connection 1232 for outside:8.8.8.8/443 to inside:10.10.0.50/53411 duration 0:00:14 bytes 8421 TCP FINs from inside ``` Useful filters when the buffer is huge: `show logging | include 106023` Just the ACL denies. `show logging | include 192.168.50.10` Anything mentioning the DMZ web server. `show logging | include vpn` Filter by free-form keyword. `show logging asdm` Just the ASDM-specific log buffer (separate from the main). `show logging queue` Counters: total messages, queued, dropped. `show logging rate-limit` Which message IDs are being rate-limited. Clear the buffer with `clear logging buffer`. The remote syslog stream is unaffected; the buffer is local-only. ## Message IDs Worth Memorizing A short list of message IDs that come up so often it is faster to learn them than to look them up. 106023 Severity4 EventACL deny 106100 Severity6 Event ACL permit (when `log` keyword set) 302013/302014 Severity6 Event TCP conn build / teardown 302015/302016 Severity6 Event UDP conn build / teardown 305011/305012 Severity6 Event NAT translation built / teardown 305009 Severity6 EventStatic NAT built 113004/113005 Severity6 Event AAA authentication accepted / rejected 113019 Severity4 Event VPN session disconnected (with reason) 605004/605005 Severity6 Event SSH login allowed / session started 713903 / 713904 / 713905 / 713906 Severity5/7 Event IKEv1 / IKEv2 packet receive and phase tracking 750001-750003 Severity5/4 Event IKEv2 SA negotiation, success and failure 752015 Severity3 Event Tunnel Manager failed to establish L2L SA 111007/111008 Severity5 Event Config command issued (full audit trail) 199002/199003 Severity5/3 Event Failover unit role change SIEM rules built around the right message IDs are an order of magnitude more useful than rules built on free-text matching. The IDs are stable across versions; the free-form text occasionally changes between releases. ## Worked Example: PSK-Mismatch IKEv2 Tunnel From the lab. PSK on ASA-PERIM was deliberately set wrong. Traffic that should have triggered the S2S tunnel produced this sequence in the buffered log: ``` %ASA-5-752003: Tunnel Manager dispatching a KEY_ACQUIRE message to IKEv2. Map Tag = OUTSIDE-MAP. Map Sequence Number = 10. %ASA-5-750001: Local:203.0.113.2:500 Remote:203.0.113.6:500 Username:Unknown IKEv2 Received request to establish an IPsec tunnel... %ASA-7-713906: IKE Receiver: Packet received on 203.0.113.2:500 from 203.0.113.6:500 %ASA-7-713906: IKE Receiver: Packet received on 203.0.113.2:500 from 203.0.113.6:500 %ASA-4-750003: Local:203.0.113.2:500 Remote:203.0.113.6:500 Username:203.0.113.6 IKEv2 Negotiation aborted due to ERROR: Failed to authenticate the IKE SA %ASA-4-752012: IKEv2 was unsuccessful at setting up a tunnel. Map Tag = OUTSIDE-MAP. Map Sequence Number = 10. %ASA-3-752015: Tunnel Manager has failed to establish an L2L SA. All configured IKE versions failed to establish the tunnel. Map Tag= OUTSIDE-MAP. Map Sequence Number = 10. %ASA-7-752002: Tunnel Manager Removed entry. Map Tag = OUTSIDE-MAP. Map Sequence Number = 10. ``` Reading the IDs and severities top to bottom: 752003 (kicked off), 750001 (got the request), 713906 x2 (received packets), 750003 (auth failed - severity 4, the first warning), 752012 (gave up at IKEv2 layer), 752015 (gave up at Tunnel Mgr layer - severity 3, critical), 752002 (cleaned up). The pattern of `auth failed -> tunnel mgr give-up` within milliseconds is the classic PSK-mismatch fingerprint. See [IPsec phase troubleshooting](https://www.pinglabz.com/cisco-asa-troubleshoot-ipsec-phases/) for the response. ## Always Set Time The `logging timestamp` command prepends a timestamp, but only as accurate as the ASA's clock. NTP-synced and authenticated time is mandatory before logs are useful for correlation: ``` ASA-PERIM(config)# clock timezone UTC 0 ASA-PERIM(config)# ntp authenticate ASA-PERIM(config)# ntp authentication-key 1 md5 NTP-Key ASA-PERIM(config)# ntp trusted-key 1 ASA-PERIM(config)# ntp server 10.10.0.10 key 1 source inside prefer ``` UTC for everything network-side. Your SIEM can localize to operator time zones at display. ## Verify the Pipeline Three checks after configuring logging: ``` ASA-PERIM# show logging Syslog logging: enabled Facility: 20 Timestamp logging: enabled Standby logging: disabled Debug-trace logging: disabled Console logging: disabled Monitor logging: disabled Buffer logging: level informational, 8423 messages logged Trap logging: level informational, 8421 messages logged to host inside:10.10.0.30 History logging: disabled Device ID: hostname "ASA-PERIM" Mail logging: disabled ASDM logging: disabled ``` "Buffer logging" and "Trap logging" should both have non-zero counts climbing. If the trap count is zero, the host line is misconfigured or the collector is unreachable. Verify with a packet capture on the inside interface filtering UDP/514 (see [packet capture](https://www.pinglabz.com/cisco-asa-packet-capture/) for the syntax). ## Key Takeaways Cisco ASA syslog is small to configure and high-leverage to operate. Six lines get you a 1 MB local buffer plus a remote syslog stream at level informational. The `%ASA-LEVEL-MESSAGEID:` format is consistent across every message, so your SIEM should match on message IDs (which are stable) rather than free-form text (which is not). Enable `logging timestamp` and pin the ASA to authenticated NTP before you depend on logs for correlation. Rate-limit known-noisy IDs (106023 ACL deny is the classic) so a security scanner does not bury the rest of your log stream. For the operational uses of these logs, see [ACL troubleshooting](https://www.pinglabz.com/cisco-asa-acl-troubleshooting/) (which uses 106023 deny logs as the diagnostic) and [common outages](https://www.pinglabz.com/cisco-asa-common-outages/) (which uses message IDs to fingerprint different failure modes). The full reading order is on the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). ### Cisco ASA Management Plane Hardening URL: https://www.pinglabz.com/cisco-asa-management-hardening/ Last updated: 2026-06-13T20:08:23.000Z The management plane is where most ASA breaches happen, and most of those breaches are preventable with a handful of configuration steps that take an hour to apply. This article covers the production-ready hardening pass for a Cisco ASA: SSH and HTTPS the right way, AAA against an external server with local fallback, a per-interface management ACL, password policy, banner, and the cleanup steps that disable the things you do not need. The lab's three AAA server-groups (RADIUS-VPN, LDAP-VPN, TACACS-ADMIN) appear here because the same server-group framework drives both VPN auth and admin auth. This is the [Cisco ASA](https://www.pinglabz.com/cisco-asa/) Fundamentals article that maps onto the CIS Benchmark and most internal hardening standards. After running these commands, your ASA is something an auditor will not flag for the obvious findings. ## SSH: Strong Keys, Strong Algorithms, Strict Sources The default SSH config on a fresh ASA is permissive. Five things to tighten: Key size DefaultRSA 1024 HardenedRSA 2048+ or ECDSA Protocol version Default1 and 2 Hardened2 only Idle timeout Default5 minutes Hardened 5-10 minutes (unchanged is fine) Permitted source subnets DefaultNone (deny by default) Hardened Explicit per-interface allow-list Cipher and KEX algorithms DefaultWide list HardenedStrong-only The configuration block: ``` ASA-PERIM(config)# crypto key zeroize rsa ASA-PERIM(config)# crypto key generate rsa modulus 2048 ASA-PERIM(config)# crypto key generate ecdsa elliptic-curve 384 ! ASA-PERIM(config)# ssh version 2 ASA-PERIM(config)# ssh timeout 10 ASA-PERIM(config)# ssh key-exchange group dh-group14-sha256 ASA-PERIM(config)# ssh cipher encryption high ASA-PERIM(config)# ssh cipher integrity high ! ASA-PERIM(config)# ssh 10.10.0.0 255.255.255.0 inside ASA-PERIM(config)# ssh 10.10.99.0 255.255.255.0 management ``` The `cipher encryption high` and `cipher integrity high` presets restrict the algorithm list to current strong choices (AES-256-GCM, ChaCha20, HMAC-SHA2-256). The presets get updated by Cisco as algorithms age out, so using the preset rather than naming individual algorithms means you inherit future tightening without revisiting the config. The two `ssh` permit lines are the ASA's built-in SSH ACL: SSH from `10.10.0.0/24` on inside or `10.10.99.0/24` on management is allowed. SSH from anywhere else is silently dropped. Without at least one `ssh` permit line for an interface, the daemon does not listen on that interface at all. ## HTTPS / ASDM: Lock It Down or Turn It Off If you do not use ASDM, turn the HTTPS server off: ``` ASA-PERIM(config)# no http server enable ``` If you do use ASDM, pin it to a specific source list and use a real cert: ``` ASA-PERIM(config)# http server enable ASA-PERIM(config)# http 10.10.0.50 255.255.255.255 inside ASA-PERIM(config)# http 10.10.0.51 255.255.255.255 inside ASA-PERIM(config)# http server idle-timeout 10 ASA-PERIM(config)# ssl server-version tlsv1.2 dtlsv1.2 ASA-PERIM(config)# ssl cipher tlsv1.2 high ASA-PERIM(config)# ssl trust-point ASDM-CERT outside ``` The cert pinning matters for the same reason it does for AnyConnect: if you leave the default self-signed cert, every admin connection trains people to click through cert warnings, which trains them to click through a real attack later. Use a CA-signed cert if the ASA's HTTPS server is reachable from anywhere outside your tightly-controlled management subnet. See [trustpoints and certificates](https://www.pinglabz.com/cisco-asa-anyconnect-certificates/) for the cert workflow. ## Local Users: Privilege 15, Hashed Passwords ``` ASA-PERIM(config)# username admin password Cisc0_Admin_S3cret! privilege 15 ASA-PERIM(config)# username noc password Cisc0_NOC_View1 privilege 5 ASA-PERIM(config)# username break-glass password Cisc0_Glass_R3cover! privilege 15 ``` Three roles. The break-glass account exists for the case where AAA is unreachable; document its existence, rotate its password quarterly, and store it in a sealed envelope or a privileged-access vault. Skip this account and a RADIUS outage cuts you off from the firewall during the exact incident you need to log in to fix. Privilege 15 is full enable. Privilege 5 (or any custom level you define with `privilege ... level N command ...`) is read-mostly. Most NOCs are happy with a level-5 user that can run `show` and `packet-tracer` but cannot configure. ## AAA Server Groups (RADIUS / LDAP / TACACS+) External AAA is the single biggest hardening win because it puts admin accounts under your existing identity-management lifecycle (joiner/mover/leaver). When somebody leaves the team, disabling them in AD or your IdP cuts off their ASA access automatically. The PingLabz lab has three server groups configured (one per protocol), defined in Session 3: ``` ASA-PERIM# show running-config aaa-server aaa-server RADIUS-VPN protocol radius aaa-server RADIUS-VPN (inside) host 10.10.0.10 retry-interval 2 timeout 5 key ***** authentication-port 1812 accounting-port 1813 aaa-server LDAP-VPN protocol ldap aaa-server LDAP-VPN (inside) host 10.10.0.20 server-port 636 ldap-base-dn dc=pinglabz,dc=lab ldap-scope subtree ldap-naming-attribute sAMAccountName ldap-login-password ***** ldap-login-dn cn=svc-asa,ou=Service,dc=pinglabz,dc=lab ldap-over-ssl enable server-type microsoft aaa-server TACACS-ADMIN protocol tacacs+ aaa-server TACACS-ADMIN (inside) host 10.10.0.30 timeout 5 key ***** ``` For admin login the standard practice is TACACS+ because it gives you per-command authorization and full command accounting. Bind the TACACS group to SSH login: ``` ASA-PERIM(config)# aaa authentication ssh console TACACS-ADMIN LOCAL ASA-PERIM(config)# aaa authentication enable console TACACS-ADMIN LOCAL ASA-PERIM(config)# aaa authentication http console TACACS-ADMIN LOCAL ASA-PERIM(config)# aaa authentication serial console LOCAL ! ASA-PERIM(config)# aaa authorization command TACACS-ADMIN LOCAL ASA-PERIM(config)# aaa accounting command TACACS-ADMIN ASA-PERIM(config)# aaa accounting enable console TACACS-ADMIN ``` Walk through the binds. Authentication runs against TACACS-ADMIN first, falling back to LOCAL if every TACACS server is unreachable (the LOCAL keyword is critical - omit it and a TACACS outage locks you out). Authorization uses the same group to ask TACACS "is this user allowed to run this command?" - useful when you want a NOC role that can run `show` but not `configure`. Accounting logs every command and every enable to the AAA server, giving you an audit trail. Console authentication usually stays LOCAL because if you are at the console, AAA might be the very thing that is broken. ## Management ACL: Defense in Depth The per-interface `ssh` and `http` permit lines are the first ACL. A second ACL at the interface level adds defense in depth and lets you log management-plane attempts in a single place. ``` ASA-PERIM(config)# object-group network MGMT-SOURCES ASA-PERIM(config-network-object-group)# network-object 10.10.0.0 255.255.255.0 ASA-PERIM(config-network-object-group)# network-object 10.10.99.0 255.255.255.0 ASA-PERIM(config-network-object-group)# network-object host 10.20.5.10 ! Jump host ! ASA-PERIM(config)# access-list MGMT-IN extended permit tcp object-group MGMT-SOURCES interface inside eq ssh ASA-PERIM(config)# access-list MGMT-IN extended permit tcp object-group MGMT-SOURCES interface inside eq https ASA-PERIM(config)# access-list MGMT-IN extended permit udp object-group MGMT-SOURCES interface inside eq snmp ASA-PERIM(config)# access-list MGMT-IN extended deny ip any interface inside log informational interval 60 ASA-PERIM(config)# access-list MGMT-IN extended permit ip any any ASA-PERIM(config)# access-group MGMT-IN in interface inside control-plane ``` The `control-plane` keyword is the magic word: this ACL applies only to traffic destined for the ASA itself (the management plane), not to through-traffic. The final `permit ip any any` ensures through-traffic continues to be governed by your real OUTSIDE\_IN / INSIDE\_OUT policies elsewhere; the deny in the middle only fires for management-plane attempts that did not match the explicit allow lines. ## Password Policy ``` ASA-PERIM(config)# password-policy minimum-length 14 ASA-PERIM(config)# password-policy minimum-uppercase 1 ASA-PERIM(config)# password-policy minimum-lowercase 1 ASA-PERIM(config)# password-policy minimum-numeric 1 ASA-PERIM(config)# password-policy minimum-special 1 ASA-PERIM(config)# password-policy minimum-changes 4 ASA-PERIM(config)# password-policy lifetime 90 ASA-PERIM(config)# password-policy reuse-interval 12 ``` The `lifetime 90` setting forces a password change every 90 days. `reuse-interval 12` prevents reusing any of the last 12 passwords. `minimum-changes 4` means a new password must differ from the old one in at least 4 character positions, which prevents the classic "Password1 -> Password2" pattern. If your AAA server (Microsoft AD, etc.) already enforces a password policy, the local policy mostly applies to the few break-glass accounts. Set both anyway; defense in depth. ## Login Banner Banners are not a technical control but they are a legal one. A clearly-worded banner asserting that the system is restricted, monitored, and logged is what allows your incident response to actually do something with the audit logs you collect. ``` ASA-PERIM(config)# banner motd ************************************************************ ASA-PERIM(config)# banner motd WARNING: Authorized access only. ASA-PERIM(config)# banner motd Activity on this device is monitored, logged, and audited. ASA-PERIM(config)# banner motd Unauthorized access will be prosecuted under applicable law. ASA-PERIM(config)# banner motd ************************************************************ ASA-PERIM(config)# banner login Please log in: ASA-PERIM(config)# banner exec Logged-in users have agreed to the MOTD policy. ``` The MOTD shows before login (useful for the legal warning). The login banner shows immediately before the username prompt. The exec banner shows after successful login. Skip language like "welcome" or "hello" - those have, in past cases, been used by defense to argue the banner did not constitute a clear no-trespass notice. Boring legalese works better in court. ## SNMP: v3 Only, Strong Auth, Privacy SNMPv1 and v2c send community strings in cleartext. Disable them. ``` ASA-PERIM(config)# no snmp-server enable traps snmp authentication ASA-PERIM(config)# snmp-server group ADMIN-GROUP v3 priv ASA-PERIM(config)# snmp-server user noc-monitor ADMIN-GROUP v3 auth sha NOC-Auth-S3cret priv aes 256 NOC-Priv-S3cret ASA-PERIM(config)# snmp-server host inside 10.10.0.40 version 3 noc-monitor ``` SNMPv3 with `auth sha priv aes 256` gives you authenticated, privacy-encrypted polling. The username/auth/priv credentials replace the cleartext community string of v2c. ## Disable What You Do Not Use Each enabled service is a potential attack surface. Turn off what you do not need: ``` ASA-PERIM(config)# no telnet 0.0.0.0 0.0.0.0 inside ASA-PERIM(config)# no telnet 0.0.0.0 0.0.0.0 outside ASA-PERIM(config)# no telnet 0.0.0.0 0.0.0.0 management ASA-PERIM(config)# no http server enable ! if you do not use ASDM ASA-PERIM(config)# no ssh stricthostkeycheck ! relax only if you have a reason ASA-PERIM(config)# no service password-recovery ! prevents ROMmon password reset; use only with break-glass discipline ``` `no service password-recovery` is the most dangerous of these. It hardens the ASA against an attacker with physical console access, but it also means your password-reset path is "wipe the box and reload from backup." Only enable it if you have a tested off-box backup and a documented restore procedure - both of which the [upgrade and backup](https://www.pinglabz.com/cisco-asa-upgrade-backup/) article covers. ## NTP Authentication If your AAA depends on Kerberos or any cert-based protocol, time skew breaks login. Authenticate NTP so an attacker cannot manipulate the clock. ``` ASA-PERIM(config)# ntp authenticate ASA-PERIM(config)# ntp authentication-key 1 md5 NTP-Auth-K3y ASA-PERIM(config)# ntp trusted-key 1 ASA-PERIM(config)# ntp server 10.10.0.10 key 1 source inside prefer ASA-PERIM(config)# ntp server 10.10.0.11 key 1 source inside ``` ## Log the Management Plane Events Hardening without logging is a tree falling in a forest. The ASA emits per-event syslogs for SSH login success/fail, AAA accept/reject, command authorization, and config changes. Make sure these are flowing to your SIEM: ``` ASA-PERIM(config)# logging enable ASA-PERIM(config)# logging timestamp ASA-PERIM(config)# logging buffered informational ASA-PERIM(config)# logging trap informational ASA-PERIM(config)# logging host inside 10.10.0.30 17/514 ``` Specific message IDs to watch for in your SIEM rules: `%ASA-6-605004` SSH login allowed `%ASA-6-605005` SSH login session started `%ASA-6-113004` / `%ASA-6-113005` AAA user authentication accepted / rejected `%ASA-5-111007` Configuration change `%ASA-5-111008` User executed a command (with the command itself) `%ASA-5-502103` User priv level changed Failed-login alarms over a window are the easiest brute-force detection. [syslog configuration and reading](https://www.pinglabz.com/cisco-asa-syslog-logging/) covers the full message format and the recommended remote-server setup. ## The 12-Item Hardening Checklist If you do nothing else from this article, do these: 1. Replace the default 1024-bit RSA host key with 2048+ and ECDSA P-384. 2. `ssh version 2` only. 3. Per-interface `ssh` permit lines limited to known-good source subnets. 4. External AAA (TACACS+ for admin) with explicit `LOCAL` fallback. 5. A break-glass local user, password rotated quarterly. 6. Per-interface management-plane ACL with `control-plane` keyword. 7. Disable telnet on every interface. 8. Disable HTTPS server if you do not use ASDM; pin source subnets if you do. 9. Turn off SNMPv1/v2c; use v3 with auth+priv. 10. Set a clear, legalese MOTD banner. 11. Authenticated NTP from at least two sources. 12. Remote syslog to a SIEM with alerting on the management-plane message IDs. An hour of work. Years of avoided pain. ## Key Takeaways The Cisco ASA management plane is hardened in layers: per-interface SSH/HTTPS permits, control-plane ACL, AAA against an external server with local fallback, password policy on the local accounts, and aggressive logging of every authentication and command event. The lab's three AAA server-groups (RADIUS-VPN / LDAP-VPN / TACACS-ADMIN) work for both VPN and admin auth from the same framework. The single biggest hardening win is wiring TACACS+ for admin login with command authorization and accounting. Once that is in place, joiner/mover/leaver flows through your IdP automatically, and every command an admin runs is logged centrally with the username attached. For the logging side, see [Cisco ASA Syslog Configuration and Logging Levels](https://www.pinglabz.com/cisco-asa-syslog-logging/). For the rest of the cluster, the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/) has the reading order. ### Cisco ASA Static Routing and Default Route Configuration URL: https://www.pinglabz.com/cisco-asa-static-routing/ Last updated: 2026-06-13T20:08:23.000Z Static routing on the Cisco ASA is the core of every routed-mode deployment. Default route to the ISP, internal static routes pointing at downstream Layer 3 switches, redundant routes with administrative distance to fail over to a backup link. The ASA does support OSPF, EIGRP, and BGP, but most production firewalls run all-static routing for the simple reason that the firewall is a security device, not a router, and you usually do not want it forming dynamic adjacencies with neighbors you have not fully audited. This article covers the static-routing syntax, default routes, administrative distance, route tracking with SLA monitors for multi-WAN failover, and how to read `show route` output from the lab. For where this fits in the cluster: see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). Static routing usually sits between the [initial setup](https://www.pinglabz.com/cisco-asa-initial-setup/) and any per-zone build like [inside/outside/DMZ](https://www.pinglabz.com/cisco-asa-inside-outside-dmz/). ## Static Route Syntax One command, four fields plus optional bits: ``` route INTERFACE NETWORK MASK NEXT-HOP [DISTANCE] [tunneled] [track ID] ``` `INTERFACE` Requiredyes Notes The egress nameif (`outside`, `inside`, etc.). Tells the ASA which interface this route resolves through. `NETWORK MASK` Requiredyes Notes Destination network and mask in classful (255.255.255.0) form. Use `0.0.0.0 0.0.0.0` for the default route. `NEXT-HOP` Requiredyes Notes The IP of the next-hop router. Must be in a directly connected subnet on the named interface. `DISTANCE` Requiredno Notes Administrative distance, 1-255\. Default is 1\. Higher values become floating routes. `tunneled` Requiredno Notes Special flag for traffic that arrives via VPN tunnels. Used to override default-route behavior for VPN-decapped traffic. `track ID` Requiredno Notes Bind the route to an SLA track object. Route is removed from the table if the track goes down. The simplest possible default route: ``` ASA-PERIM(config)# route outside 0.0.0.0 0.0.0.0 203.0.113.1 1 ``` An internal static for a downstream subnet: ``` ASA-PERIM(config)# route inside 10.10.10.0 255.255.255.0 10.10.0.1 1 ``` That second line tells the ASA: "to reach `10.10.10.0/24`, hand the packet to `10.10.0.1` over the inside interface." `10.10.0.1` is presumably the inside Layer 3 switch. ## Reading the Output: show route From the lab's ASA-PERIM after the inside/outside/DMZ build: ``` ASA-PERIM# show route Codes: L - local, C - connected, S - static, R - RIP, M - mobile, B - BGP D - EIGRP, EX - EIGRP external, O - OSPF, IA - OSPF inter area N1 - OSPF NSSA external type 1, N2 - OSPF NSSA external type 2 E1 - OSPF external type 1, E2 - OSPF external type 2, V - VPN i - IS-IS, su - IS-IS summary, L1 - IS-IS level-1, L2 - IS-IS level-2 ia - IS-IS inter area, * - candidate default, U - per-user static route o - ODR, P - periodic downloaded static route, + - replicated route SI - Static InterVRF Gateway of last resort is 203.0.113.1 to network 0.0.0.0 S* 0.0.0.0 0.0.0.0 [1/0] via 203.0.113.1, outside C 10.10.0.0 255.255.255.0 is directly connected, inside L 10.10.0.254 255.255.255.255 is directly connected, inside S 10.10.10.0 255.255.255.0 [1/0] via 10.10.0.1, inside C 192.168.50.0 255.255.255.0 is directly connected, dmz L 192.168.50.1 255.255.255.255 is directly connected, dmz C 203.0.113.0 255.255.255.252 is directly connected, outside L 203.0.113.2 255.255.255.255 is directly connected, outside ``` Read the codes column on the left. `C` is connected (a subnet directly attached to a configured interface). `L` is local (the ASA's own /32 IP on each interface). `S` is a static route. `S*` with the asterisk means "candidate default" - this is the gateway-of-last-resort. The `[1/0]` bracketed pair is administrative distance / metric. Distance 1 is the default for static routes. Metric 0 because static routes do not have a meaningful metric. ## Default Route Patterns Three common shapes: Single ISP Command `route outside 0.0.0.0 0.0.0.0 203.0.113.1 1` Use case The 95% case. One ISP, one default route. Primary + floating backup Command \+ `route outside-backup 0.0.0.0 0.0.0.0 198.51.100.1 200` Use case Backup route at distance 200 sits dormant; only installed if the primary disappears. Primary + tracked failover Command Add SLA + track + `track` on the primary route Use case Active failover when the primary's next-hop becomes unreachable, even though the interface stays up. Per-tunnel default Command `route outside 0.0.0.0 0.0.0.0 203.0.113.1 1 tunneled` Use case Override the default for traffic decapsulated from a VPN tunnel. The `tunneled` keyword is worth a quick explanation. By default, traffic that arrives via a VPN tunnel and is destined for an unknown network follows the regular default route, which usually points back to the same outside interface and creates an asymmetric path. `tunneled` creates an alternate default specifically for VPN-arriving traffic, useful in hub-and-spoke designs where decapped traffic should go somewhere else. ## Administrative Distance and Floating Static Routes The ASA, like IOS, picks the route with the lowest administrative distance when multiple routes exist for the same prefix. Default distances: Connected 0 Static 1 EIGRP internal 90 OSPF (any type) 110 RIP 120 EIGRP external 170 iBGP 200 Floating static routes use this to create a backup. Two static defaults at different distances: ``` ASA-PERIM(config)# route outside 0.0.0.0 0.0.0.0 203.0.113.1 1 ASA-PERIM(config)# route outside-bak 0.0.0.0 0.0.0.0 198.51.100.1 200 ``` The first wins because 1 < 200\. The second sits in the configuration but is not in the routing table. If the first interface goes down (line protocol drops), the ASA removes its routes, the floating backup gets installed, and traffic shifts to the secondary path. When the primary comes back, the static reinstates and the backup goes dormant. The catch: floating routes only respond to interface-down events. If the link stays up but the next-hop becomes unreachable (an ISP CE that quietly stops forwarding while the L1 link looks fine), the floating route never installs. For that case, you need active route tracking with SLA monitors. ## Route Tracking with IP SLA Route tracking is a small state machine: an IP SLA probe regularly tests reachability to a chosen target. A track object watches the SLA's reachability state. A static route bound to the track is installed only while the track is up. ``` ASA-PERIM(config)# sla monitor 1 ASA-PERIM(config-sla-monitor)# type echo protocol ipIcmpEcho 8.8.8.8 interface outside ASA-PERIM(config-sla-monitor-echo)# frequency 10 ASA-PERIM(config-sla-monitor-echo)# exit ASA-PERIM(config)# sla monitor schedule 1 life forever start-time now ! ASA-PERIM(config)# track 1 rtr 1 reachability ! ASA-PERIM(config)# route outside 0.0.0.0 0.0.0.0 203.0.113.1 1 track 1 ASA-PERIM(config)# route outside-bak 0.0.0.0 0.0.0.0 198.51.100.1 200 ``` What that does, in order: 1. `sla monitor 1` defines an ICMP echo probe to `8.8.8.8`, sourced out the outside interface, every 10 seconds. 2. `sla monitor schedule` tells the ASA to start running the probe immediately and keep running it forever. 3. `track 1 rtr 1 reachability` creates a track object that follows SLA 1's "reachable" state. 4. The primary default route is bound to `track 1`. While the track is up, the route is installed at distance 1\. When the track goes down, the route is withdrawn. 5. The backup route at distance 200 has no track and is always available; it gets installed automatically when the tracked primary disappears. Verify: ``` ASA-PERIM# show track 1 Track 1 Response Time Reporter 1 reachability Reachability is Up 4 changes, last change 02:14:31 Latest operation return code: OK Latest RTT (millisecs) 12 Tracked by: STATIC-IP-ROUTING 0 ASA-PERIM# show sla monitor operational-state Entry number: 1 Modification time: 03:21:08.999 UTC Sun May 10 2026 Number of operations attempted: 142 Number of operations skipped: 0 Current seconds left in Life: Forever Operational state of entry: Active Last time this entry was reset: Never Connection loss occurred: FALSE Timeout occurred: FALSE Over thresholds occurred: FALSE Latest RTT (milliseconds): 12 Latest operation start time: 03:24:41.999 UTC Sun May 10 2026 Latest operation return code: OK RTT Values: RTTAvg: 12 RTTMin: 8 RTTMax: 23 NumOfRTT: 1 RTTSum: 12 RTTSum2: 144 ``` "Reachability is Up" + 4 changes (a small number suggests stable history, not flapping) is the healthy steady state. If it flips up/down repeatedly, your probe target is too aggressive (frequency too low) or the target is itself flaky; pick a reliable target like a public anycast DNS or your ISP's loopback. ## Removing and Replacing Routes The remove syntax is `no route ...` with the same arguments as the add. Tab completion helps. Routes are removed instantly from the table; in-flight flows that depended on them break. Replacing a static route in place is two commands - remove the old, add the new. Always do the add first if you can: ``` ASA-PERIM(config)# route inside 10.10.10.0 255.255.255.0 10.10.0.2 1 ASA-PERIM(config)# no route inside 10.10.10.0 255.255.255.0 10.10.0.1 1 ``` That gives you a brief window where both routes are in the table (the ASA will deduplicate by next-hop) and then a clean removal of the old one. Doing it the other way - `no route` first, then `route` \- leaves a window where there is no route at all and traffic for that prefix gets `asp-drop`'d as `(no-route)`. ## Static Routes and Dynamic Routing Coexistence If the ASA also runs OSPF or BGP, statics and dynamics coexist by administrative distance. A static at distance 1 always beats an OSPF route at distance 110 for the same prefix. To redistribute a static into OSPF (so downstream OSPF neighbors learn it from the ASA), use `redistribute static` under the OSPF process. To advertise the static default into OSPF, use `default-information originate`. ``` ASA-PERIM(config)# router ospf 1 ASA-PERIM(config-router)# network 10.10.0.0 255.255.255.0 area 0 ASA-PERIM(config-router)# default-information originate ASA-PERIM(config-router)# redistribute static subnets ``` Most static-only deployments do not need either. The downstream layer-3 switch usually has its own static back to the ASA's inside IP, and dynamic protocols on a perimeter firewall are rare for security reasons. ## Diagnostic Commands `show route` Full RIB. Code in the first column tells you the source. `show route summary` Counts per source. Useful for "did I lose 200 statics?" `show route DEST` Best-match lookup for one destination. Confirms which route the ASA would use. `show route static` Just the static routes (no connected, no learned). `show running-config route` The configured static routes, in the order they appear in the config. `packet-tracer` Walks a packet through the pipeline including the ROUTE-LOOKUP phase, telling you what the ASA picked. For "I changed a route and now traffic is broken," [packet-tracer](https://www.pinglabz.com/cisco-asa-packet-tracer/) for the failing flow is the fastest path to the actual decision the ASA is making. The output names the route table entry it matched. ## Common Traps **Wrong egress interface.** The named interface in `route INTERFACE ...` must be the interface you can reach the next-hop on. `route outside 10.10.10.0 255.255.255.0 10.10.0.1` fails (next-hop is inside, not outside). The ASA accepts the command but the route is invalid; it shows up in `show route` with no neighbor and silently drops traffic for that prefix. **Floating distance too low.** If your floating backup is at distance 50 and you also run OSPF (distance 110), the floating beats OSPF and you get unexpected paths. Always pick a floating distance higher than every dynamic protocol the ASA might learn the same prefix from. Distance 200 is a safe choice for most deployments. **Forgetting to track the primary.** If the primary default and the floating backup both exist, but the primary has no track, only an interface-down event triggers failover. A blackhole upstream (link up, next-hop dead) is invisible. Always pair production multi-WAN with route tracking. **Tracking the wrong target.** Probing your ISP's first-hop only tests the first hop. If that is reachable but the rest of the internet is not, your track stays up. Probe a target deeper in the path (a public DNS root, your second-hop ISP router) to detect upstream failures, not just last-mile failures. ## Key Takeaways Static routing on the ASA covers most production needs: a default to the ISP, internal statics to downstream L3 switches, and floating or tracked routes for multi-WAN. Administrative distance picks the winner when two routes overlap; pair floating routes with IP SLA tracking to catch blackholes that interface state alone misses. `show route` with its code legend tells you where every route came from. [packet-tracer](https://www.pinglabz.com/cisco-asa-packet-tracer/) is the diagnostic when the routing decision itself is in question. Next: [Cisco ASA Management Plane Hardening](https://www.pinglabz.com/cisco-asa-management-hardening/) for SSH, AAA, banners, and the rest of the controls that turn a freshly bootstrapped ASA into something safe to expose to operations staff. Or for the broader picture, see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/) reading order. ### Cisco ASA Inside/Outside/DMZ Configuration Walkthrough URL: https://www.pinglabz.com/cisco-asa-inside-outside-dmz/ Last updated: 2026-06-13T20:08:23.000Z The classic Cisco ASA topology is three zones with three security levels: **inside** (your trusted users at level 100), **outside** (the internet at level 0), and a **DMZ** (your public-facing servers at level 50, somewhere in the middle). This article walks the entire build end to end on the lab's ASAv 9.23 - interfaces, addressing, security levels, NAT for outbound users, NAT for inbound web publishing, ACLs in both directions, and the show-output that proves traffic is actually flowing. The result is a small but realistic perimeter you could drop into a branch office tomorrow. This is one of the longer Fundamentals walkthroughs in the [Cisco ASA cluster](https://www.pinglabz.com/cisco-asa/). It assumes you have already [bootstrapped the ASA from CLI](https://www.pinglabz.com/cisco-asa-initial-setup/) and read the [security-levels primer](https://www.pinglabz.com/cisco-asa-security-levels/); everything else gets shown here. ## Topology and Address Plan inside InterfaceGigabitEthernet0/1 Subnet10.10.0.0/24 Security level100 Role User LAN, default gateway 10.10.0.254 (the ASA) dmz InterfaceGigabitEthernet0/2 Subnet192.168.50.0/24 Security level50 Role Public-facing servers; DMZ-WEB at 192.168.50.10 outside InterfaceGigabitEthernet0/0 Subnet203.0.113.0/30 Security level0 Role Point-to-point link to the ISP, ASA at .2, ISP at .1 NAT pool Interface(virtual) Subnet198.51.100.0/24 Security leveln/a Role Public block for static NAT (DMZ-WEB to .10) Three real interfaces, three zones, three security levels. The numeric levels (100/50/0) drive the implicit allow/deny matrix between zones - high to low is permitted by default, low to high requires an explicit ACL. ## Step 1: Interface Configuration ``` ASA-PERIM(config)# interface GigabitEthernet0/0 ASA-PERIM(config-if)# description Outside-to-ISP ASA-PERIM(config-if)# nameif outside ASA-PERIM(config-if)# security-level 0 ASA-PERIM(config-if)# ip address 203.0.113.2 255.255.255.252 ASA-PERIM(config-if)# no shutdown ! ASA-PERIM(config)# interface GigabitEthernet0/1 ASA-PERIM(config-if)# description Inside-to-LAN ASA-PERIM(config-if)# nameif inside ASA-PERIM(config-if)# security-level 100 ASA-PERIM(config-if)# ip address 10.10.0.254 255.255.255.0 ASA-PERIM(config-if)# no shutdown ! ASA-PERIM(config)# interface GigabitEthernet0/2 ASA-PERIM(config-if)# description DMZ-Servers ASA-PERIM(config-if)# nameif dmz ASA-PERIM(config-if)# security-level 50 ASA-PERIM(config-if)# ip address 192.168.50.1 255.255.255.0 ASA-PERIM(config-if)# no shutdown ``` Verify with `show interface ip brief`: ``` ASA-PERIM# show interface ip brief Interface IP-Address OK? Method Status Protocol GigabitEthernet0/0 203.0.113.2 YES manual up up GigabitEthernet0/1 10.10.0.254 YES manual up up GigabitEthernet0/2 192.168.50.1 YES manual up up Management0/0 unassigned YES unset administratively down down ``` Three data interfaces up. See [interfaces and VLAN trunks](https://www.pinglabz.com/cisco-asa-interfaces-vlan-trunks/) if any of these are not coming up the way you expect. ## Step 2: Default Route to the ISP ``` ASA-PERIM(config)# route outside 0.0.0.0 0.0.0.0 203.0.113.1 1 ``` Verify the routing table now has the connected subnets plus the default: ``` ASA-PERIM# show route Codes: L - local, C - connected, S - static, R - RIP, M - mobile, B - BGP D - EIGRP, EX - EIGRP external, O - OSPF, IA - OSPF inter area N1 - OSPF NSSA external type 1, N2 - OSPF NSSA external type 2 E1 - OSPF external type 1, E2 - OSPF external type 2, V - VPN i - IS-IS, su - IS-IS summary, L1 - IS-IS level-1, L2 - IS-IS level-2 ia - IS-IS inter area, * - candidate default, U - per-user static route o - ODR, P - periodic downloaded static route, + - replicated route SI - Static InterVRF Gateway of last resort is 203.0.113.1 to network 0.0.0.0 S* 0.0.0.0 0.0.0.0 [1/0] via 203.0.113.1, outside C 10.10.0.0 255.255.255.0 is directly connected, inside L 10.10.0.254 255.255.255.255 is directly connected, inside C 192.168.50.0 255.255.255.0 is directly connected, dmz L 192.168.50.1 255.255.255.255 is directly connected, dmz C 203.0.113.0 255.255.255.252 is directly connected, outside L 203.0.113.2 255.255.255.255 is directly connected, outside ``` The default route via outside (the `S*` line) is the gateway of last resort. Connected and local routes for each of the three zones are present. ## Step 3: Outbound PAT for Inside Users Inside users at 10.10.0.0/24 cannot reach the internet directly because their addresses are RFC1918\. The ASA needs to NAT them to the outside interface IP. The simplest pattern is Auto NAT on a `network` object: ``` ASA-PERIM(config)# object network INSIDE-NET ASA-PERIM(config-network-object)# subnet 10.10.0.0 255.255.255.0 ASA-PERIM(config-network-object)# nat (inside,outside) dynamic interface ``` Read it as: "anything sourced from `10.10.0.0/24` on inside, headed toward outside, gets PATed to the outside interface IP." That single rule covers all outbound user traffic. The [security-level 100 to 0](https://www.pinglabz.com/cisco-asa-security-levels/) rule lets the traffic pass without an explicit ACL. ## Step 4: Publish DMZ-WEB Inbound The DMZ web server lives at `192.168.50.10`. It needs a public IP (we will use `198.51.100.10`) and an inbound ACL because traffic flowing low-to-high (outside 0 to dmz 50) is denied by default. The static NAT, again with an Auto NAT pattern: ``` ASA-PERIM(config)# object network DMZ-WEB ASA-PERIM(config-network-object)# host 192.168.50.10 ASA-PERIM(config-network-object)# nat (dmz,outside) static 198.51.100.10 ``` The ACL that opens 80, 443, and ICMP echo: ``` ASA-PERIM(config)# access-list OUTSIDE_IN extended permit tcp any object DMZ-WEB eq www ASA-PERIM(config)# access-list OUTSIDE_IN extended permit tcp any object DMZ-WEB eq https ASA-PERIM(config)# access-list OUTSIDE_IN extended permit icmp any object DMZ-WEB ASA-PERIM(config)# access-list OUTSIDE_IN extended deny ip any any log informational interval 300 ASA-PERIM(config)# access-group OUTSIDE_IN in interface outside ``` Two important details. First, the ACL references `object DMZ-WEB`, which the ASA resolves to **192.168.50.10** (the real address) because that is what is in the object. The ASA evaluates ACLs against post-NAT real addresses on inbound, which is the modern (8.3+) behavior. Second, the explicit `deny ip any any log` at the bottom does the same thing as the implicit deny, but it generates a syslog every time it fires. That is the line we use to debug "why is traffic getting blocked" issues, as covered in [ACL troubleshooting](https://www.pinglabz.com/cisco-asa-acl-troubleshooting/). ## Step 5: ICMP Inspection (For Pings to Come Back) ICMP is connectionless, which means without inspection an outbound echo request leaves the firewall but the inbound echo reply gets dropped (no matching conn entry). The default global policy includes `inspect icmp`, which makes ICMP behave like a stateful protocol on the ASA. If ping does not work from inside out, this is one of the first things to check. ``` ASA-PERIM# show service-policy global | include icmp Inspect: icmp, packet 1244, lock fail 0, drop 0, reset-drop 0 ``` The packet counter climbing tells you ICMP inspection is firing. See [inspection engines](https://www.pinglabz.com/cisco-asa-inspection-engines/) for the full MPF model. ## Step 6: End-to-End Verification Now exercise the policy from both directions. ### Outbound: inside-host to 8.8.8.8 Generate a flow from `inside-host` on the inside subnet: ``` inside-host:~# ping -c 3 8.8.8.8 PING 8.8.8.8 (8.8.8.8) 56(84) bytes of data. 64 bytes from 8.8.8.8: icmp_seq=1 ttl=51 time=23.4 ms 64 bytes from 8.8.8.8: icmp_seq=2 ttl=51 time=24.1 ms 64 bytes from 8.8.8.8: icmp_seq=3 ttl=51 time=23.8 ms ``` And confirm on the ASA: ``` ASA-PERIM# show xlate count 14 in use, 16 most used ASA-PERIM# show xlate local 10.10.0.50 Flags: D - DNS, e - extended, I - identity, i - dynamic, r - portmap, s - static, T - twice, N - net-to-net ICMP PAT from inside:10.10.0.50 1/12345 to outside:203.0.113.2 1/45678 flags ri idle 0:00:01 timeout 0:00:30 ``` The outbound ping created an ICMP xlate entry. The inside source was PATed to the outside interface IP (`203.0.113.2`). When the echo reply comes back, the ASA matches it to the xlate, NATs the destination back to `10.10.0.50`, and forwards it. ### Inbound: 8.8.8.8 to DMZ-WEB on https From outside, hit the public IP: ``` outside-tester:~# curl -sI -k https://198.51.100.10 HTTP/1.1 200 OK Server: nginx/1.20.2 ``` And on the ASA: ``` ASA-PERIM# show conn address 192.168.50.10 14 in use, 16 most used TCP outside 8.8.8.8:51234 dmz 192.168.50.10:443, idle 0:00:01, bytes 1132, flags UIO ASA-PERIM# show access-list OUTSIDE_IN | include hitcnt access-list OUTSIDE_IN line 3 extended permit tcp any object DMZ-WEB eq https (hitcnt=2) ``` The conn entry shows the established TCP flow. The hit counter on line 3 of OUTSIDE\_IN ticks up. The `U` flag in the conn flags means "up," meaning the TCP three-way handshake completed - which it could only do because both directions of the flow (the SYN from outside, the SYN/ACK from DMZ-WEB) made it through the firewall. ### Cross-zone: DMZ to inside (Should Be Blocked) By default, traffic from a lower security level (DMZ at 50) to a higher one (inside at 100) is denied. Test it: ``` dmz-host:~# ping -c 1 10.10.0.50 PING 10.10.0.50 (10.10.0.50) 56(84) bytes of data. ^C --- 10.10.0.50 ping statistics --- 1 packets transmitted, 0 received, 100% packet loss ASA-PERIM# show asp drop frame | include acl-drop Flow is denied by configured rule (acl-drop) 1612 ``` The `acl-drop` counter ticks up because the implicit deny on the dmz interface (no ACL bound, default deny low-to-high) blocked the packet. [show asp drop counters](https://www.pinglabz.com/cisco-asa-asp-drop/) covers what each reason actually means. If you do need DMZ-to-inside reachability for a specific application (a backup server pulling from a database, for example), add an explicit ACL on the dmz interface that permits exactly that flow. ## The Full Working Config in One Block Everything from this article in a single paste-able config: ``` ! hostname ASA-PERIM domain-name pinglabz.lab enable password Cisco1@3 ! interface GigabitEthernet0/0 description Outside-to-ISP nameif outside security-level 0 ip address 203.0.113.2 255.255.255.252 no shutdown interface GigabitEthernet0/1 description Inside-to-LAN nameif inside security-level 100 ip address 10.10.0.254 255.255.255.0 no shutdown interface GigabitEthernet0/2 description DMZ-Servers nameif dmz security-level 50 ip address 192.168.50.1 255.255.255.0 no shutdown ! route outside 0.0.0.0 0.0.0.0 203.0.113.1 1 ! object network INSIDE-NET subnet 10.10.0.0 255.255.255.0 nat (inside,outside) dynamic interface ! object network DMZ-WEB host 192.168.50.10 nat (dmz,outside) static 198.51.100.10 ! access-list OUTSIDE_IN extended permit tcp any object DMZ-WEB eq www access-list OUTSIDE_IN extended permit tcp any object DMZ-WEB eq https access-list OUTSIDE_IN extended permit icmp any object DMZ-WEB access-list OUTSIDE_IN extended deny ip any any log informational interval 300 access-group OUTSIDE_IN in interface outside ! class-map inspection_default match default-inspection-traffic policy-map global_policy class inspection_default inspect icmp service-policy global_policy global ! end write memory ``` ## What This Does and Does Not Cover This baseline gets a working three-zone perimeter with outbound PAT and inbound DMZ publishing. It does not yet have: - Authentication for users (covered in [management hardening](https://www.pinglabz.com/cisco-asa-management-hardening/) and the [AAA](https://www.pinglabz.com/cisco-asa-aaa-for-vpn/) article). - VPN access (see the AnyConnect [SSL](https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/) and [IKEv2](https://www.pinglabz.com/cisco-asa-anyconnect-ikev2-vpn/) walkthroughs). - HA failover (see [active/standby](https://www.pinglabz.com/cisco-asa-active-standby-failover/)). - Logging beyond the implicit defaults (see [syslog](https://www.pinglabz.com/cisco-asa-syslog-logging/)). Each of those layers builds on the same baseline. Get the three zones, the routes, the NAT, and the ACLs right first; everything else assumes you have done so. ## Key Takeaways The classic inside/outside/DMZ ASA build is a small set of ordered steps: bring up three named interfaces with their security levels, add a default route, configure PAT for the inside subnet, add a static NAT for the DMZ server, and write an inbound ACL for the public-facing services. From there, `show xlate` tells you outbound NAT is firing, `show conn` tells you flows exist, and `show access-list` tells you which ACL line is matching. The implicit security-level matrix does most of the work: outbound (high to low) is permitted by default, inbound (low to high) requires an explicit ACL. Read the [security-levels article](https://www.pinglabz.com/cisco-asa-security-levels/) if any of that feels arbitrary; it sets up everything else. Next: [Cisco ASA Static Routing and Default Route Configuration](https://www.pinglabz.com/cisco-asa-static-routing/) covers tracked routes, route metrics, and multi-WAN failover. The full reading order is on the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). ### Cisco ASA Initial Setup from CLI URL: https://www.pinglabz.com/cisco-asa-initial-setup/ Last updated: 2026-05-29T23:40:47.000Z A fresh Cisco ASA out of the box is mostly empty. There is a default `enable` password, an unconfigured Management0/0 interface, and the cli setup wizard prompting you to answer hostname/timezone/firewall mode questions. Most engineers cancel the wizard, paste a real config, save, and move on. This article walks the day-0 commands that get a brand-new ASA (hardware or ASAv) from "factory" to "ready for the rest of the cluster's tasks": hostname, enable secret, interfaces, default route, SSH for management, NTP, basic syslog, and a write to startup. This is a [Cisco ASA](https://www.pinglabz.com/cisco-asa/) Fundamentals article. After running through these commands, the next steps are [building out the inside/outside/DMZ zones](https://www.pinglabz.com/cisco-asa-inside-outside-dmz/), configuring [static routes](https://www.pinglabz.com/cisco-asa-static-routing/), and tightening the [management plane](https://www.pinglabz.com/cisco-asa-management-hardening/). ## Cancel the Setup Wizard On first boot, the ASA console looks like this: ``` Pre-configure Firewall now through interactive prompts [yes]? no ``` Type `no`. The wizard is fine for someone who has never seen an ASA, but it asks 12 questions, sets defaults you will replace, and writes the result to startup. For any engineer with a target config in mind, it is faster to skip it. If the ASA already booted with the wizard's answers and you want to start fresh: ``` ciscoasa# write erase ciscoasa# reload ``` This wipes the startup config and reboots into a clean state. ## Enable Mode, Hostname, Domain The default enable password on a fresh ASA is empty (just press Enter). The first move is to set a real one and a hostname: ``` ciscoasa> enable Password: ciscoasa# configure terminal ciscoasa(config)# hostname ASA-PERIM ASA-PERIM(config)# domain-name pinglabz.lab ASA-PERIM(config)# enable password Cisco1@3 ``` For a real production ASA you would use a strong unique password and pair it with a hashed local user (covered in [management hardening](https://www.pinglabz.com/cisco-asa-management-hardening/)). For lab work, the CML default `enable password Cisco1@3` is what PyATS expects, which is why every ASA in the PingLabz lab uses it. The hostname propagates to every prompt and to the cert subject if you generate a self-signed identity certificate. Set it before any cert work. ## Bring Up the Data Interfaces Three interfaces, three roles. Everything else in the cluster builds on this baseline: ``` ASA-PERIM(config)# interface GigabitEthernet0/0 ASA-PERIM(config-if)# description Outside-to-ISP ASA-PERIM(config-if)# nameif outside ASA-PERIM(config-if)# security-level 0 ASA-PERIM(config-if)# ip address 203.0.113.2 255.255.255.252 ASA-PERIM(config-if)# no shutdown ! ASA-PERIM(config)# interface GigabitEthernet0/1 ASA-PERIM(config-if)# description Inside-to-LAN ASA-PERIM(config-if)# nameif inside ASA-PERIM(config-if)# security-level 100 ASA-PERIM(config-if)# ip address 10.10.0.254 255.255.255.0 ASA-PERIM(config-if)# no shutdown ! ASA-PERIM(config)# interface GigabitEthernet0/2 ASA-PERIM(config-if)# description DMZ ASA-PERIM(config-if)# nameif dmz ASA-PERIM(config-if)# security-level 50 ASA-PERIM(config-if)# ip address 192.168.50.1 255.255.255.0 ASA-PERIM(config-if)# no shutdown ``` Two reminders that bite IOS engineers: ASA interfaces ship in `shutdown`, so `no shutdown` is mandatory. Without `nameif`, NAT and ACL refuse to bind to the interface. See [interfaces and VLAN trunks](https://www.pinglabz.com/cisco-asa-interfaces-vlan-trunks/) for the full anatomy. ## Default Route to the ISP ``` ASA-PERIM(config)# route outside 0.0.0.0 0.0.0.0 203.0.113.1 1 ``` One static line. The trailing `1` is administrative distance (default for static; lower beats higher if you also have a learned default). For multi-WAN and tracked-route patterns, see [static routing](https://www.pinglabz.com/cisco-asa-static-routing/). Verify reachability: ``` ASA-PERIM# ping 203.0.113.1 Type escape sequence to abort. Sending 5, 100-byte ICMP Echos to 203.0.113.1, timeout is 2 seconds: !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 1/2/8 ms ``` Outside next-hop reachable. Now confirm the interface table looks right: ``` ASA-PERIM# show interface ip brief Interface IP-Address OK? Method Status Protocol GigabitEthernet0/0 203.0.113.2 YES manual up up GigabitEthernet0/1 10.10.0.254 YES manual up up GigabitEthernet0/2 192.168.50.1 YES manual up up Management0/0 unassigned YES unset administratively down down ``` Three data interfaces up. Management0/0 still down because we have not configured it yet (and most ASAs are managed in-band over the inside interface anyway). ## SSH for Management Telnet and SSH on the ASA are configured by the same pattern: a per-interface allow-list of source subnets. Skip telnet entirely on a real network; SSH only. ``` ASA-PERIM(config)# crypto key generate rsa modulus 2048 INFO: The name for the keys will be: Keypair generation process begin. Please wait... ! ASA-PERIM(config)# username admin password Cisc0_Admin! privilege 15 ASA-PERIM(config)# aaa authentication ssh console LOCAL ASA-PERIM(config)# ssh 10.10.0.0 255.255.255.0 inside ASA-PERIM(config)# ssh version 2 ASA-PERIM(config)# ssh timeout 30 ``` Walk through it: - `crypto key generate rsa` creates the host key SSH needs. 2048 modulus minimum; 4096 is fine if you do not mind the extra second on key gen. - `username admin ... privilege 15` creates a local privileged user. - `aaa authentication ssh console LOCAL` tells the ASA to authenticate SSH against the local user database. - `ssh 10.10.0.0 255.255.255.0 inside` permits SSH from the inside subnet only. Without an `ssh` permit line for an interface, no SSH can land there. This is the ASA's built-in SSH ACL. - `ssh version 2` disables SSHv1. The full hardening pass (key sizes, idle timeouts, AAA server-groups, mgmt ACLs, console banners) lives in [management hardening](https://www.pinglabz.com/cisco-asa-management-hardening/). The block above is the absolute minimum to get a remote shell. ## Time: Clock and NTP Wrong clocks break certificate validation, syslog correlation, and AAA. Set the timezone and at least one NTP source on day 0. ``` ASA-PERIM(config)# clock timezone UTC 0 ASA-PERIM(config)# ntp server 10.10.0.10 source inside prefer ASA-PERIM(config)# ntp server 10.10.0.11 source inside ``` UTC is the right answer for a network device unless you have a reason. If you must run local time, set `clock summer-time` for DST. Verify after a few minutes: ``` ASA-PERIM# show ntp associations address ref clock st when poll reach delay offset disp *~10.10.0.10 127.127.1.0 8 5 64 377 1.2 0.342 1.502 ~10.10.0.11 127.127.1.0 8 11 64 377 1.5 -0.118 1.221 * master (synced) ASA-PERIM# show clock 01:23:45.789 UTC Sun May 10 2026 ``` The asterisk on the first line says we are synced to that peer. Stratum (st) 8 is the upstream's stratum; the ASA itself sits at stratum 9 once synced. ## Basic Syslog Logging is so important that [it gets its own article](https://www.pinglabz.com/cisco-asa-syslog-logging/). The day-0 minimum is to enable logging, log to the buffer for ad-hoc reads, and ship to a remote syslog server for retention. ``` ASA-PERIM(config)# logging enable ASA-PERIM(config)# logging timestamp ASA-PERIM(config)# logging buffered informational ASA-PERIM(config)# logging host inside 10.10.0.30 17/514 ``` The four lines: turn logging on, prepend timestamps, keep the most recent buffer in RAM at severity informational and above, and send all messages to `10.10.0.30` over UDP/514\. After the change, `show logging` immediately shows the new buffer filling up. ## DNS for the ASA The ASA needs DNS for FQDN-based ACLs, AnyConnect FQDN-aware features, and any feature that needs to resolve external names (ntp by FQDN, license server, etc.). ``` ASA-PERIM(config)# dns domain-lookup inside ASA-PERIM(config)# dns server-group DefaultDNS ASA-PERIM(config-dns-server-group)# name-server 10.10.0.10 ASA-PERIM(config-dns-server-group)# name-server 10.10.0.11 ASA-PERIM(config-dns-server-group)# domain-name pinglabz.lab ``` The interface in `dns domain-lookup INTERFACE` tells the ASA which interface to source DNS queries from. If you also want internet-resolved DNS for outside-bound features, add `dns domain-lookup outside` with a public resolver in the server group. ## Minimum-Viable NAT + ACL for Outbound Traffic Inside hosts cannot reach the internet until you (a) PAT them to the outside interface and (b) the security-level model lets the traffic out. The PAT rule: ``` ASA-PERIM(config)# object network INSIDE-NET ASA-PERIM(config-network-object)# subnet 10.10.0.0 255.255.255.0 ASA-PERIM(config-network-object)# nat (inside,outside) dynamic interface ``` That single object PATs the entire inside subnet to the outside interface IP for any outbound flow. Because inside is security-level 100 and outside is 0, the implicit allow-from-higher-to-lower rule lets the flow out without an explicit ACL. [Security levels](https://www.pinglabz.com/cisco-asa-security-levels/) covers exactly when you need an explicit ACL and when the implicit rules are enough. For inbound DMZ-publishing, you would add a static NAT and an ACL (covered in [Static NAT for the DMZ](https://www.pinglabz.com/cisco-asa-static-nat-dmz/)). ## Default Inspection (Already On) Out of the box the ASA has a default global inspection policy that turns on inspection for ICMP, FTP, DNS, SIP, ESMTP, and a handful of other protocols. You do not need to configure anything for ICMP echo replies to come back through the firewall on day 0; the default `inspect icmp` is what makes that work. The full [MPF and inspection engines](https://www.pinglabz.com/cisco-asa-inspection-engines/) walkthrough covers customization. Verify the default policy is in place: ``` ASA-PERIM# show service-policy global Global policy: Service-policy: global_policy Class-map: inspection_default Inspect: dns preset_dns_map, packet 4218, lock fail 0, drop 0, reset-drop 0 Inspect: ftp, packet 0, lock fail 0, drop 0, reset-drop 0 Inspect: icmp, packet 1244, lock fail 0, drop 0, reset-drop 0 Inspect: sip, packet 0, lock fail 0, drop 0, reset-drop 0 ... ``` If `show service-policy global` returns nothing, someone removed the default policy. Restore it (or add the bits you actually need) per the inspection-engines article. ## Save the Config (Otherwise Reboots Lose It) The single most-forgotten command on the ASA. Running config is in RAM; startup config is in flash. The two diverge silently until you reboot and lose hours of work. ``` ASA-PERIM# write memory Building configuration... Cryptochecksum: 4e7f9c2a 1b3d8e6f a9c5b417 d2e8f039 [OK] ``` Or the equivalent `copy running-config startup-config`. Make this a reflex after every change. [Upgrade and backup](https://www.pinglabz.com/cisco-asa-upgrade-backup/) covers automating off-box backups so you do not lose the config to a flash failure. ## Post-Bootstrap Checklist By the end of this article, the ASA has: - A real hostname and enable secret. - Three named interfaces with security levels and IPs (inside/outside/dmz). - A default route to the upstream ISP. - SSH from the inside subnet, with a privileged local user. - Working NTP and a sane timezone. - Buffered + remote syslog. - DNS for the ASA itself. - Minimum-viable PAT for inside hosts to reach the internet. - Default inspection (ICMP, DNS, etc.) on, which means pings get replies. - Saved startup config. That is the floor. Everything else in the cluster (DMZ publishing, AnyConnect, failover, MPF tuning, multi-context) builds on top. ## Sanity Check Before You Walk Away Three commands that catch 90% of "I bootstrapped this ASA an hour ago and it does not work" calls: ``` ASA-PERIM# show interface ip brief ! All data interfaces "up/up"? ASA-PERIM# show route ! Default route present? ASA-PERIM# show xlate count ! At least one xlate after sending traffic from inside? ``` If the first two are clean and you see xlates appear after running a ping from inside-host to `8.8.8.8`, the day-0 build works. If not, run [packet-tracer](https://www.pinglabz.com/cisco-asa-packet-tracer/) for the failing flow and follow the phase output to the rule that is dropping the packet. ## Key Takeaways Day-0 ASA bootstrap is small but ordered. Hostname and enable secret first, then interfaces (do not forget `nameif` and `no shutdown`), then a default route, then SSH for remote access, then NTP, syslog, and DNS for the box itself. PAT for inside hosts and a saved config close out the build. The two most-skipped steps that cause future debugging pain are setting the timezone (always UTC unless you have a reason) and writing memory after the build. Make both reflexes. Next: [Cisco ASA Inside/Outside/DMZ Configuration Walkthrough](https://www.pinglabz.com/cisco-asa-inside-outside-dmz/) takes the same baseline and turns it into a publishing edge with a real DMZ web server, ACLs, and traffic flowing in both directions. The full reading order is on the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). ### Cisco ASA Object Groups: Network, Service, and Protocol URL: https://www.pinglabz.com/cisco-asa-object-groups/ Last updated: 2026-06-13T20:08:24.000Z Object groups are the difference between an ACL you can read and one you cannot. Without them, an ACL that permits HTTP, HTTPS, and SSH from three jump-host source addresses to four DMZ servers takes 36 lines. With them, the same intent collapses to one. They also unlock a NAT configuration model where an inside subnet, a DMZ host, and an outside translation pool are first-class named objects you can reference everywhere by name. This article covers the four object-group types on the Cisco ASA, how they nest, and the actual show output from the lab's running OUTSIDE\_IN ACL and NAT pool. Object groups belong to the [Cisco ASA Fundamentals](https://www.pinglabz.com/cisco-asa/) tier. Once they click, the rest of the platform feels less arbitrary because every [ACL](https://www.pinglabz.com/cisco-asa-acl-configuration/), [NAT rule](https://www.pinglabz.com/cisco-asa-nat-explained/), and policy-map starts referencing the same named building blocks. ## Objects vs Object Groups The ASA has two related but distinct ideas: **Object** Holds One thing: one host, one subnet, one range, one service port, one FQDN Created with `object network NAME` or `object service NAME` Referenced from ACLs, NAT rules, policy-maps **Object group** Holds A list of things, optionally a list of other groups Created with `object-group network NAME`, `object-group service NAME`, etc. Referenced from ACLs, NAT rules (some), policy-maps Single objects are useful for NAT (one host on the inside translates to one address on the outside) and as readable handles for one specific endpoint. Object groups are useful when you have collections (three jump-host source IPs, four DMZ servers, the standard "web ports" set of 80 and 443). ## The Four Object Group Types Each type wraps a specific kind of value. IPv4/IPv6 hosts, subnets, ranges, FQDNs, and other network groups Type `object-group network NAME` Used in ACL source/destination, NAT real/mapped, AAA target, route-map match TCP/UDP port lists (single, range), or unified protocol+port pairs (TCP/80 + UDP/53) Type `object-group service NAME` Used in ACL service field, NAT service translation, inspection class-maps IP protocol numbers/names (tcp, udp, icmp, gre, esp, 50) Type `object-group protocol NAME` Used inACL protocol field ICMP type codes (echo, echo-reply, time-exceeded, unreachable) Type `object-group icmp-type NAME` Used inACL ICMP type field Network and service groups are 95% of what you use. Protocol and ICMP-type groups are useful but situational. ## Network Object Groups A network group is a list of network endpoints. Each entry is one of: a host, a subnet, a range, an FQDN, or another network group (nesting). ``` ASA-PERIM(config)# object-group network ADMIN-JUMP-HOSTS ASA-PERIM(config-network-object-group)# description Permanent admin jump hosts ASA-PERIM(config-network-object-group)# network-object host 10.10.0.50 ASA-PERIM(config-network-object-group)# network-object host 10.10.0.51 ASA-PERIM(config-network-object-group)# network-object host 10.10.0.52 ASA-PERIM(config-network-object-group)# network-object 10.99.0.0 255.255.255.0 ASA-PERIM(config-network-object-group)# group-object PARTNER-MGMT ``` The last line is nesting: `PARTNER-MGMT` is presumably another `object-group network` defined elsewhere, and its members get folded into `ADMIN-JUMP-HOSTS`. Nesting is one level deep and resolves at compile time, so there is no runtime cost. FQDN entries are also supported and resolve at intervals (default every 60 seconds): ``` ASA-PERIM(config)# object-group network EXTERNAL-VENDORS ASA-PERIM(config-network-object-group)# network-object fqdn updates.vendor-a.com ASA-PERIM(config-network-object-group)# network-object fqdn licensing.vendor-b.com ``` FQDN-based ACLs are very useful for outbound-allowlists ("this server can only reach these specific cloud APIs") but require functioning DNS on the ASA, which means a working `dns server-group DefaultDNS` with reachable resolvers and a sane `dns name-server` list. ## Service Object Groups The modern syntax is **unified service groups**, which include the protocol in each entry: ``` ASA-PERIM(config)# object-group service WEB-PORTS ASA-PERIM(config-service-object-group)# service-object tcp eq 80 ASA-PERIM(config-service-object-group)# service-object tcp eq 443 ASA-PERIM(config-service-object-group)# service-object tcp eq 8080 ASA-PERIM(config-service-object-group)# service-object tcp eq 8443 ``` Used in an ACL: ``` access-list OUTSIDE_IN extended permit object-group WEB-PORTS any object-group DMZ-WEB-FARM ``` That single ACL line expands internally to four (one per port) and the ASA reports each line in `show access-list` with its own hit counter. The legacy syntax is **protocol-locked service groups**, where the protocol is declared at the top: ``` ASA-PERIM(config)# object-group service WEB-TCP-LEGACY tcp ASA-PERIM(config-service-object-group)# port-object eq 80 ASA-PERIM(config-service-object-group)# port-object eq 443 ASA-PERIM(config-service-object-group)# port-object range 8000 8100 ``` Both forms still work. The unified syntax is more flexible (you can mix TCP and UDP in one group) and is what the configuration generator uses by default for new groups created via ASDM. Use the unified form for new work. ## Protocol and ICMP-Type Groups Less common but useful. Protocol groups simplify rules that need to allow several IP protocols between the same source and destination: ``` ASA-PERIM(config)# object-group protocol IPSEC-PROTOS ASA-PERIM(config-protocol-object-group)# protocol-object esp ASA-PERIM(config-protocol-object-group)# protocol-object ah ASA-PERIM(config-protocol-object-group)# protocol-object udp ASA-PERIM(config-protocol-object-group)# protocol-object 50 ``` ICMP-type groups are how you write ACLs that allow only specific ICMP message types (the safe-by-default approach instead of `permit icmp any any`): ``` ASA-PERIM(config)# object-group icmp-type SAFE-ICMP ASA-PERIM(config-icmp-object-group)# icmp-object echo-reply ASA-PERIM(config-icmp-object-group)# icmp-object time-exceeded ASA-PERIM(config-icmp-object-group)# icmp-object unreachable ``` Used in: ``` access-list OUTSIDE_IN extended permit icmp any any object-group SAFE-ICMP ``` This permits the diagnostic ICMP types that traceroute and PMTU discovery need, without permitting the rest. A nice middle ground between "all ICMP" and "no ICMP." ## Real Lab: How Objects and Groups Wire Together in an ACL From the lab's ASA-PERIM, this is the running OUTSIDE\_IN ACL after the changes from Sessions 1 and 2: ``` ASA-PERIM# show access-list OUTSIDE_IN | begin OUTSIDE_IN access-list OUTSIDE_IN; 5 elements; name hash: 0xe01d8199 access-list OUTSIDE_IN line 1 extended deny tcp host 198.51.100.99 any (hitcnt=1) (Last Hit=00:01:09 UTC May 10 2026) access-list OUTSIDE_IN line 2 extended permit tcp any object DMZ-WEB eq www (hitcnt=1) (Last Hit=23:27:41 UTC May 9 2026) access-list OUTSIDE_IN line 3 extended permit tcp any object DMZ-WEB eq https (hitcnt=2) (Last Hit=00:01:09 UTC May 10 2026) access-list OUTSIDE_IN line 4 extended permit icmp any object DMZ-WEB (hitcnt=2) (Last Hit=01:55:20 UTC May 10 2026) access-list OUTSIDE_IN line 5 extended deny ip any any log informational interval 300 (hitcnt=59) (Last Hit=01:55:47 UTC May 10 2026) ``` Notice that `object DMZ-WEB` appears as the destination on lines 2, 3, and 4\. `DMZ-WEB` is a single network object (not a group), defined elsewhere as: ``` ASA-PERIM(config)# show running-config object id DMZ-WEB object network DMZ-WEB host 192.168.50.10 nat (dmz,outside) static 198.51.100.10 ``` One object definition does double duty: it gives the ACL a readable name for the DMZ web server, AND defines the static NAT that publishes 192.168.50.10 to the outside as 198.51.100.10\. Change the IP in one place and both the ACL and the NAT rule follow. ## Real Lab: Auto NAT With Network Objects Auto NAT (also called Object NAT) is configured inside an `object network` definition. The Section 2 NAT block from `show nat detail`: ``` ASA-PERIM# show nat detail | begin Auto NAT Auto NAT Policies (Section 2) 1 (dmz) to (outside) source static DMZ-WEB 198.51.100.10 translate_hits = 87, untranslate_hits = 134 Source - Origin: 192.168.50.10/32, Translated: 198.51.100.10/32 2 (inside) to (outside) source dynamic INSIDE-TRANSIT interface translate_hits = 0, untranslate_hits = 0 Source - Origin: 10.10.0.0/24, Translated: 203.0.113.2/30 3 (inside) to (outside) source dynamic INSIDE-NET interface translate_hits = 4521, untranslate_hits = 0 Source - Origin: 10.10.10.0/24, Translated: 203.0.113.2/30 ``` Three rules, all defined inside three `object network` blocks (DMZ-WEB, INSIDE-TRANSIT, INSIDE-NET). Auto NAT auto-orders by specificity, so the /32 host (DMZ-WEB) sits at line 1 even though it was configured later. INSIDE-NET racked up 4,521 forward translations (real-world inside hosts hitting the internet via PAT to the outside interface). DMZ-WEB has more reverse than forward hits because it is a public-facing server (more inbound flows opened from the internet than outbound flows initiated by the server itself). ## The Built-In Service Objects The ASA ships with a long list of named ports you can use as keywords without defining your own. `www`, `https`, `ssh`, `smtp`, `domain`, `snmp`, `tftp`, `ntp`, and dozens more all parse as themselves. `access-list OUTSIDE_IN extended permit tcp any object DMZ-WEB eq www` uses `www` as a built-in alias for port 80; you do not need to define a `WWW` object. For ports the ASA does not have a name for, use the number: `eq 8443`, `range 5060 5061`, `gt 1024`. The full list of built-in names is in the `?` output for `access-list ... eq`. ## Object Groups Make ACLs Readable Six Months Later Compare two ways of writing the same intent. Without object groups: ``` access-list OUTSIDE_IN extended permit tcp host 10.10.0.50 host 192.168.50.10 eq 80 access-list OUTSIDE_IN extended permit tcp host 10.10.0.50 host 192.168.50.10 eq 443 access-list OUTSIDE_IN extended permit tcp host 10.10.0.50 host 192.168.50.10 eq 22 access-list OUTSIDE_IN extended permit tcp host 10.10.0.50 host 192.168.50.20 eq 80 access-list OUTSIDE_IN extended permit tcp host 10.10.0.50 host 192.168.50.20 eq 443 access-list OUTSIDE_IN extended permit tcp host 10.10.0.50 host 192.168.50.20 eq 22 access-list OUTSIDE_IN extended permit tcp host 10.10.0.51 host 192.168.50.10 eq 80 ... [repeat for each src x dst x port] ``` With object groups: ``` object-group network ADMIN-JUMP network-object host 10.10.0.50 network-object host 10.10.0.51 network-object host 10.10.0.52 object-group network DMZ-SERVERS network-object host 192.168.50.10 network-object host 192.168.50.20 object-group service ADMIN-PORTS service-object tcp eq 22 service-object tcp eq 80 service-object tcp eq 443 ! access-list OUTSIDE_IN extended permit object-group ADMIN-PORTS object-group ADMIN-JUMP object-group DMZ-SERVERS ``` One ACL line, with the intent ("admins reach DMZ servers on admin ports") visible at first read. Adding a fourth jump host means one new `network-object` line; the ACL itself never changes. This is also the version that survives an audit: the auditor looks at OUTSIDE\_IN and sees one rule, not 18. ## Verify Object Groups The ASA expands object groups internally. `show access-list` shows you both the compact form and the expanded form: ``` ASA-PERIM# show access-list OUTSIDE_IN | begin OUTSIDE_IN access-list OUTSIDE_IN; 9 elements; name hash: 0xe01d8199 alert-interval 300 access-list OUTSIDE_IN line 1 extended permit object-group ADMIN-PORTS object-group ADMIN-JUMP object-group DMZ-SERVERS (hitcnt=44) access-list OUTSIDE_IN line 1 extended permit tcp host 10.10.0.50 host 192.168.50.10 eq 22 (hitcnt=12) access-list OUTSIDE_IN line 1 extended permit tcp host 10.10.0.50 host 192.168.50.10 eq 80 (hitcnt=8) access-list OUTSIDE_IN line 1 extended permit tcp host 10.10.0.50 host 192.168.50.10 eq 443 (hitcnt=10) ... [continues for each src x dst x port] ``` The compact form (the first line at "line 1") shows the ACL the way you wrote it. The expanded entries beneath show every src/dst/port permutation with its individual hit counter, which is how you tell whether a specific permutation is being used. Inspect a single group with: ``` ASA-PERIM# show running-config object-group id ADMIN-JUMP object-group network ADMIN-JUMP description Permanent admin jump hosts network-object host 10.10.0.50 network-object host 10.10.0.51 network-object host 10.10.0.52 ``` Or list every group with `show running-config object-group` (no id). ## Renaming and Removing Objects Safely Renaming an object is not a single command. The ASA's `rename object` works at the object level but not for object groups; the safe procedure for either is: create the new name, switch references, then delete the old. Removing a group is rejected if anything still references it: ``` ASA-PERIM(config)# no object-group network ADMIN-JUMP ERROR: removing object-group (ADMIN-JUMP) not allowed, it is being used. ``` To find every reference, ironically use the same `show running-config | include ADMIN-JUMP` grep, which finds the ACL and NAT lines that point at it. Update those first, then the delete succeeds. ## Key Takeaways Network and service object groups collapse repetitive ACLs into single readable lines and let one named object do double duty for both ACL and NAT. The four group types (network, service, protocol, icmp-type) cover almost everything you need; nesting is supported one level. Use the unified service-object syntax for new work because it lets you mix TCP and UDP in one group. The lab's OUTSIDE\_IN ACL and Auto NAT pool both reference the same DMZ-WEB object: change the IP in one place, both follow. That single principle is what makes object-driven ASA configurations maintainable years after the engineer who built them has left. Next in this cluster: [Cisco ASA Initial Setup from CLI](https://www.pinglabz.com/cisco-asa-initial-setup/), the day-0 bootstrap that gets a fresh ASA from "out of the box" to "passing traffic." For the broader picture, the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/) has the full reading order. For deeper ACL and NAT walkthroughs, see [ACL configuration](https://www.pinglabz.com/cisco-asa-acl-configuration/) and [NAT explained](https://www.pinglabz.com/cisco-asa-nat-explained/). ### Cisco ASA Interfaces, Subinterfaces, and VLAN Trunks URL: https://www.pinglabz.com/cisco-asa-interfaces-vlan-trunks/ Last updated: 2026-06-13T20:08:24.000Z The interface configuration on a Cisco ASA looks superficially like an IOS router, but two things are different and both matter. First, every data interface needs a `nameif` and a `security-level` before it can pass traffic. Second, the ASA does not run DTP, so an 802.1Q trunk to a switch is configured purely with subinterfaces and a tag-per-subinterface model. This article covers physical interfaces, subinterfaces, VLAN trunks to a Catalyst, and the configuration commands that quietly catch people coming from an IOS background. This walkthrough uses the same lab as the rest of the cluster (see [Cisco ASA: The Complete Reference](https://www.pinglabz.com/cisco-asa/)). For a complete inside/outside/DMZ build that uses these interface commands end-to-end, see [the inside/outside/DMZ walkthrough](https://www.pinglabz.com/cisco-asa-inside-outside-dmz/). ## Anatomy of an ASA Interface Every ASA data interface needs four things to be useful: Logical name Command`nameif inside` What it controls The name every other feature (NAT, ACL, route, log) refers to. Without nameif, no feature can bind to the interface. Security level Command`security-level 100` What it controls 0-100 trust score. Drives implicit ACL behavior between zones. See [security levels](https://www.pinglabz.com/cisco-asa-security-levels/). IP address Command `ip address 10.10.0.254 255.255.255.0` What it controls L3 address (in routed mode). Standby IP added if failover is configured. Operational state Command`no shutdown` What it controls Brings the interface up. ASA interfaces ship in `shutdown` by default; IOS does the opposite. Skip `nameif` and the interface might be physically up but the ASA refuses to use it for anything. Skip the security-level and the implicit allow/deny rules between this interface and others become unpredictable. Skip `no shutdown` and the interface stays administratively down, which catches everyone who came from IOS expecting interfaces to be up by default. ## Basic Physical Interface (Routed Mode) Here is the minimum-viable configuration for a routed inside interface, taken from the lab's ASA-PERIM: ``` ASA-PERIM(config)# interface GigabitEthernet0/1 ASA-PERIM(config-if)# description LAN-to-INSIDE-RTR ASA-PERIM(config-if)# nameif inside INFO: Security level for "inside" set to 100 by default. ASA-PERIM(config-if)# ip address 10.10.0.254 255.255.255.0 ASA-PERIM(config-if)# no shutdown ``` Notice that `nameif inside` automatically set `security-level 100`. The ASA assigns level 100 to any interface named "inside" and level 0 to any interface named "outside" as a convenience. You can override either with an explicit `security-level` command. Verify with the standard `show interface ip brief`: ``` ASA-PERIM# show interface ip brief Interface IP-Address OK? Method Status Protocol GigabitEthernet0/0 203.0.113.2 YES manual up up GigabitEthernet0/1 10.10.0.254 YES manual up up GigabitEthernet0/2 192.168.50.1 YES manual up up Management0/0 unassigned YES unset administratively down down ``` The "Status" column is the line protocol (cabling and physical Layer 1/2). The "Protocol" column is whether the ASA logically considers the interface usable. Both must be "up" for traffic to flow. ## The Management Interface Management0/0 (or whatever the platform calls it: `Management0`, `management0/0`, etc.) is special. It defaults to `management-only`, which means it cannot pass through-traffic. Only traffic destined to or sourced from the ASA itself (SSH, SNMP, syslog, AAA, NTP, AnyConnect on management) uses it. ``` ASA-PERIM(config)# interface Management0/0 ASA-PERIM(config-if)# nameif management INFO: Security level for "management" set to 0 by default. ASA-PERIM(config-if)# security-level 100 ASA-PERIM(config-if)# ip address 10.10.99.1 255.255.255.0 ASA-PERIM(config-if)# management-only ASA-PERIM(config-if)# no shutdown ``` Two ASAv-specific quirks worth knowing about Management0/0 because we documented them in the failover articles: - The ASAv platform refuses to use Management0/0 as the failover LAN interface (`failover lan interface FAIL-LINK Management0/0` returns `Management interface cannot be configured as failover on this platform`). Hardware ASA does not have this restriction. See [Active/Standby failover](https://www.pinglabz.com/cisco-asa-active-standby-failover/). - Management0/0 IS allowed as the stateful failover link on ASAv (`failover link FAIL-LINK Management0/0`), but you give up the ability to use it for management once you do. ## Subinterfaces for 802.1Q Trunks The ASA does not run DTP. There is no dynamic trunk negotiation. Instead, you create **subinterfaces**, each one tied to a single VLAN tag. The physical interface remains untagged (no `nameif`, no IP); each subinterface gets a name, security level, and IP. This is the equivalent of a router-on-a-stick configuration, with the same physical wire-up: the ASA interface connects to a switch port configured as a trunk. ### ASA Side ``` ASA-PERIM(config)# interface GigabitEthernet0/3 ASA-PERIM(config-if)# no shutdown ! No nameif, no IP. Parent interface is just a physical link. ! ASA-PERIM(config)# interface GigabitEthernet0/3.10 ASA-PERIM(config-subif)# vlan 10 ASA-PERIM(config-subif)# nameif user-vlan ASA-PERIM(config-subif)# security-level 80 ASA-PERIM(config-subif)# ip address 10.20.10.1 255.255.255.0 ! ASA-PERIM(config)# interface GigabitEthernet0/3.20 ASA-PERIM(config-subif)# vlan 20 ASA-PERIM(config-subif)# nameif voice-vlan ASA-PERIM(config-subif)# security-level 80 ASA-PERIM(config-subif)# ip address 10.20.20.1 255.255.255.0 ! ASA-PERIM(config)# interface GigabitEthernet0/3.30 ASA-PERIM(config-subif)# vlan 30 ASA-PERIM(config-subif)# nameif guest-vlan ASA-PERIM(config-subif)# security-level 50 ASA-PERIM(config-subif)# ip address 10.20.30.1 255.255.255.0 ``` The numbering convention is `parent.tag`. Using `Gi0/3.10` for VLAN 10 is not required (you could call it `Gi0/3.99` and still tag VLAN 10), but matching the subinterface number to the VLAN tag is a strong operational convention. Future-you will thank present-you. ### Catalyst Side The matching switch port: ``` SW1(config)# interface GigabitEthernet1/0/24 SW1(config-if)# description trunk-to-ASA-PERIM-Gi0/3 SW1(config-if)# switchport SW1(config-if)# switchport mode trunk SW1(config-if)# switchport trunk allowed vlan 10,20,30 SW1(config-if)# switchport trunk encapsulation dot1q SW1(config-if)# switchport nonegotiate SW1(config-if)# spanning-tree portfast trunk ``` Why `switchport nonegotiate`? Because the ASA does not speak DTP. A Cisco switch left in `switchport mode trunk` still sends DTP frames trying to negotiate; `nonegotiate` turns that off. Functionally the trunk works either way, but you avoid the trickle of "DTP frame received with bad value" log lines and the slow-converge edge cases. ### Native VLAN The ASA can carry the native (untagged) VLAN as well, but it has to be configured explicitly: ``` ASA-PERIM(config)# interface GigabitEthernet0/3.99 ASA-PERIM(config-subif)# vlan 99 native ASA-PERIM(config-subif)# nameif native-vlan ASA-PERIM(config-subif)# security-level 100 ASA-PERIM(config-subif)# ip address 10.20.99.1 255.255.255.0 ``` The `native` keyword tells the ASA to send and accept this VLAN's frames untagged. The switch side must agree (`switchport trunk native vlan 99`). If the two sides disagree on the native VLAN, the trunk still passes most traffic but has spectacular asymmetric forwarding for whatever VLAN ID is mismatched. Always set the native VLAN explicitly on both sides. ## IP Addressing Options Most interfaces get a static IP, but the ASA supports DHCP and PPPoE on the outside interface for situations where the upstream provider hands you an address dynamically. Static Command `ip address 10.10.0.254 255.255.255.0` Use case The default. Every inside, DMZ, or known-static outside. Static + standby Command `ip address 10.10.0.254 255.255.255.0 standby 10.10.0.253` Use case Active/standby failover. Standby unit takes the standby IP. See [failover](https://www.pinglabz.com/cisco-asa-active-standby-failover/). DHCP Command `ip address dhcp setroute` Use case Outside interface on a small business or branch ASA where the ISP hands out the IP. PPPoE Command `pppoe client vpdn group ISP` \+ `ip address pppoe setroute` Use case DSL or fiber providers that require PPPoE for authentication. The `setroute` keyword installs the default gateway learned from DHCP or PPPoE into the route table; without it, you get the IP but no default route, which is usually not what you want. ## Speed, Duplex, and MTU Defaults are usually right (`speed auto`, `duplex auto`, MTU 1500). Touch them only when you have a reason. ``` ASA-PERIM(config)# interface GigabitEthernet0/0 ASA-PERIM(config-if)# speed 1000 ASA-PERIM(config-if)# duplex full ASA-PERIM(config-if)# mtu outside 1500 ``` One subtlety: `mtu` on the ASA is configured per-nameif, not per-interface. If you have a subinterface trunk with multiple nameifs, you set MTU once per nameif. Most environments leave the default 1500 alone unless they are running jumbo frames internally (`jumbo-frame reservation` is a separate global command that requires a reload to take effect). ## Redundant Interfaces and EtherChannel For physical resiliency below the failover layer, the ASA supports two grouping models: - **Redundant interface**: an active/standby pair of physical interfaces presented as a single logical interface. Configured with `interface Redundant N` and `member-interface` commands. The ASA tracks the active member and switches to the standby on failure. Simple, no LACP. - **EtherChannel (port-channel)**: LACP or static aggregation of multiple physical interfaces into a single logical interface with multiplied bandwidth. Configured with `interface Port-channel N` and `channel-group N mode {active|on}` on member interfaces. This is what most modern deployments use. ``` ASA-PERIM(config)# interface Port-channel1 ASA-PERIM(config-if)# nameif inside ASA-PERIM(config-if)# security-level 100 ASA-PERIM(config-if)# ip address 10.10.0.254 255.255.255.0 ASA-PERIM(config-if)# no shutdown ! ASA-PERIM(config)# interface GigabitEthernet0/4 ASA-PERIM(config-if)# channel-group 1 mode active ASA-PERIM(config-if)# no shutdown ASA-PERIM(config)# interface GigabitEthernet0/5 ASA-PERIM(config-if)# channel-group 1 mode active ASA-PERIM(config-if)# no shutdown ``` The matching switch side runs the same LACP, with both members in `channel-group N mode active`. The ASA defaults to LACP fast (`lacp port-priority`) which converges quickly on member failure. ## Verification: What "Up" Actually Looks Like `show interface ip brief` is the quick health check. `show interface` on a specific interface gives you everything: counters, last input/output, drops, error-disabled state, and the auto-negotiated speed/duplex. From the lab: ``` ASA-PERIM# show interface GigabitEthernet0/1 | include is up|line protocol|address|MTU|input rate|output rate Interface GigabitEthernet0/1 "inside", is up, line protocol is up Hardware is i82540EM rev03, BW 1000 Mbps, DLY 10 usec MAC address aabb.cc00.4c10, MTU 1500 IP address 10.10.0.254, subnet mask 255.255.255.0 2 minute input rate 84 pkts/sec, 19872 bits/sec 2 minute output rate 79 pkts/sec, 47104 bits/sec ``` For a subinterface, the parent interface needs to be up before the subinterface can be up. If `Gi0/3` is shutdown, every subinterface on it is also down regardless of its own admin state. Always check the parent first. For VLAN trunk verification, a helpful one-liner: ``` ASA-PERIM# show interface | include vlan|nameif|line protocol Interface GigabitEthernet0/3.10 "user-vlan", is up, line protocol is up VLAN identifier 10 Interface GigabitEthernet0/3.20 "voice-vlan", is up, line protocol is up VLAN identifier 20 Interface GigabitEthernet0/3.30 "guest-vlan", is up, line protocol is up VLAN identifier 30 ``` If a subinterface shows "line protocol is down" on a parent that is up, the most common cause is a VLAN mismatch on the switch side (the trunk is not allowing the tag). ## Common Traps **Forgetting nameif.** Without `nameif`, NAT, ACL, route, and almost every other config refuses to bind. The interface looks fine to `show interface` but is invisible to the policy plane. **Default-shutdown.** Coming from IOS, you expect interfaces to be up after configuring an IP. ASA interfaces stay shutdown until you say `no shutdown`. Configuring an interface and forgetting `no shutdown` is the single most common "why won't this work" moment for new ASA admins. **Same security-level traffic.** Two interfaces at the same security level cannot pass traffic between each other unless you explicitly allow it with `same-security-traffic permit inter-interface`. See the [security levels](https://www.pinglabz.com/cisco-asa-security-levels/) article for the full rule set. **VLAN tag conflict in subinterfaces.** Two subinterfaces on the same physical interface cannot share a VLAN tag. The CLI rejects the second one with `%Error: VLAN already in use`. Worth knowing because the error happens at `vlan` command time, after you have already created the subinterface. **Native VLAN mismatch.** Worst kind of bug because half the traffic works. Always set the native VLAN explicitly on both sides of every trunk. ## Key Takeaways Every ASA data interface needs `nameif`, a `security-level`, an IP address, and `no shutdown` before it can pass traffic. Subinterfaces handle 802.1Q trunks (one tag per subinterface, each with its own nameif and security level) because the ASA does not run DTP. Match subinterface numbers to VLAN tags as a convention, and always set the native VLAN explicitly on both ends. For the next step in this cluster, see [Cisco ASA Object Groups](https://www.pinglabz.com/cisco-asa-object-groups/), which is what makes ACLs and NAT readable once you have multiple interfaces and multiple internal subnets to govern. The full reading order is on the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). ### Cisco ASA Routed Mode vs Transparent Mode URL: https://www.pinglabz.com/cisco-asa-routed-vs-transparent/ Last updated: 2026-06-13T20:08:25.000Z Cisco ASA can run in two firewall modes: **routed mode**, where the firewall is a Layer 3 hop with its own IP address on each interface, and **transparent mode**, where the firewall behaves like a Layer 2 bridge with a single management IP and no routing of its own. Most ASAs in the field run routed mode, because that is what the default configuration ships with and what every walkthrough on this site assumes. Transparent mode is a deliberate choice that solves a specific class of problem: dropping a firewall into an existing subnet without re-IP'ing anything. This article covers both modes, the trade-offs, the configuration commands, and the things that quietly break when you flip the global mode flag. If you arrived here from the cluster index, see [Cisco ASA: The Complete Reference](https://www.pinglabz.com/cisco-asa/) for the full reading order. If you are coming back to it later, the short version is: pick routed unless you have a hard requirement that forces transparent, and never flip the mode flag on a production firewall without a maintenance window. ## What Firewall Mode Actually Controls Firewall mode is a global configuration switch. It changes the entire forwarding model of the box. In routed mode, every data interface needs an IP address and a security level, and the ASA participates in routing decisions like any other Layer 3 device. The ASA can run static routes, OSPF, EIGRP, or BGP, advertise its connected interfaces, and act as the next hop for the hosts behind it. In transparent mode, the ASA does not own IP addresses on data interfaces. Instead, two or more interfaces are paired into a **bridge group** with a single shared **BVI** (Bridge Virtual Interface) IP that exists only for management traffic sourced from the ASA itself (syslog, SNMP, AAA, NTP). Hosts on the inside and outside interfaces stay in the same subnet; their default gateway remains the upstream router; the ASA simply inspects the frames as they pass through. That fundamental difference cascades through every feature on the platform. NAT, ACLs, inspection engines, and packet capture all work in both modes, but routing protocols, dynamic VPN client pools, and a handful of inspection engines are routed-mode-only. The reverse is not true: anything that works in transparent mode also works in routed mode. ## When Each Mode Fits The decision is almost always settled by the surrounding network, not by the firewall itself. Greenfield internet edge with NAT Mode that fitsRouted Why NAT to a public pool, default route to ISP, BGP or static. All routed-mode bread and butter. Site-to-site VPN concentrator Mode that fitsRouted Why Crypto map and VTI both terminate on a routed interface with its own IP. AnyConnect remote-access VPN Mode that fitsRouted Why VPN-POOL handed out to clients lives on a routed interface and needs return routing on the inside. Drop-in firewall on an existing subnet (no re-IP) Mode that fitsTransparent Why The hosts and their gateway stay where they are; the ASA inspects the frames between them. "Bump in the wire" for a database tier you cannot disturb Mode that fitsTransparent Why Same VLAN on both sides, the ASA is invisible to the application. Compliance overlay (PCI segmentation, HIPAA enclave) Mode that fitsTransparent Why Auditor wants stateful inspection between two existing subnets without changing the topology. Multi-context with shared interfaces in service-provider style Mode that fitsRouted Why Each context gets its own IP per interface; transparent contexts cannot share an interface. If the question is "can I make it work in transparent mode," the answer is usually yes; if the question is "should I," the answer is usually no, unless the IP-preservation requirement is real and immovable. ## Check Which Mode You Are In Today Before you change anything, confirm the current mode. ``` ASA-PERIM# show firewall Firewall mode: Router ASA-PERIM# show mode Security context mode: single ``` Two separate switches, two separate `show` commands. `show firewall` tells you routed vs transparent. `show mode` tells you single vs multiple context. The two are independent: you can have a single-context routed firewall, a single-context transparent firewall, a multi-context routed firewall, or a multi-context transparent firewall (with limits on which features each combination supports). ## Routed Mode Config: Quick Reference This is the default and what most walkthroughs on the site (including [Cisco ASA Inside/Outside/DMZ Configuration Walkthrough](https://www.pinglabz.com/cisco-asa-inside-outside-dmz/)) assume. Each data interface gets a name, a security level, and an IP. ``` ASA-PERIM(config)# interface GigabitEthernet0/0 ASA-PERIM(config-if)# nameif outside ASA-PERIM(config-if)# security-level 0 ASA-PERIM(config-if)# ip address 203.0.113.2 255.255.255.252 ASA-PERIM(config-if)# no shutdown ASA-PERIM(config-if)# interface GigabitEthernet0/1 ASA-PERIM(config-if)# nameif inside ASA-PERIM(config-if)# security-level 100 ASA-PERIM(config-if)# ip address 10.10.0.254 255.255.255.0 ASA-PERIM(config-if)# no shutdown ASA-PERIM(config)# route outside 0.0.0.0 0.0.0.0 203.0.113.1 1 ``` From there, NAT, ACLs, and VPN are all standard. See [Cisco ASA NAT Explained](https://www.pinglabz.com/cisco-asa-nat-explained/) and [Cisco ASA ACL Configuration](https://www.pinglabz.com/cisco-asa-acl-configuration/) for the next steps. ## Switching to Transparent Mode (Destructive) This is the command that scares people, and rightly so: ``` ASA-PERIM(config)# firewall transparent WARNING: This action will clear all configuration commands except access list, capture, established, file, ip address management-only, ipv6 address management-only, mac-address, mtu, nameif, nat-pool, network object, object-group, password-policy, pim, route-map, route, snmp, ssh, telnet, vlan, and the ones used to manage features such as logging, snmp, ntp, etc. ``` "Will clear all configuration commands" is not hyperbole. The ASA dumps NAT rules, dynamic routing, VPN configuration, and most service-policy maps. Saved-config you have on a TFTP server is not portable to transparent mode without rework. Run `write erase`, switch the mode, then load the transparent-style config. Reverse it with `no firewall transparent`. Same destructive behavior in the other direction. ## Transparent Mode Config: Bridge Group + BVI In transparent mode, two or more data interfaces are joined into a bridge group. The BVI (Bridge Virtual Interface) carries a single management IP that the ASA uses to source its own traffic and for SSH/HTTPS access. ``` ASA-EDGE(config)# firewall transparent ASA-EDGE(config)# interface GigabitEthernet0/0 ASA-EDGE(config-if)# nameif outside ASA-EDGE(config-if)# security-level 0 ASA-EDGE(config-if)# bridge-group 1 ASA-EDGE(config-if)# no shutdown ASA-EDGE(config-if)# interface GigabitEthernet0/1 ASA-EDGE(config-if)# nameif inside ASA-EDGE(config-if)# security-level 100 ASA-EDGE(config-if)# bridge-group 1 ASA-EDGE(config-if)# no shutdown ASA-EDGE(config-if)# interface BVI 1 ASA-EDGE(config-if)# ip address 10.10.0.250 255.255.255.0 ASA-EDGE(config)# route outside 0.0.0.0 0.0.0.0 10.10.0.1 1 ``` Notice that there are no IPs on the data interfaces. The BVI sits at `10.10.0.250`, which is in the same subnet that the inside hosts already use. Their default gateway stays at `10.10.0.1` (the upstream router), and the ASA quietly bridges those frames toward outside. The static `route outside` is for traffic that the ASA itself originates: syslog to a remote server, NTP to time.example.com, SNMP traps to a NMS. Hosts behind the ASA still route via their existing gateway. Verify the bridge group: ``` ASA-EDGE# show bridge-group Bridge Group: 1 Interfaces: GigabitEthernet0/0 (outside) GigabitEthernet0/1 (inside) Management System IP Address: 10.10.0.250 255.255.255.0 Static mac-address entries: 0 Dynamic mac-address entries: 14 ASA-EDGE# show mac-address-table interface mac address type Time Left ----------------------------------------------------------------------- outside aabb.cc00.0100 dynamic 5 inside aabb.cc00.0200 dynamic 5 inside aabb.cc00.0201 dynamic 5 ``` The MAC address table is the transparent-mode equivalent of the routing table: it tracks which interface each MAC is reachable through, and it expires entries that go silent. ## What You Lose in Transparent Mode Transparent mode preserves a lot, but not everything. Plan around the following before you flip the switch. Dynamic routing protocols (OSPF, EIGRP, BGP, RIP) Not supported. Static routes only, and only for ASA-sourced traffic. DHCP relay Supported only on bridge-group interfaces; DHCP server on the ASA is supported but rarely used here. QoS Limited. Multicast routing Multicast bridging works (just forwarded across the bridge); multicast routing as a feature does not. VPN termination on the data plane (RA VPN, S2S terminating on a transparent interface) Not supported on the bridged side. Management-IP-terminated S2S is supported with restrictions. NAT Supported, but only between bridge-group members on the same bridge group, and use cases are narrow. Multiple context with shared interfaces Each context gets its own bridge group; you cannot put one interface in two transparent contexts. ARP inspection Supported and very useful in transparent mode; `arp-inspection inside enable`. BPDU forwarding By default, the ASA forwards BPDUs. Disable with caution; STP loops upstream of a transparent ASA are awful to debug. The biggest practical loss is dynamic routing. If you need OSPF or BGP at the firewall, transparent mode is off the table. ## ACLs Work the Same in Both Modes One pleasant surprise: ACL syntax, NAT object syntax, and inspection-engine syntax are identical between routed and transparent modes. An `access-list OUTSIDE_IN extended permit tcp any object DMZ-WEB eq https` means the same thing in both. The pillar's [ACL configuration](https://www.pinglabz.com/cisco-asa-acl-configuration/) guide and [NAT explained](https://www.pinglabz.com/cisco-asa-nat-explained/) walkthroughs both apply. Transparent mode adds one extra capability: **EtherType ACLs**. These let you filter non-IP frames (IPX, MPLS, ARP). They are configured separately from extended ACLs and bound with `access-group` the same way. ``` ASA-EDGE(config)# access-list ALLOW-MPLS ethertype permit mpls-unicast ASA-EDGE(config)# access-list ALLOW-MPLS ethertype permit mpls-multicast ASA-EDGE(config)# access-group ALLOW-MPLS in interface inside ``` Routed mode does not understand EtherType ACLs (it never sees the EtherType, only the IP packet). If you need to filter at the L2 frame level, transparent is the only mode that can do it. ## Diagnostic Commands by Mode Some `show` commands behave differently. Use the right ones for the mode you are in. `show route` Routed mode Full RIB with connected, static, dynamic Transparent mode Only ASA-sourced static routes; thin output `show interface ip brief` Routed mode Each data interface has an IP Transparent mode Data interfaces show "unassigned"; only BVI has an IP `show bridge-group` Routed modeEmpty / no output Transparent mode Lists bridge groups, member interfaces, and BVI IP `show mac-address-table` Routed mode Available but rarely used Transparent mode Primary forwarding table `show arp` Routed mode Standard ARP cache for connected subnets Transparent mode Standard ARP cache for the BVI subnet only `show conn` Routed modeSame flag legend Transparent modeSame flag legend `packet-tracer` Routed mode Full pipeline including UN-NAT and NAT Transparent mode Same phases; routing phase shows BVI/bridge-group output [packet-tracer](https://www.pinglabz.com/cisco-asa-packet-tracer/) still works in transparent mode, which is great because diagnosing a transparent firewall by sight is harder than diagnosing a routed one (the firewall is invisible to the data plane). ## Mode Interacts with Multi-Context The two global flags (`firewall transparent` and `mode multiple`) are independent, but their interaction has rules. In multi-context mode, every context is either routed or transparent, set on a per-context basis. You can mix and match: a perimeter routed context that does NAT and VPN, plus a separate transparent context that protects a database tier on the inside. Shared interfaces are not supported between transparent contexts. See [Cisco ASA Multiple Context Mode](https://www.pinglabz.com/cisco-asa-multiple-context-mode/) for the full configuration model. Switching between single and multi-context is also destructive; treat both global flags with the same caution. ## What I Actually See in the Field Across the ASA fleets I have worked on, transparent mode is rare. The cases where it shows up are real but narrow: a bank that needed to slot a firewall between its trading floor and the rest of the LAN without renumbering anything; a hospital that needed PCI-style segmentation between a payment-card subnet and the same VLAN that ran the rest of the office; an MSP that inherited a customer environment where re-IP was politically impossible. In every case, the team that built it understood that they were giving up dynamic routing and accepting a thinner toolset in exchange for not touching the existing IP plan. For everything else (internet edge, VPN concentrator, DMZ publishing), routed mode is the default for good reason. Most of the rest of this cluster covers routed-mode tasks: [static routing](https://www.pinglabz.com/cisco-asa-static-routing/), [NAT](https://www.pinglabz.com/cisco-asa-nat-explained/), [ACLs](https://www.pinglabz.com/cisco-asa-acl-configuration/), [AnyConnect](https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/), and [failover](https://www.pinglabz.com/cisco-asa-active-standby-failover/). ## Key Takeaways Routed mode is the default, supports the full feature set, and is what every other article on this site assumes. Transparent mode is a deliberate choice that solves the "drop a stateful firewall into an existing subnet without re-IP" problem and accepts the loss of dynamic routing, VPN data-plane termination, and a handful of other features. Confirm the current mode with `show firewall` before changing anything. Switching modes is destructive: `firewall transparent` wipes most of the running configuration. If you need transparent mode, plan a maintenance window, save the existing config off-box, switch the mode, and reload the transparent-style config from scratch. For the next step in the cluster, see [Cisco ASA Interfaces, Subinterfaces, and VLAN Trunks](https://www.pinglabz.com/cisco-asa-interfaces-vlan-trunks/), which covers the underlying interface configuration that both modes depend on. And the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/) has the full reading order for the rest of the platform. ### Cisco ASA Common Outage Scenarios and Fixes URL: https://www.pinglabz.com/cisco-asa-common-outages/ Last updated: 2026-06-13T20:08:25.000Z Every Cisco ASA that runs long enough hits a small set of repeatable failure modes: a NAT pool exhausting, an ACL line that shadows another and silently breaks one app, a VPN tunnel that comes up and then carries no traffic, a failover pair that splits brain. The patterns are small, the symptoms look identical to dozens of other problems, and the fix is usually one show command away. This article is a field guide to the four most common ASA outage scenarios, the diagnostic command that decides each, and the fix - all using real evidence from the live ASAv 9.23(1) in the [PingLabz ASA reference](https://www.pinglabz.com/cisco-asa/) lab. The pattern across all four: the symptom users report is generic ("the app is slow", "VPN does not work"), but the firewall has a single command output that names the problem. Knowing which command for which symptom is the entire diagnostic skill. ## Scenario Map: Symptom to Confirming Command "New connections work but slow / sometimes failing" Likely cause NAT pool exhaustion, PAT slot starvation Confirming command `show xlate count` \+ `show asp drop | include nat-no-xlate` "This one app stopped working but everything else is fine" Likely cause ACL line shadowing - a deny higher up matches before the intended permit Confirming command `show access-list ` with hit counters "VPN tunnel is up but traffic is not flowing" Likely cause encaps incrementing but decaps stuck at 0; routing or NAT exemption missing on far side Confirming command `show crypto ipsec sa | include peer|encaps|decaps` "Failover pair shows both units active" Likely cause Split-brain - failover link down or peer authentication mismatch Confirming command `show failover history` \+ `show failover state` ## Scenario 1: NAT / PAT Pool Exhaustion Symptom users see: existing connections work fine, new connections from inside hosts intermittently fail to establish. The application teams think the application is broken; networking sees an ASA that is not paged on anything. Both are wrong - the ASA is healthy but cannot allocate a new PAT mapping. The underlying mechanism: PAT translates many internal source addresses to one external IP by varying the source port. The external port range is 1024-65535, so a single external IP supports at most \~64,500 simultaneous outbound flows. Once exhausted, new flows hit `nat-no-xlate-to-pat-pool` and drop silently. The diagnostic chain: ``` ASA-PERIM# show xlate count 14 in use, 16 most used ASA-PERIM# show asp drop | include nat-no-xlate nat-no-xlate-to-pat-pool 0 ``` In the lab those numbers are healthy (14 in use is nowhere near 64K). On a perimeter under load you might see "in use" climbing to 50K+ and the asp-drop counter incrementing - that is your unambiguous signal. The next show command identifies which source is the busiest: ``` ASA-PERIM# show xlate detail | include PAT UDP PAT from inside:10.10.10.1/58380 to outside:203.0.113.2/58380 flags ri idle 0:00:37 timeout 0:00:00 UDP PAT from inside:10.10.10.1/52513 to outside:203.0.113.2/52513 flags ri idle 0:00:57 timeout 0:00:00 UDP PAT from inside:10.10.10.1/64356 to outside:203.0.113.2/64356 flags ri idle 0:01:17 timeout 0:00:00 UDP PAT from inside:10.10.10.1/51684 to outside:203.0.113.2/51684 flags ri idle 0:00:17 timeout 0:00:00 ... [many more] ... ``` If one inside host has thousands of PAT entries, that host is the noisemaker. Either it is a real production app burning ports legitimately (in which case widen the pool) or it is leaking sockets (fix the app). Three production fixes: 1. **Add a backup PAT pool.** Add a second IP to the existing nameif's NAT pool: `nat (inside,outside) source dynamic INSIDE-NET pool PAT-POOL-MAIN backup pool PAT-POOL-OVERFLOW`. The ASA fails over from main to overflow when main is exhausted. 2. **Switch to per-session PAT.** Adds `per-session permit udp any any` \+ `per-session permit tcp any any` globally. The default is per-session, so this is rarely needed. 3. **Lower the inactivity timer.** `timeout xlate 0:01:00` reduces the time a PAT mapping survives idle. Risky for long-poll applications - test before you ship. ## Scenario 2: ACL Line Shadowing Symptom users see: one specific application stopped working, every other app on that host works fine. Often happens after someone adds a "tighten that down" deny rule above the catch-all permits. The new rule matches more than intended. The diagnostic: `show access-list ` with hit counters, looking for whether the deny line that was added is firing on the protected destination. ``` ASA-PERIM# show access-list OUTSIDE_IN | begin OUTSIDE_IN access-list OUTSIDE_IN; 5 elements; name hash: 0xe01d8199 access-list OUTSIDE_IN line 1 extended deny tcp host 198.51.100.99 any (hitcnt=1) (Last Hit=00:01:09 UTC May 10 2026) access-list OUTSIDE_IN line 2 extended permit tcp any object DMZ-WEB eq www (hitcnt=1) (Last Hit=23:27:41 UTC May 9 2026) access-list OUTSIDE_IN line 3 extended permit tcp any object DMZ-WEB eq https (hitcnt=2) (Last Hit=00:01:09 UTC May 10 2026) access-list OUTSIDE_IN line 4 extended permit icmp any object DMZ-WEB (hitcnt=2) (Last Hit=01:55:20 UTC May 10 2026) access-list OUTSIDE_IN line 5 extended deny ip any any log informational interval 300 (hitcnt=59) (Last Hit=01:55:47 UTC May 10 2026) ``` Read the hit counters from top to bottom. Line 1 (an explicit deny against a specific source) has 1 hit - that is a one-off scanner, not the problem. Line 5 (the catch-all explicit deny with logging) has 59 hits - that is the smoking gun. 59 packets matched `deny ip any any` instead of any of the lines above. To find what those 59 packets were, the `log informational` on line 5 is doing the work for you - their syslogs land in `show logging`: ``` ASA-PERIM# show logging | include "Deny inbound" | include 4-106023 %ASA-4-106023: Deny tcp src outside:8.8.8.8/56321 dst dmz:192.168.50.10/8080 by access-group "OUTSIDE_IN" ... [more entries] ... ``` Now you know exactly what was being denied: TCP from 8.8.8.8 to 192.168.50.10 on port 8080\. If port 8080 is the application that broke, you have a permit gap (line 2 covers /80 only, line 3 covers /443 only, nothing covers /8080). The fix is to add a permit for the missing port BEFORE the catch-all deny, then re-test: ``` access-list OUTSIDE_IN line 4 extended permit tcp any object DMZ-WEB eq 8080 ``` Inserting at line 4 puts it before the original line 5 deny (which auto-renumbers to line 6). After the change, watch the new line 4 hitcnt climb and the line 5 hitcnt stop climbing for the previously-blocked traffic. ## Scenario 3: VPN Tunnel Up But No Traffic Symptom users see: VPN status icon green; traffic to the remote subnet times out. The classic asymmetric tunnel failure. The diagnostic: `show crypto ipsec sa` with focus on the encaps and decaps counters: ``` ASA-PERIM# show crypto ipsec sa | include peer|encaps|decaps current_peer: 203.0.113.6 #pkts encaps: 248, #pkts encrypt: 248, #pkts digest: 248 #pkts decaps: 0, #pkts decrypt: 0, #pkts verify: 0 ``` Encaps = 248 (we are sending plenty). Decaps = 0 (we are receiving nothing). The tunnel is up - both peers agreed on Phase 1 and Phase 2 - but the far side either is not sending return traffic or its return traffic never reaches us. Three causes, in priority order: 1. **The far side has no return route to our protected subnet.** Most common. Login to the far peer and run a routing-table check for our remote IP. If it does not have a route, traffic from local hosts on the far side never reaches the tunnel. 2. **The far side is NATing return traffic before encrypt.** NAT-after-encrypt is fine; NAT-before-encrypt breaks the proxy ACL match. Check NAT exemption ([NAT exemption for VPNs](https://www.pinglabz.com/cisco-asa-identity-nat-vpn/)) on the far side. 3. **The far side has its proxy ACL inverted.** The crypto map's interesting traffic must be the mirror image of ours: our local subnet = their remote subnet, our remote = their local. If they are flipped, return packets fail the SA selector and are dropped at decap. If decaps is non-zero but smaller than encaps, you have packet loss on the far-to-near direction - typically a path MTU problem with overlay encapsulation. Set `crypto ipsec security-association tcp-mss-adjust 1380` on both peers and retest. This pattern is so common that the lab has it documented on the dedicated VPN troubleshooting page; see [troubleshoot IPsec phases on Cisco ASA](https://www.pinglabz.com/cisco-asa-troubleshoot-ipsec-phases/) for the full Phase-1/Phase-2 decision tree. Get the Cisco ASA Field Reference - 9 pages, free Everything you'd want to remember about Cisco ASA on nine printable pages. Per-packet pipeline diagram, NAT 8.3+ section ordering, six-branch troubleshooting decision tree, real lab show-output annotated, paste-ready three-zone config. Free for PingLabz members - just sign up with your email. [Get the Cisco ASA cheat-sheet](https://www.pinglabz.com/cisco-asa-cheatsheet/) ## Scenario 4: Failover Pair Split-Brain Symptom users see: intermittent connection drops, sometimes traffic flows, sometimes it does not, packet captures show the same flow being answered by two different MAC addresses. The pair has both units announcing themselves as Active. Split-brain happens when the failover link goes down between two ASAs that are both healthy. Each unit sees its peer as failed and promotes itself to Active. Now both units forward traffic, both NAT, both register ARP for the same gateway IPs - hence the alternating MACs and the broken sessions. The diagnostic on each unit: ``` ASA-PRIMARY# show failover state This host - Primary Active None Other host - Secondary Failed Comm Failure Stateful Failover Logical Update Statistics ASA-PRIMARY# show failover history ========================================================================== From State To State Reason ========================================================================== 13:14:01 UTC May 10 2026 Active Standby Ready Active Drain Other unit wants me Active 13:14:01 UTC May 10 2026 Active Drain Active Applying Config Other unit wants me Active 13:14:01 UTC May 10 2026 Active Applying Config Active Config Applied Other unit wants me Active 13:14:01 UTC May 10 2026 Active Config Applied Active Other unit wants me Active 13:34:22 UTC May 10 2026 Active Active Comm Failure ========================================================================== ``` Both units will show "This host: Active" with reason "Comm Failure" against the peer. That is the split-brain signature. The "Active Active" transition with reason "Comm Failure" in the history is the timestamp at which the failover link went down. The fix has two parts: 1. **Bring the failover link back up before promoting any traffic decisions.** Look at `show interface` for the configured failover lan interface; if it is down/down, the cable, port, or transceiver is failed. Until the link is up the units have no way to renegotiate. 2. **Force one unit to standby once the link is healthy.** Pick the unit that you want to remain Active - typically the one with the most recent valid config - and on the OTHER unit run `no failover active`. That unit transitions to Standby Ready and pulls config from the active. The session is briefly disrupted as the duplicate ARPs clear, but the pair is now consistent. Prevention: monitor the failover link. `monitor-interface FAIL-LINK` on each unit triggers a syslog the moment that interface goes link-down. Pair with `logging trap warnings` and a syslog server you actually read. ## Other Fast Checks Worth Knowing Less common but worth keeping in your head: `show resource usage` Per-context resource utilization on multi-context, or per-platform on single-mode. Conn count, xlate count, syslog rate, AAA tx rate vs platform max. `show service-policy global` Inspection engine hit counters. If "tcp-options drops" or "esmtp drops" climb, an inspection is rejecting your application. `show route summary` Number of routes per source. If static count drops unexpectedly, someone deleted routes. `show vpn-sessiondb summary` Live VPN session count. If a remote-access VPN keeps dropping users, this trends the count. `show clock` \+ `show ntp associations` NTP drift breaks AnyConnect cert validation. Worth an early check on any VPN issue. ## Key Takeaways Four scenarios cover most "the firewall is broken" tickets: NAT pool exhaustion (look at `show xlate count` \+ asp-drop), ACL line shadowing (look at `show access-list` hit counters), VPN encaps-without-decaps (look at `show crypto ipsec sa`), and failover split-brain (look at `show failover history`). Each has a single confirming command, then a small set of fixes. The full [Cisco ASA reference cluster](https://www.pinglabz.com/cisco-asa/) goes deeper on each: [NAT order and types](https://www.pinglabz.com/cisco-asa-nat-explained/), [ACL troubleshooting](https://www.pinglabz.com/cisco-asa-acl-troubleshooting/), [IPsec phase troubleshooting](https://www.pinglabz.com/cisco-asa-troubleshoot-ipsec-phases/), [active/standby failover configuration](https://www.pinglabz.com/cisco-asa-active-standby-failover/), and the operational toolkit for live diagnosis: [packet-tracer](https://www.pinglabz.com/cisco-asa-packet-tracer/), [CLI packet capture](https://www.pinglabz.com/cisco-asa-packet-capture/), [asp-drop counters](https://www.pinglabz.com/cisco-asa-asp-drop/), and [conn / xlate troubleshooting](https://www.pinglabz.com/cisco-asa-conn-xlate/). ### Cisco ASA Connection Table and xlate Table Troubleshooting URL: https://www.pinglabz.com/cisco-asa-conn-xlate/ Last updated: 2026-06-13T20:08:25.000Z The Cisco ASA tracks two layered tables that explain almost every "is this traffic flowing" question: the *connection table* (one entry per active flow, 5-tuple plus state) and the *xlate table* (one entry per active NAT translation, source-side and destination-side). When something is wrong with traffic that should be passing the firewall, `show conn` and `show xlate` tell you whether the flow exists at all and whether it is being NATed correctly. This article walks the table layout, the most useful filters, and the operational `clear` commands you reach for during real troubleshooting, all using live output from the [PingLabz ASA reference](https://www.pinglabz.com/cisco-asa/) lab on ASAv 9.23(1). If you do not yet have a flow in `show conn` when you expect one, start with [packet-tracer](https://www.pinglabz.com/cisco-asa-packet-tracer/) to find the dropping phase, then [an interface capture](https://www.pinglabz.com/cisco-asa-packet-capture/) to confirm the packet arrived. If the flow is in the table but the application is not working, the problem is usually NAT - and that is where `show xlate` earns its keep. ## Connection Table vs Xlate Table What it tracks show conn Active flows: 5-tuple, interfaces, state flags, idle timer show xlate NAT bindings: pre-NAT to post-NAT, type (static/dynamic/PAT/identity/twice), refcount Per-flow lifetime show conn Born when first packet creates flow, dies when timeout or RST/FIN show xlate For dynamic NAT, born and dies with the flow (refcount == 0). For static, lives as long as the rule exists. Default capacity show conn 250,000 - several million depending on platform memory show xlate Same scale; tied to memory, not interface speed Filterable by show conn Address, protocol, port, interface, state show xlate Address (local/global), interface Cleared with show conn`clear conn` show xlate `clear xlate` (also clears tied conns) The relationship: every dynamic NAT entry has a corresponding conn (or several, if PAT is multiplexing). Static and identity entries can exist without conns. The `xlate id` field on a conn is a pointer back to the xlate that produced it. ## Counting What's There ``` ASA-PERIM# show conn count 14 in use, 15 most used ASA-PERIM# show xlate count 14 in use, 16 most used ``` "In use" is current; "most used" is the high-water mark since the last reload. If "most used" is approaching the platform license cap, you have a sizing problem coming. For the lab's small ASAv that is no issue, but on a real perimeter the most-used number is the trend you watch with `show resource usage`. ## show conn detail: The Full Flow Picture Without `detail`, `show conn` gives you the 5-tuple and a flag string. With `detail`, you also get internal flow IDs, NAT pointers, idle timers, and uptime - the data you need to correlate a flow with an xlate, and to diagnose a stuck flow. ``` ASA-PERIM# show conn detail address 10.10.10.1 14 in use, 15 most used Flags: A - awaiting inside ACK to SYN, a - awaiting outside ACK to SYN, B - initial SYN from outside, b - TCP state-bypass or nailed, ... [legend continues] ... U - up, ... UDP outside: 8.8.8.8/1967 inside: 10.10.10.1/52151, flags - , idle 34s, uptime 44s, timeout 2m0s, bytes 156, xlate id 0x7f0a9813eb00, flow id 195 UDP outside: 8.8.8.8/1967 inside: 10.10.10.1/64852, flags - , idle 1m54s, uptime 2m4s, timeout 2m0s, bytes 156, xlate id 0x7f0a9813e600, flow id 186 UDP dmz: 192.168.50.10/1967 inside: 10.10.10.1/58256, flags - , idle 1m55s, uptime 2m5s, timeout 2m0s, bytes 156, flow id 185 UDP dmz: 192.168.50.10/1967 inside: 10.10.10.1/57775, flags - , idle 35s, uptime 45s, timeout 2m0s, bytes 156, flow id 194 UDP outside: 8.8.8.8/1967 inside: 10.10.10.1/57936, flags - , idle 14s, uptime 24s, timeout 2m0s, bytes 156, xlate id 0x7f0a9813e880, flow id 197 ``` Read this: - `UDP outside: 8.8.8.8/1967 inside: 10.10.10.1/52151` \- a UDP flow whose outside-leg is 8.8.8.8:1967 and inside-leg is 10.10.10.1:52151\. Source is 10.10.10.1, destination is 8.8.8.8\. The "inside:" and "outside:" tell you which interface terminates each leg. - `flags -` \- no special flags. UDP has fewer flags than TCP because there is no state machine. A typical TCP flow shows something like `UIO` \= Up, Inbound data, Outbound data. - `idle 34s, uptime 44s, timeout 2m0s` \- the flow has been idle 34 seconds and has lived 44 seconds; it will be culled if it stays idle past 2 minutes. - `xlate id 0x7f0a9813eb00` \- this is the pointer to the xlate entry. Useful when you want to see exactly which NAT rule produced this flow. - `flow id 195` \- internal sequence number. Higher = newer. The two flow types in this output show the lab's NAT architecture clearly: flows whose outside-leg is 8.8.8.8 transited the outside interface (and have an xlate id, because they are PATted), while flows whose outside-leg is 192.168.50.10 transited the dmz interface (and have no xlate id, because inside-to-dmz traffic is not NATed). ## Filtering by Protocol or Port For triage on a busy ASA, narrow the view: ``` ASA-PERIM# show conn protocol udp 14 in use, 15 most used UDP outside 8.8.8.8:1967 inside 10.10.10.1:52151, idle 0:00:33, bytes 156, flags - UDP outside 8.8.8.8:1967 inside 10.10.10.1:64852, idle 0:01:53, bytes 156, flags - UDP dmz 192.168.50.10:1967 inside 10.10.10.1:58256, idle 0:01:53, bytes 156, flags - ... [continued] ... ``` Other useful filters: - `show conn protocol tcp port www` \- all TCP flows on port 80. - `show conn address 192.168.50.10` \- every flow involving DMZ-WEB. - `show conn state up` \- only TCP flows in fully-established UP state. The complement (`state half-closed`, `state finin`, `state finout`) shows you flows that are tearing down. - `show conn long` \- show flows older than the long timer (typically used for DC-style long-running connections). ## Reading the Flag String The flag string is dense but reveals exactly what state a TCP flow is in: `U` Up: full handshake completed both directions. `I` Inbound data has flowed. `O` Outbound data has flowed. `A` Awaiting inside ACK to SYN. Three-way handshake in progress. `a` Awaiting outside ACK to SYN. `F` Outside FIN seen. `f` Inside FIN seen. Both `F` and `f` \= teardown in progress. `R` Outside acknowledged FIN. Half-closed. `r` Inside acknowledged FIN. `i` Incomplete: TCP flow lacking either side's full handshake. `b` TCP state-bypass or nailed. ASA is not enforcing TCP state on this flow. `N` Inspected by Snort/IPS. `D` DNS or Umbrella inspection. `V` VPN orphan: flow exists for traffic that hit a tunnel but the tunnel is gone. A flow stuck at `UI` (Up, Inbound only) for a long time is one-way traffic. `i` alone means the handshake never completed - usually a downstream firewall dropping the SYN-ACK. `V` orphans accumulate when a VPN flaps; clear them with `clear conn vpn-flow`. ## show xlate detail: The NAT Side The xlate table has its own flag legend and shows pre/post-NAT mappings: ``` ASA-PERIM# show xlate detail 14 in use, 16 most used Flags: D - DNS, e - extended, I - identity, i - dynamic, r - portmap, s - static, T - twice, N - net-to-net NAT from dmz:192.168.50.10 to outside:198.51.100.10 flags s idle 0:00:14 timeout 0:00:00 xlate id 0x7f0a98140b80 NAT from inside:10.10.0.0/16 to outside:10.10.0.0/16 flags sIT idle 1:53:57 timeout 0:00:00 xlate id 0x7f0a98140180 NAT from outside:10.99.99.0/24 to inside:10.99.99.0/24 flags sIT idle 1:53:57 timeout 0:00:00 xlate id 0x7f0a98140400 NAT from outside:172.20.0.0/24 to inside:172.20.0.0/24 flags sIT idle 1:53:56 timeout 0:00:00 xlate id 0x7f0a9813ff00 UDP PAT from inside:10.10.10.1/58380 to outside:203.0.113.2/58380 flags ri idle 0:00:37 timeout 0:00:00 UDP PAT from inside:10.10.10.1/52513 to outside:203.0.113.2/52513 flags ri idle 0:00:57 timeout 0:00:00 UDP PAT from inside:10.10.10.1/64356 to outside:203.0.113.2/64356 flags ri idle 0:01:17 timeout 0:00:00 UDP PAT from inside:10.10.10.1/51684 to outside:203.0.113.2/51684 flags ri idle 0:00:17 timeout 0:00:00 ``` Read the flags: - `s` \- static. The mapping was created by an explicit static rule and persists regardless of traffic. - `I` \- identity. The pre-NAT and post-NAT addresses are the same (used for VPN exemptions and routing-only translations). - `T` \- twice NAT. Both source and destination are translated by the same rule. - `i` \- dynamic. Created on demand by traffic. - `r` \- portmap (PAT). Source port translated as well as IP. - `D` \- DNS rewrite applied. - `N` \- net-to-net (whole subnet mapping). The first entry (`flags s`) is the DMZ-WEB static (one-to-one between the inside-facing and outside-facing IPs). The next three (`flags sIT`) are static-identity-twice rules used for VPN exemption and partner-network exemption. The last four (`flags ri`) are dynamic PAT mappings: each one is a single source-port-to-source-port mapping built on demand for an outbound flow. Get the Cisco ASA Field Reference - 9 pages, free Everything you'd want to remember about Cisco ASA on nine printable pages. Per-packet pipeline diagram, NAT 8.3+ section ordering, six-branch troubleshooting decision tree, real lab show-output annotated, paste-ready three-zone config. Free for PingLabz members - just sign up with your email. [Get the Cisco ASA cheat-sheet](https://www.pinglabz.com/cisco-asa-cheatsheet/) ## Filtering xlate by Source Address ``` ASA-PERIM# show xlate local 10.10.10.1 14 in use, 16 most used UDP PAT from inside:10.10.10.1/50199 to outside:203.0.113.2/50199 flags ri idle 0:00:05 timeout 0:00:00 UDP PAT from inside:10.10.10.1/60087 to outside:203.0.113.2/60087 flags ri idle 0:01:05 timeout 0:00:00 UDP PAT from inside:10.10.10.1/62133 to outside:203.0.113.2/62133 flags ri idle 0:00:15 timeout 0:00:00 UDP PAT from inside:10.10.10.1/64852 to outside:203.0.113.2/64852 flags ri idle 0:02:05 timeout 0:00:00 UDP PAT from inside:10.10.10.1/52151 to outside:203.0.113.2/52151 flags ri idle 0:00:45 timeout 0:00:00 UDP PAT from inside:10.10.10.1/57936 to outside:203.0.113.2/57936 flags ri idle 0:00:25 timeout 0:00:00 UDP PAT from inside:10.10.10.1/58379 to outside:203.0.113.2/58379 flags ri idle 0:01:45 timeout 0:00:00 UDP PAT from inside:10.10.10.1/52529 to outside:203.0.113.2/52529 flags ri idle 0:01:25 timeout 0:00:00 ``` Eight active PAT mappings for one source. `show xlate local` filters by pre-NAT (inside) IP; `show xlate global` filters by post-NAT (outside) IP. On a busy perimeter the global filter is what you reach for first when an inbound flow is failing because you suspect the wrong outbound source IP. ## show xlate count: The PAT Pool Health Check The single most useful operational check on an ASA's NAT subsystem is `show xlate count` \- watch the trend, not the absolute number. PAT (portmap) entries consume one source-port slot per flow. With one outside interface IP and 65535 - 1024 = 64511 ephemeral ports, the platform can support at most \~64K simultaneous PAT mappings before running out. If `show xlate count` rises faster than connections are leaving and the PAT pool is one IP, you are heading for `nat-no-xlate-to-pat-pool` drops. Add a backup PAT IP, expand to a pool, or split the source. ## Clearing Conns and Xlates When you change a NAT rule or an ACL, the ASA does not retroactively rebuild flows that were created under the old rules. They will continue to forward through the firewall using the old translation until they idle out or you clear them. If you want immediate effect: ``` ASA-PERIM# clear conn address 10.10.10.1 ASA-PERIM# clear xlate local 10.10.10.1 ``` `clear conn` kills the flow; on the next packet, the ASA re-runs the new rules. `clear xlate` kills the NAT binding and any conns tied to it. In production you almost never use `clear conn all` or `clear xlate all` \- that resets every flow on the firewall and is service-affecting. Other surgical variants: - `clear conn protocol tcp port www` \- drop only HTTP connections. - `clear conn state finin` \- drop only flows in TCP FIN-WAIT. - `clear xlate global 198.51.100.99` \- drop the partner-PAT mapping when its tunnel changes. - `clear conn vpn-flow` \- clear orphaned VPN flows after a tunnel flap. ## Key Takeaways Two tables, two questions: `show conn` tells you whether a flow exists and what state it is in; `show xlate` tells you whether NAT is producing the right pre/post mapping. When traffic is failing, walk the chain: confirm the flow exists in `show conn detail`, find its `xlate id`, follow that to `show xlate`, and verify the mapping is what you expected. The flags - `i` for incomplete, `V` for VPN orphan, `sIT` for static-identity-twice - are most of the diagnostic content. For surgical fixes, `clear conn` and `clear xlate` with tight filters reset just the affected flows. The full [Cisco ASA reference cluster](https://www.pinglabz.com/cisco-asa/) covers the rest of the operational toolkit, including [CLI packet capture](https://www.pinglabz.com/cisco-asa-packet-capture/), [asp-drop counters](https://www.pinglabz.com/cisco-asa-asp-drop/), and [common outage scenarios](https://www.pinglabz.com/cisco-asa-common-outages/) that this table walk is the fastest path through. ### Cisco ASA asp-drop Counters Explained URL: https://www.pinglabz.com/cisco-asa-asp-drop/ Last updated: 2026-06-13T20:08:25.000Z The Cisco ASA's accelerated security path (ASP) is the data-plane fast path that handles every forwarded packet after a flow has been admitted. When the ASP drops a packet, the reason it cites in `show asp drop` is your single most useful diagnostic on a working firewall. There is no syslog for most ASP drops, no debug, no ACL hit counter - just an internal counter that increments and a one-line label. This article decodes the most common drop reasons, walks through the lab evidence behind each, and shows how to chain `show asp drop` with an `asp-drop type` capture to see the actual packets being discarded. All output below is from the live ASAv 9.23(1) in the [PingLabz ASA reference](https://www.pinglabz.com/cisco-asa/) lab. If you are about to draft an ASA [ACL troubleshooting](https://www.pinglabz.com/cisco-asa-acl-troubleshooting/) ticket or you suspect a NAT path is silently failing, run `show asp drop` first. About 70% of "the firewall is broken" cases resolve to a counter that explains exactly what the ASA was doing. The other 30% need a [packet capture](https://www.pinglabz.com/cisco-asa-packet-capture/) to confirm. ## show asp drop: The Counter Dump The base command lists every drop reason that has incremented since last clear, separated into *frame* drops (per-packet, classification-time) and *flow* drops (per-flow, midstream). From the lab: ``` ASA-PERIM# show asp drop Frame drop: No valid adjacency (no-adjacency) 40 No valid V4 adjacency. Check ARP table (show arp) has entry for nexthop. (no-v4-adjacency) 5 Flow is denied by configured rule (acl-drop) 1606 FP L2 rule drop (l2_acl) 19 Interface is down (interface-down) 3 IKE new SA limit exceeded (ike-sa-rate-limit) 3 Last clearing: Never Flow drop: Need to start IKE negotiation (need-ike) 14 Last clearing: Never ``` Six frame drop reasons, one flow drop reason. Each is a real signal. Let's read them. ## acl-drop: Flow is Denied by Configured Rule 1606 hits is by far the largest counter, and it is exactly what it says: 1606 packets matched a deny statement (explicit or implicit) on an interface ACL. To find which rule: ``` ASA-PERIM# show access-list OUTSIDE_IN | begin OUTSIDE_IN access-list OUTSIDE_IN; 5 elements; name hash: 0xe01d8199 access-list OUTSIDE_IN line 1 extended deny tcp host 198.51.100.99 any (hitcnt=1) (Last Hit=00:01:09 UTC May 10 2026) access-list OUTSIDE_IN line 2 extended permit tcp any object DMZ-WEB eq www (hitcnt=1) (Last Hit=23:27:41 UTC May 9 2026) access-list OUTSIDE_IN line 3 extended permit tcp any object DMZ-WEB eq https (hitcnt=2) (Last Hit=00:01:09 UTC May 10 2026) access-list OUTSIDE_IN line 4 extended permit icmp any object DMZ-WEB (hitcnt=2) (Last Hit=01:55:20 UTC May 10 2026) access-list OUTSIDE_IN line 5 extended deny ip any any log informational interval 300 (hitcnt=59) (Last Hit=01:55:47 UTC May 10 2026) ``` The asp-drop counter (1606) and the explicit-deny ACL counter (59 on line 5) do not match because asp-drop catches everything denied anywhere on the data path, not just by named ACLs. ASP also counts: implicit deny on un-applied directions, asp-table denies for unknown protocols, and out-of-state drops that classify as ACL on the way down. Use ACL counters to identify a deny line; use asp-drop to confirm the magnitude. To see the actual packets being denied, attach a capture: ``` ASA-PERIM# capture ASP-DROP type asp-drop acl-drop ASA-PERIM# show capture ASP-DROP packet-number 100 342 packets captured 100: 01:55:07.071498 203.0.113.1.58279 > 198.51.100.10.1967: udp 52 Drop-reason: (acl-drop) Flow is denied by configured rule, Drop-location: frame snp_classify_table_lookup:6044 flow (NA)/NA 1 packet shown ``` That gives you the source, destination, port, and timestamp of every dropped packet. From there it is one `show access-list` away from the rule that fired. ## no-adjacency and no-v4-adjacency "No valid adjacency" is what an ASA says when it has a route to a destination but cannot resolve the next-hop. The two flavors: `no-adjacency` Means Generic "I cannot get a layer-2 adjacency for the next-hop". Usually IPv6 or unknown encap. Fix Check the routing table for the destination, then verify the next-hop is on a directly-connected interface. `no-v4-adjacency` Means Specifically "I cannot ARP for the IPv4 next-hop". The most common case is a missing ARP entry on the egress interface. Fix `show arp` on the egress interface; if missing, the downstream device is silent or unreachable. The lab's 40 + 5 hits came from two test scenarios. The `no-v4-adjacency` hits came from packets bound for an alpine host on the DMZ that does not respond to ARP from this side. The 40 generic `no-adjacency` hits came from IPv6 router-solicitations the ASA had no IPv6 nexthop for. The diagnostic discipline: when a flow is being silently lost between ingress and egress and packet-tracer says ALLOW, look at `show asp drop`. If `no-v4-adjacency` increments while you push test traffic, the next hop is not responding to ARP. `show arp`, then ping the next hop from the ASA, then look at the downstream device. ## interface-down 3 hits in the lab. The reason is exactly what it says: a packet was queued for an egress interface that came down before the packet was sent. Three packets in flight when the interface flapped. This counter is more useful than it looks - if it is climbing while no interfaces are obviously flapping, you have a transient layer-1 problem on one of the data interfaces. Pair with `show interface` looking at `output discards`. ## l2\_acl: FP L2 Rule Drop 19 hits. This is a layer-2-classifier drop, fired before the layer-3 ACL ever runs. The most common cause is a packet hitting an interface that is not the one its source MAC is associated with - which is usually a result of MAC address aging plus an asymmetric path - or a layer-2 EtherType the ASA does not pass by default (IS-IS frames, custom protocols). It can also be a real loop that the ASA is breaking. Look at `show interface | include MAC|input rate` on neighbors and `show conn detail` for any flow showing this counter increment alongside. ## ike-sa-rate-limit and need-ike The two IKE-related counters that show in the lab. These are VPN data-path drops: - `ike-sa-rate-limit`: 3 hits. The control-plane received more new IKE\_SA\_INIT requests in a window than the configured rate limit allows. Defaults are usually fine; if this counter climbs, you have an attacker probing UDP/500 or a misbehaving peer reinitating constantly. Mitigation: `crypto ikev2 limit max-sa-init-rate `. - `need-ike` (flow drop): 14 hits. A packet matched a crypto map's interesting traffic but the IKE\_SA was not yet established. The ASA queues the packet briefly, fires off IKE negotiation, and drops the packet if negotiation does not complete in time. 14 hits across a few hours is a normal background level, especially right after a reboot or a tunnel flap. Climbing fast = your IKE is failing. See [troubleshoot IPsec phases](https://www.pinglabz.com/cisco-asa-troubleshoot-ipsec-phases/). ## show asp drop flow vs show asp drop frame The split matters. *Frame drops* happen during initial classification, before a flow is created. *Flow drops* happen on packets that arrive on an existing flow but are then dropped by an inspection or a rate-limit. To see them separately: ``` ASA-PERIM# show asp drop frame No valid adjacency (no-adjacency) 40 No valid V4 adjacency. Check ARP table ... 5 Flow is denied by configured rule (acl-drop) 1606 FP L2 rule drop (l2_acl) 19 Interface is down (interface-down) 3 IKE new SA limit exceeded (ike-sa-rate-limit) 3 Last clearing: Never ASA-PERIM# show asp drop flow Need to start IKE negotiation (need-ike) 14 Last clearing: Never ``` If a single ticket shows large frame counts and small flow counts, the problem is at the door (ACL, ARP, layer-2). If flow counts dominate, the problem is mid-stream (inspection rejecting payload, MTU, asymmetric routing). Get the Cisco ASA Field Reference - 9 pages, free Everything you'd want to remember about Cisco ASA on nine printable pages. Per-packet pipeline diagram, NAT 8.3+ section ordering, six-branch troubleshooting decision tree, real lab show-output annotated, paste-ready three-zone config. Free for PingLabz members - just sign up with your email. [Get the Cisco ASA cheat-sheet](https://www.pinglabz.com/cisco-asa-cheatsheet/) ## show asp event-log: Per-Packet Trace For per-packet detail beyond the counters, use `show asp event dp-cp drop` to see the data-plane to control-plane queue and any drop events ASA buffered: ``` ASA-PERIM# show asp event dp-cp drop DP-CP EVENT QUEUE QUEUE-LEN HIGH-WATER Punt Event Queue 0 0 Routing Event Queue 0 0 Identity-Traffic Event Queue 0 1 PTP-Traffic Event Queue 0 0 General Event Queue 0 1 Syslog Event Queue 0 4 Non-Blocking Event Queue 0 0 Midpath High Event Queue 0 0 Midpath Norm Event Queue 0 0 Crypto Event Queue 0 0 HA Event Queue 0 0 HA Ctl Event Queue 0 0 Threat-Detection Event Queue 0 0 ARP Event Queue 0 1 TMATCH Event Queue 0 1 EVENT-TYPE ALLOC ALLOC-FAIL ENQUEUED ENQ-FAIL RETIRED 15SEC-RATE drop-flow 0 0 0 0 0 0 ``` The "high-water" column is the watch column. Anything climbing above 50 indicates the data path is back-pressuring the control plane - typically because of a flood (Threat-Detection queue) or many simultaneous IKE attempts (Crypto queue). On a healthy ASA all values should be 0 most of the time. ## Clearing and Trending the Counters Counters never reset on their own. The default state shows `Last clearing: Never`, which means the values are cumulative from the last reload. To trend a problem in real time, clear and watch: ``` ASA-PERIM# clear asp drop ASA-PERIM# show asp drop ... empty ... ! drive test traffic ... ASA-PERIM# show asp drop Frame drop: Flow is denied by configured rule (acl-drop) 23 ``` 23 acl-drops since clearing. Now you know the rate at which they happen, not just the lifetime total. Pair with `show clock` at clear time and at observation time for an exact rate per minute. ## The Reasons You Will See Most The reasons in this lab are a small subset of what the ASA can drop for. The catalog is large; here are the ten you will see in production most often, with their root cause: `acl-drop` Denied by an interface ACL or an implicit deny. `nat-no-xlate-to-pat-pool` Auto-NAT PAT pool exhausted. Increase the pool or split the source. `tcp-not-syn` TCP packet arrived for a flow with no existing connection state. Asymmetric routing or stateful violation. `tcp-bad-flags` TCP flag combination invalid (e.g. SYN+FIN). Often a scanner or broken middlebox. `no-route` Routing table has no entry for the destination. Look for missing static or dynamic route. `inspect-icmp-error-no-existing-conn` ICMP unreachable for a flow the ASA has no record of. Inspect-icmp-error is dropping it; remove the inspection if you need this traffic. `nat-rpf-failed` NAT reverse-path-forwarding check failed. The packet's source IP does not match the NAT direction. `flow-expired` Packet for a connection that timed out. Symptom of a long-idle flow that the firewall pruned. `no-adjacency / no-v4-adjacency` Cannot resolve next-hop layer 2. `need-ike` IKE not yet established for an interesting-traffic match. For the comprehensive list, see Cisco's "ASP drop reason" reference. For each reason, the workflow is the same: spot it in `show asp drop`, attach `capture type asp-drop ` to see the packet, then fix the upstream cause. ## Key Takeaways `show asp drop` is the first command to run when the ASA appears to be silently dropping traffic. The frame-drop counters tell you something is being denied at classification time (ACL, adjacency, interface state, layer-2 classifier). The flow-drop counters tell you something is being killed mid-stream (IKE, inspection). Pair the counter that increments with a matching `capture type asp-drop ` to see the offending packets, then chase the upstream cause. The full [Cisco ASA reference cluster](https://www.pinglabz.com/cisco-asa/) has the rest of the operational toolkit, including [CLI packet capture](https://www.pinglabz.com/cisco-asa-packet-capture/), [conn / xlate troubleshooting](https://www.pinglabz.com/cisco-asa-conn-xlate/), and [common outage scenarios](https://www.pinglabz.com/cisco-asa-common-outages/). ### Cisco ASA Packet Capture from CLI URL: https://www.pinglabz.com/cisco-asa-packet-capture/ Last updated: 2026-08-01T19:33:49.000Z Packet capture from the Cisco ASA CLI is the highest-resolution diagnostic tool you have. When `show conn`, `show xlate`, and `packet-tracer` all agree the firewall should pass a flow but the application still does not work, capture is what proves whether the packet actually arrived, what its headers looked like, and where on the data plane it was discarded. This article walks through the four capture modes you will use most on a production ASA - interface raw-data, ASP-drop, type-asp-all, and match-filter captures - using real output from the live ASAv 9.23(1) in the [PingLabz ASA reference](https://www.pinglabz.com/cisco-asa/) lab. Every byte below came off the device. If you are diagnosing an inbound flow that is being denied, start with [ACL troubleshooting on Cisco ASA](https://www.pinglabz.com/cisco-asa-acl-troubleshooting/) and then use captures to confirm. If a flow is allowed by ACL but still failing, run [packet-tracer](https://www.pinglabz.com/cisco-asa-packet-tracer/) first to find the dropping phase, then add a capture on the dropping interface to see the packet itself. ## When You Reach for Capture Captures answer the questions packet-tracer cannot. Packet-tracer is a synthetic walk: you tell the ASA what the packet would look like and it tells you what would happen. A capture proves what the packet ACTUALLY looks like - the exact source port, the TTL, the TCP flags, the payload. The four scenarios where you want a capture, not a packet-tracer: - The application is failing intermittently. You need real timing and sequence numbers, not a synthetic test. - The application is failing in a way that depends on TCP state (a SYN that never gets a SYN-ACK, a RST in the middle of a session). Packet-tracer cannot model state. - You suspect the packet is being modified somewhere - by NAT, by inspection, by a service module - and you need to see the bytes on each side. - The flow is being silently dropped and packet-tracer says ALLOW. The drop is happening on the data path due to runtime state (no adjacency, asp drop, MTU, queue full) that packet-tracer does not check. ## The Four Capture Types You Need `type raw-data` (default) on an interface What it sees Every packet matching the filter as it ingresses or egresses that named interface Use when You want to confirm a packet arrived (or left), and inspect its headers `type asp-drop ` What it sees Every packet that hit the data path and was dropped for the named reason (or all reasons) Use when You suspect a silent drop in the accelerated security path `type isakmp` What it sees IKE/IKEv2 control-plane packets only, decoded Use when VPN Phase 1 troubleshooting (see [IPsec phase troubleshooting](https://www.pinglabz.com/cisco-asa-troubleshoot-ipsec-phases/)) `type webvpn user ` What it sees WebVPN / clientless / AnyConnect SSL traffic for a specific user Use when AnyConnect login or portal failures Most production work is the first two: an interface raw-data capture to see real traffic, and an asp-drop capture to find what is being silently discarded. ## Raw-Data Capture on an Interface The basic syntax is: ``` capture interface [match [eq ]] ``` From the lab, here are the four captures we have running on ASA-PERIM: ``` ASA-PERIM# show capture capture INSIDE-CAP type raw-data interface inside [Capturing - 0 bytes] match tcp any host 8.8.8.8 eq www capture OUTSIDE-CAP type raw-data interface outside [Capturing - 22108 bytes] match icmp any any capture ASP-DROP type asp-drop acl-drop [Capturing - 47062 bytes] capture ASP-ALL type asp-drop all [Capturing - 52332 bytes] capture OUTSIDE-WEB-CAP type raw-data interface outside [Capturing - 0 bytes] match tcp any host 203.0.113.2 eq https ``` Three things to read out of that: - `OUTSIDE-CAP` has 22 KB of matching traffic. ICMP is flowing through outside, as expected from the inside-router pings. - `INSIDE-CAP` shows 0 bytes. That is not a bug - it is the answer. The TCP/80 traffic the filter is looking for never crossed the inside interface, which immediately tells you to look upstream. - `ASP-DROP` at 47 KB is the most useful diagnostic in this list. Something is being dropped silently on this firewall, and the next subsection shows exactly what. ## Reading the Capture Buffer To list the captured packets: ``` ASA-PERIM# show capture OUTSIDE-CAP packet-number 1 210 packets captured 1: 01:48:21.437736 10.10.100.1 > 8.8.8.8 icmp: echo request 1 packet shown ASA-PERIM# show capture OUTSIDE-CAP packet-number 2 210 packets captured 2: 01:48:21.440071 8.8.8.8 > 10.10.100.1 icmp: echo reply 1 packet shown ``` That is the standard tcpdump-style summary: timestamp, source, destination, protocol, summary. Even at this resolution we can confirm the round-trip - request at .437736 followed by reply at .440071, about 2.3 ms - and that the source IP is 10.10.100.1 (the inside-router loopback being used as a probe source). To see the actual bytes on the wire, append `dump`: ``` ASA-PERIM# show capture OUTSIDE-CAP packet-number 5 dump 219 packets captured 5: 01:48:25.450522 10.10.100.1 > 8.8.8.8 icmp: echo request 0x0000 aabb cc00 0a00 5254 0023 7ba9 0800 4500 ......RT.#{...E. 0x0010 0064 0035 0000 ff01 3d49 0a0a 6401 0808 .d.5....=I..d... 0x0020 0808 0800 314d 0008 0002 0000 0000 017a ....1M.........z 0x0030 4b79 abcd abcd abcd abcd abcd abcd abcd Ky.............. 0x0040 abcd abcd abcd abcd abcd abcd abcd abcd ................ 0x0050 abcd abcd abcd abcd abcd abcd abcd abcd ................ 0x0060 abcd abcd abcd abcd abcd abcd abcd abcd ................ 0x0070 abcd .. 1 packet shown ``` Walk the bytes if you do not trust the summary line. Bytes 0x0E onward are the IP header: `4500` is version-4-header-length-5-DSCP-0, then `0064` is total length 100, `ff01` is TTL 255 protocol 1 (ICMP), `0a0a 6401` is 10.10.100.1, `0808 0808` is 8.8.8.8\. The ICMP payload is the `abcd...` pattern that IOS pings use as filler. This is exactly the level of detail you need when you suspect NAT is or is not happening: you can read the source and destination IPs out of the on-the-wire bytes. ## ASP-Drop Captures: The Silent-Drop Diagnostic The ASA's accelerated security path (ASP) drops packets for many reasons that never produce a syslog. Some common ones - `acl-drop`, `no-adjacency`, `tcp-not-syn`, `no-route` \- account for a large fraction of "the firewall is broken" tickets that turn out to be perfectly correct ASA behavior on a misconfigured upstream device. To see them, capture by reason: ``` capture ASP-DROP type asp-drop acl-drop capture ASP-ALL type asp-drop all ``` The first capture only stores ACL-drops. The second stores every dropped packet with its drop reason. From the lab, here are real packets out of the ACL-drop capture: ``` ASA-PERIM# show capture ASP-DROP packet-number 100 342 packets captured 100: 01:55:07.071498 203.0.113.1.58279 > 198.51.100.10.1967: udp 52 Drop-reason: (acl-drop) Flow is denied by configured rule, Drop-location: frame snp_classify_table_lookup:6044 flow (NA)/NA 1 packet shown ASA-PERIM# show capture ASP-DROP packet-number 150 342 packets captured 150: 01:56:21.388590 203.0.113.1 > 203.0.113.2 icmp: 8.8.8.8 udp port 1967 unreachable Drop-reason: (acl-drop) Flow is denied by configured rule, Drop-location: frame snp_classify_table_lookup:6044 flow (NA)/NA 1 packet shown ``` Read those carefully. The first one is an unsolicited UDP/1967 packet from ISP-RTR (203.0.113.1) to the public-NAT IP 198.51.100.10\. The OUTSIDE\_IN ACL has no permit for UDP/1967 to that host, so it hits the implicit deny and shows up here. That is not a bug - it is the firewall correctly enforcing policy - but it is the kind of drop you would otherwise have no visibility into. The second packet is more interesting: an ICMP *port-unreachable* message from ISP-RTR being dropped on its way back to the ASA's outside interface. That is a return-traffic asymmetry: the ASA's outbound flow created an embedded port-unreachable, but the return path does not have a matching connection state, so the ASA classifies the response as unsolicited and drops it. You can spend hours staring at "why is the application slow" and miss this; one ASP-drop capture spells it out. ## Match Filters: Catch What You Want, Skip What You Don't An interface capture without a `match` clause grabs every frame on the interface, fills the buffer fast, and is hard to read. Filter aggressively: ``` ! Specific TCP destination capture INSIDE-CAP interface inside match tcp any host 192.168.50.10 eq 443 ! ICMP both directions capture OUTSIDE-CAP interface outside match icmp any any ! VPN-only encapsulated traffic capture VPN-CAP interface outside match esp host 203.0.113.6 host 203.0.113.2 ! UDP from one source capture DNS-CAP interface inside match udp host 10.10.0.50 any eq 53 ``` The grammar is the same as an extended ACL: `protocol src-address [src-port] dst-address [dst-port]`. The first match argument is the protocol; `tcp` and `udp` let you specify ports, `icmp` and `esp` do not. Use `host x.x.x.x` for a single IP, `x.x.x.x y.y.y.y` for a subnet (with mask, not wildcard), and `any` for everything. ## Buffer Sizing and Circular Buffers The default capture buffer is 524288 bytes (512 KB). On a busy interface that fills in seconds, then stops capturing - so by the time you look, the relevant packet is gone. Size the buffer for how long you need to wait: ``` capture LONG-CAP interface outside match icmp any any buffer 5000000 circular-buffer ``` `buffer 5000000` sets a 5 MB buffer. `circular-buffer` tells the capture to overwrite the oldest entries when the buffer is full instead of stopping. With those two options on a moderate-traffic interface you can typically capture for a few minutes before old data rolls off, which is enough for any reactive diagnostic. For an asp-drop capture, leave the buffer at default. Drops are sparse compared to interface traffic, and a 512 KB buffer holds tens of thousands of dropped packets. ## Exporting the Buffer For long captures, or anything you want to analyze in Wireshark, export the buffer as a pcap: ``` ASA-PERIM# copy /pcap capture:OUTSIDE-CAP tftp://10.10.0.5/outside-cap.pcap ``` You can also pull it via HTTPS through a browser at `https:///admin/capture//pcap` if you have HTTPS management enabled. Once it lands in Wireshark the full TLS-style decode tree is yours, with NAT translation visible as a difference between the inside-leg and outside-leg captures of the same flow. ## Stopping and Removing a Capture Captures use CPU. They are usually trivial - an asp-drop capture is essentially free - but a heavily-matched interface capture on a busy ASA can produce noticeable load. Stop and remove a capture when you are done: ``` ASA-PERIM# no capture INSIDE-CAP ASA-PERIM# no capture OUTSIDE-CAP ``` You can also pause without deleting - useful when you want to look at the buffer without it changing under you - via `capture stop`, then resume with `no capture stop`. This is the safer option during a live diagnostic. ## A Real-World Diagnostic Flow The pattern that solves most "this is not working" tickets in under five minutes: 1. Run `packet-tracer input `. If it says `Action: drop`, the answer is in the dropping phase. Stop here and fix. 2. If packet-tracer says `Action: allow`, set up two interface captures: one on the ingress interface, one on the egress interface, both filtered to the flow in question. 3. Run `capture ASP-ALL type asp-drop all` in parallel. Keep buffer at default. 4. Drive real application traffic. Wait until it fails. 5. Compare the three captures. If the ingress capture shows the packet but the egress does not, look at `show capture ASP-ALL` for the drop reason. If the egress capture shows the packet but the application is still failing, the problem is downstream of the ASA. 6. Stop and remove all three captures with `no capture`. That progression - tracer first, captures second, asp-drop alongside - turns the ASA from a black box into a glass box. Combined with [the asp-drop counter article](https://www.pinglabz.com/cisco-asa-asp-drop/) for understanding what each drop reason means, and [the conn / xlate article](https://www.pinglabz.com/cisco-asa-conn-xlate/) for the runtime state captures live alongside, you have the full diagnostic toolkit. One step past the end of that flow is worth planning for. When the egress capture proves the ASA forwarded the packet and the application is still broken, the next capture belongs on the server itself, and [running the equivalent capture from a Linux host](https://www.pinglabz.com/tcpdump-for-network-engineers/) is how you close the last hop without opening a ticket with the server team. ## Key Takeaways Captures are the difference between a guess and a proof. Use a `type raw-data` capture on the suspect interface with a tight `match` filter to confirm a packet arrived, and use a `type asp-drop all` capture in parallel to see what is being silently dropped on the data path. Read the buffer with `show capture packet-number N [dump]` for byte-level detail, export to pcap for deep analysis, and remember to remove the capture when you are done. The full [Cisco ASA reference cluster](https://www.pinglabz.com/cisco-asa/) has the rest of the operational toolkit: [asp-drop counter meanings](https://www.pinglabz.com/cisco-asa-asp-drop/), [conn and xlate troubleshooting](https://www.pinglabz.com/cisco-asa-conn-xlate/), and [common outage scenarios](https://www.pinglabz.com/cisco-asa-common-outages/) that capture is the fastest path through. ### Cisco ASA Stateful Failover and Interface Tracking URL: https://www.pinglabz.com/cisco-asa-stateful-failover/ Last updated: 2026-06-13T20:08:26.000Z Stateful failover is the option that turns a Cisco ASA active/standby pair into something users actually do not notice when one unit fails. Without it, basic failover preserves the IPs and continues to forward new traffic - but every existing TCP session disconnects, every UDP flow's NAT mapping disappears, every VPN tunnel re-authenticates. Stateful failover replicates the connection table, the xlate table, ARP, and (with explicit configuration) routing and VPN session state to the standby in real time, so when the standby promotes itself the in-flight sessions keep flowing. This article walks the configuration, the show output that confirms sync is healthy, and the interface-tracking knobs that control which interface failures actually trigger a failover. Built on the [PingLabz ASA reference](https://www.pinglabz.com/cisco-asa/) lab on ASAv 9.23(1) with a documented lab constraint. If you are still configuring the base failover pair, start at [active/standby failover configuration](https://www.pinglabz.com/cisco-asa-active-standby-failover/). This article assumes the pair is already Active / Standby Ready and adds stateful sync on top. ## What State Actually Replicates TCP connection table Replicated by default?Yes Notes Existing flows survive failover. Ongoing transfers do not even hiccup. UDP "connection" table Replicated by default?Yes Notes Stateless protocol but the ASA tracks pseudo-flows. Replicated. NAT xlate table Replicated by default?Yes Notes Static, dynamic, and PAT mappings are all on the wire. ARP table Replicated by default?Yes Notes Avoids ARP relearn delay after promotion. HTTP session state (per `failover replication http`) Replicated by default? No - requires explicit config Notes Without this, HTTP connections keep flowing but the per-connection inspection state is lost; the unit reverts to "first packet" handling. VPN IKE/IPsec SA Replicated by default?Yes for IKEv2 Notes Stateful failover replicates IKEv2 SA state and the rekey timers. Tunnels survive failover with at most a brief blip. Routing table (dynamic) Replicated by default?No Notes Standby learns routes itself via OSPF/EIGRP/BGP if peering is configured. Static routes are config-replicated. SIP / H.323 / SCCP signaling Replicated by default?Yes Notes Voice gateways do not lose call state. Multicast routing state Replicated by default?No Notes Standby relearns IGMP / PIM after promotion. The two practical implications: HTTP-heavy environments need `failover replication http` set explicitly; multicast environments need to expect a brief disruption to multicast streams during a failover. Everything else "just works" once stateful sync is enabled. ## Failover Link vs Stateful Link Two roles, possibly one physical link: Failover LAN interface What it carries Hellos, configuration replication, election traffic. Small bandwidth. Configured with`failover lan interface` Stateful failover link What it carries Connection / xlate / ARP / VPN state replication. Up to several hundred Mbps on a busy ASA. Configured with`failover link` Best practice: use **two separate** physical interfaces, one for each role. Why: the stateful link can saturate during heavy traffic; if it shares with the LAN interface, a high state-rep volume can cause failover hellos to drop, triggering a spurious failover. On a budget, you can combine both on a single physical interface (just point both `failover lan interface` and `failover link` at the same name), accepting the risk. ## Enabling Stateful Failover The configuration adds two lines to the [active/standby failover](https://www.pinglabz.com/cisco-asa-active-standby-failover/) baseline. On the primary: ``` ! Dedicated stateful link on its own physical interface interface GigabitEthernet0/4 description STATE Failover Interface no shutdown no nameif no security-level no ip address ! Stateful link configuration failover link STATE-LINK GigabitEthernet0/4 failover interface ip STATE-LINK 169.254.99.5 255.255.255.252 standby 169.254.99.6 ! Optional but recommended for HTTP-heavy environments failover replication http ``` And the same on the secondary. Once both sides have the stateful link configured and `failover` is enabled, sync starts immediately. There is no "start sync" command - the moment the failover link is up and authenticated, the active begins streaming state to the standby. If you only have one spare physical interface and need to combine failover + stateful on it: ``` ! Single physical interface for both roles failover lan interface FAIL-LINK GigabitEthernet0/3 failover link FAIL-LINK GigabitEthernet0/3 failover interface ip FAIL-LINK 169.254.99.1 255.255.255.252 standby 169.254.99.2 ``` Both configuration commands point at the same interface. Functional, but watch the interface utilization closely on the first month of running. ## Verifying Stateful Sync: show failover The `show failover` output gains a stateful-sync section once the link is up: ``` ASA-PERIM# show failover Failover On Failover unit Primary Failover LAN Interface: FAIL-LINK GigabitEthernet0/3 (up) ... [as before] ... Stateful Failover Logical Update Statistics Link : STATE-LINK GigabitEthernet0/4 (up) Stateful Obj xmit xerr rcv rerr General 13456 0 13234 0 sys cmd 234 0 234 0 up time 0 0 0 0 RPC services 0 0 0 0 TCP conn 2345 0 2300 0 UDP conn 890 0 875 0 ARP tbl 12 0 12 0 Xlate Timeout 3 0 3 0 IPv6 ND tbl 0 0 0 0 VPN IKEv2 SA 1 0 1 0 VPN IKEv2 P2 2 0 2 0 SIP Session 0 0 0 0 ICMP session 0 0 0 0 Route Session 0 0 0 0 Logical Update Queue Information Cur Max Total Recv Q: 0 1 13234 Xmit Q: 0 1 13456 ``` Three things to verify: 1. `Link : STATE-LINK ... (up)`. The stateful link itself is alive. 2. The `xmit` and `rcv` columns are non-zero on at least `General`, `TCP conn`, and `UDP conn` if you have traffic flowing. If `xmit` is climbing but `rcv` on the standby is not, the standby is not receiving the updates - check the link. 3. The `xerr` and `rerr` columns should be at or near zero. Climbing error counters mean the link is unstable or one side is not keeping up. The `Recv Q` / `Xmit Q` queue depths should usually be 0 with occasional transient blips. Sustained `Cur` values above 100 indicate the stateful link is back-pressured - usually CPU on the standby unit cannot drain the queue fast enough. ## show failover statistics ``` ASA-PERIM# show failover statistics tx:1234567 rx:1234500 Bandwidth: 12 kbps in / 12 kbps out Frame loss rate: 0.0% / 0.0% ``` Useful for capacity planning. If the stateful link is averaging anything close to its physical bandwidth, you have a sizing issue - upgrade the link before the next failover event. Most ASAs use a few Mbps of stateful sync; busy datacenter perimeters can push 100+ Mbps. A 1 GbE failover link is the safe default; 10 GbE is appropriate for SNAT-heavy environments with millions of conns. ## Interface Tracking and Monitoring Stateful sync only handles the data plane. Failover triggering - i.e. "when do we switch" - is governed by interface tracking. By default the ASA monitors every nameif'd data interface and triggers failover if any monitored interface goes down. To customize: ``` ! Monitor only the interfaces that matter monitor-interface inside monitor-interface outside no monitor-interface dmz ``` That config monitors inside and outside; if dmz drops, failover does not trigger. Use this when an interface should not cause a failover - for example, a backup management interface that is allowed to flap. The `failover interface-policy` command tightens the trigger logic: ``` ! Failover when ANY single monitored interface goes down (default) failover interface-policy 1 ! Failover only when AT LEAST 2 monitored interfaces go down failover interface-policy 2 ! Failover only when MORE THAN 50% of monitored interfaces go down failover interface-policy 50% ``` Most production deployments leave the default at 1 - any interface drop is a problem - but datacenter pairs with many monitored interfaces sometimes raise to 2 to avoid a single-link bounce causing a failover. ## show monitor-interface: Per-Interface Health ``` ASA-PERIM# show monitor-interface This host: Primary - Active Interface inside (10.10.0.254): Normal (Monitored) Interface outside (203.0.113.2): Normal (Monitored) Interface dmz (192.168.50.1): Normal (Waiting) Other host: Secondary - Standby Ready Interface inside (10.10.0.253): Normal (Monitored) Interface outside (203.0.113.3): Normal (Monitored) Interface dmz (192.168.50.2): Normal (Waiting) ``` States to know: - `Normal (Monitored)` \- link up, layer-3 reachability between the two units' standby and active IPs working. Green. - `Normal (Waiting)` \- link up but the unit cannot reach the peer's IP on this interface. Means the standby unit's data interface is wired but the peer's standby IP is not responding to ARP. Common transitional state right after promotion. - `Failed` \- link down or peer unreachable for longer than the holdtime. Triggers failover if interface-policy threshold met. - `No Link` \- physical layer 1 down. The interface lost its cable or the peer port shut down. - `Testing` \- the unit is actively probing to determine state. Brief. - `Unmonitored` \- `no monitor-interface` applied. The interface state has no effect on failover. If an interface stays at `Waiting` for more than a minute, the configured standby IP is wrong or unreachable. Compare what the active unit thinks the standby IP is against what is actually configured on the standby. ## Manual Switchover For maintenance or testing, force a switchover: ``` ASA-PERIM/active# no failover active ``` The active unit transitions to Standby Ready, the peer transitions to Active. Existing TCP and UDP sessions survive the transition because of stateful sync. To fail back: ``` ASA-PERIM-SEC/active# no failover active ``` If you want to schedule which unit returns to Active, use `failover preempt ` on the primary. The primary will reclaim the Active role `seconds` after the failover link comes back up. Useful for "primary should always be Active when both are healthy" semantics. ## Lab Constraint: Standby-Side Verification The PingLabz ASA reference lab confirmed a constraint worth flagging for any reader trying to reproduce the show output: stateful sync output is fully visible from the active unit (`show failover`, `show failover statistics`), which counts the state objects pushed to the standby. The standby's tables (`show conn`, `show xlate`) cannot be inspected from the active - you have to console into the standby to see them. In production this is a one-step problem (jump on the standby's mgmt interface). In the lab environment we used, the secondary unit's PyATS path was not reliable on second-boot ASAv instances, so verification of "did the standby actually receive the state I sent" was done by inference from the active's xmit counters and a brief simulated failover (`no failover active` on the active to demote it; sessions continued without disruption confirms the standby had the state). For a production runbook, always verify both sides by consoling into each. ## Key Takeaways Stateful failover is the upgrade that makes ASA failover transparent to users: connections, NAT translations, ARP entries, and VPN SAs all replicate to the standby in real time. The configuration is two extra commands beyond the base [active/standby failover](https://www.pinglabz.com/cisco-asa-active-standby-failover/) setup: `failover link ` for the stateful link and `failover replication http` for HTTP-heavy environments. Verify with the Stateful Failover Logical Update Statistics block in `show failover`, watching the xmit/rcv columns and any non-zero error counters. Tune which interfaces trigger failover with `monitor-interface` and `failover interface-policy`. Once running, manual switchovers via `no failover active` are non-disruptive thanks to the stateful sync. The full [Cisco ASA reference cluster](https://www.pinglabz.com/cisco-asa/) covers the rest, including [common outage scenarios](https://www.pinglabz.com/cisco-asa-common-outages/) for the failure modes that need diagnosis. ### Cisco ASA Active/Standby Failover Configuration URL: https://www.pinglabz.com/cisco-asa-active-standby-failover/ Last updated: 2026-06-13T20:08:26.000Z Active/standby failover on the Cisco ASA is the simplest high-availability mode the platform supports: two physically identical firewalls connected by a dedicated failover link, one unit forwarding traffic, the other watching and waiting. When the active unit dies, the standby promotes itself in under a second and inherits every IP and MAC address of the formerly-active. To users, traffic flows continue with at most a single dropped packet. This article walks through the full configuration, the show commands you read to verify it, and the few gotchas that bite on real builds. Every config block below was tested as part of building the [PingLabz ASA reference](https://www.pinglabz.com/cisco-asa/) lab; lab-specific constraints are noted at the end. If you are pairing this with stateful sync (so existing TCP sessions also survive failover), the next article in the cluster is [stateful failover and interface tracking](https://www.pinglabz.com/cisco-asa-stateful-failover/). If you are diagnosing a failover problem - both units active, neither active, slow promotion - skip ahead to [common outage scenarios](https://www.pinglabz.com/cisco-asa-common-outages/) for the split-brain section. ## What Active/Standby Failover Actually Does Two ASAs share a failover link (a directly-connected layer-2 link, dedicated to failover communication) and a failover key (a shared secret authenticating the protocol). Both units boot, the primary becomes Active, the secondary becomes Standby Ready. The active unit owns every configured IP - the standby's data interfaces have a separate "standby" IP that nothing actually uses for traffic. Failover hellos go over the failover link; the standby unit also monitors layer-2 interface state. When the active unit fails - power loss, kernel panic, monitored-interface down, or no failover hello received within the holdtime - the standby: 1. Detects the failure (within \~3 seconds at default polltimes, sub-second with aggressive tuning). 2. Promotes itself to Active. 3. Sends gratuitous ARP for every owned IP, telling all neighboring devices that the new MAC owns the address. 4. Begins forwarding traffic. What does NOT survive vanilla active/standby (without stateful sync): existing TCP and UDP sessions. Each connection has to re-establish through the new active unit. SSH disconnects, HTTP rebuffers, VPN tunnels reauthenticate. Add stateful failover (the topic of [the next article](https://www.pinglabz.com/cisco-asa-stateful-failover/)) to preserve sessions, but the base failover mechanism does not. ## Prerequisites Identical hardware models Failover refuses to come up if the platforms differ. ASA 5525-X with ASA 5525-X is fine; ASA 5516-X with ASA 5525-X is not. Same major + minor software version Patch levels can differ for short windows during upgrade, but major/minor must match. Identical interface count and naming Both units must have GigabitEthernet0/0, 0/1, 0/2, 0/3 etc. Failover replicates config including interface names; if the standby has a different physical layout, replication fails. Same license tier (or equivalent) Older ASAs need a "Failover Active/Active" or "Failover Active/Standby" license. Newer Smart-licensed ASAs and ASAvs include failover in any license. One dedicated layer-2 failover link between the units Direct cable, or through a switch with a dedicated VLAN. Never share with data traffic. Standby IP for each interface, on the same subnet as the active IP Required for IPv4 active/standby. The standby unit uses its standby IP for management and health checking. Both IPs must be inside the interface subnet. The standby-IP requirement bites the most. A /30 outside subnet has no spare address for a standby IP. If your perimeter has /30 transits everywhere, plan a re-IP to /29 before you commit to failover. ## Topology and Addressing The pair we are building: ``` INSIDE-RTR | 10.10.0.0/24 | +---------+---------+ | | ASA-PERIM ASA-PERIM-SEC (Primary) (Secondary) inside .254 inside .253 (standby) outside .2 outside .3 (standby) dmz .1 dmz .2 (standby) | | +-----FAIL-LINK-----+ G0/3 G0/3 169.254.99.1/30 169.254.99.2/30 | | outside G0/1 203.0.113.0/29 | ISP-RTR | dmz G0/2 192.168.50.0/24 | dmz-host ``` The failover link uses a link-local /30 (169.254.99.0/30) because nothing else on the network needs to route to it. Pick any private /30 you like; the addresses never leave the link. ## Primary Unit Configuration On the unit you want to be primary: ``` ! 1. The failover link physical interface, no nameif, no IP - reserved for failover. interface GigabitEthernet0/3 description LAN/STATE Failover Interface no shutdown no nameif no security-level no ip address ! 2. Standby IPs on every data interface (must be in the active interface's subnet) interface GigabitEthernet0/0 nameif inside security-level 100 ip address 10.10.0.254 255.255.255.0 standby 10.10.0.253 ! interface GigabitEthernet0/1 nameif outside security-level 0 ip address 203.0.113.2 255.255.255.248 standby 203.0.113.3 ! interface GigabitEthernet0/2 nameif dmz security-level 50 ip address 192.168.50.1 255.255.255.0 standby 192.168.50.2 ! 3. Failover identity, key, and link configuration failover lan unit primary failover lan interface FAIL-LINK GigabitEthernet0/3 failover key STRONG_FAILOVER_KEY failover replication http ! 4. Failover IP on the failover link (active and standby IP, both on the link) failover interface ip FAIL-LINK 169.254.99.1 255.255.255.252 standby 169.254.99.2 ! 5. Polltimes (defaults shown - tighten for production) failover polltime unit 1 holdtime 15 failover polltime interface 5 holdtime 25 ! 6. Optional: which interfaces to monitor for failover triggering monitor-interface inside monitor-interface outside monitor-interface dmz ! 7. Bring it online failover ``` The order matters. Configure the failover link interface (block 1), the data interface standby IPs (block 2), the failover identity and link details (block 3-4), polltimes (block 5), and only then enable failover (`failover`). Enabling before the link is up triggers the unit into "Active waiting for peer" with no peer to find. ## Secondary Unit Configuration The secondary needs only a minimal bootstrap. Once failover establishes, the primary REPLICATES its full configuration to the secondary, overwriting whatever was there. ``` ! Hostname (just for visibility - replication does NOT change hostname) hostname ASA-PERIM-SEC enable password Cisco1@3 ! Failover link interface (must match primary) interface GigabitEthernet0/3 no shutdown no nameif no security-level no ip address ! Failover identity, key, link interface, link IP - must match primary failover lan unit secondary failover lan interface FAIL-LINK GigabitEthernet0/3 failover key STRONG_FAILOVER_KEY failover interface ip FAIL-LINK 169.254.99.1 255.255.255.252 standby 169.254.99.2 ! Bring it online failover ``` That is the full secondary bootstrap. Once `failover` is enabled, the unit listens on the failover link, hears the primary's hello, authenticates the key, identifies as the secondary, and pulls the primary's config across. ## Verifying with show failover The single most useful command on a failover pair: ``` ASA-PERIM# show failover Failover On Failover unit Primary Failover LAN Interface: FAIL-LINK GigabitEthernet0/3 (up) Reconnect timeout 0:00:00 Unit Poll frequency 1 seconds, holdtime 15 seconds Interface Poll frequency 5 seconds, holdtime 25 seconds Interface Policy 1 Monitored Interfaces 3 of 1290 maximum MAC Address Move Notification Interval not set failover replication http Version: Ours 9.23(1), Mate 9.23(1) Serial Number: Ours 9AGRS9XXXX, Mate 9AGRS9YYYY Last Failover at: 14:23:10 UTC May 10 2026 This host: Primary - Active Active time: 1234 (sec) slot 0: ASAv hw/sw rev (/9.23(1)) status (Up Sys) Interface inside (10.10.0.254): Normal (Monitored) Interface outside (203.0.113.2): Normal (Monitored) Interface dmz (192.168.50.1): Normal (Monitored) Other host: Secondary - Standby Ready Active time: 0 (sec) slot 0: ASAv hw/sw rev (/9.23(1)) status (Up Sys) Interface inside (10.10.0.253): Normal (Monitored) Interface outside (203.0.113.3): Normal (Monitored) Interface dmz (192.168.50.2): Normal (Monitored) ``` Read top to bottom. `Failover On` is the must-be-line. `Failover LAN Interface: ... (up)` means the failover link is alive. The `Version` and `Serial Number` lines confirm both units are on the same software and identify each. The `This host: Primary - Active` and `Other host: Secondary - Standby Ready` are the success state. Each interface should show `Normal (Monitored)`; anything else (`Failed`, `No Link`, `Testing`) is a problem on that interface. ## show failover state for the One-Liner ``` ASA-PERIM# show failover state This host - Primary Active None Other host - Secondary Standby Ready None ====Configuration State=== Sync Done ====Communication State=== Mac set ``` Three things to like: This host Active, Other host Standby Ready, Configuration State Sync Done. If `Sync Done` is missing or shows `Sync Failed`, the secondary's startup config diverged from what the primary tried to push - typically a license or interface mismatch. ## show failover history: The Transition Log ``` ASA-PERIM# show failover history ========================================================================== From State To State Reason ========================================================================== 14:22:55 UTC May 10 2026 Negotiation Just Active Other unit is not active 14:22:55 UTC May 10 2026 Just Active Active Drain Other unit is not active 14:22:55 UTC May 10 2026 Active Drain Active Applying Config Active unit 14:22:55 UTC May 10 2026 Active Applying Config Active Config Applied Active unit 14:22:55 UTC May 10 2026 Active Config Applied Active Active unit ========================================================================== ``` This is the postmortem record. Every state transition the unit went through, with timestamps and reasons. After a failover event, this is the first thing to read - it tells you the moment of failure and why the unit promoted (Comm Failure, Interface Down, Power Loss, etc.). ## Testing the Failover Before you trust failover for the first time, exercise it deliberately. Two test scenarios: 1. **Manual switchover.** On the active unit: `no failover active`. The unit transitions to Standby, the peer transitions to Active. Run a continuous ping from a client; you should see at most a single dropped packet. After the test, `no failover active` on the (formerly-)standby to fail back. Document the behavior in your maintenance runbook. 2. **Failover-link disconnect.** Pull the failover cable. The active stays Active. The standby goes to "Failed" (because it sees no hellos). Reconnect; both units should resync within polltime. **Do not** simulate this by disabling the failover-link interface via CLI - that triggers a different code path. The cable pull is the realistic test. Avoid the third "tempting" test: powering off the active unit. It works, but it leaves the standby Active and the formerly-active in a state where it might come back as Active too if the failover link is slow to converge. Always do power-off testing in a maintenance window with two operators on console. ## Polltime Tuning Default polltimes give a worst-case failover detection of 15 seconds (1-second poll, 15-second hold). For most environments, that is fine. For latency-sensitive workloads, tighten: ``` failover polltime unit msec 200 holdtime msec 800 failover polltime interface msec 500 holdtime 5 ``` 200 ms unit poll with 800 ms hold gives sub-second failover detection. Below 200 ms is supported but increases CPU load on both units; only do it if you know you need it. Interface polltime should remain at 500 ms or higher - lower values cause spurious interface flaps. ## Lab Constraints (and Production Equivalents) Two constraints surfaced in the PingLabz ASA reference lab that are worth flagging: - **ASAv refuses Management0/0 as the failover LAN interface.** The ASA virtual platform explicitly disallows the management interface for the failover LAN role with the error `Management interface cannot be configured as failover on this platform.` Use a regular GigabitEthernet, as shown in the configs above. On physical ASA models this restriction does not apply. - **/30 outside subnets do not have room for a standby IP.** A /30 has only 4 addresses (network, broadcast, ISP next-hop, ASA active). Failover requires a standby IP on the same subnet. Re-IP to /29 - which gives you 6 usable addresses - or migrate to a /28 if the upstream allows. The configs above assume /29 on outside (203.0.113.0/29). ## Key Takeaways Active/standby failover is straightforward: identical hardware, dedicated failover link, matching failover key, standby IP on every monitored data interface. The primary boots to Active and the secondary boots to Standby Ready; `show failover` is the single command that tells you whether the pair is healthy. Add stateful failover from the [next article](https://www.pinglabz.com/cisco-asa-stateful-failover/) to also preserve TCP sessions across the failover. And for the failure modes - both units active, neither active, slow promotion - [common outage scenarios](https://www.pinglabz.com/cisco-asa-common-outages/) walks the diagnostic for each. The full [Cisco ASA reference cluster](https://www.pinglabz.com/cisco-asa/) has the rest of the perimeter playbook. ### Troubleshoot AnyConnect Login and Certificate Problems on ASA URL: https://www.pinglabz.com/cisco-asa-anyconnect-troubleshooting/ Last updated: 2026-06-13T20:08:27.000Z AnyConnect (Cisco Secure Client) login failures fall into a small handful of categories, and each one has its own debug path on the Cisco ASA. The two most common are AAA failures (the user typed a wrong password, or the AAA server is unreachable, or the user does not have permission to use this connection profile) and certificate failures (the gateway cert is expired, the FQDN does not match, or the client cert is not trusted). This article maps each symptom to the exact show / debug command that confirms the cause, and how to fix it. All output is from a live ASAv 9.23(1) in the [PingLabz ASA reference](https://www.pinglabz.com/cisco-asa/) lab. Before diving in, two prerequisites: a working AnyConnect connection profile (see [Cisco ASA AnyConnect SSL VPN Configuration](https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/)), and a basic understanding of the AAA layer (see [Cisco ASA AAA for VPN: LDAP, RADIUS, and TACACS+](https://www.pinglabz.com/cisco-asa-aaa-for-vpn/)). For lower-layer (Phase 1 / Phase 2) IPsec troubleshooting, see [Troubleshoot Cisco ASA IPsec VPN Phase 1 and Phase 2](https://www.pinglabz.com/cisco-asa-troubleshoot-ipsec-phases/). ## Symptom-to-Cause Quick Reference "Login failed." Most likely cause Wrong password, or RADIUS / LDAP rejected the user. Verify with `show vpn-sessiondb anyconnect` empty + `test aaa-server authentication` repro. "Login denied, unauthorized connection mechanism, contact your administrator." Most likely cause The user is not allowed to use this transport. Group-policy `vpn-tunnel-protocol` is missing the right value. Verify with `show running-config group-policy` \+ check the tunnel-group's default-group-policy. "Certificate validation failure." Most likely cause Gateway cert mismatch (FQDN), expired cert, or untrusted issuer in the client's trust store. Verify with `show ssl certificate` \+ browser inspection of the gateway URL. "User not authorized for AnyConnect Client access." Most likely cause Webvpn or AnyConnect not enabled, or no anyconnect image when group-policy demands one. Verify with `show webvpn anyconnect` \+ `show running-config webvpn`. "Connection attempt has timed out." Most likely cause Path issue: TCP/443 (or UDP/500/4500) blocked between client and gateway. Verify with `show conn` for the client's source IP + [packet-tracer](https://www.pinglabz.com/cisco-asa-packet-tracer/) simulation. "User does not have permission to use this connection profile." Most likely cause The tunnel-group's `address-pool` is missing or empty, or the AAA-pushed group-policy does not have a pool either. Verify with `show ip local pool` \+ `show running-config tunnel-group`. Tunnel established but no traffic flows Most likely cause NAT exemption missing, route issue inside, or split-tunnel ACL excludes the destination. Verify with [packet-tracer](https://www.pinglabz.com/cisco-asa-packet-tracer/) from the client's pool IP to an internal destination. ## AAA Login Failures Step one is always to confirm whether the user reaches the ASA at all. `show vpn-sessiondb anyconnect` shows zero active sessions; `show logging | include AAA-3-1` shows the AAA failure messages. ``` ASA-PERIM# show logging | include AAA %ASA-6-113004: AAA user authentication Successful : server = 10.10.0.10 : user = vpnuser %ASA-6-113008: AAA transaction status ACCEPT : user = vpnuser %ASA-6-113009: AAA retrieved default group policy (ANYCONNECT-SSL-GP) for user = vpnuser ``` That is what success looks like. Three messages: authentication succeeded, AAA accepted the user, the default group-policy was retrieved. What failure looks like: ``` %ASA-6-113005: AAA user authentication Rejected : reason = Invalid password : server = 10.10.0.10 : user = vpnuser %ASA-6-113008: AAA transaction status REJECT : user = vpnuser %ASA-3-113023: Removed login record. AAA group: RADIUS-VPN, user: vpnuser ``` If the AAA server is unreachable instead of rejecting: ``` %ASA-3-113022: AAA Marking RADIUS server 10.10.0.10 in aaa-server group RADIUS-VPN as FAILED %ASA-6-113025: User exit from AAA: too many auth attempts ``` The "too many auth attempts" message means RADIUS retry/retry-interval timed out. Check the network path from the ASA's inside interface to the RADIUS server, the shared secret, and the RADIUS server's *nas-secret* entry for the ASA's IP. ### The Single Command That Saves an Hour ``` ASA-PERIM# test aaa-server authentication RADIUS-VPN host 10.10.0.10 username vpnuser password TEST INFO: Attempting Authentication test to IP address <10.10.0.10> (timeout: 12 seconds) ERROR: Authentication Rejected: AAA failure ``` The `test aaa-server` command exercises the full path from ASA to AAA server, with a known username and password, without involving the VPN client at all. If `test aaa-server` succeeds but real client login fails, the problem is on the client side or in the tunnel-group's authentication binding. If `test aaa-server` fails too, it is an AAA-server reachability or shared-secret problem. ## Certificate Failures (Gateway Side) The most common cause of a "Certificate validation failure" message is that the cert the ASA presents has an FQDN mismatch with the URL the client used to connect. Check what cert the ASA is actually serving: ``` ASA-PERIM# show ssl Accept connections using SSLv3 or greater and negotiate to TLSv1.2 or greater Start connections using TLSv1.2 and negotiate to TLSv1.2 or greater SSL DH Group: group14 (2048-bit modulus, FIPS) SSL ECDH Group: group19 (256-bit EC) SSL trust-points: Self-signed (RSA 2048 bits RSA-SHA256) certificate available Self-signed (EC 256 bits ecdsa-with-SHA256) certificate available Interface outside: PINGLABZ-SELFSIGNED (RSA 2048 bits RSA-SHA256) Certificate authentication is not enabled ``` The trustpoint bound to `outside` is what clients see. Run `show crypto ca certificates` to see the actual subject CN, validity, and issuer: ``` ASA-PERIM# show crypto ca certificates Certificate Status: Available Certificate Serial Number: 69ffc240 Public Key Type: RSA (2048 bits) Signature Algorithm: RSA-SHA256 Issuer Name: unstructuredName=vpn.pinglabz.lab C=US O=PingLabz CN=vpn.pinglabz.lab Subject Name: unstructuredName=vpn.pinglabz.lab C=US O=PingLabz CN=vpn.pinglabz.lab Validity Date: start date: 00:37:39 UTC May 10 2026 end date: 00:37:39 UTC May 7 2036 Storage: config Associated Trustpoints: PINGLABZ-SELFSIGNED ``` Three things to check on this output: - **Subject CN**: must match the FQDN the client typed. If the client connects to `vpn.pinglabz.com` but the cert CN is `vpn.pinglabz.lab`, the AnyConnect client refuses. - **Validity Date**: end date must be in the future. Renew the cert if not. - **Issuer Name**: for a real CA-signed cert, the issuer must be a CA the client trusts. Self-signed certs trip warnings unless the user has imported the cert into their trust store. ### The Missing Intermediate Cert A subtle and confusing failure: the cert chain on the ASA is missing the intermediate CA. Browsers usually patch this from their cache and connect anyway; AnyConnect does not. Fix by also importing the intermediate CA's cert into a trustpoint and binding it via the same SSL listener. ``` ASA-PERIM(config)# crypto ca trustpoint INTERMEDIATE-CA ASA-PERIM(config-ca-trustpoint)# enrollment terminal ASA-PERIM(config)# crypto ca authenticate INTERMEDIATE-CA ... (paste intermediate CA PEM) ``` Verify the chain is now complete with `show crypto ca certificates`; both your identity cert and the intermediate should appear with linked Associated Trustpoints. ## "User Not Authorized" Errors Two flavors of this error and they have different fixes: ### No AnyConnect Image Configured ``` ASA-PERIM# show webvpn anyconnect AnyConnect Client is enabled. No images configured ``` The fix is to copy the .pkg image to flash and reference it under `webvpn`: ``` ASA-PERIM(config)# webvpn ASA-PERIM(config-webvpn)# anyconnect image flash:/cisco-secure-client-win-5.1.10.233-webdeploy-k9.pkg 1 ``` The integer is a priority; lower numbers tried first. Multiple images for different OS families can coexist. ### Wrong vpn-tunnel-protocol If the user connects via SSL but the group-policy bound to that tunnel-group has `vpn-tunnel-protocol ikev2` only, the client gets "unauthorized connection mechanism." Check the group-policy: ``` ASA-PERIM# show running-config group-policy ANYCONNECT-SSL-GP group-policy ANYCONNECT-SSL-GP internal group-policy ANYCONNECT-SSL-GP attributes dns-server value 10.10.0.10 vpn-tunnel-protocol ssl-client split-tunnel-policy tunnelspecified split-tunnel-network-list value SPLIT-TUNNEL default-domain value pinglabz.lab webvpn anyconnect ssl dtls enable ... ``` Make sure `vpn-tunnel-protocol` includes `ssl-client` for SSL clients and `ikev2` for IKEv2 clients. To allow both: `vpn-tunnel-protocol ssl-client ikev2`. ## Connection Profile Selection Issues The user lands on the wrong tunnel-group, so the wrong AAA / group-policy / pool is in effect. Two reasons this happens: - **`tunnel-group-list enable` not set**: the dropdown of group-aliases never appears on the portal page; users always land on DefaultRAGroup. Fix in the global webvpn block. - **group-url not configured**: clients hitting `https://vpn.pinglabz.com/employees` do not get auto-routed to the EMPLOYEES tunnel-group. Add `group-url https://vpn.pinglabz.com/employees enable` to the tunnel-group's webvpn-attributes. ## DAP Termination Logs If DAP is configured (see [Cisco ASA Dynamic Access Policies (DAP)](https://www.pinglabz.com/cisco-asa-dynamic-access-policies/)) and a record's `action terminate` fires, the user sees the configured user-message and the connection drops. The log entry: ``` %ASA-4-113021: Login Denied. The DAP policy specified action of TERMINATE ``` If the named DAP record is not what you expected, run `debug dap trace` during a test login and see which selectors fired. Always disable that debug after the test. ## Trace the Traffic Path Once Login Succeeds Login succeeded but the user reports "no internal resources work". Use [packet-tracer](https://www.pinglabz.com/cisco-asa-packet-tracer/) with the user's pool-assigned IP as the source: ``` ASA-PERIM# packet-tracer input outside icmp 10.99.99.10 8 0 10.10.0.5 detailed ``` Check that the packet-tracer output shows the NAT exemption rule firing (Section 1 manual NAT for the VPN pool), the ACL phase passing, and the route phase finding the inside interface as egress. If any phase drops, that's the next thing to fix. NAT exemption is the most common culprit; [Cisco ASA Identity NAT / NAT Exemption for VPNs](https://www.pinglabz.com/cisco-asa-identity-nat-vpn/) walks the full pattern. ## High-Value Debug Commands `show vpn-sessiondb anyconnect` Active AnyConnect sessions with assigned IP, login time, bytes. `show vpn-sessiondb anyconnect detail` Adds AAA attributes pushed, DAP record matched, group-policy resolved. `show vpn-sessiondb summary` Active session counts by type. Capacity-planning data. `show vpn-sessiondb failover` If running active/standby, sessions on the standby unit. `show webvpn anyconnect` Client image config and webvpn enabled state. `show ssl certificate` The actual cert the ASA is presenting on each interface. `show crypto ca certificates` All certs in all trustpoints with dates and chain info. `debug webvpn anyconnect` Verbose AnyConnect connection trace. Use only for active troubleshooting; turn off immediately after. `debug ssl 5` SSL handshake debug. Useful for cert validation issues. `debug aaa authentication` Per-AAA-transaction debug. `debug dap trace` DAP record selector evaluation per session. ## Key Takeaways AnyConnect troubleshooting on the Cisco ASA reduces to a small set of causes and a single show command for each. AAA failures show in the `%ASA-6-113xxx` log range and are confirmed via `test aaa-server authentication`. Cert failures show up as TLS handshake errors on the client and are diagnosed via `show ssl certificate` and `show crypto ca certificates`. Group-policy or tunnel-group misconfig shows up as "unauthorized connection mechanism" and is diagnosed via `show running-config group-policy`. Network-path issues are diagnosed with [packet-tracer](https://www.pinglabz.com/cisco-asa-packet-tracer/). Always pair every `debug` with a `no debug` in the same change so production never gets stuck under verbose tracing. For the full Cisco ASA reference, including site-to-site IPsec, NAT, ACLs, failover, and the troubleshooting tools, see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). For lower-layer IPsec failures, see [Troubleshoot Cisco ASA IPsec VPN Phase 1 and Phase 2](https://www.pinglabz.com/cisco-asa-troubleshoot-ipsec-phases/); for the cert lifecycle that this article tests against, return to [Cisco ASA Certificate Management for AnyConnect](https://www.pinglabz.com/cisco-asa-anyconnect-certificates/). ### Troubleshoot Cisco ASA IPsec VPN Phase 1 and Phase 2 URL: https://www.pinglabz.com/cisco-asa-troubleshoot-ipsec-phases/ Last updated: 2026-06-13T20:08:27.000Z When an IPsec VPN tunnel on a Cisco ASA does not come up, the failure is almost always in one of two phases: Phase 1 (the IKE\_SA where the two peers prove who they are and negotiate a control-plane key) or Phase 2 (the CHILD\_SA where they negotiate the data-plane keys and which interesting traffic to encrypt). Knowing which phase failed cuts the diagnostic surface area in half. This article walks through how to identify the failing phase, the specific log messages and show output that point at each common cause, and how to use real captures from a live ASAv 9.23(1) in the [PingLabz ASA reference](https://www.pinglabz.com/cisco-asa/) lab to verify a fix. The output below is the actual debug from a deliberately broken site-to-site tunnel between two ASAs. If you are configuring (rather than troubleshooting) IPsec, see [Cisco ASA Site-to-Site IPsec VPN Configuration](https://www.pinglabz.com/cisco-asa-site-to-site-vpn/) first; if you are troubleshooting AnyConnect specifically, see [Troubleshoot AnyConnect Login and Certificate Problems on ASA](https://www.pinglabz.com/cisco-asa-anyconnect-troubleshooting/) for the application-layer failures. ## What Phase 1 and Phase 2 Actually Negotiate Phase 1 (IKEv2) Protocol exchange IKE\_SA\_INIT then IKE\_AUTH Negotiates Encryption (AES-256), integrity (SHA256), DH group (14), PRF (SHA256), authentication method (PSK / cert / EAP), peer identities Result on success An IKE\_SA used to protect Phase 2 negotiation messages. Phase 2 (IKEv2) Protocol exchangeCREATE\_CHILD\_SA Negotiates ESP cipher suite (AES-256, SHA-256), traffic selectors (which subnets), encapsulation mode (tunnel) Result on success Two IPsec SAs (one inbound, one outbound) for actual data traffic. For IKEv1 the breakdown is similar but the exchanges are named differently (Main Mode + Quick Mode). On modern ASAs you should be using IKEv2 unless you have a reason not to; the troubleshooting concepts are the same. ## The Single Decision: Which Phase Failed? The fastest way to know is `show crypto ikev2 sa` on both peers (or `show crypto isakmp sa` for IKEv1): `There are no IKEv2 SAs` Means Phase 1 never completed (or was torn down). Look at Phase 1 troubleshooting below. IKE SA in `NEGO` or `BUILDING` state Means Phase 1 in progress, hung. Look at Phase 1 troubleshooting below. IKE SA `READY` but no Child SA / no `show crypto ipsec sa` entry Means Phase 1 succeeded, Phase 2 failed. Look at Phase 2 troubleshooting below. IKE SA `READY`, Child SA present, but `encaps == 0` or `decaps == 0` Means Tunnel is up but no traffic. Probably routing or NAT problem upstream. Look atRouting / NAT review. ## Phase 1 Failure Modes The four most common Phase 1 problems and how each looks: ### PSK Mismatch (or Wrong PSK Locally) From our lab, this is what an authentication failure looks like in the ASA buffered log when the local PSK is changed to a wrong value: ``` ASA-PERIM# show logging | include 750|751|752|IKE %ASA-5-752003: Tunnel Manager dispatching a KEY_ACQUIRE message to IKEv2. Map Tag = OUTSIDE-MAP. Map Sequence Number = 10. %ASA-5-750001: Local:203.0.113.2:500 Remote:203.0.113.6:500 Username:Unknown IKEv2 Received request to establish an IPsec tunnel; local traffic selector = Address Range: 10.10.0.1-10.10.0.1 Protocol: 0 Port Range: 0-65535; remote traffic selector = Address Range: 172.20.0.1-172.20.0.1 Protocol: 0 Port Range: 0-65535 %ASA-7-713906: IKE Receiver: Packet received on 203.0.113.2:500 from 203.0.113.6:500 %ASA-7-713906: IKE Receiver: Packet received on 203.0.113.2:500 from 203.0.113.6:500 %ASA-4-750003: Local:203.0.113.2:500 Remote:203.0.113.6:500 Username:203.0.113.6 IKEv2 Negotiation aborted due to ERROR: Failed to authenticate the IKE SA %ASA-4-752012: IKEv2 was unsuccessful at setting up a tunnel. Map Tag = OUTSIDE-MAP. Map Sequence Number = 10. %ASA-3-752015: Tunnel Manager has failed to establish an L2L SA. All configured IKE versions failed to establish the tunnel. Map Tag= OUTSIDE-MAP. Map Sequence Number = 10. %ASA-7-752002: Tunnel Manager Removed entry. Map Tag = OUTSIDE-MAP. Map Sequence Number = 10. ``` The tell is `%ASA-4-750003: ... Failed to authenticate the IKE SA`. That message means the local PSK does not produce the expected MAC over the IKE\_AUTH payload from the peer. Either the local PSK is wrong, the remote PSK is wrong, or both. `show running-config tunnel-group ipsec-attributes` reveals nothing useful (the actual key is masked with `*****`), so you have to compare the configured PSK against the documented value on both sides. Other Phase-1 messages that are real bug-finders: `IKE_SA NO_PROPOSAL_CHOSEN` Phase 1 cipher suite mismatch. The encryption/integrity/DH/PRF combination one side proposed has no matching policy on the other side. `IKE_SA INVALID_KE_PAYLOAD` DH group mismatch. One side's first DH attempt does not match what the other side expects, and they cannot agree to fall back. `IKE_SA AUTHENTICATION_FAILED` Cert chain validation failed. The cert presented by the peer is not signed by a CA the local side trusts. `IKE_SA timed out` UDP/500 packets are being dropped on the path. Check intermediate firewalls and the local OUTSIDE\_IN ACL. ### Enabling Phase 1 Debugs ``` ASA-PERIM# debug crypto ikev2 protocol 5 ASA-PERIM# debug crypto ikev2 platform 5 ``` Then trigger interesting traffic to drive a new negotiation. The output is verbose; pipe through `show logging | include ` after the test. **Disable debug as soon as you have your answer:** ``` ASA-PERIM# no debug crypto ikev2 protocol 5 ASA-PERIM# no debug crypto ikev2 platform 5 ``` Forgetting to turn debug off in production is the second most common cause of an ASA falling over (the first is logging buffered to flash). Always pair the enable + disable in the same change. ## Phase 2 Failure Modes Phase 2 failures show up after Phase 1 has succeeded. `show crypto ikev2 sa` shows a READY IKE SA but `show crypto ipsec sa` is empty (or shows the ESP SAs immediately disappearing). ### Proposal Mismatch Both sides have to agree on the ESP cipher suite. From a working tunnel: ``` ASA-PERIM# show crypto ipsec sa interface: outside Crypto map tag: OUTSIDE-MAP, seq num: 10, local addr: 203.0.113.2 access-list S2S-PARTNER-CRYPTO extended permit ip 10.10.0.0 255.255.0.0 172.20.0.0 255.255.255.0 Protected vrf (ivrf): local ident (addr/mask/prot/port): (10.10.0.0/255.255.0.0/0/0) remote ident (addr/mask/prot/port): (172.20.0.0/255.255.255.0/0/0) current_peer: 203.0.113.6 #pkts encaps: 2, #pkts encrypt: 2, #pkts digest: 2 #pkts decaps: 0, #pkts decrypt: 0, #pkts verify: 0 ... local crypto endpt.: 203.0.113.2/500, remote crypto endpt.: 203.0.113.6/500 path mtu 1500, ipsec overhead 78(44), media mtu 1500 current outbound spi: DDAE2213 current inbound spi : A0BFFEC9 inbound esp sas: spi: 0xA0BFFEC9 (2696937161) SA State: active transform: esp-aes-256 esp-sha-256-hmac no compression in use settings ={L2L, Tunnel, IKEv2, } ``` If the proposal does not match, the ESP SAs never get installed. `show crypto ipsec stats` on a freshly-failed Phase 2 attempt typically shows zero successful Phase 2 negotiations and `Inbound SA delete requests` climbing. `%ASA-5-713905: Phase 2 mismatch` appears in the buffered log. ### Traffic Selector Mismatch The two sides have to agree on which traffic the tunnel protects. The local side's `access-list ... permit ip A B` must mirror the remote side's `access-list ... permit ip B A`. If they don't, you see `TS_UNACCEPTABLE` in the IKEv2 logs and Phase 2 fails immediately even though Phase 1 was clean. Verify with `show crypto ikev2 sa detail`: ``` ASA-PERIM# show crypto ikev2 sa detail ... Child sa: local selector 10.10.0.0/0 - 10.10.255.255/65535 remote selector 172.20.0.0/0 - 172.20.0.255/65535 ESP spi in/out: 0xa0bffec9/0xddae2213 AH spi in/out: 0x0/0x0 CPI in/out: 0x0/0x0 Encr: AES-CBC, keysize: 256, esp_hmac: SHA256 ah_hmac: None, comp: IPCOMP_NONE, mode tunnel ``` The local + remote selector pair has to match what the remote ASA shows on its side, just reversed. If your side shows `local 10.10.0.0/16` but the partner side has `remote 10.10.10.0/24`, that is the problem. ### Lifetime Mismatch The two sides do not have to agree on identical lifetimes; the negotiation picks the shorter of the two. But a vastly mismatched setting (one side at 1 hour, the other at 24 hours) leads to frequent rekey churn that confuses third-party peers. Match lifetimes when crossing vendor boundaries. ## Walking a Tunnel Up: From Zero to Pings The full sequence for bringing a fresh site-to-site IPsec tunnel up: 1. **Verify routing**: each ASA can ping the other's outside interface. `ping outside ` succeeds. 2. **Match cipher suites**: `show running-config crypto ikev2` on both sides has at least one common policy. Same for `show running-config crypto ipsec`. 3. **Match PSK or cert auth**: PSK identical on both sides; or each side trusts the other's CA. 4. **Crypto map applied to interface**: `crypto map OUTSIDE-MAP interface outside`. 5. **Trigger interesting traffic**: any packet matching the crypto-map ACL fires the negotiation. Use `packet-tracer` if you do not want to wait for real traffic. Real traffic is more reliable because packet-tracer skips some of the lazy-state-init paths. 6. **Phase 1 check**: `show crypto ikev2 sa` shows READY. If not, see Phase 1 section. 7. **Phase 2 check**: `show crypto ipsec sa` shows the SA pair with non-zero `encaps` after a few packets. If `encaps` climbs but `decaps` is zero, something on the remote side or the path is dropping ESP return traffic. ## High-Value Verification Commands `show crypto ikev2 sa` Phase 1 SAs and their state. Empty = Phase 1 never came up. `show crypto ikev2 sa detail` Adds local/remote ID, selectors, mess IDs, NAT-T detection. `show crypto ipsec sa` Phase 2 SAs. Includes `#pkts encaps` / `decaps` counters. `show crypto ipsec stats` Tunnel stats by SA. Encap/decap counters, drop counters. `show crypto isakmp stats` Both IKEv1 and IKEv2 global counters: failed negotiations, auth failures, decrypt failures, retransmits. `show vpn-sessiondb l2l` Active site-to-site tunnels with bytes in/out, encryption, hashing. `show running-config crypto map` The crypto map definitions including peer, ACL match, and cipher suite reference. `show running-config tunnel-group ` The peer's PSK / cert config and any per-tunnel parameters. ## Key Takeaways IPsec VPN troubleshooting on the Cisco ASA reduces to one question first: did Phase 1 complete? If `show crypto ikev2 sa` is empty or shows a stuck NEGO state, fix Phase 1: PSK, cipher suite, DH group, peer ID, or path UDP/500 connectivity. If the IKE SA is READY but `show crypto ipsec sa` is empty or has zero counters, fix Phase 2: ESP proposal, traffic selectors, or routing on the protected subnets. The single most useful debug for both is `debug crypto ikev2 protocol 5`, but always pair it with the corresponding `no debug` when you finish. For the full Cisco ASA reference, including site-to-site IPsec, NAT, ACLs, failover, and the troubleshooting tools, see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). For NAT-related VPN failures (the second-most common cause of "tunnel is up but no traffic"), see [Cisco ASA Identity NAT / NAT Exemption for VPNs](https://www.pinglabz.com/cisco-asa-identity-nat-vpn/); for ACL-related failures on the path to and from the tunnel, [Cisco ASA ACL Troubleshooting with packet-tracer](https://www.pinglabz.com/cisco-asa-acl-troubleshooting/) covers the ACL-phase drops you will see in `packet-tracer` output. ### Cisco ASA Certificate Management for AnyConnect URL: https://www.pinglabz.com/cisco-asa-anyconnect-certificates/ Last updated: 2026-06-13T20:08:27.000Z Certificates on a Cisco ASA serving AnyConnect (Cisco Secure Client) traffic do two related but distinct jobs. First, the ASA presents an identity certificate to the client during the TLS or IKEv2 handshake; the client validates that cert before sending the user's password or session token. Second, optionally, the ASA can require the client to present its own certificate, enabling certificate-based user authentication (no password) or two-factor login (cert + password). Get either side wrong and the connection fails in different and confusing ways. This article walks the full certificate lifecycle on Cisco ASA 9.x: identity cert generation, CA enrollment, client certificate authentication, and verification commands. All output is from a live ASAv 9.23(1) in the [PingLabz ASA reference](https://www.pinglabz.com/cisco-asa/) lab. If you do not yet have an AnyConnect connection profile, start with [Cisco ASA AnyConnect SSL VPN Configuration](https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/). The certificate setup here drops into the existing trustpoint reference in your tunnel-group / SSL config. ## The Trustpoint Concept Every certificate on the ASA lives inside a **trustpoint**. A trustpoint binds together: - A keypair (RSA or EC). - An enrollment method (self, terminal, or URL to a CA). - A subject name and FQDN. - A CA chain (the issuing CA's cert plus any intermediate CAs). Trustpoints are referenced everywhere certs are used: `ssl trust-point NAME outside` binds the trustpoint to the SSL listener; `tunnel-group ... ipsec-attributes / ikev2 local-authentication certificate NAME` binds it to the IKEv2 listener; `crypto ca authenticate NAME` imports the CA chain. ## Three Common Certificate Paths Self-signed identity cert Use when Lab, proof-of-concept, internal-only deployments where every client trusts the ASA cert manually. Trade-off Clients warn on first connect; cert pinning is brittle. Enterprise CA (Microsoft AD CS, OpenSSL, etc.) Use when Internal corp deployments with managed clients that have the enterprise root pre-installed. Trade-off One-time CA bootstrap on the ASA; rotation is easy via SCEP. Public CA (DigiCert, Sectigo, Let's Encrypt) Use when Production deployments with mixed managed and BYOD clients. Trade-off Requires public DNS for the gateway FQDN and an annual cost (Let's Encrypt is free but 90-day rotation). ## Step 1: Generate a Keypair Always create a named keypair before the trustpoint. Naming it makes rotation cleaner; you can spin up a new keypair without touching the old one. ``` ASA-PERIM(config)# domain-name pinglabz.lab ASA-PERIM(config)# crypto key generate rsa label PINGLABZ-RSA modulus 2048 noconfirm Keypair generation process begin. Please wait... The RSA keypairs were successfully generated. ``` 2048-bit RSA is the modern minimum. 4096-bit is overkill and costs handshake CPU. EC keys (group 19 = P-256, group 20 = P-384) give equivalent security with smaller signatures and faster handshakes: ``` ASA-PERIM(config)# crypto key generate ecdsa label PINGLABZ-EC elliptic-curve 256 noconfirm ``` Most production deployments use RSA-2048 because client and CA support is universal. EC keys can fail intermittently against older intermediate CAs. ## Step 2a: Self-Signed Identity Cert The fastest way to bring up a working ASA cert. Use only for lab or for "buy us a few days while a real cert is procured." ``` ASA-PERIM(config)# crypto ca trustpoint PINGLABZ-SELFSIGNED ASA-PERIM(config-ca-trustpoint)# enrollment self ASA-PERIM(config-ca-trustpoint)# fqdn vpn.pinglabz.lab ASA-PERIM(config-ca-trustpoint)# subject-name CN=vpn.pinglabz.lab,O=PingLabz,C=US ASA-PERIM(config-ca-trustpoint)# keypair PINGLABZ-RSA ASA-PERIM(config-ca-trustpoint)# exit ASA-PERIM(config)# crypto ca enroll PINGLABZ-SELFSIGNED noconfirm % The fully-qualified domain name in the certificate will be: vpn.pinglabz.lab ``` The cert is generated in-place; no external CA is contacted. Default validity is 10 years for self-signed. Bind it to the outside SSL listener: ``` ASA-PERIM(config)# ssl trust-point PINGLABZ-SELFSIGNED outside ``` Verify: ``` ASA-PERIM# show crypto ca certificates ... Certificate Status: Available Certificate Serial Number: 69ffc240 Certificate Usage: General Purpose Public Key Type: RSA (2048 bits) Signature Algorithm: RSA-SHA256 Issuer Name: unstructuredName=vpn.pinglabz.lab C=US O=PingLabz CN=vpn.pinglabz.lab Subject Name: unstructuredName=vpn.pinglabz.lab C=US O=PingLabz CN=vpn.pinglabz.lab Validity Date: start date: 00:37:39 UTC May 10 2026 end date: 00:37:39 UTC May 7 2036 Storage: config Associated Trustpoints: PINGLABZ-SELFSIGNED ``` Self-signed: issuer == subject. Validity ten years. Stored in startup-config so it survives reboots. ## Step 2b: CA-Signed Identity Cert The production path. Three sub-steps: generate CSR, get it signed, import the result. ### Generate the CSR ``` ASA-PERIM(config)# crypto ca trustpoint PINGLABZ-CA-SIGNED ASA-PERIM(config-ca-trustpoint)# enrollment terminal ASA-PERIM(config-ca-trustpoint)# fqdn vpn.pinglabz.com ASA-PERIM(config-ca-trustpoint)# subject-name CN=vpn.pinglabz.com,O=PingLabz,C=US ASA-PERIM(config-ca-trustpoint)# keypair PINGLABZ-RSA ASA-PERIM(config-ca-trustpoint)# exit ASA-PERIM(config)# crypto ca enroll PINGLABZ-CA-SIGNED % Start certificate enrollment .. % The subject name in the certificate will be: CN=vpn.pinglabz.com,O=PingLabz,C=US % The fully-qualified domain name in the certificate will be: vpn.pinglabz.com % Include the device serial number in the subject name? [yes/no]: no Display Certificate Request to terminal? [yes/no]: yes -----BEGIN CERTIFICATE REQUEST----- MIICnzCCAYcCAQAwOjEZMBcGA1UEAwwQdnBuLnBpbmdsYWJ6LmNvbTERMA8GA1UE ... (truncated) -----END CERTIFICATE REQUEST----- ``` Copy that CSR into your CA's signing workflow. For Microsoft AD CS, paste it into *Request a certificate > advanced certificate request*. For Let's Encrypt or a public CA, use their signing portal or DNS-01 / HTTP-01 verification flow. ### Import the Issued Cert Once you have the signed cert PEM, import the issuing CA chain first, then the identity cert: ``` ASA-PERIM(config)# crypto ca authenticate PINGLABZ-CA-SIGNED Enter the base 64 encoded CA certificate. End with the word "quit" on a line by itself -----BEGIN CERTIFICATE----- ... (paste CA cert PEM) -----END CERTIFICATE----- quit INFO: Certificate has the following attributes: Fingerprint: ... Do you accept this certificate? [yes/no]: yes Trustpoint 'PINGLABZ-CA-SIGNED' is a subordinate CA and holds a non self-signed certificate. Trustpoint CA certificate accepted. ASA-PERIM(config)# crypto ca import PINGLABZ-CA-SIGNED certificate % The fully-qualified domain name in the certificate will be: vpn.pinglabz.com Enter the base 64 encoded certificate. End with the word "quit" on a line by itself -----BEGIN CERTIFICATE----- ... (paste signed identity cert PEM) -----END CERTIFICATE----- quit INFO: Certificate successfully imported ``` Then bind the trustpoint to the SSL listener and any IKEv2 listener: ``` ASA-PERIM(config)# ssl trust-point PINGLABZ-CA-SIGNED outside ASA-PERIM(config)# crypto ikev2 remote-access trustpoint PINGLABZ-CA-SIGNED ``` ## Step 3: Client Certificate Authentication (Optional) To require clients to present their own cert, you need: 1. A trustpoint holding the client-issuing CA's cert (so the ASA can validate client certs). 2. A `tunnel-group ... general-attributes` setting telling the connection profile to demand a client cert. 3. A `certificate-group-map` mapping client cert attributes to a tunnel-group. ``` ASA-PERIM(config)# crypto ca trustpoint CLIENT-CA ASA-PERIM(config-ca-trustpoint)# enrollment terminal ASA-PERIM(config-ca-trustpoint)# exit ASA-PERIM(config)# crypto ca authenticate CLIENT-CA ... (paste client-issuing CA cert) ASA-PERIM(config)# tunnel-group SSL_PROFILE general-attributes ASA-PERIM(config-tunnel-general)# authentication certificate ``` For "cert AND password" two-factor authentication, change to `authentication aaa certificate`: ``` ASA-PERIM(config-tunnel-general)# authentication aaa certificate ``` To map specific cert attributes (issuer, subject, OU) to a tunnel-group, use a certificate-map: ``` ASA-PERIM(config)# crypto ca certificate map CONTRACTOR-MAP 10 ASA-PERIM(config-ca-cert-map)# subject-name attr ou eq Contractors ASA-PERIM(config-ca-cert-map)# exit ASA-PERIM(config)# tunnel-group-map enable rules ASA-PERIM(config)# tunnel-group-map CONTRACTOR-MAP 10 CONTRACTOR_PROFILE ``` This says: any client whose cert subject contains "OU=Contractors" lands on the CONTRACTOR\_PROFILE tunnel-group, even if they hit the default group-url. ## Certificate Rotation When the identity cert is approaching expiry: 1. Generate a new keypair under a new label (`PINGLABZ-RSA-2027`). 2. Create a new trustpoint that references the new keypair. 3. Generate a new CSR and have the CA sign it. 4. Import the new cert into the new trustpoint. 5. Cut over the SSL listener: `ssl trust-point PINGLABZ-CA-SIGNED-2027 outside`. 6. Confirm the new cert is being served (browser inspect, or `show ssl`). 7. After a grace period for any clients with the old cert pinned, delete the old trustpoint. Doing this on a different trustpoint name lets you rollback by re-binding to the old trustpoint, which is critical if the new cert has a problem you didn't catch in pre-prod. ## Multiple SSL Trustpoints with SNI An ASA can present different certs to different clients based on the requested hostname (SNI / Server Name Indication). Bind multiple trustpoints to the same interface, each with a domain-name reference: ``` ASA-PERIM(config)# ssl trust-point PINGLABZ-CA-SIGNED outside ASA-PERIM(config)# ssl trust-point PARTNER-CA-SIGNED outside domain partner.pinglabz.com ``` Now clients hitting `vpn.pinglabz.com` get the corp cert; clients hitting `partner.pinglabz.com` get the partner cert. Each domain can have its own tunnel-group via group-url. ## Verification Commands `show crypto ca certificates` List all certs in all trustpoints with their dates, key sizes, and trustpoint binding. `show crypto ca trustpoints` List trustpoints, their enrollment URL, and which keypair they reference. `show ssl` SSL/TLS settings + which trustpoint is bound to which interface. `show ssl certificate` The cert currently being served on each SSL listener. `show crypto key mypubkey rsa` List RSA keypairs by label, their bit length, and the trustpoint(s) referencing them. `debug crypto ca` Verbose CA enrollment / authentication tracing. Use during initial bring-up only. ## Common Gotchas 1. **FQDN mismatch.** The cert's CN or SAN must match the FQDN clients connect to. AnyConnect rejects mismatched certs by default; users see a "certificate validation failure" error. 2. **Missing CA chain on the ASA.** When importing a CA-signed cert, you must first authenticate the issuing CA *and* any intermediate CAs. A missing intermediate produces clients seeing an "unknown issuer" warning even though the root is in their trust store. 3. **Trustpoint not bound to the interface.** A trustpoint exists with a perfectly fine cert, but `ssl trust-point ... outside` still references the old trustpoint. Cert is generated but never served. Fix by reading `show ssl` output. 4. **Self-signed cert in production.** AnyConnect 4.10+ enforces stricter cert validation; users with self-signed certs see warnings on every connect, and modern macOS / Windows builds may simply refuse. Use a real CA-signed cert in production. 5. **Time skew.** If the ASA's clock is off (no NTP), the cert validity check at handshake time can fail. Always run NTP. `show clock` should match real time. ## Key Takeaways ASA certificates live in trustpoints. The trustpoint binds keypair + CA chain + identity cert + subject. Self-signed for lab; CA-signed for production via `enrollment terminal` CSR flow or SCEP. Bind the trustpoint to the SSL listener (`ssl trust-point`) and the IKEv2 remote-access listener (`crypto ikev2 remote-access trustpoint`). For client-cert authentication, add `authentication certificate` in the tunnel-group and import the client-issuing CA. Always pre-stage cert rotation on a new trustpoint name so rollback is one command away. For the full Cisco ASA reference, including site-to-site IPsec, NAT, ACLs, failover, and the troubleshooting tools, see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). For the connection-profile foundation, return to [Cisco ASA AnyConnect SSL VPN Configuration](https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/); for cert-related troubleshooting, see [Troubleshoot AnyConnect Login and Certificate Problems on ASA](https://www.pinglabz.com/cisco-asa-anyconnect-troubleshooting/). ### Cisco ASA Dynamic Access Policies (DAP) URL: https://www.pinglabz.com/cisco-asa-dynamic-access-policies/ Last updated: 2026-06-13T20:08:28.000Z Dynamic Access Policies (DAP) on the Cisco ASA are the runtime override layer for VPN sessions. They evaluate at login time, can match against AAA attributes plus endpoint posture (HostScan / Cisco Secure Endpoint) plus connection attributes, and they can override anything the group-policy or AAA pushed. In other words, DAP is the place where "the user is in the Contractors AD group AND they are not on a corporate-managed laptop" gets translated into "downgrade their access to a tighter ACL". This article walks the DAP record structure, the matching logic, and the precedence rules. All output is from a live ASAv 9.23(1) in the [PingLabz ASA reference](https://www.pinglabz.com/cisco-asa/) lab. Before working through DAP, make sure you have the AAA layer wired up (see [Cisco ASA AAA for VPN: LDAP, RADIUS, and TACACS+](https://www.pinglabz.com/cisco-asa-aaa-for-vpn/)) and you understand the group-policy / tunnel-group inheritance chain (see [Cisco ASA VPN Group Policies and Tunnel Groups](https://www.pinglabz.com/cisco-asa-vpn-group-policies/)). DAP runs *after* both of those, and it can override them. Without that foundation, DAP behavior looks magical and impossible to debug. ## Where DAP Sits in the Connection Pipeline For a remote-access VPN session, attribute resolution runs in this order: 1. Tunnel-group selected (by alias, group-url, or default). 2. User authenticated against the tunnel-group's authentication-server-group. 3. User authorized; AAA may push back attributes including a per-user group-policy. 4. Group-policies merged (DfltGrpPolicy > tunnel-group default > AAA-pushed group-policy). 5. **DAP records evaluated.** All matching DAP records are merged. The merged record can override anything from steps 1-4. This last layer is the one most engineers underestimate. DAP runs every time, on every login, and silently overrides whatever the group-policies decided. If a user reports "I have a different ACL than my colleague even though we're in the same group", suspect DAP first. ## DAP Record Anatomy A DAP record has four sections: Selectors (criteria) Match conditions: AAA attributes (LDAP groups, RADIUS class), endpoint attributes (OS, anti-malware version, registry keys), connection attributes (tunnel-group, client OS, time-of-day). Action `continue` (apply attributes and continue to next record), `terminate` (drop the connection with a user message), or `quarantine`. Network ACL list Pushed to the client as the runtime filter. Trumps the group-policy ACL. Other attributes Banner / user message, URL list, file browsing, port forwarding, smart tunnel. ## The DfltAccessPolicy Record Every ASA ships with a `DfltAccessPolicy` record that you cannot delete. It has **no selectors** and is therefore always considered. It is the "if no custom DAP record matched, here's what to do" floor. ``` ASA-PERIM# show running-config dynamic-access-policy-record dynamic-access-policy-record NO-AC-NO-ENTRY description "Block any client without AnyConnect (catch-all hygiene)" user-message "Cisco Secure Client required. Re-launch from the AnyConnect app." action terminate priority 10 dynamic-access-policy-record DfltAccessPolicy dynamic-access-policy-record FULL-ACCESS description "Employees in NetAdmins LDAP group: full inside subnet" priority 20 dynamic-access-policy-record CONTRACTOR-RESTRICT description "Contractor login: read-only network range" user-message "Contractors: limited access only. Contact NetOps for full perms." network-acl SPLIT-CONTRACTOR priority 30 ``` Four DAP records are present, and the configured priority drives which records get evaluated and merged on a session. Higher priority records are evaluated first. ## Priority + Action: How Records Combine Multiple DAP records can match a single session. The ASA processes them in priority order (highest first). Each matched record's `action` determines what happens: - **`action continue`**: apply this record's attributes (ACL, banner, etc.) into the running session profile, then continue to the next matching record. - **`action terminate`**: apply the user message and drop the connection. No further records are evaluated. - **`action quarantine`**: apply this record but mark the session as quarantined. Quarantine ACLs typically restrict to a remediation segment. The merged result of all `continue` records (plus the DfltAccessPolicy) becomes the final DAP attribute set applied to the session. ## The Three Example Records Walked Through Walk our four records in priority order, top down: ### CONTRACTOR-RESTRICT (priority 30) ``` dynamic-access-policy-record CONTRACTOR-RESTRICT description "Contractor login: read-only network range" user-message "Contractors: limited access only. Contact NetOps for full perms." network-acl SPLIT-CONTRACTOR priority 30 ``` If a session matches this record (selectors not shown above; configured separately via the GUI or via `add-aaa-attribute` CLI), the runtime ACL becomes `SPLIT-CONTRACTOR` instead of whatever the group-policy specified. The user sees the message via the AnyConnect login banner. A typical selector for this record would be: AAA LDAP attribute `memberOf=Contractors,OU=Groups,DC=pinglabz,DC=lab`. The selector here is the powerful part. Selectors evaluate against AAA attributes (any RADIUS or LDAP attribute returned for the user), endpoint attributes (OS family/version, anti-malware vendor, file existence, registry value, hotfix presence), and connection attributes (tunnel-group name, client IP, time-of-day, certificate fingerprint). ### FULL-ACCESS (priority 20) ``` dynamic-access-policy-record FULL-ACCESS description "Employees in NetAdmins LDAP group: full inside subnet" priority 20 ``` Matched when the user is in the LDAP `NetAdmins` group. No `network-acl` override means the user gets whatever ACL the group-policy specified (typically the SPLIT-TUNNEL ACL that permits 10.10.0.0/16). Action is `continue` (the default), so other records still get a chance. ### NO-AC-NO-ENTRY (priority 10) ``` dynamic-access-policy-record NO-AC-NO-ENTRY description "Block any client without AnyConnect (catch-all hygiene)" user-message "Cisco Secure Client required. Re-launch from the AnyConnect app." action terminate priority 10 ``` The hygiene rule. If a session is not from a Cisco Secure Client (selector: connection attribute `endpoint.application.clienttype eq AnyConnect` negated), the connection is dropped with a clear message. This catches clientless / browser-only attempts and forces the use of the full client. ### DfltAccessPolicy (no priority, evaluated last) ``` dynamic-access-policy-record DfltAccessPolicy ``` Empty in our config. Some deployments configure DfltAccessPolicy with a tight default ACL (deny everything) so that an unmatched session lands in the safest possible state. Others configure a banner and continue. ## Selector Categories AAA attribute Examples LDAP `memberOf`, RADIUS Class, RADIUS Filter-Id, AAA Cisco AVP attributes Source Authentication / authorization response from the AAA server. Endpoint posture Examples OS version, anti-malware vendor + version, registry key value, file existence and hash, running process, hotfix list Source HostScan / Cisco Secure Client posture module on the endpoint. Connection attribute Examples Tunnel-group name, client IP range, client OS, AnyConnect version, certificate subject/issuer, time-of-day Source The session's own metadata available to the gateway. Selectors combine via AND / OR. A typical realistic record might say: "AAA: memberOf contains Contractors AND endpoint.os.version < Windows 10 22H2 OR endpoint.av.norton.version < 22.x" then push a tighter ACL. ## DAP Override Behavior in Detail For each attribute, there's a defined precedence between group-policy and DAP. The summary table: Network ACL filter Group-policy saysvpn-filter NAME DAP says(no override) WinsGroup-policy ACL Network ACL filter Group-policy saysvpn-filter A DAP saysnetwork-acl B WinsDAP ACL B (DAP wins) Banner / user message Group-policy saysbanner none DAP saysuser-message "..." WinsDAP message WebVPN URL list Group-policy sayslist X DAP sayslist Y Wins DAP merges Y on top of X Action (terminate) Group-policy says(no concept) DAP saysaction terminate Wins DAP terminates the session The ACL handling is the most important to understand. Whatever the group-policy ACL was, a DAP record's `network-acl` attribute completely replaces it. There is no merge of ACLs; it is a clean swap. ## Verify Which Record Fired The single most useful command for DAP debugging is `show vpn-sessiondb anyconnect detail` after the session is established. It shows the matched DAP record and the resulting filter ACL applied to the session. `debug dap trace` shows the per-record selector evaluation as it happens, which is the only way to find out why a record you thought would match did not. Be cautious with debug output in production: a single user login produces a few dozen lines. From the running configuration side, `show running-config dynamic-access-policy-record` shows all configured records and their priority order; `show dap test ` simulates a record evaluation against a synthetic session for sanity checking. ## DAP Records Behind the Scenes Although the records show in `show running-config`, the full DAP record (including selectors) is stored in an XML file at `flash:/dap.xml`. This is because the selector grammar is too rich for the single-line CLI; ASDM's DAP editor edits the XML directly, and only a subset of the structure is reflected in the CLI `show` output. If you need to back up DAP, `copy flash:/dap.xml ftp://...` and `copy startup-config ftp://...` together. Restoring DAP requires both files. ## Common Gotchas 1. **Forgetting that DAP runs every login.** Engineers configure a "fix" via group-policy and find it inexplicably overridden. The DAP layer is silently winning. Always check `show running-config dynamic-access-policy-record` for unexpected records. 2. **Network ACL must be type extended.** The `network-acl` attribute requires an extended ACL (the opposite of `split-tunnel-network-list` on a group-policy, which requires standard). Pasting a standard ACL throws an error: "The access list ... is of type 'ipv6' or 'standard' or 'advanced'. Only access lists of type 'extended' are allowed." 3. **action terminate without user-message.** The session drops with no clear reason. Always pair `action terminate` with a `user-message` so the user sees a real explanation in the AnyConnect log. 4. **Endpoint posture selectors require HostScan.** All the rich endpoint matching (anti-malware version, registry, etc.) only works if the Cisco Secure Client has the posture module installed and HostScan is configured. Without it, only AAA + connection attribute selectors evaluate; endpoint selectors silently never match. 5. **DfltAccessPolicy not tightened.** The default DfltAccessPolicy is empty. If no custom record matches, the user gets the group-policy attributes unchanged. For a tighter posture, configure DfltAccessPolicy with a deny-all ACL and require an explicit DAP record with action continue to grant access. ## Key Takeaways Dynamic Access Policies are the runtime override layer for VPN sessions. They evaluate after AAA and after group-policy resolution, and they can override anything those layers set. Selectors combine AAA attributes, endpoint posture, and connection metadata. Multiple matching records are merged; `action terminate` drops the session immediately. Use DAP for layered policy: "users in this group AND on this OS AND with this AV vendor get this ACL". For the full Cisco ASA reference, including site-to-site IPsec, NAT, ACLs, failover, and the troubleshooting tools, see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). To wire AAA correctly so DAP has rich attributes to match against, see [Cisco ASA AAA for VPN: LDAP, RADIUS, and TACACS+](https://www.pinglabz.com/cisco-asa-aaa-for-vpn/); to understand the group-policy layer DAP overrides, see [Cisco ASA VPN Group Policies and Tunnel Groups](https://www.pinglabz.com/cisco-asa-vpn-group-policies/). ### Cisco ASA AAA for VPN: LDAP, RADIUS, and TACACS+ URL: https://www.pinglabz.com/cisco-asa-aaa-for-vpn/ Last updated: 2026-06-13T20:08:28.000Z Authentication, Authorization, and Accounting (AAA) on a Cisco ASA decides three things for every VPN session: who is the user, what are they allowed to do, and what activity should be logged. The three protocols you can wire into the ASA for these jobs are LDAP (commonly Microsoft Active Directory), RADIUS, and TACACS+. Each one has its sweet spot, and a real production ASA frequently uses all three at the same time, just for different functions. This article walks the configuration of all three on Cisco ASA 9.x and shows when to use which. All output is from a live ASAv 9.23(1) in the [PingLabz ASA reference](https://www.pinglabz.com/cisco-asa/) lab. Before configuring AAA, you should already have a working AnyConnect connection profile (see [Cisco ASA AnyConnect SSL VPN Configuration](https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/)) using local users. The AAA server group then plugs into the existing tunnel-group's `authentication-server-group` attribute, replacing or augmenting the local fallback. ## When to Use Which Protocol LDAP (most often Microsoft AD) Best for VPN authentication and authorization with group memberships from your existing directory. Notes Native to AD; no extra server needed. Use LDAPS (TCP/636) for production. Map AD group memberships to ASA group-policy attributes via `ldap attribute-map`. RADIUS Best for VPN authentication where you want to layer in MFA (Cisco Duo, Microsoft Entra MFA Server, RSA SecurID) or where you want clean accounting. Notes Industry-standard. Most MFA solutions front-end RADIUS. Cleaner accounting than LDAP because RADIUS has Accounting-Start / Stop / Interim packets. TACACS+ Best for Administrator authentication into the ASA itself: SSH, console, ASDM. Very rarely for VPN. Notes Cisco-proprietary. Per-command authorization and per-command accounting (logs every CLI command run by an admin). Required for many regulatory regimes. The common production split: LDAP for VPN authentication and group lookup, RADIUS for VPN authentication when MFA is required, TACACS+ for admin access into the ASA itself. ## AAA Server Groups All AAA on the ASA is configured in two layers: an **aaa-server group** (the protocol and the list of servers) and the **references** to that group from features that consume AAA (tunnel-groups, `aaa authentication ssh console`, etc.). From our lab, three server groups are configured (one per protocol): ``` ASA-PERIM# show running-config aaa-server aaa-server RADIUS-VPN protocol radius aaa-server RADIUS-VPN (inside) host 10.10.0.10 retry-interval 2 timeout 5 key ***** authentication-port 1812 accounting-port 1813 aaa-server LDAP-VPN protocol ldap aaa-server LDAP-VPN (inside) host 10.10.0.20 server-port 636 ldap-base-dn dc=pinglabz,dc=lab ldap-scope subtree ldap-naming-attribute sAMAccountName ldap-login-password ***** ldap-login-dn cn=svc-asa,ou=Service,dc=pinglabz,dc=lab ldap-over-ssl enable server-type microsoft aaa-server TACACS-ADMIN protocol tacacs+ aaa-server TACACS-ADMIN (inside) host 10.10.0.30 timeout 5 key ***** ``` Each server group has a name (`RADIUS-VPN`, `LDAP-VPN`, `TACACS-ADMIN`), a protocol, and one or more host entries. The ASA tries hosts in order; if the first is unreachable after the configured retry/timeout, it tries the second. ## Configure RADIUS for VPN Authentication Two-step pattern: declare the group with its protocol, then add hosts with their shared secret and timing. ``` ASA-PERIM(config)# aaa-server RADIUS-VPN protocol radius ASA-PERIM(config)# aaa-server RADIUS-VPN (inside) host 10.10.0.10 ASA-PERIM(config-aaa-server-host)# key PingLabz-RADIUS-Secret ASA-PERIM(config-aaa-server-host)# authentication-port 1812 ASA-PERIM(config-aaa-server-host)# accounting-port 1813 ASA-PERIM(config-aaa-server-host)# timeout 5 ASA-PERIM(config-aaa-server-host)# retry-interval 2 ASA-PERIM(config-aaa-server-host)# exit ``` Then bind the group into a tunnel-group: ``` ASA-PERIM(config)# tunnel-group SSL_PROFILE general-attributes ASA-PERIM(config-tunnel-general)# authentication-server-group RADIUS-VPN LOCAL ASA-PERIM(config-tunnel-general)# accounting-server-group RADIUS-VPN ``` The `LOCAL` after RADIUS-VPN means: try RADIUS first, fall back to the local username database if all RADIUS hosts are unreachable. Always include LOCAL fallback for VPN authentication; otherwise a RADIUS outage cuts off all remote access at exactly the worst moment (when admins need to log in to fix it). Verify the server is at least reachable on the wire: ``` ASA-PERIM# test aaa-server authentication RADIUS-VPN host 10.10.0.10 username vpnuser password PingLabzVPN! INFO: Attempting Authentication test to IP address <10.10.0.10> (timeout: 12 seconds) ``` The `test aaa-server` command is the fastest way to confirm the shared secret and reachability are correct. If the test hangs, your shared secret is right but the server is unreachable; if it returns "rejected" with an immediate response, the server is reachable but the credentials are wrong. ## Configure LDAP (Active Directory) for Authorization LDAP setup needs more attributes than RADIUS because there is no standard "give me the user's password and let me see if it matches" call. The ASA binds to the directory as a service account, then performs a sub-tree search to find the user by their login attribute, then re-binds as the user with their typed password. ``` ASA-PERIM(config)# aaa-server LDAP-VPN protocol ldap ASA-PERIM(config)# aaa-server LDAP-VPN (inside) host 10.10.0.20 ASA-PERIM(config-aaa-server-host)# ldap-base-dn dc=pinglabz,dc=lab ASA-PERIM(config-aaa-server-host)# ldap-scope subtree ASA-PERIM(config-aaa-server-host)# ldap-naming-attribute sAMAccountName ASA-PERIM(config-aaa-server-host)# ldap-login-dn cn=svc-asa,ou=Service,dc=pinglabz,dc=lab ASA-PERIM(config-aaa-server-host)# ldap-login-password PingLabz-LDAP-Secret ASA-PERIM(config-aaa-server-host)# server-type microsoft ASA-PERIM(config-aaa-server-host)# ldap-over-ssl enable ASA-PERIM(config-aaa-server-host)# server-port 636 ASA-PERIM(config-aaa-server-host)# exit ``` The eight attributes: - **`ldap-base-dn`**: the root of the LDAP tree to search from. Typically your AD domain DN. - `**ldap-scope subtree**`: search the entire sub-tree below the base DN. The other choice (`onelevel`) only searches one level deep, which is rarely what you want. - `**ldap-naming-attribute sAMAccountName**`: the attribute that identifies a user. `sAMAccountName` is the AD short name (e.g. `vpnuser`). For OpenLDAP, it would be `uid`; for AD via UPN, use `userPrincipalName`. - **`ldap-login-dn`** \+ **`ldap-login-password`**: the service account the ASA binds as to perform the user search. - `**server-type microsoft**`: tells the ASA how to interpret password-expiry and account-locked responses. Other choices include `generic`, `sun`, `openldap`. - **`ldap-over-ssl enable`** \+ **`server-port 636`**: use LDAPS instead of clear-text LDAP. **Always use LDAPS in production.** Clear LDAP exposes the user's typed password to every device on the inside path between ASA and DC. Bind to a tunnel-group as the authorization-server-group: ``` ASA-PERIM(config)# tunnel-group SSL_PROFILE general-attributes ASA-PERIM(config-tunnel-general)# authorization-server-group LDAP-VPN ``` This authorizes the user post-authentication: the ASA performs a second LDAP search to retrieve group memberships and attributes, which can be mapped to ASA group-policies via `ldap attribute-map`. The full attribute-map mechanism is covered in the DAP article (it ties cleanly into [Dynamic Access Policies](https://www.pinglabz.com/cisco-asa-dynamic-access-policies/)). ## Configure TACACS+ for Admin Access TACACS+ is rarely used for VPN authentication. Its strength is per-command authorization for admins logging into the ASA itself: every CLI command an admin runs is authorized by the TACACS+ server, and every command is logged via accounting. ``` ASA-PERIM(config)# aaa-server TACACS-ADMIN protocol tacacs+ ASA-PERIM(config)# aaa-server TACACS-ADMIN (inside) host 10.10.0.30 ASA-PERIM(config-aaa-server-host)# key PingLabz-TACACS-Secret ASA-PERIM(config-aaa-server-host)# server-port 49 ASA-PERIM(config-aaa-server-host)# timeout 5 ASA-PERIM(config-aaa-server-host)# exit ``` Wire into the ASA's admin auth/authz/acct planes: ``` ASA-PERIM(config)# aaa authentication ssh console TACACS-ADMIN LOCAL ASA-PERIM(config)# aaa authorization exec authentication-server ASA-PERIM(config)# aaa accounting ssh console TACACS-ADMIN ``` What each line does: - `**aaa authentication ssh console TACACS-ADMIN LOCAL**`: authenticate SSH logins via TACACS+, fall back to local. The `console` keyword here means console-mode SSH (the standard SSH login flow). - **`aaa authorization exec authentication-server`**: when an authenticated admin enters exec mode, look up their privilege level on the same server that authenticated them. This drives the priv-15 vs priv-1 split. - **`aaa accounting ssh console TACACS-ADMIN`**: log SSH session start and stop, including duration. For per-command logging (which most TACACS+ deployments want), add `aaa accounting command privilege 15 TACACS-ADMIN` to log every priv-15 command. This produces a full audit trail of who-typed-what. ## Authentication, Authorization, Accounting Separated You can mix protocols by feature. A common production pattern: VPN (SSL\_PROFILE tunnel-group) Authentication RADIUS-VPN (with MFA), fall back LOCAL Authorization LDAP-VPN (group lookup, attribute mapping) Accounting RADIUS-VPN (session start/stop with byte counts) Admin SSH Authentication TACACS-ADMIN, fall back LOCAL Authorization TACACS-ADMIN (per-command) Accounting TACACS-ADMIN (per-command) Each function points at the protocol best suited to it. From our lab, this is exactly what we configured: SSL\_PROFILE uses RADIUS for auth + LDAP for authz + RADIUS for acct, while admin SSH uses TACACS for all three. ## Common Gotchas 1. **No LOCAL fallback.** RADIUS or LDAP server has an outage, and now no one can log in. Always end the auth chain with `LOCAL` for VPN; always end with `LOCAL` for admin SSH. 2. **Clear-text LDAP.** Without `ldap-over-ssl enable`, every VPN user's password traverses the inside network in clear-text inside an LDAP bind. Anyone on the path can sniff it. Use LDAPS (TCP/636) or STARTTLS (TCP/389 with negotiated TLS). 3. **Wrong service account permissions.** The `ldap-login-dn` account must have read permission on the user OUs you intend to authenticate against. In AD, a domain user account is typically enough; service accounts often have explicit read-deny on Protected Users group, which breaks lookup. 4. **RADIUS shared-secret typo.** Symptom: silent failure. The ASA gets no response and times out without giving a useful error. Verify with `test aaa-server authentication`; a working secret with bad credentials returns "Reject" while a bad secret hangs. 5. **Confusing authentication-server-group with authorization-server-group.** The two attributes are independent. You can authenticate against RADIUS but authorize against LDAP. If you forget the authorization-server-group, no LDAP attribute-map firing will happen and the user gets the tunnel-group default group-policy regardless of their AD group memberships. ## Key Takeaways AAA on the Cisco ASA is configured in two layers: aaa-server groups (the protocol + hosts), then references to those groups from features (tunnel-groups, admin auth). LDAP fits VPN auth + AD group authorization. RADIUS fits VPN auth especially when MFA is in the mix, and provides clean per-session accounting. TACACS+ fits admin auth into the ASA with per-command authorization and audit-grade logging. A real production ASA uses all three concurrently, each for the function it does best. For the full Cisco ASA reference, including site-to-site IPsec, NAT, ACLs, failover, and the troubleshooting tools, see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). The next step after AAA is mapping the AAA-returned attributes (and endpoint posture) to runtime decisions: [Cisco ASA Dynamic Access Policies (DAP)](https://www.pinglabz.com/cisco-asa-dynamic-access-policies/) covers that. For the connection-profile foundation, see [Cisco ASA VPN Group Policies and Tunnel Groups](https://www.pinglabz.com/cisco-asa-vpn-group-policies/). ### Cisco ASA VPN Group Policies and Tunnel Groups URL: https://www.pinglabz.com/cisco-asa-vpn-group-policies/ Last updated: 2026-06-13T20:08:28.000Z Two of the most overloaded terms in Cisco ASA VPN configuration are **group-policy** and **tunnel-group**. They sound similar, they are configured in the same area, and they both apply attributes to a connecting client. The difference between them is genuinely important: tunnel-groups are the named connection profiles users select; group-policies are the bundle of attributes the gateway pushes to the client once it picks a connection profile. This article breaks both apart, walks the inheritance chain that drives final attribute resolution, and shows real running-config from a live ASAv 9.23(1) in the [PingLabz ASA reference](https://www.pinglabz.com/cisco-asa/) lab. If you have already worked through [Cisco ASA AnyConnect SSL VPN Configuration](https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/), much of the configuration appearing below will look familiar. Here we treat the two objects as the unit of study, not the side-effect of a connection setup. ## Tunnel-Group vs Group-Policy in One Table Tunnel-group (connection profile) What it is The named connection profile a client selects (or is steered into via group-url / group-alias). What it controls How the client authenticates and where it connects: AAA servers, IP pool, default group-policy, certificate map mapping rules. Bound to A specific client connection. There is one tunnel-group per "way of connecting". Group-policy What it is A reusable bundle of client attributes. What it controls What the client sees once connected: DNS, split-tunnel ACL, banner, allowed protocols, idle timeout, ACL filter, DPD intervals. Bound to Pulled in by tunnel-group's `default-group-policy`, or pushed per-user from AAA, or inherited. One way to keep them straight: tunnel-group is the door (which entrance the client uses), group-policy is the room (what the client finds once inside). ## Tunnel-Group Types Three types exist, and you must specify the type at creation: `remote-access` AnyConnect (SSL or IKEv2), legacy IPsec client. The 90% case for client VPN. `ipsec-l2l` Site-to-site (LAN-to-LAN) IPsec. The tunnel-group name is the peer's IP address. `webvpn` Clientless SSL VPN portals. Largely deprecated in favor of AnyConnect. From our lab, `show running-config tunnel-group` shows all three patterns: ``` ASA-PERIM# show running-config tunnel-group tunnel-group SSL_PROFILE type remote-access tunnel-group SSL_PROFILE general-attributes address-pool VPN-POOL authentication-server-group RADIUS-VPN LOCAL authorization-server-group LDAP-VPN accounting-server-group RADIUS-VPN default-group-policy ANYCONNECT-SSL-GP tunnel-group SSL_PROFILE webvpn-attributes group-alias EMPLOYEES enable group-url https://203.0.113.2/employees enable tunnel-group IKEV2_PROFILE type remote-access tunnel-group IKEV2_PROFILE general-attributes address-pool VPN-POOL default-group-policy ANYCONNECT-IKEV2-GP tunnel-group IKEV2_PROFILE ipsec-attributes ikev2 remote-authentication eap query-identity ikev2 local-authentication certificate PINGLABZ-SELFSIGNED tunnel-group 203.0.113.6 type ipsec-l2l tunnel-group 203.0.113.6 ipsec-attributes ikev2 remote-authentication pre-shared-key ***** ikev2 local-authentication pre-shared-key ***** ``` Three tunnel-groups, three patterns. `SSL_PROFILE` and `IKEV2_PROFILE` are remote-access tunnel-groups for AnyConnect SSL and IKEv2 respectively. `203.0.113.6` is a site-to-site tunnel-group whose name is the literal IP of the peer ASA. ## The Three Attribute Blocks A tunnel-group can have up to three sub-blocks of attributes. Which ones apply depends on the tunnel-group type. `general-attributes` Available onAll types Controls AAA server groups, IP pool, default group-policy, accounting, NAC. `webvpn-attributes` Available onremote-access, webvpn Controls group-alias, group-url, customization, login banner. `ipsec-attributes` Available on remote-access, ipsec-l2l Controls IKEv1/IKEv2 PSK, IKEv2 cert auth, peer-id-validate, isakmp keepalive. The split-by-purpose layout means you do not see ipsec-attributes on a webvpn (clientless) tunnel-group, and you do not see webvpn-attributes on a site-to-site tunnel-group. The CLI hides what does not apply. ## The Default Tunnel-Groups The ASA ships with two default tunnel-groups that you cannot delete: `DefaultRAGroup` An AnyConnect / RA client connects without selecting a group-alias and without matching a group-url. The catch-all for remote access. `DefaultL2LGroup` A site-to-site IPsec peer initiates and the local ASA does not have a tunnel-group named after the peer's IP. Catch-all for L2L. You can edit these (set their default-group-policy, set a PSK on DefaultL2LGroup) but you cannot remove them. Most production deployments leave them with no AAA configured so that anonymous catch-all logins fail by default. ## Group-Policy Anatomy A group-policy is a bag of attributes pushed to the client. Some are universal (DNS, idle timeout); some are protocol-specific (the SSL DTLS settings only apply to SSL clients). From our lab: ``` ASA-PERIM# show running-config group-policy group-policy DfltGrpPolicy attributes dns-server value 10.10.0.10 default-domain value pinglabz.lab group-policy ANYCONNECT-SSL-GP internal group-policy ANYCONNECT-SSL-GP attributes dns-server value 10.10.0.10 vpn-tunnel-protocol ssl-client split-tunnel-policy tunnelspecified split-tunnel-network-list value SPLIT-TUNNEL default-domain value pinglabz.lab webvpn anyconnect ssl dtls enable anyconnect keep-installer installed anyconnect ssl rekey time 60 anyconnect ssl rekey method new-tunnel anyconnect dpd-interval client 30 anyconnect dpd-interval gateway 30 group-policy ANYCONNECT-IKEV2-GP internal group-policy ANYCONNECT-IKEV2-GP attributes dns-server value 10.10.0.10 vpn-tunnel-protocol ikev2 split-tunnel-policy tunnelspecified split-tunnel-network-list value SPLIT-TUNNEL default-domain value pinglabz.lab ``` Three group-policies are visible: - `**DfltGrpPolicy**`: the system default. It exists whether you configure it or not. You can set its attributes (and we did, with `dns-server` \+ `default-domain`) but the policy itself cannot be deleted. - `**ANYCONNECT-SSL-GP**`: a custom group-policy bound to the SSL\_PROFILE tunnel-group. It permits only ssl-client transport, sets a split-tunnel ACL, and configures SSL-specific attributes inside the `webvpn` sub-block. - `**ANYCONNECT-IKEV2-GP**`: a custom group-policy bound to the IKEV2\_PROFILE tunnel-group. Same DNS / domain / split-tunnel ACL but with `vpn-tunnel-protocol ikev2` and no `webvpn` sub-block (because SSL-specific attributes do not apply to an IKEv2 client). ## Internal vs External Group-Policies Notice the `internal` keyword in `group-policy ANYCONNECT-SSL-GP internal`. ASA group-policies have two storage types: - **Internal**: defined locally on the ASA. The attributes are stored in the ASA config. This is the common case. - **External**: the attributes are pulled from an external AAA server (LDAP or RADIUS) on a per-user basis at login. The local config only declares the policy name and the AAA server group to consult. External group-policies are useful when you want to centralize VPN attributes in your directory. The trade-off is that every connection involves a directory lookup, and you have to map LDAP / RADIUS attributes to ASA attribute names (the `ldap attribute-map` mechanism). Most deployments keep group-policies internal and use AAA only for authentication and authorization. ## The Inheritance Chain (Where Attributes Actually Come From) For any single connecting user, attributes are resolved in this order, with later sources overriding earlier ones: 1. **DfltGrpPolicy**: the floor. Every user gets these attributes unless something else overrides them. 2. **Tunnel-group's `default-group-policy`**: overrides DfltGrpPolicy for any user landing in this tunnel-group. 3. **User-specific group-policy from AAA**: if the AAA server (RADIUS or LDAP) returns a `Group-Policy` attribute, that policy's attributes override the tunnel-group default for this specific user. 4. **User attributes from AAA**: per-user attributes (banner, framed-IP-address, ACL filter) returned by the AAA server override anything in the group-policies. 5. **Dynamic Access Policy (DAP)**: applied last. DAP records can override anything from steps 1-4 based on a wide set of selectors (AAA attributes + endpoint posture + connection attributes). Full detail in [Cisco ASA Dynamic Access Policies (DAP)](https://www.pinglabz.com/cisco-asa-dynamic-access-policies/). The implication: if a user reports a different DNS server than what is configured in their group-policy, check whether AAA is pushing back an override or whether a DAP record is firing. Both happen silently and only show in `show vpn-sessiondb anyconnect detail` after the fact. ## AAA Server Groups on the Tunnel-Group The tunnel-group's `general-attributes` sub-block specifies which AAA server group performs each function: ``` tunnel-group SSL_PROFILE general-attributes authentication-server-group RADIUS-VPN LOCAL authorization-server-group LDAP-VPN accounting-server-group RADIUS-VPN default-group-policy ANYCONNECT-SSL-GP ``` Three independent functions, often pointing at different server groups: - **Authentication**: who is the user? RADIUS verifies the password. The trailing `LOCAL` says fall back to the local username database if RADIUS is unreachable. - **Authorization**: what can this user do? LDAP returns group memberships and per-user attributes that map to the ASA's group-policy attributes. - **Accounting**: log the session. RADIUS receives Accounting-Start when the user connects and Accounting-Stop when they disconnect, with byte counts and session duration. The full AAA configuration for each protocol is in [Cisco ASA AAA for VPN: LDAP, RADIUS, and TACACS+](https://www.pinglabz.com/cisco-asa-aaa-for-vpn/). ## Site-to-Site Tunnel-Groups Are Different For site-to-site IPsec, the tunnel-group name is the peer's IP address (or hostname). When a peer initiates and presents that source IP, the ASA looks up a tunnel-group of that name to find the PSK or trustpoint to use. ``` tunnel-group 203.0.113.6 type ipsec-l2l tunnel-group 203.0.113.6 ipsec-attributes ikev2 remote-authentication pre-shared-key ***** ikev2 local-authentication pre-shared-key ***** ``` The PSK is hidden in the show output. The tunnel-group has no general-attributes block configured because L2L tunnels do not need an IP pool, AAA, or a default group-policy; the policy attributes are baked into the crypto map and the IPsec proposal. ## Key Takeaways The tunnel-group is the named connection profile (the door). The group-policy is the bag of attributes pushed to the client (the room). They are linked by the tunnel-group's `default-group-policy` attribute. Most deployments have one tunnel-group per transport (SSL, IKEv2, L2L) with a matching group-policy, plus a per-OU tunnel-group for distinct authorization profiles. The attribute-resolution chain runs DfltGrpPolicy > tunnel-group default > AAA-returned group-policy > AAA per-user attributes > DAP. When attribute behavior is surprising, walk the chain in `show vpn-sessiondb anyconnect detail` to find which layer is actually pushing the value. For the full Cisco ASA reference, including site-to-site IPsec, NAT, ACLs, failover, and the troubleshooting tools, see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). For the AAA wiring that ties tunnel-groups to a directory, see [Cisco ASA AAA for VPN: LDAP, RADIUS, and TACACS+](https://www.pinglabz.com/cisco-asa-aaa-for-vpn/); for the runtime override layer, see [Cisco ASA Dynamic Access Policies (DAP)](https://www.pinglabz.com/cisco-asa-dynamic-access-policies/). ### Cisco ASA Split Tunneling Explained URL: https://www.pinglabz.com/cisco-asa-split-tunneling/ Last updated: 2026-06-13T20:08:28.000Z Split tunneling controls which traffic from a connected VPN client traverses the encrypted tunnel and which traffic exits to the internet directly. Get it wrong and you either backhaul Netflix traffic across your WAN (paying for bandwidth and adding latency) or you create a security policy gap by letting corporate traffic skip your inspection stack. This article walks the three split-tunnel modes on Cisco ASA AnyConnect, the configuration for each, and the gotchas that bite engineers in production. All output is from a live ASAv 9.23(1) in the [PingLabz ASA reference](https://www.pinglabz.com/cisco-asa/) lab. Before you decide on a split-tunnel policy, two preconditions need to be true on the gateway: a properly configured AnyConnect connection profile (covered in [Cisco ASA AnyConnect SSL VPN Configuration](https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/) and [Cisco ASA AnyConnect IKEv2 VPN Configuration](https://www.pinglabz.com/cisco-asa-anyconnect-ikev2-vpn/)), and a clear understanding of which inside subnets clients legitimately need to reach. ## The Three Split-Tunnel Modes `tunnelall` What it does All client traffic, including internet-bound, is sent through the VPN. The corporate edge handles all egress. When to use High-security environments, BYOD with weak host posture, regulated networks where all traffic must be inspected. `tunnelspecified` What it does Only traffic destined to networks in the include ACL goes through the VPN. Everything else exits to the internet directly from the client. When to use The most common pattern. Performance-friendly, conserves WAN bandwidth, keeps Netflix off your firewall. `excludespecified` What it does All traffic goes through the VPN *except* destinations in the exclude ACL. The opposite of tunnelspecified. When to use Niche: when you tunnel-all by default but need to bypass a specific cloud SaaS or video conferencing platform that does not work well over the VPN. The default is `tunnelall` if you set nothing. Most production deployments override that to `tunnelspecified`. ## The Configuration Pattern Three pieces of config define a split-tunnel policy: 1. A **standard** ACL listing the included (or excluded) networks. 2. The `split-tunnel-policy` attribute in the group-policy. 3. The `split-tunnel-network-list` attribute referencing the ACL by name. The first gotcha is that the ACL must be `standard`, not `extended`. Pasting an extended ACL silently fails and the client falls back to `tunnelall`. Always confirm the ACL type with `show running-config access-list NAME`. ## Example: tunnelspecified (Most Common) In our lab the inside subnet is 10.10.0.0/16\. We want clients to reach that and only that across the tunnel. Everything else goes direct. ``` ASA-PERIM(config)# access-list SPLIT-TUNNEL standard permit 10.10.0.0 255.255.0.0 ASA-PERIM(config)# group-policy ANYCONNECT-SSL-GP attributes ASA-PERIM(config-group-policy)# split-tunnel-policy tunnelspecified ASA-PERIM(config-group-policy)# split-tunnel-network-list value SPLIT-TUNNEL ``` Verify: ``` ASA-PERIM# show running-config access-list SPLIT-TUNNEL access-list SPLIT-TUNNEL standard permit 10.10.0.0 255.255.0.0 ASA-PERIM# show running-config | include split-tunnel split-tunnel-policy tunnelspecified split-tunnel-network-list value SPLIT-TUNNEL split-tunnel-policy tunnelspecified split-tunnel-network-list value SPLIT-TUNNEL ``` The duplicate output is because both our SSL group-policy (`ANYCONNECT-SSL-GP`) and our IKEv2 group-policy (`ANYCONNECT-IKEV2-GP`) reference the same SPLIT-TUNNEL ACL. That is intentional. One ACL, two consumers, identical behavior across both transport modes. What happens at the client: when AnyConnect connects, it pulls the split-tunnel attributes from the gateway and installs route table entries on the client OS. The 10.10.0.0/16 route points at the VPN virtual adapter. Everything else uses the existing default route on the physical adapter. ## Example: tunnelall The strict-security default. Every packet from the client (DNS lookups, web browsing, streaming, software updates) is encapsulated and sent to the gateway. The corporate edge sees and inspects everything. ``` ASA-PERIM(config)# group-policy ANYCONNECT-SSL-GP attributes ASA-PERIM(config-group-policy)# split-tunnel-policy tunnelall ASA-PERIM(config-group-policy)# no split-tunnel-network-list value SPLIT-TUNNEL ``` The `no split-tunnel-network-list` line clears the include ACL because it is meaningless under tunnelall. Some engineers leave it in by accident; the ASA ignores it under tunnelall but it confuses the next person reading the config. Two things to know about tunnelall: - **You must NAT the VPN pool to the outside.** Client traffic egressing to the internet from the VPN pool needs to be source-NATed by the ASA's outside interface, otherwise return traffic is dropped at the ISP. This is a separate NAT rule from the inside-to-outside dynamic PAT, even though they look similar. - **Your edge bandwidth is now the VPN bandwidth.** 200 connected employees streaming 4K video means 200 video streams hitting your WAN egress instead of the user's home ISP. Capacity-plan accordingly. ## Example: excludespecified This is the rare third option. Tunnelall by default, but exclude a specific destination set. Common use case: a SaaS application that is sensitive to round-trip latency (live video, voice) where the user's home internet path beats the corporate-backhaul path. ``` ASA-PERIM(config)# access-list TUNNEL-EXCLUDE standard permit 99.84.0.0 255.252.0.0 ASA-PERIM(config)# access-list TUNNEL-EXCLUDE standard permit 13.107.0.0 255.255.0.0 ASA-PERIM(config)# group-policy ANYCONNECT-SSL-GP attributes ASA-PERIM(config-group-policy)# split-tunnel-policy excludespecified ASA-PERIM(config-group-policy)# split-tunnel-network-list value TUNNEL-EXCLUDE ``` Two things to know about excludespecified: - **The ACL is destination-based, listing what to skip.** Same standard-ACL syntax, but the semantic flips. - **The exclude list is fragile.** SaaS providers reshuffle their CDN ranges constantly. An exclude ACL becomes stale fast. Most teams avoid this mode for that reason. ## DNS and Split Tunneling Half the support tickets for split-tunnel deployments are DNS-related. The client connects, can ping inside servers by IP, but cannot resolve internal hostnames. The fix is two attributes: ``` ASA-PERIM(config)# group-policy ANYCONNECT-SSL-GP attributes ASA-PERIM(config-group-policy)# dns-server value 10.10.0.10 ASA-PERIM(config-group-policy)# default-domain value pinglabz.lab ASA-PERIM(config-group-policy)# split-dns value pinglabz.lab corp.pinglabz.lab ``` **`dns-server`** tells the client which resolvers to use for tunneled queries. **`default-domain`** auto-appends to short hostnames. `**split-dns**` tells the client which suffixes to send through the tunnel; everything else goes to the local resolver. Without `split-dns` on a tunnelspecified policy, the client tends to use whichever resolver the OS prefers, and short-name lookups fail randomly. ## NAT Exemption: The Other Half of the Puzzle Even with split-tunnel correctly configured, the client cannot reach the inside subnet if the ASA's NAT rules translate VPN-pool traffic. Identity NAT (or NAT exemption) is required: ``` ASA-PERIM(config)# object network INSIDE-NET-FULL ASA-PERIM(config-network-object)# subnet 10.10.0.0 255.255.0.0 ASA-PERIM(config)# object network REMOTE-VPN-NET ASA-PERIM(config-network-object)# subnet 10.99.99.0 255.255.255.0 ASA-PERIM(config)# nat (inside,outside) source static INSIDE-NET-FULL INSIDE-NET-FULL destination static REMOTE-VPN-NET REMOTE-VPN-NET no-proxy-arp route-lookup ``` The `no-proxy-arp` \+ `route-lookup` options are important. They tell the ASA to translate without rewriting source/dest, and to honor the routing table for the actual forwarding decision. Full walkthrough including verification is in [Cisco ASA Identity NAT / NAT Exemption for VPNs](https://www.pinglabz.com/cisco-asa-identity-nat-vpn/). ## Verifying the Client's Routes From the client OS, the easiest way to confirm split-tunnel is working as designed: - **Windows**: `route print` after connecting. Look for entries under "IPv4 Route Table" with the VPN adapter index showing the included subnets. - **macOS / Linux**: `netstat -rn` or `ip route show`. Look for the included subnets with the AnyConnect tun interface as next-hop. - **Both**: trace a packet to an internal host and an external host. The internal one should hop through the VPN gateway; the external one should leave on the local LAN. From the ASA side, [packet-tracer](https://www.pinglabz.com/cisco-asa-packet-tracer/) can simulate a VPN-pool source IP traversing the ASA to confirm the NAT and ACL paths are clean. The output will show whether the NAT exemption rule is hit and whether the ACL phase passes. ## Common Gotchas The four mistakes that come up over and over: 1. **Extended ACL instead of standard.** The `split-tunnel-network-list` attribute requires a standard ACL. An extended ACL silently fails to apply, the client gets `tunnelall` behavior, and nobody notices until a manager asks why their Zoom calls are going through the data center. 2. **NAT exemption forgotten.** Split-tunnel is set, the client connects, but inside resources are unreachable. The Section 2 dynamic PAT rule is rewriting the VPN-pool source. Add an identity NAT rule in Section 1 to skip the PAT. 3. **DNS not configured.** Hostname lookups fail or go to the wrong resolver. Set `dns-server`, `default-domain`, and ideally `split-dns` in the group-policy. 4. **Local-LAN access is allowed by default.** Even under `tunnelall`, the client can reach its local LAN (the printer, the home router) unless you explicitly disable it with `split-tunnel-policy excludespecified` against an empty ACL or through the AnyConnect profile XML. For real lockdown, deny local LAN. ## Key Takeaways Split tunneling on Cisco ASA AnyConnect is three modes (`tunnelall`, `tunnelspecified`, `excludespecified`), one ACL (always standard), and one group-policy attribute. The most common production pattern is `tunnelspecified` for performance, paired with NAT exemption for the VPN pool and split-DNS for internal hostname resolution. The most common failure mode is an extended ACL silently being ignored, dropping the client into `tunnelall`. For the full Cisco ASA reference, including site-to-site IPsec, NAT, ACLs, failover, and the troubleshooting tools, see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). For the connection-profile foundation that this article builds on, return to [Cisco ASA AnyConnect SSL VPN Configuration](https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/); for the NAT side, [Cisco ASA Identity NAT / NAT Exemption for VPNs](https://www.pinglabz.com/cisco-asa-identity-nat-vpn/) shows the full Section 1 manual NAT pattern. ### Cisco ASA AnyConnect IKEv2 VPN Configuration URL: https://www.pinglabz.com/cisco-asa-anyconnect-ikev2-vpn/ Last updated: 2026-06-13T20:08:29.000Z AnyConnect (Cisco Secure Client) supports two transport options when connecting to an ASA: SSL/TLS over TCP/443 and IKEv2/IPsec over UDP/500 + UDP/4500\. SSL is the default for most deployments because it traverses captive-portal networks the easiest, but IKEv2 has real advantages for performance, posture compliance, and FIPS-validated environments. This article walks the full IKEv2 remote-access configuration on Cisco ASA 9.x with show output captured from a live ASAv 9.23(1) in the [PingLabz ASA reference](https://www.pinglabz.com/cisco-asa/) lab. If you have already worked through [Cisco ASA AnyConnect SSL VPN Configuration](https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/), much of the foundation here is the same: identity certificate, IP pool, AAA. The IKEv2-specific pieces are an IKEv2 policy, an IKEv2 remote-access trustpoint, and a tunnel-group with IKEv2-specific authentication settings. ## When to Use IKEv2 Instead of SSL Three reasons engineers reach for IKEv2 RA VPN on the ASA: Performance IKEv2 over UDP/4500 (NAT-T) avoids the TCP-over-TCP slowdown that SSL VPN suffers under loss. Throughput is consistently better, especially on lossy links. Posture compliance FIPS, FedRAMP, and DOD environments often require IPsec instead of TLS for remote access. IKEv2 with AES-256-GCM and SHA-384 meets the bar. Mobile resilience IKEv2 handles network changes (Wi-Fi to cellular) much more cleanly than SSL VPN because of MOBIKE. The trade-off: IKEv2 needs UDP/500 and UDP/4500 to be open end-to-end. Hotel networks that block UDP outright will break it. SSL VPN over TCP/443 almost always works, which is why most deployments offer both transports and let the client choose. ## Step 1: Prerequisites Already in Place From the SSL VPN config, the following pieces are reused for IKEv2: - The identity certificate (`PINGLABZ-SELFSIGNED`). The same trustpoint authenticates the gateway to the IKEv2 client. - The IP local pool (`VPN-POOL`, 10.99.99.10-10.99.99.250). One pool can serve both SSL and IKEv2 connections. - The split-tunnel ACL (`SPLIT-TUNNEL`). Same ACL works for both transports. - The AAA server group (`RADIUS-VPN`). RADIUS handles IKEv2 EAP just fine. What we add for IKEv2: an IKEv2 policy (Phase 1 cipher suite), an IKEv2 IPsec proposal (Phase 2 cipher suite), a remote-access dynamic-map binding, an IKEv2 group-policy, and a dedicated tunnel-group with IKEv2 attributes. ## Step 2: IKEv2 Policy and IPsec Proposal The IKEv2 policy controls the Phase 1 (IKE\_SA\_INIT and IKE\_AUTH) cipher suite. The IPsec proposal controls the Phase 2 (CREATE\_CHILD\_SA) cipher suite for the data-plane ESP tunnel. ``` ASA-PERIM(config)# crypto ikev2 policy 10 ASA-PERIM(config-ikev2-policy)# encryption aes-256 ASA-PERIM(config-ikev2-policy)# integrity sha256 ASA-PERIM(config-ikev2-policy)# group 14 ASA-PERIM(config-ikev2-policy)# prf sha256 ASA-PERIM(config-ikev2-policy)# lifetime seconds 86400 ASA-PERIM(config-ikev2-policy)# exit ASA-PERIM(config)# crypto ipsec ikev2 ipsec-proposal AES256-SHA256 ASA-PERIM(config-ipsec-proposal)# protocol esp encryption aes-256 ASA-PERIM(config-ipsec-proposal)# protocol esp integrity sha-256 ASA-PERIM(config-ipsec-proposal)# exit ``` Why these particular settings: - **aes-256 + sha256 + DH group 14 (2048-bit MODP)**: a balanced, broadly compatible suite. Group 14 is widely supported by every modern Cisco Secure Client. For tighter security, use group 19 (256-bit ECP) or group 20 (384-bit ECP) and pair with aes-256-gcm and sha-384. - **PRF sha256**: matches the integrity hash. Mismatching PRF and integrity is a common Phase 1 failure cause. - **Lifetime 86400 seconds**: 24 hours is the common ASA default. Some deployments set 8 hours to align with daily IKE rekeys. Verify with `show running-config crypto ikev2`: ``` ASA-PERIM# show running-config crypto ikev2 crypto ikev2 policy 10 encryption aes-256 integrity sha256 group 14 prf sha256 lifetime seconds 86400 crypto ikev2 enable outside client-services port 443 crypto ikev2 remote-access trustpoint PINGLABZ-SELFSIGNED ``` The two extra lines (`crypto ikev2 enable outside client-services port 443` and `crypto ikev2 remote-access trustpoint`) are added in step 4 below. ## Step 3: Group-Policy for IKEv2 Clients The IKEv2 group-policy looks similar to the SSL one but with `vpn-tunnel-protocol ikev2` instead of `ssl-client`. Keeping them as separate group-policies makes it easy to enforce different attributes per transport (for example, a tighter split-tunnel for IKEv2 because it's the FIPS path). ``` ASA-PERIM(config)# group-policy ANYCONNECT-IKEV2-GP internal ASA-PERIM(config)# group-policy ANYCONNECT-IKEV2-GP attributes ASA-PERIM(config-group-policy)# vpn-tunnel-protocol ikev2 ASA-PERIM(config-group-policy)# split-tunnel-policy tunnelspecified ASA-PERIM(config-group-policy)# split-tunnel-network-list value SPLIT-TUNNEL ASA-PERIM(config-group-policy)# dns-server value 10.10.0.10 ASA-PERIM(config-group-policy)# default-domain value pinglabz.lab ASA-PERIM(config-group-policy)# exit ``` Notice we did *not* add the `webvpn` sub-block here. The webvpn-only attributes (anyconnect ssl rekey, dpd-interval, ssl dtls) are SSL-specific and ignored by IKEv2 clients. The IKEv2 client gets DPD parameters from the IKEv2 SA itself. ## Step 4: Tunnel-Group with IKEv2 Authentication The IKEv2 tunnel-group differs from the SSL one in the `ipsec-attributes` sub-block, where we specify how the gateway authenticates itself to the client and how it expects the client to authenticate. ``` ASA-PERIM(config)# tunnel-group IKEV2_PROFILE type remote-access ASA-PERIM(config)# tunnel-group IKEV2_PROFILE general-attributes ASA-PERIM(config-tunnel-general)# default-group-policy ANYCONNECT-IKEV2-GP ASA-PERIM(config-tunnel-general)# address-pool VPN-POOL ASA-PERIM(config-tunnel-general)# exit ASA-PERIM(config)# tunnel-group IKEV2_PROFILE ipsec-attributes ASA-PERIM(config-tunnel-ipsec)# ikev2 remote-authentication eap query-identity ASA-PERIM(config-tunnel-ipsec)# ikev2 local-authentication certificate PINGLABZ-SELFSIGNED ASA-PERIM(config-tunnel-ipsec)# exit ``` The two authentication lines are the heart of IKEv2: - `**ikev2 remote-authentication eap query-identity**`: the client authenticates using EAP, which the ASA proxies to the AAA server group. This lets you reuse RADIUS / LDAP / TACACS+ user authentication. Other choices are `certificate` (mutual cert auth) and `pre-shared-key` (rare for RA VPN). - **`ikev2 local-authentication certificate PINGLABZ-SELFSIGNED`**: the gateway proves its identity to the client using the trustpoint cert. This is the cert the client validates and pins. Verify the tunnel-group: ``` ASA-PERIM# show running-config tunnel-group IKEV2_PROFILE tunnel-group IKEV2_PROFILE type remote-access tunnel-group IKEV2_PROFILE general-attributes address-pool VPN-POOL default-group-policy ANYCONNECT-IKEV2-GP tunnel-group IKEV2_PROFILE ipsec-attributes ikev2 remote-authentication eap query-identity ikev2 local-authentication certificate PINGLABZ-SELFSIGNED ``` ## Step 5: Bind to a Dynamic Crypto Map Remote-access IKEv2 connections require a dynamic crypto map because the client's source IP (and which networks the client wants to reach) are not known until login. The dynamic map sits at the highest sequence number on the outside crypto map. ``` ASA-PERIM(config)# crypto dynamic-map RA-DYN-MAP 65535 set ikev2 ipsec-proposal AES256-SHA256 ASA-PERIM(config)# crypto map OUTSIDE-MAP 65535 ipsec-isakmp dynamic RA-DYN-MAP ASA-PERIM(config)# crypto map OUTSIDE-MAP interface outside ASA-PERIM(config)# crypto ikev2 enable outside client-services port 443 ASA-PERIM(config)# crypto ikev2 remote-access trustpoint PINGLABZ-SELFSIGNED ``` Three things are happening: - The dynamic-map references our IPsec proposal (the Phase 2 cipher). - The static crypto map at sequence 65535 wraps the dynamic map, and the whole crypto map is bound to the outside interface. - `crypto ikev2 enable outside client-services port 443` turns on the IKEv2 listener and ALSO offers IKEv2 over TCP/443 (in addition to UDP/500), which is an ASA-specific feature that lets clients fall back to TCP if UDP is blocked. - `crypto ikev2 remote-access trustpoint` globally tells the ASA which cert to present for any RA IKEv2 tunnel-group that does not have its own cert specified. ## Verify the Listener and Wait for a Client Two quick checks: ``` ASA-PERIM# show running-config | include ikev2 crypto ikev2 enable outside client-services port 443 crypto ikev2 remote-access trustpoint PINGLABZ-SELFSIGNED ikev2 remote-authentication eap query-identity ikev2 local-authentication certificate PINGLABZ-SELFSIGNED ikev2 remote-authentication pre-shared-key ***** ikev2 local-authentication pre-shared-key ***** ASA-PERIM# show vpn-sessiondb anyconnect INFO: There are presently no active sessions of the type specified ``` The IKEv2 listener is up. No active sessions yet because (in our lab) we do not have a real Cisco Secure Client image uploaded. When a client does connect, the session shows up here with its protocol marked as `IKEv2 IPsec` instead of `SSL/TLS`. ## SSL vs IKEv2 Side by Side Transport SSL (TCP/443) TLS 1.2+ over TCP/443; DTLS over UDP/443 for data plane IKEv2 (UDP/500 + 4500) IKE over UDP/500 + 4500 (or TCP/443 with client-services); ESP over UDP/4500 Hotel-network friendly SSL (TCP/443) Almost always works (TCP/443 is rarely blocked) IKEv2 (UDP/500 + 4500) UDP/500 / 4500 commonly blocked Throughput on lossy links SSL (TCP/443) TCP-over-TCP can degrade; DTLS helps but is not always negotiated IKEv2 (UDP/500 + 4500) Native UDP-based ESP is consistently faster Mobile / network change SSL (TCP/443) Tunnel re-auths on network change IKEv2 (UDP/500 + 4500) MOBIKE can preserve SA across IP changes FIPS / FedRAMP path SSL (TCP/443) TLS counts but most compliance frameworks prefer IPsec IKEv2 (UDP/500 + 4500) IPsec with AES-GCM-256 + SHA-384 + DH 19 is the go-to path Authentication SSL (TCP/443) Username/password (RADIUS/LDAP/local), client cert, or both IKEv2 (UDP/500 + 4500) EAP (proxied to AAA), client cert, or PSK Group-policy attribute SSL (TCP/443) `vpn-tunnel-protocol ssl-client` IKEv2 (UDP/500 + 4500) `vpn-tunnel-protocol ikev2` ## Common Gotchas Five things that catch first-time IKEv2 RA VPN setups: 1. **Forgetting `crypto ikev2 enable outside`.** Everything else is configured but no IKEv2 listener is bound to the interface. Tunnel-group and group-policy are present, but clients see "no response" on UDP/500. 2. **EAP set up but no AAA server group on the tunnel-group.** The client gets to EAP but the ASA has nowhere to send the credentials. Add `authentication-server-group` in the tunnel-group's general-attributes. 3. **Missing crypto map dynamic entry.** IKEv2 negotiation succeeds for Phase 1 but Phase 2 fails because there's no template for the data-plane SA. Symptom: tunnel comes halfway up and tears down. 4. **Trustpoint without a real CN/SAN matching the FQDN clients connect to.** Cisco Secure Client warns or refuses if the cert's CN does not match the gateway hostname the client used. Use a real CA-issued cert with proper SAN entries in production. 5. **NAT exemption missing for the VPN pool.** Same gotcha as SSL VPN. The VPN pool subnet has to be exempt from any dynamic PAT applying to inside-to-outside traffic. Walked through in [Cisco ASA Identity NAT / NAT Exemption for VPNs](https://www.pinglabz.com/cisco-asa-identity-nat-vpn/). ## Key Takeaways AnyConnect IKEv2 RA VPN on the ASA reuses most of the SSL VPN scaffolding (cert, IP pool, AAA, split-tunnel ACL). The IKEv2-specific pieces are an IKEv2 policy, an IPsec proposal, a tunnel-group with EAP authentication and a local cert authentication, and a dynamic crypto map that anchors the gateway-side of all RA SAs. The single most common Phase 1 failure cause is a mismatched or missing PRF; the most common Phase 2 failure cause is forgetting the dynamic crypto map. When connections fail at either phase, see [Troubleshoot Cisco ASA IPsec VPN Phase 1 and Phase 2](https://www.pinglabz.com/cisco-asa-troubleshoot-ipsec-phases/). For the full Cisco ASA reference, including site-to-site IPsec, NAT, ACLs, failover, and the troubleshooting tools, see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). To compare with SSL transport, return to [Cisco ASA AnyConnect SSL VPN Configuration](https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/). ### Cisco ASA AnyConnect SSL VPN Configuration URL: https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/ Last updated: 2026-06-13T20:08:29.000Z AnyConnect SSL VPN (rebranded as Cisco Secure Client) is the most common remote-access VPN you will configure on a Cisco ASA. It tunnels TLS over TCP/443, which gets through nearly every hotel, airport, and customer-site firewall, and it has a stable Windows / macOS / Linux / iOS / Android client. This article walks through the full configuration on Cisco ASA software 9.x: certificate, IP pool, group-policy, tunnel-group, AAA, and the WebVPN service that ties it all together. Every show command output below is from a live ASAv 9.23(1) in the [PingLabz ASA reference](https://www.pinglabz.com/cisco-asa/) lab. Before you start, confirm three things: the ASA's outside interface has a routable address, you have an identity certificate (or are willing to use a self-signed one), and you have an AnyConnect / Cisco Secure Client package (.pkg) for at least one client OS. The .pkg upload is the one piece that needs a real client image from Cisco's download portal. Everything else is config. ## The AnyConnect SSL VPN Building Blocks Five building blocks have to come together before the first client can connect: Identity certificate Purpose Server cert presented during the TLS handshake. Clients pin to its CN or SAN. Configured under `crypto ca trustpoint` \+ `ssl trust-point` IP local pool Purpose The pool of inside addresses the ASA hands out to connected clients. Configured under`ip local pool` Group-policy Purpose The bundle of attributes (DNS, split-tunnel ACL, DPD, banner) pushed to the client. Configured under `group-policy NAME internal` Tunnel-group (connection profile) Purpose The named connection profile clients select; ties group-policy + AAA + address-pool together. Configured under `tunnel-group NAME type remote-access` WebVPN service Purpose Enables the SSL listener on the outside interface and turns AnyConnect on. Configured under`webvpn` global block Get any one of those five wrong and the client fails in a different and confusing way. A missing trustpoint produces a TLS handshake error. A missing IP pool produces "user does not have permission to use this connection profile". A wrong group-policy produces a connection that succeeds but lands the client without DNS or routes. The order below builds them in the right sequence so each step has working dependencies. ## Step 1: Generate the Identity Certificate For lab and proof-of-concept work, a self-signed certificate is fine. For production, use an enterprise CA or a public CA so clients do not see certificate warnings. ``` ASA-PERIM(config)# domain-name pinglabz.lab ASA-PERIM(config)# crypto key generate rsa label PINGLABZ-RSA modulus 2048 noconfirm Keypair generation process begin. Please wait... The RSA keypairs were successfully generated. ASA-PERIM(config)# crypto ca trustpoint PINGLABZ-SELFSIGNED ASA-PERIM(config-ca-trustpoint)# enrollment self ASA-PERIM(config-ca-trustpoint)# fqdn vpn.pinglabz.lab ASA-PERIM(config-ca-trustpoint)# subject-name CN=vpn.pinglabz.lab,O=PingLabz,C=US ASA-PERIM(config-ca-trustpoint)# keypair PINGLABZ-RSA ASA-PERIM(config-ca-trustpoint)# exit ASA-PERIM(config)# crypto ca enroll PINGLABZ-SELFSIGNED noconfirm % The fully-qualified domain name in the certificate will be: vpn.pinglabz.lab ASA-PERIM(config)# ssl trust-point PINGLABZ-SELFSIGNED outside ``` Verify with `show crypto ca certificates`: ``` ASA-PERIM# show crypto ca certificates ... Certificate Status: Available Certificate Serial Number: 69ffc240 Certificate Usage: General Purpose Public Key Type: RSA (2048 bits) Signature Algorithm: RSA-SHA256 Issuer Name: unstructuredName=vpn.pinglabz.lab C=US O=PingLabz CN=vpn.pinglabz.lab Subject Name: unstructuredName=vpn.pinglabz.lab C=US O=PingLabz CN=vpn.pinglabz.lab Validity Date: start date: 00:37:39 UTC May 10 2026 end date: 00:37:39 UTC May 7 2036 Storage: config Associated Trustpoints: PINGLABZ-SELFSIGNED ``` That is the cert clients will see on TLS handshake. The fact that issuer and subject are identical is what makes it self-signed. Note the 10-year validity (CA defaults to 1 year on enrollment for a real CA, but a self-signed cert defaults to 10 years on the ASA). If you want a CA-signed cert, the trustpoint enrollment changes from `self` to `terminal` or `url`, and you import the issuer chain. We cover the full cert workflow in [Cisco ASA Certificate Management for AnyConnect](https://www.pinglabz.com/cisco-asa-anyconnect-certificates/). ## Step 2: Create the Client IP Pool and a Local User AnyConnect clients need an inside address. Carve out a separate /24 (or larger) that is not in use anywhere else and is routable from the inside network. In our lab we use 10.99.99.0/24 because it does not collide with the inside 10.10.0.0/16. ``` ASA-PERIM(config)# ip local pool VPN-POOL 10.99.99.10-10.99.99.250 mask 255.255.255.0 ASA-PERIM(config)# username vpnuser password PingLabzVPN! privilege 0 ASA-PERIM(config)# username vpnuser attributes ASA-PERIM(config-username)# service-type remote-access ASA-PERIM(config-username)# exit ``` The `service-type remote-access` line is important. Without it, that local user can shell into the ASA via SSH (priv 0 only, but still). With it, the user can authenticate for VPN but is rejected for SSH and ASDM. Verify the pool: ``` ASA-PERIM# show ip local pool VPN-POOL Begin End Mask Free Held In use 10.99.99.10 10.99.99.250 255.255.255.0 241 0 0 ``` 241 addresses available, none assigned yet. Each connected client gets one until disconnected, at which point the address goes back to the pool after the idle timeout. ## Step 3: Build the Group-Policy The group-policy is the bag of attributes the ASA pushes to the client at login: which DNS servers, which split-tunnel ACL, what banner, DPD intervals, and most importantly which tunnel protocol(s) the client is allowed to use. ``` ASA-PERIM(config)# access-list SPLIT-TUNNEL standard permit 10.10.0.0 255.255.0.0 ASA-PERIM(config)# group-policy ANYCONNECT-SSL-GP internal ASA-PERIM(config)# group-policy ANYCONNECT-SSL-GP attributes ASA-PERIM(config-group-policy)# vpn-tunnel-protocol ssl-client ASA-PERIM(config-group-policy)# split-tunnel-policy tunnelspecified ASA-PERIM(config-group-policy)# split-tunnel-network-list value SPLIT-TUNNEL ASA-PERIM(config-group-policy)# dns-server value 10.10.0.10 ASA-PERIM(config-group-policy)# default-domain value pinglabz.lab ASA-PERIM(config-group-policy)# webvpn ASA-PERIM(config-group-webvpn)# anyconnect keep-installer installed ASA-PERIM(config-group-webvpn)# anyconnect ssl dtls enable ASA-PERIM(config-group-webvpn)# anyconnect ssl rekey time 60 ASA-PERIM(config-group-webvpn)# anyconnect ssl rekey method ssl ASA-PERIM(config-group-webvpn)# anyconnect dpd-interval client 30 ASA-PERIM(config-group-webvpn)# anyconnect dpd-interval gateway 30 ASA-PERIM(config-group-webvpn)# exit ``` A few of those lines deserve called-out attention: - **`vpn-tunnel-protocol ssl-client`**: this group-policy only permits the SSL client. If the user attempts to use IKEv2, the connection is rejected. A separate group-policy for IKEv2 lets you keep the two profiles cleanly separated. We cover IKEv2 client setup in [Cisco ASA AnyConnect IKEv2 VPN Configuration](https://www.pinglabz.com/cisco-asa-anyconnect-ikev2-vpn/). - `**split-tunnel-policy tunnelspecified**`: only traffic destined to the networks in the SPLIT-TUNNEL ACL is sent through the VPN. Everything else (Netflix, Google, the public internet) goes direct. The other choices are `tunnelall` (all traffic via VPN) and `excludespecified`. Full details in [Cisco ASA Split Tunneling Explained](https://www.pinglabz.com/cisco-asa-split-tunneling/). - **`anyconnect ssl dtls enable`**: the client opens a UDP DTLS channel in parallel to the TCP channel. UDP is much faster for video and large transfers because TCP-over-TCP gets bogged down in nested retransmits. - **`anyconnect dpd-interval`**: dead-peer-detection. Both sides probe every 30 seconds; if 4 misses, the tunnel is torn down and the client reconnects. ## Step 4: Tunnel-Group (Connection Profile) The tunnel-group is the named profile clients select. It binds together address pool, group-policy, AAA servers, and (optionally) a custom URL or display alias. ``` ASA-PERIM(config)# tunnel-group SSL_PROFILE type remote-access ASA-PERIM(config)# tunnel-group SSL_PROFILE general-attributes ASA-PERIM(config-tunnel-general)# default-group-policy ANYCONNECT-SSL-GP ASA-PERIM(config-tunnel-general)# address-pool VPN-POOL ASA-PERIM(config-tunnel-general)# authentication-server-group RADIUS-VPN LOCAL ASA-PERIM(config-tunnel-general)# exit ASA-PERIM(config)# tunnel-group SSL_PROFILE webvpn-attributes ASA-PERIM(config-tunnel-webvpn)# group-alias EMPLOYEES enable ASA-PERIM(config-tunnel-webvpn)# group-url https://203.0.113.2/employees enable ASA-PERIM(config-tunnel-webvpn)# exit ``` The `authentication-server-group RADIUS-VPN LOCAL` line authenticates against RADIUS first, falling back to the local username database if RADIUS is unreachable. We covered the AAA server setup in detail in [Cisco ASA AAA for VPN: LDAP, RADIUS, and TACACS+](https://www.pinglabz.com/cisco-asa-aaa-for-vpn/). The two webvpn-attribute lines control how clients see the connection profile: - **`group-alias EMPLOYEES enable`**: the dropdown at https://203.0.113.2/ shows "EMPLOYEES" as a selectable profile. - **`group-url https://203.0.113.2/employees enable`**: clients hitting that exact URL automatically land on this profile without choosing. A common pattern is to publish a different group-url for each org unit (employees, contractors, partners). Each maps to a different group-policy with a different split-tunnel ACL and a different banner. ## Step 5: Enable WebVPN and AnyConnect Globally Even with everything above configured, no SSL listener is running on the outside interface until you enable WebVPN. This is the global service that owns the TLS endpoint on TCP/443. ``` ASA-PERIM(config)# webvpn ASA-PERIM(config-webvpn)# enable outside INFO: WebVPN and DTLS are enabled on 'outside'. ASA-PERIM(config-webvpn)# anyconnect enable WARNING: No 'anyconnect image' commands have been issued ASA-PERIM(config-webvpn)# tunnel-group-list enable ASA-PERIM(config-webvpn)# exit ``` The warning about no `anyconnect image` command means the client package upload step is missing. In production you would copy the .pkg to flash and reference it: ``` ASA-PERIM(config-webvpn)# anyconnect image flash:/cisco-secure-client-win-5.1.10.233-webdeploy-k9.pkg 1 ``` Without the image, browsers can land on the portal page but the AnyConnect client cannot auto-download. In our lab we leave it absent because we do not have a real client image, and we are exercising the configuration plane only. The control plane is fully functional regardless of whether the .pkg is present, which is what matters for verifying the config is correct. ## Verification: Did Everything Take? Three quick sanity-check commands: ``` ASA-PERIM# show ssl Accept connections using SSLv3 or greater and negotiate to TLSv1.2 or greater Start connections using TLSv1.2 and negotiate to TLSv1.2 or greater SSL DH Group: group14 (2048-bit modulus, FIPS) SSL ECDH Group: group19 (256-bit EC) SSL trust-points: Self-signed (RSA 2048 bits RSA-SHA256) certificate available Self-signed (EC 256 bits ecdsa-with-SHA256) certificate available Interface outside: PINGLABZ-SELFSIGNED (RSA 2048 bits RSA-SHA256) Certificate authentication is not enabled ASA-PERIM# show webvpn anyconnect AnyConnect Client is enabled. No images configured ASA-PERIM# show vpn-sessiondb anyconnect INFO: There are presently no active sessions of the type specified ``` The first output confirms our trustpoint is bound to the outside interface. The second confirms AnyConnect is on but no client image is uploaded (the lab constraint). The third just shows there are no active sessions, which is the expected state until a client connects. Once a real client connects, `show vpn-sessiondb anyconnect` displays the session, the assigned IP from the pool, the encryption suite, the bytes transferred, and the duration. `show vpn-sessiondb summary` rolls those stats up for capacity planning. ## Common Gotchas The five mistakes that catch every engineer once: 1. **Trustpoint bound to the wrong interface.** If `ssl trust-point PINGLABZ-SELFSIGNED outside` is missing or points at the inside, clients see the default ASA self-signed cert and warn about a CN mismatch. Always bind the trustpoint to the interface the clients will hit. 2. **NAT exemption not configured for the VPN pool.** If 10.99.99.0/24 traffic to the inside subnet gets PAT-translated by your existing dynamic PAT rule, clients connect successfully but cannot reach inside resources because the return traffic goes to the wrong source. Fix with an identity NAT rule in Section 1\. Full walkthrough in [Cisco ASA Identity NAT / NAT Exemption for VPNs](https://www.pinglabz.com/cisco-asa-identity-nat-vpn/). 3. **Split-tunnel ACL is extended instead of standard.** The `split-tunnel-network-list` attribute requires a *standard* ACL. Pasting an extended ACL silently fails to apply, and the client gets `tunnelall` behavior by default. 4. **Group-policy default-domain or DNS missing.** Clients connect but cannot resolve `fileserver.pinglabz.lab` short names. Always set `dns-server` and `default-domain` in the group-policy. 5. **Outside ACL blocking TCP/443 from the internet.** If your `OUTSIDE_IN` ACL has an explicit deny near the top, the SSL handshake never reaches the WebVPN service. Check with [packet-tracer](https://www.pinglabz.com/cisco-asa-packet-tracer/) to see if the ACL phase is dropping the SYN. ## Key Takeaways AnyConnect SSL VPN on Cisco ASA breaks down into five building blocks: the identity certificate, the IP local pool, the group-policy, the tunnel-group connection profile, and the WebVPN service. Build them in that order and each step has the dependencies it needs. The trickiest piece for first-time setups is binding the trustpoint to the right interface and remembering to add NAT exemption for the VPN pool, both of which produce silent failures rather than clear errors. For the full Cisco ASA reference, including site-to-site IPsec, NAT, ACLs, failover, and troubleshooting tools, see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/). When connections fail, start with [Troubleshoot AnyConnect Login and Certificate Problems on ASA](https://www.pinglabz.com/cisco-asa-anyconnect-troubleshooting/); for the ACL side, [Cisco ASA ACL Troubleshooting with packet-tracer](https://www.pinglabz.com/cisco-asa-acl-troubleshooting/) covers the most common drops. ### Cisco ASA ACL Troubleshooting with packet-tracer URL: https://www.pinglabz.com/cisco-asa-acl-troubleshooting/ Last updated: 2026-06-13T20:08:29.000Z Most "the ACL is broken" tickets are not really broken ACLs. They are misunderstood ACLs: a packet you thought matched line 3 actually matched line 1 and got denied; a deny line further down absorbed traffic the user expected to permit; an implicit deny at the end caught a flow nobody had a permit for. The ASA always tells you exactly what happened, but you have to ask it the right way. This article is the ACL troubleshooting reference for the Cisco ASA on software 9.x. The diagnostic of choice is packet-tracer, backed by `show access-list` hit counters and `show asp drop`. All output below is from a live ASAv in the PingLabz reference lab. For the cluster context, see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/); for the broader ACL guide, see [Cisco ASA ACL Configuration](https://www.pinglabz.com/cisco-asa-acl-configuration/); for the packet-tracer command itself, see [Cisco ASA packet-tracer Command](https://www.pinglabz.com/cisco-asa-packet-tracer/). ## The Three Tools, in the Order You Use Them `packet-tracer input ...` Step1 Tells you For one specific 5-tuple, did the ACL allow or drop, and which line matched. `show access-list NAME` Step2 Tells you Per-line hit counters and last-hit timestamps for the named ACL. `show asp drop` Step3 Tells you Aggregate counter of every drop reason on the ASA, including `acl-drop`. Step 1 answers "what happens to this exact packet?" Step 2 answers "is the rule I just edited actually being hit?" Step 3 answers "is the ASA dropping anything, and if so, why?" ## Lab Setup The PingLabz ASA has an OUTSIDE\_IN ACL applied inbound on the outside interface. The ACL has been deliberately seeded with one bad-actor block at the top, three permits for the DMZ web server, and an explicit catch-all deny with logging at the bottom (added in addition to the implicit deny so we can see clean logging in syslog): ``` ASA-PERIM# show running-config access-list access-list OUTSIDE_IN extended deny tcp host 198.51.100.99 any access-list OUTSIDE_IN extended permit tcp any object DMZ-WEB eq www access-list OUTSIDE_IN extended permit tcp any object DMZ-WEB eq https access-list OUTSIDE_IN extended permit icmp any object DMZ-WEB access-list OUTSIDE_IN extended deny ip any any log ``` That gives us five different troubleshooting scenarios on one ACL: an explicit deny hit, three permit lines (one cold, two warm), and a logged catch-all deny. ## Scenario 1: Explicit Deny Hit An attacker IP we have blocklisted (198.51.100.99) tries to reach the DMZ web server. We expect line 1 of the ACL to drop it. ``` ASA-PERIM# packet-tracer input outside tcp 198.51.100.99 51000 198.51.100.10 443 Phase: 1 Type: UN-NAT Subtype: static Result: ALLOW Config: object network DMZ-WEB nat (dmz,outside) static 198.51.100.10 Additional Information: NAT divert to egress interface dmz Untranslate 198.51.100.10/443 to 192.168.50.10/443 Phase: 2 Type: ACCESS-LIST Subtype: Result: DROP Config: access-group OUTSIDE_IN in interface outside access-list OUTSIDE_IN extended deny tcp host 198.51.100.99 any Result: input-interface: outside input-status: up input-line-status: up output-interface: dmz output-status: up output-line-status: up Action: drop Time Taken: 52927 ns Drop-reason: (acl-drop) Flow is denied by configured rule, Drop-location: frame snp_classify_table_lookup:6044 flow (NA)/NA ``` Phase 2 is where the answer lives. The drop reason is `acl-drop`, the line that fired is printed verbatim (`access-list OUTSIDE_IN extended deny tcp host 198.51.100.99 any`), and the ACL name is shown via the `access-group` binding. No guessing required. One detail worth noting: NAT (Phase 1) ran *before* the ACL (Phase 2). The packet-tracer output shows the ACL evaluating against the post-NAT (real) destination 192.168.50.10, which is why the DMZ-WEB object reference works in the permit lines. On modern ASA software the ACL always sees real addresses on both sides of the rule. ## Scenario 2: Permit Hit (with a Lab Quirk at the End) A normal user IP reaches DMZ-WEB on TCP/443\. We expect line 3 (the HTTPS permit) to match. ``` ASA-PERIM# packet-tracer input outside tcp 198.51.100.50 51001 198.51.100.10 443 Phase: 1 Type: UN-NAT Subtype: static Result: ALLOW Config: object network DMZ-WEB nat (dmz,outside) static 198.51.100.10 Additional Information: NAT divert to egress interface dmz Untranslate 198.51.100.10/443 to 192.168.50.10/443 Phase: 2 Type: ACCESS-LIST Subtype: Result: ALLOW Config: access-group OUTSIDE_IN in interface outside Phase: 10 Type: FLOW-CREATION Subtype: Result: ALLOW New flow created with id 54, packet dispatched to next module Phase: 11 Type: INPUT-ROUTE-LOOKUP-FROM-OUTPUT-ROUTE-LOOKUP Subtype: Resolve Preferred Egress interface Result: ALLOW Found next-hop 192.168.50.10 using egress ifc dmz Result: Action: drop Drop-reason: (no-v4-adjacency) No valid V4 adjacency. Check ARP table (show arp) has entry for nexthop., Drop-location: frame snp_fp_adj_process_cb:256 flow (NA)/NA ``` Phase 2 says ALLOW. The ACL did its job. The eventual `Action: drop` at the bottom is from `no-v4-adjacency`, which is a downstream problem (no ARP entry for the DMZ host in this lab), not an ACL problem. This is the exact pattern you will hit in production all the time: the ACL is fine, the failure is somewhere else. Reading the packet-tracer in order, top-to-bottom, prevents you from blaming the ACL when the cause is one phase later. ## Scenario 3: The Explicit Catch-All Deny A legitimate user IP tries to reach DMZ-WEB on TCP/22 (SSH), which the published service does not expose. We expect line 5 (the explicit `deny ip any any log`) to fire. ``` ASA-PERIM# packet-tracer input outside tcp 198.51.100.50 51002 198.51.100.10 22 Phase: 1 Type: UN-NAT Subtype: static Result: ALLOW Config: object network DMZ-WEB nat (dmz,outside) static 198.51.100.10 Additional Information: NAT divert to egress interface dmz Untranslate 198.51.100.10/22 to 192.168.50.10/22 Phase: 2 Type: ACCESS-LIST Subtype: log Result: DROP Config: access-group OUTSIDE_IN in interface outside access-list OUTSIDE_IN extended deny ip any any log Result: Action: drop Drop-reason: (acl-drop) Flow is denied by configured rule, Drop-location: frame snp_classify_table_lookup:6044 flow (NA)/NA ``` Phase 2 says DROP, and now the `Subtype: log` field is populated because the matched line has the `log` keyword. The line that fired is the catch-all `deny ip any any log`, and the corresponding syslog message gets generated (`%ASA-4-106023`) for the security team. Lesson: if you want visibility into what is being dropped at the implicit deny, replace it with an *explicit* deny with the `log` keyword. The ASA's implicit deny does not log. ## Reading Hit Counters After running the three packet-tracer scenarios, `show access-list OUTSIDE_IN` shows the per-line counters with timestamps: ``` ASA-PERIM# show access-list OUTSIDE_IN access-list OUTSIDE_IN; 5 elements; name hash: 0xe01d8199 access-list OUTSIDE_IN line 1 extended deny tcp host 198.51.100.99 any (hitcnt=1) (Last Hit=00:01:11 UTC May 10 2026) 0xa0fb08eb access-list OUTSIDE_IN line 2 extended permit tcp any object DMZ-WEB eq www (hitcnt=1) (Last Hit=23:27:43 UTC May 9 2026) 0xf18a0028 access-list OUTSIDE_IN line 2 extended permit tcp any host 192.168.50.10 eq www (hitcnt=1) (Last Hit=23:27:43 UTC May 9 2026) 0xf18a0028 access-list OUTSIDE_IN line 3 extended permit tcp any object DMZ-WEB eq https (hitcnt=2) (Last Hit=00:01:11 UTC May 10 2026) 0x1ce50e7b access-list OUTSIDE_IN line 3 extended permit tcp any host 192.168.50.10 eq https (hitcnt=2) (Last Hit=00:01:11 UTC May 10 2026) 0x1ce50e7b access-list OUTSIDE_IN line 4 extended permit icmp any object DMZ-WEB (hitcnt=0) 0x2cc9dc3f access-list OUTSIDE_IN line 4 extended permit icmp any host 192.168.50.10 (hitcnt=0) 0x2cc9dc3f access-list OUTSIDE_IN line 5 extended deny ip any any log informational interval 300 (hitcnt=1) (Last Hit=00:01:12 UTC May 10 2026) 0x2dc51227 ``` Read off the data: - Line 1 (deny attacker): hitcnt=1\. The explicit deny scenario ran exactly once. - Line 2 (permit www): hitcnt=1\. From an earlier test session. - Line 3 (permit https): hitcnt=2\. Two HTTPS allow runs. - Line 4 (permit icmp): hitcnt=0\. Never been hit. This is informational: maybe the rule was added preemptively, maybe it should be removed, but it is not failing. - Line 5 (catchall deny): hitcnt=1\. The TCP/22 drop scenario. Each "object" line in the show output expands to its underlying real address ("host 192.168.50.10") on the next line, with the same hit count. That is how object-group expansion is rendered in the show output - one logical line, one or more underlying lines. The **real diagnostic value** of `show access-list` is the `(hitcnt=...)` column. After every change to an ACL, run a few real connections, then check the counter on the line you expected to fire. If the counter does not increment, your rule is not matching. The most common reasons are wrong source/destination, wrong port, or another line above absorbing the traffic. ## The Big Picture: `show asp drop` For aggregate visibility into *everything* the ASA has dropped, ACL or otherwise, use `show asp drop`: ``` ASA-PERIM# show asp drop frame acl-drop Flow is denied by configured rule (acl-drop) 317 Last clearing: Never ``` 317 ACL drops since the ASA last had its counters cleared. The `frame acl-drop` argument filters to just the acl-drop reason; without arguments you get the full table: ``` ASA-PERIM# show asp drop Frame drop: No valid V4 adjacency. Check ARP table (show arp) has entry for nexthop. (no-v4-adjacency) 3 Flow is denied by configured rule (acl-drop) 305 FP L2 rule drop (l2_acl) 16 Interface is down (interface-down) 3 Last clearing: Never ``` This is the place to look when a user reports "my packet is being dropped somewhere on the firewall" but cannot say where. The drop-reason names map one-to-one with what packet-tracer prints in its Result section, so once you know the reason you can run a packet-tracer and confirm. Useful asp drop reasons related to ACLs and policy: `acl-drop` Interface ACL denied the flow. `l2_acl` Layer-2 ACL denied (Ethertype, MAC). `nat-no-xlate-to-pat-pool` PAT pool exhausted; no port available. `no-v4-adjacency` No ARP entry for the next-hop on the egress interface. `rpf-violated` uRPF check failed; source routes back through a different interface. `tcp-not-syn` TCP packet without SYN arrived for a flow that does not exist; usually a stale flow on the client. If `acl-drop` is climbing without a known cause, run `capture asp_drop type asp-drop acl-drop` to capture the actual dropped frames for analysis. That is the deepest level of ACL diagnostic the ASA offers. ## Common ACL Bugs and How to Spot Them - **An earlier deny absorbs the permit.** Symptom: the permit line is in the config but its hit count is zero. Diagnosis: run packet-tracer; phase 2 will print the actual line that fired. Fix: re-order the ACL. - **Object expansion mismatch.** The permit references an object-group, but the user's traffic does not match any member of the group. Diagnosis: `show access-list` expands the group inline; compare members to the actual source/destination. Fix: add the missing member to the object-group. - **ACL applied to the wrong interface or direction.** Symptom: ACL exists but never matches. Diagnosis: `show running-config access-group` tells you what is bound where. Fix: re-bind correctly. - **Real vs mapped IP confusion (legacy configs).** ACLs from ASA software pre-8.3 used mapped IPs; modern ASA uses real IPs. Symptom: traffic works on an old ASA but breaks after upgrade. Diagnosis: packet-tracer shows the rule does not match the real address. Fix: rewrite the ACL to reference real IPs. - **Implicit deny silently dropping.** Symptom: traffic just disappears, no syslog. Diagnosis: nothing in `show access-list` matches. Fix: replace the implicit deny with an explicit `deny ip any any log` at the bottom of every ACL so future drops generate syslog. ## The Five-Step ACL Troubleshooting Checklist 1. Run `packet-tracer input INTERFACE PROTO SRC SPORT DST DPORT` with the exact 5-tuple from the user complaint. Read every phase top-to-bottom. 2. If phase 2 (ACCESS-LIST) is the drop, the printed line tells you the rule that fired. Compare to the rule you expected. 3. If the drop is somewhere else (NAT, route-lookup, adjacency), the ACL is not the problem. Stop blaming it. 4. If phase 2 says ALLOW but the user still cannot connect, run a real test connection and check `show access-list` hit counters. Counter incrementing means the ACL is letting it through; the failure is downstream. 5. For periodic auditing, run `show asp drop` to see aggregate drop counters by reason. Sudden spikes in `acl-drop` with no known cause warrant `capture asp_drop`. ## Where to Go Next - [Cisco ASA ACL Configuration](https://www.pinglabz.com/cisco-asa-acl-configuration/) for the inbound ACL syntax and object-group structure. - [Cisco ASA packet-tracer Command](https://www.pinglabz.com/cisco-asa-packet-tracer/) for the full diagnostic walk through every phase. - [Cisco ASA Packet Flow](https://www.pinglabz.com/cisco-asa-packet-flow/) for where the ACL phase sits relative to NAT, conn lookup, and forwarding. - [Cisco ASA Security Levels](https://www.pinglabz.com/cisco-asa-security-levels/) for why outside-to-DMZ even needs an ACL in the first place. - [Cisco ASA NAT Explained](https://www.pinglabz.com/cisco-asa-nat-explained/) for the NAT phase that runs before the ACL. - [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/) for the cluster index. ## Key Takeaways - ACL troubleshooting on the ASA is three commands: packet-tracer for one specific flow, `show access-list` for per-line hit counters, `show asp drop` for aggregate drop reasons. - packet-tracer phase 2 prints the exact ACL line that matched, in either ALLOW or DROP. Read it before guessing. - The implicit deny does not log. Replace it with an explicit `deny ip any any log` at the bottom of every interface ACL so silent drops become visible. - Modern ASA ACLs reference real (untranslated) addresses. Configs ported from pre-8.3 ASAs that still reference mapped IPs will fail after upgrade; rewrite them to real IPs. - If packet-tracer's phase 2 ALLOWs but the user still cannot connect, the failure is downstream of the ACL. Stop debugging the ACL. ### Cisco ASA NAT Order of Operations Cheat Sheet URL: https://www.pinglabz.com/cisco-asa-nat-order-of-operations/ Last updated: 2026-06-13T20:08:30.000Z Most Cisco ASA NAT outages are not "the rule does not work". They are "the rule works, but the wrong rule fires first". The ASA evaluates NAT rules in a fixed order, and once a packet matches, evaluation stops. If you do not know that order cold, you will eventually waste an afternoon staring at a configuration in which every rule looks correct but traffic is still going to the wrong place. This article is the cheat-sheet reference for the ASA NAT order of operations on software 9.x. It is grounded in real `show nat detail` output from the PingLabz reference lab. For the cluster context, see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/); for the broader picture, see [Cisco ASA NAT Explained](https://www.pinglabz.com/cisco-asa-nat-explained/). ## The Three Sections, in Order The ASA divides every NAT rule into one of three sections. Order is fixed and not user-configurable except by deciding which section a rule lives in. Section 1 Type Manual NAT (a.k.a. Twice NAT, Manual Before Auto) Configured how `nat (real-int,mapped-int) source ... destination ...` at global config Default orderFirst Section 2 Type Auto NAT (a.k.a. Object NAT) Configured how `nat (real-int,mapped-int) ...` nested inside `object network NAME` Default orderMiddle Section 3 TypeManual NAT after Auto Configured how `nat (real-int,mapped-int) **after-auto** source ... destination ...` at global config Default orderLast For every packet that needs NAT, the ASA scans Section 1 top to bottom; if no match, scans Section 2 top to bottom by Cisco's tie-breaking rules (more on that below); if still no match, scans Section 3\. First match wins. No more evaluation after that. ## Why the Order Exists Three reasons NAT is split this way, instead of one flat list: 1. **Specific overrides general.** Manual NAT in Section 1 lets you write a high-priority "this specific (source, destination) pair gets translated this way" rule that beats whatever broad auto NAT lives in Section 2. 2. **Auto NAT is the default chassis.** Section 2 is where the bulk of "translate this subnet to that pool" rules live. They are co-located with the network objects, easy to read, and Cisco gives them sensible automatic ordering. 3. **Section 3 is the safety net.** Manual NAT after-auto is for catch-all rules that should only fire if nothing else did. Rare in practice, common in service-provider edge configurations. ## Real `show nat detail` Output Here is the live state of the PingLabz lab ASA, which has all three section types populated except Section 3: ``` ASA-PERIM# show nat detail Manual NAT Policies (Section 1) 1 (inside) to (outside) source static INSIDE-NET-FULL INSIDE-NET-FULL destination static REMOTE-VPN-NET REMOTE-VPN-NET no-proxy-arp route-lookup translate_hits = 1, untranslate_hits = 1 Source - Origin: 10.10.0.0/16, Translated: 10.10.0.0/16 Destination - Origin: 10.99.99.0/24, Translated: 10.99.99.0/24 2 (inside) to (outside) source dynamic INSIDE-NET INSIDE-PARTNER-PAT destination static PARTNER-NET PARTNER-NET translate_hits = 1, untranslate_hits = 1 Source - Origin: 10.10.10.0/24, Translated: 198.51.100.99/32 Destination - Origin: 172.20.0.0/24, Translated: 172.20.0.0/24 Auto NAT Policies (Section 2) 1 (dmz) to (outside) source static DMZ-WEB 198.51.100.10 translate_hits = 0, untranslate_hits = 7 Source - Origin: 192.168.50.10/32, Translated: 198.51.100.10/32 2 (inside) to (outside) source dynamic INSIDE-TRANSIT interface translate_hits = 0, untranslate_hits = 0 Source - Origin: 10.10.0.0/24, Translated: 203.0.113.2/30 3 (inside) to (outside) source dynamic INSIDE-NET interface translate_hits = 46, untranslate_hits = 0 Source - Origin: 10.10.10.0/24, Translated: 203.0.113.2/30 ``` The header for each section is the diagnostic. `Manual NAT Policies (Section 1)` is two rules. `Auto NAT Policies (Section 2)` is three rules. Section 3 is empty (no header printed). Within Section 1, the order printed is the order evaluated. Within Section 2, the order is sorted by Cisco's tie-breaking rules, not by configuration order, and the printed order is what actually fires. ## How Section 1 Orders Itself Manual NAT is evaluated top-to-bottom in the order the rules appear in the running config. Newly entered rules go to the end by default. Two ways to control position: - Add an explicit line number on the `nat` command: `nat (inside,outside) **1** source static ...`. The "1" here is the position within Section 1\. New rules with that number push existing ones down. - Use `nat (inside,outside) source static ... **after** object SOMENAME` or `after-object` options on a few rule types to position relative to a named object. The default for an unnumbered manual NAT rule is "append to the end of Section 1". That is usually fine until you have several manual rules, at which point order matters. ## How Section 2 Orders Itself Auto NAT (object NAT) does not respect configuration order. The ASA sorts Section 2 rules into a deterministic priority order using these criteria, in this order: 1. **Static rules before dynamic rules.** A static auto NAT (`nat (...) static ...`) outranks any dynamic auto NAT (`nat (...) dynamic ...`). 2. **More-specific source object first.** A /32 host object outranks a /24 subnet, which outranks a /16, etc. Quantitatively: the higher the prefix length, the higher the priority. 3. **Alphabetical object name.** If specificity ties, the lexicographically lower object name wins. Look at the lab output again. Section 2 has three rules. Why are they ordered the way they are? - Rule 1 is a **static** auto NAT (DMZ-WEB). Rules 2 and 3 are dynamic. Static beats dynamic, so DMZ-WEB sorts first. - Rules 2 and 3 are both dynamic /24 subnets, same specificity. INSIDE-TRANSIT (object name starts with "INSIDE-T") sorts before INSIDE-NET? No, that is alphabetical: "INSIDE-N" < "INSIDE-T", so the lab actually got rule 2 = INSIDE-TRANSIT and rule 3 = INSIDE-NET in that order. Read the output again carefully: Section 2 rule 2 is `INSIDE-TRANSIT`, rule 3 is `INSIDE-NET`. The rule numbering printed by `show nat detail` uses the ASA's evaluation order, which is the alphabetical ordering on the object name when other criteria tie. Tie-breaker: T comes after N alphabetically, so INSIDE-NET should sort first. If that does not match what you see, run `show nat detail` and trust the printed order over any guessed mental model. The ASA prints what it actually does. ## Walking a Packet Through the Sections Let us trace what the ASA does for an outbound packet from inside (10.10.10.50) to a partner network (172.20.0.50): 1. **Section 1, Rule 1 (Identity NAT for VPN).** Source 10.10.10.50 is in 10.10.0.0/16 (matches), destination 172.20.0.50 is **not** in 10.99.99.0/24 (no match). Move on. 2. **Section 1, Rule 2 (Twice NAT for partner).** Source 10.10.10.50 is in 10.10.10.0/24 (matches), destination 172.20.0.50 is in 172.20.0.0/24 (matches). Both halves match. Apply: source rewrites to 198.51.100.99, destination unchanged. Stop. The packet never reaches Section 2\. The auto NAT PAT rule that would have rewritten the source to the outside interface IP (203.0.113.2) does not fire because Section 1 already produced a match. Now trace a different packet: inside (10.10.10.50) to the open internet (8.8.8.8): 1. **Section 1, Rule 1.** Source matches (10.10.0.0/16), destination 8.8.8.8 does **not** match REMOTE-VPN-NET. No match. 2. **Section 1, Rule 2.** Source matches INSIDE-NET (10.10.10.0/24), destination 8.8.8.8 does **not** match PARTNER-NET. No match. 3. **Section 2, Rule 1 (DMZ-WEB static).** Source 10.10.10.50 is not 192.168.50.10\. No match. 4. **Section 2, Rule 2 (INSIDE-TRANSIT PAT).** Source 10.10.10.50 is not in 10.10.0.0/24\. No match. 5. **Section 2, Rule 3 (INSIDE-NET PAT).** Source 10.10.10.50 is in 10.10.10.0/24\. Match. Apply: source rewrites to outside interface IP 203.0.113.2. That is the path for a normal outbound flow. Section 1 had no match because no manual rule had both source and destination matching, and Section 2 fell through to the catch-all PAT rule. ## Order-Related Diagnostics When you suspect a rule is not firing in the expected order, two commands answer the question quickly: - `**show nat detail**` prints every rule with its hit counters and section number. Look at `translate_hits` and `untranslate_hits`. A rule with zero hits when you expect non-zero is either not matching the traffic at all (object scope wrong) or another rule earlier in the order is winning. - **`packet-tracer input ... ...`** tells you which rule fires for a specific source/destination pair. Phase 1 (UN-NAT) and Phase 2 or 3 (NAT) both reference the rule that matched. If those phases reference a rule you did not expect, that is the answer. For the deeper troubleshooting walk, see [Cisco ASA packet-tracer Command](https://www.pinglabz.com/cisco-asa-packet-tracer/) and [Cisco ASA ACL Troubleshooting with packet-tracer](https://www.pinglabz.com/cisco-asa-acl-troubleshooting/). ## The Cheat Sheet Translate a specific (source, destination) pair specifically, beating any auto NAT Use this NAT typeManual NAT (twice NAT) Lands inSection 1 Exempt VPN traffic from NAT (identity NAT) Use this NAT type Manual NAT, source/dest both static identity Lands inSection 1 Translate a subnet to the outside interface (default outbound PAT) Use this NAT type Object NAT, dynamic interface Lands inSection 2 Publish a server with a static public IP Use this NAT typeObject NAT, static Lands inSection 2 Forward a single port (port forwarding) Use this NAT type Object NAT, static service Lands inSection 2 Catch-all NAT that should only run if nothing else matched Use this NAT typeManual NAT after-auto Lands inSection 3 And the simple rule that explains every NAT outage: the ASA evaluates Section 1 top-to-bottom, then Section 2 in priority order (static before dynamic, more-specific before less-specific, then alphabetical), then Section 3 top-to-bottom. First match wins. ## Where to Go Next - [Cisco ASA NAT Explained](https://www.pinglabz.com/cisco-asa-nat-explained/) for the broader auto/manual/object/twice taxonomy. - [Cisco ASA Dynamic PAT for Internet Access](https://www.pinglabz.com/cisco-asa-dynamic-pat/) for the standard outbound auto NAT. - [Cisco ASA Static NAT for DMZ Servers](https://www.pinglabz.com/cisco-asa-static-nat-dmz/) for inbound auto NAT. - [Cisco ASA Twice NAT Explained](https://www.pinglabz.com/cisco-asa-twice-nat/) for Section 1 manual rules. - [Cisco ASA Identity NAT for VPNs](https://www.pinglabz.com/cisco-asa-identity-nat-vpn/) for the most common Section 1 use case. - [Cisco ASA packet-tracer Command](https://www.pinglabz.com/cisco-asa-packet-tracer/) for the diagnostic walk. - [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/) for the cluster index. ## Key Takeaways - NAT rules live in three sections: 1 = Manual (twice) NAT, 2 = Auto (object) NAT, 3 = Manual after-auto. Evaluated in that order; first match wins. - Section 2 sorts itself by static-vs-dynamic, then specificity, then alphabetical name. Configuration order does not matter. - Most ASA NAT outages are ordering bugs: the wrong rule fires first. `show nat detail` hit counters and packet-tracer are the two commands that solve these in seconds. - If you need a high-priority specific rule, write it as manual NAT (Section 1). If you need a default catch-all, write it as object NAT (Section 2). Section 3 is rarely needed. ### Cisco ASA Identity NAT / NAT Exemption for VPNs URL: https://www.pinglabz.com/cisco-asa-identity-nat-vpn/ Last updated: 2026-05-29T23:40:51.000Z Identity NAT, sometimes called NAT exemption or "no-NAT", is the rule you write when you specifically do not want to translate a flow. The most common reason: site-to-site IPsec or remote-access VPN traffic. The remote VPN peer expects to see the original inside addresses, and any source NAT in the path corrupts the SA selectors. The fix is to write a rule that matches the VPN-bound flow and translates it to itself - the same address in, the same address out. This article walks identity NAT on Cisco ASA software 9.x, with real captures from the PingLabz reference lab. For the cluster context, see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/); for the broader NAT picture, see [Cisco ASA NAT Explained](https://www.pinglabz.com/cisco-asa-nat-explained/); for the two NAT types we are bypassing, see [Dynamic PAT](https://www.pinglabz.com/cisco-asa-dynamic-pat/) and [Twice NAT](https://www.pinglabz.com/cisco-asa-twice-nat/). ## Why VPN Traffic Needs to Skip NAT Two reasons engineers care about identity NAT, both rooted in how IPsec and remote-access VPN work. **IPsec selectors are absolute.** A site-to-site IPsec tunnel is defined by a crypto map ACL: traffic matching the ACL is encrypted, traffic that does not match goes in the clear. The matching is done on real (untranslated) addresses by default. If the ASA also has a dynamic PAT rule that fires before the crypto match, the source IP is rewritten to the outside interface, the crypto ACL no longer matches, and the packet leaves the ASA in the clear instead of in the tunnel. The remote peer then drops it because it expected encrypted traffic for that flow. **Remote-access VPN clients expect to see the inside subnets unchanged.** When an AnyConnect or Cisco Secure Client user from 10.99.99.0/24 reaches an inside server at 10.10.10.50, the server replies to 10.99.99.0/24\. If the ASA NATs the source on the way back, the client never sees the reply because the source IP no longer matches the local route table on the client. In both cases, the answer is the same: write an explicit identity NAT rule for VPN-bound traffic so the inside addresses arrive at the tunnel unchanged. Identity NAT is implemented as manual NAT (twice NAT) where source and destination both translate to themselves. It lives in Section 1 of the NAT table, which is evaluated before Section 2 auto NAT. That ordering is what makes the exemption work: the identity rule wins over the dynamic PAT rule that would otherwise rewrite the source. ## The Lab Topology and the Goal All output below is from a live ASAv running Cisco ASA 9.23(1). The relevant pieces: - Inside corporate space: 10.10.0.0/16 (covers all current and future inside subnets). - Outbound default: dynamic PAT to the outside interface (Section 2 auto NAT) for everything to the internet. - Remote-access VPN pool: 10.99.99.0/24 (assigned to AnyConnect / Secure Client users). The goal: when an inside host (anything in 10.10.0.0/16) talks to a VPN client (10.99.99.0/24), do not translate either address. For all other destinations, normal PAT applies. ## The Identity NAT Rule Two network objects (the inside aggregate and the VPN pool) plus one manual NAT rule: ``` object network INSIDE-NET-FULL subnet 10.10.0.0 255.255.0.0 object network REMOTE-VPN-NET subnet 10.99.99.0 255.255.255.0 nat (inside,outside) source static INSIDE-NET-FULL INSIDE-NET-FULL destination static REMOTE-VPN-NET REMOTE-VPN-NET no-proxy-arp route-lookup ``` Read it carefully: - `(inside,outside)` \- real interface, mapped interface. Outbound traffic from inside arrives here. - `source static INSIDE-NET-FULL INSIDE-NET-FULL` \- match source against object INSIDE-NET-FULL (10.10.0.0/16), translate to the same object. Same object on both sides means identity, no rewrite. - `destination static REMOTE-VPN-NET REMOTE-VPN-NET` \- match destination against the VPN pool, do not translate. Same convention. - `no-proxy-arp` \- the ASA must not answer ARP for the inside addresses on the outside interface. Default would have the ASA proxy-ARP for the entire inside aggregate, which leaks ARP traffic to the upstream and breaks the network. Always include `no-proxy-arp` on identity NAT. - `route-lookup` \- the ASA does a route lookup to determine egress interface, instead of using the rule's mapped interface. Required when the destination is reachable through a different interface than the rule names (which is the VPN case: the destination address belongs to the VPN tunnel, not literally to "outside"). The two flags at the end are not optional in practice. Skip `no-proxy-arp` and you will probably break something in production. Skip `route-lookup` and the ASA will try to forward the VPN-bound packet directly out the named outside interface, miss the tunnel, and either drop or send in the clear. ## Verifying the Rule Loaded `show running-config nat` shows the rule at the top, in Section 1: ``` ASA-PERIM# show running-config nat nat (inside,outside) source static INSIDE-NET-FULL INSIDE-NET-FULL destination static REMOTE-VPN-NET REMOTE-VPN-NET no-proxy-arp route-lookup nat (inside,outside) source dynamic INSIDE-NET INSIDE-PARTNER-PAT destination static PARTNER-NET PARTNER-NET ! object network INSIDE-NET nat (inside,outside) dynamic interface object network INSIDE-TRANSIT nat (inside,outside) dynamic interface object network DMZ-WEB nat (dmz,outside) static 198.51.100.10 ``` `show nat detail` shows what the ASA actually built from that rule: ``` ASA-PERIM# show nat detail Manual NAT Policies (Section 1) 1 (inside) to (outside) source static INSIDE-NET-FULL INSIDE-NET-FULL destination static REMOTE-VPN-NET REMOTE-VPN-NET no-proxy-arp route-lookup translate_hits = 1, untranslate_hits = 1 Source - Origin: 10.10.0.0/16, Translated: 10.10.0.0/16 Destination - Origin: 10.99.99.0/24, Translated: 10.99.99.0/24 2 (inside) to (outside) source dynamic INSIDE-NET INSIDE-PARTNER-PAT destination static PARTNER-NET PARTNER-NET translate_hits = 1, untranslate_hits = 1 Source - Origin: 10.10.10.0/24, Translated: 198.51.100.99/32 Destination - Origin: 172.20.0.0/24, Translated: 172.20.0.0/24 Auto NAT Policies (Section 2) 1 (dmz) to (outside) source static DMZ-WEB 198.51.100.10 translate_hits = 0, untranslate_hits = 7 2 (inside) to (outside) source dynamic INSIDE-TRANSIT interface translate_hits = 0, untranslate_hits = 0 3 (inside) to (outside) source dynamic INSIDE-NET interface translate_hits = 46, untranslate_hits = 0 ``` Read the Section 1 rule 1 carefully: `Source - Origin: 10.10.0.0/16, Translated: 10.10.0.0/16`. Origin and translated are identical. Same on the destination side. That is what identity means. ## Proving the Rule Wins Over Dynamic PAT The whole point of identity NAT is that it should fire before the dynamic PAT rule. packet-tracer makes that visible. Send a packet from an inside host to a VPN client address: ``` ASA-PERIM# packet-tracer input inside tcp 10.10.10.50 33000 10.99.99.10 443 Phase: 1 Type: INPUT-ROUTE-LOOKUP Subtype: Resolve Egress Interface Result: ALLOW Found next-hop 203.0.113.1 using egress ifc outside Phase: 2 Type: UN-NAT Subtype: static Result: ALLOW Config: nat (inside,outside) source static INSIDE-NET-FULL INSIDE-NET-FULL destination static REMOTE-VPN-NET REMOTE-VPN-NET no-proxy-arp route-lookup Additional Information: NAT divert to egress interface outside Untranslate 10.99.99.10/443 to 10.99.99.10/443 Phase: 3 Type: NAT Subtype: Result: ALLOW Config: nat (inside,outside) source static INSIDE-NET-FULL INSIDE-NET-FULL destination static REMOTE-VPN-NET REMOTE-VPN-NET no-proxy-arp route-lookup Additional Information: Static translate 10.10.10.50/33000 to 10.10.10.50/33000 Phase: 7 Type: NAT Subtype: rpf-check Result: ALLOW Config: nat (inside,outside) source static INSIDE-NET-FULL INSIDE-NET-FULL destination static REMOTE-VPN-NET REMOTE-VPN-NET no-proxy-arp route-lookup Phase: 11 Type: FLOW-CREATION Subtype: Result: ALLOW New flow created with id 52, packet dispatched to next module Result: input-interface: inside input-status: up input-line-status: up output-interface: outside output-status: up output-line-status: up Action: allow ``` The two NAT phases are the proof: - **Phase 2 (UN-NAT)**: `Untranslate 10.99.99.10/443 to 10.99.99.10/443` \- destination matched against the identity rule, no rewrite. - **Phase 3 (NAT)**: `Static translate 10.10.10.50/33000 to 10.10.10.50/33000` \- source matched against the identity rule, no rewrite. Both phases reference the manual NAT rule. The dynamic PAT rule (Section 2) never gets evaluated because Section 1 already produced a match. That is the entire mechanism: by virtue of being in Section 1, identity NAT pre-empts the dynamic rule. Compare to a packet from the same source going to the actual internet (8.8.8.8), which should hit the auto NAT rule: ``` ASA-PERIM# packet-tracer input inside tcp 10.10.10.50 49000 8.8.8.8 80 Phase: 2 Type: NAT Subtype: Result: ALLOW Config: object network INSIDE-NET nat (inside,outside) dynamic interface Additional Information: Dynamic translate 10.10.10.50/49000 to 203.0.113.2/49000 ``` Same source, different destination, different rule. The Section 1 identity rule does not match because 8.8.8.8 is not in REMOTE-VPN-NET, so evaluation continues into Section 2 where the dynamic PAT rule wins. Source rewrites to the outside interface IP, exactly as it should. Identity NAT only fires when the destination is VPN-relevant. ## When Identity NAT Is Required Reach for identity NAT when one of these is true: - **Site-to-site IPsec tunnel** between this ASA and a remote peer where the inside addresses must arrive at the tunnel unchanged. The crypto ACL on both sides matches real addresses. - **Remote-access VPN** (AnyConnect, Cisco Secure Client) where users in the VPN pool need to reach inside servers and the inside servers' replies must reach the VPN clients with the original IPs intact. - **Reaching a partner network through a tunnel** where the partner expects your real corporate addresses and has provisioned routing on their side accordingly. - **Carving an exception out of a broader auto NAT rule.** If your default outbound rule PATs the entire inside, and one specific destination subnet should bypass NAT, an identity NAT rule in Section 1 is the right escape hatch. ## Failure Modes You Will Actually See - **Forgetting `no-proxy-arp`.** The ASA proxy-ARPs for the entire inside aggregate on the outside interface, the upstream switch sees ARP for 10.10.0.0/16 from the ASA's outside MAC, weird forwarding loops follow. Fix: always include the flag. - **Forgetting `route-lookup`.** The ASA tries to forward the packet directly out the mapped interface named in the NAT rule, misses the VPN tunnel, packet leaves in the clear or gets black-holed. Fix: always include the flag for VPN-related identity NAT. - **Identity NAT in the wrong section.** Author wrote the rule as auto NAT (object NAT cannot match destination, so it does not actually express identity), thought it would work, did not. Manual NAT in Section 1 is the only way to write identity NAT for a (source, destination) pair. - **Identity NAT with mismatched source/destination scope.** The rule covers 10.10.0.0/16 but the actual inside subnet is 10.20.0.0/24\. Rule never fires. Symptom: `show nat detail` shows zero hits for the identity rule, and inside hosts get PATted on VPN flows. Fix: align source object with actual inside addressing. - **Order conflicts with another Section 1 rule.** Two manual NAT rules both could match a flow; the wrong one wins. Fix: use the optional rule line number on the `nat` command to control order. ## Where to Go Next - [Cisco ASA NAT Order of Operations](https://www.pinglabz.com/cisco-asa-nat-order-of-operations/) for the full Section 1/2/3 rules. - [Cisco ASA Twice NAT Explained](https://www.pinglabz.com/cisco-asa-twice-nat/) for the broader manual NAT framework. - [Cisco ASA Site-to-Site VPN](https://www.pinglabz.com/cisco-asa-site-to-site-vpn/) for the IPsec configuration that depends on identity NAT. - [Cisco ASA Dynamic PAT](https://www.pinglabz.com/cisco-asa-dynamic-pat/) for the auto NAT rule that identity NAT is exempting traffic from. - [Cisco ASA packet-tracer Command](https://www.pinglabz.com/cisco-asa-packet-tracer/) for the diagnostic walk. - [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/) for the cluster index. ## Key Takeaways - Identity NAT is a manual NAT rule where source (and usually destination) translate to themselves. It lives in Section 1 and pre-empts the dynamic PAT rule that would otherwise fire. - The classic use case is VPN: site-to-site IPsec and remote-access VPN both expect inside addresses to arrive at the tunnel untranslated. - Always include `no-proxy-arp` and, for VPN-relevant rules, `route-lookup`. They are not optional in practice. - packet-tracer is the cleanest verification: the trace will show `Static translate X to X` on the source phase. If you see a different translation, the identity rule is not winning. - Auto NAT cannot express identity for a (source, destination) pair. Identity NAT is always written as manual NAT in global config, never nested in an object. ### Cisco ASA Twice NAT Explained with Real Examples URL: https://www.pinglabz.com/cisco-asa-twice-nat/ Last updated: 2026-05-29T23:40:51.000Z Twice NAT is the rule type you reach for when one source-only or destination-only translation is not enough. The classic case: when source 10.10.10.0/24 talks to destination 172.20.0.0/24 (a partner network), the source should be PATted to a specific public address that the partner has whitelisted - but only when the destination is the partner network. For all other destinations, ordinary outbound PAT applies. That conditional ("when source is X *and* destination is Y") is what manual NAT, also known as twice NAT, exists for. This article walks the configuration and the verification on Cisco ASA software 9.x with real captures from the PingLabz reference lab. For the cluster context, see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/); for the broader NAT picture, see [Cisco ASA NAT Explained](https://www.pinglabz.com/cisco-asa-nat-explained/); for the order in which manual and auto NAT rules get evaluated, see [NAT Order of Operations](https://www.pinglabz.com/cisco-asa-nat-order-of-operations/). ## What Twice NAT Actually Means "Twice NAT" is a Cisco term for a single NAT rule that can specify both source and destination match conditions, and optionally translate either or both. Compare to object NAT (auto NAT), which can only match on source. The "twice" naming is a little misleading. Twice NAT does not necessarily translate twice. It just means the rule has two halves: a source half and a destination half. Each half can be: - **Static**: pre-existing fixed binding (one IP to one IP, or one subnet to another). - **Dynamic**: many-to-one or many-to-few (PAT pool, dynamic IP pool). - **Identity**: source/destination match without translation. Used heavily for VPN exemption (see [Identity NAT for VPNs](https://www.pinglabz.com/cisco-asa-identity-nat-vpn/)). In configuration, the keyword is `nat (real-int,mapped-int) source ... destination ...` at global config, not nested inside an object. That distinguishes it from auto NAT, which is always nested. Manual NAT lands in Section 1 of the NAT table by default, which means it gets evaluated before any auto NAT rule. That ordering is the entire reason engineers reach for it: when a more specific manual rule should win over a broader auto NAT, write the manual rule and ordering is solved. ## The Lab Topology and the Goal All output below is from a live ASAv running Cisco ASA 9.23(1). The relevant pieces: - Inside subnet INSIDE-NET: 10.10.10.0/24, behind ASA interface inside. - Outside interface: 203.0.113.2/30, default route to ISP at 203.0.113.1. - Public IP space: 198.51.100.0/24, routed to this ASA. - Default outbound rule: dynamic PAT to the outside interface IP (Section 2 auto NAT). The new requirement: when 10.10.10.0/24 talks to a partner network at 172.20.0.0/24 (reachable via the upstream provider), the partner has whitelisted only one of our public IPs - 198.51.100.99\. So we need to source-NAT to that specific address, but only on flows destined for the partner. All other outbound flows should keep using the standard PAT-to-interface rule. ## The Twice NAT Rule Two extra objects (the partner subnet and the whitelisted source IP) and one twice NAT rule does it: ``` object network INSIDE-NET subnet 10.10.10.0 255.255.255.0 object network PARTNER-NET subnet 172.20.0.0 255.255.255.0 object network INSIDE-PARTNER-PAT host 198.51.100.99 nat (inside,outside) source dynamic INSIDE-NET INSIDE-PARTNER-PAT destination static PARTNER-NET PARTNER-NET ``` Read the rule left to right: - `(inside,outside)` \- real interface, mapped interface. - `source dynamic INSIDE-NET INSIDE-PARTNER-PAT` \- match source against object INSIDE-NET (10.10.10.0/24), translate using object INSIDE-PARTNER-PAT (the host 198.51.100.99). Dynamic with a single host means PAT, source ports rewritten as needed. - `destination static PARTNER-NET PARTNER-NET` \- match destination against object PARTNER-NET (172.20.0.0/24), do not translate destination (the same object on both sides means identity, no rewrite). The "static" keyword tells the ASA the destination match is exact, not dynamic. That last detail is the most counter-intuitive piece for newcomers: `destination static PARTNER-NET PARTNER-NET` looks like it is doing destination NAT, but with the same object on both sides it just means "the destination must be in PARTNER-NET; do not change it". This is the conditional that makes the rule fire only for partner-bound flows. ## Verifying the Rule Loaded `show running-config nat` shows the manual rule at the top, before the object NAT rules: ``` ASA-PERIM# show running-config nat nat (inside,outside) source static INSIDE-NET-FULL INSIDE-NET-FULL destination static REMOTE-VPN-NET REMOTE-VPN-NET no-proxy-arp route-lookup nat (inside,outside) source dynamic INSIDE-NET INSIDE-PARTNER-PAT destination static PARTNER-NET PARTNER-NET ! object network INSIDE-NET nat (inside,outside) dynamic interface object network INSIDE-TRANSIT nat (inside,outside) dynamic interface object network DMZ-WEB nat (dmz,outside) static 198.51.100.10 ``` Two manual NAT rules at the top (an identity NAT for VPN and our partner twice NAT), three object NAT rules below (the DMZ static and the two outbound PATs). `show nat detail` tells you which Section each rule lands in, and the per-rule hit counters: ``` ASA-PERIM# show nat detail Manual NAT Policies (Section 1) 1 (inside) to (outside) source static INSIDE-NET-FULL INSIDE-NET-FULL destination static REMOTE-VPN-NET REMOTE-VPN-NET no-proxy-arp route-lookup translate_hits = 1, untranslate_hits = 1 Source - Origin: 10.10.0.0/16, Translated: 10.10.0.0/16 Destination - Origin: 10.99.99.0/24, Translated: 10.99.99.0/24 2 (inside) to (outside) source dynamic INSIDE-NET INSIDE-PARTNER-PAT destination static PARTNER-NET PARTNER-NET translate_hits = 1, untranslate_hits = 1 Source - Origin: 10.10.10.0/24, Translated: 198.51.100.99/32 Destination - Origin: 172.20.0.0/24, Translated: 172.20.0.0/24 Auto NAT Policies (Section 2) 1 (dmz) to (outside) source static DMZ-WEB 198.51.100.10 translate_hits = 0, untranslate_hits = 7 2 (inside) to (outside) source dynamic INSIDE-TRANSIT interface translate_hits = 0, untranslate_hits = 0 3 (inside) to (outside) source dynamic INSIDE-NET interface translate_hits = 46, untranslate_hits = 0 ``` Section 1 has the two manual rules. Section 2 has the auto NAT rules. Within each section, the order shown is the order in which they will be evaluated. Across sections, Section 1 is always evaluated first. ## Proving It Hits the Right Rule The most important verification on a manual NAT is "did the right rule match my traffic?", because manual NAT changes the order of operations. packet-tracer is the cleanest way to prove it. Send a hypothetical packet from inside (10.10.10.50:33000) to the partner (172.20.0.50:80): ``` ASA-PERIM# packet-tracer input inside tcp 10.10.10.50 33000 172.20.0.50 80 Phase: 1 Type: UN-NAT Subtype: static Result: ALLOW Config: nat (inside,outside) source dynamic INSIDE-NET INSIDE-PARTNER-PAT destination static PARTNER-NET PARTNER-NET Additional Information: NAT divert to egress interface outside Untranslate 172.20.0.50/80 to 172.20.0.50/80 Phase: 2 Type: NAT Subtype: Result: ALLOW Config: nat (inside,outside) source dynamic INSIDE-NET INSIDE-PARTNER-PAT destination static PARTNER-NET PARTNER-NET Additional Information: Dynamic translate 10.10.10.50/33000 to 198.51.100.99/33000 Phase: 6 Type: NAT Subtype: rpf-check Result: ALLOW Config: nat (inside,outside) source dynamic INSIDE-NET INSIDE-PARTNER-PAT destination static PARTNER-NET PARTNER-NET Phase: 10 Type: FLOW-CREATION Subtype: Result: ALLOW New flow created with id 53, packet dispatched to next module Result: input-interface: inside input-status: up input-line-status: up output-interface: outside output-status: up output-line-status: up Action: allow ``` The trace tells the entire story: - **Phase 1 (UN-NAT static)** matches the destination half of the rule. The destination 172.20.0.50 matches PARTNER-NET, untranslated to itself (no destination rewrite). The packet is diverted to egress interface outside. - **Phase 2 (NAT)** matches the source half. `Dynamic translate 10.10.10.50/33000 to 198.51.100.99/33000` \- source rewritten to the whitelisted IP. The same rule is referenced. - **Phase 6 (rpf-check)** validates the reverse path matches the same rule, so the return packet from the partner can be untranslated correctly. - **Phase 10 (FLOW-CREATION)** creates the conn entry with the source and destination as they appear after translation. Now compare a packet from the same source to a non-partner destination - the standard PAT-to-interface rule should win instead: ``` ASA-PERIM# packet-tracer input inside tcp 10.10.10.50 49000 8.8.8.8 80 Phase: 2 Type: NAT Subtype: Result: ALLOW Config: object network INSIDE-NET nat (inside,outside) dynamic interface Additional Information: Dynamic translate 10.10.10.50/49000 to 203.0.113.2/49000 ``` Same source IP, different destination, different rule. Phase 2 here matches the auto NAT rule (Section 2) instead, because the partner-network destination match in the manual rule fails. Source gets translated to the outside interface IP (203.0.113.2), exactly as it should. That side-by-side is the working proof that twice NAT does what we wanted: conditional source translation, falling through to the auto NAT default for everything else. ## When Twice NAT Is the Right Tool Reach for twice NAT (manual NAT in Section 1) when one of these is true: - **Different source translation depending on destination.** The partner-whitelist case above. Different B2B partners want to see different source IPs, all from the same set of inside hosts. - **Different destination translation depending on source.** Some inside groups should see DMZ-WEB at one IP, others at another. (Less common, but real.) - **NAT exemption for VPN traffic.** Source and destination both match VPN-relevant subnets, no translation. This is identity NAT, the topic of [Identity NAT for VPNs](https://www.pinglabz.com/cisco-asa-identity-nat-vpn/). - **Forcing rule order.** A specific match needs to win over a broader auto NAT rule. Manual NAT in Section 1 is evaluated before any Section 2 auto NAT. - **Bidirectional / same-IP scenarios.** When the same address appears on both sides of the wire (overlapping address spaces between organizations), only twice NAT can express the bidirectional rewrite. If your requirement is just "translate this one source subnet to the outside interface", reach for object NAT (Section 2) instead. It is shorter, lives next to the source object, and reads more naturally. ## Failure Modes You Will Actually See - **Wrong section.** The author wrote the rule as auto NAT (nested in an object) thinking it would beat another auto NAT, but it lands in Section 2 alongside the other rule. Fix: rewrite as manual NAT to land in Section 1. - **Source object covers more than intended.** The rule fires for traffic the author did not expect, breaking unrelated flows. Fix: `show nat detail` with hit counters tells you exactly what the rule is matching; tighten the source object. - **Forgotten `destination static` identity.** The author wrote `destination static PARTNER-NET PARTNER-PUBLIC-NET` instead of `PARTNER-NET PARTNER-NET`, accidentally rewriting the destination. Symptom: partner sees traffic at an unexpected destination IP. Fix: same object on both sides if no destination rewrite is wanted. - **Out-of-order rules within Section 1.** Multiple manual NAT rules and the wrong one wins. Fix: use the optional rule line number (`nat (inside,outside) 1 source ...`) to pin order. - **RPF check failing.** packet-tracer phase 6 fails with "rpf-check". Usually a routing or symmetry problem, not a NAT problem. Confirm the route to the destination exits via the same interface the rule uses. ## Where to Go Next - [Cisco ASA Identity NAT / NAT Exemption for VPNs](https://www.pinglabz.com/cisco-asa-identity-nat-vpn/) for the most common manual NAT use case. - [Cisco ASA NAT Order of Operations](https://www.pinglabz.com/cisco-asa-nat-order-of-operations/) for the section ordering rules. - [Cisco ASA Dynamic PAT for Internet Access](https://www.pinglabz.com/cisco-asa-dynamic-pat/) for the standard outbound auto NAT. - [Cisco ASA Static NAT for DMZ Servers](https://www.pinglabz.com/cisco-asa-static-nat-dmz/) for the inbound auto NAT. - [Cisco ASA packet-tracer Command](https://www.pinglabz.com/cisco-asa-packet-tracer/) for the diagnostic walk. - [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/) for the cluster index. ## Key Takeaways - Twice NAT (manual NAT) is the only ASA NAT type that can match on both source and destination. It is the right tool when translation depends on what destination is being reached. - Manual NAT lands in Section 1 and is always evaluated before Section 2 auto NAT. Use that ordering to your advantage. - Identity translation (same object on both sides of source or destination) is how you express "match but do not rewrite". It is the building block for both partner-conditional PAT and full VPN NAT exemption. - packet-tracer is the diagnostic of choice for confirming which rule fires for a given source/destination pair. Run two flows side by side: one that should hit the manual rule, one that should fall through to auto NAT. - If you do not need conditional behavior, use object NAT instead. Manual NAT is more flexible but also more error-prone. ### Cisco ASA Static NAT for Publishing a Server in the DMZ URL: https://www.pinglabz.com/cisco-asa-static-nat-dmz/ Last updated: 2026-05-29T23:40:51.000Z Static NAT is the rule that lets the rest of the internet reach a server you own. The classic case is a public-facing web server in the DMZ: it lives on a private IP, but customers reach it via a public IP, and the ASA quietly rewrites the destination on every inbound packet so neither side has to care. This article walks the configuration end to end on Cisco ASA software 9.x, with real captures from the PingLabz reference lab. For the cluster context, see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/); for the broader NAT picture, see [Cisco ASA NAT Explained](https://www.pinglabz.com/cisco-asa-nat-explained/); for the matching outbound rule, see [Cisco ASA Dynamic PAT for Internet Access](https://www.pinglabz.com/cisco-asa-dynamic-pat/). ## What Static NAT Does (and Why It Matters) Static NAT is a fixed, bidirectional translation between one real IP and one mapped IP. Unlike dynamic PAT, it does not need a flow to start before a translation exists: the binding is in place the moment the rule is configured, in both directions. That property is what makes it the right tool for inbound traffic. When a customer types a URL that resolves to your public IP, that connection arrives at the ASA from the outside. Without a pre-existing static translation, the ASA has no idea who that destination IP belongs to on the inside. Static NAT puts the binding on the books before the packet shows up. The ASA also installs a proxy ARP entry on the mapped interface (outside, in our case) so the ASA itself answers ARP requests for the public IP. That is why you do not have to physically configure 198.51.100.10 anywhere - the ASA is the answering machine for that IP, and it forwards the traffic inward. ## The Lab Topology All output below is from a live ASAv running Cisco ASA 9.23(1) with three interfaces: - **inside** (G0/0): 10.10.0.254/24, security-level 100\. Inside subnets behind it. - **dmz** (G0/2): 192.168.50.1/24, security-level 50\. The DMZ web server lives at 192.168.50.10. - **outside** (G0/1): 203.0.113.2/30, security-level 0\. Public IP space 198.51.100.0/24 is routed to this ASA from upstream. The customer-facing public IP for the DMZ web server is 198.51.100.10\. We never touch that address on the server itself; the ASA is the only place it lives. ## Configuration: Static NAT in One Object Like dynamic PAT, the modern way to write static NAT is object NAT. Define a network object for the real IP, attach the NAT rule to it. ``` object network DMZ-WEB host 192.168.50.10 nat (dmz,outside) static 198.51.100.10 ``` Three things to read off this rule. First, `host 192.168.50.10` is the real IP, on the real (DMZ) side. Second, `(dmz,outside)` is the interface pair: real interface first, mapped interface second. Third, `static 198.51.100.10` is the mapped IP that outside hosts will use. That is the entire NAT configuration. The ASA now answers ARP for 198.51.100.10 on the outside, untranslates inbound packets to 192.168.50.10, and translates the return traffic from the server back to 198.51.100.10. ## You Still Need an ACL Static NAT does not implicitly permit traffic. The ASA's default behavior is to deny inbound (low-security to high-security) connections, and security-level 0 to security-level 50 is exactly that direction. You need an interface ACL on outside that permits the inbound web traffic. ``` access-list OUTSIDE_IN extended permit tcp any object DMZ-WEB eq www access-list OUTSIDE_IN extended permit tcp any object DMZ-WEB eq https access-list OUTSIDE_IN extended permit icmp any object DMZ-WEB access-group OUTSIDE_IN in interface outside ``` Two important details. First, the ACL references the **real** address of the server (the object DMZ-WEB resolves to 192.168.50.10), not the mapped one. This was the opposite on ASA software before 8.3, which is why old documentation can confuse you. On modern ASA, ACLs always use real IPs. Second, we permit only the ports the server is supposed to expose. Anything else gets dropped at the implicit deny at the bottom of the ACL. For the deeper troubleshooting walk on this, see [Cisco ASA ACL Troubleshooting with packet-tracer](https://www.pinglabz.com/cisco-asa-acl-troubleshooting/). ## Verifying the Configuration The running config NAT block shows the rule lives in Section 2 (Auto NAT) along with the outbound PAT rules: ``` ASA-PERIM# show running-config nat nat (inside,outside) source static INSIDE-NET-FULL INSIDE-NET-FULL destination static REMOTE-VPN-NET REMOTE-VPN-NET no-proxy-arp route-lookup nat (inside,outside) source dynamic INSIDE-NET INSIDE-PARTNER-PAT destination static PARTNER-NET PARTNER-NET ! object network INSIDE-NET nat (inside,outside) dynamic interface object network INSIDE-TRANSIT nat (inside,outside) dynamic interface object network DMZ-WEB nat (dmz,outside) static 198.51.100.10 ``` The xlate table shows the static binding always exists, even with no traffic. Notice the `flags s` (static) and the long idle time: ``` ASA-PERIM# show xlate 1 in use, 5 most used Flags: D - DNS, e - extended, I - identity, i - dynamic, r - portmap, s - static, T - twice, N - net-to-net NAT from dmz:192.168.50.10 to outside:198.51.100.10 flags s idle 0:28:35 timeout 0:00:00 ``` Idle time keeps climbing without ever timing out (timeout is 0:00:00, meaning never). That is the static behavior: the binding is permanent until the rule is removed. `show nat detail` shows the rule with both translate and untranslate counters. For an inbound-published server the meaningful number is `untranslate_hits` \- that counts inbound packets matched against this rule: ``` ASA-PERIM# show nat detail Auto NAT Policies (Section 2) 1 (dmz) to (outside) source static DMZ-WEB 198.51.100.10 translate_hits = 0, untranslate_hits = 7 Source - Origin: 192.168.50.10/32, Translated: 198.51.100.10/32 2 (inside) to (outside) source dynamic INSIDE-TRANSIT interface translate_hits = 0, untranslate_hits = 0 3 (inside) to (outside) source dynamic INSIDE-NET interface translate_hits = 46, untranslate_hits = 0 ``` `untranslate_hits = 7` means the ASA has seen 7 packets arrive on outside destined for 198.51.100.10 and rewritten the destination to 192.168.50.10\. `translate_hits = 0` means the server itself has not initiated outbound traffic that matched this rule (the inside hosts use the PAT rules instead). ## Walking an Inbound Packet with packet-tracer The cleanest demonstration of how static NAT works is to send packet-tracer a hypothetical inbound HTTPS connection and read every phase: ``` ASA-PERIM# packet-tracer input outside tcp 198.51.100.50 51001 198.51.100.10 443 Phase: 1 Type: UN-NAT Subtype: static Result: ALLOW Config: object network DMZ-WEB nat (dmz,outside) static 198.51.100.10 Additional Information: NAT divert to egress interface dmz Untranslate 198.51.100.10/443 to 192.168.50.10/443 Phase: 2 Type: ACCESS-LIST Subtype: Result: ALLOW Config: access-group OUTSIDE_IN in interface outside Phase: 3 Type: NAT Subtype: per-session Result: ALLOW Phase: 6 Type: NAT Subtype: rpf-check Result: ALLOW Config: object network DMZ-WEB nat (dmz,outside) static 198.51.100.10 Phase: 10 Type: FLOW-CREATION Subtype: Result: ALLOW New flow created with id 54, packet dispatched to next module Phase: 11 Type: INPUT-ROUTE-LOOKUP-FROM-OUTPUT-ROUTE-LOOKUP Subtype: Resolve Preferred Egress interface Result: ALLOW Found next-hop 192.168.50.10 using egress ifc dmz ``` Read the phases in order: - **Phase 1 (UN-NAT static)** is the heart of the trace. The ASA matches the destination 198.51.100.10 against the DMZ-WEB static rule and rewrites the destination to 192.168.50.10\. It also "diverts" the packet to the DMZ egress interface so subsequent phases happen against the post-NAT destination. - **Phase 2 (ACCESS-LIST)** permits the packet using the OUTSIDE\_IN ACL. Note the ACL is matching against the real address 192.168.50.10 (via the DMZ-WEB object), which is exactly why we wrote it that way. - **Phase 6 (NAT rpf-check)** confirms the reverse-path: the ASA verifies that an answer from the server can map back to the same NAT rule. This is the protection that prevents a return packet from leaking out the wrong interface. - **Phase 10 (FLOW-CREATION)** creates the conn entry. From here on, the ASA short-circuits the slow path: subsequent packets in this flow will be processed by the existing conn, not re-evaluated against every rule. - **Phase 11** is the route lookup that decides which next-hop forwards the packet to the server. If any phase had failed, packet-tracer would stop there and tell you why. That is the value of the tool: every potential failure point on the platform is named. For the full guide, see [Cisco ASA packet-tracer Command](https://www.pinglabz.com/cisco-asa-packet-tracer/). ## Static PAT (Port Forwarding) The static NAT above maps an entire IP. If you only want to forward one port (the classic "publish only TCP/443") you can write static PAT, sometimes called port forwarding: ``` object network DMZ-WEB-HTTPS host 192.168.50.10 nat (dmz,outside) static interface service tcp 443 443 ``` That rule says: when a packet hits the outside interface IP on TCP/443, untranslate it to 192.168.50.10 on TCP/443\. The mapped address is the outside interface itself (no extra public IP burned), and only the named port comes through. Non-443 traffic to the same outside IP is unaffected and continues to hit whatever else the ASA is doing on that address. Static PAT is the right pattern when you have one public IP and several DMZ servers behind it, each on a different port. It is also the pattern you reach for in lab environments where extra public IPs are not available. ## Failure Modes You Will Actually See Top-of-list problems with inbound static NAT, ranked by frequency: - **ACL not in place or wrong direction.** NAT works, packet hits the implicit deny on outside, gets dropped. Symptom: `untranslate_hits` increments but the conn count does not. Fix: check the ACL with `show access-list`; check that `access-group ... in interface outside` is applied. - **ACL referencing the mapped IP instead of the real IP.** Common when retrofitting old ASA configs. The ACL never matches. Fix: rewrite the ACL using `object DMZ-WEB` (which is the real IP) or the literal real IP 192.168.50.10. - **Server is not actually on the DMZ subnet.** The ARP-from-DMZ on the ASA fails, packet-tracer ends with "no-v4-adjacency". Fix: check `show arp`; make sure the server is up and on the right VLAN. - **Upstream routing not pointing the public range at the ASA.** Customer can not reach 198.51.100.10 because the ISP does not route that block to this ASA's outside IP. The ASA is fine; the path to the ASA is broken. Fix: confirm with the upstream provider; check from a host on the ISP side. - **Conflicting NAT rule.** A second NAT rule earlier in the table also matches. Fix: `show nat detail` tells you which rule the ASA chose, and the order is Section 1, Section 2, Section 3\. See [NAT Order of Operations](https://www.pinglabz.com/cisco-asa-nat-order-of-operations/). ## Where to Go Next - [Cisco ASA Dynamic PAT for Internet Access](https://www.pinglabz.com/cisco-asa-dynamic-pat/) for the outbound counterpart. - [Cisco ASA Twice NAT Explained](https://www.pinglabz.com/cisco-asa-twice-nat/) when source and destination both need translating. - [Cisco ASA ACL Configuration](https://www.pinglabz.com/cisco-asa-acl-configuration/) for the inbound permit rules and object groups. - [Cisco ASA Security Levels](https://www.pinglabz.com/cisco-asa-security-levels/) for why outside-to-DMZ needs an explicit permit in the first place. - [Cisco ASA packet-tracer Command](https://www.pinglabz.com/cisco-asa-packet-tracer/) for the full diagnostic walk. - [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/) for the rest of the cluster. ## Key Takeaways - Static NAT is the inbound-publication tool: a fixed bidirectional binding between one real IP and one mapped IP, present from the moment the rule is configured. - Object NAT keeps the configuration to two lines: a network object and a nested `nat ... static` rule. It lands in Section 2 of the NAT table. - An interface ACL on outside is still required. The ACL references the real IP, not the mapped IP, on modern ASA software. - Use `show xlate` to confirm the binding, `show nat detail` for the hit counters, and packet-tracer with the public IP and port to confirm an inbound packet would land on the right server. - Static PAT (port forwarding) is the same idea limited to a single port - useful when you have one public IP and many servers. ### Cisco ASA Dynamic PAT Configuration for Internet Access URL: https://www.pinglabz.com/cisco-asa-dynamic-pat/ Last updated: 2026-06-13T20:08:30.000Z Almost every Cisco ASA in the world runs the same outbound NAT rule: take everything coming from the inside subnets and translate it to the outside interface IP, port-mapped. That is dynamic Port Address Translation, and it is the rule that keeps the entire user side of the network connected to the internet. This article is the field reference for that rule on Cisco ASA software 9.x. We cover what dynamic PAT is doing, the exact configuration on a real ASAv, the show commands that prove it is working, and the failure modes you will eventually see in production. For the broader cluster context, see the [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/); for what NAT looks like across the whole platform, see [Cisco ASA NAT Explained](https://www.pinglabz.com/cisco-asa-nat-explained/). ## What Dynamic PAT Actually Does Dynamic PAT (Port Address Translation, sometimes called NAT overload) takes many private source addresses and rewrites them to a single public address, using the Layer 4 source port to keep flows distinct. Two different inside hosts can both look like the outside interface IP because their port numbers are different. The ASA tracks each translation in the xlate table and the matching connection in the conn table. The xlate is what stitches the return packet back to the right inside host: the ASA looks at (destination IP, destination port) on the return, finds the matching xlate, rewrites the destination back to the original inside (IP, port), and forwards. If you take one thing away from this article, make it that: PAT is a stateful operation, and the xlate is the state. No xlate, no return path. ## The Lab Topology Every output below comes from a live ASAv running Cisco ASA 9.23(1) in the PingLabz reference lab. The relevant pieces: - **inside** interface (G0/0): 10.10.0.254/24, security-level 100\. Inside subnets 10.10.0.0/24 and 10.10.10.0/24 sit behind it. - **outside** interface (G0/1): 203.0.113.2/30, security-level 0\. Default route points to 203.0.113.1 (the ISP). - **dmz** interface (G0/2): 192.168.50.1/24, security-level 50\. Not relevant to PAT, mentioned for completeness. ## Configuration: Object NAT, the One-Liner The cleanest way to write dynamic PAT on a modern ASA is object NAT (also called Auto NAT, lives in Section 2 of the NAT table). Define a network object for the source subnet, attach the NAT rule to it, done. ``` object network INSIDE-NET subnet 10.10.10.0 255.255.255.0 nat (inside,outside) dynamic interface object network INSIDE-TRANSIT subnet 10.10.0.0 255.255.255.0 nat (inside,outside) dynamic interface ``` Two things to notice. First, the NAT rule is nested inside the object. The keyword is `dynamic interface`, which means "translate to whatever IP is currently configured on the named outside interface". You do not have to know or hardcode the public IP. If your ISP changes it, the NAT rule does not need to change. Second, the rule names two interfaces: `(inside,outside)`. The first is the real (pre-NAT) interface, the second is the mapped (post-NAT) interface. Get the order wrong and the rule will not match traffic in the direction you expect. ## Verifying the Configuration Loaded Pull the running config NAT block to confirm both pieces of the rule (the object and the nested NAT) are in place. ``` ASA-PERIM# show running-config nat nat (inside,outside) source static INSIDE-NET-FULL INSIDE-NET-FULL destination static REMOTE-VPN-NET REMOTE-VPN-NET no-proxy-arp route-lookup nat (inside,outside) source dynamic INSIDE-NET INSIDE-PARTNER-PAT destination static PARTNER-NET PARTNER-NET ! object network INSIDE-NET nat (inside,outside) dynamic interface object network INSIDE-TRANSIT nat (inside,outside) dynamic interface object network DMZ-WEB nat (dmz,outside) static 198.51.100.10 ``` The two object NAT lines at the bottom (`INSIDE-NET` and `INSIDE-TRANSIT`) are the dynamic PAT rules. The other entries are unrelated (manual NAT for VPN exemption, manual twice NAT for a partner network, static NAT for the DMZ web server) and will become familiar across the rest of the NAT articles in this cluster. ## Watching PAT in Action With inside hosts generating outbound traffic, the xlate table fills up. Here is the live xlate from the same ASA after a series of pings sourced from 10.10.10.1 to 8.8.8.8: ``` ASA-PERIM# show xlate 5 in use, 5 most used Flags: D - DNS, e - extended, I - identity, i - dynamic, r - portmap, s - static, T - twice, N - net-to-net NAT from dmz:192.168.50.10 to outside:198.51.100.10 flags s idle 0:28:35 timeout 0:00:00 ICMP PAT from inside:10.10.10.1/2 to outside:203.0.113.2/2 flags ri idle 0:00:08 timeout 0:00:00 ICMP PAT from inside:10.10.10.1/0 to outside:203.0.113.2/26309 flags ri idle 0:00:19 timeout 0:00:00 ICMP PAT from inside:10.10.10.1/3 to outside:203.0.113.2/3 flags ri idle 0:00:08 timeout 0:00:00 ICMP PAT from inside:10.10.10.1/1 to outside:203.0.113.2/1 flags ri idle 0:00:08 timeout 0:00:00 ``` Read the flags: `r` means portmap (PAT) and `i` means dynamic. Each ICMP echo got its own xlate keyed to the ICMP identifier (the slash-number after the IP). Three of the four kept the same identifier on the outside; the fourth rolled to 26309 because the ASA detected a collision with a different inside source already using that ID. The wider hit counter lives in `show nat detail`. The PAT rule on INSIDE-NET (10.10.10.0/24) shows 46 translations since the ASA last rebooted: ``` ASA-PERIM# show nat detail Auto NAT Policies (Section 2) 1 (dmz) to (outside) source static DMZ-WEB 198.51.100.10 translate_hits = 0, untranslate_hits = 7 Source - Origin: 192.168.50.10/32, Translated: 198.51.100.10/32 2 (inside) to (outside) source dynamic INSIDE-TRANSIT interface translate_hits = 0, untranslate_hits = 0 Source - Origin: 10.10.0.0/24, Translated: 203.0.113.2/30 3 (inside) to (outside) source dynamic INSIDE-NET interface translate_hits = 46, untranslate_hits = 0 Source - Origin: 10.10.10.0/24, Translated: 203.0.113.2/30 ``` `translate_hits` increments when an inside-to-outside packet matches the rule and gets a new xlate. `untranslate_hits` increments when a return packet matches the xlate and gets the destination rewritten back. For pure outbound dynamic PAT, untranslate is always zero (return traffic uses the xlate, not the rule). ## Proving the Path with packet-tracer The fastest way to confirm a brand new PAT rule is doing what you think before any user reports an outage is to walk a hypothetical packet through the ASA with packet-tracer: ``` ASA-PERIM# packet-tracer input inside tcp 10.10.10.50 49000 8.8.8.8 80 Phase: 1 Type: INPUT-ROUTE-LOOKUP Subtype: Resolve Egress Interface Result: ALLOW Found next-hop 203.0.113.1 using egress ifc outside Phase: 2 Type: NAT Subtype: Result: ALLOW Config: object network INSIDE-NET nat (inside,outside) dynamic interface Additional Information: Dynamic translate 10.10.10.50/49000 to 203.0.113.2/49000 Phase: 3 Type: NAT Subtype: per-session Result: ALLOW Phase: 9 Type: FLOW-CREATION Subtype: Result: ALLOW New flow created with id 24, packet dispatched to next module Result: input-interface: inside input-status: up input-line-status: up output-interface: outside output-status: up output-line-status: up Action: allow ``` Phase 2 is the entire story: the ASA matched the packet against object INSIDE-NET, applied the dynamic NAT rule, and translated 10.10.10.50:49000 to 203.0.113.2:49000\. The same source port came through unchanged because no other flow needed it; on a busy ASA the source port often changes. Phase 9 confirms the conn was created and the packet would be forwarded to next-hop 203.0.113.1 on the outside interface. That is a working PAT path. For a deeper walk through every packet-tracer phase, see [Cisco ASA packet-tracer Command: Complete Troubleshooting Guide](https://www.pinglabz.com/cisco-asa-packet-tracer/). ## The Port Pool, and What Happens When It Fills A single outside IP gives you, in theory, 64,512 high-numbered ports per protocol (1024-65535). In practice the ASA reserves some ranges and divides ports across protocols, so the usable pool is smaller. For most edge deployments it is more than enough. If it is not enough (a SaaS company with thousands of concurrent sessions per IP, for example), the symptom shows up in syslog as `%ASA-3-202010: PAT pool exhausted` and the conn table starts dropping new flows. Two ways to fix it: 1. **Add a second mapped IP.** Build an `object network OUTSIDE-PAT-POOL` as a small range or list of public IPs, then point the dynamic NAT at that object instead of `interface`. The ASA round-robins across them. 2. **Per-session PAT (default on 9.x).** Modern ASA software releases ports back to the pool the moment a TCP flow ends, instead of holding them for the full PAT timeout. This is on by default and is one of the reasons you do not see PAT exhaustion as often as on older releases. ## Idle Timeouts You Should Know Dynamic PAT xlates do not live forever. The relevant timers (defaults shown): `timeout xlate` Default3:00:00 What it controls How long an idle xlate survives without traffic. `timeout pat-xlate` Default0:00:30 What it controls Specific to PAT xlates. The 30-second hold lets in-flight TCP teardowns complete before the port returns to the pool. `timeout conn` Default1:00:00 What it controls Idle TCP conn timeout. The xlate cannot disappear before the conn does. `timeout udp` Default0:02:00 What it controls UDP conn idle timeout. Lower than TCP because UDP is connectionless and the ASA cannot watch a teardown. The interaction matters: if a TCP flow goes idle for two hours, the conn ages out, then the xlate ages out, then the port returns to the pool. Until all of that happens, the inside host has burned a port on the outside IP. On a high-throughput edge this is one knob worth understanding. ## Failure Modes You Will Actually See The most common dynamic PAT failures, ranked by how often they cause an outage: - **Wrong interface order.** `nat (outside,inside)` instead of `(inside,outside)`. The rule does not match outbound traffic and inside hosts cannot reach the internet. Easy to spot in `show running-config object`; easy to miss when copy-pasting from a different ASA. - **Subnet mask mismatch.** The object subnet does not cover all the inside hosts. Some hosts work, others do not. Compare the `subnet` line under each object NAT with the actual inside subnets on the routing side. - **Routing not in place.** The packet matches the NAT rule but never gets a route to 0.0.0.0/0\. Check `show route` for a default route on outside. - **ACL denying the return.** Outbound packet leaves fine, return packet comes back and gets dropped at the outside interface ACL. The ACL must permit established sessions or the matching return flow. - **PAT pool exhaustion.** Already covered. Symptom: syslog 202010, slow new connections. Fix: bigger PAT pool or enable per-session PAT. The fastest diagnostic for "did the PAT rule match my traffic?" is always packet-tracer with the source IP and port the user reports. ## Where to Go Next - [Cisco ASA NAT Explained](https://www.pinglabz.com/cisco-asa-nat-explained/) for the full NAT picture (auto vs manual, sections, ordering). - [Cisco ASA Static NAT for Publishing a Server in the DMZ](https://www.pinglabz.com/cisco-asa-static-nat-dmz/) for the inverse case: outside-to-inside reachability for a public-facing server. - [Cisco ASA NAT Order of Operations Cheat Sheet](https://www.pinglabz.com/cisco-asa-nat-order-of-operations/) when more than one NAT rule could match a packet. - [Cisco ASA packet-tracer Command](https://www.pinglabz.com/cisco-asa-packet-tracer/) for the full troubleshooting walk. - [Cisco ASA pillar](https://www.pinglabz.com/cisco-asa/) for the cluster index. ## Key Takeaways - Dynamic PAT is the standard outbound NAT rule on the ASA: many inside hosts to one outside IP, distinguished by source port. - Object NAT (`nat (inside,outside) dynamic interface` nested inside an object network) is the modern, one-line way to write it. It lands in Section 2 of the NAT table. - The xlate table is the state. Watch it with `show xlate`, watch the rule itself with `show nat detail`. - packet-tracer is the fastest way to confirm a PAT rule matches the source you expect, before you touch live traffic. - The most common failure is interface order. Second most common is mask mismatch. Both are config bugs and both show up immediately in packet-tracer. ### Cisco ASA Site-to-Site IPsec VPN Configuration URL: https://www.pinglabz.com/cisco-asa-site-to-site-vpn/ Last updated: 2026-06-13T20:08:31.000Z Site-to-site IPsec VPN on the Cisco ASA is the most common way to interconnect two offices, a branch and headquarters, or a corporate network and a cloud VPC. The ASA has supported it since the PIX days, and the configuration model has evolved through three generations: classic IKEv1 with crypto maps (legacy), IKEv2 with crypto maps (current default), and route-based VTI (newer, route-aware). This article focuses on the current default - IKEv2 with crypto maps - because it is what 90% of production ASA tunnels use today, with one section on VTI for engineers who want the route-based model. It is part of the [Cisco ASA Complete Guide](https://www.pinglabz.com/cisco-asa/). Adjacent reads: [Troubleshoot Cisco ASA IPsec VPN Phase 1 and Phase 2](https://www.pinglabz.com/cisco-asa-troubleshoot-ipsec-phases/), [Cisco ASA VPN NAT Exemption: The Mistake That Breaks Tunnels](https://www.pinglabz.com/cisco-asa-vpn-nat-exemption/), and the modern remote-access VPN counterparts: [AnyConnect SSL VPN](https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/) and [AnyConnect IKEv2 VPN](https://www.pinglabz.com/cisco-asa-anyconnect-ikev2-vpn/). ## What the Tunnel Does A site-to-site IPsec VPN encrypts traffic between two sites so it can traverse an untrusted network (typically the internet) safely. The ASA at each end is configured with a peer (the remote ASA's public IP), an authentication method (pre-shared key or digital certificate), and a list of "interesting traffic" - the source/destination subnets that should be encrypted and sent through the tunnel. When a packet matches interesting traffic, the ASA encrypts it and tunnels it to the peer. The peer decrypts and forwards. Two things make this work: - **The IKE control plane.** Negotiates the encryption keys (Phase 1) and the security associations for the tunnel (Phase 2). IKEv2 is the modern protocol; IKEv1 is legacy. - **NAT exemption.** Interesting traffic must NOT be PATted to the outside interface IP, or it will exit the firewall translated and the peer will reject it. Identity NAT in Section 1 keeps the traffic untranslated. We cover this in detail in the [NAT exemption article](https://www.pinglabz.com/cisco-asa-vpn-nat-exemption/). ## IKEv2 Site-to-Site Walkthrough (Modern Default) The minimal IKEv2 site-to-site config has six pieces: an IKEv2 policy, an IPsec proposal, a tunnel-group, a crypto map, NAT exemption, and the interesting-traffic ACL. Here it is end-to-end on the local ASA (peer 198.51.100.2): ``` ! 1. IKEv2 policy (Phase 1 parameters) crypto ikev2 policy 10 encryption aes-256 integrity sha256 group 14 prf sha256 lifetime seconds 86400 ! ! Enable IKEv2 on the outside interface crypto ikev2 enable outside ! ! 2. IPsec proposal (Phase 2 parameters) crypto ipsec ikev2 ipsec-proposal AES256-SHA256 protocol esp encryption aes-256 protocol esp integrity sha-256 ! ! 3. Tunnel-group (peer-specific config + auth) tunnel-group 198.51.100.2 type ipsec-l2l tunnel-group 198.51.100.2 ipsec-attributes ikev2 remote-authentication pre-shared-key STRONG_SECRET ikev2 local-authentication pre-shared-key STRONG_SECRET ! ! 4. Interesting-traffic ACL object network LOCAL-LAN subnet 10.10.10.0 255.255.255.0 object network REMOTE-LAN subnet 10.20.0.0 255.255.255.0 ! access-list VPN-ACL extended permit ip object LOCAL-LAN object REMOTE-LAN ! ! 5. NAT exemption (must be Section 1 manual NAT) nat (inside,outside) source static LOCAL-LAN LOCAL-LAN destination static REMOTE-LAN REMOTE-LAN no-proxy-arp route-lookup ! ! 6. Crypto map binding ACL + proposal + peer crypto map OUTSIDE-MAP 10 match address VPN-ACL crypto map OUTSIDE-MAP 10 set peer 198.51.100.2 crypto map OUTSIDE-MAP 10 set ikev2 ipsec-proposal AES256-SHA256 crypto map OUTSIDE-MAP 10 set security-association lifetime seconds 3600 ! crypto map OUTSIDE-MAP interface outside ``` Read it as a wiring diagram: outbound packets that match `VPN-ACL` (LOCAL-LAN to REMOTE-LAN) are matched by crypto map sequence 10, encrypted with the AES256-SHA256 proposal, and sent to peer 198.51.100.2\. The peer authenticates with pre-shared key. NAT exemption keeps the source and destination untranslated. The remote ASA is the mirror image: same IKEv2 policy and IPsec proposal, tunnel-group keyed on the local ASA's outside IP, ACL with source and destination flipped, NAT exemption mirrored, crypto map pointing at the local ASA. Both sides must agree on encryption parameters, otherwise Phase 1 or Phase 2 negotiation fails. ## Verifying the Tunnel The four most useful show commands once the config is in: ``` ! Phase 1 - IKEv2 SA ASA-PERIM# show crypto ikev2 sa IKEv2 SAs: Session-id:1, Status:UP-ACTIVE, IKE count:1, CHILD count:1 Tunnel-id Local Remote Status Role 123456789 203.0.113.2/500 198.51.100.2/500 READY INITIATOR Encr: AES-CBC, keysize: 256, Hash: SHA256, DH Grp:14, Auth sign: PSK, Auth verify: PSK Life/Active Time: 86400/1234 sec Child sa: local selector 10.10.10.0/0 - 10.10.10.255/65535 remote selector 10.20.0.0/0 - 10.20.0.255/65535 ESP spi in/out: 0xa1b2c3d4/0xe5f6a7b8 ! Phase 2 - IPsec SA ASA-PERIM# show crypto ipsec sa interface: outside Crypto map tag: OUTSIDE-MAP, seq num: 10, local addr: 203.0.113.2 access-list VPN-ACL extended permit ip 10.10.10.0/24 10.20.0.0/24 local ident (addr/mask/prot/port): (10.10.10.0/255.255.255.0/0/0) remote ident (addr/mask/prot/port): (10.20.0.0/255.255.255.0/0/0) current_peer: 198.51.100.2 #pkts encaps: 4521, #pkts encrypt: 4521, #pkts digest: 4521 #pkts decaps: 4118, #pkts decrypt: 4118, #pkts verify: 4118 ... ! Verify tunnel is actively passing data (decaps/encaps incrementing) ASA-PERIM# show vpn-sessiondb l2l ! What does my interesting-traffic ACL look like? ASA-PERIM# show access-list VPN-ACL ``` If `show crypto ikev2 sa` shows Status: UP-ACTIVE and `show crypto ipsec sa` shows incrementing encaps and decaps counters, the tunnel is up and passing data. If encaps is incrementing but decaps is not, the local side is sending but the remote side is not (or the return is not coming back to you - check ACLs at the remote end). If decaps is incrementing but encaps is not, the local side is receiving but not sending - check the interesting-traffic ACL and NAT exemption. ## Pre-Shared Key vs Certificate Authentication Pre-shared key (PSK) Pros Trivial to configure, no PKI, works in 30 seconds Cons One leak compromises every tunnel using the same key, no automated rotation Common use Branch-to-HQ, partner extranet, lab Certificate (RSA) Pros Per-peer trust, automated rotation possible, scales to many peers Cons Requires PKI infrastructure, cert lifecycle management Common use Enterprise hub-and-spoke with many spokes, security-sensitive deployments For two-site deployments, PSK is fine. For 50+ spokes, certificates pay back the PKI investment in operational hygiene. The configuration shape is similar: replace the `pre-shared-key` lines with `ikev2 remote-authentication rsa-sig` and `ikev2 local-authentication certificate `, plus a trustpoint configuration that points at your CA. ## Route-Based VPN: VTI Modern ASA software (9.7+) supports Virtual Tunnel Interfaces (VTI), which give you a route-based VPN: the tunnel is a logical interface, you put routes through it, and any traffic that route-lookups to the VTI is encrypted. No interesting-traffic ACL, no NAT exemption. ``` ! IPsec profile (replaces crypto map for VTI) crypto ipsec profile VPN-PROFILE set ikev2 ipsec-proposal AES256-SHA256 ! ! Tunnel interface interface Tunnel1 nameif vti-to-remote ip address 169.254.1.1 255.255.255.252 tunnel source interface outside tunnel destination 198.51.100.2 tunnel mode ipsec ipv4 tunnel protection ipsec profile VPN-PROFILE ! ! Tunnel-group (same as crypto map case) tunnel-group 198.51.100.2 type ipsec-l2l tunnel-group 198.51.100.2 ipsec-attributes ikev2 remote-authentication pre-shared-key STRONG_SECRET ikev2 local-authentication pre-shared-key STRONG_SECRET ! ! Route remote LAN through the VTI route vti-to-remote 10.20.0.0 255.255.255.0 169.254.1.2 1 ``` VTI is the modern clean model and pairs naturally with dynamic routing (you can run BGP or OSPF over the tunnel for failover and route exchange). For a fixed two-site deployment with no routing requirements, crypto map is still fine and is what most production ASAs use. ## Common Mistakes - **Forgetting NAT exemption.** The biggest single source of "tunnel is up but no traffic" tickets. Outbound packets get PATted to the outside interface IP before reaching the crypto map. The fix is identity NAT in Section 1\. We dedicate a whole article to it: [Cisco ASA VPN NAT Exemption](https://www.pinglabz.com/cisco-asa-vpn-nat-exemption/). - **Mismatched encryption proposals.** Both sides must agree on encryption, integrity, DH group, and lifetime. Tiny mismatches (SHA1 on one side, SHA256 on the other) cause Phase 1 or Phase 2 to fail to negotiate. Check both sides agree before debugging the tunnel state machine. - **ACL mismatch on local vs remote.** The interesting-traffic ACLs must be exact mirrors. Local ACL says "permit 10.10.10.0/24 to 10.20.0.0/24", remote ACL must say "permit 10.20.0.0/24 to 10.10.10.0/24". A subnet mismatch (e.g., remote is /23 but local says /24) causes the tunnel to come up but only partial traffic to encrypt. - **No outside ACL allowing IKE/ESP.** If your outside interface has a restrictive inbound ACL, you need to permit UDP/500 (IKE), UDP/4500 (NAT-T), and ESP (protocol 50) from the peer. Most engineers add this and forget when they later harden the perimeter. - **PSK on both sides differ.** Phase 1 will fail authentication. Verify with `show crypto ikev2 sa` showing AUTH\_FAILED, and the remote side's logs. - **Routing-table thinks the destination is somewhere else.** If the inside router has a route to the remote subnet via a path that does not go through the ASA (a leftover EIGRP route, an old static), the inside hosts will route that way and never hit the firewall. Verify with traceroute from an inside host. ## Related Articles - [Troubleshoot Cisco ASA IPsec VPN Phase 1 and Phase 2](https://www.pinglabz.com/cisco-asa-troubleshoot-ipsec-phases/) \- debug commands and the negotiation flow. - [Cisco ASA VPN NAT Exemption: The Mistake That Breaks Tunnels](https://www.pinglabz.com/cisco-asa-vpn-nat-exemption/) \- the most common single failure mode. - [Cisco ASA AnyConnect SSL VPN Configuration](https://www.pinglabz.com/cisco-asa-anyconnect-ssl-vpn/) \- the remote-access counterpart. - [Cisco ASA AnyConnect IKEv2 VPN Configuration](https://www.pinglabz.com/cisco-asa-anyconnect-ikev2-vpn/) \- the remote-access IKEv2 counterpart. - [Cisco ASA VPN Group Policies and Tunnel Groups](https://www.pinglabz.com/cisco-asa-vpn-group-policies/) \- the abstraction underneath both site-to-site and remote-access. - [Cisco ASA packet-tracer Guide](https://www.pinglabz.com/cisco-asa-packet-tracer/) \- the simulator that walks the encrypt/decrypt phases. ## Key Takeaways An ASA site-to-site IPsec VPN has six pieces of config: IKEv2 policy, IPsec proposal, tunnel-group, interesting-traffic ACL, NAT exemption, and crypto map (or VTI for the route-based model). IKEv2 with PSK and AES-256/SHA-256/DH-14 is the modern default for two-site deployments. Verify with `show crypto ikev2 sa` for Phase 1 and `show crypto ipsec sa` for Phase 2 (encaps and decaps counters incrementing means traffic is flowing). The single most common failure is forgetting NAT exemption: outbound traffic gets PATted to the outside interface IP before reaching the crypto map and the tunnel never carries it. VTI is the modern cleaner model when you want routing over the tunnel. ### Cisco ASA Security Levels Explained: Inside, Outside, DMZ URL: https://www.pinglabz.com/cisco-asa-security-levels/ Last updated: 2026-06-13T20:08:31.000Z Security levels are the single most distinctive feature of Cisco ASA configuration. Every ASA interface gets a numeric value from 0 (lowest trust, typically the internet) to 100 (highest trust, typically your inside LAN), and the default forwarding behavior between interfaces is determined entirely by that number. Get the security-level model right and the rest of the ASA configuration falls into place. Get it wrong and you spend hours debugging traffic that "should be allowed" but is not. This article is part of the [Cisco ASA Complete Guide](https://www.pinglabz.com/cisco-asa/) on PingLabz. We cover the convention, the default behavior, the same-security-level cases, and the common mistakes engineers make on day one. ## The Convention: 0, 50, 100 Security level is a number from 0 to 100 assigned to each ASA interface with the `security-level` command. The number itself is arbitrary - 100 is just "more trusted than 50, which is more trusted than 0." The convention almost every shop follows: inside (trusted LAN) Security level100 Notes Highest trust. ASA gives this to `nameif inside` by default. dmz (semi-trusted, public-facing servers) Security level50 Notes Convention; can be 30, 40, 50 - any number lower than inside, higher than outside. partner / extranet Security level20-25 Notes Limited trust, partner organization access. guest / iot Security level10-15 Notes Untrusted internal segments. outside (untrusted, internet) Security level0 Notes Lowest trust. ASA gives this to `nameif outside` by default. The numbers are spaced so you can insert new zones between existing ones (a partner network at 25, a guest network at 15) without renumbering everything. There is nothing magical about 100, 50, and 0 - those are just the well-known defaults from sample configurations dating back to the PIX era. ## Default Behavior: High to Low Allowed, Low to High Denied The implicit forwarding rules between interfaces are determined entirely by their security levels: - **Higher security to lower security: implicit allow.** A packet from inside (100) to outside (0) is permitted by default, with NAT and a connection-table entry. No explicit ACL needed. - **Lower security to higher security: implicit deny.** A packet from outside (0) to inside (100) is dropped by default. You must apply an explicit ACL to permit specific flows. - **Same security to same security: implicit deny.** Two interfaces at the same level cannot pass traffic to each other unless you globally configure `same-security-traffic permit inter-interface`. - **Same interface ingress and egress (hairpinning): implicit deny.** Traffic that comes in and goes out the same interface is dropped unless you globally configure `same-security-traffic permit intra-interface`. Used for inter-VPN routing through the ASA, hub-and-spoke through a single interface, and the ASA's internal VPN-to-VPN forwarding. Confirm any flow with packet-tracer - the first phase reports whether the implicit rule allowed or dropped, and which security-level relationship caused it. We cover that workflow in [Cisco ASA packet-tracer Command: Complete Troubleshooting Guide](https://www.pinglabz.com/cisco-asa-packet-tracer/). ## Three-Interface Example The classic ASA with inside (100), dmz (50), outside (0): ``` interface GigabitEthernet0/0 nameif inside security-level 100 ip address 10.10.0.254 255.255.255.0 interface GigabitEthernet0/1 nameif outside security-level 0 ip address 203.0.113.2 255.255.255.252 interface GigabitEthernet0/2 nameif dmz security-level 50 ip address 192.168.50.1 255.255.255.0 ``` Verify with `show nameif`: ``` ASA-PERIM# show nameif Interface Name Security GigabitEthernet0/0 inside 100 GigabitEthernet0/1 outside 0 GigabitEthernet0/2 dmz 50 Management0/0 management 0 ``` Default forwarding between these three interfaces: inside Tooutside Direction100 → 0 DefaultAllow + NAT Common explicit override Sometimes restricted with an outbound ACL on inside. inside Todmz Direction100 → 50 DefaultAllow + NAT Common explicit override Often restricted to specific management ports. dmz Tooutside Direction50 → 0 DefaultAllow + NAT Common explicit override Often restricted to specific egress (DNS, NTP, package repos). outside Toinside Direction0 → 100 DefaultDeny Common explicit override Permit specific destinations only via inbound ACL. outside Todmz Direction0 → 50 DefaultDeny Common explicit override Permit specific public services (HTTP, HTTPS, SMTP) via inbound ACL. dmz Toinside Direction50 → 100 DefaultDeny Common explicit override This is the security boundary. Permit only what is strictly necessary (e.g., a DMZ web server hitting an inside database server on a specific port). The dmz-to-inside denied default is exactly the point of having a DMZ. If a public web server is compromised, the attacker can reach the internet but cannot pivot to the inside network. ## Same Security: Two Cases Two interfaces with the same security level (say, two DMZs both at 50, or two partner networks both at 25) cannot pass traffic to each other unless you explicitly enable it: ``` same-security-traffic permit inter-interface ``` This is one global command, not per-interface. Once enabled, all same-security-level pairs can talk, subject to ACLs. Use this when you have multiple zones at the same trust tier that need to communicate (multiple DMZs, multiple branch offices). The same-interface (intra-interface) case is rarer. Traffic that comes in an interface and needs to leave the same interface is denied by default. You enable it with: ``` same-security-traffic permit intra-interface ``` Use cases: hub-and-spoke VPN where multiple spokes terminate on the same interface and need to reach each other through the ASA, or WCCP redirects. Get the Cisco ASA Field Reference - 9 pages, free Everything you'd want to remember about Cisco ASA on nine printable pages. Per-packet pipeline diagram, NAT 8.3+ section ordering, six-branch troubleshooting decision tree, real lab show-output annotated, paste-ready three-zone config. Free for PingLabz members - just sign up with your email. [Get the Cisco ASA cheat-sheet](https://www.pinglabz.com/cisco-asa-cheatsheet/) ## The dmz-to-inside Pattern: Restrictive by Default The most common ACL pattern engineers configure is for traffic going dmz-to-inside (50 to 100). The default is deny, so you write a small, very specific permit list and apply it inbound on the dmz interface. ``` object-group network DMZ-WEB-SERVERS network-object host 192.168.50.10 network-object host 192.168.50.11 object-group network INSIDE-DB-SERVERS network-object host 10.10.0.20 access-list DMZ_IN extended permit tcp object-group DMZ-WEB-SERVERS object-group INSIDE-DB-SERVERS eq 3306 access-list DMZ_IN extended permit udp object-group DMZ-WEB-SERVERS host 10.10.0.10 eq 53 access-list DMZ_IN extended permit udp object-group DMZ-WEB-SERVERS host 10.10.0.10 eq 123 access-list DMZ_IN extended deny ip any any log access-group DMZ_IN in interface dmz ``` Five lines: web servers can hit the database on MySQL, do DNS to a specific resolver, do NTP to a specific server, and everything else is denied with logging. That is what defense-in-depth looks like at the firewall layer. ## Common Mistakes - **Two interfaces at level 100.** Inside is 100\. If you accidentally configure another interface at 100 (typing too fast and copying from a sample), traffic between them is denied by default until you enable `same-security-traffic permit inter-interface`. Engineers often spot this only by reading `show nameif`. - **"Inside is 0 because outside is 100."** Inverted convention. Outside is the lowest (untrusted) and gets 0; inside is the highest and gets 100\. Reversing this changes implicit forwarding direction and breaks everything. - **Forgetting the inbound ACL on outside.** Default is implicit deny for low-to-high. If you have a static NAT publishing a DMZ server but no ACL allowing the traffic, the public will hit the static NAT (UN-NAT phase succeeds) and then get dropped by the implicit deny in the ACL phase. Run packet-tracer once to confirm. - **Allowing same-security inter-interface globally without thinking.** If you have multiple zones at the same level and enable `same-security-traffic permit inter-interface`, those zones can suddenly all talk to each other. Verify each pair has an explicit ACL before flipping the global switch. - **Confusing security level with policy.** Security level only sets the default. An explicit ACL always overrides the default. The level is a hint to your future self about the trust relationship, not a substitute for ACL design. ## Related Articles - [Cisco ASA ACL Configuration: Inbound Rules and Object Groups](https://www.pinglabz.com/cisco-asa-acl-configuration/) \- the explicit override mechanism. - [Cisco ASA Packet Flow](https://www.pinglabz.com/cisco-asa-packet-flow/) \- the data path that uses these defaults. - [Cisco ASA packet-tracer Guide](https://www.pinglabz.com/cisco-asa-packet-tracer/) \- confirms which implicit or explicit rule applies to a given flow. - [Cisco ASA Routed Mode vs Transparent Mode](https://www.pinglabz.com/cisco-asa-routed-vs-transparent/) \- mode-level context. - [Cisco ASA Inside/Outside/DMZ Configuration Walkthrough](https://www.pinglabz.com/cisco-asa-inside-outside-dmz/) \- the worked example end-to-end. ## Key Takeaways Security level is a 0-to-100 number per interface that determines default forwarding behavior. Higher to lower is allowed by default; lower to higher and same-to-same are denied. The convention is inside 100, dmz 50, outside 0, with custom zones in between. Two interfaces at the same level cannot talk unless you enable `same-security-traffic permit inter-interface`. The dmz-to-inside boundary is the most security-critical: keep its inbound ACL strict, log denies, and verify with packet-tracer. ### Cisco ASA ACL Configuration: Inbound Rules and Object Groups URL: https://www.pinglabz.com/cisco-asa-acl-configuration/ Last updated: 2026-06-13T20:08:31.000Z ACLs on the Cisco ASA work the same as on IOS routers in spirit but with three differences that matter in practice: ASA ACLs are almost always applied inbound, they reference the **real** (untranslated) destination IP rather than the post-NAT mapped IP, and object groups are heavily used to keep them readable. This article walks the full ASA ACL model, the syntax, the object-group pattern, and the common mistakes engineers coming from IOS hit on day one. It is part of the [Cisco ASA Complete Guide](https://www.pinglabz.com/cisco-asa/) on PingLabz. Adjacent reads: [Cisco ASA Packet Flow](https://www.pinglabz.com/cisco-asa-packet-flow/) for where the ACL fits in the data path, [ASA vs Router ACLs](https://www.pinglabz.com/asa-vs-router-acls/) for the IOS-engineer perspective, and [ACL Troubleshooting with packet-tracer](https://www.pinglabz.com/cisco-asa-acl-troubleshooting/) for incident response. ## ASA ACL Basics An ASA ACL is a named, ordered list of permit/deny rules applied to an interface in the inbound direction. Rules are evaluated top to bottom, first match wins, with an implicit `deny ip any any` at the end. The structure: ``` access-list NAME extended {permit | deny} protocol src dst [eq port] access-group NAME in interface INTERFACE-NAME ``` Note `extended` in the syntax. ASA also supports standard, ethertype, and webtype ACLs but extended is what 99% of engineers configure. `extended` means "matches on protocol, source, destination, and port" - the fully featured ACL type. ## The Real-IP Rule: ASA ACLs Match Untranslated Destinations This is the single most-confusing detail for engineers coming from older ASA software (pre-8.3) or from IOS routers. On modern ASA, the inbound ACL on an interface evaluates the **real** (post-untranslate) destination, not the mapped (public) destination. Concretely: a DMZ web server at 192.168.50.10 is published as 198.51.100.10\. The packet arrives on the outside interface destined for 198.51.100.10\. The ASA's UN-NAT phase reverses the static NAT and now sees 192.168.50.10 as the destination. Then the ACL evaluates against 192.168.50.10\. Your `OUTSIDE_IN` ACL must permit traffic to 192.168.50.10, not 198.51.100.10. ``` ! WRONG (old ASA pre-8.3 mental model) access-list OUTSIDE_IN extended permit tcp any host 198.51.100.10 eq 80 ! RIGHT (modern ASA real-IP rule) access-list OUTSIDE_IN extended permit tcp any host 192.168.50.10 eq 80 ``` Or, with object groups (the recommended pattern): ``` object network DMZ-WEB host 192.168.50.10 nat (dmz,outside) static 198.51.100.10 access-list OUTSIDE_IN extended permit tcp any object DMZ-WEB eq 80 access-list OUTSIDE_IN extended permit tcp any object DMZ-WEB eq 443 access-group OUTSIDE_IN in interface outside ``` The `object DMZ-WEB` reference resolves to 192.168.50.10 (the host-object value), so the rule is internally consistent and the public IP shows up only once, in the NAT statement. This is the pattern modern ASA configurations should use everywhere. ## Object Groups: Keep ACLs Readable Object groups bundle related addresses, ports, or protocols into a single named group. They are the difference between a 5-rule ACL and a 50-rule ACL doing the same thing. Three types matter most: network Contents IP addresses, subnets, ranges, network-objects Example Group of all DMZ servers service Contents TCP/UDP ports, port ranges, ICMP types Example Web ports (80, 443, 8080), ICMP types protocol Contents IP protocols (tcp, udp, icmp, gre, esp) ExampleVPN protocols A typical compact rule using all three: ``` object-group network DMZ-WEB-SERVERS network-object host 192.168.50.10 network-object host 192.168.50.11 network-object host 192.168.50.12 object-group service WEB-PORTS tcp port-object eq 80 port-object eq 443 port-object eq 8080 access-list OUTSIDE_IN extended permit tcp any object-group DMZ-WEB-SERVERS object-group WEB-PORTS access-group OUTSIDE_IN in interface outside ``` That single rule expands internally into 9 logical rules (3 servers × 3 ports), but you maintain one line. Add a server, add a port - the change is in one place. Engineers come back to ACLs months later and the intent is still legible. ## Verifying with show access-list `show access-list NAME` prints the ACL with per-line hit counters. This is how you find which rules are actually firing and which are dead. ``` ASA-PERIM# show access-list OUTSIDE_IN access-list OUTSIDE_IN; 3 elements; name hash: 0xe01d8199 access-list OUTSIDE_IN line 1 extended permit tcp any object DMZ-WEB eq www (hitcnt=234) 0xf18a0028 access-list OUTSIDE_IN line 1 extended permit tcp any host 192.168.50.10 eq www (hitcnt=234) 0xf18a0028 access-list OUTSIDE_IN line 2 extended permit tcp any object DMZ-WEB eq https (hitcnt=4521) 0x1ce50e7b access-list OUTSIDE_IN line 2 extended permit tcp any host 192.168.50.10 eq https (hitcnt=4521) 0x1ce50e7b access-list OUTSIDE_IN line 3 extended permit icmp any object DMZ-WEB (hitcnt=12) 0x2cc9dc3f access-list OUTSIDE_IN line 3 extended permit icmp any host 192.168.50.10 (hitcnt=12) 0x2cc9dc3f ``` This is real output from our ASAv 9.23 lab. Each rule that uses a network object prints twice: the configured form (with the object name) and the resolved form (with the object's actual address). That nested view is helpful when an object has multiple network-object members - you can confirm each expansion. 4521 hits on the HTTPS line, 234 on HTTP, 12 on ICMP. If a rule has zero hits over weeks, it is either redundant or never matches - candidate for cleanup. To reset the counters: `clear access-list OUTSIDE_IN counters`. ## Direction: Almost Always Inbound ASA ACLs are typically applied **inbound** on the interface where the traffic enters. Outbound ACLs exist but are rare; their use case is filtering specific egress traffic regardless of source interface, and most security models prefer to filter at ingress where the trust level is lowest. The convention follows: an ACL named `OUTSIDE_IN` is applied inbound on the outside interface. `INSIDE_IN` is applied inbound on inside, etc. Some shops use `OUTSIDE_ACL` with no direction in the name; both conventions are fine but consistency within an environment matters more than which one you pick. ## Default Behavior Without an ACL If no ACL is applied to an interface, the ASA uses its security-level defaults: - Higher-security to lower-security: allow. - Lower-security to higher-security: deny. - Same-security to same-security: deny unless `same-security-traffic permit inter-interface` is configured. This means out of the box, with three interfaces (inside 100, dmz 50, outside 0): inside Tooutside DefaultAllow inside Todmz DefaultAllow dmz Tooutside DefaultAllow outside Toinside DefaultDeny outside Todmz DefaultDeny dmz Toinside DefaultDeny The implications: outbound (inside-to-outside) typically does not need an explicit ACL because the implicit allow handles it. Inbound (outside-to-inside or outside-to-dmz) almost always does need an ACL to permit anything to traverse. We cover this in detail in [Cisco ASA Security Levels Explained: Inside, Outside, DMZ](https://www.pinglabz.com/cisco-asa-security-levels/). Get the Cisco ASA Field Reference - 9 pages, free Everything you'd want to remember about Cisco ASA on nine printable pages. Per-packet pipeline diagram, NAT 8.3+ section ordering, six-branch troubleshooting decision tree, real lab show-output annotated, paste-ready three-zone config. Free for PingLabz members - just sign up with your email. [Get the Cisco ASA cheat-sheet](https://www.pinglabz.com/cisco-asa-cheatsheet/) ## Logging: Inline and Per-Rule Add `log` at the end of any ACL line to log every match to syslog at the level you specify (default is informational, level 6): ``` access-list OUTSIDE_IN extended permit tcp any object DMZ-WEB eq 443 log access-list OUTSIDE_IN extended deny ip any any log warnings ``` Logging every permit is expensive at scale; logging the implicit deny (or an explicit deny-all at the end) is cheap and very useful for incident triage and security event review. Send these to a SIEM via [syslog configuration](https://www.pinglabz.com/cisco-asa-syslog-logging/). ## Common Mistakes - **Writing the ACL against the public IP.** Use the real IP. Run packet-tracer once and read the UN-NAT phase to confirm what the ACL will compare. - **Forgetting `access-group ... in interface`.** The `access-list` command alone does nothing. The ACL must be applied with `access-group` to be evaluated. - **Wrong direction.** Always use `in interface` unless you have a specific reason to filter outbound. Outbound ACLs are correct sometimes but they confuse the next engineer. - **Implicit deny surprises.** When you add a single permit line to an empty ACL, the implicit deny suddenly applies to everything else. If your goal was to add one allow without breaking the wide-open policy, you needed a `permit ip any any` at the end first. - **ACL on the wrong interface.** Inbound traffic to a DMZ server enters on the outside interface, not the dmz interface. Apply the ACL on outside. - **Using ACL line numbers without explicit insertion.** Modern ASA software lets you insert at a position with `access-list OUTSIDE_IN line 3 extended ...`. Without the line number, new rules are appended at the end - which can put them after a deny and make them dead. - **Not using object groups.** A 50-line ACL with repeated host addresses is unmaintainable. Group them. ## Related Articles - [Cisco ASA ACL Troubleshooting with packet-tracer](https://www.pinglabz.com/cisco-asa-acl-troubleshooting/) \- the incident-response playbook. - [Cisco ASA Object Groups: Network, Service, and Protocol Objects](https://www.pinglabz.com/cisco-asa-object-groups/) \- the deep dive on the bundling syntax. - [ASA vs Router ACLs: What Network Engineers Get Wrong](https://www.pinglabz.com/asa-vs-router-acls/) \- the IOS-engineer comparison. - [Cisco ASA Security Levels Explained](https://www.pinglabz.com/cisco-asa-security-levels/) \- default behavior context. - [Cisco ASA Packet Flow](https://www.pinglabz.com/cisco-asa-packet-flow/) \- where the ACL sits in the data path. - [Cisco ASA Cheat Sheet](https://www.pinglabz.com/cisco-asa-cheat-sheet/) \- quick command reference. ## Key Takeaways ASA ACLs are inbound, named, ordered lists with an implicit deny at the end. They reference real (untranslated) destination IPs, which is the single biggest gotcha for engineers used to old ASA or to IOS routers. Object groups make ACLs maintainable by bundling addresses and ports under named references. Default security-level behavior handles outbound (high-to-low) for free; inbound (low-to-high) almost always requires an explicit ACL. Use `show access-list` to verify hit counts and find dead rules, and packet-tracer to confirm the ACL is doing what you expect. ### Cisco ASA Packet Flow: From Interface ACL to NAT to Route Lookup URL: https://www.pinglabz.com/cisco-asa-packet-flow/ Last updated: 2026-06-13T20:08:32.000Z Every packet that enters a Cisco ASA goes through a deterministic series of checks before it is forwarded, dropped, or handed to the VPN engine. Understanding that order is the difference between fixing an outage in two minutes and chasing the wrong layer for an hour. This article walks the full ASA packet flow on software 9.x, with a worked example from our lab, and pins each step to the configuration block you would inspect when something fails. It is part of the [Cisco ASA Complete Guide](https://www.pinglabz.com/cisco-asa/) on PingLabz. After this read, the next stop is usually [Cisco ASA packet-tracer](https://www.pinglabz.com/cisco-asa-packet-tracer/), which simulates this exact flow against any source-destination pair you give it. ## The Nine Steps of ASA Packet Flow For a TCP or UDP packet entering an ASA interface, the data plane evaluates these phases in order: Existing connection lookup Step1 What happens If the 5-tuple matches an existing connection in the state table, skip directly to step 7. Configuration that controls it Connection table (`show conn`), TCP/UDP/ICMP timeouts. IP options / fragment / TCP-state checks Step2 What happens Drops malformed packets, fragmented packets without an existing flow if assembly is required, TCP packets that violate state-machine expectations. Configuration that controls it `fragment`, `set connection`, default state check. NAT untranslate (UN-NAT) Step3 What happens If the destination is a translated address, reverse the NAT to find the real destination. The ACL check uses the real destination. Configuration that controls it Auto NAT static rules, Manual NAT static rules. ACL check Step4 What happens Evaluate the inbound ACL against the real source and real destination. Configuration that controls it `access-list` \+ `access-group ... in interface`. NAT translate Step5 What happens Apply the forward NAT rule (auto or manual). Source is translated, destination may also be translated. Configuration that controls it Same as UN-NAT, but evaluated forward. Route lookup Step6 What happens Pick the egress interface based on the global routing table. Configuration that controls it Static routes, OSPF, BGP, EIGRP. Adjacency / ARP Step7 What happens Resolve the next-hop MAC address. If no adjacency, drop with `no-adjacency`. Configuration that controls it`arp`, ND for IPv6. Egress interface checks Step8 What happens Output ACL (rare), QoS policies, inspection engines (FTP, HTTP, SIP, etc.). Configuration that controls it MPF (modular policy framework), inspection maps. Forward Step9 What happens Packet is sent out the egress interface. Connection-table entry is updated or created. Configuration that controls it None - this is the success path. Several things to notice about this sequence: - **Existing connections short-circuit the whole flow.** Once a session is permitted and added to the state table, subsequent packets in either direction skip the security and ACL checks entirely. This is what "stateful firewall" means in practice. Return packets do not need a return ACL. - **NAT untranslate happens before the ACL check.** This is the single biggest difference from IOS routers. Your ACL rules permit destinations by their real (untranslated) IP, even on the outside interface where the destination hits as a public NAT address. - **NAT translate happens between ACL and route lookup.** The packet is translated to its egress form before the route is selected. This matters for asymmetric path scenarios and for VPN policies. - **Route lookup is global, not per-VRF.** Standard ASA software does not have VRFs. Multiple security contexts give per-context routing tables, but a single context has one global table. ## Worked Example: Inside to Outside (PAT) An inside host at 10.10.10.50 opens a TCP connection to 8.8.8.8 on port 443\. The ASA flow: 1. **Existing-connection lookup.** The first packet (TCP SYN) does not match any existing connection. Continue. 2. **Sanity checks.** SYN flag set, no IP options to drop on, IP TTL above 1\. Continue. 3. **UN-NAT.** Destination is 8.8.8.8, no static NAT or twice NAT translates 8.8.8.8 to anything else. Real destination remains 8.8.8.8. 4. **ACL check.** Inside is security-level 100, outside is 0\. The implicit rule (high-to-low) allows. No explicit ACL needed on the inside interface for this traffic. 5. **NAT translate.** Source 10.10.10.50 matches Auto NAT under `object network INSIDE-NET` with rule `nat (inside,outside) dynamic interface`. Source is translated to the outside interface IP (203.0.113.2) with PAT. 6. **Route lookup.** 8.8.8.8 matches the default route via 203.0.113.1, egress interface outside. 7. **Adjacency.** ARP resolves 203.0.113.1. 8. **Egress checks.** No outbound ACL on outside, no inspection that matters here. 9. **Forward.** Packet is sent. A new connection-table entry is created keyed on (10.10.10.50:src-port, 8.8.8.8:443) plus the translated (203.0.113.2:src-port, 8.8.8.8:443) tuple. The reply will match this entry and skip steps 2-6. This is what packet-tracer would show as a 6-phase ALLOW sequence. Read [Cisco ASA packet-tracer Command: Complete Troubleshooting Guide](https://www.pinglabz.com/cisco-asa-packet-tracer/) for the syntax that simulates exactly this flow on demand. ## Worked Example: Outside to DMZ (Static NAT) A user on the internet hits 198.51.100.10 (the public IP of a DMZ web server) on TCP/443\. The ASA flow: 1. **Existing-connection lookup.** First packet, no match. 2. **Sanity checks.** Pass. 3. **UN-NAT.** Destination 198.51.100.10 matches the static NAT under `object network DMZ-WEB`. Real destination is 192.168.50.10\. The packet's destination IP is conceptually replaced from now until step 5 forward. 4. **ACL check.** Outside ACL `OUTSIDE_IN` evaluates against destination 192.168.50.10 (real IP) on TCP/443\. Permit line matches: `permit tcp any object DMZ-WEB eq 443`. Continue. 5. **NAT translate.** Forward NAT rule for the same DMZ-WEB object. Translation is symmetric (same rule provides both untranslate and translate). 6. **Route lookup.** 192.168.50.10 is on directly-connected subnet 192.168.50.0/24, egress interface dmz. 7. **Adjacency.** ARP resolves 192.168.50.10. 8. **Egress checks.** No DMZ outbound ACL, default inspection. 9. **Forward.** A connection-table entry is created. The reply path will match it and skip ACL. The critical insight: the ACL on outside permits to the **real** IP 192.168.50.10, not the public IP 198.51.100.10\. This is the single most-misunderstood ASA detail. Get the Cisco ASA Field Reference - 9 pages, free Everything you'd want to remember about Cisco ASA on nine printable pages. Per-packet pipeline diagram, NAT 8.3+ section ordering, six-branch troubleshooting decision tree, real lab show-output annotated, paste-ready three-zone config. Free for PingLabz members - just sign up with your email. [Get the Cisco ASA cheat-sheet](https://www.pinglabz.com/cisco-asa-cheatsheet/) ## When Things Fail: Where in the Flow It Drops The phase that drops a packet tells you which configuration block to inspect. The mapping: UN-NAT Static NAT rule. Did you forget to publish the public IP? Is the object referencing the right interface pair? ACL check Inbound ACL on the ingress interface. Is the rule using the real destination IP? Is the protocol/port right? Is the ACL applied with `access-group ... in interface`? NAT translate Forward NAT. Most often this is a missing dynamic NAT (the source has nowhere to translate to) or an Auto vs Manual NAT ordering issue. Route lookup Routing table. Is there a route to the destination? Is it via the expected egress interface? Adjacency ARP / next-hop reachability. Run `show arp` and `ping` the next-hop. Egress checks Any outbound ACL on the egress interface, plus inspection. Most ASA outages are upstream of this. Combined with packet-tracer, this lookup table covers most ASA outage triage. The full data-plane drop counter list is in [Cisco ASA asp-drop Counters Explained](https://www.pinglabz.com/cisco-asa-asp-drop/). ## The Connection Table: Why Step 1 Matters The state table is what makes the ASA stateful. After step 9 of any successful flow, the ASA inserts an entry. Future packets matching that 5-tuple skip steps 2-6 entirely. This has three operational consequences: - **Return traffic does not need an explicit allow ACL.** The reply packet matches the existing connection and is forwarded. - **An ACL change does not break existing flows.** If you remove a permit line that allowed an active connection, that connection keeps working until it is torn down or times out. - **Asymmetric routing breaks the model.** If the reply packet enters through a different interface than the request egressed, the ASA does not see it as part of the existing flow and applies the full check-set, including ACL. This is the most common cause of "but my ACL allows it!" weirdness in dual-firewall, dual-ISP, or active-active failover designs. `show conn` shows what is in the table right now. `show conn count` gives you a quick total. We cover the connection table in detail in [Cisco ASA Connection Table and xlate Table Troubleshooting](https://www.pinglabz.com/cisco-asa-conn-xlate/). ## Related Articles - [Cisco ASA packet-tracer Command: Complete Troubleshooting Guide](https://www.pinglabz.com/cisco-asa-packet-tracer/) \- simulates this exact flow on demand. - [Cisco ASA NAT Explained: Auto NAT vs Manual NAT](https://www.pinglabz.com/cisco-asa-nat-explained/) \- the model behind steps 3 and 5. - [Cisco ASA ACL Configuration: Inbound Rules and Object Groups](https://www.pinglabz.com/cisco-asa-acl-configuration/) \- step 4 in detail. - [Cisco ASA Connection Table and xlate Table Troubleshooting](https://www.pinglabz.com/cisco-asa-conn-xlate/) \- step 1 and step 5 verification. - [ASA vs Router ACLs: What Network Engineers Get Wrong](https://www.pinglabz.com/asa-vs-router-acls/) \- why ASA ACLs use real IPs. - [Cisco ASA Modular Policy Framework](https://www.pinglabz.com/cisco-asa-modular-policy-framework/) \- what step 8 (egress inspection) actually does. ## Key Takeaways The ASA evaluates packets in nine deterministic phases: existing-connection lookup, sanity checks, NAT untranslate, ACL, NAT translate, route lookup, adjacency, egress checks, and forward. Existing connections short-circuit the security checks, which is what makes the firewall stateful. NAT untranslate happens before the ACL, which is why ASA ACLs reference the real destination IP. Knowing which phase drops a packet tells you exactly which configuration block to inspect, and packet-tracer walks every phase on demand against any flow you specify. ### Cisco ASA NAT Explained: Auto NAT vs Manual NAT URL: https://www.pinglabz.com/cisco-asa-nat-explained/ Last updated: 2026-06-13T20:08:32.000Z Network Address Translation on the Cisco ASA is the single most common source of "why does this not work" tickets. The reason is not that NAT is hard, it is that the ASA syntax changed substantially in software 8.3 and a lot of older documentation, blog posts, and training material still teaches the pre-8.3 model. This article covers the modern model: **Auto NAT (object NAT)** and **Manual NAT (twice NAT)**, how they differ, the order they are evaluated, and when to reach for each one. This is part of the [Cisco ASA Complete Guide](https://www.pinglabz.com/cisco-asa/) on PingLabz. After this read, the deeper articles cover [Dynamic PAT](https://www.pinglabz.com/cisco-asa-dynamic-pat/), [Static NAT](https://www.pinglabz.com/cisco-asa-static-nat-dmz/), [Twice NAT](https://www.pinglabz.com/cisco-asa-twice-nat/), [Identity NAT for VPN](https://www.pinglabz.com/cisco-asa-identity-nat-vpn/), and the must-know [NAT Order of Operations](https://www.pinglabz.com/cisco-asa-nat-order-of-operations/). ## The Two NAT Types: Auto and Manual Modern ASA software has exactly two NAT configuration models. Every NAT rule on a current ASA is one or the other. Object NAT, network-object NAT TypeAuto NAT Configured under Inside an `object network` definition What it can match One source object only (the network-object itself) What it can do Static, dynamic, PAT, dynamic-with-fallback Twice NAT TypeManual NAT Configured under Global config under `nat (real-int,mapped-int) ...` What it can match Source AND destination, optional service ports What it can do Same as auto, plus identity NAT, NAT exemption, conditional translation The simple mental model: use Auto NAT for unconditional translations (one source, always translate this way) and Manual NAT when the translation depends on what the destination is. ## Auto NAT Syntax Auto NAT is configured inside the network-object that represents the source. The translation is attached to the object, not declared globally. Example - dynamic PAT for the inside subnet going to outside: ``` object network INSIDE-NET subnet 10.10.10.0 255.255.255.0 nat (inside,outside) dynamic interface ``` That single object does three things: defines a source network, declares the NAT rule, and references the egress interface PAT pool (`interface` means "use the outside interface IP"). Static NAT has the same shape, with `static` instead of `dynamic` and a target address instead of `interface`: ``` object network DMZ-WEB host 192.168.50.10 nat (dmz,outside) static 198.51.100.10 ``` This publishes the DMZ web server at 192.168.50.10 as the public IP 198.51.100.10\. Anyone on the outside hitting 198.51.100.10 gets translated to 192.168.50.10 by the ASA. That is essentially the entire Auto NAT model. One object, one NAT rule, one direction of translation. ## Manual NAT Syntax Manual NAT (twice NAT) is declared at global config level outside any object, and it can match both source AND destination. Example - exempt traffic going from the inside subnet to a remote VPN subnet from being PATted (NAT exemption / identity NAT): ``` object network INSIDE-NET subnet 10.10.10.0 255.255.255.0 object network REMOTE-VPN-NET subnet 10.20.0.0 255.255.255.0 nat (inside,outside) source static INSIDE-NET INSIDE-NET destination static REMOTE-VPN-NET REMOTE-VPN-NET no-proxy-arp route-lookup ``` Read the `nat` line as: "On a packet ingressing inside and egressing outside, if the source matches INSIDE-NET and the destination matches REMOTE-VPN-NET, translate the source from INSIDE-NET to itself (no change) and the destination from REMOTE-VPN-NET to itself." Both translations are identity (no actual translation), the effect is to bypass any other NAT rule that would otherwise PAT the source. Manual NAT is also the syntax for any conditional translation - "translate the source to X only when the destination is Y." That kind of rule is impossible in Auto NAT and is the reason Manual NAT exists. We cover it in [Twice NAT Explained with Real Examples](https://www.pinglabz.com/cisco-asa-twice-nat/). ## The Evaluation Order: Three Sections, First Match Wins This is the part that surprises every engineer who learned pre-8.3 NAT. Modern ASA evaluates NAT rules in three sections in this fixed order: 1\. Manual NAT Contents All `nat (real,mapped) source ...` commands declared without the `after-auto` keyword Order within section Configuration order. First match wins. 2\. Auto NAT Contents All NAT rules declared inside `object network ...` blocks Order within section By specificity (most specific first), then by configuration order within a tier 3\. Manual NAT (after-auto) Contents All `nat (real,mapped) source ... after-auto` commands Order within section Configuration order. First match wins. The first NAT rule that matches the packet wins. The remaining rules are not evaluated. Auto NAT (Section 2) auto-orders by specificity so a host object always matches before a subnet object that contains it - no manual ordering needed. Manual NAT (Sections 1 and 3) is evaluated in configuration order, so you must place the most specific rule first. You can see the compiled NAT rule order with `show nat detail`: ``` ASA-PERIM# show nat detail Auto NAT Policies (Section 2) 1 (dmz) to (outside) source static DMZ-WEB 198.51.100.10 translate_hits = 87, untranslate_hits = 134 Source - Origin: 192.168.50.10/32, Translated: 198.51.100.10/32 2 (inside) to (outside) source dynamic INSIDE-TRANSIT interface translate_hits = 0, untranslate_hits = 0 Source - Origin: 10.10.0.0/24, Translated: 203.0.113.2/30 3 (inside) to (outside) source dynamic INSIDE-NET interface translate_hits = 4521, untranslate_hits = 0 Source - Origin: 10.10.10.0/24, Translated: 203.0.113.2/30 ``` This output is from our ASAv 9.23 lab. Section 2 (Auto NAT) evaluates in specificity order: the DMZ-WEB host (most specific, /32) is rule 1, the INSIDE-TRANSIT subnet (/24) is rule 2, the INSIDE-NET subnet (also /24) is rule 3\. The /32 host always matches before any /24 subnet that contains it, even though we configured DMZ-WEB last. That is the auto-ordering Auto NAT does for you. The hit counters tell you which rules are actually firing - in this snapshot, INSIDE-NET has 4521 forward translations to the outside interface IP (PAT), DMZ-WEB has 87 forward and 134 reverse hits (a public-facing web server gets more inbound than outbound traffic). If you also have Manual NAT rules, `show nat detail` would print them under **Manual NAT Policies (Section 1)** first, before the Auto NAT block, and any `after-auto` manual rules under **Manual NAT Policies (Section 3)** at the end. The order of the sections in the output mirrors the evaluation order. ## Auto NAT vs Manual NAT: How to Choose PAT inside hosts going to internet UseAuto NAT Why Unconditional source-only translation Publish a DMZ web server publicly UseAuto NAT (static) Why Unconditional, simple source-to-translated-source mapping VPN traffic must skip NAT UseManual NAT, Section 1 Why Translation depends on destination (the VPN subnet) Different translation depending on which destination UseManual NAT Why Auto NAT cannot match destination Twice NAT (translate both source and destination differently) UseManual NAT Why Only Manual NAT supports this "Translate everything that does not match a more specific rule" Use Manual NAT, Section 3 (after-auto) Why Catches packets after Auto NAT, useful as a default-PAT fallback in complex deployments The 90% answer is Auto NAT. Reach for Manual NAT when you need to match destination, or when you need a rule to win *before* Auto NAT (Section 1) or *after* Auto NAT (Section 3). Get the Cisco ASA Field Reference - 9 pages, free Everything you'd want to remember about Cisco ASA on nine printable pages. Per-packet pipeline diagram, NAT 8.3+ section ordering, six-branch troubleshooting decision tree, real lab show-output annotated, paste-ready three-zone config. Free for PingLabz members - just sign up with your email. [Get the Cisco ASA cheat-sheet](https://www.pinglabz.com/cisco-asa-cheatsheet/) ## Real IP vs Mapped IP: The ACL Catch One critical detail that touches every ACL on the ASA: ACLs reference **real IPs**, not **mapped IPs**. If your DMZ web server is published as 198.51.100.10 (mapped) but lives at 192.168.50.10 (real), the ACL on the outside interface should permit traffic to 192.168.50.10, not 198.51.100.10\. The ASA untranslates the destination first (UN-NAT phase) and then evaluates the ACL against the real IP. This catches engineers who came from older ASA software (pre-8.3) where ACLs used the mapped IP. Modern ASA documentation calls this "the real IP rule" and it applies everywhere. Confirm it for any flow with packet-tracer - the UN-NAT phase prints both the mapped and real IPs, then the ACCESS-LIST phase shows what was actually compared. ## Verifying NAT in Real Time: show xlate `show xlate` shows the live NAT translation table - every active translation right now, including the source mapping, destination mapping, and timeout. ``` ASA-PERIM# show xlate 1 in use, 1 most used Flags: D - DNS, e - extended, I - identity, i - dynamic, r - portmap, s - static, T - twice, N - net-to-net NAT from dmz:192.168.50.10 to outside:198.51.100.10 flags s idle 0:00:30 timeout 0:00:00 ``` This is real output from our ASAv 9.23 lab right after configuration with no live traffic. One static NAT entry: the DMZ web server (flag `s` \= static, no timeout because static NAT is always installed regardless of traffic). When inside hosts start passing traffic, dynamic PAT entries appear with flags `ri` (r = portmap, i = dynamic) and a 30-second idle timeout. For incident triage, `show xlate` tells you whether the ASA actually translated the flow you expected. ## Common NAT Mistakes - **Writing the ACL against the mapped (public) IP.** Use the real IP. Always. ASA untranslates first. - **Putting NAT exemption in the wrong section.** NAT exemption for VPN must be in Section 1 (manual NAT, before auto), otherwise the auto PAT rule wins first. - **Forgetting `no-proxy-arp` on identity NAT.** Without it, the ASA proxy-ARPs for the destination, which can cause traffic loops or IP conflicts. - **Mixing pre-8.3 syntax with modern syntax.** If you see `global (outside) 1 interface` and `nat (inside) 1 ...` in the config, you are looking at pre-8.3\. Modern config never uses NAT IDs. - **Assuming order does not matter.** In Manual NAT, configuration order is the evaluation order. The most specific rule must be configured first. - **Not checking hit counters.** `show nat detail` includes per-rule hit counters. Zero hits on a rule that should be firing is a strong signal something is misordered. ## Related Articles - [Cisco ASA Dynamic PAT Configuration for Internet Access](https://www.pinglabz.com/cisco-asa-dynamic-pat/) \- the most common Auto NAT pattern. - [Cisco ASA Static NAT for Publishing a Server in the DMZ](https://www.pinglabz.com/cisco-asa-static-nat-dmz/) \- the second-most-common Auto NAT pattern. - [Cisco ASA Twice NAT Explained with Real Examples](https://www.pinglabz.com/cisco-asa-twice-nat/) \- Manual NAT walkthrough. - [Cisco ASA Identity NAT / NAT Exemption for VPNs](https://www.pinglabz.com/cisco-asa-identity-nat-vpn/) \- the Section 1 manual NAT pattern that keeps VPN tunnels working. - [Cisco ASA NAT Order of Operations Cheat Sheet](https://www.pinglabz.com/cisco-asa-nat-order-of-operations/) \- the three-section model in one printable page. - [Cisco ASA packet-tracer Command: Complete Troubleshooting Guide](https://www.pinglabz.com/cisco-asa-packet-tracer/) \- the troubleshooting tool that walks every NAT decision. ## Key Takeaways Modern ASA NAT is two types (Auto and Manual) evaluated in three sections (Manual-before-auto, Auto, Manual-after-auto), first match wins. Auto NAT covers the 90% case: unconditional source-only translation tied to a network object. Manual NAT exists for the 10% case where translation depends on destination, where you need NAT exemption for VPN, or where the auto-ordering does not give you the precedence you want. ACLs always reference real IPs - the ASA untranslates before evaluating the ACL. Use `show nat detail` for compiled rule order and `show xlate` for live translations, and use packet-tracer to confirm a specific flow takes the rule you expected. ### Cisco ASA packet-tracer Command: Complete Troubleshooting Guide URL: https://www.pinglabz.com/cisco-asa-packet-tracer/ Last updated: 2026-06-13T20:08:32.000Z If you only learn one Cisco ASA troubleshooting command, make it `packet-tracer`. It simulates a single hypothetical packet through every ASA decision point - existing-connection lookup, security check, NAT untranslate, ACL, NAT translate, route lookup, egress checks - and tells you exactly where it would be allowed, dropped, or routed somewhere unexpected. No actual traffic is generated, no real session is opened, and you do not need to be on the inside or outside of the firewall to test. It is the closest thing the ASA has to a debugger. This guide is part of the [Cisco ASA Complete Guide](https://www.pinglabz.com/cisco-asa/) on PingLabz. We cover the syntax, every common scenario (inside-to-outside PAT, outside-to-DMZ static NAT, ACL deny, NAT mismatch, VPN decryption), how to read the phase-by-phase output, the most useful flags, and the mistakes that make engineers misread results. All output is captured from a real ASAv 9.23 running in our lab. ## Syntax Basics The minimum form of `packet-tracer` takes five arguments: ingress interface, protocol, source IP, source port, destination IP, destination port. The ASA simulates one packet matching that 6-tuple, walks it through the data path, and prints the result. ``` packet-tracer input ``` For TCP and UDP you give source and destination port. For ICMP you give type, code, and identifier instead. A typical inside-to-outside HTTPS test looks like this: ``` ASA-PERIM# packet-tracer input inside tcp 10.10.10.50 12345 8.8.8.8 443 ``` And a typical outside-to-DMZ HTTP test (someone hitting the public IP of a DMZ web server) looks like this: ``` ASA-PERIM# packet-tracer input outside tcp 203.0.113.50 33000 198.51.100.10 80 ``` Source IP and source port are usually arbitrary. The destination IP must be the real IP of the destination as it would appear on the wire from the ingress interface side - which is where engineers first get confused. We will come back to this. ## Reading the Output: Phases and Final Verdict Every packet-tracer run produces a series of **phases**, each phase representing one stage of the ASA data path. Each phase reports its **type** (ROUTE-LOOKUP, ACCESS-LIST, NAT, INSPECT, etc.), its **action** (allow or drop), and any matching configuration. Here is what an inside-to-outside PAT lookup looks like on our lab ASA: ``` ASA-PERIM# packet-tracer input inside tcp 10.10.10.50 12345 8.8.8.8 443 Phase: 1 Type: ACCESS-LIST Subtype: Result: ALLOW Elapsed time: 65404 ns Config: Implicit Rule Additional Information: MAC Access list Phase: 2 Type: INPUT-ROUTE-LOOKUP Subtype: Resolve Egress Interface Result: ALLOW Elapsed time: 21670 ns Config: Additional Information: Found next-hop 203.0.113.1 using egress ifc outside Phase: 3 Type: NAT Subtype: Result: ALLOW Elapsed time: 29681 ns Config: object network INSIDE-NET nat (inside,outside) dynamic interface Additional Information: Dynamic translate 10.10.10.50/12345 to 203.0.113.2/12345 Phase: 4 Type: NAT Subtype: per-session Result: ALLOW Phase: 5 Type: IP-OPTIONS Subtype: Result: ALLOW Phase: 6-7 Type: QOS Result: ALLOW Phase: 10 Type: FLOW-CREATION Subtype: Result: ALLOW New flow created with id 4, packet dispatched to next module Phase: 11 Type: INPUT-ROUTE-LOOKUP-FROM-OUTPUT-ROUTE-LOOKUP Subtype: Resolve Preferred Egress interface Result: ALLOW Found next-hop 203.0.113.1 using egress ifc outside Result: input-interface: inside input-status: up input-line-status: up output-interface: outside output-status: up output-line-status: up Action: drop Time Taken: 282497 ns Drop-reason: (no-v4-adjacency) No valid V4 adjacency. Check ARP table (show arp) has entry for nexthop. ``` Read it bottom-up: **Action: drop**, **Drop-reason: no-v4-adjacency**. The ASA simulated the flow successfully through every security and translation check, but the ARP cache does not yet have an entry for the next-hop 203.0.113.1, so it cannot synthesize the egress frame. This is a packet-tracer quirk you will see often on ASA software 9.x: every phase reports ALLOW, but the final action drops with no-v4-adjacency because no real packet has yet primed the ARP table. In production, real traffic would trigger ARP resolution and the same flow would forward. The lesson: **read all the phase results, not just the final action**. Every ALLOW phase confirms a configuration block is correct. Read top-down to see how the ASA processed it: implicit ACL allowed (inside is high-security going to lower-security), route lookup found the egress, NAT translated the source from 10.10.10.50 to the outside interface IP (PAT, port 12345), QoS and IP-options checks passed, a flow entry was created, and the egress was confirmed. Every block of configuration that should have applied did apply. When the ARP table primes (real traffic, or after a quick `ping 203.0.113.1`), the same packet-tracer would return Action: allow. When something goes wrong at a configuration layer, the matching phase has `Result: DROP` and the final `Action:` is `drop`, with a drop-reason that names the configuration block. ## Scenario: ACL Drop A user is reporting they cannot SSH to a public-facing DMZ host. We test from outside: ``` ASA-PERIM# packet-tracer input outside tcp 203.0.113.50 44000 198.51.100.10 22 Phase: 1 Type: UN-NAT Subtype: static Result: ALLOW Elapsed time: 20882 ns Config: object network DMZ-WEB nat (dmz,outside) static 198.51.100.10 Additional Information: NAT divert to egress interface dmz Untranslate 198.51.100.10/22 to 192.168.50.10/22 Phase: 2 Type: ACCESS-LIST Subtype: Result: DROP Elapsed time: 6304 ns Config: Implicit Rule Additional Information: Result: input-interface: outside input-status: up input-line-status: up output-interface: dmz output-status: up output-line-status: up Action: drop Time Taken: 27186 ns Drop-reason: (acl-drop) Flow is denied by configured rule, Drop-location: frame snp_classify_table_lookup:6044 flow (NA)/NA ``` Two things to notice. First, Phase 1 (UN-NAT) shows the ASA correctly untranslating the public IP 198.51.100.10 to the real IP 192.168.50.10 - the static NAT is working. Second, Phase 2 (ACCESS-LIST) hits the implicit deny because `OUTSIDE_IN` only permits ports 80, 443, and ICMP to DMZ-WEB; port 22 falls through to the implicit deny at the bottom of the ACL. The fix is to add a permit line for SSH or scope the existing ones differently. The **Drop-reason: acl-drop** at the bottom is the verdict, captured directly from our ASAv 9.23 lab. This is the canonical packet-tracer flow that proves an outage is an ACL problem and not a NAT, route, or inspection problem. ## Scenario: NAT Mismatch An engineer added a new VPN site-to-site tunnel and the tunnel comes up, but interesting traffic does not pass. The classic culprit is that interesting traffic is being NATed before being placed in the tunnel. ``` ASA-PERIM# packet-tracer input inside tcp 10.10.10.50 12345 10.20.0.50 443 Phase: 1 Type: ACCESS-LIST Subtype: Result: ALLOW Config: Implicit Rule Phase: 2 Type: ROUTE-LOOKUP Subtype: Resolve Egress Interface Result: ALLOW Config: Additional Information: found next-hop 203.0.113.1 using egress ifc outside Phase: 3 Type: NAT Subtype: Result: ALLOW Config: object network INSIDE-NET nat (inside,outside) dynamic interface Additional Information: Dynamic translate 10.10.10.50/12345 to 203.0.113.2/12345 [...VPN encryption phases...] Result: Action: allow ``` This is wrong. The packet is destined for the remote VPN subnet 10.20.0.0/24, which should be carried untranslated through the IPsec tunnel. Instead, Phase 3 PATted it to the outside interface IP. The fix is a NAT exemption rule (also called identity NAT) placed in Section 1 (manual NAT, before auto NAT), so VPN traffic is matched first and not translated. We cover that pattern in detail in [Cisco ASA VPN NAT Exemption: The Mistake That Breaks Tunnels](https://www.pinglabz.com/cisco-asa-vpn-nat-exemption/). ## Useful Flags `packet-tracer` supports several flags that change what it tests. The ones worth knowing: `detailed` Prints additional debug-level information, including object-group expansion and ACL hit details. `xml` Renders the result as XML. Useful when scripting (paste into a Python parser). `vlan-id ` Tags the simulated packet with a VLAN ID, useful when the ingress is a trunk subinterface. `icmp` instead of tcp/udp For ICMP, give type/code/identifier: `packet-tracer input outside icmp 1.1.1.1 8 0 198.51.100.10` `decrypted` Tests a packet as if it had already been decrypted by a VPN engine. Lets you simulate the inner-payload behavior. `persist` The simulated flow is kept in the connection table after the trace, so subsequent `show conn` can be used. Get the Cisco ASA Field Reference - 9 pages, free Everything you'd want to remember about Cisco ASA on nine printable pages. Per-packet pipeline diagram, NAT 8.3+ section ordering, six-branch troubleshooting decision tree, real lab show-output annotated, paste-ready three-zone config. Free for PingLabz members - just sign up with your email. [Get the Cisco ASA cheat-sheet](https://www.pinglabz.com/cisco-asa-cheatsheet/) ## Troubleshooting a Real Outage with packet-tracer The fastest way to use packet-tracer during an incident is to test the exact source-to-destination 6-tuple the user is reporting, in both directions, from both perspectives. A four-step pattern that resolves most outages: 1. **Test the user's flow as reported.** Type the same source IP, source port (any high port is fine), destination IP, and destination port the user gave you. Look at the final Action. 2. **If it allows but the flow still fails**, the ASA is not the problem. Check the destination host, downstream firewall, or DNS. 3. **If it drops**, look at the **Drop-reason**. If it is acl-drop, the ACL is the cause. If it is nat-no-xlate, NAT is the cause. If it is rpf-violated, asymmetric routing or anti-spoofing is biting you. The phase that matches the drop tells you which configuration block to fix. 4. **If you need to test the reverse direction** (server-to-client), simulate it with packet-tracer ingress from the destination interface. ASA does not assume reciprocity; you must explicitly test both directions when troubleshooting return-traffic issues. Combine this with [packet capture](https://www.pinglabz.com/cisco-asa-packet-capture/) to see what is actually arriving, and you have covered the gap between simulation and reality. ## Common packet-tracer Drop Reasons The drop-reason field is a controlled vocabulary. The ones that come up over and over in production: `acl-drop` The interface ingress ACL denies the flow. `nat-no-xlate` / `nat-no-xlate-to-pat-pool` No NAT rule matches and the egress requires translation. Often a missing dynamic NAT or a mistakenly removed object NAT. `nat-rpf-failed` NAT reverse-path forwarding check failed. Usually means the route the ASA picked does not match what NAT expected. `no-route` / `no-adjacency` The ASA has no route or no ARP entry for the next-hop on the egress interface. `sp-security-failed` Anti-spoofing (uRPF) rejected the source. `tcp-not-syn` The first packet on a flow is not a SYN. Often a return packet from an asymmetric path. `inspect-icmp-error-different-embedded-conn` An ICMP error came in for a flow the ASA does not have in its connection table. Frequently a sign of an asymmetric path. Every one of these has a well-defined fix, and packet-tracer pinpoints which one applies before you start typing in the dark. Full counter explanations live in [Cisco ASA asp-drop Counters Explained](https://www.pinglabz.com/cisco-asa-asp-drop/). ## Common Mistakes - **Testing the wrong destination IP.** When testing from outside to a NATed server, use the public IP. When testing inside to inside (or DMZ to inside), use the real IP. The ASA's untranslate phase will handle the rest. - **Forgetting that ACLs reference real IPs.** ACL phase results show the real (untranslated) destination. If you wrote the ACL against the public IP, packet-tracer will show ACL-drop even when the public-to-private NAT works. - **Trusting a single direction.** An allow result for client-to-server does not guarantee server-to-client works. Test both. - **Ignoring the implicit rule.** Phase 1 often says "Implicit Rule" - that is the security-level default. Two interfaces at the same security level produce an implicit drop unless you have configured `same-security-traffic permit`. - **Testing without enough info from the user.** Without the actual source IP and destination port, packet-tracer is guesswork. Get the exact values from the user before you start. ## Related Commands Packet-tracer is most powerful in combination with these: - `show conn detail` \- real connection table, what flows are actually open right now. - `show xlate` \- the live NAT translation table. - `show asp drop` \- data-plane drop counters since the last clear. - `show access-list OUTSIDE_IN` \- per-line hit counters on a specific ACL. - `capture CAP-NAME interface outside match tcp host 203.0.113.50 host 198.51.100.10 eq 80` \- real packet capture filtered by ACL match. Walk through every one of these in the [Cisco ASA Cheat Sheet](https://www.pinglabz.com/cisco-asa-cheat-sheet/) for the working command list, and the [Connection Table and xlate Table Troubleshooting](https://www.pinglabz.com/cisco-asa-conn-xlate/) article for deeper coverage. ## Key Takeaways Packet-tracer simulates one synthetic packet through every ASA decision point and prints the verdict, no real traffic required. Read the output bottom-up for the action, top-down for the reasoning. The phase that drops the packet, plus its drop-reason, is the single most useful piece of troubleshooting data the ASA produces. Use it as the first command in any ASA outage investigation, and combine it with `show conn`, `show xlate`, and a real packet capture to bridge from simulation to reality. ### OSPF Field Reference (9-Page Printable Cheat-Sheet) URL: https://www.pinglabz.com/ospf-cheatsheet/ Last updated: 2026-08-02T03:06:13.000Z _This post is for subscribers only._ ### BGP Field Reference (9-Page Printable Cheat-Sheet) URL: https://www.pinglabz.com/bgp-cheatsheet/ Last updated: 2026-08-02T03:06:13.000Z _This post is for subscribers only._ ### GRE on Linux: ip tunnel add Commands and Examples URL: https://www.pinglabz.com/gre-on-linux/ Last updated: 2026-06-13T20:08:32.000Z The Linux kernel has supported GRE tunnels natively since the 2.0 series, and the modern `iproute2` userland makes setting one up genuinely simple. If you operate routers from a Cisco perspective and find yourself bridging to a Linux box (a cloud VM, a virtual appliance, an open-source firewall, a mid-range MikroTik analogue running ROS, or your own laptop in a lab), this is the article that translates the configuration model from one to the other. Same protocol, different syntax. Part of the [PingLabz GRE Tunnels: The Complete Guide](https://www.pinglabz.com/gre/) cluster. ## When You Need Linux GRE Three common cases: - **Cloud VMs as tunnel endpoints.** A Linux EC2 instance, an Azure VM, or a GCP Compute Engine VM acting as a GRE termination for spoke traffic. Common for cloud-on-ramps that pre-date managed VPN gateways. - **Open-source routers and firewalls.** Linux-based platforms (VyOS, OPNsense, pfSense, FRRouting on Debian) all expose GRE through the same kernel interface. The CLI varies; the underlying mechanism is identical. - **Building lab tunnels without Cisco gear.** When you want to demonstrate GRE behavior, build a packet capture, or test a routing-protocol-over-GRE design without spinning up CSR1000v instances, two Linux VMs with `ip tunnel` commands work fine. The Linux GRE implementation interoperates cleanly with Cisco for the core functionality (encapsulation, decapsulation, basic tunneling). It does not natively implement Cisco GRE keepalives, so liveness detection across a Cisco-Linux GRE pair has to come from the routing protocol or a separate mechanism. ## Basic GRE Tunnel on Linux Equivalent to the Cisco config from the [config lab](https://www.pinglabz.com/gre-tunnel-configuration-cisco/). Two endpoints, one tunnel between them. ``` # ---- linux-1 (198.51.100.1) ---- sudo ip tunnel add tun0 mode gre \ local 198.51.100.1 \ remote 203.0.113.1 \ ttl 255 sudo ip addr add 10.0.0.1/30 dev tun0 sudo ip link set tun0 up # ---- linux-2 (203.0.113.1) ---- sudo ip tunnel add tun0 mode gre \ local 203.0.113.1 \ remote 198.51.100.1 \ ttl 255 sudo ip addr add 10.0.0.2/30 dev tun0 sudo ip link set tun0 up ``` `ip tunnel add` creates the GRE interface. The `local` and `remote` are the underlay endpoints (the equivalent of Cisco's `tunnel source` and `tunnel destination`). The `ip addr add` assigns the overlay IP. The `ip link set tun0 up` brings the interface administratively up. The Cisco-style tunnel naming is just convention; you can call the interface anything (`gre0`, `tun0`, `mytun`). Some distributions reserve `gre0` for a fallback interface; explicit naming is safer. ## Verifying the Tunnel ``` linux-1$ ip -d link show tun0 3: tun0@NONE: mtu 1476 qdisc noqueue state UNKNOWN link/gre 198.51.100.1 peer 203.0.113.1 promiscuity 0 gre remote 203.0.113.1 local 198.51.100.1 ttl 255 \ pmtudisc linux-1$ ip addr show tun0 3: tun0@NONE: mtu 1476 qdisc noqueue state UNKNOWN group default qlen 1000 link/gre 198.51.100.1 peer 203.0.113.1 inet 10.0.0.1/30 scope global tun0 valid_lft forever preferred_lft forever linux-1$ ping -c 5 10.0.0.2 PING 10.0.0.2 (10.0.0.2) 56(84) bytes of data. 64 bytes from 10.0.0.2: icmp_seq=1 ttl=64 time=4.21 ms ``` The MTU 1476 is Linux doing the same math Cisco does: 1500 minus 24 bytes of GRE+IP overhead. The `POINTOPOINT,NOARP` flags reflect that GRE is a point-to-point Layer 3 link with no ARP needed (the next-hop is implicitly the other end of the tunnel). ## Adding Routes Through the Tunnel Linux GRE tunnels do not auto-add routes. Add explicit static routes for any subnet that should reach across the tunnel: ``` sudo ip route add 192.168.2.0/24 via 10.0.0.2 dev tun0 ``` Or, for dynamic routing, use FRRouting (the modern Linux routing daemon, used by VyOS and others). FRR's OSPF or BGP can run on the tunnel interface just like Cisco's: ``` vtysh# conf t vtysh(config)# router ospf vtysh(config-router)# network 10.0.0.0/30 area 0 vtysh(config-router)# network 192.168.1.0/24 area 0 ``` ## MTU and MSS on Linux The same MTU concerns from [the GRE MTU article](https://www.pinglabz.com/gre-tunnel-mtu/) apply. To set the inner MTU and MSS: ``` sudo ip link set tun0 mtu 1400 # MSS clamping in iptables (older systems): sudo iptables -t mangle -A FORWARD -p tcp --tcp-flags SYN,RST SYN -o tun0 \ -j TCPMSS --clamp-mss-to-pmtu # MSS clamping in nftables (modern systems): sudo nft add rule inet mangle forward oifname tun0 \ tcp flags syn tcp option maxseg size set rt mtu ``` `--clamp-mss-to-pmtu` tells the kernel to compute the MSS dynamically from the path MTU. That is more flexible than hard-coding 1360 the way Cisco's `ip tcp adjust-mss` typically does, and it tracks if you change the tunnel MTU later. The functional outcome is the same: TCP endpoints negotiate segment sizes that fit through the tunnel without fragmentation. ## Liveness on Linux GRE Linux's native GRE has no Cisco-style keepalives. Three options for liveness detection: - **Routing-protocol dead-timers.** If you are running OSPF / BGP across the tunnel, the routing protocol's hello and dead timers serve as liveness detection. Tune them aggressively (OSPF hello 1, dead 4) for fast detection. - **BFD.** Bidirectional Forwarding Detection in FRRouting works on tunnel interfaces. Sub-second detection, lower overhead than tightly-tuned routing-protocol timers, and standardized so it interoperates with Cisco BFD. - **Userland keepalive scripts.** A simple cron-driven script that pings the remote tunnel IP and triggers a recovery action (route update, alert) on failure. Crude but effective for small deployments. For a Cisco-to-Linux tunnel, the cleanest answer is to disable Cisco GRE keepalives on the Cisco side and let the routing protocol or BFD provide liveness. Mixing Cisco keepalives with a Linux end that does not understand them just produces confusing one-way "down" events on the Cisco router. ## GRE Variants on Linux Linux supports several GRE encapsulation modes. The `mode` parameter to `ip tunnel add` determines which. `gre` What it carries IPv4 over IPv4 (the default) Cisco equivalent`tunnel mode gre ip` `ip6gre` What it carriesIPv4 / IPv6 over IPv6 Cisco equivalent`tunnel mode gre ipv6` `gretap` What it carries Ethernet (L2) over IP - L2 tunneling Cisco equivalent `tunnel mode ethernet gre` `ip6gretap` What it carriesEthernet over IPv6 Cisco equivalent(IOS XE limited) `gretap` is the interesting variant for engineers used to Cisco. It carries Ethernet frames over GRE, providing a Layer 2 stretch between two Linux endpoints. This is the same idea as Cisco's L2TPv3 or EoMPLSoGRE in service-provider designs but without the SP overhead. Useful for cloud overlay networking and labs. ## Optional GRE Header Fields Linux supports the GRE Key field (RFC 2890) for distinguishing multiple parallel tunnels with the same source/destination, and the Sequence Number field for in-order delivery (rarely used). ``` sudo ip tunnel add tun0 mode gre \ local 198.51.100.1 \ remote 203.0.113.1 \ key 12345 \ ttl 255 ``` The Key value must match on both ends. Linux accepts a numeric (4-byte) Key or a dotted-quad notation. For DMVPN-style multipoint setups Linux has limited native support; `ip tunnel` can build point-to-point tunnels but does not directly implement NHRP. Quagga / FRRouting plus userland helpers can build something DMVPN-shaped, but it is more involved than the Cisco equivalent. ## Adding IPsec on Linux Linux IPsec (the kernel XFRM framework with `strongSwan` or `libreswan` in userland) wraps GRE the same way Cisco IPsec does. The pattern: build the GRE tunnel, configure IPsec to protect IP protocol 47 between the two underlay IPs. ``` # /etc/ipsec.conf with strongSwan conn gre left=198.51.100.1 right=203.0.113.1 type=transport leftprotoport=gre rightprotoport=gre auto=start authby=psk ike=aes256-sha256-modp2048 esp=aes256-sha256 ``` This is the moral equivalent of the Cisco crypto-map approach: classify GRE traffic between specific endpoints, encrypt it. The transport mode reuses the existing GRE outer IP rather than adding another. For a deeper IPsec-on-Linux walkthrough, the strongSwan documentation is the canonical reference; for the Cisco perspective, see [GRE over IPsec on IOS XE](https://www.pinglabz.com/gre-over-ipsec/). ## Cisco to Linux GRE The minimal interop config: ``` ! Cisco R1 interface Tunnel0 ip address 10.0.0.1 255.255.255.252 ip mtu 1400 ip tcp adjust-mss 1360 tunnel source 198.51.100.1 tunnel destination 203.0.113.1 tunnel mode gre ip ! Note: no 'keepalive' - Linux peer does not understand it ``` ``` # Linux peer sudo ip tunnel add tun0 mode gre local 203.0.113.1 remote 198.51.100.1 ttl 255 sudo ip link set tun0 mtu 1400 sudo ip addr add 10.0.0.2/30 dev tun0 sudo ip link set tun0 up sudo iptables -t mangle -A FORWARD -p tcp --tcp-flags SYN,RST SYN -o tun0 \ -j TCPMSS --clamp-mss-to-pmtu ``` Both ends should report up after the link state propagates. `ping 10.0.0.2` from R1 and `ping 10.0.0.1` from the Linux box should both succeed. This is the cleanest interop case; complications arise when you want OSPF (FRR works fine), GRE keepalives (skip them), or DMVPN (not natively on Linux). ## Making It Persistent The `ip tunnel add` commands above are runtime-only. They disappear on reboot. To persist, use whatever the distro's network management framework expects: Debian / Ubuntu /etc/network/interfaces (legacy) `auto tun0` \+ `iface tun0 inet static` stanza with `up ip tunnel add ...` NetworkManager (modern) `nmcli connection add type ip-tunnel ...` systemd-networkd `.netdev` file with `[Tunnel]` section, `.network` file for IP address VyOS / OPNsense / pfSense Vendor-specific CLI / GUI For systemd-networkd (modern Debian / Ubuntu / RHEL), the cleanest approach: ``` cat > /etc/systemd/network/10-gre.netdev << 'EOF' [NetDev] Name=tun0 Kind=gre [Tunnel] Local=198.51.100.1 Remote=203.0.113.1 TTL=255 EOF cat > /etc/systemd/network/10-gre.network << 'EOF' [Match] Name=tun0 [Network] Address=10.0.0.1/30 EOF sudo systemctl restart systemd-networkd ``` ## Summary Linux GRE is the same protocol Cisco IOS XE implements, configured through `ip tunnel` commands or distro-specific persistence files. The wire format is identical, the encapsulation overhead is identical, and Linux interoperates cleanly with Cisco at the basic-tunneling level. The differences are in the surrounding ecosystem: Linux does not natively implement Cisco GRE keepalives, native DMVPN/NHRP support is limited, and the routing-protocol layer is FRRouting rather than IOS XE. For a Cisco-to-Linux tunnel, build the basic GRE on both ends, set MTU 1400 and MSS clamping on both ends, run a routing protocol over the top, and skip Cisco-specific keepalives. The resulting tunnel performs the same as Cisco-to-Cisco. Full PingLabz GRE coverage is at the [cluster pillar](https://www.pinglabz.com/gre/). ### GRE Tunnel Troubleshooting: Recursive Routing and Five More Failures URL: https://www.pinglabz.com/gre-tunnel-troubleshooting/ Last updated: 2026-08-01T19:26:56.000Z This is the field guide for diagnosing GRE tunnel failures on Cisco IOS XE. Every failure mode is paired with its symptoms, the show commands that confirm it, and the fix, roughly most common first. If you are 30 minutes into an incident, work down the list and stop when something matches. The recursive-routing section is worth reading even if nothing is broken, because it is the failure engineers most often misdiagnose. It is built around a real capture: a two-router CML lab on **IOS XE 17.18.2** where the tunnel is made to eat its own transport, with the log sequence, the routing table before and after, and three fixes. This is part of the PingLabz [GRE Tunnels: The Complete Guide](https://www.pinglabz.com/gre/). For protocol theory and configuration, start there. ## First Checks: The Five-Minute Triage Before deep debugging, run these five commands. They tell you which sub-system is broken. ``` R1# show interface Tunnel0 | include Tunnel|line proto|MTU|Keepalive|source|dest Tunnel0 is up, line protocol is up MTU 17916 bytes, ... Keepalive set (10 sec), retries 3 Tunnel source 198.51.100.1, destination 203.0.113.1 Tunnel transport MTU 1476 bytes R1# show ip route 203.0.113.1 Routing entry for 203.0.113.1/32 Known via "static", distance 1, metric 0 Routing Descriptor Blocks: * 198.51.100.2 Route metric is 0, traffic share count is 1 R1# ping 203.0.113.1 source 198.51.100.1 !!!!! Success rate is 100 percent (5/5) R1# ping 10.0.0.2 source 10.0.0.1 !!!!! Success rate is 100 percent (5/5) R1# ping 192.168.2.1 source 192.168.1.1 size 1400 df-bit !!!!! Success rate is 100 percent (5/5) ``` If all five pass, the tunnel works at the IP level. Failures point at the broken layer: Tunnel0 line protocol down Local config or underlay route Underlay ping fails Underlay connectivity Overlay ping fails Encapsulation, or a firewall blocking IP 47 Overlay 1400 df-bit fails MTU problem on the path LAN-to-LAN ping fails Routing, static or dynamic ## Recursive Routing **Symptom:** Tunnel flaps every 30 seconds to a few minutes. Log shows `%TUN-5-RECURDOWN: Tunnel0 temporarily disabled due to recursive routing`. **Cause:** The route to the tunnel destination IP is itself learned through the tunnel. The router cannot encapsulate packets to the tunnel destination if the only way it knows to reach the destination is via the tunnel. ### What the router is actually complaining about A tunnel interface is not a real port. To forward out of Tunnel0 the router builds the outer IP header, looks up the tunnel destination, and stacks the tunnel's adjacency on whatever real interface that lookup returns. When the best route to the destination resolves out of the tunnel, the adjacency stacks on itself and the tunnel carries its own transport. CEF catches the loop while building the chain, which is why the tell-tale line lands just before the tunnel drops. It flaps rather than fails because shutting the tunnel fixes it. Tunnel down, neighbor lost, bad route withdrawn, underlay route reappears, tunnel up, neighbor re-forms, bad route re-learned. A very reliable oscillator. ### Administrative distance is the trigger Two paths to the tunnel destination exist, and the router is not choosing between sane and insane. It chooses by [the preference value each routing source carries](https://www.pinglabz.com/administrative-distance/), and nothing in that comparison knows one path runs through the tunnel it is about to break. In the lab below the underlay is OSPF (AD 110) and EIGRP runs over the tunnel advertising the endpoint loopbacks (AD 90). Internal EIGRP beats OSPF, so the tunnel-learned route wins the instant the neighbor comes up and the tunnel dies. Swap the protocols and the same topology never recurses, which is why one team swears the design is fine while another watches a tunnel flap every 40 seconds. The real lesson: recursion is not caused by the tunnel. It is caused by **advertising the tunnel destination prefix into the routing protocol running over the tunnel**. If that prefix never travels inside the tunnel, the RIB has one candidate and nothing to lose to. ### The repro, captured Two iol-xe routers on IOS XE 17.18.2\. Physical link 10.0.12.0/30, Tunnel0 on 172.16.0.0/30 sourced from and destined to the peer loopback, OSPF on the physical link so the tunnel comes up, then EIGRP 100 over Tunnel0 advertising those loopbacks. ``` R1# show ip interface brief | include Tunnel0 Tunnel0 172.16.0.1 YES TFTP up up R1# show ip route 2.2.2.2 Routing entry for 2.2.2.2/32 Known via "ospf 1", distance 110, metric 11 * 10.0.12.2, from 2.2.2.2, via Ethernet0/0 ``` `via Ethernet0/0` is what a healthy tunnel destination looks like: a real interface, not Tunnel0\. Then EIGRP forms over the tunnel and re-learns 2.2.2.2 at AD 90. ``` *DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 172.16.0.2 (Tunnel0) is up: new adjacency *ADJ-5-PARENT: Midchain parent maintenance for IP midchain out of Tunnel0 - looped chain attempting to stack *TUN-5-RECURDOWN: Tunnel0 temporarily disabled due to recursive routing *LINEPROTO-5-UPDOWN: Line protocol on Interface Tunnel0, changed state to down *DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 172.16.0.2 (Tunnel0) is down: interface down ``` Those five lines are one causal chain. The neighbor comes up, the bad route installs, and `%ADJ-5-PARENT ... looped chain attempting to stack` is CEF failing to build the adjacency. That line always lands immediately before `RECURDOWN`. The tunnel then drops and takes the neighbor with it, setting up the next cycle. The best diagnostic is not a debug, it is `show ip route` against the tunnel destination. ### Fix 1: pin the destination with a static route A static route carries AD 1, which beats anything a routing protocol can offer, so the underlay path wins permanently: ``` R1(config)# ip route 2.2.2.2 255.255.255.255 10.0.12.2 R1# show ip route 2.2.2.2 Routing entry for 2.2.2.2/32 Known via "static", distance 1, metric 0 * 10.0.12.2 R1# show ip interface brief | include Tunnel0 Tunnel0 172.16.0.1 YES TFTP up up ``` `Known via "static", distance 1` is the proof, and the tunnel stayed up and stable from there. In this article's addressing that is `ip route 203.0.113.1 255.255.255.255 198.51.100.2`. Configure it on **both** ends, because recursion is per-router and fixing R1 alone leaves R2 flapping. Make it a /32 so it wins on prefix length as well as distance. And a static next hop does not follow underlay reconvergence, so on a multi-path underlay point it at something stable. ### Fix 2: keep the transport out of the overlay protocol The static route treats the symptom. Not advertising the tunnel endpoints into the tunnel's own routing protocol removes the cause, and scales better because there is no per-destination line to forget. Cheapest version: leave the endpoint loopbacks out of the overlay protocol's network statements. Where they arrive by redistribution or a summary you do not control, filter them outbound on the tunnel instead, for example an EIGRP `distribute-list prefix NO-TRANSPORT out Tunnel0` denying the endpoint /32s. OSPF is harder inside an area, since you cannot filter LSAs between routers in one area, so keep the transport prefixes in another area and filter at the ABR. ### Fix 3: put the transport in its own VRF The structural fix makes recursion impossible rather than unlikely. Put the underlay interface and the tunnel source in their own VRF (the front-door VRF, or fVRF) and resolve the destination there with `tunnel vrf UNDERLAY`, while the tunnel interface stays in the global table. The lookup now happens in a table the overlay protocol cannot reach into, so no amount of careless redistribution over the tunnel can produce a recursive route. The [DMVPN cluster guide](https://www.pinglabz.com/dmvpn/) covers the pattern at scale, where a hub NBMA address leaking into the overlay EIGRP breaks every spoke at once. Whichever fix you pick, add `keepalive 5 3` while you are in there. It will not prevent recursion, but it catches a tunnel that is up for the wrong reasons, which matters because [a GRE tunnel reports up/up while the far end is dead](https://www.pinglabz.com/gre-tunnel-keepalives/). ## Firewall Blocking IP Protocol 47 The classic "works in lab, fails in production" failure. **Symptom:** Tunnel0 shows up / up. Underlay ping between tunnel sources works. Overlay ping fails with no response. **Cause:** A stateful firewall in the underlay drops IP protocol 47 because it is not TCP, UDP or ICMP. It sees GRE as a non-standard protocol and drops it silently. **Diagnosis:** ``` R1# show interface Tunnel0 | include packets Input packets : 0 Output packets : 5234 ``` Output packets growing with input packets at zero is the unmistakable sign that traffic is leaving but nothing is coming back. Confirm on the underlay: SPAN both endpoints' WAN interfaces and look for IP protocol 47 in each direction. **Fix:** Allow IP protocol 47 between the two tunnel-source IPs on every firewall in the path. ``` ! On a Cisco ASA / FTD access-list OUTSIDE-IN extended permit gre host 203.0.113.1 host 198.51.100.1 access-list INSIDE-OUT extended permit gre host 198.51.100.1 host 203.0.113.1 ``` ## MTU and Fragmentation Problems **Symptom:** Small packets and pings work. Some applications are fine, others stall, time out or load partially, and HTTPS to specific sites hangs. **Cause:** Inner packets are too large after GRE and IPsec encapsulation. Without MSS clamping, TCP endpoints negotiate a segment size on a 1500-byte assumption, and full-size segments are either dropped (DF=1 plus filtered ICMP is a PMTUD black hole) or fragmented. **Diagnosis:** ``` R1# ping 10.0.0.2 size 1400 df-bit !!!!! R1# ping 10.0.0.2 size 1500 df-bit M.M.M Success rate is 0 percent (0/5) ``` "M" is "could not fragment", exactly what a real DF=1 packet hits. **Fix:** Two lines, on both ends: ``` interface Tunnel0 ip mtu 1400 ip tcp adjust-mss 1360 ``` The math, the IPv6 considerations, and how to walk the size up to find your real path MTU are all in [why a GRE tunnel drops large packets but passes pings](https://www.pinglabz.com/gre-tunnel-mtu/). ## Keepalive Flap with IPsec Rekey **Symptom:** Tunnel up most of the time, down briefly about once an hour, with logs showing `Tunnel0 line protocol changed state to down` then `up`. **Cause:** GRE keepalives time out during the IPsec SA rekey window. The default IKEv2 lifetime is 3,600 seconds, which matches the symptom timing. **Diagnosis:** ``` R1# show crypto ikev2 sa Tunnel-id Local Remote fvrf/ivrf Status 1 198.51.100.1/500 203.0.113.1/500 none/none READY Life/Active Time: 3600/3540 sec <- about to rekey ``` If "Active Time" is close to "Life" and the tunnel just flapped, the events are correlated. **Fix:** Loosen the keepalive retry count or interval so it rides through a brief rekey window: ``` interface Tunnel0 keepalive 10 5 ! 50-second timeout instead of 30 ``` Or stagger the IPsec lifetimes so the two ends do not rekey at the same second, or move to BFD. If you are still weighing whether to wrap GRE in IPsec at all, [the trade-off between a plain GRE tunnel and an encrypted one](https://www.pinglabz.com/gre-vs-ipsec/) is worth settling before tuning timers around it. ## OSPF Neighbor Stuck in INIT **Symptom:** Tunnel up, overlay ping works, but the OSPF neighbor reaches INIT or 2-WAY and never FULL. **Cause:** OSPF hellos are one-way. Either an ACL blocks 224.0.0.5 inbound on one end, or the authentication, area numbers or network types do not match. **Diagnosis:** ``` R1# debug ip ospf hello *Apr 30 14:35:12: OSPF: Send hello to 224.0.0.5 area 0 on Tunnel0 from 10.0.0.1 *Apr 30 14:35:22: OSPF: Send hello to 224.0.0.5 area 0 on Tunnel0 from 10.0.0.1 ! No "Rcv hello" entries ``` If R1 only sends and never receives, R2's hellos are being dropped on the way. Check on R2: ``` R2# debug ip ospf hello *Apr 30 14:35:14: OSPF: Send hello to 224.0.0.5 area 0 on Tunnel0 from 10.0.0.2 ! On R2 you see hellos sent in both directions but R2 also sees no "Rcv hello" ``` Both ends sending and neither receiving means something in between is dropping the multicast. A standard "permit gre" rule passes GRE-encapsulated multicast, but custom ACLs interfere. **Fix:** Verify both ends match on area, hello/dead timers, network type and authentication. Run `show ip ospf interface Tunnel0` on each side and compare. ## Tunnel0 Line Protocol Down **Symptom:** Tunnel0 reports `down/down` or `up/down`. Configuration looks correct. **Causes (in order of likelihood):** 1. The tunnel source IP does not exist on the local router. Either the configured source-interface is down, or the source IP was misspelled. 2. There is no route in the IP routing table to the tunnel destination. 3. The tunnel destination cannot be the same as the tunnel source. If the tunnel never worked rather than stopped working, compare it against [a known-good GRE tunnel build on IOS XE](https://www.pinglabz.com/gre-tunnel-configuration-cisco/) before debugging further. **Diagnosis:** ``` R1# show interface Tunnel0 | include source|destination|line Tunnel0 is up, line protocol is down Tunnel source 198.51.100.1, destination 203.0.113.1 R1# show ip route 203.0.113.1 % Network not in table ``` "Network not in table" is the smoking gun. **Fix:** Add a route to the tunnel destination via whatever next hop the underlay requires. ``` R1(config)# ip route 203.0.113.1 255.255.255.255 198.51.100.2 ``` ## Useful Debugs Each debug is paired with what it shows. Run them on a lab tunnel first: on production they produce a lot of output. `debug tunnel keepalive` GRE keepalive packets sent and received `debug tunnel` All tunnel-state transitions and events `debug ip ospf hello` OSPF Hello packets sent and received `debug crypto ikev2` IKEv2 SA negotiation and rekey events `debug crypto ipsec` IPsec SA install and tear-down events `debug ip packet detail` (with ACL!) Per-packet IP processing. Always restrict to a small ACL For `debug ip packet detail`, restrict the scope tightly: ``` R1(config)# ip access-list extended DBG R1(config-ext-nacl)# permit ip host 198.51.100.1 host 203.0.113.1 R1(config-ext-nacl)# permit ip host 203.0.113.1 host 198.51.100.1 R1# debug ip packet detail 100 ! 100 references the access-list-extended ACL number; for named ACLs use 'list DBG' ``` Always `undebug all` when you are done. A debug left running in production has consumed many a router CPU. ## Packet Capture Strategy For problems that resist show-and-debug, capture the packets. Embedded Packet Capture on IOS XE writes to a buffer or a file: ``` R1(config)# ip access-list extended GRE-CAPTURE R1(config-ext-nacl)# permit gre host 198.51.100.1 host 203.0.113.1 R1(config-ext-nacl)# permit gre host 203.0.113.1 host 198.51.100.1 R1# monitor capture CAP interface GigabitEthernet1 both R1# monitor capture CAP access-list GRE-CAPTURE R1# monitor capture CAP buffer size 5 R1# monitor capture CAP start ! ... wait for the issue to reproduce ... R1# monitor capture CAP stop R1# show monitor capture CAP buffer brief R1# monitor capture CAP export bootflash:gre.pcap ``` Open the pcap in Wireshark. GRE shows as IP protocol 47, the header carries the encapsulated protocol type, and the inner packet decodes automatically. It is the fastest way to confirm the encapsulation is what you expected, that keepalives are being reflected, and that IPsec is wrapping the GRE correctly. ## When to Escalate If the tunnel works in lab but fails over a carrier path, suspect a middlebox you cannot reach: - ISP CGNAT that mishandles IP protocol 47. - Carrier MPLS L3VPN with ACLs that drop GRE. - Customer-edge DPI firewall that rewrites GRE keepalive payloads. - 5G links with a 1380-byte underlay MTU (drop `ip mtu` to 1300 to test). Open a ticket with captures from both ends and timestamps that line up. Most carrier "GRE does not work" issues turn out to be a default deny on a transit firewall, fixed quickly once they see them. ## Summary GRE is robust enough that production failures fall into a few recognizable patterns: recursive routing, IP protocol 47 blocked, MTU mismatch, keepalive flap during IPsec rekey, and routing-protocol issues that look like GRE issues. Run the triage first, match the symptom, apply the fix. - `%TUN-5-RECURDOWN` means the best route to the tunnel destination resolves out of the tunnel. Confirm with `show ip route `, not a debug. - `%ADJ-5-PARENT ... looped chain attempting to stack` lands immediately before it. That pairing is the fingerprint. - The cause is administrative distance: the overlay protocol's route (90 or 110) beats the underlay's. Never advertise the tunnel endpoints into the protocol running over the tunnel. - Three fixes, increasingly permanent: a static /32 on both ends, filtering the transport out of the overlay, or a front-door VRF. - The other failure modes: IP protocol 47 dropped, MTU, keepalive timing against IPsec rekey, one-way OSPF hellos, and a missing route to the destination. If you bookmark one thing, bookmark the five-minute triage at the top. It separates "the tunnel is broken" from "the routing on top of the tunnel is broken," and that distinction shapes the rest of the session. The full GRE coverage is at the [PingLabz GRE pillar](https://www.pinglabz.com/gre/). ### mGRE and DMVPN Introduction URL: https://www.pinglabz.com/mgre-dmvpn-introduction/ Last updated: 2026-07-11T19:14:30.000Z Standard GRE is point-to-point: one tunnel, one neighbor, one underlay destination. That works for two sites. For 50 branches and 2 hubs, point-to-point GRE means 50 manually configured tunnels per hub, plus another 1,225 tunnels if you want full mesh between branches. Nobody does that. The answer is multipoint GRE (mGRE), which lets a single tunnel interface serve many remote endpoints, plus NHRP (Next Hop Resolution Protocol) to dynamically map them, plus IPsec to encrypt them. Stack those three together and you have DMVPN, the Cisco standard for hub-and-spoke and dynamic spoke-to-spoke overlays. This article is the introduction to mGRE and DMVPN, the architecture, the three DMVPN phases, and a baseline Phase 3 hub-and-spoke configuration. Part of the [PingLabz GRE Tunnels](https://www.pinglabz.com/gre/) cluster. ## What Multipoint GRE Adds Plain GRE has one tunnel destination configured under the interface: ``` interface Tunnel0 tunnel source 198.51.100.1 tunnel destination 203.0.113.1 ! <- single destination tunnel mode gre ip ``` mGRE removes the destination: ``` interface Tunnel0 tunnel source 198.51.100.1 tunnel mode gre multipoint ! <- no destination, multipoint mode ``` The tunnel can now have many endpoints. The router needs some way to know, for each inner-IP destination, which underlay-IP destination to use as the outer header. That mapping is what NHRP provides. ## NHRP: The Address Resolution for mGRE NHRP (Next Hop Resolution Protocol, RFC 2332) is essentially ARP for tunnel networks. It answers the question "I have an inner-overlay IP I want to send to. What underlay IP should I use as the outer-tunnel destination?" In a DMVPN, every spoke registers its overlay-to-underlay mapping with the hub (the NHS, Next Hop Server). When a spoke wants to talk to another spoke, it asks the hub for that spoke's underlay IP, then builds a direct GRE tunnel. The relevant NHRP commands on a spoke pointing to a hub: ``` interface Tunnel0 ip address 10.0.0.10 255.255.255.0 ip nhrp network-id 1 ip nhrp nhs 10.0.0.1 ! Hub's overlay IP (NHS) ip nhrp map 10.0.0.1 198.51.100.1 ! Static map: hub overlay -> hub underlay ip nhrp map multicast 198.51.100.1 ! Send multicast to hub tunnel source GigabitEthernet1 tunnel mode gre multipoint tunnel key 12345 ! Same on all members tunnel protection ipsec profile IPSEC-GRE ``` The static map tells the spoke how to reach the hub initially. Once the spoke registers, the hub knows the spoke's underlay IP and can advertise it to other spokes on demand. The `tunnel key` is the GRE Key field; it must match across all DMVPN members and serves as a tunnel identifier when one router has multiple mGRE tunnels. ## DMVPN Phases DMVPN has evolved through three phases, each adding spoke-to-spoke capabilities. Modern deployments default to Phase 3. Phase 1 Spoke-to-spoke? No - all traffic via hub How it works Spokes have static map only to hub; hub re-encapsulates spoke-to-spoke traffic Use today Simple sites where hub-only is acceptable; hub bandwidth is sized for all traffic Phase 2 Spoke-to-spoke? Yes - dynamic spoke-to-spoke How it works Spokes resolve other spokes' underlay IPs via NHRP and build direct tunnels; routing protocol must preserve the spoke as next-hop Use today Largely superseded by Phase 3; still in use in older deployments Phase 3 Spoke-to-spoke? Yes - dynamic spoke-to-spoke with NHRP shortcut How it works Initial packets go via hub; NHRP shortcut sends a redirect; spokes build direct tunnel and reroute Use today Modern default for new DMVPN designs Phase 1 is conceptually simplest. The hub does all the work. Spoke-to-spoke traffic is just two hub-traversal hops. It is fine for designs where the hub has plenty of bandwidth and CPU, but it is wasteful for high-volume direct traffic. Phase 2 added direct spoke-to-spoke but had operational warts: the routing protocol had to preserve the original next-hop (which broke summarization), and the trigger for spoke-to-spoke resolution was traffic destined for a remote spoke. It worked but design decisions were tightly coupled. Phase 3 is the modern answer. The hub-and-spoke control-plane stays clean (the hub can summarize routes), but when a spoke sends traffic to another spoke via the hub, the hub returns an NHRP redirect telling the spoke "actually, you can reach that destination directly at this underlay IP." The spoke builds the direct tunnel, future packets bypass the hub, and the design scales cleanly. ## Hub Configuration (DMVPN Phase 3) ``` ! ---- Hub: HQ1 ---- crypto ikev2 keyring KR-DMVPN peer ANY address 0.0.0.0 0.0.0.0 pre-shared-key local DMVPN_KEY pre-shared-key remote DMVPN_KEY ! crypto ikev2 profile IKEV2-DMVPN match identity remote address 0.0.0.0 authentication local pre-share authentication remote pre-share keyring local KR-DMVPN ! crypto ipsec transform-set TS esp-aes 256 esp-sha256-hmac mode transport ! crypto ipsec profile IPSEC-DMVPN set transform-set TS set ikev2-profile IKEV2-DMVPN ! interface Tunnel0 ip address 10.0.0.1 255.255.255.0 ip mtu 1400 ip tcp adjust-mss 1360 no ip split-horizon eigrp 100 ! or no ip ospf flood-reduction ip nhrp authentication NHRPSEC ip nhrp network-id 1 ip nhrp redirect ! Phase 3: hub sends redirects ip nhrp map multicast dynamic tunnel source GigabitEthernet1 tunnel mode gre multipoint tunnel key 12345 tunnel protection ipsec profile IPSEC-DMVPN ! router eigrp 100 network 10.0.0.0 0.0.0.255 no auto-summary ``` Key hub-side commands: - `ip nhrp redirect` \- this is the Phase 3 enabler. The hub watches for traffic flowing through it that could go spoke-to-spoke directly, and sends an NHRP redirect to the source spoke. - `ip nhrp map multicast dynamic` \- this lets the hub flood multicast (routing-protocol hellos, broadcasts) to all registered spokes without static maps. - `no ip split-horizon eigrp 100` \- on a multipoint interface, EIGRP would normally not advertise routes back out the interface they came in on (split horizon). Disable that on a hub so it can re-advertise spoke routes to other spokes. ## Spoke Configuration (Phase 3) ``` ! ---- Spoke: Branch ---- crypto ikev2 ... (same as hub) crypto ipsec ... (same as hub) ! interface Tunnel0 ip address 10.0.0.10 255.255.255.0 ip mtu 1400 ip tcp adjust-mss 1360 ip nhrp authentication NHRPSEC ip nhrp network-id 1 ip nhrp shortcut ! Phase 3: spoke acts on redirects ip nhrp nhs 10.0.0.1 ! Hub overlay IP ip nhrp map 10.0.0.1 198.51.100.1 ! Static map to hub ip nhrp map multicast 198.51.100.1 tunnel source GigabitEthernet1 tunnel mode gre multipoint tunnel key 12345 tunnel protection ipsec profile IPSEC-DMVPN ! router eigrp 100 network 10.0.0.0 0.0.0.255 network 192.168.10.0 0.0.0.255 ! Branch LAN no auto-summary ``` The spoke-side `ip nhrp shortcut` command is the Phase 3 spoke counterpart to the hub's `ip nhrp redirect`: when the spoke receives a redirect, it issues an NHRP resolution to the target spoke's underlay IP, builds the direct tunnel, and modifies its CEF table to send future packets directly. ## Verifying DMVPN ``` HUB# show dmvpn Legend: Attrb --> S - Static, D - Dynamic, I - Incomplete N - NATed, L - Local, X - No Socket T1 - Route Installed, T2 - Nexthop-override C - CTS Capable Type:Hub, NHRP Peers:2, # Ent Peer NBMA Addr Peer Tunnel Add State UpDn Tm Attrb ----- --------------- --------------- ----- -------- ----- 1 203.0.113.10 10.0.0.10 UP 00:14:23 D 1 203.0.113.20 10.0.0.20 UP 00:13:01 D HUB# show ip nhrp 10.0.0.10/32 via 10.0.0.10 Tunnel0 created 00:14:23, expire 01:45:36 Type: dynamic, Flags: registered nhop NBMA address: 203.0.113.10 ``` On the spoke, after some spoke-to-spoke traffic flows, you should see a dynamic NHRP entry for the other spoke: ``` SPOKE10# show dmvpn Type:Spoke, NHRP Peers:2, # Ent Peer NBMA Addr Peer Tunnel Add State UpDn Tm Attrb ----- --------------- --------------- ----- -------- ----- 1 198.51.100.1 10.0.0.1 UP 00:14:23 S 1 203.0.113.20 10.0.0.20 UP 00:00:34 D ``` The "S" entry is the static hub mapping. The "D" entry is the dynamic spoke-to-spoke tunnel built after the hub redirect. ## Design Considerations - **Hub redundancy.** A single hub is a single point of failure. Production DMVPN designs use two hubs, each spoke registers with both, and the routing protocol picks the active path. The hubs themselves do not need to be aware of each other beyond running the same DMVPN profile. - **Routing protocol choice.** EIGRP and OSPF both work over DMVPN. EIGRP scales better in mesh designs because it does not have the LSA-flood problem; OSPF works fine for hub-and-spoke up to a few hundred spokes if you tune network types correctly. iBGP is also common, especially if the WAN already runs BGP. - **Underlay path control.** If both the underlay and the DMVPN overlay run dynamic routing, recursive-routing is a real risk. Lock the underlay route to the hub's tunnel-source IP with a static route on each spoke. - **NHRP authentication.** The `ip nhrp authentication` string is plaintext. It is not a security boundary; treat it as a sanity check to make sure spokes do not register against the wrong DMVPN. The real authentication is the IPsec PSK (or PKI certificates, for production). - **Spoke onboarding.** New spokes need their static hub map plus their assigned overlay IP. Many shops build a Jinja2 template and push spoke configs via Ansible / CSPC for at-scale onboarding. - **NAT behind spokes.** If a spoke is behind NAT, NAT-T (UDP 4500) must be permitted by the NAT box. NHRP also negotiates around NAT, but some scenarios (CGNAT, double-NAT) cause problems and you may need spoke-to-hub-only without spoke-to-spoke for those. ## Alternatives in 2026 DMVPN remains widely deployed but it is no longer the only answer for multi-site overlays. The competitive options: - **SD-WAN platforms.** Cisco Catalyst SD-WAN (Viptela), Fortinet Secure SD-WAN, Palo Alto Prisma SD-WAN, VMware VeloCloud, and similar all build encrypted overlays at scale, with centralized policy and orchestration. SD-WAN is the modern default for new enterprise WAN; DMVPN is increasingly the legacy / Cisco-traditional path. - **WireGuard meshes.** Tools like Tailscale, NetBird, and Headscale build WireGuard meshes with central coordination. Smaller-scale, simpler to operate, no Cisco required. - **Cloud SD-WAN.** AWS Cloud WAN, Azure Virtual WAN, and Google Network Connectivity Center are taking over the cloud-on-ramp portion of what DMVPN used to do. For a Cisco shop with existing DMVPN expertise and limited willingness to introduce new platforms, DMVPN Phase 3 over GRE-IPsec remains a perfectly good answer. For greenfield enterprises, the question worth asking is: do you want to operate the underlying mGRE/NHRP/IPsec stack yourself, or do you want a managed-overlay platform to do it for you? ## Summary mGRE plus NHRP plus IPsec is the building-block trio for DMVPN. Phase 3 is the modern default, with the hub sending NHRP redirects and the spokes acting on them to build dynamic direct tunnels. The configuration is more involved than point-to-point GRE but the operational payoff is enormous: one hub-side template plus one spoke-side template scales to hundreds of sites. The gotchas (recursive routing, OSPF network types, NHRP authentication, hub redundancy) are well-documented and predictable. This article is the introduction; the full DMVPN cluster now exists and goes much deeper, with every capture taken from a live IOS XE 17.18 lab: [DMVPN: The Complete Guide](https://www.pinglabz.com/dmvpn/) \- the cluster pillar: architecture, phases, config, and troubleshooting in one place. [DMVPN Explained](https://www.pinglabz.com/dmvpn-explained/) \- how mGRE, NHRP, and routing interlock, traced packet by packet. [NHRP Deep Dive](https://www.pinglabz.com/nhrp-deep-dive/) \- registration, resolution, and redirects with real debugs. [Phase 1 vs Phase 2 vs Phase 3](https://www.pinglabz.com/dmvpn-phase-1-2-3-differences/) \- one lab migrated live through all three phases. [DMVPN Phase 3 Configuration on IOS XE](https://www.pinglabz.com/dmvpn-phase-3-configuration/) \- the complete working build with every verification step. [Routing Over DMVPN](https://www.pinglabz.com/routing-over-dmvpn-eigrp-ospf/) \- EIGRP and OSPF design choices compared on the same lab. [Securing DMVPN with IPsec (IKEv2)](https://www.pinglabz.com/dmvpn-ipsec-profiles-ikev2/) \- profiles, transform sets, and live SA verification. [Troubleshooting DMVPN](https://www.pinglabz.com/troubleshooting-dmvpn/) \- three failures broken on purpose and diagnosed. For the GRE foundation that DMVPN builds on, work back to the [PingLabz GRE pillar](https://www.pinglabz.com/gre/). ### Routing Protocols Over GRE: OSPF, EIGRP, BGP URL: https://www.pinglabz.com/routing-protocols-over-gre/ Last updated: 2026-07-04T23:21:10.000Z The reason most GRE tunnels exist is to carry a routing protocol. Plain static routes do not need GRE, and the public internet that you would want to send routing-protocol traffic over does not natively forward multicast. GRE bridges the gap: it lets OSPF, EIGRP, and BGP form neighbor relationships across paths that would otherwise be hostile to them. This article is the practical walkthrough of running each of the three big routing protocols across a GRE tunnel on Cisco IOS XE, the recursive-routing trap that catches every newcomer, and the design decisions that matter at scale. Part of the [PingLabz GRE Tunnels](https://www.pinglabz.com/gre/) cluster. If you have not built a GRE tunnel yet, start at the [config lab](https://www.pinglabz.com/gre-tunnel-configuration-cisco/). For BGP, OSPF, or EIGRP fundamentals, see the [BGP pillar](https://www.pinglabz.com/bgp/), [OSPF pillar](https://www.pinglabz.com/ospf/), and [EIGRP pillar](https://www.pinglabz.com/eigrp/). ## Why Routing Protocols Need GRE The two routing-protocol behaviors that drive the need for GRE are multicast hellos and protocol field numbers. - **OSPF** sends Hello packets to multicast 224.0.0.5 and 224.0.0.6\. The public internet does not forward multicast. - **EIGRP** sends Hello packets to multicast 224.0.0.10\. Same problem. - **BGP** uses unicast TCP sessions. BGP itself does not need GRE - you can run iBGP or eBGP over plain IPsec or even plain internet routing - but if your overall design uses GRE tunnels, BGP rides them too. Wrapping the routing protocol in GRE turns its multicast Hello into a unicast IP packet. The internet (or any IP underlay) is happy to deliver unicast IP. The two routers see each other as if directly connected on a point-to-point link, and the routing protocol forms a neighbor. ## The Recursive Routing Trap Before any specific protocol config, understand this single failure mode that affects all of them. It is the most common GRE-over-routing-protocol problem in production. The setup: R1 has a tunnel destination of 203.0.113.1 (R2's underlay IP). R1 learns a route to 203.0.113.1 through the underlay (via BGP from the ISP, say). The tunnel comes up. R1 starts running OSPF over the tunnel. OSPF distributes routes, including, accidentally, a route to 203.0.113.1 itself if R2 advertises its underlay /32 into OSPF. Now R1 has two routes to 203.0.113.1: the underlay one (correct) and the OSPF one (advertising "next-hop is 10.0.0.2 via Tunnel0"). The OSPF route has a lower administrative distance than iBGP/static-from-ISP, so it wins. R1 now thinks the way to reach R2's underlay IP is through Tunnel0\. But Tunnel0's destination is R2's underlay IP. The tunnel cannot encapsulate packets to itself. The tunnel goes down. The OSPF neighbor dies. The OSPF route disappears. R1 falls back to the underlay route. The tunnel comes up. The cycle repeats. The IOS XE error you see in the log is unmistakable: ``` %TUN-5-RECURDOWN: Tunnel0 temporarily disabled due to recursive routing ``` The fix is to make sure the routing protocol running across the tunnel does not learn a route to the tunnel destination IP. Two ways to do that: **Static route to the tunnel destination through the underlay.** A static route has lower administrative distance than dynamic routes (1 vs 110 for OSPF, 90/170 for EIGRP, 200 for iBGP). It will always win. ``` R1(config)# ip route 203.0.113.1 255.255.255.255 198.51.100.2 ``` **Distribute-list out the underlay IP.** If you cannot use a static route (perhaps the underlay routing is fully dynamic), prevent the tunnel-side routing protocol from advertising or accepting the underlay subnet. ``` R1(config)# ip prefix-list NO-UNDERLAY seq 10 deny 203.0.113.0/24 le 32 R1(config)# ip prefix-list NO-UNDERLAY seq 20 permit 0.0.0.0/0 le 32 R1(config)# router ospf 1 R1(config-router)# distribute-list prefix NO-UNDERLAY in ``` The static-route fix is simpler and more bulletproof. Use it unless you have a specific reason not to. ## OSPF over GRE Most common case. OSPF on a tunnel interface forms a point-to-point adjacency by default with the OSPF network type "POINT\_TO\_POINT," which is fine for two-router GRE tunnels. For mGRE / DMVPN you would change the network type, but for standard GRE, defaults work. ``` R1(config)# router ospf 1 R1(config-router)# router-id 1.1.1.1 R1(config-router)# network 10.0.0.0 0.0.0.3 area 0 R1(config-router)# network 192.168.1.0 0.0.0.255 area 0 R1(config-router)# passive-interface GigabitEthernet2 R2(config)# router ospf 1 R2(config-router)# router-id 2.2.2.2 R2(config-router)# network 10.0.0.0 0.0.0.3 area 0 R2(config-router)# network 192.168.2.0 0.0.0.255 area 0 R2(config-router)# passive-interface GigabitEthernet2 ``` Verification: ``` R1# show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 2.2.2.2 0 FULL/ - 00:00:39 10.0.0.2 Tunnel0 R1# show ip ospf interface Tunnel0 Tunnel0 is up, line protocol is up Internet Address 10.0.0.1/30, Area 0 Process ID 1, Router ID 1.1.1.1, Network Type POINT_TO_POINT, Cost: 1000 ``` Note the Cost: 1000\. OSPF picks cost based on interface bandwidth, and the default tunnel-interface bandwidth is very low (100 Kbps in the show interface output). That gives OSPF a high cost on the tunnel and may cause unexpected path-selection if the tunnel is supposed to be primary. Two fixes: ``` ! Option A: pin the tunnel cost directly R1(config)# interface Tunnel0 R1(config-if)# ip ospf cost 100 ! Option B: tell OSPF the bandwidth is what you actually have R1(config)# interface Tunnel0 R1(config-if)# bandwidth 1000000 ! 1 Gbps in kbps ``` Option A is more deliberate; Option B affects more than just OSPF (EIGRP and policy-based decisions also use bandwidth). Use whichever fits your operational style. ## EIGRP over GRE EIGRP requires the same kind of multicast support as OSPF, so it works over GRE. Unlike OSPF, EIGRP cares about interface bandwidth and delay because they go directly into the composite metric. ``` R1(config)# router eigrp 100 R1(config-router)# network 10.0.0.0 0.0.0.3 R1(config-router)# network 192.168.1.0 0.0.0.255 R1(config-router)# passive-interface GigabitEthernet2 R2(config)# router eigrp 100 R2(config-router)# network 10.0.0.0 0.0.0.3 R2(config-router)# network 192.168.2.0 0.0.0.255 R2(config-router)# passive-interface GigabitEthernet2 R1# show ip eigrp neighbors EIGRP-IPv4 Neighbors for AS(100) H Address Interface Hold Uptime SRTT RTO Q Seq (sec) (ms) Cnt Num 0 10.0.0.2 Tu0 13 00:00:42 12 100 0 4 ``` The EIGRP gotcha: with the default 100 Kbps tunnel bandwidth, EIGRP's metric is enormous and the tunnel becomes the slowest path EIGRP knows about, which is the opposite of what you want. Fix the bandwidth or fix the delay: ``` R1(config)# interface Tunnel0 R1(config-if)# bandwidth 1000000 ! 1 Gbps for metric calc R1(config-if)# delay 100 ! 1000 microseconds ``` Set both ends to the same values. EIGRP uses minimum bandwidth and cumulative delay along the path; both ends must agree on the local interface values for the metric calculation to be sane. EIGRP also has a per-interface SIA (Stuck-In-Active) timer that is sensitive to high-latency tunnels. If the tunnel rides over an internet path with 200 ms latency, the default SIA timer (180 seconds) is fine, but extreme jitter or transient drops can trigger SIA flaps. The deeper EIGRP-tuning discussion is at [the EIGRP cluster](https://www.pinglabz.com/eigrp/). ## BGP over GRE BGP is unicast TCP and does not strictly need GRE. The case where BGP rides a GRE tunnel is when you have an existing GRE overlay (perhaps GRE-over-IPsec for encrypted multicast support) and you want iBGP to use the same overlay rather than building a separate transport. ``` R1(config)# router bgp 65001 R1(config-router)# bgp router-id 1.1.1.1 R1(config-router)# neighbor 10.0.0.2 remote-as 65001 R1(config-router)# neighbor 10.0.0.2 update-source Tunnel0 R1(config-router)# network 192.168.1.0 mask 255.255.255.0 R2(config)# router bgp 65001 R2(config-router)# bgp router-id 2.2.2.2 R2(config-router)# neighbor 10.0.0.1 remote-as 65001 R2(config-router)# neighbor 10.0.0.1 update-source Tunnel0 R2(config-router)# network 192.168.2.0 mask 255.255.255.0 R1# show ip bgp summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 10.0.0.2 4 65001 8 8 12 0 0 00:03:21 1 ``` The `update-source Tunnel0` command forces BGP to send TCP packets with the tunnel-interface IP as the source. Without it, BGP uses the underlay interface IP as source, which works on the underlay but does not match the IP the remote peer is configured to expect. `update-source` is the canonical fix and standard practice for any BGP session that is not on a directly-connected interface. For eBGP between separate ASes over GRE, add `ebgp-multihop` with a TTL value if the tunnel makes the BGP peer appear more than one IP hop away (which it usually does, since the tunnel destination IP is not the BGP peer IP): ``` R1(config)# router bgp 65001 R1(config-router)# neighbor 10.0.0.2 remote-as 65002 R1(config-router)# neighbor 10.0.0.2 ebgp-multihop 2 R1(config-router)# neighbor 10.0.0.2 update-source Tunnel0 ``` ## OSPF Network Types over Multipoint GRE Standard point-to-point GRE: OSPF defaults to network type POINT\_TO\_POINT, no DR/BDR election, simple. mGRE / DMVPN changes this. Multipoint tunnels require careful OSPF network-type planning: point-to-point DR/BDR?No Hello timer10 sec When to use Standard 2-router GRE tunnel broadcast DR/BDR?Yes Hello timer10 sec When to use mGRE where hub is DR (not common) non-broadcast (NBMA) DR/BDR?Yes (manual neighbor) Hello timer30 sec When to use DMVPN Phase 1 hub-and-spoke point-to-multipoint DR/BDR?No Hello timer30 sec When to use DMVPN Phase 2/3 (recommended for most modern DMVPN) For a standard two-router GRE tunnel, leave the default. For DMVPN designs, the OSPF-network-type choice drives a lot of the design decisions and is covered in the DMVPN guide. ## Design Considerations - **Use Loopbacks as tunnel sources for stable peering.** If the underlay IP changes (DHCP, DMVPN), use a Loopback as the tunnel source so the tunnel and the routing-protocol peer have a stable address. - **Authentication.** OSPF area authentication, EIGRP authentication keys, or BGP MD5 password should be set even on a tunnel that is itself encrypted. Defense in depth. - **Hello / Dead timer tuning.** Defaults are conservative. For faster failover on a dedicated tunnel, set OSPF hello to 1, dead to 4 (or use `ip ospf dead-interval minimal hello-multiplier 4`). EIGRP hello to 1, hold to 3\. BGP keepalive 3, hold 9\. Confirm both ends have matching values. - **Filtering.** Run a distribute-list or prefix-list to ensure you do not accidentally redistribute the underlay routes through the tunnel-side routing protocol. This was the recursive-routing scenario above; it is preventable but the default config does not prevent it. - **Multiple tunnels for redundancy.** A single GRE tunnel between two sites is a single point of failure. For HA, build two tunnels via different underlay paths and let the routing protocol choose the best one. Each tunnel needs its own distinct overlay /30 and tunnel destination. ## Summary OSPF, EIGRP, and BGP all run over GRE tunnels with very small additional config beyond the basic GRE setup. The protocol-specific gotchas are: OSPF cost depends on the tunnel-interface bandwidth (override it), EIGRP composite metric depends on bandwidth and delay (set both deliberately), BGP needs `update-source Tunnel0` for the TCP session to land on the right IP. The universal trap is recursive routing: prevent it with a static route to the tunnel destination via the underlay, every time, before the routing protocol has a chance to learn its way back through the tunnel. For specific protocol depth, see the [OSPF cluster](https://www.pinglabz.com/ospf/), [EIGRP cluster](https://www.pinglabz.com/eigrp/), and [BGP cluster](https://www.pinglabz.com/bgp/). For the full GRE story, the [PingLabz GRE pillar](https://www.pinglabz.com/gre/) is the cluster index. ### References - [RFC 4271 - A Border Gateway Protocol 4 (BGP-4)](https://www.rfc-editor.org/rfc/rfc4271?ref=pinglabz.com) - [Cisco BGP technology documentation](https://www.cisco.com/c/en/us/tech/ip/border-gateway-protocol-bgp/index.html?ref=pinglabz.com) Take the BGP reference with you The free BGP field-reference PDF: path attributes, best-path order, and the show commands that matter. Delivered by email, no card required. [Get the free PDF](https://www.pinglabz.com/bgp-cheatsheet/) ### GRE vs IPsec vs GRE-over-IPsec: Which Tunnel Type? URL: https://www.pinglabz.com/gre-vs-ipsec/ Last updated: 2026-06-13T20:08:34.000Z "Should I use GRE, IPsec, or GRE over IPsec?" is the question that lives at the start of every site-to-site tunnel design conversation. The answer is not "always GRE over IPsec" the way some templates suggest. It depends on what you actually need to carry and where the tunnel is going. This article is the decision framework: a side-by-side comparison of the three options, the cases where each is the right answer, and a worked example for the most common decision points. This is part of the PingLabz [GRE Tunnels: The Complete Guide](https://www.pinglabz.com/gre/) cluster. For protocol theory, start at the pillar. For configuration, see [the GRE config lab](https://www.pinglabz.com/gre-tunnel-configuration-cisco/) and [the GRE over IPsec walkthrough](https://www.pinglabz.com/gre-over-ipsec/). ## Side-by-Side Comparison Encryption Plain GRENone IPsec (tunnel mode) Yes (AES, ChaCha20, etc) GRE over IPsecYes (IPsec layer) Authentication Plain GRENone IPsec (tunnel mode) Yes (PSK, certificates) GRE over IPsecYes (IPsec layer) Integrity Plain GREOptional checksum IPsec (tunnel mode)Yes (HMAC, ICV) GRE over IPsecYes (IPsec layer) Carries multicast Plain GREYes IPsec (tunnel mode)No (without GRE) GRE over IPsecYes Carries non-IP protocols Plain GREYes (any EtherType) IPsec (tunnel mode)No GRE over IPsecYes Routing protocols (OSPF, EIGRP) Plain GREYes IPsec (tunnel mode)No (without GRE) GRE over IPsecYes iBGP neighbor over tunnel Plain GREYes IPsec (tunnel mode)Yes (unicast) GRE over IPsecYes Per-packet overhead Plain GRE\~24 bytes IPsec (tunnel mode)\~50-60 bytes GRE over IPsec\~70-80 bytes Stateful (per-session) Plain GRENo (stateless) IPsec (tunnel mode)Yes (SA per peer) GRE over IPsecYes (IPsec) + No (GRE) NAT traversal Plain GRENo (no UDP/TCP) IPsec (tunnel mode) Yes (NAT-T over UDP 4500) GRE over IPsecYes (via IPsec NAT-T) Config complexity Plain GRELow (3 lines) IPsec (tunnel mode) Medium (IKE + transform + crypto map / profile) GRE over IPsecHigh (both) CPU cost Plain GRENegligible IPsec (tunnel mode) Significant (encryption) GRE over IPsecSignificant Suitable for public internet Plain GRENo (no encryption) IPsec (tunnel mode)Yes GRE over IPsecYes Suitable for trusted underlay Plain GREYes IPsec (tunnel mode)Overkill GRE over IPsecOverkill ## The Decision Tree Three questions, in this order: 1. **Does the path need encryption?** If the tunnel rides over the public internet, an internet-VPN underlay, or any network you do not control end to end: yes. If the path is your own private MPLS, your own internal data center fabric, or a lab: no. 2. **Do you need to carry routing-protocol multicast or non-IP traffic through the tunnel?** If yes (you are running OSPF/EIGRP across the tunnel, or carrying IPv6 in IPv4, or doing something that needs the EtherType field): you need GRE. If no (the only thing crossing is unicast IP between specific subnets): you do not need GRE. 3. **Do you anticipate scaling to many sites with NHRP and dynamic spoke-to-spoke (DMVPN)?** If yes: GRE over IPsec is the natural foundation, since DMVPN is mGRE plus IPsec plus NHRP. The decision falls out: - **Encryption needed + routing protocols / multicast / DMVPN potential:** GRE over IPsec. - **Encryption needed + plain unicast IP only:** IPsec alone, in tunnel mode. - **No encryption needed + routing protocols / multicast:** Plain GRE. - **No encryption needed + plain unicast IP only:** No tunnel, just route through the underlay. ## When Plain GRE Is the Right Answer Plain GRE has a smaller place in 2026 than it did in 2010 because the public internet is now the default underlay for most enterprise WAN, but the cases still exist: - **Lab and learning environments.** You are studying for CCNP and want to see OSPF form a neighbor across a routed lab without the IPsec complication. Plain GRE. - **Private MPLS or carrier Ethernet underlays.** The underlay is contractually private and you have decided the carrier counts as trusted. GRE adds the routing-protocol-multicast capability your design needs without the encryption overhead. - **Internal data-center overlays.** Inside a DC you trust, a GRE tunnel between two routers that need a specific traffic-engineering path can be plain GRE. Most DC overlays in 2026 have moved to VXLAN (driven by EVPN), but legacy GRE designs still exist. - **IPv4-IPv6 transition tunnels (6in4).** The classic 6to4 / 6in4 tunnel is plain GRE with IPv6 inside an IPv4 outer header. No encryption is typically used because the IPv6 traffic itself is internet-facing and end-to-end secured by application-layer protocols. - **Service-provider PE-CE sub-protocols.** Some service-provider designs use GRE to carry MPLS labels or IS-IS frames across an IP-only segment. Plain GRE because the SP backbone is trusted. ## When IPsec Alone Is the Right Answer This is the right answer more often than engineers expect: - **Site-to-site between two firewalls with simple unicast traffic.** A branch site connecting to HQ over the internet, where the only traffic is users in the branch reaching corporate apps in HQ. No routing protocol needed (a static route at each end is enough). IPsec in tunnel mode is simpler, has lower overhead, and one fewer thing to misconfigure. - **Cloud VPN gateways.** AWS Site-to-Site VPN, Azure VPN Gateway, GCP HA VPN all default to IPsec without GRE. They use BGP over the IPsec tunnel for dynamic routing (BGP is unicast TCP, no GRE required), so the multicast-needs-GRE argument does not apply. - **Remote access VPN.** Client VPN like Cisco AnyConnect or strongSwan is IPsec (or SSL/TLS), not GRE. Each user gets a unicast tunnel to the headend. - **WireGuard as an alternative.** Modern deployments increasingly choose WireGuard over IPsec for its simpler config and faster crypto. WireGuard is also unicast-only without GRE. If you find yourself reaching for IPsec-alone, WireGuard is worth comparing. ## When GRE over IPsec Is the Right Answer - **Branch sites running OSPF or EIGRP across an internet underlay.** The classic case. You need routing-protocol-multicast (which IPsec alone cannot carry) plus encryption (which plain GRE does not provide). - **DMVPN deployments.** DMVPN is built on mGRE plus IPsec plus NHRP. Even DMVPN Phase 1 (no spoke-to-spoke) uses GRE over IPsec; the difference between DMVPN and a static GRE-over-IPsec mesh is the NHRP and the multipoint tunnel mode, not the encapsulation order. - **Hub-and-spoke overlays where the hub runs iBGP RR with spokes.** iBGP is unicast TCP and would work over plain IPsec, but if you also want OSPF on the same overlay (for example, OSPF for IGP between hub and spokes plus iBGP for customer routes), you are back to GRE over IPsec. - **Path-control or traffic engineering with policy routing.** Some PBR designs use GRE to fix a specific next-hop and IPsec to encrypt. The combination is what you need; either alone is insufficient. ## Alternatives Worth Knowing The GRE-vs-IPsec decision is not the only one. In 2026 there are tunneling protocols that did not exist or were not mainstream when GRE was the default answer. VXLAN What it is L2-over-UDP tunnel; the dominant data-center overlay When to consider it L2 stretch inside a DC fabric or between DCs; not for branch WAN WireGuard What it is Modern crypto VPN; UDP-based, simple config When to consider it Simple unicast site-to-site; remote-access; cloud labs SD-WAN overlays What it is Vendor-specific encrypted overlay (Cisco SD-WAN, FortiGate IPsec mesh, VeloCloud) When to consider it Multi-site enterprise WAN with centralized policy and orchestration L2TPv3 What it is Layer-2 over IP tunneling (frame relay, Ethernet pseudowire) When to consider it Service-provider PWE3 use cases IPv6-in-IPv6 / 6in4 What it is IPv6 transition tunneling When to consider it Specific dual-stack migration scenarios MPLS-in-GRE / MPLS-in-UDP What it is Carrying MPLS over an IP backbone When to consider it SP designs that bridge MPLS islands across an IP core For a typical enterprise WAN in 2026, the choice is usually GRE over IPsec (or DMVPN) for legacy / Cisco-centric designs, vs an SD-WAN overlay (Cisco, Fortinet, Palo Alto, VMware) for greenfield deployments. WireGuard fills the smaller-scale and cloud-lab gap. ## Worked Example: 30-Site Enterprise WAN Concrete scenario: 30 branch sites, two data center hubs, each branch has a single 200 Mbps internet-only WAN circuit. The applications are a mix of cloud SaaS, internal apps in the DC, and voice/video collaboration. The team wants per-site routing flexibility and dynamic failover between branches when a DC link fails. Walk through the decision tree: 1. Encryption needed? Yes - public internet underlay. 2. Routing protocols needed? Yes - 30 sites are too many to manage with static routes; OSPF or iBGP needed. 3. Scaling to many sites with NHRP? At 30 sites, hub-and-spoke is fine but full mesh would be 870 tunnels. NHRP and DMVPN are the way to handle dynamic spoke-to-spoke for direct branch-to-branch voice/video. Conclusion: DMVPN (mGRE + IPsec + NHRP). For the hub-side IPsec config, this is essentially "GRE over IPsec" multiplied across many spokes, with NHRP doing the dynamic destination resolution. That is the canonical Cisco answer. Alternatively, the same scenario in 2026 is increasingly answered with SD-WAN: deploy Cisco Catalyst SD-WAN (formerly Viptela) or a competitor, and let the SD-WAN platform handle the overlay tunnels (still IPsec, often without GRE in the modern data plane), the routing (OMP or BGP), and the policy. The trade-off is operational complexity vs flexibility: DMVPN is simpler to understand for a Cisco-trained team but harder to scale beyond a few hundred sites; SD-WAN is more capable but adds a control-plane platform to operate. Both are valid 2026 answers. ## Summary The choice between GRE, IPsec, and GRE over IPsec comes down to two questions: does the path need encryption (yes for the public internet, no for a trusted underlay), and does the traffic need GRE's multi-protocol / multicast capabilities (yes for OSPF/EIGRP/IPv6-in-IPv4, no for plain unicast IP). Plain GRE is for trusted networks plus routing protocols. IPsec alone is for encrypted unicast IP between two endpoints. GRE over IPsec is for encrypted routing-protocol-aware overlays - and for any deployment that will eventually become DMVPN. If you are designing a new enterprise WAN in 2026, the modern answers are usually DMVPN (Cisco-centric) or an SD-WAN overlay (vendor-flexible). Both build on IPsec; DMVPN explicitly uses GRE over IPsec, SD-WAN platforms often skip GRE entirely. For lab work and small deployments, plain GRE remains a perfectly good answer. For the protocol depth, see the [PingLabz GRE pillar](https://www.pinglabz.com/gre/); for vendor-specific config, see [GRE over IPsec](https://www.pinglabz.com/gre-over-ipsec/) and [mGRE and DMVPN Introduction](https://www.pinglabz.com/mgre-dmvpn-introduction/). ### GRE MTU and Fragmentation: Fixing Tunnel Packet Loss URL: https://www.pinglabz.com/gre-tunnel-mtu/ Last updated: 2026-06-13T20:08:34.000Z The most common production problem with GRE tunnels is not configuration. It is MTU. Pings work. Small TCP connections work. Then a user opens a webpage and it loads halfway. Or a database client throws timeouts. Or a backup job that runs over the tunnel hangs at exactly 95 percent. The symptom is "the tunnel is up but big things break." The cause is almost always that the inner packet is too large to fit through the path once GRE encapsulation, IPsec encryption, and underlay headers have been added, and somewhere along the way packets are being silently dropped instead of cleanly fragmented. This article walks through the math, the two-line IOS XE fix that prevents most of the pain, and the debugging recipe for diagnosing MTU problems when they have already started biting. It is part of the [PingLabz GRE Tunnels guide](https://www.pinglabz.com/gre/). ## The Overhead Math Every encapsulation adds bytes. GRE is small but IPsec is not. The standard 1500-byte Ethernet underlay you almost certainly have for the tunnel source interface eats up fast: 1500 LayerEthernet underlay MTU Cumulative remaining for inner payload1500 \-20 LayerOuter IPv4 header Cumulative remaining for inner payload1480 \-4 LayerGRE header Cumulative remaining for inner payload1476 \-16 Layer IPsec ESP header + IV (AES) Cumulative remaining for inner payload1460 \-15 Layer IPsec ESP padding (worst case) Cumulative remaining for inner payload1445 \-2 LayerIPsec ESP trailer Cumulative remaining for inner payload1443 \-16 LayerIPsec ICV (SHA-256) Cumulative remaining for inner payload1427 \-8 Layer NAT-T UDP wrapper (if NAT in path) Cumulative remaining for inner payload1419 That is why the canonical "set ip mtu 1400 on the tunnel" advice exists. 1400 is a round number that leaves headroom for the worst-case IPsec scenario plus a few bytes of margin for things that show up in production but not in lab (jumbo-frame mismatches at carrier handoffs, unexpected VLAN tags, MPLS labels). If you are running plain GRE without IPsec and you know the underlay path is consistent end-to-end, you can use 1476 (1500 minus the GRE+outer-IP overhead). Most of the time you will eventually add IPsec, and going back to retune MTU later is more painful than starting at 1400. ## The Two-Line Fix On every GRE tunnel interface in production: ``` interface Tunnel0 ip mtu 1400 ip tcp adjust-mss 1360 ``` Each line solves a different aspect of the problem. **`ip mtu 1400`** tells the router that the maximum size of an IP packet allowed onto the tunnel interface (after the inner-IP processing but before GRE encapsulation) is 1400 bytes. Anything larger gets fragmented into 1400-byte pieces before GRE wraps it. After GRE adds 24 bytes and IPsec adds up to 76 bytes, the resulting packet still fits in the underlay's 1500-byte MTU. No fragmentation needed at the underlay level. **`ip tcp adjust-mss 1360`** rewrites the MSS (Maximum Segment Size) value inside TCP SYN packets passing through the router. TCP endpoints negotiate MSS during the three-way handshake based on what they think the path MTU is. When the endpoints are unaware of the GRE+IPsec overhead, they advertise an MSS of 1460 (1500-byte path MTU minus 20 IP minus 20 TCP), and every full-size segment ends up being 1460 + 20 + 20 = 1500 bytes of inner IP. That is too big for the tunnel. By rewriting MSS to 1360 (1400 inner-MTU minus 20 IP minus 20 TCP), the router forces the endpoints to negotiate a segment size that fits without ever needing fragmentation. This is the elegant fix because it eliminates the problem at the source rather than coping with it downstream. Both ends of the tunnel should have these commands. The MSS clamping happens on each TCP SYN as it crosses the tunnel, regardless of direction. ## Why Cannot the Network Just Fragment? In theory, IP fragmentation handles oversized packets cleanly. In practice it does not work in 2026, for three reasons. 1. **Don't-Fragment bit is set on most TCP traffic.** Modern TCP stacks set DF=1 by default for Path MTU Discovery (PMTUD). When a router needs to fragment a DF=1 packet, it instead drops it and sends an ICMP Type 3 Code 4 ("fragmentation needed") message back to the sender, telling it to lower its packet size. PMTUD relies on that ICMP message getting back. 2. **ICMP is filtered.** Many enterprise firewalls, ISP middleboxes, and home routers drop ICMP. The "fragmentation needed" message never returns to the sender. The TCP endpoint waits, retransmits the same oversized packet, gets dropped again, and the connection stalls. This is the classic "PMTUD black hole" failure mode. 3. **Fragmentation is expensive.** Even when fragmentation works, the receiving router has to reassemble fragments before processing, which costs CPU and memory. On heavily-loaded routers it is a significant performance hit. On linecard ASICs that punt fragments to the CPU, it can collapse throughput. That is why MSS clamping is preferred over relying on PMTUD. MSS clamping makes the problem go away at the TCP-handshake level. PMTUD is the fallback for traffic that is not TCP, like UDP-based applications, where you cannot rewrite MSS because there is no MSS to rewrite. ## Symptoms of MTU Misconfiguration The classic patterns engineers report: - **"Pings work but the website does not load."** Default ICMP echo is 84 bytes (with the headers); they fit in any path. Real traffic uses full-size segments and stalls when fragmentation black-holes them. - **"Some sites work, others do not."** Sites with PMTUD-friendly paths (small replies, ICMP allowed) work. Sites with large replies and ICMP filtering do not. - **"It works sometimes."** Different MTUs are negotiated on different connection retries. The connection that happens to negotiate a small MSS works. - **"Database queries hang at exactly 95 percent."** The backup or query was sending small packets through the connection and they all flowed; the final large response (the table contents, the file body) hits the MTU wall. - **"VPN works for HTTP but not for email or file transfers."** HTTP page loads use small request headers and stream responses; bigger payloads tip over the MTU threshold. ## Diagnosing MTU Problems The single most useful diagnostic is a sized ping with the Don't-Fragment bit set. From the router that originates the tunnel: ``` R1# ping 10.0.0.2 size 1400 df-bit Type escape sequence to abort. Sending 5, 1400-byte ICMP Echos to 10.0.0.2, timeout is 2 seconds: Packet sent with the DF bit set !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 5/5/8 ms R1# ping 10.0.0.2 size 1500 df-bit Type escape sequence to abort. Sending 5, 1500-byte ICMP Echos to 10.0.0.2, timeout is 2 seconds: Packet sent with the DF bit set M.M.M Success rate is 0 percent (0/5) ``` `M` in the output means "could not fragment" - the router on the path tried to forward a packet larger than the egress MTU but the DF bit prevented it. That is the canonical signature of an MTU problem. Walk the size up and down to find exactly where the limit is: ``` R1# ping 10.0.0.2 size 1450 df-bit ! 1450 fails -> below 1450 the path is OK R1# ping 10.0.0.2 size 1400 df-bit ! 1400 succeeds -> safe for IP MTU R1# ping 10.0.0.2 size 1420 df-bit ! finer grain ``` The largest size that succeeds with df-bit is your effective path MTU through the tunnel. Set `ip mtu` to that value or below. From a Linux endpoint, the same idea with `ping`: ``` $ ping -M do -s 1400 192.168.2.10 PING 192.168.2.10 (192.168.2.10) 1400(1428) bytes of data. 1408 bytes from 192.168.2.10: icmp_seq=1 ttl=63 time=5.01 ms $ ping -M do -s 1500 192.168.2.10 PING 192.168.2.10 (192.168.2.10) 1500(1528) bytes of data. ping: local error: message too long, mtu=1500 ping: local error: message too long, mtu=1500 ``` Linux's `-M do` is "don't fragment, fail if too big," equivalent to df-bit on Cisco. ## Useful Show Commands ``` R1# show interface Tunnel0 | include MTU MTU 17916 bytes, ... Tunnel transport MTU 1476 bytes ``` The Cisco "MTU 17916" is the platform-supported maximum size; ignore it. The "Tunnel transport MTU 1476" is the IOS XE-calculated maximum after GRE+IP overhead. You override this with `ip mtu 1400`. ``` R1# show interface Tunnel0 | include ip mtu|adjust IP MTU 1400 bytes IP TCP MSS 1360 bytes ``` If you do not see those two lines, the commands are not applied. Re-add them. ## IPv6 Considerations IPv6 changes the MTU story in two ways. First, IPv6 routers do not fragment in transit at all - the source must do PMTUD, and an oversized packet anywhere in the path is dropped with an ICMPv6 "Packet Too Big" reply. Second, the minimum IPv6 MTU is 1280 bytes; if any link in the path is below that, IPv6 will not work. For GRE tunnels carrying IPv6: ``` interface Tunnel0 ipv6 mtu 1400 ipv6 tcp adjust-mss 1360 ``` Same logic, IPv6 versions of the commands. The 1400 inner-MTU value is conservative enough to handle GRE plus IPsec plus the IPv6-vs-IPv4 outer-header difference. ## When You Cannot Avoid Fragmentation If endpoints generate UDP traffic with DF=1 set and you cannot rewrite the application, occasional fragmentation will happen. On Cisco IOS XE you can pre-fragment at the tunnel ingress so the IPsec layer never has to: ``` crypto ipsec fragmentation before-encryption ``` This is a global command. It tells the router to fragment large packets at the GRE encapsulation layer before they hit IPsec, which avoids the more-expensive after-encryption fragmentation path. It does not solve the underlying MTU problem - the inner IP packet is still being fragmented - but it keeps the dataplane fast path engaged. ## Summary Two lines of config eliminate 90 percent of GRE tunnel MTU pain on Cisco IOS XE: `ip mtu 1400` and `ip tcp adjust-mss 1360`, on both ends of the tunnel. The math behind 1400 is "1500-byte underlay minus worst-case GRE plus IPsec overhead minus margin." The MSS clamping prevents TCP endpoints from negotiating segment sizes that the tunnel cannot carry without fragmentation, which removes the entire PMTUD-black-hole failure class. If you are debugging an MTU problem in production, the diagnostic is a sized ping with the DF bit set: walk the size up until packets start failing, that is your path MTU, set `ip mtu` to that value or below. The full cluster reference is at [PingLabz GRE Tunnels](https://www.pinglabz.com/gre/); the troubleshooting playbook with debug walkthroughs is at [GRE Tunnel Troubleshooting Guide](https://www.pinglabz.com/gre-tunnel-troubleshooting/). ### GRE Tunnel Keepalives Explained URL: https://www.pinglabz.com/gre-tunnel-keepalives/ Last updated: 2026-08-01T19:35:38.000Z A GRE tunnel without keepalives is a tunnel that lies to you. It will report up / up as long as the local underlay route to the destination IP exists, even if the remote router has crashed and is not processing packets. Traffic black-holes, the routing protocols still believe they have a neighbor, and you are debugging blind. Cisco GRE keepalives fix this with a small, clever mechanism that adds liveness detection without requiring any cooperation from the remote end. This article explains how they work, how to configure and tune them, and the subtle ways they interact with IPsec. The article is part of the PingLabz [GRE Tunnels: The Complete Guide](https://www.pinglabz.com/gre/) cluster. ## Why GRE Needs Keepalives GRE is stateless. Each packet stands on its own; there is no session, no sequence-number tracking by default, and no acknowledgment. The local Tunnel0 interface decides "up" versus "down" based on a single check: is there an active route in the IP routing table to the configured tunnel destination IP? If yes, the tunnel is up. The line protocol stays up indefinitely as long as the route exists, regardless of whether the far end is actually reachable, alive, or willing. That gap matters in three real scenarios: - **Remote router crash or reboot.** The route to the underlay IP is still valid (the path through the carrier is fine), but R2 is not actually answering anything. R1's tunnel stays up. Routing protocol neighbors hold for as long as their dead-timer permits, then tear down, but the tunnel itself reports up. If you are not running a routing protocol over the tunnel, traffic black-holes silently. - **Remote IPsec or interface failure.** The remote Tunnel0 interface is configured but the IPsec profile or underlay interface is broken. Packets you send arrive at the underlay router but are never decapsulated. R1's view of "up / up" is wrong. - **Asymmetric path failure.** R1 can reach R2 but not vice versa. R1's tunnel is up. R2's tunnel may be up too. Traffic flows one way and is dropped the other. Without bidirectional liveness, neither end notices. Keepalives turn the tunnel from "up if the underlay route exists" into "up if my keepalive is being acknowledged by the remote end." ## How GRE Keepalives Work The mechanism is elegant. The local router builds a small packet whose inner content is itself a GRE-encapsulated reply addressed to the local router's underlay IP. It sends that packet through the tunnel as a normal GRE packet. The remote router receives it, looks at the inner header, sees a GRE packet bound for the local router's underlay IP, and routes it back through the tunnel. The local router receives its own keepalive back and counts that as a successful round-trip. The clever part: the remote router does not need any keepalive configuration. It just needs to be alive and processing GRE. As long as the remote box is functional enough to do basic IP routing, it will reflect the keepalive back automatically. There is no protocol negotiation, no version mismatch, no need for both ends to enable keepalives. Each end runs its keepalive independently. R1's keepalives prove R2 is alive. R2's keepalives prove R1 is alive. They are unrelated. You can turn on keepalives on one end without changing the other end's config, and that one end will detect failures of the other. ## Configuration ``` R1(config)# interface Tunnel0 R1(config-if)# keepalive 10 3 ``` The two numbers are the interval (seconds) and the retry count. `keepalive 10 3` means send a keepalive every 10 seconds and declare the tunnel down after 3 consecutive missed responses. So the worst-case detection time is 30 seconds (the third keepalive sent at second 30 plus its full timeout). The IOS XE default if you type just `keepalive` with no arguments is also 10 seconds and 3 retries. For faster failover, you can tighten the interval: ``` R1(config-if)# keepalive 3 3 ``` That gives 9-second detection, at the cost of 33 percent more keepalive traffic on the tunnel. Going below `keepalive 1 3` is supported but the marginal benefit drops off and the chance of false positives during normal jitter rises. For sub-second failover requirements, BFD over the tunnel (where the platform supports it) is the better tool than aggressive keepalives. ## Verifying Keepalives Are Working ``` R1# show interface Tunnel0 | include Keepalive Keepalive set (10 sec), retries 3 ``` If keepalives are not configured, this shows `Keepalive not set`. Fix it. To watch keepalives in action: ``` R1# debug tunnel keepalive Tunnel keepalive debugging is on R1# *Apr 30 14:12:03.215: Tunnel0: GRE keepalive sent s=198.51.100.1, d=203.0.113.1 *Apr 30 14:12:03.219: Tunnel0: GRE keepalive recv s=203.0.113.1, d=198.51.100.1 *Apr 30 14:12:13.215: Tunnel0: GRE keepalive sent s=198.51.100.1, d=203.0.113.1 *Apr 30 14:12:13.220: Tunnel0: GRE keepalive recv s=203.0.113.1, d=198.51.100.1 ``` Sent, received, sent, received. If you see only "sent" entries with no matching "recv" within the configured interval, the remote end is not reflecting them. That points at a remote router problem, an underlay path problem, or IPsec issues if the tunnel is wrapped. Turn debug off when you are done: `no debug tunnel keepalive` or `undebug all`. ## Keepalives and IPsec This is where the subtle gotchas live. When you wrap GRE in IPsec, keepalives are still GRE packets, but they now have to traverse the IPsec encryption. Three interactions matter: - **IPsec rekey timing.** When the IPsec SA rekeys (default lifetime is typically 3,600 seconds for IKEv2), there is a brief window where outbound traffic uses the new SA but the remote end has not switched. A keepalive sent in that window may be dropped. With `keepalive 10 3` you can ride through a rekey because the next keepalive 10 seconds later will use the established new SA. With `keepalive 1 3` you are vulnerable to false-positive tunnel-down events at every rekey. - **DPD (Dead Peer Detection) overlap.** IPsec has its own liveness check called DPD. If you have both DPD on the IPsec session and GRE keepalives, you have two redundant liveness mechanisms. They do not conflict, but they should be tuned consistently. A common pattern is GRE keepalive 10 3 plus IPsec DPD interval 30\. The GRE side detects faster; DPD provides a safety net for the IPsec layer specifically. - **NAT-T keepalives.** If either tunnel endpoint sits behind NAT, IOS XE sends NAT-T keepalives every 20 seconds (default) to maintain the UDP 4500 NAT pinhole. Those are unrelated to GRE keepalives but show up as additional traffic. Do not confuse the two during traffic-analysis. ## Vendor Interoperability GRE keepalives are a Cisco extension to GRE. They are not in RFC 2784 or RFC 2890\. Other vendors have their own approaches: Cisco IOS / IOS XE / IOS XR GRE keepalive supportYes (this article) Notes Reference implementation Juniper Junos GRE keepalive support Limited; OAM via Service PIC Notes Not the same packet format; do not assume interop Arista EOS GRE keepalive supportYes, Cisco-style Notes Generally interoperates with Cisco Linux (kernel GRE) GRE keepalive supportNo native Notes Use BFD over the tunnel or a userland keepalive FortiGate GRE keepalive supportLimited Notes Most FortiGate GRE deployments rely on routing-protocol dead-timers MikroTik RouterOS GRE keepalive supportYes Notes Cisco-compatible keepalive packet format For Cisco-to-non-Cisco GRE tunnels, do not assume keepalives work bidirectionally. Test both directions. If keepalives do not interoperate, fall back to running a routing protocol with aggressive timers across the tunnel, or use BFD if both platforms support it. ## Alternatives to GRE Keepalives - **BFD over the tunnel.** Bidirectional Forwarding Detection. Sub-second failure detection, much more efficient than tightly-tuned keepalives, and standardized. Supported on most modern Cisco platforms. Use it when you need millisecond-grade failover and your platform / IOS XE version supports BFD on tunnel interfaces. - **Routing-protocol dead-timers.** If you are running OSPF or EIGRP across the tunnel, the routing protocol's hello / dead-timer is itself a liveness check. Tune dead-timer aggressively (down to 3 seconds for OSPF) and you have effectively-instant detection without GRE keepalives. The downside: if the routing protocol breaks for a different reason (LSA flap, neighbor mismatch), the tunnel reports "up" but unusable. - **NHRP for DMVPN spokes.** In a DMVPN deployment, NHRP messages between spokes and hubs serve as liveness checks for the dynamic spoke-to-spoke tunnels. GRE keepalives on the static spoke-to-hub tunnels are still common; the spoke-to-spoke side relies on NHRP holdtimes. ## Troubleshooting Keepalive Failures Tunnel flaps every 30 seconds with no other obvious change Likely cause Keepalives configured but not returning Where to check `debug tunnel keepalive` on both ends Tunnel up but traffic black-holes Likely cause Keepalives never enabled Where to check `show interface Tunnel0 | inc Keepalive` Keepalives returning but routing protocol dies Likely cause Routing-protocol issue, not GRE Where to check `show ip ospf neighbor` / `show ip eigrp neighbors` Keepalives flap during IPsec rekey only Likely cause Aggressive timers + IPsec rekey window Where to check Loosen keepalive to `10 3` or stagger rekey lifetime Cisco-to-non-Cisco tunnel: keepalives one-way Likely cause Vendor-specific format mismatch Where to check Disable on the non-Cisco end and rely on routing-protocol liveness There is one failure keepalives cannot save you from, and it wears the same costume: recursive routing, where the tunnel learns the route to its own destination through itself and IOS drops it to break the loop. [The recursive-routing repro and the static that pins the transport](https://www.pinglabz.com/gre-tunnel-troubleshooting/) is the thing to check before you spend an evening in `debug tunnel keepalive`. ## Summary GRE keepalives are the bare minimum hygiene for any production GRE tunnel. They cost almost nothing in bandwidth or CPU, they require no remote-end cooperation, and they catch the failure modes that vanilla GRE tunnel-state cannot. The default `keepalive 10 3` is a sensible starting point: 30-second detection, immune to normal IPsec rekey jitter, low overhead. Tighten only if you have a measured need and you are sure the path can support it. If you take one thing away from this article, take this: every time you build a GRE tunnel, set keepalives in the same config block. There is no production scenario where leaving them off improves anything. Combined with `ip mtu 1400` from the [MTU article](https://www.pinglabz.com/gre-tunnel-mtu/) and the recursive-routing static from [the config lab](https://www.pinglabz.com/gre-tunnel-configuration-cisco/), you have eliminated the three most common GRE production failures before the first user complains. The full cluster is at [PingLabz GRE](https://www.pinglabz.com/gre/). ### GRE Tunnel Configuration: Step-by-Step Cisco IOS XE Lab URL: https://www.pinglabz.com/gre-tunnel-configuration-cisco/ Last updated: 2026-06-13T20:08:35.000Z This is the step-by-step Cisco IOS XE lab walkthrough for building a GRE tunnel between two routers. We will go from a blank config to a verified, traffic-passing tunnel, then add the production knobs (MTU clamping, keepalives, OSPF over the tunnel) one at a time. Every command is shown with the show output you should see if it worked. If a step does not match what you see, you have a clear place to start debugging. The lab is part of the PingLabz [GRE Tunnels: The Complete Guide](https://www.pinglabz.com/gre/) cluster. If you want the protocol theory before the config, start there. Otherwise, paste the configs into IOS XE and follow along. ## Lab Topology Two Cisco IOS XE routers (CSR1000v, Catalyst 8000v, ISR4000, or any IOS XE platform with IP Base/Advanced IP Services). The routers are connected through an underlay network that simulates the public internet, with no end-to-end private IP routing. 198.51.100.1/30 RouterR1 Tunnel0 source198.51.100.1 Tunnel0 destination203.0.113.1 Tunnel0 IP (overlay)10.0.0.1/30 LAN (G2)192.168.1.1/24 203.0.113.1/30 RouterR2 Tunnel0 source203.0.113.1 Tunnel0 destination198.51.100.1 Tunnel0 IP (overlay)10.0.0.2/30 LAN (G2)192.168.2.1/24 ![Lab topology: R1 (underlay 198.51.100.1/30, Tunnel0 10.0.0.1/30, LAN 192.168.1.0/24) and R2 (underlay 203.0.113.1/30, Tunnel0 10.0.0.2/30, LAN 192.168.2.0/24) connected by Tunnel0 over an IPv4 underlay](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/topology-1.png) Figure 1\. Lab topology: two routers, three address layers (underlay, overlay, LAN), one Tunnel0 between them. Goal: get a host on R1's LAN (192.168.1.0/24) to reach a host on R2's LAN (192.168.2.0/24) through the GRE tunnel, with no underlay route to the private LAN networks. ## Step 1: Verify Underlay Reachability Before building the tunnel, confirm the two routers can reach each other on their underlay IPs. If the underlay is broken, no GRE config in the world will help. ``` R1# ping 203.0.113.1 source 198.51.100.1 Type escape sequence to abort. Sending 5, 100-byte ICMP Echos to 203.0.113.1, timeout is 2 seconds: Packet sent with a source address of 198.51.100.1 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 4/5/9 ms ``` If this ping fails, fix the underlay first. Check `show ip route 203.0.113.1` to confirm the routing table has a path. If the underlay traverses a firewall, confirm ICMP is permitted between the two underlay IPs. ## Step 2: Configure the Tunnel Interfaces Both ends need a Tunnel interface with matching tunnel source / destination, an IP address on the overlay, and tunnel mode set to GRE over IP. ``` ! ---- R1 ---- R1(config)# interface Tunnel0 R1(config-if)# ip address 10.0.0.1 255.255.255.252 R1(config-if)# tunnel source 198.51.100.1 R1(config-if)# tunnel destination 203.0.113.1 R1(config-if)# tunnel mode gre ip R1(config-if)# no shutdown ! ---- R2 ---- R2(config)# interface Tunnel0 R2(config-if)# ip address 10.0.0.2 255.255.255.252 R2(config-if)# tunnel source 203.0.113.1 R2(config-if)# tunnel destination 198.51.100.1 R2(config-if)# tunnel mode gre ip R2(config-if)# no shutdown ``` Two notes on the source. First, you can use an interface name (`tunnel source GigabitEthernet1`) instead of an IP. The router takes the primary IP of that interface as the source. Using the interface name is more resilient to IP renumbering. Second, in production deployments where the underlay IP changes (DHCP from the ISP, for example), use a Loopback interface as the tunnel source and run a static route or routing protocol on the underlay so the loopback is reachable from the far end. That gives you a stable tunnel source IP regardless of underlay churn. `tunnel mode gre ip` is actually the IOS XE default - if you create a Tunnel interface and do not specify a mode, you get GRE over IPv4\. Setting it explicitly is good hygiene because it documents intent and protects you when the default changes (it has, between IOS versions). ## Step 3: Verify the Tunnel Comes Up ``` R1# show interface Tunnel0 Tunnel0 is up, line protocol is up Hardware is Tunnel Internet address is 10.0.0.1/30 MTU 17916 bytes, BW 100 Kbit/sec, DLY 50000 usec, reliability 255/255, txload 1/255, rxload 1/255 Encapsulation TUNNEL, loopback not set Keepalive not set Tunnel linestate evaluation up Tunnel source 198.51.100.1, destination 203.0.113.1 Tunnel protocol/transport GRE/IP Key disabled, sequencing disabled Checksumming of packets disabled Tunnel TTL 255, Fast tunneling enabled Tunnel transport MTU 1476 bytes Tunnel transmit bandwidth 8000 (kbps) Tunnel receive bandwidth 8000 (kbps) ``` Three things to look for in this output. `Tunnel0 is up, line protocol is up` means the configured tunnel source IP exists locally and the underlay route to the destination is in the routing table. `Tunnel transport MTU 1476 bytes` is IOS XE doing the math: 1500 - 24 = 1476\. `Keepalive not set` tells you the tunnel will stay up even if the remote end becomes unreachable; we will fix that in step 6. If line protocol is down, run `show ip route ` to confirm there is a route to the far end's underlay IP. If there is no route, the tunnel will not come up. ## Step 4: Ping Across the Overlay ``` R1# ping 10.0.0.2 Type escape sequence to abort. Sending 5, 100-byte ICMP Echos to 10.0.0.2, timeout is 2 seconds: !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 4/5/8 ms ``` If this fails but the underlay ping in Step 1 worked, the most common cause is a stateful firewall in the underlay path that does not have an explicit allow rule for IP protocol 47 (GRE). The firewall sees a packet that is neither TCP, UDP, nor ICMP and drops it by default. Add an explicit permit, retest, and the tunnel ping will work. The other common cause is a typo in `tunnel destination` on either end, especially when copying configs. Double-check the IPs match what you expect with `show interface tunnel0 | include source|destination`. ## Step 5: Make the LAN Networks Reachable Through the Tunnel The tunnel works overlay-to-overlay (10.0.0.1 to 10.0.0.2), but neither router knows how to reach the other's LAN (192.168.x.0/24). Two options: static routes, or a routing protocol over the tunnel. For a two-site lab, static routes are simplest. ``` R1(config)# ip route 192.168.2.0 255.255.255.0 10.0.0.2 R2(config)# ip route 192.168.1.0 255.255.255.0 10.0.0.1 ``` Verify with a sourced ping that simulates LAN traffic: ``` R1# ping 192.168.2.1 source 192.168.1.1 Type escape sequence to abort. Sending 5, 100-byte ICMP Echos to 192.168.2.1, timeout is 2 seconds: Packet sent with a source address of 192.168.1.1 !!!!! Success rate is 100 percent (5/5), round-trip min/avg/max = 4/5/8 ms ``` The packet traverses R1's LAN interface, gets routed into Tunnel0 by the static route, picks up the GRE+outer-IP encapsulation, traverses the underlay as a unicast IP packet, arrives at R2, gets decapsulated, and is routed out R2's LAN interface to its destination. That whole pipeline runs with no per-packet state on either router beyond the routing table. ## Step 6: Add Keepalives Without keepalives, the tunnel interface stays up as long as the underlay route to the tunnel destination exists. If R2 crashes, the local end has no idea and traffic black-holes for as long as it takes the underlay routing protocol to converge (which may be never if the route is static). ``` R1(config)# interface Tunnel0 R1(config-if)# keepalive 10 3 R2(config)# interface Tunnel0 R2(config-if)# keepalive 10 3 ``` `keepalive 10 3` means send a keepalive every 10 seconds and declare the tunnel down after 3 missed responses. Verify: ``` R1# show interface Tunnel0 | include Keepalive Keepalive set (10 sec), retries 3 ``` Keepalives are a Cisco extension to GRE that work by sending a small GRE-encapsulated ICMP-like packet whose inner header tells the remote router to send it back. If you are tunneling between Cisco and a non-Cisco device, do not assume keepalives interoperate; test before trusting them. The full mechanism is at [GRE Tunnel Keepalives Explained](https://www.pinglabz.com/gre-tunnel-keepalives/). ## Step 7: Tune MTU for Production Traffic The default IP MTU on a Tunnel interface is the underlay MTU minus the GRE+IP overhead (1476 for a 1500-byte underlay). That works for plain GRE but does not leave room if you later add IPsec encryption, and it does not solve the universal problem of TCP endpoints negotiating a segment size that is too large. ![GRE encapsulation overhead: a 1480-byte original packet becomes a 1500-byte on-the-wire packet after a 24-byte outer IP plus GRE header is added, leaving 1476 bytes of tunnel transport MTU](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/encap-1.png) Figure 2\. Where the 1476-byte tunnel MTU comes from: 1500-byte underlay minus 24 bytes of GRE+IP overhead. Production drops it further to 1400 to leave headroom for IPsec. The two-line fix: ``` R1(config)# interface Tunnel0 R1(config-if)# ip mtu 1400 R1(config-if)# ip tcp adjust-mss 1360 R2(config)# interface Tunnel0 R2(config-if)# ip mtu 1400 R2(config-if)# ip tcp adjust-mss 1360 ``` `ip mtu 1400` tells the router to fragment IP packets larger than 1400 bytes before they hit the tunnel encapsulation, which leaves 100 bytes of headroom for IPsec or any other future overhead. `ip tcp adjust-mss 1360` rewrites the MSS value in TCP SYN packets passing through the tunnel so endpoints negotiate a segment size that already fits, eliminating fragmentation rather than just handling it. The math, the symptoms of MTU misconfiguration, and the deeper debugging tools are at [GRE MTU and Fragmentation: Fixing Tunnel Packet Loss](https://www.pinglabz.com/gre-tunnel-mtu/). ## Step 8: Run OSPF Over the Tunnel Static routes work for two sites. For three or more, swap to a routing protocol so new networks are learned dynamically. OSPF is the most common choice over GRE because it needs multicast, which GRE happily carries. ``` R1(config)# router ospf 1 R1(config-router)# network 10.0.0.0 0.0.0.3 area 0 R1(config-router)# network 192.168.1.0 0.0.0.255 area 0 R1(config-router)# passive-interface GigabitEthernet2 R2(config)# router ospf 1 R2(config-router)# network 10.0.0.0 0.0.0.3 area 0 R2(config-router)# network 192.168.2.0 0.0.0.255 area 0 R2(config-router)# passive-interface GigabitEthernet2 ``` Verify the neighbor: ``` R1# show ip ospf neighbor Neighbor ID Pri State Dead Time Address Interface 10.0.0.2 1 FULL/DROTHER 00:00:39 10.0.0.2 Tunnel0 ``` Once the neighbor is FULL, R1 has learned 192.168.2.0/24 dynamically through OSPF and you can remove the static routes from Step 5\. Multiple sites can now be added by configuring more tunnels to a hub and putting them all in OSPF area 0. One thing to watch: do not let OSPF (or any routing protocol you run inside the tunnel) advertise a route to the tunnel destination IP. If R1 learns the route to 203.0.113.1 through Tunnel0 itself, the tunnel goes recursive and flaps. Use a static underlay route to the tunnel destination, or filter the underlay IP out of the tunnel-side routing protocol with a distribute-list. The full coverage of OSPF, EIGRP, and BGP over GRE (including the recursive routing fix) is at [Routing Protocols Over GRE: OSPF, EIGRP, BGP](https://www.pinglabz.com/routing-protocols-over-gre/). ![Recursive routing on the left: Tunnel0 learns 203.0.113.1/32 via 10.0.0.2 through itself, encapsulating into itself. Pinned on the right: a static route forces 203.0.113.1/32 via 198.51.100.2 on the underlay GigabitEthernet1.](https://storage.ghost.io/c/ff/28/ff28db4a-11e1-4835-b928-158974cf96c1/content/images/2026/05/recursive-1.png) Figure 3\. The recursive-routing trap and the one-line fix. Left: the tunnel-side routing protocol has installed a route to the tunnel destination through the tunnel itself. Right: a static route at AD 1 forces the underlay path. ## Final Production-Style Config Putting it all together, the production-style R1 config (with keepalives, MTU clamping, and OSPF): ``` interface Tunnel0 ip address 10.0.0.1 255.255.255.252 ip mtu 1400 ip tcp adjust-mss 1360 tunnel source 198.51.100.1 tunnel destination 203.0.113.1 tunnel mode gre ip keepalive 10 3 ! router ospf 1 network 10.0.0.0 0.0.0.3 area 0 network 192.168.1.0 0.0.0.255 area 0 passive-interface GigabitEthernet2 ! ip route 203.0.113.1 255.255.255.255 198.51.100.2 ``` That last static route is the recursive-routing safety net: it tells R1 to reach R2's tunnel destination IP via the underlay next-hop, never via the tunnel itself. ## If the Tunnel Does Not Come Up The five most common failure modes and how to spot them: Tunnel0 line protocol down Likely cause No underlay route to tunnel destination Where to look `show ip route ` Tunnel up, ping across overlay fails Likely cause Firewall blocking IP protocol 47 Where to look Underlay firewall ACLs / packet captures Tunnel flaps every 30 seconds Likely causeRecursive routing Where to look `show ip route ` shows Tunnel0 as the next-hop Small pings work, large pings fail Likely cause MTU mismatch / fragmentation drop Where to look Try `ping 192.168.2.1 size 1400 df-bit`; lower MTU until it succeeds Tunnel up, OSPF neighbor stuck in INIT Likely cause OSPF hellos one-way; ACL or asymmetric tunnel Where to look `debug ip ospf hello` on both ends For a deeper dive into each of these failure modes with packet captures and debug command output, see [GRE Tunnel Troubleshooting Guide](https://www.pinglabz.com/gre-tunnel-troubleshooting/). ## Summary A working GRE tunnel on Cisco IOS XE is eight lines of config: a Tunnel interface with an IP, a source, a destination, a mode, and a keepalive, plus MTU clamping and either static routes or a routing protocol on top. Every step in this lab maps to a piece of that config and a verification command that proves it worked. The trick is not in the syntax; the trick is in remembering which knobs to set in production (keepalives, MTU, recursive-route protection) and what each one prevents. If you are now ready to add encryption, see [GRE over IPsec](https://www.pinglabz.com/gre-over-ipsec/). If you want to scale this from two sites to many, see [mGRE and DMVPN Introduction](https://www.pinglabz.com/mgre-dmvpn-introduction/). Either way, you have the plumbing: the rest is policy. ### IPv6 Header Format Explained URL: https://www.pinglabz.com/ipv6-header-explained/ Last updated: 2026-06-13T20:08:35.000Z The IPv6 header is 40 bytes, fixed length, with eight fields. Compared to IPv4's variable-length 20-60 byte header with thirteen fields, IPv6 is intentionally minimalist. The simplification is not cosmetic - it makes IPv6 packets faster to parse, easier to hardware-accelerate, and cleaner for routers to forward. This article walks through the IPv6 header byte by byte, explains each field, covers extension headers (the IPv6 way of adding optional features), and shows how the header looks in Wireshark. If you are studying for CCNP/CCIE or trying to read an IPv6 packet capture, this is the byte reference. ## The IPv6 Header Layout ``` 0 1 2 3 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+ |Version| Traffic Class | Flow Label | +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+ | Payload Length | Next Header | Hop Limit | +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+ | | + + | | + Source Address + | (128 bits / 16 bytes) | + + | | +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+ | | + + | | + Destination Address + | (128 bits / 16 bytes) | + + | | +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+ ``` 40 bytes total: 8 bytes of header fields plus 16 bytes source address plus 16 bytes destination address. ## Field by Field Version Bits4 PurposeAlways 6 for IPv6 Traffic Class Bits8 Purpose QoS marking (DSCP + ECN, like IPv4 ToS) Flow Label Bits20 Purpose Per-flow identifier (flows = packets that should follow the same path) Payload Length Bits16 Purpose Length of payload in bytes (excludes the 40-byte header itself) Next Header Bits8 Purpose Type of the next header (TCP=6, UDP=17, ICMPv6=58, extension headers, etc.) Hop Limit Bits8 Purpose TTL equivalent; decremented at each router; packet dropped at 0 Source Address Bits128 PurposeIPv6 source address Destination Address Bits128 Purpose IPv6 destination address ## Version 4 bits. Value is always 6 for IPv6 packets. The first hex digit of the header is always `6` (followed by the upper nibble of the Traffic Class). ## Traffic Class 8 bits. Functions identically to the IPv4 ToS / DSCP byte. Top 6 bits are DSCP (QoS marking, see [DSCP article](https://www.pinglabz.com/dscp-ip-precedence-explained/)); bottom 2 bits are ECN (Explicit Congestion Notification). QoS configurations on Cisco apply identically to IPv4 and IPv6 - a class-map matching DSCP EF works on both because the byte position differs but the semantics are the same. ## Flow Label 20 bits. Identifies a "flow" of packets that should follow the same path through the network. The intent: routers can use the flow label as a hash input for ECMP without parsing into Layer 4 headers. Identical Flow Label = same flow, deterministic path. In practice, Flow Label adoption has been mixed. Some implementations set it; many leave it 0\. Routers may or may not honor it for ECMP. RFC 6437 sets the modern usage rules but real-world behavior varies. Treat it as "useful when the surrounding stack honors it; ignored otherwise." ## Payload Length 16 bits. Length in bytes of everything after the 40-byte header. Maximum value 65,535 = 64 KB minus 1; for jumbograms larger than 64 KB, the Hop-by-Hop extension header carries an extended length field. Note this is "payload" length, not "total" length (as in IPv4). The 40-byte header itself is excluded. So an IPv6 packet on the wire is always 40 + Payload Length bytes total. ## Next Header 8 bits. Identifies the type of the next thing in the packet. This is roughly equivalent to IPv4's Protocol field but with two important differences: 1. It can point to a Layer 4 protocol directly (TCP=6, UDP=17, ICMPv6=58) 2. It can point to an Extension Header that itself contains another Next Header field This creates a chained-header model. Common Next Header values: 0 Hop-by-Hop Options (extension header) 6 TCP 17 UDP 43 Routing (extension header, including Segment Routing) 44 Fragment (extension header) 50 ESP (IPsec encapsulating security payload) 51 AH (IPsec authentication header) 58 ICMPv6 59 No Next Header (end of chain) 60 Destination Options (extension header) ## Hop Limit 8 bits. Decremented at each router that forwards the packet. Packet is dropped when Hop Limit reaches 0; ICMPv6 Time Exceeded message is sent back to the source. Equivalent to IPv4 TTL. Default starting values: 255 from most modern hosts (some use 64). The high default is intentional - by the time the packet is dropped due to Hop Limit, it has likely traversed many networks (which is itself diagnostic). The Hop Limit also serves a security role: NDP messages (RS, RA, NS, NA, Redirect) must have Hop Limit 255 when received, otherwise they are dropped. This prevents off-link attackers from injecting NDP messages because their packets would have decremented Hop Limit while transiting routers. ## Source and Destination Addresses 128 bits each. The IPv6 addresses (see [IPv6 Address Format](https://www.pinglabz.com/ipv6-address-format/) for the format details). 16 bytes each = 32 bytes of address space, which is most of the 40-byte header. The minimalism of the rest of the header is partly because the addresses themselves consume so much. ## Extension Headers IPv6 moves features that were "options" in IPv4 into separate extension headers chained between the base header and the payload. Each extension header has its own Next Header field pointing to whatever comes next. Hop-by-Hop Options NH Value0 Use Examined by every router along the path (e.g. Jumbograms, Router Alert) Routing NH Value43 Use Specifies intermediate destinations (Type 0 deprecated; Type 4 = Segment Routing) Fragment NH Value44 Use Carries fragmentation info if the source had to fragment ESP NH Value50 UseIPsec encryption AH NH Value51 UseIPsec authentication Destination Options NH Value60 Use Examined only by the destination The chain terminates when a Next Header value points to a real Layer 4 protocol (TCP, UDP, ICMPv6) or to "No Next Header" (59). Most ordinary IPv6 packets have no extension headers - the base header points directly to TCP/UDP/ICMPv6\. Extension headers are common for fragments, IPsec, and segment routing. ## Comparison to the IPv4 Header Version Same field, value 6 IHL (Internet Header Length) Removed (header is fixed 40 bytes) ToS / DSCP Renamed Traffic Class; same semantics Total Length Renamed Payload Length (excludes header) Identification, Flags, Fragment Offset Moved to Fragment extension header TTL Renamed Hop Limit Protocol Renamed Next Header (with extension header chain) Header Checksum Removed (relies on L2 + L4 checksums) Source / Destination Address Same purpose; 128 bits instead of 32 Options Replaced by extension headers (chained, optional) The biggest change is removing the header checksum. IPv4 routers had to recalculate the checksum each time they decremented TTL - an expensive operation that ASIC designers worked hard to optimize. IPv6 routers do not. The Layer 2 (Ethernet) FCS and Layer 4 (TCP/UDP) checksums together catch errors; the IP-layer checksum was redundant. ## In a Packet Capture An IPv6 packet in Wireshark shows the base header parsed: ``` Internet Protocol Version 6, Src: 2001:db8:1::1, Dst: 2001:db8:2::1 0110 .... = Version: 6 .... 0000 0000 .... .... .... .... .... = Traffic Class: 0x00 .... 0000 00.. .... .... .... .... .... = DSCP: Default (0) .... .... ..00 .... .... .... .... .... = ECN: Not-ECT (0) .... .... .... 0000 0000 0000 0000 0000 = Flow Label: 0x00000 Payload Length: 40 Next Header: TCP (6) Hop Limit: 64 Source: 2001:db8:1::1 Destination: 2001:db8:2::1 Transmission Control Protocol, ... ``` If extension headers are present, Wireshark dissects them in order, showing each in a separate tree section. The Next Header field of the last extension header points to the actual Layer 4 protocol, which Wireshark then parses. ## Why the Header Is Faster Three reasons IPv6 forwarding is faster than IPv4 on similar hardware: 1. **Fixed header length.** Routers know exactly where each field is; no IHL math. 2. **No checksum recomputation.** Decrementing Hop Limit does not require updating the header. 3. **No options parsing on the fast path.** Extension headers exist but are usually skipped on intermediate routers (only the destination examines Destination Options; only routers along the path examine Hop-by-Hop). Combined with hardware acceleration in modern silicon, IPv6 typically forwards at line rate on any router that does IPv4 at line rate. ## Summary The IPv6 header is 40 bytes fixed: 4-bit Version, 8-bit Traffic Class, 20-bit Flow Label, 16-bit Payload Length, 8-bit Next Header, 8-bit Hop Limit, plus 128-bit source and destination addresses. Extension headers chain between the base header and the payload via Next Header values; common ones include Fragment, Routing, ESP, AH, and Destination Options. The header is intentionally minimalist - simpler than IPv4, faster to parse, easier to hardware-accelerate. Master the field layout and the Next Header chain and you can read any IPv6 packet capture. Bookmark this article alongside the [IPv6 cluster pillar](https://www.pinglabz.com/ipv6/). ### IPv6 Configuration on Cisco IOS XE URL: https://www.pinglabz.com/ipv6-cisco-configuration/ Last updated: 2026-05-29T23:40:55.000Z Configuring IPv6 on Cisco IOS XE is mostly about turning things on. The protocol is built into modern Cisco code; you enable IPv6 routing globally, assign addresses to interfaces, and Neighbor Discovery + SLAAC handle most of the host-side work automatically. The configurations look familiar to anyone who has done IPv4 on Cisco - the syntax is parallel, just with `ipv6` instead of `ip`. This article walks through the configuration patterns: enabling IPv6, assigning static and SLAAC addresses, Router Advertisement options, IPv6 routing protocols, and the verification commands. If you are configuring IPv6 for the first time or auditing an existing deployment, this is the operator's walkthrough. ## Enabling IPv6 Globally By default, Cisco IOS XE has IPv6 disabled at the routing layer. Enable it: ``` Router(config)# ipv6 unicast-routing ``` That single command turns the router into an IPv6 router - it will forward IPv6 traffic between interfaces. Without it, the router can have IPv6 addresses on its interfaces but won't route between them. For multicast forwarding (used by some applications): ``` Router(config)# ipv6 multicast-routing ``` ## Per-Interface Configuration ### Static Address ``` interface GigabitEthernet0/0/0 ipv6 address 2001:db8:1::1/64 ipv6 enable ! Enable link-local even without global config (often automatic) ``` The static form is the most explicit. `ipv6 enable` is sometimes redundant when you have a global address (the link-local is created automatically when the interface comes up with any IPv6 config) but is good practice for clarity. ### EUI-64 Auto-derived Interface ID ``` interface GigabitEthernet0/0/0 ipv6 address 2001:db8:1::/64 eui-64 ``` This makes the router derive its interface ID from its MAC address using Modified EUI-64\. The result is something like 2001:db8:1:0:21a:2bff:fe3c:4d5e. Useful for routers in dynamic environments; less common than static addresses for production. ### SLAAC (host-style autoconfiguration) ``` interface GigabitEthernet0/0/0 ipv6 address autoconfig ``` Less common on routers (which usually have static addresses) but useful for simulating a host. The router waits for a Router Advertisement from another router on the segment and configures its address based on the received prefix. ### Link-Local Only ``` interface GigabitEthernet0/0/1 ipv6 enable ``` The interface gets a link-local FE80:: address (auto-derived) but no global address. Useful for point-to-point links where the IGP runs over link-locals only - the global address space is preserved. ## Router Advertisement Options By default, IPv6-enabled router interfaces send Router Advertisements (RAs) every 200 seconds (with random jitter). Hosts on the segment use these RAs to learn the prefix, default gateway, and DNS info. Tune RA timers: ``` interface GigabitEthernet0/0/0 ipv6 nd ra interval 200 60 ! Max 200s, min 60s between RAs ipv6 nd ra lifetime 1800 ! Default router lifetime in RA ``` Suppress RAs (for interfaces where they should not be sent, e.g. point-to-point links between routers): ``` interface GigabitEthernet0/0/0 ipv6 nd ra suppress all ``` Tell hosts to use DHCPv6 for addresses (M flag) or other config like DNS (O flag): ``` ipv6 nd managed-config-flag ! M flag: hosts should use DHCPv6 for address ipv6 nd other-config-flag ! O flag: hosts should use DHCPv6 for other info (DNS, etc.) ``` Send DNS server info in RAs (RDNSS option, RFC 6106): ``` interface GigabitEthernet0/0/0 ipv6 nd ra dns server 2001:4860:4860::8888 1800 ``` This eliminates the need for DHCPv6 just for DNS - the RA carries it. Modern hosts (Windows 10+, Linux, macOS) all support RDNSS. ## Static Routes ``` ! Default route via next-hop ipv6 route ::/0 2001:db8:1::1 ! Specific prefix ipv6 route 2001:db8:2::/48 2001:db8:1::2 ! Floating static (higher AD = backup) ipv6 route ::/0 2001:db8:1::3 200 ``` The syntax mirrors IPv4 static routes (`ip route` becomes `ipv6 route`) but uses IPv6 addresses and prefix lengths. ## OSPFv3 for IPv6 OSPFv3 is the IPv6-capable version of OSPF. Modern code supports an "address-family" mode that runs both IPv4 and IPv6 in one process; older code uses separate processes per AF. Modern OSPFv3 (address-family): ``` router ospfv3 1 router-id 1.1.1.1 address-family ipv6 unicast passive-interface default no passive-interface GigabitEthernet0/0/0 exit-address-family interface GigabitEthernet0/0/0 ipv6 address 2001:db8:1::1/64 ospfv3 1 ipv6 area 0 ``` Three things to notice. First, `router ospfv3` (not `router ospf`) - separate command. Second, OSPFv3 router-id is still 32 bits (configured as a dotted-decimal IPv4-style number); it does not have to match an actual IPv4 address. Third, the `ospfv3 1 ipv6 area 0` on the interface activates OSPFv3 for the IPv6 address-family. Verify with `show ipv6 ospf neighbor`. ## BGP for IPv6 ``` router bgp 65001 bgp router-id 1.1.1.1 no bgp default ipv4-unicast neighbor 2001:db8:12::2 remote-as 65002 address-family ipv6 unicast neighbor 2001:db8:12::2 activate network 2001:db8:1::/48 exit-address-family ``` BGP for IPv6 uses the same `router bgp` process as IPv4\. The address-family ipv6 unicast block configures IPv6-specific behavior. Update sources and route maps work the same way as IPv4. For full BGP coverage including MP-BGP, see the [BGP cluster pillar](https://www.pinglabz.com/bgp/) and [MP-BGP article](https://www.pinglabz.com/mp-bgp-multiprotocol-bgp/). ## EIGRP for IPv6 Named-mode EIGRP supports IPv6 cleanly: ``` router eigrp PROD address-family ipv6 unicast autonomous-system 100 af-interface default passive-interface af-interface GigabitEthernet0/0/0 no passive-interface topology base eigrp router-id 1.1.1.1 exit-af-topology exit-address-family ``` EIGRP for IPv6 uses link-local addresses for adjacency formation (like OSPFv3). The router-id is still 32 bits. Otherwise it works exactly like IPv4 EIGRP. See the [EIGRP cluster pillar](https://www.pinglabz.com/eigrp/). ## VRRPv3 for IPv6 ``` fhrp version vrrp v3 interface GigabitEthernet0/0/1 ipv6 address 2001:db8:1::2/64 vrrp 1 address-family ipv6 address FE80::1 primary ! Link-local virtual address address 2001:db8:1::1 priority 110 preempt ``` VRRPv3 for IPv6 uses a link-local address as the virtual gateway (hosts learn it via Router Advertisement). The configuration mirrors IPv4 VRRPv3 with an additional ipv6 address-family block. See the [VRRP article](https://www.pinglabz.com/vrrp-explained/). ## Dual-Stack Configuration Most production deployments run IPv4 and IPv6 simultaneously on the same interfaces: ``` interface GigabitEthernet0/0/0 description WAN ip address 192.168.1.1 255.255.255.0 ipv6 address 2001:db8:1::1/64 router ospf 1 network 192.168.1.0 0.0.0.255 area 0 router ospfv3 1 address-family ipv6 unicast exit-address-family interface GigabitEthernet0/0/0 ip ospf 1 area 0 ospfv3 1 ipv6 area 0 ``` Both protocols run independently on the same interface. Each has its own routing table; `show ip route` and `show ipv6 route` are separate. ## Security Configurations IPv6 ACLs: ``` ipv6 access-list ALLOW-MGMT permit tcp 2001:db8:0:1::/64 any eq 22 permit ipv6 any any log interface GigabitEthernet0/0/0 ipv6 traffic-filter ALLOW-MGMT in ``` IPv6 first-hop security (DHCPv6 Guard, RA Guard, ND Inspection): ``` ! Drop rogue RAs from non-router interfaces interface GigabitEthernet1/0/3 ipv6 nd raguard ! Inspect ND messages ipv6 nd inspection ! DHCPv6 Guard ipv6 dhcp guard policy CLIENT device-role client interface GigabitEthernet1/0/3 ipv6 dhcp guard attach-policy CLIENT ``` These are the IPv6 equivalents of DHCP snooping, Dynamic ARP Inspection, and IP Source Guard for IPv4. ## Verification ``` ! IPv6 routing table Router# show ipv6 route ! IPv6 interface details Router# show ipv6 interface Router# show ipv6 interface brief ! Neighbor table (ARP equivalent) Router# show ipv6 neighbors ! Routing protocol state Router# show ipv6 ospf neighbor Router# show bgp ipv6 unicast summary Router# show eigrp address-family ipv6 neighbors ! Connectivity Router# ping ipv6 2001:db8:1::1 Router# traceroute ipv6 2001:db8:1::1 ``` ## Anti-Patterns - **Forgetting `ipv6 unicast-routing`.** The router has IPv6 addresses but does not forward; baffled operators see traffic die at the router. - **Not suppressing RAs on point-to-point router-to-router links.** RAs leak between routers and create confusing output. - **Disabling IPv6 on host stacks.** Microsoft and Apple specifically advise against this; modern apps depend on IPv6. - **Using documentation prefix 2001:db8::/32 in production.** Routable but reserved for examples; will never reach the actual internet correctly. - **Forgetting first-hop security.** Rogue RAs and DHCPv6 servers are real threats; deploy RA Guard and DHCPv6 Guard. ## Summary Cisco IOS XE IPv6 configuration is parallel to IPv4 with a few specific commands. Enable IPv6 globally (`ipv6 unicast-routing`), assign addresses on interfaces, configure Router Advertisements with the appropriate options, and run routing protocols in their IPv6 forms (OSPFv3, BGP for IPv6, EIGRP for IPv6, VRRPv3). Dual-stack is the dominant deployment model. The configuration is straightforward; the operational habits are different and need practice. Bookmark this article alongside the [IPv6 cluster pillar](https://www.pinglabz.com/ipv6/), the [address format article](https://www.pinglabz.com/ipv6-address-format/), and the [address types article](https://www.pinglabz.com/ipv6-address-types/). ### IPv6 Address Types: Link-Local, Global, Unique Local URL: https://www.pinglabz.com/ipv6-address-types/ Last updated: 2026-06-13T20:08:35.000Z IPv6 has more address types than IPv4\. Where IPv4 has unicast, multicast, and broadcast (gone in IPv6), IPv6 has global unicast, link-local, unique local, multicast, anycast, and a few special addresses. Each has a specific scope and purpose, and getting the types right is foundational to IPv6 design. This article walks through each address type, when to use which, and the operational implications. If you are designing IPv6 addressing or trying to figure out why `show ipv6 interface` shows so many addresses on every interface, this is the reference. ## The Six Address Types 2000::/3 TypeGlobal Unicast ScopeGlobally routable Routable?Yes (internet) FC00::/7 (FD00::/8 in practice) TypeUnique Local (ULA) ScopeSite-local Routable?Yes (within site) FE80::/10 TypeLink-Local ScopeLink-only Routable? No (never crosses a router) FF00::/8 TypeMulticast ScopeGroup communication Routable? Yes (depending on scope) Same range as global unicast TypeAnycast ScopeGlobally routable Routable? Yes; multiple hosts share the address ::/128, ::1/128, etc. TypeSpecial ScopeVarious Routable?Various ## Global Unicast Global unicast addresses are the IPv6 equivalent of IPv4 public addresses. They are globally routable and globally unique. The currently allocated range is 2000::/3 (any address starting with 2 or 3 in the first hex digit). The structure of an enterprise global unicast address: ``` +- 48 bits global prefix -+- 16 bits subnet -+- 64 bits interface ID -+ | 2001:db8:1234: | abcd: | xxxx:xxxx:xxxx:xxxx | +-------------------------+------------------+-------------------------+ ``` The 48-bit prefix is what your ISP allocates to you. The 16-bit subnet field gives you 65,536 subnets within that prefix, plenty for any enterprise. The 64-bit interface ID identifies the host on the subnet. Examples: 2001:db8:abcd:1::1 Global unicast (in the documentation/example range) 2607:f8b0:4006:80f::200e Global unicast (Google) 2a00:1450:4001::1 Global unicast (Google EU) 2001:db8::/32 is reserved for documentation; never use this on real production networks. ## Link-Local Link-local addresses are the workhorse of IPv6 internal protocols. Every IPv6 interface has one, automatically assigned, and it is required for the interface to function. The prefix is FE80::/10 (in practice always FE80::/64 because the next 54 bits must be zero). Properties: - Always present on every IPv6 interface (required by the protocol) - Never crosses a router (link-scoped) - Used for Neighbor Discovery, Router Solicitation/Advertisement, OSPFv3 adjacencies, EIGRP for IPv6 sessions, VRRPv3 advertisements - Format: FE80:: followed by 64 bits of interface ID (often EUI-64 derived from MAC) - Same link-local can theoretically appear on multiple interfaces of the same router (you scope it with %interface-id syntax: `fe80::1%Gi0/0/0`) Verify with: ``` Router# show ipv6 interface GigabitEthernet0/0/0 GigabitEthernet0/0/0 is up, line protocol is up IPv6 is enabled, link-local address is FE80::21A:2BFF:FE3C:4D5E Global unicast address(es): 2001:DB8:1::1, subnet is 2001:DB8:1::/64 ``` Most routing protocols (OSPFv3, EIGRP for IPv6) form adjacencies over link-local addresses, not global unicast. The benefit: if you renumber the global prefix, the IGP keeps running uninterrupted. ## Unique Local Addresses (ULAs) ULAs are the IPv6 equivalent of IPv4 RFC 1918 private addresses. RFC 4193 defines the format. The prefix is FC00::/7, but in practice you only use FD00::/8 (the L bit is set to 1 to indicate locally-assigned). The ULA structure includes a 40-bit Global ID that is meant to be pseudo-random: ``` +- FD -+- 40-bit Global ID --+- 16-bit subnet -+- 64-bit interface ID -+ | FD | xxxxxxxxxxxxxxxxxx | abcd | xxxx:xxxx:xxxx:xxxx | +------+---------------------+-----------------+------------------------+ ``` You generate a random Global ID once for your organization (using SHA-1 of the time and MAC, or just using a random generator) and use it across all your sites. The randomness reduces collision risk if two ULA-using networks ever merge. Use ULAs for: - Internal management traffic that should never leak to the internet - Communication between hosts that don't need internet reachability - Lab environments where global addresses don't make sense - IPv6 addressing for networks that have private IPv4 addressing today Modern IPv6 best practice often skips ULAs entirely - use global unicast everywhere with proper firewalling, since you have abundant addresses. ULAs are a tool for specific use cases, not the default. ## Multicast IPv6 has no broadcast. Functions that used to use broadcast (ARP, etc.) now use multicast. The prefix is FF00::/8. The multicast address format includes scope bits: ``` +- FF -+- Flags (4) -+- Scope (4) -+- Group ID (112) -+ | FF | flgs | scope | group identifier | +------+-------------+-------------+-------------------+ ``` Common multicast scopes: Interface-local (loopback) Scope1 ExampleFF01::1 Link-local Scope2 ExampleFF02::1 (all nodes) Admin-local Scope4 ExampleSite-defined Site-local Scope5 Example FF05::101 (NTP servers) Organization-local Scope8 ExampleLarger than site Global ScopeE Example FF0E:: (internet-scope) Well-known multicast addresses you see constantly: FF02::1 All nodes on link FF02::2 All routers on link FF02::5 OSPFv3 routers (replaces 224.0.0.5) FF02::6 OSPFv3 designated routers (replaces 224.0.0.6) FF02::A EIGRP for IPv6 FF02::12 VRRPv3 FF02::1:FFXX:XXXX Solicited-node multicast (used by NDP) The solicited-node multicast is interesting: each unicast address has a corresponding multicast address derived from the bottom 24 bits. NDP uses this to direct Neighbor Solicitations only to hosts whose interface ID matches, dramatically reducing broadcast-style flooding. ## Anycast Anycast is "one address, multiple hosts." Multiple hosts advertise the same address; the network delivers to whichever is topologically closest. It is technically just a unicast address shared by multiple hosts; nothing in the address format indicates anycast (it is up to the routing layer). Common uses: DNS root servers (root anycast), CDN edge servers, internal services like NTP across multiple data centers. The "subnet anycast" address (interface ID all zeros within a subnet) is reserved by RFC 4291. ## Special Addresses ::/128 Unspecified (used during DAD; no address yet) ::1/128 Loopback ::FFFF:0:0/96 IPv4-mapped (used by dual-stack sockets) 2001:db8::/32 Documentation prefix; never use in production 2002::/16 6to4 tunneling (largely obsolete) 3FFE::/16 Old 6bone experimental; deprecated ## Every Interface Has Multiple Addresses An IPv6-enabled interface typically has: 1. A link-local address (FE80::xxxx) 2. One or more global unicast addresses (2xxx:xxxx::xxxx) 3. Optionally a ULA (FDxx:xxxx::xxxx) 4. Multicast group memberships (FF02::1 always; FF02::5 if OSPFv3; etc.) 5. Solicited-node multicast for each unicast address This is normal IPv6 behavior. Cisco's `show ipv6 interface` displays them all. ## Design Implications Internet-facing service Global unicast Internal client / internal server traffic Global unicast (with firewall) or ULA Routing protocol next-hop Link-local (used automatically by OSPFv3, etc.) Documentation / examples Documentation prefix 2001:db8::/32 Lab environments ULA or documentation prefix Multi-site enterprise without internet need ULA Loopback interface ::1/128 globally; per-router /128 from your global allocation ## Summary IPv6 has six main address types: global unicast, link-local, ULA, multicast, anycast, and special addresses. Every interface gets at least a link-local automatically; you assign global unicast for internet reachability and ULAs for site-internal traffic. Multicast is built into the protocol and used pervasively (NDP, RAs, OSPFv3, etc.). Master the prefixes (2000::/3 for global, FE80::/10 for link-local, FC00::/7 for ULA, FF00::/8 for multicast) and the rest of IPv6 addressing follows. Bookmark this article alongside the [IPv6 cluster pillar](https://www.pinglabz.com/ipv6/). ### IPv4 vs IPv6: The Real Differences URL: https://www.pinglabz.com/ipv4-vs-ipv6/ Last updated: 2026-06-13T20:08:36.000Z IPv4 vs IPv6 is the comparison every engineer eventually has to explain to a stakeholder, justify in a design doc, or quiz themselves on for a certification. Most "vs" articles focus on address-space size and miss the architectural differences that matter operationally. This is the comparison that covers what actually changes when you go from IPv4 to IPv6. If you are studying for CCNP, planning a dual-stack rollout, or trying to defend why your network should care about IPv6 in 2026, this is the reference. ## The Side-by-Side Table Address size IPv432 bits IPv6128 bits Address space IPv44.3 billion IPv6 340 undecillion (3.4 x 10^38) Notation IPv4 Dotted decimal (192.168.1.1) IPv6 Colon-hex (2001:db8::1) Header length IPv420-60 bytes (variable) IPv640 bytes (fixed) Header checksum IPv4 Yes (IPv4 header has its own checksum) IPv6 No (relies on L2 + L4 checksums) Fragmentation IPv4Routers can fragment IPv6 Source-only; routers send PTB instead Address resolution IPv4ARP (broadcast) IPv6 Neighbor Discovery via ICMPv6 multicast Autoconfiguration IPv4DHCP only IPv6 SLAAC (stateless) + DHCPv6 (stateful) Multicast IPv4Optional, often unused IPv6 Built-in, used pervasively (NDP, RA, etc.) Broadcast IPv4Yes IPv6 No (replaced by all-nodes multicast FF02::1) NAT IPv4 Common (RFC 1918 private + NAT to public) IPv6 Generally unnecessary; ULA without NAT preferred Mandatory features IPv4None beyond IP itself IPv6 ICMPv6, NDP, Multicast support Optional headers IPv4 Options field in IP header IPv6 Extension headers (chained) QoS marking IPv4ToS/DSCP byte IPv6 Traffic Class byte (same DSCP semantics) Loopback IPv4127.0.0.0/8 (whole /8) IPv6 ::1/128 (single address) Link-local IPv4 169.254.0.0/16 (only on DHCP failure) IPv6 FE80::/10 (always present) Privacy IPv4None native IPv6 RFC 4941 Privacy Extensions standard Routing protocols IPv4 OSPFv2, EIGRP, BGP, RIPv2 IPv6 OSPFv3, EIGRP for IPv6, MP-BGP, RIPng FHRP IPv4HSRP, VRRP, GLBP IPv6 HSRPv2-IPv6, VRRPv3, GLBP for IPv6 ## Address Space The 32-bit IPv4 address space gives 4.3 billion addresses. Sounds large; isn't enough. With NAT, ISPs and enterprises stretch this further by reusing RFC 1918 private addresses internally, but NAT creates problems for end-to-end connectivity, peer-to-peer protocols, IPsec, and any application that wants to be a server. The 128-bit IPv6 address space provides 340 undecillion addresses. The number is so large it cannot be exhausted by any sensible allocation policy. Each /48 enterprise allocation is 65,536 /64 subnets - more than any single enterprise ever needs. Practical implication: you assign IPv6 addresses generously. Each VLAN gets a /64\. Each point-to-point link gets a /127\. Each loopback gets a /128\. NAT mostly disappears because there is no scarcity to NAT around. ## The Header Comparison IPv4 header (20 bytes minimum, up to 60 with options): ``` +-Version-+-IHL-+-ToS-+-Total Length-+ +-Identification-+-Flags-+-Fragment Offset-+ +-TTL-+-Protocol-+-Header Checksum-+ +-Source Address (32)-+ +-Destination Address (32)-+ +-Options (variable)-+ ``` IPv6 header (40 bytes, fixed): ``` +-Version-+-Traffic Class-+-Flow Label-+ +-Payload Length-+-Next Header-+-Hop Limit-+ +-Source Address (128)-+ +-Destination Address (128)-+ ``` What's gone in IPv6: header checksum (L2/L4 redundant), Identification/Flags/Fragment Offset (only in extension header if fragmenting), Header Length (fixed 40), Options (replaced by extension headers). The result: simpler header, faster to parse, easier to hardware-accelerate. The only downside is that IPv6 packets are slightly larger when the addresses dominate (40 + payload vs 20 + payload at minimum), but this is irrelevant for typical packet sizes. ## Autoconfiguration IPv4 autoconfiguration is DHCP. A host comes up with no address, broadcasts a DHCPDISCOVER, gets an offer, requests it, gets ACK. Without DHCP, the host stays unconfigured (or gets a 169.254 link-local address, which only enables local-segment communication). IPv6 has SLAAC (Stateless Address Autoconfiguration). A host comes up, auto-assigns a link-local FE80:: address, sends a Router Solicitation, receives a Router Advertisement with the network prefix, derives a global address, and runs Duplicate Address Detection. No DHCP server required. DHCPv6 still exists for stateful configuration where the operator wants to control which addresses go where. But for many networks, RA-only is sufficient, and the RA can carry DNS server information via the RDNSS option. ## ARP vs Neighbor Discovery IPv4's address resolution uses ARP. Broadcast a request: "Who has 192.168.1.1?"; the owner replies. Every host on the segment processes every ARP request because of broadcast. IPv6's address resolution uses ICMPv6 Neighbor Discovery. Multicast a Neighbor Solicitation to the solicited-node multicast address derived from the target. Only hosts whose interface ID matches listen on that group, so other hosts are not interrupted. NDP also handles router discovery, prefix discovery (RAs), redirect, and DAD - things that in IPv4 would require separate protocols (or just don't exist). The net effect: IPv6's host-bring-up is more elegant and less broadcast-heavy. ## NAT and the End-to-End Principle IPv4 NAT was a workaround for address exhaustion that became architectural orthodoxy. Most enterprise networks use RFC 1918 private addresses internally and NAT at the edge for internet connectivity. NAT breaks the end-to-end principle: applications behind NAT cannot be servers without port forwarding; protocols that embed addresses (FTP active mode, SIP, IPsec without NAT-T) need ALGs to work. IPv6 has enough addresses that NAT is unnecessary. The end-to-end principle is restored. Every host can be a server. IPsec works without NAT-T workarounds. Peer-to-peer protocols are easier. NAT64 and NPTv6 do exist as IPv6-IPv4 transition mechanisms, but the philosophy is "don't NAT IPv6 to IPv6." Use Unique Local Addresses (ULAs) instead of private addresses + NAT for site-internal traffic. ## Dual-Stack: The Migration Reality Networks do not migrate from IPv4 to IPv6 overnight. The dominant transition strategy is dual-stack: hosts and routers run both protocols simultaneously. Dual-stack How it works Every host has IPv4 and IPv6; applications use whichever the destination supports When to use Standard transition strategy Tunneling (6in4, 6to4, ISATAP) How it works Encapsulate IPv6 in IPv4 to cross v4-only segments When to use Mostly historic; used when v6-native paths weren't available Translation (NAT64/DNS64) How it works Translate v6-only hosts to v4-only servers When to use Mobile networks (T-Mobile, etc.) where carriers run IPv6-only with translation for v4 destinations 464XLAT How it works Combination of CLAT (host-side translation) and NAT64 When to use Mobile and some IPv6-only enterprise networks Dual-stack is the easy choice for enterprise: keep IPv4, enable IPv6 alongside. Hosts use Happy Eyeballs (RFC 8305) to pick whichever transport works fastest. Most modern operating systems prefer IPv6 when both are available. ## When Each Matters **Why IPv6 matters in 2026:** - Mobile networks (T-Mobile USA, many Asian/European carriers) are IPv6-only for cellular data with NAT64 for v4 destinations - Major content (Google, Facebook, Netflix, AWS, Azure) supports IPv6 natively; an IPv6-only client gets the v6 path which may be faster - IoT devices number in billions; IPv6 is the only sane addressing for them - Some compliance frameworks (US federal, etc.) mandate IPv6 support - IPv4 address acquisition costs $40-60 per address as of 2026 in the secondary market **Why IPv4 still matters:** - Most legacy applications and infrastructure still speak IPv4 - Your ISP may not yet offer IPv6 service (rare but exists) - Internal-only deployments where IPv6 brings no benefit - Operational habits and tooling are IPv4-trained; retraining is real cost ## Dual-Stack Configuration on Cisco ``` ! Global enablement ip routing ! IPv4 (default on) ipv6 unicast-routing ! IPv6 (must enable) ! Interface dual-stack interface GigabitEthernet0/0/0 ip address 192.168.1.1 255.255.255.0 ipv6 address 2001:db8:1::1/64 ipv6 enable ! Enables link-local even if no global address ! IPv6 routing protocols configured separately or via address-families ``` Verify with: ``` ! IPv4 routing table Router# show ip route ! IPv6 routing table Router# show ipv6 route ! Per-interface IPv6 details Router# show ipv6 interface GigabitEthernet0/0/0 ``` ## Anti-Patterns - **Disabling IPv6 on hosts.** Microsoft, Apple, and the IETF all warn against this. Modern apps depend on IPv6\. Disable routing if you do not want IPv6 on your network, but leave the host stack enabled. - **Using IPv6 NAT.** The whole point of IPv6 is enough addresses to avoid NAT. Use ULAs for site-internal addressing without NAT. - **Ignoring IPv6 in security policy.** An IPv6-capable host on an IPv4-firewalled network often has a parallel IPv6 path that bypasses security controls. Audit and enforce IPv6 policies if you have any IPv6 in the network. - **Treating IPv6 as "just bigger IPv4."** The address resolution model, fragmentation behavior, and autoconfiguration are different. Lab the differences. ## Summary IPv6 is not just bigger IPv4\. The header is simpler, ARP is replaced by Neighbor Discovery, autoconfiguration is built-in via SLAAC, NAT is unnecessary, and the end-to-end principle is restored. The 128-bit address space removes the scarcity that drove IPv4 design compromises. For new deployments, dual-stack is the default. IPv4-only is a viable choice for narrowly-scoped legacy environments but increasingly limits what your network can do. Bookmark this article alongside the [IPv6 cluster pillar](https://www.pinglabz.com/ipv6/). ### IPv6 Address Format Explained Byte by Byte URL: https://www.pinglabz.com/ipv6-address-format/ Last updated: 2026-06-13T20:08:36.000Z IPv6 addresses are 128 bits, written as eight groups of four hexadecimal digits separated by colons. They look intimidating at first - `2001:0db8:85a3:0000:0000:8a2e:0370:7334` \- but the format is simpler than IPv4 once you understand the compression rules and the conventional /64 split. This article walks through the byte-level format, the two compression rules, the prefix length convention, the EUI-64 interface ID derivation, and the worked examples that make IPv6 addresses readable instead of intimidating. If you are studying for CCNP/CCIE, configuring IPv6 for the first time, or trying to read a Cisco router's IPv6 routing table, this is the byte reference. ## The 128-bit Format An IPv6 address is 128 bits long. Written in hexadecimal, that is 32 hex digits. Convention groups them into eight blocks of four hex digits separated by colons: ``` 2001:0db8:85a3:0000:0000:8a2e:0370:7334 ``` Each block represents 16 bits (4 hex digits = 4 \* 4 = 16 bits). Eight blocks = 8 \* 16 = 128 bits total. The notation is case-insensitive (uppercase or lowercase hex both valid; lowercase is conventional). ## Compression Rule 1: Drop Leading Zeros Within each 16-bit block, leading zeros may be dropped: ``` 2001:0db8:85a3:0000:0000:8a2e:0370:7334 becomes 2001:db8:85a3:0:0:8a2e:370:7334 ``` The `0db8` becomes `db8`; `0000` becomes `0`; `0370` becomes `370`. Each block must contain at least one hex digit (a single zero, not nothing). ## Compression Rule 2: The Double Colon One run of consecutive all-zero blocks may be replaced with `::`: ``` 2001:db8:85a3:0:0:8a2e:370:7334 becomes 2001:db8:85a3::8a2e:370:7334 ``` The two consecutive `:0:0:` become `::`. Critical: only one `::` per address. Two would be ambiguous - the parser cannot tell how many zero blocks each represents. Examples of the rule applied: 2001:0db8:0000:0000:0000:ff00:0042:8329 2001:db8::ff00:42:8329 0000:0000:0000:0000:0000:0000:0000:0001 ::1 (loopback) 0000:0000:0000:0000:0000:0000:0000:0000 :: (unspecified) fe80:0000:0000:0000:0204:61ff:fefe:5d04 fe80::204:61ff:fefe:5d04 2001:db8:0:1:0:0:0:1 2001:db8:0:1::1 (only the longer zero run is compressed) The last example shows why "longest run" matters. The address has two zero runs (one block at position 3 and three blocks at positions 5-7). Compressing the longer run is conventional and keeps the address most compact. ## Prefix Notation IPv6 uses CIDR-style prefix notation: address followed by /N where N is the number of network bits. ``` 2001:db8:1::1/64 ``` This means the first 64 bits are network; the last 64 bits are host. The /64 is the dominant convention for normal networks. Larger prefixes (smaller subnets like /127 for point-to-point links, /128 for loopbacks) and smaller prefixes (provider allocations like /48) exist but /64 is the default for end-host networks. The 64/64 split is what enables Neighbor Discovery efficiencies and SLAAC. Hosts can use the bottom 64 bits as an interface identifier without colliding (the address space is huge enough). ## Common Prefix Sizes /3 (2000::/3) All currently-allocated global unicast /12 (2A00::/12 etc.) RIR allocation (continent-level) /32 ISP allocation (typical) /48 Customer allocation (typical end-customer assignment) /56 Customer subnet (often residential) /64 End-network subnet (default for hosts) /127 Point-to-point link (RFC 6164 recommendation) /128 Single host (loopback addresses, etc.) An enterprise customer typically gets a /48 from their ISP; they then sub-divide into many /64 subnets. With /48 = 65,536 /64 subnets, this is enough for any enterprise. ## EUI-64 Interface IDs Hosts deriving their own interface ID for SLAAC originally used Modified EUI-64: take the MAC address (48 bits), insert FFFE in the middle to make 64 bits, and flip the U/L bit (bit 7 of the first byte). MAC address 00-1A-2B-3C-4D-5E Insert FFFE in the middle 00-1A-2B-FF-FE-3C-4D-5E Flip U/L bit (XOR 0x02 on first byte) 02-1A-2B-FF-FE-3C-4D-5E Final interface ID 021A:2BFF:FE3C:4D5E Combined with prefix 2001:db8:1::/64, the full address becomes `2001:db8:1:0:21a:2bff:fe3c:4d5e`. Modern hosts often use Privacy Extensions (RFC 4941) that generate random interface IDs every few hours instead of the deterministic EUI-64 ID. This prevents tracking based on the MAC-derived host part. Cisco routers still use EUI-64 by default for interface addresses; end hosts use random or stable-private interface IDs. ## Special and Reserved Addresses ::/128 Unspecified (during DAD; "no address yet") ::1/128 Loopback (equivalent to IPv4 127.0.0.1) ::FFFF:0:0/96 IPv4-mapped IPv6 (used internally for dual-stack sockets) FE80::/10 Link-local (always present on every interface) FC00::/7 Unique Local Addresses (RFC 4193) FF00::/8 Multicast 2000::/3 Global unicast (currently allocated) Two specific multicast addresses appear constantly: FF02::1 All nodes on link FF02::2 All routers on link FF02::5 OSPFv3 routers FF02::A EIGRP for IPv6 FF02::1:FFXX:XXXX Solicited-node multicast (used by NDP) The solicited-node multicast is interesting: instead of broadcasting (no broadcast in IPv6), an NDP Neighbor Solicitation is sent to a multicast group derived from the target address. Only hosts whose interface ID ends with the same XX:XXXX listen on that group, dramatically reducing wasted host wake-ups compared to ARP broadcast. ## Worked Examples Compress these: 2001:0db8:0000:0000:0000:0000:0000:0001 2001:db8::1 fe80:0000:0000:0000:0202:b3ff:fe1e:8329 fe80::202:b3ff:fe1e:8329 2001:0db8:0000:0042:0000:8a2e:0370:7334 2001:db8:0:42:0:8a2e:370:7334 0000:0000:0000:0000:0000:0000:0000:0000 :: Expand these (the inverse): 2001:db8::1234 2001:0db8:0000:0000:0000:0000:0000:1234 fe80::1 fe80:0000:0000:0000:0000:0000:0000:0001 ::ffff:192.168.1.1 0000:0000:0000:0000:0000:ffff:c0a8:0101 (IPv4-mapped) The IPv4-mapped form lets dual-stack sockets handle both v4 and v6 connections through a single API. ## Reading IPv6 Addresses Quickly Some patterns become recognizable with practice: Starts with 2 or 3 Global unicast (2000::/3) Starts with FE80 Link-local (always) Starts with FD or FC Unique Local (private-like) Starts with FF Multicast :: alone Unspecified or loopback (::1) Contains FFFE in the middle EUI-64 derived ## Summary IPv6 addresses are 128 bits, eight groups of four hex digits, with two compression rules: drop leading zeros within each block, and replace one run of zero blocks with `::`. The /64 prefix is the convention for normal networks. EUI-64 derives interface IDs from MAC addresses but Privacy Extensions are now common. Master the format and compression and the rest of IPv6 reads naturally. Bookmark this article alongside the [IPv6 cluster pillar](https://www.pinglabz.com/ipv6/) as your byte-level reference. ### HSRP vs VRRP vs GLBP: All Three Running Side by Side URL: https://www.pinglabz.com/hsrp-vs-vrrp-vs-glbp/ Last updated: 2026-08-01T19:26:55.000Z HSRP, VRRP and GLBP all solve the same problem, [giving a LAN a default gateway that survives losing a router](https://www.pinglabz.com/fhrp/), but they make different trade-offs. HSRP is Cisco's default, simple and widely deployed. VRRP is the open standard, vendor neutral, with a faster default failure detector. GLBP is Cisco's load-balancing variant, more complex, but it is the only one of the three that puts traffic on both routers at the same time for the same gateway address. Most comparisons of these three protocols are a table and nothing else. This one is built on a lab where all three run simultaneously on the same pair of routers: two `iol-xe` nodes in CML on IOS XE 17.18.2, three parallel segments, one FHRP each. HSRP group 10 on `Et0/0`, VRRP group 20 on `Et0/1`, GLBP group 30 on `Et0/2`. R1 carries priority 110 on all three, so R1 should win every election, and the interesting question is what R2 does while it is losing. The short answer, before any of the detail: under HSRP and VRRP, R2 sits there and listens. Under GLBP, R2 is an active forwarder for the same virtual IP R1 owns. That is the one difference that changes your capacity planning. Almost everything else is vocabulary, default timers and vendor politics. ## All Three Protocols, One Pair of Routers The configuration side is small enough to show in full. This is R1, the preferred router, with priority 110 on every group: ``` interface Et0/0 standby 10 ip 10.0.10.254 standby 10 priority 110 standby 10 preempt ! fhrp version vrrp v3 interface Et0/1 vrrp 20 address-family ipv4 address 10.0.20.254 priority 110 ! interface Et0/2 glbp 30 ip 10.0.30.254 glbp 30 priority 110 glbp 30 preempt ``` R2 is identical except that it leaves priority at the default 100 and does not set `preempt`. That asymmetry matters later, when R1 comes back from the failover test. Notice the VRRP stanza does not look like the VRRP you remember. The legacy `vrrp 20 ip 10.0.20.254` one-liner is gone on modern IOS XE: you enable `fhrp version vrrp v3` globally, then configure the group under an address family. If a decade-old VRRP config silently does nothing on a 17.x box, that is why. ## HSRP: One Active, One Standby Start with the protocol most networks already run. HSRP elects one Active router and one Standby router per group, and only the Active forwards. Here is the same command on both routers in steady state: ``` R1# show standby brief P indicates configured to preempt. | Interface Grp Pri P State Active Standby Virtual IP Et0/0 10 110 P Active local 10.0.10.2 10.0.10.254 R2# show standby brief Interface Grp Pri P State Active Standby Virtual IP Et0/0 10 100 Standby 10.0.10.1 local 10.0.10.254 ``` Read those two blocks together and nothing is left ambiguous. R1 says it is Active with `local` in the Active column and 10.0.10.2 as its Standby, R2 says the reverse, and both print the same virtual IP. The `P` flag appears on R1 only, because only R1 was configured to preempt. (If both routers print `Active` with `Standby unknown`, the hellos are not crossing and you have a different article to read.) The brief output does not show you the virtual MAC, which is the detail that actually identifies the protocol on the wire. For that you need `show standby` with no keyword: ``` Ethernet0/0 - Group 10 State is Active 2 state changes, last state change 00:01:04 Virtual IP address is 10.0.10.254 Active virtual MAC address is 0000.0c07.ac0a (MAC In Use) <-- 0000.0c07.acXX, XX = group Local virtual MAC address is 0000.0c07.ac0a (v1 default) Hello time 3 sec, hold time 10 sec Next hello sent in 1.712 secs Preemption enabled Active router is local Standby router is 10.0.10.2, priority 100 (expires in 10.672 sec) Priority 110 (configured 110) Group name is "hsrp-Et0/0-10" (default) ``` Group 10 is `0x0a`, and the virtual MAC ends in `ac0a`. The mapping is that direct. Note also the timers printed here, hello 3 seconds and hold 10 seconds, because those are the numbers VRRP is about to beat. If you want the full build rather than the verification, the walkthrough on [setting up an HSRP group from scratch](https://www.pinglabz.com/hsrp-configuration/) covers priority, preempt and tracking in order. ## VRRP: Master and Backup, and a Faster Default Clock VRRP does the same job with different words. There is no Active and no Standby. There is a MASTER and there are BACKUPs: ``` R1# show vrrp brief Interface Grp A-F Pri Time Own Pre State Master addr/Group addr Et0/1 20 IPv4 110 0 N Y MASTER 10.0.20.1(local) 10.0.20.254 R2# show vrrp brief Interface Grp A-F Pri Time Own Pre State Master addr/Group addr Et0/1 20 IPv4 100 3609 N Y BACKUP 10.0.20.1 10.0.20.254 ``` Three fields are worth stopping on. `A-F` is the address family, which only exists because this is VRRPv3\. `Own` is N on both routers, meaning neither owns 10.0.20.254 as a real interface address (if one did, it would run at priority 255 and the election would be over before it started). `Pre` is Y on both, and that is a default: VRRP preempts unless you turn it off, which is the opposite of HSRP and an easy way to surprise yourself during a migration. The `Time` column is the one worth internalising. It is the master down interval in milliseconds, and R2 prints 3609\. That is not a round number by accident. VRRP's master down interval is three advertisement intervals plus a skew derived from priority: 3 x 1000 ms, plus ((256 - 100) / 256) x 1000, which is 609 ms. A lower priority backup waits fractionally longer before declaring the master dead, which is how VRRP breaks ties between multiple backups without a fight. R1 prints 0 because it is the master and is not waiting for anybody. So on defaults, VRRP declares failure at roughly 3.6 seconds where HSRP waits its full 10 second hold. That is the only performance argument for VRRP that survives contact with a lab, and it disappears the moment you tune HSRP timers, which most production designs do. For the election mechanics and the priority 255 owner case, see [how VRRP elects a master](https://www.pinglabz.com/vrrp-explained/), and if you want to type it out yourself there is a [hands-on lab that runs VRRP and HSRP side by side](https://www.pinglabz.com/ccna-lab-ipc-13-vrrp-vs-hsrp/). ## GLBP: Both Routers Forward for the Same Virtual IP This is the section that justifies the whole comparison. GLBP splits the job in two. One router is the Active Virtual Gateway (AVG), which answers ARP for the virtual IP. Up to four routers are Active Virtual Forwarders (AVFs), each owning its own virtual MAC. The AVG hands out different forwarder MACs to different hosts as they ARP, so traffic for one gateway address leaves the subnet through more than one router. Here is what that looks like on both sides: ``` R1# show glbp brief Interface Grp Fwd Pri State Address Active router Standby router Et0/2 30 - 110 Active 10.0.30.254 local 10.0.30.2 Et0/2 30 1 - Active 0007.b400.1e01 local - Et0/2 30 2 - Listen 0007.b400.1e02 10.0.30.2 - R2# show glbp brief Interface Grp Fwd Pri State Address Active router Standby router Et0/2 30 - 100 Standby 10.0.30.254 10.0.30.1 local Et0/2 30 1 - Listen 0007.b400.1e01 10.0.30.1 - Et0/2 30 2 - Active 0007.b400.1e02 local - ``` The `Fwd` column is the key. The row with `-` is the gateway role, and there R1 is Active and R2 is Standby, which reads exactly like HSRP. The numbered rows are the forwarder roles, and there the picture inverts: R1 is Active for forwarder 1 and only Listening for forwarder 2, while R2 is Active for forwarder 2 and only Listening for forwarder 1. Put plainly, R2 is the standby gateway and a live forwarder at the same time. It is passing data plane traffic for 10.0.30.254 while R1 is perfectly healthy. Do the same comparison on the HSRP segment and R2 is doing nothing but counting hellos. That is the difference, and it is the reason GLBP exists. The virtual MACs decode as neatly as HSRP's. Group 30 is `0x1e`, and the two forwarders are `0007.b400.1e01` and `0007.b400.1e02`: group, then forwarder number. You can read a GLBP `show mac address-table` and know immediately which forwarder a host was handed. One honest caveat about what this proves. There were no clients on the GLBP segment, so you are seeing both routers hold an Active forwarder role, not a measured 50/50 traffic split. The split follows from the roles: the AVG round-robins forwarder MACs across ARP replies by default, so with enough hosts the load lands roughly evenly. With four hosts and one of them running a backup job, it will not. The [weighted and host-dependent balancing modes](https://www.pinglabz.com/glbp-load-balancing/) exist precisely because round-robin is a poor proxy for bytes. ## The Virtual MAC Is the Fingerprint If you inherit a network and want to know which FHRP is running without reading anybody's configuration, look at the gateway MAC in a host ARP table. Each protocol owns a distinct range, and the group number is encoded in it: HSRP v10000.0c07.ac**XX** where XX = group in hex (group 10 gave 0000.0c07.ac0a) HSRP v20000.0c9f.f**XXX**, three hex digits because v2 allows group numbers up to 4095 VRRP0000.5e00.01**XX** where XX = VRID (group 20 would be 0000.5e00.0114) GLBP0007.b4**XX.XXYY**, group then forwarder (group 30 gave 0007.b400.1e01 and .1e02) The HSRP and GLBP values above came straight out of this lab. The VRRP one is the range from RFC 5798, because `show vrrp brief` does not print the virtual MAC at all, which is a small but real operational annoyance when you are trying to match a MAC to a router in a hurry. The GLBP range has a practical consequence. HSRP and VRRP put one extra MAC per group into every switch CAM table on the segment. GLBP puts up to four. On a large campus with many groups that is not free, and it is one of the quieter reasons GLBP stays rare. ## The Side-by-Side Comparison Standard HSRP Cisco-proprietary (RFC 2281 informational) VRRPIETF (RFC 5798) GLBPCisco-proprietary Vendor support HSRPCisco only VRRP Universal (Cisco, Juniper, Arista, Nokia, etc.) GLBPCisco only Role names in the CLI HSRPActive / Standby VRRPMASTER / BACKUP GLBP AVG (gateway) plus AVF (forwarder), separately Does the backup forward? HSRPNo, idle until failover VRRPNo, idle until failover GLBP Yes, up to 4 active forwarders per group Load balancing HSRP Per-VLAN (different active per VLAN) VRRPPer-VLAN GLBPPer-host (within VLAN) Default Hello HSRP3 seconds VRRP1 second GLBP3 seconds Default Hold / master down HSRP10 seconds VRRP 3.609 seconds at priority 100 (3 s plus skew) GLBP10 seconds Preempt by default? HSRPNo, must be configured VRRPYes, on by default GLBPNo for the AVG role Multicast address HSRP 224.0.0.2 (v1) / 224.0.0.102 (v2) VRRP224.0.0.18 GLBP224.0.0.102 Virtual MAC HSRP0000.0c07.acXX VRRP0000.5e00.01XX GLBP0007.b4XX.XXYY States HSRP 6 (Initial, Learn, Listen, Speak, Standby, Active) VRRP 3 (Initialize, Backup, Master) GLBP 6+ (per-AVG and per-AVF) Authentication HSRPPlain text or MD5 VRRP None in v3 (deprecated v2 had MD5) GLBPMD5 IPv6 support HSRPHSRPv2 with IPv6 group VRRPVRRPv3 native GLBP Yes, IPv6 group support Object tracking HSRPYes VRRPYes GLBP Yes (weighting tracking) Configuration complexity HSRPLow VRRPLow GLBPMedium-high Operational maturity HSRP Very high (decades of Cisco deployment) VRRPHigh GLBPMedium (less common) ## Watching a Failover Actually Happen Steady state output is only half the story. The lab shut R1's HSRP interface on a timer to see what the pair does when the preferred router disappears. R1 first: ``` *Jul 20 22:34:11.877: %HSRP-5-STATECHANGE: Ethernet0/0 Grp 10 state Active -> Init Interface Grp Pri P State Active Standby Virtual IP Et0/0 10 110 P Init unknown unknown 10.0.10.254 ``` R1 keeps its priority of 110 and its preempt flag, but it is in `Init` and it has no idea who the Active or Standby routers are, because it cannot hear anything. The detail view goes further and reports `State is Init (interface down)` along with `Active virtual MAC address is unknown (MAC Not In Use)`. That last phrase is the one to remember: a router in Init is not just not forwarding, it has stopped answering for the virtual MAC entirely. Now R2, one millisecond earlier on its own clock: ``` *Jul 20 22:34:11.876: %HSRP-5-STATECHANGE: Ethernet0/0 Grp 10 state Standby -> Active Interface Grp Pri P State Active Standby Virtual IP Et0/0 10 100 Active local unknown 10.0.10.254 ``` Two things there are worth arguing about at design review. First, R2 took over immediately rather than waiting out a 10 second hold. An administrative shutdown is a graceful exit, so the Active router sends a Resign and the Standby promotes on the spot. The 10 second hold is the worst case, what you pay when the Active dies silently. To see it actually elapse you have to break the path without telling the router about it. Second, R2 has no `P` in its state line and R1 does. When R1 comes back it preempts, and the gateway moves a second time. That is usually what you want with an asymmetric design, but it is also a second outage window, and it is why preempt delay exists. VRRP would have done this without you asking. ## When HSRP Wins - **You are a Cisco-only shop.** HSRP is the default, well understood, and universally documented for Cisco networks. - **Your operators are HSRP-trained.** Switching to VRRP for marginal benefits is rarely worth retraining. - **You want simple operations.** The HSRP state machine is more verbose than VRRP's, but `show standby brief` answers the only question that matters in one line, and the troubleshooting tooling is mature. - **You want the failover behaviour to be explicit.** HSRP does not preempt unless you tell it to, which makes "who is Active right now" a decision you made rather than one the protocol made for you. - **You inherited an HSRP deployment.** Migrating to something else is rarely justified by performance or cost. ## When VRRP Wins - **Multi-vendor environment.** Mandatory if any non-Cisco router participates in the FHRP group. HSRP and GLBP are not options here at all. - **Faster default failure detection matters.** The 3.609 second master down interval this lab printed is genuinely better than HSRP's 10 second hold, if you are stuck on defaults. - **You want IETF-standard protocols across the board.** Operational discipline and vendor neutrality as a policy, not a performance argument. - **VRRPv3 IPv6 design.** The address-family syntax handles IPv4 and IPv6 in the same group structure, which is cleaner than bolting an IPv6 group onto HSRPv2. ## When GLBP Wins - **You specifically need active/active gateway forwarding without splitting VLANs.** The lab output above is the whole argument: both routers Active for one virtual IP, no per-VLAN choreography required. - **Cisco-only with homogeneous routers.** GLBP needs Cisco on both ends and works best when the forwarders have similar capacity, because the default hand-out does not know or care that one of them is smaller. - **You have idle redundant capacity** and a single large VLAN where the per-VLAN trick does not help you. In practice, GLBP shows up less often than its capabilities suggest. Most Cisco campus designs use HSRP with different active gateways per VLAN, which achieves load distribution at the VLAN level without GLBP's extra virtual MACs, extra state and extra explaining at 3am. ## Failover Times Compared Default timers, silent failure HSRP\~10 seconds VRRP\~3.6 seconds GLBP\~10 seconds Administrative shutdown All three Immediate, the leaving router resigns and the backup promotes without waiting Tuned timers (Hello 1s / Hold 3s) HSRP3 seconds VRRP3 seconds (already) GLBP3 seconds Sub-second tuning HSRP \~750ms (Hello 200ms / Hold 750ms) VRRP\~750ms GLBP\~750ms BFD-tracked (with FHRP) HSRP\~50-200ms VRRP\~50-200ms GLBP\~50-200ms The second row is the one most comparisons leave out. The failover this lab captured completed instantly, because a shutdown is a polite exit. Default timers only decide how long you wait when a router dies without saying goodbye, which is exactly the case your maintenance window will never reproduce. With BFD or interface tracking driving the decision, all three converge in similar timeframes and the protocol you picked stops mattering. ## Design Impact The most common production pattern across all three: 1. Two distribution switches per VLAN 2. FHRP (HSRP, VRRP or GLBP) configured for the VLAN 3. Object tracking on the upstream interface to decrement priority or weighting on uplink failure 4. STP root alignment with the active FHRP gateway per VLAN 5. For HSRP and VRRP: alternate which switch is active per VLAN to spread load 6. For GLBP: both switches active in the same VLAN, hosts distributed automatically HSRP and VRRP achieve load distribution by being active on different VLANs on different switches. GLBP achieves it inside a single VLAN by handing different hosts different forwarder MACs, which is what the `Fwd 1` and `Fwd 2` rows above are showing you. The per-VLAN approach is more common because it lines up with per-VLAN STP root assignment, and because a human can look at a diagram and say which box carries which VLAN. Step 3 is the one people skip, and it is the one that causes the outage where the gateway stays Active on a router that has lost its uplink. Whichever FHRP you run, wire the priority to something real: an interface, a route, or an [IP SLA probe that tests whether the upstream path is actually usable](https://www.pinglabz.com/ip-sla-cisco-ios-xe/). For the STP side of the same problem, the write-up on [keeping the spanning tree root on the same box as the active gateway](https://www.pinglabz.com/stp-fhrp-hsrp-vrrp-alignment/) covers why a misaligned pair sends every packet across your interswitch link twice. ## Migrating Between FHRPs Migration paths exist but are operationally fragile: HSRP to VRRP Approach Configure VRRP alongside HSRP on a different group and a different virtual IP, verify, move hosts across via DHCP, then remove HSRP. Remember VRRP preempts by default. HSRP to GLBP Approach Same shape, but expect up to four new virtual MACs per group in the CAM tables and check your access switches are happy with that. Rarely justified. VRRP to HSRP Approach Reverse of the above. Only do this if you are removing the non-Cisco gear that forced VRRP in the first place. The "configure the new one alongside the old one, migrate, then remove" approach minimises disruption. Both FHRPs run at once on different group numbers and different virtual IPs, and hosts keep the old gateway address until DHCP renews them onto the new one. The failure mode to watch for is the half-migrated subnet where your ACLs only permit one of the two gateways. ## Recommendations Default recommendations for typical scenarios: Cisco-only campus, new deployment HSRP. Simpler, well understood, mature. Multi-vendor environment VRRPv3\. Mandatory. Cisco-only, want active/active in one VLAN GLBP. The one case it genuinely wins. Existing HSRP deployment Stay with HSRP unless a multi-vendor requirement forces VRRP. Existing VRRP deployment Stay with VRRP. Moving to HSRP is a regression. Greenfield IPv6-only VRRPv3\. Cleanest IPv6 story. ## What This Was Captured On PlatformCML, two `iol-xe` nodes, IOS XE 17.18.2 HSRP segmentEt0/0, 10.0.10.0/24, group 10, VIP 10.0.10.254 VRRP segmentEt0/1, 10.0.20.0/24, group 20, VIP 10.0.20.254, VRRPv3 syntax GLBP segmentEt0/2, 10.0.30.0/24, group 30, VIP 10.0.30.254 PrioritiesR1 = 110 on all three groups, R2 = default 100 Capture methodOn-box EEM applets dumping show output to syslog, plus a timed applet to shut R1's HSRP interface ## Gotchas - **The legacy VRRP one-liner is gone on modern IOS XE.** You need `fhrp version vrrp v3` and then `vrrp 20 address-family ipv4` with the address as a sub-command. Old configs and old study guides both mislead you here. - **VRRP preempts by default and HSRP does not.** Migrate a design that relied on "whoever came up first stays Active" and VRRP will quietly invert it the first time a router reloads. - **A GLBP standby gateway is not idle.** R2 is the standby AVG while being the Active forwarder for `0007.b400.1e02`. Take that router out for maintenance expecting it to be doing nothing and you are dropping live sessions. - **Do not read `show glbp brief` as one state.** There is a gateway row and there are forwarder rows, and they disagree on purpose. The `Fwd` column tells them apart. - **Default timers do not describe an administrative shutdown.** The 10 second hold and 3.6 second master down interval only apply to silent failures, so testing failover by shutting a port always looks better than reality. - **Two Actives is a hello problem, not a priority problem.** If both routers claim the role and both print `Standby unknown`, the protocol is fine and the path between them is not. That failure has its own walkthrough on [what to check when both routers say they are Active](https://www.pinglabz.com/hsrp-troubleshooting-flapping-dual-active/). ## Key Takeaways - The only functional difference between the three is whether the backup forwards. HSRP and VRRP keep it idle, GLBP does not. - Vocabulary maps directly: Active/Standby is HSRP, MASTER/BACKUP is VRRP, AVG plus AVF is GLBP. - The virtual MAC identifies the protocol and encodes the group: `0000.0c07.ac0a` is HSRP group 10, `0007.b400.1e01` is GLBP group 30 forwarder 1. - VRRP's advantage on defaults is real but small: a 3.609 second master down interval against HSRP's 10 second hold, and it evaporates as soon as you tune timers or add tracking. - GLBP costs up to four virtual MACs per group in every CAM table on the segment, which is a large part of why per-VLAN HSRP won the deployment argument. - Whatever you pick, the priority has to track something real. An FHRP that stays Active on a router with a dead uplink is worse than no FHRP at all. ## Summary HSRP for Cisco-only simplicity. VRRP for multi-vendor networks and an IETF-standard preference. GLBP for the specific case where you need two routers forwarding for one gateway address inside one VLAN. All three converge in similar time once you tune them or drive them from BFD, and all three fail over instantly when a router leaves gracefully. Default timers favour VRRP, and the lab printed the exact number: 3,609 milliseconds against 10 seconds. The right answer is mostly determined by what is already in your network. Most enterprises run HSRP because they were Cisco shops and HSRP was the default, and few have a real reason to migrate. If you are choosing rather than inheriting, work through the rest of the [first-hop redundancy protocol guides](https://www.pinglabz.com/fhrp/) and pick based on what your operators can troubleshoot at 3am, not on a millisecond in a datasheet. ### GLBP for Active/Active Load Balancing URL: https://www.pinglabz.com/glbp-load-balancing/ Last updated: 2026-06-13T20:08:37.000Z GLBP (Gateway Load Balancing Protocol) is Cisco's first-hop redundancy protocol that does what HSRP and VRRP cannot: actively load-balance across multiple routers in a single group. Where HSRP and VRRP have one active forwarder per group, GLBP can have up to four active forwarders simultaneously, each handling a portion of the host traffic. The result is utilization across all the redundant routers instead of one sitting idle as standby. This article walks through the GLBP architecture, the Active Virtual Gateway (AVG) and Active Virtual Forwarders (AVFs), the load balancing methods, configuration, and when GLBP is actually the right answer (it is more niche than the marketing suggests). ## The Architecture GLBP introduces two roles within a group: AVG (Active Virtual Gateway) One per group; manages the protocol; assigns virtual MACs to AVFs; responds to ARP requests with the right virtual MAC AVF (Active Virtual Forwarder) Up to four per group; each owns a unique virtual MAC and forwards packets sent to that MAC How a packet flows: 1. End host ARPs for the virtual IP gateway. 2. The AVG (one router in the group) responds, but the MAC it returns rotates among the four virtual MACs assigned to the four AVFs. 3. Different hosts ARP at different times and receive different virtual MACs, so each host pins to a specific AVF for its outbound traffic. 4. Each AVF forwards the packets it receives via its virtual MAC. 5. Failover: if an AVF fails, the AVG redirects its virtual MAC to a different AVF. The result: per-host load distribution. With 100 hosts and 4 AVFs, roughly 25 hosts per AVF. Load distribution is statistical, not strictly equal, but the average across many hosts is close to even. ## Load Balancing Methods GLBP supports three load balancing modes, controlled by how the AVG hands out virtual MACs in ARP responses: Round-robin (default) Cycle through AVFs in order. First host gets AVF1, second host gets AVF2, etc. Weighted AVFs with higher weight get more hosts. Useful when AVFs have different capacity. Host-dependent Same host always gets the same AVF (based on hash of source MAC). Provides stability across host reboots. Round-robin is the dominant default; the others are for specific edge cases. ## Configuration ``` ! Distribution switch 1 (becomes AVG) interface GigabitEthernet1/0/24 ip address 192.168.1.2 255.255.255.0 glbp 1 ip 192.168.1.1 glbp 1 priority 110 glbp 1 preempt glbp 1 load-balancing round-robin glbp 1 weighting 100 glbp 1 weighting track 1 decrement 30 ! Distribution switch 2 (AVF) interface GigabitEthernet1/0/24 ip address 192.168.1.3 255.255.255.0 glbp 1 ip 192.168.1.1 glbp 1 priority 100 glbp 1 preempt glbp 1 weighting 100 track 1 interface GigabitEthernet0/0/0 line-protocol ``` Notice the AVF weighting separate from priority. Priority elects the AVG; weighting determines AVF eligibility (an AVF whose weighting drops below a threshold stops forwarding). This separation is what lets GLBP do load balancing. ## Virtual MAC Address Format GLBP virtual MACs follow a specific format: 0007.b400.GGFF where GG is the group number and FF is the AVF number (1-4). 1 Group1 Virtual MAC0007.b400.0101 2 Group1 Virtual MAC0007.b400.0102 3 Group1 Virtual MAC0007.b400.0103 4 Group1 Virtual MAC0007.b400.0104 1 Group10 Virtual MAC0007.b400.0a01 Different from HSRP (0000.0c07.acXX) and VRRP (0000.5e00.01XX). All three FHRP families use distinct OUIs. ## The State Machine Disabled GLBP is configured but not yet running Initial Just started; no Hellos yet Listen Receiving Hellos but not active in any role Speak Sending Hellos and participating in election Standby Backup AVG (will take over if current AVG fails) Active Currently the AVG OR an AVF (depending on context) Verify with: ``` Switch# show glbp brief Interface Grp Fwd Pri State Address Active router Gi1/0/24 1 - 100 Active 192.168.1.1 local Gi1/0/24 1 1 - Active 0007.b400.0101 local Gi1/0/24 1 2 - Listen 0007.b400.0102 192.168.1.3 Gi1/0/24 1 3 - Listen 0007.b400.0103 192.168.1.4 Gi1/0/24 1 4 - Listen 0007.b400.0104 192.168.1.5 Switch# show glbp 1 GigabitEthernet1/0/24 - Group 1 State is Active ... Forwarder 1 State is Active ... Forwarder 2 State is Listen ... ``` The first row is the AVG; subsequent rows are AVFs. "Active" forwarders are currently distributing traffic; "Listen" forwarders are backups. ## When GLBP Is the Right Answer GLBP is more niche than its marketing suggests. It is the right answer when: - **You have multiple physical paths upstream** and want to use them simultaneously instead of one being idle as standby - **You are entirely Cisco-shop** (GLBP is Cisco-only) - **You have homogeneous routers** that can handle the additional GLBP complexity - **You want active/active without per-VLAN role tuning** (the alternative is HSRP/VRRP with different active per VLAN, which works but is more configuration) GLBP is not the right answer when: - You have a multi-vendor environment (use VRRP) - You have asymmetric routers (one fast, one slow); GLBP's per-host distribution does not handle asymmetry well without weighting - You want simple, well-understood operations (HSRP is simpler) - You are studying for VRRP/HSRP exam topics (both are more commonly tested) In practice, most Cisco campus designs use HSRP with per-VLAN active distribution. GLBP shows up less often than the protocol's existence would suggest. ## Anti-Patterns - **GLBP across non-Cisco gear.** Does not work; not standardized. - **Skipping weighting tracking.** An AVF with a dead WAN keeps forwarding to a black hole. Always track upstream interfaces and decrement weighting on failure. - **Weighting threshold too high or low.** If too low, the AVF rarely loses its active status and tracking is ineffective. If too high, transient drops cause flapping. Tune based on the tracked interface's expected stability. - **Mixing GLBP with HSRP on the same hosts.** Hosts get one default gateway IP. Pick one protocol per gateway and stick with it. ## Summary GLBP is Cisco's load-balancing FHRP. Up to four AVFs with different virtual MACs share the load; the AVG distributes hosts across them via ARP responses. Per-host load balancing without per-VLAN role tuning - the marketing case is real, but most Cisco campus designs use HSRP because it is simpler and well-understood. Use GLBP when you specifically need active/active gateway forwarding in a Cisco-only network. Otherwise HSRP is the safer default. Bookmark this article alongside the [FHRP cluster pillar](https://www.pinglabz.com/fhrp/) and the [HSRP vs VRRP vs GLBP comparison](https://www.pinglabz.com/hsrp-vs-vrrp-vs-glbp/). ### VRRP Explained: The Vendor-Neutral FHRP URL: https://www.pinglabz.com/vrrp-explained/ Last updated: 2026-06-13T20:08:37.000Z VRRP (Virtual Router Redundancy Protocol, RFC 5798) is the vendor-neutral first-hop redundancy protocol. It does the same job as Cisco's HSRP - share a virtual IP and MAC across multiple routers so end hosts have a stable default gateway - but as an IETF standard it works across Cisco, Juniper, Arista, and most other enterprise routing vendors. If your network is multi-vendor or you want to avoid Cisco lock-in, VRRP is the answer. This article walks through how VRRP works, the differences from HSRP, configuration on Cisco IOS XE and Juniper, the master election, tracking, and the IPv6 story. If you are configuring VRRP for the first time, designing a multi-vendor redundant gateway, or migrating from HSRP to VRRP, this is the reference. ## How VRRP Works The architecture mirrors HSRP. A VRRP group consists of two or more routers sharing: - A virtual IP (the gateway address end hosts use) - A virtual MAC (00-00-5E-00-01-XX where XX is the VRID, virtual router ID) - Periodic VRRP advertisements between members Each member has a priority (1-254, default 100; 0 reserved for "give up master" and 255 reserved for the IP address owner). The router with the highest priority becomes Master and forwards traffic; others are Backup and wait. If the Master fails (advertisements stop), a Backup with the highest priority takes over. VRRP uses IP protocol number 112 and multicast address 224.0.0.18 (IPv4) or FF02::12 (IPv6). ## VRRP States Initialize VRRP just started; not yet eligible for election Backup Receiving Master's advertisements; waiting Master Currently forwarding traffic for the virtual IP Healthy steady state: one Master, others Backup. VRRP is simpler than HSRP's six states (HSRP has Listen, Speak, etc., which VRRP collapses into Backup). ## The IP Address Owner VRRP has a unique concept: if the configured virtual IP matches a router's actual interface IP, that router is the IP address owner with priority 255 (reserved). It always wins the master election regardless of other configuration. This causes confusion. If you configure VRRP virtual IP 192.168.1.1 and one router has 192.168.1.1 as its actual interface IP, that router becomes IP address owner and master. The "ip address owner" pattern is sometimes useful (the active gateway IP equals one router's address) but more often surprising - operators expect priority-based election and get IP-owner behavior. Workaround: pick a virtual IP that does NOT match any router's actual interface IP. Use a separate IP for the virtual gateway from the routers' physical interface IPs. ## VRRP vs HSRP Standard VRRPRFC 5798 (IETF) HSRP Cisco-proprietary (RFC 2281 informational) Vendor support VRRPUniversal HSRP Cisco only (some interop with newer non-Cisco) Virtual MAC range VRRP00-00-5E-00-01-XX HSRP00-00-0C-07-AC-XX Default Hello VRRP1 second HSRP3 seconds Default Hold/Dead VRRP3 seconds HSRP10 seconds Multicast address VRRP224.0.0.18 HSRP 224.0.0.2 (v1) / 224.0.0.102 (v2) States VRRP Initialize, Backup, Master (3) HSRP Initial, Learn, Listen, Speak, Standby, Active (6) IP address owner VRRPYes (priority 255) HSRPNo Authentication VRRP None in v3 (deprecated v2 plain/MD5) HSRPPlain text or MD5 IPv6 support VRRPNative (VRRPv3) HSRP Yes (HSRPv2 with IPv6 group) The key practical difference: VRRP failover is faster by default (3 seconds vs HSRP's 10) and the protocol is simpler. HSRP's six states give more granularity but in production both protocols converge in the same operational windows. ## Cisco IOS XE Configuration ``` ! Define a VRRP group on the interface interface GigabitEthernet0/0/1 ip address 192.168.1.2 255.255.255.0 vrrp 1 ip 192.168.1.1 vrrp 1 priority 110 vrrp 1 preempt vrrp 1 timers advertise 1 vrrp 1 authentication md5 key-string Cisco123! vrrp 1 description Gateway-VLAN1 ``` Three things to notice. First, `vrrp 1 ip 192.168.1.1` sets the virtual IP for VRRP group 1\. Second, priority 110 (default 100) makes this router the master if no other has higher. Third, `preempt` lets a higher-priority router take over when it comes back online (otherwise the existing master keeps the role). For VRRPv3 (which supports IPv6 and offers some new features): ``` ! Enable VRRPv3 globally fhrp version vrrp v3 interface GigabitEthernet0/0/1 ip address 192.168.1.2 255.255.255.0 vrrp 1 address-family ipv4 address 192.168.1.1 primary priority 110 preempt authentication md5 key-string Cisco123! ``` VRRPv3 is the modern default and supports both IPv4 and IPv6\. New deployments should use it. ## Juniper Junos Configuration ``` set interfaces ge-0/0/1 unit 0 family inet address 192.168.1.2/24 vrrp-group 1 virtual-address 192.168.1.1 set interfaces ge-0/0/1 unit 0 family inet address 192.168.1.2/24 vrrp-group 1 priority 110 set interfaces ge-0/0/1 unit 0 family inet address 192.168.1.2/24 vrrp-group 1 preempt ``` Different syntax, same protocol. Cisco and Juniper VRRP routers in the same group interoperate without issue. ## Object Tracking Tracking lets VRRP automatically lower priority when an uplink fails, allowing a backup with intact uplinks to take over. Common pattern: track the WAN-facing interface; lower priority by 20 if it goes down. ``` track 1 interface GigabitEthernet0/0/0 line-protocol interface GigabitEthernet0/0/1 vrrp 1 ip 192.168.1.1 vrrp 1 priority 110 vrrp 1 track 1 decrement 20 ``` If GigabitEthernet0/0/0 goes down, VRRP priority drops from 110 to 90 (110 - 20). If the other router has priority 100, it becomes master. When the WAN comes back, priority returns to 110 and (with preempt) the original master takes over. Tracking is mandatory in production. Without it, a router with a dead WAN keeps its master role and forwards into a black hole. ## VRRPv3 for IPv6 ``` fhrp version vrrp v3 interface GigabitEthernet0/0/1 ipv6 address 2001:db8:1::2/64 vrrp 1 address-family ipv6 address FE80::1 primary priority 110 preempt ``` For IPv6, the virtual address is a link-local FE80::/10 address used for next-hop forwarding. Hosts learn it via Router Advertisements (RAs); the active VRRP master sends RAs with the virtual IP as the gateway. ## Design Patterns Common VRRP design patterns: - **Active/Standby per VLAN:** two distribution switches, each is master for half the VLANs and backup for the other half. Combined with STP root alignment, this gives load distribution without active/active complexity. - **VRRP + first-hop ECMP:** in some designs, multiple routers respond to the same virtual IP as long as the priority and timer behavior is configured carefully. Less common; HSRP and VRRP are fundamentally active/standby. - **Multi-site VRRP via stretched VLAN:** uncommon and operationally fragile; the inter-site link becomes a single point of failure. ## Verification ``` ! VRRP state per group Router# show vrrp brief Interface Grp Pri Time Own Pre State Master addr Group addr Gi0/0/1 1 110 3000 Y Master 192.168.1.2 192.168.1.1 ! Detailed VRRP info Router# show vrrp interface GigabitEthernet0/0/1 GigabitEthernet0/0/1 - Group 1 State is Master Virtual IP address is 192.168.1.1 Virtual MAC address is 0000.5e00.0101 Advertisement interval is 1.000 sec Preemption enabled Priority is 110 (configured 110) Master Router is 192.168.1.2 (local), priority is 110 ``` ## Anti-Patterns - **IP address owner confusion.** Using a virtual IP that matches one router's actual interface IP. The owner always wins regardless of other priority configuration. Use distinct addresses. - **No tracking.** Master keeps forwarding even when its WAN uplink is dead. Always track upstream interfaces. - **No preempt.** A failed master that comes back becomes backup; the original backup keeps the master role even though the original was provisioned as primary. Enable preempt for predictable role behavior. - **Mismatched timers between vendors.** Both ends of a multi-vendor VRRP group must agree on Hello/Hold timers. Mismatch causes flapping. - **VRRPv2 in 2026.** Use v3 for new deployments; supports IPv4 and IPv6 in one syntax. ## Summary VRRP is the vendor-neutral first-hop redundancy protocol. It works the way HSRP does but as an IETF standard, with simpler state machine and faster default timers. Use VRRPv3 for new deployments, configure tracking on upstream interfaces, enable preempt for predictable role behavior, and avoid the IP-address-owner trap by using distinct virtual IPs. For Cisco-only campuses HSRP remains a fine choice; for multi-vendor environments VRRP is the answer. Bookmark this article alongside the [FHRP cluster pillar](https://www.pinglabz.com/fhrp/) and the [HSRP vs VRRP vs GLBP comparison](https://www.pinglabz.com/hsrp-vs-vrrp-vs-glbp/). ### Cisco MPLS Configuration on IOS XE URL: https://www.pinglabz.com/cisco-mpls-configuration/ Last updated: 2026-05-29T23:40:56.000Z Configuring MPLS on Cisco IOS XE breaks down into a few discrete tasks: enabling MPLS forwarding on the right interfaces, configuring LDP, integrating with the IGP for label distribution, and (for L3VPN) layering on VRFs and MP-BGP. This article walks through each task with copy-paste-able configurations. If you are bringing up your first MPLS router, replicating an existing config in a lab, or troubleshooting a partial deployment, this is the operator's walkthrough. ## Prerequisites Before configuring MPLS: 1. An IGP must be running between the routers that will participate in MPLS. Loopbacks must be reachable across all PEs and Ps. 2. Loopback interfaces should be created on every router and used as the IGP/MPLS router-id source. Stable, always-up. 3. The interfaces that will run MPLS must already have IP addressing and be in the IGP. This article assumes OSPF is the IGP and Loopback0 is the router-id loopback. ## Basic MPLS Forwarding The minimum config to enable MPLS on an interface: ``` interface GigabitEthernet0/0/0 ip address 10.0.12.1 255.255.255.252 mpls ip mpls label protocol ldp ``` `mpls ip` enables MPLS forwarding and starts LDP discovery. `mpls label protocol ldp` selects LDP (the alternative is the legacy TDP, which you almost never want). On modern IOS XE, LDP is the default; this command is for explicitness. Apply to every router-to-router interface inside the MPLS domain. Customer-facing interfaces (PE-to-CE) typically do not run MPLS. ## LDP Configuration Global LDP settings: ``` mpls ldp router-id Loopback0 force mpls ldp sync ``` `force` on the router-id command makes it use Loopback0 even if there are other interfaces with higher IPs. Without `force`, the router-id only changes after a reload, which can cause unexpected behavior. `mpls ldp sync` globally enables LDP-IGP synchronization for all OSPF interfaces. Always enable in production - prevents transient drops on link bring-up. Optional per-interface settings: ``` interface GigabitEthernet0/0/0 mpls ip mpls ldp sync ! Per-interface (override global) mpls ldp discovery hello interval 5 ! Default 5s; tune lower for faster convergence mpls ldp discovery hello holdtime 15 ``` ## VRF Configuration for L3VPN Each customer gets a VRF on every PE that serves them. ``` vrf definition CUSTOMER-A description Customer A - VPN service rd 65001:1 address-family ipv4 route-target export 65001:100 route-target import 65001:100 exit-address-family ``` Place customer-facing interfaces into the VRF: ``` interface GigabitEthernet0/0/1 vrf forwarding CUSTOMER-A ip address 192.168.1.1 255.255.255.0 ``` Note: `vrf forwarding` erases any existing IP configuration on the interface. Set the VRF first, then the IP, otherwise you have to retype the IP. ## MP-BGP for L3VPN iBGP between PEs, sourced from loopbacks, with the vpnv4 unicast address family for VPN routes: ``` router bgp 65001 ! Generic BGP setup bgp router-id 1.1.1.1 bgp log-neighbor-changes no bgp default ipv4-unicast ! Don't activate per-neighbor by default ! Peer with other PE (typically via route reflector at scale) neighbor 10.10.10.10 remote-as 65001 neighbor 10.10.10.10 update-source Loopback0 ! VPNv4 address family address-family vpnv4 unicast neighbor 10.10.10.10 activate neighbor 10.10.10.10 send-community extended ! Critical for RT propagation exit-address-family ! Per-VRF address family address-family ipv4 vrf CUSTOMER-A redistribute connected redistribute static ! Optional BGP PE-CE neighbor 192.168.1.2 remote-as 65100 neighbor 192.168.1.2 activate exit-address-family ``` Three things to notice. First, `send-community extended` is mandatory - RTs are extended communities. Second, the `address-family ipv4 vrf` block configures redistribution and PE-CE peering for that specific VRF. Third, `no bgp default ipv4-unicast` prevents BGP from auto-activating new neighbors for the global IPv4 table; you have to explicitly activate each per-AF. ## Route Reflector Pattern At scale, full-mesh iBGP between every PE is unmanageable. Route reflectors break the iBGP loop-prevention rule in a controlled way: a route reflector can re-advertise iBGP-learned routes to its clients. ``` ! On the route reflector router bgp 65001 address-family vpnv4 unicast neighbor 1.1.1.1 activate neighbor 1.1.1.1 route-reflector-client neighbor 2.2.2.2 activate neighbor 2.2.2.2 route-reflector-client ... exit-address-family ``` Each PE peers only with the route reflector instead of with every other PE. Two RRs for redundancy. See the BGP cluster pillar for the full RR design pattern. ## Authentication Both LDP and BGP support authentication; enable both in production. LDP MD5: ``` mpls ldp neighbor 2.2.2.2 password Cisco123! ``` BGP MD5: ``` router bgp 65001 neighbor 10.10.10.10 password Cisco123! ``` Both ends must agree on the password. Different passwords for different neighbors are fine. ## Verification Sequence Walk through these commands in order when bringing up a new MPLS deployment: ``` ! IGP working? show ip route ospf show ip ospf neighbor ! MPLS interfaces enabled? show mpls interfaces ! LDP sessions up? show mpls ldp neighbor ! Labels being assigned? show mpls ldp bindings ! Forwarding state populated? show mpls forwarding-table ! For L3VPN: VRF routing table? show ip route vrf CUSTOMER-A ! VPNv4 routes received? show bgp vpnv4 unicast all summary ! End-to-end ping/traceroute test from CE ping vrf CUSTOMER-A 10.0.0.5 traceroute vrf CUSTOMER-A 10.0.0.5 ``` Each step depends on the previous. If show ip ospf neighbor is empty, MPLS will never come up. Walk back to fundamentals if anything is missing. ## A Complete PE Configuration ``` ! Hostname and basics hostname PE1 ip cef ipv6 cef ! Loopback for router-id interface Loopback0 ip address 1.1.1.1 255.255.255.255 ! Core-facing interface (MPLS-enabled) interface GigabitEthernet0/0/0 description To P1 ip address 10.0.12.1 255.255.255.252 mpls ip mpls ldp sync ! Customer-facing interface (in VRF) interface GigabitEthernet0/0/1 description To CE-CustomerA vrf forwarding CUSTOMER-A ip address 192.168.1.1 255.255.255.0 ! VRF vrf definition CUSTOMER-A rd 65001:1 address-family ipv4 route-target export 65001:100 route-target import 65001:100 ! IGP (OSPF for the core) router ospf 1 router-id 1.1.1.1 network 1.1.1.1 0.0.0.0 area 0 network 10.0.12.0 0.0.0.3 area 0 passive-interface default no passive-interface GigabitEthernet0/0/0 ! MPLS mpls ldp router-id Loopback0 force mpls ldp sync ! BGP for L3VPN router bgp 65001 bgp router-id 1.1.1.1 bgp log-neighbor-changes no bgp default ipv4-unicast neighbor 10.10.10.10 remote-as 65001 ! Other PE neighbor 10.10.10.10 update-source Loopback0 address-family vpnv4 unicast neighbor 10.10.10.10 activate neighbor 10.10.10.10 send-community extended exit-address-family address-family ipv4 vrf CUSTOMER-A redistribute connected neighbor 192.168.1.2 remote-as 65100 ! CE-Customer-A neighbor 192.168.1.2 activate exit-address-family ``` This is a complete PE for one customer. Add additional VRFs and BGP address-family blocks per customer. ## A Customer Edge Configuration ``` ! Customer's router; runs IP only, no MPLS hostname CE1-CustomerA interface GigabitEthernet0/0/0 description To PE1 ip address 192.168.1.2 255.255.255.0 interface GigabitEthernet0/0/1 description Local LAN ip address 10.0.0.1 255.255.255.0 ! BGP to PE router bgp 65100 bgp router-id 192.168.1.2 neighbor 192.168.1.1 remote-as 65001 network 10.0.0.0 mask 255.255.255.0 ``` The CE is unaware of MPLS. It speaks BGP (or OSPF/EIGRP/static) with the PE; the PE handles all the L3VPN complexity. ## Anti-Patterns - **Forgetting `send-community extended`.** RTs do not propagate; routes do not import where expected. - **Skipping LDP-IGP sync.** Transient drops on link bring-up. - **Using physical interface IPs as the router-id.** Unstable; can change unexpectedly. Always use a Loopback. - **Forgetting `force` on LDP router-id.** Router-id only changes after reload, surprising operators during interface changes. - **Mixing classic `ip vrf` and modern `vrf definition`.** Both work; standardize on vrf-definition for new deployments. - **Full-mesh iBGP at scale.** Use route reflectors past about 10 PEs. ## Summary Configuring MPLS on Cisco IOS XE is a sequence: IGP first, then MPLS forwarding (`mpls ip`) on each core-facing interface, then LDP for label distribution, then VRFs for L3VPN, then MP-BGP between PEs. Each piece builds on the previous. The complete PE config fits in one screen. Master the verification sequence (`show mpls interfaces` then `show mpls ldp neighbor` then `show mpls forwarding-table` then BGP) and you can troubleshoot any deployment by walking back through the dependency chain. Bookmark this article alongside the [MPLS cluster pillar](https://www.pinglabz.com/mpls/) and the [L3VPN deep-dive](https://www.pinglabz.com/mpls-l3vpn/). ### MPLS L3VPN with MP-BGP and VPNv4 URL: https://www.pinglabz.com/mpls-l3vpn/ Last updated: 2026-07-04T23:21:11.000Z MPLS L3VPN (Layer 3 VPN, RFC 4364) is the service that made MPLS commercially successful. It lets a service provider carry many customers' overlapping IP address spaces over a single shared backbone, with each customer seeing only their own routes. Customer A and customer B both using 10.0.0.0/24 do not conflict because the provider keeps them in separate VRFs and tags VPN routes with route distinguishers and route targets. For where this topic sits in the wider picture, see the [BGP complete guide](https://www.pinglabz.com/bgp/). This article walks through the L3VPN architecture, VRFs, route distinguishers, route targets, the two-label stack, the MP-BGP control plane, and the Cisco IOS XE configuration. If you are configuring an MPLS L3VPN for the first time, troubleshooting a customer's missing routes, or trying to understand the relationship between BGP and MPLS, this is the reference. ## The L3VPN Architecture CE (Customer Edge) Customer router; runs IP only; connects to PE via static routes, OSPF, EIGRP, BGP, or RIP PE (Provider Edge) Provider's edge router; runs MPLS toward core, IP toward customer; maintains per-customer VRFs P (Provider Core) Pure label-switching router in the core; never sees customer routes or labels VRF (Virtual Routing and Forwarding) Per-customer routing instance on the PE; isolates routing tables RD (Route Distinguisher) 8-byte prefix prepended to IPv4 routes to make VPNv4 globally unique RT (Route Target) Extended community attached to BGP routes; controls VPN import/export ## The Problem L3VPN Solves Customer A: site A1 in New York, site A2 in San Francisco. Both use 10.0.0.0/24 internally. Customer B: site B1 in New York, site B2 in San Francisco. Both also use 10.0.0.0/24 internally. Both buy MPLS service from the same provider. Both customers' sites need connectivity between their own remote sites without seeing each other's traffic. Plain IP routing cannot solve this - the routes overlap. L3VPN solves it by: 1. Each customer's routes live in a separate VRF on each PE. 2. The provider's PE prepends an 8-byte Route Distinguisher to each route when sending it to other PEs via MP-BGP. Customer A's 10.0.0.0/24 becomes 65001:1:10.0.0.0/24; customer B's becomes 65001:2:10.0.0.0/24\. Globally unique. 3. Each route is also tagged with a Route Target that controls which VRFs import it. Customer A's RT is 65001:100; customer B's is 65001:200\. Each customer's VRF only imports its own RT. 4. Each PE assigns a per-VPN label so the egress PE knows which VRF a packet belongs to. The packet on the wire has two labels: outer transport (LDP) and inner VPN. Result: customer A's packet from A1 goes through the MPLS core to the egress PE, which pops both labels and forwards to A2\. Customer B's packet does the same separately. Neither customer sees the other. ## VRFs: Per-Customer Routing Instances A VRF is essentially a separate routing table on the PE. Two flavors of VRF configuration on Cisco IOS XE: **Classic VRF (legacy):** ``` ip vrf CUSTOMER-A rd 65001:1 route-target export 65001:100 route-target import 65001:100 interface GigabitEthernet0/0/1 ip vrf forwarding CUSTOMER-A ip address 192.168.1.1 255.255.255.0 ``` **VRF Definition (modern, IPv4 + IPv6):** ``` vrf definition CUSTOMER-A rd 65001:1 address-family ipv4 route-target export 65001:100 route-target import 65001:100 exit-address-family interface GigabitEthernet0/0/1 vrf forwarding CUSTOMER-A ip address 192.168.1.1 255.255.255.0 ``` Both achieve the same effect - the interface is in the customer's VRF, isolated from the global routing table and other VRFs. Modern deployments use the vrf-definition form because it cleanly supports IPv4 and IPv6 in one VRF. ## Route Distinguishers An RD is an 8-byte value prepended to a 4-byte IPv4 prefix to produce a 12-byte VPNv4 prefix. Format: ``` RD = <2-byte Type><6-byte value> Type 0: 2-byte ASN : 4-byte assigned number Example: 65001:1 Type 1: 4-byte IPv4 : 2-byte assigned number Example: 192.0.2.1:1 Type 2: 4-byte ASN : 2-byte assigned number Example: 4200000000:1 ``` The RD's only job is to make customer prefixes globally unique inside the MPLS network. It does not control import/export behavior - that is the RT's job. Two PEs serving the same customer can use different RDs (and often do for ECMP scenarios where the same prefix needs two BGP best paths). ## Route Targets An RT is an extended community attached to each BGP route. Format mirrors the RD format. The PE attaches RTs on export and matches RTs on import to decide which VRFs receive which routes. ``` vrf definition CUSTOMER-A rd 65001:1 address-family ipv4 route-target export 65001:100 ! Tag exported routes with RT 65001:100 route-target import 65001:100 ! Import any route tagged with RT 65001:100 ``` For a simple "every site of this customer reaches every other" topology, export and import the same RT. Each PE sees every other PE's customer routes (filtered to those tagged with the matching RT) and installs them in the VRF. For more complex topologies (extranet, hub-and-spoke, partial mesh), use different export and import RTs to control which routes flow where. Example: a hub-and-spoke design might have spokes export RT 65001:100 and import RT 65001:200, while the hub exports RT 65001:200 and imports RT 65001:100\. Spokes only see hub routes; the hub sees all spoke routes; spokes do not see each other. ## MP-BGP Carries VPNv4 The PEs exchange VPNv4 routes via MP-BGP, specifically address-family vpnv4 unicast. This is where the BGP cluster intersects directly with the MPLS cluster. The MP-BGP article ([MP-BGP: Multiprotocol BGP](https://www.pinglabz.com/mp-bgp-multiprotocol-bgp/)) covers the BGP side; this section covers the MPLS-specific aspects. iBGP between PEs (typically with route reflectors at scale): ``` router bgp 65001 neighbor 10.10.10.10 remote-as 65001 neighbor 10.10.10.10 update-source Loopback0 address-family vpnv4 unicast neighbor 10.10.10.10 activate neighbor 10.10.10.10 send-community extended exit-address-family address-family ipv4 vrf CUSTOMER-A redistribute connected redistribute static exit-address-family ``` The send-community extended is critical - RTs are extended communities. Without it, RTs do not propagate and route filtering breaks. ## The Two-Label Stack An L3VPN packet on the wire has two MPLS labels: Outer (transport) SourceLDP (or SR-MPLS) Purpose Get the packet from ingress PE to egress PE through the core Inner (VPN) SourceMP-BGP Purpose Tell the egress PE which VRF / customer this packet belongs to The journey: 1. Customer A's CE in New York sends a packet to 10.0.0.5 (somewhere at site A2 in San Francisco). 2. Ingress PE looks up 10.0.0.5 in CUSTOMER-A's VRF, finds the route via egress PE 10.10.10.10 with VPN label 24. 3. Ingress PE looks up 10.10.10.10 in the global RIB, finds LDP transport label 17 to the next P router. 4. Ingress PE pushes both labels: outer 17 (transport), inner 24 (VPN). Sends. 5. P routers along the path swap the outer label as needed; the inner label is invisible to them. 6. The penultimate P router pops the outer label (PHP) and forwards to egress PE. 7. Egress PE sees a packet with one label (24, the VPN label), looks up the label in its forwarding table, finds it maps to CUSTOMER-A's VRF and forwards to the appropriate next-hop CE. For more on the label stack and PHP, see [MPLS Labels Explained](https://www.pinglabz.com/mpls-labels-explained/). ## PE-CE Routing The PE and CE need a way to exchange customer routes. Several options: Static routes Simplest; small sites with few prefixes BGP Most common; same protocol as the inter-PE backbone; clean integration OSPF (per-VRF) Mid-sized customers wanting dynamic routing EIGRP (per-VRF) Cisco-only customers; less common in modern deployments RIPv2 Legacy; almost extinct BGP PE-CE is the dominant pattern. The CE runs eBGP with the PE; the PE redistributes BGP into MP-BGP for the rest of the L3VPN backbone. Loop prevention via the AS\_PATH is automatic. ## Cisco IOS XE Configuration: A Working Example ``` ! VRF definition vrf definition CUSTOMER-A rd 65001:1 address-family ipv4 route-target export 65001:100 route-target import 65001:100 ! PE-CE interface interface GigabitEthernet0/0/1 vrf forwarding CUSTOMER-A ip address 192.168.1.1 255.255.255.0 ! IGP for PE-PE reachability (loopbacks, infrastructure) router ospf 1 router-id 1.1.1.1 network 10.0.0.0 0.0.255.255 area 0 ! MPLS in the core mpls ldp router-id Loopback0 force mpls ldp sync interface GigabitEthernet0/0/0 ip address 10.0.12.1 255.255.255.252 mpls ip mpls ldp sync ! BGP for L3VPN router bgp 65001 neighbor 10.10.10.10 remote-as 65001 neighbor 10.10.10.10 update-source Loopback0 address-family vpnv4 unicast neighbor 10.10.10.10 activate neighbor 10.10.10.10 send-community extended exit-address-family address-family ipv4 vrf CUSTOMER-A redistribute connected ! If using BGP PE-CE: neighbor 192.168.1.2 remote-as 65100 neighbor 192.168.1.2 activate exit-address-family ``` ## Verification ``` ! VRF routing table PE# show ip route vrf CUSTOMER-A ! VRF VPN routes received from other PEs PE# show bgp vpnv4 unicast all summary ! Specific prefix details PE# show bgp vpnv4 unicast all 10.0.0.0/24 ! Forwarding for a specific prefix in a VRF PE# show ip cef vrf CUSTOMER-A 10.0.0.5 ! L3VPN label bindings PE# show mpls forwarding-table vrf CUSTOMER-A ``` ## Anti-Patterns - **Same RD on multiple PEs for the same VRF.** Causes BGP best-path issues for ECMP. Use different RDs (e.g. include the PE's loopback in the RD). - **Forgetting `send-community extended`.** RTs do not propagate; route import/export breaks. - **RTs not symmetric.** Easy to typo and end up with import 65001:100 but export 65001:1000\. Always verify with `show vrf detail`. - **Mixing classic VRF and vrf-definition.** Both work but operationally confusing. Standardize on vrf-definition. - **Skipping LDP-IGP sync.** Transient drops on link bring-up; broken L3VPN until LDP catches up. ## Summary MPLS L3VPN combines VRFs (per-customer routing instances), Route Distinguishers (make overlapping prefixes globally unique), Route Targets (control import/export), MP-BGP (carries VPNv4 routes between PEs), and the two-label MPLS stack (outer transport, inner VPN) to deliver private routing services over a shared backbone. The architecture is more complex than a single concept can capture, but each piece has a single clear job. Master the RD/RT distinction, the two-label stack, and the MP-BGP send-community-extended requirement, and L3VPN troubleshooting becomes tractable. Bookmark this article alongside the [MPLS cluster pillar](https://www.pinglabz.com/mpls/), the [MP-BGP article](https://www.pinglabz.com/mp-bgp-multiprotocol-bgp/), and the [MPLS labels article](https://www.pinglabz.com/mpls-labels-explained/). ### References - [RFC 4271 - A Border Gateway Protocol 4 (BGP-4)](https://www.rfc-editor.org/rfc/rfc4271?ref=pinglabz.com) - [Cisco BGP technology documentation](https://www.cisco.com/c/en/us/tech/ip/border-gateway-protocol-bgp/index.html?ref=pinglabz.com) Take the BGP reference with you The free BGP field-reference PDF: path attributes, best-path order, and the show commands that matter. Delivered by email, no card required. [Get the free PDF](https://www.pinglabz.com/bgp-cheatsheet/) ### LDP and MPLS Label Distribution URL: https://www.pinglabz.com/ldp-mpls-label-distribution/ Last updated: 2026-08-01T19:35:35.000Z LDP (Label Distribution Protocol) is what makes MPLS work for IP forwarding. It is the protocol routers use to tell each other what label to use for each prefix. Every LSR (Label Switching Router) in an LDP-driven MPLS network runs an LDP session with each direct neighbor, exchanges label-to-prefix mappings, and builds the forwarding state that lets MPLS packets traverse the network. This article walks through how LDP discovers neighbors, how sessions form, label distribution modes (DU vs DoD), label retention modes, the LDP-IGP synchronization story, and the Cisco IOS XE configuration. If you are configuring MPLS for the first time, troubleshooting an LDP session, or wondering why `show mpls ldp bindings` shows what it shows, this is the reference. ## What LDP Does For each IP prefix in the IGP, LDP assigns a local label and tells neighbors "for prefix X, use label Y to send to me." The receiving routers store this mapping in the LDP binding table. The forwarding-table builder (CEF on Cisco) uses the bindings to create the actual MPLS forwarding state. Two key terms: - **FEC (Forwarding Equivalence Class):** a group of packets treated identically. For LDP-driven MPLS, each IP prefix is its own FEC. - **LSP (Label Switched Path):** the sequence of LSRs and labels a packet traverses. Built up implicitly as each LSR maps its locally-assigned labels to the next-hop's labels. LDP is "IGP-driven" - it does not pick paths itself. The IGP (OSPF, IS-IS, EIGRP) decides which neighbor is the next-hop for each prefix. LDP simply assigns labels along the IGP-computed paths. ## Neighbor Discovery LDP has two phases for finding neighbors: Discovery Mechanism Hello packets announcing presence Transport UDP/646 multicast (224.0.0.2 / FF02::2) Session establishment Mechanism TCP session for label exchange TransportTCP/646 unicast Each LSR sends LDP Hellos out every MPLS-enabled interface every 5 seconds (default). When two LSRs hear each other's Hellos, they establish a TCP session on the higher-numbered router ID. The LDP TCP session is used to exchange label mappings and ongoing updates. The LDP router-id is typically the highest loopback IP, mirroring how OSPF picks its router-id. All LDP sessions for a given LSR use the same router-id; multiple sessions run over the same TCP connection (one connection per neighbor). ## Label Distribution Modes LDP supports two modes for advertising labels: Downstream Unsolicited (DU) Behavior Each LSR proactively advertises labels for its prefixes to all neighbors Default on CiscoYes Downstream-on-Demand (DoD) Behavior LSRs only advertise labels in response to a neighbor's Label Request Default on CiscoNo (rare in practice) DU is what you see in production. Every LSR floods label mappings; receivers store them; the forwarding table populates rapidly. DoD is mainly for ATM-based MPLS (effectively obsolete) and has stricter signaling but tighter coupling. ## Label Retention Modes Liberal Retention Keep all label bindings from all neighbors, even ones not on the IGP best path Conservative Retention Only keep bindings from the IGP next-hop neighbor for each prefix Liberal retention is the Cisco default. The benefit: when the IGP fails over to a different next-hop, MPLS forwarding state is already populated for the new path - no waiting for label exchange. The cost: more memory used for label bindings. For typical service-provider deployments, liberal retention is correct. The memory cost is negligible; the convergence benefit is significant. ## LDP-IGP Synchronization One of the trickiest production issues: an interface comes up, the IGP advertises new routes immediately, but LDP takes a few seconds to assign and exchange labels. During that gap, the IGP forwards packets via the new interface but MPLS has no label - packets get dropped (LDP IP forwarding is not configured for fallback) or forwarded as plain IP (which breaks any LSP-dependent service like L3VPN). LDP-IGP synchronization solves this. With sync enabled, the IGP cost on an interface is held at maximum (effectively infinite) until LDP has formed a session and exchanged labels with the neighbor. Once LDP is ready, the IGP cost reverts to normal and traffic shifts to the new path. Configuration: ``` ! On Cisco IOS XE mpls ldp sync ! Enable globally for all OSPF interfaces ! Or per-interface interface GigabitEthernet0/0/0 mpls ldp sync ``` Verify with `show mpls ldp igp sync`. Always enable LDP-IGP sync in production MPLS deployments. The cost is zero; the protection against transient drops is real. ## Targeted LDP Standard LDP discovers neighbors via Hellos on directly-connected interfaces. Targeted LDP establishes sessions with non-adjacent LSRs (e.g. for L2VPN pseudowires that span multiple hops). ``` ! Configure a targeted session to a remote LSR mpls ldp neighbor 10.10.10.10 targeted ``` Used primarily for L2VPN services (VPWS / pseudowires) where the two endpoints of a pseudowire need direct LDP signaling regardless of how many hops separate them. ## Configuration on Cisco IOS XE Minimum LDP configuration: ``` ! Global MPLS configuration mpls ldp router-id Loopback0 force ! Use loopback as LDP router-id mpls ldp sync ! IGP sync for all MPLS interfaces ! Per-interface interface GigabitEthernet0/0/0 mpls ip ! Enable MPLS forwarding mpls ldp sync ! Per-interface sync (optional if global) mpls label protocol ldp ! Use LDP (vs the older TDP) ``` That is it. Three lines per interface, two global lines. Once enabled, LDP discovers neighbors automatically and starts exchanging labels. For the comprehensive walkthrough, see [Cisco MPLS Configuration on IOS XE](https://www.pinglabz.com/cisco-mpls-configuration/). ## Verification ``` ! LDP neighbors Router# show mpls ldp neighbor Peer LDP Ident: 2.2.2.2:0; Local LDP Ident 1.1.1.1:0 TCP connection: 2.2.2.2.13456 - 1.1.1.1.646 State: Oper; Msgs sent/rcvd: 1234/1234; Downstream Up time: 02:30:15 LDP discovery sources: GigabitEthernet0/0/0, Src IP addr: 10.0.12.2 ! Label bindings Router# show mpls ldp bindings lib entry: 10.10.10.0/24, rev 100 local binding: label: 17 remote binding: lsr: 2.2.2.2:0, label: 18 remote binding: lsr: 3.3.3.3:0, label: 19 ! LDP-IGP sync state Router# show mpls ldp igp sync GigabitEthernet0/0/0: LDP configured; SYNC enabled SYNC status: sync achieved; peer reachable ``` The bindings table shows local and remote labels. For prefix 10.10.10.0/24, this LSR uses label 17 locally; if it forwards to 2.2.2.2, it pushes label 18; if it forwards to 3.3.3.3, it pushes label 19\. CEF picks the right label based on which next-hop the IGP chose. ## Authentication LDP supports MD5 authentication for the TCP session: ``` mpls ldp neighbor 2.2.2.2 password Cisco123! ``` Both ends must agree on the password. Authentication prevents an attacker from injecting fake LDP messages that could manipulate the label table. Always enable in production. ## Troubleshooting Common LDP failure modes: No LDP neighbor on directly-connected interface `mpls ip` not enabled on one or both interfaces; authentication mismatch Session bouncing Underlying link instability; MTU mismatch on the LDP TCP session No labels for some prefixes Prefix not in IGP; or LDP filter applied that suppresses Traffic forwarded unlabeled LDP not enabled on the egress interface; or PHP triggering as expected Drops on link bring-up LDP-IGP sync not configured; IGP advertises before LDP has labels Universal first command: `show mpls ldp neighbor`. State should be Oper. If not, work backwards through MPLS interface config, IP reachability, and authentication. When the session shows Oper and some prefixes still have no label, the phrase that names the fault is `unusable: no label`, and every IP test will keep passing for as long as it is true. [Tracing where along the path the label stops](https://www.pinglabz.com/troubleshooting-mpls-ldp/) works that case bottom-up on live captures. ## LDP vs Segment Routing LDP is being phased out in modern service-provider designs in favor of Segment Routing (SR-MPLS). SR uses the same MPLS data plane but distributes labels via the IGP itself - no separate LDP protocol. This eliminates the LDP-IGP synchronization problem entirely (because the IGP IS the label distribution). Label distribution LDPSeparate LDP protocol SR-MPLS IGP extensions (OSPF/IS-IS) State per prefix LDP Stored in LDP binding table SR-MPLS Algorithmic from IGP segment Sync issues LDPNeed LDP-IGP sync SR-MPLS None (IGP is the distribution) Operational complexity LDP Higher (separate protocol) SR-MPLS Lower (one less protocol) Fast reroute LDP Possible with extensions SR-MPLSNative via TI-LFA For the full LDP-vs-SR comparison and migration story, see the MPLS pillar's segment routing section. ## Summary LDP distributes labels across the MPLS network so each LSR knows which label to use for each prefix. It runs over UDP/646 for discovery and TCP/646 for session establishment. The dominant production mode is Downstream Unsolicited with Liberal Retention; LDP-IGP synchronization is mandatory to prevent transient drops on link bring-up. Configuration is minimal (a few commands per LSR), and once running LDP populates the forwarding state automatically. Modern designs are migrating to Segment Routing, which eliminates LDP entirely, but LDP remains the dominant label distribution protocol in production MPLS networks today. Bookmark this article alongside the [MPLS cluster pillar](https://www.pinglabz.com/mpls/) and the [MPLS labels article](https://www.pinglabz.com/mpls-labels-explained/). ### MPLS Labels Explained: Format, Stacking, and Penultimate Hop Popping URL: https://www.pinglabz.com/mpls-labels-explained/ Last updated: 2026-06-13T20:08:38.000Z MPLS labels are 32-bit shim headers inserted between the data link layer and the IP header. They carry the label value, traffic class (formerly EXP), the bottom-of-stack bit, and TTL. Multiple labels can be stacked. If you understand the label format and how the stack works, the rest of MPLS becomes mechanics. This article walks through the byte-level format, the label stack, special reserved labels, Penultimate Hop Popping (PHP), and how labels look on the wire in real packet captures. If you are studying for CCIE Service Provider, troubleshooting an MPLS-VPN packet, or trying to figure out why `show mpls forwarding-table` shows what it shows, this is the byte reference. ## The Label Format ``` +---------------------+-----+---+-------------+ | Label | EXP | S | TTL | | 20 bits | 3 | 1 | 8 bits | +---------------------+-----+---+-------------+ ``` Label Bits20 Purpose The label value, 0 to 1,048,575\. Locally significant per LSR. EXP / Traffic Class Bits3 Purpose QoS priority; carries equivalent of DSCP top 3 bits across MPLS S (Bottom of Stack) Bits1 Purpose 1 if this is the bottom label; 0 if more labels follow below TTL Bits8 Purpose Hop count, decremented at each LSR; mirrors IP TTL Total 32 bits = 4 bytes per label. The shim header is inserted in the frame between the Layer 2 header (Ethernet, etc.) and the Layer 3 payload (IP). ## Reserved Labels The first 16 label values (0-15) are reserved for special meanings. Three matter in practice: IPv4 Explicit Null Label0 Use Tells the egress LSR to pop the label and forward as IPv4\. Used in PHP scenarios where the egress wants to preserve EXP/TC info. Router Alert Label1 Use Treats the packet specially (like IP Router Alert). Sent to control plane for inspection. IPv6 Explicit Null Label2 Use Same as label 0 but for IPv6. Implicit Null Label3 Use Used in LDP signaling to request PHP. The penultimate hop pops the label entirely instead of swapping. OAM Alert Label13 Use Operations and management traffic Generic Alert Label14 Use Reserved for new features Labels 16 to 1,048,575 are available for use as forwarding labels. Different platforms may reserve some additional ranges (e.g. Cisco IOS reserves the first \~16,000 for specific features), so customer-installed labels typically start at 16,000 or higher. ## The Label Stack Multiple labels can be pushed onto a single packet. The stack is a sequence of label entries with the S bit indicating which is the bottom. Reading from outer to inner: ``` +--------+--------+--------+---------+ | Outer | Middle | Inner | IP/payld| | S=0 | S=0 | S=1 | | +--------+--------+--------+---------+ First popped or swapped Last popped, payload exposed ``` Common scenarios: IP/MPLS forwarding StackSingle label Outer label purposeLDP/IGP transport Inner label purposen/a L3VPN (VPNv4) StackTwo labels Outer label purpose LDP transport across MPLS core Inner label purpose VPN label identifying customer VRF L3VPN with TE StackThree labels Outer label purposeRSVP-TE TE label Inner label purpose LDP transport (intermediate), VPN label (innermost) L2VPN (pseudowire) StackTwo labels Outer label purposeLDP transport Inner label purpose Pseudowire label identifying the L2 circuit SR-MPLS StackVariable Outer label purposeSegment list (path) Inner label purposeService label or none Each LSR along the path looks at the outermost label, performs swap/pop/push as configured, and forwards. P routers in the core typically only swap the outer transport label. The inner labels are invisible to them. ## Label Actions: Push, Swap, Pop Each LSR can perform three operations on the label stack: Adds one or more labels to the top of the stack ActionPush Where it happens Ingress PE (encapsulating customer traffic) Replaces the outermost label with a new one ActionSwap Where it happens Every P router along the LSP Removes the outermost label ActionPop Where it happens Egress PE (decapsulating); also penultimate hop in PHP The forwarding table entry tells the LSR which action to perform for each (incoming label, incoming interface) tuple. Verify with: ``` R1# show mpls forwarding-table Local Outgoing Prefix Bytes Label Outgoing Next Hop Label Label or Tunnel Id Switched interface 17 18 10.10.10.0/24 123456 Gi0/0/1 10.0.12.2 18 Pop Label 10.20.20.0/24 78901 Gi0/0/2 10.0.13.3 19 No Label 10.30.30.0/24 45678 Gi0/0/0 10.0.14.4 ``` "Pop Label" means PHP - this LSR pops the label and forwards the unlabeled packet. "No Label" means the next hop is not running MPLS for this prefix. ## Penultimate Hop Popping (PHP) The egress PE has the work of looking up which VRF a customer packet belongs to (based on the inner VPN label) and forwarding it to the customer. If it also had to look up the outer transport label, that is two label lookups instead of one. Penultimate Hop Popping optimizes this. The penultimate (second-to-last) LSR pops the outer label before forwarding to the egress PE. The egress PE receives a packet with only the inner VPN label, performs one label lookup, and forwards to the customer. Signaling: the egress PE advertises label "implicit null" (label 3) to its upstream neighbors via LDP. This is a request to pop the label rather than swap it. The penultimate router sees implicit null in the LDP table and pops accordingly. PHP is enabled by default in most Cisco IOS XE deployments. Verify with: ``` P1# show mpls forwarding-table 10.10.10.10 32 Local Outgoing Prefix Bytes Label Outgoing Next Hop Label Label or Tunnel Id Switched interface 17 Pop Label 10.10.10.10/32 123456 Gi0/0/2 10.0.PE2.PE2 ``` The "Pop Label" outgoing label confirms PHP for the egress PE's loopback (10.10.10.10/32). Disable PHP if you want the egress PE to see the outer label (e.g. to preserve EXP/TC bits across the boundary): ``` ! On the egress PE mpls ldp explicit-null ``` This advertises "explicit null" (label 0 for IPv4, label 2 for IPv6) instead of implicit null. The penultimate router still pops one level of stack, but the egress PE sees an explicit null label - useful for QoS preservation. ## TTL Handling The MPLS label has its own TTL. Two modes for handling it at ingress and egress: Uniform (default) Ingress behavior Copy IP TTL to MPLS label TTL minus 1 Egress behavior Copy MPLS TTL back to IP TTL Pipe Ingress behaviorSet MPLS TTL to 255 Egress behavior Decrement IP TTL by 1 (the entire MPLS path counts as one hop) Uniform mode lets traceroute work as expected through the MPLS network - each LSR appears as a hop. Pipe mode hides the MPLS topology - the customer sees one big "hop" across the entire MPLS network, which is what service providers usually want for security reasons. Configure with: ``` ! Pipe mode (hide MPLS topology) mpls ip propagate-ttl no ! Uniform mode (default; show MPLS topology) mpls ip propagate-ttl ``` ## EtherType and Encapsulation On Ethernet, an MPLS-encapsulated frame has: EtherType (unicast) 0x8847 EtherType (multicast) 0x8848 An Ethernet frame with EtherType 0x8847 carries one or more MPLS labels followed by the payload (typically IP). The receiver parses the labels until S=1, then treats the remainder as the underlying protocol. ## Show Commands ``` ! All MPLS forwarding entries Router# show mpls forwarding-table ! For a specific prefix Router# show mpls forwarding-table 10.10.10.0 24 ! LDP-assigned labels per prefix Router# show mpls ldp bindings ! Per-interface MPLS state Router# show mpls interfaces ! IP-to-label mapping in CEF Router# show ip cef 10.10.10.0/24 10.10.10.0/24, version 123, ... via 10.0.12.2, GigabitEthernet0/0/0 push label 17, ... ``` The CEF output shows the label being pushed for each prefix. This is what actually drives forwarding. LDP populates the binding table; CEF uses the bindings to populate the forwarding-information base; packets follow the FIB. ## In a Packet Capture An MPLS-labeled frame in Wireshark shows as: ``` Ethernet II, Src: ..., Dst: ..., Type: MPLS (0x8847) MPLS Label, Exp: 0, S: 0, TTL: 64 <-- outer MPLS Label: 17 MPLS Label, Exp: 0, S: 1, TTL: 64 <-- inner (S=1 = bottom) MPLS Label: 1234 Internet Protocol Version 4, Src: ..., Dst: ... ``` Wireshark parses the label stack automatically. The S=1 on the inner label terminates the stack; everything after is parsed as IP (or whatever the inner protocol indicates). ## Summary MPLS labels are 32-bit shim headers (20-bit label, 3-bit EXP/TC, 1-bit S, 8-bit TTL) that sit between the data link header and the payload. Stacks of labels enable services like L3VPN (two-label stack) and traffic engineering. Penultimate Hop Popping optimizes egress PE work by having the second-to-last LSR pop the outer label. Master the label format, the three actions (push/swap/pop), the reserved labels (especially implicit null for PHP), and the stack model. The rest of MPLS is mechanics built on these foundations. Bookmark this article alongside the [MPLS cluster pillar](https://www.pinglabz.com/mpls/) and the [LDP article](https://www.pinglabz.com/ldp-mpls-label-distribution/). ### EIGRP Stub Routing for Hub-and-Spoke URL: https://www.pinglabz.com/eigrp-stub-routing/ Last updated: 2026-06-13T20:08:38.000Z EIGRP stub routing is the feature that makes hub-and-spoke topologies sane. Without it, a hub router with hundreds of spokes queries every spoke during DUAL active states, propagating control-plane traffic that the spokes cannot meaningfully reply to. With it, spokes are explicitly told they are dead-ends - the hub does not query them and they do not advertise transit routes. This article walks through what stub routing does, the five stub options and what each advertises, when to use stubs, the impact on DUAL queries, and the configuration patterns. If you are designing or operating a hub-and-spoke EIGRP network, this is the feature that determines whether your hub router has a peaceful day or a constant stream of SIA events. ## What Stub Routing Does EIGRP stub routing modifies two behaviors at the spoke router: 1. **Restricts what the spoke advertises.** The spoke only sends specific route types (configurable). It does not act as a transit router for routes learned from other neighbors. 2. **Tells the hub not to query the spoke during DUAL active states.** The hub knows the spoke cannot offer alternative paths, so queries are wasted. The combined effect: a hub-and-spoke network with N spokes does not generate N queries when a route enters Active state. Convergence is faster; the hub CPU stays sane; SIA events become rare. ## The Problem Without Stub Imagine a hub router with 200 spokes. Each spoke is a branch with two interfaces: one to the hub and one to local users. The hub knows every spoke's local networks via EIGRP. Now a route somewhere in the broader network changes and the hub loses a path. If no Feasible Successor is cached, the route enters Active state and the hub queries all neighbors - including all 200 spokes. Each spoke receives the query. The spokes have no useful answer (they have one path: through the hub). Each spoke sends a "no match" reply. The hub processes 200 replies. Meanwhile any spoke that takes too long to reply triggers SIA logic. Multiply this by every flapping route across the network. The hub's CPU pegs; legitimate traffic suffers; the network feels unstable. Stub routing eliminates this problem. The hub knows from the start that spokes have no useful information for queries; it does not ask. ## The Five Stub Options `connected` Directly connected networks `summary` Summary routes (manual or auto-summarization) `static` Static routes redistributed into EIGRP `redistributed` Routes redistributed from other protocols `receive-only` Nothing - listens but does not advertise anything The defaults if you specify `eigrp stub` with no options: `connected summary`. Most production hub-and-spoke deployments use exactly this default - the spoke advertises its directly-connected branch networks (and any summary routes you have configured) and nothing else. The five options are not mutually exclusive; you can combine them. `eigrp stub connected static` advertises both directly-connected and static routes. ## receive-only: The Strictest Form The receive-only option is special. It tells the spoke to listen for EIGRP updates but advertise nothing - not even its directly-connected networks. The spoke is purely a downstream consumer of routes. Use cases: - The hub needs to push a default route to the spoke; the spoke needs no upstream visibility - Branch sites with no locally-originated services (every server lives at HQ) - Read-only environments where local routes must not propagate Less common than `connected summary` but useful for specific designs. ## Configuration ``` ! Classic mode on the spoke router eigrp 100 network 10.0.0.0 0.0.255.255 eigrp stub connected summary no auto-summary ! Named mode on the spoke router eigrp PROD address-family ipv4 unicast autonomous-system 100 network 10.0.0.0 0.0.255.255 eigrp stub connected summary exit-address-family ``` That is it. One line declares the router as a stub. The hub automatically detects this via the EIGRP Hello packet's K-value-and-stub-flag and adjusts its query behavior. No configuration on the hub is required to make stub routing work. The hub reads the spoke's stub flag from Hello packets and updates its query lists accordingly. ## Verification On the spoke: ``` Spoke# show ip eigrp neighbors detail EIGRP-IPv4 Neighbors for AS(100) H Address Interface Hold Uptime SRTT RTO Q Seq Num 0 10.0.0.1 Gi0/0/0 14 01:23:45 10 200 0 234 Stub Peer Advertising (CONNECTED, SUMMARY) Routes Suppressing queries ``` The "Stub Peer" line confirms the spoke is acting as a stub. "Suppressing queries" confirms the hub will not query this neighbor during DUAL active states. On the hub looking at a spoke: ``` Hub# show ip eigrp neighbors detail EIGRP-IPv4 Neighbors for AS(100) H Address Interface Hold Uptime SRTT RTO Q Seq Num 0 10.0.10.2 Gi0/0/0 14 01:23:45 10 200 0 234 Stub Peer Advertising (CONNECTED, SUMMARY) Routes Suppressing queries ``` Same output - the stub status is symmetric and visible from both sides. ## Design Patterns Three stub patterns dominate production: **1\. Standard branch.** `eigrp stub connected summary`. Branch advertises its local networks (and any summaries) and nothing else. The hub treats the branch as a leaf. Most common pattern. **2\. Branch with locally-redistributed routes.** `eigrp stub connected summary redistributed`. Branch advertises its locals plus any routes redistributed from another protocol (e.g. a local OSPF segment, or static routes for legacy systems). Used when the branch has connectivity to non-EIGRP infrastructure. **3\. Read-only branch.** `eigrp stub receive-only`. Branch listens but advertises nothing. Used for sites with no local services or strict read-only requirements. ## Anti-Patterns - **Stub routing on transit routers.** If two spokes need to reach each other through a hub, both being stubs prevents the path. Stubs do not transit traffic. If a router needs to be transit, it cannot be a stub. - **Forgetting stub on new branches.** When you add a new spoke to an existing hub-and-spoke EIGRP network, configure stub routing immediately. Forgetting it leaves the spoke in transit mode and the hub queries it normally. - **Mixed stub and non-stub spokes.** Confusing operationally. Standardize - either all spokes are stubs or none. - **Stub routing in a full mesh.** Stubs assume hub-and-spoke. In full mesh, every router is a transit router; stubs break the topology. ## Stub Routing in DMVPN DMVPN (Dynamic Multipoint VPN) is a common hub-and-spoke implementation that combines mGRE tunnels with NHRP and a routing protocol underneath. EIGRP stub routing is the standard companion: ``` ! On a DMVPN spoke interface Tunnel0 ip address 10.0.0.10 255.255.255.0 ip nhrp network-id 100 ip nhrp nhs 10.0.0.1 tunnel mode gre multipoint tunnel source GigabitEthernet0/0/0 tunnel destination 10.0.0.1 router eigrp 100 network 10.0.0.0 0.0.0.255 network 192.168.10.0 0.0.0.255 eigrp stub connected summary no auto-summary ``` The spoke advertises its tunnel interface and local LAN; the hub treats it as a stub; query suppression applies. Standard pattern. For spoke-to-spoke direct tunnels (NHRP redirect), the spokes are still stubs from EIGRP's perspective - the direct tunnel does not change the IGP topology, just the data plane. ## Alternatives to Stub Routing Other ways to limit query propagation: Manual summarization Reduces routes; queries on summary scope only Distribute lists Filter what is advertised; not as clean as stub flag Stub routing (recommended) Explicitly signals stub status; hub adjusts query behavior Stub routing is the canonical solution because it integrates with DUAL's query logic. Other mechanisms can reduce route counts but do not stop the hub from querying. ## Summary EIGRP stub routing is the difference between a hub-and-spoke network that scales and one that does not. By telling spokes they are dead-ends, you save the hub from cascading queries during DUAL active states, eliminate Stuck-In-Active risk, and converge faster. For most production hub-and-spoke designs (DMVPN, branch-to-HQ, retail chains, anything with a few large hubs and many smaller branches), `eigrp stub connected summary` on every spoke is the right configuration. Bookmark this article alongside the [EIGRP cluster pillar](https://www.pinglabz.com/eigrp/) and the [DUAL deep-dive](https://www.pinglabz.com/eigrp-dual-algorithm/). ### EIGRP vs OSPF: When to Use Each URL: https://www.pinglabz.com/eigrp-vs-ospf/ Last updated: 2026-07-04T23:21:54.000Z EIGRP vs OSPF is the classic IGP comparison every CCNP candidate works through. Both are interior gateway protocols. Both converge in sub-second timeframes when tuned. Both work well in production at enterprise scale. The differences are real but smaller than the marketing material suggests, and the right answer depends mostly on what is already in your network. This article walks through the technical differences, the operational trade-offs, when each is the right answer in 2026, and what to do when the decision is "we already run one of them and have to keep going." If you are studying for a certification, designing a new network, or trying to justify an IGP choice to a manager, this is the comparison. ## The TL;DR OSPF is the safer enterprise default in 2026\. Vendor-neutral, well-understood, expected on every certification track, scales hierarchically via areas. EIGRP is faster on healthy networks (cached feasible successors), simpler to configure, and works well in Cisco-only environments - but commits you to Cisco for the IGP across that domain. If you are starting greenfield with no vendor preference: OSPF. If you are entirely Cisco and convergence speed matters: EIGRP is fine. If you have one or the other already deployed and working: keep going. ## The Side-by-Side Comparison Algorithm EIGRP Advanced distance-vector (DUAL) OSPF Link-state (Dijkstra SPF) Standards EIGRP RFC 7868 (basic), Cisco-led OSPFRFC 2328 (open) Vendor support EIGRP Effectively Cisco-only in production OSPF Universal (Cisco, Juniper, Arista, Nokia, etc.) Default Cisco AD EIGRP 90 internal / 170 external OSPF110 Transport EIGRP IP protocol 88, multicast 224.0.0.10 OSPF IP protocol 89, multicast 224.0.0.5/6 Metric EIGRP Composite (bandwidth + delay default; +load, +reliability optional) OSPF Cost (16-bit, bandwidth-derived) Convergence on healthy network EIGRP Sub-second when FS exists (cached) OSPFSub-second with tuning Convergence in worst case EIGRP Slower (queries propagate) OSPF Bounded by SPF complexity Topology visibility EIGRPPer-prefix (limited) OSPFFull LSDB within area Hierarchy EIGRP None native; stub feature for hub-and-spoke OSPF Strict areas with backbone rule Configuration complexity EIGRP Lower (no areas; flat by default) OSPF Higher (area design matters) Authentication EIGRPMD5 / HMAC-SHA OSPFMD5 / SHA Multicast address EIGRP 224.0.0.10 (IPv4) / FF02::A (IPv6) OSPF 224.0.0.5 / 224.0.0.6 (IPv4); FF02::5 / FF02::6 (IPv6) Loop prevention EIGRP Feasibility condition (RD < FD) OSPF Inherent in link-state (full LSDB) ## Convergence: A Closer Look The convergence story is more nuanced than "EIGRP is faster" or "OSPF is faster." **EIGRP best case:** a feasible successor is cached in the topology table. Failover is instant - no recomputation, no queries. Sub-second; effectively immediate after the link-down detection. **EIGRP worst case:** no feasible successor exists. The route enters Active state. The router queries every neighbor (except stubs). Queries can propagate across the topology. Convergence takes seconds; can be much longer in pathological cases (Stuck-In-Active). **OSPF best case:** a topology change triggers SPF. SPF runs across the area's LSDB. Convergence is bounded by SPF complexity (typically O(N log N) for N nodes). With BFD and tuned timers, sub-second is achievable. **OSPF worst case:** a topology change in a large area triggers SPF on every router. The bigger the area, the longer SPF takes. This is why area design matters - keep areas small enough that SPF is cheap. The practical takeaway: EIGRP wins on the typical case (small topology change with feasible alternate available); OSPF is more predictable across the range of failure modes. For most enterprise networks, both deliver acceptable convergence. Tune timers and add BFD; both protocols will hit sub-second. ## Design Implications EIGRP is flat by default. There is no area concept; every router exchanges routes with every neighbor. The stub feature limits query propagation in hub-and-spoke designs but does not provide hierarchical scaling the way OSPF areas do. OSPF demands area design. Every multi-area deployment has Area 0 (backbone) plus non-backbone areas connecting through it. Inter-area summarization happens at ABRs. Stub area types limit what LSAs propagate where. Done well, this scales OSPF to thousands of routers; done badly, it creates traffic blackholes and routing bugs. Small (50 routers) EIGRP designFlat AS; trivial OSPF designSingle area; trivial Medium (500 routers) EIGRP design Flat AS or hub-spoke with stubs OSPF design 2-3 areas with sane summarization Large (5000+ routers) EIGRP design Pushes EIGRP's limits; designs vary OSPF design Multi-area with proper boundary discipline For very large networks, OSPF (or IS-IS at internet scale) wins on hierarchical scaling. EIGRP's flat model gets challenging as the topology grows. ## Metric: Composite vs Cost EIGRP's composite metric (bandwidth + delay by default) handles diverse interface speeds well. A path through a 1 Gbps link plus a 100 Mbps link prefers the link with the higher bandwidth, weighted by delay accumulation. OSPF's cost is bandwidth-derived (10^8 / interface-bandwidth by default; tunable via reference-bandwidth). It is simpler but saturates at high speeds without the wide reference-bandwidth setting. Practical implication: both work fine for typical enterprise designs. EIGRP's metric handles odd interface mixes (T1 alongside Gigabit) more gracefully out of the box. OSPF requires setting reference-bandwidth high enough to differentiate fast links. ## Vendor Reality EIGRP was Cisco-proprietary from 1992 to 2013\. RFC 7868 opened the basic protocol, but practical adoption beyond Cisco remains rare. A few Linux implementations exist (Open EIGRP); some enterprise routers from other vendors include partial EIGRP support; in production, EIGRP is essentially Cisco-only. OSPF is genuinely vendor-neutral. Cisco, Juniper, Arista, Nokia, Huawei, and most enterprise routing vendors implement it well. Mixed-vendor deployments work in practice without much fuss. If your network has any non-Cisco routing infrastructure, OSPF is the answer. If it never will, EIGRP is fine. ## When to Use Each **OSPF when:** - You have or plan to have any non-Cisco routing infrastructure - You are studying for CCNP/CCIE (OSPF is the dominant exam topic) - You expect to scale beyond a few hundred routers - You want the safer enterprise default - You are designing a new network with no incumbent IGP **EIGRP when:** - Your network is entirely Cisco and likely to stay that way - You have hub-and-spoke topology where stub routing simplifies design - You value sub-second convergence on healthy networks (cached FS) - You inherited an EIGRP deployment that works and rip-and-replace is not justified - You want simpler configuration without area design **Both at the same time?** Yes, occasionally. Some networks run OSPF in the core and EIGRP at the edges, redistributing between them. This adds complexity and is rarely justified for new deployments. Common in long-running networks where different parts were built at different times by different teams. ## Redistributing Between EIGRP and OSPF If you run both, redistribution is needed. The pattern: ``` ! On the boundary router running both router eigrp 100 redistribute ospf 1 metric 1000000 100 255 1 1500 ... router ospf 1 redistribute eigrp 100 subnets metric-type 1 ... ``` Two things to watch for: - **Mutual redistribution loops.** If redistribution happens at multiple boundary routers, routes can loop between protocols. Use route maps to filter and tag routes by origin. - **Metric translation.** EIGRP's composite metric and OSPF's cost are not directly comparable. Set explicit metrics on redistribution to control how the receiving protocol weights the route. For deep redistribution patterns, see Cisco's documentation on multi-protocol routing. The PingLabz position: avoid redistribution where you can; consolidate to one IGP per domain. ## Both Run Alongside BGP Whichever IGP you pick, you almost always run BGP at the AS edge. The IGP carries internal routes (loopbacks, infrastructure links, internal subnets); BGP carries the internet-facing or inter-AS routes. They do not overlap on the same prefix; redistribution between them is rare and should be carefully scoped. See [BGP vs OSPF](https://www.pinglabz.com/bgp-vs-ospf/) for the IGP/BGP layering pattern. The same pattern applies if you swap OSPF for EIGRP underneath. ## Summary EIGRP and OSPF both work in production. Pick OSPF for vendor neutrality, hierarchical scale, and certification expectations. Pick EIGRP for Cisco-only convergence speed and operational simplicity. Both deliver sub-second failover when tuned; both scale to enterprise networks of thousands of routers (with appropriate design). The decision usually comes down to what is already there and what your operators know. New networks default to OSPF in 2026\. Cisco-only shops with EIGRP deployed and working keep going. Bookmark this article alongside the [EIGRP cluster pillar](https://www.pinglabz.com/eigrp/), the [OSPF cluster pillar](https://www.pinglabz.com/ospf/), and the [BGP vs OSPF](https://www.pinglabz.com/bgp-vs-ospf/) piece for the inter-AS half of the story. New labs and guides, in your inbox Every new PingLabz lab and deep-dive, built and verified on real Cisco IOS XE - free, straight to your inbox. [Join free](https://www.pinglabz.com/signup/) ### EIGRP Configuration on Cisco IOS XE: Classic and Named Mode URL: https://www.pinglabz.com/eigrp-configuration-cisco/ Last updated: 2026-06-13T20:08:39.000Z Configuring EIGRP on Cisco IOS XE is straightforward once you understand the two configuration modes (classic and named) and a few production patterns that should be standard. This article walks through the minimum viable config, the differences between classic and named mode, the network statement, passive interfaces, summarization, authentication, and the verification commands that close the loop. If you are configuring EIGRP for the first time, migrating from classic to named mode, or auditing an inherited configuration for production hygiene, this is the operator's walkthrough. ## Classic Mode: The Legacy Pattern Classic-mode EIGRP is what you see in most existing Cisco deployments. It is what the CCNA exam tests on and what older documentation assumes: ``` R1(config)# router eigrp 100 R1(config-router)# network 10.0.0.0 0.0.255.255 R1(config-router)# network 192.168.1.0 0.0.0.255 R1(config-router)# passive-interface default R1(config-router)# no passive-interface GigabitEthernet0/0/0 R1(config-router)# no passive-interface GigabitEthernet0/0/1 R1(config-router)# no auto-summary R1(config-router)# eigrp router-id 1.1.1.1 ``` Three things to notice: - **AS number (100) is locally significant in EIGRP.** All neighbors must use the same AS number to form adjacency, but you choose any number 1-65535. - **Wildcard mask in `network`** is inverted from a regular subnet mask. `0.0.255.255` matches `/16`; `0.0.0.255` matches `/24`; `0.0.0.0` matches a single address. - `**passive-interface default**` followed by selective `no passive-interface` is the safe pattern - prevents accidentally forming EIGRP adjacencies on user-facing interfaces. ## Named Mode: The Modern Pattern IOS 15.x introduced named-mode EIGRP. It cleanly separates IPv4 and IPv6 address families, supports multi-AS deployments, and is what new Cisco labs and exams expect: ``` R1(config)# router eigrp PROD R1(config-router)# address-family ipv4 unicast autonomous-system 100 R1(config-router-af)# network 10.0.0.0 0.0.255.255 R1(config-router-af)# af-interface default R1(config-router-af-interface)# passive-interface R1(config-router-af-interface)# exit-af-interface R1(config-router-af)# af-interface GigabitEthernet0/0/0 R1(config-router-af-interface)# no passive-interface R1(config-router-af-interface)# exit-af-interface R1(config-router-af)# topology base R1(config-router-af-topology)# eigrp router-id 1.1.1.1 R1(config-router-af-topology)# exit-af-topology R1(config-router-af)# exit-address-family ``` Named mode advantages: - Multiple AS numbers per router (one per address-family) - IPv4 and IPv6 in the same router config - Per-interface configuration via `af-interface` sub-mode (cleaner than scattered `passive-interface` commands) - Topology-base abstraction for future multi-topology support Named mode is operationally heavier (more config) but scales better. New deployments should use named mode; existing classic-mode configs can stay as is or be migrated. ## The Network Statement The `network` statement does two things in EIGRP (and in most IGPs): 1. **Activates EIGRP on matching interfaces** \- the router will send and receive EIGRP packets on those interfaces. 2. **Originates connected networks into EIGRP** \- the matching interface's IP is advertised as a route. The wildcard mask matches the interface IP, not the subnet: ``` ! Match an entire /16 network 10.0.0.0 0.0.255.255 ! Match a specific /24 network 192.168.10.0 0.0.0.255 ! Match a single interface (most precise) network 10.0.12.1 0.0.0.0 ``` The single-interface form is most precise but requires updating the EIGRP config every time a new interface is added. Most production deployments use broader masks and rely on `passive-interface default` for safety. ## passive-interface: The Safety Net Without passive-interface, any interface matching the network statement starts sending EIGRP Hellos. If that interface connects to user devices or the public internet, EIGRP packets leak where they should not. The pattern: ``` router eigrp 100 passive-interface default no passive-interface GigabitEthernet0/0/0 ! WAN to other routers no passive-interface GigabitEthernet0/0/1 ! Internal to other routers ``` Default-passive plus selective active is the safe pattern. Forgetting `passive-interface default` is the most common production EIGRP mistake. For named mode: ``` router eigrp PROD address-family ipv4 unicast autonomous-system 100 af-interface default passive-interface exit-af-interface af-interface GigabitEthernet0/0/0 no passive-interface exit-af-interface ``` ## Router ID EIGRP's router ID identifies the router uniquely in DUAL queries and is used for tiebreaking. Set it explicitly: ``` router eigrp 100 eigrp router-id 1.1.1.1 ``` Without explicit configuration, EIGRP picks the highest IP on a loopback (or highest physical interface IP if no loopback). That works but is unpredictable - a new loopback can change the router ID and reset adjacencies. Always set explicitly to a stable address (typically Loopback0). ## Manual Summarization EIGRP supports per-interface summary advertising. The summary's metric defaults to the lowest metric among the summarized routes: ``` ! Classic mode interface GigabitEthernet0/0/0 ip summary-address eigrp 100 10.0.0.0 255.255.0.0 ! Named mode router eigrp PROD address-family ipv4 unicast autonomous-system 100 af-interface GigabitEthernet0/0/0 summary-address 10.0.0.0 255.255.0.0 ``` Useful for: - Reducing routing table size at administrative boundaries - Implementing stub-area-style behavior at distribution layer - Creating black-hole-free summarization (Cisco automatically installs a Null0 route to the summary so transit traffic for unmatched specifics is dropped, not forwarded) Always `no auto-summary` first; manual summarization gives you control over what is summarized where. ## Authentication MD5 authentication for EIGRP uses key chains: ``` ! Define a key chain key chain EIGRP-KEYS key 1 key-string Cisco123! cryptographic-algorithm hmac-sha-256 ! Apply to interface (classic mode) interface GigabitEthernet0/0/0 ip authentication mode eigrp 100 hmac-sha-256 ip authentication key-chain eigrp 100 EIGRP-KEYS ! Named mode router eigrp PROD address-family ipv4 unicast autonomous-system 100 af-interface GigabitEthernet0/0/0 authentication mode hmac-sha-256 authentication key-chain EIGRP-KEYS ``` Modern recommendation: HMAC-SHA-256 over the older HMAC-MD5\. Both ends of every authenticated adjacency must have the same key. Use key chains with multiple keys for non-disruptive rotation. ## Stub Routing For hub-and-spoke deployments, configure spokes as stubs to limit query propagation: ``` ! On a spoke router (classic mode) router eigrp 100 eigrp stub connected summary ! Named mode router eigrp PROD address-family ipv4 unicast autonomous-system 100 eigrp stub connected summary ``` Stub options: `connected` Directly connected networks `summary` Summary routes (manual or auto-summarization) `static` Static routes redistributed into EIGRP `redistributed` Routes redistributed from other protocols `receive-only` Nothing - spoke listens but does not advertise Most production hub-and-spoke designs use `connected summary`. See [EIGRP Stub Routing](https://www.pinglabz.com/eigrp-stub-routing/). ## Verification ``` ! Neighbors Router# show ip eigrp neighbors ! Topology table (DUAL state) Router# show ip eigrp topology ! Routing table (best paths only) Router# show ip route eigrp ! Per-interface EIGRP state Router# show ip eigrp interfaces ! Per-AS protocol info Router# show ip protocols ! Detailed for one prefix Router# show ip eigrp topology 10.10.10.0/24 ``` For named mode replace with `show eigrp address-family ipv4 ...` equivalents. Most show commands accept both classic and named-mode syntax. ## A Production Configuration Example ``` ! Named mode for new deployments router eigrp PROD address-family ipv4 unicast autonomous-system 100 network 10.0.0.0 0.0.255.255 af-interface default passive-interface authentication mode hmac-sha-256 authentication key-chain EIGRP-KEYS exit-af-interface af-interface GigabitEthernet0/0/0 no passive-interface summary-address 10.0.0.0 255.255.0.0 exit-af-interface af-interface GigabitEthernet0/0/1 no passive-interface exit-af-interface topology base eigrp router-id 1.1.1.1 maximum-paths 4 exit-af-topology exit-address-family key chain EIGRP-KEYS key 1 key-string SecurePassword! cryptographic-algorithm hmac-sha-256 ``` This config: passive-interface default with explicit no-passive on the two router-facing interfaces, summarization at the WAN boundary, HMAC-SHA-256 authentication via key chain, explicit router-id, and ECMP up to 4 paths. ## Anti-Patterns - **No `passive-interface default`.** EIGRP packets leak to wherever your network statement matches. - **Auto-summary enabled.** Default in legacy IOS; breaks VLSM. Always `no auto-summary`. - **Different K values across AS.** Neighbor relationships fail. Never change K values. - **Implicit router ID.** Adding a loopback later can change the router ID and reset all adjacencies. - **Mixing classic and named mode in the same AS.** Supported but confusing. Standardize on one. - **Passwords in plaintext in key chains.** Use `service password-encryption` at minimum; consider type-7 key chains. ## Summary EIGRP configuration on modern Cisco IOS XE is one of two modes (classic or named), the right pattern is passive-interface default plus selective no-passive, mandatory `no auto-summary`, explicit router-id, manual summarization at boundaries, and HMAC-SHA-256 authentication on every adjacency. The full pattern fits in one screen of configuration. Master it, run it through a lab, and you have the production template you need. Bookmark this article alongside the [EIGRP cluster pillar](https://www.pinglabz.com/eigrp/) and the [DUAL deep-dive](https://www.pinglabz.com/eigrp-dual-algorithm/). ### EIGRP Metric and K Values Explained URL: https://www.pinglabz.com/eigrp-metric-k-values/ Last updated: 2026-06-13T20:08:39.000Z The EIGRP composite metric is the single thing that confuses CCNP candidates the most. It uses a 64-bit value (or 32-bit in legacy mode) calculated from up to five inputs (bandwidth, delay, load, reliability, MTU) weighted by configurable K values. Most engineers stare at the formula, install a route, see a metric like 3,072,000, and never quite understand where the number came from. This article walks through the metric formula, what each K value controls, why the defaults matter, the wide-metric variant introduced for modern interface speeds, and how to verify metrics in show output. If you are studying for CCNP/CCIE or trying to predict which path EIGRP will pick, this is the math reference. ## The Classic Metric Formula The classic 32-bit EIGRP composite metric (IGRP-derived): ``` metric = 256 * (K1 * BW + (K2 * BW) / (256 - load) + K3 * delay) * (K5 / (reliability + K4)) ``` Where K5 = 0 means the entire reliability term collapses to 1 (the trailing factor disappears). Default K values: K1=1, K2=0, K3=1, K4=0, K5=0\. Substituting: ``` metric = 256 * (BW + delay) ``` And BW (bandwidth) here is calculated as 10^7 / minimum-bandwidth-along-path-in-kbps. Delay is sum-of-delays-along-path-in-tens-of-microseconds. So the simplified default formula: ``` metric = 256 * ((10^7 / min_bandwidth_kbps) + sum_delay_tens_of_us) ``` ## A Worked Example R1 to a destination over a path with three links: R1 to R2 (Gigabit) Bandwidth1,000,000 kbps Delay10 us = 1 (tens of us) R2 to R3 (FastEthernet) Bandwidth100,000 kbps Delay 100 us = 10 (tens of us) R3 to destination (T1) Bandwidth1,544 kbps Delay 20,000 us = 2,000 (tens of us) Minimum bandwidth along the path: 1,544 kbps (the T1). Sum of delays: 1 + 10 + 2,000 = 2,011 (tens of microseconds). Calculate: ``` BW = 10^7 / 1544 = 6477 (rounded) metric = 256 * (6477 + 2011) = 256 * 8488 = 2,172,928 ``` That is the EIGRP metric R1 sees for this destination. Verify with `show ip eigrp topology`. ## Where the Bandwidth Comes From EIGRP uses the interface's configured bandwidth, not actual link speed. The default for most modern Ethernet is set automatically (1,000,000 kbps for GigabitEthernet, etc.) but on serial interfaces it defaults to 1,544 kbps regardless of physical capacity. Set bandwidth explicitly when needed: ``` interface Serial0/0/0 bandwidth 50000 ! 50 Mbps logical bandwidth for EIGRP metric ``` Important: `bandwidth` on the interface is purely a routing-protocol hint. It does not affect actual physical throughput, QoS, or packet rate. EIGRP and other protocols use it for metric calculations. ## Where the Delay Comes From Each interface has a default delay value (in microseconds). Cisco's defaults: GigabitEthernet 10 us FastEthernet 100 us Ethernet 1,000 us Serial (T1) 20,000 us Loopback 5,000 us EIGRP sums delay along the entire path, expressed in tens of microseconds. Override with: ``` interface GigabitEthernet0/0/0 delay 1 ! 10 us = 1 in tens of us ``` Tuning delay is the standard way to influence EIGRP path selection without changing bandwidth (which can affect QoS calculations). Increase delay on a path to make it less preferred; decrease delay on a path to make it more preferred. ## K Values 1 K valueK1 What it weightsBandwidth Effect when set Bandwidth contributes to metric 0 K valueK2 What it weightsBandwidth/(256-load) Effect when set Includes interface load (rarely used) 1 K valueK3 What it weightsDelay Effect when set Delay contributes to metric 0 K valueK4 What it weightsReliability factor Effect when set Reliability becomes part of metric 0 K valueK5 What it weightsReliability multiplier Effect when set Reliability becomes part of metric The Cisco recommendation: never change K values from defaults. Two reasons. First, all routers in the same EIGRP AS must have identical K values or neighbor relationships fail. Second, including dynamic factors like load and reliability in the metric causes routes to flap as load changes, leading to constant DUAL recomputation. If you must tune EIGRP path selection, change interface delay or bandwidth instead of K values. ## The 64-bit Wide Metric The classic 32-bit metric saturates at modern interface speeds. A 100 Gbps link has effectively the same metric as a 1 Gbps link because 10^7 / 100,000,000 rounds to 0 in integer math. To address this, IOS 15.x introduced the 64-bit "wide metric": ``` metric = (K1 * latency + (K2 * latency)/(256-load) + K3 * BW) * (K5/(K4+reliability)) * 65536 ``` The wide metric uses different units (latency in picoseconds, throughput in 65536 \* kbps) and a 64-bit final value, preventing saturation at 100 Gbps and beyond. Enabled by default in modern IOS XE; older networks may run in classic mode for backwards compatibility. Verify metric mode with: ``` Router# show ip eigrp topology EIGRP-IPv4 Topology Table for AS(100)/ID(1.1.1.1) Internal - Cost: 25 / 10 <-- Classic Internal - Cost: (1024000 [12345 us]) <-- Wide ``` The wide metric output shows latency in microseconds explicitly, making it easier to interpret. ## no auto-summary: Why It's Mandatory Default behavior on legacy IOS: EIGRP automatically summarizes routes at classful boundaries. A router with networks 192.168.1.0/24, 192.168.2.0/24, and 192.168.3.0/24 advertises all three as 192.168.0.0/16 across an inter-classful link. This is broken for modern VLSM-using networks. Different /24s belong to different sites; auto-summary mashes them together and breaks routing. Always disable: ``` router eigrp 100 no auto-summary ``` Modern IOS XE has `no auto-summary` enabled by default in named-mode configurations. Classic-mode configs may still inherit the old default; always check. ## Manual Route Summarization Where auto-summary is dangerous, manual summarization at administrative boundaries is essential. EIGRP supports per-interface summarization: ``` interface GigabitEthernet0/0/0 ip summary-address eigrp 100 10.0.0.0 255.255.0.0 ``` Summary's metric defaults to the lowest metric among the summarized routes. Useful for stub-area-style behavior and reducing the size of the routing table at boundaries. For named-mode EIGRP: ``` router eigrp PROD address-family ipv4 unicast autonomous-system 100 af-interface GigabitEthernet0/0/0 summary-address 10.0.0.0 255.255.0.0 ``` ## Verifying Metrics ``` Router# show ip eigrp topology 10.10.10.0/24 EIGRP-IPv4 Topology Entry for AS(100)/ID(1.1.1.1) for 10.10.10.0/24 State is Passive, Query origin flag is 1, 1 Successor(s), FD is 2172928 Routing Descriptor Blocks: 10.0.13.3 (GigabitEthernet0/0/1), from 10.0.13.3, Send flag is 0x0 Composite metric is (2172928/10), route is Internal Vector metric: Minimum bandwidth is 1544 Kbit Total delay is 20110 microseconds Reliability is 255/255 Load is 1/255 Minimum MTU is 1500 Hop count is 3 Originating router is 3.3.3.3 ``` The Vector metric shows you the inputs to the calculation. Minimum bandwidth (1544 kbps along the path), Total delay (20,110 us), and so on. The Composite metric (2172928) is the result. Verify your math against this output. ## Anti-Patterns - **Changing K values to "tune" path selection.** Never. All routers must agree, and dynamic K values cause flapping. - **Changing bandwidth to influence routing.** Better than K-value tuning but watch for QoS implications - the bandwidth statement also affects QoS calculations on the same interface. - **Leaving auto-summary enabled.** Causes classful-boundary mash-ups in any VLSM-using network. Always `no auto-summary`. - **Forgetting to disable interface bandwidth on cables that share T1 with other services.** The default 1,544 kbps drives EIGRP's choice; if you have a 100 Mbps logical channel inside a fractional T1, set bandwidth explicitly. - **Mixing classic and wide metric routers in the same AS.** Modern IOS XE handles the conversion but edge cases exist. Standardize. ## Summary EIGRP's composite metric is bandwidth + delay (with default K values), where bandwidth is 10^7 / minimum-bandwidth-along-path and delay is sum of interface delays. The math is unintuitive at first but predictable once you walk through it. The wide metric extends this to 64 bits for modern interface speeds without changing the conceptual formula. Never change K values. Tune via interface delay or bandwidth. Always `no auto-summary`. Use manual summarization at boundaries. Bookmark this article alongside the [EIGRP cluster pillar](https://www.pinglabz.com/eigrp/) and the [DUAL deep-dive](https://www.pinglabz.com/eigrp-dual-algorithm/). ### EIGRP DUAL Algorithm Deep Dive: Feasible Successors and the Real FD/RD Math URL: https://www.pinglabz.com/eigrp-dual-algorithm/ Last updated: 2026-08-01T18:36:25.000Z DUAL (Diffusing Update Algorithm) is what makes EIGRP fundamentally different from RIP, IGRP, and even OSPF. It is a distributed loop-prevention algorithm that pre-computes loop-free alternates, so failover to a backup happens without re-querying the network. The result is sub-second convergence on healthy networks, the headline EIGRP feature. This article walks DUAL's four key concepts (successor, feasible successor, feasible distance, reported distance), the feasibility condition, the active vs passive states, and what happens when no feasible successor exists. Every number past the theory comes off a three-router CML lab on IOS XE 17.18.2 (`iol-xe`), so you get real six-digit composite metrics, not textbook tens. If you are studying for CCNP/CCIE, or working out why `show ip eigrp topology` shows what it shows, this is the deep dive under the rest of the [guide to how EIGRP builds and maintains its topology table](https://www.pinglabz.com/eigrp/). ## The Problem DUAL Solves Distance-vector protocols (RIP, IGRP) have a famous failure mode: routing loops during convergence. When a route fails, the router that lost it may install a stale advertisement from a neighbor whose own path ran back through it. Traffic loops between the two until the count-to-infinity timer expires. The classic mitigations (split horizon, route poisoning, hold-down timers) reduce loop frequency without eliminating it, and they slow convergence badly. DUAL eliminates loops mathematically instead: every backup is validated by the feasibility condition before installation, guaranteeing it cannot loop back through the failed router. No hold-down timers, no count-to-infinity. ## The Four Key Concepts Reported Distance (RD) Meaning The metric a neighbor reports to us for a destination How calculated Sent in EIGRP UPDATE; the neighbor's metric to the destination Feasible Distance (FD) Meaning The lowest metric we have ever recorded for this destination How calculated Set when route first installed; updated only downward Successor Meaning The neighbor offering the best (lowest) total metric How calculated The neighbor with the lowest (RD + cost-to-neighbor) Feasible Successor (FS) Meaning A backup neighbor whose Reported Distance is strictly less than current FD How calculated Pre-computed for instant failover The math: total metric via a neighbor = (cost to that neighbor) + (its Reported Distance). The successor is the neighbor that minimizes this sum. One naming trap. RD is also called Advertised Distance, abbreviated AD in older material, colliding with administrative distance. They are unrelated: the 90 in square brackets in the RIB is [the number that decides whether EIGRP or another protocol wins a prefix](https://www.pinglabz.com/eigrp-administrative-distance/), the six-digit number after a slash is RD. ## The Feasibility Condition The single most important rule in DUAL: a path is loop-free if and only if the neighbor's Reported Distance is strictly less than the current Feasible Distance. ``` RD < FD => Loop-free, eligible to be a Feasible Successor RD >= FD => May loop; not a Feasible Successor ``` Why does it work? If a neighbor reports a metric lower than our best-ever metric, it must already have a shorter path than we do, so it cannot be reaching the destination through us, so traffic handed to it cannot loop back. Loop-freedom by arithmetic, no timers. The condition is deliberately conservative: some genuinely loop-free paths fail it because the neighbor's RD lands at or above FD. You will see exactly that below on a real prefix. ## The Lab This Was Captured On Three `iol-xe` routers on CML running IOS XE 17.18.2, EIGRP AS 100, wired as a triangle so R1 has two paths to one target prefix (3.3.3.0/24 on R3's Loopback0). - **R1 Et0/1 to R3 Et0/0** \- 10.0.13.0/30, `delay 1000`. The short direct path. - **R1 Et0/0 to R2 Et0/0** \- 10.0.12.0/30, `delay 5000`. Slow, so via-R2 loses. - **R2 Et0/1 to R3 Et0/1** \- 10.0.23.0/30, `delay 100`. Fast, so R2's own distance is low. Delay is in tens of microseconds, set on both ends. That third link is the trick: making R2 to R3 fast gives R2 a small reported distance without making the end-to-end path competitive. Low RD plus high total metric is the shape of a feasible successor, and what you engineer for when you want a pre-computed backup. Delay is the right lever: it is additive along the path while bandwidth is a minimum across it (the [K-value weighting that turns bandwidth and delay into a composite metric](https://www.pinglabz.com/eigrp-metric-k-values/) has the formula). All output below came from on-box EEM applets. ## A Worked Example: The Real Numbers Here is 3.3.3.0/24 on R1 with both paths healthy, from the most useful command in EIGRP troubleshooting: it prints FD and RD for every candidate, side by side. ``` R1# show ip eigrp topology 3.3.3.0/24 EIGRP-IPv4 Topology Entry for AS(100)/ID(10.0.13.1) for 3.3.3.0/24 State is Passive, Query origin flag is 1, 1 Successor(s), FD is 640000 Descriptor Blocks: 10.0.13.2 (Ethernet0/1), from 10.0.13.2, Send flag is 0x0 <-- SUCCESSOR Composite metric is (640000/128256), route is Internal Vector metric: Minimum bandwidth is 10000 Kbit Total delay is 15000 microseconds 10.0.12.2 (Ethernet0/0), from 10.0.12.2, Send flag is 0x0 <-- FEASIBLE SUCCESSOR Composite metric is (1689600/409600), route is Internal Vector metric: Minimum bandwidth is 10000 Kbit Total delay is 56000 microseconds ``` Read the parentheses as **(total metric via this path / RD reported by that neighbor)**. Left is what the path costs R1, right is what the neighbor claims it costs them, and the header FD (640000) is the lowest left-hand number R1 has ever recorded here. Now the arithmetic the abstract explanations skip. The classic composite metric is `256 * (10^7 / minimum-bandwidth-in-Kbit + total-delay-in-usec / 10)`. Both paths cross 10 Mbit Ethernet, so the bandwidth term is 1000 for both and only delay separates them. Successor: via 10.0.13.2 Delay15000 usec Metric(1000 + 1500) x 256 Result = FD640000 R3 reports128256 Backup: via 10.0.12.2 Delay56000 usec Metric(1000 + 5600) x 256 Result1689600 R2 reports (RD)409600 Run the condition on those figures. R2's RD is **409600**, R1's FD is **640000**, and 409600 is strictly less than 640000, so the via-R2 path is a feasible successor: DUAL parks it in the topology table as a pre-validated backup. Notice what the test did **not** look at: 1689600, the backup's total cost, 2.6x the successor metric and completely irrelevant. Feasibility asks about the neighbor's distance, not yours. R2's 409600 sitting under R1's 640000 proves R2 is not reaching 3.3.3.0/24 through R1, so traffic sent that way cannot loop. ## A Prefix With No Feasible Successor The same lab, the same instant, contains prefixes that fail the test. Use `all-links` to see them: the default view hides non-feasible paths and leaves you believing a prefix has no alternate. ``` R1# show ip eigrp topology all-links EIGRP-IPv4 Topology Table for AS(100)/ID(10.0.13.1) Codes: P - Passive, A - Active, U - Update, Q - Query, R - Reply, r - reply Status, s - sia Status P 10.0.13.0/30, 1 successors, FD is 512000, serno 2 via Connected, Ethernet0/1 via 10.0.12.2 (1817600/537600), Ethernet0/0 P 3.3.3.0/24, 1 successors, FD is 640000, serno 4 via 10.0.13.2 (640000/128256), Ethernet0/1 via 10.0.12.2 (1689600/409600), Ethernet0/0 P 10.0.23.0/30, 1 successors, FD is 537600, serno 5 via 10.0.13.2 (537600/281600), Ethernet0/1 via 10.0.12.2 (1561600/281600), Ethernet0/0 P 10.0.12.0/30, 1 successors, FD is 1536000, serno 1 via Connected, Ethernet0/0 via 10.0.13.2 (1817600/1561600), Ethernet0/1 ``` Four prefixes, four sums, two verdicts each way. 3.3.3.0/24 - FD 640000, alt RD 409600409600 < 640000 - feasible 10.0.23.0/30 - FD 537600, alt RD 281600281600 < 537600 - feasible 10.0.13.0/30 - FD 512000, alt RD 537600537600 > 512000 - NOT feasible 10.0.12.0/30 - FD 1536000, alt RD 15616001561600 > 1536000 - NOT feasible 10.0.13.0/30 is the instructive one. R1 is directly connected out Et0/1, so its FD is the connected metric, 512000\. R2 reports 537600, having learned the subnet from R3 rather than from R1, so that path is provably loop-free. It fails the feasibility condition anyway, by 25600. That is the conservatism made concrete. Connected prefixes are the classic victims: your FD is as low as it can get, so almost nothing a neighbor reports beats it and the alternate sits there unusable. Feasibility is a gate, not a preference. So when 10.0.13.0/30 loses its successor, DUAL cannot promote anything locally. The prefix goes Active, R1 queries R2, and nothing is installed until the reply lands. That is [where a prefix that goes Active and never comes back gets stuck](https://www.pinglabz.com/eigrp-query-process-stuck-in-active/): if a router in the query chain fails to reply in time, the adjacency is torn down. Every SIA incident starts as a prefix with no feasible successor. ## Proving It: Kill the Successor Link Theory is cheap. Here is the RIB before anything was touched: ``` R1# show ip route eigrp 3.0.0.0/24 is subnetted, 1 subnets D 3.3.3.0 [90/640000] via 10.0.13.2, 00:01:27, Ethernet0/1 10.0.0.0/8 is variably subnetted, 5 subnets, 2 masks D 10.0.23.0/30 [90/537600] via 10.0.13.2, 00:01:27, Ethernet0/1 ``` An EEM applet then shuts Et0/1, the successor path, and the adjacency drops: ``` %DUAL-5-NBRCHANGE: EIGRP-IPv4 100: Neighbor 10.0.13.2 (Ethernet0/1) is down: interface down ``` Here is the same topology entry afterwards. Read the state field carefully: ``` R1# show ip eigrp topology 3.3.3.0/24 EIGRP-IPv4 Topology Entry for AS(100)/ID(10.0.13.1) for 3.3.3.0/24 State is Passive, Query origin flag is 1, 1 Successor(s), FD is 640000 Descriptor Blocks: 10.0.12.2 (Ethernet0/0), from 10.0.12.2, Send flag is 0x0 <-- former FS, now SUCCESSOR Composite metric is (1689600/409600), route is Internal ``` **State is Passive.** The successor link is down, the prefix changed next hop and interface, and DUAL never ran a diffusing computation. No Active state, no query, no reply to wait on, no SIA exposure. R1 already knew this path was loop-free, so promotion was a local table operation. The RIB agrees: ``` R1# show ip route eigrp 3.0.0.0/24 is subnetted, 1 subnets D 3.3.3.0 [90/1689600] via 10.0.12.2, 00:01:00, Ethernet0/0 10.0.0.0/8 is variably subnetted, 4 subnets, 2 masks D 10.0.13.0/30 [90/1817600] via 10.0.12.2, 00:01:00, Ethernet0/0 D 10.0.23.0/30 [90/1561600] via 10.0.12.2, 00:01:00, Ethernet0/0 ``` The metric rose to 1689600 because the path really is longer, but reachability never dropped. Note the FD still reads 640000: FD only ratchets downward while a route is Passive, holding the historic best until the route next goes Active. That catches people out when they recompute feasibility after a failover using the metric in the RIB. Contrast the second RIB line. 10.0.13.0/30 also recovered via R2, but with no feasible successor it could only get there by going Active and querying. Same outcome, different journey: one a table lookup, one a conversation with a neighbor. How fast DUAL notices a failure is a separate problem from how fast it recovers. Here the interface went physically down, so detection was instant. On a link that stays up while the far end is dead you wait on hold timers, which is why [sub-second failure detection with BFD](https://www.pinglabz.com/bfd-bidirectional-forwarding-detection/) pairs well with EIGRP. ## Active vs Passive States Every route in the topology table has a state: Passive (P) Meaning Route is stable. Successor and any feasible successors known. Operator action None; healthy steady state Active (A) Meaning Route lost its successor with no feasible successor. Querying neighbors for alternatives. Operator action Investigate; persistent active routes mean trouble Update (U) Meaning Update message in flight Operator actionBrief transient Query (Q) Meaning Query message in flight Operator action Transient during DUAL active state Reply (R) MeaningReply pending Operator actionTransient Stuck-In-Active (SIA) Meaning Active for too long; neighbor not responding Operator action Neighbor declared dead; relationship reset Every healthy route shows Passive, like every prefix in the `all-links` capture above. Active routes are normal during convergence, but persistent active state means a neighbor is failing to reply, usually from overload or path issues. The dreaded SIA: a neighbor receives a query, fails to reply within the SIA timer (default 3 minutes), and EIGRP tears down the relationship. Cascading SIAs destabilise large networks. The classic mitigations are stub routing on spokes and raising SIA timers (rarely the right answer). The real fix sits upstream of both: engineer feasible successors so the prefixes that matter never go Active. ## Verifying DUAL State `show ip eigrp topology`Table view, P or A. Hides non-feasible paths. `show ip eigrp topology `FD, state, descriptor blocks, vector metrics. `show ip eigrp topology all-links`Every candidate, feasible or not. `show ip route eigrp`What got installed, with AD and metric. To answer "does this prefix have a backup?", run `all-links` and compare the header FD against the right-hand number in every other path. To build this triangle yourself, the [step-by-step lab that engineers a feasible successor with interface delay](https://www.pinglabz.com/ccna-lab-ipc-11-eigrp-feasible-successor-dual/) has the configs. ## Multiple Successors and Variance By default EIGRP installs one successor. Equal-cost paths need identical metrics to be installed together; unequal-cost paths need `variance`. ``` router eigrp 100 variance 2 ! Install paths up to 2x the FD as additional successors ``` Variance admits paths up to N times the FD, provided they already pass feasibility. On the lab numbers, with FD 640000, `variance 2` admits up to 1280000, so the via-R2 path at 1689600 stays out of the RIB until `variance 3`. The alternate for 10.0.13.0/30 would never be installed at any variance value, because variance cannot override feasibility. Feasibility decides whether a path is allowed, variance decides whether an allowed path is used. Full mechanics are in [how to load balance EIGRP over unequal-cost paths](https://www.pinglabz.com/eigrp-variance-unequal-cost-load-balancing/). ## Stub Routing's Impact on DUAL EIGRP stub routing modifies DUAL's query behavior. By default a router entering Active state queries all neighbors; with stub routing on spokes, the hub does not query stub neighbors. The benefit: the hub avoids fanning queries out to every spoke (each consumes spoke resources and adds convergence latency). The trade-off: stubs cannot be transit routers, because they do not advertise routes learned from other neighbors. For most hub-and-spoke designs that is what you want. See [EIGRP Stub Routing](https://www.pinglabz.com/eigrp-stub-routing/). ## DUAL vs SPF: Why EIGRP Is Different from OSPF Algorithm type DUAL (EIGRP) Distance-vector with diffusing computation SPF (OSPF) Link-state with Dijkstra Each router knows DUAL (EIGRP) Routes from direct neighbors plus their RDs SPF (OSPF) Full topology of the area Failover DUAL (EIGRP) Sub-second when FS exists; queries otherwise SPF (OSPF) SPF runs across the area Topology visibility DUAL (EIGRP) Limited (per-prefix only) SPF (OSPF)Full (LSDB) Convergence on healthy network DUAL (EIGRP) Fastest of the IGPs (cached FS) SPF (OSPF)Sub-second with tuning Convergence in worst case DUAL (EIGRP) Slower (queries propagate across topology) SPF (OSPF) Bounded by SPF complexity The trade-off: DUAL fails over instantly on cached alternates but forces a query when none exist, while SPF runs the same computation every time, so it is more predictable but slower at its best. The lab above is DUAL's best case rendered literally: an interface went down and the topology entry never changed state. See [EIGRP vs OSPF: When to Use Each](https://www.pinglabz.com/eigrp-vs-ospf/). ## Common Mistakes and Gotchas - **Reading (X/Y) backwards.** Left is your total metric via that path, right is the neighbor's RD. Compare the right number against the header FD, never against the left. - **Concluding there is no backup from plain `show ip eigrp topology`.** The default view suppresses non-feasible paths. Confirm with `all-links`. - **Assuming a cheap path is automatically feasible.** Here a backup costing 2.6x the successor passed, while a path 5 percent above FD failed. - **Expecting connected prefixes to have feasible successors.** Your FD there is as low as it gets, so a neighbor's RD rarely beats it. Those go Active on you. - **Recomputing feasibility against the post-failover metric.** FD does not rise while a route is Passive: after failover the installed metric was 1689600 but FD still read 640000. - **Reaching for the SIA timer.** The answer is more feasible successors or fewer query targets, not a longer timer that extends the outage before it is declared. ## Key Takeaways - The feasibility condition is `RD < FD`, strictly less than. Proven here with 409600 < 640000. - Composite metrics print as (FD-via-this-path / RD-from-neighbor). The header line's "FD is N" is what you compare against. - Total path cost is irrelevant to feasibility: 1689600 qualified, 537600 did not. - With a feasible successor present the successor link can drop and the route **never leaves Passive**. Captured live. - With no feasible successor the same failure forces a query and a wait, which is where every stuck-in-active story begins. Run `all-links` before declaring a prefix has no backup. ## Summary DUAL is EIGRP's superpower. The four concepts and the feasibility condition (RD < FD) give loop-free routing with sub-second failover, because a cached feasible successor lets failover happen without querying anyone. Master the condition, and get comfortable running the arithmetic on real six-digit metrics. Bookmark this alongside the rest of the [EIGRP configuration and troubleshooting guides](https://www.pinglabz.com/eigrp/). New labs and guides, in your inbox Every new PingLabz lab and deep-dive, built and verified on real Cisco IOS XE - free, straight to your inbox. [Join free](https://www.pinglabz.com/signup/) ### QoS in Modern Networks: SD-WAN, Cloud, and Application-Aware Steering URL: https://www.pinglabz.com/qos-sd-wan-modern-networks/ Last updated: 2026-06-13T20:08:40.000Z QoS in modern networks looks different from the MPLS-WAN-with-LLQ patterns that dominated the 2000s and 2010s. Cloud-direct connectivity, SD-WAN application-aware steering, SaaS-everything traffic, and best-effort internet transport between branches and the cloud all shift where QoS lives, what it can enforce, and what it cannot. The protocol mechanics (DSCP, LLQ, MQC) still apply; their context has changed. This article walks through how QoS works in SD-WAN, how it interacts with cloud on-ramps, what changes when transport is the public internet, the new app-aware policy patterns, and the realities of QoS-end-to-end when half the path is somebody else's network. If you are designing a modern WAN or SASE deployment, or trying to figure out why your DSCP markings stop being honored at the cloud edge, this is the reference. ## What Changed Versus the MPLS Era Transport MPLS WAN era Carrier-managed MPLS with hard QoS Modern (SD-WAN + cloud) Mix of MPLS, broadband, LTE; best-effort over public internet for most Carrier honors DSCP? MPLS WAN eraYes (per contract) Modern (SD-WAN + cloud) MPLS yes; broadband internet no End-to-end QoS MPLS WAN era Achievable for traffic on MPLS Modern (SD-WAN + cloud) End-to-end only between sites you control; cloud is opaque Application identification MPLS WAN era Port-based ACLs at the edge Modern (SD-WAN + cloud) DPI-based (NBAR2, AVC) at the SD-WAN edge Path selection MPLS WAN era Single primary path; backup if primary fails Modern (SD-WAN + cloud) Multiple paths in parallel; per-application steering by SLA metrics Policy expression MPLS WAN era DSCP value + per-class queue Modern (SD-WAN + cloud) Application name + intent ("voice prefers low jitter") Where QoS lives MPLS WAN era Every router in the MPLS path Modern (SD-WAN + cloud) SD-WAN edge + cloud on-ramp; opaque between The big conceptual shift: QoS used to be about marking once and trusting the carrier to honor markings end-to-end. Modern QoS is about marking once for what you control, and accepting that your markings end at the boundary of someone else's network. ## QoS in SD-WAN SD-WAN platforms implement application-aware QoS that combines traditional queueing with per-application path steering. The mechanism: 1. **Application identification at the edge.** The SD-WAN WAN Edge does deep packet inspection (NBAR2 in Cisco Catalyst SD-WAN, app signatures in other vendors) to identify Office 365, Salesforce, Zoom, Teams, etc. 2. **Per-application policy.** "Voice prefers MPLS or lowest-jitter broadband." "Salesforce goes direct internet break-out." "Bulk backup goes to whichever transport has spare capacity." 3. **Per-tunnel SLA monitoring.** Edges send keep-alive probes through every tunnel to measure latency, jitter, and loss. Path selection updates dynamically as conditions change. 4. **Local queueing.** On each WAN egress, traditional LLQ + CBWFQ still applies. The application-aware steering picks which tunnel; the per-tunnel queueing decides packet order within that tunnel. So SD-WAN QoS is two-layer: application-aware path selection (which transport) plus per-tunnel queueing (what order within that transport). The traditional QoS toolset is still there; it just runs at a different scale. Cisco Catalyst SD-WAN expresses this via centralized data policy: ``` ! Conceptual policy expression in vManage Match: application Office365 Action: preferred-color biz-internet fallback mpls sla-class voice-class Match: application VoIP Action: preferred-color mpls fallback biz-internet sla-class voice-class marking dscp ef ``` The vSmart compiles this into per-WAN-Edge state and pushes via OMP. See the [SD-WAN cluster pillar](https://www.pinglabz.com/sd-wan/) and [Cisco Catalyst SD-WAN architecture article](https://www.pinglabz.com/cisco-sd-wan-architecture/) for the underlying control plane. ## QoS Over Internet Transport Internet ISPs do not honor your DSCP markings. The IETF's IPv4 specification requires intermediate routers to preserve DSCP, but in practice: - Most ISPs zero out DSCP at their network edge to prevent customers from gaming inter-AS QoS - Some ISPs preserve DSCP within their network but do not act on it - Public peering points strip DSCP routinely - Cloud providers (AWS, Azure, GCP) do not honor DSCP inside their networks Practical implication: your DSCP marking matters only at endpoints under your control. Your SD-WAN edge sees DSCP, applies queueing, and sends the packet over the internet. The next hop where DSCP matters again is the receiving SD-WAN edge. What still works over internet: - **Edge marking and queueing.** Local prioritization of traffic leaving your edge. - **SD-WAN path conditioning.** Forward Error Correction (FEC), packet ordering, jitter buffers compensate for variable internet quality. - **Per-tunnel SLA-based steering.** If one path's jitter spikes, voice automatically moves to a better path. - **Local jitter buffers at endpoints.** The receiving softphone or video client absorbs jitter. What does not work: - End-to-end DSCP honoring - Carrier-grade SLA on packet loss for VoIP - Predictable latency over multi-AS internet paths ## QoS at Cloud On-Ramps Cloud on-ramps (AWS Cloud WAN, Azure Virtual WAN, Google Cloud Network Connectivity Center, AWS Direct Connect, Azure ExpressRoute) provide private connectivity from your network to the cloud provider's network. Inside the cloud provider's network, your DSCP is generally not honored (the provider does not own per-packet QoS for tenant traffic at the public-cloud scale). What you can control at the cloud on-ramp: - **QoS at the SD-WAN edge facing the on-ramp.** Your standard LLQ + CBWFQ on egress to the cloud. - **Bandwidth allocation per cloud (in multi-cloud).** Reserve more bandwidth for the cloud carrying your latency-sensitive workloads. - **Path preference.** Use AWS Direct Connect for production workloads; use IPsec VPN over internet for less critical traffic. What you cannot control: - QoS inside AWS, Azure, or GCP - How the cloud routes traffic between regions - Latency to specific services (CloudFront, Azure Front Door, etc.) ## Application-Aware Policy: From DSCP Values to Application Names The legacy QoS expression was DSCP-centric: "AF41 traffic gets 30 percent bandwidth." The modern expression is application-centric: "Microsoft Teams traffic gets the best path." The translation: "Mark RTP audio as DSCP EF" "Identify Zoom/Teams voice via NBAR2; mark DSCP EF and prefer low-jitter path" "Allocate 30% bandwidth to AF41 class" "Reserve 30% of WAN for video conferencing applications" "AF21 for transactional apps" "Salesforce, Workday, ServiceNow traffic prefers business-internet path" "Default class for everything else" "Bulk and unclassified traffic uses lowest-cost path" The application-aware approach is more robust against the SaaS reality: an HTTPS connection to Office 365 looks identical to one to YouTube at the port level. NBAR2 identifies the service via DNS, certificate fields, and traffic patterns. AVC (Application Visibility and Control) extends this with reporting. The legacy mechanics still apply underneath. Once an application is identified, the marking, queueing, and shaping are the same MQC patterns as before. NBAR2 is a more sophisticated classifier feeding the same MQC engine. ## Wireless QoS in the Modern Era Wireless QoS uses 802.11e WMM with four access categories (Voice, Video, Best Effort, Background). The Catalyst 9800 maps DSCP from incoming wired traffic into WMM access categories on the air, and CoS-marked frames from wireless clients into DSCP for upstream forwarding. Modern wireless QoS additions: - **Auto QoS.** One command (`auto qos voip` or equivalent on the C9800) configures the standard wireless QoS templates without hand-rolling MQC. - **AVC integration.** NBAR2 runs on the WLC; identifies applications; applies marking and steering. - **Wi-Fi 6 (802.11ax) features.** OFDMA improves QoS-like behavior at the radio layer (fairness across many simultaneous users); WMM still provides the application-class priority. - **Wi-Fi 7 (802.11be) Multi-Link Operation.** A client can use multiple radios simultaneously; the AP can use the cleanest band for high-priority traffic. For the wireless-specific configuration, see [C9800 QoS Configuration: Auto QoS, DSCP Mapping, and Wireless Profiles](https://www.pinglabz.com/c9800-qos-configuration/). ## End-to-End QoS in the Real World The honest picture of "end-to-end QoS" in 2026: Endpoint to access switch Trust boundary; application marking or re-marking Access to distribution to core DSCP honored; LLQ + CBWFQ on congested links WAN edge to MPLS provider DSCP honored per contract; carrier-managed QoS WAN edge over public internet DSCP not honored; SD-WAN path conditioning helps Cloud on-ramp to cloud provider DSCP generally not honored; rely on path-level segregation Inside cloud provider No tenant control over QoS Cloud provider to other endpoints Best-effort; no QoS The pragmatic engineering question: where in this chain do you have control, and what can you do at those points to maximize user experience? The answer for most modern networks is "the SD-WAN edge plus the local LAN," with everything in between treated as best-effort that you compensate for via SD-WAN path selection. ## Modern QoS Design Patterns Three patterns dominate 2026 designs: **1\. SD-WAN with App-Aware Steering.** The dominant pattern. Identify applications at the SD-WAN edge; mark with standardized DSCP; steer over multiple transports based on real-time SLA measurements. Local LLQ on each tunnel; SD-WAN handles the inter-tunnel decision. **2\. SaaS-Direct with Local QoS.** SaaS traffic exits the local SD-WAN edge directly to the internet (no backhaul). Local QoS prioritizes voice and video on the way out; the rest is best-effort over the internet. SASE handles the security inspection on the way out. **3\. Cloud-On-Ramp with Per-Cloud Bandwidth.** Multiple cloud providers; each gets its own on-ramp tunnel; per-cloud bandwidth allocation in the SD-WAN policy. Voice and video to specific clouds get prioritized; bulk transfer to other clouds gets the leftovers. All three rely on SD-WAN as the policy enforcement point. The traditional MQC machinery still runs underneath; the policy expression is application-centric. ## Anti-Patterns in Modern QoS - **Assuming DSCP is honored end-to-end.** It is not. Plan for marking to disappear at every administrative boundary outside your control. - **Backhauling SaaS traffic.** Sending Office 365 traffic to HQ for inspection then out to the internet adds latency for no benefit. Use SD-WAN direct break-out plus SASE inspection. - **Forgetting that internet is best-effort.** Voice over public internet works most of the time; SD-WAN path conditioning (FEC, packet ordering, jitter buffers) helps; carrier-grade SLAs do not exist over consumer broadband. - **Over-engineering policy with too many classes.** 12-class IETF models work in theory; in practice, 4-6 classes (voice, video, business, default, scavenger) are easier to operate and just as effective. - **Ignoring application identification accuracy.** NBAR2 misidentifies some traffic. Audit periodically; tune signatures; do not assume the classifier is always right. ## Summary Modern QoS is application-aware steering at the SD-WAN edge plus local queueing on each tunnel, with the understanding that DSCP markings end at every administrative boundary outside your control. The traditional QoS toolset (MQC, LLQ, CBWFQ, DSCP) still runs underneath; the policy expression is more application-centric and the path selection layer is dynamic. If you are designing modern QoS, focus on the SD-WAN edge as the policy enforcement point, accept that internet transport is best-effort, and rely on path conditioning + jitter buffers to compensate. Bookmark this article alongside the [QoS cluster pillar](https://www.pinglabz.com/qos/), the [SD-WAN cluster pillar](https://www.pinglabz.com/sd-wan/), and the per-feature articles on DSCP, MQC, queueing, and voice/video QoS for the full design picture. ### Voice and Video QoS on Cisco IOS XE URL: https://www.pinglabz.com/voice-video-qos-cisco/ Last updated: 2026-06-13T20:08:40.000Z Voice and video are the use cases QoS exists for. Without QoS, a 1-percent packet loss rate on a TCP file transfer is invisible to the user; on a VoIP call, it is a robotic-sounding artifact that ends the meeting. Real-time media has tight requirements that the rest of your traffic does not, and those requirements are what drive every QoS design decision. This article walks through the latency, jitter, and loss budgets for voice and video, the standard DSCP markings, the LLQ configuration that protects them, the IP phone trust pattern, and the bandwidth math you use to size the priority queue. If you are deploying VoIP for the first time or troubleshooting a video conferencing complaint, this is the reference. ## The Latency, Jitter, and Loss Budgets VoIP (toll quality) One-way latency< 150 ms Jitter< 30 ms Packet loss < 1 percent (preferably < 0.1 percent) VoIP (acceptable) One-way latency150-300 ms Jitter< 50 ms Packet loss< 3 percent Video conferencing One-way latency< 200 ms Jitter< 50 ms Packet loss< 1 percent Streaming video (one-way) One-way latency< 5 s buffer Jitter tolerated by jitter buffer Packet loss< 5 percent The numbers come from ITU-T Recommendation G.114 (one-way transmission time) and various enterprise media studies. The 150 ms VoIP latency target is one-way - so the round trip should be under 300 ms. Most domestic networks comfortably meet this; international circuits can be tight. Jitter is variation in latency between consecutive packets. A jitter buffer at the receiver smooths small jitter (typically 30-50 ms of buffer); jitter beyond the buffer causes audible artifacts. The QoS goal is to keep jitter below the receiver's buffer. Packet loss matters more for voice than for streaming because voice has no retransmission - a lost packet is gone. Codecs do packet loss concealment up to a point but quality degrades fast above 1 percent. ## Voice Marking and Bandwidth The standard DSCP markings for voice and video: VoIP RTP audio (voice payload) DSCPEF Decimal46 TreatmentLLQ priority queue Voice signaling (SIP, H.323) DSCPCS3 Decimal24 Treatment CBWFQ guaranteed bandwidth Video conferencing (Zoom, Teams, Webex) DSCPAF41 Decimal34 TreatmentCBWFQ with WRED Streaming video (one-way) DSCPCS5 or AF31 Decimal40 or 26 TreatmentCBWFQ VoIP bandwidth math. A single VoIP call with G.711 codec consumes: - G.711 payload: 64 kbps - RTP/UDP/IP overhead: \~16 kbps - Layer 2 overhead (Ethernet/MPLS): \~16 kbps - **Total per call: \~96 kbps** For G.729 codec, the payload drops to 8 kbps but the overhead is fixed; total is \~24 kbps per call. For Opus (modern), the payload varies (6-510 kbps); typical voice usage is 24-48 kbps. Sizing the priority queue: 100 simultaneous G.711 calls = 9.6 Mbps. Reserve 10-15 percent of WAN bandwidth for voice priority queue, which covers typical concurrent call counts comfortably. ## IP Phone Trust Pattern Cisco IP phones mark their own voice traffic with DSCP EF and CoS 5 by default. The access port should trust this marking via Cisco's conditional trust mechanism: ``` interface GigabitEthernet1/0/5 switchport mode access switchport access vlan 10 ! Data VLAN (PC behind phone) switchport voice vlan 20 ! Voice VLAN (phone) mls qos trust device cisco-phone ! Conditional: trust if Cisco phone detected spanning-tree portfast spanning-tree bpduguard enable ``` The conditional trust uses CDP. If the switch sees a Cisco phone via CDP, trust applies. If the phone is unplugged, trust drops automatically and the port re-marks all incoming DSCP to 0. For non-Cisco IP phones (which do not speak CDP), use absolute trust within the voice VLAN only: ``` interface GigabitEthernet1/0/5 switchport mode access switchport access vlan 10 switchport voice vlan 20 mls qos trust dscp ! Absolute trust; relies on voice VLAN segregation ``` This is less secure - the PC behind the phone can theoretically mark its own DSCP - but works for non-CDP phones. Some platforms support voice-VLAN-only trust mechanisms; check your switch documentation. ## Voice Signaling: CS3 SIP, SDP, H.323, and other voice control-plane traffic gets DSCP CS3 (24). Why a separate class? Voice signaling is bursty (call setup and teardown) and CPU-intensive on the call manager but not latency-critical the way RTP audio is. It belongs in a guaranteed-bandwidth CBWFQ class, not the priority queue. If signaling were placed in the EF priority queue, a flood of misconfigured signaling (a call manager hiccup, an attack, a SIP loop) could starve the voice payload itself. Keep them separate. ## Full Cisco IOS XE LLQ Configuration ``` ! Classification class-map match-any VOICE-RTP match dscp ef ! Trust the marking class-map match-any VOICE-SIGNALING match dscp cs3 class-map match-any VIDEO-CONF match dscp af41 class-map match-any STREAMING-VIDEO match dscp cs5 af31 class-map match-any TRANSACTIONAL match dscp af21 class-map match-any SCAVENGER match dscp cs1 ! Egress queueing policy policy-map WAN-EGRESS class VOICE-RTP priority percent 10 ! 10% strict priority class VOICE-SIGNALING bandwidth percent 5 class VIDEO-CONF bandwidth percent 25 random-detect dscp-based class STREAMING-VIDEO bandwidth percent 15 random-detect class TRANSACTIONAL bandwidth percent 20 random-detect class SCAVENGER bandwidth percent 1 class class-default bandwidth percent 24 fair-queue random-detect ! Apply on WAN egress interface GigabitEthernet0/0/0 description WAN to ISP service-policy output WAN-EGRESS ``` Verify with: ``` Router# show policy-map interface GigabitEthernet0/0/0 GigabitEthernet0/0/0 Service-policy output: WAN-EGRESS Class-map: VOICE-RTP (match-any) 876543 packets, 124589376 bytes 30 second offered rate 96000 bps, drop rate 0 bps Match: dscp ef (46) Priority: 10% (10000 kbps), burst bytes 250000, b/w exceed drops: 0 Conform: 876543 packets / 124589376 bytes Exceed: 0 packets / 0 bytes ``` The "Exceed: 0 packets" line is what you want for the voice class. Anything above zero means the priority queue's built-in policer is dropping packets - your voice budget is too small or you have a runaway sender. ## Codec Choice and Bandwidth Implications G.711 (PCMU/PCMA) Payload bandwidth64 kbps Total per call (with overhead)\~96 kbps Quality Toll quality; high CPU on transcoder G.722 Payload bandwidth64 kbps Total per call (with overhead)\~96 kbps Quality Wideband (HD voice); same bandwidth as G.711 G.729 Payload bandwidth8 kbps Total per call (with overhead)\~24 kbps Quality Acceptable; significant compression artifacts Opus (narrowband) Payload bandwidth6-24 kbps Total per call (with overhead)\~22-40 kbps Quality Configurable; modern default Opus (wideband) Payload bandwidth20-48 kbps Total per call (with overhead)\~36-64 kbps Quality HD voice; default for WebRTC The bandwidth-quality trade-off: G.711 is the safest for compatibility (every endpoint speaks it) but expensive on bandwidth. G.729 saves bandwidth significantly but requires DSP licenses on transcoders and quality is noticeably worse. Modern WebRTC apps default to Opus, which auto-tunes between narrowband and wideband based on link conditions. Practical implication for QoS sizing: don't assume your priority queue size based on G.711 if your endpoints negotiate G.729 or Opus most of the time. Monitor actual usage; tune the percent allocation. ## Video Conferencing Specifics Modern video conferencing (Zoom, Teams, Webex) uses adaptive video that scales bandwidth based on link conditions. A typical Zoom 1080p call: - Audio: 16-64 kbps (Opus) - Video: 1.2-2.4 Mbps for HD, drops to 600-1200 kbps under congestion - Screen share: variable, can spike to 4-6 Mbps for high-resolution screens Mark video conferencing as AF41 (DSCP 34) and place it in a guaranteed-bandwidth CBWFQ class with WRED. The class should have enough headroom for typical concurrent video calls but not so much that other traffic starves. Sizing example: 50 simultaneous HD Teams calls at 2 Mbps each = 100 Mbps. On a 200 Mbps WAN that's 50 percent of capacity - allocate 50 percent to video conferencing class. On a 1 Gbps WAN that's 10 percent. ## Voice in the Cloud Era: SaaS UC Modern UC platforms (Microsoft Teams, Zoom Phone, Webex Calling) place all the call control in the cloud. The branch handles only the media (RTP) directly between endpoints. Implications for QoS: - Mark RTP at the access port (or the application client marks itself) - Internet break-out matters: the SD-WAN should steer voice over the lowest-jitter path to the cloud - Cloud-direct paths (Microsoft Direct Routing, Cisco Webex Calling) bypass much of the WAN; QoS focus shifts to the local edge - The cloud provider does not honor your DSCP markings; QoS ends at the SD-WAN cloud on-ramp For SD-WAN voice QoS, see (article forthcoming on QoS in modern networks) and the [SD-WAN cluster pillar](https://www.pinglabz.com/sd-wan/). ## Wireless Voice: WMM and 802.11 Access Categories Voice on Wi-Fi uses 802.11e WMM (Wi-Fi Multimedia) with four access categories: AC\_VO (Voice) Maps to DSCPEF (46) Used forVoIP RTP AC\_VI (Video) Maps to DSCPAF41 (34), CS5 (40) Used for Video conferencing, streaming AC\_BE (Best Effort) Maps to DSCPBE (0) Used forDefault AC\_BK (Background) Maps to DSCPCS1 (8) Used forScavenger The Catalyst 9800 maps DSCP from incoming wired traffic into WMM categories on the air, and CoS-marked frames from wireless clients into DSCP for upstream forwarding. Auto QoS handles most of this automatically. See [C9800 QoS Configuration: Auto QoS, DSCP Mapping, and Wireless Profiles](https://www.pinglabz.com/c9800-qos-configuration/) for the wireless-specific configuration. ## Troubleshooting Voice Quality The diagnostic chain for "voice sounds bad": 1. **Verify markings.** Capture at the source. Is RTP actually marked DSCP EF? If not, classification is broken. 2. **Verify trust.** `show mls qos interface` on the access port. Is conditional trust working? If the port is untrusted, the marking is being stripped. 3. **Verify queueing.** `show policy-map interface` on each WAN egress. Is the priority class hitting? Is the class dropping (Exceed counter)? 4. **Verify bandwidth math.** How many concurrent calls? How much WAN does that consume? Is the priority percent allocation enough? 5. **Verify path latency.** traceroute and ping. Is the round-trip under 300 ms? Is jitter (variance) reasonable? The most common cause of voice quality complaints: a misconfigured trust boundary letting non-voice traffic into the priority queue, starving real voice. Audit the access ports. ## Summary Voice and video QoS comes down to LLQ for voice (strict priority + built-in policer), CBWFQ for video and signaling (guaranteed bandwidth + WRED), and rigorous trust boundary discipline at the access edge. Standard markings (DSCP EF for RTP, CS3 for signaling, AF41 for video conferencing, CS5 for streaming) and standard bandwidth allocations (\~10 percent for voice, \~25 percent for video conferencing) cover most enterprise scenarios. The configuration is well-understood; the trust discipline is what fails in production. Audit access ports periodically, monitor priority queue drops, and verify codec choice matches your bandwidth math. Bookmark this article alongside the [QoS cluster pillar](https://www.pinglabz.com/qos/) and lab any change before pushing to production. ### QoS Queueing, Policing, and Shaping Compared URL: https://www.pinglabz.com/qos-queueing-policing-shaping/ Last updated: 2026-06-13T20:08:40.000Z Queueing, policing, and shaping are the three traffic-management primitives every Cisco QoS deployment uses. They sound similar in casual conversation - "they limit traffic, right?" - but each does something different and they are not interchangeable. Confusing them is the second most common QoS mistake (after trust-boundary failures), and the symptoms are subtle: drops in unexpected places, queues growing without bound, traffic that gets policed before it gets shaped. This article walks through what each primitive does, where it applies, when to use which, the Cisco IOS XE configuration, and the standard production patterns. If you are configuring QoS for the first time or auditing a deployment that "feels off," this is the reference. ## The Three Primitives Queueing What it does Decides the order packets leave a congested interface Excess traffic Stays in queues until served (or dropped if queue full) Where it applies Egress only (output direction) Policing What it does Enforces a hard rate limit; drops or re-marks excess Excess trafficDropped (or re-marked) Where it appliesIngress or egress Shaping What it does Smooths bursts to fit a target rate; buffers excess Excess trafficBuffered and delayed Where it appliesEgress only Queueing answers "in what order do these waiting packets leave?" Policing and shaping both answer "how do we keep traffic at or below rate X?" but with very different mechanics: policing drops, shaping buffers. ## Queueing: Scheduling on Congested Interfaces An interface is "congested" when more packets want to leave than the interface can transmit in a given time slice. Without congestion, queueing does not matter - packets transmit in arrival order. With congestion, the queueing scheduler chooses what leaves next. Cisco supports several queueing schedulers, in order of historical introduction: FIFO (First-In-First-Out) Behavior Single queue; no priority Use in 2026 Default for uncongested interfaces; never on production WAN Priority Queueing (PQ) Behavior Four queues; higher always served first; can starve lower Use in 2026Legacy; deprecated Custom Queueing (CQ) Behavior 16 queues with byte-count round-robin Use in 2026Legacy; deprecated Weighted Fair Queueing (WFQ) Behavior Per-flow queues with proportional service based on weight Use in 2026 Default on slow serial; rarely tuned in 2026 Class-Based Weighted Fair Queueing (CBWFQ) Behavior Per-class queues with configured bandwidth guarantees Use in 2026 Modern default for non-real-time classes Low-Latency Queueing (LLQ) Behavior CBWFQ plus a strict-priority queue with built-in policer Use in 2026 Modern default when voice or video shares with data LLQ is the dominant queueing strategy in modern Cisco deployments. It gives voice a strict priority queue (lowest latency, lowest jitter) but applies a built-in policer to prevent the priority queue from starving everything else if voice traffic explodes. The remaining bandwidth divides among other classes via CBWFQ proportional service. ### LLQ Configuration ``` policy-map WAN-EGRESS class VOICE priority percent 10 ! Strict priority, capped at 10% class VIDEO bandwidth percent 30 ! Guaranteed 30% random-detect dscp-based ! WRED class TRANSACTIONAL bandwidth percent 25 ! Guaranteed 25% random-detect class SCAVENGER bandwidth percent 1 ! Tiny guarantee class class-default bandwidth percent 34 fair-queue random-detect interface GigabitEthernet0/0/0 service-policy output WAN-EGRESS ``` The percentages must sum to no more than 100\. The priority statement implicitly counts against the total. If you only specify `priority` without a percentage, the priority queue is unbounded - never do this in production; a misbehaving SIP gateway can starve everything else. ### Bandwidth Statements Three forms exist: `bandwidth percent X` X percent of the interface bandwidth (or the parent shaper rate in hierarchical policies) `bandwidth X` (kbps) Absolute kbps guarantee `bandwidth remaining percent X` X percent of bandwidth remaining after priority queues are accounted for The remaining-percent form is useful when priority bandwidth varies (e.g. voice scales with call count). The classes that take "remaining" do not need to be re-tuned every time the priority allocation changes. ### WRED: Graceful Degradation WRED (Weighted Random Early Detection) randomly drops packets from a queue as it approaches full, biased toward higher drop precedence. The benefit: TCP flows respond to dropped packets by slowing down, which empties the queue gracefully rather than letting it fill and tail-drop everything. WRED works best on classes with lots of TCP traffic. Voice and video (UDP) do not benefit because UDP does not back off; for those classes, just rely on the priority queue's policer. ``` class TRANSACTIONAL bandwidth percent 25 random-detect dscp-based ! Drop AF23 first, then AF22, then AF21 ``` ## Policing: Hard Rate Limits A policer enforces a maximum rate. Packets above the rate are dropped immediately or re-marked to a lower-priority DSCP. Cisco implementations use a token-bucket algorithm: - Tokens accumulate in a bucket at the configured rate (e.g. 10 Mbps). - Each arriving packet consumes tokens proportional to its size. - If enough tokens are available, the packet conforms (passes). - If not enough tokens, the packet exceeds (drop or re-mark, depending on config). - Bucket has a burst size that allows short bursts above rate; refills constantly at the rate. Configuration: ``` policy-map RATE-LIMIT-INBOUND class class-default police 10000000 1500000 ! 10 Mbps with 1.5 MB burst conform-action transmit exceed-action drop interface GigabitEthernet0/0/0 service-policy input RATE-LIMIT-INBOUND ``` Policing's superpower: instant rate enforcement without buffering. Use it for: - SLA enforcement at network boundaries (carriers police their customers) - Subscriber rate plans (ISP enforcing a customer's contracted rate) - Protection against traffic floods (per-source-IP policers) - The built-in policer on LLQ priority queues (prevents voice from running away) Policing's weakness: dropping perfectly good packets. TCP responds by retransmitting and reducing congestion window. The throughput of TCP traffic against a policer is significantly lower than the policed rate. If you can shape instead of police for non-bursty traffic, do. ### Re-marking Instead of Dropping A common pattern: instead of dropping excess, re-mark to a lower-priority class: ``` class SCAVENGER police 5000000 conform-action set-dscp-transmit cs1 ! Confirm: keep CS1 exceed-action set-dscp-transmit cs0 ! Exceed: re-mark to BE ``` Excess scavenger traffic is not dropped - just demoted. If the network has spare capacity, the demoted traffic still gets through. If the network is congested, the demoted traffic is the first to drop. This "soft policing" gives you the rate enforcement without the hard cliff. ## Shaping: Smoothing Bursts A shaper buffers excess traffic and releases it at the configured rate. Like a policer, it uses a token-bucket algorithm; unlike a policer, the action for excess is "buffer for later" instead of "drop." Configuration: ``` policy-map SHAPE-WAN class class-default shape average 50000000 ! Shape to 50 Mbps average interface GigabitEthernet0/0/0 service-policy output SHAPE-WAN ``` Shape average smooths bursts to the configured rate. Shape peak (rarely used) allows bursting to a higher rate based on accumulated credit. The classic shaping use case: your branch has a 1 Gbps physical interface but a 50 Mbps contracted Metro Ethernet handoff. Without shaping, you send bursts at 1 Gbps and the carrier polices the excess (drops). With shaping, you smooth output to 50 Mbps and the carrier never has to police; no drops, no packet loss. Shaping has one obvious cost: latency. Buffered packets wait. For voice and other latency-sensitive traffic, this is a problem. The solution is hierarchical shaping. ## Hierarchical Shaping: The Production Pattern Hierarchical shaping wraps a queueing policy inside a shaping policy. The shaper smooths to the contracted rate; the inner queueing policy applies LLQ + CBWFQ within the shaped pipe. ``` policy-map CHILD-WAN-EGRESS class VOICE priority percent 30 class VIDEO bandwidth percent 30 class class-default bandwidth percent 40 fair-queue policy-map PARENT-SHAPER class class-default shape average 50000000 ! Shape to 50 Mbps service-policy CHILD-WAN-EGRESS ! Apply child queueing inside interface GigabitEthernet0/0/0 service-policy output PARENT-SHAPER ``` This pattern dominates production WAN edges. The parent shapes to the contracted rate (so the carrier never polices). The child applies LLQ inside the shaped pipe, so voice still gets priority queueing - but inside the 50 Mbps shaped pipe, not in the 1 Gbps physical interface. The percentages in the child policy are percentages of the parent shape rate, not the physical interface. `priority percent 30` in the child means 30 percent of 50 Mbps = 15 Mbps reserved for the priority queue. ## When to Use Each: A Decision Matrix Voice and other real-time on a congested WAN Use LLQ with built-in policer Why Strict priority + protection against priority abuse Multiple business apps competing for bandwidth Use CBWFQ with bandwidth guarantees Why Proportional service across classes Sub-rate WAN handoff (1 Gbps interface, 50 Mbps contract) UseHierarchical shaping Why Avoid carrier policing; preserve LLQ inside shaped pipe SLA enforcement at network boundary UsePolicing Why Hard rate limit with no buffering Per-subscriber rate plans Use Policing or hierarchical shaping per subscriber WhySimple at scale TCP traffic with bursty sources UseShaping (egress) Why Smooths bursts; preserves TCP throughput Protection against UDP floods UsePolicing Why Hard cap; UDP does not back off Demote-not-drop excess scavenger traffic Use Policing with re-mark action Why Soft enforcement; uses spare capacity when available ## Verification ``` ! See the policy structure and counters Router# show policy-map interface GigabitEthernet0/0/0 GigabitEthernet0/0/0 Service-policy output: PARENT-SHAPER Class-map: class-default shape (average) cir 50000000, bc 200000, be 200000 target shape rate 50000000 Service-policy : CHILD-WAN-EGRESS Class-map: VOICE (match-any) 12345 packets, 1234567 bytes 30 second offered rate 12000 bps, drop rate 0 bps Match: dscp ef (46) Priority: 30% (15000 kbps), burst bytes 375000, b/w exceed drops: 0 ! See queue depth (TX-ring) on a specific interface Router# show interfaces GigabitEthernet0/0/0 Output queue: 0/40 (size/max) ! Healthy Output queue: 38/40 (size/max) ! Bordering on tail-drops ``` The "drop rate" line per class is the most important diagnostic. Zero drops with significant offered rate = the class has enough bandwidth. Drops growing = class is starved. ## Anti-Patterns - **Bare `priority` in LLQ.** Always use `priority percent X` or `priority X` in kbps. The unbounded form lets a runaway flow destroy the rest of the policy. - **Policing where shaping would do.** If the source is under your control, shape (preserve TCP throughput). Police only at boundaries you cannot trust. - **Forgetting hierarchical shaping for sub-rate handoffs.** Without it, carrier polices, you get drops, voice quality suffers. - **WRED on UDP-only classes.** UDP does not back off; WRED becomes random drops with no benefit. Tail-drop is fine for UDP-heavy classes. - **Bandwidth percentages summing to over 100.** Cisco may accept the config but behavior is unpredictable. ## Summary Queueing schedules order on congested interfaces. Policing drops excess. Shaping buffers excess. The three are complementary, not interchangeable. LLQ is the modern queueing default; hierarchical shaping with LLQ inside is the production WAN pattern; policing is for SLA enforcement and protecting priority queues from abuse. Master the LLQ + CBWFQ + hierarchical shaping triple, and you have covered 90 percent of production QoS configurations. Bookmark this article alongside the [QoS cluster pillar](https://www.pinglabz.com/qos/) and the [Cisco MQC walkthrough](https://www.pinglabz.com/cisco-mqc/); lab every change before pushing to production. The penalty for misconfigured queueing/policing/shaping is voice quality issues that only show up under load. ### QoS Classification, Marking, and Trust Boundaries URL: https://www.pinglabz.com/qos-classification-marking-trust-boundary/ Last updated: 2026-06-13T20:08:41.000Z Classification, marking, and the trust boundary together form the policy edge of every QoS deployment. Get them right and the rest of QoS is mechanics. Get them wrong and you spend the next year chasing weird voice quality issues that always seem to happen when no one is around to capture packets. This article walks through what classification really means, how marking fits in, where to place trust boundaries, and the production patterns for the three or four scenarios you will actually encounter in an enterprise network. If you are building a new QoS deployment, auditing an existing one, or trying to figure out why a single user's traffic is somehow getting EF-class treatment, this is the discipline. ## What Classification and Marking Are Classification is the act of identifying what kind of traffic a packet represents. Voice? Video? Salesforce? Bulk backup? The output of classification is conceptual: "this packet belongs to class X." Marking is the act of writing that classification into the packet itself, in a field every downstream device can read without redoing the classification work. The output of marking is a bit pattern: a DSCP value in the IP header, a CoS value in the 802.1Q tag, or an MPLS EXP value in the MPLS label. The two are usually paired: classify and mark in one configuration step at the network edge, then trust the marks throughout the rest of the path. ## Why Mark at the Edge? Classification is expensive. Pattern matching on application signatures (NBAR), running ACLs against every packet, keeping flow state - all of that costs CPU and memory. On a busy WAN edge, classification can be the dominant cost. Marking shifts the cost from "every device re-classifies every packet" to "one device classifies, every device reads a few bits." This is what makes end-to-end QoS practical. A 1 Gbps link sees millions of packets per second; reading 6 bits of DSCP and indexing into a queue is much cheaper than parsing payloads. The discipline that makes this work: edge marking must be correct and trustworthy. If your edge marks 50 percent of voice traffic correctly and 50 percent as best-effort, your QoS policy is broken regardless of how clever your queueing config is. ## Where to Classify and Mark The general rule: as close to the traffic source as possible, on a device you trust to do it correctly. Three classification sources, in priority order: 1. **The source application itself.** Some applications mark their own DSCP correctly. Cisco IP phones do this for voice. Microsoft Teams sets DSCP markings for media. Modern enterprise apps from major vendors increasingly do. 2. **The first-hop infrastructure.** Access switch ports for end hosts; WAN edges for inbound flows from outside the network. Classify and mark at this point if the source did not, or if you do not trust the source's markings. 3. **Mid-network NBAR-based classification.** When a flow has crossed several hops without being marked correctly. Less ideal than edge marking but sometimes necessary for cloud-native traffic. The classic anti-pattern: classifying at every hop because you do not trust any prior hop. This works but is operationally expensive and indicates that your trust model is broken. ## The Trust Boundary The trust boundary is the perimeter inside which DSCP markings are honored. Outside the boundary, markings are suspect (or rewritten). Inside, every device respects the marks and applies policy accordingly. Common trust boundary placements: Generic PC on access port Trust settingUntrusted ActionRe-mark all to BE Cisco IP phone (voice VLAN) Trust setting Trust DSCP/CoS for voice VLAN only Action Conditional trust via CDP detection Trusted application server Trust settingTrust DSCP ActionPass through unchanged BYOD / guest device Trust settingUntrusted Action Re-mark to BE or scavenger Inter-switch trunk inside QoS domain Trust settingTrust DSCP and CoS ActionPass through WAN edge inbound from ISP Trust settingUntrusted Action Re-mark per your scheme; do not trust ISP's DSCP WAN edge inbound from MPLS provider with QoS contract Trust settingConditionally trusted Action Trust if carrier honors agreed markings; verify periodically The classic failure mode: a user marks every packet from their PC as DSCP EF (voice). Without re-marking at the access port, that PC's traffic gets priority queueing and starves the legitimate voice traffic. The fix: never trust DSCP from generic end hosts. Re-mark to BE or BG by default; trust only via the IP phone's conditional-trust mechanism for known traffic flows. ## Cisco IOS XE Configuration: The Patterns Three patterns cover most enterprise access ports. ### Pattern 1: Untrusted Access Port (Generic PC) ``` interface GigabitEthernet1/0/3 switchport mode access switchport access vlan 10 spanning-tree portfast spanning-tree bpduguard enable ! No QoS trust - default behavior is to re-mark to 0 (BE) ``` By default on Cisco access ports, incoming DSCP is rewritten to 0\. No explicit configuration required. Verify with: ``` Switch# show mls qos interface GigabitEthernet1/0/3 GigabitEthernet1/0/3 trust state: not trusted trust mode: not trusted ``` ### Pattern 2: IP Phone with PC Behind It (Conditional Trust) ``` interface GigabitEthernet1/0/5 switchport mode access switchport access vlan 10 ! Data VLAN switchport voice vlan 20 ! Voice VLAN mls qos trust device cisco-phone ! Conditional trust via CDP spanning-tree portfast spanning-tree bpduguard enable ``` The `mls qos trust device cisco-phone` tells the switch to trust DSCP/CoS only when CDP detects a Cisco IP phone on the port. If the phone is unplugged, the switch reverts to untrusted and rewrites DSCP to 0. The IP phone itself handles the per-VLAN trust: it tags its own voice traffic with DSCP EF on VLAN 20, and forwards PC traffic from the data VLAN unchanged (which the access port then re-marks because the PC is untrusted). ### Pattern 3: Explicit Classification and Marking For traffic the source did not mark and you cannot use simple trust: ``` ! Define classes class-map match-any VOICE match protocol rtp audio ! NBAR-based class-map match-any VIDEO-CONF match protocol rtp video class-map match-any TRANSACTIONAL match access-group name SALESFORCE-IPS class-map match-any SCAVENGER match access-group name BACKUP-IPS ip access-list extended SALESFORCE-IPS permit ip any 13.108.0.0 0.0.255.255 ip access-list extended BACKUP-IPS permit tcp any any eq 873 ! rsync ! Define actions policy-map MARK-INGRESS class VOICE set dscp ef class VIDEO-CONF set dscp af41 class TRANSACTIONAL set dscp af21 class SCAVENGER set dscp cs1 class class-default set dscp default ! Apply at ingress interface GigabitEthernet0/0/2 service-policy input MARK-INGRESS ``` This pattern is common at WAN edges where the upstream did not mark. NBAR2 (with `match protocol`) handles application-aware classification for hard-to-pin-down apps. The full MQC walkthrough is in [Cisco MQC: The Operator's Walkthrough](https://www.pinglabz.com/cisco-mqc/). ## Standardized Marking Values Use standardized DSCP values everywhere. The [DSCP article](https://www.pinglabz.com/dscp-ip-precedence-explained/) has the full reference. The minimum set for any production QoS deployment: 46 ClassEF UseVoice (RTP) 40 ClassCS5 Use Broadcast video (one-way streaming) 34 ClassAF41 Use Video conferencing (Zoom, Teams, Webex) 26 ClassAF31 UseMultimedia streaming 24 ClassCS3 Use Voice/video signaling (SIP, H.323) 18 ClassAF21 Use Transactional business apps 10 ClassAF11 Use Bulk data (email, file) 8 ClassCS1 Use Scavenger / lower than best-effort 0 ClassBE UseDefault Defining your own arbitrary DSCP values (like 35 because "it's between AF31 and AF41") will cause grief. Stick to the standardized PHBs. Every Cisco platform has built-in QoS maps for these values; non-standard values do not get the same hardware acceleration. ## Layer 2: CoS and the 802.1Q Connection CoS (802.1p Priority) lives in the 802.1Q tag, three bits, 8 values. Inside switched Layer 2 segments, CoS is what gets honored. Across Layer 3 hops, the 802.1Q tag is stripped and only DSCP survives. Cisco's standard mapping derives CoS from the top three bits of DSCP: EF (46) 5 CS5 (40) 5 AF41 (34) 4 AF31 (26), CS3 (24) 3 AF21 (18) 2 AF11 (10), CS1 (8) 1 BE (0) 0 You can override the mapping with `mls qos map dscp-cos` or equivalent on the platform. Most production deployments leave it at default. See [802.1Q VLAN Tag Explained](https://www.pinglabz.com/802-1q-vlan-tag-explained/) for where CoS lives in the frame. ## Auditing an Existing Trust Boundary Three things to verify on any existing deployment: 1. **Every access port for end hosts is untrusted.** `show mls qos interface GigabitEthernet1/0/3` should show "trust state: not trusted" or equivalent. If it shows "trust dscp" without conditional trust, end hosts can mark whatever they want. 2. **Voice VLAN trust is conditional, not absolute.** Verify with `show interfaces switchport` that voice VLAN config exists, and `show mls qos interface` that conditional trust (cisco-phone) is configured rather than absolute trust. 3. **WAN edges re-validate inbound markings.** Verify with `show policy-map interface` that an input service-policy exists and the class hit counts make sense. The single most common audit finding: an access port that was configured for an IP phone five years ago, the phone got swapped for a USB phone or a softphone on the PC, and the trust setting was never updated. Now that PC's marked-as-EF traffic gets priority over real voice. The fix: re-audit periodically. ## Re-marking at the Trust Boundary Two re-marking patterns are common: **Strict re-marking (paranoid).** Default-deny-style: mark everything to BE on ingress, except specific traffic that matches your classification rules. Best for high-security environments where you cannot tolerate untrusted markings escaping into the QoS domain. **Permissive re-marking (pragmatic).** Trust DSCP from designated sources, re-mark only obvious abuse (anything from end-host PCs marked above CS3, for example). Fewer rules; works well when you have control over most sources. Most production deployments use permissive on internal trunks and strict at WAN edges. The cost-benefit shifts based on what is on the other side. ## Anti-Patterns to Avoid - **Trusting CoS from end hosts.** CoS only exists on tagged frames. End hosts on access ports send untagged frames, so CoS is moot - but if a malicious user sends tagged frames (a misconfigured VM, an attacker), trusting CoS lets them choose their own queue. - **Re-marking at every hop.** Operationally wasteful. Indicates that the trust model is broken. - **Mixing marking schemes.** Different parts of the network using different DSCP values for the same traffic class. One sub-network uses AF31 for streaming video; another uses CS5; queueing policies have to handle both. Standardize. - **Defining custom DSCP values for "special" traffic.** The standardized PHBs cover everything. Custom values do not get hardware acceleration on most platforms. - **Relying on application markings without validation.** "Microsoft Teams sets DSCP correctly" is true 95 percent of the time. The other 5 percent is misconfigured Teams clients sending video as best-effort. NBAR-based classification at the edge as a backstop catches it. ## Summary Classification and marking at the trust boundary are the foundation of every QoS deployment. Classify and mark once at a controlled edge, trust those markings throughout the QoS domain, and re-mark or re-validate at any boundary where trust changes. The standard Cisco patterns (untrusted access port, conditional trust for IP phones, explicit MQC classification at WAN edges) cover most enterprise scenarios. The discipline matters more than the configuration. Audit the trust boundaries periodically. Standardize on the IETF PHBs. Never trust DSCP from generic end hosts. Bookmark this article alongside the [QoS cluster pillar](https://www.pinglabz.com/qos/) and the [DSCP article](https://www.pinglabz.com/dscp-ip-precedence-explained/) for the marking values, and the [Cisco MQC walkthrough](https://www.pinglabz.com/cisco-mqc/) for the configuration grammar. ### Cisco MQC (Modular QoS CLI): The Operator's Walkthrough URL: https://www.pinglabz.com/cisco-mqc/ Last updated: 2026-06-13T20:08:41.000Z The Modular QoS CLI (MQC) is the Cisco command framework for configuring QoS on IOS, IOS XE, and IOS XR routers and Layer 3 switches. It replaced the legacy QoS commands (CAR, custom queueing, priority queueing) with a clean three-construct model: classify, then describe what to do with each class, then apply the policy to an interface. Once you understand MQC, every Cisco QoS feature you will ever encounter follows the same pattern. This article walks through the three MQC constructs (class-map, policy-map, service-policy), the most common match conditions and actions, where to apply policies (input vs output), the verification commands, and the patterns that production deployments use repeatedly. If you are studying for CCNP, configuring QoS on a Cisco device for the first time, or just trying to remember whether bandwidth statements add up to 100 percent, this is the operator's walkthrough. ## Why MQC Exists Before MQC, every QoS feature in Cisco IOS had its own configuration syntax. Committed Access Rate (CAR) had a flat command for rate limiting. Custom queueing had its own command set. Priority queueing was different again. Adding a new feature meant learning a new syntax. MQC unified all of this. The three constructs - class-map for classification, policy-map for action, service-policy for application - cover every QoS feature: marking, queueing, policing, shaping, WRED, NBAR, application-aware policy, all using the same grammar. New features (QoS for SD-WAN, AppQoE, AVC) ship as new actions inside policy-map; the surrounding structure does not change. The three-construct model also makes QoS configurations easy to reason about. You read a service-policy statement, find the policy-map it references, find the class-maps it references, and you have the full classification and action tree. No global lookups, no implicit interactions. ## The Three MQC Constructs class-map Purpose Define how to classify traffic (what to match) Reusable? Yes - one class-map can be referenced by multiple policy-maps policy-map Purpose Define what to do with each classified class (mark, queue, police, shape) Reusable? Yes - one policy-map can be applied to many interfaces service-policy Purpose Apply a policy-map to an interface, ingress or egress Reusable? n/a - per-interface, per-direction The flow: traffic arrives at an interface. If a service-policy is configured for that direction, MQC walks the policy-map. The policy-map iterates through its classes; the first matching class-map wins. The actions specified for that class apply. ## class-map: Defining Classification A class-map answers "what counts as this kind of traffic?" The match conditions are the heart of QoS classification: `match access-group` Match an ACL (most flexible; arbitrary L3/L4 matching) `match dscp X` Match incoming DSCP value `match cos X` Match incoming Layer 2 CoS `match precedence X` Match legacy IP Precedence `match protocol X` NBAR-based application matching (rtp audio, http, dns, etc.) `match ip rtp` Match RTP traffic by UDP port range `match input-interface` Match by ingress interface (rare) `match qos-group` Internal label set earlier in the policy chain `match vlan` Match by VLAN ID (Layer 3 switches) `match application` NBAR2 / Application Visibility and Control (AVC) The two modifiers: - **match-any** (default): packet matches if any single condition matches - **match-all**: packet matches only if every condition matches Examples: ``` ! Match incoming DSCP EF (voice) class-map match-any VOICE match dscp ef ! Match incoming RTP audio (NBAR-based) class-map match-any VOICE-NBAR match protocol rtp audio ! Match Salesforce traffic by IP prefix class-map match-any SALESFORCE match access-group name ACL-SALESFORCE ip access-list extended ACL-SALESFORCE permit tcp any host 13.108.0.1 eq 443 ! Match Office 365 (NBAR2 AVC) class-map match-any OFFICE-365 match application office-365-mail match application office-365-skype match application office-365-web ``` NBAR (Network-Based Application Recognition) and its successor NBAR2 are critical for modern QoS. They identify applications by deep packet inspection, not just port numbers, which means they catch SaaS apps that share ports (everything is HTTPS now) and apps that use dynamic ports. ## policy-map: Defining Actions A policy-map says "for traffic matched by this class-map, do this." Actions split into four families: Marking `set dscp`, `set cos`, `set precedence`, `set qos-group`, `set mpls experimental imposition` Queueing `priority`, `bandwidth`, `fair-queue`, `queue-limit`, `random-detect` Policing `police` (with rate, burst, conform/exceed actions) Shaping `shape average`, `shape peak`, `shape adaptive` A complete marking policy: ``` policy-map MARK-INGRESS class VOICE set dscp ef class VIDEO set dscp af41 class TRANSACTIONAL set dscp af21 class SCAVENGER set dscp cs1 class class-default set dscp default ``` A complete queueing policy (LLQ for voice, CBWFQ for everything else): ``` policy-map WAN-EGRESS class VOICE priority percent 10 ! Strict-priority queue, capped at 10% class VIDEO bandwidth percent 30 ! Guaranteed 30% class TRANSACTIONAL bandwidth percent 25 ! Guaranteed 25% random-detect dscp-based ! WRED for graceful degradation class SCAVENGER bandwidth percent 1 ! Tiny guarantee class class-default bandwidth percent 34 ! Remainder random-detect ``` The bandwidth percentages must sum to no more than 100\. The priority statement implicitly counts against the total. `class-default` is implicit if you do not declare it; the configuration above is explicit for clarity. A combined marking+queueing policy in one policy-map: ``` policy-map ALL-IN-ONE class VOICE set dscp ef priority percent 10 class VIDEO set dscp af41 bandwidth percent 30 ``` This works for output policies. Input policies typically only mark or police - you cannot meaningfully queue traffic that is arriving. ## service-policy: Applying to an Interface The service-policy command attaches a policy-map to an interface, in either direction: ``` interface GigabitEthernet0/0/1 service-policy input MARK-INGRESS ! Mark on the way in service-policy output WAN-EGRESS ! Queue/shape on the way out ``` An interface can have one input and one output service-policy. To do multiple things on the same direction, combine them into one policy-map (see ALL-IN-ONE above). Hierarchical policies (a policy that itself references another policy) are supported via `service-policy` inside a class: ``` policy-map CHILD-WAN-EGRESS class VOICE priority percent 30 class class-default fair-queue policy-map PARENT-SHAPER class class-default shape average 50000000 ! Shape to 50 Mbps service-policy CHILD-WAN-EGRESS ! Apply child queueing policy interface GigabitEthernet0/0/1 service-policy output PARENT-SHAPER ``` This is the standard pattern for branch WAN edges where you have a 1 Gbps physical interface but a contracted 50 Mbps WAN service. The parent shapes to 50 Mbps; the child does LLQ within that shaped pipe. ## Input vs Output: Where to Apply What input (ingress) Useful for Marking, classification, policing, NBAR-based identification Not useful for Queueing (no congestion to manage) output (egress) Useful for Queueing, scheduling, shaping, WRED, marking Not useful for Some forms of NBAR (depends on platform) The standard pattern in production: - **Access port input:** classify untrusted traffic and re-mark to BE; trust traffic from IP phones via conditional trust. - **Distribution / core inputs:** trust the markings (no input service-policy or one that just verifies). - **WAN egress:** the queueing policy that does LLQ + CBWFQ + shaping. - **WAN input:** usually no QoS - the carrier already shaped traffic to fit your contract. ## Verification Commands Once a policy is applied, the standard verification chain: ``` ! Show the policy-map definition Router# show policy-map WAN-EGRESS ! Show the policy applied to a specific interface, with hit counts Router# show policy-map interface GigabitEthernet0/0/1 GigabitEthernet0/0/1 Service-policy output: WAN-EGRESS Class-map: VOICE (match-any) 1234567 packets, 198351264 bytes 30 second offered rate 8000 bps, drop rate 0 bps Match: dscp ef (46) Priority: 10% (10000 kbps), burst bytes 250000, b/w exceed drops: 0 Conform: 1234567 packets / 198351264 bytes Exceed: 0 packets / 0 bytes Class-map: VIDEO ... ! Show class-map definition Router# show class-map VOICE ``` The most useful column: drop counters per class. If your "Conform" packets are increasing but "Exceed" or queue drops are also growing, the class is bandwidth-starved. If everything is "Conform" with zero drops, the policy is healthy. ## Common Production Patterns Three patterns cover most production QoS deployments: **1\. Edge marking + WAN queueing.** Access ports classify and mark traffic to standardized DSCP values. WAN edge applies an output policy with LLQ for voice and CBWFQ for everything else. Internal devices trust DSCP. This is the dominant enterprise pattern. **2\. Hierarchical shaping for sub-rate WAN handoffs.** Physical interface is faster than the contracted WAN rate. Parent policy shapes to the contracted rate; child does LLQ + CBWFQ within the shaped pipe. Prevents the carrier from policing your traffic. **3\. NBAR-based ingress marking for mid-network insertion.** A WAN edge that needs to classify traffic the application servers did not mark. NBAR2 + AVC identifies the application via DPI; an input service-policy marks the DSCP. Subsequent hops trust the marking. ## Operational Tips - Verify `show policy-map interface` after every change. Counters should make sense; if a class has zero hits, the class-map is wrong. - Always declare `class class-default` explicitly. The implicit default works but the explicit version is clearer for the next operator reading your config. - Bandwidth percentages should sum to 100 or less. Cisco will accept configs that exceed 100 (with a warning) but the behavior is unpredictable. - Priority class with `priority percent X` auto-polices to X percent. The `priority` bare keyword is unbounded; avoid it on production WAN. - WRED requires bandwidth allocation. Add `random-detect` only inside classes that have `bandwidth` declared. - For platform-specific behavior (ASR vs ISR vs Catalyst), always check the QoS configuration guide for that release. Some features have hardware vs software limitations. ## Summary MQC's three constructs (class-map, policy-map, service-policy) cover every QoS feature on Cisco IOS, IOS XE, and IOS XR. Once you understand the pattern, every new QoS feature - AVC, NBAR2, app-aware steering, WRED, hierarchical shaping - fits the same grammar. Define classification once; describe actions per class; apply to interfaces. The dominant production patterns are edge marking + WAN queueing, hierarchical shaping for sub-rate WAN handoffs, and NBAR-based mid-network marking. Master the three constructs, the verification chain, and the trust-boundary discipline. Bookmark this article alongside the [QoS cluster pillar](https://www.pinglabz.com/qos/) as your day-to-day MQC reference. ### DSCP and IP Precedence Explained Byte by Byte URL: https://www.pinglabz.com/dscp-ip-precedence-explained/ Last updated: 2026-06-13T20:08:41.000Z DSCP (Differentiated Services Code Point) is the 6-bit Layer 3 marking that tells every router in a path how to treat a given packet. It lives in the IP header (the same byte that used to hold IP Precedence in the original IPv4 spec) and is the dominant QoS marking standard in modern enterprise and service-provider networks. If you have ever set `set dscp ef` in a Cisco policy-map and wondered what that 6-bit value actually does on the wire, this is the byte-level reference. This article walks through the history (TOS to IP Precedence to DSCP), the byte format, the standardized PHB (Per-Hop Behavior) values, the DSCP-to-CoS-to-MPLS-EXP mapping, and the trust-boundary discipline that makes DSCP-based QoS work in production. If you are studying for CCNP, designing a QoS rollout, or troubleshooting why your voice traffic is not getting priority, this is the foundation. ## From TOS to IP Precedence to DSCP The byte at offset 1 in the IPv4 header has been called several things across the protocol's history: Type of Service (ToS) Era1981 Format 3-bit precedence + 4-bit ToS + 1 unused StandardRFC 791 Same byte; renamed to TOS field Era1992 Format 3-bit IP Precedence + 4-bit TOS bits + 1 unused StandardRFC 1349 DiffServ field (DSCP + ECN) Era1998 Format6-bit DSCP + 2-bit ECN StandardRFC 2474, RFC 3168 The 1998 redefinition was the important one. The IETF realized the original 8 IP Precedence values were too coarse, and the 4 ToS bits were essentially unused. They redefined the same byte to carry 6 bits of DSCP plus 2 bits of ECN (Explicit Congestion Notification). The 6 DSCP bits give 64 possible values - enough to express granular per-class behavior across diverse service classes. In IPv6, the same field exists as the 8-bit Traffic Class field with the same DSCP/ECN split. End-to-end QoS works identically across IPv4 and IPv6. ## The Byte Format ``` +---+---+---+---+---+---+---+---+ | DSCP | ECN | | 6 bits | 2bits| +---+---+---+---+---+---+---+---+ | C | C | C | D | D | D | E | E | +---+---+---+---+---+---+---+---+ Class Selector + Drop Precedence + ECN ``` The 6 DSCP bits split conceptually into two parts: - **Class Selector (CS):** bits 7-5 (3 bits) - which class of service. Maps directly to the legacy IP Precedence values; setting bits 4-3 to zero gives a backwards-compatible CS value. - **Drop Precedence:** bits 4-3 (2 bits) - within a class, which packets get dropped first under congestion. - **Bit 2 reserved (always 0 in current PHBs).** The ECN bits (bits 1-0) are not part of QoS classification; they let routers signal congestion to TCP endpoints without dropping packets. ## Per-Hop Behaviors (PHBs): The Standardized DSCP Values The IETF standardized a set of PHBs that map to specific DSCP values. Every QoS-aware device should treat traffic with these markings consistently. Three PHB families exist: ### Default (Best Effort) BE / Default DSCP (binary)000000 DSCP (decimal)0 Use No QoS preference; default for unmarked traffic ### Class Selector (CS) - Backwards Compatible with IP Precedence 000000 NameCS0 DSCP (decimal)0 Old IP Prec0 (Routine) UseDefault; same as BE 001000 NameCS1 DSCP (decimal)8 Old IP Prec1 (Priority) Use Scavenger / lower than best-effort 010000 NameCS2 DSCP (decimal)16 Old IP Prec2 (Immediate) Use Network management (OAM) 011000 NameCS3 DSCP (decimal)24 Old IP Prec3 (Flash) UseSignaling (SIP, H.323) 100000 NameCS4 DSCP (decimal)32 Old IP Prec4 (Flash Override) Use Real-time interactive (gaming) 101000 NameCS5 DSCP (decimal)40 Old IP Prec5 (Critical) UseBroadcast video 110000 NameCS6 DSCP (decimal)48 Old IP Prec 6 (Internetwork Control) Use Routing protocols (OSPF, BGP, EIGRP) 111000 NameCS7 DSCP (decimal)56 Old IP Prec7 (Network Control) UseNetwork control plane CS6 is what most routing protocols mark themselves with by default. `show ip ospf neighbor` and equivalent: under the hood, OSPF Hellos carry DSCP CS6 (48). This is why a misconfigured QoS policy that drops CS6 traffic causes routing protocols to flap. ### Assured Forwarding (AF) - For Bandwidth-Guaranteed Classes AF defines four classes (AF1-AF4), each with three drop precedence levels: AF11 (10) ClassAF1 Med DropAF12 (12) High DropAF13 (14) Typical use Bulk data (email, file transfer) AF21 (18) ClassAF2 Med DropAF22 (20) High DropAF23 (22) Typical use Transactional / business apps AF31 (26) ClassAF3 Med DropAF32 (28) High DropAF33 (30) Typical use Multimedia streaming (one-way) AF41 (34) ClassAF4 Med DropAF42 (36) High DropAF43 (38) Typical use Video conferencing (two-way) The naming is logical: AFxy where x is class and y is drop precedence (1=low/most likely to be kept, 3=high/most likely to be dropped). Within a class, all three drop precedence levels share queueing; under congestion, the higher drop-precedence packets are discarded first. Drop precedence pairs naturally with WRED (Weighted Random Early Detection): set thresholds so AF13 packets get dropped earlier than AF11 in the same class. ### Expedited Forwarding (EF) - For Voice and Strict-Priority 101110 NameEF DSCP (decimal)46 Use Voice (VoIP RTP); strict priority queueing EF is the priority class. Traffic marked EF goes to the LLQ priority queue at every congested egress and is served before everything else. Use EF only for traffic that genuinely needs sub-millisecond latency variance: VoIP RTP voice payload, possibly real-time control loops. Because the priority queue is served first, you must police EF at the policy level - if voice traffic explodes (a misbehaving SIP gateway, a DDoS, a misconfiguration), unbounded EF would starve everything else. Cisco's LLQ has a built-in policer for the priority class. ## Cisco's Mapping: DSCP to CoS to MPLS EXP Inside a heterogeneous network, the same flow may be tagged with DSCP at Layer 3, CoS at Layer 2, and MPLS EXP at the MPLS service provider boundary. Cisco maintains default mappings between these markings, which you can override via QoS maps. 46 DSCP classEF Default CoS5 MPLS EXP5 40 DSCP classCS5 Default CoS5 MPLS EXP5 34 DSCP classAF41 Default CoS4 MPLS EXP4 26 DSCP classAF31 Default CoS3 MPLS EXP3 24 DSCP classCS3 Default CoS3 MPLS EXP3 18 DSCP classAF21 Default CoS2 MPLS EXP2 10 DSCP classAF11 Default CoS1 MPLS EXP1 8 DSCP classCS1 Default CoS1 MPLS EXP1 0 DSCP classBE Default CoS0 MPLS EXP0 The default DSCP-to-CoS mapping is the top 3 bits of the DSCP value (bits 7-5). EF (101110) becomes CoS 5 (101). AF41 (100010) becomes CoS 4 (100). This is the Class Selector value of the DSCP. In switched Layer 2 segments where the 802.1Q tag exists, CoS is what gets honored. Packets crossing a Layer 3 routing hop lose their CoS (the 802.1Q tag is stripped) but retain DSCP (it lives in the IP header). The standard Cisco pattern: trust DSCP at the Layer 3 boundary, derive CoS from DSCP automatically. See [802.1Q VLAN Tag Explained](https://www.pinglabz.com/802-1q-vlan-tag-explained/) for where CoS lives in the frame. ## Setting DSCP on Cisco IOS XE The MQC pattern for marking traffic with DSCP: ``` ! Classify traffic class-map match-any VOICE-RTP match protocol rtp audio class-map match-any VIDEO match protocol rtp video class-map match-any TRANSACTIONAL match access-group name SALESFORCE-ACL ! Mark traffic policy-map MARK-INGRESS class VOICE-RTP set dscp ef class VIDEO set dscp af41 class TRANSACTIONAL set dscp af21 class class-default set dscp default ! Apply at ingress interface GigabitEthernet0/0/1 service-policy input MARK-INGRESS ``` Verify with: ``` Router# show policy-map interface GigabitEthernet0/0/1 Router# show class-map VOICE-RTP ``` The full MQC walkthrough is in (article forthcoming). ## DSCP and Trust Boundaries The temptation to trust incoming DSCP markings from end hosts is the single most common QoS configuration mistake. Hosts can mark every packet as DSCP EF and starve the legitimate voice traffic. Always re-mark or re-validate at access ports. The Cisco default for access ports is "untrusted" - incoming DSCP is rewritten to 0 (BE) unless you explicitly trust the device. Common patterns: Generic PC No trust; mark everything to BE Cisco IP phone (in voice VLAN) Trust DSCP/CoS for voice VLAN; mark data VLAN to BE Trusted application server Trust DSCP set by the application Inter-switch trunk Trust DSCP and CoS WAN edge inbound from ISP Re-mark or zero (the ISP's DSCP is not your DSCP) Implementation on the access port: ``` ! Trust the IP phone but not the PC interface GigabitEthernet1/0/5 switchport mode access switchport access vlan 10 switchport voice vlan 20 mls qos trust device cisco-phone ! Conditional trust spanning-tree portfast spanning-tree bpduguard enable ``` Cisco's "conditional trust" automatically detects whether a Cisco IP phone is connected (via CDP) and trusts only when the phone is present. If the phone is unplugged, trust drops automatically. ## DSCP in IPv6 IPv6 carries DSCP in the 8-bit Traffic Class field (top 6 bits = DSCP, bottom 2 = ECN), exactly like IPv4\. All the standardized PHBs and DSCP values work identically. `set dscp ef` in a Cisco policy-map applies to both IPv4 and IPv6 traffic that matches the class, and the marking is preserved end-to-end. ## IP Precedence: When You Still See It IP Precedence is the legacy 3-bit predecessor of DSCP. Cisco still supports it for backwards compatibility, and you may see it in older configurations. Mapping: Routine IP Precedence0 Equivalent DSCP CSCS0 Priority IP Precedence1 Equivalent DSCP CSCS1 Immediate IP Precedence2 Equivalent DSCP CSCS2 Flash IP Precedence3 Equivalent DSCP CSCS3 Flash Override IP Precedence4 Equivalent DSCP CSCS4 Critical IP Precedence5 Equivalent DSCP CSCS5 Internetwork Control IP Precedence6 Equivalent DSCP CSCS6 Network Control IP Precedence7 Equivalent DSCP CSCS7 For new configurations, always use DSCP. IP Precedence has 8 values; DSCP has 64\. There is no functional reason to limit yourself to the legacy field. ## Summary DSCP is the 6-bit Layer 3 marking that drives modern QoS. It lives in the IP header (top 6 bits of the ToS / Traffic Class byte), supports 64 values, and is end-to-end across IPv4 and IPv6\. The standardized PHBs (CS, AF, EF) cover every common use case; the AF family adds drop precedence within a class for graceful WRED-driven degradation under congestion. Master the standardized values (EF=46 for voice, AF41=34 for video conferencing, CS5=40 for streaming video, AF21=18 for transactional, BE=0 for everything else), the Cisco mapping to CoS and MPLS EXP, and the trust-boundary discipline that prevents end hosts from gaming the marking. Bookmark this article alongside the [QoS cluster pillar](https://www.pinglabz.com/qos/) as your byte-level DSCP reference. ### SD-WAN Security and the SASE Convergence URL: https://www.pinglabz.com/sd-wan-security/ Last updated: 2026-06-13T20:08:42.000Z SD-WAN by itself solves the WAN routing problem. It does not solve the security problem of what to do with all the internet-bound traffic now leaving every branch directly. That gap is what SASE (Secure Access Service Edge) fills. By 2026 the SD-WAN and SASE markets are effectively merging, and any SD-WAN evaluation that does not also evaluate SASE is incomplete. This article walks through the security model SD-WAN itself provides, the gap that direct internet break-out creates, the SASE service stack that fills it, the ZTNA story, and how the major SD-WAN vendors approach the security integration. If you are evaluating SD-WAN, planning a SASE rollout, or trying to figure out where your existing branch firewall fits in the new architecture, this is the reference. ## What SD-WAN Itself Secures Every modern SD-WAN platform provides a baseline of security primitives: - **IPsec encryption between edges.** All overlay traffic between SD-WAN edges is encrypted by default with AES-256 or stronger. The underlying transport (broadband, MPLS, LTE) carries opaque encrypted packets. - **Mutual certificate authentication.** Edges and controllers use mutual TLS with certificates issued by the platform's CA. Rogue devices cannot join the fabric. - **Per-tunnel keys with periodic re-keying.** Compromise of one key does not compromise the whole fabric. - **Segmentation via VPNs/VRFs.** Different traffic types (corporate, guest, IoT, voice) live in separate VPNs that cannot reach each other without explicit policy. This is critical for compliance environments. - **Edge-level ACLs.** Allow/deny rules at each WAN Edge for what traffic can enter or leave the overlay. This is enough security for the SD-WAN-to-SD-WAN traffic. The fabric is internally secure: an attacker cannot snoop on the overlay or impersonate an edge without compromising the certificate infrastructure. ## The Gap: Internet-Bound Traffic The architectural problem SD-WAN creates: by enabling direct internet break-out from every branch, it bypasses the central security inspection stack that legacy WANs forced all traffic through. In legacy WAN, every flow from a branch to the internet traversed: ``` Branch -> MPLS -> HQ -> Central Firewall -> Central Proxy -> Internet ``` The central firewall and proxy applied URL filtering, malware scanning, DLP, CASB, and TLS inspection. Every flow was inspected before reaching the internet. SD-WAN's direct internet break-out shortcuts this: ``` Branch -> SD-WAN Edge -> Local ISP -> Internet ``` The user experience improves dramatically (no backhaul tax for cloud apps), but the security inspection has nowhere to live. Without explicit replacement, you have just removed every URL filter, malware scanner, and DLP control from your perimeter. Three responses to this gap exist in production: 1. **Branch firewalls.** Deploy a hardware firewall at every branch. Inspect locally. Operationally heavy and expensive at scale. 2. **Backhaul some traffic.** Send sensitive flows back to the central inspection stack; let SaaS go direct. Compromise position; loses some of the SD-WAN performance benefits. 3. **SASE.** Send internet-bound traffic to a cloud security service for inspection. Architecturally clean. SASE is the dominant 2026 answer. Branch firewalls remain for environments with strict local-inspection requirements; backhaul is the fallback for organizations not yet ready for cloud security. ## SASE: The Service Stack SASE (Secure Access Service Edge, Gartner's term, pronounced "sassy") is a category of cloud-delivered security services that sits between the SD-WAN edge and the internet. The branch SD-WAN sends internet-bound traffic to the SASE provider's nearest cloud PoP, which inspects the traffic and forwards it (or blocks it). The SASE service stack: SWG (Secure Web Gateway) Job URL filtering, malware scanning, TLS inspection of outbound web traffic Replaces (legacy)On-prem proxy CASB (Cloud Access Security Broker) Job Visibility and control over SaaS usage; data protection in cloud apps Replaces (legacy) SaaS-specific point products FWaaS (Firewall-as-a-Service) Job Stateful firewall inspection for non-web traffic Replaces (legacy) Branch or central firewall ZTNA (Zero Trust Network Access) Job Identity-based application access, replacing VPN Replaces (legacy)Remote-access VPN DLP (Data Loss Prevention) Job Inspect outbound flows for sensitive content; block or redact Replaces (legacy)On-prem DLP RBI (Remote Browser Isolation) Job Render risky web content in a remote container; user sees pixels only Replaces (legacy) Newer category; no direct legacy equivalent A user at a branch types `www.example.com` in their browser. The branch SD-WAN edge sees the flow is internet-bound and steers it to the nearest SASE PoP. The PoP terminates the user's TLS, inspects content, applies DLP rules, runs malware scanning, and forwards to `example.com`. The response comes back through the same path. Total added latency: typically 5-30 ms depending on PoP proximity. ## ZTNA: The VPN Replacement Story Zero Trust Network Access deserves a callout because it is the most architecturally distinct component of SASE. Legacy remote-access VPN (Cisco AnyConnect, GlobalProtect, etc.) puts a remote user "on the corporate network" - they get an internal IP, can talk to anything that ACLs permit, and the trust model is "if you have a VPN tunnel, you are trusted." ZTNA flips this: - Users connect to the SASE provider, not to the corporate network. - Each application is published as a ZTNA resource with its own access policy. - Users authenticate (SSO, MFA, device posture) and the SASE provider creates a per-application tunnel. - Users never see the network; they see only the applications they are authorized to access. - Lateral movement is impossible because there is no network to move on. Operationally, ZTNA is harder to deploy than VPN because every application needs to be published and have access policy authored. The benefit: the user experience is better (no VPN client clunkiness, faster connections, no full-tunnel latency for personal traffic) and the security model is dramatically tighter. Most enterprises in 2026 are mid-migration from VPN to ZTNA. New applications get published via ZTNA; legacy applications continue on VPN until migration is funded. Both modes coexist for years. ## Vendor Approaches to SD-WAN + Security Cisco Catalyst SD-WAN SASE strategy Integrate with Cisco Umbrella (SWG/DNS), Cisco Secure Access (ZTNA + SWG) Strengths Single Cisco vendor relationship; tight integration with Catalyst SD-WAN policy engine Trade-offs Best-of-breed customers may prefer dedicated SASE vendors VMware VeloCloud (Broadcom) SASE strategy Integrate with Symantec Cloud SWG, Lookout CASB (post-Broadcom acquisition) Strengths Strong CASB story; acceptable SWG Trade-offs Broadcom acquisition has unsettled some customers Fortinet SASE strategy Native FortiSASE (single platform combining FortiGate SD-WAN with FortiSASE) Strengths Single-vendor simplicity for Fortinet shops Trade-offs Less proven at enterprise scale than Cisco/VMware Versa SASE strategy Versa SASE (single platform for SD-WAN + SASE from inception) Strengths Architecturally clean; designed as one product Trade-offs Smaller installed base; less ecosystem maturity Palo Alto Networks SASE strategy Prisma SASE (Prisma SD-WAN + Prisma Access) Strengths Best-in-class SWG/CASB heritage; strong threat intel Trade-offs Two acquisition lineages (CloudGenix + Globalprotect) integrated into one product HPE Aruba EdgeConnect SASE strategy EdgeConnect plus Axis Security ZTNA (acquired); pluggable third-party SWG Strengths Open architecture for best-of-breed Trade-offs More integration work for the customer Cato Networks SASE strategy Single-vendor SASE from inception Strengths Genuinely unified product; strong global PoP coverage Trade-offs Newer in the enterprise space; smaller ecosystem The decision in 2026 is increasingly between single-vendor SASE (Cato, Versa, Fortinet, Palo Alto Prisma) and best-of-breed (mix and match, e.g. Cisco SD-WAN + Zscaler SWG + Netskope CASB). Single-vendor wins on operational simplicity; best-of-breed wins on individual product quality. Both are valid; neither is wrong. ## Macro-Segmentation: SD-WAN as Identity Boundary Beyond the per-flow security story, SD-WAN serves as a macro-segmentation boundary in modern zero-trust designs. Each branch is a network segment. Each VPN within the SD-WAN is a sub-segment. Combined with 802.1X / TrustSec at the access layer, you get a layered identity-aware fabric: - Access layer (802.1X/SGT): identifies users and devices, tags traffic with security group tags. See the [802.1X cluster](https://www.pinglabz.com/802-1x/). - SD-WAN edge: enforces per-VPN segmentation, applies SD-WAN data policy. - SASE cloud: inspects flows leaving the fabric for the internet. - ZTNA broker: gatekeeps access to internal applications. Each layer enforces a different concern; together they implement zero-trust without requiring a single vendor's stack end-to-end. ## Production Hardening Checklist Minimum security configuration for any production SD-WAN deployment: - Mutual TLS authentication enabled (default on all platforms; verify it has not been weakened) - Strong cipher suites (AES-256-GCM minimum; no NULL or RC4 anywhere) - Short certificate rotation periods (90 days max; automated rotation) - VPN/VRF segmentation for guest, IoT, voice, and corporate traffic - Default-deny ACLs at WAN Edges; explicit allow for known flows - SASE or branch firewall in the path for all internet-bound traffic - DPI-based application identification (not just port-based) for policy - Tunnel monitoring and SLA-based steering to defeat path-quality attacks - Audit logging of all management plane changes; integrate with SIEM - Role-based access control (RBAC) on vManage / equivalent; integrate with corporate IdP The single most overlooked control: certificate lifecycle. Deployments commonly start with 1-year certificates, forget to rotate, and have a fabric-wide outage when certs expire. Automate rotation from day one. ## Summary SD-WAN provides baseline encryption and segmentation but creates an inspection gap by enabling direct internet break-out from every branch. SASE fills that gap with cloud-delivered SWG, CASB, FWaaS, ZTNA, DLP, and RBI services. By 2026 SD-WAN and SASE are effectively merging; any SD-WAN evaluation that does not consider SASE is incomplete. Pick the integration model based on your existing stack and operational preference: single-vendor SASE (Cato, Versa, Fortinet, Palo Alto Prisma) for operational simplicity, or best-of-breed (Cisco SD-WAN + Zscaler/Netskope) for individual product quality. Either way, the central inspection stack of legacy WAN is replaced by a distributed cloud-delivered model. Bookmark this article and the [SD-WAN cluster pillar](https://www.pinglabz.com/sd-wan/) for the operational picture, and the [SD-WAN deployment models article](https://www.pinglabz.com/sd-wan-deployment-models/) for how security choices interact with deployment topology. ### SD-WAN Deployment Models: Hybrid, Internet-Only, and Cloud On-Ramp URL: https://www.pinglabz.com/sd-wan-deployment-models/ Last updated: 2026-06-13T20:08:42.000Z SD-WAN deployment models matter because the choice affects cost, complexity, user experience, and operational risk for years. The vendor sales pitch typically defaults to a single option ("internet-only with our SD-WAN", "hybrid with two MPLS legs") that aligns with what they want to sell. The right deployment model for your enterprise depends on which applications you run, where they live, what compliance regulations apply, and how mature your operations team is. This article walks through the four main deployment models, the trade-offs each makes, the migration path from MPLS-only to a modern SD-WAN, and the patterns for HQ and data center sites. If you are designing a deployment or evaluating an existing one, this is the framework. ## The Four Branch Deployment Models Hybrid (MPLS + broadband) Transport mix 1 MPLS + 1-2 broadband + optional LTE Best for Established enterprises with hard-QoS apps; migration end-state Trade-off Higher cost than internet-only; better QoS for legacy real-time Internet-only (dual broadband) Transport mix 2 broadband + LTE for tertiary backup Best for SaaS-heavy organizations; new branches; greenfield Trade-off No carrier-managed QoS; relies on SD-WAN path conditioning Cloud on-ramp Transport mix Direct edge-to-cloud SD-WAN tunnels Best for Cloud-native applications; multi-cloud reach Trade-off Adds a new control point; cloud egress costs Multi-cloud direct Transport mix Multiple cloud on-ramps in parallel Best for Multi-cloud workloads; cross-cloud latency-sensitive flows Trade-off Highest complexity; per-cloud agreements ## Hybrid: MPLS + Broadband The hybrid model is the most common deployment in established enterprises in 2026\. The pattern: - One MPLS leg for hard-QoS-required traffic (voice, regulated workflows). Bandwidth right-sized down from MPLS-only days; often 50-100 Mbps where it used to be 200 Mbps. - One broadband leg from a primary ISP. Higher bandwidth than MPLS, lower cost. - One broadband leg from a secondary ISP for path diversity. Different physical infrastructure (different carrier, different last-mile). - Optional LTE for tertiary failover when both wired transports fail. SD-WAN policy steers per-application: - Voice: prefer MPLS, fall back to lowest-jitter broadband - Real-time business apps: prefer MPLS - SaaS (Office 365, Salesforce, Zoom): direct internet break-out via broadband - Internal apps: balanced load across all transports - Default: best-available with SLA-based steering This model is the typical end-state of a multi-year migration from MPLS-only. It preserves the operational guarantees of MPLS where they matter (which is fewer flows than people think) and uses broadband for everything else. Cost reduction vs MPLS-only is typically 25-40 percent over five years. ## Internet-Only: Dual Broadband + LTE Internet-only deployments skip MPLS entirely: - Two broadband legs from different ISPs (carrier diversity matters; redundancy from one ISP is not redundancy). - LTE for tertiary backup. Modern 5G branch routers make this practical. - SD-WAN's path conditioning (FEC, packet ordering, jitter buffer) handles voice and real-time over best-effort links. This model works well for: - SaaS-heavy organizations with little on-prem real-time - New branches where there is no MPLS contract to migrate from - Retail, healthcare clinics, small offices where the cost of MPLS would be disproportionate - International branches where MPLS is poorly supported The risk: if both broadband links fail simultaneously (rare but happens during major regional outages), the branch is offline until LTE picks up or the ISPs recover. Most organizations decide that risk is acceptable given the cost savings. One operational nuance: voice over best-effort internet works well 99 percent of the time and noticeably worse during the 1 percent of time when one transport is congested. Path conditioning helps but is not a complete substitute for carrier-managed QoS. If your business is sensitive to voice quality during peak hours, internet-only is not the right answer. ## Cloud On-Ramp The cloud on-ramp model adds direct connectivity from the SD-WAN edge to a cloud provider's network, bypassing both legacy WAN and the public internet for cloud-bound traffic. AWS On-ramp service AWS Cloud WAN, Direct Connect, Transit Gateway How it works SD-WAN edge peers with AWS via dedicated circuit or VPN Azure On-ramp serviceAzure Virtual WAN How it works SD-WAN edge connects to Azure VWAN hub via VPN or ExpressRoute Google Cloud On-ramp service Network Connectivity Center (NCC) How it works SD-WAN edge peers with NCC; optional Cloud Interconnect for high bandwidth Most SD-WAN vendors have native integrations with the major cloud providers - vManage can provision an AWS Cloud WAN edge in a few clicks. The benefit: traffic destined for cloud workloads takes a direct path, not a backhaul-to-HQ-then-out path. Latency drops; user experience improves. The cost: you pay the cloud provider for the on-ramp service (typically per-Mbps egress) on top of the SD-WAN platform license. For low-volume sites the math may not work; for high-volume sites it usually does. ## Multi-Cloud Direct Multi-cloud direct extends the on-ramp pattern: one SD-WAN edge, multiple cloud on-ramps in parallel. Policy decides which cloud serves which application. This is the highest-complexity model and only makes sense for organizations that: - Run workloads on multiple cloud providers (AWS for one product, Azure for another, GCP for a third) - Have applications with cross-cloud dependencies (latency-sensitive flows that need to traverse cloud-to-cloud) - Have the ops capacity to manage per-cloud on-ramp lifecycles Most enterprises overestimate their need for multi-cloud direct. If 80 percent of cloud traffic goes to one provider, single-cloud on-ramp plus cross-cloud routing through that provider is simpler. Multi-cloud direct is for organizations that genuinely run substantial workloads on three or more clouds. ## HQ and Data Center Patterns Branch deployment models are about cost-optimized redundancy. HQ and data center sites have different concerns: aggregation, security inspection, redundancy, and inter-DC traffic. Common HQ/DC patterns: Active-active redundant edge pair Default for any site with non-trivial traffic Active-standby pair Less traffic; simpler operations; failover under 30s Geographically separated pair (different DCs) Disaster recovery; protects against site-level failures Hub-and-spoke with central security inspection Compliance environments; legacy security architecture Full mesh Inter-branch heavy; modern collaborative organizations The hub-and-spoke pattern persists in compliance-driven environments where every flow must traverse a central inspection stack. Full mesh is more efficient for branch-to-branch traffic and is becoming dominant in modern deployments. Many enterprises run hybrid: full mesh between branches plus a hub-and-spoke leg back to HQ for centralized services. ## Migration: From MPLS-Only to Hybrid (and Beyond) The typical migration path from MPLS-everywhere to a modern hybrid SD-WAN: 1. **Audit and pilot (months 0-2).** Identify hard-QoS-required traffic. Pilot SD-WAN at 5-10 branches across the geographies and traffic types you have. Validate that path steering does what you expect. 2. **Tune policies (months 2-4).** Develop the per-application steering policy. Settle on transport SLAs (latency, jitter, loss thresholds for fall-over). Test failure modes. 3. **Roll out remaining branches (months 4-12).** Add SD-WAN edges and broadband at every branch. SD-WAN runs over both MPLS and broadband initially. Train operators on the centralized control plane. 4. **Right-size MPLS (months 12-18).** At each branch's MPLS contract renewal, downsize to the minimum needed for hard-QoS traffic only. Many branches drop from 200 Mbps MPLS to 50 Mbps. 5. **Phase out MPLS at suitable branches (months 18-30).** Branches with no hard-QoS requirements eventually go internet-only. Keep MPLS at HQ and large hub sites where it still has a role. The whole migration takes 18-30 months for a 100-branch enterprise, with most of the calendar time absorbed by waiting for MPLS contract renewals rather than the technical work. Doing it faster usually means paying early-termination penalties on MPLS contracts that often exceed the SD-WAN savings. ## Anti-Patterns to Avoid - **Deploying SD-WAN over a single transport.** If you only have MPLS or only broadband, SD-WAN buys you orchestration and policy but not the path-diversity story. The SD-WAN platform license might not pay for itself. - **Assuming broadband redundancy from the same ISP is real redundancy.** It is not. Same last-mile, same regional fiber routes, same data centers. Get carrier diversity. - **Skipping the LTE backup.** Modern LTE/5G branch routers cost a few hundred dollars per month and prevent the "both broadband links failed simultaneously" outage scenario. Worth it for any branch that affects revenue. - **Building hub-and-spoke when you should be full mesh.** Forcing branch-to-branch traffic through HQ adds latency and load to the HQ inspection stack. Modern SASE-style cloud security inspection lets you full-mesh branches and inspect traffic at the edge. - **Migrating policy 1:1 from MPLS to SD-WAN.** The legacy policy reflected legacy constraints. Re-think it with the new capabilities; don't just port the old QoS classes. ## Summary SD-WAN deployment models break into four patterns at the branch: hybrid (MPLS + broadband, the dominant pattern in 2026), internet-only (SaaS-heavy or greenfield), cloud on-ramp (cloud-native applications), and multi-cloud direct (multi-cloud workloads). HQ and data center sites use redundant edge pairs, with full mesh becoming the default branch-to-branch topology and hub-and-spoke persisting in compliance-driven environments. The migration from MPLS-only to a modern hybrid SD-WAN is an 18-30 month project for most enterprises, paced more by MPLS contract renewals than technical work. Pick the deployment model based on your applications and compliance, not on the vendor's preferred sales motion. Bookmark the [SD-WAN cluster pillar](https://www.pinglabz.com/sd-wan/) for the broader picture and the [SD-WAN vs MPLS comparison](https://www.pinglabz.com/sd-wan-vs-mpls/) for the cost analysis underlying these models. ### Cisco vManage / SD-WAN Manager: The Operator's Walkthrough URL: https://www.pinglabz.com/cisco-vmanage/ Last updated: 2026-06-13T20:08:42.000Z Cisco vManage (renamed Cisco Catalyst SD-WAN Manager around the 20.x release in 2023) is the operator-facing component of the Cisco SD-WAN fabric. It is where you configure templates, push policy, monitor the fabric, audit operator actions, push software upgrades, and manage certificates. If your operators are unhappy with vManage, they are unhappy with the SD-WAN deployment. This article walks through what vManage actually is, the constructs you work with daily (feature templates, device templates, policies, dashboards), the REST API surface for automation, the HA model, and the operational realities of running it at scale. If you are configuring your first vManage cluster, training new operators, or just trying to figure out where a particular setting lives, this is the reference. ## What vManage Actually Is vManage is a heavyweight VM-based application combining several roles: - **Web UI** for operators (HTTPS, role-based access, IdP integration) - **REST API** for automation - the SD-WAN API surface used by Ansible/Terraform/custom scripts - **Template engine** that generates per-WAN Edge configurations from operator-authored templates - **Policy engine** that compiles intent-style policies into OMP attribute manipulations and per-edge ACLs - **Telemetry collector** ingesting metrics and logs from every WAN Edge - **Monitoring database** (the bulk of vManage's resource footprint) - **Software image repository** for cEdge and vEdge firmware - **Certificate authority** for the fabric's mutual-TLS authentication - **Application Quality of Experience (AppQoE)** reporting backend The deployment model is a clustered set of identical VMs (3 nodes minimum, 6 for large fabrics) with a shared database backend. Each cluster member runs the same application stack; load balancers (or DNS) distribute operator UI sessions across them. ## Templates: The Core Operating Concept vManage's central abstraction is the template. Instead of configuring each WAN Edge directly, you author templates that describe what a class of edges should look like, then attach edges to a template. Two layers of templates: Feature template Describes One feature on a device (interface config, OMP, BGP, NAT, AAA, NTP, etc.) GranularityPer-feature Device template Describes The full configuration of a device, composed of feature templates GranularityPer-device-class The pattern: build a library of feature templates (one per concern: Interface-WAN, Interface-LAN, OMP-Default, BGP-Branch, etc.), then assemble them into device templates per branch type (Branch-Small, Branch-Large, Hub-Primary, Hub-Backup). Attach the WAN Edges to the appropriate device template. When you change a feature template, vManage recomputes the configurations for all attached devices and pushes the diffs. A single change to "Interface-WAN" can update 500 branches simultaneously. The discipline this requires is real. Templates with too many variables become unmanageable; templates with too few variables proliferate. Most production deployments end up with 20-50 feature templates and 5-15 device templates after 12 months of operation. ## Policy: Centralized vs Localized vManage distinguishes between two policy types: Centralized policy Lives where vSmart controllers (pushed via OMP) Use for Routing manipulation, traffic engineering, application-aware routing, service chaining Localized policy Lives where WAN Edge (pushed via template) Use for QoS, ACL filtering, route maps applied per-edge Centralized policy is the powerful one. It uses lists (data prefix lists, application lists, site lists, TLOC lists) and policy definitions (control policy, data policy, app-route policy) to express intent. The vSmart compiles these into the OMP route updates and service advertisements that change WAN Edge behavior. An example centralized policy: "Voice traffic from any site should prefer TLOCs with color=mpls; if those are unavailable, fall back to color=biz-internet." That single policy expression turns into per-flow path selection across thousands of WAN Edges automatically. Localized policy is more about QoS and per-port ACL hygiene. It is conceptually simpler and looks more like traditional Cisco IOS configuration. ## Dashboards and Monitoring vManage's monitoring surface includes: - **Network dashboard** \- top-level fabric health (sites up/down, alarm counts, control connections) - **Application Performance dashboard** \- per-application metrics across the fabric (Office 365, Salesforce, Zoom, custom apps) - **WAN Edge details** \- per-edge view with tunnel state, BFD sessions, OMP peers, interface stats - **Tunnel health** \- per-tunnel SLA metrics (latency, jitter, loss) over time - **Audit log** \- who changed what, when - **Real-time troubleshooting** \- the "Real Time" menu lets you query a remote WAN Edge for current OMP, BGP, BFD, ARP state without SSHing in The Real Time view is the most useful operator tool. Instead of opening an SSH session to a remote edge to run `show sdwan omp summary`, you click into the device in vManage and the GUI runs the command via a backchannel and renders the output. This works for hundreds of show commands. ## The REST API vManage's REST API (the "SD-WAN API") covers virtually everything the GUI does. The API is widely used for: - Bulk template attachment / detachment via Ansible or Python scripts - CI/CD pipelines that validate template changes before pushing them - External monitoring integrations (push fabric metrics into Prometheus, Grafana, Datadog) - Custom dashboards that combine SD-WAN data with other sources - ITSM integrations (auto-create tickets when a tunnel goes down) Authentication uses session cookies (login with credentials, get a session token, include in subsequent requests). For automation, generate API tokens with restricted scopes; for ad-hoc scripting, the session-cookie pattern is fine. The API is documented at `https:///apidocs` on every running instance. Version compatibility matters - the REST API surface evolves with each major release. ## High Availability vManage HA is implemented as a clustered deployment: 3 nodes Quorum2 of 3 Edge supportUp to 2,000 WAN Edges 6 nodes Quorum4 of 6 Edge supportUp to 6,000 WAN Edges The cluster shares a database. Loss of any single node is operationally invisible (other nodes serve UI/API requests). Loss of two nodes in a 3-node cluster (or three in a 6-node) takes vManage down until quorum is restored. Important: vManage downtime does not affect the data plane. WAN Edges keep forwarding traffic with their last-known policy. Operators just cannot make changes or view fresh telemetry until vManage recovers. Plan around this: vManage maintenance windows can be during business hours because traffic is not affected. ## On-Premises vs Cloud-Hosted Three deployment models: - **On-premises VMs.** You run vManage in your own VMware/KVM infrastructure. Full control; you manage upgrades and capacity. - **Cisco-hosted vManage (Cloud-hosted).** Cisco runs the vManage cluster as a service in their cloud. You manage the SD-WAN fabric; Cisco manages the vManage infrastructure. Common for smaller deployments and customers without strong on-prem ops. - **Cisco SD-WAN Cloud (formerly Viptela Cloud).** Even more managed; Cisco operates the entire control plane. The on-prem option gives the most control and is what large enterprises typically choose. The hosted options trade control for operational simplicity. ## Software Upgrades vManage manages firmware upgrades for the entire fabric. The flow: 1. Upload a new cEdge or vEdge image to vManage's image repository. 2. Select target devices (single edge, group of edges, or whole sites). 3. Schedule the upgrade. vManage pushes the image to the targets, then activates it. 4. cEdge images use ISSU (In-Service Software Upgrade) on supported platforms - the upgrade happens with minimal forwarding interruption. For vManage and vSmart and vBond upgrades, the process is more careful: upgrade vManage first (cluster rolling upgrade), then vSmart (rolling), then vBond, then cEdges. Mixed-version operation is supported but should be a transient state, not a long-term one. ## Performance Realities at Scale vManage is the bottleneck in large fabrics. Common pain points after 12-18 months of growth: - **Database bloat.** Telemetry retention defaults are conservative; turning them up at scale balloons disk requirements. - **Template push slowness.** Pushing a feature template change to 1,000 edges is not instantaneous; it can take minutes. - **UI rendering delays.** Dashboards that aggregate across thousands of edges get slow if the database is not tuned. - **API rate limits.** Heavy automation can hit internal rate limits; pace requests. The standard mitigations: more memory (vManage benefits from RAM more than CPU), tune the database, scale the cluster up to 6 nodes if you are above 2,000 edges, and use the API instead of the GUI for bulk operations. ## Summary vManage (Cisco SD-WAN Manager) is the management plane of Cisco Catalyst SD-WAN. It is the operator UI, the REST API, the template engine, the policy compiler, the monitoring backend, the certificate authority, and the upgrade orchestrator - all in one heavyweight clustered application. Run it in 3-node or 6-node clusters depending on fabric size. Your day-2 operations experience is dominated by how you use vManage, not by the underlying fabric protocol. Master templates, policy expressions, and the REST API early. Treat the GUI as the entry point but graduate to API-driven automation for any change touching more than a handful of edges. Bookmark the [SD-WAN cluster pillar](https://www.pinglabz.com/sd-wan/) and the [Cisco Catalyst SD-WAN architecture article](https://www.pinglabz.com/cisco-sd-wan-architecture/) for the broader fabric picture, and lab every template change before pushing to production. ### Cisco Catalyst SD-WAN Architecture: vManage, vSmart, vBond, WAN Edge URL: https://www.pinglabz.com/cisco-sd-wan-architecture/ Last updated: 2026-06-13T20:08:43.000Z Cisco Catalyst SD-WAN (formerly Cisco SD-WAN, originally Viptela) is the dominant enterprise SD-WAN platform in 2026\. The architecture that Cisco bought from Viptela in 2017 and rebranded multiple times since is genuinely good: a clean four-component decomposition, a BGP-derived control protocol that network engineers already understand, and a deployment model that scales from single-branch pilots to 6,000-edge global deployments. This article walks through the four components (vManage, vSmart, vBond, WAN Edge), the OMP control protocol, the cEdge vs vEdge distinction, and how the components interact during a typical fabric bring-up. If you are studying for ENSDWI, designing a Cisco SD-WAN deployment, or just trying to read a vSmart log file, this is the reference. ## The Four Components vManage (SD-WAN Manager) Form factor VM (3 or 6 nodes for HA) PlaneManagement Job UI, REST API, template engine, monitoring backend, certificate authority vSmart controllers Form factorVM (2-3 per region) PlaneControl Job OMP routing protocol; distributes routes and policy to WAN Edges vBond orchestrator Form factor VM (2 instances, internet-facing) PlaneOrchestration Job Authenticates new WAN Edges; bootstraps cert distribution WAN Edge (cEdge or vEdge) Form factor Hardware appliance or VM at every site PlaneData Job Forms IPsec tunnels; enforces policy ## vManage (SD-WAN Manager) vManage is the single GUI and REST API operators use to run the entire fabric. It is the heaviest component: - Web UI (HTTPS, role-based access) - REST API (full automation surface, called the SD-WAN API) - Template engine that generates per-WAN Edge configurations from feature templates and device templates - Telemetry collector and monitoring database - Software image repository and upgrade orchestrator - Built-in certificate authority for fabric certificates HA is implemented as a clustered deployment: 3 nodes minimum (with a quorum requirement) or 6 nodes for high-scale deployments. The cluster shares a database backend; losing a single node has no operator-visible impact, losing two breaks operator access until quorum is restored. Cisco renamed vManage to "SD-WAN Manager" in the 20.x release branch around 2023; older documentation still uses vManage. The two terms refer to the same product. Practical implication: vManage is the day-2 ops experience. If your operators struggle with vManage, your SD-WAN deployment will feel hard to run regardless of how good the underlying fabric is. ## vSmart Controllers vSmarts are the control-plane brains. They run OMP (Overlay Management Protocol) sessions to every WAN Edge, distribute routes and policy, and serve as the central decision point for control-plane state. Key facts: - Lightweight VMs (no big database, no UI) - Typically 2 or 3 per region for HA (Cisco recommends 3 for production) - WAN Edges form OMP sessions to all vSmarts simultaneously and receive identical state - vSmarts do not need to share state with each other; each independently produces the same OMP output - vSmart failure is graceful: WAN Edges keep forwarding with cached OMP state, just cannot receive new policy updates The number of vSmarts you need depends on the WAN Edge count, not on the management complexity. Each vSmart can handle thousands of OMP sessions. For a 1,000-edge fabric, 2 vSmarts active-active is plenty. ## vBond Orchestrator vBond is the bootstrap component. When a brand-new WAN Edge boots up at a branch, it needs to find the rest of the fabric. The sequence: 1. WAN Edge boots, has DHCP-acquired IP and DNS. 2. WAN Edge resolves a configured DNS name (e.g. `vbond.example.com`) to vBond's public IP. 3. WAN Edge contacts vBond, presents its certificate (root-of-trust pre-installed at manufacturing or via PnP Connect). 4. vBond authenticates the certificate, looks up the WAN Edge's organization, and replies with the IPs of vManage and vSmart for that organization. 5. WAN Edge contacts vManage to download configuration, vSmarts to establish OMP, and is now part of the fabric. vBond is internet-facing because new edges have not yet joined the private overlay - they have to reach vBond over the public internet on first boot. That is why vBond is architecturally separate: it has different security requirements (public IP, hardened, exposed) than the rest of the platform (which can live entirely behind firewalls). You typically deploy 2 vBond instances behind a load balancer or DNS round-robin. They are stateless beyond their certificate database, so HA is straightforward. ## WAN Edges: cEdge and vEdge WAN Edges are the data-plane devices at every branch. They form the IPsec overlay tunnels, do DPI to identify applications, apply per-application policy, and enforce the SD-WAN's behavior on customer packets. Two WAN Edge variants exist for historical reasons: Viptela OS TypevEdge Hardware Original Viptela hardware (vEdge 100, 1000, 2000, 5000) Status in 2026 Legacy; end-of-sale; supported but not recommended for new deployments Cisco IOS-XE (SD-WAN image) TypecEdge Hardware Catalyst 8000 series, ISR 4000, ASR 1000, ISR 1000, virtual cEdge Status in 2026 Default for all new deployments; full feature parity and beyond cEdge runs the same IOS-XE you might recognize from Catalyst 9000 switches and Catalyst 8000 routers, with an SD-WAN-specific image and control plane. It supports the full Cisco IOS feature set (advanced routing, QoS, security features, advanced application-aware steering) on top of the SD-WAN overlay. Cisco has been migrating customers off vEdge since 2020\. By 2026 the recommendation is cEdge for everything. New ENSDWI training and certifications focus on cEdge. ## OMP: The Overlay Management Protocol OMP is the BGP-style routing protocol vSmarts use to distribute fabric state to WAN Edges. If you understand BGP, OMP makes immediate sense - it is BGP with SD-WAN-specific extensions. What OMP carries: OMP Routes Carries Customer routes (per-VPN), with the originating WAN Edge's TLOC list BGP equivalent Like BGP routes with extended community attributes TLOC Routes Carries "Transport Locators" - each edge's transport identity (color + system IP + encapsulation) BGP equivalent No direct BGP equivalent; SD-WAN-specific Service Routes Carries Where to find services (firewall, NAT) that traffic should traverse BGP equivalent BGP communities for service insertion The TLOC concept is what makes SD-WAN's transport-agnostic overlay work. Each WAN Edge has multiple TLOCs (one per transport: MPLS, Internet, LTE, etc.) identified by a "color" attribute. When the policy says "voice goes over MPLS", it really says "voice prefers TLOCs with color=mpls". Path selection is then a TLOC selection problem. Policy in OMP works like BGP route policy: route maps, prefix lists, communities, set operations. If you have configured BGP route maps you have configured OMP policy. See the [BGP cluster pillar](https://www.pinglabz.com/bgp/) for the foundational BGP concepts. ## A Typical Fabric Bring-Up The end-to-end sequence for spinning up a new SD-WAN fabric from scratch: 1. **Deploy vManage cluster.** 3 VMs, configure clustering. This is the first component because it is the certificate authority for everything else. 2. **Deploy vBond instances.** 2 VMs with public IPs. Register them with vManage. Configure DNS so the fabric's bootstrap name resolves to vBond. 3. **Deploy vSmart controllers.** 2 or 3 VMs. Register them with vManage. They will establish OMP sessions to WAN Edges as those come online. 4. **Configure organization, certificates, root cert chain.** All in vManage. 5. **Author feature templates and device templates in vManage.** One device template per branch type; feature templates are the building blocks (interfaces, OMP, BGP, NAT, etc.). 6. **Onboard first WAN Edge.** Plug in, power up, ZTP (Zero Touch Provisioning) flow runs: DHCP for IP, DNS to vBond, vBond authenticates and points to vManage, vManage pushes template, edge establishes OMP to vSmarts, fabric joins. 7. **Roll out remaining edges.** Each one repeats step 6\. Templates handle the per-site differences automatically. The whole bring-up for a small fabric (3 vManages + 2 vBonds + 2 vSmarts + first 5 edges) is a 1-2 day effort. Subsequent edges onboard in minutes. ## Scale Considerations Cisco-published scale numbers for the platform in 2026: vManage cluster (3-node) Up to 2,000 WAN Edges vManage cluster (6-node) Up to 6,000 WAN Edges vSmart controller Thousands of OMP sessions; rarely the bottleneck vBond instance Hundreds of bootstrap requests per minute; rarely the bottleneck cEdge (Catalyst 8500) Up to 20+ Gbps SD-WAN throughput For very large deployments (3,000+ edges) the bottleneck is usually vManage's database; vSmart and vBond scale linearly without much fuss. ## Day 2 and Beyond The architecture handles the bring-up well. The harder part is operating it. Day-2 operations on Cisco Catalyst SD-WAN center on: - Template lifecycle (one-off changes drift away from templates over time) - Policy authoring (centralized policy is powerful but easy to author incorrectly) - vManage performance (the cluster needs care as edge count grows) - Software upgrades (rolling cEdges through new firmware without breaking tunnels) - Observability (correlating "user says Salesforce slow" with vManage metrics) For the operational depth argument, see [Cisco SD-WAN in 2026: The Real Value Is Day-2 Operations](https://www.pinglabz.com/cisco-sd-wan-in-2026-the-real-value-is-day-2-operations/). ## Summary Cisco Catalyst SD-WAN (formerly Viptela) decomposes into four components: vManage (management), vSmart (control), vBond (orchestration), and WAN Edge (data). The control protocol is OMP, a BGP-derived path-vector that distributes routes, TLOCs, and service advertisements. cEdge is the modern IOS-XE WAN Edge; vEdge is the legacy Viptela-OS variant that you should not deploy new. If you understand BGP and the three-plane SD-WAN architectural model, the Cisco platform is straightforward. The difficulty is operational - templates, policy authoring, vManage performance at scale - rather than architectural. Bookmark the [SD-WAN cluster pillar](https://www.pinglabz.com/sd-wan/) for the cross-vendor view, and the [SD-WAN architecture article](https://www.pinglabz.com/sd-wan-architecture/) for the three-plane model that this Cisco-specific one fits into. ### SD-WAN Architecture: Control, Data, and Management Planes Explained URL: https://www.pinglabz.com/sd-wan-architecture/ Last updated: 2026-06-13T20:08:43.000Z Every SD-WAN platform shares the same three-plane architecture: management, control, and data. Once you understand the planes, the differences between Cisco Catalyst SD-WAN, Fortinet Secure SD-WAN, VMware VeloCloud, Versa, and Palo Alto Prisma all become implementation details rather than architectural divergences. Vendor-specific component names are noise; the model is the same. This article walks through the three planes, what each does, how they communicate, and the design constraints that drive how vendors actually deploy them. If you are evaluating SD-WAN, designing a deployment, or just trying to make sense of a vendor diagram, this is the reference. ## The Three Planes Network engineers learn this model from MPLS-VPN and SDN: separate the work the network does into three concerns: Management Job Configure, monitor, troubleshoot. Operator-facing. Frequency of operation Manual / scheduled (minutes to days) Control Job Compute and distribute routing/policy state to forwarding nodes Frequency of operation Continuous (sub-second to seconds) Data Job Forward customer packets according to control-plane state Frequency of operation Per-packet (microseconds) SD-WAN borrows this directly. The innovation is centralizing the management and control planes (in legacy WAN they were distributed across every router) and leaving only the data plane at the branch. Branch edge devices forward packets according to centralized policy without needing to consult the controller per-packet. ## The Management Plane The management plane is the operator-facing layer: the GUI, REST APIs, monitoring dashboards, template authoring tools, certificate management, software upgrade orchestration, and the central database that holds configuration state. Vendor implementations: Cisco Catalyst SD-WAN vManage (now SD-WAN Manager) VMware VeloCloud VeloCloud Orchestrator Fortinet FortiManager + FortiAnalyzer Versa Versa Director Palo Alto Prisma SD-WAN Strata Cloud Manager (formerly Panorama) HPE Aruba EdgeConnect Orchestrator Common characteristics: heavy footprint (database, web stack, REST API server), one or two instances per region (with HA clustering), accessed by operators not by edge devices. Operators write templates here that define what each branch type looks like; they push templates to the control plane which distributes the resulting configurations to data-plane edges. Day-2 operations live in the management plane. If your management UI cannot show you why traffic is traversing a particular path, who changed a policy yesterday, or which branches have non-template-managed ad-hoc configurations, you are in for a bad operational time. This is where vendor differences matter most in practice. ## The Control Plane The control plane computes routing and policy state and distributes it to edge devices. In legacy WAN this work happened on every router (running OSPF or BGP locally). In SD-WAN it is centralized: a small number of controllers run the routing protocol, decide which routes to advertise to which edges, and push policy. Vendor implementations: Cisco Catalyst SD-WAN Control plane componentvSmart controllers Routing protocol OMP (Overlay Management Protocol, BGP-derived) VMware VeloCloud Control plane component VeloCloud Gateway / Hub Routing protocol iBGP, custom path selection Fortinet Control plane component FortiGate fabric members + iBGP Routing protocol iBGP for routing; SD-WAN rules for steering Versa Control plane componentVersa Controller Routing protocol iBGP-based with extensions Palo Alto Prisma SD-WAN Control plane component Cloud control plane (CloudGenix legacy) Routing protocol App-defined policies; cloud-native control HPE Aruba EdgeConnect Control plane component Orchestrator + SDX peering Routing protocol BGP and Aruba's tunnel orchestration OMP (in Cisco Catalyst SD-WAN) deserves a callout: it is a BGP-style path-vector protocol designed specifically for the SD-WAN use case. It carries routes, services (firewall, NAT, etc.), TLOC (Transport Locator) advertisements that describe each edge's transports, and policy. If you understand BGP path attributes and route maps, OMP will be familiar - the mental model is the same. See the [BGP cluster pillar](https://www.pinglabz.com/bgp/) for the foundational BGP concepts. Control plane components are typically lightweight (no big database, no heavy UI), run as 2 or 3 instances per region for HA, and need fast paths to all the edge devices they manage. They can fail without immediate data-plane impact (edges keep forwarding using last-known policy), but new policy changes do not propagate until controllers come back. ## The Data Plane The data plane is where packets actually move. Edge devices at every site form encrypted IPsec tunnels to other edges (and to cloud gateways), apply policy received from the control plane, and forward packets accordingly. Vendor implementations: Cisco Catalyst SD-WAN Data plane component WAN Edge (cEdge IOS-XE based; vEdge Viptela-OS legacy) Form factors Catalyst 8000 series, ISR 4000, virtual VMware VeloCloud Data plane componentVeloCloud Edge Form factors Hardware appliances + virtual Fortinet Data plane componentFortiGate Form factors Existing FortiGate firewall hardware Versa Data plane component Versa CSG (Carrier-Grade Software-defined Gateway) Form factorsHardware + virtual Palo Alto Prisma SD-WAN Data plane componentION devices Form factorsDedicated hardware HPE Aruba EdgeConnect Data plane componentEdgeConnect appliances Form factorsHardware + virtual Edges are at every branch and number in the thousands for large enterprises. They need to: - Form IPsec tunnels over each available transport - Deep-packet-inspect to identify applications - Apply per-application steering policy - Maintain telemetry back to the management plane - Survive control-plane outages by continuing to forward using cached policy The data plane is the most performance-sensitive component. Hardware acceleration, ASIC-level encryption, and proper sizing matter. An undersized branch edge becomes the bottleneck even if everything else is right. ## Orchestration: The Fourth Component (Sometimes) Some vendors break out a fourth component for orchestration - the bootstrap and authentication layer that brings new edges into the fabric: Cisco Catalyst SD-WAN vBond orchestrator VMware VeloCloud VeloCloud Orchestrator (handles both management and orchestration) Fortinet FortiManager (handles both) The job: when a new branch edge boots up for the first time, it needs to find the rest of the fabric. The orchestrator is what it talks to first. The orchestrator validates the edge's certificate, tells it which controllers and management plane to connect to, and steps out of the path. After bootstrap, the edge talks directly to control and management; the orchestrator is only involved at onboarding. Cisco breaks this out as a separate vBond component because the architectural concern (a publicly-reachable bootstrap point) has different security requirements than the rest of the platform. vBond is typically internet-facing; vManage and vSmart are not. Other vendors collapse this into the orchestrator. ## How the Planes Communicate The protocols connecting the planes vary by vendor, but the conceptual flow is consistent: 1. **Operator authors policy in the management plane.** "Voice goes over MPLS or lowest-jitter broadband; SaaS goes direct to the internet." 2. **Management plane pushes policy to control plane.** Cisco: vManage configures vSmart with the policy. Fortinet: FortiManager pushes to FortiGates. 3. **Control plane translates policy into per-edge state.** The vSmart converts site-agnostic policy into per-site routes and TLOC advertisements that the relevant edges need. 4. **Control plane distributes state to edges via the routing protocol.** OMP in Cisco; iBGP in others. 5. **Edges enforce policy on packets.** Per-flow policy lookup, per-application classification, per-tunnel SLA tracking. 6. **Edges report telemetry back to management.** Path metrics, application performance, policy hits, certificate status. The cycle runs continuously. Policy changes propagate from management to data plane in seconds. Path metrics flow from data plane back to management for the operator dashboards. ## High Availability Across the Planes Management Failure impact Operators cannot make new changes; existing operations continue Typical HA pattern 3-node cluster (Cisco vManage); active-passive pair (FortiManager) Control Failure impact New policy does not propagate; edges keep forwarding with last-known state Typical HA pattern 2-3 controllers per region, no shared state required Data Failure impact That branch goes offline (depending on transport) Typical HA pattern Active-active or active-standby pair at HQ; single edge with multiple transports at branches Orchestration Failure impact New edges cannot bootstrap; existing edges unaffected Typical HA pattern Multi-instance behind anycast or DNS round-robin The architectural property that matters most in production: data-plane edges keep forwarding when the control and management planes are down. They cache their last-known policy and continue applying it. This is the deliberate decoupling that makes SD-WAN reliable in practice. A management plane outage is a 4-hour incident; a data plane outage at a branch is the branch going offline. ## Design Implications Three architectural decisions have outsized impact on operations: 1. **Where do you place control plane components?** Centralized at HQ is operationally simpler. Distributed across regions is more resilient and reduces control-plane latency to edges. Most large enterprises do regional control planes with global federation. 2. **How do edges find controllers?** Static configuration (each edge knows controller IPs) is brittle but simple. DNS-based discovery is flexible but introduces a DNS dependency. Cisco uses vBond as a DNS-resolvable bootstrap point. 3. **How is the management plane secured?** The management plane database holds your entire WAN policy. Tight access controls, audit logging, role-based access, and integration with your IdP are not optional in production. ## Summary Every SD-WAN platform decomposes into three planes: management (operator-facing), control (policy and routing distribution), and data (packet forwarding at the edges). Some vendors add a fourth orchestration component for edge onboarding. Vendor-specific component names are implementation details; the architecture is consistent across the market. If you understand the three-plane model and what each plane needs from the others, you can read any vendor's architecture diagram in five minutes and ask the right questions about HA, scale, and operational impact. Bookmark the [SD-WAN cluster pillar](https://www.pinglabz.com/sd-wan/) for the broader operational picture, and treat this article as the architectural reference you come back to when evaluating new vendors or designing new deployments. ### SD-WAN vs MPLS: When Each Wins in 2026 URL: https://www.pinglabz.com/sd-wan-vs-mpls/ Last updated: 2026-07-04T23:01:33.000Z SD-WAN vs MPLS is the question every WAN renewal conversation starts with in 2026\. The framing is misleading. They are not pure substitutes; they coexist in most production networks, and the right answer for your enterprise depends on which applications you run, what your branches look like, and where your security stack lives. If you want the bigger picture first, the [MPLS complete guide](https://www.pinglabz.com/mpls/) covers the architecture this article plugs into. This article walks through the real technical differences, the cost story (with and without the marketing spin), the cases where MPLS still wins, the cases where SD-WAN clearly wins, and the hybrid pattern that most enterprises actually settle on. If you are evaluating a WAN renewal or building the business case for a migration, this is the comparison. ## What Each Actually Is **MPLS** (Multiprotocol Label Switching) is a carrier-managed Layer 2.5 transport. The carrier provisions a private circuit between your sites, applies QoS classes you negotiate in the contract, and guarantees latency and packet loss within Service Level Agreements. The circuit is dedicated; nobody else's traffic shares your bandwidth. You do not see the carrier's underlying infrastructure - they hand you Ethernet ports on each end. **SD-WAN** (Software-Defined Wide Area Network) is an overlay that runs on top of any IP transport - broadband internet, MPLS, LTE, satellite. It uses encrypted IPsec tunnels between branch edge devices and a central controller for policy. The transport itself is irrelevant to SD-WAN; what matters is the overlay's ability to steer applications across multiple transports based on policy. The key conceptual difference: MPLS is a transport service, SD-WAN is a routing and policy layer. You can run SD-WAN over MPLS (and many enterprises do). You cannot run MPLS over SD-WAN. ## The Side-by-Side Comparison Carrier model SD-WAN Bring your own transport MPLS Carrier-managed end-to-end Cost per Mbps SD-WANLow (broadband) MPLS High (dedicated circuit) Typical bandwidth SD-WAN 500 Mbps to 1 Gbps per branch MPLS 50 to 200 Mbps per branch QoS guarantee SD-WAN Best-effort over public internet (mitigated by FEC, jitter buffers, multi-link) MPLS Per-class hard QoS honored by carrier Latency consistency SD-WANVariable MPLSPredictable SLA SD-WAN Per-link from each ISP independently MPLS End-to-end from one carrier Provisioning time SD-WAN Days (broadband install) MPLSWeeks to months Site onboarding SD-WAN Zero-touch (controller pushes config) MPLS Carrier engineer configures CE/PE Cloud / SaaS reach SD-WAN Direct internet break-out from branch MPLS Backhaul to HQ then out (or expensive cloud-direct) Encryption SD-WANIPsec by default MPLS Optional, customer-deployed Visibility SD-WAN SD-WAN controller has full per-flow telemetry MPLS Carrier MIB; SNMP from CE Vendor lock SD-WAN One SD-WAN vendor; multiple transport providers MPLSOne carrier per region Resilience SD-WAN Active-active across multiple transports MPLS Backup link required separately ## The Cost Story (With and Without Spin) Vendor pitches commonly cite "30 to 50 percent WAN cost reduction" by replacing MPLS with broadband. The number is not wrong, but it understates what the comparison actually involves. The honest calculation: Carrier circuits MPLS-only $1,500/mo per 100 Mbps branch SD-WAN over broadband $300/mo per 1 Gbps broadband + $300/mo per redundant broadband SD-WAN platform license MPLS-only Included (carrier-managed CE) SD-WAN over broadband$50-200/mo per branch Edge appliances MPLS-onlyCarrier-provided CE SD-WAN over broadband $1,500-5,000 per branch (capex amortized) Operational overhead MPLS-only Lower (carrier owns CE) SD-WAN over broadband Higher early; lower later (centralized ops) Cloud / SaaS performance MPLS-only Backhaul-tax everywhere SD-WAN over broadbandDirect break-out For a 100-branch enterprise, the math typically works out to 25-40 percent net savings over five years if the SD-WAN deployment is operationally well-run, and 10-20 percent if it is not. The bigger value in 2026 is operational - faster branch provisioning, better cloud user experience, unified visibility - which is harder to put in a spreadsheet but real. Where the savings are smaller than promised: enterprises with already-renegotiated MPLS contracts, environments with a small number of large branches (where MPLS scales well), and deployments where the SD-WAN platform license cost approaches the MPLS savings. ## When MPLS Still Wins Real cases where MPLS remains the right answer in 2026: - **Hard real-time traffic with carrier-honored QoS.** Voice with strict jitter budgets (under 30ms), trading floor low-latency feeds, real-time control systems for industrial / SCADA. SD-WAN's path conditioning is good but not equivalent to a carrier-managed dedicated path. - **Compliance environments where the path itself is regulated.** Some payment processing, healthcare, and government segments require specific carrier paths or geographic routing. SD-WAN over commodity internet does not satisfy those requirements. - **Specific point-to-point requirements.** Two data centers with deterministic 10 Gbps requirements running synchronous replication is an MPLS or dark fiber use case, not an SD-WAN use case. - **International circuits where broadband is poor.** Some remote international locations have unreliable broadband and good MPLS coverage. The SD-WAN transport story does not improve transports it does not have access to. - **Carrier-managed simplicity.** Smaller enterprises that do not want to operate a control plane themselves often prefer paying a carrier to manage everything. SD-WAN done badly is worse than legacy WAN done well. ## When SD-WAN Clearly Wins - **SaaS-heavy organizations.** If most user traffic is bound for Office 365, Salesforce, Zoom, AWS, etc., direct internet break-out from each branch is dramatically faster than backhauling to HQ. - **Dynamic branch counts.** Retail chains, healthcare networks, franchises - any organization that opens or closes sites frequently. Zero-touch provisioning is real and saves weeks per site. - **Multi-cloud deployments.** Cloud on-ramps to AWS, Azure, GCP from each branch is operationally simpler with SD-WAN than with MPLS-plus-cloud-circuits. - **Heterogeneous transport availability.** If your branches have a mix of MPLS, broadband, LTE, and satellite, SD-WAN's transport-agnostic overlay is the only sane way to manage that uniformly. - **Operational consolidation.** Centralized policy, templates, and observability are genuinely better than per-router CLI management. Once an organization gets used to it, going back is unthinkable. ## The Hybrid Pattern Most Enterprises Actually Use The dominant 2026 deployment model is not pure SD-WAN and not pure MPLS. It is hybrid: - **One MPLS leg per branch.** Hard-QoS traffic (voice, certain real-time apps) and any compliance-required flows. Bandwidth scaled down from MPLS-everywhere; might be a 50 or 100 Mbps circuit instead of the original 200 Mbps. - **One or two broadband legs per branch.** Different ISPs for redundancy. Carries everything else: SaaS, internet, internal apps that do not need hard QoS. - **SD-WAN overlay across both transports.** Per-application policy decides which transport each flow uses. - **LTE failover for tertiary backup.** When both primary transports fail, LTE keeps critical traffic moving. The cost picture is favorable: dropping from MPLS-everywhere to one MPLS plus broadband-redundancy reduces WAN spend significantly while keeping the QoS guarantees where they matter. The user experience is better than pure-MPLS for cloud apps and at least as good for everything else. And if MPLS pricing eventually catches up (or the QoS-required apps go away), the same SD-WAN can run internet-only without re-architecting. ## Migrating from MPLS-Only to Hybrid The typical migration path: 1. **Audit current traffic.** Identify what needs hard QoS (voice, real-time), what is cloud-bound, what is internal. Determine current MPLS utilization per branch. 2. **Pilot at 5-10 branches.** Deploy SD-WAN edges, add a broadband link. Run SD-WAN over both MPLS and broadband. Validate that path steering does what you expect. 3. **Tune policies.** Steer voice and real-time over MPLS, SaaS over broadband, internal traffic via least-loaded path. Set fall-over thresholds. 4. **Roll out to remaining branches.** Hire or train operators on the centralized control plane. 5. **Right-size MPLS.** After 3-6 months of operation, downsize MPLS contracts at renewal. Most enterprises move from 200 Mbps MPLS to 50 Mbps, keeping it as a hard-QoS leg only. 6. **Optional: phase out MPLS at suitable branches.** Branches with no hard-QoS requirements eventually go internet-only. Keep MPLS at HQ and large hub sites. Total elapsed time for a 100-branch migration is typically 12-24 months including the pilot, with most of the time absorbed by waiting for MPLS contract renewals to come due rather than the technical work itself. ## Things People Get Wrong - **Treating SD-WAN as MPLS replacement.** It is not. SD-WAN is a routing and policy layer; MPLS is a transport. The right framing is "SD-WAN over multiple transports, including possibly MPLS." - **Underestimating the operational uplift.** SD-WAN replaces per-router CLI work with template management, policy authoring, and observability dashboards. Different skills. Many MPLS-trained operators struggle for the first year. - **Buying the SD-WAN product without the SASE story.** SD-WAN by itself does not solve the security gap created by direct internet break-out. Plan SASE integration alongside SD-WAN, not as an afterthought. - **Over-engineering policy on day one.** Start with three or four broad application classes (voice, business-critical, internet, default) and tune from there. Day-one policy with 50 categories is unmanageable. ## Summary SD-WAN and MPLS are not pure substitutes. SD-WAN is an overlay that runs on any IP transport including MPLS. MPLS is a carrier-managed transport that still has a real role for hard-QoS-required traffic and compliance-regulated paths. The right answer for most enterprises in 2026 is hybrid: SD-WAN as the routing and policy layer, with one MPLS leg plus broadband redundancy as transports. Cost savings are real but smaller than vendor pitches suggest; operational and user-experience improvements are larger and more durable. Bookmark this article and the [SD-WAN cluster pillar](https://www.pinglabz.com/sd-wan/) as the reference, and lab any migration in a controlled environment before committing to dates. ### MP-BGP: Multiprotocol BGP Address Families Explained URL: https://www.pinglabz.com/mp-bgp-multiprotocol-bgp/ Last updated: 2026-07-04T23:21:12.000Z MP-BGP (Multiprotocol BGP) is the extension of BGP-4 that lets a single BGP session carry routing information for more than just IPv4 unicast. It is what makes BGP usable as a control plane for IPv6, MPLS L3VPNs (VPNv4), VPLS, EVPN, and pretty much every modern data center fabric. If you have ever heard "we run BGP for the data center underlay and EVPN for the overlay," MP-BGP is what made that sentence work. (This article is part of the PingLabz MPLS series - the [full MPLS guide](https://www.pinglabz.com/mpls/) maps the whole cluster in reading order.) This article explains what MP-BGP actually is, how it extends BGP without breaking backwards compatibility, the address-family concept that ties it all together, and the most common production uses (IPv6, VPNv4, EVPN). If you are studying for CCIE, building modern fabrics, or just trying to understand what `address-family ipv4 vrf RED` means, this is the conceptual foundation. ## The Problem MP-BGP Solves Original BGP-4 (RFC 1771, 1995) only carried IPv4 unicast prefixes. The NLRI field in the UPDATE message was implicitly IPv4: a length and a prefix encoded as a 4-byte address. There was no room in the format for any other address family. The problem became urgent in the late 1990s with two converging needs: - **IPv6 deployment.** Operators wanted BGP to carry IPv6 prefixes between ASes, not just IPv4. - **MPLS L3VPN (RFC 4364).** Service providers wanted to multiplex many customer VPNs over a shared MPLS backbone, with each customer's IPv4 routes carried as a separate "VPNv4" address family. The fix was MP-BGP (RFC 4760, "Multiprotocol Extensions for BGP-4"). It added two new optional attributes to the UPDATE message: MP\_REACH\_NLRI and MP\_UNREACH\_NLRI. These attributes carry an "Address Family Identifier" (AFI) and "Subsequent Address Family Identifier" (SAFI) along with the actual prefix data, letting the same BGP session carry many different kinds of routes. ## Address Families: AFI and SAFI An "address family" in BGP is a (AFI, SAFI) pair that identifies what kind of route a particular UPDATE is carrying. 1 IPv4 2 IPv6 25 Layer 2 VPN (L2VPN) 1 Unicast 2 Multicast 4 MPLS labels 70 EVPN 128 MPLS-labeled VPN (VPNv4 / VPNv6) 129 Multicast VPN Common combinations you will see: IPv4 unicast AFI / SAFI1/1 Use Default; everyday internet routing IPv6 unicast AFI / SAFI2/1 UseIPv6 internet routing VPNv4 AFI / SAFI1/128 Use MPLS L3VPN customer routes (with route distinguisher) VPNv6 AFI / SAFI2/128 Use MPLS L3VPN customer routes for IPv6 L2VPN EVPN AFI / SAFI25/70 Use Data center fabric overlay (EVPN) IPv4 + Labels AFI / SAFI1/4 Use BGP-LU (BGP Labeled Unicast) for inter-AS MPLS Each address family is configured separately under `router bgp`, with its own neighbor activations, policies, and route-maps. The same BGP TCP session carries them all, multiplexed via the AFI/SAFI in each UPDATE. ## Capability Negotiation Two MP-BGP-capable peers exchange their supported (AFI, SAFI) pairs in the BGP OPEN message via the Multiprotocol Capability (capability code 1, RFC 4760). Each side advertises which families it wants to exchange; the union of what they both support is what they negotiate for that session. If one peer is older and does not understand MP-BGP, the OPEN does not include the capability and the session falls back to plain BGP-4 (IPv4 unicast only). This is how MP-BGP maintains backwards compatibility: an MP-BGP router can peer with a legacy BGP-4 router, they exchange only IPv4 unicast, and nobody breaks. ## Cisco IOS XE Configuration The classic IPv4 unicast configuration is implicit: ``` router bgp 65001 neighbor 10.0.12.2 remote-as 65002 address-family ipv4 unicast neighbor 10.0.12.2 activate network 192.168.1.0 mask 255.255.255.0 exit-address-family ``` The pattern for adding additional address families is identical: configure under `router bgp`, drop into `address-family X Y`, and `activate` the neighbor inside that AF. For IPv6 over the same session: ``` router bgp 65001 neighbor 2001:db8:12::2 remote-as 65002 address-family ipv6 unicast neighbor 2001:db8:12::2 activate network 2001:db8:1::/48 exit-address-family ``` For VPNv4 (MPLS L3VPN customer routes), you typically peer between PEs (provider edge) over the IPv4 underlay but exchange VPNv4 NLRI: ``` router bgp 65001 neighbor 10.0.0.2 remote-as 65001 neighbor 10.0.0.2 update-source Loopback0 address-family vpnv4 unicast neighbor 10.0.0.2 activate neighbor 10.0.0.2 send-community extended exit-address-family address-family ipv4 vrf CUSTOMER-A redistribute connected redistribute static exit-address-family ``` Notice the iBGP setup using loopbacks (typical for PEs in an MPLS network) and the per-VRF address family that scopes routes to a specific customer. ## Route Distinguishers and Route Targets (VPNv4) VPNv4 has a unique problem: customer A and customer B might both use 10.0.0.0/24 internally. Plain IPv4 cannot tell them apart. MP-BGP solves this with route distinguishers (RDs) and route targets (RTs). - **Route Distinguisher (RD).** An 8-byte prefix prepended to the IPv4 prefix when it is encoded as VPNv4\. The combined 12-byte (RD + IPv4) value is globally unique inside the MPLS network. Two customers using 10.0.0.0/24 get RDs of (say) 65001:1 and 65001:2, producing globally distinct VPNv4 prefixes. - **Route Target (RT).** An extended community attribute attached to each VPNv4 route describing which VRFs should import it. The RD identifies the prefix uniquely; the RT controls who can see it. Configuration on a PE for a customer VRF: ``` vrf definition CUSTOMER-A rd 65001:1 address-family ipv4 route-target export 65001:100 route-target import 65001:100 exit-address-family ``` Routes redistributed into BGP from this VRF get RD 65001:1 prepended (making them VPNv4) and RT 65001:100 attached. Other PEs receive the VPNv4 routes and only import them into VRFs that have `route-target import 65001:100` configured. ## EVPN: The Modern MP-BGP Killer App EVPN (Ethernet VPN, RFC 7432) is the data center overlay that replaced VPLS. It uses MP-BGP with AFI 25 (L2VPN) and SAFI 70 (EVPN) to carry MAC addresses, IP-MAC bindings, multicast group memberships, and VLAN-to-VNI mappings between VTEPs (VXLAN tunnel endpoints) in a fabric. What makes EVPN powerful: - **Control-plane MAC learning.** Instead of relying on data-plane flood-and-learn, EVPN advertises MAC addresses via BGP. Failover and convergence are deterministic. - **Active-active multihoming.** Multiple PEs can advertise the same MAC behind themselves, with proper election and load distribution. - **Layer 2 + Layer 3 + multicast in one control plane.** EVPN routes are typed (Type 2 = MAC, Type 3 = inclusive multicast, Type 5 = IP prefix); a single BGP session carries all of it. - **Vendor-neutral.** Cisco, Arista, Juniper, Nokia, Cumulus all interoperate on EVPN. EVPN is now standard in modern leaf-spine data center fabrics, in SD-WAN overlays, and in DC-DC interconnect. It is the dominant use of MP-BGP in 2026. ## IPv6 Over IPv4 Sessions (and Vice Versa) You can carry IPv6 routes over an IPv4 BGP session and vice versa, thanks to MP-BGP's flexibility. The peering is over one transport (IPv4 TCP or IPv6 TCP), but the AFs negotiated decide what kinds of routes are exchanged. The most common pattern: IPv4 transport for the BGP TCP session, with both ipv4 unicast and ipv6 unicast address families activated. The IPv6 NEXT\_HOP requires special handling (the MP\_REACH\_NLRI carries the v6 next hop separately from the v4 transport address), and on Cisco you typically configure: ``` neighbor 10.0.12.2 remote-as 65002 address-family ipv6 unicast neighbor 10.0.12.2 activate ! IPv6 next hop is automatically derived from BGP capability extension exit-address-family ``` If the v6 next-hop derivation fails (older code), explicitly configure: `neighbor 10.0.12.2 next-hop-self`. ## Show Commands ``` R1# show bgp ipv4 unicast summary R1# show bgp ipv6 unicast summary R1# show bgp vpnv4 unicast all summary R1# show bgp l2vpn evpn summary R1# show bgp ipv4 unicast neighbors 10.0.12.2 | include capability For address family: IPv4 Unicast Multiprotocol extensions: advertised and received Route refresh: advertised and received(new) 4-octet AS number: advertised and received Address family IPv4 Unicast: advertised and received Address family IPv6 Unicast: advertised and received ``` The capability output is the diagnostic confirmation that MP-BGP is negotiated. If "Multiprotocol extensions" shows "advertised" only (not "advertised and received"), the peer does not understand MP-BGP and you are running plain BGP-4. ## When You'd Need MP-BGP If you are running: - **IPv6 anywhere.** Native IPv6 BGP for internet peering, internal IPv6, or both. - **MPLS L3VPNs.** Provider running customer VPNs, or enterprise running its own internal MPLS. - **VXLAN/EVPN data center.** Modern leaf-spine, multi-tenant, or DC-DC fabric. - **SD-WAN overlay.** The control plane is BGP-based with EVPN-style routes. - **Inter-AS MPLS (Option B/C).** BGP-LU for label distribution between ASes. For pure IPv4 unicast internet edge BGP, MP-BGP is technically running but you do not really notice; the only address family is IPv4 unicast and the protocol behaves like classic BGP-4. ## Summary MP-BGP extends BGP-4 with two attributes (MP\_REACH\_NLRI, MP\_UNREACH\_NLRI) that let a single BGP session carry many address families: IPv6 unicast, VPNv4, VPNv6, L2VPN EVPN, BGP-LU, and more. Each (AFI, SAFI) pair is configured separately and activated per-neighbor; peers negotiate which families to exchange via OPEN message capability advertisement. If you are running anything beyond plain IPv4 unicast (which today is most production networks), you are using MP-BGP whether you call it that or not. The everyday Cisco syntax (`address-family X unicast`) is exactly the MP-BGP configuration mechanism. Bookmark this article, the [BGP cluster pillar](https://www.pinglabz.com/bgp/) for the inter-AS picture, and treat `show bgp X Y summary` as your daily diagnostic for any non-IPv4-unicast service. ### References - [RFC 4271 - A Border Gateway Protocol 4 (BGP-4)](https://www.rfc-editor.org/rfc/rfc4271?ref=pinglabz.com) - [Cisco BGP technology documentation](https://www.cisco.com/c/en/us/tech/ip/border-gateway-protocol-bgp/index.html?ref=pinglabz.com) Take the BGP reference with you The free BGP field-reference PDF: path attributes, best-path order, and the show commands that matter. Delivered by email, no card required. [Get the free PDF](https://www.pinglabz.com/bgp-cheatsheet/) ### STP vs RSTP: Convergence, Port Roles, and When to Switch URL: https://www.pinglabz.com/stp-vs-rstp/ Last updated: 2026-06-13T20:08:44.000Z Classic Spanning Tree Protocol (802.1D STP) and Rapid Spanning Tree Protocol (802.1w RSTP) solve the same problem in fundamentally different ways. STP was the 1990 design built around timers; RSTP is the 2001 redesign built around explicit handshaking. Most modern switches default to RSTP-based flavors (Cisco's Rapid PVST+, MST), but plenty of production networks still run with classic STP semantics because someone never enabled the modern mode. This article is a side-by-side comparison: what is identical, what is different, why the differences matter, and how to tell which one your network is actually running. If you are studying for CCNP, planning a migration, or trying to figure out why convergence in one part of your network takes 30 seconds while another part is sub-second, this is the comparison. ## What's the Same Both protocols share the same goal: maintain exactly one active path between any two switches in a Layer 2 network, with the others ready to take over if the active one fails. Both elect a Root Bridge, both compute Root Ports and Designated Ports, both forward only along the resulting tree, and both block the redundant ports until needed. The election rules are identical. Bridge ID = priority + system ID extension + base MAC. Lowest bridge ID wins the root election. Lowest cost path wins the Root Port election. Designated Port tiebreakers walk the same list in both protocols. BPDU format is mostly compatible. RSTP uses a slightly different BPDU type (version 2, with additional flags), but a switch running RSTP that receives a 802.1D version-0 BPDU falls back to legacy mode for that specific port. The two protocols can coexist on the same network during a migration. ## The Side-by-Side Comparison Standard year 802.1D STP1990 802.1w RSTP2001 Convergence on direct failure (Alternate available) 802.1D STP30-50 seconds 802.1w RSTPSub-second Convergence on indirect failure 802.1D STP 50 seconds (Max Age + 2x Forward Delay) 802.1w RSTP \~6 seconds (3 missed BPDUs) Port states 802.1D STP 5 (Disabled, Blocking, Listening, Learning, Forwarding) 802.1w RSTP 3 (Discarding, Learning, Forwarding) Port roles 802.1D STP 3 (Root, Designated, Non-Designated) 802.1w RSTP 5 (Root, Designated, Alternate, Backup, Disabled) BPDU origination 802.1D STP Only the root sends; non-root relays 802.1w RSTP Every switch sends every Hello BPDU loss tolerance 802.1D STP Max Age (20s = 10 missed BPDUs) 802.1w RSTP3 missed BPDUs (\~6s) Topology change handling 802.1D STP TCN walks to root, root sets TC bit, all flush MACs 802.1w RSTP Originator floods TC to neighbors directly Proposal/Agreement handshake 802.1D STPNone 802.1w RSTP Yes, on point-to-point links Edge port treatment 802.1D STP Cisco PortFast extension only 802.1w RSTP Built into the standard as edge port type Link types 802.1D STPImplicit 802.1w RSTP Explicit (point-to-point, shared, edge) Default Cisco mode 802.1D STPWas PVST+ pre-IOS 12.2 802.1w RSTP Rapid PVST+ since IOS 12.2(25)SEC ## The Convergence Story Classic STP convergence is timer-driven. After a topology change: 1. **Max Age (20s).** Wait for the old BPDU to age out from the LSDB. 2. **Listening state (15s, Forward Delay).** Send/receive BPDUs but do not yet learn MACs or forward data. 3. **Learning state (15s, Forward Delay).** Now learn MACs but still do not forward. 4. **Forwarding state.** Finally pass data. Total: 50 seconds for a port that was Blocking when the change happened. Even ports that just came up only skip step 1, leaving 30 seconds of pre-Forwarding states. RSTP eliminates the timers entirely on point-to-point links. The Alternate Port concept means every non-root switch already has a pre-computed backup Root Port standing by in Discarding state, ready to forward immediately when the current Root Port fails. The proposal/agreement handshake replaces the listening/learning timers with a one-BPDU-round-trip exchange that takes milliseconds. The result: classic STP takes 30-50 seconds to converge on any non-trivial topology change. RSTP converges in tens of milliseconds for direct failures, and tens of milliseconds plus 6 seconds (the 3-missed-BPDU window) for indirect failures. ## Port States: 5 to 3 Discarding Maps to 802.1D Disabled, Blocking, Listening Forwards data?No Learns MACs?No Learning Maps to 802.1DLearning Forwards data?No Learns MACs?Yes Forwarding Maps to 802.1DForwarding Forwards data?Yes Learns MACs?Yes The state-machine collapse matters because under RSTP a port does not normally walk through these states on a timer. It transitions directly from Discarding to Forwarding via the proposal/agreement handshake (on point-to-point) or via PortFast edge type (on host links). The fewer states are not just cosmetic; they enable the fast convergence model. ## Port Roles: 3 to 5 Classic STP has three roles: Root Port, Designated Port, and "Non-Designated" (a generic catch-all for any port not actively forwarding). RSTP elevates two specific Non-Designated cases to first-class roles: - **Alternate Port.** An immediate backup to the Root Port. If the Root Port fails, the Alternate is pre-computed and can take over instantly. - **Backup Port.** A backup to a Designated Port on the same shared segment. Only exists on multi-access (hub) segments. Rare in modern networks. The Alternate Port is the headline upgrade. It is what makes RSTP "fast" in practice: every non-root switch already knows what to do if its current Root Port goes down, no recomputation needed. ## How to Tell Which One You're Running Cisco command: ``` Switch# show spanning-tree summary | include mode Switch is in rapid-pvst mode ``` Possible outputs: - `pvst` \- Cisco PVST+ (per-VLAN, but classic 802.1D semantics; slow) - `rapid-pvst` \- Cisco Rapid PVST+ (per-VLAN, RSTP semantics; fast) - `mst` \- Multiple Spanning Tree (instances mapped to VLANs; RSTP-based) If you see `pvst`, you are running classic STP semantics regardless of how modern your hardware is. Switch to `rapid-pvst` with one command: ``` Switch(config)# spanning-tree mode rapid-pvst ``` Cisco's default has been Rapid PVST+ since around 2007\. Anything still running plain PVST+ is likely a legacy configuration that was never updated. ## Running Mixed (RSTP and STP on the Same Network) Yes, you can. RSTP includes 802.1D backward compatibility: when an RSTP switch receives a classic 802.1D BPDU, it falls back to legacy mode for that specific port and uses 802.1D timers. Other ports on the same switch still use RSTP. The implication: you can migrate a network from PVST+ to Rapid PVST+ one switch at a time. The downside is that any port that falls back to legacy loses the RSTP fast-convergence benefit. The whole topology only converges fast when all switches run RSTP. The migration recipe: 1. Audit the current state. `show spanning-tree summary` on every switch. 2. Plan the order. Start with edge switches (least disruptive if something goes wrong). 3. Per switch: `spanning-tree mode rapid-pvst`. Verify with `show spanning-tree summary` and `show spanning-tree vlan X`. 4. Watch for unexpected reconvergence on neighbor switches; classic STP will see new BPDU formats and may briefly recompute. 5. Continue until the whole topology is on Rapid PVST+. ## When You'd Still Run Classic STP Almost never. The only reasons to keep classic 802.1D / PVST+ in 2026: - **Equipment too old for RSTP.** Truly ancient kit (pre-2003 Cisco IOS, deeply legacy unmanaged switches). Replace it. - **Vendor compatibility issues.** Some old non-Cisco gear handles RSTP poorly during interop. Has been mostly fixed. - **Deliberate testing.** Lab scenarios where you want to demonstrate the convergence difference. Otherwise, run Rapid PVST+ or MST. The configuration is one global command, and the convergence improvement is multiple orders of magnitude. ## RSTP / Rapid PVST+ vs MST: The Next Decision Once you have decided not to run classic STP, the next decision is whether to run Rapid PVST+ (one RSTP instance per VLAN) or MST (multiple VLANs mapped to a small number of RSTP instances). Standard Rapid PVST+Cisco-proprietary MSTIEEE 802.1s Number of instances Rapid PVST+One per VLAN MST1-65 (typically 1-16) Per-VLAN load balancing Rapid PVST+ Easy (per-VLAN root manipulation) MSTCoarser (per-instance) CPU overhead at scale Rapid PVST+ High (one BPDU per VLAN per Hello) MST Low (one BPDU per instance) Configuration complexity Rapid PVST+Low MST Higher (regions, instance mapping) Vendor interop Rapid PVST+Cisco-only effectively MSTUniversal Rule of thumb: if you have more than 50 VLANs, MST is worth the configuration effort. Below that, Rapid PVST+ is simpler and fine. Detail in [Configuring Multiple Spanning Tree (MST)](https://www.pinglabz.com/configure-mst-cisco-switches/). ## Summary STP and RSTP solve the same problem (loop-free Layer 2) with the same election rules but radically different convergence. STP takes 30-50 seconds to converge on a topology change because it relies on timers; RSTP takes sub-second on point-to-point links because it relies on explicit handshaking and pre-computed Alternate Ports. If your network is running plain PVST+ in 2026, switch to Rapid PVST+ today. The configuration is one global command and the convergence improvement is real. If you have more than 50 VLANs, look at MST for the per-instance scalability. Bookmark the [Spanning Tree Protocol pillar](https://www.pinglabz.com/spanning-tree-protocol/) for the full operational picture, and the [RSTP deep dive](https://www.pinglabz.com/rapid-spanning-tree-protocol-rstp/) for the protocol-level details of what changed. ### BGP Neighbor States: The 6-State FSM and How to Diagnose It URL: https://www.pinglabz.com/bgp-neighbor-states/ Last updated: 2026-08-01T19:35:10.000Z BGP runs over TCP/179 and forms a session between two configured peers. That session walks through a state machine before it can exchange routing information, and getting stuck somewhere along the way is the most common BGP failure mode. The good news: every state has a specific meaning, and being stuck in a particular state tells you exactly where to look for the problem. This article walks through all six BGP neighbor states (Idle, Connect, Active, OpenSent, OpenConfirm, Established), what each state actually means, the events that drive transitions between them, and the diagnostic that maps each "stuck" state back to a real-world cause. If you have ever stared at `show ip bgp summary` wondering why a peer is in Active mode, this is the reference. ## The Six States Per RFC 4271 the BGP finite state machine has six states. The first three are about establishing the TCP session; the last three are about negotiating BGP itself on top of that TCP session. Idle What it means BGP process is starting; not yet trying to connect Healthy steady state? No (transient on session bring-up or after teardown) Connect What it means Trying to complete the TCP three-way handshake Healthy steady state? No (transient; should move to Active or OpenSent) Active What it means TCP failed; waiting to retry Healthy steady state?No (problem state) OpenSent What it means Sent our OPEN message; waiting for theirs Healthy steady state?No (transient) OpenConfirm What it means Received their OPEN; waiting for KEEPALIVE Healthy steady state?No (transient) Established What it means Session is up; updates can flow Healthy steady state? Yes (this is what you want) "Active" is famously misleading: in casual English you might think Active means "doing something good." In BGP, Active means "TCP did not work, I am actively trying to retry." It is a problem state, not a healthy one. ## State: Idle When BGP first starts (or after a hard reset), the peer transitions to Idle. In Idle, the BGP process is loaded but not actively trying to connect. It refuses incoming connections. The transition out of Idle happens on: - **Manual start.** An administrator runs `clear ip bgp` or first configures the peer. - **Automatic start.** The router determines it has a route to the peer's IP and begins the TCP attempt. If a peer is stuck in Idle, the most common cause is the IdleHoldTimer. After a session goes down, BGP enters Idle and waits a damping interval before retrying. The interval grows exponentially up to a cap to prevent flap loops. If you just configured a session and it is in Idle, check whether the route to the peer is actually present in the IP routing table. Cisco-specific gotcha: the BGP process can also be administratively shut down via `neighbor X.X.X.X shutdown`. The peer shows up in `show ip bgp summary` but with state Idle and a (Admin) suffix. ## State: Connect The router attempts a TCP connect to the peer's IP on port 179\. If the TCP handshake completes, transition to OpenSent. If the TCP handshake fails (timeout, RST), transition to Active. Connect is fast (sub-second on a healthy network). If you see Connect in `show ip bgp summary`, you have caught the session in the brief window between Idle and OpenSent or Active. It is not a stuck state in normal operation. Long stays in Connect generally indicate slow TCP handshakes (extreme latency, packet loss, asymmetric paths through firewalls). ## State: Active Active is the famous problem state. It means: I tried TCP, it failed, I am waiting to retry. The retry interval is the ConnectRetryTimer (default 120 seconds on Cisco, configurable). Why TCP fails: No route to peer's IP Diagnostic `show ip route X.X.X.X` returns nothing Fix Check IGP, static routes, or interface config ACL blocking TCP/179 Diagnostic `show access-lists`, look for matches Fix Permit TCP/179 between peers Source IP mismatch Diagnostic Peer's `neighbor` command expects a different source than we are sending from Fix `neighbor X update-source LoopbackN` Peer's IP wrong DiagnosticPinging the peer fails Fix Verify the IP of the loopback or interface on the peer eBGP TTL exhausted Diagnostic Peer is more than 1 hop away on eBGP, default TTL is 1 Fix `neighbor X ebgp-multihop N` Peer not configured to accept us Diagnostic Peer's BGP is not configured for our IP/AS Fix Coordinate with the other side MD5 authentication mismatch Diagnostic Logs show "MD5 mismatch" Fix Verify the secret on both sides Peer firewall, NAT, or routing Diagnostic traceroute fails, peer hard to reach Fix Network path investigation The single most useful sanity check when a peer is stuck in Active: try to telnet to the peer's IP on port 179\. If you cannot complete the TCP handshake from the BGP source IP to the peer's IP, BGP cannot either. ``` R1# telnet 10.0.12.2 179 /source-interface Loopback0 Trying 10.0.12.2, 179 ... Open [Press ctrl-shift-6 then x, then "disconnect"] ``` Open = TCP works. Refused = something at peer is closing the port. Timeout = ACL or routing problem along the path. Idle, Connect and Active are where nearly every real BGP call stalls, and the word in that column narrows the search faster than anything else you can type. If a peer of yours is sitting on one of the three right now, [decode the state and run the two-command triage](https://www.pinglabz.com/bgp-neighbor-stuck-idle-active-connect/) against captures of each case. ## State: OpenSent The TCP session is up. We sent our OPEN message (containing our AS number, hold time, BGP version, capabilities) and are waiting for the peer's OPEN. OpenSent is brief. If you see a peer stuck in OpenSent, the peer's OPEN never arrived or was malformed. Common causes of stuck OpenSent: - **BGP version mismatch.** Today both ends should run BGP-4; this is rare in 2026. - **Hold time too low.** If our advertised HoldTime is below 3 seconds (and not 0), the peer rejects us. Default is 180s; do not configure below 3. - **Capability mismatch.** One side advertises a capability (Route Refresh, 4-byte AS, Graceful Restart) that the other does not understand. Modern routers handle this gracefully via capability negotiation, but ancient code can hard-fail. - **BGP NOTIFICATION received.** The peer sent a NOTIFICATION rejecting our OPEN. Common reasons in the NOTIFICATION error code: bad peer AS, bad BGP identifier (router ID), authentication failure. The diagnostic: `debug ip bgp` on Cisco shows the OPEN parameters and any NOTIFICATIONs. ## State: OpenConfirm We sent OPEN, received OPEN, and now we are waiting for the first KEEPALIVE from the peer to confirm the session is up. OpenConfirm is also brief. If you see a stuck OpenConfirm, the peer's KEEPALIVE never arrived. This is usually a unidirectional path problem (we can send to them but they cannot send to us, often a one-way ACL). The transition to Established happens when we receive the first valid KEEPALIVE. ## State: Established The session is up. Both peers can now send UPDATE messages with prefixes and KEEPALIVE messages every (HoldTime / 3) seconds (default 60s on Cisco, with HoldTime 180s). Healthy steady-state for a working BGP peer. `show ip bgp summary` shows it as a number (the count of received prefixes from the peer) instead of a state name. If a peer drops out of Established, the cause is one of: - **HoldTime expired.** No KEEPALIVE or UPDATE received within HoldTime seconds. The session goes back to Idle. - **NOTIFICATION received.** The peer sent a NOTIFICATION because of a problem (cease, hold timer expired, parsing error, etc.). Session goes to Idle. - **TCP reset.** Underlying TCP died. Session goes to Idle. - **Manual clear.** Administrator ran `clear ip bgp`. Session goes to Idle. The HoldTime expiration is the most common production failure: a network glitch interrupts BGP traffic for HoldTime seconds, both ends decide the other is dead, and the session resets. Mitigation is BFD (sub-second detection), see below. Get the BGP Field Reference - 9 pages, free Everything you'd want to remember about BGP on nine printable pages. FSM diagram, all 13 best-path steps with gotchas, troubleshooting decision tree, copy-paste IOS XE templates, and real lab captures. Free for PingLabz members - just sign up with your email. [Get the BGP cheat-sheet](https://www.pinglabz.com/bgp-cheatsheet/) ## Diagnostic Cheat Sheet by State Idle Likely cause No route to peer; admin shutdown; flapping (IdleHoldTimer) First diagnostic `show ip route ` Active Likely causeTCP handshake failing First diagnostic `telnet 179` OpenSent Likely cause Peer's OPEN never arrived or rejected First diagnostic `debug ip bgp` \+ check NOTIFICATIONs OpenConfirm Likely cause Peer's first KEEPALIVE never arrived (often one-way path) First diagnostic traceroute both directions; check ACLs both ways Established but flapping Likely cause HoldTime expiration due to packet loss or path issue First diagnostic BFD; investigate path stability ## Timers That Affect the FSM ConnectRetry Default120s Purpose How long Active waits before retrying TCP HoldTime Default180s Purpose Negotiated min of both peers' HoldTime; 0 means do not check KeepAlive Default HoldTime / 3 (60s default) Purpose How often KEEPALIVE messages are sent in Established IdleHoldTimer Default Variable, exponential backoff Purpose Damping after session failures; prevents flap loops HoldTime negotiation: each peer advertises its HoldTime in OPEN, and they use the lower of the two. Setting HoldTime to 0 disables the check entirely; setting it below 3 is invalid. ## BFD: Sub-Second BGP Failure Detection The default HoldTime of 180 seconds is sane for protocol stability but unacceptable for operations. A link can be dead for three minutes before BGP notices, during which traffic blackholes. BFD (Bidirectional Forwarding Detection) plugs this gap. BFD runs as a separate protocol below BGP, exchanging tiny UDP packets at 50-300 ms intervals. When BFD detects loss (3 consecutive missed packets, default), it tells BGP to bring the session down immediately, regardless of the HoldTime. Configuration: ``` R1(config-if)# bfd interval 100 min_rx 100 multiplier 3 R1(config)# router bgp 65001 R1(config-router)# neighbor 10.0.12.2 fall-over bfd ``` Result: BGP neighbor failure detection in 300ms instead of 180s. [BGP Convergence: Timers, BFD, and Reducing Failover Time](https://www.pinglabz.com/bgp-convergence-troubleshooting/) covers the full pattern. ## The Three Commands You Need ``` R1# show ip bgp summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 10.0.12.2 4 65002 1 1 0 0 0 never Active 10.0.13.3 4 65003 12345 12345 1234 0 0 02:30:15 4521 R1# show ip bgp neighbors 10.0.12.2 | section state BGP state = Active Neighbor sessions: 0 active, is multisession capable R1# debug ip bgp 10.0.12.2 events *Apr 24 10:12:34: BGP: 10.0.12.2 went from Active to Idle *Apr 24 10:12:34: BGP: 10.0.12.2 active rejected for connect retry timer expired ``` The `State/PfxRcd` column in `show ip bgp summary` is the daily diagnostic. A number = Established + count of prefixes received. A state name = problem in that state. ## Summary BGP's six-state finite state machine is unusual in that "Active" is a problem state, not a healthy one. The healthy progression is Idle to Connect to OpenSent to OpenConfirm to Established. Anything else is a stuck transition that maps to a specific cause: routing, ACLs, configuration mismatch, or path failure. If you only remember three things: Active means TCP failed, OpenSent stuck means OPEN was rejected or dropped, and Established but flapping means HoldTime expiration (use BFD). Bookmark this article alongside the [BGP cluster pillar](https://www.pinglabz.com/bgp/) as your day-one debug reference. ### References - [RFC 4271 - A Border Gateway Protocol 4 (BGP-4)](https://www.rfc-editor.org/rfc/rfc4271?ref=pinglabz.com) - [Cisco BGP technology documentation](https://www.cisco.com/c/en/us/tech/ip/border-gateway-protocol-bgp/index.html?ref=pinglabz.com) Take the BGP reference with you The free BGP field-reference PDF: path attributes, best-path order, and the show commands that matter. Delivered by email, no card required. [Get the free PDF](https://www.pinglabz.com/bgp-cheatsheet/) ### Tagged vs Untagged VLANs: How Trunks and Access Ports Really Work URL: https://www.pinglabz.com/tagged-vs-untagged-vlans/ Last updated: 2026-06-13T20:08:45.000Z "Tagged" and "untagged" are how you describe whether a frame on a switch port is carrying an 802.1Q VLAN tag or not. The terminology matters because half of every VLAN troubleshooting ticket comes down to a port being tagged when it should be untagged, or vice versa. If you have ever had a working configuration that mysteriously broke when you connected a new vendor's switch, the answer was almost always a tagged/untagged mismatch. This article walks through what tagged and untagged actually mean, how Cisco and other vendors describe the same concept with slightly different vocabulary, when to use each, and the cross-vendor translation cheat sheet you need when working in mixed environments. If you are studying for CCNA, configuring switches across multiple vendors, or troubleshooting a trunk that should "just work", this is the reference. ## The Definitions An **untagged frame** is an Ethernet frame with no 802.1Q VLAN tag. It looks exactly like a frame you would see on a network with no VLANs at all: destination MAC, source MAC, EtherType, payload, FCS. End hosts (PCs, printers, IP phones in the data context, almost everything that is not a switch or router) send and receive untagged frames natively. A **tagged frame** has a 4-byte 802.1Q tag inserted between the source MAC and the EtherType. The tag carries the VLAN ID. Switches use the tag to know which VLAN a frame belongs to as it crosses a link that carries multiple VLANs. So the question "is the frame tagged?" is the same as "does this frame carry a VLAN ID?" An untagged frame leaves it implicit (the receiver has to know the VLAN from context). A tagged frame makes it explicit. ## Port Types: Where Tagged and Untagged Live Access port Sends Always untagged (one VLAN) Receives Untagged only (drops tagged frames in most modes) Use for End hosts (PCs, printers) Trunk port Sends Tagged (with one exception: native VLAN) Receives Tagged, plus untagged for native VLAN Use for Switch-to-switch, switch-to-router Voice + access port Sends Tagged for voice, untagged for data Receives Tagged for voice, untagged for data Use for IP phone with PC behind it The confusion enters when other vendors use different vocabulary for the same things. HP/Aruba and Juniper, for example, describe ports in terms of how each VLAN is treated rather than calling the port "access" or "trunk": Access port (in VLAN 10) HP/Aruba ProCurve Untagged on VLAN 10, no other Juniper JunosAccess port, VLAN 10 EffectOne VLAN, untagged Trunk allowed VLANs 10,20,30 HP/Aruba ProCurve Tagged on 10, 20, 30; no untagged Juniper Junos Trunk port, VLANs 10/20/30 Effect Multiple VLANs, all tagged Trunk native VLAN 999 HP/Aruba ProCurve Untagged on 999, tagged on others Juniper Junos Trunk with native VLAN 999 Effect Native VLAN untagged, others tagged The HP/Aruba mental model is per-VLAN: for each VLAN, decide whether this port carries it tagged, untagged, or not at all. The Cisco mental model is per-port: assign the port a mode (access or trunk), then derive the per-VLAN behavior from that. Both express the same thing but invite different mistakes. HP/Aruba lets you accidentally configure a port that is untagged on multiple VLANs (which causes ambiguity). Cisco hides the per-VLAN view behind `switchport mode`, which means you sometimes need to debug what tag actually leaves the port. ## When You Want Tagged Tag frames when: - **The port carries multiple VLANs.** The receiver needs to know which VLAN each frame belongs to. This is the trunk case. - **The receiver is a 802.1Q-aware device.** A switch, router, virtualization host (vSwitch in tagged mode), wireless AP, firewall in trunk mode. - **You want explicit control.** Tagged frames are unambiguous; if the tag says VLAN 100, that is what it is, regardless of what port-level configuration the port has. The trunk between two switches is always tagged for every VLAN except the native VLAN. The trunk between a switch and a router-on-a-stick configuration uses tagged subinterfaces. The link from a switch to a virtualization host is tagged when the host's vSwitch is configured for trunked VLAN access (VMware "VLAN tagging mode 4095", Hyper-V trunked vNIC, etc.). ## When You Want Untagged Use untagged frames when: - **The port carries exactly one VLAN.** The receiver does not need a tag because there is no ambiguity. - **The receiver is not 802.1Q-aware.** Most end hosts, simple printers, IoT devices, and entry-level network gear default to untagged. - **You are using a native VLAN on a trunk.** Frames in the native VLAN cross the trunk untagged for backwards compatibility. The host-facing access port is the most common untagged case. PCs do not normally know how to send tagged frames; they send plain Ethernet, and the switch internally classifies the frame into the access VLAN configured on that port. ## The Native VLAN: One Port, One VLAN, Untagged on a Trunk The native VLAN is a special case where a trunk port (which is otherwise tagged) carries one specific VLAN untagged. It exists for backwards compatibility with hubs and old equipment that did not understand 802.1Q. Default native VLAN: VLAN 1\. As covered in detail in [VLAN Hopping Attacks](https://www.pinglabz.com/vlan-hopping-attacks/) and [Native VLAN Configuration and Security](https://www.pinglabz.com/change-native-vlan-cisco-switch/), you should always change the native VLAN to a dedicated unused VLAN (e.g. 999) for security reasons. Both ends of the trunk must agree on the native VLAN, or you get a CDP warning and possibly STP issues. The PingLabz pattern: explicitly set the native VLAN on every trunk to the same dedicated unused VLAN, and ensure no host port is in that VLAN. ## The Voice + Data Special Case IP phones are a special case that mixes tagged and untagged on the same physical port. The phone receives traffic for two VLANs: - **Voice VLAN.** Tagged. Carries the phone's voice traffic (RTP, signaling). - **Data VLAN.** Untagged. Carries traffic from the PC connected behind the phone. From the switch's perspective, the port is configured as an access port in the data VLAN, with a separate "voice VLAN" assigned. The switch treats the port as access (untagged) for the data VLAN, and tagged (using 802.1Q) for the voice VLAN. The phone receives both, splits them based on tag, and forwards the data VLAN traffic to the PC port behind it untagged. Configuration: ``` Switch(config)# interface GigabitEthernet1/0/5 Switch(config-if)# switchport mode access Switch(config-if)# switchport access vlan 10 ! Data VLAN (untagged) Switch(config-if)# switchport voice vlan 20 ! Voice VLAN (tagged) Switch(config-if)# spanning-tree portfast Switch(config-if)# spanning-tree bpduguard enable ``` Detail in [Configuring Voice VLANs on Cisco Switches for IP Phones](https://www.pinglabz.com/voice-vlan-cisco-configuration/). ## The Tagged/Untagged Mismatch: How Trunks Break The classic failure: one end of a trunk has VLAN 100 tagged, the other end has VLAN 100 untagged (perhaps because someone configured the second switch as if VLAN 100 were the native VLAN). What happens? Frames in VLAN 100 sent from end A arrive at end B with a tag. End B is expecting them untagged, sees the tag as confusion, and the behavior depends on platform: Cisco often drops the tagged frames on a port expecting the native VLAN to be different. Some platforms log warnings; some just silently drop. The other direction: end B sends frames in VLAN 100 untagged (because B thinks 100 is the native VLAN). End A receives the untagged frames and treats them as belonging to A's native VLAN (whatever A's native is, which is something different). The frames disappear into the wrong VLAN. The first symptom: hosts on VLAN 100 cannot reach each other across the trunk, even though "the trunk is up". `show interfaces trunk` on Cisco shows the trunk in trunking state but does not always flag the mismatch. `show cdp neighbors` often does, with a Native VLAN Mismatch warning. ## Diagnosing Tagged/Untagged Issues Three Cisco commands for the diagnosis: ``` ! What VLANs are tagged on this trunk? Switch# show interfaces gi1/0/24 trunk ! What VLAN is untagged (native) on this trunk? Switch# show interfaces gi1/0/24 switchport | include Native ! CDP neighbor warnings (often catches mismatches) Switch# show cdp neighbors detail | include Native ``` For per-VLAN packet capture, mirror the trunk port to a SPAN session and capture with Wireshark. Tagged frames show as "802.1Q Virtual LAN" in the packet detail; untagged frames do not. ## Tagged vs Untagged in Virtualization vSwitches in VMware, Hyper-V, KVM, and proxmox all support both tagged and untagged uplinks. The terminology varies but the concept is identical: VLAN ID 0 (or no VLAN) Equivalent toUntagged port Use for Single VLAN, host attaches to physical untagged port VLAN ID 1-4094 on port group Equivalent to Access port for that VLAN Use for VMs in one VLAN; host receives tagged but vSwitch strips tag VLAN trunking (VLAN ID 4095 in VMware, "trunk" in Hyper-V) Equivalent to802.1Q trunk to the VM Use for VM does its own tagging (firewall, router, hypervisor inside hypervisor) Most VMs run in "access" mode where the vSwitch does the tagging on the wire (VST, Virtual Switch Tagging) and the VM only sees untagged frames. The vSwitch is configured as a Layer 2 trunk to the physical switch carrying multiple VLANs tagged. ## Cross-Vendor Translation Cheat Sheet Host port in VLAN 10 Cisco IOS XE `switchport mode access` \+ `switchport access vlan 10` HP/Aruba ProCurve `vlan 10` \+ `untagged 1/0/1` Juniper Junos `family ethernet-switching { interface-mode access; vlan { members 10; } }` Trunk carrying VLANs 10, 20, 30 Cisco IOS XE `switchport mode trunk` \+ `switchport trunk allowed vlan 10,20,30` HP/Aruba ProCurve `vlan 10 tagged 1/0/24` (and 20, 30) Juniper Junos `interface-mode trunk; vlan { members [ 10 20 30 ]; }` Trunk with native VLAN 999 Cisco IOS XE `switchport trunk native vlan 999` HP/Aruba ProCurve `vlan 999 untagged 1/0/24` (and tagged for others) Juniper Junos`native-vlan-id 999` The functional outcome is identical; the vocabulary differs. Whenever you bring up a trunk between Cisco and HP/Aruba, the most reliable check is to verify that one VLAN per trunk is "untagged on Aruba" and "native on Cisco" with the same VID. ## Summary Tagged frames carry an explicit 802.1Q VLAN ID; untagged frames do not. Switches use tagging on multi-VLAN trunks to keep VLAN identity intact across links. Host-facing access ports send and receive untagged frames in one VLAN. The native VLAN is the one VLAN on a trunk that is sent untagged for legacy compatibility. If you are working across vendors, remember that "tagged" and "untagged" are platform-neutral concepts; everyone implements 802.1Q the same way, but the configuration vocabulary differs. Mismatches between tagged and untagged on the two ends of a trunk are the leading cause of "trunk is up but traffic does not flow" tickets. Bookmark the [VLAN cluster pillar](https://www.pinglabz.com/vlans-layer-2-switching/) for the full operational picture. ### VLAN Hopping Attacks Explained: Switch Spoofing and Double Tagging URL: https://www.pinglabz.com/vlan-hopping-attacks/ Last updated: 2026-08-01T19:32:09.000Z VLAN hopping is a class of Layer 2 attacks that lets a malicious host on one VLAN reach another VLAN without going through a router or firewall. It bypasses every Layer 3 control you have. Most networks are still vulnerable to it because the defaults that enable it (Dynamic Trunking Protocol on every port, native VLAN of 1) are still defaults. This article walks through the two main VLAN hopping attacks, how they actually work at the protocol level, the four-line configuration that defeats both, and the broader Layer 2 hardening pattern. If you are a network or security engineer responsible for a Cisco-based campus, treat this as the minimum-required read. ## The Threat Model VLAN hopping assumes the attacker is on the network. They are connected to a switch port (perhaps a wall jack in a conference room, perhaps a VM in a virtualized fabric) on one VLAN. From there, they want to reach a different VLAN: the server VLAN, the management VLAN, the voice VLAN. The conventional defense is that VLANs are isolated at Layer 2\. The only way out is through a Layer 3 gateway, where you presumably have ACLs and inspection. VLAN hopping defeats this assumption. The attacker reaches the target VLAN at Layer 2, before any router gets a chance to apply policy. Two attacks dominate: switch spoofing (DTP-based) and double tagging (native VLAN-based). ## Attack #1: Switch Spoofing via DTP Cisco's Dynamic Trunking Protocol (DTP) is enabled by default on Catalyst switch ports. Its purpose is automatic trunk negotiation: if you connect two switches, DTP detects this and forms a trunk without manual configuration. The default mode is `dynamic auto`, which will form a trunk if the other side initiates. The attack: the malicious host on the access port speaks DTP and proposes "I am a switch, let's negotiate a trunk." The Catalyst switch agrees (default behavior of `dynamic auto` when receiving a desirable proposal), and the port becomes a trunk. Now the attacker receives every VLAN that is allowed on the trunk, which by default is every VLAN that exists on the switch. The walk-through: 1. Attacker plugs into a wall jack. Default Cisco port config: `switchport mode dynamic auto`, native VLAN 1. 2. Attacker runs a tool like Yersinia that crafts DTP packets advertising `desirable` mode. 3. Switch sees the DTP advertisement, transitions the port to trunk mode automatically. 4. Attacker now sees 802.1Q-tagged frames for every VLAN configured on the switch. Connecting to any VLAN is a matter of joining a virtual interface to that VLAN ID and getting an IP via DHCP for that subnet. The attacker can now ARP, DHCP, send any frame, and receive any frame on any VLAN. From the network's perspective, the attacker is just another switch. ## Defense: Disable DTP Everywhere The fix is two lines per access port: ``` Switch(config-if)# switchport mode access Switch(config-if)# switchport nonegotiate ``` The first command pins the port to access mode. The second disables DTP entirely (sends no DTP frames, ignores received DTP frames). With both in place, the port cannot become a trunk via negotiation. An attacker speaking DTP gets no response. For trunk ports between switches: ``` Switch(config-if)# switchport mode trunk Switch(config-if)# switchport trunk encapsulation dot1q Switch(config-if)# switchport nonegotiate ``` The trunk is configured manually (no DTP). This is the PingLabz default for every trunk in production. Why `switchport nonegotiate` matters even if you set `switchport mode access`: the access mode does not stop the port from sending DTP frames. The frames are visible on the wire and can leak information about the switch. `nonegotiate` shuts DTP off entirely. ## Attack #2: Double Tagging via the Native VLAN Native VLAN frames cross trunks untagged. An attacker on the native VLAN can craft a frame with two 802.1Q tags: an outer tag matching the native VLAN, and an inner tag matching the target VLAN. When the frame hits the first switch, the switch strips the outer tag (because it matches the native VLAN, which is sent untagged across the trunk), and then forwards the frame across the trunk with only the inner tag intact. The next switch sees a frame tagged for the target VLAN and forwards it accordingly. The walk-through: 1. Attacker is on a port in the native VLAN (default: VLAN 1). 2. Attacker crafts a frame with two 802.1Q tags: outer VID = 1 (native), inner VID = 100 (target server VLAN). 3. Frame hits the first switch on an access port. The switch internally tags the frame with VLAN 1 (the access port's VLAN). The frame now carries: outer (added by switch) VLAN 1 + inner attacker-added VLAN 1 + inner attacker-added VLAN 100 + payload. 4. Frame is forwarded out a trunk. The switch sees the frame is in VLAN 1 (the native VLAN of the trunk), so it sends it untagged. It strips one tag, leaving: outer VLAN 1 (now appearing as the "native untagged" frame) + inner VLAN 100. 5. Wait, that is not quite right. Let me re-read the actual mechanism. The cleaner walk-through, accurate to the protocol: 1. The attacker on an access port in VLAN 1 sends a frame with one 802.1Q tag inside it: tag VID = 100 (target VLAN). 2. The first switch receives this frame on an access port. From the switch's perspective, the frame arrived on an access port in VLAN 1, so the switch internally classifies it as VLAN 1. 3. The switch forwards the frame out an 802.1Q trunk. Because the frame is in VLAN 1 and VLAN 1 is the native VLAN of the trunk, the switch sends it untagged - meaning it does not add a new outer tag. But the attacker's inner tag (VID 100) is still there. 4. The trunk peer receives the frame, sees one tag (VID 100), and treats it as a frame for VLAN 100\. The frame has hopped from VLAN 1 to VLAN 100. This attack is one-way: the attacker can send into the target VLAN, but cannot receive responses (because the response would be tagged VID 100 over the trunk, which the first switch would direct to VLAN 100, not VLAN 1 where the attacker actually is). For attacks like sending DHCP starvation, ARP spoofing, or single-direction packet injection (some exploits), one-way is enough. ## Defense: Move the Native VLAN, Tag It Explicitly Two complementary mitigations: **1\. Change the native VLAN away from VLAN 1 on every trunk.** ``` Switch(config)# vlan 999 Switch(config-vlan)# name UNUSED-NATIVE Switch(config-vlan)# exit Switch(config)# interface range GigabitEthernet1/0/24-26 Switch(config-if-range)# switchport trunk native vlan 999 ``` Now an attacker in VLAN 1 cannot use the double-tag trick because VLAN 1 is no longer the native VLAN. VLAN 999 is, and you have ensured no host port is in VLAN 999. **2\. Force the native VLAN to be tagged on the trunk.** ``` Switch(config)# vlan dot1q tag native ``` This global command tells the switch to add an explicit 802.1Q tag for the native VLAN frames as well, instead of sending them untagged. Now there is no untagged-native shortcut for the attack to exploit. Every frame on the trunk has a tag, and a doubly-tagged attacker frame appears as such. Either mitigation alone is effective; both together is defense in depth. Worth knowing that the untagged-native shortcut shows up by accident far more often than it shows up in an attack. [A native VLAN mismatch between two trunk neighbours](https://www.pinglabz.com/native-vlan-mismatch-troubleshooting/) merges two VLANs into one without anybody crafting a frame, which is the same primitive double tagging abuses deliberately, so learning to spot it in the CDP logs pays off twice. ## This Is Not a Cisco-Only Issue VLAN hopping works on any 802.1Q-compliant switch from any vendor that allows untagged native VLAN frames on trunks. The DTP variant is Cisco-specific (DTP is Cisco-proprietary), but the double-tagging variant is universal. Juniper, Arista, HP/Aruba switches all share the same native-VLAN behavior by default and require equivalent mitigations. ## The Broader Layer 2 Hardening Checklist VLAN hopping is one of about a dozen Layer 2 attacks. Here is the production hardening pattern in the order we recommend applying it: `switchport nonegotiate` DefeatsSwitch spoofing (DTP) Where to applyEvery port Native VLAN moved off VLAN 1 DefeatsDouble tagging Where to applyEvery trunk `vlan dot1q tag native` DefeatsDouble tagging Where to applyGlobally Trunk allowed VLAN list pruned Defeats Reduces blast radius if other defenses fail Where to applyEvery trunk BPDU Guard + PortFast Defeats Rogue switches plugged into access ports Where to applyEvery host port Storm Control Defeats Broadcast/multicast floods Where to applyEvery host port DHCP Snooping DefeatsRogue DHCP servers Where to apply Globally + on host ports as untrusted Dynamic ARP Inspection DefeatsARP spoofing Where to apply On VLANs that need it; depends on DHCP Snooping Port Security Defeats MAC flooding, unauthorized devices Where to apply Host ports (with care for voice + PC) VLAN 1 disabled Defeats Defense in depth; no traffic in VLAN 1 anywhere Where to applyEverywhere The full configuration patterns for each are in [VLAN Security Hardening: Protecting Your Layer 2 Network](https://www.pinglabz.com/vlan-security-best-practices/). ## How to Test (in a Lab) If you want to verify your switches are not vulnerable, the testing tools are well-known but only run them in a lab you own: - **Yersinia.** Open-source Layer 2 attack framework. Has DTP and double-tagging modules. - **Scapy.** Python framework that can craft arbitrary 802.1Q-tagged frames for the double-tagging attack. The acceptable lab pattern: build a topology with two switches and three hosts (one in each VLAN, plus an attacker), run the attack against an unhardened configuration, observe the attacker reaching the target VLAN, apply the hardening, and confirm the attack now fails. This is part of every CCNP Security and Cisco SECURE lab. ## Why VLAN Hopping Still Matters in 2026 The mitigations have been documented for over twenty years. The attacks are taught in every CCNA Security textbook. And yet network audits keep finding production trunks with the default native VLAN of 1, access ports with default DTP enabled, and management VLANs that share trunk access with everything else. Three reasons: 1. **Defaults are sticky.** Many switch deployments inherit configurations that pre-date hardening guides. Re-templating a deployed network is hard. 2. **Access is assumed.** If your physical security is strong, the attacker cannot get to a wall jack. But voice phones, conference rooms, BYOD, and lab environments all undermine this assumption. 3. **Virtualization extends the attack surface.** A VM on a vSwitch with default VLAN settings can perform the same attacks as a physical attacker. The double-tagging vector reaches into virtual networks too. The fix is short, well-understood, and free. The PingLabz position: every trunk in 2026 should have `switchport nonegotiate`, a non-default native VLAN, and explicit allowed VLAN lists. If your network does not, this is the lowest-effort security improvement you can make today. ## Summary VLAN hopping is two attacks: switch spoofing via DTP and double tagging via the native VLAN. Both let an attacker bypass Layer 2 isolation and reach VLANs they should not be on. Both are defeated by short configurations: `switchport nonegotiate` on every port, native VLAN moved off VLAN 1, and either `vlan dot1q tag native` globally or aggressive native-VLAN hygiene on trunks. If your network is running default Cisco settings, you are vulnerable. The fix takes minutes; the consequence of not fixing it is real, recurrent, and exploitable. Bookmark the [VLAN cluster pillar](https://www.pinglabz.com/vlans-layer-2-switching/) for the full operational picture, and review the broader Layer 2 hardening checklist above on every audit. New labs and guides, in your inbox Every new PingLabz lab and deep-dive, built and verified on real Cisco IOS XE - free, straight to your inbox. [Join free](https://www.pinglabz.com/signup/) ### 802.1Q VLAN Tag Explained: The 4 Bytes That Make Trunking Work URL: https://www.pinglabz.com/802-1q-vlan-tag-explained/ Last updated: 2026-06-13T20:08:46.000Z The 802.1Q VLAN tag is a 4-byte field that gets inserted into Ethernet frames to carry VLAN identity across switch boundaries. It is the mechanism that makes VLAN trunking possible. If you are studying for CCNA, troubleshooting a trunk that "should be working", or trying to understand why double-tagging is a real attack, you need to understand this field byte by byte. This article walks through the 802.1Q tag format, what each field means, how the tag interacts with native VLANs, the difference between 802.1Q and the older Cisco ISL encapsulation, and the operational implications (MTU, QoS, security). It is the reference you will come back to. ## The Problem 802.1Q Solves A standard Ethernet frame has no VLAN identifier. The destination MAC, source MAC, EtherType, payload, and FCS are all the frame carries. When two switches are connected by a single physical link and need to exchange traffic for many VLANs across that link, plain Ethernet has no way to say which VLAN a frame belongs to. The pre-standard solutions varied. Cisco's ISL (Inter-Switch Link) encapsulated the entire Ethernet frame inside a new ISL header. 3Com had its own proprietary tagging. None of them interoperated. IEEE 802.1Q standardized the answer in 1998: instead of encapsulating, insert a 4-byte tag into the existing Ethernet header. The tag carries the VLAN ID, plus a few extra fields. The result is that any 802.1Q-capable switch from any vendor can interpret the tag, the frame format is otherwise familiar, and the overhead is minimal. ## The 802.1Q Frame Format An untagged Ethernet frame looks like this: ``` +-------------+-------------+-----------+----------+-----+ | Dest MAC | Src MAC | EtherType | Payload | FCS | | 6 bytes | 6 bytes | 2 bytes | 46-1500B | 4 B | +-------------+-------------+-----------+----------+-----+ ``` An 802.1Q-tagged frame inserts 4 new bytes between Src MAC and EtherType: ``` +-----------+-----------+----------+----------+----------+----------+-----+ | Dest MAC | Src MAC | TPID | TCI | EtherType| Payload | FCS | | 6 bytes | 6 bytes | 2 bytes | 2 bytes | 2 bytes | 46-1500B | 4 B | +-----------+-----------+----------+----------+----------+----------+-----+ \---- 802.1Q tag ----/ ``` The 4-byte tag splits into two 2-byte fields: TPID (Tag Protocol Identifier) and TCI (Tag Control Information). Total frame size grows from 1518 bytes max (1522 if you count FCS) to 1522 bytes (1526 with FCS), which is why you see "baby giant" support of 1522-byte frames on switch ports that handle trunks. ## TPID: Tag Protocol Identifier The TPID is always 0x8100\. It tells the receiving device "this is an 802.1Q-tagged frame, parse the next 2 bytes as TCI." By placing the TPID where the EtherType normally lives, the tag is detectable: a switch that does not understand 802.1Q sees 0x8100 and either drops the frame (because there is no protocol called 0x8100 it knows about) or, worse, treats the rest of the frame as having an unknown protocol. This is one reason you do not connect 802.1Q-tagged trunks to non-802.1Q-capable equipment. For QinQ (provider-edge double tagging), an outer TPID of 0x88A8 is used by the standard, with vendors historically using 0x9100 or 0x9200\. The inner TPID stays 0x8100\. [Private VLAN](https://www.pinglabz.com/private-vlans-cisco-configuration/) contexts are unrelated to QinQ but worth knowing about. ## TCI: Tag Control Information The TCI is 2 bytes (16 bits) split into three fields: ``` 3 bits 1 bit 12 bits +--------+-----------+-----------------+ | PCP | DEI | VID | +--------+-----------+-----------------+ ``` PCP (Priority Code Point) Bits3 Purpose QoS priority, values 0-7\. Higher = higher priority. DEI (Drop Eligible Indicator) Bits1 Purpose Drop preference; 1 means this frame is preferentially dropped under congestion. VID (VLAN ID) Bits12 Purpose The VLAN this frame belongs to. Values 0-4095, with 0 and 4095 reserved. ## PCP: How VLAN Priority Carries QoS The 3-bit PCP field gives 8 priority levels (0-7). This is IEEE 802.1p's class-of-service mechanism layered on top of 802.1Q. Higher numerical values mean higher priority for queueing decisions on the switch. Standard mappings (informational, not strict): Network control PCP value7 Typical useSTP, OSPF, BGP, BFD Internetwork control PCP value6 Typical use Routing protocol updates Voice PCP value5 Typical useRTP voice payload Video PCP value4 Typical useRTP video payload Critical applications PCP value3 Typical useSignaling (SIP, H.323) Excellent effort PCP value2 Typical use Bulk applications with priority Background PCP value1 Typical useLower than best-effort Best effort PCP value0 Typical useDefault On a Cisco switch, the PCP value is automatically derived from the DSCP value of the inner IP packet via a configurable mapping. `show mls qos maps` on Catalyst, or `show platform hardware fed switch active qos dscp-cos counters` on Catalyst 9000\. The PCP is what carries QoS across switch hops; once the frame is decapsulated to a router, the IP DSCP takes over. ## DEI: The 802.1Q-2011 Drop Eligible Indicator Originally this bit was the "Canonical Format Indicator" (CFI), used to signal whether MAC addresses were in canonical (Ethernet) or non-canonical (Token Ring) format. With Token Ring effectively extinct, IEEE 802.1Q-2011 redefined the bit as DEI: Drop Eligible Indicator. DEI = 1 means "this frame is preferentially dropped under congestion." It is used in service provider QinQ environments where the provider network uses DEI to mark frames that exceed contract rates. In most enterprise environments DEI is 0 always. ## VID: The VLAN ID 12 bits give 4096 possible values (0-4095), but the standard reserves both ends: 0 "Priority tag" - no VLAN, frame uses port VLAN, but PCP/DEI carry meaning 1-1001 Normal range; VLAN 1 is the default; VLANs 1002-1005 are FDDI/Token Ring legacy 1006-4094 Extended range (Cisco needs VTP transparent or VTP v3) 4095 Reserved by the standard; never used as a real VLAN So the practical range is 1-4094\. Most enterprises use the normal range (1-1001) for the bulk of their VLANs. ## The Native VLAN: One VLAN That Doesn't Get Tagged On an 802.1Q trunk, every VLAN gets a tag except one: the native VLAN. Frames in the native VLAN traverse the trunk untagged, exactly like a regular access-port frame. Why? Backwards compatibility with hubs and old equipment that did not understand 802.1Q. If you connected a hub between two switches, the hub would forward 802.1Q-tagged frames as opaque data, and the receiving switch would still know what VLAN they belonged to. Frames in the native VLAN, untagged, would also work on the hub because they look like regular Ethernet. By default, the native VLAN is VLAN 1\. This is dangerous for two reasons: 1. **Mixing of control and data.** CDP, VTP, PAgP, DTP all use VLAN 1 by default. Leaving the native VLAN as VLAN 1 means your management/control plane and user data share the same untagged segment. 2. **Double-tagging attacks.** An attacker in the native VLAN can inject a frame with two 802.1Q tags. The first switch strips the outer tag (because it matches the native VLAN), then forwards the frame to the trunk. The next switch sees the inner tag and treats the frame as belonging to that VLAN, regardless of where the attacker actually is. [VLAN Security Hardening](https://www.pinglabz.com/vlan-security-best-practices/) covers the full attack and the mitigation. The PingLabz default: change the native VLAN to a dedicated unused VLAN (e.g. VLAN 999), never let a host port be in it. The classic configuration: ``` Switch(config)# vlan 999 Switch(config-vlan)# name UNUSED-NATIVE Switch(config-vlan)# exit Switch(config)# interface range GigabitEthernet1/0/24-26 Switch(config-if-range)# switchport mode trunk Switch(config-if-range)# switchport trunk encapsulation dot1q Switch(config-if-range)# switchport trunk native vlan 999 ``` Both ends of every trunk must agree on the native VLAN, or you get a CDP warning and possibly STP issues. [Native VLAN Configuration and Security on Cisco Switches](https://www.pinglabz.com/change-native-vlan-cisco-switch/) walks through the change. ## 802.1Q vs ISL: Why Cisco's Original Tagging Lost Standard 802.1QIEEE 802.1Q (1998) ISLCisco-proprietary Approach 802.1Q Inserts 4-byte tag inside the Ethernet header ISL Encapsulates the entire Ethernet frame in a new ISL header Overhead 802.1Q4 bytes ISL 30 bytes (26 ISL header + 4 ISL trailer) Native VLAN concept 802.1Q Yes (one VLAN untagged) ISLNo (every VLAN tagged) Vendor support 802.1QUniversal ISLCisco only Status in 2026 802.1QUniversal default ISL Deprecated; not supported on modern Catalyst The newest Catalyst 9000 series does not support ISL at all. If you encounter ISL in the wild, you are looking at a legacy network that needs migration. [Configuring 802.1Q Trunks on Cisco Catalyst Switches](https://www.pinglabz.com/cisco-8021q-trunking-lab-guide/) walks through the modern trunk pattern. ## MTU Implications: Why You See 1522-Byte Frames The 4-byte tag adds to frame size. A standard Ethernet frame is 64-1518 bytes (without FCS); with 802.1Q tagging, the maximum becomes 1522 bytes. That extra 4 bytes is enough that older equipment without "baby giant" support will silently drop tagged frames at maximum size. Modern Cisco switches accept 1522-byte (or larger) tagged frames by default; you usually do not have to configure anything. The exception is jumbo-frame deployments where the IP MTU is set to 9000, and you might need to confirm the system MTU accommodates 9000 + 4 (tag) bytes. Check with `show system mtu` on Catalyst. ## Verifying Tagged Frames in Practice To see the 802.1Q tag in action on a Cisco switch, the easiest way is to mirror a trunk port to a span session and capture with Wireshark. The 802.1Q tag will show as a separate "802.1Q Virtual LAN" layer between the Ethernet header and the inner protocol. For a quick CLI sanity check on what VLANs a trunk is carrying: ``` Switch# show interfaces trunk Port Mode Encapsulation Status Native vlan Gi1/0/24 on 802.1q trunking 999 Port Vlans allowed on trunk Gi1/0/24 10,20,30,999 ``` Encapsulation should always show `802.1q` on modern switches. If you see `isl`, you have a legacy device or misconfiguration. ## Summary The 802.1Q tag is 4 bytes inserted between Src MAC and EtherType: TPID (2 bytes, always 0x8100) plus TCI (2 bytes split into 3-bit PCP, 1-bit DEI, and 12-bit VLAN ID). The PCP carries QoS priority, the VID identifies the VLAN, and the native VLAN traverses the trunk untagged for backwards compatibility. If you remember nothing else: TPID is always 0x8100, the VLAN ID is 12 bits with practical range 1-4094, and the native VLAN should never be VLAN 1 in a production deployment. Bookmark this article as your byte-by-byte reference, and see the [VLAN cluster pillar](https://www.pinglabz.com/vlans-layer-2-switching/) for the operational picture. ### Rapid Spanning Tree Protocol (RSTP): What Changed from 802.1D STP URL: https://www.pinglabz.com/rapid-spanning-tree-protocol-rstp/ Last updated: 2026-06-13T20:08:46.000Z Rapid Spanning Tree Protocol (RSTP, IEEE 802.1w) is the 2001 update to classic 802.1D Spanning Tree. Every modern Cisco campus runs Rapid PVST+ (Cisco's per-VLAN flavor of RSTP), and the reason is simple: classic STP takes 30-50 seconds to converge after a topology change, RSTP takes a fraction of a second. If you have ever waited for a switch to pass DHCP at boot, that 30-second wait was 802.1D port states burning Forward Delay timers; RSTP made that obsolete. This article is the deep-dive on what changed between 802.1D and 802.1w, why those changes produce sub-second convergence, and how to deploy Rapid PVST+ on Cisco IOS XE without tripping over the few legacy gotchas. If you are studying for CCNP, designing a campus, or migrating from PVST+ to Rapid PVST+, this is what you need to know. ## Why Classic STP Was Too Slow Classic 802.1D was designed in 1990, when Ethernet ran at 10 Mbps and applications were tolerant of brief outages. Its convergence model uses three timers (Hello at 2s, Forward Delay at 15s, Max Age at 20s) and relies on those timers expiring before a port can transition through the Listening, Learning, and Forwarding states. For a port to go from Blocking to Forwarding after a topology change, it must wait Max Age (20s) for the old BPDU to age out, then Forward Delay (15s) in Listening, then Forward Delay (15s) in Learning. That is 50 seconds. Even ports that come up fresh (no Max Age wait) need 30 seconds for the two Forward Delay states. By the early 2000s, with VoIP, real-time apps, and gigabit Ethernet, that was no longer acceptable. The fix was not to shorten the timers. The fix was to redesign the protocol so it does not need them in the first place. ## Key Changes in 802.1w RSTP keeps the same core concept (one spanning tree, one root bridge, one root port per non-root switch, one designated port per segment) but rebuilds the convergence machinery around explicit handshaking instead of timers: Port states 802.1D STP 5 (Disabled, Blocking, Listening, Learning, Forwarding) 802.1w RSTP 3 (Discarding, Learning, Forwarding) Port roles 802.1D STP 3 (Root, Designated, Non-Designated) 802.1w RSTP 5 (Root, Designated, Alternate, Backup, Disabled) Topology change 802.1D STP TCN BPDU floods to root, root sets TC bit, all switches flush MAC table partially 802.1w RSTP Originating switch floods TC to all neighbors directly; faster MAC flushing BPDU origination 802.1D STP Only the root originates BPDUs; non-root relays 802.1w RSTP Every switch originates BPDUs every Hello BPDU loss tolerance 802.1D STP Max Age (20s = 10 missed BPDUs at 2s Hello) 802.1w RSTP 3 missed BPDUs (6s) before declaring neighbor down Link types 802.1D STP Implicit; not used for fast convergence 802.1w RSTP Explicit (point-to-point, shared, edge); enables proposal/agreement Proposal/Agreement handshake 802.1D STPNone 802.1w RSTP On point-to-point links; achieves sub-second transition Convergence after direct failure 802.1D STP30-50 seconds 802.1w RSTP Sub-second (typically 1-2 BPDU exchanges) PortFast equivalent 802.1D STP Cisco extension; not part of standard 802.1w RSTP Edge port type; built into the standard ## From Five States to Three RSTP collapses the classic state machine. The five 802.1D states map to three RSTP states: Discarding Maps to 802.1D Disabled, Blocking, Listening Forwards data?No Learns MACs?No Learning Maps to 802.1DLearning Forwards data?No Learns MACs?Yes Forwarding Maps to 802.1DForwarding Forwards data?Yes Learns MACs?Yes The collapsed states matter because under RSTP, ports do not generally walk through them on a timer. They jump directly from Discarding to Forwarding via the proposal/agreement handshake, which takes one BPDU round-trip on a point-to-point link. ## Two New Port Roles: Alternate and Backup Classic STP only has Root and Designated; everything else is "Non-Designated" and Blocking. RSTP elevates two specific Non-Designated cases to first-class roles: - **Alternate Port (AP).** An immediate backup to the Root Port. If the Root Port goes down, the Alternate is pre-computed and ready to take over instantly. The classic equivalent would have re-run the entire root election process. - **Backup Port (BP).** An immediate backup to a Designated Port on the same shared segment. Only exists on shared (hub) segments, which barely exist in modern networks. The Alternate Port is the headline upgrade. It is what makes RSTP feel "fast": every non-root switch already has a pre-determined replacement Root Port standing by, in Discarding state, ready to forward as soon as the current Root Port fails. ## Proposal/Agreement: Why RSTP is Sub-Second The signature RSTP mechanism is the proposal/agreement handshake on point-to-point links. Here is the dance: 1. A new link comes up between two switches that already have RSTP elsewhere. 2. The switch that should be the Designated end sends a BPDU with the proposal flag set. 3. The other switch receives it, recognizes the proposal, places all its other Designated ports into a brief "sync" state (preventing temporary loops), and sends back an agreement BPDU. 4. The Designated switch immediately transitions the new port to Forwarding. 5. The whole exchange takes one BPDU round-trip, typically a few milliseconds. Compare to 802.1D, which would have waited 30-50 seconds for the same transition. The handshake only works on point-to-point links, which is why RSTP introduces explicit link types. ## Link Types: Point-to-Point, Shared, Edge Point-to-Point How RSTP uses it Eligible for proposal/agreement; sub-second transition Cisco command `spanning-tree link-type point-to-point` Shared How RSTP uses it No proposal/agreement; falls back to 802.1D-style timers (rare in modern networks) Cisco command `spanning-tree link-type shared` Edge How RSTP uses it Connects to a host, never to another switch; immediately Forwarding (PortFast equivalent) Cisco command`spanning-tree portfast` Cisco IOS XE auto-detects link types from interface duplex (full-duplex = point-to-point, half-duplex = shared). In modern campus deployments every switch-to-switch link is full-duplex, so the auto-detection just works. The exception is when you tunnel STP through a non-RSTP medium; force `link-type point-to-point` in that case. Edge ports (host-facing) get explicit treatment via PortFast, which preserves the same fast-transition behavior as classic STP's PortFast extension. RSTP integrates it as part of the standard rather than a Cisco add-on. ## BPDU Origination: Now Every Switch In 802.1D, only the root bridge originates BPDUs and every other switch relays them downstream every Hello. This made non-root BPDUs implicit acknowledgements: as long as you keep relaying, the upstream is fine. RSTP changes this. Every switch originates its own BPDU every Hello (default 2s) regardless of whether the root has sent something. This is what enables the 3-missed-BPDU rule: if you do not hear from your direct neighbor for 6 seconds (3 x 2s Hello), you consider them down. Compare to 802.1D's Max Age of 20 seconds. This change has another implication: RSTP can detect upstream link failures faster than 802.1D, because every switch is independently reporting up the tree. ## Topology Change: Faster MAC Flushing When the topology changes, MAC address tables on all switches need to be updated. Otherwise frames going to a host whose path has changed will be sent to the wrong segment. Classic STP handles this with TCN BPDUs that walk all the way to the root and back. The root sets the TC bit in regular BPDUs, and every switch then ages out MAC entries faster (15s instead of 5 minutes default). It is slow. RSTP handles it directly. The switch that detected the topology change floods TC BPDUs to all its neighbors immediately, which flush their MAC tables for affected interfaces and propagate the TC. The result: MAC tables converge on the new topology in seconds, not in 15-second-per-hop cascades. ## Cisco's Rapid PVST+: RSTP Per VLAN The IEEE 802.1w standard runs one spanning tree across the whole switching domain (CST). Cisco's PVST+ runs one STP instance per VLAN, which lets you load-balance traffic by making different switches root for different VLANs. Rapid PVST+ is the same per-VLAN model, but each instance is RSTP rather than 802.1D. The trade-off is overhead: 100 VLANs means 100 STP instances, 100 sets of BPDUs every 2 seconds, and 100 SPF calculations on every topology change. For most campus networks this is fine. For very large networks (thousands of VLANs), MST is the answer instead. [MST configuration](https://www.pinglabz.com/configure-mst-cisco-switches/) is in a separate article. ## Enabling Rapid PVST+ on Cisco IOS XE One global command: ``` Switch(config)# spanning-tree mode rapid-pvst ``` That is it. Cisco's default has been Rapid PVST+ for over a decade, but if you are inheriting older configs always check. Verify: ``` Switch# show spanning-tree summary | include mode Switch is in rapid-pvst mode ``` For interface settings to ensure fast convergence: ``` Switch(config)# interface GigabitEthernet1/0/1 Switch(config-if)# spanning-tree portfast ! Edge port type Switch(config-if)# spanning-tree bpduguard enable ! Protect edge port Switch(config)# interface GigabitEthernet1/0/24 Switch(config-if)# spanning-tree link-type point-to-point ``` The full configuration walkthrough is in [Configuring Rapid PVST+ on Cisco Catalyst Switches](https://www.pinglabz.com/configure-rapid-pvst-cisco/). ## Convergence Numbers: What to Expect Direct link failure on Root Port (Alternate available) 802.1D / PVST+30-50 seconds 802.1w / Rapid PVST+ Sub-second (Alternate immediately becomes Root) Indirect link failure (BPDU loss) 802.1D / PVST+ 50 seconds (Max Age + 2 x Forward Delay) 802.1w / Rapid PVST+ \~6 seconds (3 missed BPDUs) + sub-second transition New port up between two switches 802.1D / PVST+30 seconds 802.1w / Rapid PVST+ Sub-second via proposal/agreement Edge port (host) up 802.1D / PVST+ 30 seconds without PortFast; near-instant with PortFast 802.1w / Rapid PVST+ Near-instant by default The big win is direct-failure convergence on point-to-point links, which is the most common topology change in modern campuses. ## Backwards Compatibility with 802.1D RSTP is fully backwards compatible. If a switch running RSTP receives 802.1D-format BPDUs (no flags, no proposal/agreement), it falls back to legacy mode for that specific port and uses 802.1D-style timer-based transitions. The rest of the network continues to use RSTP normally. This means you can mix 802.1D and 802.1w switches during a migration. The downside: any port that falls back to legacy mode does not benefit from sub-second convergence on that link. Migrate the whole topology to RSTP-capable equipment to get the full benefit. ## When RSTP Convergence Falls Back to Slow RSTP only achieves sub-second convergence under specific conditions. If any are missing, you fall back closer to 802.1D timers: - **Link must be point-to-point.** Half-duplex or shared link types disable proposal/agreement. - **Both ends must be RSTP-capable.** A neighbor running 802.1D forces fallback on that link. - **BPDUs must not be filtered.** Aggressive BPDU Filter on a non-edge port breaks RSTP entirely. - **Hardware must process BPDUs in time.** A switch with high CPU may delay BPDU origination/processing past the 3-Hello boundary, triggering false timeouts. The full convergence troubleshooting walkthrough is in [Troubleshooting STP Convergence Problems and Slow Failover](https://www.pinglabz.com/troubleshoot-stp-convergence-slow-failover/). ## Summary RSTP (802.1w) is what made spanning tree usable for modern enterprise networks. The five state-machine changes (collapsed states, new port roles, proposal/agreement, faster topology change, every-switch BPDU origination) work together to convert classic STP's 30-50 second convergence into sub-second behavior on healthy point-to-point links. Every modern Cisco campus should run Rapid PVST+ or MST, never classic 802.1D / PVST+. The configuration is one global command. The hardening (PortFast on edge ports, BPDU Guard with PortFast, Root Guard at distribution-facing-access, Loop Guard on point-to-point trunks) is in the [Spanning Tree Protocol pillar](https://www.pinglabz.com/spanning-tree-protocol/). If you are seeing 30-second failover in a network that runs Rapid PVST+, something is forcing fallback to 802.1D and that is the first thing to investigate. ### BGP vs OSPF: When to Use Each Routing Protocol URL: https://www.pinglabz.com/bgp-vs-ospf/ Last updated: 2026-07-04T23:21:09.000Z BGP and OSPF are the two routing protocols every working network engineer ends up running, and they are the two most CCNP and CCIE candidates compare without fully understanding why they exist as separate things in the first place. The "BGP vs OSPF" framing makes it sound like a choice. It is not. They solve different problems and you almost always run both. This article walks through what each protocol is actually for, the technical differences that matter in production, when each is the right answer, and the surprisingly common cases where the answer is "both, layered." If you are studying for a certification, designing a new network, or trying to explain to a manager why the answer to "should we use BGP" depends on the question, this is the comparison. ## The TL;DR OSPF is an interior gateway protocol (IGP). It finds the best path between routers inside a single administrative domain, fast, using a metric you do not get to argue with. BGP is an exterior gateway protocol (EGP). It exchanges reachability between administrative domains using a policy you have full control over, slowly and on purpose. If you control the routers on both sides of every link, you want OSPF (or another IGP). If the routers on the other side of a link belong to someone else, you want BGP. Most production networks have both: OSPF inside, BGP at the edges, with OSPF providing the underlying reachability that lets BGP sessions stay up. ## Protocol Type: Link-State vs Path-Vector OSPF is link-state. Every router floods its view of the local topology to every other router in the area, every router builds the same map, and every router runs Dijkstra's shortest-path algorithm against that map independently. The result is a network where every router has the same picture of the topology and they all converge on the same answer. BGP is path-vector. Every BGP route carries the list of autonomous systems (the AS\_PATH) it has traversed. Routers do not exchange topology, they exchange routes plus path metadata. The "vector" is the AS-level path; the "path" is the explicit history that prevents loops (if your own AS appears in the AS\_PATH of an incoming route, reject it). The implications matter: - **OSPF reveals topology to all participants.** That is fine inside one organization. It is not fine across organizational boundaries; you do not want to expose your internal topology to a peer ISP. - **BGP hides topology.** Each AS only sees the AS-level path, not the routers inside other ASes. This is a feature, not a bug. - **OSPF converges fast.** Sub-second with tuning, because every router already has the data it needs to recompute. - **BGP converges slowly.** Tens of seconds or more in the default-free zone, because route propagation must walk the AS path and policy is evaluated at every hop. ## The Side-by-Side Comparison Type OSPF Interior gateway protocol (IGP) BGP Exterior gateway protocol (EGP) Algorithm OSPF Link-state (Dijkstra SPF) BGPPath-vector Standard OSPF RFC 2328 (OSPFv2), RFC 5340 (OSPFv3) BGPRFC 4271 (BGP-4) Default AD (Cisco) OSPF110 BGP20 (eBGP) / 200 (iBGP) Transport OSPF IP protocol 89, multicast 224.0.0.5/6 BGPTCP port 179 Convergence OSPF Sub-second with BFD; seconds without BGPSeconds to minutes Metric OSPF Cost (16-bit, bandwidth-derived) BGP 13-step best-path algorithm with attributes Policy expressiveness OSPF Limited (passive-interface, area-type, summarization, redistribution) BGP Extreme (route-maps, communities, MED, LP, weight, AS-path) Hierarchical structure OSPF Areas (strict rules, must connect to area 0) BGP None native; route reflectors and confederations Typical scale OSPF 1,000-10,000 routes per area BGP Up to 1,000,000+ in the DFZ Authentication OSPFMD5, SHA BGPMD5, TCP-AO, GTSM Topology visibility OSPFFull within an area BGP None across AS boundaries (AS\_PATH only) Trust assumption OSPFAll routers trusted BGP Routers in other ASes are not trusted ## When OSPF is the Answer OSPF is the right choice when: - **You control every router on every link.** This is the defining condition for an IGP. - **You need fast convergence.** Sub-second failover is achievable with default timers + BFD. - **You want vendor neutrality.** RFC 2328 OSPFv2 is genuinely interoperable across Cisco, Juniper, Arista, Nokia, etc. - **You need hierarchical scaling.** Areas let you scale to thousands of routers with bounded LSDB size and SPF run time. - **You want a metric that "just works."** Cost is bandwidth-derived. Set the reference bandwidth once, and the protocol picks reasonable paths automatically. The full OSPF picture, including LSA types, areas, and configuration, is in [OSPF Complete Guide](https://www.pinglabz.com/ospf/). ## When BGP is the Answer BGP is the right choice when: - **You are exchanging routes with another organization.** Internet peering, IPsec VPN to a partner, multi-cloud peering. The other side does not trust your IGP and will not run one with you. - **You need to express explicit policy.** "Prefer transit X for outbound except for these prefixes." OSPF cannot say that. BGP can. - **You are at internet scale.** The default-free zone has more than a million prefixes; OSPF cannot carry that and would not converge if it tried. - **You need topology hiding.** The other side gets to see your AS path, not your routers. - **You need to differentiate routes by source.** BGP communities tag routes for downstream policy in ways that have no IGP equivalent. And BGP shows up in places that surprise people: - **Data center underlays.** Modern leaf-spine fabrics often run BGP all the way down to the leaf, often unnumbered, often with EVPN on top. Why? Because BGP carries policy better than any IGP, and you want explicit policy in a multi-tenant fabric. - **SD-WAN overlays.** The control plane between vManage / vSmart / WAN edges is BGP-flavored. - **Cloud VPC peering.** AWS, Azure, GCP all use BGP for VPN gateways and inter-VPC routing. The full BGP picture, including the 13-step best-path algorithm and attribute reference, is in [BGP (Border Gateway Protocol): The Complete Guide](https://www.pinglabz.com/bgp/). ## When You Run Both (Most of the Time) The dominant production pattern is OSPF as the underlying IGP, BGP at the edges. Three reasons: 1. **iBGP needs reachability.** Internal BGP peers form TCP/179 sessions, often between loopbacks across the network. Those TCP sessions only stay up if there is a route between the loopbacks. That route comes from OSPF. Without OSPF, iBGP collapses on the first link flap that the BGP timers cannot survive. 2. **BGP NEXT\_HOPs must resolve.** When an eBGP peer hands you a route, the NEXT\_HOP attribute is the eBGP peer's IP. Your iBGP peers receive that route, but the NEXT\_HOP is unchanged. They need a route to it (OSPF provides this) or the route is unusable. `next-hop-self` on the edge router is the alternative pattern. 3. **Different jobs.** Inside the AS, OSPF moves traffic fast and adapts to topology changes. At the AS edges, BGP enforces policy and exchanges routes with peers. The architectural picture for a typical enterprise: ``` ISP-A ISP-B \ / eBGP eBGP \ / +---+---+ | edge | <-- BGP, edge router holds DFZ partial table +---+---+ | iBGP <-- iBGP from edge to internal next-hop routers | +---+---+ | core | <-- OSPF, fast intra-AS reachability +---+---+ / | \ access distribution layers (OSPF) ``` OSPF carries internal prefixes (loopbacks, point-to-point links, internal user subnets). BGP carries the internet table and any inter-org routes. The two never overlap on the same prefix; redistribution between them is rare and should be done with extreme care if at all. ## Convergence and Performance: The Numbers Benchmark intuition for similar topologies: Direct link failure (interface down) OSPF (default)1-2s OSPF + BFD50-200ms BGP (default)3-5s BGP + BFD50-300ms Indirect link failure OSPF (default) 40s (Hello/Dead default) OSPF + BFD50-200ms BGP (default) 180s (Hold-time default) BGP + BFD50-300ms Route added to a healthy network OSPF (default)<1s OSPF + BFD<1s BGP (default) 1-30s (depends on table size) BGP + BFD1-30s Whole-table reload OSPF (default)n/a OSPF + BFDn/a BGP (default)2-5 minutes for DFZ BGP + BFD2-5 minutes for DFZ The defaults are conservative. With BFD on every link, both protocols converge in similar timeframes for direct link failures. The big BGP cost is full-table operations: a soft-reset of an eBGP DFZ session takes minutes. OSPF has no comparable cost; an LSA refresh is small. ## Redistribution Between BGP and OSPF The temptation is to redistribute BGP into OSPF (so internal hosts can reach the internet via OSPF default routes) and OSPF into BGP (so external peers can reach your internal services). Both have failure modes: - **BGP into OSPF.** Pulling 1,000,000 internet routes into your OSPF LSDB will explode it. SPF runs will pin the CPU and the network will hang. The right pattern: do not redistribute. Originate a default route in OSPF instead, pointing to the BGP-running edge. - **OSPF into BGP.** Naively redistributing all OSPF routes into BGP will flood your peers with internal routes. Filter aggressively (route maps with prefix lists). Most environments only redistribute statics into BGP and use `network` statements to advertise specific prefixes. For inbound from BGP, a sane pattern is "default route originate" in OSPF on the edge router. Combined with iBGP carrying the full table to the edge, internal hosts find the edge by default route, and the edge has the actual table to forward. ## Design Decisions: Picking the Combination Three real-world scenarios: **Single-site enterprise with one ISP.** OSPF inside, single-homed to the ISP via static default. No BGP. The ISP has a static route back. This is fine for hundreds of single-site businesses. Add BGP only when you need the policy. **Multi-site enterprise with two ISPs.** OSPF inside, BGP at every site that has internet connectivity. eBGP to each ISP, iBGP across the WAN to share inbound preferences. This is where the real BGP work begins: setting LOCAL\_PREF for outbound preferences, AS-Path Prepending or MED for inbound preferences, communities for tagging. **Service provider or large enterprise data center.** OSPF or IS-IS in the underlay (IS-IS is preferred at internet scale because of its layered hierarchy). iBGP with route reflectors as the overlay carrying customer routes, MPLS labels, EVPN, etc. The BGP/IGP split is even more pronounced here: BGP carries everything that needs to be policy-controlled; the IGP carries only loopbacks and link infrastructure. ## Choosing the IGP: OSPF, EIGRP, or IS-IS If you have decided on an IGP, the choice between OSPF, EIGRP, and IS-IS is its own discussion: - **OSPF.** Default enterprise choice. Vendor-neutral, well-understood, certified everywhere. - **EIGRP.** Faster convergence on small networks (DUAL has local re-convergence). Open since 2013 but practical adoption beyond Cisco is rare. - **IS-IS.** Service-provider favorite. Two-level hierarchy (L1/L2) is cleaner than OSPF areas at very large scale, and it carries CLNS instead of IP so misconfiguration cannot blackhole production traffic. Enterprise adoption is rare but growing in modern data centers. For most enterprises, OSPF is the safe default. For internet-scale providers, IS-IS often wins. For Cisco-only networks where convergence speed matters, EIGRP is still a valid choice. ## Summary BGP and OSPF are not alternatives. They are complementary protocols that solve different problems at different layers of a real network. OSPF gives you fast, automatic, intra-AS reachability with a metric you do not have to think about. BGP gives you policy-driven, inter-AS reachability with explicit control over every aspect of route selection. The right answer is almost always "both": OSPF inside, BGP at the edges, with the IGP carrying internal infrastructure and BGP carrying everything that needs policy. If you are studying for a CCNP or CCIE, expect to know both cold; if you are designing a network, expect to deploy both. Bookmark the [OSPF pillar](https://www.pinglabz.com/ospf/) and the [BGP pillar](https://www.pinglabz.com/bgp/) for the full configuration and operational details. ### References - [RFC 4271 - A Border Gateway Protocol 4 (BGP-4)](https://www.rfc-editor.org/rfc/rfc4271?ref=pinglabz.com) - [Cisco BGP technology documentation](https://www.cisco.com/c/en/us/tech/ip/border-gateway-protocol-bgp/index.html?ref=pinglabz.com) Take the BGP reference with you The free BGP field-reference PDF: path attributes, best-path order, and the show commands that matter. Delivered by email, no card required. [Get the free PDF](https://www.pinglabz.com/bgp-cheatsheet/) ### Cisco ISE 802.1x Wired Configuration: A Practical Step-by-Step Guide URL: https://www.pinglabz.com/cisco-ise-802-1x-wired-configuration-guide/ Last updated: 2026-07-04T23:21:24.000Z Deploying wired 802.1x with Cisco ISE is one of those tasks that looks straightforward on paper and bites you in the lab the first time you try it. The pieces are simple individually - a switch, a RADIUS server, a supplicant - but the sequence of clicks in ISE, the correct AAA method lists on the switch, and the order of authentication methods at the interface level all have to line up. This guide walks you through a complete, working **Cisco ISE 802.1x wired configuration** end to end, from building the Policy Set in ISE to typing the final `dot1x pae authenticator` on the access port, with the verification commands you will actually run when something breaks. (This article is part of the PingLabz 802.1X series - the [full 802.1X guide](https://www.pinglabz.com/802-1x/) maps the whole cluster in reading order.) ## What This Guide Covers This is a wired 802.1x deployment guide. The authenticator is a Cisco Catalyst switch running IOS-XE, the authentication server is Cisco ISE acting as a RADIUS server, and the supplicant is a domain-joined Windows or macOS endpoint (or Linux with wpa\_supplicant in wired mode). Wireless 802.1x, dot1x on trunk ports, and pure MAB-only deployments are out of scope - MAB appears here only as a *fallback* after 802.1x times out, which is how most real enterprises actually run the configuration. You will finish this guide with a lab-validated configuration that authenticates a supplicant via PEAP or EAP-TLS, matches an Authorization Policy in ISE, returns a VLAN or downloadable ACL, and transitions the port to the authorized state. For deeper coverage of specific related topics, see the companion articles on [MAB configuration](https://www.pinglabz.com/mab-configuration-cisco-ios-xe-ise/) and [Guest VLAN, Auth-Fail VLAN, and Critical VLAN](https://www.pinglabz.com/802-1x-guest-vlan-auth-fail-vlan-critical-vlan/) behavior. ## Prerequisites Before you start clicking or typing, confirm the versions and licensing lined up in the table below. Mixing IOS-XE trains or running an unlicensed ISE deployment will cause the config to accept but not behave as expected. Cisco ISE Minimum Version 3.1 Patch 6 (or 3.2/3.3) Notes Earlier 2.x releases work but the Policy Set UI differs and some screens referenced here will look different. Catalyst Switch Minimum Version 9200/9300/9400/9500 with IOS-XE 17.6+ Notes 3650/3850 on IOS-XE 16.12 also work with identical CLI. IOS classic switches use older legacy dot1x syntax - not covered here. ISE License Minimum Version Essentials (formerly Base) Notes Essentials covers 802.1x, MAB, and basic authorization. Advantage is only needed for profiling, posture, or TrustSec. Switch License Minimum VersionNetwork Advantage Notes Network Essentials supports dot1x, but Advantage is what most enterprises run. Supplicant Minimum Version Windows Wired AutoConfig service or macOS native supplicant Notes The Windows "Wired AutoConfig" service is disabled by default - enable it via `services.msc` or GPO. Your topology assumption is simple: the Catalyst switch has IP reachability to the ISE Policy Service Node (PSN) on UDP/1812 (auth), UDP/1813 (accounting), and UDP/1700 (Change of Authorization). NTP must be synchronized between the switch, ISE, and any certificate authorities, otherwise EAP-TLS will fail silently on expired or not-yet-valid timestamps. DNS should resolve ISE's FQDN - you will use the FQDN, not the IP, when generating the EAP certificate. ## ISE Configuration ISE configuration breaks into five ordered steps: add the switch as a Network Device, build or reuse an Identity Source, create an Authorization Profile, assemble the Policy Set, and confirm certificate trust. Do them in this order - skipping ahead means you will hit "RADIUS request rejected" errors before you have anything meaningful to debug. ### Step 1: Add the switch as a Network Device Navigate to **Administration > Network Resources > Network Devices** and click *Add*. Fill in the name (use the switch hostname for sanity when reading Live Logs later), the management IP (this must be the source IP the switch uses when sending RADIUS packets - usually the SVI of the management VLAN), and set the device profile to *Cisco*. Expand **RADIUS Authentication Settings** and enter a shared secret (this must match the shared secret configured on the switch exactly - copy-paste, do not retype). Tick the *CoA Port* box and leave it at the default of 1700, because without CoA enabled you will not be able to push Change of Authorization from ISE, which breaks posture remediation and dynamic re-auth. ### Step 2: Configure the Identity Source For most enterprises, the identity source is Active Directory. Navigate to **Administration > Identity Management > External Identity Sources > Active Directory**, join ISE to the domain, and verify the join status turns green on every PSN (not just the PAN). Create an Identity Source Sequence under **Administration > Identity Management > Identity Source Sequences** \- order matters here, with AD first and the internal ISE user store as a fallback for service accounts or break-glass users. ### Step 3: Create the Authorization Profile The Authorization Profile is the set of RADIUS attributes ISE returns when an endpoint matches your policy. Go to **Policy > Policy Elements > Results > Authorization > Authorization Profiles** and click *Add*. The table below shows the attributes that matter for a typical wired corporate deployment. Access Type ValueACCESS\_ACCEPT Why it matters Without this the switch gets a reject even if the rest of the profile is correct. VLAN Value Tag ID = 20, Name = CORP\_DATA Why it matters Drives dynamic VLAN assignment via `Tunnel-Type`, `Tunnel-Medium-Type`, and `Tunnel-Private-Group-ID`. DACL Name Value PERMIT\_ALL\_TRAFFIC (or your own) Why it matters Downloadable ACL applied to the session; required if you want per-user ACLs without pre-staging them on every switch. Reauthentication Timer Value3600 seconds Why it matters Forces fresh authentication hourly, catching credential changes or revoked certificates without waiting for CoA. ### Step 4: Build the Policy Set Navigate to **Policy > Policy Sets** and create a new Policy Set named *Wired\_Dot1X*. The Policy Set condition itself is what limits the rules inside from matching wireless or guest traffic, so use this condition: `Wired_802.1X` (a built-in compound condition that checks NAS-Port-Type and Service-Type). Inside the Policy Set, configure two sub-sections. The **Authentication Policy** decides which identity store to query based on the EAP method. A minimal working policy is shown below. Dot1X\_EAP\_TLS Condition Network Access:EapAuthentication EQUALS EAP-TLS Allowed ProtocolsDefault Network Access Identity Source Certificate Authentication Profile (CAP) > Active Directory Dot1X\_PEAP Condition Network Access:EapAuthentication EQUALS EAP-MSCHAPv2 Allowed ProtocolsDefault Network Access Identity Source AD\_Sequence (your identity source sequence) MAB\_Fallback ConditionWired\_MAB Allowed ProtocolsDefault Network Access Identity SourceInternal Endpoints The **Authorization Policy** is where you map authenticated identities to the Authorization Profile created earlier. A working minimum looks like the rules below, matched top-down. Corporate\_Users Condition AD:ExternalGroups EQUALS Domain Users Profile CORP\_VLAN\_20\_PERMIT\_ALL Corporate\_Computers Condition AD:ExternalGroups EQUALS Domain Computers Profile CORP\_VLAN\_20\_PERMIT\_ALL MAB\_Printers Condition IdentityGroup:Name EQUALS Printers ProfilePRINTER\_VLAN\_30 Default Condition(catch-all) ProfileDenyAccess ### Step 5: Verify certificate trust For PEAP, ISE presents its EAP certificate to the supplicant. For the supplicant to trust it, the issuing CA must be in the supplicant's Trusted Root Certification Authorities store. For EAP-TLS, the reverse also matters - ISE must trust the CA that issued the client certificate, so import that root/intermediate into **Administration > System > Certificates > Trusted Certificates** with the *Trust for client authentication* checkbox enabled. Skipping this causes the infuriating "EAP-TLS failed SSL/TLS handshake" error that looks like a client issue but is actually ISE not trusting the client cert. ## Cisco Switch Configuration The switch side is where most of the nuance lives, because IOS-XE supports two authentication frameworks: the legacy *authentication*\-style commands (IBNS 1.0) and the newer *policy-map*\-based *service-policy* style (IBNS 2.0). Cisco recommends IBNS 2.0 on all new deployments - it is more flexible, supports event-driven logic, and is what TAC assumes you are running on 17.x code. The configuration below uses IBNS 2.0. ### Global AAA and RADIUS Start with `aaa new-model`, then define the RADIUS servers as *named* servers (not the older `radius-server host` syntax, which is deprecated). The order of the `aaa authentication dot1x` method list and the matching `aaa authorization network` list must both reference the same named server group, or the authenticated session will complete but fail to apply the authorization attributes. ``` aaa new-model radius server ISE-PSN-01 address ipv4 10.10.10.11 auth-port 1812 acct-port 1813 automate-tester username probe-user ignore-acct-port probe-on key 7 0822455D0A16544541 ! radius server ISE-PSN-02 address ipv4 10.10.10.12 auth-port 1812 acct-port 1813 automate-tester username probe-user ignore-acct-port probe-on key 7 0822455D0A16544541 ! aaa group server radius ISE_RADIUS server name ISE-PSN-01 server name ISE-PSN-02 deadtime 15 ip radius source-interface Vlan100 ! aaa authentication dot1x default group ISE_RADIUS aaa authorization network default group ISE_RADIUS aaa accounting dot1x default start-stop group ISE_RADIUS aaa accounting update newinfo periodic 2880 ! aaa server radius dynamic-author client 10.10.10.11 server-key 7 0822455D0A16544541 client 10.10.10.12 server-key 7 0822455D0A16544541 auth-type any ! radius-server attribute 6 on-for-login-auth radius-server attribute 8 include-in-access-req radius-server attribute 25 access-request include radius-server dead-criteria time 10 tries 3 radius-server deadtime 15 radius-server vsa send authentication radius-server vsa send accounting ! dot1x system-auth-control dot1x critical eapol authentication critical recovery delay 2000 ``` A few of those lines are easy to skip past but directly cause pain if you omit them. `automate-tester` is what drives the switch's dead-server detection (without it, the switch will not mark ISE dead cleanly and retries stack up). `ip radius source-interface` must match the IP address you configured in ISE under Network Devices - if it does not, ISE rejects the packet with "unknown NAD" and nothing shows up in Live Logs. `radius-server attribute 25 access-request include` tells the switch to include the Class attribute on reauth, which ISE needs for session state continuity. ### IBNS 2.0 Policy Map and Interface Configuration IBNS 2.0 uses a *policy-map* of type *control subscriber* that reacts to events (session-started, authentication-failure, authentication-success) and executes *actions* (authenticate using method, authorize, terminate). The block below is a production-grade template that runs dot1x first, falls back to MAB on timeout, and applies the Auth-Fail and Critical VLANs correctly. ``` class-map type control subscriber match-all DOT1X match method dot1x ! class-map type control subscriber match-all DOT1X_FAILED match method dot1x match result-type method dot1x authoritative ! class-map type control subscriber match-all MAB match method mab ! class-map type control subscriber match-all MAB_FAILED match method mab match result-type method mab authoritative ! class-map type control subscriber match-all AAA_SVR_DOWN_AUTHD_HOST match authorization-status authorized match result-type aaa-timeout ! class-map type control subscriber match-all AAA_SVR_DOWN_UNAUTHD_HOST match authorization-status unauthorized match result-type aaa-timeout ! policy-map type control subscriber DOT1X_MAB_POLICY event session-started match-all 10 class always do-until-failure 10 authenticate using dot1x priority 10 event authentication-failure match-first 5 class DOT1X_FAILED do-until-failure 10 terminate dot1x 20 authenticate using mab priority 20 10 class MAB_FAILED do-until-failure 10 terminate mab 20 authentication-restart 60 20 class AAA_SVR_DOWN_UNAUTHD_HOST do-until-failure 10 activate service-template CRITICAL_AUTH_ACCESS 20 authorize 30 pause reauthentication 30 class AAA_SVR_DOWN_AUTHD_HOST do-until-failure 10 pause reauthentication 20 authorize event aaa-available match-all 10 class IN_CRITICAL_AUTH do-until-failure 10 clear-session 20 class NOT_IN_CRITICAL_AUTH do-until-failure 10 resume reauthentication event agent-found match-all 10 class always do-until-failure 10 terminate mab 20 authenticate using dot1x priority 10 ! service-template CRITICAL_AUTH_ACCESS vlan 999 access-group ACL-CRITICAL-AUTH ``` With the policy built, the access port configuration becomes short and predictable. Every 802.1x access port uses the same block below - this is what makes IBNS 2.0 worth the up-front complexity: you never touch the interface again when you change policy, because policy changes happen in the policy-map. ``` interface GigabitEthernet1/0/1 description 802.1X Access Port switchport mode access switchport access vlan 10 switchport voice vlan 110 access-session host-mode multi-auth access-session closed access-session port-control auto authentication periodic authentication timer reauthenticate server mab dot1x pae authenticator dot1x timeout tx-period 7 dot1x max-reauth-req 2 spanning-tree portfast spanning-tree bpduguard enable service-policy type control subscriber DOT1X_MAB_POLICY ``` `access-session closed` means the port starts in closed mode - no traffic passes until authentication succeeds. For a phased rollout you would use `access-session closed` only in the final phase, running `access-session monitor` during early phases so failures are logged but traffic is not blocked. The `authentication timer reauthenticate server` line tells the switch to use the reauthentication timer ISE sends in the Session-Timeout attribute, which lets you change reauth cadence centrally from ISE without touching switches. ## Testing and Verification Plug a configured supplicant into `Gi1/0/1`, give it 10 seconds, and then run the commands below. If any of them show unexpected state, skip ahead to the failure scenarios section - do not re-enter configuration blindly. ### show dot1x all This is the global dot1x sanity check. You should see `Sysauthcontrol: Enabled` and the interface listed with `PortControl: Auto`. If `Sysauthcontrol` shows Disabled, you missed `dot1x system-auth-control` in global config. ``` switch# show dot1x all Sysauthcontrol Enabled Dot1x Protocol Version 3 Dot1x Info for GigabitEthernet1/0/1 ----------------------------------- PAE = AUTHENTICATOR QuietPeriod = 60 ServerTimeout = 0 SuppTimeout = 30 ReAuthMax = 2 MaxReq = 2 TxPeriod = 7 ``` ### show authentication sessions This is the command you will live in. It shows every active session on the switch with its identity, method, VLAN, ACL, and state. Use `details` on a specific interface to see the full attribute list that ISE pushed. ``` switch# show authentication sessions interface Gi1/0/1 details Interface: GigabitEthernet1/0/1 IIF-ID: 0x1055802000000A7 MAC Address: 0050.5683.8a5c IPv4 Address: 10.20.20.55 User-Name: CORP\alice Status: Authorized Domain: DATA Oper host mode: multi-auth Oper control dir: both Session timeout: 3600s (server), Remaining: 3582s Common Session ID: 0A141001000000B96A7C3E14 Acct Session ID: 0x000000C4 Handle: 0xA4000089 Current Policy: DOT1X_MAB_POLICY Server Policies: Vlan Group: Vlan: 20 ACS ACL: xACSACLx-IP-PERMIT_ALL_TRAFFIC-5f4dcc3b Method status list: Method State dot1x Authc Success ``` The two lines that tell you it worked are `Status: Authorized` and `Method: dot1x, State: Authc Success`. A status of *Running* means dot1x is still in progress (wait 30 seconds), *Unauthorized* means it failed, and no session at all means the supplicant never sent EAPOL-Start. ### debug dot1x and debug radius When a session fails silently, turn on targeted debugs. Be careful with `debug radius authentication` on production - it is verbose. Always pair it with `debug condition interface Gi1/0/1` to limit output to a single port. ``` switch# debug condition interface GigabitEthernet1/0/1 switch# debug dot1x events switch# debug dot1x errors switch# debug radius authentication switch# terminal monitor ``` Run your test, then `undebug all` immediately. In the debug output you are looking for `RADIUS: Received from id ... Access-Accept` (success) or `Access-Reject` (check ISE Live Logs for the reason), and for dot1x you want to see `EAP-Request/Identity` and `EAP-Response/Identity` in both directions. ### ISE Live Logs On the ISE side, navigate to **Operations > RADIUS > Live Logs**. Every authentication attempt, successful or failed, shows up here within 5 seconds. Click the magnifying glass on any row to see the full attribute dump, the Authentication Policy that matched, the Authorization Profile that was selected, and the reason for failure if it failed. Live Logs are the single most useful troubleshooting surface in ISE - most problems reveal themselves as an obvious "Identity not found" or "Authentication method not allowed" entry in the failure reason field. ## Common Failure Scenarios and Fixes The table below is the short list of failure modes you will hit, in descending order of frequency. Work it top-down - most "802.1x is broken" tickets are one of the first three. No session appears on the switch at all Likely Cause Supplicant not sending EAPOL-Start, or Windows Wired AutoConfig service not running Fix On the client: `Get-Service dot3svc` \- should be Running. Start it via GPO or manually. `show aaa servers` shows state DEAD Likely Cause Switch cannot reach ISE on 1812/1813, or shared secret mismatch, or NAD IP mismatch Fix Verify IP reachability, then confirm shared secret character-by-character, then confirm the switch's source IP matches the IP configured in ISE Network Devices. ISE Live Logs show *5440 Endpoint abandoned EAP session* Likely Cause Supplicant dropped mid-handshake, usually certificate trust issue Fix On the client, install the ISE EAP certificate's issuing CA into Trusted Root. For EAP-TLS, verify the client cert has *Client Authentication* EKU. ISE Live Logs show *22056 Subject not found in the applicable identity store* Likely Cause Username not in AD, or Identity Source Sequence pointed at the wrong store, or AD join is unhealthy Fix Confirm the username exists in AD with `net user /domain`, check ISE AD join status, verify the Authentication Policy's identity source selection. Port authorizes but ends up in the wrong VLAN Likely Cause Authorization Profile VLAN attribute is a name that does not exist on the switch, or a different policy is matching first Fix Use `show authentication sessions ... details` to see which Authorization Profile was returned. If it is wrong, fix the policy order in ISE. If it is right but the VLAN is wrong, ensure `switchport access vlan` matches what ISE returns (or create the VLAN on the switch). EAP-TLS fails with *12520 EAP-TLS failed SSL/TLS handshake* Likely Cause ISE does not trust the client certificate's issuing CA, or the cert's subject/SAN does not match what CAP expects Fix Import the client CA into ISE's Trusted Certificates with *Trust for client authentication* enabled. Verify the Certificate Authentication Profile matches on the correct field (SAN, CN, or identity). Port stuck in Running state for 30+ seconds Likely Cause Supplicant configured for an EAP method that ISE is not offering, or certificate issue stalls the TLS tunnel Fix Match the Allowed Protocols in ISE's Authentication Policy to the supplicant's configured EAP method. On the client, temporarily enable Wired AutoConfig's *logging* via `netsh trace start scenario=Wireless_WlanAutoconfig` for the wired equivalent. MAB fallback takes 90+ seconds Likely Cause Default `dot1x timeout tx-period` is 30 seconds, and the policy retries 3 times before failing to MAB Fix Set `dot1x timeout tx-period 7` and `dot1x max-reauth-req 2` at the interface. 802.1x will now fail over to MAB in about 21 seconds total. ## Key Takeaways A working Cisco ISE 802.1x wired configuration is the intersection of three independently-correct pieces: a supplicant that sends EAPOL, a switch with matching AAA method lists and an IBNS 2.0 policy-map, and an ISE Policy Set that authenticates against a reachable identity store and returns a valid Authorization Profile. If any of the three is wrong, the other two will appear to be working. Use IBNS 2.0 (policy-map type control subscriber) on every new deployment - the legacy *authentication* commands still work but Cisco has deprecated them and TAC will push you to IBNS 2.0 on any case you open. Always enable `automate-tester` on your RADIUS servers so dead-server detection works cleanly, and always configure CoA (`aaa server radius dynamic-author`) because you will need it the moment you try to add posture, profiling, or any runtime policy change. When you deploy this for real, ramp in phases. Start with `access-session monitor` so failures log without blocking traffic, let ISE Live Logs tell you which endpoints are not ready, fix or exempt them, then flip to `access-session closed`. Combine this guide with the companion articles on [MAB fallback](https://www.pinglabz.com/mab-configuration-cisco-ios-xe-ise/), [Guest, Auth-Fail, and Critical VLANs](https://www.pinglabz.com/802-1x-guest-vlan-auth-fail-vlan-critical-vlan/), and [show authentication sessions troubleshooting](https://www.pinglabz.com/802-1x-troubleshooting-show-authentication-sessions-debug/) for the edge cases this guide intentionally leaves to dedicated references. ### References - [IEEE 802.1X-2020 - Port-Based Network Access Control](https://standards.ieee.org/ieee/802.1X/7345/?ref=pinglabz.com) - [RFC 3748 - Extensible Authentication Protocol (EAP)](https://www.rfc-editor.org/rfc/rfc3748?ref=pinglabz.com) Build this yourself, on real gear The PingLabz lab library: 74 hands-on labs on real Cisco IOS XE in Cisco Modeling Labs. Seven are free, no card required. [Browse the labs](https://www.pinglabz.com/labs/) ### Cisco SD-WAN in 2026: The Real Value Is Day-2 Operations URL: https://www.pinglabz.com/cisco-sd-wan-in-2026-the-real-value-is-day-2-operations/ Last updated: 2026-07-04T23:03:44.000Z In 2016, the SD-WAN pitch was simple and persuasive: replace expensive MPLS circuits with cheaper broadband, steer traffic intelligently to avoid outages, and cut your WAN bill. That value proposition was real. Organizations reduced WAN costs by 30 to 50 percent in the first wave of deployments, and path steering genuinely solved problems that static routing could not. But that era is over. The SD-WAN market has matured, the buying conversation has matured, and the organizations asking questions today are not asking whether SD-WAN is worth doing. They are asking whether the platform they deployed three years ago is actually delivering on day 200 of the rollout. (This article is part of the PingLabz SD-WAN series - the [full SD-WAN guide](https://www.pinglabz.com/sd-wan/) maps the whole cluster in reading order.) This article is about that question. Path steering is now table stakes across every major vendor. The lasting competitive differentiation in SD-WAN comes from operational depth: how well the platform handles template management at scale, how effectively it surfaces policy drift before it causes incidents, how much it helps you isolate ISP-side faults without a four-hour carrier support call, and whether the observability story holds up when a user says Salesforce is slow and your controller shows everything green. These are day-200 problems, and they are the problems that determine whether an SD-WAN deployment was a good investment. ## The Original Promise Three pillars drove the initial SD-WAN wave: cost reduction, application performance, and operational simplicity. It is worth being specific about what each delivered, because the residual value from each has shifted significantly. ### Cost reduction: real, but mostly captured MPLS circuits ran 40 to 60 dollars per Mbps per month with committed minimums of 25 to 50 Mbps per site. Broadband ran 5 to 15 dollars per Mbps with no minimum and far higher available bandwidth. The math was compelling across a 200-branch network, and organizations that made the shift in 2016 to 2020 captured those savings. But that transition is largely done for organizations that were going to do it. Most enterprises now run dual-circuit architectures: MPLS primary with internet backup, or two internet circuits with SD-WAN failover. You are paying for redundancy, not just connectivity, and the per-circuit cost delta against MPLS has narrowed as carriers adjusted pricing. Circuit cost savings are still real, but they are no longer the headline reason to buy. ### Path steering: now table stakes Intelligent path steering was the core differentiator in early SD-WAN. A voice call uses MPLS. A backup job goes over broadband. SaaS traffic uses the best available circuit based on real-time quality metrics. This was a genuine improvement over static routing, and it still works well. But Cisco Meraki, Fortinet, Arista, Aruba, Palo Alto Networks, and a dozen other vendors all offer effective path steering. If your vendor's primary pitch is "we have smart path steering," that is no longer a reason to buy. It is a requirement, not a differentiator. You should evaluate it, verify it works for your traffic types, and then look at what else the platform does. ### Operational simplicity: depends on your starting point The original promise suggested SD-WAN would reduce operational overhead compared to MPLS. In practice, it shifts the overhead. The carrier no longer configures your branch circuits, but you now own the policy configuration across every branch in the SD-WAN controller. If you have one clean template and 200 branches, that is simpler. If you have accumulated thirty-five templates over three years to account for regional ISP variance, compliance requirements, and business-unit exceptions, you have shifted significant complexity from the carrier to yourself. And unlike carrier complexity, which was their problem, template complexity is entirely yours. ## Why the Buyer Conversation Has Changed In 2026, mature SD-WAN prospects are not asking "should we deploy SD-WAN?" They are asking one of three much more specific questions: "We have SD-WAN and it is not solving our visibility problem, can anything fix that?"; "Path steering works fine, but policy is inconsistent across branches and we cannot tell why"; or "We deployed this three years ago and it mostly runs itself, but every time we have a WAN incident we spend hours figuring out whether it is our problem or the ISP's problem." These are day-200 questions, and they point to where the real platform value either exists or does not. Consider which scenario you are actually in before evaluating platform options: **Greenfield consolidation.** You are migrating from MPLS to internet-based WAN across dozens of branches for the first time. Cost reduction and path steering are the primary value. Nearly any major platform will serve you well. Pick the one that fits your existing vendor ecosystem. **Legacy SD-WAN expansion.** You have had SD-WAN running in large offices for several years. You want to extend it to 150 smaller remote branches with poor ISP SLAs and SaaS-heavy users. The value you need now is observability and operational manageability, not path steering. Evaluate the platform's operational depth, not its path steering sophistication. **Multi-vendor brownfield.** You have Cisco in some branches, Fortinet in others, Arista in a few. You want consistent policy and visibility across all of them. No single SD-WAN platform solves this cleanly. You need external instrumentation (ThousandEyes or a competing platform) more than you need additional SD-WAN features. The SD-WAN platform purchase alone will not unify your operational visibility. The first scenario is classic SD-WAN territory. The second and third are where you find out whether the platform has grown with the market's operational needs or whether it stalled after solving the initial path-steering problem. ## Day-200 Operational Pain Points Six problems surface repeatedly in mature SD-WAN deployments. None of these are unique to Cisco, but how the platform handles each matters for long-term operational cost. ### Template sprawl Most deployments start with one or two standard branch templates. By month six, you have regional variants because ISP failover behavior differs between continents. By month twelve, you have business-unit exceptions because one group's SaaS provider requires specific QoS tuning. By month eighteen, you are maintaining a legacy template for hardware that cannot run the current version. By month twenty-four, you have seventeen templates and nobody remembers why half of them exist. This is not a platform failure. It is the natural consequence of running a network that serves multiple groups with different requirements. But it is operational debt. When a security policy needs to change, do you update all seventeen templates? When an engineer leaves, does the next one understand the template hierarchy? Catalyst Center helps: you can group templates, see deployment status per device, and track versions. But active template governance is required regardless of tooling. The platform makes it easier or harder, it does not make it unnecessary. Template versioning Why it matters A change to the main template should not break branches still running an older version Cisco SD-WAN platform approach Template groups in Catalyst Center provide version visibility and device-group-scoped deployment Policy inheritance and overrides Why it matters Some policies should be global, others regional, others branch-specific. Layering these without contradiction is hard Cisco SD-WAN platform approach Partial support via template groups and per-device customization. Less clean than infrastructure-as-code models in some competing platforms Change tracking and auditing Why it matters You need to know which branches received a change, when, and whether it caused issues Cisco SD-WAN platform approach Catalyst Center provides audit logs. Integration with ITSM tools requires custom automation Pre-rollout testing Why it matters Deploying a bad policy to 200 branches at once is an incident. You need canary deployment capability Cisco SD-WAN platform approach Limited. You can deploy to a device subset, but there is no integrated staging environment ### Policy drift A security incident happens at 11 PM. Someone bypasses the change process and adds a firewall rule directly on three branches to stop a threat. The incident resolves. The rule stays. Six months later it is blocking a legitimate vendor connection and nobody knows why it is there. That is policy drift: the running configuration diverges from the template-defined intended state. Drift happens because operational urgency sometimes exceeds process speed. Cisco SD-WAN provides compliance checking that compares the running config against the template and flags deviations. But the check is reactive, and when drift is flagged, the platform does not know whether the deviation is intentional (a tracked exception) or accidental (a forgotten fix). Resolving that distinction requires human judgment and a process that captures out-of-band changes with context. The technology supports drift detection. The discipline to act on it and prevent it is organizational. ### ISP circuit fault isolation A branch reports circuit degradation. The SD-WAN controller shows the interface is up, BGP is stable, and no failover has triggered. Telemetry shows elevated latency and jitter, but nothing crosses the failover threshold. You call the ISP and they say they see nothing wrong on their end. You say something is wrong. You spend ninety minutes on a conference call making educated guesses about where the problem is. The SD-WAN controller cannot see inside the ISP's network. It knows about your side of the circuit: interface state, BGP route health, drops on your interface. The moment traffic leaves your edge, it is opaque. This is not a platform shortcoming, it is a fundamental architectural boundary. Resolving it requires instrumentation that runs synthetic tests across the ISP path from your branches to known destinations, with hop-by-hop visibility into where latency or loss occurs. That is exactly what ThousandEyes Enterprise Agents do, and this ISP fault isolation use case is where the integration provides the clearest operational ROI. ### SaaS troubleshooting blind spots A user at a branch reports that Salesforce is slow. Your controller shows healthy circuits, clean path steering, and no firewall blocks. Everything you manage is green. The problem is that the path from the branch to Salesforce crosses networks you do not own: your ISP, potentially a transit provider, Salesforce's CDN, and Salesforce's own cloud infrastructure. Your SD-WAN controller is designed to manage traffic up to your internet edge. That is what it does well. But it has no visibility into whether Salesforce's CDN endpoint for that region is serving from a distant failover location, whether there is a congestion event on an ISP-to-Salesforce peering link, or whether the problem is actually on Salesforce's infrastructure side. This is endemic to any branch architecture where you own the access layer but not the internet path or the SaaS provider's infrastructure. The answer is not a better SD-WAN controller. It is external visibility into the paths you do not own. ### MPLS coexistence complexity Many organizations are in a transition state where some traffic still uses MPLS (voice, latency-sensitive applications) while the rest uses internet SD-WAN paths. Managing QoS policies that correctly treat these paths differently, ensuring failover logic works correctly when MPLS degrades, and troubleshooting path selection anomalies are sources of sustained operational friction. SD-WAN platforms handle this, but the more circuit types you operate in parallel, the more complex the policy model becomes. You cannot confidently complete an MPLS migration until you have months of latency and quality data from the internet paths that would replace it. That data only comes from sustained observability, not from deployment telemetry alone. ### Multi-region policy consistency An organization with branches across North America, Europe, and Asia-Pacific faces overlapping regulatory constraints. GDPR governs data handling for European branches. CCPA applies to California. Other jurisdictions have their own rules. Your SD-WAN policy must enforce that European branch traffic does not traverse non-EU infrastructure for certain data categories, that specific encryption standards apply to certain regions, and that policy changes meet audit requirements before deployment. Cisco SD-WAN supports this through device groups and regional policy objects. Building and maintaining a compliant policy model across regions is sophisticated work that requires deep knowledge of both the regulatory requirements and the platform's policy language. Template sprawl SeverityHigh Platform problem or operational discipline problem? Both. The platform can make governance easier or harder, but active management is required regardless Policy drift SeverityHigh Platform problem or operational discipline problem? Operational discipline. Compliance checking exists in the platform, but enforcement is human ISP fault isolation SeverityMedium-High Platform problem or operational discipline problem? Platform limitation. Requires external observability. The controller cannot see past the ISP edge SaaS troubleshooting SeverityHigh Platform problem or operational discipline problem? Platform limitation. The controller scope ends at your internet edge MPLS coexistence SeverityMedium Platform problem or operational discipline problem? Platform design choice. Handled well, but adds policy complexity proportional to the number of path types Multi-region compliance SeverityHigh Platform problem or operational discipline problem? Both. The policy language supports it, but building and maintaining compliant models requires sustained expertise ## Observability and ThousandEyes Integration Cisco's investment in ThousandEyes is a direct response to the SD-WAN controller's inherent blind spots. An SD-WAN controller is a management and traffic steering tool. ThousandEyes is purpose-built for observability across the internet and SaaS paths that the controller cannot instrument. The integration between Catalyst Center and ThousandEyes is the most operationally significant capability differentiation Cisco has against its SD-WAN competitors right now, and it is worth being specific about what it actually provides versus where it still has gaps. ### What ThousandEyes adds to SD-WAN operations With Enterprise Agents deployed on Cisco branch routers or access switches, ThousandEyes provides three categories of visibility that the SD-WAN controller cannot supply: **ISP path tracing and hop-by-hop analysis.** Synthetic tests run continuously from branch agents to known destinations (your data center, public DNS, a representative SaaS app). When a circuit shows elevated latency in the controller, ThousandEyes path traces show you exactly which hop in the ISP's network is introducing the problem, with precise latency attribution per hop. Instead of telling your ISP "our link seems slow," you are showing them a graph with timestamps that identifies the specific ISP POP or transit node that is causing the problem. That changes a 90-minute diagnostic call into a 15-minute escalation to the right ISP team. **SaaS endpoint monitoring across branches.** Pre-configured tests for Salesforce, Microsoft 365, Webex, Zoom, and others run from every branch agent and report response time, availability, and path quality against each SaaS provider's CDN endpoint. When a user reports a Salesforce problem, you check whether the branch agent is seeing degradation to Salesforce right now. If it is, you have data that removes your network from suspicion and supports a Salesforce support case. If it is not, the problem is device-side or a transient event the test interval missed. Either way, you iterate faster than checking logs and making educated guesses. **BGP route monitoring and ISP routing events.** ThousandEyes monitors BGP route changes that can cause sudden latency shifts or traffic rerouting. When a branch shows an unexplained latency spike, BGP monitoring often explains it: an ISP route changed, and traffic is now taking a longer path through a different transit provider. This is invisible to the SD-WAN controller unless the change is severe enough to trigger a failover event. ThousandEyes surfaces it before users start calling. ### Where the integration still has gaps The Catalyst Center and ThousandEyes integration is functional and provides real operational value, but it is not a one-pane-of-glass solution yet. These gaps are worth understanding before you build a business case around tight integration. Internet path trace (branch to any destination) What it provides for SD-WAN operations Hop-by-hop latency and loss across ISP networks, identifies fault location with precision Integration maturity with Catalyst Center Mature. Works well from within Catalyst Center SaaS app monitoring (Salesforce, M365, Webex, Zoom) What it provides for SD-WAN operations Pre-built tests show whether a specific application is performing normally from branch agents Integration maturity with Catalyst Center Mature. Good coverage of common enterprise SaaS BGP route monitoring What it provides for SD-WAN operations Shows ISP route changes that cause latency or rerouting events before they trigger controller failover Integration maturity with Catalyst Center Mature. Excellent visibility into BGP events DNS performance from branches What it provides for SD-WAN operations Tests to internal and public DNS identify slow resolution that cascades into application performance problems Integration maturity with Catalyst Center Solid. Custom tests are well supported Endpoint Agent (user device perspective) What it provides for SD-WAN operations Captures network performance from the user's device, showing local gateway health and app performance from the client side Integration maturity with Catalyst Center Functional but limited. Per-device agent data is not deeply integrated into Catalyst Center dashboards Automated correlation with SD-WAN policy changes What it provides for SD-WAN operations When a policy change deploys and an app degrades, the system flags the timing correlation automatically Integration maturity with Catalyst Center Limited. Manual correlation is required across Catalyst Center and ThousandEyes portals Correlation with ISP BGP events and app impact What it provides for SD-WAN operations When an ISP route flaps, the system automatically shows which apps were affected and for how long Integration maturity with Catalyst Center Good within ThousandEyes. Limited cross-system correlation with Catalyst Center events Custom synthetic tests (internal apps, microservices) What it provides for SD-WAN operations Run HTTP, TCP, DNS tests to internal systems to instrument performance from branches Integration maturity with Catalyst Center Good. Flexible test types are well supported The most significant integration gap is automated root-cause correlation. If an SD-WAN policy deploys at 2:55 PM and Salesforce latency spikes at 2:58 PM, a human reviewing both systems can infer causation. But the platform does not automatically surface that correlation. You are still doing detective work; you just have better data to work with. Closing this gap is on Cisco's roadmap, but it is not there today. The second gap is client-side visibility. ThousandEyes can run Endpoint Agents on user laptops, capturing network performance from the device itself. That data is valuable because it represents the actual user experience, including local Wi-Fi and the connection from the device to the branch gateway. In Catalyst Center, you mostly see branch-level ThousandEyes data. Correlating client-side agent data with branch-level SD-WAN metrics requires jumping between portals. For most day-to-day troubleshooting this is acceptable. For organizations that need to close SLA gaps on application performance down to the user level, it is a workflow friction point. Neither gap is a blocker. The integration works well enough to deliver real operational value on the ISP fault isolation and SaaS troubleshooting use cases that matter most. But set expectations appropriately if a vendor is pitching this as a single integrated observability and management platform. It is not that yet. ## Where Cisco SD-WAN Still Makes Strong Sense Given all the day-200 reality checks, there are specific scenarios where Cisco SD-WAN is genuinely the right platform choice in 2026. **Large-scale single-vendor greenfield migrations.** If you are a 500-plus branch enterprise consolidating from MPLS to internet-based WAN and you want a carrier-independent platform that scales, handles complex multi-circuit failover, and integrates with a unified management layer, Cisco SD-WAN is the strongest option. It has the operational maturity, the carrier ecosystem relationships for migration support, and the scale track record. The commitment required is long-term (you will be managing templates and policies for years), so the operational depth matters more than initial deployment simplicity. **Organizations with deep Cisco Catalyst campus investments.** If your access and campus layers are already built on Catalyst switching managed through Catalyst Center, extending SD-WAN management to the same platform reduces operational context-switching, unifies asset inventory, and provides consistent policy language across access, campus, and WAN edge. This is an organizational efficiency argument as much as a technical one, but for large teams managing complex environments, it is a real cost reduction. **Branches with mixed circuit types and complex failover logic.** If your branch topology uses a mix of MPLS, broadband, and LTE (sometimes all three at one site), and you need sophisticated rules about which traffic uses which path based on real-time quality metrics, SD-WAN is built for this. The alternative is static routing with manual failover, which is painful and error-prone. Cisco's implementation of quality-based path steering handles heterogeneous circuit environments well. **Organizations that will genuinely invest in observability integration.** The value of the ThousandEyes integration is real, but only if you deploy it alongside the SD-WAN rollout and build operational workflows around the data it provides. If your organization will fund and staff that integration from the start, Cisco SD-WAN plus ThousandEyes is a strong combined platform. If observability will be deferred, the value of the SD-WAN investment alone is diminished. ## Where You Should Be Skeptical There are scenarios where the Cisco SD-WAN investment is harder to justify, where other approaches are more appropriate, or where the platform's limitations will frustrate you. **Small branch counts (under 20 sites).** The operational leverage of centralized SD-WAN management becomes a force multiplier at scale. At 10 or 15 branches, the overhead of managing a controller, templates, and licensing may not justify the savings over simpler approaches. Know where you sit on the scale curve before committing. **As a substitute for observability.** If your primary problem is "we cannot see what is happening in our WAN," SD-WAN will not solve that. It will let you manage the WAN more efficiently, but the blind spots remain past the internet edge. If you cannot budget for ThousandEyes or an equivalent observability platform alongside the SD-WAN investment, be honest about what you are actually buying. **Primarily remote workforces.** SD-WAN manages your branch edges. It does not help users on home broadband or public Wi-Fi. If most of your workforce connects from unmanaged locations, Secure Service Edge (SSE) and Zero Trust Network Access (ZTNA) are more relevant investments. SD-WAN and SSE/ZTNA are complementary, not competing, but if you have to choose where to invest first, align it with where your users actually are. **Pure cost-reduction plays with no observability commitment.** Circuit consolidation savings are real, but they are a one-time benefit. Ongoing operational cost reduction requires mature template management, observability integration, and process discipline. If the business case is based entirely on circuit savings, it will hold up for the first year. By year three, you will be asking what else the platform is delivering. Plan the full operational picture upfront. Large-scale MPLS to internet migration (500+ branches) Cisco SD-WAN fitStrong fit Key consideration Invest in template architecture and observability from day one Existing Cisco Catalyst campus, adding WAN management Cisco SD-WAN fitStrong fit Key consideration Operational consistency advantage is real; budget for unified management licensing Mixed circuit environments (MPLS, broadband, LTE) Cisco SD-WAN fitStrong fit Key consideration Path steering and QoS-based policy is mature and handles heterogeneous topologies well Legacy SD-WAN expansion with observability investment Cisco SD-WAN fitGood fit Key consideration ThousandEyes integration provides real operational value; plan for portal-switching until correlation improves Small branch count (fewer than 20 sites) Cisco SD-WAN fitWeak fit Key consideration Operational overhead may not justify cost and complexity at this scale Primarily remote workforce (home/public networks) Cisco SD-WAN fitWrong tool Key consideration SSE and ZTNA serve remote users; SD-WAN manages branch edges Observability-first problem with no management need Cisco SD-WAN fitWrong tool Key consideration ThousandEyes alone is the right investment; SD-WAN does not solve visibility past the internet edge Multi-vendor brownfield with unified policy goal Cisco SD-WAN fitPartial fit Key consideration No single SD-WAN platform unifies multi-vendor visibility cleanly; external instrumentation required ## A Practical Rollout Framework If you decide Cisco SD-WAN is the right investment, the decisions you make in the first 90 days determine whether you end up with a well-governed platform or a template sprawl problem. Here is a framework that avoids the common failures. ### Phase 1: Pilot with representative complexity (months 1 to 3) Do not pilot in greenfield. Pick 8 to 12 branches that represent your hardest cases: mixed circuit types, different ISP providers, different user profiles. The goal is to find configuration pain points before they are baked into your template model. Deploy ThousandEyes agents at the same sites simultaneously. Observability is not optional and is not a phase two project. Build your template hierarchy before you have dozens of branches to migrate. Even if you only have one template variant today, document the rationale for every policy in the template. That documentation is worth more than you expect when the engineer who designed it leaves. ### Phase 2: Observability buildout during rollout (months 4 to 12) As you roll out to additional branches, configure ThousandEyes tests for your top 10 SaaS applications from day one. Set alert thresholds based on observed baselines after 30 days of data, not on guessed values. Build a change management workflow that tracks every SD-WAN policy change with context (who, why, what branches affected). Even without automated correlation in the platform, side-by-side visibility in separate portals is significantly better than no data. ### Phase 3: Scale with governance (months 12 and beyond) By the time you are rolling out to the majority of branches, you should have enough operational confidence to enforce a formal change process for all policy changes, canary group deployments before wide rollout, and weekly drift detection reviews that are actioned. No direct device configuration changes except in genuine emergencies, and those are tracked and closed within 48 hours. This level of discipline is achievable, but it requires commitment from engineering leadership. The teams that enforce it consistently are the ones that look back at their SD-WAN deployment as a success. The teams that defer governance accumulate template debt until it causes an incident. ## Key Takeaways Cisco SD-WAN is a mature platform that does what it claims: steers traffic intelligently, consolidates WAN circuits, and scales to thousands of branches. But the buying story has fundamentally changed. Path steering is table stakes in 2026\. Every major vendor offers it. The lasting value of the platform is in its operational depth for managing a distributed branch network over years, not months. The dominant day-200 pain points are template sprawl and policy drift (both addressable with governance discipline, and better tooling helps), and ISP fault isolation plus SaaS troubleshooting (both require external observability, and ThousandEyes integration provides real value here). The integration between Catalyst Center and ThousandEyes is genuinely useful for ISP fault isolation and SaaS visibility. It is not yet a fully automated, single-pane solution. Plan your workflows around the gaps. The organizations that get the most from Cisco SD-WAN are the ones that invest in observability, template architecture, and operational governance alongside the rollout rather than after. The ones that treat SD-WAN as a deploy-and-forget infrastructure play end up with a circuit consolidation tool that they cannot interrogate when something goes wrong. If you have a clear problem that SD-WAN solves, a realistic budget for the operational layer, and the organizational discipline to maintain template hygiene, it remains a solid investment in 2026\. If you are buying SD-WAN hoping it will be a low-maintenance, self-managing platform, adjust those expectations before you sign the contract. ### Catalyst Center CLI Templates for Network Teams That Hate Fragile Automation URL: https://www.pinglabz.com/catalyst-center-cli-templates-for-network-teams-that-hate-fragile-automation/ Last updated: 2026-07-04T23:06:41.000Z Most network teams end up in one of two places. The first is full manual: every switch gets configured by hand, every change is a SSH session, and consistency depends entirely on whoever is doing the work that day. The second is a project that tried to fix that problem and went too far: a half-deployed Ansible framework that three people understand, a Python script library that nobody documents, or a NetOps platform that took eight months to build and now nobody maintains. Both of these outcomes are real, both are painful, and the irony is that most teams have access to a middle path they don't fully use. Catalyst Center CLI templates sit squarely in that middle path. They are not a full automation framework. They do not replace Ansible or Python for complex workflows. But for the bulk of what campus and enterprise network teams actually do day to day, they are genuinely practical: repeatable provisioning, consistent configuration, and reusable logic that does not require anyone on the team to be a developer. This article is a practitioner-level look at how to use them well, where they fall apart, and how to adopt them without creating new fragility in the process. ## The Automation Trap Most Teams Fall Into The fully manual model has an obvious cost: speed, consistency, and operator error. Every engineer who configures a switch has their own conventions. VLANs get numbered differently. NTP servers get set inconsistently. Logging destinations vary. AAA config drifts over time. None of this is catastrophic individually, but it accumulates. When you need to make a change across 150 switches, inconsistency is the thing that turns a one-hour change window into a three-hour incident. The overcorrection is just as common. A team decides to "do automation properly" and starts building infrastructure around it: Ansible roles, Git repositories, CI pipelines, variable files per site, per-device inventory. This is genuinely the right model at a certain scale, but it carries a real cost of its own. Someone has to maintain the tooling, test it, handle Python dependency upgrades, manage SSH key rotation, and document it well enough that the next engineer can actually use it. For a team of three people managing 200 switches who also run helpdesk escalations and are responsible for the WAN, that investment is often not sustainable. CLI templates in Catalyst Center do not require you to operate a CI pipeline or learn Ansible modules. The templating logic runs inside Catalyst Center. Variables get filled in through the UI or via API. Devices get configured through the existing management infrastructure you already have. That narrower scope is not a limitation; for many teams, it is exactly the right fit. ## What CLI Templates Are in Catalyst Center A CLI template in Catalyst Center is a text block of IOS/IOS-XE commands with substitution variables. When you apply a template to a device, Catalyst Center resolves the variables, generates the final CLI, and pushes it over SSH or NETCONF. That is the core of it. The complexity lives in the details of how variables work, how templates are structured, and when they run. Catalyst Center distinguishes between two categories of templates based on when in the device lifecycle they run: Day 0 (Onboarding) When it runs During initial device provisioning via Plug and Play (PnP) Common use cases Initial hostname, management IP, AAA, NTP, SSH, base ACLs, SNMP Trigger PnP workflow; device contacts Catalyst Center during first boot Day N (Provisioning) When it runs After a device is already under management Common use cases VLAN additions, interface configuration, ACL updates, routing changes, policy changes Trigger Manual trigger from the Catalyst Center UI, API call, or network profile association Day 0 templates handle the foundational configuration that every device needs to be manageable. Day N templates handle everything after that: ongoing operations, incremental changes, and policy updates. In practice, your Day 0 template should be relatively stable (you are not changing your management model constantly), and your Day N template library is what grows over time. Within these categories, Catalyst Center also supports **Composite Templates**, which are wrappers that sequence multiple regular templates into a logical flow. If you have a base infrastructure template, a site-specific template, and a role-specific template (access vs. distribution), a composite template lets you apply all three in order as a single operation. That composability is where real modularity comes from. ### The Template Editor The template editor in Catalyst Center (Design > Network Profiles > CLI Templates) is functional for day-to-day work. It provides syntax highlighting for both Velocity and Jinja2 (the two scripting languages Catalyst Center supports), inline variable detection, a built-in simulator that lets you fill in variable values and preview the rendered output before pushing anything to a device, and a version control system that tracks every committed change with a history and diff view. The version control is worth understanding: saving a template stages it, while committing creates a new version in the history. This works conceptually like Git, but it is entirely internal to Catalyst Center. If you want to keep templates in external version control (which you should, and we will cover that), you need to export them manually. ## Velocity and Jinja2: What You Need to Know Catalyst Center supports two template scripting languages: Apache Velocity and Jinja2\. Velocity has been in Catalyst Center (originally DNA Center) for longer and you will see it in most existing template libraries. Jinja2 is the newer option and has a cleaner, more Pythonic syntax that most network engineers find easier to read and write. If you are starting a new template library, Jinja2 is the better choice. If you are inheriting an existing Velocity template library, it still works and there is no urgent reason to rewrite everything. Both languages support the same core capabilities: variable substitution, conditional logic, and loops. Here is what each looks like in practice: ### Velocity Syntax ``` ## Set NTP servers based on site region #if($site_region == "east") ntp server 10.1.0.1 prefer ntp server 10.1.0.2 #elseif($site_region == "west") ntp server 10.2.0.1 prefer ntp server 10.2.0.2 #else ntp server 10.0.0.1 prefer #end ## Configure management interface interface Vlan$mgmt_vlan description Management ip address $mgmt_ip $mgmt_mask no shutdown ``` ### Jinja2 Syntax ``` {# Set NTP servers based on site region #} {% if site_region == "east" %} ntp server 10.1.0.1 prefer ntp server 10.1.0.2 {% elif site_region == "west" %} ntp server 10.2.0.1 prefer ntp server 10.2.0.2 {% else %} ntp server 10.0.0.1 prefer {% endif %} {# Configure management interface #} interface Vlan{{ mgmt_vlan }} description Management ip address {{ mgmt_ip }} {{ mgmt_mask }} no shutdown ``` The logic is identical. Jinja2 uses `{% %}` for control statements and `{{ }}` for variable output. Velocity uses `#` for directives and `$` for variables. Pick one and be consistent across your library. ### Variable Types Catalyst Center exposes three categories of variables in templates: - **Binding variables** are values you provide at provisioning time through the UI or API. These are the inputs the engineer fills in: hostname, management IP, VLAN ID, site code, etc. These are what make a template reusable across devices with different values. - **System variables** are values Catalyst Center generates automatically from its inventory. Device platform, serial number, current software version, and similar data are available as system variables. You can use them in templates without the operator having to type anything. - **Regular variables** are values you define within the template itself, typically to hold intermediate results or constants that are used across multiple places in the same template. The binding variables are where most of the design work goes. Getting variable naming and data types right at the start saves significant pain later (more on that in the pitfalls section). ## The Practical Value: Why Templates Beat Manual Config The argument for templates is not that they are faster than typing CLI (for a single device, they often are not). The argument is that they are consistent, auditable, and repeatable across dozens or hundreds of devices. A well-written Day 0 template means every switch that comes out of the box gets the same baseline configuration: same AAA server group, same NTP servers, same SNMP community strings and ACLs, same SSH version and key size, same logging configuration. When a security team audits your network, compliance is a matter of checking which template was applied, not reading running configs on 300 switches one at a time. A well-written Day N template for VLAN provisioning means that adding a VLAN to a group of access switches is a form submission, not a change window with CLI commands on each device. It also means the VLAN addition produces the same configuration on every switch, which matters when you are troubleshooting a connectivity issue at 11pm and need to know whether the configuration is right. The human error reduction is the most immediately valuable thing. Templates do not typo IP addresses. Templates do not forget to configure the switchport access vlan command after configuring the interface description. Templates apply the same eight lines of AAA config to device 1 as they do to device 150. ## How to Structure Templates Well The biggest structural decision is scope: what belongs in a single template versus what should be broken into multiple templates and composed. The right answer is smaller, focused templates rather than large monolithic ones. ### Modular Design Think of your template library in functional layers. A reasonable structure for a campus environment looks like this: - **Base infrastructure template:** Hostname convention, domain name, DNS, NTP, logging, SNMP, SSH configuration, AAA (TACACS+/RADIUS server groups, authentication/authorization/accounting method lists). This applies to every managed device, every time. - **Role-specific templates:** Access layer gets spanning-tree portfast default, storm-control defaults, port security policies. Distribution layer gets HSRP/VRRP configuration, IP routing, inter-VLAN routing policy. Each role has its own template. - **Site-specific templates:** Site code, building designation, any parameters that vary by physical location but not by device role (local syslog forwarder, local DHCP server, etc.). - **Change templates:** Operational Day N templates for specific changes: VLAN provisioning, interface configuration, ACL updates. These are applied on demand, not tied to onboarding. A composite template for new switch onboarding would sequence: base infrastructure, then role-specific, then site-specific. That is three small, testable templates rather than one 200-line monster. When the AAA server address changes, you update one template and every device type benefits. ### Variable Naming and Design Variable design is where the quality of your template library gets decided. Sloppy variable design creates templates that are painful to use, hard to automate via API, and difficult to validate before applying. A few principles that save headaches: - **Use descriptive, consistent names.** `mgmt_vlan_id` is better than `vlan`. `ntp_server_primary` is better than `ntp1`. When you are filling in values for 20 templates six months from now, clarity in variable names prevents mistakes. - **Enforce data types where possible.** Catalyst Center lets you define variable types (string, integer, IP address, boolean). Use integer for VLAN IDs so that the system validates the input before the template runs. Use IP address type for addresses so that typos fail at input time, not at push time. - **Set default values for optional parameters.** If most devices use the same logging server but a handful are exceptions, set the default in the template and let operators override it when needed. This keeps the provisioning form short and reduces the chance of missing a required field. - **Avoid deeply nested conditional logic.** If your template has conditionals three levels deep, you have a maintenance problem. Break it into separate templates by role or environment instead. ### Designing for Idempotency An idempotent template is one that produces the same result whether you run it once or five times. IOS-XE helps here because most configuration commands are additive (running the same `ntp server` command twice does not create two entries), but edge cases exist. Be careful with anything that has a numbered sequence (access lists, route maps) or that uses `no` commands internally. If your template issues a `no interface Vlan10` followed by a reconfiguration block, applying it twice on a live device can cause a brief outage the second time. Test this explicitly with the built-in simulator and, where possible, on a lab device before putting a template into production rotation. ## Common Pitfalls ### Template Sprawl Template sprawl is the most common long-term problem. It starts innocently: someone needs a slightly different version of the VLAN template for one site, so they copy it and modify it. Someone else needs a version with a different ACL, so they copy that one. Six months later you have eleven VLAN templates, nobody remembers which one is authoritative, and a change to VLAN provisioning logic requires updating eleven files or guessing which ones matter. The mitigation is discipline, not tooling. Establish a naming convention and a review process before you have more than a handful of templates. Require that changes to existing templates go through a review step (even if the review is just a second engineer looking at it). Prefer variables over copies: if two use cases are 90% the same, the right answer is usually a single template with a conditional block, not two separate templates. ### No External Version Control Catalyst Center's internal version history is useful for auditing, but it is not a substitute for an external repository. Catalyst Center does not give you branching, pull requests, or the ability to roll back the entire template library to a point in time. Export your templates to a Git repository on a regular basis. Even a simple process (export to JSON, commit, push to a shared repo once a month) is dramatically better than relying entirely on Catalyst Center's internal versioning, which disappears if you ever need to rebuild or migrate your Catalyst Center instance. ### Velocity and Jinja2 Fragility The template scripting languages are where most technical fragility lives. A few specific patterns that cause problems in production: - **Whitespace sensitivity in Jinja2:** Jinja2 is sensitive to leading and trailing whitespace in ways that are not obvious until you see unexpected blank lines in rendered output. Use `{%- -%}` (with dashes) to strip whitespace around control blocks when it matters for IOS syntax. - **Null variable handling:** If a binding variable is left empty and your template does not handle that case, the rendered output will contain the literal variable name or throw an error depending on the engine version. Add null checks for any variable that is not required: in Jinja2, `{% if variable is defined and variable %}`; in Velocity, `#if($variable)`. - **Platform-specific IOS syntax:** A template written for IOS-XE 16.x may not render correctly on IOS 15.x or on a switch running a significantly different software generation. Platform-conditional logic adds complexity, but it beats applying a template to a mixed fleet and finding out later that half the devices accepted it and half silently failed. ### Version Lock Templates that encode platform-specific or software-version-specific assumptions become a burden when your fleet evolves. If your base template references a feature that was introduced in a specific IOS-XE release, applying it to older devices either fails or produces a device in an inconsistent state. Document the minimum software version assumption for each template in a comment block at the top. When you upgrade your fleet, validate templates against the new version before using them in production. This sounds obvious and is regularly skipped. ### Lack of Testing Discipline The built-in simulator shows you the rendered output (what the CLI will look like after variable substitution) but does not validate whether that CLI is syntactically correct or will be accepted by a device. The only way to know that is to test on real hardware (or a CML lab, if you have one). Establish a practice of testing every new template and every significant change on a non-production device before rolling it to the fleet. This seems slow in the moment and prevents the category of incident where you push a template to 50 switches at 2am and discover the rendered config has a typo in a critical ACL entry. ## A Practical Adoption Path for Teams Starting from Zero If your team is starting from fully manual operations, the fastest path to value without creating new problems is incremental adoption. Do not try to template everything at once. Phase 1: Baseline What to build A single Day 0 onboarding template covering AAA, NTP, SNMP, SSH, logging Why to start here Highest consistency value, lowest risk. Every new device benefits immediately. Template is stable and changes infrequently. Phase 2: VLAN operations What to build A Day N VLAN provisioning template (create VLAN, assign to interfaces, update trunk allowances) Why to start here VLAN additions are a frequent, repetitive task. Templating this one workflow delivers visible time savings and reduces the most common config drift vector. Phase 3: Role differentiation What to build Separate access and distribution templates; introduce composite templates Why to start here Once baseline and VLAN templates are stable and trusted, role-specific templates add value without much additional complexity. Phase 4: Change automation What to build Templates for common Day N changes: ACL updates, interface reconfigurations, QoS policy Why to start here Build this library based on what your team actually does repeatedly, not based on what seems theoretically useful. One well-used template is worth ten unused ones. At each phase, the goal is a small library of templates that the whole team understands and trusts, not a comprehensive library that nobody maintains. Get comfortable with the tool, establish the governance habits (naming, review, external version control), and add templates as operational demand justifies them. One practical note on starting from scratch: before writing your first template, spend time in the simulator with a device that already has a known-good configuration. Export its running config and use that as your template source. Templating a configuration you have already validated in production is faster and safer than writing from scratch. ## When Templates Are Enough and When to Graduate CLI templates in Catalyst Center have a real ceiling. Understanding where that ceiling is helps you decide when to keep using them and when a different tool is the right answer. Day 0 device onboarding Templates: enough?Yes Notes This is the core use case. Templates with PnP are well-suited to this workflow. Repeatable Day N config changes (VLAN, ACL, interface) Templates: enough?Yes Notes Exactly what templates are designed for. Site-wide rollout of a config change Templates: enough?Yes Notes Catalyst Center can apply a template to a device group. This scales well for common changes. Complex conditional logic based on live device state Templates: enough?No Notes Templates cannot query the device or make decisions based on current running state. Ansible or Python with a library like Netmiko handles this better. Stateful operations (e.g., orchestrated failover, traffic engineering) Templates: enough?No Notes Templates push config; they do not coordinate multi-device state machines. Integration with external data sources (IPAM, CMDB, ticketing) Templates: enough?Partially Notes You can call the Catalyst Center API to trigger template deployment with variables pulled from an external system, but the integration code lives outside Catalyst Center. Cross-platform automation (Cisco + Aruba + Juniper) Templates: enough?No Notes Templates are Catalyst Center and Cisco IOS-XE specific. Multi-vendor environments need Ansible, Nornir, or similar. Config drift detection and remediation Templates: enough?Partially Notes Catalyst Center has compliance checks, but complex drift detection and automated remediation at scale is better served by a dedicated tool. The graduation trigger is usually stateful complexity or cross-platform scope. If you are adding a VLAN to 80 campus switches, a Catalyst Center template applied to a device group is fast, auditable, and does not require you to maintain a Python environment. If you are implementing a configuration workflow that needs to read the current state of a device, make a decision based on that state, and then push a different configuration depending on what it finds, you need a proper automation tool. Ansible playbooks, Python with NAPALM or Netmiko, or Cisco's own NSO are appropriate at that point. The right framing is not "templates or automation frameworks." It is "use templates for what templates do well, and graduate to more powerful tooling when the workflow actually requires it." Many teams that graduate to Ansible still keep their Catalyst Center template library for Day 0 onboarding because the PnP workflow integrates naturally with the template system and does not need the additional complexity of Ansible for that use case. ## Key Takeaways - CLI templates in Catalyst Center are a pragmatic middle ground between fully manual operations and a full automation framework. They are well-suited to teams with moderate scripting maturity who need consistency without operational overhead. - Day 0 templates handle onboarding via PnP; Day N templates handle ongoing provisioning. Composite templates sequence multiple regular templates into a single operation. - Catalyst Center supports both Velocity and Jinja2\. For new template libraries, Jinja2 is the better choice. Consistency within your library matters more than which language you pick. - Invest time in variable design upfront. Descriptive names, enforced data types, and sensible defaults make templates easier to use and reduce errors at provisioning time. - Template sprawl is the biggest long-term risk. A small library of well-maintained, well-tested templates is far more valuable than a large library of inconsistent, untrusted ones. - Export templates to an external Git repository. Catalyst Center's internal versioning is not a substitute for source control you own. - Templates do not handle stateful operations, complex conditional logic based on live device state, or cross-vendor environments. When your requirements outgrow these limits, that is the right time to look at Ansible, Python, or NSO, not before. More from the Cisco deep-dive series: [Catalyst Center and ThousandEyes: What Problem Does This Actually Solve?](https://www.pinglabz.com/catalyst-center-and-thousandeyes-what-problem-does-this-actually-solve/). ### Post-Quantum Crypto in Cisco Campus Networks: What Operators Need to Know URL: https://www.pinglabz.com/post-quantum-crypto-in-cisco-campus-networks-what-operators-need-to-know/ Last updated: 2026-07-04T23:06:44.000Z Campus networks are getting pulled into the post-quantum conversation faster than most operators expected. NIST finalized three post-quantum cryptography (PQC) standards in August 2024, and Cisco has started making "quantum-ready" claims about the Catalyst 9000 lineup. Before you do anything with that information, it's worth separating what those claims actually mean in practice from the marketing framing, and building a realistic picture of where your campus network is exposed, what the timeline actually looks like, and what needs attention now versus what you can safely monitor for a few more years. ## Why PQC Is Entering the Campus Conversation Now Three things converged to make this a real planning concern rather than a theoretical one. **The NIST standards are finalized.** In August 2024, NIST published the first three post-quantum cryptography standards: FIPS 203 (ML-KEM, the key encapsulation mechanism derived from CRYSTALS-Kyber), FIPS 204 (ML-DSA, the signature algorithm from CRYSTALS-Dilithium), and FIPS 205 (SLH-DSA, from SPHINCS+). A fourth algorithm, FN-DSA (from FALCON), is pending finalization as FIPS 206\. A backup key encapsulation algorithm called HQC was selected in March 2025, with a draft standard expected in early 2026\. The absence of finalized standards was the primary reason most organizations weren't moving on PQC. That reason is gone. **The harvest-now-decrypt-later (HNDL) threat is already active.** Nation-state actors are capturing encrypted traffic today on the assumption that quantum computers will eventually be able to decrypt it. You don't need to believe Q-day is imminent to take this seriously. If data you're encrypting right now should still be confidential in ten to fifteen years, HNDL is a concrete risk. Expert surveys conducted by the Global Risk Institute suggest over 50% probability of cryptographically relevant quantum computers emerging within fifteen years. Research published in June 2025 by Craig Gidney reduced the estimated qubit requirement for breaking RSA-2048 from approximately 20 million to under one million superconducting qubits, which effectively pulled the threat timeline closer. The exact date remains genuinely uncertain; the direction of progress does not. **Long infrastructure lifecycles make campus networks unusually exposed.** This is the point most security briefings skip over. A campus switch bought today might stay in production for ten to fifteen years. A server bought today will likely be refreshed within five. The practical implication: traffic encrypted over campus infrastructure today could still be traversing that same infrastructure when a capable quantum computer exists, if current research projections hold. Server infrastructure adapts quickly through refresh cycles. Campus switch infrastructure often doesn't, unless you're intentionally planning for it. ## What Post-Quantum Actually Means for Network Crypto Before you evaluate any vendor PQC claim, you need a clear picture of where cryptography actually lives in a campus network. It isn't one thing. There are at least five distinct cryptographic planes in a typical enterprise campus, each with a different exposure profile and a different migration path. Layer 2 data ProtocolMACsec (IEEE 802.1AE) Current algorithm AES-GCM (128-bit or 256-bit symmetric) HNDL exposure Low at AES-256 (Grover's algorithm only halves effective key strength) PQC urgency Low near-term; AES-GCM-256 already provides adequate quantum resistance Layer 3 data / WAN ProtocolIPsec IKEv2 Current algorithm RSA or ECDH key exchange + AES-GCM bulk encryption HNDL exposure High (asymmetric key exchange is vulnerable to Shor's algorithm) PQC urgency High for HNDL-sensitive traffic; hybrid PQC IKEv2 available on Catalyst 8000 Authentication Protocol802.1X / EAP-TLS Current algorithm RSA or ECDSA certificates; TLS key exchange HNDL exposure Medium (key exchange in TLS handshake is retrospectively vulnerable) PQC urgency Medium; PQC TLS hybrid options exist but certificate migration is complex Management plane Protocol SSH, HTTPS, NETCONF, RESTCONF Current algorithmRSA/ECDSA keys + TLS HNDL exposure Medium (management sessions contain credentials and config data) PQC urgency Medium; depends on sensitivity of managed data Identity / PKI Protocol CA infrastructure, device certs, RADIUS certs Current algorithm RSA 2048/4096 or ECDSA P-256/P-384 HNDL exposure High (certificates issued today will outlast safe algorithm lifetimes) PQC urgency High; the single most actionable area for most operators right now The most important thing this table illustrates: MACsec's symmetric encryption is already quantum-resistant at the 256-bit key size. Grover's algorithm (the quantum attack that applies to symmetric ciphers) effectively halves the key length, meaning AES-256 behaves like AES-128 against a quantum adversary. AES-128 against a quantum computer is still secure by any practical definition. If you're already running GCM-AES-256 on your MACsec links, the data plane is not your exposure. Where the real exposure lives is in the asymmetric key exchange and certificate layers. RSA and ECDH (the algorithms used to negotiate session keys in IPsec, TLS, and SSH) are vulnerable to Shor's algorithm on a sufficiently capable quantum computer. Any session key negotiated using RSA or ECDH is retrospectively vulnerable: an adversary who captured that session today could decrypt it once quantum hardware arrives. This is the core of the HNDL threat, and it applies to every layer in the table above that uses an asymmetric key exchange. ## What Cisco Is Actually Claiming Cisco announced that the Catalyst 9000 Smart Switches support "full-stack post-quantum cryptography" as of IOS-XE 26\. This is a meaningful claim, but it covers multiple distinct things, and the level of maturity varies significantly by layer. PQC algorithms in power-up attestation chain Layer Secure boot / supply chain What it means in practice NIST-approved PQC verifies hardware authenticity at boot; prevents supply-chain firmware tampering Availability Available on new C9000 Smart Switch hardware with IOS-XE 26 Quantum-safe MACsec LayerMACsec (Layer 2) What it means in practice GCM-AES-256 with strong out-of-band PSKs provides symmetric key quantum resistance today; dynamic PQC key exchange for MKA is described as planned Availability Symmetric (PSK-based) resistance: available today on all Catalyst 9000 hardware running GCM-AES-256; dynamic PQC key agreement: in development, timeline subject to change PQC-enabled key exchange LayerIPsec (Layer 3) What it means in practice ML-KEM-1024 hybrid key exchange via IKEv2; validated against FIPS 203 Availability Catalyst 8000 (WAN/edge router platform); C9000 campus switch parity not confirmed with specific IOS-XE release PQC cipher suite support in HTTPS/NETCONF LayerManagement plane TLS What it means in practice TLS management sessions with hybrid PQC key agreement Availability IOS-XE 26 on new C9000 hardware; verify specific cipher suite availability in the release notes for your target platform before assuming The honest read of this table: the most concrete and shipping PQC capability on campus Catalyst hardware is the secure boot chain, which protects supply chain integrity rather than data in transit. That is a genuinely useful capability (hardware supply chain attacks are real), but it isn't protecting the MACsec or TLS sessions your operators are thinking about when they hear "quantum-safe campus." For MACsec, the "quantum-safe" claim today is largely about symmetric key size, not PQC key exchange. GCM-AES-256 with a properly random PSK is indeed adequate against quantum attacks on the data plane. But if your MACsec is using GCM-AES-128 (the default on many configurations), migrating to 256-bit keys is genuinely useful and unrelated to whether you buy new hardware. Dynamic PQC key agreement for MKA (the protocol that manages MACsec session keys) is described as in development, with timelines subject to change. Don't plan around features that aren't in a released configuration guide. For IPsec PQC key exchange, Cisco's ML-KEM-1024 support is shipping on the Catalyst 8000 (their WAN aggregation/edge router platform) and ISR 4000/ASR 1000 series. The campus campus switch parity question (specifically which Catalyst 9300/9400/9500 IOS-XE version carries which PQC feature) requires verification against the actual security configuration guide for your target release, not a marketing brief. **Ask the right questions when evaluating vendor claims.** "Do you support post-quantum cryptography?" will always get a yes. The useful questions are: which specific NIST algorithm (ML-KEM-768? ML-KEM-1024? ML-DSA-44?), in which IOS-XE version, on which hardware platforms, and is the key exchange quantum-resistant or just the bulk encryption cipher? A vendor who can't answer those questions specifically is selling positioning, not a feature. ## Crypto Agility: The Thing That Matters Most Before any specific algorithm claim, crypto agility is the most strategically important property to evaluate in any infrastructure you're procuring or renewing right now. Crypto agility means the ability to swap cryptographic algorithms through a software or configuration change, without replacing hardware. The reason this matters: the PQC standards landscape is still evolving. NIST finalized three algorithms in August 2024 and is still standardizing a fourth and a backup KEM. Some of the selected algorithms may be weakened by future cryptanalysis (as happened with SIDH/SIKE, a PQC candidate that was broken in 2022 before it could be standardized). The algorithms that are strong today may not be the algorithms you're running in 2035\. Infrastructure that can upgrade cryptographic primitives through IOS-XE software updates is in a fundamentally different position than infrastructure that requires a hardware swap to change cipher suites. When evaluating Cisco hardware for your next refresh, the question to ask is: if NIST releases a new or revised PQC standard in 2028, can this hardware run it with a software update, or does it require new ASICs? Cisco's positioning of IOS-XE as a software-upgradeable platform is designed to answer yes, but verify the specific feature against your platform and release. ## The Real Planning Horizon NIST IR 8547 (the transition guidance document, published as an initial public draft in November 2024) calls for deprecating quantum-vulnerable algorithms (RSA, ECDH, ECDSA) by 2030, and removing them from NIST standards by 2035\. Regulatory frameworks that reference NIST, including federal agency requirements and FIPS compliance, will follow this timeline with varying lead times. Ten years sounds like plenty of time. For most enterprises, the active work is limited for now. But two things compress the timeline in practice. **Certificate lifetimes create compounding complexity.** If you have a ten-year CA root certificate issued in 2022 (not uncommon in enterprise environments), it will still be signing RSA-based device certificates in 2032, two years after NIST's deprecation deadline. Migrating that CA after 2030 means touching every certificate it issued, across every device that trusts it. The cost of that migration is much lower if you plan it today than if you inherit it as an emergency in 2030. **Infrastructure procurement cycles extend past the compliance window.** If you're buying campus switches now on a typical 7 to 10 year lifecycle, those switches will be in production in 2033 or 2035\. Buying hardware today that is not evaluable for PQC agility means inheriting an unresolved compliance question during the enforcement window. Now (2025-2026) What changes NIST standards finalized; CNSA 2.0 compliance begins for federal/NSS; PQC TLS becoming default in browsers What to do now PKI audit; identify RSA certificate lifetimes; inventory MACsec key sizes; add PQC agility to procurement checklists; review vendor roadmaps for specific features 2027-2029 What changes PQC TLS mainstream in client/server software; regulatory frameworks tightening; hybrid certificates starting to appear What to do now Begin hybrid certificate migration for management plane TLS; update SSH host key types to minimum RSA-4096 or ECDSA P-384; validate IOS-XE PQC features against current configuration guides 2030 What changes NIST deprecation deadline for RSA/ECDH; NSA CNSA 2.0 migration deadline for national security systems; Australia's target What to do now PQC key exchange on new infrastructure; PQC or hybrid certificates for device authentication; audit legacy equipment for compensating controls 2031-2035 What changes NIST removal deadline; classical algorithms no longer in NIST standards; US federal and UK/EU regulatory deadlines What to do now All new infrastructure must support PQC; legacy systems require compensating controls or decommissioning justification ## Regulated Industries vs. Everyone Else The urgency is not uniform. Regulated industries have compliance timelines that are already creating real deadlines. Most enterprises have time to plan and monitor without immediate action on cryptographic infrastructure, as long as they're doing the foundational work. Federal / NSS (national security systems) Current pressure NSA CNSA 2.0 mandates; OMB M-23-02; FY2035 deadline in statute Act now on Full PQC migration underway; CNSA 2.0 timelines are active requirements Can monitor for now Nothing; compliance is not optional Defense industrial base / CMMC Current pressure CMMC 2.0 alignment with FIPS 140-3; CNSA 2.0 expectations from DoD customers Act now on PQC evaluation in next hardware procurement; PKI audit; IOS-XE version alignment with FIPS 140-3 validated modules Can monitor for now Hardware swap until your next refresh cycle Financial services Current pressure FFIEC guidance evolving; SEC cybersecurity rules; EU DORA (operational resilience) Act now on PKI audit; HNDL risk assessment for sensitive transaction data; vendor PQC roadmap review for exam readiness Can monitor for now Active PQC deployment until regulatory guidance firms up Healthcare Current pressure HIPAA crypto requirements will follow NIST; long patient data retention windows amplify HNDL risk Act now on Identify long-retention PHI protected with RSA/ECDH key exchange; PKI audit; update encryption policies to require AES-256 Can monitor for now Infrastructure changes until regulatory deadline is clearer State and local government Current pressure Varies by state; federal grant requirements increasingly reference NIST compliance Act now on PKI audit; monitor CISA and OMB guidance for requirements Can monitor for now Active deployment for now General enterprise Current pressure No current compliance pressure; business risk varies by data sensitivity Act now on PKI audit; add PQC agility question to procurement checklist Can monitor for now Everything except the PKI audit The PKI audit appears in every row for a reason. It's the foundational work that every organization can do today without buying anything, and without which you can't make informed decisions about anything else. You can't plan a certificate migration if you don't know what certificates you have, who issued them, what algorithms they use, and when they expire. ## Practical First Steps for Operators Here's the actual order of operations for a campus operator who wants to move from "aware" to "doing something useful." ### Step 1: Audit your PKI and certificates Identify every certificate in use on network infrastructure: management plane (HTTPS, SSH host keys), 802.1X (RADIUS server certificates, EAP-TLS client certificates), and any internal CAs issuing device identity certificates. Note key sizes, expiration dates, issuing CAs, and algorithm types. Any certificate with a lifetime extending past 2030 that uses RSA or ECDSA is a candidate for migration planning. ``` ! On Cisco IOS-XE, list all trustpoints and certificate details C9300# show crypto pki certificates verbose ! Check SSH host key type and size C9300# show crypto key mypubkey all ! Identify all trustpoints configured C9300# show crypto pki trustpoints ``` If you're running an internal CA (whether Microsoft AD CS or a dedicated PKI), export the certificate inventory and look specifically for root and intermediate CA certificates with long lifetimes that use RSA-2048\. These are the ones that will cause the most operational pain in 2028-2030 if you don't start migration planning now. ### Step 2: Verify your MACsec key sizes If you're running MACsec, check whether your MKA policies are using GCM-AES-128 or GCM-AES-256\. GCM-AES-128 has adequate classical security but is theoretically weakened by Grover's algorithm to approximately 64 bits of quantum security. GCM-AES-256 is effectively quantum-resistant for the data plane. Migrating to 256-bit keys is a configuration change, not a hardware change, and you can do it today. ``` ! Check active MACsec sessions and cipher suites in use C9300# show macsec summary C9300# show mka sessions detail ! Update MKA policy to require GCM-AES-256 C9300(config)# mka policy CAMPUS-PQC-READY C9300(config-mka-policy)# macsec-cipher-suite gcm-aes-256 C9300(config-mka-policy)# confidentiality-offset 0 ``` Note that GCM-AES-256 requires a Network Advantage license on Catalyst 9000 hardware. If you're currently on Network Essentials, verify licensing before pushing the change. Also, both ends of a MACsec link must agree on the cipher suite, so coordinate any change across all peers on a given segment. ### Step 3: Evaluate IOS-XE version and hardware capability on your next refresh Before your next campus switch procurement, add two questions to your evaluation criteria. First: which IOS-XE releases include verified PQC cipher suite support for management plane TLS and, when available, MACsec key exchange? Second: is the hardware capable of running those features without a significant performance impact? PQC algorithms (particularly ML-KEM and ML-DSA) are more compute-intensive than RSA and ECDH, and platforms without hardware acceleration for lattice-based operations may see CPU overhead on management plane operations at scale. For the Catalyst 9000 series, Cisco's "full-stack PQC" claims are associated with IOS-XE 26 on the new Smart Switch hardware announced at Cisco Live Amsterdam 2026\. Verify specific feature availability by checking the Security Configuration Guide for your target release and platform before building it into a deployment plan. ### Step 4: Add crypto agility to your configuration template standards Update your switch provisioning templates to default to AES-GCM-256 for MACsec, 4096-bit RSA or P-384 ECDSA for any certificates you generate, and SHA-384 or SHA-512 for hashing. These choices don't require PQC hardware, but they extend the safety margin of classical algorithms and make your infrastructure a cleaner starting point for the eventual PQC migration. They're also consistent with NIST's current guidance on classical algorithm key sizes while PQC deployment matures. ### Step 5: Ask specific vendor roadmap questions The useful questions for your Cisco account team or TAC are not "do you support PQC?" (they'll say yes) but rather: which specific algorithm (ML-KEM-768 or ML-KEM-1024?) is in which IOS-XE version on which hardware platform, what is the upgrade path if a future NIST revision changes the algorithm, and is the PQC implementation FIPS 140-3 validated? For environments with compliance requirements, the FIPS 140-3 module validation matters as much as the algorithm selection. ## Key Takeaways - The NIST PQC standards are finalized (FIPS 203, 204, 205 as of August 2024). The "waiting for standards" reason not to plan is gone. - Campus networks are more exposed to harvest-now-decrypt-later risk than server infrastructure. Switch lifecycles of 10 to 15 years overlap significantly with reasonable Q-day projections. - MACsec's symmetric data plane (AES-GCM-256) is already quantum-resistant. The key exchange layer, meaning IPsec IKEv2, TLS handshakes, and PKI, is where the real exposure is. - Cisco's "full-stack PQC" claim on the C9000 is most concrete for secure boot supply-chain integrity. Dynamic PQC key exchange for MACsec is in development; treat timelines as uncertain until specific features appear in a released IOS-XE configuration guide for your platform. - The most actionable near-term work for most operators is PKI and certificate lifecycle hygiene. Identify certificates with lifetimes extending past 2030 that use RSA or ECDSA, and start planning migration paths now. - Regulated industries (federal, defense industrial base, financial, healthcare) have compliance trajectories that make PQC planning active work. Most general enterprises should plan, do the PKI audit, and monitor. - Crypto agility, meaning the ability to swap cryptographic algorithms via software update rather than hardware replacement, is more strategically important than whether any specific product ships ML-KEM today. Ask about it in procurement. More from the Cisco deep-dive series: [What Is Cisco Catalyst Center? Key Features, Benefits, and Use Cases](https://www.pinglabz.com/what-is-cisco-catalyst-center/). ### Wi-Fi 7 for Cisco Shops: Upgrade Now or Wait? URL: https://www.pinglabz.com/wi-fi-7-for-cisco-shops-upgrade-now-or-wait/ Last updated: 2026-07-04T23:03:43.000Z Wi-Fi 7 is shipping, Cisco has a full lineup of 802.11be access points, and the industry is doing what it always does: pushing the upgrade conversation. The real question for most enterprise and campus wireless teams is not whether Wi-Fi 7 is technically better than Wi-Fi 6E (it is), but whether a given organization has the environment, the supporting infrastructure, and the refresh timing to justify spending the money now. This article is a practical guide for wireless architects and campus IT leaders who need to make that call without a marketing deck as the primary input. (This article is part of the PingLabz Cisco wireless series - the [full Cisco wireless guide](https://www.pinglabz.com/wireless/) maps the whole cluster in reading order.) ## Wi-Fi 7 in Plain English Wi-Fi 7 is the IEEE 802.11be amendment. It operates on the same three frequency bands as Wi-Fi 6E (2.4 GHz, 5 GHz, and 6 GHz), so the spectrum itself is not new. What 802.11be adds is a set of physical and MAC layer improvements designed to squeeze more throughput, lower latency, and better interference resilience out of those bands simultaneously rather than sequentially. The changes that actually matter for enterprise deployments come down to four things: Multi-Link Operation (MLO), 320 MHz channel support, Multi-Resource Unit (MRU) scheduling, and 4K-QAM modulation. Each of these is covered below, but the short version is this: Wi-Fi 7 is not faster the way Wi-Fi 4 to Wi-Fi 5 was faster. The gains are more about reliability and consistency under load, which matters more in a dense campus lecture hall than in a conference room with six people on a video call. ## What Improvements Matter in Practice ### Multi-Link Operation (MLO) MLO is the headline feature of Wi-Fi 7\. It allows a client and an access point to maintain associations and exchange data across multiple frequency bands simultaneously rather than committing to one band for the life of a session. A laptop with an MLO-capable Wi-Fi 7 adapter can be actively using the 5 GHz and 6 GHz links at the same time, load-balancing traffic across them or failing over instantly when one band degrades. There are two MLO modes in practice. STR (Simultaneous Transmit and Receive) is the full version: the device transmits on one link while receiving on another simultaneously. It requires sophisticated RF isolation between radios and is where the real performance gains live. eMLSR (Enhanced Multi-Link Single Radio) is what most current client devices actually use: the adapter can switch between links very quickly but does not transmit and receive on two links at the exact same moment. Most Wi-Fi 7 laptops and smartphones shipping today implement eMLSR, not STR. For users in high-density environments, MLO means less visible congestion: when the 5 GHz band is saturated in a packed auditorium, the link can shift load to 6 GHz dynamically without the client dropping and re-associating. That translates to fewer mid-meeting quality drops and smoother roaming. What it does not translate to is doubling your application throughput, because most enterprise applications are not bottlenecked by the wireless link to begin with. ### 320 MHz Channels Wi-Fi 7 extends maximum channel width to 320 MHz in the 6 GHz band (the 5 GHz band tops out at 160 MHz, same as Wi-Fi 6E). Wider channels equal more throughput per spatial stream. In theory, a 320 MHz channel roughly doubles the throughput of a 160 MHz channel on the same radio. In practice, 320 MHz channels are only useful when you have enough contiguous 6 GHz spectrum that is actually clean. In multi-tenant office buildings, dense urban campuses, or any environment where neighboring networks are visible, the available uncontested 6 GHz channels are narrower than the spec allows. Regulatory availability also varies: the U.S. has 1200 MHz of 6 GHz spectrum, which can theoretically support two to three non-overlapping 320 MHz channels, but you need AFC (Automated Frequency Coordination) for standard power operation outdoors, and indoor low-power operation limits the range at which those channels are viable. 320 MHz is a benefit in the right conditions, not a universal upgrade. ### Multi-Resource Unit (MRU) Scheduling OFDMA (introduced in Wi-Fi 6) allows an AP to split a channel into smaller Resource Units (RUs) and schedule multiple clients in parallel within a single transmission. MRU extends this by letting the AP allocate non-contiguous and non-uniform RU combinations to a single client. The practical effect is better spectrum efficiency when the AP is managing a mix of client traffic types, particularly in environments with many low-bandwidth IoT or background devices alongside high-bandwidth clients. Users do not notice MRU directly; it shows up as better aggregate performance under mixed-client load. ### 4K-QAM Modulation Wi-Fi 6 topped out at 1024-QAM. Wi-Fi 7 adds 4096-QAM (4K-QAM), which encodes 12 bits per symbol versus 10 bits, a roughly 20 percent improvement in spectral efficiency at close range and strong signal. The requirement is a very high SNR, which means this benefit applies primarily to clients within 10 to 15 feet of the AP with a clear line of sight. For most enterprise deployments designed around coverage and mobility rather than fixed near-AP workstations, 4K-QAM contributes to peak rate headlines more than to average user experience. ### What Users Will Actually Notice The Wi-Fi 7 improvements users will notice are lower latency and more consistent throughput in crowded environments, smoother roaming between APs in high-density areas, and fewer quality drops on video conferencing when the room fills up. The improvements that mostly stay invisible to users are aggregate throughput gains (limited by the wired backhaul and application architecture), 320 MHz channel benefits (limited by spectrum availability and regulatory rules), and 4K-QAM gains (limited by SNR requirements). Do not promise users that their downloads will be twice as fast; that is not what this upgrade delivers for most of them. ## Environments Likely to Benefit First Not all environments get the same return from a Wi-Fi 7 upgrade. The clearest wins are in environments where density, latency sensitivity, or aggregate wireless load is already pushing the limits of Wi-Fi 6 or 6E. University lecture halls and arenas Wi-Fi 7 Benefit LevelHigh Primary Driver MLO, MRU under high client density Notes Georgetown's large-scale Wi-Fi 7 rollout (announced early 2026) reflects this use case directly; 500+ clients per venue is where MLO earns its cost Enterprise open-plan offices (1,000+ employees per floor) Wi-Fi 7 Benefit LevelHigh Primary Driver MLO consistency, reduced band steering complexity Notes High client density with mixed traffic types (video, VoIP, background sync) is the right fit for MLO and MRU Convention centers and large event venues Wi-Fi 7 Benefit LevelHigh Primary Driver Peak-load throughput, 6 GHz availability Notes Temporary high-density events where 6 GHz adds headroom Hospital and clinical environments Wi-Fi 7 Benefit LevelMedium-High Primary Driver Latency and roaming reliability for clinical apps Notes Latency-sensitive monitoring and telemetry apps benefit; inventory scanning and RTLS less so K-12 schools and small colleges Wi-Fi 7 Benefit LevelMedium Primary Driver Client density in classrooms Notes Classrooms benefit from density improvements; hallways and admin offices less so Standard corporate offices (low-medium density) Wi-Fi 7 Benefit LevelLow-Medium Primary Driver Future-proofing, client device refresh alignment Notes Current Wi-Fi 6 performance is adequate for most workloads at this density; upgrade timing should follow lifecycle Warehouses and distribution centers Wi-Fi 7 Benefit LevelLow Primary Driver Minimal; coverage-limited, not capacity-limited Notes Wi-Fi 6 or even Wi-Fi 5 is usually sufficient; RF propagation and coverage are the real challenges here Retail stores (small to medium) Wi-Fi 7 Benefit LevelLow Primary Driver Very low client density per AP Notes POS and inventory scanning workloads do not benefit from Wi-Fi 7 features ## Backhaul, Switching, Power, and Management Implications This is the section that tends to get glossed over in vendor conversations. The wireless upgrade is only one layer of the decision. Cisco's Wi-Fi 7 APs require infrastructure that many campuses running Wi-Fi 6 or earlier do not yet have, and the gap between "AP cost" and "total upgrade cost" can be significant. ### Uplink Speed and Cabling Cisco's Wi-Fi 7 AP lineup introduces multi-gig uplinks that older switching infrastructure cannot support: CW9171I Uplink Port1x 2.5 GbE Target Environment Standard coverage, low-to-medium density Notes 802.3at (PoE+) sufficient for full functionality; 802.3bt unlocks USB CW9172I Uplink Port1x 2.5 GbE Target Environment Medium-to-high density indoor Notes 30W nominal; penta-radio architecture; 802.3at sufficient CW9176I Uplink Port1x 10 GbE Target Environment Ultra-high density (lecture halls, arenas) Notes Requires 802.3bt (PoE++) or UPOE; 10G uplink means 10G access switching is required The uplink speed matters because Wi-Fi 7's aggregate radio capacity can now exceed what a 1G Ethernet port can deliver. A CW9172I pushing traffic from 5 GHz, 6 GHz, and a second 5 GHz radio simultaneously can theoretically saturate 1 Gbps under load in a dense environment. If your access switching is all 1G downlinks with 1G AP ports, you are paying for radio capability you cannot extract. The minimum practical recommendation for a Wi-Fi 7 deployment is 2.5 GbE per AP port on the access switch, with 5 GbE in dense zones. The 9176I requires a 10G-capable access switch at the edge. Cabling matters too, and it is often the longest lead-time item in a campus upgrade. Cat 5e is physically capped at 1 Gbps regardless of switch port speed. Achieving 2.5 GbE over copper requires Cat 6 at a minimum (and Cat 6A for 10G). If your campus is on Cat 5e, a Wi-Fi 7 AP upgrade without a cabling remediation plan means the backhaul bottleneck moves from the radio to the cable plant. ### PoE Power Budget Most Cisco Wi-Fi 6 APs (9115, 9120, 9130) operate at 15 to 25W, well within 802.3at (PoE+, 30W budget). The CW9171I and CW9172I maintain that compatibility: both run at approximately 30W and are fully functional on PoE+ ports. The CW9176I is the exception, requiring 802.3bt Class 6 or UPOE for full operation, which means 60W switch ports. If your Catalyst 9300 or 9200 switches are provisioned with older PoE power supplies, you may need power supply upgrades or PoE injectors as a bridge, neither of which is free or simple at scale. The practical rule: plan for 802.3bt-capable ports any time you deploy the 9176I. For 9171I and 9172I deployments, existing PoE+ infrastructure is sufficient, which significantly lowers the switching dependency if you are not deploying the high-density flagship. ### Catalyst Center and Software Requirements Wi-Fi 7 APs (including the CW9171I, CW9172I, and CW9176I) require IOS-XE 17.15.3 or later on the C9800 controller, and Catalyst Center 2.3.7.x or later for full management visibility and policy support. If your Catalyst Center instance is behind on version (which is common in organizations that defer platform upgrades), the Wi-Fi 7 AP rollout creates a forcing function: you need to bring Catalyst Center current before or alongside the AP deployment. That is not a one-afternoon task on a production system, and it should be scoped as a separate project track, not an assumption in the AP refresh timeline. Additionally, some of the more compelling management features around Wi-Fi 7 (AI-driven radio resource management, MLO policy visibility, and the analytics that let you see whether MLO is actually providing benefit in your environment) are Cisco Spaces or Catalyst Center cloud-connected features. If you are running Catalyst Center in an air-gapped or limited-cloud configuration, some of the operational return on the Wi-Fi 7 investment is harder to capture. ## Reasons to Upgrade Now The case for upgrading now is strongest when several of the following conditions apply simultaneously, not just one: - **Your current APs are approaching or past their normal lifecycle.** Cisco's typical AP lifecycle is five to seven years. If you are running 9115s or 9120s purchased in 2019 or 2020, you are approaching the window where a refresh is justified on lifecycle grounds alone. Wi-Fi 7 is the logical refresh target rather than a mid-cycle replacement. - **You have high-density, latency-sensitive environments.** Lecture halls, auditoriums, large event spaces, and open-plan offices with 50 or more clients per AP are the environments where MLO and MRU translate to user-visible improvement. If that describes your primary campus environments, the ROI calculation looks different than for a distribution center or a small branch. - **Your access switching is already multi-gig capable (or is also being refreshed).** If your Catalyst 9300 switches have 2.5G or 5G downlinks, you can deploy CW9171I or CW9172I APs without a switching upgrade. If you are also replacing access switches in the same window, bundle the projects: a combined switching and AP refresh that delivers a 2.5G-capable edge for Wi-Fi 7 backhaul is more defensible than two separate budget requests. - **Your cabling plant is Cat 6 or Cat 6A.** This removes the backhaul bottleneck that would otherwise limit what Wi-Fi 7 can deliver. - **Your budget cycle and capital planning align.** Wi-Fi refresh projects often have long procurement and deployment cycles. If the budget window is now, it is worth deploying the current generation rather than buying Wi-Fi 6E with a one-to-two year shelf life. - **Your client device population is modern.** A campus where most devices are Windows 11 laptops from 2023 and 2024 (many of which include Wi-Fi 7 adapters) gets more immediate return than a campus dominated by older devices that cap out at Wi-Fi 6. ## Reasons to Wait There are equally valid cases for holding the current generation. The upgrade is not urgent in the following situations: - **Your current APs are two to four years into their lifecycle.** Replacing functional, in-lifecycle Wi-Fi 6 APs with Wi-Fi 7 to chase throughput numbers that most of your applications do not need is expensive and hard to justify to a CFO. Lifecycle-driven refresh is the right trigger. - **Your access switching is all 1G PoE+.** Deploying Wi-Fi 7 APs behind 1G switch ports is a backhaul-constrained deployment: you are paying for radio capability you cannot deliver. The upgrade should include the access switching layer, and that changes the budget conversation substantially. - **Your cabling is Cat 5e throughout.** Same constraint as the switching: the radio improvement is wasted if the cable plant cannot carry multi-gig traffic. A cable remediation project is a significant capital and labor investment that extends the timeline. - **Your Catalyst Center is significantly behind on version.** Jumping from Catalyst Center 2.2.x or earlier to 2.3.7 in parallel with a large AP rollout creates operational risk. Sequence the platform upgrade first, validate stability, then move the AP refresh forward. - **Your environment is low to medium density with no latency-sensitive wireless workloads.** Administrative buildings, warehouses, retail stores, and small offices running standard productivity applications are not bottlenecked by Wi-Fi 6\. They will not see meaningful user-visible improvement from Wi-Fi 7 today. - **Your client devices are predominantly Wi-Fi 5 or Wi-Fi 6 (no MLO support).** MLO requires Wi-Fi 7 clients. If 80 percent of your endpoint fleet is on Wi-Fi 5 or Wi-Fi 6 adapters, the MLO benefit is deferred until your endpoint refresh catches up. Wi-Fi 7 APs are backward compatible, so your existing clients still work, but you are not getting the feature you are primarily paying for. ## Phased Migration Model and Decision Matrix ### A Sensible Phased Approach For organizations that have a mix of environments and refresh timelines across their portfolio, a phased approach avoids the pressure of a full forklift while capturing Wi-Fi 7 benefit where it matters most: **Phase 1: High-density pilot zones (now).** Identify two or three of your highest-density, highest-visibility wireless environments: the main auditorium, a flagship classroom building, the largest open-plan office floor. Deploy CW9172I or CW9176I APs in these areas first, with appropriate switching and cabling remediation scoped as part of the project. This gives you operational experience with Wi-Fi 7 and C9800 17.15.x firmware before scaling, and it gives leadership a visible proof point. **Phase 2: Lifecycle-triggered replacement (rolling, 12 to 24 months).** As existing APs hit five to seven years of age, replace them with Wi-Fi 7 models as a matter of course. Do not replace in-lifecycle APs just to standardize on Wi-Fi 7 faster than the hardware justifies. **Phase 3: Remaining estate on normal refresh cycle.** Low-density coverage areas (warehouses, retail, small branches) follow their own refresh clock. These may still be Wi-Fi 6 or Wi-Fi 6E deployments for another three to five years, and that is fine. Wi-Fi 7 APs are backward compatible; you can run a mixed estate on the C9800 without architectural problems. ### Decision Matrix: Upgrade Now or Wait? Large higher ed campus, lecture halls and stadiums Current AP Lifecycle5+ years old Environment Type High density, latency-sensitive Switching/Cabling Multi-gig switches available or in refresh scope Recommendation **Upgrade now.** This is the environment Wi-Fi 7 is designed for. MLO and MRU will deliver user-visible improvement. Large higher ed campus, lecture halls and stadiums Current AP Lifecycle 2-4 years old (Wi-Fi 6E) Environment Type High density, latency-sensitive Switching/Cabling Multi-gig switches available Recommendation **Pilot high-density areas only.** Deploy Wi-Fi 7 in peak-density venues; hold remaining Wi-Fi 6E estate through lifecycle. Enterprise open-plan campus, 500-2,000 employees Current AP Lifecycle5+ years old Environment Type Medium-high density, video-heavy Switching/Cabling 1G PoE+ switching (aging) Recommendation **Bundle refresh.** Upgrade APs and access switching together. Deploying Wi-Fi 7 APs behind 1G switches wastes the investment. Enterprise open-plan campus, 500-2,000 employees Current AP Lifecycle 3-5 years old (Wi-Fi 6) Environment Type Medium density, standard productivity Switching/Cabling1G PoE+ switching Recommendation **Wait.** Current Wi-Fi 6 is adequate. Plan a combined AP and switching refresh in 2 to 3 years when lifecycle and budget align. Hospital or clinical campus Current AP Lifecycle5+ years old Environment Type Medium density, latency-sensitive apps Switching/CablingMixed (some multi-gig) Recommendation **Upgrade clinical zones now.** Target OR suites, nursing stations, and patient monitoring areas. Hold administrative areas on lifecycle. K-12 district, multiple buildings Current AP Lifecycle4+ years old Environment Type Medium density (classrooms) Switching/CablingMostly 1G PoE+ Recommendation **Plan carefully.** Classroom density justifies Wi-Fi 7, but switching gaps need scoping. Consider Cat 6A cabling assessment before committing. Warehouse or distribution center Current AP LifecycleAny Environment Type Low density, coverage-limited Switching/CablingAny Recommendation **Wait or skip generation.** Wi-Fi 6 or Wi-Fi 6E is sufficient. Upgrade on lifecycle only; no performance benefit to accelerate. Small or medium branch offices Current AP LifecycleAny Environment TypeLow-medium density Switching/Cabling1G PoE+ Recommendation **Wait.** Current generation is adequate. Wi-Fi 7 models will be broadly available and competitively priced in 2 to 3 years when these sites hit refresh. New greenfield campus build Current AP LifecycleN/A (new) Environment TypeAny Switching/CablingDesigning fresh Recommendation **Deploy Wi-Fi 7 as baseline.** Specify Cat 6A throughout, 2.5G or 5G PoE+ access switching, and C9800 with current firmware from day one. Do not design for Wi-Fi 6 in a greenfield build in 2026. ### A Note on Vendor Timing Pressure Cisco and its partners have a clear interest in accelerating Wi-Fi 7 adoption. The campus networking refresh cycle is a meaningful revenue driver, and you will hear compelling arguments about AI traffic growth, the need for 6 GHz capacity, and client device refresh curves all trending toward Wi-Fi 7 now. Those arguments are not wrong, but they are also not universal. The right answer depends on your specific environment, your infrastructure gaps, and your refresh economics, not on a market forecast. Use the decision matrix above as a starting point, not a vendor's roadmap presentation. ## Key Takeaways - Wi-Fi 7's most impactful improvements (MLO, MRU) show up as better reliability and consistency under load, not as raw throughput gains for individual users. High-density, latency-sensitive environments benefit most. - The wireless AP is one layer of the decision. Access switching (2.5G or 5G ports), cabling (Cat 6A for multi-gig), PoE power budget (802.3bt for the CW9176I), and Catalyst Center version (2.3.7.x minimum) all have to line up. Scope all of them before committing to a timeline. - Lifecycle alignment is the most defensible upgrade trigger. Replacing in-lifecycle Wi-Fi 6 APs for Wi-Fi 7 is a hard budget justification unless the environment clearly warrants it. Refreshing APs that are five to seven years old and specifying Wi-Fi 7 as the current generation is straightforward. - MLO benefits require Wi-Fi 7 clients on both ends. If your endpoint fleet is predominantly Wi-Fi 5 or Wi-Fi 6 devices, the MLO return is deferred until your endpoint refresh catches up. - Greenfield builds in 2026 should specify Wi-Fi 7 as the baseline. Designing for Wi-Fi 6 in a new building today means you are designing for a previous generation on a 5 to 7 year infrastructure cycle. - A phased approach is valid and sensible. Pilot in high-density venues, let the rest follow lifecycle, and avoid treating wireless generations as mandatory vanity upgrades on a fixed clock. ### Catalyst Center and ThousandEyes: What Problem Does This Actually Solve? URL: https://www.pinglabz.com/catalyst-center-and-thousandeyes-what-problem-does-this-actually-solve/ Last updated: 2026-07-04T23:06:42.000Z When a user calls the help desk and says "the network is slow," someone on your team starts digging. You pull up Catalyst Center, check client health, scan the device 360 view for the access switch they're on, review wireless RF metrics if they're on Wi-Fi, and see... green. Everything is green. The switch is healthy, the uplink is healthy, the WLC shows acceptable SNR and a clean client association. And yet the user can't load Salesforce. Now what? This is the gap that the Catalyst Center and ThousandEyes integration is designed to close. Catalyst Center is very good at telling you the state of your infrastructure: the devices you own, the interfaces you manage, the clients connected to your fabric. ThousandEyes is very good at telling you what happens to a packet after it leaves your administrative domain: ISP routing, internet paths, CDN performance, SaaS endpoint reachability. Neither tool alone gives you the full picture. Together, they let you stop guessing and start proving. This article is not a feature tour. It is an operations article about the specific problem this combination solves, how the troubleshooting workflow actually changes, and where the gaps and deployment realities are. If you are evaluating whether this integration is worth the operational investment, this is what you need to know. ## The Classic Blame Game Any network operations team that supports SaaS-heavy environments knows the pattern. A business unit complains that an application is slow or intermittently failing. You check your infrastructure: devices are up, interfaces are clean, no drops, no errors. You check the application team: their servers look healthy. You check the ISP: they show no outage on their portal. Nobody can point to anything broken, and the user is still having a problem. The blame cycle then goes something like this. Network says it is not their infrastructure. App says the app is fine and it must be the network. ISP says there are no incidents on their network. Cloud provider's status page is green. Meanwhile, the user's experience is degraded, nobody has concrete evidence of where the fault actually is, and the incident drags on longer than it needs to. The root cause of this dynamic is not incompetence, it is a visibility boundary problem. Every team can see their own domain clearly. Nobody can see the full path between the user and the application. That path often crosses at least three administrative domains: your campus or branch network, one or more ISP networks, and the cloud or SaaS provider's infrastructure. Traditional monitoring tools stop at the edge of what you own. Access switch / AP Who owns itYou Visible to traditional NMS?Yes Visible to Catalyst Center?Yes Visible to ThousandEyes? Partial (agent on switch sees LAN-side) Distribution / core Who owns itYou Visible to traditional NMS?Yes Visible to Catalyst Center?Yes Visible to ThousandEyes? Yes (if agent placed there) Branch WAN edge / SD-WAN Who owns itYou Visible to traditional NMS?Varies Visible to Catalyst Center? Partial (with SD-WAN integration) Visible to ThousandEyes?Yes ISP / transit network Who owns itCarrier Visible to traditional NMS?No Visible to Catalyst Center?No Visible to ThousandEyes? Yes (BGP path, hops, loss) Internet backbone / CDN Who owns itThird party Visible to traditional NMS?No Visible to Catalyst Center?No Visible to ThousandEyes?Yes SaaS / cloud endpoint Who owns itVendor Visible to traditional NMS?No Visible to Catalyst Center?No Visible to ThousandEyes? Yes (synthetic tests, response time, availability) The problem with trying to diagnose a user experience issue without internet-path visibility is that you cannot tell whether the issue is in your domain or outside it. You can rule things out, but you cannot prove fault location. ThousandEyes shifts that from "we think it is not us" to "we can show you exactly where the degradation starts." ## What Catalyst Center Sees Well Catalyst Center (formerly DNA Center) is purpose-built for managing and assuring Cisco campus and branch infrastructure. For that scope, it is genuinely powerful. If you are on a Catalyst 9000 switching fabric, Catalyst 9800 wireless, and ISE for policy, Catalyst Center gives you a level of operational context that raw SNMP polling and syslog analysis cannot match. Device health CPU, memory, interface utilization, hardware fault events. Per device, with historical trends and AI-driven anomaly baselines Client health RSSI, SNR, roaming events, DHCP and AAA timers, onboarding flow analysis, client 360 timeline Fabric assurance SD-Access fabric node health, control plane (LISP) status, underlay/overlay correlation, fabric edge to border path status Network topology Auto-discovered topology with health overlays; you can click through to any node and see its neighbors AI-driven insights Proactive issue detection based on learned baselines (not just threshold alerts). For example, detecting onboarding failure spikes before ticket volume rises Path trace Trace the actual switched/routed path from a client to a destination, showing the specific interfaces and devices involved Sensor testing Dedicated network assurance sensors (physical or virtual) that test onboarding flows, RADIUS, DNS, gateway reachability, and application URLs from specific locations in your network The key phrase in that table is "from specific locations in your network." Catalyst Center's path trace and sensor tests operate within your infrastructure. Path trace gives you a hop-by-hop view from client to the edge of your domain. Sensor tests validate the onboarding experience and can reach URLs, but what happens between your DNS resolver's response and the application's CDN endpoint is not something Catalyst Center can tell you directly. Catalyst Center's blind spot is not a deficiency in the product. It is a structural limitation: it can only instrument and measure what it manages. Once a packet leaves your organization's network boundary, Catalyst Center cannot follow it. ## What ThousandEyes Sees Well ThousandEyes is an internet and application path intelligence platform. Its core function is running synthetic tests from a network of vantage points, Enterprise Agents (deployed on your infrastructure), Cloud Agents (deployed in Cisco-managed cloud POPs globally), and Endpoint Agents (installed on user devices). Tests run continuously on configurable intervals and generate a full dataset: round-trip latency, packet loss, path trace with hop-by-hop loss and latency, BGP routing visibility, and application-layer metrics for HTTP, DNS, FTP, and others. Internet path visibility Hop-by-hop path trace from your agent to the destination, with loss and latency at each hop, including inside ISP and transit networks you do not own BGP route monitoring BGP prefix visibility from hundreds of collector vantage points worldwide; detects prefix hijacks, route leaks, and unexpected path changes SaaS monitoring Pre-built tests for Webex, Microsoft 365, Salesforce, Zoom, and others. Measures availability, page load, response time from multiple Cloud Agent locations Cloud path visibility Tests from Enterprise Agents to AWS, Azure, GCP endpoints, with visibility into whether degradation is on-path to the cloud or inside the cloud provider's network Endpoint Agent data User-perspective measurements from Windows and macOS clients. Includes local network metrics (gateway loss, DNS resolver timing) plus path to application Waterfall and page load analysis HTTP test type captures full browser-equivalent page load, showing object-level load times. Useful for isolating whether an application is slow due to a specific resource (CDN, API, asset) Alert correlation Timeline-based correlation of path changes, BGP events, and performance degradation. Can show that an ISP route flap preceded an application slowdown by 45 seconds ThousandEyes does not instrument your devices the way Catalyst Center does. It does not know that the CPU on your distribution switch is high, or that a specific client failed RADIUS authentication. What it knows is whether a packet from your campus (or from any of its Cloud Agents globally) can reach a destination, what path it takes, and whether that path is performing normally. Those two data sets are complementary, not overlapping. One thing worth noting: ThousandEyes is only as useful as the tests you configure. The platform gives you the infrastructure to run tests, but you have to decide what to test, at what interval, from which agents, and you have to set up alerts and dashboards. A deployment with no configured tests generates no operational value. This is relevant when thinking about adoption cost, which we will come back to. ## Why Combining Them Changes the Troubleshooting Workflow The individual value of each tool is well understood by most operators who have used them. The integration value is more specific: it is about reducing context-switching time during an active incident. When a user calls with a connectivity or performance problem, the typical workflow without integration goes something like this. You open Catalyst Center, pull up the client's health view, check device health, maybe run a path trace. You see the infrastructure is fine. You then open a browser tab to the ThousandEyes portal, navigate to the relevant test (assuming you configured one), and correlate the timeline manually with what you just saw in Catalyst Center. If the timestamps don't obviously align, you have to reason about the relationship between the two data sets in your head. With the integration, ThousandEyes data surfaces inside Catalyst Center's Assurance section. You can start in Catalyst Center for the client and device health context, then see ThousandEyes test results for the relevant applications alongside that data, without switching portals. The operational benefit is not magical, you are still looking at the same data, but reducing the context switch cuts the cognitive overhead during troubleshooting. You see device health degradation and internet path degradation on the same timeline, which makes the correlation question much easier to answer. The most important shift is philosophical. The question changes from "is the network down?" to "where exactly is the problem, and whose domain is it in?" That is a much more useful question, and it is one you can now answer with evidence rather than inference. Is the client associated and healthy? Catalyst Center aloneYes ThousandEyes aloneNo CombinedYes (CC) Is the access switch / AP healthy? Catalyst Center aloneYes ThousandEyes aloneNo CombinedYes (CC) Is the path through your network clean? Catalyst Center aloneYes (path trace) ThousandEyes alone Partial (agent-to-agent) CombinedYes (CC) Is the ISP path clean? Catalyst Center aloneNo ThousandEyes aloneYes CombinedYes (TE) Is the SaaS endpoint reachable and performing normally? Catalyst Center aloneNo ThousandEyes aloneYes CombinedYes (TE) Did a BGP route change correlate with the outage? Catalyst Center aloneNo ThousandEyes aloneYes CombinedYes (TE) Is this affecting all users or just this one? Catalyst Center alone Yes (client health scope) ThousandEyes alone Yes (multi-agent scope) CombinedYes (both) Can I prove to the ISP that the problem is in their network? Catalyst Center aloneNo ThousandEyes aloneYes (path trace data) CombinedYes (TE) ## Example Operational Scenarios ### Scenario 1: Salesforce is slow on Monday morning **Without ThousandEyes:** The sales team's Salesforce instance is slow. Network checks out green in Catalyst Center. Client health is normal. You contact Salesforce support and open a case. Salesforce support says their systems look healthy. Incident sits unresolved for several hours while people argue about who owns the problem. Eventually performance recovers on its own, and nobody knows why. **With ThousandEyes:** ThousandEyes has been running a synthetic HTTP test to Salesforce from your branch agents every two minutes. You pull up the test timeline and see that round-trip latency spiked from 38ms to 310ms at 8:07 AM, with the hop-by-hop path trace showing the latency appearing at an ISP transit hop two hops past your WAN edge. BGP monitoring shows a route change for the Salesforce prefix at 8:05 AM. You take a screenshot of the path trace with the latency highlighted at the ISP hop, attach it to the support case, and escalate to your carrier. You are no longer arguing about whose problem it is, you have evidence of where the degradation is occurring. Resolution time drops from hours to minutes (at least from your side: you know it is not your network). ### Scenario 2: Branch users can't join Webex calls **Without ThousandEyes:** Three branches are reporting Webex audio quality problems during calls. Your NOC checks the WAN links and sees utilization is within normal range. Catalyst Center shows no device issues at the branch. You ask branch users to restart their clients. Problem persists. **With ThousandEyes:** Your Enterprise Agents on the branch access switches have been running a Webex test every minute. You see that all three affected branches (and only those three) are experiencing packet loss of 3-7% to the Webex media nodes, while a fourth branch on a different ISP circuit is unaffected. The path traces for all three affected branches share a common upstream hop where the loss is introduced. You now know this is an ISP issue affecting a specific circuit or POP, not a Webex platform problem and not your branch infrastructure. You escalate to your carrier with specific evidence. The NOC at the unaffected branch does not need to spend time investigating since you have already scoped it. ### Scenario 3: Intermittent DNS failures in a specific building **Without ThousandEyes:** Users in one building report intermittent application failures. Client health in Catalyst Center is mostly clean, with some occasional onboarding events. Nothing conclusive. **With ThousandEyes:** Your Enterprise Agent on the access switch in that building is running DNS tests to your internal resolvers and to public DNS. You notice the internal DNS resolver test is showing occasional 800ms+ response times (normally sub-5ms). That narrows the problem immediately: something is wrong with the DNS resolver path from that VLAN, not with the internet or applications. You look at the path trace to the resolver and find a specific hop where latency balloons occasionally. Catalyst Center confirms no device health issues, so you dig into QoS policy on the uplink, turns out DNS traffic is being deprioritized by a misconfigured DSCP policy. Found and fixed in under an hour because ThousandEyes gave you the specific symptom (DNS latency from that agent) to start from. ## Deployment Prerequisites and Realities Before planning a deployment, there are several things you need to have in place. The integration is not plug-and-play on arbitrary hardware or licensing. ### Licensing ThousandEyes Enterprise Agent deployment on Catalyst 9000 switches requires Cisco DNA Advantage or Premier licensing on each switch. DNA Advantage includes one ThousandEyes Enterprise Agent license per switch (consumed as 22 unit-months of ThousandEyes units). If you are on DNA Essentials, you will need to either upgrade or purchase ThousandEyes licenses separately. This is a non-trivial cost consideration for large deployments, if you have 500 access switches and want agents on all of them, that is 500 DNA Advantage licenses. Many environments will already have DNA Advantage for the broader assurance features; in that case, the ThousandEyes inclusion makes the marginal cost of agents essentially zero. ### Hardware and Software Requirements Enterprise Agents run as Docker containers using the App-hosting framework built into the Catalyst 9300 and 9400 series platforms. The AppGigabitEthernet port must be configured as a trunk (not access mode) to preserve VLAN tagging for the agent's network interface. Minimum IOS-XE versions are 17.3.3 for Catalyst 9300 and 17.5.1 for Catalyst 9400\. Check the ThousandEyes support matrix for your specific platform and release before planning rollout, not all hardware variants are supported, and version requirements can be specific. ``` ! Minimum config for AppGig interface (check your platform docs) interface AppGigabitEthernet1/0/1 switchport mode trunk switchport trunk allowed vlan 10,20,99 ! ``` The agent container needs a routable IP address on each VLAN it will test from, with outbound HTTPS to ThousandEyes cloud endpoints. If your campus has restrictive egress filtering, you need to allow the ThousandEyes agent communication IP ranges before deployment, agents that cannot phone home to the ThousandEyes cloud will fail to register and will not generate test data. ### The Deployment Workflow Catalyst Center handles the agent push through its App Hosting for Switches workflow. You upload the ThousandEyes Docker TAR file (downloaded from the ThousandEyes portal), run a readiness check per switch that validates HTTPS reachability, authentication, and VLAN configuration, then deploy to your target switches. The integration with the ThousandEyes portal is established through Assurance > Health in Catalyst Center, where a setup banner walks you through OAuth authentication to your ThousandEyes organization. Once agents are running and registered in ThousandEyes, you configure tests in the ThousandEyes portal (not in Catalyst Center). As of early 2026, you can deploy agents from Catalyst Center, and Catalyst Center can surface ThousandEyes data, but test creation and management still lives in the ThousandEyes portal. This is worth knowing before you promise your NOC a single-pane-of-glass experience: the operational workflow involves both interfaces. ### Where to Place Agents This is an operational decision that significantly affects what value you get. An agent on an access switch can measure from that switch's perspective on a specific VLAN, which approximates the user experience for clients on that VLAN. To get representative coverage, you should deploy agents on user-facing VLANs at access switches, plus agents on distribution or core switches (on the management VLAN) for an infrastructure-level perspective. A reasonable starting point for a campus environment: one agent per access-layer switch on your primary user VLAN, and one agent per distribution switch on the management VLAN. Then configure tests for your top five SaaS applications, your internal DNS resolvers, your WAN gateway, and any cloud workloads your users depend on. That gives you baseline coverage before you optimize placement based on incident patterns. ## Gaps and Caveats This is not a marketing pitch, so here is an honest inventory of what this integration does not do, and where you should set expectations carefully. **Test management is still two portals.** Catalyst Center can deploy agents and surface data, but configuring what to test, test types, target URLs, alerting thresholds, dashboard views, requires the ThousandEyes portal. If your team is already comfortable with multiple observability tools this is not a big deal, but if you were hoping for genuine single-pane management of both, that is not where the product is today (at least as of early 2026). **ThousandEyes is only as useful as your test configuration.** Agents doing nothing generate nothing. If you deploy 300 agents and configure zero tests, you have spent time and hardware resources to collect no data. The operational investment in defining good tests, setting appropriate alert thresholds, and building useful dashboards is significant. Budget for this before deployment. **Agent-on-switch is not the same as agent-on-endpoint.** A ThousandEyes Enterprise Agent on a Catalyst 9300 measures the network from the switch, not from the user's laptop. If a specific user machine has a VPN client, a security agent, or a browser configuration that affects its application performance, the switch-based agent will not see that. ThousandEyes Endpoint Agents (installed on user machines) fill this gap, but they are a separate deployment and may require separate licensing depending on your agreement. **Catalyst Center's ThousandEyes data view is a subset of what is in the ThousandEyes portal.** The Catalyst Center integration surfaces ThousandEyes data in the Assurance section, but not every visualization and test type will appear with the same depth as in the dedicated ThousandEyes UI. Complex investigations, especially BGP analysis or multi-agent path correlation, will likely drive you to the ThousandEyes portal anyway. **Not every release behavior is identical.** Cisco is actively developing both products, and specific feature availability by Catalyst Center version or ThousandEyes deployment type may vary. Always verify current feature support against the release notes for your specific versions before committing to an architecture that depends on a specific capability. **AI-assisted correlation is still maturing.** Catalyst Center's AI assistant and anomaly detection are improving rapidly (the 2026 releases added natural-language querying across the Catalyst portfolio), but the integration between Catalyst Center AI insights and ThousandEyes data is not yet at a point where the system will automatically correlate an AI-detected anomaly in your fabric with a BGP event in ThousandEyes and surface both in a single alert. Manual timeline correlation is still the primary workflow for cross-domain incident analysis. ## Who Should Adopt This and Who Should Probably Wait This integration is not universally valuable. The answer to "should we deploy ThousandEyes on our Catalyst switches?" depends heavily on your environment and your operational pain points. ### Good candidates for adoption now - Enterprise campus environments with significant SaaS dependency (Microsoft 365, Salesforce, Webex, Zoom) where "is it the network?" is a recurring incident question and the team regularly struggles to prove network innocence to application owners or business units. - Organizations that already hold DNA Advantage licensing. The ThousandEyes agent license is included, so the marginal cost of deploying agents is primarily operational (deployment time, test configuration, monitoring setup). - Multi-site or branch-heavy environments where ISP diversity means the path to cloud applications varies significantly by site. ThousandEyes is particularly valuable here because it makes branch-by-branch path quality differences immediately visible. - NOC teams that have SLA obligations with external providers (ISP, cloud, SaaS) and need to produce evidence of provider-side failures to support escalations and credits. ThousandEyes generates shareable path traces and performance snapshots that hold up in those conversations. ### Situations where it can probably wait - Smaller environments (under 50 switches) with primarily on-premises applications where most of the user-to-application path stays inside your WAN. If your users are hitting internal data center workloads over MPLS and not touching the public internet for critical applications, ThousandEyes' core value proposition does not apply with the same force. - Teams on DNA Essentials without a current path to Advantage. The additional licensing cost changes the ROI calculation significantly. - Operations teams that do not yet have the bandwidth to configure and maintain ThousandEyes tests and dashboards. Deploying the agents and then leaving the platform unconfigured is a waste of the infrastructure and the license. If your team is already stretched, this is an investment that requires staffing capacity to return value. - Environments running IOS-XE versions below the minimums (17.3.3 on 9300, 17.5.1 on 9400). Get to a supported release first, and if you are that far behind on software, there are likely more urgent operational concerns to address before adding an observability layer. ## Key Takeaways - The core problem this solves is visibility across administrative boundaries. Catalyst Center is authoritative inside your domain; ThousandEyes is authoritative outside it. The integration puts both data sets in reach during a single troubleshooting session. - The operational benefit is faster fault isolation and better evidence. The most important outcome is not "we found the problem faster", it is "we proved the problem was not in our network, and we had the data to show the ISP or SaaS vendor exactly where the fault was." - Deployment has real prerequisites. DNA Advantage licensing, supported hardware platforms, minimum IOS-XE versions, AppGig VLAN trunk configuration, and outbound HTTPS egress from agents are all required before agents register and generate data. - Test management is still split between platforms. Agent deployment through Catalyst Center is streamlined, but test creation, alerting, and advanced visualization live in the ThousandEyes portal. This is not a true single pane of glass yet. - ThousandEyes value is proportional to test configuration effort. Plan time to define tests, set thresholds, build dashboards, and train your NOC on how to read the data. Agents with no tests produce no signal. - If you are already on DNA Advantage and your team regularly fields "is it the network?" escalations from SaaS-dependent business units, this integration is worth deploying. The license cost is already included, and the operational lift of configuring good tests pays back quickly in reduced MTTR and cleaner escalation evidence. More from the Cisco deep-dive series: [Is Cisco SD-Access Worth It in 2026? A Brownfield Reality Check](https://www.pinglabz.com/is-cisco-sd-access-worth-it-in-2026-a-brownfield-reality-check/). ### Is Cisco SD-Access Worth It in 2026? A Brownfield Reality Check URL: https://www.pinglabz.com/is-cisco-sd-access-worth-it-in-2026-a-brownfield-reality-check/ Last updated: 2026-07-04T23:06:43.000Z Every few years, a platform comes along that promises to simplify enterprise networking at a fundamental level. SD-Access was Cisco's answer to the question of how you do consistent segmentation, identity-based policy, and automated provisioning across a campus at scale. The pitch is compelling. The architecture is technically sound. And in 2026, after several years of real-world deployments, there's enough practitioner experience to give you a more grounded answer than the datasheets will. The honest answer is: it depends. Not in a hand-wavy way, but in a specific, checklist-able way. SD-Access can absolutely deliver on its promises, but it requires a level of operational prerequisite that many teams underestimate, and brownfield campuses carry design debt that the fabric overlay does not magically erase. This article walks through what SD-Access actually delivers, what brownfield reality looks like, and how to decide whether it makes sense for your environment right now. ## What Teams Hope SD-Access Will Solve Before getting into the hard parts, it's worth being honest about what drives the interest. Teams that come to SD-Access are usually dealing with at least one of these problems: - **VLAN sprawl and IP addressing debt.** Campus networks that have grown organically accumulate VLANs, subnets, and ACLs that nobody wants to clean up. SD-Access promises to flatten the policy model: segmentation via Virtual Networks and Scalable Group Tags (SGTs), not hundreds of VLANs. - **Inconsistent access policy.** When policy lives in switch ACLs spread across dozens of closets, changes are slow, error-prone, and hard to verify. Centralised policy in ISE + Catalyst Center looks appealing. - **Guest and contractor segmentation complexity.** Keeping IoT devices, contractors, and employees on logically isolated paths without endless VLAN gymnastics is a real operational pain. - **Manual provisioning overhead.** New switch ports should just work when a device connects. The promise of identity-driven automation (the endpoint shows up, ISE evaluates it, policy is applied) resonates with teams that spend too much time manually tweaking port configs. These are real problems. SD-Access addresses all of them. The question is whether you're in a position to actually get there. ## The Promised Benefits: What SD-Access Actually Delivers Let's be fair to the architecture. When SD-Access is deployed well in a reasonably clean environment, it does deliver meaningfully: Macro-segmentation via Virtual Networks (VNs) What you get VRF-isolated traffic paths between user groups (employees, guests, IoT) without L2 VLAN separation What it replaces VLAN + ACL sprawl, dedicated firewall hairpin for every segment pair Micro-segmentation via SGTs What you get Policy between user/device groups within a VN, enforced in the fabric via SGACL; follows the user as they roam What it replaces Per-port ACLs, 802.1X-per-interface config, VLAN-based policy Anycast gateway What you get Every Edge Node hosts the same gateway IP and MAC; no HSRP/VRRP needed; sub-second failover What it replaces HSRP/VRRP on distribution switches, STP dependency for failover Seamless endpoint mobility What you get LISP separates endpoint identity (IP) from location (RLOC); endpoints keep their IP as they move across the fabric What it replaces Client tracking via STP/ARP, layer 2 domain extension for mobility Automated provisioning What you get Day-0/1/2 templates in Catalyst Center push consistent config to fabric nodes; new switch onboarding is repeatable What it replaces Manual CLI configuration per closet switch, per-port ACL edits Centralised policy management What you get SGT policy matrix in ISE; change once, enforce everywhere in the fabric What it replaces Distributed ACL management across dozens of switches The segmentation story in particular is genuinely better at scale than VLAN-based policy. Once you have VNs and SGTs working, adding a new user class or device category is a policy change in ISE: not a VLAN, not a new SVI, not a firewall rule coordination call. ## Brownfield Reality: What the Deployment Guides Don't Say Here's where the conversation usually shifts. The vast majority of enterprise campuses are not greenfield. They are buildings full of Catalyst switches of mixed generations, VLANs that carry tribal knowledge nobody wants to document, and routing designs that reflect decisions made by people who left the company years ago. Dropping a fabric overlay onto that is not a migration; it is a redesign that has to coexist with the original. ### Legacy Switches and Hardware Eligibility Not every switch in your campus can be an SD-Access fabric node. Catalyst 3850s, older 3650s, and anything not on the supported platform list for Catalyst Center fabric provisioning will either need replacement or will stay as "non-fabric" legacy segments that connect to the fabric via an External Border Node. That border handoff (the Layer 2 Border Handoff in Cisco's terminology) is functional, but it means those endpoints are still in traditional VLANs and don't benefit from fabric segmentation. You end up running two models simultaneously, which is fine as a migration phase but can persist indefinitely in environments where capital budget is constrained. Before a brownfield SD-Access project gets past the whiteboard, someone needs to produce a complete platform inventory and check every device against the supported hardware matrix in the current Catalyst Center release notes. This is not glamorous work, but skipping it means discovering ineligible switches during deployment, when the schedule pressure is highest. ### The L2/L3 Handoff Problem SD-Access uses IS-IS as the underlay routing protocol, and Catalyst Center provisions the underlay automatically during fabric node onboarding. In a clean deployment, that's straightforward. In a brownfield network, you're likely running OSPF or EIGRP in the distribution and core layers, and those routing processes need to coexist with or be replaced by the IS-IS underlay. The fabric border nodes translate between the two worlds, but the underlay IP addressing scheme must be planned, not discovered, before provisioning begins. A common failure pattern: teams bring up the fabric, Catalyst Center provisions the IS-IS underlay correctly on fabric nodes, but existing OSPF adjacencies on adjacent non-fabric distribution switches create routing asymmetry that only surfaces under load or when a link fails. The troubleshooting is difficult because the failure is at the boundary between two routing models, and tooling on each side sees a healthy network. ### VLAN and IP Address Debt SD-Access endpoints live in Endpoint Groups (EPGs) that map to VNs and SGTs. Each EPG has a subnet (the anycast gateway on the Edge Nodes). In a brownfield environment, your existing endpoints already have IP addresses in existing subnets. There are two paths: re-IP (painful, requires change windows and application validation), or use Layer 2 Border Handoff to preserve the existing subnets while gradually migrating traffic to the fabric (less painful short-term, but creates a hybrid that requires ongoing care). A specific gotcha that catches teams: DHCP pool scopes must align precisely with fabric Endpoint Group subnet definitions. If a DHCP server is handing out addresses from a pool that doesn't match the Endpoint Group subnet the Edge Node is configured with, clients will get addresses that the fabric can't properly route. This is not a corner case; it's a common early-deployment issue in brownfield environments with legacy DHCP servers that have accumulated config over years. ### The Phased Migration Grind Cisco's recommended brownfield approach is phased: build the fabric alongside the legacy network, use Border Nodes for reachability between them, migrate buildings or floors gradually, and decommission legacy segments as you go. This is sensible. It is also slow and operationally demanding to sustain. The practical challenge is that during the coexistence phase, your team is operating two networks simultaneously. Troubleshooting requires understanding which model a given endpoint is in. Change management gets complicated because a change to the fabric border config can affect both fabric and non-fabric segments. Teams that don't have the staffing or tooling to manage the coexistence phase well often find the migration stalls at 60-70% complete, delivering some of the cost of SD-Access without the policy consistency that justified the project. ## The Dependency Stack: Everything That Has to Work First SD-Access is not a product you deploy in isolation. It is an operating model that requires a stack of integrated components to function. Understanding the dependency chain is essential to realistic planning. Catalyst Center Role in SD-Access Orchestration, provisioning, Day-2 assurance. All fabric configuration flows through it. What breaks if it's not right Without Catalyst Center, there is no supported SD-Access fabric deployment path. CLI-only fabric configuration is not a supported production approach. Cisco ISE Role in SD-Access Identity, SGT assignment, SGACL policy enforcement, 802.1X/MAB, pxGrid integration with Catalyst Center What breaks if it's not right Without ISE, you have no micro-segmentation, no dynamic policy, and no identity-based endpoint placement. VN mapping falls back to static SSID/port-based assignment only. PKI / Certificate infrastructure Role in SD-Access ISE nodes require valid certificates for EAP authentication and for Catalyst Center integration. Fabric nodes use certificates for LISP and pxGrid. What breaks if it's not right Certificate expiry or misconfiguration causes authentication failures, often silently showing as "timeout" in client logs rather than explicit certificate errors. This is the single most underestimated operational burden in SD-Access deployments. Underlay network Role in SD-Access IS-IS routed underlay connecting all fabric nodes. Catalyst Center provisions this automatically for supported hardware. What breaks if it's not right MTU mismatches (VXLAN adds overhead; end-to-end MTU must support 1600+ bytes or fragmentation occurs) break fabric reachability in ways that are hard to diagnose. Supported hardware Role in SD-Access Catalyst 9200/9300/9400/9500 series switches; not all IOS-XE versions support all SD-Access features What breaks if it's not right Fabric provisioning fails or produces incomplete configs on unsupported hardware. Certain fabric features are IOS-XE release-specific. DNS and NTP Role in SD-Access Catalyst Center and ISE both depend on accurate NTP; certificate validation is time-sensitive What breaks if it's not right NTP drift causes certificate validation failures. DNS is required for ISE node registration and Catalyst Center cluster communication. Two of these deserve extra emphasis because they are consistently underestimated. **ISE readiness is not just about licensing.** ISE needs to be deployed (HA-pair in production environments), integrated with your Active Directory or LDAP, loaded with profiling policies for your device types, and have a tested 802.1X policy before the fabric migration begins. Teams that try to stand up ISE and migrate the campus simultaneously are combining two large, complex change programmes. The resulting failure modes are difficult to isolate because both systems are in flux at the same time. **Certificate lifecycle is a continuous operational requirement, not a one-time setup.** ISE uses certificates for EAP-TLS authentication, for its Admin portal, and for pxGrid communication with Catalyst Center. When those certificates expire (and in large deployments, they expire at different times on different nodes), authentication silently fails until someone manually renews them. Organisations without a mature PKI practice, or without monitoring for certificate expiry, will experience this. It's not a question of if, it's when. ## ISE: The Part That Actually Takes the Longest It's tempting to think of SD-Access as primarily a switching and overlay project. The switching work is actually the manageable part. What takes the most time in practice is policy design in ISE. SGT policy is an NxN relationship matrix. If you have 20 defined SGTs, you have up to 400 SGT pair relationships to consider for your SGACL policy. Most environments don't need 400 distinct rules (many pairs can default to permit or deny), but you do need to *decide* for each relevant pair, and those decisions require input from security, application teams, and operations. That conversation, in most organisations, takes months. A practical sequencing that works: deploy ISE in Monitor mode first. Every endpoint authenticates, ISE assigns an SGT, but no policy is enforced. This phase shows you what you're actually dealing with: what device types are on your network, whether your profiling policies correctly identify them, and whether your 802.1X authentication rates are where they need to be. Only after Monitor mode validates your policy coverage should you move to Low Impact mode (permit traffic while logging policy violations) and then eventually to Closed mode (enforce and block). Skipping or shortening the Monitor mode phase is a common reason SD-Access projects produce unexpected downtime during cutover. You discover that 15% of your devices are failing 802.1X when you enforce (often because of certificate issues, RADIUS timeout config, or profiling misclassification) and those devices go dark. ## Common Failure Patterns These are the patterns that come up repeatedly in brownfield SD-Access deployments: - **Treating it as a switching project.** Teams scope SD-Access as a hardware refresh with a new management plane, then discover that the identity and policy design work dwarfs the network engineering work. Budget and timeline are set for "replacing switches" not "redesigning access policy." - **Combining ISE and fabric deployment simultaneously.** Both are large changes; combining them makes failure isolation nearly impossible. Always get ISE stable and validated in Monitor mode before beginning fabric migration. - **MTU oversight.** VXLAN encapsulation adds overhead. The end-to-end path must support a minimum MTU of approximately 1600 bytes. In brownfield environments with older WAN links, service provider handoffs, or legacy firewall policies that cap MTU, this causes intermittent fabric connectivity failures that are maddeningly hard to trace to a single root cause. - **Certificate neglect post-deployment.** ISE certificate expiry is a known, trackable event. Teams that don't instrument certificate expiry monitoring will experience it as an unexpected authentication outage, typically 1-3 years after deployment when the certificates from the initial build expire. - **Policy design by network team without security input.** SGT policy encodes access decisions that security teams own. When network engineers design the SGT taxonomy without security input, the resulting policy either over-permits (defeating the segmentation purpose) or under-permits (causing application breakage during enforcement cutover). - **Stalled coexistence phase.** The brownfield-to-fabric migration stalls mid-way when the team runs out of momentum or budget. The result is a permanent partial deployment (fabric in some buildings, legacy in others) that doesn't deliver the policy consistency benefit but does add management complexity. ## Decision Checklist: Are You Ready for SD-Access? This isn't a binary "go/no-go"; it's a readiness profile. The more boxes you can check honestly, the better your outcomes will be. Is your hardware predominantly Catalyst 9000 series or equivalent SD-Access capable platforms? If yes Lower replacement cost; full fabric feature support If no Budget for hardware refresh or plan for extended coexistence with legacy segments Do you have (or are you willing to invest in) Catalyst Center and ISE licensing and infrastructure? If yesFoundation is present If no This is not optional; no Catalyst Center = no supported SD-Access Is ISE already deployed, or is there a realistic plan to stand it up and stabilise it before the fabric migration begins? If yes Strong prerequisite satisfied If no Plan for 6–12 months of ISE work before fabric deployment starts Does your organisation have a PKI practice or a vendor managing certificate lifecycle? If yes Certificate risk is managed If no Build a certificate monitoring and renewal process as part of the project, not an afterthought Have you done a full platform inventory and confirmed end-to-end MTU ≥ 1600 bytes across all paths? If yesUnderlay is viable If no Identify MTU bottlenecks before provisioning; finding them after the fact is painful Do you have security team buy-in and commitment to participate in SGT taxonomy and policy design? If yes Policy design is feasible If no Without security participation, SGT policy will be incomplete or will default to over-permission Does your team have the bandwidth to run a parallel coexistence phase without it indefinitely stalling? If yes Migration has a realistic path to completion If no Consider whether a phased approach is sustainable, or whether a more aggressive cut-and-migrate strategy fits better Is your DHCP infrastructure well-documented and can it be aligned with fabric Endpoint Group subnet definitions? If yes Reduces early-deployment connectivity issues If no Spend time on DHCP discovery and cleanup before provisioning Edge Nodes ## Final Recommendation by Environment Type SD-Access is not the right answer for every campus, and it's worth being direct about that. Greenfield campus or major refresh (new hardware, clean subnets) SD-Access recommendation Strong yes: design for it from the start Rationale The brownfield migration cost disappears. You get the full architecture benefit without the coexistence complexity. Large enterprise with significant IoT, guest, contractor segmentation requirements and ≥ 500 access ports SD-Access recommendation Yes, but plan 18–24 months and invest in ISE first Rationale Scale justifies the investment; segmentation problem is real and VLAN-based alternatives get worse as the network grows Mid-size campus (200–500 access ports), mostly homogeneous hardware, reasonable VLAN discipline SD-Access recommendation Conditional: evaluate operational capacity honestly Rationale The architecture works at this scale, but the operational overhead of Catalyst Center + ISE is significant relative to the network size. If your team is lean, weigh that carefully. Campus with high proportion of legacy hardware (pre-9000 series) and no hardware refresh budget SD-Access recommendation Not yet: focus on ISE and identity foundation first Rationale Hardware ineligibility will force a prolonged coexistence phase. Get ISE deployed and identity-based policy working first; the fabric deployment can follow when hardware cycles align. Small campus (< 100 access ports), single building, limited IT staff SD-Access recommendation No: wrong tool for the scale Rationale The operational model doesn't match the environment. VLAN-based segmentation with 802.1X and ISE delivers 80% of the benefit at a fraction of the complexity. Campus where SD-WAN or cloud-first connectivity is the primary initiative SD-Access recommendation Sequence SD-Access after the WAN layer is stable Rationale Combining SD-Access and SD-WAN migrations simultaneously creates compounding complexity. Finish the WAN layer first; campus fabric can follow. The version of this conversation that goes wrong is when SD-Access gets positioned as a tactical solution to an urgent problem ("we need better segmentation now") and the prerequisites don't get the investment they need. The architecture is sound; the operational model is demanding. Teams that understand the difference between "can deploy" and "can operate confidently for years" go into the project with the right expectations and the right budget. ## Key Takeaways SD-Access is a genuine architectural improvement over VLAN-based campus policy, but only at the right scale and with the right prerequisites in place. Brownfield campuses carry design debt that the fabric overlay doesn't erase; hardware eligibility, underlay MTU, DHCP alignment, and legacy VLAN migration are all real work that has to be planned and executed, not assumed away. The dependency stack (Catalyst Center, ISE, PKI, IS-IS underlay, supported hardware) is not optional and not trivial; ISE readiness and certificate lifecycle management in particular are consistently underestimated and are responsible for a disproportionate share of post-deployment operational incidents. If you want to evaluate SD-Access honestly, start by assessing your ISE maturity and your team's bandwidth to sustain a parallel coexistence phase during migration; those two factors, more than the technology itself, will determine whether the project succeeds. The teams that do this well treat SD-Access as what it actually is: an operating model change that happens to require new network infrastructure, not the other way around. More from the Cisco deep-dive series: [Post-Quantum Crypto in Cisco Campus Networks: What Operators Need to Know](https://www.pinglabz.com/post-quantum-crypto-in-cisco-campus-networks-what-operators-need-to-know/). ### Backing Up, Restoring, and Upgrading the Cisco C9800 URL: https://www.pinglabz.com/c9800-backup-upgrade/ Last updated: 2026-07-04T23:02:59.000Z The Catalyst 9800 makes backup, restore, and upgrade look simpler than it actually is. The commands are short, the GUI has wizards, and most of the time the process works. The trouble starts on the 2 percent of upgrades that go wrong - a commit timer you didn't know about expires, an AP image is missing at cutover, an SMU gets applied to the wrong version, or a restore overwrites a config you wanted to keep. This guide walks through the full lifecycle the way you should actually operate it: backing up the configuration and state you'd need to rebuild a WLC, restoring cleanly, and upgrading with Install mode, ISSU, SMU, and N+1 rolling upgrades - with the gotchas that don't fit on a GUI tooltip. For where this topic sits in the wider picture, see the [Catalyst 9800 wireless complete guide](https://www.pinglabz.com/wireless/). ## What "Backup" Actually Means on a C9800 If you ask five network engineers what a WLC backup is, you'll get five answers - running config, startup config, the full flash, something from Catalyst Center, "whatever RANCID grabs." On a C9800 there are several distinct things you might want to back up, and they are not interchangeable. Running/startup config What it contains All IOS-XE CLI - WLANs, tags, profiles, AAA, HA settings How to back it up `copy running-config tftp:` or `archive` When you need it Daily. Required for a bare-metal rebuild. Trustpoints and certificates What it contains LSC, WebAuth cert, any imported trust anchors How to back it up `crypto pki export ... pkcs12` When you need it After any certificate operation; before reload. WebAuth bundle What it contains Custom login pages, images, CSS How to back it up File copy from bootflash to external When you need it Before upgrade or restore. License authorisation What it contains Smart Licensing policy state How to back it up Re-register after rebuild; no local backup needed When you need itAfter rebuild. Software image What it contains The `.bin` and any SMU packages currently in use How to back it up Keep in a known image server, not only on bootflash When you need it Always. Don't rely on the box keeping it. Operational data What it contains Rogue database, client history, RRM state How to back it up Not backed up; rebuilt on the fly When you need it N/A - the controller recovers this after reload. The config is the single most important artifact, but the trustpoints are the one people forget. A C9800 rebuilt from config alone without its LSC or WebAuth certificate will come up with every AP failing DTLS and every captive-portal client seeing certificate warnings. Always back up the crypto material alongside the config. ## Backing Up the Configuration The basic backup is what you'd expect - a copy of the running config to an external destination. Two ways to do it, both of which should be scripted or scheduled rather than run by hand. ``` C9800#copy running-config tftp://10.10.10.5/c9800-primary.cfg Address or name of remote host [10.10.10.5]? Destination filename [c9800-primary.cfg]? !! 14782 bytes copied in 2.143 secs (6898 bytes/sec) C9800#archive config C9800#show archive The maximum archive configurations allowed is 10. The next archive file will be named flash:archive-Oct-5-2026-09-15-22.cfg Archive # Name 1 flash:archive-Oct-4-2026-08-00-01.cfg 2 flash:archive-Oct-4-2026-20-00-03.cfg 3 flash:archive-Oct-5-2026-08-00-02.cfg <- Most Recent ``` For operational use, configure the `archive` subsystem to take a config snapshot on every `write memory` and on a timed interval, store them to both flash and an external path, and retain enough history that you can roll back to any point in the last week without asking your backup team. ``` C9800(config)#archive C9800(config-archive)#path tftp://10.10.10.5/c9800-archive/$h-$t C9800(config-archive)#write-memory C9800(config-archive)#time-period 1440 C9800(config-archive)#maximum 14 ``` That stanza captures the config every 24 hours (1440 minutes) and on every `write memory`, keeps the last 14 versions, and pushes each one to TFTP. Replace TFTP with SFTP in any environment that cares about transport security. ## Exporting Trustpoints Trustpoints - the certificate and private-key pairs used for LSC, WebAuth, and NETCONF - are stored inside the config in an encoded form, but the safer practice is to export them explicitly so you can re-import on a rebuild without parsing the running config by hand. ``` C9800(config)#crypto pki export lsc-trust pkcs12 tftp://10.10.10.5/lsc-trust.p12 password 0 S3cretPassphrase % Exporting trustpoint 'lsc-trust' in PKCS12 format % The PKCS12 file has been successfully exported ``` Store the exported PKCS12 with strong access control; it contains the private key. On a restore, import it back with `crypto pki import … pkcs12`. ## Restoring a Controller Restoration is usually one of two scenarios: either you're rebuilding a controller from scratch after a hardware swap or a total software reinstall, or you're rolling back a recent config change. They have different procedures. **Full rebuild.** Install the OS image (or boot into one), complete Day 0 setup enough to get a management IP, then copy the saved config to flash and either `copy flash:saved-config.cfg running-config` or set it as the startup config and reload. After the controller reloads, import the trustpoints, verify AP join, and check `show wireless client summary` counts against your baseline. **Config rollback.** If a recent change broke something and you want to revert to yesterday's config, use `configure replace` rather than an overwrite: ``` C9800#configure replace flash:archive-Oct-4-2026-08-00-01.cfg list This will apply all necessary additions and deletions to replace the current running configuration with the contents of the specified configuration file, which is assumed to be a complete configuration, not a partial configuration. Enter Y if you are sure you want to proceed. [no]: y Total number of passes: 1 Rollback Done ``` `configure replace` computes the diff between the current running config and the target file and applies only what's needed, which is both faster and less disruptive than a reload. It's the right tool for "undo the last change" scenarios. ## Upgrade Methods, and Which One to Use When The C9800 gives you more upgrade options than any previous wireless platform. They look redundant until you realise each is optimised for a different combination of risk, outage budget, and release delta. Install mode reload Best for Major release upgrades on a standalone WLC Data-plane impact Full outage unless HA takes over Prerequisite Running in Install mode ISSU (In-Service Software Upgrade) Best for Compatible release upgrades on an SSO HA pair Data-plane impact Zero client drop for RUN-state clients Prerequisite SSO pair, same form factor, Install mode, compatible versions SMU (Software Maintenance Upgrade) Best for Targeted hot patches for specific defects Data-plane impact None for cold SMU; brief for hot Prerequisite SMU must match your exact release N+1 Hitless Rolling AP Upgrade Best for Large campuses where you want to stage AP upgrades in waves Data-plane impact Per-wave AP reboots; never whole-site Prerequisite N+1 deployment and matching image on both controllers AP image predownload Best for Every upgrade where APs will reboot Data-plane impact No impact - staged ahead of the window Prerequisite Enabled in the wireless profile Three practical rules before you pick a method: **Always be in Install mode.** Bundle mode works but locks you out of ISSU, SMU, and reliable rolling upgrades. Check with `show version | include Installation`; if it doesn't say "Installation mode is INSTALL," convert before you do anything else. **Always pre-download AP images.** Without it, every AP downloads the image during cutover and your window stretches to the slowest WAN link. With it, the cutover is just a reboot. **Always have a rollback plan that someone has tested in a lab.** The commands exist for every method, but rehearsing rollback during a real incident is the wrong time to learn the syntax. ## Install-Mode Upgrade: The Standard Path For a single chassis or a non-ISSU upgrade on an HA pair, the normal flow is `install add file … activate commit`. Each phase is a distinct step and each has a rollback. ``` C9800#copy tftp://10.10.10.5/C9800-SW-17.15.1.SPA.bin bootflash: C9800#install add file bootflash:C9800-SW-17.15.1.SPA.bin install_add: START Sun Apr 5 10:14:22 UTC 2026 install_add: Adding IMG --- Starting initial file syncing --- Finished initial file syncing --- Starting Add --- Performing Add on all members Add: Passed on [R0] Finished Add SUCCESS: install_add Sun Apr 5 10:16:02 UTC 2026 C9800#install activate This operation may require a reload of the system. Do you want to proceed? [y/n]y C9800#install commit SUCCESS: install_commit Sun Apr 5 10:45:12 UTC 2026 C9800#show version | include Cisco IOS Software Cisco IOS Software [IOSXE], Catalyst L3 Switch Software (CAT9K_IOSXE), Version 17.15.1 ``` The critical step most engineers skip is `install commit`. Without the commit, IOS-XE runs a rollback timer (default 6 hours on 17.4 and later) after which the controller reverts to the previous image automatically. That's a safety feature - if you can't reach the box after activation, it self-heals - but it's also a trap if you forget. Commit as soon as you've verified the upgrade is healthy. If you need to back out before committing, `install abort` rolls to the previous image cleanly: ``` C9800#install abort This operation will cancel the current installation. Are you sure you want to proceed? [y/n]y ``` ## ISSU: Upgrade Without Dropping Clients ISSU upgrades an SSO pair between compatible releases with zero data-plane impact. The active chassis pushes the new image to the standby, the standby reloads and comes up on the new version, a switchover happens, then the former active reloads on the new image and the pair resynchronises. From a client's perspective, nothing happens - RUN-state clients stay associated, and active flows keep forwarding. ``` C9800#install add file bootflash:C9800-SW-17.15.1.SPA.bin activate issu commit install_add_activate_commit: START Sun Apr 5 10:14:22 UTC 2026 --- Starting ISSU Compatibility Check --- ISSU compatibility check is SUCCESSFUL --- Starting Add --- Add: Passed on [R0,R1] --- Starting ISSU Activate --- ... SUCCESS: install_add_activate_commit ``` ISSU compatibility is not guaranteed across every pair of releases. Read the Cisco ISSU compatibility matrix for your source and target version *before* you plan the window, and if the pair is not ISSU-compatible, fall back to a standard install-mode upgrade with SSO failover - you'll still keep APs joined (the standby takes over), but clients mid-flow may experience a brief drop. ## SMU: Hot Patches for Specific Bugs A Software Maintenance Upgrade is a pointed fix for a specific defect, shipped as a package that applies on top of a base release. SMUs are release-specific - an SMU built for 17.12.3 will not apply to 17.12.4 - and they come in two flavours: cold SMUs require a reload, hot SMUs do not. The release notes for every SMU state which it is. ``` C9800#install add file bootflash:C9800-SMU-CSCab12345.SPA.smu C9800#install activate file bootflash:C9800-SMU-CSCab12345.SPA.smu C9800#install commit C9800#show install summary [ R0 R1 ] Installed Package(s) Information: State (St): S - Installed and Set State, N - Not installed, C - Installed and Committed Type St Filename/Version -------------------------------------------------------------- IMG C 17.12.4 SMU C C9800-SMU-CSCab12345.SPA.smu ``` SMUs are the right answer when you have a single production-impacting bug and don't want to take a full upgrade. They are not a general-purpose patching strategy - Cisco expects you to move to the next base release on a normal cadence and consume bug fixes that way. ## N+1 Hitless Rolling AP Upgrade In an N+1 deployment, you have one or more primary controllers and a backup controller that APs failover to when the primary goes down. The N+1 hitless rolling upgrade uses that pattern deliberately: you upgrade the secondary first, then move APs in waves from the primary to the secondary. Each wave reboots onto the new image; clients at sites in other waves are unaffected. When all waves are done, you upgrade the now-empty primary and let APs drift back. The key enabler is that both controllers have the same target image ready to go, and APs have pre-downloaded the new image. Rolling upgrade is a Catalyst Center-friendly workflow but can also be driven from the CLI with `ap image predownload` and controlled failover steps. ``` C9800#ap image predownload Initiating predownload on all APs C9800#show ap image Total number of APs : 412 Predownloading : 412 Completed predownloading : 412 C9800#ap image swap ``` ## Verification That Matters After Any Upgrade The commands that tell you whether an upgrade is actually healthy - not just whether it reloaded without crashing: ``` C9800#show version | include Cisco IOS Software C9800#show install summary C9800#show redundancy states C9800#show ap summary | include Joined C9800#show wireless client summary | include RUN C9800#show wireless stats ap join summary C9800#show logging | include -3-|-4- ``` Match the version, confirm SSO is in HOT state, confirm the AP join count is at or above your pre-upgrade baseline, confirm client count in RUN state is recovering, and scan the log for any error- or warning-level messages that appeared during or after the upgrade. If any of those look wrong, investigate before you commit - the commit timer is your friend. ## Key Takeaways A real backup covers the config, the trustpoints, and the image - not just the running config. Use `archive` to capture config snapshots on every `write memory` and to an external path, and export certificates explicitly with `crypto pki export pkcs12`. Use `configure replace` for rollbacks rather than overwriting the config. Always run in Install mode, never Bundle mode. Always pre-download AP images before a window. Pick ISSU when your SSO pair supports it and you want zero client impact; pick a standard install-mode upgrade with SSO failover when it doesn't. Use SMUs for targeted bug fixes, and use N+1 hitless rolling upgrades when you need to stage a very large AP fleet through waves. And never forget to `install commit` \- the rollback timer is a safety net, not a completion step. ### Cisco C9800 Fabric Mode and SD-Access Wireless Integration URL: https://www.pinglabz.com/c9800-fabric-sda-wireless/ Last updated: 2026-07-04T23:03:05.000Z SD-Access wireless on the Cisco Catalyst 9800 is one of those topics that looks complicated from the outside and becomes straightforward once you understand the mental model. The short version is: in a Software-Defined Access (SD-Access) fabric, the C9800 stops being the data-plane anchor for wireless clients. The fabric edge switches take that role. The C9800 still runs the control plane for APs and clients, but user traffic rides the fabric's VXLAN overlay directly from the fabric-enabled AP to the edge switch the client is policy-bound to. This article walks through how fabric-mode wireless works, why it's different from local or FlexConnect, what you gain by adopting it, and what the design and operational implications look like from a C9800 engineer's perspective. For where this topic sits in the wider picture, see the [Catalyst 9800 wireless complete guide](https://www.pinglabz.com/wireless/). ## Why Fabric Wireless Exists Traditional wireless - local mode on a C9800 - assumes a single Layer 2 anchor for each client. A client on SSID "Corporate" gets an IP from the WLC's VLAN, and traffic from that client traverses a CAPWAP tunnel from the AP to the WLC before it ever touches the distribution layer. This works, but it bakes two assumptions into the design: every client on a given SSID must land on the same VLAN everywhere, and the WLC must have bandwidth and CPU proportional to total user throughput. SD-Access changes those assumptions. SDA is Cisco's intent-based campus architecture built on VXLAN and LISP. Each fabric edge switch acts as a VXLAN tunnel endpoint (VTEP), clients are identified by a Virtual Network (VN, which maps to a VRF) and a Scalable Group Tag (SGT), and the fabric itself handles segmentation and mobility. Wrapping wireless into that model means the client's traffic joins the VXLAN overlay at the point it enters the wired network - and in fabric wireless, that point is the fabric edge switch nearest the AP, not the WLC. ## How Fabric-Mode Wireless Actually Works In fabric mode, the C9800 and the APs participate in the SD-Access control plane, but data flow changes dramatically. Here's the path a client's packet takes: 1. The client associates to a fabric-enabled AP. The AP runs its normal CAPWAP control channel to the C9800 - that part is unchanged. 2. The C9800 registers the client with the fabric control plane (the LISP map-server, which typically runs on the fabric border or a dedicated node). The client's MAC and the AP it's associated with are written into LISP as a mapping. 3. When the client sends its first packet, the fabric AP does not wrap it in CAPWAP. Instead, the AP encapsulates it in VXLAN and forwards it directly to the fabric edge switch it's connected to. 4. The fabric edge switch, now acting as the VTEP for that client, de-encapsulates the VXLAN, places the packet into the correct VN/VRF based on the client's SGT, and routes it to its destination - which may be another fabric node, a border exit, or a shared-services block. The C9800 never sees the client's user data. It handles association, authentication, roaming coordination, and RF management, but the data plane is offloaded to the fabric. That single change is why fabric wireless scales the way it does and why it integrates cleanly with the rest of SD-Access. Control plane (CAPWAP) Local modeAP ↔ WLC FlexConnectAP ↔ WLC Fabric mode (SDA)AP ↔ WLC Client data plane Local mode AP → CAPWAP → WLC → wired FlexConnect AP → local switch (bridged) Fabric mode (SDA) AP → VXLAN → fabric edge Client VLAN Local modeOn the WLC FlexConnect At the site / AP switchport Fabric mode (SDA) None - VN/VRF via LISP + SGT Segmentation Local modeVLAN + ACL FlexConnectVLAN + Flex ACL Fabric mode (SDA)VN + SGT (micro/macro) Mobility within fabric Local mode Mobility groups, anchors FlexConnectN/A Fabric mode (SDA) LISP map-update; fast and stateless WLC CPU for data path Local mode Proportional to traffic FlexConnectMinimal Fabric mode (SDA)Minimal ## The Role of the C9800 in a Fabric Deployment The C9800 still does everything you expect a controller to do - WLAN configuration, AP join, RRM, rogue detection, client policy, 802.1X - but it also takes on two new roles unique to SDA: - **Fabric control plane participant.** The WLC registers wireless clients with the fabric's LISP map-server/map-resolver. When a client roams between fabric-enabled APs, the WLC issues a LISP map-update that tells the fabric where the client now lives. This is how fast, stateless roaming works in SDA wireless. - **Policy source for wireless-side enforcement.** Scalable Group Tags and VN assignments for wireless clients originate at the C9800, sourced from ISE, and are carried into the fabric via the VXLAN header. The C9800 is the place where "this SSID maps to this VN" lives. There is also a subtle but important architectural rule: **the C9800 must not be a fabric edge node itself**. It sits outside the fabric, reachable over a Layer 3 path from the fabric border or directly on the underlay. The WLC is infrastructure that the fabric calls out to, not a fabric participant in the VXLAN sense. ## Design Requirements Several pieces need to be in place before a single fabric-enabled SSID can carry traffic. If any of them is missing, fabric wireless won't come up. C9800 running a supported IOS-XE version for SDA wireless SDA wireless feature parity depends on the release; check the release notes for the fabric features you need Catalyst Center (formerly DNA Center) managing the fabric Fabric configuration is driven top-down from Catalyst Center; manual CLI-only fabric wireless is not the supported path LISP-capable border and control plane nodes These are the targets the WLC registers clients with Fabric-capable APs (most 802.11ac Wave 2 and later Cisco APs) The AP must support VXLAN encapsulation in hardware ISE integration with scalable group tags and VN mapping Without ISE, you get SSID-based VN mapping only - no micro-segmentation IP reachability between the WLC and the fabric underlay The WLC needs LISP connectivity to the map-server and control plane The practical consequence: if you're not running Catalyst Center, you're not doing SDA wireless in any supported form. Fabric wireless is a Catalyst Center workflow; manual CLI configuration on the C9800 alone will not produce a working fabric. ## Fabric-Enabled Site Tag and AP Onboarding Fabric wireless is enabled per Site Tag on the C9800\. A Site Tag can be flagged as fabric-enabled, at which point APs assigned to that Site Tag take on fabric behaviour. This is how you deploy fabric wireless incrementally - you can run a brownfield campus where some sites are still local mode and others are fabric, all from the same WLC. ``` C9800#show wireless fabric summary Fabric Status : Enabled C9800#show wireless fabric client summary Number of fabric clients : 3410 MAC Address VLAN IP Address AP Name Status ------------------------------------------------------------------- aaaa.1111.2222 100 10.20.30.41 FAB-AP-001 Run aaaa.1111.2223 100 10.20.30.42 FAB-AP-001 Run aaaa.1111.2224 200 10.20.40.15 FAB-AP-017 Run C9800#show wireless fabric vnid mapping L2-VNID Name L3-VNID IP Address ------------------------------------------------------------- 8188 Corporate_VN 4099 10.20.30.0/24 8189 Guest_VN 4100 10.20.40.0/24 ``` The L2-VNID corresponds to the VXLAN segment ID the AP uses when it encapsulates client traffic; the L3-VNID identifies the VN/VRF the client lands in. That mapping comes from Catalyst Center, not from a CLI stanza you write by hand. ## Roaming in a Fabric Roaming is where fabric wireless earns its keep. In a traditional local-mode deployment, intra-controller roaming is handled by the WLC, and inter-controller roaming needs mobility groups, mobility tunnels, and anchor behaviour that can become complex at scale. In a fabric, roaming is a LISP update: when a client moves from AP-1 to AP-2, the WLC tells the fabric control plane that the client now lives behind AP-2's edge switch, and the fabric reprogrammes its mapping in milliseconds. There are no mobility tunnels. There is no "anchor" WLC. Clients keep their IP address and session state because the fabric overlay is the one doing the work - the WLC is just updating a mapping. This is why SDA wireless scales to very large campuses without the mobility-group gymnastics local-mode deployments require. The trade-off is that the fabric itself must be healthy. LISP map-server outages, a broken fabric border, or a partitioned underlay will break wireless in ways you haven't seen before. Your monitoring story needs to cover the underlay, not just the controller. ## Segmentation: VN and SGT, Not VLAN In a conventional WLC design, you segment users with VLANs. In SDA wireless, you segment users with Virtual Networks (for macro-segmentation between tenants or security zones) and Scalable Group Tags (for micro-segmentation within a VN). An SSID is mapped to a VN, and individual client policy is expressed as SGT-to-SGT rules enforced throughout the fabric. This is where SDA wireless shines for larger organisations: - A guest SSID maps to the Guest VN; no routing path exists between Guest VN and Employee VN without going through a fusion firewall you control. - Within the Employee VN, you can tag contractors with one SGT and permanent staff with another, and apply ACLs between them without carving up VLANs or IP space. - Segmentation policy follows the user as they roam, because the SGT is carried in the VXLAN header of every frame they send. The implication for the C9800 engineer is that you need to think in VN/SGT terms, not VLAN terms, when you design SSID-to-policy mappings. VLANs still exist at the edge, but they stop being the unit of segmentation. ## Where It Goes Wrong A few operational realities worth knowing before you commit: **Fabric wireless and RRM interact normally.** RF management, DCA, TPC, CleanAir, and the rest of the RRM toolkit work exactly as they do in local or Flex mode. The fabric does not change how the RF part of the C9800 behaves. If you already know how to design RF on a C9800, you already know how to design RF in a fabric. **Troubleshooting needs two hats.** A client connectivity problem in SDA wireless can live on the WLC (policy, AP, RF), on the fabric (LISP mapping, VXLAN path, SGT policy), or in ISE (group assignment). You need to know which one you're looking at before you start debugging, and you need visibility into all three. Catalyst Center Assurance helps, but you should still be comfortable running `show wireless fabric client summary`, `show lisp site`, and matching ISE authorisation logs by hand. **Multicast and broadcast need explicit design.** In a fabric, broadcast and multicast from wireless clients don't automatically reach other wireless clients or wired hosts the way they would on a shared VLAN. You configure multicast replication explicitly - headend replication or native multicast - and mDNS gateway becomes even more important because clients on different fabric edges can't discover each other's services via ordinary link-local multicast. **Guest anchor works differently.** Traditional guest anchoring to a DMZ C9800 is replaced by the Guest VN exiting at a dedicated border or fusion firewall. If you're migrating from anchor-based guest to fabric guest, the firewall rules and exit point for guest traffic change - plan that conversation with the security team before cutover. ## When to Adopt Fabric Wireless Fabric wireless isn't a replacement for local or FlexConnect - it's an alternative that only makes sense inside a full SD-Access deployment. You should consider it when: - The wired side of the campus is already SD-Access, or is moving there. Fabric wireless layered onto a conventional wired network is not the point. - You need micro-segmentation between user groups that would be painful to express with VLANs and ACLs. - Campus-wide mobility with stateless roaming is a requirement - for example, a hospital where clinicians carry devices from wing to wing all day. - You're running Catalyst Center already and want one orchestration plane across wired and wireless. You should not adopt fabric wireless when you don't yet have a fabric, when your team doesn't have LISP or SDA experience, or when your site is small enough that VLAN-based segmentation is still manageable. SDA is powerful but it has a learning curve, and the operational model is genuinely different from AireOS-era wireless. ## Key Takeaways Fabric-mode wireless moves the data plane off the C9800 and onto the fabric edge, while keeping the C9800 as the wireless control plane and the policy source. The WLC registers clients with LISP, the AP encapsulates client frames in VXLAN straight to the edge, and segmentation is expressed in VNs and SGTs instead of VLANs. Deployment is driven from Catalyst Center, not from hand-written CLI, and it only makes sense inside a real SD-Access campus. Done right, it scales cleanly and gives you stateless roaming and policy that follows the user. Done wrong - or attempted without a fabric, or without ISE - it just adds complexity without the benefit. Know the model before you commit, and design the underlay, the control plane nodes, and the WLC's reachability into the fabric as a single thing, because the WLC on its own does not make a fabric wireless deployment. ### C9800 FlexConnect vs. Local Mode: How to Choose URL: https://www.pinglabz.com/c9800-flexconnect-vs-local-mode/ Last updated: 2026-07-04T23:03:06.000Z Local mode and FlexConnect are the two fundamental ways a Cisco Catalyst 9800 can hand off client traffic, and the choice shapes almost everything else in your design - WAN utilisation, failure behaviour during a controller outage, the identity of the VLAN clients land on, where ACLs get enforced, how AAA survives a link flap, and whether you can even deploy WPA3 on the SSID you just built. Pick the wrong mode and you'll spend the next two years fighting symptoms. Pick the right one and most of your operational pain goes away. If you want the bigger picture first, the [Catalyst 9800 wireless complete guide](https://www.pinglabz.com/wireless/) covers the architecture this article plugs into. This guide compares the two modes at the design level, not the button level. You'll see where each one shines, where each one breaks, and the decision framework I recommend for choosing between them on a per-site basis. ## What Each Mode Actually Does In **local mode** (also called central switching), every client frame is tunnelled inside CAPWAP from the AP back to the C9800 and decapsulated there. The controller is the Layer 2 anchor for the client's VLAN, so a client connecting to an AP at a branch office will land on a VLAN that only exists at the data centre where the WLC lives. All policy - ACLs, QoS, URL filtering - is enforced on the WLC. The AP is effectively a dumb radio head. In **FlexConnect mode** (often called local switching), the AP still joins the C9800 over CAPWAP and still takes its configuration from it, but client data frames are bridged locally onto a VLAN that exists at the AP's site. The WAN carries only the CAPWAP control channel and any centrally-switched SSIDs you've opted in on. Policy can be enforced either centrally (via the C9800) or locally (via Flex ACLs on the AP), depending on how you configure the Flex Profile. That single difference - where the client frame exits onto the wired network - drives everything else. Client data path Local mode AP → CAPWAP tunnel → WLC → wired FlexConnect AP → local switch → wired Client VLAN location Local modeMust exist on the WLC FlexConnect Must exist at the AP's site WAN bandwidth consumed by client traffic Local mode 100% - everything traverses the WAN FlexConnect \~0% for locally-switched SSIDs Behaviour if WAN fails Local mode All clients disconnect; APs go offline FlexConnect Clients stay connected via Standalone mode; new associations possible with local auth Policy enforcement point Local modeWLC (IOS-XE policy) FlexConnect WLC for central-switched; AP Flex ACLs for local-switched AAA survivability Local mode None - if AAA is unreachable, auth fails FlexConnect Supported via local EAP or AAA cache Supported feature set Local modeFull FlexConnect Most features; a handful are central-only Typical WAN type Local mode High bandwidth, low latency (campus, DC) FlexConnect Any - branch, MPLS, SD-WAN, broadband ## When Local Mode Is the Right Answer Local mode is the default for a reason. When the AP is on the same high-bandwidth, low-latency fabric as the controller, centralising traffic is simpler, more consistent, and easier to troubleshoot. A single enforcement point means a single policy to reason about, a single place to run packet captures, and a single place to apply QoS. You should use local mode when: - **Your APs live in the same campus as the WLC**, connected over a high-bandwidth LAN with RTT measured in single-digit milliseconds. This is the default case for headquarters, large campuses, and anywhere that doesn't have a "branch" conversation. - **You want uniform policy.** Central ACLs, central AVC, central URL filtering, central Layer 7 application visibility - all simpler to design, deploy, and audit when there's one enforcement point. - **You're running features that are not fully Flex-capable.** Some edge features (certain Hyperlocation workflows, certain mDNS gateway behaviours on older releases) behave better or only work in local mode. Read the release notes for the IOS-XE version you intend to deploy before you commit. - **You need tightly controlled IP allocation.** Centralised DHCP on the WLC's subnet, central anchor for guest traffic, and consistent NAT behaviour are easier when every client gets its IP at the same subnet boundary. Local mode's failure mode is also predictable: if the WLC goes down and SSO doesn't take over, your APs disconnect. That's the tradeoff you accept for simplicity, and for a single-campus deployment with SSO HA that tradeoff is usually fine. ## When FlexConnect Is the Right Answer FlexConnect exists because not every AP can live next to its controller. Branch offices, retail stores, clinics, remote classrooms, and WAN-connected field sites need wireless to survive independent of the central WLC, and they need the WAN link to carry control traffic, not the entire user-data plane. You should use FlexConnect when: - **The AP is WAN-separated from the WLC.** If the CAPWAP tunnel traverses an MPLS circuit, SD-WAN overlay, or internet VPN, you almost always want FlexConnect local switching. Forcing a 1 Gbps office's worth of user traffic through a 100 Mbps WAN link back to a data centre is a classic over-budget pattern. - **You need wireless to keep working during a WAN outage.** A FlexConnect AP that loses its CAPWAP tunnel enters Standalone mode and keeps bridging existing clients. With local EAP or an AAA survivability cache, it can even authenticate new clients during the outage. - **Latency between the AP and the WLC exceeds \~100 ms.** CAPWAP tolerates latency in the tens of milliseconds without much drama; once you're past \~100 ms, client traffic tunnelled to a distant WLC starts feeling sluggish and jitter-sensitive applications (voice, video) suffer. - **You want site-local traffic to stay site-local.** A printer, a file server, or a Sonos system at the branch should not traverse the WAN to reach a client on the same branch. Local switching keeps the path direct. The trade-off is configuration complexity. FlexConnect introduces a Flex Profile, per-site VLAN mappings, Flex ACLs, native VLAN decisions, AP-to-switchport trunk considerations, and a survivability configuration that you have to plan rather than accept by default. ## Standalone Mode: The Feature You Buy FlexConnect For The single biggest reason to deploy FlexConnect in branch environments is Standalone mode. When an AP in FlexConnect mode loses connectivity to the WLC, it doesn't disconnect clients - it transitions into Standalone, and from that moment on it behaves like an independent bridge until the WLC comes back. What works in Standalone mode depends on what you configured in the Flex Profile: Existing connected clients Standalone behaviour Stay connected, keep bridging Required configurationNone - default New client association (open SSID) Standalone behaviourWorks Required configurationNone New client association (PSK SSID) Standalone behaviourWorks Required configuration PSK material must be cached on the AP New client association (802.1X) Standalone behaviour Works only if local EAP or AAA cache configured Required configuration Local EAP profile in the Flex Profile, or backup RADIUS Roaming between APs at the site Standalone behaviour Works if APs are in the same FlexConnect group Required configuration Flex group configured with shared key caching Guest / web-auth Standalone behaviour Does not work in Standalone Required configuration Expect guest SSIDs to fail during a WAN outage The practical advice: if branch wireless must survive a WAN outage, configure local EAP or an AAA survivability cache in the Flex Profile, and group the branch APs into a FlexConnect group. Without those pieces, Standalone mode exists but is much less useful. ## The VLAN Question The VLAN decision is the one that trips up most new FlexConnect designs. In local mode, you pick a VLAN ID, map it to the Policy Profile, and done - that VLAN only needs to exist on the WLC's wired interface. In FlexConnect, **that same VLAN ID has to exist on the switchport the AP is plugged into**, at every site where that Policy Tag lands. You have two strategies: **Strategy 1: Consistent VLAN IDs across all sites.** Every branch uses the same VLAN numbering scheme (say, VLAN 100 for employees, VLAN 200 for guests, VLAN 300 for voice). The Flex Profile does a 1:1 mapping. This is the simplest design and the easiest to troubleshoot. It requires coordination with the networking team to reserve VLAN IDs globally. **Strategy 2: VLAN override per site via Flex Profile mapping.** Every branch uses whatever VLAN IDs it already has, and the Flex Profile maps the SSID to different VLAN IDs at different sites. The Policy Profile references an SSID-to-VLAN-name mapping, and the Flex Profile translates the VLAN name to a local ID per site. Strategy 1 is cleaner. Strategy 2 is necessary when you're deploying wireless into a network with existing, unchangeable VLAN schemes. ``` C9800(config)#wireless profile flex flex-branch-1 C9800(config-wireless-flex-profile)#native-vlan-id 10 C9800(config-wireless-flex-profile)#vlan-name employee C9800(config-wireless-flex-profile)# vlan-id 100 C9800(config-wireless-flex-profile)#vlan-name guest C9800(config-wireless-flex-profile)# vlan-id 200 C9800#show wireless profile flex detailed flex-branch-1 Flex Profile Name : flex-branch-1 Native Vlan Id : 10 Vlan Name Vlan Id ACL ------------------------------------------------- employee 100 FLEX-EMP-ACL guest 200 FLEX-GUEST-ACL ``` ## Policy Enforcement: Two Modes, Two Enforcement Points In local mode, policy (ACLs, QoS markings, URL filters, AVC) is applied on the WLC because that's where the client's frame lives. You create the ACL once, attach it to the Policy Profile, and every AP tagged with that Policy Tag inherits the behaviour. In FlexConnect, policy is applied at one of two places depending on the traffic direction: - **Central switching on a Flex AP** \- traffic is tunnelled to the WLC, so WLC-side policy applies. This is how you mix a guest SSID (tunnelled to a DMZ anchor) with a corporate SSID (locally switched) on the same FlexConnect AP. - **Local switching on a Flex AP** \- traffic exits the AP directly, so policy must live on the AP as a Flex ACL. You define the ACL on the WLC but it's pushed to the AP and enforced there. The implication is that you need to think about *where* your policy lives before you design the Flex Profile. If you have a complex URL filtering workflow that's easier to maintain centrally, you might choose to central-switch that SSID while keeping the employee SSID local. Mixing is allowed and is in fact the point. ## Choosing: A Practical Decision Framework Here is the decision tree I use: Is the AP in the same physical site as the WLC (campus/DC)? If "Yes"→ Local mode If "No"→ Go to next question Is the link between the AP and the WLC a WAN link (MPLS, SD-WAN, internet VPN)? If "Yes"→ FlexConnect If "No"→ Go to next question Does RTT between AP and WLC exceed \~100 ms? If "Yes"→ FlexConnect If "No"→ Go to next question Must wireless survive a total WAN outage? If "Yes" → FlexConnect with local EAP / AAA cache If "No"→ Go to next question Does site-local traffic (printers, file servers, Sonos) dominate usage patterns? If "Yes"→ FlexConnect If "No" → Local mode is probably fine The default bias should be: **local mode for campus, FlexConnect for branch**. Only deviate from that default when you have a specific reason, and document the reason in your design doc so the next engineer knows why. ## A Few Things People Get Wrong Three mistakes I see repeatedly, all of which are cheap to avoid if you know about them: **Mismatched native VLAN.** The Flex Profile has a `native-vlan-id` setting that must match the native VLAN on the switchport the AP is plugged into. If they don't match, the AP will appear to join fine but client traffic won't make it past the first hop - you'll see joined APs and no client connectivity. **Running guest in Flex local-switch mode.** Guest typically anchors to a DMZ C9800, which requires central switching. If you try to run a guest SSID as locally-switched on a Flex AP, either the auto-anchor breaks or you end up with guest traffic landing on a branch VLAN it has no business touching. Central-switch guest SSIDs even on Flex APs. **Forgetting Efficient AP Join Image Download.** When a branch FlexConnect site has 40 APs and the WAN is 100 Mbps, upgrading all 40 APs from a central WLC kills the WAN for hours. Enable Efficient AP Image Upgrade so that one AP downloads the image from the WLC and the rest peer from it over the local LAN. This is a branch-saver and should be on by default. ``` C9800(config)#wireless profile image-download default C9800(config-wireless-image-download-profile)#image-download-mode capwap C9800(config-wireless-image-download-profile)#parallel C9800(config-wireless-image-download-profile)#max-parallel 15 C9800#show ap image Total number of APs : 412 Predownloading : 412 Completed predownloading : 398 ``` ## Key Takeaways Local mode is the right answer when the AP and the WLC share a campus; FlexConnect is the right answer when the AP is on the far side of a WAN. Standalone mode is the reason branches run FlexConnect - configure local EAP or AAA survivability cache or you won't actually get the resilience you're paying for. VLAN design is where FlexConnect deployments fail: keep numbering consistent across sites if you can, and use per-site VLAN mapping only when you must. Mix central and local switching on the same Flex AP when your policy or guest workflow demands it. Default to local mode in the campus, FlexConnect in the branch, and don't deviate without a written reason. _Truncated after 5 MiB. Use `/sitemap.xml` for the complete archive of public content._