netdev - Re: NIU - Sun Neptune 10g - Transmit timed out reset (2.6.24)

lists.openwall.net		lists / announce owl-users owl-dev john-users john-dev passwdqc-users yescrypt popa3d-users / oss-security kernel-hardening musl sabotage tlsify passwords / crypt-dev xvendor / Bugtraq Full-Disclosure linux-kernel linux-netdev linux-ext4 linux-hardening linux-cve-announce PHC
Open Source and information security mailing list archives

Hash Suite for Android: free password hash cracker in your pocket

[<prev] [next>] [<thread-prev] [thread-next>] [day] [month] [year] [list]

Message-ID: <483B239D.4090402@krogh.cc>
Date:	Mon, 26 May 2008 22:54:53 +0200
From:	Jesper Krogh <jesper@...gh.cc>
To:	David Miller <davem@...emloft.net>
CC:	yhlu.kernel@...il.com, linux-kernel@...r.kernel.org,
	netdev@...r.kernel.org, matheos.worku@....com
Subject: Re: NIU - Sun Neptune 10g - Transmit timed out reset (2.6.24)

David Miller wrote:
> From: David Miller <davem@...emloft.net>
> Date: Mon, 26 May 2008 12:33:38 -0700 (PDT)
> 
>> From: Jesper Krogh <jesper@...gh.cc>
>> Date: Mon, 26 May 2008 21:03:34 +0200
>>
>>> Ok. Now I also hit it in production with the NFS-server, so this
>>> is definately a real bug somewhere in the driver. Should I register it
>>> at bugzilla?
>> Please feel free to do that.
> 
> BTW, I did stare at some of the transmit code of the NIU driver
> while flying from Tokyo to Seattle a few hours ago, and I
> found one possible theory on the transmit timeouts.
> 
> Can you try the patch below and let us know if the symptoms
> continue?
> 
> [ Note to Matheos: The IRQ marking scheme of the NIU doesn't mesh
>   well with how things work under Linux.  We really needs a
>   "TX queue empty" interrupt status in order to handle all cases
>   properly.  Otherwise we really cannot decide not mark some TX
>   descriptors without potentially entering a deadlock condition.  ]
> 
> diff --git a/drivers/net/niu.c b/drivers/net/niu.c
> index 918f802..7ab7f8e 100644
> --- a/drivers/net/niu.c
> +++ b/drivers/net/niu.c
> @@ -6165,7 +6165,7 @@ static int niu_start_xmit(struct sk_buff *skb, struct net_device *dev)
>  	rp->tx_buffs[prod].mapping = mapping;
>  
>  	mrk = TX_DESC_SOP;
> -	if (++rp->mark_counter == rp->mark_freq) {
> +	if (1 /*++rp->mark_counter == rp->mark_freq*/) {
>  		rp->mark_counter = 0;
>  		mrk |= TX_DESC_MARK;
>  		rp->mark_pending++;

Applied and running.. I've now pushed 400GB of data through it trying to
get it to hit the bug but it is still running.

So without saying that it solved the problem, it definately seems so.
2.6.26-rc4 + above patch.

Jesper
-- 
Jesper
--
To unsubscribe from this list: send the line "unsubscribe netdev" in
the body of a message to majordomo@...r.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html