Thursday, 14 December 2017

Weblogic STUCK & Hogging Threads, How to deal with STUCK & HOGGING Threads.



What is STUCK thread?

How to deal with STUCK thread?

What is HOGGING thread?

How Weblogic determine a threadto declare as Hogging?


What is a STUCK thread?

We know that a STUCK thread is a thread which is processing a request for more than maximum time defined for a thread to complete the request which is default 600 and can be configured from admin console.  Based on different technical circumstances like due to some intermittent issues with network, database, application server etc a STUCK can be release after some time, but most of the time is certain thread or threads declared as STUCK then there would be some problem either temporary or permanent which need some fix.

How to deal with STUCK thread?

As of now there is no way to deal with a STUCK thread like, sometime end users ask if there any way to kill STUCK thread. No, there is no way to deal with STUCK thread.

To deal with STUCK thread –

·                     Take multiple thread dumps immediately.
·                     Review thread dumps or from console (managed server > monitoring > threads) Where it exactly got stuck?
·                     See how many threads got stuck?
·                     If the stuck thread count is increasing or constant?
·                     If constant then if got stuck on same area (application code etc ) or at different places ?
·                     If getting increase then there would be some serious problem and you have to do a quick health check of youapplication server, database and other integrated technologies wherever your application reaching like ldap server for authentication, some other API’s or web services etc, and in parallel review thread dumps for STUCK threads and share same with your developers to analyze quickly.
·                     If you have one, two or few constant STUCK threads and it’s not increasing then you can monitor it for some more time to check if they get clear or not, if not then to clear them you have only option to restart your managed server(s), and its better to restart and clear them before they make further any impact.

  
What is Hogging thread?  


I am sure if you are going to read this post then you must aware about what is hogging thread. Ok, let me define it again in a single line, “I hogging thread is a thread which is taking more than usual time to complete the request and can be declared as STUCK”.


How Weblogic determine a thread to declare as Hogging?

As we know a thread declared as STUCK if it runs over 600 secs (default configuration which you can increase or decrease from admin console).

Now, How Weblogic determines a thread to declare as Hogging? ok, here is the logic which 
I had learn from some of the Oracle internal portal note.

1.             There is an internal WebLogic polar which runs every 2 secs  (by default 2 secs and can be alter)
2.             It checks for the number of requests completed in last two minutes
3.             Then it checkhow much times each took to complete
4.             Then it takes the average time of all completed request (completed in last 2 sec)
5.             Then multiply average time with 7, and the value came consider as “usual time to complete the request”
6.             Now weblogiccheck each current executed thread in last 2 secs and compare with above average time, if for any of the thread it’s above this value then that thread will declare as Hogged thread.


For example –

1.             At a particular moment,  total number of completed requests in last two seconds – 4
2.             Total time took by all 4 requests – 16 secs
3.             Req1 took – 5 secs, Req2 took – 3 secs, Req3 took – 7 secs, Req4 took – 1 sec
4.             Average time = 16/4 = 4 secs
5.             7*4 = 28 secs
6.             Now weblogic check all executed threads to see which taking more than 28 secs, if any then that thread(s) declared as Hogged Thread.



Only the thing you can change with respect to hogging threads configuration is Polar time (Stuck Thread Timer Interval parameter) which is 2 secs by default. You can change this polar value to some different value like 4 secs if you want polar to run in every 4 secs instead of 2 secs.


Logs Rotation


Each WebLogic Server instance writes all messages from its subsystems and applications to a server log file that is located on the local host computer. By default, the server log file is located in the logsdirectory below the server instance root directory; for example, DOMAIN_NAME\servers\SERVER_NAME\logs\SERVER_NAME.log, where DOMAIN_NAME is the name of the directory in which you located the domain and SERVER_NAME is the name of the server.


In addition to writing messages to the server log file, each server instance forwards a subset of its messages to a domain-wide log file.The domain log file provides a central location from which to view the overall status of the domain. The domain log resides in the Administration Server logs directory. The default name and location for the domain log file is DOMAIN_NAME\servers\ADMIN_SERVER_NAME\logs\DOMAIN_NAME.log, where DOMAIN_NAME is the name of the directory in which you located the domain and ADMIN_SERVER_NAME is the name of the Administration Server.




You can rotate log files - 


By size ( default)
By Time




By default, when you start a WebLogic Server instance in development mode, the server automatically renames (rotates) its local server log file as SERVER_NAME.log.n. For the remainder of the server session, log messages accumulate in SERVER_NAME.log until the file grows to a size of 500 kilobytes.
Each time the server log file reaches this size, the server renames the log file and creates a new SERVER_NAME.log to store new messages. By default, the rotated log files are numbered in order of creationfilenamennnnn, where filename is the name configured for the log file. You can configure a server instance to include a time and date stamp in the file name of rotated log files; for example, server-name-%yyyy%-%mm%-%dd%-%hh%-%mm%.log.
By default, when you start a server instance in production mode, the server rotates its server log file whenever the file grows to 5000 kilobytes in size. It does not rotate the local server log file when you start the server.



You can change these default settings for log file rotation. 


For example
you can change the file size at which the server rotates the log file or 
you can configure a server to rotate log files based on a time interval. 
You can also specify the maximum number of rotated files that can accumulate. After the number   of log files reaches this number, subsequent file rotations delete the oldest log file and create a new log file with the latest suffix.


To change log size for rotation or by time -

Login to Administration Console
click on the server
On right hand side click on logging tab
under default selected general tab

https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEjTvLHYi1J24GvttfA-Ynhq4QVymmcmqI2lpBrhK1V6iCqngHm5LWQS3EIzo9MyGrR2S51h1QWJHsj7NyWcIDJPL39j8GNXKZg5AurUa-xFXmg_Pnj1UYtfDjqr1YseCw9k2rXK3PDU0uo/s400/logs+-+2.JPG

Or select By time option under rotation type drop box and enter time in begin rotation time at the you want your server rotate log file

To update for access logs

Login to Administration Console
click on the server
On right hand side click on logging tab
Select HTTP tab and follow the same above instructions.
https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEjdF_lmw-XOYnGhb4OENvVDTFpmFbgZGn93uGGv1DF5uyl-Hy8SvFJaeP5owh0U91IU1JijHSG5FVVNkm4v1YFAa-pFlerfm9DzNHgFCgmKVZK67cy31uBOxqQidqul6raAGntu75w8FKs/s400/logs+-+1.1.JPG

https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEjzqqHQdnJj-dvqP_fk-d1C5kINO9OqUWqujA6Q3YJ4rwG7RYfv1exwTq-RqZa4OJzs9lfW1gdKg81o-hB-qiJKVBVUMnBXg-yY3POjF_kOFeIjD7iV8-hFRtSAwh2DmhzGEolR6lbVnJs/s400/logs+-+3.JPG


Thread Dump:



1.            What is Thread dump?
2.            When we will take Thread dump? (Scenarios)
3.            How  Many  ways take Thread Dumps
4.            Thread Dump Generating Procedure
5.            What can I Analysis with Thread Dump?
6.            How can I analysis thread dump?
7.            Actions taken for Issue resolving
8.            References

 Coming to step by step learning:
--------------------------------

What is Thread dump?

Thread Dump is a textual dump of all active threads and monitors of Java apps running in a Virtual Machine.

When we will take Thread dump? (Scenarios)

1.  Scenario 1: when server is hang Position, i.e. that time server will not respond to coming requests.

2.  Scenario 2: While sever is taking more time to restart

3.  Scenario 3: When we are Getting exception like “java.lang.OutOfMemoryException”

4.  Scenario 4: Process running out of File descriptors. Server cannot accept further requests because sockets cannot be created

5.  Scenario 5: Infinite Looping in the code


How many ways take Thread Dumps?

Many types we have to take a Thread dumps. As per your flexibility you can choose one Procedure. For analyzing take dumps some Intervals (like every 10mins, 10mins etc.).



Generating Dump Talking Procedures

1. Take Thread dump from Console by Using of below command
    $kill -3 PID
   (For Getting PID, Use this Command ps –ef | grep “java”)
Here The Output of the Thread Dump will be generated in the Server STDOUT.
(Note: If a process is not responding to kill -3 then it’s a JVM bug.)


2. Generation Thread Dump via Admin Console

a.  login to Admin Console(with Admin Username/Password)
b.  Click on Server, after choose your server
c.  Goto Monitoring TAB
d.  Goto Threads TAB, after click on “Dump Thread Stack” Button
e.  Now you can view the all the Threads in Same page
f.  Copy and paste in a txt file.

3.  We can Collect Thread Dump Using “WebLogic.Admin” which is deprecated, but still available or may be available in near future as well As i think because it is one of the best debugging utility for Admins.

  java WebLogic.Admin -url t3://hostname: port -username Weblogic -password Weblogic THREAD_DUMP


This Thread Dumps will be generated in Servers STDOUT file


4. Getting Thread Dumps by using Jstack Utility

    a.jstack –m (to connect to a live java process)

    b. jstack –m [server_id@]
                (to connect to a remote debug server)
    (-m Means print both java and native frames (mixed mode)) 
process ID: jps

[soaosb]-->jps

19234 Server

17786 Server

11653 TMMain

18087 NodeManager

23300 Jps

18414 Server

18886 Server


example: -->$JAVA_HOME/bin/jstack -l 19234 > osb1.tdump

5. By Using WLST Script, can contain extension of (.py)

connect(‘weblogic’,'weblogic’,'t3://hostname:port′) cd (”Servers’) ls()cd (‘AdminServer’) ls() threadDump()

 Execute this Script in console. 

What can I Analysis with Thread Dumps?
We need to analyze the thread dumps for analyzing running threads and their states to identifying.

How can I analysis thread dumps?

For analyze thread dumps we have lots of tools to understand easily thread states

1.  samurai tool :

    In this tool you can identify all the Thread states by     identifying colors. We need to take care about Deadlocks and waiting state threads.

   More Details:

    $ java -jar samurai.jar

     After running we will get a Screen like below

      Goto Thread dump tab
    When Samurai detects a thread dump in your log, a tab named "Thread Dump" will appear.

 You can just click "Thread dumps" tab to see the analysis result. Samurai colors idle threads in gray, blocked threds in red and running threds in green. There are three result views and Samurai shows "Table view" by default. In many case you are just interested in the table view and the sequence view. Use the table view to decide which thread needs be inspected, the sequence view to understand the thread's behavior. You should takecare especially threds always in red.


2.  TDA Tool :


Actions taken for Issue resolving

1.  Classic Dead Locks : Look for the threads waiting for monitor entry

For Example :

"ExecuteThread: '95' for queue: 'default'" daemon prio=5 tid=0x411cf8 nid=0x6c waiting for monitor entry [0xd0f80000..0xd0f819d8]
    at weblogic.common.internal.ResourceAllocator.release(ResourceAllocator.java:766)
    at weblogic.jdbc.common.internal.ConnectionEnv.destroy(ConnectionEnv.java:590)
Reason: The above thread is waiting to acquire lock on Resource Allocator object. The next step is to identify the thread that is holding the Resource Allocator object
"ExecuteThread: '0' for queue: '__weblogic_admin_rmi_queue'" daemon prio=5 tid=0x41b978 nid=0x77 waiting for monitor entry [0xd0480000..0xd04819d8]
    at weblogic.jdbc.common.internal.ConnectionEnv.getPrepStmtCacheHits(ConnectionEnv.java:174)
    at weblogic.common.internal.ResourceAllocator.getPrepStmtCacheHitCount   (ResourceAllocator.java:1525)
Reason: This thread is holding lock on source Allocator object, but is waiting for Connection Env object. This is a classic deadlock.

      
2.  Threads in wait() state:
   A sample dump:

"ExecuteThread: '10' for queue: 'SERV_EJB_QUEUE'" daemon prio=5 tid=0x005607f0 nid=0x30 in Object.wait() [83300000..83301998]
  at java.lang.Object.wait(Native Method)
  - waiting on (a weblogic.ejb20.pool.StatelessSessionPool)
at weblogic.ejb20.pool.StatelessSessionPool.waitForBean(StatelessSessionPool.java:222)

Reason: The above thread would come out of wait() under two conditions
 (Depending on application logic)
1) One of the thread available in the execute queue pool would call notify() on this object when an instance is available. (If the wait() is indefinite).
  This can cause the thread to hang for ever if server never does a notify() to this object.

2) If the timeout exceeds, the thread would throw an exception and back to execute queue thread pool.

What is lok file ?how many types of lok files are there?



lok files are for locking some action which should be run or made only by one user or holding process. You can find config.lok invoked for serial update of config.xml, then each server has its own lok file for server and lok file for embedLDAP to avoid start same manged server in time twice. Then you can find edit.lok invoked by Edit actions for one user editing domain configuration in time. I am not sure if other exist.

It is good to know it, because sometimes when server crash you cant start servers and it writes something like "Instance is already running" Or second issue server begin to start and then after line in log with text "....IIOP..." next step is load embedLDAP but server stuck for long time in this step because ldap is locked or corrupted. Remove of these lok files usually can help you to start server perfectly again

How to change / reset weblogic admin user password



Some engineers think it's just a single step to change the weblogic admin user password from console under realm option, but it's not really a single step because if you do change the admin user password from console only then you would able to logout with existing session and login with new password but you would not able start your server once you will brought it down untill and unless you will do some more workaround which is the part of weblogic admin user password change procedure.


if you will only change the admin user password from console and after that try to start your admin server you will get below error

*********************************************************************************

hentication denied: Boot identity not valid; The user name and/or password from the boot identity file (boot.properties) is not valid. The boot identity may hav
e been changed since the boot identity file was created. Please edit and update the boot identity file with the proper values of username and password. The firs
t time the updated boot identity file is used to start the server, these new values are encrypted.
weblogic.security.SecurityInitializationException: Authentication denied: Boot identity not valid; The user name and/or password from the boot identity file (bo
ot.properties) is not valid. The boot identity may have been changed since the boot identity file was created. Please edit and update the boot identity file wit
h the proper values of username and password. The first time the updated boot identity file is used to start the server, these new values are encrypted. at weblogic.security.service.CommonSecurityServiceManagerDelegateImpl.doBootAuthorization(CommonSecurityServiceManagerDelegateImpl.java:959)at weblogic.security.service.CommonSecurityServiceManagerDelegateImpl.initialize(CommonSecurityServiceManagerDelegateImpl.java:1050)
        at weblogic.security.service.SecurityServiceManager.initialize(SecurityServiceManager.java:873)
        at weblogic.security.SecurityService.start(SecurityService.java:141)
        at weblogic.t3.srvr.SubsystemRequest.run(SubsystemRequest.java:64)
        Truncated. see log file for complete stacktrace
Caused By: javax.security.auth.login.FailedLoginException: [Security:090304]Authentication Failed: User weblogic javax.security.auth.login.FailedLoginException:
 [Security:090302]Authentication Failed: User weblogic denied
        at weblogic.security.providers.authentication.LDAPAtnLoginModuleImpl.login(LDAPAtnLoginModuleImpl.java:261)
        at com.bea.common.security.internal.service.LoginModuleWrapper$1.run(LoginModuleWrapper.java:110)
*********************************************************************************

To avoid this you have to update the admin server boot.properties file also
So, here is the procedure to change the weblogic admin user password

Part A.
Login to admin  console
Under Domain Structure, select “Security Realms” option
Click on “myrealm”
Click on tab “Users and Groups"
Click on your admin user
Click on the Passwords tab
Update the password

Part B.
Logout and login again with new password to make sure you are able to login with new password. ( if you are not able login with new password then it means you have updated something else and trying with something else :)  )
Ok, now


1. Stop your admin server
2. Go to your_domain/servers/you_admin_server/security directory
3. Take backup of existing boot.properties file
4. Create new boot.properties file with below contents
username=your_admin_user
password=your_new_password

5. Now start your admin server

Wait, its not over, If you have managed servers in your domain then you have to do some more workaround for them to boot up properly during next restart

Important :
If you always start your managed servers from console and never started using command line ( using startManagedserver command ) by you or by anyone since provisioning ( means setup of  environment ) then you will not see any boot.properties file under your managed server(s) staging security directory ( your_domain/servers/your_managed_server/security ) and if will try to start managed servers using script then you will be prompt for username and password always untill and unless you will create boot.properties manually under your_domain/servers/your_managed_server/security directory.
If you have changed admin user password ( using the way I have mentioned above ) then you would able to stop start login admin console successfully but you will not able to start managed servers once you will stop them ( you will get same above highlighted exception in logs ) untill and unless you will do below work around

Workaround - 1
1. Go to "your_domain/servers/your_managed_server/data" for each managed server you have    
    and rename ldap folder to ldap.old and nodemanager folder to nodemanager.old
2. Start managed server(s) from console

Workaround - 2
if you still getting same authentication exception then including workaround-1 first step, follow below steps also
1. Change the nodemanager password from admin console also
Login to admin console
Click on your domain name ( on left hand tree under Domain Structure )
Click on security tab
Click on advance option link
Change "NodeManager Password:"

2. Go to your WL_HOME/common/nodemanager folder and rename nm_data.properties file as nm_data.properties.old
3. Restart node manager
4. Start your managed servers